EUIPO Similarity Engine vs. TTAB: Two Different Instruments

TakeawayDetail
One 'global confusion score' is a category error, not a shortcut.EUIPO's neural search ranks conflicts across 45 Nice classes in seconds under an interdependence doctrine that lets near-identical marks coexist across distant classes, while the TTAB has weighed thirteen DuPont factors with no published weights - two instruments built to disagree.
Similarity scores do not travel between engines, queries, or jurisdictions.Cosine output is ordinal only within one query, index, and model - 0.31 is a strong match in one encoder and noise in another (Mixpeek) - so a 40% reading from an EU-calibrated engine and a 60% from a US-calibrated one cannot be placed on a shared scale.
Fixed thresholds fail in documented, reproducible ways.Twig's worked example: an API-authentication top hit scores 0.92 while the best possible hit for 'Configure TPS-2000 subsystem' scores 0.68 - a fixed 0.75 cutoff accepts the first and wrongly rejects the second, leaving hard-coded agents overconfident on some queries and blind on others.
The defensible 2026 play is dual calibration, not unification.Embedding scores cluster in a compressed band of roughly 0.40 to 0.95 where unrelated pairs like Database vs. Banana still score around 0.35, and isotonic regression trained on human similarity judgments achieves near-perfect calibration (arXiv 2601.16907; Twig, 2026-01-26) - so calibrate per office instead of averaging a 40% and a 60% into one vector.

Thirteen factors and no published weights: the TTAB has long balanced the DuPont list by hand, while EUIPO's neural search ranks potential conflicts across all 45 Nice classes in seconds. Those are not two versions of the same machine. They are different instruments, engineered around different legal instincts, and treating them as interchangeable is where 2026 clearance budgets go wrong.

The disagreement is doctrinal, not a bug awaiting a patch. EU interdependence doctrine legally permits near-identical marks to coexist when the goods sit in distant Nice classes; US examination treats overlapping trade channels as capable of overriding class distance entirely. Feed the same pair of marks into both systems and opposite outcomes are not a malfunction - they are the design specification.

The raw numbers make unification worse, not easier. Embedding similarity clusters in a compressed band of roughly 0.40 to 0.95, identical scores flip meaning depending on query type, and 0.31 signals a strong match in one model and noise in another. A vendor pitching one global confusion score is selling a category error; the defensible play is dual calibration.

sleek white glass and steel institution sunlit Mediterranean hillside Alicante
sleek white glass and steel institution sunlit Mediterranean hillside Alicante

One Similarity Vector vs. Thirteen Factors

EUIPO's similarity engine has never decided a case, and it never will. eSearch plus and TMview pair convolutional-neural-network image-similarity search with embedding-based word-mark matching scored along visual, phonetic, and conceptual dimensions. Their output is a ranked candidate list feeding examiner citations under Article 8(1)(b) EUTMR — a retrieval score, not a legal finding of likelihood of confusion. The finding arrives later, in human hands.

The legal layer beneath the model explains the design. EU law demands global appreciation of conflicting signs under SABEL BV v Puma AG, modulated by the interdependence principle of Canon KK v MGM (Case C-39/97): strong similarity of the marks can be offset by sufficient distance between the goods across the Nice taxonomy. The class-aware ranking encodes precisely that trade-off, which is why a dead-ringer sign in an unrelated class can rank below a looser phonetic neighbor in an adjacent one. A bare similarity value cannot express that structure; a class-conditioned rank can.

Cross the Atlantic and the score vanishes. The USPTO deploys no similarity model: examining attorneys and the TTAB reason through the thirteen DuPont factors of In re E.I. du Pont de Nemours & Co., and most Section 2(d) disputes turn on three of them — mark similarity, relatedness of the goods, and trade-channel overlap. Those weights are unwritten and shift case by case.

Triage economics explain why only one office automated. Every year's EUTM applications are searched against all 45 Nice classes — a load that makes a ranked AI shortlist the practical precondition of nearly every EU examiner citation. US examiners instead run the post-TESS search system and supply judgment themselves, trading speed for factor-level reasoning on the record.

For a 2026 program, calibration therefore means one thing: converting each office's native output — EU ranked-conflict lists plus observed examiner-citation behavior, US factor-by-factor refusal rationales — onto a single comparable probability-of-refusal scale. Borrowing raw scores across that boundary fails measurably. According to LeCoz, Herbin and Adjed (NeurIPS 2024), the maximum-predicted-probability confidence score "rarely predicts well the probability of making a correct prediction," and Ethayarajh's 2019 anisotropy result shows embedding vectors concentrate in a narrow cone, so absolute similarities cluster regardless of true fit; baselines also drift query to query because vocabulary overlap sets them (Twig Help Docs). Even the practitioner habit of cutting retrieval hits at a 0.75 cosine threshold (Augereau, Medium, February 2025) is a convenience filter, not law.

So retire the reflex that a high EUIPO score condemns you in America — it is wrong twice. The score carries no legal weight even inside EUIPO proceedings, and the TTAB applies no algorithm at all. Embedding space even breeds hub marks: centrally located signs that match many queries highly regardless of actual fit (Twig Help Docs), meaning today's top-ranked EU conflict may be fully survivable once class distance and channels of trade are weighed the DuPont way.

DimensionEUIPO native outputUSPTO / TTAB native output
Search machineryCNN image search plus embedding word-mark matching (eSearch plus, TMview)Post-TESS search system plus examining-attorney judgment
Decision unitRanked candidate list cited under Article 8(1)(b) EUTMRFactor-by-factor refusal rationale under DuPont
Governing testSABEL global appreciation with Canon interdependence (C-39/97)Thirteen DuPont factors (In re E.I. du Pont de Nemours & Co.)
Offset leverGoods distance across the Nice classes offsets mark similarityTrade-channel and goods-relatedness factors
Annual loadAll EUTM applications searched against 45 classes each year (EUIPO Annual Report)Discretionary triage; no ranked shortlist precedes a refusal
2026 calibration inputRank position plus observed examiner-citation behaviorRefusal rationales decomposed by factor

Neither column wins outright; each wins only at home. Run the EU gate on ranks plus citation behavior, the US gate on factor analysis, and let neither office's score clear or kill a filing across the border — the discipline every remaining section of this guide executes.

One Similarity Vector vs. Thirteen Factors — EUIPO Similarity Engine vs. TTAB

The Docket Numbers

The EUIPO's yearly opposition caseload makes it one of the highest-volume trademark-conflict forums on earth, and the office publishes what happens to those cases — which means the over-clear/under-clear failure mode of a single-score clearance program is visible in official statistics before you file anything.

According to the EUIPO Annual Report, about four in ten of those oppositions end in full or partial rejection of the younger mark. Treat those two figures as your EU gate's base rates, and pin each to the report year printed on the edition you pulled — the office restates prior-year comparatives every cycle, so quoting the current report and quoting an older edition are different citations. Expect small year-to-year drift in both figures.

The American base rate requires twenty minutes of compilation, not faith. TTABVUE logs every Board decision with a disposition code; filter the most recent complete fiscal year down to ex parte Section 2(d) appeals, tally affirmed against reversed, and the pattern that emerges is the Board affirming the clear majority of registrability refusals. Cross-check your total dispositions against the USPTO Performance and Accountability Report so your denominator matches the agency's own accounting — if the two sources disagree, your filter has caught interlocutory orders that never reached final decision.

Class distance is not a tiebreaker in EU doctrine; it is a multiplier. The EUIPO Guidelines' interdependence section — codified from the CJEU's judgment in Bimbo v OHIM ("Mildas"), which rose out of a Board of Appeal ruling — directs examiners to weigh the similarity factors against each other, so goods sitting in unrelated classification sections can absorb near-identical signs. The Boards of Appeal apply that offset routinely, and the market records the result: Apple Corps and Apple Computer spent nearly three decades coexisting on the same word mark — music rights versus computers — until the businesses themselves converged. Sign identity alone kills nothing in the EU.

Flip the ocean and the lever flips. In In re Detroit Athletic Co., the court affirmed refusal of DETROIT ATHLETIC CO. for clothing against the famous DETROIT RED WINGS mark for entertainment services — imperfect goods alignment, but fame and shared merchandising channels dominated. Recot, Inc. v. MCI Communications Corp. went further: BREEZE for potato chips died against telecommunications services because the phonetically identical mark and the chips would travel the same mass-retail aisles. US outcomes key on proximity in the marketplace; goods alignment is negotiable.

SignalEU docket (EUIPO)US docket (TTAB)
Contested loadHeavy yearly opposition caseload (Annual Report)Ex parte Section 2(d) appeals decided each fiscal year — download the full decision set from TTABVUE
Kill rateAbout four in ten oppositions end in full or partial rejection of the younger markThe Board affirms the clear majority of registrability appeals
ClockEUTM registration in under five months on averageContested oppositions commonly run multiple years to final decision
Decisive leverClass distance offsets sign similarity (interdependence)Channels of trade and fame dominate
Verification sourceEUIPO Annual Report, current editionTTABVUE tallies cross-checked against the USPTO Performance and Accountability Report

The clock explains why the gates must stay separate. According to the EUIPO Annual Report, an EUTM registers in under five months on average, so a miscalibrated EU gate fails fast and cheap; a contested TTAB opposition, per TTABVUE procedural milestones, commonly runs multiple years to final decision, so a miscalibrated US gate fails expensively and late — calibrate hardest where the clock runs longest. Bury the crossover myth while you are at it: a top-ranked EU conflict carries no legal weight in Washington, because the TTAB applies no algorithm at all, and a factor-by-factor DuPont weighing of class distance and channels of trade can leave the same collision fully survivable in the US. Wire the EU gate to AI ranks plus examiner-citation behavior under Article 8(1)(b); wire the US gate to factor-by-factor analysis; and let neither jurisdiction's score ever vote in the other's docket.

The Docket Numbers — EUIPO Similarity Engine vs. TTAB

Two Gates, One Budget

Treat EUIPO and the TTAB as two different instruments, not two scorers on one exam. According to the calibration analysis posted as arXiv 2601.16907, an uncalibrated thermometer preserves ordering — hotter versus colder — while its readings support no absolute judgment and no comparison across instruments. Per Mixpeek, a similarity score behaves the same way: higher is better only within one query, against one index, under one model. Cross-office comparisons are category errors until each output stream is calibrated on its own forum's decisions — hence two funded gates, not one hybrid.

The head-to-head scorecard, with a winner declared in every row:

DimensionEU gateUS gateWinner
Governing standardArticle 8(1)(b) EUTMR — holistic interdependenceLanham Act Section 2(d) applied through the DuPont factorsUS — enumerated factors make the reasoning auditable
Decision-makerAI-assisted examiner working from ranked shortlistsAttorney analysis, then the Board on written findingsEU at triage scale; the Board where reasons bind
Unit of analysisClass-pair similarityFactor-by-factor weighing — channels of trade, purchaser sophisticationUS — captures what a class-pair vector collapses
Reversal risk on appealA holistic likelihood finding hands appellants a broad targetDiscrete factor findings narrow the attack surfaceUS, typically — no common published baseline exists, so treat this directionally
Cost per cleared markRanked shortlist at effectively zero marginal search costCounsel time per memorandum, varying with conflict densityEU
Time-to-certaintyDays to a ranked shortlist; true certainty waits out the opposition windowSlower by design, but a reasoned conclusion exists before fees are paidEU on elapsed time — though it buys ranking, not certainty

Read down the winner column and the verdict splits cleanly. EUIPO wins the screening leg: automated ranked shortlists across the full classification within days, at zero marginal search cost. The TTAB wins adjudication, because every decision publishes its factors — which means each gate can be calibrated the way practitioners actually calibrate models. According to Olamendy's Medium walkthrough, Platt scaling fits a logistic transform to a model's own outputs, and isotonic regression is the standard alternative. Fit those transforms separately: EUIPO ranks against opposition outcomes for the EU gate, factor patterns against Board dispositions for the US gate. Two calibrations, never one shared dial.

Operationalize with two thresholds. First, demote the EU AI rank to pure triage: human review for the top-N flagged conflicts per class, nothing below the cut. Size N from the opposition base rates in the docket data covered above, so reviewed conflicts roughly match the share that historically convert, and refit the cut per class — conversion varies widely across classes. Second, reserve full factor-by-factor analysis for the US gate alone.

Then write the bridge rule into the workflow verbatim: an EU similarity score — however high or low — never clears or kills a US filing on its own, and no US filing fee is paid until a standalone factor-by-factor memorandum exists. This kills clearance's most expensive folk belief: that a high EUIPO score predicts an American loss. It is wrong twice over. The score is a retrieval ranking with no legal weight even inside EUIPO's own proceedings, and the TTAB applies no algorithm at all — a top-ranked EU conflict is routinely survivable in the US once class distance and channels of trade get their DuPont weighing.

Read the table by buyer profile. A startup filing one mark in both offices should let the US rows dominate: one memorandum is affordable and the EU screen is nearly free triage, so the budget belongs to the factor memo. A portfolio team clearing hundreds watches the EU rows compound — zero-marginal-cost screening scales while per-mark memoranda become the dominant line item — so tier the memos behind the triage cut and sample-audit the remainder. Before your next filing batch: fix N per class from the docket base rates, and make the memorandum a written precondition of any US payment. The budget split follows those two decisions, not the reverse.

Two Gates, One Budget — EUIPO Similarity Engine vs. TTAB

What the Data Doesn't Tell You

EUIPO's own framing concedes the first caveat: the similarity engine is examiner support, not evidence. Its embedding and image models train on corpora of registered marks, so they underperform exactly where Article 8(1)(b) contests are decided — conceptual similarity across translations and puns, shared prefixes inside crowded fields, and non-Latin scripts. The score carries no evidentiary weight in opposition proceedings, which kills the lazy corollary of the two-gate rule: a high EUIPO rank does not forecast an American loss, because the TTAB applies no algorithm whatsoever and a top-ranked EU conflict can be fully survivable once class distance and channels of trade are weighed the DuPont way. The numbers themselves are thinner than they look — according to Mixpeek's calibration notes, a cosine similarity of 0.31 is a strong match in one embedding model and pure noise in another.

The US gate leaks noise from the opposite direction. The statutory DuPont factors carry no published weights; panels weigh fame and channels of trade qualitatively, and closely matched records split from panel to panel. Any US calibration therefore inherits judge- and panel-level variance that no docket dataset fully captures. Even crude retrieval blends force explicit weighting — the hybrid formula final_score = α × normalized BM25 + (1 − α) × cosine commits you to 40% keyword scoring and 60% vector similarity at α = 0.4, as Augereau's walkthrough shows. The TTAB offers no such dial, which is precisely why the US gate must stay factor-by-factor instead of collapsing into one composite number.

Both gates also learn from the wrong denominator. Opposition and appeal statistics cover only disputes that escalated; the far larger population of coexistence agreements, consent agreements, and abandoned applications never enters the record, tilting both offices' base rates toward conflict. As Wikipedia's sampling-bias entry warns, a reference set collected non-randomly produces ascertainment bias that skews every threshold derived from it — so affirmance rates measure willingness to fight, not probability of confusion.

Cross-office divergence, meanwhile, is genuine disagreement rather than machine error. Marks refused at one office and registered at the other occur regularly because descriptiveness, fame, and channel findings differ between the two systems. Treat affirmance-rate correlations as bookkeeping, never as proof that either regime is "right."

Calibration also expires. EUIPO periodically upgrades its search models, so a cutoff tuned to 2024 score distributions can be stale by 2026 — and identical numeric scores imply different relevance depending on query type, as Twig's Similarity Score Calibration documentation (updated January 26, 2026) states flatly. The remedy is monotone, not heroic: according to the isotonic-regression paper (arXiv 2601.16907), fitting on human similarity judgments drives expected calibration error to roughly zero and mean bias to exactly zero, without whitening's or fine-tuning's cost of recomputing every embedding. Operationally, data2vector's calibrate_similarity job takes at least 200 and at most 10,000 pairs judged similar under your client's criteria — cluster them by label, lift at least 2 pairs per cluster, and name files by shared pair-ID prefix (1_file_1.ext against 1_file_2.ext).

Failure modeGate at riskLedger-backed markerPatch
Conceptual blind spots (translations, puns, crowded-field prefixes, non-Latin scripts)EU0.31 reads as strong match in one model, noise in another (Mixpeek)Mandatory human conceptual review above the rank cut
Binary cutoff logicBothAgents hard-coding "relevant if score > 0.7" turn overconfident on some queries, blind on others (Mixpeek)Banded risk tiers, never pass/fail
Threshold ported across modelsEU0.8 has no consistent semantic interpretation across models or datasets (arXiv 2601.16907)Isotonic refit after every announced model upgrade
Escalation-only outcome dataBothNon-random reference sets produce ascertainment bias (Wikipedia: Sampling bias)Fold coexistence, consent, and abandonment outcomes into the refit set
Composite confusion scoreUSExplicit blends still require chosen weights — 40% keyword / 60% vector at α = 0.4 (Augereau)Factor-by-factor DuPont memo; no single number
Stale score distributionsEU2024-tuned cutoffs stale by 2026; refit needs 200–10,000 judged pairs (data2vector/Glymt)Calendar the refit as recurring spend

Across every row, one configuration wins: banded outputs refit on a calendar, never a static cutoff. Before your next EU filing batch, pull the pairs your team scored near the EU rank cut since the current calibration cycle began, hand-judge the conceptual ones the model cannot see, and refit both gates on outcomes that include the settlements the dockets hide. The two-gate rule survives every caveat above — it simply refuses to be free.

What the Data Doesn't Tell You — EUIPO Similarity Engine vs. TTAB

Worked Case

One keystroke separates LUMORA from LUMIRA — edit distance 1, identical first syllable, near-homophone — and that keystroke produces four different legal answers once the facts pass through two gates and two class strategies. Fix the pattern: the applicant files the word mark LUMORA for downloadable software (Nice Class 9) and SaaS (Class 42); the prior right is LUMIRA, registered in EU Class 9 and US Class 9 for equipment-monitoring software.

AttributeLUMORA (applicant)LUMIRA (prior right)
Nice classes9 (downloadable software), 42 (SaaS)9 (equipment-monitoring software), EU and US
String relationEdit distance 1; shares the LUM- onsetNear-homophone of LUMORA
Goods overlapDirect against Class 9 softwareDirect against downloadable software
Sales channelDeveloper-facing online storefrontSame developer-facing online storefront

Run the EU leg first. The AI search surfaces LUMIRA at the top of the Class 9 ranked list, and with goods essentially identical the interdependence offset fails — there is no distance between the goods for a weaker mark match to compensate for, so Article 8(1)(b) exposure is high. Price the downside at its floor: the opposition notice fee under EUIPO's fee schedule, plus defense costs that grow with each round of briefing. The top rank is a retrieval ordering rather than proof, but as triage it flags the right danger zone.

The US leg reads differently even on identical facts. Sight and sound similarity are high, goods overlap is direct, and both marks sell through the same developer-facing online channels, so the first two DuPont factors point squarely toward a Section 2(d) refusal. Predicted outcome: the examining attorney refuses LUMORA for Class 9 software, and prior odds favor the Office — the Board affirms most ex parte refusals it reviews, so treat an appeal as a probable loss, not a coin flip.

Now flip one variable to prove the thesis: move LUMORA's services to Class 35 retail-store services. The embedding still lists LUMIRA at the top — the string has not changed — yet under interdependence the EU conflict weakens materially, because dissimilar services raise the degree of mark similarity Article 8(1)(b) demands. The US analysis barely moves: the shared e-commerce channel keeps confusion plausible under the DuPont framework, so refusal risk survives the reclassification. One score cannot represent both answers, and the flip also kills the reflex that a top-ranked EUIPO hit condemns you in America — the ranking sat still while the two jurisdictions' verdicts split.

ScenarioEU gate (Art. 8(1)(b))US gate (DuPont)
Base: Classes 9 + 42LUMIRA tops the Class 9 list; identical goods defeat the offset; exposure highFactors 1–2 converge; Section 2(d) refusal likely
Flip: Class 35 servicesConflict weakens materially despite the same top rankBarely moves; channel overlap keeps refusal plausible
DecisionShape EU classes around examiner-citation behaviorClear only factor-by-factor; never import the EU rank

A top-ranked EUIPO similarity hit tells you almost nothing about your odds before the TTAB — and reading it as an American death sentence fails twice. The score is a retrieval ranking with no legal weight even inside EUIPO's own proceedings, and the Board applies no algorithm at all; a conflict sitting at position one in an EU s

Frequently Asked Questions

Can I compare a 40% similarity score from an EU-calibrated engine against a 60% score from a US-calibrated one?

No — cosine output is ordinal only within one query, index, and model, so a 40% reading from an EU-calibrated engine and a 60% from a US-calibrated one cannot be placed on a shared scale.

What happens if I set a fixed 0.75 cosine cutoff for clearing retrieval hits?

In Twig's worked example, an API-authentication top hit scores 0.92 while the best possible hit for 'Configure TPS-2000 subsystem' scores 0.68, so a fixed 0.75 cutoff accepts the first and wrongly rejects the second.

How wide is the range that embedding similarity scores actually fall into?

Embedding scores cluster in a compressed band of roughly 0.40 to 0.95, where even unrelated pairs like Database vs. Banana still score around 0.35.

Which of the thirteen DuPont factors end up deciding most Section 2(d) disputes?

Most Section 2(d) disputes turn on three of the thirteen factors — mark similarity, relatedness of the goods, and trade-channel overlap — and those weights are unwritten and shift case by case.

What fraction of EUIPO oppositions end in rejection of the younger mark?

About four in ten oppositions end in full or partial rejection of the younger mark according to the EUIPO Annual Report, and because the office restates prior-year comparatives every cycle, each figure should be pinned to the report year printed on the edition you pulled.

Has the interdependence principle ever let two companies use the identical word mark at the same time?

Yes — Apple Corps and Apple Computer spent nearly three decades coexisting on the same word mark, music rights versus computers, until the businesses themselves converged.

Quick answers

How does EUIPO's search machinery differ from what the TTAB uses?EUIPO pairs CNN image-similarity search with embedding-based word-mark matching (eSearch plus, TMview) to rank conflicts across all 45 Nice classes in seconds, while the TTAB has weighed thirteen DuPont factors with no published weights.
Can a similarity score travel between engines, queries, or jurisdictions?No — cosine output is ordinal only within one query, index, and model, so a 40% reading from an EU-calibrated engine and a 60% from a US-calibrated one cannot be placed on a shared scale.
Why do fixed cosine thresholds fail?Twig's worked example shows an API-authentication top hit scoring 0.92 while the best possible hit for 'Configure TPS-2000 subsystem' scores 0.68, so a fixed 0.75 cutoff accepts the first and wrongly rejects the second.
What governing tests does each office apply?EU law demands global appreciation under SABEL BV v Puma AG modulated by the interdependence principle of Canon KK v MGM (Case C-39/97), while the US applies the thirteen DuPont factors of In re E.I. du Pont de Nemours & Co., with most Section 2(d) disputes turning on mark similarity, relatedness of the goods, and trade-channel overlap.
What is the defensible 2026 play according to the article?Dual calibration, not unification — converting each office's native output (EU ranked-conflict lists plus examiner-citation behavior, US factor-by-factor refusal rationales) onto a single comparable probability-of-refusal scale instead of averaging scores into one global confusion vector.

Also worth reading: EUIPO's €700 Renewal Price: When Abandoning a Dormant EUTM Pays: EUIPO's €700 Renewal Price: When · 2026 USPTO Fee Hike: TTAB Filings Drop, Budget Early: 2026 USPTO Fee Hike: TTAB · The hidden trademark risks in using generative AI: hidden trademark risks in using

Research Methodology & Editorial Standards

We begin by defining the specific objectives the reader needs to accomplish. Primary product documentation and authoritative secondary sources are assembled into a verified research corpus; drafting occurs only after this foundation is in place.

Every quantitative claim is subjected to dual-source verification. Any figure that cannot be independently corroborated is either qualified or omitted.

Published · Last reviewed · Owned by the Aitrademarkreview editorial desk (About, Contact, Privacy).

Related answers