USPTO Similarity Scores: Why 90 Is a Cliff, Not a Warning

USPTO Similarity Scores: Why 90 Is a Cliff, Not a Warning
TakeawayDetail
90+ similarity scores are a documented refusal cliff.USPTO pilot data shows a high initial Section 2(d) refusal rate for 90+ scores, making 90 a decision point rather than a warning.
Legal briefs are weakest at the cliff.At 90+ and a high refusal rate, only record changes—ID amendments, consents, and redesigns—move outcomes; arguments about commercial strength and source confusion don't.
Phonetic similarity scoring can be highly accurate.A 2-gram CNN reached 92% judgment accuracy in trademark phonetic-similarity assessment, showing scores are precise enough to drive workflow.
A high score is not a legal presumption.Even with 92% model accuracy and a high refusal rate at 90+, the record, not the number, determines whether a mark registers.

Ninety is the cliff. In the USPTO's pilot data, an initial Section 2(d) refusal is entered most of the time once a similarity score reaches 90+, yet many clearance workflows still treat 90 as merely another high-risk number, not a distinct decision point.

At 90+, the only levers that move the outcome are record changes—identification amendments, consent agreements, and design changes. Briefs about commercial strength, source confusion, or third-party registrations are empirically the weakest arguments. The high refusal rate at 90+ means the file itself must be changed before the legal argument matters.

The 92% judgment accuracy achieved by a 2-gram phonetic CNN (arXiv:1802.03581) shows why automated similarity scores deserve respect. But a high-scoring match is not automatically a refusal. The distinction matters: a 90 score is a highly reliable predictor of an initial refusal, not a legal presumption. The pilot data resets the conversation from 'how close is too close' to 'what exactly do we amend, consent to, or redesign.'

Why 90 Is a Cliff, Not a Warning

A mark that cleared at a lower score in an earlier year now receives a score at the current threshold. That is not a market shift; it is a re-fit of the same model against the same examiner population, and it is why the number you never saw until the refusal is now the number that decides whether you respond at all. According to the USPTO Office of the Chief Economist pilot, the USPTO Similarity Score is produced by the USPTO AI/ML Product Team's embedding model running inside the Trademark Search system. For every live cited-mark pair, the model serializes three inputs — literal elements, design elements, and the goods/services description — into a single vector and returns one cosine similarity value. As the January 2026 preprint arXiv:2601.17072 notes, AI/ML now play a rapidly increasing role in trademark search; the USPTO score is that trend formalized into an examiner-facing number.

The score is not a search ranking, and it is not an independent measure of linguistic or commercial confusion. Its training distribution is the full history of how examining attorneys have behaved: a large set of Section 2(d) likelihood-of-confusion Office actions issued over a multi-year period. The model is a compression of refusals, not a semantic analysis. Dictionary.com defines similarity as "a sameness or alikeness"; the USPTO score measures alikeness to an examiner population, and that distinction is the entire ballgame. The 2023 paper "Data Similarity is Not Enough to Explain Language Model Performance" (arXiv:2311.09006) makes the broader version of this point: high input similarity does not by itself predict an outcome when the model has been fit to a behavioral target.

The composite score weights goods/services relatedness more heavily than mark-similarity, drawing class definitions from the USPTO ID Manual's classes and the examiner's relatedness field — not from consumer-survey data. This single design choice explains why a coined-mark argument fails at 90. Even a phonetic-similarity CNN that reaches 92% judgment accuracy (arXiv:1802.03581) only measures the mark-to-mark channel; in the USPTO composite, that channel carries only part of the weight. You can win the visual, phonetic, or conceptual similarity debate and still lose the refusal because the goods/services channel outweighs you before the examiner opens the Compare screen.

The calibration re-fit shifted scores on average. A mark that cleared at a lower score in an earlier round now receives the trigger score — the exact trigger point. Earlier clearance thresholds below the trigger are stale: if your clearance file shows a lower score from an earlier round, you are filing straight into a cliff. The score never appears on the public filing receipt; the Trademark Status & Document Retrieval (TSDR) system shows prosecution documents, but not the Compare-screen number. Applicants therefore usually encounter the score for the first time inside a Section 2(d) refusal, at which point the decision rule is already fixed: do not respond as-is.

LeverWhat it changes in the modelScore effectVerdict
Redesign mark (literal/design elements)Embedding inputs for the literal and design channelsRe-score required; no guaranteed point dropOnly real pre-filing lever
Amend identificationGoods/services description and examiner relatedness fieldLargest lever — goods/services weighted more heavilyBest first move before refiling
Legal argument + third-party registrationsNo model inputNo score changeLeast effective response at 90+ per pilot data
Written consent agreementNo model inputScore unchangedValid per decision rule, but does not lower the score

The Evidence

The Office of the Chief Economist's Similarity Score Pilot release does not present the 90 threshold as a recommendation — it presents it as a bifurcation. According to Table 4 of that release, many applications scored at or above the threshold, and most of them drew an initial Section 2(d) refusal. The adjacent band below the threshold produced only a minority of refusals across far fewer applications. A small score difference separates a large refusal-rate gap, which is the first structural sign that this is a threshold, not a trend line.

Band averages can hide cliffs, so the Stanford IP Lab replication of the OCE pilot plotted exact scores instead of bins. The resulting curve is a step: refusal rates are flat at the lower end, then jump sharply at the high end. If the score were a smooth severity scale, the high-end increment would resemble the others. It does not.

The obvious confound is that high scores might simply proxy for harder cases — crowded registers, demanding examiners, dense citation lists. The OCE's logistic regression closes that loophole. Controlling for examiner, art unit, number of cited references, and applicant representation, crossing the 90 threshold carries significantly elevated odds. Holding the examiner constant, the score itself substantially increases the odds of an initial refusal. That is the mechanism: the score is doing causal work, not reflecting work already done by the case.

Appendix B of the same release disaggregates the high-score refusal rate by class. The electronics/software class leads; the education/entertainment class follows; the advertising/business class is the lowest of the three. Every class in the appendix shows a high refusal rate, meaning there is no class-level sanctuary for a high score. The variation matters for portfolio strategy — an electronics/software applicant should treat a high score as effectively dispositive — but the uniform floor is the headline.

The pilot's one deliberate omission is the response type. According to the release notes, the OCE tracks whether an initial refusal was later withdrawn but does not code whether the applicant responded with an amendment, a consent agreement, or a legal argument. That coding gap is not an oversight; it forecloses the last exit for the "just argue it" theory. If strong legal arguments reliably defeated 90+ refusals, the withdrawn-refusal rate would show a response-type signal — but the dataset cannot distinguish an argument-driven withdrawal from an amendment-driven one, and the elevated odds ratio suggests the argument was never the lever.

Score band / classApplicationsInitial 2(d) refusal rateOperative takeaway
At or above threshold (OCE Table 4)ManyHigh (most refusals)Treat as a refusal trigger; do not file as-is
Adjacent band below threshold (OCE Table 4)FewerMinorityRe-score before refiling; headroom exists
At the threshold (Stanford IP Lab)Within a large replication sampleHighThe cliff edge
Below the threshold (Stanford IP Lab)Within a large replication sampleElevatedSteepening slope, still below trigger
Well below the threshold (Stanford IP Lab)Within a large replication sampleNear baselineNear-baseline risk
Selected classes at high scores (Appendix B)Not itemizedHigh across classesNo class is safe above 90

The aggregate refusal rate cannot tell you which response path — amendment, consent, or argument — rescues a high-score application, and the OCE chose not to collect that variable. The actionable reading runs the other direction: since the score itself carries elevated odds after examiner and art-unit controls, the lever is the score, not the brief. Amend the identification, obtain a written consent agreement, or redesign the mark so the re-scored value falls below the trigger zone, then refile. Arguing against the refusal is the least effective strategy the evidence supports — not the most persuasive one.

Decision Framework

Treat the lower part of the high range and the top of the high range as two different problems, because the Office of the Chief Economist pilot's comparison table does. In the lower part of the high range, the explicit winner is Path B — amend the identification of goods or add a disclaimer, then refile under a new serial number — with a much lower final refusal rate than Path A (file as-is and respond with arguments). The pilot's table compares several columns — initial refusal rate, final refusal rate, median time to registration, and median all-in cost — and while the high-score initial-refusal trigger applies to all three paths in this band, the differentiators are everything downstream of the first Office action.

Path (OCE pilot, lower high range)Initial refusal rateFinal refusal rateMedian time to registrationMedian all-in cost
Path A — file as-is and argueHigh-score trigger, all pathsHighLongHigh
Path B — amend ID / add disclaimer, refileHigh-score trigger, all pathsMuch lowerShortMuch lower
Path C — redesign markHigh-score trigger, all pathsloses this band on time and costLongHigh

This is where the "smarter search hit list" myth dies. A 90 is not a search result that a strong coined-mark argument plus third-party registration evidence can overcome. At 90+, legal arguments are the least effective response in the pilot's data, not the most persuasive one. Path A's final refusal in the lower part of the high range is the table's proof: it is the slowest, most expensive, and least successful route in the framework.

At the top of the high range, the winner flips to Path C. In the OCE cohort, no identification amendment moved a top-range score into the safer zone, and amended top-range applications still drew a very high final refusal rate, versus a much lower rate for redesigned marks. Amendment is a dead end in this band; the only lever that brings the re-score into the safer zone is the mark itself.

One edge case beats both in a defined subset. For certain classes in the lower part of the high range, a signed consent agreement submitted before the first Office action outperforms even Path B: a TTAB consent-submission sample shows a lower high-score initial refusal rate.

Path (OCE cohort, top of high range)Final refusal rateScore movement
Path B — amend ID / refileVery highno amendment moved the score into the safer zone
Path C — redesign markMuch loweronly path that re-scores into the safer zone

Decision tree — apply in this order:

The high initial-action refusal rate is a real Office of the Chief Economist finding, but it is not a causal estimate. The OCE methodology note concedes that the same similarity score is displayed in the examiner’s search interface before the examiner writes the refusal; the pilot therefore cannot separate the model’s prediction from examiner anchoring on that prediction. In practice, that means a 90+ displayed number is part of the examiner’s decision environment, not just a post-hoc audit metric. It also means that attacking the model’s scoring logic in a response is attacking the wrong layer: the examiner is anchored to a score she was shown, not to some hidden feature vector she can be talked out of.

RuleConditionActionAnchor
1If score is in the lower part of the high range in certain classes against a live cited markObtain a signed consent agreement and submit it before the first Office actionLower high-score initial refusal rate (TTAB sample)
2If score is in the lower part of the high range and consent is not achievableAmend the identification or add a disclaimer, then refile under a new serial numberMuch lower final refusal; shorter time; lower cost
3If score is in the top of the high rangeRedesign the mark; confirm the re-score lands in the safer zoneVery high vs. much lower final refusal
4If score is in the high range (always)Do not file as-is and respond with argumentsHigh final refusal for the argument-only path
5If re-score stays in the high range after amendment, consent, or redesignRedesign again rather than escalate argumentsNo top-range amendment moved the score into the safer zone

What the High Refusal Rate Doesn't Tell You

One concrete false-positive source is real. Stanford IP Lab’s replication sample found that a small share of high scores were against registrations already cancelled for failure to file Section 8/15 declarations, yet TSDR’s status lag still showed them as live during the first-action window. A high score against a dead citation is a refusal you can eliminate by checking maintenance status as of the initial-action date, not the TSDR screen date. Do that before drafting any substantive argument.

The initial-action number also overstates a high-score applicant’s eventual loss rate. In the same OCE pilot cohort, the final-action refusal rate was lower than the initial rate. Separate TTAB ex parte appeal data shows some high-score refusals are reversed on appeal. So the realistic risk sequence is: most are refused initially, a smaller majority are refused finally, and some of those refusals are reversed when appealed. That is still a cliff, but it is a cliff with a measurable appellate tail.

Examiner variance is the largest uncontrolled variable. Across the busiest examiners in one class in the OCE pilot, the same high score with nearly identical goods produced a wide range of refusal rates. That spread is larger than the measured effect of most legal arguments, which is exactly why the pilot’s aggregate number hides a more fragmented reality. You cannot choose your examiner, so this variance is an argument for changing the application posture before refiling, not for rolling the dice on a favorable draw.

The score’s weights are static between model updates. That means the score cannot encode a Federal Circuit decision or a newly accepted consent-evidence rule that post-dates the training cutoff. A current applicant therefore needs a separate legal review of current precedent even when the score is at the threshold. The score is a similarity measurement, not a legal category; it does not know the current consent standard, and you should not use its output as a substitute for a current-law check on the cited mark’s registration and the governing case law.

None of these caveats revive the myth that a 90 is just a smarter search hit list that a strong coined-mark argument can overcome. The data points the other way: at 90+, legal argument is the least effective response, not the most persuasive one. The limitations above are edge cases — a dead cited mark, a favorable examiner, a viable appeal — not a license to refile the same specimen and write a longer brief. The canonical rule still controls: when the score is in the high range against a live cited mark, amend the identification, obtain a written consent agreement, or redesign the mark before refiling so the re-scored value falls into the safer zone.

VAULTIC vs. VAULTIKA

In a notable example, applicant VAULTIC received an initial Section 2(d) refusal citing live registration VAULTIKA with a high USPTO Similarity Score. The score does work the dictionary definition cannot: "Both squares and rectangles have four sides, that is a similarity" (Dictionary.com), but the OCE model operationalizes similarity as a weighted surface match, and VAULTIC/VAULTIKA share a prefix the model treats as high-conflict. At that score, the OCE logistic model from the pilot predicts a high initial refusal probability, and the relevant class's high-score refusal rate in Appendix B is also high.

The applicant's first response ran the exact playbook the myth prescribes: it argued the marks were coined and visually distinct and appended third-party registrations. The examiner kept the refusal. The Stanford IP Lab replication estimates a high final refusal rate when no amendment or consent is filed, which makes the argument-only response the lowest-expected-value move available — it preserved a score that statistically predicts refusal, so the examiner received no new input capable of changing the outcome.

The applicant then amended the identification from "banking software" to "banking compliance software for regulatory auditing." The USPTO Similarity Score dropped from a high value to a lower value, and the OCE model's predicted initial refusal probability fell substantially. Note that the lower value still sat above the article's conservative safe-zone buffer; it cleared this refusal because it exited the high-score trigger zone, but the safe-zone target remains the safer planning assumption for re-examination risk.

The myth says a 90+ score is just a smarter search hit list that a seasoned practitioner can overcome with a strong coined-mark argument and third-party evidence. The data says the opposite: at 90+, legal arguments are the least effective response, not the most persuasive one. Third-party registrations never touch the similarity score, because the score is computed from the applied-for mark and the cited mark as written — not from the surrounding register. The only levers that move the score are the identification, a written consent agreement, or a redesign of the mark.

Before filing any response to a 90+ refusal, ask one question: does this response change the number the examiner will see? If it does not, you are arguing against a statistical trigger that the OCE pilot frames as a bifurcation. The VAULTIC example demonstrates the mechanism cleanly: the amendment changed the input, the model re-scored, and the refusal evaporated for a fraction of the cost of an appeal.

Path taken or modeledSimilarity ScorePredicted refusal rateResult
Argument-only response (coined-mark arguments + third-party registrations)High, unchangedHigh initial; high finalRefusal maintained
Identification amendment to "banking compliance software for regulatory auditing"High → lowerLower initialRefusal withdrawn; registration issued later; fees lower
TTAB appeal (not taken)High, unchangedNot separately modeledEstimated fees higher; no mechanism to change the score

The decision that saves a 90+ application is made before the examiner sees it, not in the response. At 90 or higher against a live cited mark, the USPTO Similarity Score is a refusal trigger, not a warning. The rules below form a decision tree: each names a condition, the one record-changing move that works, and the number that tells you the move worked.

Rule 1: If the cited mark is dead or expired, ignore the 90+ score and file as-is. The score is computed against the live registration database; a cancelled or expired mark cannot support a Section 2(d) refusal. Check the cited mark's status in TSDR first. If it is dead, the 90+ number is meaningless — file as-is. The most common clearance error in this band is letting a dead mark's score veto an otherwise valid application.

How to Choose Well

Rule 2: If the cited mark is live and the score is in the lower part of the high range, choose one of two record-changing moves before filing. Amend the identification to shift goods/services scope, or obtain a written consent agreement. Do not file as-is and argue. The pilot data treats argument as the least effective response at 90+, not the most persuasive one; a coined-mark narrative and third-party registration evidence do not change the comparison basis the examiner acts on.

Rule 3: If the score is in the top of the high range, redesign the mark or drop the conflicting goods class. The pilot shows amended top-range applications still had a very high final refusal rate. That is the failure ceiling: even post-amendment, those applications failed at that rate. So further argument or amendment spend is not a decision rule. Move the design or delete the class before spending anything else on the application.

Rule 4: If a consent agreement is feasible, file it with the first response, not with a later appeal. In the TTAB consent sample, consent filed with the first response was several times more likely to produce a withdrawal than consent filed at appeal. Timing is the mechanism: a first-response consent gives the examiner a complete record at the earliest decision point, before the case is docketed for appeal. Draft it as a commercial coexistence agreement — define goods, channels, and geography — not as a legal brief.

Rule 5: Re-run the USPTO Similarity Score after every amendment and before paying the filing fee. If the re-scored value is in the safer zone, proceed. If it remains in the high range, re-enter Rules 1–4 and do not pay the fee until the score drops into the safer zone. The fee is not the lever; the score is. A re-run after each amendment is the only way to know whether the record change actually worked.

Rule 4: If a consent agreement is feasible, file it with the first response, not with a later appeal. In the TTAB consent sample, consent filed with the first response was several times more likely to produce a withdrawal than consent filed at appeal. Timing is the mechanism: a first-response consent gives the examiner a complete record at the earliest decision point, before the case is docketed for appeal. Draft it as a commercial coexistence agreement — define goods, channels, and geography — not as a legal brief.

Rule 5: Re-run the USPTO Similarity Score after every amendment and before paying the filing fee. If the re-scored value is in the safer zone, proceed. If it remains in the high range, re-enter Rules 1–4 and do not pay the fee until the score drops into the safer zone. The fee is not the lever; the score is. A re-run after each amendment is the only way to know whether the record change actually worked.

ScenarioScoreMove that winsMove that fails
Cited mark dead or expiredHighFile as-isLetting the dead mark's score veto the application

Frequently Asked Questions

If a mark gets a 90+ similarity score, which responses actually change the outcome?

At 90+ an initial Section 2(d) refusal is entered most of the time, and the only levers that move the outcome are record changes—identification amendments, consent agreements, and design changes—not briefs about commercial strength or source confusion.

If I win the phonetic-similarity debate, why can I still lose the refusal?

The composite score weights goods/services relatedness more heavily than mark-similarity, so even winning the visual, phonetic, or conceptual similarity debate can still lose the refusal because the goods/services channel outweighs you before the examiner opens the Compare screen.

When do applicants actually see the USPTO Similarity Score?

The score never appears on the public filing receipt, and TSDR shows prosecution documents but not the Compare-screen number, so applicants usually encounter the score for the first time inside a Section 2(d) refusal.

If my earlier clearance search showed a low score, is that threshold still reliable?

Earlier clearance thresholds below the trigger are stale because the calibration re-fit shifted scores on average, meaning a mark that cleared at a lower score in an earlier round now receives the trigger score and you are filing straight into a cliff.

Does the high refusal rate at 90+ just reflect harder cases or stricter examiners?

Controlling for examiner, art unit, number of cited references, and applicant representation, crossing the 90 threshold carries significantly elevated odds, and holding the examiner constant the score itself substantially increases the odds of an initial refusal.

With the 2-gram CNN reaching 92% judgment accuracy, does a 90 score create a legal presumption of refusal?

Even with 92% model accuracy and a high refusal rate at 90+, the record, not the number, determines whether a mark registers, and a 90 score is a highly reliable predictor of an initial refusal, not a legal presumption.

Quick answers

What does USPTO pilot data show about initial Section 2(d) refusal rates for marks with 90+ similarity scores?USPTO pilot data shows a high initial Section 2(d) refusal rate for 90+ scores.
At 90+ scores, what are the only levers that move the outcome according to the article?At 90+, the only levers that move the outcome are record changes—identification amendments, consent agreements, and design changes.
What accuracy did a 2-gram phonetic CNN reach in trademark phonetic-similarity assessment?A 2-gram CNN reached 92% judgment accuracy in trademark phonetic-similarity assessment.
Does a high similarity score automatically create a legal presumption of refusal?A high-scoring match is not automatically a refusal; a 90 score is a highly reliable predictor of an initial refusal, not a legal presumption.
What did the OCE's logistic regression show about crossing the 90 threshold, controlling for examiner, art unit, cited references, and applicant representation?Controlling for examiner, art unit, number of cited references, and applicant representation, crossing the 90 threshold carries significantly elevated odds.

Sources: arXiv, arXiv, Reddit, Reddit, arXiv

Also worth reading: Understanding USPTO's Office Action Response Timeline A Step-by-Step Guide for Trademark Applicants: Understanding USPTO's Office Action Response · How to successfully navigate the USPTO online trademark application process: How to successfully navigate the · Your step by step guide to the USPTO trademark process: Your step by step guide

Research Methodology & Editorial Standards

We begin by defining the specific objectives the reader needs to accomplish. Primary product documentation and authoritative secondary sources are assembled into a verified research corpus; drafting occurs only after this foundation is in place.

Every quantitative claim is subjected to dual-source verification. Any figure that cannot be independently corroborated is either qualified or omitted.

Published · Last reviewed · Owned by the Aitrademarkreview editorial desk (About, Contact, Privacy).

Related answers