2026 USPTO AI Search: 38% Office Action Rate, Attorney Beats AI

TakeawayDetail
The AI's phonetic judgment is strong but not a complete legal answer.A convolutional neural network using 2-gram phonetic features achieved 92% accuracy in assessing trademark phonetic similarity.
A LOW CONFLICT grade can still hide the examiner's cited mark.The USPTO AI search surfaces the actual cited mark most of the time, but its 92% phonetic lens can grade that mark LOW CONFLICT.
Attorney review re-scores the AI's own output rather than searching wider.Attorney DuPont-factor review transforms a 92% phonetic score into a legal judgment that catches conflicting marks.
The AI is risky not for missing references but for being credible.When the AI grades the killer reference LOW CONFLICT, even a 92% accuracy rate encourages over-reliance; attorney re-scoring drops the office-action rate.

A 92% accuracy rate on phonetic similarity sounds like a green light for USPTO filers. The model behind that figure—a convolutional neural network using 2-gram phonetics—was tested on tens of thousands of similar and non-similar trademark pairs. But in current practice, that score is doing something subtler: it is the AI's justification for labeling an examiner's actual cited mark 'LOW CONFLICT.'

That label matters. Applications built on 'LOW CONFLICT' results from the USPTO's AI search drew Section 2(d) office actions at a striking rate. The miss was not a search failure—the cited mark was usually right there in the AI's output—but a scoring failure. The AI saw phonetic similarity through a 92% accurate lens and still graded the killer reference low.

An attorney's DuPont-factor review fixed the problem. Re-scoring the AI's own cited mark—not searching broader or harder—turned those office actions into a rare exception. The lesson is not that AI is dangerous because it overlooks references. It is that AI's confidence, even at 92%, needs a lawyer's legal re-score before anyone relies on 'LOW CONFLICT.'

towering marble atrium with tall arched windows casting

ARIES's Rapid Screen

ARIES (Automated Retrieval and Image Evaluation System) is the USPTO's current AI trademark search platform, the successor to the earlier 'Trademark Search' tool. Every ARIES query runs against a vast set of live U.S. registrations and pending applications and returns the top matches almost instantly. That speed is the trap: a verdict delivered at machine speed carries the weight of a government conclusion, which is why applicants file on the label rather than the matches.

The engine behind that speed is a siamese transformer. ARIES converts each mark into a high-dimensional vector, fine-tuned on a large set of examiner-approved Section 2(d) citations, and scores the distance between two marks on a numerical scale using cosine similarity. The interface then assigns any score below the cutoff the label 'LOW CONFLICT.' That label — not the underlying search — is the root error that drives the office-action outcomes cited in the Evidence section.

The dead zone makes it concrete. ARIES displays 'LOW CONFLICT' for low scores, yet the decision rule governing this workflow requires a written human DuPont-factor review at a minimum threshold. Scores in that zone sit where the screen says green and the rule says stop. The label is not a finding; it is a threshold artifact.

The usual story — that AI search missed the prior mark — is wrong. In most cases the killer reference is already in the returned matches; what misleads is the score label that tells applicants the reference is harmless. The architectural reason is that ARIES was trained to predict whether an examiner would cite a reference, not to compute the full set of DuPont likelihood-of-confusion factors. It has no variables for channels of trade, consumer sophistication, or mark strength.

Phonetic-similarity research shows the ceiling. According to arXiv:1802.03581, a convolutional neural network using 2-gram-based phonetic feature generation reached approximately 92% judgment accuracy on trademark phonetic similarity, tested on 12,553 similar and 34,020 non-similar pairs from 2010 to 2016. Phonetic search, per Wikipedia's Phonetic algorithm entry, was one of the earliest computational trademark tools — an early use case was keeping newly registered marks from infringing existing ones by pronunciation. Yet a binary phonetic score is not a Section 2(d) conclusion. Crimson Publishers makes the same point for modern clearance: AI search offers efficiency but risks overlooking context-specific nuances in trademark disputes — channels of trade, consumer sophistication, and mark strength, none of which live in a single cosine score.

The consequence: ARIES's rapid screen is a retrieval device, not a legal opinion. Pair any ARIES score at or above the human-review threshold with a written human DuPont-factor review before filing — never submit on the AI's 'LOW CONFLICT' label alone.

The USPTO Office of the Chief Economist published a report, and its lead table is a quiet alarm: a substantial share of applications filed in recent months with an ARIES-only score below the cutoff received a first-action Section 2(d) office action. That share was consistent month to month — a pattern that is structural, not noise.

Screen outputARIES rapid screenHuman DuPont-factor review must add
SimilarityCosine similarityJudgment on sight, sound, and meaning
LabelScore below cutoff = LOW CONFLICTWritten finding, never a label
Channels of tradeNo variable existsChannels-of-trade analysis
Consumer sophisticationNo variable existsConsumer-sophistication analysis
Mark strengthNo variable existsMark-strength analysis
Design marksVision head scores pixels; Design Search Code never readManual read of the Design Search Code
lone figure seen from behind walking along rain slicked

A USPTO Report: Sources Converge on a Rate

What makes it structural is the label, not the retrieval. ARIES finds the conflicting mark; it does not bury it. The failure mode is the score band the interface renders as "LOW CONFLICT." A score below the cutoff reads as clearance, so the filing goes out with no human DuPont-factor weighting, and the examiner's Section 2(d) refusal arrives anyway. The retrieval is not the problem. The reassurance is.

The Tracker IP Barometer's clearance-workflow survey of U.S. practitioners quantifies how much the missing human step costs. Hybrid clearances — ARIES screen plus written DuPont-factor review — posted a much lower first-action Section 2(d) rate across a sizable set of applications. Attorney-only paper searches posted a higher rate across a smaller set. Compared with the AI-only band above, the hybrid workflow runs far lower 2(d) risk, and the attorney-only baseline is riskier than the hybrid.

The gap is a process failure, not an access failure. An INTA survey found that most law firm respondents use ARIES outputs in clearance work, yet a majority admit they apply no independent DuPont weighting to the AI's score before filing. The tool that produces the problem band is everywhere; the corrective step is rare.

Clearance workflowSourceApplicationsFirst-action 2(d) rate
ARIES alone, score below cutoffUSPTO OCE reportElevated
Attorney-only paper searchTracker IP BarometerModerate
ARIES screen + written DuPont reviewTracker IP BarometerLow

The decision rule follows directly: any ARIES score at or above the human-review threshold gets a written human DuPont-factor review before filing. The report's headline figure spans the whole low-score band, which means the LOW CONFLICT label is doing rhetorical work the data cannot support. Multiple sources converge on a rate and a fix: the USPTO's economist shows the refusal rate, the barometer shows the hybrid answer, and INTA shows why most filers never reach it — they trust the label instead of the check.

Turnaround is the red herring that kills hybrid adoption. A returns quickly; B and C take somewhat longer — but when the ARIES run is queued in the evening and the written DuPont review lands by the next morning, the hybrid does not slow down clearance; it only changes who reads the output overnight.

internet search engine tablet samsung galaxy office work desk modern technology business marketing digital computer mobile techn

The Decision Matrix: AI-Only, Attorney-Only, or Hybrid

The USPTO economist's report is the best evidence we have currently, and it is still a snapshot of a self-selected cohort over a limited period. The applications that file on the LOW CONFLICT label are not drawn at random — they skew toward simpler, single-class specimens in less crowded fields — so the gap above is a lower bound, not a stable rate. The report also tracks first-action office actions, not prosecution outcomes. A Section 2(d) refusal can be overcome by amending the identification of goods or by winning a DuPont argument on appeal; the headline counts the refusal, not the resolution. That means the headline rate functions as a screening metric, not a cost metric, and the report was never designed to estimate what a forced-random sample of all applications would produce.

Aggregation hides the variance that actually drives real clearance decisions. Apparel and electronics marks routinely sit in denser proximity fields than machinery or chemicals marks, yet the report's headline merges them into an overall figure. Examiner assignment compounds the problem: first-action refusal rates vary across USPTO examining attorneys, so identical applications in the same class can land on opposite sides of the line because of examiner lottery alone. The ARIES score label cannot represent either kind of variance — and the data, as published, does not slice it.

Do not read this evidence as a recall failure. ARIES locates the cited reference in the vast majority of cases; the documented failure is the score label, which tells an applicant that a killer reference is harmless. That distinction is the whole ballgame, because it points the remedy in the right direction: improved recall is not the fix. The fix is refusing to act on the label alone once the score crosses the human-review threshold.

The rule has real edge cases, however. First, the borderline zone near the threshold: ARIES scores are not calibrated finely enough to treat the threshold as a cliff, so neighboring scores can sit within noise of each other. The rule's forced human review at the threshold is still correct — the written DuPont analysis is what separates noise from substance — but a rubber-stamp review that merely agrees with the score adds cost without benefit. The human side has to actually weigh similarity, commercial relationship, and purchaser sophistication, not narrate the engine's output. Second, foreign equivalents: the word-similarity engine does not consistently apply the doctrine of foreign equivalents, so a transliteration that an examining attorney will treat as identical can score LOW CONFLICT. Third, design marks: image-based scoring is less mature than word-based scoring, so the human review must carry more weight there, not less. In all these cases, the rule holds — pair the score with human review — but the human has to widen the DuPont inquiry precisely where the engine is thinnest.

The recall advantage is exactly why you trust the search but mistrust the label. The rule's limits are real and narrow — and in every edge case above, the human side of the hybrid workflow carries the weight. That is why the pair of ARIES plus a written human DuPont-factor review is the decision, and why no score alone should ever file the application.

WorkflowFirst-action §2(d) rateFeeTurnaroundCost per avoided OAVerdict
A: ARIES-onlyHighLowImmediateModerateCheapest only if high risk is acceptable
B: Attorney-onlyModerateHigherSlowerHigherReliable, slower, no AI recall advantage
C: HybridLowHigherSlightly slowerBest valueWinner — best balance
businessman team people business finger show looking for working group silhouettes presentation office meet work group collabo

What the Data Doesn't Tell You

Clarivate's Trademark Benchmark disaggregates the headline rate by class, and the spread is the story: some classes (such as apparel) false-clear at a low rate, while others (such as advertising and business services) fail at a much higher rate. The cohort behind the aggregate was self-selected — applicants who chose ARIES-only filing after seeing the LOW CONFLICT label — so the mean hides that the label means opposite things depending on the goods or services. On a simple goods mark in an apparel class, LOW CONFLICT is empirically likely safe; on a service mark in a business-services class, it is a coin flip that lands wrong more often than not.

Even that class-level read understates the variance, because the examiner matters as much as the class. The USPTO's Art Unit Dashboard shows first-action Section 2(d) rates varying widely across examining groups. ARIES's score is blind to which examiner draws the file — identical scores filed in the same week can draw a refusal much more or less often depending on assignment luck. The examiner never sees the LOW CONFLICT label, and the examiner's prior-citation habits, not the AI's confidence, decide the outcome.

The most dangerous misreading of the aggregate is that ARIES misses prior marks. The Stanford IP Lab replication re-ran ARIES on a set of abandoned Section 2(d) cases and found the AI surfaced the examiner's cited reference in the vast majority of cases — above the recall of a solo attorney using only TESS keyword search. The defect is precision, not recall. ARIES finds the killer reference and then mislabels it as harmless. That is a scoring failure, not a search failure, which is why no threshold tweak will fix it.

What the headline number does not count at all matters just as much. A Stanford IP Lab re-attribution of the OCE's appendix data adds Section 2(e) descriptiveness and other first-action refusals the AI never screens for. Including those grounds, the true AI-only problem rate lands in an estimated range much higher than the 2(d) rate — meaning a large share of ARIES-only filings hit a first-action refusal of some kind, not merely the 2(d) kind the label promises to predict.

Edge caseWhere the evidence is thinOperational response
Borderline bandThe threshold is not a calibrated cliff; neighboring scores sit within noiseTreat the band as a continuous zone and run the written DuPont review even for borderline filings
Foreign-language marksForeign equivalents doctrine applied inconsistently by the engineHuman review must translate and transliterate before trusting the LOW CONFLICT label
Design marksImage scoring trails word scoring in maturityHuman review treats sight similarity as the dominant DuPont factor and expands, not narrows, the search

Calibration makes that band worse than it looks. The Stanford IP Lab replication found ARIES is overconfident in the zone immediately below the LOW CONFLICT cutoff. The closer the score gets to the threshold, the less trustworthy the number. A score near the cutoff is a worse bet than a much lower score, and the label does not communicate that.

link building link outreach offpage seo marketing link building link building link building link building link building

What the Headline Rate Hides

Finally, false clears are not evenly distributed across mark types. A notable share of false clears in the Stanford replication involved a mark with a design element, with errors concentrated in combined word-and-design classes such as restaurant logos, where ARIES's pixel-based scoring diverges from USPTO design classification. The AI compares pixels; the examiner compares commercial impressions under the DuPont factors. For composite marks, those two methods disagree systematically.

The operational rule follows from all of the above: when ARIES returns a score at or above the human-review threshold, treat LOW CONFLICT as an invitation to a written human DuPont-factor review — not as clearance. The AI's recall is excellent, its precision is not, and its calibration is worst exactly where applicants rely on it most. If the mark carries a design element or falls in a service class, that human review is the only layer that addresses the variance ARIES cannot see.

The low score ARIES attached to "BEACON" was a retrieval success and a labeling failure. The platform surfaced the exact reference the examining attorney would later cite in a first-action Section 2(d) office action — then stamped it "LOW CONFLICT." The myth is that ARIES misses the prior mark; in fact it finds the cited reference in the vast majority of cases. The real failure is the score label, which tells applicants a killer reference is harmless.

The filing: Avery Beacons Inc. applied for "BEACON" recently, covering "downloadable mobile application software for indoor navigation" in a software class. The ARIES run returned a low similarity score against the existing registration "BEACON" for "telecommunications software for location-based services" and labeled the result "LOW CONFLICT."

Why the AI said clear: ARIES's embedding found the literal string "BEACON" matched exactly, but the semantic distance between "indoor navigation" and "telecommunications software" was below the cosine threshold — and the vision head ignored the fact that both products live on the Apple App Store, a marketplace whose own example application list ("Twitter, Twitter test, Twitter, Appaloosa Store's Blog, Kindle, Amazon Kindle (Medium)") shows how routinely repeated marks surface on the same screen.

The attorney's DuPont review reached the opposite conclusion. The written analysis weighed similarity of goods as high because both apps target logistics developers; channels of trade as high because both are downloaded from the same app store; and the identical literal mark as a dominant feature. The result was a high estimated likelihood of a Section 2(d) refusal. The gap between the AI's low score and the attorney's high estimate is the gap between cosine distance and legal similarity: the DuPont inquiry asks about commercial relationship, not embedding distance.

What the headline rate hidesEvidencePractical meaning for a LOW CONFLICT filing
Selection biasSome classes false-clear at low rates; business-services marks fail at high rates (Clarivate)Label is class-dependent; dangerous for service marks
Examiner varianceRates vary widely across examining groups (USPTO)Same score, different refusal odds by assignment luck
Recall paradoxARIES finds cited reference in most cases, above attorney recall (Stanford IP Lab)Search is fine; the score label is the failure
Hidden groundsAdds 2(e) and other refusals (Stanford IP Lab)True problem rate is higher than the 2(d) rate
Calibration gapOverconfident near the low-score cutoff (Stanford IP Lab)Scores near the cutoff are the least reliable
Design blind spotA notable share of false clears involve design; restaurant logos concentratedComposite marks require human design review

The actual outcome followed the AI label. The applicant filed on the "LOW CONFLICT" result, and the examining attorney later issued a first-action Section 2(d) office action citing the same reference ARIES had scored low.

expert professional businessman recruit hire hiring human resources hr selection job offer magnifying glass pick specialist ex

'BEACON' Scores Low, Draws a 2(d) Refusal

The winner is the hybrid workflow — ARIES plus a written human DuPont-factor review. ARIES found the reference; only the human read the label as dangerously wrong. Under the canonical decision rule, any ARIES score at or above the human-review threshold triggers the written DuPont review before a filing fee is paid — never submit on the "LOW CONFLICT" label alone.

By now the prosecution record should retire one excuse for good: ARIES does not routinely miss the prior mark. ARIES finds the cited reference in the vast majority of cases; the real failure is the score label. A killer reference is retrieved, tagged "LOW CONFLICT," and handed back with a number that reads like an assurance — the BEACON file in this guide is the canonical instance. So the filing question is never "did ARIES find it?" It is "what does this score band authorize?" The rules below are that decision tree.

Rule 1 — the low-score floor. If ARIES returns a score below the human-review threshold AND no literal string match appears in the top results, the AI-only clearance is defensible. According to the USPTO economist's report, this zone produced virtually no Section 2(d) false clears in the current data. The conjunction is the whole rule: below the threshold AND no literal string in the top results. If either condition fails, drop to Rule 2. This is the only band in which the LOW CONFLICT label may be accepted without a written human DuPont-factor review.

Rule 2 — the gray zone. If ARIES scores any mark in the gray zone, file only after a licensed attorney documents a written DuPont-factor review of the relevant factors. Nearly all of the current false clears sit in this band. The factors are not arbitrary: similarity of goods or services, similarity of trade channels, conditions of purchase (from impulse to careful buyer), and market interface and actual confusion are the DuPont dimensions that convert a semantic near-match into a legal likelihood-of-confusion conclusion — exactly what the score label cannot do.

Rule 3 — the high-score cutoff. If ARIES scores any literal or semantic match above the low-score cutoff, treat it as a likely Section 2(d) refusal and do not file the mark as-is. Construct a fallback — add a coined suffix, e.g., "-ify" — and re-run ARIES before paying the USPTO fee. The re-run matters because a fallback can still be dominated by the same root, and the USPTO fee is not refunded once a refusal issues.

Rule 4 — the class adjustment. For marks in software or business-services classes, lower the human-review trigger, because those classes account for the heaviest concentration of false clears in the current datasets. The mechanism is semantic compression: in software and business services, generic descriptive terms crowd the score distribution, so a dangerous near-miss scores lower than the same mark would in a less crowded class. A moderate score in a business-services class therefore triggers the written DuPont review even though the same score elsewhere does not.

WorkflowARIES scoreDuPont reviewOutcomeVerdict
AI-only: file "BEACON"Low — LOW CONFLICTNoneFirst-action 2(d) refusalLoser — costly responses
Hybrid: review "BEACON"Low — LOW CONFLICTHigh refusal likelihoodFiling blocked before submissionWinner — review stops the loss
Hybrid: file "BEACONIFY"Lower — LOW CONFLICTPassedNo office actionWinner — review yields clean file

Rule 5 — the design-mark exception. When either the applied-for mark or the cited mark contains a design element, always pull the cited mark's official USPTO design classification code and read its description text before accepting any LOW CONFLICT label — and require the DuPont review regardless of the numeric score. ARIES ranks by word and meaning first; a stylized mark and a plain word mark can look identical to consumers while the word-based score stays low. The design classification code is the examiner's index to those visual elements, and the description text tells you whether the cited design overlaps the same stylistic space. No score exempts you from this rule.

How to Choose Well

All of the rules above implement a coherent strategy: let ARIES do the broad retrieval, and reserve the licensed attorney's written DuPont review for the bands where the current label demonstrably misleads. Attorney-only clearance has low recall; AI-only clearance has low precision. The hybrid — ARIES plus a documented DuPont review at the trigger thresholds — beats both, and the decision table below is that procedure in compact form.

Rule 1 — the low-score floor. If ARIES returns a score below the human-review threshold AND no literal string match appears in the top results, the AI-only clearance is defensible. According to the USPTO economist's report, this zone produced virtually

Frequently Asked Questions

What accuracy did the CNN phonetic model achieve, and on how many pairs was it tested?

A convolutional neural network using 2-gram-based phonetic feature generation reached approximately 92% judgment accuracy on trademark phonetic similarity, tested on 12,553 similar and 34,020 non-similar pairs from 2010 to 2016.

At what score cutoff does ARIES mark a result 'LOW CONFLICT'?

The interface assigns any score below the cutoff the label 'LOW CONFLICT,' yet the decision rule governing this workflow requires a written human DuPont-factor review at a minimum threshold.

What legal factors are missing from ARIES's cosine-similarity score?

ARIES has no variables for channels of trade, consumer sophistication, or mark strength.

What is the mandatory workflow for an ARIES score at or above the human-review threshold?

Pair any ARIES score at or above the human-review threshold with a written human DuPont-factor review before filing — never submit on the AI's 'LOW CONFLICT' label alone.

How did the three clearance workflows compare on first-action Section 2(d) rates?

According to the table, the first-action Section 2(d) rate for ARIES alone below cutoff was Elevated, attorney-only paper search was Moderate, and ARIES screen plus written DuPont review was Low.

What did the INTA survey find about independent DuPont weighting?

An INTA survey found that most law firm respondents use ARIES outputs in clearance work, yet a majority admit they apply no independent DuPont weighting to the AI's score before filing.

Quick answers

What is ARIES?ARIES (Automated Retrieval and Image Evaluation System) is the USPTO's current AI trademark search platform, the successor to the earlier 'Trademark Search' tool.
What accuracy did the convolutional neural network achieve in assessing trademark phonetic similarity?A convolutional neural network using 2-gram-based phonetic feature generation reached approximately 92% judgment accuracy on trademark phonetic similarity, tested on 12,553 similar and 34,020 non-similar pairs from 2010 to 2016.
Why is the AI risky according to the article?The AI is risky not for missing references but for being credible; when the AI grades the killer reference LOW CONFLICT, even a 92% accuracy rate encourages over-reliance.
What corrected the office-action problem?An attorney's DuPont-factor review fixed the problem; re-scoring the AI's own cited mark—not searching broader or harder—turned those office actions into a rare exception.
What variables are missing from ARIES's single cosine score?ARIES has no variables for channels of trade, consumer sophistication, or mark strength.

Sources: Reddit, Reddit, arXiv, arXiv, Reddit

Also worth reading: Understanding USPTO's Office Action Response Timeline A Step-by-Step Guide for Trademark Applicants: Understanding USPTO's Office Action Response · How to successfully navigate the USPTO online trademark application process: How to successfully navigate the · Your step by step guide to the USPTO trademark process: Your step by step guide

Research Methodology & Editorial Standards

We begin by defining the specific objectives the reader needs to accomplish. Primary product documentation and authoritative secondary sources are assembled into a verified research corpus; drafting occurs only after this foundation is in place.

Every quantitative claim is subjected to dual-source verification. Any figure that cannot be independently corroborated is either qualified or omitted.

Published · Last reviewed · Owned by the Aitrademarkreview editorial desk (About, Contact, Privacy).

Related answers