Authenticating AI survey evidence before the Trademark Trial and Appeal Board (TTAB) has become one of the most contested evidentiary issues in trademark litigation since 2023. The short answer is this: an AI-assisted survey is admissible only if the proponent can establish, through competent testimony and documentation, that (1) the survey was designed and supervised by a qualified human expert, (2) the underlying data was collected from real respondents through verifiable means, (3) the AI tools used were disclosed, reproducible, and did not fabricate or distort responses, and (4) the methodology satisfies the reliability standards of Federal Rule of Evidence 702 as applied by the Board. A survey that cannot be tied to a testifying human expert who can be cross-examined about every stage of its creation will almost certainly be excluded or given no weight.
The Direct Answer: What Authentication Requires
Also worth reading: How do I register a trademark for my AI-generated voice or digital persona in 2026? · What are the rules regarding AI generated trademark registration eligibility in the United States? · How can lawyers implement C2PA standards to protect AI-generated content in trademark disputes?
Authentication under Rule 901 of the Federal Rules of Evidence, which the TTAB applies through TBMP Section 704, requires evidence sufficient to support a finding that the item is what the proponent claims it is. For a survey, that means proving the survey actually reflects responses from actual consumers, gathered under the conditions described in the report. When artificial intelligence enters the workflow — whether for questionnaire drafting, respondent screening, response coding, or synthetic data generation — each point where AI touched the process becomes a potential authentication gap.
The Board's position, consistent with Federal Circuit precedent on expert testimony, is that a survey must rest on scientifically reliable methodology. In practice, this means the party offering the survey should produce a declaration or deposition testimony from the survey expert who personally directed the work. That expert must explain the sampling frame, the screening criteria, the interview script, the data collection platform, quality-control measures such as attention checks and speeder removal, and any role played by machine learning models or large language models in generating or processing content. If the AI component cannot be explained by a human witness under oath, the evidence fails authentication regardless of how polished the final report looks.
A second requirement is disclosure. Under TBMP rules and the discovery obligations applicable to inter partes proceedings, parties must identify their experts and produce the materials relied upon. An undisclosed AI tool used to analyze open-ended responses, for example, may be treated as an unproduced testing instrument, giving the opponent grounds to move to strike the related testimony. Practitioners who treat AI as a mere internal convenience rather than a discoverable method do so at their peril.
Why AI Surveys Face Heightened Scrutiny
Traditional online surveys already carry known weaknesses: panel fraud, bot traffic, professional respondents, and straight-lining. The TTAB has repeatedly discounted surveys with poor incidence rates, inadequate screening, or suspicious completion times. Generative AI magnifies these concerns in three distinct ways.
First, AI can generate synthetic respondents. Some vendors market "synthetic samples" as a cheap substitute for live fielding. The Board has never accepted purely synthetic respondent data as probative of consumer perception, because trademark likelihood-of-confusion and genericness findings depend on evidence of how real marketplace participants actually think. A survey built on simulated personas tests the model's assumptions, not the market.
Second, large language models can contaminate question design. If a model drafts leading questions or embeds assumptions drawn from training data about the very marks at issue, the resulting bias may be invisible in the final instrument. Opposing counsel increasingly demand the prompts and drafts used during design to probe for this contamination.
Third, AI-based response analysis — sentiment classification, clustering of open-ended answers, automated coding — introduces error rates that must be validated. Without a documented human audit comparing AI coding against manual coding on a sample of responses, the accuracy of the analysis is unverifiable, and the Board may exclude it under Rule 702 reliability principles.
The Legal Framework: Rules 702, 901, and TBMP Standards
The TTAB applies the Federal Rules of Evidence in inter partes proceedings under 37 C.F.R. § 2.122. Rule 702, as amended in December 2023, requires that expert testimony be based on sufficient facts or data, be the product of reliable principles and methods, and reflect a reliable application of those methods to the case. The amendment's emphasis on the proponent's burden to demonstrate reliability by a preponderance makes it harder to get marginal surveys admitted.
Rule 901(a) requires authentication through evidence sufficient to support a finding of authenticity. For survey evidence, courts and the Board typically look for testimony from the person who conducted or supervised the research, business records from the fielding vendor, and metadata showing when and how interviews occurred. Rule 902(13) and (14), added in 2017, allow self-authentication of records generated by electronic processes and of data copied from electronic devices, but only when accompanied by a certification from a qualified person. These provisions are useful for authenticating raw response files, yet they do not substitute for expert testimony about methodology.
The TBMP also imposes proportionality and relevance limits. A survey must measure something legally relevant — confusion, secondary meaning, genericness, fame — using a recognized format such as Eveready-style formats, Squirt-style formats, or Thermos-style formats adapted to the issue. An AI-drafted survey that blends formats or asks compound questions invites objections that go to both admissibility and weight.
Comparison: Traditional Fielded Survey vs. AI-Assisted Survey vs. Synthetic Sample
| Feature | Traditional Human-Fielded Survey | AI-Assisted Survey (Human Respondents) | Fully Synthetic AI Sample |
|---|---|---|---|
| Respondent source | Live panel or intercept recruits | Live panel, AI used for design/coding | LLM-generated personas |
| Authentication path | Expert testimony + vendor records | Expert testimony + AI disclosure + validation | Generally not authenticatable |
| Typical cost | $30,000–$150,000+ | $20,000–$100,000 | $500–$10,000 |
| Fielding time | 2–8 weeks | 1–6 weeks | Hours to days |
| Risk of exclusion | Low if well-designed | Moderate; depends on disclosure | Very high; near-certain exclusion |
| Cross-examination exposure | Methodology only | Methodology plus AI pipeline | Cannot survive cross-examination |
| Weight given by TTAB | Full weight if reliable | Potentially full weight with validation | None or minimal |
Practical Steps to Authenticate an AI-Assisted Survey
Begin with expert selection and engagement. Retain a survey methodologist with credentials in consumer research or statistics who will personally supervise the project and testify. The engagement letter should make clear that the expert reviews and approves every stage, including any AI-generated drafts. This creates the testimonial anchor that Rule 901 demands.
Second, document the AI pipeline contemporaneously. Preserve the prompts used, the versions of any models involved, the outputs at each iteration, and the human edits made along the way. Maintain a change log. If opposing counsel later alleges that a model hallucinated a question or embedded bias, the log allows your expert to walk through exactly what happened. Undocumented pipelines are nearly impossible to defend.
Third, validate AI-driven analysis against human benchmarks. If a language model codes 5,000 open-ended responses into confusion categories, have two trained human coders independently code a random sample of at least 10 percent, calculate inter-coder agreement (Cohen's kappa above roughly 0.80 is a common benchmark), and compare the AI output against the human consensus. Report the agreement statistics in the expert report. This converts an opaque black box into a validated instrument.
Fourth, harden respondent verification. Use panel vendors with device fingerprinting, duplicate detection, geolocation confirmation, and post-survey verification questions. Remove speeders (commonly defined as completing in less than half the median time) and failed attention checks, and disclose the exact numbers removed. In one frequently cited pattern, boards discount surveys where more than 15 to 20 percent of completes are discarded without explanation, so transparency about attrition matters.
Fifth, prepare the expert for cross-examination on AI specifically. Expect questions about whether the model was trained on data containing the parties' marks, whether outputs were cherry-picked, and whether the expert could reproduce the results. Rehearse direct answers grounded in the documentation file.
Common Mistakes That Sink AI Survey Evidence
The most frequent fatal mistake is submitting a survey report without a sponsoring expert. A declaration from a paralegal or attorney who merely received the report does not authenticate methodology, and hearsay objections typically succeed. Every survey offered in a TTAB case needs a living, breathing witness.
The second mistake is hiding the AI. Discovery requests now routinely ask whether generative AI was used in preparing expert materials. Failure to disclose invites motions to compel, motions to strike, and in extreme cases adverse credibility findings that taint the entire case. Disclosure costs little; concealment can cost the motion.
Third, parties over-rely on sample size as a proxy for quality. A survey of 1,200 respondents recruited through weak screening is worth less than a clean study of 300 properly screened target-market consumers. Boards evaluate the sampling frame first: respondents must plausibly represent the relevant class of purchasers for the goods at issue.
Fourth, litigants confuse convenience with relevance. An AI tool that rapidly produces a "confusion rate" of 47 percent means nothing if the question format does not map onto the legal standard. Courts have long warned that survey percentages are not direct measurements of legal conclusions, and AI-generated dashboards can create a false aura of precision around flawed instruments.
Fifth, some parties field surveys too late. A survey commissioned after the close of fact discovery may be challenged as untimely expert disclosure, and the Board may exclude it or reopen discovery with cost-shifting. Plan survey work to fit the protective order and scheduling order deadlines, which in TTAB proceedings typically allocate 30 days for initial expert disclosures after fact discovery closes.
Timing, Cost, and Budgeting Considerations
Survey timing should be decided at the case-strategy stage, not after pleadings. In an opposition or cancellation, fact discovery generally runs several months, followed by expert periods. Commissioning the survey consultant early — ideally within the first 60 days of the proceeding — leaves room for pilot testing, which experienced methodologists treat as mandatory. A pilot of 25 to 50 respondents often reveals ambiguous questions, broken skip logic, or unintended AI artifacts before full fielding.
Budget expectations as of 2026: a defensible likelihood-of-confusion survey with 250 to 400 respondents typically costs between $40,000 and $120,000 all-in, including expert fees for design, supervision, report preparation, and deposition testimony. AI assistance can compress drafting and coding time by 30 to 50 percent, translating into modest savings of perhaps $5,000 to $20,000, but it does not eliminate the dominant costs of recruiting, incentives, and expert hours. Beware of vendors quoting under $15,000 for a "litigation-ready" survey; those prices usually signal thin samples, recycled panels, or synthetic augmentation that will not survive challenge.
Weigh cost against stakes realistically. In a dispute over a mark supporting a nine-figure brand, a $75,000 survey is routine. In a small opposition, consider whether other evidence — sales figures, advertising spend, consumer declarations, third-party registrations — might carry the day without survey expense. The Board decides cases on the totality of the record, and a weak survey can affirmatively hurt by highlighting what the record lacks.
Strategic Alternatives and Complementary Evidence
Surveys are not always necessary. Fame and strength can be shown through revenue data, advertising expenditures spanning multiple years, media coverage, and social media metrics, all authenticated through business-records foundations under Rule 902(11). Consumer declarations obtained through counsel-directed outreach can support secondary meaning, though they carry limited weight compared to randomized surveys because of selection bias.
Where budget permits only one major evidentiary investment, many practitioners prioritize a Teflon-format or Eveready-format confusion survey over voluminous documentary evidence, because likelihood-of-confusion is the dispositive issue in most oppositions. Conversely, in genericness disputes, a properly constructed genus-focused survey remains close to indispensable, and no amount of dictionary excerpts fully substitutes for it.
A hybrid strategy is often optimal: use AI internally to organize and summarize discovery documents, draft preliminary questionnaires for expert review, and code open-ended pilot responses, while reserving formal evidence for a conventionally fielded, expert-supervised study. This captures efficiency gains without placing the case's fate on an authentication fight. Parties on the receiving end of suspect AI surveys should serve targeted interrogatories about AI usage, request the prompt logs and model versions, and retain their own methodologist to critique the pipeline — attacks on authentication and reliability are frequently more effective than competing surveys costing twice as much.
Ultimately, the Board rewards rigor and candor. An AI-assisted survey that is fully disclosed, human-supervised, validated against benchmarks, and anchored to a credible testifying expert stands on equal footing with traditional research. Anything less risks exclusion, and with it, the argument the survey was meant to win.