The output that is already in exhibits
Machine-generated output reached this field before any rule was written for it. Four kinds recur.
- Similarity and duplication scores. A number expressing how much of one page's text matches another's, offered to support copying or scraping allegations.
- Language-model classification. A large language model asked to sort pages, listings, or reviews into categories — genuine or fabricated, machine-written or human-written.
- "Toxic" or "spam" link scores. A vendor's rating of an inbound link — a hyperlink from another site — or of a whole profile, offered as evidence that someone built harmful links at a rival.
- Visibility and share-of-voice indices. A composite of estimated positions, estimated volumes, and an assumed click curve, offered as a measure of what a party lost.
Every one is an opinion expressed as a number. It looks like a measurement, prints like a measurement, and charts like a measurement. It is the output of a model whose definition of the thing it scores is the vendor's, whose training data is undisclosed, and whose threshold was set by someone with no view of the litigation. An expert who pastes that number into an exhibit has adopted all of it.
Where proposed Rule 707 actually stands
This needs stating carefully, because much commentary has it wrong. Proposed Federal Rule of Evidence 707 is not law and is not pending adoption. It would require machine-generated evidence offered without a supporting expert to satisfy Rule 702(a)–(d). It was published for comment in August 2025; comment closed 16 February 2026; the Advisory Committee on Evidence Rules discussed the comments on 7 May 2026 and identified greater overall concerns; and on 3–4 June 2026 the Standing Committee declined to advance it, returning it for revision. It carries no effective date, and the June 2026 agenda book lists it as an information item, not an action item.
Two errors circulate widely. The first is a headline reporting that Rule 707 was approved by the Judicial Conference; what the Standing Committee approved in June 2025 was publication for comment — the start of the process, not the end. The second is an effective date of 1 December 2027; that date belongs to a different amendment.
The text as published read: When machine-generated evidence is offered without an expert witness and would be subject to Rule 702 if testified to by a witness, the court may admit the evidence only if it satisfies the requirements of Rule 702(a)–(d). This rule does not apply to the output of simple scientific instruments. The committee note was revised to say these standards will be difficult to meet, and sometimes impossible, without presenting expert testimony — the opposite of the assumption that an automated tool lets a party skip the expert.
The companion proposal on fabricated and altered evidence, a new Rule 901(c), has not been published for comment either. It is a burden-shifting draft: the opponent would first present evidence supporting a claim that an item is fabricated or altered, and the proponent would then establish authenticity by a preponderance under Rule 104(a). Treat it as a proposal under study, never as a rule.
Why the analysis does not wait for a new rule
The absence of Rule 707 changes little about how a court handles an AI-derived number in a search case, because here the output nearly always arrives sponsored by an expert — and the moment it does, Rule 702 reaches it.
A similarity score an expert relies on is part of the facts or data the testimony rests on, a Rule 702(b) question. A model classification adopted as a finding becomes part of the principles and methods, a Rule 702(c) question. A conclusion the output cannot bear — a probabilistic classification treated as a determination — is the Rule 702(d) problem the 2023 amendment was written to catch.
Authentication runs in parallel. Rule 901(a) requires evidence sufficient to support a finding that the item is what the proponent claims, and Rule 901(b)(9) fits automated output: evidence describing a process or system and showing that it produces an accurate result. A vendor's marketing page is neither a description nor a showing. Where the system is one a party controls, Rule 902(13) offers a certification route for a record generated by an electronic process that produces an accurate result.
What 707 would add is a rule for the exhibit nobody sponsors: a tool's screen with a number on it.The five things an expert has to be able to state
Whether the challenge comes under 702(c), 901(b)(9), or someday 707, the same five facts decide whether there is an answer.
- The model. Which system produced the output, named specifically, not "AI" or "our proprietary algorithm." If it is a vendor's score, what the vendor publishes about how the score is computed and what it means.
- The version. Models are revised and retired. A score produced by one version is not necessarily reproducible on the next, and a report that omits it has recorded a result nobody can return to.
- The inputs. Exactly what was submitted: which URLs, which text, which capture date, which fields, with what preprocessing. Stripping navigation, normalizing whitespace, or truncating to a token limit can move a similarity score substantially.
- The prompts. Where a language model was used, the prompt is part of the method: verbatim text, system instructions included, with the temperature setting and whether outputs were re-generated and selected.
- The validation. What was done to establish that the output is accurate for this task on this material — not what the vendor claims in general.
The first four are recordkeeping. The fifth is work, and it is the one usually missing.
Validation is the part that is usually missing
Validation means comparing the model's output against a known answer on material like the material at issue, and reporting how often it was right and how it was wrong.
In practice that is a ground-truth exercise. Take a sample from the same population — the reviews, pages, or links at issue. Have a human classify them under a written rule, without seeing the model's answer. Run the model and compare. Report the false positive and false negative rates separately: a classifier that flags genuine reviews as fake is a different problem from one that misses fabricated ones, and which error matters depends on which side you are on. State who labeled, under what criteria, and how the sample was drawn. If it is too small to support a population estimate, say so rather than produce one.
This converts an unfalsifiable number into an opinion with a stated error rate, which is what the reliability inquiry asks about for a method with no peer-reviewed literature behind it. It has a practical benefit too: I have run this exercise and abandoned a classification approach because the error rate was too high to support the opinion. Better to learn that before the report is served than in deposition.
The reproducibility problem nobody has solved
Machine output introduces a difficulty ordinary technical evidence does not have: the same input can produce a different answer on the second run. A generative model's output varies with sampling settings, with session context, and sometimes with nothing observable. Vendors update models continuously and retire old versions, so a run performed in one quarter may be impossible to repeat in the next, and retrieval-based systems change their answers when the retrieved material changes.
Three practices follow, and should be adopted before there is a dispute. Capture the raw output when it is produced, in full, including any identifiers the system returns, and preserve it unaltered alongside the inputs and the prompt. Record the run conditions: date, time, model, version, settings, account. Run it more than once and report the variation, because the stability of the output is itself a finding. In disputes over what an AI system says about a business, establishing that an output recurs across sessions, accounts, and settings is the difference between a screenshot and evidence.
Worked example: a vendor's toxic link score
Take the exhibit I see most often, because it shows how fast the questions run out. A report states that a competitor built thousands of toxic backlinks to the plaintiff's site, and attaches a table from a commercial tool listing each link with a toxicity score and a threshold. The technical work is real: the links exist, a crawl found them, and the crawl is reproducible. The problem is the toxicity column.
The cross-examination writes itself. What does the score measure? Who set the threshold, on what basis? What does the vendor publish about the inputs? Was it validated against any known set of manipulative links? Does a search engine use it for anything? Have these links caused any demonstrable effect — a manual action in the Search Console account, measurable movement for the pages they point to?
The defensible version abandons the score and keeps the observable facts. These links appeared in this window at this rate. They share these registration, hosting, template, or anchor-text characteristics. They point disproportionately to these pages. This is what the engine's published documentation says about links of this character, in the edition current at the time. Here is what the site's first-party data shows for those pages over the same period. That opinion rests on records and re-runnable queries rather than an undisclosed classifier, and it does not need Rule 707 to survive.
What to do now, whether or not the rule arrives
Rule 707 may be revised, narrowed, or never adopted. None of that changes what a party should do with machine-derived material today. For a party offering it: document the model, version, inputs, prompts, and validation contemporaneously, preserve the raw output, and keep derived files separate from originals. Prefer observable facts to vendor scores, and where a score is the only available measure, present it as what it is — a vendor's model output, with its definition and limits stated in the report rather than extracted in deposition.
For a party opposing it: ask the five questions early enough that the answers land in the deposition, request the prompts and raw outputs in discovery specifically because they are often absent from a production containing only the summary table, and ask whether the run can be repeated today, on the current version, with the same result.
The through-line is the one the Advisory Committee's own note identified: reliability standards for machine output are difficult to meet without an expert who can explain the system. That points toward more expert involvement, not less.
Frequently Asked Questions
Is Federal Rule of Evidence 707 in effect?
No. Proposed Rule 707 is not law and is not pending adoption. It was published for public comment in August 2025, the comment period closed on 16 February 2026, the Advisory Committee identified greater overall concerns with it in May 2026, and in June 2026 the Standing Committee declined to advance it and returned it for revision and continued study. No effective date has been set. Commentary reporting that the Judicial Conference approved Rule 707, or that it takes effect on 1 December 2027, is inaccurate; what was approved in 2025 was publication for comment.Can AI-generated analysis be used in an expert report?
It can, but the model becomes part of the method and is examined as such. If an expert relies on a model's output, that output is part of the facts or data under Rule 702(b); if the expert adopts the model's classification as a finding, the model is part of the principles and methods under Rule 702(c); and a conclusion stated more firmly than a probabilistic output can support is the Rule 702(d) problem. The practical requirement is that the expert be able to state the model, its version, the inputs, the prompts, and what validation was performed.What would proposed Rule 707 have required?
As published for comment in August 2025, it read that when machine-generated evidence is offered without an expert witness and would be subject to Rule 702 if testified to by a witness, the court may admit the evidence only if it satisfies the requirements of Rule 702(a)–(d), with a carve-out stating that the rule does not apply to the output of simple scientific instruments. The committee note was revised to observe that these reliability standards will be difficult, and sometimes impossible, to meet without presenting expert testimony. The proposal has not been adopted.Are commercial "toxic link" scores admissible?
The score is a vendor's model output, not a measurement, and an expert who adopts it has adopted a classifier whose definition, inputs, and threshold are set by the vendor and generally undisclosed. Absent a validation against known material and an account of how the score is computed, it is exposed under Rule 702(c). The stronger version of the same opinion drops the score and rests on observable facts: when the links appeared, what characteristics they share, which pages they point to, and what the site's own first-party data shows over the same period.How do you authenticate the output of an automated tool?
Through the process, not the output. Rule 901(a) requires evidence sufficient to support a finding that the item is what the proponent claims, and Rule 901(b)(9) provides for authentication by evidence describing a process or system and showing that it produces an accurate result. That means a description of how the system works and a showing about its accuracy. Where the system is one a party controls, Rule 902(13) allows a certification that a record was generated by an electronic process or system that produces an accurate result, subject to the notice requirements it incorporates.Is there a federal rule addressing deepfakes and fabricated screenshots?
Not yet. A proposed Rule 901(c) exists in draft and has not been published for public comment. It is a burden-shifting approach: the opponent would first present evidence supporting a claim that an item is fabricated or altered, and the proponent would then establish authenticity by a preponderance under Rule 104(a). The Advisory Committee questioned whether there is yet enough real-world litigation to justify publication and requested a Federal Judicial Center survey on how often such disputes arise before deciding. Existing Rule 901 governs in the meantime.Why does reproducibility matter for a language model's output?
Because the same input can produce a different answer on a second run, and because vendors revise and retire models, so a run performed in one quarter may be impossible to repeat in the next. The practices that address it are capturing the raw output in full at the time it is produced, recording the run conditions including model, version, and settings, and running the query repeatedly to report the variation. An output that recurs across sessions, accounts, and settings supports a materially stronger opinion than one captured once.Published