What the document set is, and what it is not
On 13 March 2024 an automated process committed internal Google documentation to a public repository on GitHub, a hosting service for software projects where every change is recorded as a timestamped, cryptographically identified commit. Removal was attempted on 7 May 2024, by which time third parties had mirrored it. Erfan Azimi located the material and gave it to Rand Fishkin, who verified it with former Google employees and published on 27 May 2024; Michael King published a technical analysis on 28 May 2024. On 29 May 2024 Google confirmed the documents' authenticity through a spokesperson, while cautioning against drawing conclusions from them.
The set contains 2,596 modules and 14,014 attributes — internal reference documentation describing the named data fields Google's content storage systems can hold about a page or a site, organized into groups. That is the whole of it. There is no source code, no ranking weights, and no indication of which attributes are live, deprecated, experimental, or used only for internal analysis.
Both halves matter. The set is real internal Google documentation, confirmed as such by Google. It is also a field dictionary rather than a specification of the ranking system, and every question below turns on holding both facts at once.
Authentication: the strongest fact is a party admission
Federal Rule of Evidence 901(a) requires evidence sufficient to support a finding that the item is what its proponent claims. Note the claim: that these are internal Google documents, not that their contents are accurate. Four routes are available and they reinforce one another.
Google's own confirmation of 29 May 2024. This is the strongest authentication fact in the subject: the party whose documents these are said publicly that they are its documents. Offered against Google, that statement is an opposing-party statement under Rule 801(d)(2) and not hearsay at all; elsewhere, it is still a published statement by the entity best placed to deny the documents, which did not.
Distinctive characteristics under 901(b)(4). The rule permits authentication from appearance, contents, substance, internal patterns, or other distinctive characteristics taken with the circumstances. Internal naming conventions, structural consistency across thousands of fields, and references to internal systems are exactly that kind of showing.
Process or system under 901(b)(9). The rule permits evidence describing a process or system and showing it produces an accurate result. Version control is such a system: the commit history records who committed what, when, and with what content hash, and a witness competent to describe that mechanism supports a finding about when the material was published.
Hash-based certification under 902(13) and 902(14). Copies taken from the mirrored repositories can be authenticated by digital identification — hash values — under 902(14), or as records generated by an electronic process under 902(13), each on a certification from a qualified person satisfying the 902(11) requirements, including written notice to opponents before trial. Be precise about what a hash proves: that the copy is identical to the file the examiner processed. It says nothing about where the original came from. Hashing is the last link in the chain, not the chain.
The hearsay posture when offered for the truth
Offered to prove what Google's ranking system actually does, the documents are hearsay unless an exception applies, and there are two realistic candidates.
Business records, Rule 803(6). The record must have been made at or near the time by, or from information transmitted by, someone with knowledge; kept in the course of a regularly conducted activity; and made as a regular practice — shown by a custodian or other qualified witness, or by a certification meeting Rule 902(11). That requirement is the practical obstacle, because the custodian is Google. Where Google is not a party and has no reason to certify, the route is largely theoretical, and an expert who assumes it is planning around a foundation nobody has laid.
Opposing-party statements, Rule 801(d)(2). Where Google is a party, a statement offered against it that was made by it, adopted by it, made by a person it authorized, or made by an agent or employee on a matter within the scope of that relationship, is defined as not hearsay. No exception is needed and no custodian is required. That is precisely the posture these documents occupy in the publisher and online-education antitrust actions filed against Google in 2025, where its internal documentation about what its systems record is a party statement rather than a hearsay problem.
The asymmetry drives strategy. In litigation against Google the documents are comparatively easy to get in for their truth; in the ordinary commercial case between two private parties they are not, and a proponent who needs them for their truth needs another theory.
The impeachment use, which is the real one
The alternative theory is better than the primary one, and it is why this material matters outside antitrust.
For years Google's public guidance said things the internal documentation does not support. Public statements minimized or denied the use of click behavior as a ranking signal; the documentation names click-based systems and fields distinguishing good clicks from bad. Public statements denied the use of Chrome browsing data, denied a sandbox suppressing new sites, and denied that domain age was used; the documentation contains Chrome-related attributes, describes a suppression mechanism, and records field-level data on age. A site-level authority metric appears in it after years of public denial that any such thing existed.
Offered not for the truth of what the system does, but to show that Google's public guidance was an unreliable description of the system, these documents are not hearsay at all. The proposition proved is not "Google uses click data." It is "Google's public statements about its system were, on Google's own internal documents, an incomplete account of it." For that purpose the documents need only exist and say what they say. Whether any attribute is live is irrelevant — the impeachment use survives even if the entire set were deprecated tomorrow.
This matters wherever a standard of care was built from Google's public guidance, which is how it is usually constructed in this field. There is no licensure in search work and no published professional standard, so an expert asked whether conduct fell below the standard reaches for the public guidance in force at the time. That construction is not destroyed by the leak, but it stops being self-evidently authoritative. An opinion that conduct was substandard because Google said so now has to reckon with documented evidence that what Google said and what Google's system recorded were not the same thing — a rebuttal argument with real force, available in matters that have nothing to do with antitrust.
The Rule 702(d) trap
Rule 702(d), as amended effective 1 December 2023, requires the proponent to show it is more likely than not that the opinion reflects a reliable application of the principles and methods to the facts. It targets overstatement, and this document set is a near-perfect instrument for producing one.
The overstatement is this: an expert testifies that because an attribute appears in the documentation, it is a ranking factor, or carries a particular weight, or caused the site's problem. The set contains no weights and no live indicator, and nothing in it distinguishes a field written to daily in production from one abandoned years ago and never removed. That is not a reliable application of anything; it is a leap from a name to a conclusion about behavior, dismantled in three questions on cross-examination.
The defensible opinion is narrower, and more useful than the overstated one because it survives:
- The documentation, authenticated by Google's own confirmation, records the existence of a named attribute of a stated description, in a stated module.
- It contains no weights, no source code, and no indication of which attributes are in production use, so it does not establish that any listed attribute influences ranking.
- It is therefore inconsistent with public statements that no such data was collected or held — a proposition about what Google said, not about what Google's system does.
- Where an independent source corroborates a particular system, that corroboration — not the document set — carries the weight.
An expert who says that much has said something true and load-bearing. An expert who says the leak proves a ranking factor has handed the other side a Rule 702(d) motion, and in my experience that is the opinion that draws one.
Corroboration from the antitrust record
One system in the documentation does not rest on the documentation alone. NavBoost, a click-driven ranking system, had already been described under oath in the government's search monopolization case against Google, where testimony placed its use of click data at roughly 2005 onward. That is sworn testimony in a public record, and it corroborates the leaked material on that point independently of the leak's provenance.
Two cautions. Secondary coverage paraphrases that testimony loosely, so pull the transcript before quoting or characterizing it. And corroboration of one system does not corroborate 14,014 attributes; an expert who uses it to vouch for the set as a whole has repeated the overstatement in a more sophisticated form.
The surrounding matter changes the evidentiary environment. Liability was found on 5 August 2024, reported at 747 F. Supp. 3d 1 (D.D.C. 2024); the remedies opinion issued 2 September 2025; Final Judgment was entered 5 December 2025 and took effect about sixty days later; Google's notice of appeal followed on 3 February 2026 and the appeal is pending. The judgment runs six years, requires search index and user-side data to be made available to qualified competitors on stated terms, and installs a five-member technical committee with access to source code and algorithms under confidentiality.
The contrast with the leak is the point. A court-supervised regime contemplates access to the actual systems, under confidentiality, with a mechanism behind it. The leaked documentation is a field dictionary with no weights that arrived through a bot commit. Treating the second as though it were the first is the error this page exists to name.
How the other side attacks it, and how the record is built
Assume the opposing expert has read this page. The attacks that land are narrow ones.
- Completeness. Which copy is the exhibit, from which mirror, obtained on what date, and how does it differ from others? A partial copy authenticates as what it is, not as the original set.
- Currency. The documentation reflects a state as of March 2024, and says nothing about the system in a damage window elsewhere in time.
- Live status. Which of the attributes you relied on are in production use, and how do you know that from this document?
- Interpretation. Field names are internal shorthand written for engineers. An expert's gloss on what a name means is an inference and should be identified as one.
- Purpose. Where the documents come in for impeachment rather than for their truth, a limiting instruction under Rule 105 is available, and an expert who then reasons from them as though they described the system has exceeded the purpose for which they were received.
Building the record is correspondingly practical. Obtain the copy from a mirror with a documented retrieval date and hash it on receipt; preserve the surrounding metadata, including commit records, retrieval logs, and the contemporaneous published analyses with their dates; capture Google's confirmation as published and quote its qualifying language as well as the confirmation; and state in the report which use — authentication, truth of contents, or impeachment of public guidance — each conclusion depends on. The governing rules are worth reading in full: Rule 901, Rule 801(d)(2), and Rule 702.
Frequently Asked Questions
Can the 2024 Google documentation leak be used as evidence in court?
It can, but the use determines everything. Authentication is unusually strong, because Google publicly confirmed the documents were genuine on 29 May 2024 — a party admission on the authenticity question. Offered to prove what Google's ranking system does, the documents are hearsay and need Rule 803(6) or, where Google is a party, Rule 801(d)(2). Offered to show that Google's public guidance was an unreliable description of its system, they are not offered for their truth and no hearsay exception is required. That third use is the one available in ordinary commercial litigation.How are the leaked documents authenticated?
Four routes, which reinforce each other. Google's public confirmation of authenticity on 29 May 2024 is the strongest single fact and functions as an admission by the party whose documents these are. Rule 901(b)(4) supports authentication from distinctive characteristics — internal naming conventions, structural consistency across thousands of fields, and references to internal systems. Rule 901(b)(9) supports it through the version-control process that recorded the commit. Copies from mirrored repositories can be certified under Rule 902(13) or 902(14) with hash values, subject to the 902(11) notice requirement. A hash proves the copy matches the file examined, not where the file came from.Are the leaked documents hearsay?
It depends entirely on what they are offered to prove. Offered for the truth of what Google's ranking system does, yes — they are out-of-court statements offered for their truth, admissible only as business records under Rule 803(6) or, against Google, as opposing-party statements under Rule 801(d)(2). The business-records route requires a custodian or a compliant certification, and the custodian is Google, which makes it impractical where Google is not a party. Offered to show that Google's public statements were an unreliable account of its own system, they are not offered for their truth and hearsay does not apply.Can an expert testify that a leaked attribute is a Google ranking factor?
That is the overstatement the amended Rule 702(d) is aimed at. The set contains 2,596 modules and 14,014 attributes, but no source code, no weights, and no indication of which attributes are in production use. Nothing in it distinguishes a field written to constantly from one abandoned years ago and never deleted. Testifying that a listed attribute is a ranking factor moves from the existence of a name to a conclusion about system behavior with nothing in between. The defensible opinion is that the documentation records the existence of the named attribute and does not establish its effect.Does the leak change the standard of care in an SEO malpractice case?
It complicates the way that standard is normally built. There is no licensure in search work and no published professional standard, so a standard of care is usually constructed from Google's public guidance in force at the time of the conduct. The documentation shows that public guidance and internal documentation diverged on several points, including click signals, Chrome data, new-site suppression, and domain age. That does not make the guidance irrelevant, and it remains the best contemporaneous public statement of expected practice. It does mean an opinion resting on Google's public statements as an authoritative description of the system now has a documented answer waiting for it.Does the antitrust record corroborate the leaked documents?
On one point, independently and under oath. NavBoost, a click-driven ranking system named in the documentation, was described in testimony in the government's search monopolization case as having used click data since roughly 2005. That is sworn testimony in a public record and it corroborates the documentation on that specific system without relying on the leak's provenance at all. It does not corroborate 14,014 attributes, and using it that way repeats the overstatement in a more sophisticated form. Pull the transcript before quoting the testimony, because secondary coverage paraphrases it loosely.What is not in the leaked document set?
Everything that would make it a description of the ranking system. There is no source code. There are no signal weights, so nothing indicates how much any attribute matters or whether it matters at all. There is no live-or-deprecated flag, so nothing distinguishes production fields from abandoned ones. There is no query-level or site-level data about any party. And the material reflects a state as of March 2024 only, which says nothing about the system in a damage window elsewhere in time. Each of those gaps is a cross-examination question, and each has a correct answer that concedes the gap.Published