Search evidence and expert testimony
Is Search Testimony Admissible?

How Search Engines Work, for a Court

Crawling, indexing, ranking, and the generated answer on top — described at the level an expert has to defend

Why the mechanism has to be explained at all

Most people, including most judges and every jury, carry one model of a search engine: a directory of the web, in which a page that exists can be found. That model is wrong in three places, and each is where search cases are decided.

A page can exist and never be discovered. It can be discovered and deliberately kept out of the index. It can be in the index and never returned for the query that matters. All three failures are invisible to anyone who types the address into a browser, because the page loads perfectly. That is the gap specialized testimony fills, and why helpfulness under Rule 702(a) is rarely contested here.

What follows is the mechanism in four stages, at the level an expert has to defend — including where the honest answer is that the details are not public.

Stage one — discovery and crawling

A crawler is automated software that requests web pages the way a browser does, reads what comes back, extracts the links, and queues the addresses it finds. Google's is Googlebot. Crawling is how the engine learns a page exists.

Discovery happens mostly through links — a page nothing links to and no sitemap lists may never be requested. A sitemap is a file, usually XML, in which a site lists its own addresses.

Two controls matter in litigation, and they are constantly confused:

  • robots.txt is a text file at the root of a site telling compliant crawlers which paths not to request. Major engines honor it; anything else may ignore it. Critically, it controls crawling, not indexing: a URL blocked in robots.txt can still appear in results if the engine learns of it elsewhere, because the engine has been told not to look at the page rather than not to list it.
  • noindex is a directive in the page's own HTML or in an HTTP header telling the engine not to include the page in the index. It works only if the engine may fetch the page and see it — which is why blocking a URL in robots.txt while adding noindex to it is self-defeating, and why a noindex left in place after a staging site goes live removes real pages from search while the site looks normal to a visitor.

Crawling is budgeted. A large site is not fully re-crawled on demand, so a change may be reflected for some pages within hours and for others over weeks. That lag is why the date a change was made and the date its effect appears are two different dates, and why server logs — the server's own record of every request it answered — are the best evidence of what the engine did and when.

Stage two — indexing, and why a real page may not be in it

The index is the engine's stored, processed copy of the pages it has decided to keep. Being crawled does not mean being indexed. Indexing is a selection.

Between fetching and storing, the engine renders the page — it executes the JavaScript to see what a browser would display. Content appearing only after script runs may be seen later than content in the initial HTML response, or, where the script fails, not at all. That is why a page can look complete to a person and be substantially empty to an engine, and why crawl configuration matters: a crawl with rendering disabled reports content missing that a rendered crawl would find.

The engine also deduplicates. Where several addresses serve near-identical content, it selects one to represent the group. A canonical tag — a <link rel="canonical"> element in the page's head — declares which address the site considers primary. It is a signal the engine weighs, not an instruction it must obey, and misconfigured tags are a recurring cause of pages vanishing after a redesign, because a template can point every page in a section at one address.

Then the 301, a permanent redirect: a server response saying an address has moved permanently. Redirects carry a site's accumulated standing from old addresses to new during a migration, and a redirect map sending retired URLs to a homepage, to a 404 not-found response, or through long chains of hops is the most common technical failure in these disputes.

Finally, an engine may crawl a page, render it, and decline to index it anyway. There is no obligation to index anything, and “discovered but not indexed” is an ordinary state Search Console reports in volume.

Stage three — ranking, and the line between known and unknown

This is where testimony most often overreaches, so state precisely what is public and what is not.

What is publicly known. When a query arrives the engine interprets it — correcting spelling, expanding synonyms, inferring language and location. It retrieves candidate documents from the index, scores them using many signals derived from the page, the site, the links pointing at it, and the query, reranks with further systems, and assembles a results page. Google publishes documentation describing its ranking systems at a conceptual level and announces broad core updates — engine-wide changes to those systems — several times a year, with dates.

What is not public. The weights. No published document states how much any signal contributes, how signals interact, which are active for a given query type, or how that has changed. There is no published error rate and no way to hold the index constant while testing a variable. An expert who testifies that a named factor accounts for a share of a ranking outcome asserts something no public source supports.

The leaked material, handled correctly. In March 2024 internal Google documentation was committed by an automated process to a public code repository, mirrored by third parties, and published that May; Google confirmed its authenticity on 29 May 2024 while cautioning against drawing conclusions. It contains no source code, no signal weights, and no indication of which attributes are in live use. The position that survives cross-examination is that it shows attributes existing in an internal system — not that any named attribute is a ranking factor whose effect can be quantified. Treating an attribute name as a confirmed factor is the overstatement the 2023 amendment to Rule 702 was written to catch.

One narrow point stands on independent footing: testimony under oath in the government's search antitrust case described a system using click data in ranking since roughly 2005. That is the model — say what a sworn record or the operator's own published statement establishes, and stop.

Stage four — the results page is not a list of ten links

A jury shown a screenshot will assume the first organic listing is at the top of the page. Frequently it is not, and the difference matters when the dispute is about visibility.

A modern results page assembles several kinds of block: paid advertisements labeled as such, which may occupy the top of the screen; the organic listings; and features drawn from the index and elsewhere — a panel of facts about an entity, a local pack with a map, image and video blocks, a product grid, related questions, and a featured snippet, being a passage extracted from one page and displayed above the rest.

Two consequences follow. Position is not location: ranking first organically can mean appearing below the visible area of a phone screen, and average position in Search Console is an average rank among organic results, not a measure of prominence. And an impression is not a view: Search Console counts an impression when a result appears for a query, subject to its own definitions for results requiring scrolling. Testimony treating impressions as people who saw something is imprecise in a way opposing counsel will find.

The generated layer: AI Overviews and AI Mode

Above the results, for many queries, Google now displays an AI Overview: an answer generated by a language model from material retrieved at the time of the query, with links to some sources. AI Mode is a separate conversational interface in which the whole interaction is generated rather than listed. Three properties matter to a court, and none is intuitive.

  • It is not a fixed document. The same query can produce different text on different days, from different locations, for different accounts, and sometimes on consecutive runs. A screenshot records one output at one moment, not a stable published statement — which is why establishing that an output recurs across sessions and accounts is a methodological question rather than a matter of taking another screenshot.
  • It does not appear everywhere. Overviews are shown for some queries and not others, so whether one appeared for a given query on a given date is a fact to establish rather than assume.
  • It changes the click. The honest presentation carries the caveats with the numbers. Pew Research Center, tracking 900 consenting U.S. adults' browsing, found that where an AI summary appeared users clicked a traditional result in 8% of visits against 15% where none appeared, and clicked a link inside the summary in 1%; Google disputes the study. Ahrefs, comparing 300,000 informational keywords, found a 34.5% lower clickthrough rate for the top organic result where an Overview appeared, and stated expressly that the finding is correlational.

Those figures describe an environment. They do not establish what happened to any particular site, and an expert who applies an industry average to a plaintiff has substituted a statistic for an analysis.

Why two people searching the same thing see different results

There is no single results page for a query. What is returned depends on inferred location, device, browser, language and country setting, time, the account signed in, and whatever experiments are running on some fraction of traffic. Results also change as the index and the ranking systems change, so a page that ranked third in March may rank ninth in June with no change to the page.

This is why an unlabeled screenshot proves little. What makes a capture evidence is the surrounding record: the exact query string, the date and time with a time zone, how the location was set, the device and browser, whether the session was signed in, and whether the search was repeated elsewhere to see whether the result held. Capturing systematically — the same queries, on a schedule, from stated locations, with raw responses preserved — produces something an opposing expert can test rather than dismantle.

What all of this means for the evidence in the case

Three translations, where the explanation earns its place at trial.

The site's own records are the best evidence of what the engine did. Server logs record every request Googlebot made and what the server answered. Search Console reports the engine's record of impressions, clicks, average position, and queries for a verified property on a rolling sixteen-month window — so its record of an older period expires whether or not anyone is in litigation. Analytics records what visitors did after arriving. The three answer different questions and are frequently confused.

The failure has to be located at a stage. “The site lost traffic” is not a finding. Whether pages stopped being crawled, stopped being indexed, or stayed indexed and stopped ranking are three states with three bodies of proof and usually three different responsible parties.

The unknowns cut both ways, and saying so is the credible position. Because ranking weights are not public, no expert on either side can testify to what a change was worth in isolation. That constrains a plaintiff's causal claim and equally constrains a defendant's assertion that the loss came from somewhere else. The testimony that holds up describes what can be observed — what was crawled, indexed, and returned, and what the site's own records show — and states where the observable evidence stops.

Frequently Asked Questions

How does a search engine decide what to show?

In four stages. A crawler discovers and requests pages, mostly by following links. The engine renders and processes what it retrieves and selects some of it into an index, deduplicating near-identical addresses. When a query arrives it interprets the query, retrieves candidate documents from the index, scores them against many signals, reranks, and assembles a results page. What is public is the structure and a conceptual description of the systems involved. What is not public is the weighting: no document states how much any signal contributes or which are active for a given query.

Why would a page that exists not appear in search results?

Because there are three separate failure points, all invisible to a person who types the address. The page may never have been discovered, if nothing links to it and no sitemap lists it. It may have been discovered but excluded from the index by a noindex directive, by a canonical tag pointing at a different address, or by the engine simply declining to index it. Or it may be indexed and not returned for the query at issue. Distinguishing which occurred requires the site's server logs and Search Console data, not a browser.

What is the difference between robots.txt and a noindex directive?

robots.txt is a file at the root of a site that asks compliant crawlers not to request certain paths. It controls crawling, not listing, so a blocked URL can still appear in results if the engine learns of it from elsewhere. A noindex directive, in the page's HTML or in an HTTP header, tells the engine not to include the page in its index — but the engine has to be allowed to fetch the page to see it. Blocking a URL in robots.txt while adding noindex to it is therefore self-defeating, and it is a common real-world error.

Are Google's ranking factors public?

The existence of many signals is described publicly at a conceptual level, and broad core updates are announced with dates. The weights are not public. No published source states how much any signal contributes, how signals interact, which are active for a given query type, or how that has changed. Internal documentation that became public in 2024 and whose authenticity Google confirmed contains no source code, no weights, and no indication of which attributes are in live use, so it shows that attributes existed in an internal system — not that any named attribute is a ranking factor whose effect can be quantified.

What is an AI Overview and how does it affect traffic?

It is an answer generated by a language model above the search results, synthesized from material retrieved at query time, with links to some sources. Measured effects exist, with caveats. Pew Research Center found users clicked a traditional result in 8% of visits where an AI summary appeared against 15% where none did, and Google has disputed the study. Ahrefs reported a 34.5% lower clickthrough rate for the top organic result where an Overview appeared, stating the finding is correlational. Those describe an environment; they do not establish what happened to any particular website.

Why do two people get different search results for the same query?

Because results depend on inferred location, device, browser, language and country settings, time, whether an account is signed in, and any experiments running on a fraction of traffic. Results also change as the index and ranking systems change. That is why an unlabeled screenshot is weak evidence: what makes a capture usable is the surrounding record — the exact query string, the date and time with a time zone, how the location was set, the device and browser, whether the session was signed in, and whether the result was reproduced from another location or account.

What records show what a search engine actually did to a website?

Three, and they answer different questions. Server logs record every request the crawler made and what the server returned, which is the best evidence of crawling and of the status codes served. Search Console reports the engine's own record of impressions, clicks, average position, and queries for a verified property, on a rolling sixteen-month window — so the engine's record of an older period expires on its own. Analytics records what visitors did after they arrived. Confusing the three is one of the more common errors in reports in this field.
Keep reading

The entries behind this guide

Every rule, method and dispute type named here has its own entry: the authority that governs it, the question it answers, and the evidence it runs on.

Top