Search evidence and expert testimony
Building the Evidentiary Record

Capturing and Authenticating a Web Page

A capture that survives a foundation fight records the page, the response, the process, and the moment — not just the picture

A screenshot answers one question, and only the one you already asked

The exhibit that arrives in most search matters is a screenshot pasted into a Word document, sometimes with the browser's date at the bottom. It records that something looked a particular way to somebody. It does not record which URL was actually served, what status code came back, how many redirects the browser followed, what the response headers said, what timezone the clock was in, which tool produced the image, or which user agent was used — the identifying string a client sends with each request, which often changes what a server returns.

None of that is fatal to admission on its own. Rule 901(a) requires only evidence sufficient to support a finding that the item is what the proponent claims, and a witness who took the screenshot can often say so. The problem is more practical: a screenshot answers one question, the one you knew to ask on the day. Every other question arising over the following two years has to be answered from something else, and by then the page has changed.

Decide the scope before the first request

Capture is cheap; deciding what to capture requires judgment. Three questions settle it.

Which URLs, and in which states? The disputed page, the pages linking to it, the canonical alternates, and the URLs redirecting into it. Desktop and mobile user agents, if the site serves different markup to each. The page as served and as rendered after JavaScript runs, because the two are frequently not the same document.

How many times? A single capture proves a moment. Where the fact in issue is a change, a duration, or a pattern, one capture is the weakest possible evidence of it. Capture on a schedule and keep the negatives.

Who is doing it? Whoever captures may become the witness with knowledge under Rule 901(b)(1), and where that person is the testifying expert the capture becomes part of the facts or data considered, disclosable with the report. Neither answer is wrong; drifting into one is.

What a capture has to record, item by item

For each URL captured, store all of the following together as one unit:

  1. The rendered page — a full-length image as displayed after scripts have executed, not a viewport crop.
  2. The raw HTML as served — the response body exactly as it came off the wire, before the browser modified it.
  3. The rendered document — the markup after script execution, where the two differ, saved separately and labeled as such.
  4. The HTTP response headers and status code — the complete header block and the numeric status, so a 200, a 301, a 404, and a 410 are distinguishable on the face of the record.
  5. The full response chain — every hop from the requested URL to the final one, each with its status code and location header.
  6. Subresources — the images, stylesheets, and scripts the page loaded, ideally in a web archive container storing the request and response bytes for every resource fetched.
  7. Timestamps in a stated timezone — the start and end of the capture, in UTC with the local offset stated. A time with no timezone is not a time.
  8. The capturing tool and its version, plus any configuration that affects output.
  9. The requesting user agent, verbatim, plus the network path if the capture ran through a proxy or from a stated location.
  10. The operator who initiated the capture.

Items four, five, and eight get omitted most often, and they are the ones that matter when the dispute is about what a crawler was served rather than what a person saw.

The hash, computed at the moment of capture

Compute a SHA-256 digest — a fixed-length cryptographic fingerprint of a file, where any change to the file produces an entirely different value — over every stored artifact at the moment it is written, and record the digests in a manifest with the capture metadata. Hash the raw HTML, the rendered image, the header block, and the archive container.

The reason is narrow. The Advisory Committee note to Rule 902(14) states that data copied from devices, media, and files are ordinarily authenticated by hash value, and that identical values for the original and the copy reliably attest that they are exact duplicates. That converts a later argument about alteration into arithmetic anyone can repeat.

Timing is the whole point. A hash computed when the exhibit is printed for trial proves only that the file has not changed since the printing. A hash computed at capture, recorded contemporaneously, and carried through every copy proves that the file produced in discovery is the file that came off the wire on the date stated. The difference shows up the first time the other side suggests the exhibit was edited.

Chain of custody, recorded as you go

Custody documentation is unglamorous and it is what separates a collection from a pile of files. The log is contemporaneous, appended to rather than reconstructed, and records:

  • Who collected it, from what system or account, using what tool.
  • When, in a stated timezone, and from what network location.
  • What was collected, by filename, size, and hash.
  • Where it has been stored since, including every copy and transfer, with dates.

Two practices do most of the work. Store the originals write-once and work only from copies, so the collected artifact is never the one anyone opened, renamed, or re-saved. And keep the naming mechanical — URL, timestamp, sequence — because files named after what somebody thought they showed are a folder of conclusions. The log is also what a qualified person needs to sign a certification under Rule 902(13) or 902(14).

Why this combination lines up with the rules

The procedure above is designed around what the authentication rules ask for, not around what looks thorough.

Rule 901(b)(9) permits authentication by evidence describing a process or system and showing that it produces an accurate result. A capture recording the tool, the version, the user agent, the request, the unmodified response, the status codes, and the time is a description of a process. Add a demonstration that it reproduces the same output on a known input and you have the accuracy showing too. A screenshot supports none of this, because the process it documents is a person pressing a key.

Rules 902(13) and 902(14) make records self-authenticating on certification by a qualified person. 902(13) covers a record generated by an electronic process or system producing an accurate result; 902(14) covers data copied from a device, medium, or file where the copy is authenticated by a process of digital identification — in practice, a hash. Both incorporate Rule 902(11)'s procedure: reasonable written notice of intent to offer the record, and making the record and certification available for inspection. The text is at Cornell.

The practical effect is that a contested foundation becomes paperwork — worth saying plainly, and qualifying just as plainly. Authentication is not admission. A properly captured page can still be hearsay offered for the truth of what it says, and remains subject to Rule 403.

Capturing a search results page is a different exercise

A results page is not a document sitting on a server. It is generated per request and varies with the query, the location specified or inferred, the device and user agent, the language and country parameters, personalization, and time. Two people capturing the same query in the same hour can legitimately get different results, and the opposing expert will demonstrate that.

So record the inputs as carefully as the output — the exact query string, every parameter, the stated location, the device, whether the session was signed out, and the time in a stated timezone — and capture the full page rather than the visible portion, since position is measured from the top. Then repeat, because a claim about visibility over a period needs many captures on a stated schedule.

The honest comparison with archive retrieval

Archive material does something forensic capture cannot: it reaches backwards. When the conduct is already over, a snapshot may be the only evidence of what a page said, and archived pages are routinely admitted. But they come in on foundation, with limits worth knowing before you rely on them.

The foundation routes are established. In Telewizja Polska USA, Inc. v. Echostar Satellite Corp., No. 02 C 3293, 2004 WL 2367740 (N.D. Ill. Oct. 15, 2004), an affidavit from an Internet Archive representative satisfied Rule 901's prima facie showing. In United States v. Gasperini, 894 F.3d 482 (2d Cir. 2018), testimony from the Archive's office manager explaining how content is captured and preserved, with comparison against its true copies, satisfied Rule 901(a). In United States v. Bansal, 663 F.3d 634 (3d Cir. 2011), a witness explained how the archive works and the screenshots were compared with previously authenticated images. And in Valve Corp. v. Ironburg Inventions Ltd., 8 F.4th 1364 (Fed. Cir. 2021), the Federal Circuit held that a printout could be authenticated by comparison under Rule 901(b)(3) with an authenticated specimen, and that expert testimony was not required — a permissive holding, and a useful one.

The judicial-notice shortcut is another matter. In Weinhoffer v. Davie Shoring, Inc., 23 F.4th 579 (5th Cir. 2022), the Fifth Circuit held that a court may not take judicial notice of an archived page, because a private internet archive is not a source whose accuracy cannot reasonably be questioned under Rule 201(b)(2), and observed that archived pages are not self-evidently reliable in the way Rule 902 records are. The evidence still comes in, but under Rule 901 and ordinarily through a witness with personal knowledge of the archive's process. The opinion is at the Fifth Circuit's site.

Dates are the other constraint, and the principle is not confined to archives. In ATEN International Co. v. Uniclass Technology Co., 932 F.3d 1364 (Fed. Cir. 2019), the Federal Circuit held that merely establishing that material existed in the same year as the critical date was insufficient, the reference having been placed in a year without a day or month. Snapshots carry a precise crawl timestamp, but the interval between them is not: a page captured in March and again in September proves nothing about July. Disclose the gaps.

What an archive affidavit cannot supply after the fact

An affidavit from an archive is a statement about the archive's process and what it holds. It is not a statement about the original response and cannot become one later. Specifically, it cannot retroactively supply a hash bound to the original crawl. Nobody computed a digest of the bytes the server returned at that moment and preserved it with the request and response metadata, because the crawl was not made for this case. What can be attested is that the copy matches the archive's holdings — a different and weaker proposition than that the holdings match what the server sent.

Three further limits travel with archive material. Subresources are often missing or captured at different times, so a rendered archived page can show a layout that never existed. Content can be withdrawn from retrieval afterward by a change to a site's robots.txt file, so the absence of a snapshot proves nothing. And the crawl schedule was the archive's, not your case's.

Which leads to the hybrid I would describe as sensible practice. Forensic capture is the primary evidence for anything currently observable and anything likely to be contested, because it is the only route producing a hash bound to the original response. Archive material is historical corroboration for the period before anyone was watching. An archive affidavit is what you obtain when a court wants formal foundation, and it is obtained early, because it takes time and costs money.

Frequently Asked Questions

Is a PDF print of a web page good enough evidence?

It is often enough to authenticate, and it is rarely enough to work with. A print records what one person saw once, and omits the status code, the redirect chain, the response headers, the raw HTML as served, the timezone, the capturing tool, and the user agent. When the fact in issue is what a search engine's crawler was served, or whether a URL returned a redirect on a particular date, none of that is answerable from a print. Capture the underlying response alongside the visual record, and hash both at collection.

Why compute a hash at the time of capture rather than later?

Because a hash proves only that a file has not changed since the hash was computed. Computed at trial preparation, it establishes nothing about the eighteen months before. Computed at capture and recorded contemporaneously in a manifest, it establishes that the file produced in discovery is the file collected on the date stated. The Advisory Committee note to Rule 902(14) treats hash values as the ordinary means of authenticating copied data, on the basis that identical values reliably attest the copy and original are exact duplicates.

Does a hash prove the web page itself was genuine?

No, and overstating this is a straightforward way to be embarrassed. A hash proves file integrity — that these bytes are the bytes that were stored. It says nothing about whether the server returned honest content, whether the page was cloaked to show something different to a different requester, or whether the capture was made under conditions that represent what anyone else would have seen. Those are separate questions answered by the process record, by repeat captures under stated conditions, and by corroborating sources such as logs.

How many captures of the same page are enough?

It depends entirely on the proposition. One capture supports a statement about one moment. A claim that a page said something for six months, or that a redirect was in place on a given date, or that a result appeared consistently for a query, requires captures across the period on a stated schedule — including captures that show no change, which are evidence of stability. Where the period predates the engagement, the answer is corroboration from other sources rather than more captures, because you cannot capture the past.

What has to be recorded when capturing a search results page?

The inputs, as carefully as the output. A results page is generated per request and varies with the query string, the location specified or inferred, the device and user agent, the language and country parameters, personalization, and time. Record the exact query, every parameter, the stated location, the device, whether the session was signed out, and the time in a stated timezone, and capture the full page rather than the visible portion, since position is measured from the top. One capture supports a statement about one request, not about visibility over a period.

Can an Internet Archive affidavit fix a capture that was never made?

Only partly. An affidavit attests to the archive's process and to what the archive holds, which is enough to authenticate a snapshot under Rule 901 in most courts. It cannot retroactively supply a hash bound to the original crawl, because nobody computed a digest of the server's response at the time and preserved it with the request metadata. It also cannot fill the gaps between snapshots, restore missing subresources, or explain content withdrawn from retrieval afterward. It is corroboration, obtained early because it takes time and costs money.

Does the expert have to be the person who captures the pages?

No, but the decision has consequences and should be made deliberately. Whoever captures may become the witness with knowledge under Rule 901(b)(1) and must be able to describe the process. Where the testifying expert captures, the material becomes part of the facts or data considered and is disclosable with the report. Where a client employee or a separate technician captures, the process documentation has to be good enough that someone else can describe it accurately. What fails is when nobody decided, and the capture has no documented operator at all.
Keep reading

The entries behind this guide

Every rule, method and dispute type named here has its own entry: the authority that governs it, the question it answers, and the evidence it runs on.

Top