Search evidence and expert testimony
Abstract angled shard illustration representing Archive Gaps and Link Rot

IssueFoundationIs the exhibit received?

Archive Gaps and Link Rot

Governing authority
FRE 901 authentication of archived and live web content
Question at issue
Does the page in the exhibit show what existed on the stated date, and nothing else?
Primary evidence
Capture timestamps for the page and each embedded resource, lookups across archives, contemporaneous forensic captures
When it arises
Raised when the original page has changed or disappeared and an archived copy is the best surviving evidence

An archived page is a sample of the web, often assembled from captures that were not taken on the same day

Why a missing page needs a base rate

Search disputes are full of pages that no longer exist: the product page a competitor took down, the blog post that carried the disputed links, the landing page from before a migration. A missing page invites an inference, usually that someone removed it, and sometimes that someone removed it because of the dispute. That inference may be right. But the web loses content at a rate most litigators would not guess, and a page that is gone is, more often than not, gone for ordinary reasons.

The same research that measures how much disappears also measures how much was archived, and how faithfully. Both numbers belong in any opinion that relies on an archived page, because the gaps in an archive are not random and the composite nature of an archived page is not visible on its face. This entry collects the empirical record. The foundation routes for archived material, and the circuit split on judicial notice, are covered in the entry on Wayback Machine evidence.

How much of the web disappears

The most recent large study is from Pew Research Center, which drew roughly a million pages from Common Crawl, an open repository of web crawl data, and tested them in October 2023. Some 38% of pages that existed in 2013 were no longer available, compared with 8% of pages that existed in 2023. Across all pages collected from 2013 through 2023, 25% were inaccessible: 16% were individually missing from a domain that still worked, and 9% sat on a root domain that no longer worked at all. About one in five pages from 2021 was gone two years later.

That split between a deleted page and a dead domain is worth carrying into a report. The first is a decision by someone who still controls the site. The second is often an expiration or an abandoned business. They support different inferences, and an opinion that treats them alike has discarded information.

Legal citation is not immune. A study of URLs cited in United States Supreme Court opinions from 1996 through the 2009 to 2010 term found 29% of them invalid. A separate Harvard study found that 50% of URLs in Supreme Court opinions, and more than 70% in three Harvard law journals, no longer produced the information originally cited. A study of links in New York Times articles found 25% of deep links completely inaccessible, rising to 72% for links from 1998.

A live address is not the same content

The gap between those two Supreme Court figures is the point of this section. The lower number counts dead links; the higher one counts links that resolve but no longer show what was cited. In the Harvard study, only 76% of Supreme Court links returning an HTTP 200 status (the server's signal that the request succeeded) still led to the cited material. The New York Times study found a further 13% of reachable links had drifted significantly since publication.

The practical lesson is narrow and frequently missed. An expert who confirms that a page still exists by loading the URL today has confirmed the address. Whether the content on that address matches the content on the date in dispute is a separate question, answered by comparison against a dated capture, and the report should say which question it answered.

How much is archived, and how often

Archive coverage depends heavily on where a page came from. A study by researchers at Old Dominion University found that between 35% and 90% of URLs had at least one archived copy, depending on the sample: URLs found through search engines had roughly a two-in-three chance of being archived, and URLs shared through a link shortener just under one in three. No more than 31.3% of URLs in any sample were archived more than once a month. That data is from around 2012 and crawling has expanded since, but the shape remains: popular, well-linked pages are captured often, and obscure pages are captured rarely or never.

The Supreme Court study tested the same question from the other direction. Of the dead links in Supreme Court opinions, 68% were accessible in the Internet Archive when checked in March 2013. About a third had no archived copy at all, and the study did not check whether the copies that existed matched what the Court had cited.

Between any two captures, the state of a page is unknown. You can bound it; you cannot observe it. In a dispute about when content changed, the capture interval is the precision limit of the evidence, and an opinion that dates a change more precisely than the captures allow has outrun its data.

An archived page is a composite

A web page is not one file. It is an HTML document that calls for images, stylesheets, and scripts, each of which an archive captures separately and possibly on a different day. The Internet Archive's own standard affidavit says so: the date the Archive assigns applies to the HTML file but not to the images linked in it, so images on a page may not have been archived on the same date. The same affidavit explains that a link clicked on an archived page serves the capture with the closest available date, not necessarily one from the same day.

Researchers measured how often that matters. In a study published at ACM Hypertext 2015, Ainsworth, Nelson and Van de Sompel found that at most 38.7% of archived pages were temporally coherent, meaning every embedded resource could have existed at the moment the page was captured, and at most 17.9%, roughly one in five, were both coherent and complete. The spread between the oldest and newest resource on a single archived page was commonly weeks or months, sometimes a year or more, and in a few cases more than ten years.

An earlier study of 45,341 Internet Archive captures found only 46% complete, with 15% of embedded resources missing. For a dispute about text, a missing image may not matter. For a dispute about how a page looked, what an advertisement displayed, or whether a disclosure appeared next to a claim, it is the whole question.

Timeline diagram of one archived web page: the Wayback Machine date applies to the HTML file only, while images, stylesheets and scripts each carry their own capture dates, and a clicked link serves the closest available capture. Research found at most 17.9% of archived pages were both temporally coherent and complete.One archived page, several capture datesWhat the Wayback Machine timestamp covers, per the Archive's affidavit.earlierlaterHTML filethe date in the addressImageits own capture dateStylesheetits own capture dateScriptits own capture dateLinked pageclosest available capturePositions are illustrative. In the 2015 study, spreads of weeks or monthsbetween the oldest and newest resource were common, a year or moreoccurred, and a few exceeded ten years. At most 17.9% of archived pageswere both temporally coherent and complete.
The Wayback Machine date covers the HTML file; every other resource carries its own.

What a crawler never saw

Archives capture what their crawlers retrieve, and crawlers have historically struggled with content built by JavaScript in the browser after the page loads. A 2015 study found that a headless browser (a browser run by software without a visible window) discovered 1.75 times as many resources as Heritrix, the Internet Archive's crawler, across the same 10,000 starting pages. Crawlers have improved since, but the risk has not gone away for sites whose content, prices, or reviews load through scripts.

A capture of such a page can show a template with empty containers where the disputed content would have been. That is not evidence the content was absent. It is evidence the crawler did not execute the code that would have displayed it, and the difference should be stated before opposing counsel states it for you.

The integrity of the replay

Archived pages are replayed, not simply displayed, and replay can be manipulated. Researchers at the University of Washington, Lerner, Kohno and Roesner, found that 74% of Wayback Machine snapshots of the Top 500 websites over twenty years contained a vulnerability that could expose the snapshot to complete control by an attacker. In a set of 840 archive URLs drawn from 991 legal documents, 57 snapshots were vulnerable and 37 of those allowed complete control.

Read that carefully. A vulnerability is not a manipulation. The attack depended on control of a relevant domain, and the authors reported they were not aware of attackers having exploited it. The finding does not impeach archived evidence generally. It is a reason to corroborate an archived page that carries a case against a second source, and a reason to be specific about which resources on the page the opinion relies on.

Building and testing an archive exhibit

Several practical steps follow from the record, and they apply on either side.

  • Record every timestamp, not just one. The date in the archive address covers the HTML. Record the capture date of each image, stylesheet, and script visible in the exhibit, and flag any that fall outside the period in dispute.
  • Check more than one archive. In preliminary results presented to the Library of Congress in 2014, mean completeness of recomposed pages rose from 76.1% using a single archive to about 80% using several.
  • Separate deletion from expiration. Whether a page vanished from a working domain or the domain itself stopped resolving supports different inferences about who acted and why.
  • Prefer a contemporaneous capture. Where content is known to be contested, a forensic capture made at the time, with hash values and a documented process, is stronger than any archive retrieved later.

None of this makes archived evidence weak. It makes it specific. An archived page is good evidence of what a particular crawler retrieved at particular moments, and an opinion confined to that is difficult to attack. The vulnerable opinion is the one that treats the composite as a photograph.

Frequently Asked Questions

How common is link rot?

Very common. Pew Research Center found that 38% of web pages that existed in 2013 were no longer accessible in October 2023, and that one in five pages from 2021 was gone two years later. Across all pages collected from 2013 through 2023, 25% were inaccessible, split between pages deleted from working domains and domains that stopped working. Studies of legal citations report similar or higher rates. A missing page is therefore weak evidence, standing alone, that anyone removed it deliberately.

Does the Wayback Machine date cover everything on an archived page?

No. The Internet Archive's standard affidavit states that the date it assigns applies to the HTML file but not to linked images, which may have been archived on different dates. Links clicked within an archived page serve the capture with the closest available date. Research published in 2015 found that at most 17.9% of archived pages were both temporally coherent and complete. An exhibit that depends on images, layout, or embedded content should record the capture date of each resource it relies on.

If a URL still loads today, does that prove the content is unchanged?

No. A successful response proves the address resolves. In a Harvard study of links cited in Supreme Court opinions, only 76% of links returning a successful status still led to the cited material, and a study of New York Times links found a further 13% of reachable links had drifted significantly. Whether current content matches content on a past date requires comparison against a dated capture or another authenticated copy.

What share of the web is archived?

It depends on the kind of page. A study from around 2012 found that between 35% and 90% of URLs had at least one archived copy, depending on how the sample was drawn, and that no more than 31.3% were archived more than once a month. Popular, well-linked pages are captured often; obscure pages may never be captured. Of dead links in Supreme Court opinions, 68% had an Internet Archive copy when checked in 2013, leaving about a third with none.

Can archived pages be altered?

Research has shown that they can be vulnerable. A 2017 study found that 74% of Wayback Machine snapshots of the Top 500 websites contained a vulnerability that could expose the snapshot to complete control by an attacker, including some snapshots cited in legal documents. The authors noted that exploitation required control of a relevant domain and that they knew of no attacks. The practical response is corroboration against a second source for archived evidence that carries significant weight.

Why might an archived page show empty spaces where content should be?

Often because the content was loaded by JavaScript and the crawler did not run the script. A 2015 study found a headless browser discovered 1.75 times as many resources as the Internet Archive's crawler across the same pages. A capture showing an empty container is evidence that the crawler did not retrieve the content, not evidence that the content did not exist on the live page.

Keep reading

Read the guides

An entry states what a rule requires or what a dispute turns on. A guide walks the sequence — what you do, in what order, before the evidence is gone.

Top