What a yardstick is for
A yardstick or benchmark analysis measures the loss against something other than the claimant's own prior performance. It is the accepted alternative where the party's history is unavailable, too short, or too contaminated to serve as a baseline, and it substitutes comparable businesses, undamaged units, or industry-level measures.
It matters more in search litigation than in most fields because of who brings these claims. Many search damages claimants are young e-commerce or lead-generation businesses with eighteen months of history and a steep growth curve, which is the fact pattern the old new business rule was built to exclude. That rule — historically barring lost-profits recovery by enterprises without a track record — has been abandoned in most states but survives in modified form in several. Georgia retains it in its strongest form; Illinois carves out exceptions for products identical to existing ones in known markets and for acquisitions of going concerns; Iowa, Washington, and Pennsylvania apply versions functionally close to the reasonable-certainty standard. Where the claimant has no usable history, a benchmark is not a stylistic preference. It is the only route to a number a court can properly estimate damages from.
The second reason has nothing to do with missing history: search supplies unusually good benchmarks, and using them converts a correlation into a design.
The four controls search gives you that other fields do not
In a typical lost-profits case the expert must go outside the plaintiff to find a comparator, and comparability is then the whole fight. Search evidence is different: the dataset that shows the loss also contains several independent controls.
The party's own other channels
A website records how each visit arrived: organic search, paid search, direct, email, referral, social. If organic sessions fell 45% while direct, email, and paid held steady, demand did not collapse and measurement did not break; something specific to search happened. If every channel fell in the same proportion, the cause is demand, brand, or tracking, and no ranking theory explains it. This is the cheapest test in the discipline and the one most often skipped.
Unaffected page cohorts on the same site
A sitewide algorithmic event and a targeted defect leave different footprints. A cohort of pages the alleged conduct did not touch — a different template, directory, or content type — sits on the same domain, under the same brand, with the same seasonality and the same measurement, and provides a within-firm control no external comparator can match.
Competitor visibility over the same window
Third-party visibility indices track how prominently a set of domains appears for a keyword set over time. Used as a market-level control rather than as a measurement of a competitor's business, they answer whether the market moved on the date at issue.
Query-level demand
Public query-volume tools and Search Console impression data show whether the underlying searches were still being made. Without this control, a category-wide collapse in interest is easily and expensively mistaken for a collapse in ranking.
Difference-in-differences, stated as a method
Put two of those controls together with an event date and you have a design. A difference-in-differences comparison takes two groups — one affected by the event, one not — and two periods, before and after, and estimates the effect as the change in the affected group minus the change in the control group. What is claimed is not that traffic fell after the conduct, but that traffic in the affected cohort fell relative to a control cohort that experienced everything else it experienced.
Three properties make this the strongest available approach in search matters.
- It is testable. Run the same estimator across a pre-event date on which nothing happened. It should return approximately zero, and if it does not, the design has a problem you now know about.
- It has an assessable error rate. The estimate comes with dispersion across the cohort, so the opinion can state a range rather than a point, and can state what portion of the movement remains unexplained.
- It states its own assumption. The design rests on parallel pre-trends — the two cohorts moving together before the event — and that assumption is shown from the data rather than asserted.
This is what Kumho Tire Co. v. Carmichael, 526 U.S. 137 (1999), asks of a technical witness. The Court held that the gatekeeping obligation extends beyond scientific testimony to technical and other specialized knowledge, that the Daubert factors are neither mandatory nor exhaustive, and that those questions can help evaluate the reliability even of experience-based testimony. An expert in this field is a Kumho expert. Twenty years of practice is a qualification, not a method. A stated estimator, a stated control cohort, a stated event date, and a placebo result is a method — and it is what amended Rule 702(b) and (d) require: sufficient data, and a fit between what the method supports and what the opinion claims.
How cohorts are actually defined
The design is only as good as the assignment rule, so the rule has to be written down before the outcomes are examined. I would keep the artifacts of each step below.
- Enumerate the URL inventory as it existed on the event date, from a crawl, an archived sitemap, or a contemporaneous export — not from the site as it stands today.
- Assign by a mechanism-derived rule, not by outcome. Which URLs the redirect map touched, which templates carry the changed markup, which pages the disallowed directory contains. The rule must be statable in one sentence to someone who has never seen the data.
- Match the control cohort on pre-period characteristics: click volume, query intent, template, and conversion behavior. A cohort of thin category pages does not control for anything happening to product pages.
- Freeze the assignment and record when you froze it. A file, a timestamp, a hash. This answers the accusation that the cohorts were chosen to produce the answer.
- Fix the URL set to pages present in both periods, so pages added or removed mid-window do not create movement unrelated to the event.
- Document every exclusion and the reason for it, including the ones that hurt.
All of it should be reproducible from what is served with the report. Rule 26(a)(2)(B)(ii) requires the facts or data considered — not merely relied on — which here means the exports, the crawl files, the date ranges, the filters, and the queries or code that produced the cohorts. A benchmark opinion whose cohort assignment cannot be rebuilt by the other side is worth less than the chart it replaced.
How a cohort assignment gets attacked
Assume the assignment is where the challenge lands, because it is. Six attacks recur, and each has an answer that has to be built in advance rather than improvised.
- Selection on the outcome. "You put the pages that declined in the affected group." The answer is the ex ante rule and its timestamp, stated in terms of the mechanism rather than the result.
- Contamination. Control pages are linked from affected pages, share a template, or compete for the same crawl attention, so the event reached them too. That biases the estimate downward, and the honest response is to measure it or move the cohort boundary.
- Non-parallel pre-trends. If the cohorts were already diverging before the event, the design's assumption fails. Show the pre-period series, and where they diverge, say the design does not support the inference rather than running it anyway.
- Composition change. Pages published or removed mid-period create movement unrelated to the event, which is why the URL set is fixed to pages present throughout.
- Thin cohorts. Twelve URLs do not support a stated error rate. Give the count and the dispersion, and where the cohort is too small, say so.
- Post-hoc reclassification. Moving a page between cohorts after seeing the result is the most damaging admission available in a deposition on this subject, and it is why the frozen assignment file exists.
Answering all six is the difference between a method and an argument. Where one cannot be answered, the limitation belongs in the report rather than in the other side's brief.
External yardsticks and their foundation problem
Internal controls are the strongest part of this method. External ones — competitor visibility indices, keyword-volume estimates, traffic estimators — are weaker, and the weakness has to be stated rather than papered over.
These products are modeled estimates derived from proprietary panels and clickstream samples. They are not measurements of a competitor's traffic, and nobody outside the vendor can audit how they are produced. That has consequences under Rule 702(b), which asks whether the testimony rests on sufficient facts or data. An opinion resting on a third-party rank tracker or traffic estimator as its sole quantitative basis is a genuinely weak opinion, and I would say so from either side of a case.
What external data does well is establish direction and timing across a market: whether comparable domains moved on the same dates, in the same direction, by roughly the same magnitude. That is corroboration for a market-level control, and it is what internal cohorts cannot supply for a sitewide event. Where it is used, the report should state the tool, the keyword set, the location, the device, the sampling frequency, and the date the data was pulled.
What a benchmark model cannot do
Four limits, all better stated by the expert than extracted from him.
It does not establish who caused the divergence. The estimate shows that the affected cohort moved relative to a control. Tying that movement to the defendant still requires the mechanism — the directive, the redirect, the links, the disabled account — and a timeline that fits.
It cannot rescue a wrong event date. The design pivots on the date, and if the conduct's date is genuinely uncertain, no amount of cohort discipline repairs that.
It does not produce dollars. The output is a counterfactual traffic or visibility series. Conversion to revenue and incremental profit is a separate discipline with separate methods.
It can hurt the party who commissioned it. If the demand control shows paid and direct fell in step with organic, the claim has a problem — and that is the same check that would have supported the claim had it come out the other way. The symmetry is why the method is credible. An expert who runs it only when it helps is running something else.
Frequently Asked Questions
What is a yardstick damages analysis in a search case?
It measures the loss against a comparator rather than against the claimant's own history: a control group of pages on the same site, the party's other marketing channels, comparable competitor domains, or query-level demand over the same window. It is the accepted approach where the party's own history is missing, too short, or contaminated by its own changes. In search matters it is often the better method even when history exists, because the controls sit inside the same dataset and can be shown to move together before the event.Can another company's traffic be used as the benchmark?
Only with care. Nobody outside a company can measure that company's traffic. Competitor figures come from modeled estimates built on proprietary panels, and they are not auditable by the other side or by the court. Used as a market-level control for direction and timing — did comparable domains move on these dates, in this direction — they are informative. Used as a measurement of what the claimant would have earned, they invite a Rule 702(b) challenge to the sufficiency of the underlying data, and that challenge is usually well founded.What is difference-in-differences and why would a court care?
It compares the change in an affected group against the change in an unaffected control group across the same event date, so that anything affecting both groups — a core update, seasonality, market demand, a measurement change — drops out of the estimate. Courts care because it converts an experience-based assertion into a method with a stated assumption, a stated estimator, and an assessable error rate. That is what a reliability inquiry into technical testimony looks for, and it is the difference between an opinion that survives Rule 702 and one that reads as correlation.How do you decide which pages count as affected?
By a rule derived from the mechanism, written down and frozen before the post-event outcomes are examined. Which URLs the redirect map touched, which templates carry the changed markup, which directory the disallow covered, which queries returned a given feature on dated snapshots. The control cohort is then matched on pre-period volume, intent, template, and conversion behavior, and the URL set is fixed to pages present in both periods. The timestamp on that frozen assignment is what answers the accusation that the groups were picked to produce the answer.What if the plaintiff's other marketing channels fell at the same time?
Then the demand control is telling you something, and it belongs in the report. A proportional decline across organic, paid, direct, and email points to demand, brand, or measurement rather than to a ranking event, and a claim attributing the whole decline to search conduct will not survive that fact. The check cuts both ways: where the other channels held steady while organic collapsed, it is among the strongest evidence available that something specific to search occurred. An expert who runs the test only when it helps has no method.Is a rank-tracking tool enough to support a benchmark opinion?
Not as the sole basis. Rank trackers sample a fixed keyword set from fixed locations and devices on a fixed schedule, none of which matches the claimant's actual users, and the sampling and weighting are proprietary. They are useful corroboration that a market-wide movement occurred on a date, and useless as a measurement of what a specific site earned. An opinion whose only quantitative foundation is third-party rank data is exposed on the sufficiency of facts or data, and that is a fair criticism rather than a technicality.Does a benchmark model work for a business with almost no history?
It is usually the only model that can. The before-and-after approach needs a clean prior period, which a young business does not have, and several states still apply a modified version of the new business rule to enterprises without a track record. A benchmark built on comparable units, comparable page cohorts once the site is live, or documented industry conversion behavior gives a court data from which damages can be estimated. It does not lower the reasonable-certainty standard, and a hockey-stick projection dressed as a benchmark will not survive.Published