The False Positive in the Iceberg Lettuce


Have you ever made a decision based on a pile of evidence…

…and then had everyone fixate on the one number in the pile?


Britain has already established that an iceberg lettuce can provide a surprisingly effective measure of institutional stability.

In October 2022, the Daily Star placed one beside a photograph of Liz Truss and began a livestream to discover which would last longer.

The lettuce won.

Truss resigned six days into the experiment, while the increasingly decorated vegetable remained sufficiently intact to receive a golden crown and a celebratory glass of prosecco.

Four years later, another member of the lettuce family has found itself at the centre of a rather more consequential experiment.

This time, the question was not whether it could outlast a prime minister.

It was whether it had made 1,644 people ill.

In October 2022, an iceberg lettuce famously outlasted Liz Truss as Prime Minister. It remains one of Britain’s more memorable experiments in institutional durability.

Imagine that you have just completed an A/B test.

  • The result is significant.
  • The variant wins.
  • The dashboard is green.

Someone has already pasted the uplift into a presentation and added a projected annual revenue figure with rather more confidence than the situation warrants.

There it is: proof.

Except the following day, your analysts examine the result again. And the win disappears.

A significant uplift, a projected revenue figure and a reassuring green interface can make an experiment feel conclusive—until the mechanism beneath the result is examined.
  • Not because the effect has gradually weakened.
  • Not because a later segment behaved differently.

Because the original positive result should never have been declared positive at all.

What would that tell you? That the result was wrong? Certainly.

That the entire hypothesis was wrong? Not necessarily.

And that distinction sits at the centre of the Taylor Farms lettuce saga.

The apparently conclusive result

US health authorities had been investigating a large outbreak of Cyclospora, a parasite that can cause prolonged and deeply unpleasant gastrointestinal illness.

The cases had something in common.

They involved people who had eaten at Taco Bell restaurants in Indiana, Kentucky, Michigan, Ohio and West Virginia. Investigators examined detailed food histories from 190 cases, and 90% of those people reported eating iceberg lettuce.

The FDA then traced the lettuce served by the affected restaurants backwards through the supply chain.

Those routes converged on one supplier: Taylor Farms de Mexico.

On 17 July 2026, Taylor Farms began removing iceberg lettuce sourced from central Mexico from the US market and initiated a recall.

Then, on 18 July, the investigation appeared to acquire the piece of evidence everyone had been waiting for.

A sample of lettuce supplied by Taylor Farms tested positive for Cyclospora.

The interviews pointed to lettuce.

The supply-chain records pointed to Taylor Farms.

And now the product itself appeared to contain the parasite.

Hypothesis confirmed. Winner declared.

Except the next day, FDA laboratory specialists reviewed the result and concluded that the signal did not represent genuine amplification.

The positive result was a false positive.

As of 19 July, none of the tested product samples had produced a confirmed positive result for Cyclospora.

The apparently conclusive result was not conclusive.

It was not even positive.

The A/B test interpretation

For experimenters, this is a familiar fear.

A false positive is what happens when a test tells us that an effect exists when it does not

  • The variant appears to have beaten the control.
  • The new checkout appears to have increased conversion.
  • The rewritten headline appears to have improved sign-ups.

But the apparent difference may have been created by chance, faulty measurement, contamination, an analytical mistake or some other feature of the testing process.

The result looks real. It passes the mechanism intended to protect us from imaginary wins. And yet it is still wrong.

That is why statistical significance has never meant: We have proved that the variant is better.

It means something much narrower:

Assuming the test and its underlying conditions are valid, the observed result would be relatively unlikely under the null hypothesis.

There are quite a few escape hatches hidden inside that sentence.

  • The tracking must be working.
  • The samples must be comparable.
  • The stopping rule must not have been improvised halfway through.
  • The metric must mean what we think it means.
  • The analysis must have been performed correctly.
  • And the test result must survive closer examination.

A green dashboard does not repeal uncertainty. It merely gives uncertainty a more reassuring interface.

So was the lettuce innocent?

Once the FDA withdrew the result, Taylor Farms understandably emphasised that no product had produced a confirmed positive test.

And on the surface, the reversal appeared to destroy the case.

Positive lettuce test: the supplier was responsible.
False-positive lettuce test: the supplier had been wrongly accused.

But that assumes the laboratory result was the original reason for suspecting the lettuce.

It wasn’t.

The recall had already begun before the positive test was announced.

Investigators had reached Taylor Farms through a combination of evidence:

  • a large group of people had become ill;
  • they had eaten at the same restaurant chain;
  • iceberg lettuce was the ingredient reported by 90% of the cases examined in detail;
  • and the supply routes from the affected restaurants converged on the same supplier.
The laboratory result was withdrawn, but the illness patterns, restaurant data and supply-chain tracing did not disappear with it. One failed data point weakened the case; it did not erase the wider investigation.

The laboratory result arrived later. It did not create the hypothesis. It appeared to validate it.

The FDA therefore maintained that withdrawing the result did not overturn the wider investigation, while Taylor Farms stressed that there was still no confirmed contaminated product sample.

Both positions contain something important.

The false positive weakened the case.
It did not erase every other piece of evidence.

Nor did the wider pattern magically make the faulty result valid.

The test was wrong. The hypothesis might still be right.

The result becomes the story

This is where the lettuce saga stops being a food-safety curiosity and becomes an experimentation problem. Because most organisations say they use multiple forms of evidence.

  • They combine analytics with research.
  • They examine behavioural patterns.
  • They speak to customers.
  • They review operational data.
  • They investigate what happened before and after the visible moment.
  • Then an A/B test produces a significant result.

And suddenly the experiment becomes the evidence.

Everything else is demoted to interesting background material.

  • The result appears in the board deck.
  • The uplift acquires a currency symbol.

A neat causal story is constructed around it: We ran the test. The variant won. Therefore the change caused the improvement.

That story is easier to explain than the truth:

We observed a result within a system containing users, interfaces, acquisition sources, technical dependencies, operational changes, external events and a substantial amount of uncertainty.

One version fits comfortably on a slide. The other tends to ruin it.

So we remember the green result and forget the evidence chain underneath it, until the result is questioned.

  • Perhaps the test cannot be replicated.
  • Perhaps the uplift disappears when novelty fades.
  • Perhaps the primary metric improved while complaints, returns or cancellations worsened.
  • Perhaps someone discovers that the tracking fired twice.

Or perhaps, like the lettuce sample, the positive result was never positive in the first place.

At that point, teams often swing from excessive certainty to excessive dismissal.

  • The test was wrong, therefore the idea was wrong.
  • The metric failed, therefore the research was wrong.
  • The result was a false positive, therefore there was never anything worth investigating.

But a failed experiment result tells us something about the result. It does not automatically settle every question surrounding the hypothesis.

The Corpus reframe

At Corpus, we treat experiments as part of an evidence chain rather than machines that manufacture verdicts. A result becomes useful when we can connect it to the surrounding reality:

  • What did we observe?
  • What else changed?
  • What mechanism could explain the result?
  • Does qualitative evidence support or contradict it?
  • What happened to the measures that were not designated as the primary metric?
  • Can the finding survive another test, another segment or another context?

The aim is not to undermine experimentation. It is to stop demanding more certainty from an experiment than it can provide.

A false positive matters. It can waste money, misdirect strategy and persuade an organisation to scale something that never worked.

But the greater failure is often narrative. We allow the cleanest result to replace the entire body of evidence. Then, when that result fails, we discover that nobody can remember why the decision made sense in the first place.

The FDA may eventually establish that Taylor Farms lettuce caused the outbreak. It may establish that it did not.

For now, the honest position is more awkward: The product test was wrong.

The wider investigation still points towards the lettuce.

And one iceberg has already taught Britain that sometimes the most revealing part of an experiment is simply seeing what remains standing at the end.

Abi Hough

Abi Hough is the founder of Corpus and a UX, experimentation and accessibility strategist. Her work examines the gap between what organisations claim, what their systems communicate and what people ultimately understand.

Talk about how this applies in your organisation.

If a field note resonates and you want to talk about how the same patterns are showing up where you work, a conversation can help.
Typical first conversations last 45 to 60 minutes and focus on understanding your current situation and constraints.
Upstream optimisation for zero click and AI search.
Contact
[email protected]

Typical first conversations last 45 to 60 minutes and focus on your current situation, constraints and goals
We Are Corpus is an Abi Hough-led consultancy and the trading name of UU3 Ltd, registered in England and Wales. Company number 06272638.