Five Sigma
When Physics Agrees to Believe Something
The Bump That Was Not There
In December 2015, both big detectors at the Large Hadron Collider saw the same strange thing. Pairs of photons were coming out of collisions carrying a particular combined energy slightly more often than they should have. Two independent experiments, the same bump, in the same place. Theorists noticed within hours. Over the following months, more than five hundred papers were written explaining what the new particle might be.
There was no particle. By August 2016, with roughly four times as much data in hand, the bump was gone. It had been noise that happened to line up in two places at once. Nobody had cheated and nobody had blundered. The collaborations had reported their numbers correctly, and the numbers had simply been a coincidence of the kind that a large enough dataset produces on schedule.
This is the ordinary condition of experimental physics. Every dataset large enough to be interesting contains bumps, dips, and coincidences that mean nothing. The hard problem is never seeing something unusual. The hard problem is deciding how surprised to be. Physics has an answer to that, and the answer is a number: five sigma.
What Sigma Counts
Repeat any measurement many times and the answers scatter. Sigma is the name for the width of that scatter – the typical distance between one measurement and the average of all of them. It is a ruler made out of the experiment’s own noise, which is what makes it useful. An excess of two hundred events means nothing until you know whether the ordinary wobble is ten events or a thousand.
So a result quoted at 3σ is not a claim that the effect is large. It is a claim that the effect is three times the size of the experiment’s own random jitter. Sigma answers exactly one question, and it is worth stating precisely: if there were nothing there at all, how often would noise alone produce something at least this striking?
The answers fall off with brutal speed. A one-sigma excess turns up by chance about one time in six, which is to say constantly. Two sigma comes up roughly one time in forty-four. Three sigma, one time in seven hundred and forty. Five sigma is one time in three and a half million. Each step out is not a small tightening. From three sigma to five sigma the odds against a fluke improve by a factor of nearly five thousand.
One detail matters here and is often skipped. Particle physics counts only excesses, not deficits, because a new particle can only add events to a spectrum. That one-sided convention is why the number quoted for five sigma is one in three and a half million rather than one in one and three quarter million. Different fields count differently, so the same symbol does not always mean the same odds.
Why Five and Not Three
Five sigma is not a law of nature. It is a scar. The threshold sits where it does because physics spent decades watching three-sigma results die, and eventually stopped believing them.
The graveyard is well populated. A pentaquark reported in 2003 was seen by roughly ten experiments and then unseen by better ones. A signal for primordial gravitational waves announced in 2014 turned out to be dust in our own galaxy. The two-photon bump of 2015 evaporated. Each of these was reported honestly, and each looked, at the moment of announcement, like the beginning of something.
The reason so many die has nothing to do with dishonesty. Modern experiments examine enormous numbers of possible signals. When you look at a hundred thousand places where something could be hiding, a one-in-seven-hundred coincidence is not rare. It is guaranteed, over and over. Three sigma is roughly the level at which interesting things start appearing by accident faster than they appear for real.
So the field moved the line out to a place where accidents essentially stop. The convention hardened during the 1990s and became near-universal after it was applied to the top quark and later the Higgs boson. It costs something real. A genuine discovery sitting at four sigma has to wait, sometimes for years, while more data accumulates. Physics decided that being slow is cheaper than being wrong in public.
Real Signals Grow, Flukes Wash Out
There is one test that separates a real effect from a lucky accident better than any threshold, and it costs nothing but patience. Take more data and watch what the bump does.
A real signal produces events at a steady rate. Collect four times the data and you collect four times as many signal events. The background under it also grows, but background is random, so its wobble only grows as the square root. The signal outruns its own noise. Significance climbs roughly with the square root of the data collected, which is why a genuine three-sigma hint becomes a six-sigma discovery once you have quadrupled the dataset.
A fluke has no such engine. It was an accident of a particular set of collisions, and the next set of collisions has no memory of it. As data accumulates, the excess does not grow. It gets diluted, then swallowed. Watching significance fall while data rises is the unmistakable signature of something that was never there.
This is the honest reason physicists are calm about anomalies. They are not refusing to be excited. They are waiting for the one measurement that costs nothing to interpret, because a curve that climbs and a curve that sags mean entirely different things and no argument can disguise which one you are looking at.
The Look-Elsewhere Effect
Suppose you predict in advance that a new particle will show up at one specific energy, and it does, at three sigma. That is a one-in-seven-hundred coincidence and genuinely surprising. Now suppose instead that you scanned a whole spectrum, found a bump somewhere, and only then pointed at it. Those are completely different claims, and only one of them is impressive.
The arithmetic is unforgiving. Search one place and the odds of being fooled at three sigma are about one in seven hundred and forty. Search five hundred independent places and the odds that at least one of them throws a three-sigma bump are close to even. You have not done anything wrong. You have simply bought that many chances to be fooled, and the price shows up in the answer.
Physics handles this by quoting two numbers. Local significance is how striking the bump is where it sits. Global significance is what remains after accounting for every place you could have found one. The gap between them is often large. The 2015 two-photon bump reached about 3.9σ locally at one experiment, and roughly 2σ globally once the search range was counted. In hindsight the global number was telling the truth the whole time.
When you read about an anomaly and only one number is given, it is almost always the local one, because it is the larger and more exciting of the two. Which of the two is being quoted is the first thing worth asking.
When the Apparatus Itself Is Wrong
Everything so far assumes the apparatus is telling the truth and only randomness stands between you and the answer. That assumption is where most famous mistakes actually live.
In 2011, OPERA, an experiment measuring neutrinos sent through the earth from Geneva to a detector in Italy, reported that they arrived about sixty billionths of a second earlier than light would have. The statistical significance was around six sigma – past the discovery threshold, on a result that would have broken relativity. The collaboration did not claim a discovery. They published the measurement and asked the community to find their mistake.
The mistake was a fiber optic cable that was not fully seated, plus a clock oscillator running slightly fast. Both shifted every single timestamp by the same amount. This is what makes such errors deadly. Random noise averages away when you collect more data, but a constant offset does not. It survives every repetition perfectly and comes out looking like a rock-solid signal that only gets more significant.
Physicists call these systematic uncertainties, and estimating them is the least glamorous and most consequential part of an experiment. Statistical error tells you how much the answer wobbles. Systematic error tells you how far the whole scale might be shifted, and nothing in the data itself can reveal it. That is why a stated significance is only as trustworthy as the list of systematic effects somebody remembered to check.
Hiding the Answer From Yourself
There is one more failure mode, and it is the most human. Analyzing data involves hundreds of small choices about which events to keep and how to model the background. Every one of those choices nudges the answer. If you can see the answer while you are choosing, you will drift toward the version you were hoping for, without ever noticing that you did.
The defense is blind analysis. The region of the data where the signal would appear is hidden from the people doing the work. They tune and validate everything on the surrounding regions and on simulations, then write down the complete procedure and freeze it. Only then is the box opened, and whatever comes out is the result. Some collaborations go further and add a secret artificial offset, so that even the shape of the answer stays unknown until the number is unhidden at the end.
Gravitational wave astronomy took the idea to its logical end. A tiny group had the power to inject a fake signal into the detectors without telling anyone. The full collaboration would find it, analyze it for months, and write the discovery paper. Only at the moment before submission was the envelope opened to reveal whether the event had been real. Twice this happened before the first genuine detection, and the rehearsal is a large part of why the 2015 announcement was trusted immediately.
Two Detectors, One Answer
On the fourth of July 2012, ATLAS and CMS announced independently that each had found a new particle near a mass of one hundred and twenty five giga-electronvolts. Both quoted five sigma. That pairing is the reason the announcement was believed on the day rather than a year later.
The two detectors are built on different principles, with different magnets, different materials, and separately written software. They were kept deliberately independent, and each was blinded to the other’s results while the analyses were being finalized. A systematic effect that fools one has little reason to fool the other in the same direction by the same amount.
This is the part that five sigma alone cannot supply. The threshold protects against randomness. Independent replication with different hardware is the only real protection against the errors that randomness cannot reach. When physicists say a result is solid, they usually mean both conditions were met, not just the first one.
Reading Today’s Anomalies
Several results currently sit in the uncomfortable zone, and the framework above is enough to read them properly.
The muon’s magnetism was for years the most quoted anomaly in physics, with experiment and prediction separated by more than four sigma. It has since deflated, and it deflated from the direction fewer people were watching. The measurement only got better, and Fermilab’s final result is the most precise ever made. What moved was the prediction. Working out what the Standard Model expects requires a hard piece of quantum chromodynamics, and as lattice computations converged and became the reference, the predicted value shifted toward the measurement and the case for new physics faded. The two ways of doing that calculation still disagree with each other, which is now a question about the calculation rather than about the muon.
The Hubble tension is a different animal. Two well-understood ways of measuring how fast space is expanding give answers that differ by around five sigma, and the gap has widened as both sides got more precise. Because the two methods share almost no equipment and no assumptions, the usual explanation of a single overlooked systematic has become progressively harder to sustain. Nobody has demonstrated an error, and nobody has demonstrated new physics.
A cluster of anomalies in the decays of B mesons offers the cleanest recent lesson. Around 2021 several measurements pointed the same way at a combined significance that had theorists writing seriously about a new force. Improved analyses with more data brought the headline results back into agreement with the Standard Model. The bumps did not survive. That outcome is the system working, not the system failing.
A Promise Made in Advance
Five sigma is not really a statement about statistics. It is a promise made in advance, by a community that knows exactly how good it is at finding patterns in noise. Human beings are superb pattern detectors, and that talent does not switch off in a control room. The threshold exists because physicists decided not to trust their own enthusiasm, and wrote the distrust down as a number before the data arrived.
Other fields drew the line somewhere else. Much of medicine and psychology settled on a standard near two sigma, counted in both directions, which corresponds to being fooled roughly one time in twenty rather than one in forty-four. Combine that with thousands of researchers testing thousands of hypotheses and publishing mainly the ones that worked, and a large fraction of published findings fail when someone tries to repeat them. The mathematics is the same everywhere. The difference is where each field put its threshold and whether it demanded independent replication before believing.
None of this makes physics immune. Five sigma says nothing about a loose cable, and it says nothing about a mistake that everyone in the field shares. What the convention does buy is a shared, public, and boring standard, applied the same way to a result you love and a result you hate. That is a modest thing to build a science on, and it has turned out to be enough.




