Flaky Tests: Why They Happen and How to Kill Them
Shalini works as a Technical Writer at LambdaTest. She loves to explore recent trends in test automation. Other than test automation, Shalini is always up for travel & adventure.
TL;DR: A flaky test passes and fails on the same code, with nothing changed. It isn't random, it only feels random. Luo et al. (FSE 2014) mapped the sources: async wait issues (45%), concurrency (20%), test order dependency (12%). The workflow for flaky tests is three moves: quarantine it so it stops poisoning your suite, diagnose the real root cause instead of adding a retry, then fix or delete it under a hard deadline. Retries are a painkiller, not a cure.
I've spent too many hours staring at a red CI run, hitting re-run, watching it go green, and shipping anyway. That reflex is the most expensive habit in test automation. A big part of my job at TestMu AI (formerly LambdaTest) is helping teams pull reliable signal out of large parallel runs, and flakiness wrecks that signal more than any single bug does. Here's the sequence I use to kill it. No AI, no tooling hype, just mechanics.
What is a flaky test, exactly?
A flaky test is one that produces both a passing and a failing result against the same code, with no change to the code under test, the test itself, or the environment. You run it, it fails. You run it again, it passes. Nothing changed.
That last part is what makes flakiness so corrosive. A test that always fails is useful, it's pointing at a bug. A test that flips is worse than useless, because it trains you to ignore it. Martin Fowler put it best back in 2011: non-deterministic tests have two problems, they're useless as a signal, and they're a virulent infection that can ruin your entire suite. Once your team learns to shrug at red builds, a real failure hiding among the flakes sails straight through.
Why should you care? The cost is bigger than the re-run
The obvious cost is wasted CI compute from re-running builds. The real cost is trust, and it compounds.
Even at the highest end of engineering maturity, this is unsolved. Google has publicly reported that about 1.5% of their test runs are flaky, and that flakiness affects a striking share of their overall suite. When you're running millions of tests, 1.5% is an enormous, permanent tax on developer productivity, enough that Google built dedicated tooling and a dedicated team just to manage it.
Here's the mechanism that makes it dangerous, in Google's own framing: if a meaningful fraction of results are flaky, engineers start reflexively dismissing failures as "just a flake." That's the same psychology as pilots learning to ignore a false alarm. Eventually a legitimate failure gets waved through, and broken code ships. The flakiness didn't just waste time, it defeated the entire point of having tests.
What actually causes flaky tests?
This is the part most teams skip. They treat each flake as a one-off annoyance instead of an instance of a known pattern. But the research is remarkably consistent. Luo et al. analyzed 201 flakiness-fixing commits across 51 open-source projects and built a taxonomy of ten root causes that everyone still cites. Later work, including Eck et al.'s 2019 study of 200 flaky tests from Mozilla's tracker, confirmed the same top offenders. Here are the ones that matter in practice.
Async wait (roughly 45%, the big one)
Nearly half of all flakiness comes from a single mistake: the test performs an action and asserts on the result before the application has finished processing it. The test is racing the app, and sometimes it loses.
The classic tell is a hardcoded sleep. sleep(2) works on your machine on a good day and fails on a loaded CI runner where the operation took 2.3 seconds. You are guessing at a duration instead of waiting for a condition.
The fix: replace fixed sleeps with explicit waits that poll for the actual condition you care about. Wait for the element to be visible, for the network request to resolve, for the state to become true, with a timeout as a backstop. Modern frameworks bake this in: Playwright's auto-waiting, for example, addresses this single largest root cause at the framework level, which is a big part of why teams migrating to it report lower flake rates.
# Flaky: guessing at timing
click(submit_button)
sleep(2)
assert success_banner.is_visible()
# Stable: waiting for the actual condition
click(submit_button)
wait_until(lambda: success_banner.is_visible(), timeout=10)
assert success_banner.is_visible()
Concurrency (roughly 20%, the hard one)
Race conditions, deadlocks, and atomicity violations between threads or processes. These are the flakes that make you question your career choice, because the failure depends on the exact interleaving of operations, which changes run to run.
The fix: there's no one-liner here. You isolate shared mutable state, make critical sections genuinely atomic, and where possible test the concurrent logic through a deterministic seam rather than by spinning up real threads and hoping. This is code-level investigation, not a test tweak, and it's why concurrency flakes take the longest to resolve.
Test order dependency (roughly 12%)
Test B passes when it runs after Test A, and fails when it runs alone or first, because it silently depends on state that A left behind. A shared database row, a /tmp file, a global variable, an authenticated session. Run your suite in a new order (or in parallel) and these surface immediately.
The fix: make every test set up its own preconditions and tear them down after. No test should ever assume another test ran first. Randomizing test order in CI is a cheap way to flush these out before they bite you.
Environment, time, and resource leaks (the rest)
The remaining causes cluster around the machine and the clock. Tests that depend on the current date break at midnight, at month boundaries, or in a different timezone. Tests that leak file handles or memory pass in isolation and fail in a long suite once the resource runs out. Tests that hit real external services inherit every bit of that network's unreliability.
The fix: inject the clock instead of reading it directly, so "now" is controllable in tests. Clean up resources deterministically. And stub external services at your boundary rather than calling them for real inside a unit or integration test.
How do you find which tests are flaky?
You can't fix what you can't see, and flakiness is invisible in a single run by definition. You need history.
The core metric is simple: over a rolling window (7 to 30 days is typical), count the tests that produced both a pass and a fail without a code change, and divide by your total test count. That's your flake rate. Teams like GitHub and Spotify go further and assign each test a per-test flakiness score so they can rank fixes by impact rather than chasing whichever flake screamed loudest today.
The trap Fowler and others keep warning about is the spreadsheet. Teams start tracking flakes manually, nobody enjoys it, and within a month everyone quietly stops. Whatever you use, detection has to be automatic and continuous, wired into your CI, or it won't survive contact with a busy sprint. Google's own research pushed this further with tooling to automatically locate the root cause of a flake in the code, studied across 428 of their projects, precisely because manual triage doesn't scale.
Should you just retry flaky tests?
This is the question every team eventually asks, usually at 5pm before a release. The honest answer: retries are a legitimate tool and a dangerous crutch, and the difference is entirely in whether you also investigate.
Google uses re-runs as a mitigation. They even have a mechanism where a test is only reported as failing if it fails three times in a row. But they're clear-eyed about the cost: this reduces false positives while quietly encouraging engineers to ignore flakiness in their own tests until it gets bad enough to fail three times. It also means a genuinely broken 15-minute integration test won't be flagged until three executions, 45 minutes, later.
So my rule is: a retry is allowed to keep the build moving, but it is never allowed to close the ticket. Every retry that saved a build should also file (or increment) a record that the test is flaky, so it enters the fix queue. Retry to survive the day; diagnose to actually win.
What workflow keeps them from coming back?
Killing flaky tests is a process, not a one-time cleanup. This is the loop I run:
1. Quarantine on sight. The moment a test is confirmed flaky, move it out of the blocking suite into a separate quarantined suite. This is Fowler's central advice, and it's right: a flaky test in your main suite is actively degrading the signal of every healthy test around it. Quarantine stops the bleeding.
2. Quarantine with a hard SLA. This is the step everyone forgets, and it's the one that matters. A quarantined test is a test you've decided not to run, which means it's no longer protecting you from regressions. Quarantine without a deadline is just deletion with extra steps. Every quarantined test gets an owner and a fix-or-delete deadline, one sprint, no exceptions.
3. Diagnose against the taxonomy. Don't debug blind. Walk the known causes in order of likelihood: is it an async wait (start here, it's usually this), an order dependency (run it in isolation), a concurrency issue, a time or environment assumption? The taxonomy turns "this test is haunted" into a checklist.
4. Fix the root cause, or delete the test. If the test earns its place, fix the actual cause using the treatments above. If nobody will own it and it's not covering anything critical, delete it. A deleted flaky test is strictly better than a quarantined one you'll never fix, because at least it stops consuming attention.
5. Assign ownership deliberately. Don't dump the fix on whoever tripped over the failure, they're frustrated and lack context. A better default: the last person who touched that test. Best of all, treat CI health as a shared team responsibility with real time budgeted for it, because a suite nobody maintains will always drift back toward flaky.
FAQ
Are flaky tests actually random?
No. They feel random because the failures are intermittent, but every flaky test has a deterministic root cause, it's racing against something, depending on something it shouldn't, or assuming something that isn't always true. The Luo et al. taxonomy catalogs the ten recurring causes.
What is the most common cause of flaky tests?
Async wait, by a wide margin, roughly 45% of cases in the foundational research. The test asserts on a result before the application has finished producing it. The usual culprit is a hardcoded sleep() instead of an explicit wait for a specific condition.
Is it OK to retry flaky tests in CI?
As a stopgap, yes; as a solution, no. Retries keep the build moving but hide the underlying problem, and over-relying on them trains your team to ignore failures. Google uses retries but pairs them with detection, quarantine, and tracking. A retry should never close the investigation.
What does "quarantining" a flaky test mean?
Moving a confirmed flaky test out of your blocking suite into a separate, non-blocking one so it stops corrupting the signal from your healthy tests. It's Martin Fowler's core recommendation, but only works if quarantine comes with a strict fix-or-delete deadline. A quarantined test is a test not protecting you from regressions.
Can a flaky test be hiding a real bug?
Yes, and this is why you can't just delete them all reflexively. Sometimes the non-determinism in the test mirrors non-determinism (a real race condition) in production code. In the Luo et al. data, a meaningful share of flakiness fixes required changing production code, not just the test. Diagnose before you dismiss.
How do I measure my flake rate?
Over a rolling 7-to-30-day window, count tests that produced both a pass and a fail with no code change, then divide by total test count. Track it continuously in CI, not in a manual spreadsheet, which teams inevitably abandon.



![How To Run JUnit Tests In Jupiter? [JUnit Jupiter Tutorial]](https://cdn.hashnode.com/res/hashnode/image/upload/v1677484137604/1d4533db-9c1e-4b1c-b808-6b981f7ebdf6.jpeg)