Flaky Tests: Why They Happen and How to Fix Them

Flaky Tests: Why They Happen and How to Fix Them

Learn what causes flaky tests, how to diagnose unreliable automation and when retries or self-healing help - or simply hide the problem.

Back to Publications

The build fails. Someone reruns it without changing the code. This time, everything passes. The immediate problem appears to be gone, but the team has learned nothing. Was there a real regression? Did the test start too early? Did another test change the data? Was the environment temporarily unavailable?

A test that sometimes passes and sometimes fails under the same intended conditions is flaky. Over time, unreliable results change how people treat the suite. Engineers rerun failures instead of investigating them, and genuine defects become easier to dismiss. Fixing flaky tests starts with identifying the source of uncertainty - a broken selector needs a different remedy from shared test data, an overloaded environment or an incorrect wait. Retries and self-healing can help in specific situations, but neither should become a blanket solution.

Code Displaying an Intermittent Test Failure

What Is a Flaky Test?

A flaky test produces inconsistent results without a relevant change to the application or test. It may pass on one run and fail on the next because its outcome depends on timing, execution order, shared state, infrastructure or another uncontrolled condition. This differs from a repeatable failure, which usually points to a product defect, an outdated expectation, or a persistent test problem. Flakiness is harder to diagnose because a rerun can make the evidence disappear.

The goal should not be to keep every build green. It should be to make failures meaningful. When a test fails, the team needs enough confidence to treat that result as information rather than noise.

Common Causes of Flaky Tests

The same symptoms can have different causes. A missing checkout button might indicate a changed locator, slow page state, failed API response or insufficient permission - treating all four as one problem creates the wrong fix.

  • Unstable selectors - the test sometimes finds the wrong element or fails after a UI change; prefer stable roles, labels or test IDs and use controlled self-healing for genuine locator changes.
  • Timing and synchronisation - an element, request or background job is not ready when the assertion runs; wait for a meaningful condition instead of a fixed delay.
  • Shared test data - tests update or delete the same user, order or account; create isolated data or assign unique records per test.
  • Environment instability - failures cluster around deployments, resource pressure or unavailable services; monitor environment health separately from product results.
  • Test-order dependency - a test passes in the suite but fails alone, or the reverse; give every test its own setup and cleanup.
  • Network and API issues - requests time out or return temporary errors; capture request details and use mocks only where they preserve the purpose of the test.
  • Parallel-execution conflicts - tests pass sequentially but fail when workers run together; isolate accounts, files, ports and other shared resources.
  • Weak assertions - the test checks before the system reaches the expected state; use condition-based assertions tied to the real outcome.

Replace Fixed Waits With Observable Conditions

Timing failures often begin with assumptions such as "the page will be ready in two seconds." Increasing every timeout makes the suite slower and may only reduce how often the problem appears. Wait instead for evidence that the system has reached the required state: an element is visible and enabled, a network request has completed, a status has changed, a record appears in the response, a loading indicator has disappeared, or the expected business result is available. Modern frameworks support condition-based waiting - Cypress, for example, retries linked queries and assertions until they succeed or time out.

Isolate Tests, Data and Parallel Runs

A reliable test should control its own state. It should not assume another test created a user, left an item in a basket or updated an account - shared accounts cause similar problems when one worker changes data another worker still needs. Use unique data where possible and establish required state during setup; browser states should also be isolated so cookies, local storage and sessions do not leak between tests.

A suite that becomes unreliable only in parallel usually depends on a resource incorrectly assumed to be exclusive - reusing the same account or email address, writing to the same file or storage location, updating the same database record, depending on a shared queue, or using fixed ports. Run suspected tests with different worker counts; if failures appear only under concurrency, look for shared resources before adjusting timeouts.

Separate Product Failures From Environment Failures

An overloaded environment, expired credential, unavailable dependency or incomplete deployment can make valid checks fail. Capture the evidence needed to distinguish these causes: application and service logs, network requests and responses, screenshots or traces, environment version and deployment status, test data identifiers, worker and retry number, and timestamps around the failure. "Element not found" is rarely enough evidence - the page may never have received the data required to render it.

When Retries Help - and When They Hide Problems

Retries can collect more evidence or keep one intermittent failure from blocking a diagnostic run, but a retry does not fix the cause. If the first attempt fails and the second passes, the test should still be investigated - Playwright explicitly categorises this result as "flaky." Use retries with limits and visibility: record the initial failure and retry result, track which tests require retries, avoid repeatedly retrying destructive actions, never count "passed after retry" as equivalent to a first-run pass, and assign frequently retried tests for investigation. Temporary quarantine needs an owner and review date, or important coverage can quietly disappear.

Where Self-Healing Is Appropriate

Self-healing is relevant when the application changes but expected behaviour does not - for example, when a button keeps the same role and purpose after its ID changes. It cannot fix slow data, failed APIs, incorrect results, or order dependency; locator healing in those situations treats only a symptom. A trustworthy workflow records the original locator, replacement, confidence and subsequent assertion result, and low-confidence changes should require review.

How to Measure and Reduce Flakiness

Track flakiness at the individual-test level instead of relying only on the pipeline pass rate. Useful measures include tests that pass only after retry, flaky failures by cause and component, first-run pass rate, time spent investigating non-product failures, tests currently quarantined, reopened flakiness issues, and failures that cannot be reproduced.

  • Identify the tests producing the most unreliable runs
  • Preserve traces, logs, network evidence and test data from failures
  • Reproduce under different orders, worker counts and environment conditions
  • Classify the cause before choosing a fix
  • Verify the fix through repeated clean runs and continue monitoring it

How Nogrunt Helps Teams Improve Test Reliability

Nogrunt's self-healing capability can address recoverable UI changes while keeping actions visible for review. Its failure-analysis capabilities help teams examine execution evidence rather than relying on a one-line message, and parallel execution and delivery integrations keep results connected to the wider quality process. The objective is not to turn unreliable tests green - it is to reduce maintenance noise without weakening the suite's ability to expose genuine defects.

Make Every Failure Worth Investigating

Flaky tests are not just an automation inconvenience. They weaken the feedback loop that teams depend on to make release decisions. Start with the tests that fail most often or protect the highest-risk journeys. Capture enough evidence to classify each failure, fix the source of uncertainty and verify the result across repeated runs.

A reliable suite does not need to pass all the time. It needs to fail for reasons the team can trust.