A CI Retry Needs a Failure Policy

Published · Programming

A green check after a retry can mean several things. A shared service recovered. A runner briefly lost its network connection. A test has a race that happened not to fire twice. Or the code is broken, and the rerun happened against a different state. Treating all four as “CI was flaky” makes the build look healthier while making the engineering system harder to trust.

Retries are useful. They can distinguish a transient incident from a repeatable failure and restore progress after an external service hiccup. But a retry is an experiment, not an eraser. Its value depends on preserving what failed first, controlling what changed, and deciding what the second result is allowed to mean.

Separate the failures before choosing a response

Start with the failure layer. The same red check can describe very different problems.

Failure Useful immediate response What a retry establishes
Runner never starts or loses its workspace Inspect runner and service status; retry the job once after recovery The execution environment can now run the job; the original result was not a code verdict
Dependency download or external test service times out Record the dependency and incident window; retry when it is available The external dependency recovered, if the inputs stayed fixed
One test assertion fails intermittently Keep its name, log, seed, and failure output; rerun under the same inputs A pass suggests nondeterminism but does not prove correctness
Compiler error or deterministic test failure Fix the change or configuration Usually nothing useful unless the failure was traced to a changing environment
Entire job fails after an unrelated commit lands Check the exact commit and workflow inputs A new run on new code answers a different question

This classification should be quick enough for a developer under deadline. It is not a demand for a formal incident before every rerun. It is a guard against clicking “rerun” until the result turns green and then forgetting what happened.

Keep the first attempt in the record

The useful unit of evidence is the whole attempt history: commit, workflow version, runner or image, failing test or step, logs, and each subsequent outcome. A status badge that reports only the last green attempt loses the part you need to improve reliability.

GitHub Actions reruns use the original event's commit and ref, which makes a rerun different from pushing a new commit and starting a fresh run. Its interface can rerun all jobs, failed jobs, or one job. Choose the smallest unit that still tests the suspected failure. If the runner itself failed before tests started, rerunning that job may be enough. If setup changed an external resource used by several jobs, inspect the broader workflow before assuming one isolated retry is meaningful.

For a test-level retry, record at least the test identifier, first failure message, attempt count, and final outcome. Some test tooling can express “passed on retry” in reports. Gradle's JVM testing documentation discusses flaky-test reporting and retry support. The important policy is independent of the tool: a pass-on-retry must remain discoverable in CI output and in the team's reliability view.

Put a limit on automatic retries

An unlimited retry loop turns a failing build into a slow lottery. A small, explicit limit is easier to reason about: for example, one automatic retry for a known transient setup step, or one diagnostic test rerun that is clearly marked. Choose the limit from your actual failure modes and CI budget, not from a desire to make red checks disappear.

Avoid a blanket policy that retries every test failure. It multiplies CI time, masks new regressions that happen to be intermittent, and makes a green branch mean “eventually passed.” If a suite has a known flaky test, give that test a temporary, named exception with an owner and a removal condition. A quarantine can protect the main signal, but it should still run somewhere and report failures. Otherwise the exception becomes silent deletion.

The policy should also say whether a pass-on-retry is allowed to unblock a merge. There is no universal answer. A low-risk documentation change and a release branch touching payments may warrant different decisions. Make the rule explicit at the repository or path level, then let the reviewer see the first failure before deciding. A green badge alone is insufficient context.

Distinguish a flaky test from a flaky environment

“Flaky” is an observation: the same intended check produced different outcomes. It is not a root cause. For a test, look for shared mutable state, time-dependent assertions, ordering, random seeds, parallel execution, and asynchronous waits. For infrastructure, look at runner capacity, DNS, artifact storage, credentials, rate limits, and external services.

The investigation should follow the evidence. If a test fails only with parallel workers, reproducing it as a single test may hide the problem. If an artifact fetch failed before tests loaded, rewriting assertions is wasted effort. The companion guide How To Design CI Output That Humans Can Actually Debug explains how to make the failed layer visible without a log excavation.

Capture just enough context to repeat the failure: command, test shard, environment, seed where applicable, and a stable link to logs or artifacts. Do not dump secrets or huge diagnostics into a public log. If the only way to understand a failure is to rerun it repeatedly and watch for luck, the system needs better observability before it needs more retries.

Give recurring failures an owner and a deadline

A retry policy is incomplete without a way to remove the need for retries. Review pass-on-retry counts by test, service, and failure class. Rank them by developer interruption and risk, not only raw count. A rare failure in a release gate may deserve attention before a frequent annoyance in a nonblocking check.

For each recurring case, record:

  1. The named owner who can change the test or environment.
  2. The smallest credible diagnosis and the evidence still missing.
  3. The temporary merge or quarantine rule.
  4. A review date or condition that ends the exception.

That list should live where the team already manages engineering work. It need not become a new dashboard project. Developer Platform Health Signals That Lead to Action makes the broader point: measure only what can change an investment or operating decision. Pass-on-retry volume is useful when it helps fund a fix or retire a bad gate.

A practical policy for the next failed run

Before pressing retry, ask what hypothesis it tests. Is the runner healthy now? Was the external service unavailable? Is this a specific test suspected of nondeterminism? Save the first failure and check that the same commit and relevant inputs will be used. Retry once at the narrowest meaningful level. Then classify the result: environment recovered, test passed on retry, failure repeated, or evidence inconclusive.

If the second attempt passes, the first failure still happened. Report that fact in the review or CI summary, and open an owned reliability task when the pattern repeats. If it fails again, stop using retries as a substitute for diagnosis. The purpose of the policy is to make CI faster to interpret and more trustworthy, not to make the status page greener.

For more practical engineering systems writing, visit Slaptijack.

Slaptijack's Koding Kraken