A Slow Local Test Is a Product Decision, Not Just a CI Problem

Published · Programming

A team can have a green CI pipeline and still make every developer wait six minutes to learn whether a small edit worked. The build passes eventually. The pull request merges. Yet the expensive part of the day happens before either event: a person changes code, waits, loses context, checks a message, and has to reconstruct what they were testing.

It is tempting to call this a CI problem because tests are involved. Often CI is behaving exactly as designed. The local loop has a different product requirement: give someone useful feedback while the change is still in their head. If that loop is slow, the engineering organization has made an investment decision, whether or not anyone remembers making it.

A useful response starts by treating local feedback as an internal product. Identify its users and their most common jobs, measure where their time goes, choose a smaller useful result, and give someone responsibility for keeping it reliable. "Make tests faster" is a wish. "Give service owners a trustworthy two-minute changed-package check" is an outcome a team can fund and verify.

Measure the iteration, not only the command

The number that matters is not simply the elapsed time of pytest, go test, or a build target. It is the delay between a developer making a plausible change and receiving feedback that can change the next edit. That delay includes setup, dependency resolution, compilation, test selection, execution, failure explanation, and occasional reruns.

Start with a few representative workflows rather than an impressive repository-wide benchmark:

Workflow Question the developer needs answered Useful local feedback
Edit a small pure function Did I break its behavior? Focused unit tests and a clear failure
Change a service contract Will callers still compile and agree on the schema? Affected packages and contract checks
Modify configuration Is the generated result valid? Fast validation against realistic inputs
Refactor across modules What else depends on this interface? Dependency-aware build and broader tests

Record cold and warm runs. A command that looks fast after a warm cache may be miserable on a fresh checkout. A fast pass with an unreadable failure may require several extra iterations. Include how frequently the workflow occurs and how many people encounter it. Ten seconds on a daily loop used by thirty engineers may matter more than a minute on a rare release test.

A small measurement sheet is enough: task, machine class, clean or warm state, time to first actionable result, failure clarity, and observed frequency. Do not turn this into surveillance of individual developers. The unit of analysis is the workflow, and the purpose is to find where the system wastes attention.

Separate the fast question from the complete answer

A full suite often answers a valid question: can we ship this change with confidence? It is usually the wrong first question after every edit. The local loop should offer layers. The fastest layer catches likely mistakes in the part being changed. A broader layer runs before review or merge. CI remains the independent confirmation on a clean environment.

The boundary must be explicit. If the fast command silently skips tests that users believe it covers, its speed is counterfeit. Document what each layer proves, how to request more coverage, and which checks remain required in CI. A useful default might be:

edit → focused test or validation → affected-package checks → full CI

This is a contract, not a claim that every repository needs exactly four commands. Some systems need an integration environment to answer any meaningful question. In those cases, invest in quicker provisioning, a realistic local substitute, or a hosted test environment with a short queue. The question remains: what is the earliest honest signal for this edit?

Making Local CI Commands Boring Enough for Humans and AI Agents covers the command interface. A Safe Developer Feedback Loop for AI Agents covers safe use of that interface. The investment decision here is who makes the fast path trustworthy and how the organization knows the work paid off.

Find the actual source of delay

Before buying more CI capacity, profile the local path. The bottleneck might be compilation, test process startup, a container image pull, a remote dependency, fixture creation, database migration, serial execution, or a test that sleeps for a real timeout. Each calls for a different change.

A few useful experiments:

  • Run a representative test twice to distinguish cold setup from repeat execution.
  • Time discovery, build, fixture setup, and test execution separately.
  • Check whether a single integration dependency forces every small unit test into a container.
  • Inspect whether test selection follows actual dependency edges or an overbroad directory boundary.
  • Compare failure output with the question a developer was trying to answer.

A cache can help repeated compilation but cannot repair a flaky assertion. Parallel workers can reduce elapsed time while exhausting a laptop and making every other task slower. Replacing an integration test with a mock can shorten a command while removing the only check that catches a contract break. Optimize the whole feedback loop, including confidence and machine usability.

When Can You Finally Retire a Developer Tool? is a useful companion when the slow path is maintained by an old tool nobody owns. A migration may be worthwhile, but first distinguish the tool's inherent cost from the team's accumulated configuration and unsupported dependencies.

Fund the smallest improvement that changes behavior

The proposal should name a user group, a costly workflow, a measured baseline, an intervention, and a success condition. For example: "Service developers wait four to seven minutes for the first contract failure after a schema change. We will add an affected-package validation command, keep full integration checks in CI, and measure whether the median first actionable result drops below two minutes without missed contract failures."

Those numbers are an example of a proposal, not a benchmark or a claim about this site. The actual target must come from the repository and its users. It might be more valuable to make a thirty-second failure explain itself than to cut ten seconds from a passing test.

Decide who owns the result after launch. The platform team might own the common runner while service teams own their fixtures and test boundaries. A build team might own dependency graph accuracy, with a release team owning the full CI gate. Without this split, a fast path decays into a collection of exceptions. Developers return to "run everything," and the original investment disappears.

Define a small maintenance budget as part of the decision: version updates, cache health, runner compatibility, flaky-test routing, and documentation. Otherwise the initial project is funded while the product it creates is not.

Watch for the tradeoffs that create false wins

A faster local command is useful only if people can trust what it says. Review three risks after the first rollout:

  1. Coverage drift. Affected-test selection misses a real dependency or stops tracking new code paths. Compare local results with CI failures and investigate escapes.
  2. Resource cost. More parallelism speeds a benchmark but makes laptops hot, drains batteries, or slows editors. Measure the developer's whole machine, not just the test process.
  3. Complexity transfer. A complicated command saves run time but costs more human time to choose flags and interpret output. Prefer a simple default and an obvious broader option.

Do not claim success from a single benchmark run. Look for repeated use, shorter time to actionable feedback, understandable failures, and no meaningful increase in defects escaping to CI. Ask developers whether they trust the command enough to use it before they open a pull request. Usage without trust is compliance; trust without usage usually means the command is too hard to find or too awkward to run.

Make the decision visible

A slow local test is rarely the most dramatic problem on a roadmap. That is why it can survive for years. Its cost is spread across small waits and interrupted thoughts, while its fix requires one team to make an explicit capacity choice.

Put the proposal beside other product investments. Show the affected workflows, the best available evidence, the smallest honest improvement, the owner, and the review date. If the team decides not to fund it now, record that decision and the condition that would change it. That is better than calling the delay "just how CI works" while every developer keeps paying for it.

For more on the engineering systems behind these choices, visit Slaptijack.

Slaptijack's Koding Kraken