A Developer Productivity Metric Needs a Decision It Can Change

Published · Technology Management / Leadership

A dashboard can tell you that review time rose last month. It cannot tell you whether to fund another platform engineer, change review routing, reduce work in progress, or leave the system alone. Those choices require judgment about what is waiting, why it matters, and what you are willing to trade to improve it.

That is the missing contract in many developer productivity programs. The organization collects a number and assumes someone will eventually turn it into a useful decision. Meanwhile, the chart becomes a recurring slide, its movement becomes an explanation for performance, and nobody can remember what action it was supposed to support.

Before collecting another measure, finish this sentence: “If we learn this, we can decide whether to do that.” The decision does not need to be dramatic. It might be a four-week experiment with review coverage. But it needs an owner, alternatives, and evidence that could change the answer. Otherwise, you have a reporting habit wearing an engineering hat.

Give the measure a specific job

“Improve productivity” is not a decision. “Should we reserve daily review capacity for small service changes?” is. The second question names something a team can try, keep, adjust, or stop. It also gives the data a boundary: small service changes, their waiting time, and the consequences of changing review coverage.

The SPACE research argues that developer productivity cannot be represented by one activity count or one dimension. That is a useful constraint on the brief: choose evidence that informs the decision without pretending it measures an engineer's entire contribution. A workflow signal needs context and a counterweight.

How To Measure AI-Assisted Engineering Productivity Without Turning Developers Into Telemetry covers the broader measurement system. Here the narrower question is how a particular signal earns its place in an actual allocation or policy choice, whether or not AI is involved.

Try three questions before instrumenting anything:

  • What decision is currently blocked by uncertainty?
  • Which observation would make one alternative more plausible than another?
  • Who can act on the result, and when will they do it?

If every possible result leads to “keep collecting data,” the proposal is not ready. Sometimes the missing input is a conversation with engineers, not a more detailed event stream.

Define the clock before arguing about the number

A measure called “review time” can refer to time until the first comment, time until approval, or time until merge. Those clocks answer different questions. A bot comment is not necessarily useful review. A requested change may send the author back to implementation. A deployment freeze can delay merge without saying anything about review capacity.

For a review-coverage decision, define the interval as “ready for human review to first substantive human response.” Exclude drafts, record the ready event consistently, and decide how reopened changes are handled. Keep time zones and working days visible. Calendar elapsed time and working-hour delay are both useful, but they are not interchangeable.

Then record the population. Which repository, risk class, and team are included? Does “small” mean a bounded behavior change, or merely a line-count threshold? Generated files can distort that threshold. A five-line authorization change can require more judgment than a hundred-line documentation patch.

Use a sample count and a distribution rather than a lonely average. A median describes an ordinary case; a high percentile can expose a painful tail, though it is unstable with a small sample. Read several long-wait examples to understand the mechanism. Do not quietly discard difficult cases because they make the chart untidy.

Work an example all the way to a choice

Consider a hypothetical team with forty eligible service changes over four weeks. Its median wait for a substantive response is eighteen working hours. Several changes wait much longer. Engineers report that review requests land on a few specialists while other qualified reviewers never see them. These numbers are illustrative, not a benchmark or a report about a real team.

The team has at least three plausible responses:

  1. Add a short daily rotation that routes ordinary changes to qualified reviewers.
  2. Invest in clearer ownership and review-request metadata before reserving capacity.
  3. Keep the current process because the apparent wait is mostly deliberate sequencing or specialist risk review.

Inspect the sample before choosing. If the slow changes are mostly waiting for scarce domain expertise, a generic rotation may only produce faster acknowledgments. If ownership is clear but requests arrive while reviewers are deeply booked, protected capacity may be the right experiment. The same headline metric can support different actions depending on the cases beneath it.

Suppose the evidence favors the rotation. Give it a bounded budget: one reviewer protects a short daily window, with complex changes routed to the appropriate specialist. Keep existing approval requirements. Review the experiment after another four weeks, with an explicit option to stop.

The team chooses a provisional target of a median under eight working hours for the same class of work. It also checks review rework, escaped defects, and the effect on reviewers' planned work. The target comes from this team's decision about useful feedback, not an industry productivity quota.

Write a decision brief people can challenge

A short brief is more useful than a dashboard specification that never mentions a choice. The following version can fit in an issue or project proposal:

Decision: Keep, revise, or stop protected review coverage.
Population: Ready, bounded service changes in the named repository.
Baseline: Four weeks; 40 changes; median wait 18 working hours.
Intervention: Four-week daily coverage rotation with risk-based routing.
Owner: Service engineering manager, with reviewer and author input.
Primary signal: Wait to first substantive human response.
Guardrails: Review rework, escaped defects, reviewer workload.
Budget: Named daily window; no reduction in approval requirements.
Review date: Four weeks after launch; calendar date set before starting.
Keep condition: Useful wait reduction with acceptable guardrail results.
Revise condition: Gains confined to a subset or burden shifted elsewhere.
Stop condition: No useful gain, weaker review, or unsustainable workload.
Limitations: Small sample; work mix and staffing may change.

Replace every illustrative value before using it. Add links to the metric definition and sampled cases. Record what would count as unacceptable rework or workload before the experiment; a vague guardrail can be reinterpreted to defend any result.

The owner must have authority over the intervention. An analyst can maintain the query, but cannot necessarily change staffing or review policy. If the decision needs a director's capacity allocation, name that dependency rather than giving the team a target it cannot influence.

Use delivery outcomes as a cross-check

A local improvement can move delay to a different queue. Faster first responses may be followed by slower revisions, more release rework, or less time for difficult design work. Check the downstream result before calling the intervention successful.

DORA's software delivery performance guidance treats throughput and instability together. Use that principle to check whether the workflow improvement survives contact with delivery. A review-wait signal is not itself a DORA metric, and neither is a verdict on an individual's productivity.

Keep the cross-check proportional. A four-week review experiment does not need a company-wide measurement platform. A few sampled changes, release observations, and conversations may be enough to reveal a shifted bottleneck. Rare defects may require a longer observation window; zero incidents in a tiny sample proves little.

A Slow Local Test Is a Product Decision, Not Just a CI Problem applies a similar investment discipline to feedback speed. The transferable habit is to budget for the capability and its maintenance, not merely celebrate a faster command or a greener chart.

Let the evidence change the plan

At the review date, show the sample size, definition, work mix, and limitations beside the result. If the team changed repositories, lost a specialist, or shifted from maintenance to a major migration, explain that before attributing the movement to the intervention. A before-and-after comparison can support a decision without establishing causality.

Then make the choice. Keep the rotation if it removes useful waiting at an acceptable cost. Narrow it if ordinary changes benefit but specialist work needs a different path. Stop it if quicker responses are mostly acknowledgments and reviewers now lose the focus time needed for their own commitments.

Stopping is a valid outcome. It means the experiment answered a question instead of becoming permanent ceremony. Preserve the decision and useful limitations, then retire collection that no longer supports an active choice. Retain only the evidence needed for the next review under the team's normal data practices.

A developer productivity metric earns its cost when it changes what the organization funds, supports, or stops doing. Start with that choice, make the observation honest, and put a date on the decision. The chart should help engineering judgment do its job.

More at Slaptijack.

Slaptijack's Koding Kraken