“Your benchmarks don't apply to us"
Why benchmark trends matter more than you think.
This post was originally published in Engineering Enablement, DX’s newsletter dedicated to sharing research and perspectives on developer productivity. Subscribe to be notified when we publish new issues.
Many engineering leaders are right when they tell me that industry benchmarks don’t apply to them. They’re just wrong about what that means.
Most have already looked at a benchmark, compared it to their own numbers, and concluded the comparison isn’t useful. The mistake isn’t recognizing that their organization is different—it’s expecting a benchmark to answer a question it was never designed to answer.
Organizational context genuinely matters. Smaller engineering organizations consistently outperform larger ones on many metrics. Technology companies spend more time on new features than traditional enterprises. Mobile engineering has sufficiently different workflows that it warrants its own benchmark segment. Even survey response styles differ systematically across regions, making some absolute comparisons misleading.
The mix of factors goes far beyond things we can easily segment. Every engineering organization has its own governance model, release process, architecture, regulatory requirements, engineering culture, and history. Some require five approvals before deployment; others deploy continuously. Some invest heavily in internal platforms;others rely on commercial tooling. These choices can dramatically affect developer metrics, making two organizations within the same industry or size cohort look very different from each other.
Those differences are real, but they don’t make benchmarking useless. They simply mean we’ve been asking benchmarks to answer the wrong question.
Exec summary
- Engineering leaders are often right that benchmark values don’t directly apply to them, but they are wrong to dismiss benchmarks entirely.
- Benchmarks, internal trends, and benchmark trends each answer a different question:
- Benchmark levels → Are we normal?
- Internal trends → Are we improving?
- Benchmark trends → Are we improving faster than everyone else?
- The third question (“Are we improving faster than everyone else?”) is often the most important when evaluating investments like AI tools or process changes.
- Benchmark trends help separate your results from broader industry tailwinds (e.g., AI adoption, economic shifts)
- Benchmark values are still useful for identifying unusual performance and areas worth investigating.
- DX data shows metrics consistently move in the same direction across very different organizations year over year, suggesting organizations are less unique than they may believe.
- Changing metrics, survey instruments, or org structure mid-window makes it challenging to use benchmarks to measure change
- Benchmark trends won’t prove causation, but they meaningfully reduce uncertainty about what would have happened anyway.
- The real value of benchmarks isn’t knowing if you’re average, it’s reasoning more carefully about change.
Benchmarks and trends answer different questions
Benchmarks and trends answer fundamentally different questions.
- A benchmark answers: Is this normal? That question is often more valuable than we give it credit for. Knowing that your review latency or deployment frequency is unusual can help identify where deeper investigation is warranted, even if the benchmark itself doesn’t explain why. It tells you where you sit within a distribution. That’s a question about levels, and levels are influenced by things you may never be able to change like your industry, your size, your regulatory environment, your architecture, and the countless organizational decisions that shape how engineering gets done.Your own historical trend answers a different question: Are we improving? That’s a question about change. Because you’re comparing yourself to yourself, most of the differences that make your organization unique simply cancel out.
Both questions are valuable, but neither is the question engineering leaders usually care about most. The question they really want answered is: *Are we improving faster or slower than everyone else?*That’s a fundamentally different question, and it’s one that neither a benchmark nor an internal trend can answer on its own.
Suppose your deployment frequency improves by 15% over the next year. Is that good? If the rest of the industry improved by only 5%, you’re pulling ahead. If everyone else improved by 30%, you’re falling behind. In other words,improvement alone can’t tell you whether you’re pulling ahead or just keeping pace. Likewise, a benchmark can’t answer it by itself either. Knowing you’re at the 60th percentile today doesn’t tell you whether you’ve been gaining ground or losing it.
To answer the question leaders actually care about, you need both.That’s where benchmarks show their value, even for organizations that genuinely are unique. They don’t require your absolute metrics to be directly comparable to another company’s. They simply show you that the change in your organization can be interpreted alongside the change in everyone else’s.
And that, it turns out, is a far easier condition to satisfy.
Benchmark trends as an observational control group
The reason I find benchmark trends so valuable has very little to do with benchmarking. It has to do with causal inference.
Now suppose that 15% improvement in deployment frequency followed an investment in an internal developer platform. Was the platform responsible? Maybe. But maybe AI coding assistants became dramatically better during the same period. Maybe developer workflows improved across the industry. Maybe a slowing economy reduced feature work and increased engineering capacity everywhere.
A simple before-and-after comparison can’t distinguish between those explanations. That’s where benchmark trends really begin to shine. They don’t tell you what your metrics should have been; they tell you what happened to comparable organizations over the same period. They become an observational control group, a way of estimating the background improvement that would likely have occurred even if you had done nothing.
They’re not perfect. Organizations aren’t randomly assigned to different engineering strategies, and no benchmark population is identical to yours. My point is that they don’t have to be. If your deployment frequency improves 15% while comparable organizations improve 5%, that’s evidence that something beyond the broader industry trend may be happening inside your organization.
This is actually how measurement teams already communicate internally, even if they don’t describe it this way. When one engineering organization of roughly 3,000 developers adopted an AI agent to reduce live-site toil, the result wasn’t reported as the absolute changes in incident mitigation time. It was reported that mitigation time improved 2.4x faster than the company as a whole.
Nobody cared whether that organization’s services looked like the company average. The claim wasn’t about absolute performance. It was about relative improvement.
Your organization is less unusual than you think
At this point, there is still a reasonable objection.
Organizations don’t just differ in their absolute metrics, they also respond differently to new tools, new processes, and new ways of working. If every company has its own architecture, engineering culture, governance model, and technical debt, why should we expect benchmark trends to tell us anything useful at all?
I don’t think there’s a complete answer to that question, but I do think there is strong evidence that benchmark trends are more transferable than we expect.
Part of the reason is psychological. Researchers have studied our tendency to believe that our own situation is more unusual than it actually is. It’s called uniqueness * * bias. Organizations can fall into the same trap. Every company has a list of reasons why their engineering organization is unlike everyone else’s, and many of those reasons are legitimate. Organizational theorists have argued for decades that organizations facing similar environments tend to converge in their structures and practices, despite many local differences.
I’ve seen this in DX’s own benchmark data, and the pattern keeps repeating year after year. In our 2025 benchmarks, change confidence improved across every segment, with the median rising more than 12 points. Cross-team collaboration declined across every segment. In 2026, different metrics told the same kind of story. Documentation improved across every segment and customer focus rose across all of them, while review turnaround and incremental delivery declined across most.
These segments differ by size, by sector, and by geography, and their absolute values differ substantially. Yet year after year they move in the same direction at roughly the same time. If organizational context dominated the way the uniqueness objection assumes, we would expect these trajectories to diverge. Mostly they don’t. Those organizations weren’t identical, but they were responding to many of the same underlying forces.
It should be noted that direction and timing are shared, magnitude isn’t. Segments improved on the same metrics in the same years without improving by the same amounts, and individual organizations within a segment vary more still. Benchmark trends need both properties. Shared direction is what makes the baseline trustworthy. Dispersion around it is where your own signal lives. A control group is useful precisely because it behaves predictably, and the same is true here. Co-movement isn’t evidence that there’s nothing left to detect. It’s what makes detection possible.
This isn’t unique to engineering metrics, either. Economists routinely compare countries with different political systems, cultures, and industrial structures. Healthcare researchers compare hospitals serving very different patient populations. Education researchers compare schools with very different student demographics. None of these comparisons are randomized experiments, and none produce perfect control groups. Yet they’re still valuable because the alternative is to assume that nothing else in the world changed while your intervention took place.
I don’t think benchmark trends eliminate that uncertainty, but I do think they substantially reduce it. They’re not a replacement for understanding your own organization. They’re a way of putting your organization’s improvement into context.
You can’t trend against a moving ruler
Aligning to this way of thinking does change some of the advice we’ve historically given. We’ve encouraged organizations to improve their benchmark position over time, and I still think that’s good advice. Absolute benchmark values provide valuable context. They help answer whether your organization looks unusual relative to similar organizations, identify potential areas of opportunity, and highlight where deeper investigation might be worthwhile.
But once you understand where you stand today, the more interesting question becomes whether you’re improving faster than the background trend.
That distinction matters because organizations don’t improve in isolation. New tools, changing engineering practices, AI adoption, and broader industry shifts all influence engineering metrics over time. Benchmark trends help separate improvements that are happening everywhere from improvements that may reflect something unique about your organization.
This also explains why organizations that don’t trust absolute benchmark values shouldn’t dismiss benchmark data entirely. Even if your architecture, culture, or regulatory environment makes direct comparisons difficult, organizations facing similar external forces can still provide valuable context for understanding how quickly the world around you is changing.
That said, comparing changes only works for the things that hold still. If you reorganized, acquired a company, or changed how you count engineers during the same window, those differences don’t cancel either, and your trend becomes as hard to interpret as your levels were. Changing a metric definition or a survey instrument partway through does the same damage. You can’t trend against a moving ruler, which is why, for trend analysis, consistency in how you measure matters more than precision in what you measure.
None of this makes benchmark cohorts less important. The better your comparison group, the better your estimate of that background trend.
Absolute benchmarks tell you where you are. Internal trends tell you whether you’re improving. Benchmark trends help answer whether you’re improving faster than you would reasonably have expected.
What we still don’t know
I can’t yet prove that benchmark trends consistently produce better decisions than benchmark levels alone. I also don’t know which comparison groups produce the most useful estimates of background improvement, or how similar two organizations need to be before their trends become informative.
Those are empirical questions, and I think they’re worth studying.
My intuition is that benchmark trends won’t eliminate uncertainty. They simply provide another source of evidence. Like any observational control, they’re imperfect. But imperfect evidence can still improve decision making when it’s interpreted appropriately.
Final thoughts
Benchmark trends aren’t a replacement for internal metrics, experiments, or randomized trials. Whenever we can run controlled experiments, we should.
But most engineering organizations don’t make decisions under laboratory conditions. We launch AI coding assistants to everyone at once. We reorganize teams. We change review processes. We invest in developer platforms. We rarely have the luxury of a true control group.
That means engineering leaders spend much of their time making decisions from imperfect evidence.
Benchmark trends make that evidence meaningfully stronger. They don’t eliminate uncertainty. They don’t prove causation. They don’t guarantee that one intervention caused another outcome. They simply provide a better estimate of what might have happened anyway.
To me, that’s the real value of benchmarks. Not because they tell us whether we’re average. But because they help us reason more carefully about change.