Skip to content
Podcast

Research briefing with Brian Houck: Measuring AI agents and revisiting the Core 4

AI coding agents are changing how software gets built, but they're also forcing organizations to rethink how engineering effectiveness is measured. Traditional developer experience metrics were designed for humans, not AI agents, so how should engineering leaders adapt? In this webinar, Justin Reock, Deputy CTO at DX, is joined by Brian Houck, Distinguished Scientist at DX and co-author of the SPACE framework, to explore the emerging field of agent experience and how it builds on developer experience rather than replacing it. They discuss how organizations can prepare for AI-assisted software development, how the DX Core 4 applies in the age of AI, why metrics like token usage and PR throughput don't tell the whole story, and the growing importance of documentation. They also examine the impact AI-driven pressure is having on burnout and cognitive overload. Throughout the conversation, they share practical guidance for building engineering organizations where both developers and AI agents can do their best work.

Show notes

Agent experience builds on developer experience

  • Agent experience focuses on creating the conditions for AI agents to succeed. Brian defines agent experience as the environment surrounding AI agents, including the quality of context, documentation, validation, and feedback they receive. As agents become part of software teams, improving those conditions becomes increasingly important.
  • Model quality is only one part of successful AI adoption. Organizations often focus on choosing the best model, but Brian argues that context, clear intent, and effective workflows often have a greater impact on outcomes than incremental improvements in model capability.
  • The same systems that help developers often help AI agents. Investments in documentation, development workflows, and engineering platforms create a stronger foundation for both humans and AI to produce high-quality work.

Developer experience and agent experience don’t always align

  • Many improvements benefit both developers and AI agents. Better documentation, clearer context, and stronger engineering practices improve outcomes across the board, making existing developer experience investments even more valuable.
  • Optimizing for one doesn’t automatically optimize for the other. Brian explains that organizations will increasingly encounter situations where workflows that help AI agents introduce friction for developers, or vice versa.
  • Organizations should measure both independently. Rather than assuming every AI optimization improves the developer experience, engineering leaders should evaluate where the two reinforce each other and where they diverge.

Preparing for AI requires organizational change

  • Successful AI adoption requires more than coding tools. Justin and Brian describe AI readiness as a combination of developer tooling, engineering platforms, and organizational practices rather than a single technology decision.
  • The DX Core 4 still provides a useful foundation. Instead of abandoning existing engineering metrics, organizations should reinterpret them for AI-assisted development while continuing to focus on business outcomes rather than activity.
  • Validation becomes more important as generation becomes easier. As AI produces more code, engineering organizations need stronger review, testing, and verification processes to ensure quality keeps pace with productivity.

Documentation becomes infrastructure for AI agents

  • Documentation is no longer just for people. AI agents rely on high-quality documentation to understand systems, follow conventions, and complete work accurately, making documentation a core engineering asset rather than an afterthought.
  • Not all documentation delivers equal value. Brian highlights that the biggest returns come from documenting information that helps agents understand systems, architecture, and engineering intent rather than simply producing more documentation.
  • Capturing organizational knowledge improves both human and AI performance. Teams that make important context explicit reduce repeated questions, improve onboarding, and enable AI agents to work more effectively.

Engineering metrics need to evolve with AI

  • Token usage is a cost metric, not a productivity metric. Brian cautions against treating token consumption as a measure of engineering effectiveness because it reflects AI usage rather than business value or software quality.
  • PR throughput tells only part of the story. Larger pull requests and faster code generation may indicate increased AI adoption, but they can also increase review complexity and cognitive load if organizations measure throughput in isolation.
  • Outcome metrics matter more than activity metrics. Justin emphasizes measuring whether engineering teams deliver value, improve quality, and create better developer experiences instead of rewarding raw AI utilization.

AI changes how engineering work feels—not just how it’s done

  • AI pressure is contributing to burnout and cognitive overload. Brian describes growing pressure to move faster with AI while simultaneously reviewing larger code changes and maintaining confidence in increasingly AI-generated systems.
  • Software engineering is much more than writing code. Even as AI accelerates code generation, engineers remain responsible for judgment, communication, system design, validation, and building trust in what gets shipped.
  • The long-term challenge is balancing speed with confidence. Organizations that move faster than their ability to verify AI-generated work risk increasing technical debt, developer stress, and uncertainty rather than creating sustainable productivity gains.

Timestamps

(00:00) Intro

(01:26) Justin’s new role at DX

(03:53) What agent experience is and why engineering leaders should care

(08:55) How to improve agent experience at the platform level

(11:25) How agent experience and developer experience influence each other

(14:41) Preparing engineering teams for agentic work

(21:38) Why the DX Core 4 still matters in the age of AI

(27:23) What PR throughput actually measures

(32:47) The limits of token metrics

(37:10) What the data shows about documentation and developer experience

(39:32) Improving documentation for AI agents

(40:43) AI-washing, burnout, and cognitive overload

(45:35) Brian’s upcoming research on agent experience

Listen to this episode on:

Transcript

Justin Reock:

Thanks everybody for joining us today. I am really excited about this conversation. Brian and I have been in this space looking at developer experience and productivity for a number of years. In fact, I was just pulling up an old podcast that we did together almost five years ago now, and it was amazing to me just the difference in the context of the things that we were discussing back then, looking at tying developer joy to developer productivity. The first principles didn’t shift very much. Some of the metrics that we thought were important and certainly the spirit of developer experience. Who would’ve thought that almost five years later, here we are, we’re going to have a conversation where we’re going to delve into agent experience today. Anyway, thanks for joining. We’re going to go through some of the newer research that we’ve seen with DX.

I’ll be kind of playing host, and Brian will be like a panelist. We’ve got some questions that we’re going to go through, but we both have some opinions on this, so we’ll be sharing our perspectives and insights. I’m Justin Reock, Deputy CTO at DX. Nice to see you all. And Brian, why don’t you go ahead and introduce yourself, and then we’ll get things kicked off.

Brian Houck:

Awesome. Thanks, Justin. Yeah, so I’m Brian Houck. I’m an applied scientist on the DX research team. I’m fairly new to the team. I take a very broad view to the study of developer productivity. I do a lot of research exploring everything from how tooling factors, like AI, are impacting developer experience to how collaboration patterns change developer experience, even to wild things, like studying how sunlight impacts developer productivity. A lot of my work centers on measurement frameworks and metric design. And so, I’m just super excited to be here geeking out about developer productivity with you.

Justin Reock:

Oh, we’re going to have a blast. I always learn so much from you, Brian. And yeah, you talk about things like sunlight, so many variables. It’s what makes this such a nuanced problem. Well, let’s dive right into it. We will be watching the QA and chat for questions and we do have somebody helping to moderate those. So as much as possible, I’ll do my best to move to audience questions as they come in. Feel free to interact with us. We like that. It makes us remember that we’re not just screaming into our laptops here. Let’s kick it off talking about agent experience. We spent a lot of time talking about developer experience. That’s what we measure. That’s the acronym for our company. But what, in your mind, Brian, is agent experience as a research area specifically, and why should engineering leaders care about it, not just researchers like us?

Brian Houck:

I think this is such an interesting emerging area. I think of agent experience as a research area that explores whether agents have the optimal environment to do their best work. Do they have clear requirements? Are they working with accurate data? Do they know what good output even looks like? I think it’s the natural extension of developer experience in a world where agents are becoming first-class participants in the engineering system themselves. And so, if we want to maximize the human agent collaboration, developer experience tackles that from the developer side and agent experience tackles it from the agent side. Engineering leaders should care because if agents aren’t set up for success, then humans can’t produce their best work. And I think part of the challenge is many leaders evaluate AI tools on tool capability. Does it write good code?

Agent experience I think helps answer the harder question. Does the human agent collaboration produce good outcomes? And I think one thing that has stuck with me is I recently published some findings from a great researcher, Sarah Chisari in a paper, The Space of AI. Sarah found that the strongest predictor of multi-agent success was not tool capability but specification quality, like how well humans set context. And so, the agent capabilities is what the vendors sell you. The agent experience is what the developers actually get, and those aren’t the same things and that’s what I’ve been looking at there.

Justin Reock:

That’s super interesting because there’s so many facets to that too. I think that when we talk about what’s good for humans is also good for agents. And so, we can certainly, to a degree, use what we’ve considered to be good developer experience in the past as a foundation for building good agent experience. But there is another dimension. There is, like how skilled is the engineer in steering the agent? How well have we provided context? How prepared is the platform to provide that context? What are some things that you’re seeing in terms of efforts to really improve the agent experience? What have we seen so far in terms of the improved outcomes that you mentioned from that?

Brian Houck:

Yeah. Honestly, I think we need a lot more research in this area, but to figure it out, we should borrow from the playbook that worked for improving developer experience for years. It’s define it, measure it, improve it. And so, if we try to break that down is, can we come up with shared definitions, shared principles for what a great agent experience even means? That’s the starting point. And then, how do we measure it? It’s much like a great first step in measuring developer experience is to ask developers. I think the first step of measuring agent experience is to ask agents and trying to figure out how to effectively survey agents about their agent experience is like something I’m spending a lot of time trying to wrap my mind around right now in fact. But ultimately, that is to give us the signal for how to improve it.

And I think that context quality is where the big lever lives right now. Is the context that agents are working from clear? Are the tasks actionable? Is the data accurate? And I think things like stored skills and shared prompts help, data validation is critical, but also clarity of intent. Have you really thought about what you’re actually trying to accomplish? And I do think it’s important as we try to improve agent experience. You can’t do it simply by trying to improve the agent itself. You improve it by investing in the conditions around the agent, just like with developer experience.

Justin Reock:

That’s interesting too, because certainly, the models are continuing to improve, and we’re able to store more parameters, and we’re able to shove more data into these things. The big epic shifts in the capabilities of these models have not really been in improving the underlying LLM, but more like how are we providing context, things like RAG and MCP and all these other architectural adjustments. And so, I think that’s right. When we really see leaps in improvements here, it’s more about how can we efficiently provide this additional context to the agent. What should people be thinking about? I think those who are very familiar with improvement in developer experience are probably clued in into what some of these dimensions of improvement are. Can you give me an example of a few of these things that we should really be focusing on at the platform level when it comes to improving agent experience?

Brian Houck:

I touched on some of these things on, do you have a library of shared prompts and things like that. As I think of some of the dimensions, I’ve already touched on them, like clarity is the context you are giving to agents. Is it unambiguous or is it up for interpretation? Because if it’s up for interpretation, you’re going to get a lot of randomness in the results. Like I said, is the data accurate? And so, you should have at the platform level, how do you do data validation? You mentioned things, like RAG, and I think the efficiency of our human agent collaboration on things. Are we giving it the right amount of context, the right scope or something that I am as guilty of as anyone is like, “Oh, I’m really asking about this chunk of code, but I pass in my entire code base because why not?” Those sorts of things where you can put protections and guardrails in place at the platform level, I think will be super helpful. I’m actually curious, Justin, if you have heard from customers across the industry what others might be doing.

Justin Reock:

These are things like concise and up-to-date documentation, that’s a big one. Thinking about areas of delivery, like fast CI. That’s not so much necessarily something that the agent is going to be “aware of” as much as it is improving the efficiency of agentic work. If we can generate code instantly, but we have a 45-minute build time, well, who cares that we can generate code instantly? The bottleneck is going to be the build system. If we have flaky tests and those flaky tests are difficult for validating the code that we’re sticking in for human engineers, guess what? The same thing is going to extend the agent except now we’re generating 10 times as many PRs with twice the PR size that we’re seeing in some of our more recent data. And so, it’s just going to exacerbate these existing problems.

Brian Houck:

The importance of having good documentation, if only we hadn’t known this for years and years as it relates to humans and it’s just compounded now as it relates to agents.

Justin Reock:

That’s the irony, isn’t it? For me, and I’m sure for you too, we’ve been clawing at this space for years. We’ve been really trying to convince executives to invest in a better developer experience. And now, we’re finally doing it because of all this spending that we’re doing on AI. I’ve come to grips with it. It’s like, “Okay, I don’t really care what the catalyst is as long as we’re finally going to take care of these aspects of the developer experience that are now extending to agents.” I think we’ve already hit on then this relationship between developer experience and agent experience. But do you think, maybe more broadly, this does become, even if it’s accidental, kind of a virtuous thing. Because if we’re going to be improving this for agents and we’re making developer experience better, what are your thoughts on that?

Brian Houck:

Oh, my gosh, absolutely. I think they are deeply intertwined and increasingly inseparable. The agent is now part of the system that the developer works in, operates in. That, by definition, means it’s a critical part of the overall developer experience. But I do think that there are some signals in research that I’ve read recently that shows that that relationship has some complicated implications. And so, a phenomenal researcher, Annie Vella, recently released this really cool longitudinal study that found what they called the productivity experience paradox. What she had found is that developers, as they are increasingly likely to report that AI is making them more productive, they are also increasingly likely to say that their developer experience is getting worse. For those of us, like you and I, who have spent years working with things like frameworks like space and DevX, it’s pretty mind-blowing to say that for some of these AI-assisted workflows, productivity and developer experience are decoupling and they were always so tightly coupled.

And so the ramification there is, as we are optimizing for productivity, maybe we’re getting short-term wins. We’re seeing PR throughput spike and things of that nature. We might be degrading other aspects of the developer experience and that could have ill tidings for longer term ramifications, like retention. It could have negative impact on code quality. It might not set you up for sustainable productivity long-term. And so, I think you can’t separate the two and you definitely, definitely, definitely can’t measure only one of them. And then I think on the flip side is, as you improve agent experience, that’s now a developer experience intervention as well.

Justin Reock:

Now, that’s really interesting, and obviously, without naming any names, but some of the more AI forward companies that we even see in our platform have some of the lowest developer experience index scores. And so, I think maybe you’re right. Sometimes we push past some of these more human dimensions of platform improvement. I think that’s really important. I’m glad that you brought that up because you’re right, we can’t disentangle them. By the way, I’m seeing some questions come in. We are recording those and we will be making some time at the end here to go through each of those questions. Thank you so much for submitting those. Okay.

Maybe finally then something a little bit more actionable. What can companies do to ensure that they are ready specifically for deploying agents. When we shift from this agent experience, which I think is the outcome of providing better readiness, AI readiness for our platforms, beyond the things that we’ve discussed like documentation and better CI, what can we be focused on and how really should we gauge that companies are ready for all of this agentic work?

Brian Houck:

I certainly have a perspective. I’m, honestly, going to be more interested in your perspective. I think you might have a better pulse on this than I do. But I think of AI readiness as sort of hitting on three different layers. And I think a challenge in this conversation is that most organizations only focus on one tooling. But I think, the two deeper layers are culture and infrastructure, and the research is actually fairly unambiguous on what matters most. If I deconstruct that, tooling readiness, it’s things like, do you have the licenses? Do your employees have access? Is it integrated? Are the tools integrated into your existing workflows? Completely necessary, but it’s the easiest part. And I think that, oftentimes, organizations overindex on it. Cultural readiness is, I think, by far the highest leverage layer.

Again, the space of AI study, one of the things we found there is that developers in organizations where leadership actively and strongly advocates for the usage of AI found that daily adoption was 7x what it is for developers who are in organizations that aren’t as strongly advocating for the uses. 7x is this massive change. And I think that cultural readiness, it applies to also things like do organizations have robust commitments to training efforts, things like peer mentoring.

A thing that we published in, again, The Space of AI study, was this notion of social proof. When a developer watches a trusted colleague solve a real world problem with AI, that was the single most consistently cited adoption catalyst. Training programs are incredibly important, but it is a distant second to this like, “Can I see what my peers around me are actually doing this social proof?” And that’s all part of cultural readiness. And then, the last thing is infrastructure readiness, and this comes down to measurement. Before rolling out new AI tools or workflows, do you have a baseline of where your system is today in order to compare against? Do you have feedback loops with your developers who can hear how it’s going? And I think most organizations or these many organizations skip this, and then you can’t tell if anything is working. And I guess the last thing I said there are three layers, but I’m going to have a bonus layer here for the agent side of this, for agents specifically, which is context readiness.

Agents amplify whatever sort of specification quality already exists. And if a team operates without clear specs, agents are going to produce a nightmare of results. And I think those three plus one dimensions are what ladder up to readiness. But I’d love to hear what you’ve seen and heard.

Justin Reock:

I think that you gave a near perfect answer there, honestly. Well done. We can go home, we’re done. This is great. No, I think that apart from the platform readiness, which are these things that we’ve already mentioned, good DevOps practices, things that we’ve learned first principles over the last decade or so. There are certain agent specific things too, like having agent markdown present separate from the human-readable documentation, making sure that we’ve really got our linting and our style checks and things like that. You mentioned setting up these feedback loops. I think that’s really important. One of the most important feedback loops is having somebody shepherding and maintaining the durable prompts that are being sent, the system prompts that are being sent alongside other prompts that engineers are putting in. And it’s really like, I don’t care how you do that, the feedback loop is what’s most important.

So whether you’ve got a ticketing system or a Slack channel or even having opening up some of those durable prompts to source control so that engineers can make suggestions to what the durable prompts should be looking like. It’s more about having somebody gatekeeping that, making sure that there’s an open channel for when the agents are doing something that they’re not supposed to do, that we then put that as part of the system prompts that the agent can correct its behavior going forward. And I love that you mentioned cultural readiness, specifically around the areas like psychological safety, which we know when you look at Google’s Project Aristotle from mid-2010s where they found they had a hypothesis that high-performing teams would be mostly comprised of experienced engineers and really good leaders and unlimited access to compute power, and they were totally wrong.

The number one factor was psychological safety overwhelmingly. And so, being proactive about alleviating fears of displacement, making sure that we’re tying adding … Leaders are thinking that we should be tying these skillsets to employee success because these skills are going to benefit employees for the remainder of their careers. So I think, yeah, I would agree with everything that you said in terms of these three layers, plus the bonus layer, and I think that we can’t ignore the cultural aspect and the psychological safety aspect of this. There was a recent straw poll that I saw executed, where engineers were asked to sum up in three words how they felt about AI in the workplace. And it was very consistent to have engineers putting both excited and terrified in the same breath. It’s like, “Let’s see if we can reduce the terror a little bit and push the excitement.”

Brian Houck:

And in fact, I’ve seen on the call, cognitive psychologists are going to have so much interesting work for a long time around. What is just the psychological ramifications of all of these new tools and experiences?

Justin Reock:

Oh, totally. Yeah, how is it going to change our cognition? Because our brain science hasn’t changed.

Brian Houck:

How is it going to change the way our brains work potentially? Much like social media has rewired our brains in certain ways, is AI going to rewire our brains?

Justin Reock:

We’ve spent so much time in this space warning against context switching and we’re really trying to encourage building platforms that encourage flow state.

Brian Houck:

Now, it’s a feature.

Justin Reock:

Now, you have 10 CLIs open with 10 different agents. There’s a term for it already. AI brain fry is what we’re calling it. And now, I think the orchestration layer is going to become even more important if we look at projects like Gastown, Gas City, which are not ready to be used, but I think you squint a little bit and you get a glimpse of where this is going to go in terms of orchestration. And then, talking about some of those first principles then, let’s talk about the Core 4. Let’s talk about the DX Core 4 framework a little bit. So obviously, it predates by several years this current hype cycle that we’re in right now.

But I would love to hear, you were a co-author of the space framework, you’ve spent so much time researching good measurement frameworks. What still holds in the Core 4 in your opinion today? What should we be reinterpreting? And is there anything as part of the Core 4 that’s just broken now?

Brian Houck:

This is a question I am particularly passionate about and I feel like I’m getting it and increasing the amount now from everyone I know. And I feel, as an industry, we are afflicted with this collective impulse to throw away all of the tried and true developer experience metrics that have been helping us for years. It makes me want to pull my hair out. I am a firm believer that almost everything in the Core 4 holds by design, because it’s a framework designed around measuring outcomes. And so, the Core 4, it has four top-level dimensions, speed, effectiveness, quality, and business impact. And those outcome dimensions describe what engineering teams, what engineering organizations are trying to accomplish. AI doesn’t change what we’re trying to accomplish. It just changes the way we get there. AI is a means to an end. AI is an incredibly powerful tool, set of tools, but we’re not using AI for the sake of using AI. We’re using it to accomplish something.

Frameworks, durable frameworks, are those that have an eye towards the outcome. And I think what really changes is not those top-level dimensions, but what are the diagnostic metrics underneath those dimensions powering them? When I think about what holds, the four dimensions of the Core 4 still hold. I think those are still the right dimensions. They each have a key outcome metric that ladders up to them. These are things like DXI or change failure rate. And I think that protects this intellectual lineage tying DORA, which really is about delivery outcomes to space, which is really just sort of a set of principles on like, hey, measure across lots of dimensions to DevX, which is really around lived experiences. I think that holds, I will say there’s a little bit of an asterisk around PR throughput.

I think that still holds. I think that’s still valuable, but that might be changing a bit, which we can certainly get into. Now, I do think what requires some reinterpretation are diagnostic metrics, those metrics that help explain what we are seeing in our higher level, top level outcome metrics. These are things like PR merge rate, time to 10th PR. The definitions haven’t changed, but the causal relationships around them have changed. And so, an example is let’s say we have a 95% merge rate. Well, does that mean that AI is producing flawless code? Does it mean that humans are just rubber-stamping everything? Does it mean that we’re not pushing agents and finding where the bounds of their capacities, capabilities are? We’re playing it too safe.

It’s the same number, but it can have very different implications on system health now than it might have otherwise before. And so, I guess in closing, I would say my very strong view, and I’m going to release an entire article about this, is that the framework is stable. It’s just the interpretation is where the real work begins and understanding these new emerging patterns.

Justin Reock:

I completely agree too. I think that these foundational metrics, the way that we’ve been kind of positioning it, the new metrics that we can look at, the telemetry metrics coming from the APIs showing us cohorts of users and things, these are useful for segregating cohorts that we can then turn around and cross reference to our foundational productivity metrics. But they’re really only telling us what’s happening with the tech, not like, are we actually shipping more value? Are we actually getting more features out to market faster, which is really what we want to know. That’s where the foundational metrics, I think, then tell us whether or not these investments are actually working. So I know-

Brian Houck:

They’re input metrics. They describe how we get the work done, and then we have engineering system metrics that describe how that work flows through the system and then we have outcome metrics.

Justin Reock:

And I think it’s easy to get lost in the hype and just forget that this is still about developer experience and developer productivity. And it’s interesting too, when we look at some of the data that we source in our Q1 AI impact report and our upcoming Q2 report, which is coming out in just a couple of weeks, you look at these patterns, like we’ve looked at per company impacted quality metrics, and it’s very volatile. We see some people going up in quality, others going down in quality. If we peel back the utilization from that, you see basically the same pattern. You just don’t see the amplitude that we see now. We’re an order of magnitude higher now, whereas we might’ve shifted two points, now we’re shifting 20 points. But the underlying pattern is still representative of the health of the platform, the health of the engineering team, all the foundational metrics that we’ve been able to trust up and host.

Brian Houck:

Totally. All of the core principles of what good development patterns have always meant is just like AI supercharging, and that’s such a cool finding.

Justin Reock:

It really is interesting. When you look at the sway between companies, it should be a little bit alarming, just kind of a reminder that we really do need to measure and that we need to just stay vigilant as we turn up the gas in all these things. I think that we will build some new metrics too. Like we were already talking about agent experience, I think that there are metrics to be found in there, the effectiveness and quality of our prompting and our context engineering. And these, ultimately, will be tied to performance. I think we’ve been thrown around the statement that AI is not going to take your job, but somebody really good at AI will probably take your job, and so it’s good to be good at AI. Some performance metrics there I think will be helpful as well.

Interesting. Along that same thread, we’ve had people reach out and ask questions about whether PR throughput even makes sense anymore when agents are acting as extra humans on the team. How should we be interpreting PR throughput down? Is this one of those metrics that we should be reinterpreting? What are your thoughts there?

Brian Houck:

I will acknowledge, PR throughput has always been a controversial metric. I will say, I am a strong defender of PR throughput historically as a great measure of how quickly and easily can we flow work through the system. I think it’s a garbage measure of individual performance, there’s too much normalization that has to happen. But if we zoom out as a system level metric, I think it has always told us a lot about the friction in our developer experience. Now, that being said, I do think that PR throughput occupies a unique position amongst the key Core 4 metrics because it is directly tied to the mechanics of software delivery, and that’s what AI is changing the most. And as more and more code is being written by AI, which is something that you’ve been looking at a lot lately, it’s going to be an important measure of, can we still flow that code through the system? But it’s not going to tell us very much about developer work.

And I think, particularly in the future, where the atomic unit of work starts shifting away from being PR-centric to being what some people are calling intent-centric or specification-centric, the unit of work changes and so does this metric’s meaning. I think long term, we’re not there yet, and so I think PR throughput is still both valuable as sort of diagnostic, but I think it’s still valuable as an outcome. Are we flowing work through our system successfully? But I think, longer term, this gets moved to a purely diagnostic supporting metric and our top-level speed metric looks at the flow of innovation more broadly. And so, I recently published a paper called Engineering Thrive, and there, I proposed a metric called idea to customer, which is like, can we, instead of measuring how quickly we flow our PRs, it’s how quickly can we flow our ideas?

From our first planning constructs till it is in the hands of customers, how long does that take? And I think pull request is a sort of code first artifact that’s incredibly important as part of that loop, but we’ll tell that broader story. And so, I think PR throughput in summary is less a measure of overall speed and more a measure of system flow. Still useful, just different.

Justin Reock:

So a measure of friction in the system, but we’ve always been very careful to make sure that people are looking at multiple metrics in tension with one another. Goodhart’s laws in full effect, PR throughput is an easy metric to gain if that’s all that we’re looking at. But the idea of the concept in manufacturing would be called concept to cash, idea to value. That’s always been the holy grail of what we really want to understand. We’re putting in more R&D units on this side. Are we actually extracting more value for our customers, and then hopefully, more revenue as a result of that?

Brian Houck:

You actually hit on something that’s super interesting. You hit on lots of things that are interesting, Justin, but one of them was particularly interesting. You talk about gaming? One of the reasons I like PR throughput is I actually think it’s fairly resistant to gaming. Historically, if you really want to game PR throughput, the easiest way is you just break your PRs up into a bunch of smaller chunks. And it turns out that’s the right thing to do anyway. They’re easier to review, they’re easier to test, easier to revert. They flow through the system much quicker. The act of gaming leads to what we want to accomplish. The newsletter you published yesterday shows that AI is doing the exact opposite of that, is doubling PR sizes. I think that’s really scary.

Justin Reock:

It’s all that same [inaudible 00:30:35] psychology in place too. No, that might be one of those rare times when gamification actually forces a good function in terms of breaking into more incremental PRs and more incremental delivery. But yeah, the PR size thing, if you didn’t read the newsletter, we just published it yesterday. We did find longitudinally over a year that PR size has almost doubled, and that should be concerning. When we were engineers, it was part of our subtle art was the ability to pull off the same use case in as little code as possible because every line of code means a potential bug, a potential vulnerability, less portability for the code. But then, there’s psychology on top of this too, which is that we saw this with build times when engineering organizations have long build times. Engineers were more likely than to shove in multiple features into a single PR because they know that they have to wait for that build to execute.

Now, if we have an agent instantly creating the code, but we still have that same delay in the build time, well, of course, we’re going to try to shove more features in because we don’t have to spend as much time actually even writing that code. So multiple angles where this could be hurting us. We’ve got about 15 minutes left in the webinar and I want to make sure we’ve got some really good questions coming in from the live Q&A. I want to ask a question first about token maxing, and then I want to jump into the audience questions. Speaking of metrics, measuring token use is a new metric here trying to understand, I think utilization. What’s your view on the role of tokens with respect to the way that we should be measuring AI impact?

Brian Houck:

Oh, my gosh. It’s just like talk about places where as an industry we have lost our minds. It’s like token usage, in my mind, it is the quintessential diagnostic metric. It tells you about cost. It maybe tells you a little something about engagement intensity. It tells you nothing about whether developers were able to do good work quickly and easily. I don’t even think it tells you anything about how sophisticated your AI usage is. It’s a cost metric. And I think reporting it up to executives without context, it is the modern day equivalent of reporting lines of code. It’s precise, but it’s meaningless.

As a cost metric that is a legitimate use, if you’re spending money on tokens and we are spending increasing amounts of dollars on tokens, yeah, you should track tokens. But if you’re using tokens as evidence of high usage, high sophistication, high better outcomes, that’s when you get token maxing. That logic is backwards because tokens are about consumption. And I think, the right way to use tokens is, again, triangulating across lots of metrics and use it as a bottom layer diagnostic that helps you contextualize the movement you see in higher layers. And so, again, talk about idea to customer, which I just mentioned. If token usage goes up, an idea to customer improves, it gets faster. Great, you have a pretty compelling story. If token usage goes up and outcome metrics stay flat, you have a cost problem. And so, I think tokens, they are a cost metric often pretending to be an outcome metric, track them for what they are.

Justin Reock:

No, when I started seeing those token leaderboards, it made my stomach turn. I’m like, “What are we doing here?” Because you’re right, to compare it to lines of code, I think is very apt on multiple levels, like what we just said before about the art of good engineering, pulling off the same use case with the lowest number of lines of code that we actually have to make as part of that. Same with tokens. We should be able to pull off the same use case through good prompt engineering, good context engineering, good agent experience, using less tokens to achieve the same goal. So yeah, no, I think we’re very much aligned there as we are mostly.

Brian Houck:

One of the things I’m so excited about doing, like researching, is exploring what is that empirical relationship between agent experience and token costs. As you improve the environment around the agents, yes, ideally the primary driver is can we get better outcomes through improving that human agent collaboration, but it should also improve our token efficiency. And so, we’ll see. Let’s quantify it.

Justin Reock:

Hundred percent. And I think you brought this up before, but looking at innovation ratio, which is the way that Core 4 looks at this path.

Brian Houck:

A very stable outcome metric. Are we spending more time on innovation?

Justin Reock:

And that’s what we want to do. If we’re increasing the capacity of any engineer, is that actually translating then to that engineer being able to spend more time working on new features and creating new value? And spoiler, we’re not really seeing that. When we put out our Q2 impact report, we’ve actually seen innovation ratio not fluctuating all that much. And I absolutely talk to engineers who are like, “Oh, I love ClaudeCode. I can load up a spec and go play PlayStation for half an hour and I come back and my work is done.” And it’s like, “Well, okay. That’s maintaining the status quo while doing less work. That’s not really increasing value capabilities.” And so I think it’s going to be really important to look at that metric. I think we’re focusing here on, obviously, we never want to hyperfocus on any single metric, but token utilization, cross reference to innovation ratio, looking at agent experience, and then using some of these proxy metrics to understand friction in the system is-

Brian Houck:

If you ever find yourself reporting on a single metric in isolation, alarm bells should go off.

Justin Reock:

Yes. Absolutely. Okay. I want to move to the audience Q&A. We’ve got some really great questions that have come in here. This first one, I’m not going to read the whole thing. I’ll get to the gist of the question. Concise and up-to-date documentation has been a goal for decades, something that few teams actually achieve. I think that’s fair. Is there any realistic hope that teams finally document things and keep it all up-to-date, and do we have any hard data indicating that that investment and effort in this is actually improving?

Brian Houck:

That’s a large part of what I’m researching now. I think that yes, there is hard evidence that having better documentation has led to better developer experiences. An example is that from some of my previous unpublished research, I found that as developers, development teams that have higher satisfaction with documentation quality, the new hire developers on those teams onboard about twice as fast. And so, you would expect agents to onboard twice as fast. But I haven’t proven that yet. I sure would like to. To your point, it has been one of the top pain points for developers forever. Something like 88% of developers say that they regularly have to spend an hour or more, they waste an hour or more searching for what ends up being out of-date documentation. And it’s just like that’s just like you’re lighting time on fire. It’s so frustrating.

Everyone wants good documentation. No one wants to write good documentation. Now, I do think what constitutes good documentation for humans will be different than what it does for AI. They’ll be able to infer a lot more things. In fact, it should be a lot more succinct. And I don’t know that anyone has solved it yet, but it’s something I’m looking at. If anyone has any great research papers that they’ve seen or anecdotes, please, put them in the chat because I’d love to dive in.

Justin Reock:

It’s kind of the inverse, but one thing that we’ll be publishing in the Q2, I just looked at this data today in the Q2 impact report is we looked at individual DXI drivers and how they’ve shifted. Documentation has improved the most out of any of the other DXI drivers, so that’s a qualitative indicator. But that’s an interesting signal saying that maybe there is hope, because we can do more automated documentation. I think that’s very good advice. We should be splitting agent memory. So agent markdown, as well as human-readable documentation, we should bifurcate that, which we can do easily. The agent should always be updating documentation when it does work anyway. We should just be making sure it’s part of a durable prompt or whatever that we’re instructing it to update agent markdown separately, treating that more like agent memory, which to your point, may look and be formatted differently than what’s easy for humans to consume.

Okay, moving on. My company has just spun up a company-wide initiative to add documentation to code basis to enable agents. Time has been carved out from product work to make that happen. Given we now have the time, what would you say are the two highest value areas to incorporate into the AI harnesses?

Brian Houck:

I don’t know. I’d have to think about that. Justin, does anything hop out to you?

Justin Reock:

I think only what I just said, that making sure that there’s that workflow where we’re updating agent markdowns separately from human-readable markdown. And I think, thinking about the substrate for agent memory is going to be a big discussion this year too. Right now, it’s just files sitting in a repo, but I think we can do better than that.

Brian Houck:

We’re just the bloat of MD files going everywhere. Thinking about the validation layer, I think the verification layer, how do you know that it worked I think is something I would focus a lot on. If I invested a huge initiative in improving context quality, documentation quality, it’s having a detailed plan on how do we ensure that works, and what can you build into the platform itself so that you can detect regressions and things like that could potentially be interesting.

Justin Reock:

That’s a really good point. Anything we can do for better validation loops as well. Yeah, no, that’s good advice. Let’s see. We got about five minutes left here, several questions. We probably won’t get through them all, but we will follow up with some of our answers to the questions that we didn’t get to as part of the follow-up with this webinar. This is a spicy one, psychological safety. Can we comment on the impact of AI washing of layoffs?

Brian Houck:

I don’t know that I’m actually familiar with the term AI washing.

Justin Reock:

Yeah. Effectively, this is a statement that means we’re laying people off, but not really because of AI, but it’s because of AI.

Brian Houck:

Yeah. Okay. No, which I agree. I think that I’ve talked about this in numerous forums at length. AI right now is good at writing code, but writing code is a fairly small part of what a developer does. It’s 14% of their day. Software engineering is about so much more than coding, and agents aren’t nearly as good at those other parts of the development responsibility set. And so, I definitely agree that a lot of what we see is AI washing. It’s like we have different monetary conditions, market conditions, all of these things that are complicated from an economic standpoint, but I do think it contributes to something you were talking about earlier, like AI brain fry. I think that I worry that there’s a growing burnout epidemic that we’re going to have to reckon with, where we feel more and more pressure to just sprint, sprint, sprint, sprint, sprint, try things with AI, and we are getting ahead of our ability to verify and validate.

PRs are twice as big, but can we cognitively handle that? Can our test systems handle that? That’s all another one of these knock-on effects from we have all this pressure to go, go, go, go, go. And I think, it is going to change how we think of what our roles are, it’s going to change our sense of identity, and I think it’s definitely going to impact our wellbeing. Maybe that’s some of that pressure from Annie’s paper talking about productivity and developer experience seem to be decoupling a little bit and we might be seeing quantifying some of that over time.

Justin Reock:

I think organizations need to be very careful. We have plenty of compelling data now that shows us that this technology is not and may never be ready to replace human engineers for many of the reasons that you just said. But certainly, we can augment capabilities. Certainly, you look at a company like Zapier, they’re hiring more than they ever have in the history of their whole company right now, because they know that they get more value out of any investment in one single engineer. I think that there’s a balance. If it’s a matter of trying to populate your company with people who are really good at AI, not necessarily replacing them with agents but treating this as now a very valuable skillset, I think that that’s reasonable. I think anytime that we’ve had a new bit of technology and a new thing to learn, certainly, people who spend time learning that are going to have an edge. But I think that if trying to go down this path of replacing humans with agents, I don’t know that we’ll ever get there. I don’t think it’s the right attitude at all.

Brian Houck:

I’ve waxed and waned even on how far does the line between roles blur. And it’s just like, at one point, November came around, Claude 45 came out. I’m like, “Oh, everyone’s going to be a developer.” And I quickly backtracked on that, because of some of the questions I’ve been seeing in chat around verification. You need someone who can recognize the failure patterns. You also need someone who has good tastes on what should we be building and what is good enough. But also, how do you understand the complex systems? How do you operate them and deploy them and manage them and maintain them? That requires a lot of very specialized expertise, and I think there will always be a need for as many developers as we already have, if not significantly more.

Justin Reock:

I think we will totally induce the demand for more engineers. We’ve been here before. When COBOL was released, everybody thought we wouldn’t need software developers anymore, because we were going to write human language for code and business. That didn’t happen. We ended up needing way more engineers. It’s law of induced demand. We have a four-lane highway, too much traffic. We build an eight lane highway. What do we get? More traffic. I think there’s an infinite amount a software to be written.

Brian Houck:

I do want to acknowledge, just because I believe that that is the truth doesn’t mean that non-engineering executives always interpret the data that way. That’s why I think we see all of this volatility in the job market. And I think that it is both misguided and certainly scary to be a part of.

Justin Reock:

I’m really glad you said that because just because we know that this tech may not ever be ready to replace humans doesn’t mean that a CEO, a misguided CEO somewhere looking at the market research or whatever and the promises of 10x engineering could misinterpret that, and that I agree is the real-

Brian Houck:

That’s why people like us are trying to help contextualize the data.

Justin Reock:

Yes, yes. We’re doing our best. We are at time. We have so many great questions. I feel like we definitely need a follow-up session here because we could get through so much more content together. Can you tell people before we close out, what’s on your horizon? What are you working on right now? What can people expect from you in the next few months?

Brian Houck:

I have a variety of papers coming out on a wide range of topics, but the big area of focus for me is agent experience. How do we define it? What are the dimensions that represent a good healthy agent experience? What is the space equivalent for agents? And then, how can we measure it? How can we improve it? And so, a lot of work on that as I look at how can we optimize the human agent collaboration?

Justin Reock:

Brilliant. I can’t wait to see it. Brian, it’s always such a pleasure having these conversations. Thanks, everybody. I hope this was a good use of your time. We’ll be sending out follow-up. We’ll do our best to answer the other questions, and I think we definitely need a repeat session here too. Thanks, everybody, for the time. Take care.

Brian Houck:

Thanks, everyone.