Skip to content
Podcast

How Meta reduced diff authoring time by 40%

Meta researcher Moritz Beller joins Brian Houck to discuss how Meta measures developer productivity. Moritz explains diff authoring time and how Meta uses it to evaluate tools and guide engineering decisions. They explore how AI is changing engineers’ work and revealing gaps in traditional metrics. They also revisit Moritz’s “Mind the Gap” study and discuss the promise and risks of AI agents for testing.

Show notes

Diff authoring time enables more precise productivity measurement

  • Diff authoring time measures the active work involved in creating, testing, and reviewing a specific code change. This gives Meta more precision than metrics averaged across an entire day.
  • No single diff metric provides a complete picture of productivity. Meta considers authoring time alongside throughput, diff size, and quality to make changes in the data easier to interpret.
  • Meta does not use diff authoring time to evaluate individual developers. It is used at the team, organizational, and company levels to identify trends, regressions, and opportunities for improvement.

A/B experiments help Meta prioritize engineering investments

  • Diff-level measurement allows Meta to quantify whether changes to developer tools and frameworks save engineers time. This helps leaders decide which improvements deserve further investment.
  • Automatic memoization in the React compiler reduced diff authoring time by roughly 30% compared with implementing caching manually. Results of that magnitude suggest that foundational improvements can be more valuable than small interface optimizations.

AI is changing how engineers spend their time

  • Meta’s year-over-year diff authoring time has fallen by more than 40%, while developers are producing more and larger diffs. Moritz views the company-wide pattern as a strong signal of AI’s impact.
  • It is not yet clear where all the saved authoring time goes. Moritz believes engineers are spending more time gathering context and performing work adjacent to coding, while relatively little time goes into writing prompts.
  • As code becomes cheaper to generate, intent becomes more valuable. The important question is increasingly whether the finished code accurately reflects what the developer intended to build.

Traditional telemetry misses important engineering work

  • Activities like whiteboarding, brainstorming, and architectural alignment are difficult to connect to code changes through conventional telemetry. Meeting transcripts and other AI-generated records could help make that work more visible.
  • Cheaper code generation does not eliminate the need for architecture. Skipping a design document can simply transfer the cost to reviewers, who must reconstruct the intended design from the implementation.

Developer productivity is both measured and perceived

  • Moritz’s “Mind the Gap” study connected automatically measured activity with developers’ perceptions of their own productivity. Time spent coding was an important predictor, but sleep, interruptions, and on-call responsibilities also mattered.
  • A modern version of the study would need to account for agentic work. Moritz would examine how developers interact with agents, how much they trust them, and whether managing parallel work creates cognitive overload.
  • Higher output does not necessarily improve developer experience. Brian points to research showing that output can rise while flow, cognitive load, and overall experience remain flat or worsen.

AI makes testing easier but not necessarily safer

  • Agents make it inexpensive to generate unit and end-to-end tests. This could fill gaps left by developers who previously did little testing.
  • Passing tests can create false confidence when both the implementation and tests reflect the same misunderstanding. The agent may satisfy its interpretation of an underspecified request without delivering what the developer actually intended.
  • Specification-driven development could become a new form of test-driven development. Defining correctness before implementation may matter more as agents take on more of the coding.

Collaboration remains one of the hardest parts of productivity to measure

  • Interpersonal dynamics, alignment, and knowledge sharing are central to engineering productivity but poorly captured by existing metrics. Faster implementation also makes it easier for teams to unknowingly duplicate one another’s work.
  • Reducing low-quality meetings can produce major throughput gains, but eliminating collaboration creates different problems. Teams still need enough interaction to generate ideas, share context, and stay aligned.
  • Small human interactions can have measurable value. Brian’s research found that informal conversation before and after meetings predicted self-reported productivity better than internet quality.

Timestamps

(00:00) Intro

(02:10) Moritz’s role at Meta

(04:01) Measuring diff authoring time

(08:20) Measuring A/B experiments

(12:11) What diff authoring time reveals

(14:45) Planning and leadership reporting

(18:36) Why Meta measures teams, not individuals

(22:33) Where developer time goes with AI

(26:06) AI’s impact on junior and senior developers

(27:18) Capturing invisible work with AI

(29:32) The “Mind the Gap” study

(33:40) Revisiting “Mind the Gap” in 2026

(38:32) AI agents for software testing

(41:49) What remains hard to measure

Listen to this episode on:

Transcript

Brian Houck (00:01.774)

Hello, welcome to the Engineering Enablement Podcast. I’m your host, Brian Houk. Today’s guest is Moritz Beller, a software engineering researcher at Meta whose work sits at the intersection of empirical software engineering and real-world developer productivity. Over the years, Moritz has helped advance our understanding of software testing, developer workflows, continuous integration, and more recently, not surprisingly, AI-assisted software engineering.

Today we’re going to dig into some of the work he’s been leading at Meta, including new ways of measuring engineering productivity and what those findings tell us about the future of software development. Now, on a more personal note, I will say that Maritz did a internship with Microsoft Research many, many years ago. And I first met him as my team helped on one of his studies.Which was foundational in my transition into a more formal software engineering researcher. And so Moritz has played a huge role in shaping my career. And so I am incredibly excited to have him on the podcast this week. Moritz, welcome.

Moritz (01:12.671)

Thank you, Brian. Well it’s it’s an honor to talk to you.

Brian Houck (01:16.91)

So just to you know help set set some background context, could you tell us a little bit about your your role at Meta?

Moritz (01:24.629)

Yeah. so like you were saying, I’m basically doing sort of or wearing three different hats in in a sense. I have this academic background on doing software engineering research. So naturally a lot of the problems that we then tackle at Meta also have a research component to them. So it’s not like just taking a requirement, right, and seeing it end to end through to production, but actually

usually, you know, kind of ironing out what it actually is that that we are interested in and then designing some studies to further identify that. That typically then yields some empirical results. we do that if we if we think right we we can do it ourselves then then we do that. But if we think that’s a more you know rigor is needed, we maybe pull in folks like you to actually get us good statistics.

And once we have that, then with a with a team of actual engineers, right, we then we then decide based on the data, okay, what should be the best approach to implementing this or just you know kind of surfacing it as a dashboard, be it for leadership or for direct personal consumption. so it’s really like a basically a research head, kind of software engineer head, and then kind of tech lead and kind of orchestrating between the different roles.

Brian Houck (02:49.858)

I I love it. We we’ve sort of overloaded the term applied scientist. typically now it means people actually like building ML models. But like in my mind, like that’s what real applied science is, right? Like you’re doing actual research and then you’re you’re doing it not just for the sake of advancing sort of knowledge, which is a a a a valid cause, but how do you put it in into practice? And so speaking of that, some of the the work that I’ve I’ve seen from from you

you know, over the last couple of years, talks about a metric that I think is incredibly interesting called diff authoring time. And so just like to help us start with the basics, what exactly is diff authoring time?

Moritz (03:37.001)

Yeah, it’s we should maybe prefix this by asking the question, what is a diff for those that are not super familiar with terminology. So it’s essentially the equivalent of a pull request, right? So it’s a self contained code change together with a summary of the changes that you’re making or the agent is making in today’s world, with a test plan and then reviewers have the possibility to leave comments on this. and you ha may have multiple rounds, right? Multiple versions of that.

Brian Houck (03:43.63)

Please, yes.

Moritz (04:06.976)

Code change. And so basically, I would say until quite recently, the currency at Meta was code. So diffs played an extremely central element at the company, not just for kind of performance reasons, but simply because this is how we were talking about things and this is what moved the company forward from an engineering perspective. And so everyone talked about diffs all all the time, basically.

And it is then natural that the business wants to know, hey, how much time are we actually spending on that very central thing that we’re talking about all the time, which is diffs? And thus diff authoring time was born, which basically measures the active time that humans spend engineering, reviewing, maybe even preparing, testing diff.

Brian Houck (05:00.632)

So like that’s that’s super interesting to me. A lot of my work has looked into how are developers spending their time, which of those activities are sort of productive, what aren’t, what are the, you know, how do they relate to each other? And you know, I I definitely have looked at

Well, like how much time are developers spending sort of like actively writing code? And I know that it is, you know, maybe changing today because of AI, but in general, like that’s the activity developers want to do the most. Now, are you able to tie that sort of authoring time to specific diffs? Are you just like averaging out over all of the diffs that that someone may participate in?

Moritz (05:43.241)

It started out as practically being an average. This was a kind of cruder metric that predated DAT to fill the immediate business need. But then we realized, hey, this leaves a lot of gaps, right? There are a lot of things that we can’t do. One of the more important ones is that we can’t actually run controlled A B experiments much in the same way that we do in product. So of course, any anyone who’s interested in software engineering knows this that

Brian Houck (06:07.982)

Mm-hmm.

Moritz (06:13.408)

Sites like Facebook, Instagram, but also others run lots of A A B experiments all the time to find out, hey, which which of these features gives a better user experience, which of these features optimizes for certain other metrics, right? So performance, could could be other metrics as well. but internally we weren’t quite at that point because we didn’t quite have that metric for our own internal development.

and how we could assess basically the development experience of our own engineers at meta. So these cruder metrics wouldn’t allow us to do that because, say, if we had introduced a certain feature or people were using a certain feature in a diff, but were also not using that feature on say on the same day where or at the same period where we were averaging over, we we wouldn’t get clean signal. And so it was pretty apparent that if we wanted to make

Our internal development more scientific and basically bring it on par with the standards that we had in product, we would need the fidelity that diff authoring type allows us to do. And so that’s exactly what you say. Pretty much on a you know millisecond basis, we can tell sort of what diff you were working on at that point in time. If you were working

Brian Houck (07:37.923)

So that’s like th that’s wild to me, sort of the the level of of precision there. And

But but you said something there that like really piqued my curiosity around like using this for A-B testing. And so I I think probably like the natural reaction to that is, well, like using different development tools. And it’s like, like development tool, you know, if I’m using VS Code versus Emacs, like what does that do to my my diff authority time? if we do different sorts of like build optimizations, all of these these these sorts of things. But I’m curious, like, do you ever run wild A-B experiments like

you know, playing with our environment. Like, well what happens if we put a bunch of plants in? What happens if we put like UV replacement? Red walls versus blue walls versus cream walls. Like d like what sorts of A B tests might you be running?

Moritz (08:28.018)

Yeah, I think it’s it’s closer to your first than the second one because there is a a cost and we don’t have I don’t think we have a team right that optimizes some like wall color in development environments, although it’s a certain it’s a it’s an interesting thought. but so one example would be and actually I think I I ba basically have two really good examples that are controlled A B experiments.

Brian Houck (08:31.79)

Yeah, I I suspect it.

Moritz (08:56.22)

And one of them is using auto memorization in the hack compiler, where basically people had to previously hand roll their own kind of, if you will, caching mechanism in the UI. so if the state of something changed, then you would need to kind of you know recalculate for all of its children whether their cha whether their state also had to change and then possibly re-render this.

that is both error prone and pretty costly to do because also to be redone for every single UI element that you wanted to have caching enabled. and so then the React compiler team set out to basically make this a supported default in the React forget compiler is is the name of that. And so we were able to compare diffs.

that were using this feature of the new React for Get compiler versus diffs that were hand rolling their own memoization slash caching in the past. And so and we didn’t even look at the sort of you know quality benefits that you would get from that but just on the on a diff authoring scale we saw quite massive improvements with like I think around thirty percent fewer

DAT for the diffs that were using you know, the the memoization. Of course it’s it was a big lift from the React team to get that in. So you would have to eventually, you know, trade off those costs against the one time investment. But if you have enough reuse of that feature later on, it’s it’s certainly paying off.

Brian Houck (10:44.366)

I mean I certainly acknowledge that

you know use like running A B experiments using diff authoring time as sort of the outcome variable to look at practical changes in engineering tools and engineering workflows like that makes sense like that is you know I I I think probably a more useful useful use case. But the mad scientists in me wants you to run things like, hey if we take a bunch of people who work remotely, what happens if we drop a dog into their like environment versus a cat? What’s more disruptive to their diff authoring time?

But maybe we’ll we’ll save the mad sciencey experiments for for some future collaboration, you and I.

Moritz (11:22.602)

Yeah, I mean I I know Microsoft Research had some interesting stuff on on that and I think are still doing it, right? I think we are a little bit more conservative in that approach, but but I do agree with you. It it would be interesting and I think there’s certainly factors that are maybe much more influential than what we’re currently measuring.

Brian Houck (11:43.883)

Yeah, we we are products of complex environments. Absolutely. So there are like

Moritz (11:47.089)

Exactly. Yes.

Brian Houck (11:51.797)

Everyone is passionate about their own developer experience metrics, right? Like there’s lots of opinions on what metrics are best for capturing developer productivity. And I’m curious, you know, what do you think was missing from the metrics that Meta was already using that diff authoring time sort of solves that makes it so suitable for running experiments like you you just mentioned?

Moritz (12:18.506)

I think it was so it’s a velocity metric, right? in a sense it it has to be embedded in if you’re interested in kind of drawing a relatively complete picture, it has to be counterbalanced against some measure of quality if you care about that. And of course different businesses would a associate different weight to that and also through basically the throughput.

and if you I think these are like if this is basically the kind of triangle that it works in. so we have another metric called DDM, diffs per developer per month that was basically pioneered by my good friend Kareem Nakat. who you should maybe also invite on this guy. Yes, exactly. And so

Brian Houck (13:02.12)

Our our shared good friend Kareem said.

Moritz (13:08.448)

DDM is interesting because it gives you gives you basically a throughput metric in a given time frame that’s hopefully large enough to you know factor out some minor changes in in in like day-to-day life, if you take a PDO or something like that. and then I think there’s a there’s a quality aspect, but then you could like if you want to stay sort of in the in the diff ecosystem.

Another thing that would be interesting is to look at like how large are those diffs, right? And those three f kind of factors, like throughput, size, and duration, they would naturally balance each other out. Say if you wanted to game one of these, then you would see movements in in in the others. so but I think maybe coming back a bit to your question, what what what were we

missing in that landscape that we had was that we had this relatively crude metric that averaged on a daily basis and we found that it just didn’t provide the fidelity that we needed. And in fact, now that DAT has existed for some three years, we actually can we still benefit from from the investments that we were making because we see that it applies much better.

in a in an AI world than our previously averaged metrics. It’s much more sensitive to movements that we were that we would expect than these averaged out metrics. So there’s actually continued investment in it.

Brian Houck (14:45.688)

So that like that that’s very interesting to me because I suspect when people hear, you know, something like diffs per developer or even diff diff offer offering time, there might be some concern on, well, like

Not every diff is created equal. Like there are different expectations for you know more more senior developers versus more junior developers. So if these things are being used for performance assessment, like it might not be truly capturing what I’m doing. And I I can understand teams like yours, where you have a rich background in rigorous research, like using this responsibly for things like running A B tests. I’m curious like how you actually see this used on, you know.

develop you know frontline development teams in the wild. Like what are the different real world use cases for for this?

Moritz (15:36.287)

Right. So I think we covered I think the scientifically most exciting one already, which is running A B experiments. even if you had an intuition, right, of like, hey, this helped our developers, it really is beneficial to be able to quantify it. and that then allows you know, organizations to plan for it and to prioritize certain changes that they might otherwise not have. So for example, the magnitude of those

DAT savings that we saw in this experiment, but also in the in the second one that I was briefly referring to earlier, were unexpectedly large. And so it seemed to us that we needed to invest more into kind of foundational changes in the in frameworks and in tooling. Rather than maybe micro-optimizing.

U ice a little bit, right? Which is a bit different than what you see in the product landscape, where anecdotally, I think there’s a paper by Airbnb where they did they change the button by a couple of pixels or they changed the color, right? And it yielded a static huge improvement for for them. We we haven’t actually seen this with DAT, right? Maybe it’s a different population. but but the this is interesting. So it’s basically experiments and then they informed how we as an organization

think about planning and which changes are being funded. And then I think even more so it also informs a way of developers for even if they’re not using DAT, which they really don’t have to for every single feature, how to think about designing their products and their changes. So it’s instills a you know kind of different I would I would say state of mind when when you’re developing.

more like scientist. and then this is one class. And then I think the second one is for basically leadership reporting, right? Making sure that basically at the team org and then company level, how like how are our metrics doing? Do we see kind of unexpected crazy regressions in in something? or are we hitting our goals if we if we set certain goals because they’re associated with different initiatives.

Brian Houck (17:59.555)

I mean, thinking about goals as it relates to diff authoring time is really interesting too, because not all diffs, certainly not all sort of class of activities, are created the same. And I’m I’m curious, like, do you try to sort of bucketize diffs by like, is this new innovation versus sort of running the business versus some sort of like, you know, paying down some sort of tech debt? how do you like think about that?

Moritz (18:24.126)

Yeah, exact. Yeah, yeah, exactly. It’s it’s exactly what we’re doing. We’re categorizing diffs a certain way there. And what’s also interesting is you would expect different distribution of those categories per org or team, right? A team that’s an infrastructure, if most of their diffs are internal, that’s totally expected and fine. However, if you’re a product feature team and most of your diffs are internal, something’s

Strange, right? It could be our classifier, it could be like the work you’re doing.

Brian Houck (18:56.98)

It’s like I I think that research is pretty clear, like the literature is pretty clear that developers want to be spending time actually building, right? Like they they want to be spending time building. We certainly as an industry want engineers, and we talk about this as part of the DX Core 4, like

We want to be spending more time delivering innovation. And so, how do we spend more time on delivering innovation? And like diff authoring time, especially if you can categorize the diffs, like I see how it’s it’s a good measure for that. But it’s not just a velocity metric in my mind, right? Like there’s also sort of an efficiency standpoint. So is it really about

Can we drive total diff author and time up or down? Or do you actually look at it as like time per diff? And are we getting more of are we creating more diffs per unit of time that we’re we’re we’re spending?

Moritz (19:48.726)

Yeah, so we’ve never used it for performance reasons, right? Unlike other metrics that were surfaced in the tools that were used to assess performance, diff authoring time never appeared in there and that that was a conscious choice on our part. that said

especially now, right? But also for meta the year of efficiency was twenty twenty two, which is when diff authoring time development started. So that was not a coincidence, obviously. So we are interested in that. even if it’s not maybe on an individual basis or or on a performance base. But but yes. So we basically look at okay how many just are being created in a certain time period, like per month. and what was the what was the

DAT for that, right? And the the the the reason for that pair is basically well you can trivially, you know, 2x your DDM if you half your DAT, but the end product is the same. We wanna avoid that. probably it’s not exactly that relation because there’s some overhead associated with creating more diffs. so actually it’s like the incentive is not to do that for yourself, right? For your own review because

In theory what you’re evaluated on is is the impact you shipped and whether that’s done in ten or twenty discs is in in in isolation would make a difference. Now

Brian Houck (21:15.21)

Interesting. So like that certainly looks at the you know the gameability of it as like what can individuals do to to to game it but i was thinking even like more more high level which is just like if you want to look at system efficiency you know it’s even if if you you sort of normalize the diffs in some way for size and complexity and all these things it’s if I can create 10 diffs

with an hour of diff authoring time versus I can create twenty. And maybe that comes from like my build tools are faster or more predictable. So like I’m not spending as much time waiting around. Our code review processes are better because we’re like automating code review with, you know, radar or or something. It’s like, I I’m curious if if it’s used in that way at all.

Moritz (22:01.961)

Certainly. And I mean the biggest game changer that we’ve seen there obviously was AI, right? And so what we’ve seen there is basically our year over year DAT is down by over forty percent. At the same time, DDM has increased substantially. While and now you could say, well maybe you just made your s your diffs smaller, right? That would explain that. But actually diffs are getting larger as well. So

We’ve never had and this is at the company level, right? Which is obviously the strongest indication that you could see for this. So we were, you know, occasionally observing something like this for individual teams when they had, you know, different initiatives. Sometimes also, to be honest, I think some of those are just noise because at the end of the day it’s a a very large company with so many teams, you’re always gonna have some outliers. So you need to be careful not to draw conclusions from this. But if you get such a clean and strong signal like we’ve seen with

AI adoption, where we can also ascribe certain events in time to certain reactions in the graphs. it is it is very clear.

Brian Houck (23:13.134)

Okay, this is all wild to me. So this is all wild to me. It’s first off, I saw that in fact I recently

published a a newsletter where I talked about a study from Microsoft, from Alex Halman at Microsoft that similarly found that the number of pull requests per hour of active coding had increased about 40% in Microsoft. And it seems like you know we’re we’re seeing very similar orders of magnitude. And I’m I’m curious, you know, if we’re having more and more code, but having to spend less and less time actively authoring it. Do you have some

idea of where does that time go? Like is it going to other things that are still building activities, but we don’t really think of them as diff offering time because it’s not actually writing code, it’s generating specs. Is it going to you know validation time? Like where is that time going?

Moritz (24:08.479)

Yeah, this is this is only an informed opinion that I’m gonna usher now. It’s it’s not actually scientific. We we’re still looking into this, right? I think it’s context gathering. I think you’re also right in pointing out that a lot of the activity that we traditionally considered not coding are now perhaps or should now perhaps be considered as kind of adjacent part two coding. I think it’s also pretty obvious that not a lot of time goes into prompt writing. I mean yes.

some some time goes into it, right? but I think it’s more so like I s I think initially we had a point where like, yeah, I’m gonna just gonna stuff my prompt, right, with like all of the information I can get. But now it’s I think more efficient to just link the information because the agent can then go ahead and retrieve that itself. and so I th I think there’s there’s some of that going on, for sure. but I don’t think we have

quite found where all of that you know time is exactly spent. I think there’s also a narrative that you could say, well maybe we need to spend more time reviewing now. Right. And I’ve seen that. I think GitHub has made some similar findings. on the other hand I I wonder do you think that you know a line of code still has the same value? Do we even need to review it that thoroughly?

Brian Houck (25:36.889)

Well, I mean I I think there’s two questions in there, right? Like does a line of code have the same value? And I mean to the business maybe, but to sort of developer platform teams people still need developer productivity will will no, because the cost of generating that line is so much less. I mean you said at the very

very beginning, you know, historically code was the currency of software engineers, the the currency you cared about. And now really intent is the currency we care about. Like what is the throughput of our our ideas? And I think how do we measure those things are like the going to be increasingly the important thing.

Moritz (26:14.731)

Yes, yes. And is is the idea that we put into the system actually what’s coming out in the in the diff, right? Or is there some drift between the two going?

Brian Houck (26:23.746)

I mean, yes. And so much around, you know, you mentioned so much time now is sort of going into gathering context and like how do we think about the quality of that context, which would include all of the prompts you are, you know, it’s not just the data you’re feeding in, but like are you able to clearly express your ideas and then you know from from harness and model capabilities, does that translate all the way through to finished product?

and there’s lots of places where the human agent collaboration can break down as it as it turns out, I think. So, you know, as we look at like you have this like this rich data set on how much time are developers actually spending to create these atomic units of work. And that maybe these are now the atomic units of work of of agents moving forward, and sort of intent is the atomic unit of work for developers.

Moritz (26:53.163)

Yeah.

Brian Houck (27:15.34)

And you’re seeing how patterns are changing sort of broadly. Do you see differences as you slice into different cohorts? Like for example, are junior devs seeing their diff authoring time drop faster than you know because of using AI than you know very senior devs are.

Moritz (27:35.81)

We have initially observed the opposite of that, much to our surprise as well. where it seemed like Yeah. I think sort of this finding’s been corroborated in in the industry, where it seemed like you know, seniors were able to more effectively use these AI tools. but I think that has evened out some over time now. so if you at least from a metric perspective.

Brian Houck (27:41.324)

That is definitely to my surprise.

Brian Houck (28:05.646)

Like as I think about all of this, it’s you know, I I’m I’m struggling with this idea that traditional telemetry misses a lot of the invisible work that goes into software engineering, sort of whiteboarding and brainstorming and talking with colleagues to sort of do architectural alignment. And I’m curious, like, can we start leveraging

LLMs to sort of like synthesize our our meeting transcripts or our meeting audio to capture some of that like missing piece that is going to become increasingly important and have that sort of feedback into the definition of of diff authoring time.

Moritz (28:51.307)

Yeah, I played around with this idea at the very beginning of when we started at Meta to record our meetings internally. now we’re basically auto-transcribing everything if you want it to be auto-described, right? So then it’s pretty easy to do a similarity match, I think, between this meeting and the diffs that are out. I I think that that would be very interesting to kind of strengthen these links that that are latent.

But do exist right now. I also think it’s you know it’s there’s also an interesting shift. So on the one hand, we say, well, it’s much cheaper to generate the code. Therefore, maybe we don’t need to be as careful in architecting it. But you could have the the you could have the counter argument, actually for a very recent iteration of of DAT, we’re now at version 8. the

person who implemented it didn’t have a design doc. But then what that cost is that I needed to spend a lot more time reviewing this and basically extracting the design from the details of the code where I didn’t need to look at the code. Where I’m thinking, well if you trust the agent to implement your architecture, then it would have been much easier to just get the architecture because that’s the piece maybe we care about. And so maybe just because we

Can skip the architecture skip, we we can skip the architecture step, we we shouldn’t, right? So

Brian Houck (30:22.002)

Yeah. I think I think Doctor Peggy’s story’s recent triple debt framework talking about the rise of not just technical debt but cognitive debt and intent debt would argue pretty strongly that we shouldn’t just skip the architecture phase.

So this this notion of how developers are spending their time and how it’s changing sort of makes me reflect back on where you and I first met, where you were working on a study called Mind the Gap.

which was hugely influential in my own work and multiple of my future studies relied on sort of the concepts of mind your gap very, very heavily. And so could you give us like a two-minute version of what you were trying to study and what you learned from from that project?

Moritz (31:12.427)

Yeah, thanks Ryan. pretty simple. this was about developer productivity way before AI. basically we had these traditional productivity measures such as lines of code per hour worked, number of features, maybe even something like function points on the one hand, and then there was a relatively nascent research trend to look at.

perceived or self-perceived productivity. But those two lived in isolation, basically. And what we wanted to do in that study with Thomas Zimmerman, who was my mentor at Microsoft Research, was to basically bridge the gap right between the two. And so we just wanted to assess both at the same time and basically counterbalance the drawbacks that each of them has with the strengths of the other.

and so yeah, this is like I was super excited about doing that.

Brian Houck (32:17.9)

I mean, I am a huge proponent of mixed methods research. And I think, you know, measuring large-scale telemetry and blending it with sort of developer perception is so valuable. And I think one of the, for me, really important findings of your study is that there is a relationship between those two. Like the developers actually are, you know, when they are assessing their own productivity.

Moritz (32:41.109)

Yes.

Brian Houck (32:47.214)

Like that has a relationship in the actual activities they are doing that you can measure with telemetry.

Moritz (32:52.907)

Th that’s right. The number one thing we saw there, which you you I think mentioned here and there in this episode already, was time spent coding. And we were also low and behold, it was pretty low back then. at Microsoft Rides, it was like, Yeah, you’re a software developer, so like most of your time should be spent developing, but it’s actually not the case. Like the core development time was was not the number one activity in the day. And so it turned out that on

Brian Houck (33:17.994)

much to the dismay of many developers. And like and now lots of studies have sort of come along and reconfirmed that yeah, time submit coding like you know it may only be 13, 14, 15% of the day for many developers. So

Moritz (33:30.655)

That’s right. Yeah. what’s also interesting, we we did find that I think overall this was a large driver, but we also saw things like whether you’re on call or I think Microsoft calls the DRI designated response individual, I believe. it’s it’s been a while. that that made that had an influence, but things such as like how many interruptions you had or

also counted. So some of these were like attributes that people had to enter into the system manually. with the all of the telemetry that we have in place at at meta now we would be able to automate this, which you could say is a more truthy signal in that case.

Brian Houck (34:13.07)

I mean, if I remember correctly, one of the strongest predictors you had w of of someone’s self-reported productivity is whether or not they had a good or a bad night’s sleep before, which has influenced a a future study of mine as well. It’s just like th these sorts of things matter, which is like harkens back to the crazy experiments we should be running, the crazy AP.

Moritz (34:24.457)

Yeah.

Moritz (34:35.081)

Yes. That’s awesome. How how did you measure whether someone had a good night’s sleep or not?

Brian Houck (34:39.756)

So in in Mind the Gap, I wrote I know that it was self-reported in this study that hopefully we will be publishing soon. We used fitness trackers. We used Fitbit First of Force to actually quantify sleep quality. So I’m curious, like the world, you know, Mind the Gap at this point in time has gotta be, I don’t know, six or seven years old, right? And it’s the world has changed so much since then. And so like

Moritz (34:50.388)

Awesome. that’s great.

Brian Houck (35:08.768)

If you were to rerun it today, what activities would you add? What changes do you think you would see in the behaviors you measured back then? Like how would it be different? How would mind the gap be different in 2026?

Moritz (35:22.327)

Okay, I’ll I’ll give you an an answer to a question you didn’t raise for a second, and then I’ll ask actually answer your question. If I were to rerun it then, I think the one thing that we were really missing was number little the lines of code. So we had the number of pull requests, but it wasn’t it was more a binary feature because it was in the Windows devices group where we launched this and people were basically at max like putting out one pull request a day because you know, kernel complicated and stuff.

So I I would have like this I think would have really, you know, put the silver lining on the paper, but we Microsoft had just changed to GitHub and so this feature wasn’t available, sadly. So this is the thing that I

Brian Houck (36:03.406)

I mean, do you think that lines of code would have had a measurable relationship to self reported productivity, or do you think you would have found that like no lines of code doesn’t actually like that has nothing to do?

Moritz (36:11.659)

I I think I’m I’m yeah, I’m pretty certain it would have had one. I don’t know just how strongly you write. And the reason for that is is like the more time you spend coding, the more code you write. And so therefore there definitely relationship. But yeah, how strong that would be.

Brian Houck (36:20.62)

That’s true, yes, of course. interesting. Like in general, I think of lines of code as a pretty garbage measurement. But it’s like you’re you are right. It’s like as you code more, like time spent coding matters, and there’s got to be a relationship between time spent coding and the lines of code you generate.

Moritz (36:39.741)

Yeah, yeah. Now as to the other qu as to your actual question. what would we be doing different today? we would have to incorporate a agentic behavior much more. some some measure, some form of how well are you able to interact or even operate in in today’s world, right? Which is very

Brian Houck (36:44.086)

The actual question.

Moritz (37:09.591)

parallel and where I don’t think there are very clear patterns yet on what really works and what doesn’t. I think I would also like to have something like a sort of more opinionated question on like, hey, what’s your opinion on on agents, right? And so I would be very interested in people that are like on different ends of the spectrum of like, hey, I’m trusting this a lot versus I’m, you know, I’m not trusting that so much. And is there like

Where is the because obviously if you trust everything, you’re gonna commit stuff that doesn’t work and worst case causes problems. And we’ve we’ve seen, you know, these outages or already. I think pretty much every large tech company has had one of one agent cost outage yet. So so I wonder if there’s if there’s you know sort of a burrito frontier somewhere where like hey, I’m like I’m eighty percent trusting this.

So I wonder if if if there’s something like that. I think other than that, I’d be very curious to read your study on the you know, health aspects of job performance.

Brian Houck (38:22.904)

Well, like makes me think of a recent study from Annie Vella that showed that sort of productivity and developer experience may be decoupling. And DX found in our most recent AI impact report that even though the amount of work that we are producing is you know skyrocketing because of using AI and agentec experiences, our sort of developer experience index

is flat at best. And so I am curious i also then if if we were to rerun mind your gap if the relationship between things like diff authoring time coding time and self-reported productivity if that is even weakening now what do you like what what would you expect there

Moritz (39:12.555)

I I mean I only have opinions right at this point. I think I I think it may not weaken. I probably suspect that if you run too much in parallel that that could be weakening. I think the same influences for for kind of, you know, why you might have a bad day are still true, which is on the one hand the all of those health reasons, right? And then the other ones kind of environmental ones.

Brian Houck (39:16.536)

Course.

Moritz (39:41.048)

And I think the third one was still pretty much like I don’t see those change at all, which is like, hey, I have too many meetings, we’re in endless alignment meetings, we’re not like we’re getting pulled in one direction and then you know the following week in the other one. so don’t think that changes at all. But I could see that there’s now like a fourth category, if you will, which is like how satisfied and how cognitively not overloaded am I with my setup of using agents.

Brian Houck (40:10.328)

Mm-hmm.

Moritz (40:10.399)

And I think that’s actually I think that’s a that’s a relative benchmark to your peers, right? You see what what they produce and so I think that’s that’s an interesting aspect, yeah.

Brian Houck (40:16.205)

Yeah.

Brian Houck (40:23.246)

Yeah, we’ve we’ve talked a lot about sort of coding and act active coding and

But that the value of a line of code may be dropping because it’s so much easier to produce. And lots of people across the industry are talking about sort of shifting bottlenecks. And now planning’s a bottleneck and validation’s a bottleneck. And you have another in cr you have many incredible papers, but another one of them is when, how, and why developers do not test in their IDEs. And if I recall correctly, you’re a a a big finding is that like,

Moritz (40:54.367)

Exactly.

Brian Houck (40:58.806)

A lot of developers just don’t do any testing. And I’m curious, like, are you breathing a sigh of relief at this point in time because it’s like, well now AI can fill that gap? Or are you even more worried because it’s like we weren’t great at validation before and now we are gonna understand it even less?

Moritz (41:17.185)

think both are true for different reasons. I think it’s very easy to add tests now. We can have unit tests, you can have in some cases even end to end tests because you have a harness that allows you to screenshot the result and then you know the agent can either operate on that and improve or include it as part of the test plan. So it’s at least it’s e it’s easy, right? If if for nothing else, then we can either include it in our system prompt or we just

if there are no tests, we just you know tell the agent to add a couple tests and people understand that testing is more important now because we don’t quite trust what the agents are doing. on the other hand, I also think that the presence of tests, something that’s green, right, like that might you know lure you into a false sense of security of like, hey, the agent actually did what we are asking, but there really was a drift in what it did versus what we asked it to do. Or maybe

Not even a drift, right? But just under specified on our end, and you can have different interpretations of of what was asked. And I think

Brian Houck (42:23.874)

Yeah, that’s like it it’s so interesting. And now I think I’m I’m more worried than before I I asked the question on this like notions like it’s so cheap to show I mean that’s like that’s that’s why we’re having these conversations. It’s so cheap to generate tests, but if you don’t really care about testing before, then you’re not like really going to know what you should be validating that the agent actually like wrote the right tests and

Moritz (42:31.095)

Sorry.

Moritz (42:49.993)

Yeah, yeah.

Brian Houck (42:50.518)

I I wonder if that feeds into this sort of triple debt notion that I mentioned before.

Moritz (42:55.135)

Yeah, maybe we need a you know, some sort of resurrection of T D D in the age of AI.

Brian Houck (43:02.134)

I like, I’ll be honest, I loved TDD. And like I’m curious if like, you know, what is sort of s like is spec driven development? Is that sort of like conceptually very similar? And it’s as we think about generating our context, how much do we think about, well, what validates the correctness of our our the actions in our goals? and how much how do we specify that ahead of time? Like is that the new version of TDD?

Moritz (43:28.127)

Yeah. Yeah. And and so far, right, what we see often is all well we’ll just throw, you know, a another agent with maybe a different agent harness and a different model at it. But y you know, I like don’t quite think that that’s the answer.

Brian Houck (43:43.147)

I also I don’t think that ultimately what’s going to be holding us back in many cases is sort of the capabilities of the models themselves. So you’ve spent years studying developer behavior through telemetry, through sort of self reported perceptions. As you look ahead, and it’s hard to predict this changing world.

Moritz (43:52.055)

Yep.

Brian Houck (44:09.228)

What is one aspect of software engineering that you still think we lack a good way to measure?

Moritz (44:23.137)

I think it’s really qu basically how you interact with your colleagues, right? How you can argue for certain projects, just like all of these interpersonal dynamics. I I think and and so the assuming that there are still people working, right? apart from from that and

somewhat r related to it though is basically arguing or being able to see through the changes that you make. I think what we’re seeing now basically is that w so many new features get implemented, so much new stuff is happening. It’s it’s very hard, even even if the functionality is there and we’re perfect, which actually doubt is always the case, but for the business to think harness

what is there, right? Because people move on so quickly, so there’s no real sort of consumption of of of what has been implemented, unless it’s very obvious and and and you know, very high priority. So I think for for this latter part and then also feeds into hey, let’s not redo the same thing multiple times because it’s so easy to do that now, right? And everybody’s kind of working in the same environment.

it’s easy for people to redo the same thing again. And I think we need a solution at least at a company level for, you know, flagging this or or avoiding it if we if we want to. Not not maybe, you know and not maybe for the sake of avoiding, but at least making other people aware of like, hey, Brian is also working on productivity metrics. Maybe Brian and Moritz should, you know, talk to each other for the like at least.

Brian Houck (46:18.574)

So I I I love that concept, you know, in the space framework sort of

Communication and collaboration is explicitly part of the very definition of productivity. Like, how well are we actually collaborating with our our our colleagues? And yet it is by far the least measured and the least understood. Like we can I we can see interesting patterns between you know collaboration and other outcomes. And it’s I’ve recently written about a case study where we saw that like

A organization effectively all they did was cancel their low quality meetings and they had a bigger throughput gain from that than they did from adopting AI. And so we know that these things are incredibly important, but we don’t really know how to measure like are they good or not. Like we can see after the fact, like, was this bad if we got rid of it? But like if we could measure that more effectively and then use it to, you know, increase innovation velocity, I think it would be so powerful.

Moritz (47:18.389)

That’s right. I I think this kind of meeting cancelling, I think that’s an that’s an interesting point. A colleague of mine wrote about this during COVID where I think initially we saw a huge productivity boost. because I think basically people’s time got freed up so that they could implement all of these things that were on their backlog. But at some point that backlog was drained, right? And so then they argued that we were in a situation where we were actually, you know,

basically suffering from the lack of meetings or lack of alignment. so I’m I’m not sure if I, you know, buy a hundred percent into this, but I think it’s it’s interesting thought.

Brian Houck (47:56.655)

Yeah. Well, there need there needs to be balance. It’s, you know, I I found that canceling low quality meetings had a bigger impact on like positive impact to throughput than AI did. And so it’s like, let’s cancel all meetings. Now on the flip side, I again sort of going back to COVID days where it’s like, yeah, we found sort of the height of the COVID pandemic days. You know, we found that, you know, having more, more uninterrupted time like led to more throughput. Now on the flip side of that,

Moritz (48:09.547)

Yes.

Brian Houck (48:25.85)

I worked on a study, Tale of Two Cities, where we found that the best predictor of a developer’s productiv self-reported productivity was whether or not they frequently engaged in informal small talk before or after meetings. And so like those small human moments of like, hey, how are you doing? Like, how was your weekend? Like that was more predictive of a developer’s productivity than the quality of their internet connections. And so, like, we we need these human interactions.

But it has to be balanced. And so it’s like how do we understand that I think would be so valuable.

Moritz (48:57.397)

It was a positive correlation. The more they checked in with their colleagues, the higher their productivity was. that’s cool. Fascinating.

Brian Houck (49:03.106)

Yep. Yep. Yes. Yeah. Yeah. It’s like more the more they checked in on just like sort of just like trivial like human moments. and so it’s like these things matter. And so it’s like not all collaboration is bad, bad collaboration is bad. And so how do we understand that more? Like I I think that would be that would be awesome.

So I think like with that, as as we look to the future and all the future things to discover, that is that is a a a phenomenal closing note for us. Moritz, this was a phenomenal conversation. Thank you so much for for joining us.

Moritz (49:38.433)

Thank you, Brian. I’ve had a blast.

Brian Houck (49:41.304)

Cheers.