Skip to content
Podcast

How Microsoft sees engineering bottlenecks changing with AI

Tim Bozarth is a Corporate Vice President in Microsoft CoreAI, where he leads engineering for next-generation developer experiences and Microsoft's Engineering Thrive initiative. Throughout his career at Microsoft, Google, Netflix, and Box, he has focused on developer productivity, engineering systems, and organizational effectiveness. In this episode of Engineering Enablement, Tim joins host Brian Houck to discuss Engineering Thrive, Microsoft's framework for measuring and improving engineering productivity. They explore why AI makes outcome-based metrics more important than ever, where new bottlenecks are emerging in the software development lifecycle, and why verification and confidence may become more valuable than code generation itself. Tim also shares why the purpose of engineering remains the same despite rising levels of abstraction, the skills that remain durable in an AI-driven world, and why engineering leaders should focus on outcomes rather than activity metrics.

Show notes

Engineering Thrive gives Microsoft a common language for improving productivity

  • Engineering Thrive defines productivity through speed, ease, and quality. Its goal is to make it fast and easy to do great work by identifying friction across the developer experience rather than optimizing isolated activities.
  • Speed, ease, and quality should not be treated as opposing goals. Engineering Thrive applies a “do no harm” principle: an improvement in one dimension should not be celebrated if it makes another materially worse.
  • The framework helps Microsoft identify bottlenecks and invest where they will have the greatest impact. Tim argues that developer time is one of the company’s most valuable resources, making productivity improvements a strategic investment.

AI makes outcome-based metrics more important, not obsolete

  • The industry is returning to activity metrics it rejected years ago. Tim sees renewed attention to lines of code, pull request counts, and similar measures as a search for easy answers to how AI is affecting productivity.
  • More engineering activity does not necessarily mean more value. Engineering Thrive instead measures outcomes such as product quality, end-to-end speed, and the amount of time engineers can devote to innovation.
  • A web of outcome metrics is harder to game than any individual measure. Looking across speed, ease, quality, and innovation time creates a more durable picture of whether an organization is actually improving.

Planning and validation are becoming the new bottlenecks

  • Before AI, most engineering time was spent creating and operating software. Tim estimates that those phases historically accounted for more than 90% of engineering time, with operations alone consuming roughly 70% to 80%.
  • On frontier teams, the bottlenecks have already shifted to planning and validation. Code creation is taking dramatically less time, while deciding what to build and determining whether the result is trustworthy are consuming a greater share of the work.
  • The create phase may continue to shrink as models and agent harnesses improve. Tim expects planning and validation to remain durable constraints over the next several years, even as other parts of the software development lifecycle become increasingly automated.

The code review bottleneck is really a confidence bottleneck

  • Higher PR throughput has increased the amount of change humans must evaluate. Enterprise software still requires a level of trust that teams cannot achieve by simply accepting AI-generated code without review.
  • Code review is only one way to establish confidence. Testing, continuous rollouts, feature flags, canaries, and deployment practices all contribute signals about whether a change is reliable.
  • The next generation of verification should help humans ask higher-level questions. Rather than inspecting every branch or class, engineers should be able to assess the scope, complexity, impact, and fundamental purpose of a change.

More abstraction does not change the purpose of engineering

  • AI will reduce the time engineers spend on low-level implementation details. Tim sees that as another step in the long history of abstractions that allow engineers to spend more time describing system behavior and intended outcomes.
  • Great engineers have never been defined by their ability to write a line of code. Their value comes from systems thinking, understanding objectives, and expressing functional and nonfunctional requirements coherently.
  • Producing more software increases the need for strong engineering judgment. The faster organizations can build, the more important it becomes to ensure that their systems remain coherent, valid, and reliable.

A maker’s mindset and ability to experiment effectively remain durable advantages

  • The maker’s mindset is relentlessly oriented toward producing a valuable final outcome. It combines a clear vision of what should be built with the ability to understand the needs of the customer using it.
  • Engineers will need to continuously experiment and evaluate changes in outcomes. Tim argues that this discipline is no longer limited to model trainers; everyone using AI tools must determine whether a new approach actually improves the result.
  • Using AI is not itself the goal. The goal is to accomplish more, and sometimes the best way to do that will be not to use AI at all.

Engineering leaders should measure idea-to-value, not PR velocity

  • PR velocity says little about the success of an engineering organization. Tim urges leaders to stop evaluating teams through activity measures the industry already recognized as inadequate years ago.
  • Idea-to-value connects engineering work to meaningful outcomes. Leaders should also examine how much time engineers spend innovating compared with keeping the lights on or handling corporate overhead.
  • Product quality and business results should remain the ultimate measures of success. Focusing on outcomes can improve developer happiness, productivity, margins, and the value delivered by the organization.

Timestamps

(00:00) Intro

(02:13) What Engineering Thrive is and the problem it solves

(04:46) Why Engineering Thrive isn’t specific to Microsoft

(09:10) The impact of Engineering Thrive at Microsoft

(14:31) Why AI makes outcome-based productivity metrics more important

(18:22) Where AI is creating new bottlenecks in the SDLC

(24:37) Why more abstraction doesn’t change the purpose of engineering

(27:25) The durable skills of good engineers

(33:03) The changing economics of software development

(36:56) Advice for leaders: measure outcomes, not activity

Listen to this episode on:

Transcript

Brian Houck:

Welcome to the Engineering Enablement Podcast. I’m your host, Brian Houck. Today’s guest is Tim Bozarth, corporate vice president within core AI at Microsoft, where he leads engineering for next generation developer experiences and Microsoft’s Engineering Thrive initiative. Tim has led engineering organizations at Google, Netflix, Box, and now Microsoft, and that gives him a unique perspective on what helps engineering teams move quickly while maintaining quality and creating an environment where developers can do their best work.

Now in full disclosure and on a personal note, Tim was my manager at Microsoft. And so we had the opportunity to collaborate very closely together on Engineering Thrive research, including our recent ACMQ article In Thrive: Make It Fast and Easy to Do Great Work. So I am particularly excited to have him today on the podcast to talk not just about Engineering Thrive as a framework, but also the leadership philosophies behind it.

Tim, welcome to the show.

Tim Bozarth:

Excellent. Thank you so much, Brian. It’s such a pleasure to be here with you.

Brian Houck:

All right, let’s dig in. So Tim, what even is Engineering Thrive and what are the problems you were trying to solve that led to its creation?

Tim Bozarth:

Engineering Thrive is really all about understanding developer productivity and understanding developer experience end to end. And the name sort of describes the intent. The point of this entire focus was to figure out where bottlenecks are on the developer experience, what developer experience outcomes we really care about and connect to what makes successful engineers and successful organizations. And to use consistent frameworks and consistent measures to drive change in these engineering outcomes.

And I keep using the word outcome over and over and over because this is such one of the most absolutely essential essences of really understanding the nature of productivity is it is not about the activities that you take. It’s not about the specific things you do at a point in time. It is really about our ability to drive outcomes because that is ultimately what leads to success. That’s the reason why we are all engineers is to actually build things and to make things happen end to end, not just do one particular part of a process over and over or faster or faster.

And so Engineering Thrive was how to bring into Microsoft a consistent approach to talking about engineering productivity end to end, talking about engineering experience end to end, and really optimizing for the three foundational and most essential aspects of productivity. And the things that I believe strongly are the best ways of looking at productivity, which is speed, ease and quality. Or said in a simple sentence, you make it fast and easy to do great work. And I don’t think there’s an engineer on earth that when they hear this concept of like, I want it to be fast and easy to do great work, it is skeptical of that intent or of that goal.

And again, this is what EngThrive is really all about. We actually want to create and build environments where engineers are thriving because it is fast and easy to do great work. And that really is what gives you the platform and the focus and the time to be able to do great engineering.

Brian Houck:

I love that. I love that. Coming up with sort of a common language in order to identify bottlenecks, identify points of friction, remove those and to allow everyone to do their best work. And so as I think about that, you’ve led engineering organizations at Google, at Netflix, at Box, now at Microsoft. And so I’m curious, how much is EngThrive addressing uniquely Microsoft challenges versus industry challenges?

Tim Bozarth:

So this is a great question. And it is not a framework that is optimized or designed for Microsoft. It’s a framework that’s designed for understanding and measuring engineering experience and productivity durably regardless of what you’re trying to do. I mean, interestingly enough, the framework, as we talked about it in the paper that we wrote and the dives into these details, it is optimized for engineering, but it’s really not just engineering specific.

The concept of making it fast and easy to do great work is durable across the industry no matter what you are. If you can be a doctor and you can think about those same things and you will immediately start to understand and identify what are the sources of friction, what are the things that slow you down? What are the things that consume your time as an individual human? What are the things that slow you down so your speed? And then what are the things that reduce the quality of those outcomes? No matter what job type you’re in, those things instantaneously resonate. And that’s because that idea of making it fast and easy to do great work is in fact the essence of productivity, not just Engine Thrive, but that’s the thing that we anchor towards.

Now, the nuance in the answer to your question is that Microsoft’s bottlenecks are unique to Microsoft in the same way that Google’s bottlenecks were unique to Google and in the same way that Netflix’s bottlenecks were unique to Netflix. And in the same way that everybody’s bottlenecks are unique to their own organization and/or system. Because in the end, productivity optimization and understanding the shape of these things is all systems analysis and systems thinking in the same way that designing great application frameworks and software and platforms is an essence of system thinking. This is the same thing. And so our bottlenecks at Microsoft are unique, interesting, sometimes surprising, just in the same way that they were everywhere else, at Google and Netflix.

And so the powerful thing is that we are able to, by applying these techniques, this framing and this consistency across the business, able to have a common language and then use that common language, have a common language and then common measures. Because again, I think it was EM Goldratt that maybe said this, which was show me how you measure me and I’ll show you how I behave. That’s a non-perfect rendition of the quote, but that really is the truth.

When you show people how they’re being measured, you can see how they will behave. If you say, I’m going to measure you based on the number of lines of code you write, or I’m going to measure you based on the number of bugs you submit. Instantaneously, our minds know that there are negative side effects to that. Cool. I’m just going to put forward slop, or cool, I’m going to stop committing code. There’s all of these obvious side effects.

However, when you start talking about measuring on outcomes, I’m going to measure on are you delivering good, great quality products? Is your experience end to end fast? Do you have the time to actually spend on innovation? Do you have more innovation time? When you create these outcome-based metrics, which are effectively immune to gaming, especially as they form into a web, any individual metrics, how beautiful it is not entirely immune to gaming, but small systems of metrics woven together can be.

This is again, sort of the point and what is actually durable. It’s not Microsoft specific, our implementation, and we’ve done some really great stuff around how we have rolled out EngThrive in the company and how we continue to use it as performance metrics for senior leaders and leaders across the board. Target metrics for our Eng systems teams. So engineering systems groups which are investing in infrastructure use these metrics and the ability to drive these outcome metrics as their success metrics. It’s incredibly powerful because it orients the business around the right goals. It creates the right conversation and the right visibility and ultimately then the right behaviors.

Brian Houck:

I love that. And certainly measuring for the sake of measuring might be interesting to me as a researcher, but does it actually make the lives of developers better? And for my time on the EngThrive team, I certainly remember EngThrive is not just a measurement framework, it’s an entire operating model. And so I am curious, what has changed in the way that Microsoft runs engineering because of having this shared language to align to?

Tim Bozarth:

I mean, if I compared where we are now versus three years ago basically, I’ve been here at Microsoft for about four years now and we really deployed the first version of EngThrive about a year into that. So in about three years, there has been a massive change in a couple of things. One, where and how we have invested in engineering infrastructure. And two, culture. There’s a real culture shift that’s propagated across Microsoft and now there is common language. I can’t say that this is absolutely universal and consistent across the business. Of course, in any large organization there are pockets. The graph is multimodal, but there is really a change in the nature of the conversation.

For starters, developer productivity, the belief that developer time, the belief in the actions that developer time is one of our most valuable commodities as a technical company has truly changed. Because Microsoft has a deep and powerful history in product quality and security and compliance and many of the things that make us so successful in the enterprise in the same way that Netflix was focused on speed and innovation and these other aspects of things and Google had their own focuses as well. Netflix had some of the, or excuse me, Microsoft has these cultural elements, which are really important to how we work.

But one of the things that would often or could be traded was you could trade some aspects of quality for aspects of productivity. And they were treated as a teeter-totter rather than things that would allow you to rise all boats where sort of using this thing like you know this, this is when we talk about these metrics across the three dimensions of speed, ease, and quality, there’s a do no harm principle in there, which is you cannot lift one and celebrate it at the cost of another. We can only lift one and leave the others the same or lift others at the same time.

The mechanisms and the framework by which we run this, we create the visibility and we create the accountability, have actually made a real material shift in that experience in the way that we invest in those systems and how decisions are made. No longer do you roll out changes which really do lift security or lift a quality component at the instantaneous cost of many, many engineers across the board. It’s now done so with an understanding that there have to be mitigations for those costs, that it doesn’t increase overhead and taxes, that there’s central efforts to reduce and automate. And AI has really transformed the way that we do migrations here.

But even prior to that, there was a change in the nature of how we invest. The impacts in the last four years on both culture and the way that we invest in edge systems are really profound. And without the measures, without consistent measures, without the framework that we run this and the program that we use to run this to create visibility, accountability, these things would have never happened. Or the probability that they would have happened would be much, much lower.

Brian Houck:

I love hearing you say that. A few years ago, Dr. Thomas Zimmerman, Dr. Margaret-Anne Story and I did a paper where amongst other things, we asked developers how often do you have to trade off quality for productivity? And the majority of developers say like, “Yeah, I have to make that trade off.” And I was like, “That’s a trick question. In my mind, if you’re having to trade off quality, you’re not being more productive.” And so I love that lens of like, no, these don’t actually have to be trade-offs by definition. We can sort of rise them off.

Tim Bozarth:

That’s exactly right. And it’s funny, our purpose is to break the iron triangle of fast, cheap, and good. And then you would rightfully say like, “Well, there’s no such thing as breaking the iron triangle. There’s always a trade-off.” And you are actually right when you zoom the system out. So the trade-off is done by you’re investing money in infrastructure. You’re investing human time, thus money in infrastructure, but you’re creating leverage in the way, again, it’s all about leverage.

In the end, developer infrastructure, developer systems, and investments in productivity are leverage plays. You invest a small amount of money, you build a fulcrum and a lever, and then you move a mountain with it. And that is what allows you to, for such a large system and the consumers, the customers of those systems, which are also the engineers building the systems, it should be at least, that is how you create what looks like the improvement of speed, ease and quality at the same time, the breaking of the iron triangle. It’s just because you built a lever to move the whole thing, but you have to invest in that lever and you have to invest smart. You have to understand and measure what those end-to-end values are and then use those to identify your bottlenecks, and then put that lever right at the bottleneck and then move the bottleneck.

Tim Bozarth:

And I think that’s actually maybe a lead into some of the stuff that we’ll hopefully talk about later, which is the nature of our bottlenecks are fundamentally changing as AI is changing the nature of the software development lifecycle, and that is incredibly exciting. It is completely durable even as the nature of software is fundamentally changing.

Brian Houck:

So this is actually a great opportunity to sort of segue into the AI discussion. And so EngThrive certainly is not unique for AI or specific for AI, but obviously applying the EngThrive lenses to AI is, I’m sure, something you are asked to do a lot. And you had talked about measuring outcomes, and I feel like AI has caused us all to collectively lose our minds and forget that very central principle.

Tim Bozarth:

I love everything that we need to learn in the engineering industry and leaders, all of these lessons about what is the nature of productivity? How do we think about outcomes? How do we think about developer behaviors and activities? We’re in a collective hallucination across the industry. Somehow now it is 2026. Somehow right now I am hearing people talk about lines of code and PR counts again. It has been 15 years since somebody seriously came to me, like a leader across the organization. And at the time, it’s just like, “I love that you’re thinking about this, but you’re taking the wrong approach. You have to look at it differently.” We in the industry knew this. Somehow we’ve forgotten it all because we’ve shaken up the system and the nature of the behaviors are different and the industry is looking for easy answers.

But here’s the thing. Right now, these ideas in EngThrive and the essence of how we measure productivity is more important than ever. And it has come to the forefront more than ever because a new tool has emerged on the scene, a new way of working that very, very clearly is breaking bottlenecks. It’s changing where our energy goes. And so the conversation of productivity and the question of ROI, the question for how differently are we working? This question is at the front of literally everybody’s mind. Every CEO in the industry is asking this question. I see it. I hear it all the time as I work with our customers, as I advise and I’m connected into the startup community and in the VC community as well. This is the most universal question right now that’s being asked, which is fascinating because a productivity question was not the most universal question being asked a couple of years ago, but suddenly it is.

So these principles of looking at outcomes right now are the way that we can actually understand the impact and the transition of AI and we can use it to tackle our most important bottlenecks. And as I mentioned, the bottlenecks are changing. I think that any engineer or any engineering leader, understands already the AI, the current frontier models and harnesses are as good as us at writing a line of code already at this point. In fact, they’re probably better than us. But I said that carefully, at writing a line of code. I didn’t say at building a system or a distributed system or a complete application. They are still not great at design. They really require human in the loop for the systems thinking aspects of how these are going. And that will probably change and improve.

Brian Houck:

Now could not agree more. And it’s one of these things where I wish that engineering leaders and business leaders had always been asking what are the outcomes we care about and how do we get them? As you pointed out, developer time has always been an expensive, precious resource, but it is interesting that AI is sort of accelerating the need to get to those questions. And a lot of that is because the fundamentals of our workflows are changing so rapidly. We want to understand that. And if EngThrive sort of gives you the lenses to look at bottlenecks and sources of friction, then I’m curious, where do you see new bottlenecks forming because of the ways in which AI are changing our workflows?

Tim Bozarth:

And the answer to this is consistent across the industry. So I’ll use a five point version of the SDLC, plan, create, validate, deploy, operate. The vast majority of engineering time historically pre-AI was spent in the create and operate phases. 90 plus percent was in those two with operate being the largest of those by far, 70 to 80% measured across the industry in general with create being some percent of that and the rest sort of just fall in wherever you can.

Those bottlenecks have changed, especially in our frontier teams. The new bottlenecks are planning and validating. So creating is already dramatically, dramatically decreasing the amount of time that’s there. Again, and I don’t mean the whole of it. It’s not disappeared. It’s not zero. There still is the need for great engineering principles and understanding to how to guide the code creation right now, but those systems continue to improve day by day. The systems that we build, how we adapt and modify the harnesses, and then the models themselves continue to improve. If you roll the clock forward a couple of years, you can think of the create phase here largely evaporating, where planning and validating are already the new bottlenecks and they continue to be the new bottlenecks. And in fact, I think those bottlenecks will continue to durably be the bottlenecks for even the next couple of years as sort of the model and the AI ecosystem evolves and improves in its shape and in the feedback loop and in the way that these things work.

So validate is it. I’ll give you just an example of an experience again that any engineer that’s listening has seen and felt, whether you’re at a startup or whether you’re at a large company. PR throughput’s increased. We’re writing more and more commits. The challenge is that human review is still an essential part of what we’re doing. Now you could be YOLO-ing it and just not paying attention to the code at all. And this is sort of the vibe coding approach. That’s great. The challenge of course is that there’s a ceiling on that and that ceiling has been pretty broadly discovered across the industry. If you’re trying to really build enterprise grade software where you are betting and we sort of use this like how much are you willing to bet? Are you willing to bet lunch? Are you willing to bet your paycheck? Are you willing to bet your job? Are you willing to bet your house? It’s just this framework for how much do you really deeply trust the thing?

Most vibe coded software people aren’t willing to bet that much on. Certainly not their houses. But when you’re building software for the enterprise, people that are building their businesses on it where trust is everything. The trust is the essence of what you do. The nature of validation is very different. And so code review is an absolutely an essential part of it. We see the new bottleneck now as we increase, like, the amount of change is the ability to gain confidence in those changes. And code review is one of the essential aspects of that. But actually what I would posit is that we don’t talk about this as the code review bottleneck. We talk about this as the confidence bottleneck because that is the thing we have drawn forward and it’s the thing that we’ve lasered in on.

We use code review as a way of gaining dimensions of confidence, but we also use testing as ways of gaining confidence. We use things like continuous roll-outs and feature flags and canaries and approaches to deployment. We use all of these as mechanisms to gain confidence. And that is where the new bottleneck is. And I think the conversation around this is really fascinating. In fact, this is a conversation I’ve been really driving on in even just the last few weeks, which is how do we refocus our energies on end systems to address this confidence bottleneck? And you could easily call it just like a code review agent. A code review agent is an aspect of that. In GitHub, we’ve invested heavily in the code review agent and continue to invest in this because we know it’s important. And that is one aspect of how we build confidence, but really there are so many other signals that we have to bring in and consolidate.

And I mean, I could say shift left because boy has that been a popular way of describing this in the industry for a long time. But I don’t actually think that’s entirely the purpose. What I think it is, it’s about getting all of those signals into this one place so that when the human does come in, when the engineer comes in with that perspective, that holistic perspective of taste, of system understanding, of being able to look at and make judgment calls on the scope of the change, the complexity of the change, the impact of that change, to really be asking questions like, “Should we be doing this? Is this touching or modifying the system in the way that we need?” Where they’re not asking questions of like, “Is this particular logical branch implemented correctly? Is this nuanced thing here? Is the structure of this class in the most optimal and maintainable way, maintainable for human way, but not necessarily maintainable for agents way?” The nature of the questions where we need human intervention in the confidence building stage are changing.

And I think that is actually really important. This idea of don’t make the next generation of verification look the same as the previous generation of verification. We have to start rethinking it in a way that builds on where agents are the most powerful and how we utilize them to their fullest. That really is the essence of this new bottleneck that’s forming and where I see so much investment in the industry and need across the industry, across all dimensions.

Brian Houck:

So that’s super interesting. In that world, I imagine that you really don’t need humans to be doing sort of like line level code inspection of changes in the PR review process. They’re looking at maybe summarizations of what these changes mean and what are the ramifications, what are all the systems, the components they touch. I’m curious in that world, are you afraid then that humans are going to lose sort of the understanding of how the systems actually work if they haven’t dug in themselves?

Tim Bozarth:

I love this. So this is such a great question. I’ve been asked questions like this over the years almost anytime we build an abstraction. To your question right now, I see a future. I absolutely believe that the future will have us as engineers spending less time on specific logic implementations in specific logical branches and defining the nature of the very low level implementation inside of a system. We as builders, we want this world where we can spend more of our time on describing the behaviors of systems and seeing, experiencing those, having those come out of the other end of the box, but also needing to have confidence that they work predictably, they work consistently and they work reliably.

So your question is a great one. Is it going to disconnect us from much of what we have done, much of where our day-to-day time has gone as engineers? Yes, it will, but it also gives us more time to do the thing that brought us to be engineers to begin with.

The things that differentiate those great engineers, it’s not how good they were at writing a line of code. It’s how good they were at thinking about systems and understanding the objectives of the system, modeling and describing functional and non-functional requirements in coherent, logically coherent ways.

This is the essence of engineering. The greatest engineers can swap over to another language and pick it up fast and easy. Because it’s all about how we express and capture intent. So the nature of how we express and interact with intent is going to evolve. The modalities by which we do it, they’re going to evolve. They’ve already evolved. Already we’re having back and forth agentic conversations. We’re starting to have multimodal interactions with these sources of truth. That is only going to change and it’s only going to evolve more, evolve or transform more as time goes on.

But if there’s one thing I am absolutely confident in, it’s that the need for great systems thinkers, the need for great engineers is more important now than ever because we are already seeing that we can produce more faster, but the need to have coherent, valid, reliable systems increases with each of these new products and each of these new ambitions that we’re creating.

Brian Houck:

I love that because I spend a lot of time thinking about what even is the role of a software engineer in three years, in five years? What do we even call this role? And as we move more towards system thinking is ideation, planning, deciding what it is to build to begin with, I am curious what you think it even means to be a software engineer five years down the road and what are the skills that people should be focused on building right now to set them up for success then?

Tim Bozarth:

I love this. This is such a great question. This is a place where inside of Microsoft, we’ve been spending a lot of time and energy thinking about.
Now, the thing that I can say confidently right now is no one knows. We don’t know. If you think you know, you’re probably wrong. If you think you can right now build a new job profile and job ladder for the builder or the new maker, you have great intentions, but you’re definitely going to get it wrong. The reason why is because we are in a transitional state right now. The snow globe, the little particles of snow are still flying around. They haven’t settled yet.

So with that in mind, I do, however, believe there’s something very durable. There’s probably one or two ideas in here, which I think are really durable. One, the maker’s mindset.

The maker’s mindset is the focus on creating something, on creating a final product and being willing to do whatever it takes to get to that goal. The maker’s mindset of being relentlessly oriented towards an objective and having a vision and views for the nature of that objective. You are thinking about the product that comes out the other end. And that also requires you having a customer thinking mindset. You have to place yourself in the mind of the person that uses the product in order to build a great product. That is the maker’s mindset in its most essential quality. That is probably the single most durable thing for us to be thinking about as we look at the future of engineering and the future of job roles.

Now, I mentioned there’s two. The second one is very adjacent to the first one, which is expertise and experimentation and evaluation. I say experimentation and evaluation. They’re the same thing. We really, really need people who, as they are going, doing things every day and picking up these new tools and using them in new ways are constantly thinking about measuring and evaluating the changes to outcomes. This is not just a thing that model trainers and post-trainers need to be thinking about. This is a thing that everybody has to be thinking about, which is how do you experiment and evaluate the outcome and the change to the goal and the objective that you’re targeting with the work that you’re doing? This is so important because using AI in a new way is not the goal. Using AI is not the goal. Using AI to do more is the goal. And in fact, if you can do more by not using AI, then do more by not using AI. And in fact, I think this is one of the places where… This actually connects back to a bunch of metrics in some of these big conversations that we’ve been having that have been driving across Microsoft and that we’ve really, really internalized inside of core AI, is the power use of AI is not using AI for everything. It’s using AI for the right things.

Brian Houck:

Yeah, absolutely. You had said something earlier that I have certainly found to be true, and that is anyone who thinks that they know the future for sure is assuredly wrong. It’s like every prediction I’m making is either wrong in timeline or actually wrong in substance. And I’m curious, three years from now, if we were to look back on this conversation, where do you think we will realize that we completely misunderstood AI?

Tim Bozarth:

It’ll be in one of two directions. It will be in, one, that we assumed the agentic systems and the ecosystem that’s evolving had limitations that it doesn’t have. Or two, it will be that we assumed it would decrease an entire class of effort to zero and it doesn’t.

Now, I have an optimist’s view, which is the first one. And why that’s an optimistic view is because I love making. I love outcomes. I like actually figuring out how to produce the end state product. The thing I have in my mind that I know has a whole bunch of complexity and a million steps in the middle, I see the ability to need to do less manually in the pursuit of that objective as a great thing because the outcome is the thing that has the value. The toil in the process or the process is a part of the story of it, but ultimately, when it comes to creating value, that thing is in the end state.

So maybe in three years, what we will really see is the evolution of some of the aspects of the SDLC dropping to zero. So again, plan, create, validate, deploy, operate. We already know, we already see the shift in frontier teams that plan and validate become the place where they spend the vast majority of their time. Maybe the industry holistically will get there in three years. Maybe in fact, validate or some aspect of validate will continue to diminish and we really go back to ideas and really focused on planning. We’re truly the bottleneck are our ideas and the ability to generate fully formed thoughts, ideas and the composition of these things. I mean, that would be really, really, really incredible. I hope that my views on where we’ll be in three years are wrong and that we’re closer in that direction.

Brian Houck:

I find for me that my optimism of the future is somewhat moderated by sort of uncertain economic realities. What I mean by that is like token costs. Tokens are expensive. And this is where you had mentioned you were thinking about doing some blog posts about like, “Well, how do we use AI for the things it’s great at and not use AI for the things we don’t need to use AI for?” And that’s where thinking like that will help us break through where some of those economic barriers may end up being.

Tim Bozarth:

So you are absolutely right. Fundamentally, the economics of software are changing. For the first time ever, we’ve introduced marginal cost in the software development and software operation process, which is a fundamental economic shift in the way that software is built. But I’m going to say that actually this is great.

So token cost is real. All of us have been seeing an exponential curve and the pop culture has said AI capabilities and the systems of these AI systems, not just the models, but the systems around them have been on an exponential curve of capability over the last few years. And that’s correct, right? That exponential curve has actually been pretty clearly mappable in terms of capabilities.

Now, the challenge with exponential curves is that sometimes there’s other exponential curves that are attached to them. So that exponential capability curve also comes with an exponential cost curve. And Jevons paradox: The simple version of saying it is like, “If you add another lane to a freeway, traffic doesn’t get worse, just more cars go on the freeway.” The usage of the resources in a system expands with the availability of the resources in that system is sort of the nature of that paradox. So, yes, exponential capability curve, also exponential cost curve.

When I mentioned that idea of what a real AI power user is doing right now and that we see in these teams is not using AI for everything, it’s using it for the right things.

Now, in order to achieve that, you have to really think about AI use differently. You want to basically have agents which then build out reusable components and it can be tools, it could be platform-like layers, dependencies, libraries, whatever the case may be. So the next thing only knows that it has those available to them and what domain it has to operate on so that it can then build out the next thing so that the next layer needs to use less and less inference in order to get the job done.

This is still an AI maximalist view, but it is a AI maximalist view where we are thinking about how do you actually create efficiency and scale and leverage inside of that AI maximalist system. So I really do believe that in order for us to address some of the fundamental economic challenges of AI use, we have to start thinking about effectively how we build, redo what we did in the last 15 years with platform engineering, but in a way that is agent-based and agent-optimized where the layers and interfaces are optimized for agents.

And in fact, the good news is I think we can actually reuse a lot of the abstractions that are already familiar to us as we think about application frameworks and application run times the way that we do this. I think that those abstractions will evolve in the future because, again, those are optimized for the human mind.

Agents have different thresholds of what they’re capable of doing. We can and will discover those. And in the end, that leads to sublinear and ideally logarithmic token consumption for working on complex systems over time. I think that is a goal we have to achieve that in order to make the economics make sense in the long term.

Brian Houck:

Love it. Just like we built platforms to provide great developer experience to allow developers to do their best work. How do we build platforms to give a great agent experience to allow agents to do their best work.

Tim Bozarth:

Exactly. Which then allows humans to focus on what they’re making, on the objective, on the goal, on the creative aspect of defining what we’re trying to do and helping get it done. To me, this is really optimistic. It’s really exciting.

Yes, our job is changing. It’s not scary. It’s just a change. Well, it might be scary, but it’s not fatalistic. It’s just a change in the essence of how we are productive and how we get our job done and how we make.

Brian Houck:

I have one last question for you, Tim. So what advice would you give to a VP of engineering trying to navigate the next 18 months?

Tim Bozarth:

Yes, a piece of advice if I could stop hearing this question, it is remove PR velocity from your mind. Stop talking about PR velocity. Stop thinking about the success of your organization and of your people via activity metrics. Start looking at outcome metrics. Start thinking about idea to value. Refocus all of that energy that you’re wasting looking at these trivial metrics that we already know are meaningless, that we made that conclusion 15 years ago. Look at idea to value. Look at the time that the humans in your team, the engineers in your team have to focus on innovation versus keeping the lights on versus corporate overhead, those three buckets together and focus on product quality outcomes. That is it.

If that’s the one piece of advice I can give, I am absolutely confident that so much more of the industry would move in directions that created excitement and happiness in their engineers and would create better business outcomes, create better business outcomes, whatever those business outcomes may be, product development, margin, profit, and/or developer productivity outcomes when it comes to speed and quality. Those things, that is the single most consistent piece of advice I feel like I’m giving right now across the industry to leaders across the industry who are churning in this change.

Brian Houck:

I love it. Outcomes, not activity. Love it.

Tim Bozarth:

Absolutely. Absolutely.

Brian Houck:

So with that, Tim, thank you so much. Anytime we talk, my brain hurts, but I love every minute of it. This has been phenomenal. Thank you so much for joining us.

Tim Bozarth:

Absolutely. No, and again, it’s such a pleasure to be here. I hope your audience finds this interesting and I look forward to continue pushing on this and sharing and helping share with the rest of the industry the things that we’re tackling inside of Microsoft and broadly.

Brian Houck:

Wonderful. Thank you everyone for listening. Cheers.