2026 State of AI Engineering — Barr Yaron, Amplify Partners

summarized

TLDR

The 2026 State of AI Engineering survey reveals that AI engineering is becoming a discipline with widespread adoption of agents, cost as a first-class constraint, and a multi-model stack. Image generation usage doubled, audio has the highest intent to adopt, and open-weight models augment rather than replace closed models. Teams are standardizing on tools but not models, and agents are increasingly write-enabled, though control methods remain primitive.

Key points

  • Text remains the dominant modality, but audio has the strongest intent to adopt, with 56% of non-users planning to use it, up from 37% last year.
  • Image generation usage doubled from 18% to 36% due to improved models like DALL-E and ChatGPT images.
  • 94% of respondents use closed models, 45% use open-weight models, and over 90% of open-weight users also use closed models, indicating augmentation not replacement.
  • Cost has become a first-class engineering constraint, with 76% of respondents adjusting AI usage based on cost, and cost/token usage is the second most monitored metric after quality.
  • Agent adoption surged to 95% of teams, with 89% of agents having write access, tripling the share of write-enabled agents compared to last year.
  • Evaluation remains the top stack challenge, with 'vibe review' as the most common method, and inference/model serving is the most bought layer while prompt management is mostly built in-house.
  • AI has a net positive effect on organizations (97%), enabling cheaper failure and more experimentation, but also causing erosion of deep technical skills and blurring role boundaries.
  • 67% of respondents expect a leading lab to declare AGI within five years, and only 9% believe Transformers will still be state-of-the-art in five years.

Techniques

  • routing by task type
  • human-in-the-loop approvals
  • gating permissions
  • vibe review
  • task decomposition
  • retrieval
  • memory
  • sandboxing
  • fine-tuning
  • prompt management
  • eval
Transcript (captions)
[music] Now joining us on stage is the partner at Amplify Barren. [music] Fantastic. You did a great job practicing. I feel very very loved. Um, let's get started. So, like you just heard, my name is Bar. I run a survey every year on the state of AI engineering. And the funny thing about running a survey on the state of AI engineering is that the field changes as you make the slides. Just in the past week, we've had Frontier releases treated like national security events. Meta reportedly exploring selling AI compute. By the time I get off stage, maybe something else will happen. So, if I miss a major announcement while I'm up here, please come find me after. But that's exactly why we run the survey every year to cut through the noise, take a moment, step back and understand what AI engineers are actually doing. Uh for the first time this year, we were thrilled to partner with Notion and Verscell to run this survey. Very quickly on me, uh this is the least interesting slide. I'm an investment partner at Amplify. Very lucky to invest in companies built by and for AI engineers. And I'll make the same promise that I make every single year, which is short time on bar, long time on bar charts. So, let's get right into it with lots of bar charts. First, let's talk about well, maybe raise your hand. Did you fill out the survey? This is a very large group. Okay. Yes, I see you in the front. Um, if the answer is you, thank you so much. If the answer is not you, I will find you in 2027. But genuinely, this only exists because a thousand of you gave your time. So, thank you. We had 1,048 respondents this year, which is a lot of AI engineers. And to be precise, this is not just AI engineers, as I'm sure you see at the conference. Every year, we see that AI engineering is more of a discipline than a job title. It touches founders, CTO's, engineers, product people, folks across company sizes and experience levels. And that range shows up in experience too. Um for the third year running we see the same pattern which is skew towards senior engineers but newer to AI. Of those with over 10 years of software experience over half have three years or less of AI experience which tracks uh these are very experienced engineers learning a new paradigm in real time. And the newest cohort, the ones who just started uh engineering, the median new engineer has nearly as much AI experience as the median 10-year software veteran. Uh so the newest engineers have never known software without this. But doing AI doesn't mean one thing. We talked about all these different titles, all these different roles. Before we get into models and agents, I have a more basic question, which is when people say they're doing AI at work, what are they actually doing? So, first up, like to start with the modalities. We asked, which modalities are you actively building with at work? Can anyone take a guess? Text dominates. I know. Hold your applause. Um, but one piece of this chart that I always find very interesting and I always look at is the ratio of nope, I'm not using this modality to I'm not using it, but I do plan to. I call this the intent to adopt ratio. Of the people who are not building with a modality today, how many say they plan to use it? And audio has the strongest intent to adopt this year. Among AI engineers who are not building with audio today, a whopping 56% say they plan to adopt it in the AI applications they build. And this is not a brand new signal. Last year audio also had the highest intent to adopt across modalities but 37%. So audio continues to take the lead and have high interest but that interest is accelerating. Now there has been an audio swing but if we look at what changed most from the last year in the survey the biggest jump is actually in people using image generation. The share of respondents using generative AI for images and feeling really good about it doubled from 18% last year to 36% this year. Makes sense if you look at what we launched in the same window. Over the past year plus survey time, uh we've had models Nano Banana, Nano Banana 2, Chat GPT images 2.0. The products have gotten much better. What used to feel like an efficient way to generate cursed hands is just increasingly becoming a part of real work. Audio may have the strongest intent to adopt, but image generation shows us what happens when a modality crosses that threshold. So, I'm excited to continue watching these adoption curves every single year. I think we're going to see a lot this year. Uh, now models. If you've Who here spends time on Twitter? All right. Yes. I imagine this is a very Twitter pilled uh crowd. If you spend any time on Twitter in this uh in this circle, you've seen a lot written about openweight models these past few months and I think we'll see it even more in the next year. Um so we asked what models are you actually using in production. 94% use closed models. 45% are using openweight models. But here's the thing, you know, openweight models are not replacing closed models for the most part. at least not yet. The respondents using openweight models, over 90% of them are also using closed models. So they're looking like an augmentation. Teams are mixing and matching. We also asked just to double click on this for the top three considerations when choosing a model. If you're choosing a model, what is important to you? Um and despite the airtime of the open versus closed, it's not what drives model choice. It was a top three consideration for only 5% of the respondents. What matters is actually more straightforward. It's quality. Quality dominates. Followed by agentic capabilities like tool calling and cost tied right with it. We'll money money. We'll get back to that. Um, one thing that I found very interesting is that reliability is not near the top. Only one in five named reliability. That doesn't mean teams stopped caring about reliability. Uh there are different ways to interpret this data. My guess is that it's more likely to become a threshold requirement and the models they're choosing are reliable enough so the decision moves up the stack outside of certain circumstances to quality, capability, cost, but we could talk after. All right, so here's where the model story all comes together. Like I said, teams are not choosing one model and calling it a day. Earlier I showed that 87% of teams are using more than one model. Uh the model that's the opposite of standardization. Uh and the way that they choose models for given tasks varies. Most popular is routing by task type. Some run multiple models compare outputs. Some route based on cost. Uh but models are good at different things. What was interesting was that more than half of respondents said that their organizations starting to standardize on fewer AI tools. They're trading standardiz flexibility for standardization. A share of those are mixed. They say they're standardizing on some layers while staying flexible on others. But the headline here is that there's we're in the early great standardization of the platform and tools, not the models. All right, this is the slide where anyone who's opened an AI bill in the last year starts nodding. So it turns out that infinite intelligence still comes with a usagebased bill. Once teams are managing many models and AI workflows, the next question becomes cost. Cost is now a first class engineering constraint. We see this in the data. 40% of respondents say that cost regularly shapes how ambitiously they use AI and another 36% say that it sometimes does. Well, this is pretty straightforward. So all in about uh three out of four respondents are adjusting their AI usage based on cost and maybe the fourth has a company card. That might be surprising or maybe it's obvious but 12 months ago it was not. Token maxing is cool. Being able to find real use cases is amazing but cost is becoming a real big part of the product decision today. And it shows up in monitoring too. What folks are monitoring in production includes cost and token usage as the number two thing they watch for. It's being monitored like an SLA right under quality itself. Which brings us to the biggest line item of them all. Agents. We've been talking about agents for a while. Uh this year, as you've seen, as you'll see today, as you've seen in previous days, you're going to talk a lot about harness engineering. They're escaping demo world. So, we asked respondents what level of tool permissions their agents typically have. And this is where agents start to look more real. There are two things happening at once. First, and I don't think this is surprising, relative to last year, there are far more teams using agents. This year, 95%, this seems high to me, 95% say they're using agents, roughly double last year. Second, amongst the teams that are using agents, those agents are much more likely to have write access. Last year, 52% of folks building with agents said their agents could actually write data. This year, that number is 89%. So when you combine these two shifts, more teams using agents and more of those agents having write permissions, the share of all the respondents and again it's a survey using write enabled agents is up more than three times relative to last year. So this is really the big shift. Agents are no longer reading, summarizing, drafting. They're taking actions inside of systems. And that raises the obvious question, how are we controlling all of this? um with pretty blunt instruments. Uh very there are many ways that folks are controlling agents today. The top two are human in the loop approvals and gating permissions which are the right instincts but kind of the same toolkit you'd use to manage an intern. Below that the results scatter. Task decomposition, retrieval, memory, sandboxing. People are trying everything. Nobody has settled the control layer for agents. uh memory and persistent context is one that I'm watching very carefully right now. I think it's going to evolve a lot in the next year. And when agents fail or when people complain about agents failing to be more precise, it's usually the thinking, not the plumbing. So, uh you know, like twothirds say that hallucination or losing context mid task is what frustrates them the most. All right, so agents are out in the wild, which makes it a good time to look at what everyone's actually running underneath. So let's take a peek at the stack. Um, we asked, what is the biggest challenge in your stack? Every single year that I ask this, the answer, the number one answer is eval. Um, so Eval's lead here is same as always, but by a very thin margin, like that margin is getting smaller. And I'll say the quiet part here, which is that 96% of the people in the survey in this room have a problem with the stack. You just can't agree on which one. Um, so if you're deciding what to build next, if you're interested in infrastructure, that scatter is the map. And the leading challenge, how to evaluate your AI outputs requires many different methods, but as always, the vibe review is number one. So there are some consistent things that we'll see if they change over the time, but they they have not changed. Okay, this is interesting. So across eight layers of the stack, we asked what do people build versus buy. Um again, maybe the corporate card is is going to play a part in this, but there is a wide range and mix for every layer of the stack and a few clear takeaways. So the first is that inference and model serving is the layer that people buy the most. Many people don't want to build inference infrastructure and fair enough. Uh prompt management is the opposite. 61% build it themselves. Um apparently everyone's prompts are special. And this is true of a lot of the product logic prompts rag eval. They tend to stay inhouse on a relative basis. fine-tuning is the clearest not yet. Like most people don't have it at all and uh folks are pretty locked in. So those who bought aren't looking as much to build. Those who built aren't looking as much to buy. Uh but those are those are the core takeaways from the usage in our stack. So many of you work on teams and like we said at the start these range from solo founders to large enterprises. What is this doing to teams? And remember this is a builderheavy sample. But among builders the vibes are good which you know I'm sure if you look to your left and your right you're feeling that the vibes are pretty good. 97% report a net positive effect on their organization. The top effect isn't really just speed. It's cheaper failure, more experimentation, more prototypes, more bets. It didn't just make engineers faster, but it made trying things nearly free. And so there's some happy campers as a result of that. But it's not free free. You know, there's no free lunch as nothing is. So the same tool that increases experimentation also increases review burden. Both can be true. And um you know o over nine and 10 respondents are feeling negative downstream effects in some way. The most common ones being wid you know widely discussed at this conference uh online and anywhere that you see AI engineers erosion of deep technical skills and understanding of the codebase. And these are consequences of cheap code generation. And the org chart is really feeling it. So many folks, 81% are saying that AI is blurring the line between their role as engineers and product design and marketing. These stats shocked me. Um, where you feel it the most is shipping software once exclusively the engineers domain. I know folks talk about vibe coding and how that's accessible to more folks than ever before in different roles, but today over a third of teams have non-developers shipping features, which was pretty wild to me. Mostly smaller, mostly internal, but 17% say that non-developers are regularly shipping customerf facing features across the stack. And even when non-developers aren't shipping, a third of teams see them building really useful things, prototypes, front-end mocks, and more. So, shipping software is not gated on being an engineer. We knew this, but uh the extent to which it's being pushed is is higher than I expected. All right, so where does all of this go? We always ask people to place bets rapid fire. So, let's talk about those results. Um, so present tense first. 76% say AI boosted their job satisfaction. So that's good for most of this crowd. I hope you're uh as uh Alphaba and Glenda say, I hope you're happy now. Um, that's great. But 59% fear today's AI code creates long-term liabilities. Only a third call software engineering a solved problem. Although uh when I have conversations with folks sometimes the way in which they define software engineering is different. So you can read into that stat as you will. Um happier faster but embracing the maintenance build is the TLDDR and people are unsure what's going to happen with hiring. And for the five-year bets we have 67% expect a leading lab will declare AGI in the next five years. Note the wording. We said will, we asked about the press release, not the achievement. So, will they declare it? Yes. What does that mean? Not sure. Uh, only 9% bet on Transformers being state-of-the-art in 5 years. Most are unsure. Uh, but that was interesting. And then my favorite, will there be more AI compute in space or on land? 36 yes. 38 no. The most divisive question in the survey is about outer space. I promised you a lot of bar charts and that was a lot of information. So a review or our 2026 wrapped impact is overwhelmingly positive. Image genen doubled or happy image gen doubled while audio has the highest adoption intent the same as last year. cost really became a first class constraint and we see that everywhere in monitoring in how ambitious folks that are going out and building AI products are behaving. Open weights augment but they don't replace. So we're seeing a multimodel future with a consolidation of the stack. Agents got right access more than ever before, tripling relative to last year. While the guardrails stayed pretty primitive, and inference is the buy market, everything closer to product logic tends to relatively stay more in-house. It is a very exciting time to be an AI engineer. I cannot wait to see how the next year unfolds. So, you can find the full report in the link up here. Every chart plus some cuts that we didn't have time for today. Um, I won't ask you to fill out a survey about the survey, but if there's something that you want on the books for 2027, something you're curious about, you can come find me here on the internet. I'm easy to spot. Thank you so much. Uh, we will see you next year or per 36% of you, maybe in orbit. Thank you.

Frontier News · by Hyperjump Technology