The R&D engineering workflow has been structurally rewritten in 18 months. AI tools and agents can not only make engineers faster — they can change what engineers spend their time on, how long it takes to ship, what it costs per seat, and how reliable the output is. Wave 3 benchmarks indicate this functional transformation is scaling quickly, with many aspects now majority-adopted in production.
Finding 1
Automated coding is now mainstream
Automated coding and debugging in production jumped to 67% — up 30 points from Wave 2, one of the largest wave-over-wave jumps in our R&D benchmarks.
Automated coding & debugging in production
67% of technical decision makers now run automated coding and debugging in production — up 30 points from Wave 2 in just nine months.
0 pts
Finding 2
Universal access, uneven intensity
At 95% adoption, agentic-coding-tool access is reaching saturation. What Wave 3 reveals is a bifurcation in usage intensity that universal adoption numbers obscure: 53% of Runners have most or all engineers running parallel coding agents versus 8% of Walkers — a 6.6× gap. We believe this benchmark helps track the transition from individual engineering productivity gains to full R&D team workflow productivity gains.
R&D decision makers using agentic coding tools
95% of R&D decision makers use agentic coding tools at some level — adoption is near-universal; the real divide is single versus parallel use.
0%
“How would you describe your organization’s use of agentic coding tools?” Technical Decision Makers · n=252 · Wave 3 · Significant (95%)
“How would you describe your organization’s use of agentic coding tools?” Technical Decision Makers · n=252 · Wave 3 · Significant (95%)
Finding 3
The productization cycle has compressed
The conventional 12-to-18-month B2B SaaS productization cycle appears to no longer apply. 71% of technical decision makers ship AI features from pilot to production in under six months; 25% do it in under three. Growth-stage organizations surveyed move fastest — 34% ship quarterly or faster versus 16% of enterprise peers. The engineering pipeline — evaluation, testing, deployment — has been compressed, likely due to the presence of AI automated tools often coding in parallel.
Ship pilot-to-production in under six months
71% of technical decision makers ship AI features from pilot to production in under six months, and 25% do it in under three.
0%
“How long does it typically take to move an AI feature from pilot to production?” Technical Decision Makers · n=252 · Wave 3 · Significant (95%)
“How long does it typically take to move an AI feature from pilot to production?” Technical Decision Makers · n=252 · Wave 3 · Significant (95%)
Finding 4
The per-seat budget model is breaking
The per-seat software budget model is breaking in the age of AI. 33% of technical decision makers report per-engineer agentic-coding spend in the $251–$1,000 band, and 23% spend above $1,000 per engineer per month. Only 9% remain below $100. A 200-person engineering team at the modal band is spending an estimated $50K–$200K per month on agentic tooling alone — a P&L line item that did not exist 18 months ago. And 63% expect that spend to at least double in the next 90 days. In our view, it’s unlikely most engineering budgets have modeled this curve. We explore the P&L impacts of AI in Chapter 6: AI Meet P&L.
The new cost curve
“How much does your organization spend per engineer per month on agentic coding tools?” Technical Decision Makers · n=252 · Wave 3 · Significant (95%)
“How much does your organization spend per engineer per month on agentic coding tools?” Technical Decision Makers · n=252 · Wave 3 · Significant (95%)
Finding 5
Reliability is catching up to velocity
In Wave 2, agentic coding was clearly buying speed and leaving reliability behind. That gap has narrowed. Every engineering metric we track moved in the same direction this wave: the share of technical decision makers reporting a positive AI impact rose +17 points on development velocity and +15 on deployment frequency, but it rose fastest on the two reliability measures — +24 points on change-failure rate and +19 on mean-time-to-restore.
Reliability still trails in absolute terms. Development velocity is the standout at 87% positive, while change-failure rate (49%) and mean-time-to-restore (48%) remain the only metrics where fewer than half of teams report improvement. The direction of travel, though, is that the tooling is closing its own gap — and Chapter 7: The Governance Gap examines what is still missing from the controls around it.
“How has AI adoption impacted each of the following engineering metrics?” Technical Decision Makers · n=252 · Direct W2 → W3 · % reporting positive impact · the W2 scale (“increased/decreased by X%”) and W3 scale (“positive/negative impact”) are directionally comparable but not precisely equivalent
“How has AI adoption impacted each of the following engineering metrics?” Technical Decision Makers · n=252 · Direct W2 → W3 · % reporting positive impact · the W2 scale (“increased/decreased by X%”) and W3 scale (“positive/negative impact”) are directionally comparable but not precisely equivalent
Finding 6
Many teams ship AI without a feedback loop
11% of technical decision makers have no formal link between AI evaluation and optimization; another 4% perform no formal evaluation at all. Combined, 15% of engineering teams are deploying AI at production scale with no systematic way to detect degradation, capture failures, or improve model performance over time. At the current 67% production rate, that is not a small tail risk — it is, in our view, a meaningful share of the market operating without a feedback loop that every other production software system takes for granted.
Running production AI with no evaluation loop
15% of technical decision makers run AI in production with no formal evaluation-to-optimization loop (11% no link, plus 4% with no formal evaluation at all).
0%
“Which best describes your organization’s approach to AI evaluation and optimization?” Technical Decision Makers · n=252 · Wave 3 · Significant (95%)
“Which best describes your organization’s approach to AI evaluation and optimization?” Technical Decision Makers · n=252 · Wave 3 · Significant (95%)