Chapter 4

The Agent Engineer

AI rewrote the R&D workflow in 18 months — what engineers do, how fast they ship, and what it costs.

Georgian + NewtonX AI, Applied Benchmarks, Wave 3 (n=252 Technical Decision Makers) · March–April 2026

The R&D engineering workflow has been structurally rewritten in 18 months. AI tools and agents can not only make engineers faster — they can change what engineers spend their time on, how long it takes to ship, what it costs per seat, and how reliable the output is. Wave 3 benchmarks indicate this functional transformation is scaling quickly, with many aspects now majority-adopted in production.

Finding 1

Automated coding is now mainstream

Automated coding and debugging in production jumped to 67% — up 30 points from Wave 2, one of the largest wave-over-wave jumps in our R&D benchmarks.

Automated coding & debugging in production

Automated coding & debugging in production

67% of technical decision makers now run automated coding and debugging in production — up 30 points from Wave 2 in just nine months.

0 pts

As of early 2026, AI-authored code makes up 26.9% of all production code — up from 22% the prior quarter. (Shift Magazine / LinearB, Feb 2026)

Finding 2

Universal access, uneven intensity

At 95% adoption, agentic-coding-tool access is reaching saturation. What Wave 3 reveals is a bifurcation in usage intensity that universal adoption numbers obscure: 53% of Runners have most or all engineers running parallel coding agents versus 8% of Walkers — a 6.6× gap. We believe this benchmark helps track the transition from individual engineering productivity gains to full R&D team workflow productivity gains.

R&D decision makers using agentic coding tools

R&D decision makers using agentic coding tools

95% of R&D decision makers use agentic coding tools at some level — adoption is near-universal; the real divide is single versus parallel use.

0%

Agentic coding usage intensity by tier

“How would you describe your organization’s use of agentic coding tools?” Technical Decision Makers · n=252 · Wave 3 · Significant (95%)

Runner53%27%18%Jogger26%44%27%Walker8%41%41%Crawler24%65%0%100%
Most/all parallelSome parallelSingle agent only

“How would you describe your organization’s use of agentic coding tools?” Technical Decision Makers · n=252 · Wave 3 · Significant (95%)

Finding 3

The productization cycle has compressed

The conventional 12-to-18-month B2B SaaS productization cycle appears to no longer apply. 71% of technical decision makers ship AI features from pilot to production in under six months; 25% do it in under three. Growth-stage organizations surveyed move fastest — 34% ship quarterly or faster versus 16% of enterprise peers. The engineering pipeline — evaluation, testing, deployment — has been compressed, likely due to the presence of AI automated tools often coding in parallel.

Ship pilot-to-production in under six months

Ship pilot-to-production in under six months

71% of technical decision makers ship AI features from pilot to production in under six months, and 25% do it in under three.

0%

Time from pilot to production

“How long does it typically take to move an AI feature from pilot to production?” Technical Decision Makers · n=252 · Wave 3 · Significant (95%)

Total sample25%46%23%Enterprise16%49%28%Growth-stage34%43%18%0%100%
Under 3 months3–6 months7–12 monthsOver 12 months

“How long does it typically take to move an AI feature from pilot to production?” Technical Decision Makers · n=252 · Wave 3 · Significant (95%)

Finding 4

The per-seat budget model is breaking

The per-seat software budget model is breaking in the age of AI. 33% of technical decision makers report per-engineer agentic-coding spend in the $251–$1,000 band, and 23% spend above $1,000 per engineer per month. Only 9% remain below $100. A 200-person engineering team at the modal band is spending an estimated $50K–$200K per month on agentic tooling alone — a P&L line item that did not exist 18 months ago. And 63% expect that spend to at least double in the next 90 days. In our view, it’s unlikely most engineering budgets have modeled this curve. We explore the P&L impacts of AI in Chapter 6: AI Meet P&L.

The new cost curve

The new cost curve

Modal monthly spend per engineer
33% of teams fall in this band
$251–$1,000
Expect spend to at least double in 90 days
A line item that didn’t exist 18 months ago
63%
Per-engineer monthly spend on agentic coding tools

“How much does your organization spend per engineer per month on agentic coding tools?” Technical Decision Makers · n=252 · Wave 3 · Significant (95%)

0%10%20%30%40%Under $100Under $100 — % of technical decision makers: 9%9%$101–$250$101–$250 — % of technical decision makers: 25%25%$251–$1,000$251–$1,000 — % of technical decision makers: 33%33%$1,001–$5,000$1,001–$5,000 — % of technical decision makers: 15%15%More than $5,000More than $5,000 — % of technical decision makers: 8%8%

“How much does your organization spend per engineer per month on agentic coding tools?” Technical Decision Makers · n=252 · Wave 3 · Significant (95%)

Finding 5

Reliability is catching up to velocity

In Wave 2, agentic coding was clearly buying speed and leaving reliability behind. That gap has narrowed. Every engineering metric we track moved in the same direction this wave: the share of technical decision makers reporting a positive AI impact rose +17 points on development velocity and +15 on deployment frequency, but it rose fastest on the two reliability measures — +24 points on change-failure rate and +19 on mean-time-to-restore.

Reliability still trails in absolute terms. Development velocity is the standout at 87% positive, while change-failure rate (49%) and mean-time-to-restore (48%) remain the only metrics where fewer than half of teams report improvement. The direction of travel, though, is that the tooling is closing its own gap — and Chapter 7: The Governance Gap examines what is still missing from the controls around it.

Positive AI impact on engineering metrics, Wave 2 vs Wave 3

“How has AI adoption impacted each of the following engineering metrics?” Technical Decision Makers · n=252 · Direct W2 → W3 · % reporting positive impact · the W2 scale (“increased/decreased by X%”) and W3 scale (“positive/negative impact”) are directionally comparable but not precisely equivalent

0%20%40%60%80%100%Development velocityDevelopment velocity — Wave 2: 70%70%Development velocity — Wave 3: 87%87%Deployment frequencyDeployment frequency — Wave 2: 53%53%Deployment frequency — Wave 3: 68%68%Change failure rateChange failure rate — Wave 2: 25%25%Change failure rate — Wave 3: 49%49%Mean time to restoreMean time to restore — Wave 2: 29%29%Mean time to restore — Wave 3: 48%48%
Wave 2Wave 3

“How has AI adoption impacted each of the following engineering metrics?” Technical Decision Makers · n=252 · Direct W2 → W3 · % reporting positive impact · the W2 scale (“increased/decreased by X%”) and W3 scale (“positive/negative impact”) are directionally comparable but not precisely equivalent

Finding 6

Many teams ship AI without a feedback loop

11% of technical decision makers have no formal link between AI evaluation and optimization; another 4% perform no formal evaluation at all. Combined, 15% of engineering teams are deploying AI at production scale with no systematic way to detect degradation, capture failures, or improve model performance over time. At the current 67% production rate, that is not a small tail risk — it is, in our view, a meaningful share of the market operating without a feedback loop that every other production software system takes for granted.

Running production AI with no evaluation loop

Running production AI with no evaluation loop

15% of technical decision makers run AI in production with no formal evaluation-to-optimization loop (11% no link, plus 4% with no formal evaluation at all).

0%

Approach to AI evaluation and optimization

“Which best describes your organization’s approach to AI evaluation and optimization?” Technical Decision Makers · n=252 · Wave 3 · Significant (95%)

0%10%20%30%40%Benchmarking & baseliningBenchmarking & baselining — Share: 25%25%Continuous production optimizationContinuous production optimization — Share: 21%21%System-level agent optimizationSystem-level agent optimization — Share: 17%17%Interactive / playground tuningInteractive / playground tuning — Share: 15%15%Simulated & adversarial optimizationSimulated & adversarial optimization — Share: 6%6%No formal link, eval → optimizationNo formal link, eval → optimization — Share: 11%11%No formal evaluationNo formal evaluation — Share: 4%4%

“Which best describes your organization’s approach to AI evaluation and optimization?” Technical Decision Makers · n=252 · Wave 3 · Significant (95%)