Benchmark 2.1
AI use case penetration
Survey question
For each of the following AI use cases, what is your organization’s current deployment status?
- Audience
- Technical Decision Makers · n=252
- In depth
- Ch 4 — The Agent Engineer
67%
have automated coding and debugging in production — the highest use-case rate in the benchmark
Technical Leadership · n=252 · Wave 3 · vibe coding is a Wave 3 debut
| Category | % in production |
|---|---|
| Automated coding and debugging | 67% |
| Customer support and chatbots | 58% |
| Data analytics and insights | 57% |
| Product design and development | 52% |
| Vibe coding (low-code/no-code) | 47% |
| Automated testing and QA | 41% |
| Cybersecurity and fraud detection | 32% |
| Personalization and recommendation | 29% |
| Sentiment analysis | 29% |
| Predictive maintenance | 23% |
| Image and video recognition | 23% |
| Risk management and compliance | 21% |
Technical Leadership · n=252 · Wave 3 · vibe coding is a Wave 3 debut
Benchmark 2.2
Agentic coding instances
Survey question
Which best describes how your engineering team currently uses agentic coding tools?
- Audience
- Technical Decision Makers · n=252
- In depth
- Ch 2 — The Agentic Divide Quantified Ch 3 — What Makes a Runner Ch 4 — The Agent Engineer
95%
of engineering teams surveyed use agentic coding tools — adoption is statistically near-universal across surveyed B2B software organizations
Technical Leadership · n=252 · Wave 3
| Currently use | Do not use | |
|---|---|---|
| Share of engineering teams | 95% | 5% |
Technical Leadership · n=252 · Wave 3
Benchmark 2.3
Agentic coding spend per engineer
Survey question
Approximately how much does your organization spend on agentic coding tools per engineer per month?
- Audience
- Technical Decision Makers · n=252
- In depth
- Ch 4 — The Agent Engineer Ch 6 — AI Meet P&L
$251–$1,000
is the modal monthly spend on agentic coding tools per engineer — 33% of engineering teams sit in this band, with 23% spending above $1,000
Technical Leadership · n=252 · Wave 3 · the survey used spend bands, so figures are reported as a modal band, not a median
| Category | % of teams |
|---|---|
| More than $5,000 | 8% |
| $1,001–$5,000 | 15% |
| $251–$1,000 (modal) | 33% |
| $101–$250 | 25% |
| $0–$100 | 9% |
| Unsure / don't know | 10% |
Technical Leadership · n=252 · Wave 3 · the survey used spend bands, so figures are reported as a modal band, not a median
Benchmark 2.4
Agentic coding spend trajectory
Survey question
How do you expect your per-engineer agentic coding spend to change in the next 3 months?
- Audience
- Technical Decision Makers · n=252
- In depth
- Ch 4 — The Agent Engineer Ch 6 — AI Meet P&L
63%
expect their per-engineer agentic coding spend to at least double in the next 90 days — with only 3% expecting a decrease
Technical Leadership · n=252 · Wave 3
| Category | % of teams |
|---|---|
| Increase by 10× or more | 3% |
| More than double | 19% |
| Approximately double | 41% |
| Stay the same | 31% |
| Decrease | 3% |
| Unsure / don't know | 3% |
Technical Leadership · n=252 · Wave 3
Benchmark 2.6
Engineering metrics impact
Survey question
How has AI adoption impacted each of the following engineering metrics?
- Audience
- Technical Decision Makers · n=252
- In depth
- Ch 4 — The Agent Engineer
87%
report positive AI impact on development velocity — the highest positive rate of any engineering metric tracked
Technical Leadership · n=252 · Wave 3
| Category | % positive | % negative |
|---|---|---|
| Development velocity | 87% | 4% |
| Cycle time | 74% | 4% |
| Commit frequency | 73% | 5% |
| Lead time | 69% | 5% |
| Documentation completeness | 69% | 6% |
| Deployment frequency | 68% | 5% |
| Pull request activity | 68% | 3% |
| Code quality metrics | 63% | 10% |
| Change failure rate | 49% | 10% |
| Mean time to restore | 48% | 6% |
Technical Leadership · n=252 · Wave 3
Benchmark 2.7
Pilot-to-production velocity
Survey question
How long does it typically take your organization to move an AI feature from pilot to production?
- Audience
- Technical Decision Makers · n=252
- In depth
- Ch 4 — The Agent Engineer
71%
ship AI features from pilot to production in under 6 months — 25% do it in under 3 months
Technical Leadership · n=252 · Wave 3
| Under 3 months | 3–6 months | 7–12 months | Over 12 months | |
|---|---|---|---|---|
| Time to production | 25% | 46% | 23% | 5% |
Technical Leadership · n=252 · Wave 3
Benchmark 2.8
AI evaluation and optimization
Survey question
Which best describes your organization’s approach to AI evaluation and optimization?
- Audience
- Technical Decision Makers · n=252
- In depth
- Ch 4 — The Agent Engineer
15%
have no formal eval-to-optimization loop — running AI in production with no systematic process to turn evaluation results into improvements
Technical Leadership · n=252 · Wave 3
| Category | % of teams |
|---|---|
| Benchmarking and baselining | 25% |
| Continuous production optimization | 21% |
| System-level agent optimization | 17% |
| Interactive / playground tuning | 15% |
| Simulated and adversarial optimization | 6% |
| No formal link from eval to optimization | 11% |
| No formal AI evaluations at all | 4% |
Technical Leadership · n=252 · Wave 3
Benchmark 2.9
AI techniques in use
Survey question
Which of the following AI techniques does your organization currently use?
- Audience
- Technical Decision Makers · n=252
- In depth
- Ch 3 — What Makes a Runner Ch 4 — The Agent Engineer
68%
use prompt & context engineering — now the most widely used AI technique
Technical Leadership · n=252 · Wave 3 · question restructured between waves, not wave-comparable
| Category | % in use |
|---|---|
| Prompt & context engineering | 68% |
| Agentic orchestration | 61% |
| Classical ML & statistical techniques | 55% |
| Retrieval augmented generation (RAG) | 53% |
| Data clustering | 33% |
| Reinforcement learning (RL / RLHF) | 31% |
| Fine-tuning (LoRA, QLoRA) | 29% |
| Memory & state management | 27% |
| Transformers | 23% |
| Convolutional neural networks | 20% |
| Data dimensionality reduction | 17% |
| Ensemble learning | 14% |
Technical Leadership · n=252 · Wave 3 · question restructured between waves, not wave-comparable
Benchmark 2.10
AI cost impact
Survey question
Where are you seeing increases or decreases in costs associated with AI functionality in the past 12 months?
- Audience
- Technical Decision Makers · n=252
- In depth
- Ch 6 — AI Meet P&L
42%
report a net decrease in labor costs — the sharpest cost reduction of any category, offset by rising API, data storage, and training costs
Technical Leadership · n=252 · Wave 3 · neutral responses excluded
| Category | % increased | % decreased |
|---|---|---|
| API / off-the-shelf model costs | 58% | 16% |
| Data storage and processing | 56% | 15% |
| Training and skills development | 54% | 20% |
| Hardware costs | 41% | 13% |
| Implementation and integration | 40% | 30% |
| Software maintenance and support | 37% | 31% |
| Operational costs | 38% | 37% |
| Labor costs | 26% | 42% |
Technical Leadership · n=252 · Wave 3 · neutral responses excluded
Benchmark 2.11
AI cost management strategies
Survey question
Which of the following cost management strategies are you employing?
- Audience
- Technical Decision Makers · n=252
- In depth
- Ch 6 — AI Meet P&L
45%
choose lower-cost foundation models as their primary cost management strategy — the most common tactic, ahead of reducing headcount (39%) and restricting scope (37%)
Technical Leadership · n=252 · Wave 3
| Category | % employing |
|---|---|
| Choosing lower-cost foundation models | 45% |
| Purchasing third-party solutions instead of building | 43% |
| Reducing headcount | 39% |
| Restricting AI to most important initiatives | 37% |
| Employing simpler techniques | 33% |
| None | 4% |
Technical Leadership · n=252 · Wave 3
Benchmark 2.12
Defensible data advantage for AI
Survey question
Where does your organization have its strongest defensible data advantage for AI?
- Audience
- Technical Decision Makers · n=252
- In depth
- Ch 6 — AI Meet P&L
33%
name proprietary customer-interaction data as their strongest defensible data advantage for AI — 6 points ahead of domain-specific ontologies and knowledge graphs
Technical Leadership · n=252 · Runner n=66 / Jogger n=94 / Walker n=75 / Crawler n=17 · Wave 3
| Category | Total | Runner | Jogger | Walker | Crawler |
|---|---|---|---|---|---|
| Proprietary user / customer interaction data | 33% | 27% | 32% | 37% | 41% |
| Domain-specific ontologies / knowledge graphs | 27% | 36% | 26% | 21% | 24% |
| System and product telemetry | 15% | 17% | 15% | 12% | 24% |
| High-quality human annotations / labels | 10% | 9% | 15% | 8% | 0% |
| No clear defensible data advantage | 9% | 6% | 7% | 13% | 12% |
| Unique third-party / partner data assets | 5% | 5% | 4% | 8% | 0% |
Technical Leadership · n=252 · Runner n=66 / Jogger n=94 / Walker n=75 / Crawler n=17 · Wave 3
Benchmark 2.13
Barriers to production scale for AI
Survey question
Which of the following is the #1 biggest barrier or challenge your organization faces in trusting, scaling, and productizing agentic AI?
- Audience
- Technical Decision Makers · n=252
29%
name inconsistent or unreliable performance as the #1 barrier to trusting, scaling and productizing agentic AI — 14 points higher than the second most cited barrier, regulatory and compliance challenges
Technical Leadership · n=252 · Wave 3
| Category | % ranked #1 |
|---|---|
| Inconsistent or unreliable performance | 29% |
| Regulatory and compliance challenges | 15% |
| Responsible / trustworthy AI concerns | 13% |
| Privacy and data security issues | 11% |
| Lack of explainability / interpretability | 8% |
| Lack of internal expertise or AI talent | 6% |
| Limited control and governance over AI actions | 4% |
| Ambiguous business case or unclear ROI | 4% |
| High cost of scaling, training, deploying | 3% |
| Lack of standardized frameworks | 2% |
| Legacy systems make integration hard | 2% |
| Resistance from employees or leadership | 1% |
| Job displacement and workforce impact | 1% |
Technical Leadership · n=252 · Wave 3
Benchmark 2.14
AI model employment depth
Survey question
Which of the following AI model types are your team and your organization currently employing, and at what level?
- Audience
- Technical Decision Makers · n=252
- In depth
- Ch 3 — What Makes a Runner
53%
employ large language models at an advanced or fundamental level — the only model type a majority has integrated deeply. Large reasoning models are next at 18%, 35 points behind
Technical Leadership · n=252 · Wave 3 · fundamental integration = the technique is fundamental to the business and integrated into core operations · advanced implementation = actively implementing the technique across multiple areas to enhance products and services
| Category | Any level | Advanced or fundamental |
|---|---|---|
| Large language models (LLMs) | 84% | 53% |
| Multimodal models | 51% | 27% |
| Large reasoning models (LRMs) | 44% | 18% |
| Open-source / open-weight models | 41% | 6% |
| Voice models | 25% | 6% |
| Model routers | 17% | 10% |
| Diffusion LLMs | 16% | 15% |
| Vision foundation models (VFMs) | 14% | 4% |
| Vision-language-action (VLA) models | 14% | 2% |
| World models | 9% | 5% |
| Video foundation models (ViFMs) | 9% | 3% |
Technical Leadership · n=252 · Wave 3 · fundamental integration = the technique is fundamental to the business and integrated into core operations · advanced implementation = actively implementing the technique across multiple areas to enhance products and services
Benchmark 2.15
AI resourcing status, technical teams
Survey question
Which of the following AI-related roles does your organization currently have in-house, and which are you hiring or outsourcing in the next 12 months?
- Audience
- Technical Decision Makers · n=252
- In depth
- Ch 8 — The Human Side of the AI Stack
62%
have an in-house ML/AI engineer — the highest of the nine AI roles evaluated, while model evaluation and red-teaming is unresourced at 60%
Technical Leadership · n=252 · Wave 3
| Category | In-house | Outsourced | Not resourced yet |
|---|---|---|---|
| Machine learning engineer / AI engineer | 62% | 12% | 27% |
| Data quality lead / data governance lead | 52% | 13% | 35% |
| AI product manager | 51% | 7% | 42% |
| AI solutions architect | 50% | 10% | 40% |
| MLOps / LLMOps engineer | 46% | 15% | 39% |
| Vibe coder | 45% | 2% | 53% |
| Prompt engineer / conversation designer | 42% | 9% | 48% |
| AI security / AI risk & compliance lead | 37% | 16% | 47% |
| Model evaluator / red team | 23% | 17% | 60% |
Technical Leadership · n=252 · Wave 3
Benchmark 2.16
AI impact on company metrics
Survey question
Which of the following best describes the impact of AI on your company’s metrics in the past 12 months?
- Audience
- Technical Decision Makers · n=252
- In depth
- Ch 6 — AI Meet P&L
3–9%
is the full range of negative impact reported across all six company KPIs measured — very low against the positive impacts reported from AI
Technical Leadership · n=252 · Wave 3 · for CAC a decrease is the positive direction
| Category | Positive | Neutral / don't know | Negative |
|---|---|---|---|
| Cost savings | 66% | 25% | 9% |
| Customer satisfaction | 62% | 35% | 3% |
| Revenue | 53% | 43% | 4% |
| Lifetime value (LTV) | 48% | 49% | 3% |
| Customer acquisition cost (CAC) | 42% | 50% | 8% |
| Net retention | 42% | 52% | 6% |
Technical Leadership · n=252 · Wave 3 · for CAC a decrease is the positive direction