Benchmarks

R&D Benchmarks

Use-case penetration, agentic coding, spend, engineering metrics, the model stack, AI resourcing, and the cost picture for technical teams.

Georgian + NewtonX AI, Applied Benchmarks, Wave 3 · Technical Leadership n=252 unless noted · March–April 2026

Benchmark 2.1

AI use case penetration

Survey question

For each of the following AI use cases, what is your organization’s current deployment status?

Audience
Technical Decision Makers · n=252
In depth
Ch 4 — The Agent Engineer
AI use cases in production, Wave 3

67%

have automated coding and debugging in production — the highest use-case rate in the benchmark

Technical Leadership · n=252 · Wave 3 · vibe coding is a Wave 3 debut

Data for AI use cases in production, Wave 3
Category% in production
Automated coding and debugging67%
Customer support and chatbots58%
Data analytics and insights57%
Product design and development52%
Vibe coding (low-code/no-code)47%
Automated testing and QA41%
Cybersecurity and fraud detection32%
Personalization and recommendation29%
Sentiment analysis29%
Predictive maintenance23%
Image and video recognition23%
Risk management and compliance21%
0%20%40%60%80%Automated coding and debuggingAutomated coding and debugging — % in production: 67%67%Customer support and chatbotsCustomer support and chatbots — % in production: 58%58%Data analytics and insightsData analytics and insights — % in production: 57%57%Product design and developmentProduct design and development — % in production: 52%52%Vibe coding (low-code/no-code)Vibe coding (low-code/no-code) — % in production: 47%47%Automated testing and QAAutomated testing and QA — % in production: 41%41%Cybersecurity and fraud detectionCybersecurity and fraud detection — % in production: 32%32%Personalization and recommendationPersonalization and recommendation — % in production: 29%29%Sentiment analysisSentiment analysis — % in production: 29%29%Predictive maintenancePredictive maintenance — % in production: 23%23%Image and video recognitionImage and video recognition — % in production: 23%23%Risk management and complianceRisk management and compliance — % in production: 21%21%

Technical Leadership · n=252 · Wave 3 · vibe coding is a Wave 3 debut

Benchmark 2.2

Agentic coding instances

Survey question

Which best describes how your engineering team currently uses agentic coding tools?

Audience
Technical Decision Makers · n=252
In depth
Ch 2 — The Agentic Divide Quantified Ch 3 — What Makes a Runner Ch 4 — The Agent Engineer
Agentic coding tool usage, Wave 3

95%

of engineering teams surveyed use agentic coding tools — adoption is statistically near-universal across surveyed B2B software organizations

Technical Leadership · n=252 · Wave 3

Data for Agentic coding tool usage, Wave 3
Currently useDo not use
Share of engineering teams95%5%
Share of engineering teams95%0%100%
Currently useDo not use

Technical Leadership · n=252 · Wave 3

Benchmark 2.3

Agentic coding spend per engineer

Survey question

Approximately how much does your organization spend on agentic coding tools per engineer per month?

Audience
Technical Decision Makers · n=252
In depth
Ch 4 — The Agent Engineer Ch 6 — AI Meet P&L
Monthly agentic coding spend per engineer, Wave 3

$251–$1,000

is the modal monthly spend on agentic coding tools per engineer — 33% of engineering teams sit in this band, with 23% spending above $1,000

Technical Leadership · n=252 · Wave 3 · the survey used spend bands, so figures are reported as a modal band, not a median

Data for Monthly agentic coding spend per engineer, Wave 3
Category% of teams
More than $5,0008%
$1,001–$5,00015%
$251–$1,000 (modal)33%
$101–$25025%
$0–$1009%
Unsure / don't know10%
0%10%20%30%40%More than $5,000More than $5,000 — % of teams: 8%8%$1,001–$5,000$1,001–$5,000 — % of teams: 15%15%$251–$1,000 (modal)$251–$1,000 (modal) — % of teams: 33%33%$101–$250$101–$250 — % of teams: 25%25%$0–$100$0–$100 — % of teams: 9%9%Unsure / don't knowUnsure / don't know — % of teams: 10%10%

Technical Leadership · n=252 · Wave 3 · the survey used spend bands, so figures are reported as a modal band, not a median

Benchmark 2.4

Agentic coding spend trajectory

Survey question

How do you expect your per-engineer agentic coding spend to change in the next 3 months?

Audience
Technical Decision Makers · n=252
In depth
Ch 4 — The Agent Engineer Ch 6 — AI Meet P&L
Expected change in per-engineer agentic coding spend (next 3 months)

63%

expect their per-engineer agentic coding spend to at least double in the next 90 days — with only 3% expecting a decrease

Technical Leadership · n=252 · Wave 3

Data for Expected change in per-engineer agentic coding spend (next 3 months)
Category% of teams
Increase by 10× or more3%
More than double19%
Approximately double41%
Stay the same31%
Decrease3%
Unsure / don't know3%
0%10%20%30%40%50%Increase by 10× or moreIncrease by 10× or more — % of teams: 3%3%More than doubleMore than double — % of teams: 19%19%Approximately doubleApproximately double — % of teams: 41%41%Stay the sameStay the same — % of teams: 31%31%DecreaseDecrease — % of teams: 3%3%Unsure / don't knowUnsure / don't know — % of teams: 3%3%

Technical Leadership · n=252 · Wave 3

Benchmark 2.5

AI share of IT spend

Survey question

Approximately what percentage of your organization’s total IT spend in the past 12 months was allocated to AI tools, infrastructure, and related services?

Audience
Technical Decision Makers · n=252 asked · n=205 answered · n=47 unsure (19%)
In depth
Ch 3 — What Makes a Runner Ch 4 — The Agent Engineer Ch 6 — AI Meet P&L
AI share of total IT spend, Wave 3

15%

median AI share of IT spend among B2B technical decision makers — 19% could not quantify their AI spend

Technical Leadership · n=252 asked · n=205 answered · 19% unsure · Wave 3 · the distribution is right-skewed, so the headline is reported as a median

Data for AI share of total IT spend, Wave 3
Category% of answered
1–10%35%
11–20%21%
21–30%12%
31–50%10%
51–100%3%
0%10%20%30%40%1–10%1–10% — % of answered: 35%35%11–20%11–20% — % of answered: 21%21%21–30%21–30% — % of answered: 12%12%31–50%31–50% — % of answered: 10%10%51–100%51–100% — % of answered: 3%3%

Technical Leadership · n=252 asked · n=205 answered · 19% unsure · Wave 3 · the distribution is right-skewed, so the headline is reported as a median

Benchmark 2.6

Engineering metrics impact

Survey question

How has AI adoption impacted each of the following engineering metrics?

Audience
Technical Decision Makers · n=252
In depth
Ch 4 — The Agent Engineer
AI impact on engineering metrics, Wave 3

87%

report positive AI impact on development velocity — the highest positive rate of any engineering metric tracked

Technical Leadership · n=252 · Wave 3

Data for AI impact on engineering metrics, Wave 3
Category% positive% negative
Development velocity87%4%
Cycle time74%4%
Commit frequency73%5%
Lead time69%5%
Documentation completeness69%6%
Deployment frequency68%5%
Pull request activity68%3%
Code quality metrics63%10%
Change failure rate49%10%
Mean time to restore48%6%
0%20%40%60%80%100%Development velocityDevelopment velocity — % positive: 87%87%Development velocity — % negative: 4%4%Cycle timeCycle time — % positive: 74%74%Cycle time — % negative: 4%4%Commit frequencyCommit frequency — % positive: 73%73%Commit frequency — % negative: 5%5%Lead timeLead time — % positive: 69%69%Lead time — % negative: 5%5%Documentation completenessDocumentation completeness — % positive: 69%69%Documentation completeness — % negative: 6%6%Deployment frequencyDeployment frequency — % positive: 68%68%Deployment frequency — % negative: 5%5%Pull request activityPull request activity — % positive: 68%68%Pull request activity — % negative: 3%3%Code quality metricsCode quality metrics — % positive: 63%63%Code quality metrics — % negative: 10%10%Change failure rateChange failure rate — % positive: 49%49%Change failure rate — % negative: 10%10%Mean time to restoreMean time to restore — % positive: 48%48%Mean time to restore — % negative: 6%6%
% positive% negative

Technical Leadership · n=252 · Wave 3

Benchmark 2.7

Pilot-to-production velocity

Survey question

How long does it typically take your organization to move an AI feature from pilot to production?

Audience
Technical Decision Makers · n=252
In depth
Ch 4 — The Agent Engineer
Time from pilot to production, Wave 3

71%

ship AI features from pilot to production in under 6 months — 25% do it in under 3 months

Technical Leadership · n=252 · Wave 3

Data for Time from pilot to production, Wave 3
Under 3 months3–6 months7–12 monthsOver 12 months
Time to production25%46%23%5%
Time to production25%46%23%0%100%
Under 3 months3–6 months7–12 monthsOver 12 months

Technical Leadership · n=252 · Wave 3

Benchmark 2.8

AI evaluation and optimization

Survey question

Which best describes your organization’s approach to AI evaluation and optimization?

Audience
Technical Decision Makers · n=252
In depth
Ch 4 — The Agent Engineer
Approach to AI evaluation and optimization, Wave 3

15%

have no formal eval-to-optimization loop — running AI in production with no systematic process to turn evaluation results into improvements

Technical Leadership · n=252 · Wave 3

Data for Approach to AI evaluation and optimization, Wave 3
Category% of teams
Benchmarking and baselining25%
Continuous production optimization21%
System-level agent optimization17%
Interactive / playground tuning15%
Simulated and adversarial optimization6%
No formal link from eval to optimization11%
No formal AI evaluations at all4%
0%10%20%30%40%Benchmarking and baseliningBenchmarking and baselining — % of teams: 25%25%Continuous production optimizationContinuous production optimization — % of teams: 21%21%System-level agent optimizationSystem-level agent optimization — % of teams: 17%17%Interactive / playground tuningInteractive / playground tuning — % of teams: 15%15%Simulated and adversarial optimizationSimulated and adversarial optimization — % of teams: 6%6%No formal link from eval to optimizationNo formal link from eval to optimization — % of teams: 11%11%No formal AI evaluations at allNo formal AI evaluations at all — % of teams: 4%4%

Technical Leadership · n=252 · Wave 3

Benchmark 2.9

AI techniques in use

Survey question

Which of the following AI techniques does your organization currently use?

Audience
Technical Decision Makers · n=252
In depth
Ch 3 — What Makes a Runner Ch 4 — The Agent Engineer
AI techniques in use, Wave 3

68%

use prompt & context engineering — now the most widely used AI technique

Technical Leadership · n=252 · Wave 3 · question restructured between waves, not wave-comparable

Data for AI techniques in use, Wave 3
Category% in use
Prompt & context engineering68%
Agentic orchestration61%
Classical ML & statistical techniques55%
Retrieval augmented generation (RAG)53%
Data clustering33%
Reinforcement learning (RL / RLHF)31%
Fine-tuning (LoRA, QLoRA)29%
Memory & state management27%
Transformers23%
Convolutional neural networks20%
Data dimensionality reduction17%
Ensemble learning14%
0%20%40%60%80%Prompt & context engineeringPrompt & context engineering — % in use: 68%68%Agentic orchestrationAgentic orchestration — % in use: 61%61%Classical ML & statistical techniquesClassical ML & statistical techniques — % in use: 55%55%Retrieval augmented generation (RAG)Retrieval augmented generation (RAG) — % in use: 53%53%Data clusteringData clustering — % in use: 33%33%Reinforcement learning (RL / RLHF)Reinforcement learning (RL / RLHF) — % in use: 31%31%Fine-tuning (LoRA, QLoRA)Fine-tuning (LoRA, QLoRA) — % in use: 29%29%Memory & state managementMemory & state management — % in use: 27%27%TransformersTransformers — % in use: 23%23%Convolutional neural networksConvolutional neural networks — % in use: 20%20%Data dimensionality reductionData dimensionality reduction — % in use: 17%17%Ensemble learningEnsemble learning — % in use: 14%14%

Technical Leadership · n=252 · Wave 3 · question restructured between waves, not wave-comparable

Benchmark 2.10

AI cost impact

Survey question

Where are you seeing increases or decreases in costs associated with AI functionality in the past 12 months?

Audience
Technical Decision Makers · n=252
In depth
Ch 6 — AI Meet P&L
Cost increases vs decreases from AI, Wave 3

42%

report a net decrease in labor costs — the sharpest cost reduction of any category, offset by rising API, data storage, and training costs

Technical Leadership · n=252 · Wave 3 · neutral responses excluded

Data for Cost increases vs decreases from AI, Wave 3
Category% increased% decreased
API / off-the-shelf model costs58%16%
Data storage and processing56%15%
Training and skills development54%20%
Hardware costs41%13%
Implementation and integration40%30%
Software maintenance and support37%31%
Operational costs38%37%
Labor costs26%42%
0%20%40%60%API / off-the-shelf model costsAPI / off-the-shelf model costs — % increased: 58%58%API / off-the-shelf model costs — % decreased: 16%16%Data storage and processingData storage and processing — % increased: 56%56%Data storage and processing — % decreased: 15%15%Training and skills developmentTraining and skills development — % increased: 54%54%Training and skills development — % decreased: 20%20%Hardware costsHardware costs — % increased: 41%41%Hardware costs — % decreased: 13%13%Implementation and integrationImplementation and integration — % increased: 40%40%Implementation and integration — % decreased: 30%30%Software maintenance and supportSoftware maintenance and support — % increased: 37%37%Software maintenance and support — % decreased: 31%31%Operational costsOperational costs — % increased: 38%38%Operational costs — % decreased: 37%37%Labor costsLabor costs — % increased: 26%26%Labor costs — % decreased: 42%42%
% increased% decreased

Technical Leadership · n=252 · Wave 3 · neutral responses excluded

Benchmark 2.11

AI cost management strategies

Survey question

Which of the following cost management strategies are you employing?

Audience
Technical Decision Makers · n=252
In depth
Ch 6 — AI Meet P&L
AI cost management strategies, Wave 3

45%

choose lower-cost foundation models as their primary cost management strategy — the most common tactic, ahead of reducing headcount (39%) and restricting scope (37%)

Technical Leadership · n=252 · Wave 3

Data for AI cost management strategies, Wave 3
Category% employing
Choosing lower-cost foundation models45%
Purchasing third-party solutions instead of building43%
Reducing headcount39%
Restricting AI to most important initiatives37%
Employing simpler techniques33%
None4%
0%20%40%60%Choosing lower-cost foundation modelsChoosing lower-cost foundation models — % employing: 45%45%Purchasing third-party solutions instead ofbuildingPurchasing third-party solutions instead of building — % employing: 43%43%Reducing headcountReducing headcount — % employing: 39%39%Restricting AI to most importantinitiativesRestricting AI to most important initiatives — % employing: 37%37%Employing simpler techniquesEmploying simpler techniques — % employing: 33%33%NoneNone — % employing: 4%4%

Technical Leadership · n=252 · Wave 3

Benchmark 2.12

Defensible data advantage for AI

Survey question

Where does your organization have its strongest defensible data advantage for AI?

Audience
Technical Decision Makers · n=252
In depth
Ch 6 — AI Meet P&L
Strongest defensible data advantage, by maturity tier

33%

name proprietary customer-interaction data as their strongest defensible data advantage for AI — 6 points ahead of domain-specific ontologies and knowledge graphs

Technical Leadership · n=252 · Runner n=66 / Jogger n=94 / Walker n=75 / Crawler n=17 · Wave 3

Data for Strongest defensible data advantage, by maturity tier
CategoryTotalRunnerJoggerWalkerCrawler
Proprietary user / customer interaction data33%27%32%37%41%
Domain-specific ontologies / knowledge graphs27%36%26%21%24%
System and product telemetry15%17%15%12%24%
High-quality human annotations / labels10%9%15%8%0%
No clear defensible data advantage9%6%7%13%12%
Unique third-party / partner data assets5%5%4%8%0%
0%10%20%30%40%50%Proprietary user / customer interactiondataProprietary user / customer interaction data — Total: 33%33%Proprietary user / customer interaction data — Runner: 27%27%Proprietary user / customer interaction data — Jogger: 32%32%Proprietary user / customer interaction data — Walker: 37%37%Proprietary user / customer interaction data — Crawler: 41%41%Domain-specific ontologies / knowledgegraphsDomain-specific ontologies / knowledge graphs — Total: 27%27%Domain-specific ontologies / knowledge graphs — Runner: 36%36%Domain-specific ontologies / knowledge graphs — Jogger: 26%26%Domain-specific ontologies / knowledge graphs — Walker: 21%21%Domain-specific ontologies / knowledge graphs — Crawler: 24%24%System and product telemetrySystem and product telemetry — Total: 15%15%System and product telemetry — Runner: 17%17%System and product telemetry — Jogger: 15%15%System and product telemetry — Walker: 12%12%System and product telemetry — Crawler: 24%24%High-quality human annotations / labelsHigh-quality human annotations / labels — Total: 10%10%High-quality human annotations / labels — Runner: 9%9%High-quality human annotations / labels — Jogger: 15%15%High-quality human annotations / labels — Walker: 8%8%High-quality human annotations / labels — Crawler: 0%0%No clear defensible data advantageNo clear defensible data advantage — Total: 9%9%No clear defensible data advantage — Runner: 6%6%No clear defensible data advantage — Jogger: 7%7%No clear defensible data advantage — Walker: 13%13%No clear defensible data advantage — Crawler: 12%12%Unique third-party / partner data assetsUnique third-party / partner data assets — Total: 5%5%Unique third-party / partner data assets — Runner: 5%5%Unique third-party / partner data assets — Jogger: 4%4%Unique third-party / partner data assets — Walker: 8%8%Unique third-party / partner data assets — Crawler: 0%0%
TotalRunnerJoggerWalkerCrawler

Technical Leadership · n=252 · Runner n=66 / Jogger n=94 / Walker n=75 / Crawler n=17 · Wave 3

Benchmark 2.13

Barriers to production scale for AI

Survey question

Which of the following is the #1 biggest barrier or challenge your organization faces in trusting, scaling, and productizing agentic AI?

Audience
Technical Decision Makers · n=252
Top barrier to scaling agentic AI, Wave 3

29%

name inconsistent or unreliable performance as the #1 barrier to trusting, scaling and productizing agentic AI — 14 points higher than the second most cited barrier, regulatory and compliance challenges

Technical Leadership · n=252 · Wave 3

Data for Top barrier to scaling agentic AI, Wave 3
Category% ranked #1
Inconsistent or unreliable performance29%
Regulatory and compliance challenges15%
Responsible / trustworthy AI concerns13%
Privacy and data security issues11%
Lack of explainability / interpretability8%
Lack of internal expertise or AI talent6%
Limited control and governance over AI actions4%
Ambiguous business case or unclear ROI4%
High cost of scaling, training, deploying3%
Lack of standardized frameworks2%
Legacy systems make integration hard2%
Resistance from employees or leadership1%
Job displacement and workforce impact1%
0%10%20%30%40%Inconsistent or unreliable performanceInconsistent or unreliable performance — % ranked #1: 29%29%Regulatory and compliance challengesRegulatory and compliance challenges — % ranked #1: 15%15%Responsible / trustworthy AI concernsResponsible / trustworthy AI concerns — % ranked #1: 13%13%Privacy and data security issuesPrivacy and data security issues — % ranked #1: 11%11%Lack of explainability / interpretabilityLack of explainability / interpretability — % ranked #1: 8%8%Lack of internal expertise or AI talentLack of internal expertise or AI talent — % ranked #1: 6%6%Limited control and governance over AIactionsLimited control and governance over AI actions — % ranked #1: 4%4%Ambiguous business case or unclear ROIAmbiguous business case or unclear ROI — % ranked #1: 4%4%High cost of scaling, training, deployingHigh cost of scaling, training, deploying — % ranked #1: 3%3%Lack of standardized frameworksLack of standardized frameworks — % ranked #1: 2%2%Legacy systems make integration hardLegacy systems make integration hard — % ranked #1: 2%2%Resistance from employees or leadershipResistance from employees or leadership — % ranked #1: 1%1%Job displacement and workforce impactJob displacement and workforce impact — % ranked #1: 1%1%

Technical Leadership · n=252 · Wave 3

Benchmark 2.14

AI model employment depth

Survey question

Which of the following AI model types are your team and your organization currently employing, and at what level?

Audience
Technical Decision Makers · n=252
In depth
Ch 3 — What Makes a Runner
AI model employment, any level vs advanced or fundamental

53%

employ large language models at an advanced or fundamental level — the only model type a majority has integrated deeply. Large reasoning models are next at 18%, 35 points behind

Technical Leadership · n=252 · Wave 3 · fundamental integration = the technique is fundamental to the business and integrated into core operations · advanced implementation = actively implementing the technique across multiple areas to enhance products and services

Data for AI model employment, any level vs advanced or fundamental
CategoryAny levelAdvanced or fundamental
Large language models (LLMs)84%53%
Multimodal models51%27%
Large reasoning models (LRMs)44%18%
Open-source / open-weight models41%6%
Voice models25%6%
Model routers17%10%
Diffusion LLMs16%15%
Vision foundation models (VFMs)14%4%
Vision-language-action (VLA) models14%2%
World models9%5%
Video foundation models (ViFMs)9%3%
0%20%40%60%80%100%Large language models (LLMs)Large language models (LLMs) — Any level: 84%84%Large language models (LLMs) — Advanced or fundamental: 53%53%Multimodal modelsMultimodal models — Any level: 51%51%Multimodal models — Advanced or fundamental: 27%27%Large reasoning models (LRMs)Large reasoning models (LRMs) — Any level: 44%44%Large reasoning models (LRMs) — Advanced or fundamental: 18%18%Open-source / open-weight modelsOpen-source / open-weight models — Any level: 41%41%Open-source / open-weight models — Advanced or fundamental: 6%6%Voice modelsVoice models — Any level: 25%25%Voice models — Advanced or fundamental: 6%6%Model routersModel routers — Any level: 17%17%Model routers — Advanced or fundamental: 10%10%Diffusion LLMsDiffusion LLMs — Any level: 16%16%Diffusion LLMs — Advanced or fundamental: 15%15%Vision foundation models (VFMs)Vision foundation models (VFMs) — Any level: 14%14%Vision foundation models (VFMs) — Advanced or fundamental: 4%4%Vision-language-action (VLA) modelsVision-language-action (VLA) models — Any level: 14%14%Vision-language-action (VLA) models — Advanced or fundamental: 2%2%World modelsWorld models — Any level: 9%9%World models — Advanced or fundamental: 5%5%Video foundation models (ViFMs)Video foundation models (ViFMs) — Any level: 9%9%Video foundation models (ViFMs) — Advanced or fundamental: 3%3%
Any levelAdvanced or fundamental

Technical Leadership · n=252 · Wave 3 · fundamental integration = the technique is fundamental to the business and integrated into core operations · advanced implementation = actively implementing the technique across multiple areas to enhance products and services

Model type definitions. **Large language models (LLMs)** — core text models used for summarization, code generation and general reasoning (e.g. GPT-4o, Claude 3.5, Gemini). **Multimodal models** — general-purpose models that process text, images and audio interchangeably in one system (e.g. GPT-4o, Gemini 1.5). **Large reasoning models (LRMs)** — models that use a deliberate “thinking” phase to solve complex, multi-step problems (e.g. OpenAI o1, DeepSeek-R1). **Open-source / open-weight models** — publicly accessible models that allow for self-hosting and private fine-tuning (e.g. Llama 3, Qwen, Mistral). **Voice models** — “end-to-end” systems designed for real-time speech recognition and lifelike audio generation (e.g. Whisper, Deepgram). **Model routers** — traffic controllers that select the best model based on cost, quality and latency (e.g. OpenRouter, Martian, TrueFoundry). **Diffusion LLMs** — an emerging class that generates large chunks of text simultaneously via iterative denoising (e.g. Gemini Diffusion, Mercury). **Vision foundation models (VFMs)** — systems built for visual analysis, including object detection and image segmentation (e.g. Meta SAM 2, Florence-2). **Vision-language-action (VLA) models** — models that translate visual scenes and text instructions directly into robotic actions (e.g. Google RT-2, π0.5). **World models** — predictive simulations that understand physical laws and causality to guide agent planning (e.g. Meta V-JEPA, World Labs). **Video foundation models (ViFMs)** — models capable of analyzing events over time or generating realistic video clips (e.g. Sora, Veo, Runway).

Benchmark 2.15

AI resourcing status, technical teams

Survey question

Which of the following AI-related roles does your organization currently have in-house, and which are you hiring or outsourcing in the next 12 months?

Audience
Technical Decision Makers · n=252
In depth
Ch 8 — The Human Side of the AI Stack
AI role resourcing, Wave 3

62%

have an in-house ML/AI engineer — the highest of the nine AI roles evaluated, while model evaluation and red-teaming is unresourced at 60%

Technical Leadership · n=252 · Wave 3

Data for AI role resourcing, Wave 3
CategoryIn-houseOutsourcedNot resourced yet
Machine learning engineer / AI engineer62%12%27%
Data quality lead / data governance lead52%13%35%
AI product manager51%7%42%
AI solutions architect50%10%40%
MLOps / LLMOps engineer46%15%39%
Vibe coder45%2%53%
Prompt engineer / conversation designer42%9%48%
AI security / AI risk & compliance lead37%16%47%
Model evaluator / red team23%17%60%
0%20%40%60%Machine learning engineer / AI engineerMachine learning engineer / AI engineer — In-house: 62%62%Machine learning engineer / AI engineer — Outsourced: 12%12%Machine learning engineer / AI engineer — Not resourced yet: 27%27%Data quality lead / data governance leadData quality lead / data governance lead — In-house: 52%52%Data quality lead / data governance lead — Outsourced: 13%13%Data quality lead / data governance lead — Not resourced yet: 35%35%AI product managerAI product manager — In-house: 51%51%AI product manager — Outsourced: 7%7%AI product manager — Not resourced yet: 42%42%AI solutions architectAI solutions architect — In-house: 50%50%AI solutions architect — Outsourced: 10%10%AI solutions architect — Not resourced yet: 40%40%MLOps / LLMOps engineerMLOps / LLMOps engineer — In-house: 46%46%MLOps / LLMOps engineer — Outsourced: 15%15%MLOps / LLMOps engineer — Not resourced yet: 39%39%Vibe coderVibe coder — In-house: 45%45%Vibe coder — Outsourced: 2%2%Vibe coder — Not resourced yet: 53%53%Prompt engineer / conversation designerPrompt engineer / conversation designer — In-house: 42%42%Prompt engineer / conversation designer — Outsourced: 9%9%Prompt engineer / conversation designer — Not resourced yet: 48%48%AI security / AI risk & compliance leadAI security / AI risk & compliance lead — In-house: 37%37%AI security / AI risk & compliance lead — Outsourced: 16%16%AI security / AI risk & compliance lead — Not resourced yet: 47%47%Model evaluator / red teamModel evaluator / red team — In-house: 23%23%Model evaluator / red team — Outsourced: 17%17%Model evaluator / red team — Not resourced yet: 60%60%
In-houseOutsourcedNot resourced yet

Technical Leadership · n=252 · Wave 3

Benchmark 2.16

AI impact on company metrics

Survey question

Which of the following best describes the impact of AI on your company’s metrics in the past 12 months?

Audience
Technical Decision Makers · n=252
In depth
Ch 6 — AI Meet P&L
AI impact on company metrics, Wave 3

3–9%

is the full range of negative impact reported across all six company KPIs measured — very low against the positive impacts reported from AI

Technical Leadership · n=252 · Wave 3 · for CAC a decrease is the positive direction

Data for AI impact on company metrics, Wave 3
CategoryPositiveNeutral / don't knowNegative
Cost savings66%25%9%
Customer satisfaction62%35%3%
Revenue53%43%4%
Lifetime value (LTV)48%49%3%
Customer acquisition cost (CAC)42%50%8%
Net retention42%52%6%
0%20%40%60%80%Cost savingsCost savings — Positive: 66%66%Cost savings — Neutral / don't know: 25%25%Cost savings — Negative: 9%9%Customer satisfactionCustomer satisfaction — Positive: 62%62%Customer satisfaction — Neutral / don't know: 35%35%Customer satisfaction — Negative: 3%3%RevenueRevenue — Positive: 53%53%Revenue — Neutral / don't know: 43%43%Revenue — Negative: 4%4%Lifetime value (LTV)Lifetime value (LTV) — Positive: 48%48%Lifetime value (LTV) — Neutral / don't know: 49%49%Lifetime value (LTV) — Negative: 3%3%Customer acquisition cost (CAC)Customer acquisition cost (CAC) — Positive: 42%42%Customer acquisition cost (CAC) — Neutral / don't know: 50%50%Customer acquisition cost (CAC) — Negative: 8%8%Net retentionNet retention — Positive: 42%42%Net retention — Neutral / don't know: 52%52%Net retention — Negative: 6%6%
PositiveNeutral / don't knowNegative

Technical Leadership · n=252 · Wave 3 · for CAC a decrease is the positive direction