From Pilot Graveyard to Production Lane: Innovation Gates That Ship
Production-lane gates grounded in McKinsey 2025 AI maturity data (62% pre-scale), AIDev agent PR rejection rates, and eval sovereignty research—not composite vignettes.
Executive Summary
Who this is for: Innovation leads, platform directors, and finance partners stuck in pilot sprawl who need defensible gates—not another demo.
Most enterprises are not failing to start AI programs—they are failing to scale them. McKinsey’s 2025 global AI survey finds 88% of organizations use AI in at least one business function, yet at the enterprise level 32% are still experimenting, 30% are piloting, and only 31% have begun scaling—with 7% reporting AI fully deployed and integrated [1]. That 62% pre-scale cohort is the pilot graveyard in survey data, not anecdote.
The agent layer repeats the pattern. 62% of respondents say their organizations are at least experimenting with AI agents; 23% report scaling an agentic system somewhere in the enterprise—but in any single business function, no more than 10% say agents have moved beyond piloting [1]. Production lanes need gates—owner, workflow ID, unit economics, eval, rollback, retirement—not more demos.
Parallel evidence from software engineering quantifies wasted capacity when “almost production” work never ships. On the AIDev dataset, 46.41% of agent-generated fix pull requests from Copilot, Devin, Cursor, and Claude are rejected after review, CI, and validation [2]. Rejected fixes carry a median code churn of 81–293 lines—review effort spent on work that does not merge.
Cross-domain lens: Manufacturing Andon cords stop the line on defect signals; innovation teams need the same signal for pilots that cannot show unit economics or an owner.
The Challenge
The McKinsey data describes a structural mismatch: adoption is nearly universal (88% use AI somewhere), but enterprise-wide scaling is not. Larger organizations scale faster—nearly half of respondents from firms with more than $5 billion in revenue report reaching the scaling phase, versus 29% at companies below $100 million (Singla et al., 2025)—but even at the top end, full deployment remains rare (7%).
For agentic systems specifically, curiosity outruns production. 39% of respondents say they have begun experimenting with agents; 23% are scaling at least one agentic system—but scaling is narrow. Most organizations scaling agents do so in only one or two functions (Singla et al., 2025). IT, knowledge management, and marketing lead; everything else lags. Demos and experiments are rewarded; production discipline is not.
Engineering teams see the same waste in measurable form. Abujadallah et al. (2026) analyze 306 rejected agentic fix PRs (95% confidence, 5% margin of error) from 1,497 rejections across 2,807 GitHub projects in AIDev. They identify 14 rejection sub-themes in four categories: relevance (e.g., inactivity 17.3%, superseded 5.9%), implementation (incorrect fix 5.6%, wrong approach 2.6%), technical (CI failure 6.9%), and provider failure (7.5%). Enterprise AI pilots fail under different labels but the same economics: capacity spent, nothing shipped.
Research baseline
Source: McKinsey Global Survey on AI (2025) · Sample / scope: Enterprise respondents · Finding: 32% experimenting · 30% piloting · 31% scaling · 7% fully deployed · Operator read: 62% pre-scale at enterprise level
Source: McKinsey (2025) — agents · Sample / scope: Same survey · Finding: 62% experimenting with agents · 23% scaling somewhere · ≤10% scaling per function · Operator read: Agent pilot graveyard is function-local
Source: McKinsey (2025) — adoption · Sample / scope: Same survey · Finding: 88% use AI in ≥1 function (up from 78% in 2024) · Operator read: Adoption ≠ production
Source: AIDev / MSR ’26 (Abujadallah et al., 2026) · Sample / scope: 3,225 agent fix PRs; 1,497 rejected · Finding: 46.41% rejection rate · Operator read: Nearly half of agent fix attempts discarded
Source: AIDev / MSR ’26 · Sample / scope: 306 rejected PR sample · Finding: Median churn 81–293 lines; median 1–4.5 comments per rejection · Operator read: Review cost on non-shipping work
Source: Instructions-as-Code (Arabat & Sayagh, 2026) · Sample / scope: 15,549 agentic PRs · 148 projects · Finding: 27.7% projects ↑ merge rate ≥20% after instruction files; 26.35% ↓ ≥20% · Operator read: Documentation alone does not gate production
Source: Evaluation sovereignty (arxiv 2606.13436) · Sample / scope: Hierarchical metadata classification · Finding: Micro-F1 ~0.54 under operational labels vs ~0.03 under independent gold eval · Operator read: Demo metrics can collapse under audit
Figure 1: Where enterprises sit on the AI maturity curve
Survey-weighted enterprise stages—experimentation and piloting still dominate scaling (Singla et al., 2025).
The Approach
The production-lane checklist below translates survey findings into enforceable gates. It does not replace the five-gate QBR framework in Agent Sandbox to Production—it operationalizes when a pilot earns capacity.
Gate: Owner · Evidence required: Named P&L plus engineering owner · Research anchor: McKinsey: scaling correlates with workflow redesign, not bolt-on pilots
Gate: Workflow ID · Evidence required: Tagged telemetry, chargeback or showback row · Research anchor: FinOps allocation; unattributed spend matches pre-scale sprawl
Gate: Unit economics · Evidence required: Cost per decision or task with ceiling · Research anchor: Required before agent spend scales beyond one function
Gate: Eval · Evidence required: Revenue- or safety-weighted golden set in CI · Research anchor: Eval sovereignty: operational vs independent label regimes
Gate: Runback · Evidence required: Rollback tested; kill switch demonstrated · Research anchor: AIDev: 6.9% of rejections tied to CI failure—gates must be executable
Gate: Retirement · Evidence required: Kill date if not promoted · Research anchor: Counter 17.3% inactivity rejections in agentic PRs
Enforcement mechanics: Pilots failing any gate within 60 days enter graveyard review—archive, merge into an existing lane, or kill with published rationale. Duplicate pilots on the same workflow merge monthly before promotion. Innovation fund tranches require checklist score; kills are recorded in council minutes with hours and budget returned.
Figure 2: Agent curiosity vs function-level scaling
Most organizations experiment with agents; few scale them in any one function (Singla et al., 2025).
The Analysis
Documentation is not production readiness
Arabat and Sayagh (2026) compare 15,549 agentic pull requests across 148 projects before and after instruction-file creation. 27.7% of projects increased merge rate by at least 20%; 26.35% decreased by at least 20%—nearly symmetric. Projects that improved tended to have longer, better-structured instruction files (median 976 vs 569 words), but adding files alone is insufficient. Pilots with wiki checklists but no telemetry, eval, or chargeback rows are the enterprise analog: Instructions-as-Code without gates is checklist theater.
Eval sovereignty: slide-deck metrics vs production truth
Evaluation sovereignty research shows models scoring well under operational “silver” labels can collapse under independent “gold” evaluation—Micro-F1 falling from approximately 0.54 to 0.03 in one hierarchical classification setting (arxiv 2606.13436). Pilot eval gates must require golden sets tied to production failure classes, not demo accuracy on team-controlled datasets. See The Eval That Pays Rent for eval economics; here the checklist treats eval as pass/fail.
Rejection economics mirror pilot graveyards
Abujadallah et al. (2026) report that rejected agentic fixes waste tokens, compute, premium queries, and human review—median churn up to 293 lines for technical-issue rejections. 17.3% of rejections are inactivity closures—parallel to pilots that linger without owner or kill date. Production lane promotion without chargeback is still a pilot with better branding. Graveyard archives should list killed workflow IDs so teams cannot resurrect the same demo under a new name.
Key Takeaways
Metric (McKinsey): track enterprise stage—target movement from piloting (30% cohort) into scaling (31%), not pilot count alone.
Metric (gates): percent of pilots with owner, workflow ID, and kill date—target 100% at day 30.
Metric (AIDev analog): track rejected-or-killed experiments with logged review hours—treat inactivity like the 17.3% PR inactivity bucket.
Promoted workflows earn chargeback rows and on-call cards in the promotion week—not next quarter.
What Operators Miss
McKinsey’s 88% adoption figure hides the 62% still experimenting or piloting at enterprise level. Agent surveys show the same gap: 62% experimenting, ≤10% scaling per function. Innovation councils should publish graveyard reviews with sourced baselines—not vanity pilot counts. Spot-audit checklist claims with production traces; box-checking without working eval recreates the graveyard with better documentation.
Canonical scope
Archive note (June 2026): Differentiated sibling in the production gates cluster. Research-backed production-lane gates grounded in McKinsey 2025 maturity data and AIDev rejection economics; extends Agent Sandbox to Production without re-stating the five QBR gates.
Methodology & limitations
This case log synthesizes published survey and repository-mining data—McKinsey Global Survey on AI (2025), AIDev/MSR ’26 papers, and evaluation-sovereignty research. Figures plot reported percentages from those primary sources (see chart footnotes). This is editorial analysis for operators, not legal, financial, or investment advice. McKinsey figures are survey-weighted enterprise self-report; AIDev statistics describe open-source agentic PRs and may not match your internal tooling mix. Validate gates and baselines in your organization before production decisions.
References
Singla, R., Chui, M., & Hall, B. (2025). The state of AI: Global survey 2025. McKinsey & Company, QuantumBlack AI. https://www.mckinsey.com/capabilities/quantumblack/our-insights/the-state-of-ai
Abujadallah, M., Arabat, A., & Sayagh, M. (2026). Understanding the rejection of fixes generated by agentic pull requests—insights from the AIDev dataset. Proceedings of MSR ’26. arXiv:2606.13468. http://arxiv.org/abs/2606.13468v1
Arabat, A., & Sayagh, M. (2026). Toward Instructions-as-Code: Understanding the impact of instruction files on agentic pull requests. MSR ’26. arXiv:2606.13449. http://arxiv.org/abs/2606.13449v1
Evaluation sovereignty in metadata-driven classification. arXiv:2606.13436. http://arxiv.org/abs/2606.13436v1
Li, X., et al. (2025). AIDev dataset (cited in [2])—33,596 curated agentic PRs across 2,807 projects.
The AI Operator. (2025). Agent Sandbox to Production — five QBR-grade production gates. https://theaioperator.com/articles/agent-sandbox-production-gates
The AI Operator. (2025). The Eval That Pays Rent — eval economics and golden-set gates. https://theaioperator.com/articles/eval-pays-rent
The AI Operator. (2025). The Chargeback Imperative — P&L attribution companion.
National Institute of Standards and Technology. (2023). AI Risk Management Framework (AI RMF 1.0). https://www.nist.gov/itl/ai-risk-management-framework
Monday Morning Checklist
[ ] Pull your org’s McKinsey-class baseline: count active pilots vs production lanes with owner + workflow ID.
[ ] Require owner, workflow ID, and kill date on every pilot by day 30—target 100% coverage.
[ ] Schedule graveyard review for pilots missing unit economics or eval gates; publish kills in council minutes.
[ ] Add top five production failures to a golden set before the next model or agent release [7].
[ ] Tag agent and ML spend with
workflow_id; tie promoted lanes to chargeback rows in the promotion week [8].[ ] Walk Figure 1 with innovation council: are you measuring stage movement, not pilot count alone?
Editorial transparency. Essays at The AI Operator may use AI-assisted research, drafting, and editing tools under staff editorial review. Facts, figures, and recommendations are checked before publication; we correct the record when evidence changes. Questions: hello@theaioperator.net.
Published on [Substack](https://theaioperator2.substack.com/p/pilot-graveyard-to-production).




