The demonstration lasts forty minutes. An agent reads a loan file, pulls the covenant schedule, checks the borrower's latest financials against policy and drafts the credit memo. The pilot is approved. Nine months later the agent is still in a sandbox, a credit officer rewrites every memo it drafts, and the project has been renamed a learning initiative. In financial services this sequence is now common enough to have a shape.
The 2026 numbers describe the stall precisely. In the financial services cut of Deloitte's enterprise survey, 573 leaders surveyed in August and September 2025 and published in March, only 24 percent said their organization had moved 40 percent or more of its AI experiments into production, while 53 percent expected to reach that level within three to six months. Twenty-one percent were using agentic AI at least moderately and 71 percent expected to within two years.1 KPMG's second-quarter pulse of 204 US banking leaders found 51 percent piloting agents, 24 percent scaling them across several functions, 15 percent orchestrating several agents across workflows and 10 percent still exploring (Exhibit 1).2 In the Cambridge Centre for Alternative Finance's global survey of 628 organizations, 52 percent of industry respondents were piloting agentic AI or beyond, and 55 percent, rising to 76 percent of large financial institutions, found the value of AI deployment difficult to measure.3
The public record says the same. Evident's tracker of the 50 banks in its index found agentic applications at 31 percent of new use cases in the first quarter of 2026, up from 15 percent a quarter earlier.4 In the second quarter the banks announced 93 new use cases, and only 27 percent came with a disclosed outcome, down from 41 percent.5 Across the index, 12 percent of use cases report an effect on operational KPIs and barely 1 percent disclose a financial return.6 In McKinsey's August survey, large organizations scaling agents rose from 27 to 40 percent in a year, while those attributing any EBIT effect to AI stayed at 37 percent.7 Pilots are multiplying faster than production.
Exhibit 1

Our argument is that the stall has one cause, and most programs are not chasing it. The demo runs on public knowledge and a curated file. Production requires the institution's own definitions, access policy, precedent and decision rights, and in almost every firm those exist only in people and in prose. What a senior person supplied from the back of the room during the demo, nobody supplies in production. This paper sets out the evidence and a decision framework for the CIO and COO who must decide which pilots to fund, which to stop, and what to build first.
The demo and the desk are different tests
The demo tests whether an agent can perform the steps; the desk tests whether it performs them the way this institution does, every time. On the Finance Agent Benchmark v2, updated on October 7 with 76 models scored against 927 expert-reviewed questions, Gemini 4 Argon leads on partial credit at 65.4 percent; Claude Fable 5.1, Claude Opus 5.5 and Claude Sonnet 5.5 sit between 58.1 and 58.9 percent; and the open-weight GLM 5.3 Flash, Mistral Large 4, DeepSeek V4.1 Flash, Kimi K3 and Qwen 3.8 Max score between 50.6 and 57.9 percent, with GPT-5.6 Luna and GPT-6 Astra among them. On the stricter all-pass score, which credits a question only when every check is met, the best model reaches 50.9 percent. The benchmark's authors note that models "still struggle to perform reliably on harder, multi-step financial work."8 Stanford's AI Index, published in April, found the top three model providers within 9 points of each other on the Arena leaderboard.9 Choosing the model is no longer the decision that matters.
BigFinanceBench, published in June with 928 tasks written by 52 financial-research professionals, locates the failure. The three best agents at publication, Claude Opus 4.7, GPT-5.5 and Claude Sonnet 4.6, scored 58.8, 58.8 and 58.5 percent on the expert rubric and 41.5, 44.3 and 38.4 percent on the final answer (Exhibit 2). When the authors isolated trajectories in which the setup was already right, calculation scores across nine models fell within 3.6 points of each other, between 84.1 and 87.6 percent. Their conclusion is that "retrieval and setup dominate residual failures once arithmetic is instrumented."10 An agent that knows which definition and adjustment the house uses gets the arithmetic right. One that guesses gets the memo rewritten.
Exhibit 2

The pattern holds on enterprise tasks. EnterpriseRAG, an August benchmark of 13 models under retrieval noise, knowledge gaps and factual conflicts, found that models satisfy 80 percent of individual constraints but meet all of a task's requirements at once in only 26.8 percent of responses.11 AlphaEval, built from 94 real tasks supplied by seven companies, scored the best configuration at 64.41 out of 100 and named a production failure mode in which "agents optimize explicitly stated objectives while violating implicit constraints."12 The implicit constraints are the institution's. Nobody wrote them down, so the agent could not read them.
What the demo borrowed and production has to be given
In KPMG's banking pulse, the top barriers to deploying agents were data readiness and access at 63 percent, the complexity of agentic systems at 49 percent and "human-in-the-loop judgment and escalation skills" at 41 percent. Accountability for AI-informed or AI-executed decisions sat with a named C-suite executive at 49 percent of banks, the CEO or executive committee at 34 percent, a business unit leader at 12 percent and a centralized AI governance committee at 3 percent.2 KPMG's third-quarter pulse of 314 US leaders across industries again found data readiness and access the top barrier, even as the share building multi-agent systems climbed to 25 percent from 6 percent.13 Deloitte's financial services respondents named governance capabilities and oversight as a top AI risk at 48 percent.1 Among 728 EU securities firms surveyed by ESMA, 82 percent cited at least one data-related challenge, and only 426 of 847 reported use cases were in production.14
The override question in the KPMG survey is the most revealing. Asked what triggers a person to overrule an AI output, 52 percent of banks named a conflict with regulatory or legal requirements, 47 percent a breach of predefined risk thresholds and 40 percent a confidence or quality score below a limit. But 24 percent said overrides happen case by case without formal criteria, and only 19 percent recognized a difference between business judgment and the AI's recommendation as a trigger (Exhibit 3).2 That is the stall in one chart: an agent in production has to know, before it acts, which outputs a person must see. In a quarter of banks that rule is unwritten, and in most the trigger experienced staff actually use, their own judgment, is not an official one.
Exhibit 3

The table lists five things a demo borrows without anyone noticing and production must be given explicitly.
Definitions
The demo's terms meant what the presenter meant. In production the agent meets exposure, adjusted EBITDA and material breach as each function defines them, and must apply the one that governs.
Access policy
The demo ran on a curated extract. Production requires rules about who may see what, applied before the query runs.
Precedent
Production requires the exceptions, the grandfathered cases, and last year's committee decision with its reason. Recorded in minutes, not in a form a system can match.
Decision rights
Nothing was at stake in the demo. Production requires who may approve, who must be told, and what the agent may do alone or only propose. Rarely written down.
Escalation and evidence
Production requires the cases that go to a human to be specified in advance, and a record of the definition, policy and authority each action applied.
Data readiness is the barrier executives cite first. Precedent and decision rights are a smaller problem that almost nobody has started, and they are the rows the agent cannot run without. Agent projects stall not at a committee meeting but at the absence of an artifact the committee could approve.
Writing the institution down changes what the model does
Evidence that an explicit layer of definitions changes model behavior now comes from current models. In an April study by the semantic-layer vendor Cube, three frontier models answered the same 99 analytical questions about a retail business twice: once with only the database schema, and once with a short document of business definitions. Claude Opus 4.7 went from 50.5 to 67.7 percent, Claude Sonnet 4.6 from 46.5 to 68.7, and GPT-5.4 from 45.5 to 68.7.15 A July study by a single industry author separated definitions from governance. On a 100-question enterprise analytics benchmark, a model writing SQL directly produced at least one hallucination in 79 percent of answers and violated row-level security in 78 percent. Exact business definitions cut hallucination to 40 percent, yet the model still leaked data across entity boundaries in 35 percent of cases. Only when definitions were combined with an enforced access policy and validation before execution did both rates fall to zero, with strict accuracy rising from zero to 79.3 percent (Exhibit 4). Row-level security, the author concludes, "was not recovered from metric definitions alone; it required an explicit, enforced policy."16
Exhibit 4

dbt Labs' April benchmark on an insurance dataset found the same asymmetry: on modeled data, a governed semantic layer answered 98.2 percent of questions correctly with Claude Sonnet 4.6 and 100 percent with GPT-5.3 Codex, against 90.0 and 84.1 percent for direct text-to-SQL; with text-to-SQL, the authors note, "failure looks like a plausible but incorrect answer."17 The same lesson applies to what an agent says about its own work. In a September study, six models including GPT-5.6 Terra and Claude Sonnet 5 faced 100 tasks in which a tool failed partway, with 3,600 responses annotated by people. With no instruction, 22.8 percent of responses claimed success without evidence and 28.3 percent contained fabricated detail. Given a structured evidence contract, a fixed format that ties any claimed status to its evidence and names the next action, both rates fell to 0.8 percent, and useful responses rose from 74.9 to 98.8 percent.18 The model did not change; the rule did.
Three caveats belong here. Two of these studies are vendor-authored, one has a single author, all use modest question sets, and none is a bank. But the direction is the same in every one, and it matches what practitioners report. In Anthropic's June Economic Index, drawing on about 9,700 surveyed users, people with fifteen or more years of experience rated the share of their work AI can do about ten points lower than first-year workers did, citing "the judgment, contextual awareness, and situational reasoning that their work requires."19 Rule-like work will be done the house's way once the rules are stated; judgment-dependent work must be routed, with evidence, to the person who holds the judgment. Both require the institution to write itself down.
The vendor's agent is not the institution's capability
The common response to a stalled pilot is to buy one built. Gartner predicted on September 29 that by 2028, 70 percent of enterprises will abandon agentic AI built by vendors' forward-deployed engineers, because such engagements "often fail structurally before they fail technically" and customers "fail to build internal capability." Its prescription reads like a governance checklist: contract for knowledge transfer and "transition or exit responsibilities," "define decision rights," and ensure the organization "develops the capabilities, governance, and operational ownership" to run the solution alone. In April the same firm predicted that by 2028 over half of all enterprises will stop paying for assistive AI, and observed that "in the execution era, control of enterprise context is economic power."20
Vendors have noticed. In the Cambridge survey, 28 percent of financial-services-specific AI products operate fully autonomously against 22 percent of general-purpose tools, which the authors attribute to workflow vendors embedding "enough domain logic (regulatory rules, risk thresholds, compliance guardrails) to be trusted with higher autonomy."3 Ninety-one percent of Deloitte's respondents expect to customize agents.1 The Bank of England's February roundtables recorded that substitution between AI providers "may become more challenging."21 The question for the CIO is whether the institution's rules live in a vendor's product or in an asset the institution owns.
Nor are the economics the obstacle. The IMF's April note on AI records inference prices for certain frontier models dropping by over 99 percent and concludes that the transition from capability to impact "is constrained primarily by institutional and organizational frictions."22 On the Finance Agent Benchmark, a test run costs $0.05 with GLM 5.3 Flash and $9.22 with Claude Opus 5.5, for scores within a point of each other.8 In the FinOps Foundation's 2026 survey of 1,192 practitioners, 98 percent now manage AI spend, up from 63 percent in 2025 and 31 percent in 2024.23 Only 31 percent of KPMG's banks say AI operating costs are fully visible, and about one in five of McKinsey's respondents say operating costs, including tokens, have constrained their AI use.2,7 A pilot with no cost model is as hard to approve as one with no decision rights.
Supervisors are describing the same missing object
Regulators spent 2026 saying their frameworks do not yet cover agents, then describing what they will ask for anyway. On April 17 the Federal Reserve, OCC and FDIC issued SR 26-2, revised model risk guidance, and stated in a footnote that "generative AI and agentic AI models are novel and rapidly evolving. As such, they are not within the scope of this guidance," leaving a bank's own "risk management and governance practices" to determine the controls.24 The Bank of England's roundtables recorded that "the concept of having a 'human-in-the-loop' was also challenged by the rise of agentic AI."21 In February the ECB's Pedro Machado said more than 85 percent of large supervised banks already use AI in some form, flagged "fragmented ownership" across IT, data science, business lines and control functions as a supervisory concern, and said that if "a bank cannot explain why an AI model behaves the way it does," then "it cannot truly control that model."25
The Financial Stability Board's twelve sound practices, consulted on in June with the final report due this month, describe inventories that record for each agent "the tools the agent can access, its components, and guardrails to manage risk," and in some cases an agent "certification," in which it is "reviewed and approved for use within defined boundaries." Effective oversight, the FSB says, is "meaningful, i.e. it involves humans having sufficient ability, authority, and incentive to intervene, as opposed to oversight that is nominal."26 The Monetary Authority of Singapore's October 7 guidelines expect an inventory, a risk materiality assessment for each use case and accountability for third-party AI from October 2027, with an agentic AI consultation to follow in 2027.27 These are requests for the artifact the pilot never produced: which definitions the agent applied, under whose authority, within what boundaries, and which cases a person must take.
What the people shaping the next four years expect
Forecasts from people with a mandate at stake converged in 2026.
- Bought agents will not stick. Gartner's forecast that 70 percent of enterprises will abandon vendor-built agentic AI by 2028 rests on the cause behind the stalled pilot: customers "fail to build internal capability."20
- Adoption will run ahead of the organization. Jamie Dimon's April letter to JPMorganChase shareholders said the pace of adoption "will likely be far faster than prior technological transformations, like electricity or the internet."28 In Accenture's January banking outlook, 57 percent of banking executives expected agents to be fully embedded in risk, compliance, audit, fraud and transaction monitoring within three years, and 56 percent expected broad adoption in credit assessment, loan processing and KYC.29 These functions carry the most decision rights and the least written down.
- The capital will not wait for the evidence. The BIS General Manager, Pablo Hernández de Cos, said in September that global AI-related investment is expected to rise "from around $500 billion today to between $3 trillion and $4 trillion by 2030," against a median productivity estimate of around half a percentage point a year, and warned that "should the returns to AI disappoint," a pullback "could turn today's capital expenditure boom into a bust."30 For a bank, that return depends on pilots crossing into production.
- The rules will matter more than the model. The IMF's scenario-planning exercise concluded that even as the frontier advances, economic gains "will be delayed or concentrated among organizations with higher readiness," because adoption is limited "largely by regulatory uncertainty, compliance burdens, organizational inertia, and trust issues."22 In the Cambridge survey, 81 percent of industry respondents expect agentic AI to be "meaningfully achieved" by 2030, yet the accountability framework for these systems "remains unresolved to some extent."3
Our own expectation follows from the evidence. The models have already converged: proprietary and open-weight leaders sit within eight points of each other on financial agent work, and a near-equal score can cost less than one percent as much. The institutions that reach production by 2028 will be those that made their definitions, access policy, precedent and decision rights explicit enough for an agent to run on and a supervisor to inspect, and own that layer themselves. We expect the first supervisory findings on agentic AI, in 2027 and 2028, to turn on whether an institution can state which rule an agent applied and who authorized the action. By 2030 we expect boards to ask for the pilot-to-production rate, and the decision-rights register to be as ordinary as the model inventory is today.
A decision framework for the CIO and COO: four moves
The evidence points to four moves, in roughly this order. Together they are also a test for every pilot in the portfolio.
1. Run the demo-to-desk test before funding the pilot
For each candidate, list what the demo borrowed: the definitions it assumed, the access policy it bypassed, the precedent it skipped and the decision rights it did not need. Then ask whether each exists in a form a system can read. If none does, it is a knowledge project first and an agent project second, and should be budgeted that way; otherwise the 53 percent of Deloitte's respondents expecting to move 40 percent of experiments into production within six months will discover this slowly.1
- Concrete marker: Every pilot's charter names the decisions it touches, the owner of each, and whether the agent prepares, proposes or acts.
- Concrete marker: No pilot proceeds to production design until the definitions and evidence standard for its decisions exist in writing and have an owner.
2. Build the institution's context as an asset no vendor owns
Once an agent has the right setup, the models calculate within a few points of each other.10 So the institution's definitions, with scope, effective date and owner, and its policy, precedent and access rules, should be the context every agent reads, held outside any one agent, model or vendor. That is what Gartner means by control of enterprise context.20
- Concrete marker: Definitions, policy, precedent and access rules live in a governed layer that any agent and any model can read, not in one agent's prompt or one vendor's product.
- Concrete marker: Vendor contracts specify knowledge transfer, ownership of the encoded rules, and an exit in which the institution's context stays behind.
3. Write decision rights and escalation before the first production run
A quarter of banks override agents case by case with no formal criteria, and only a fifth treat a difference with business judgment as a trigger.2 Reverse that. Decide in advance which cases go to a person, by materiality, confidence, novelty against precedent and conflict with policy. Require the agent to attach evidence and say what it could not do; a structured evidence contract cut unsupported claims of success from 22.8 to 0.8 percent with no change to the model.18 The FSB's test is the right one: oversight counts only if the person has the "ability, authority, and incentive to intervene."26
- Concrete marker: Each agent's escalation rules are written, versioned and tested before go-live, and its response format makes a missing escalation visible.
- Concrete marker: Overrides are logged with the reason and reviewed quarterly; recurring reasons become new rules, definitions or precedent.
4. Run the agent estate through a register, and report the rate that matters
Accountability for AI decisions is split four ways across banks, and sits with a central governance committee at only 3 percent.2 What works is narrower than a committee: a named owner for each decision and for the knowledge layer, an agent register recording tools, components and guardrails as the FSB describes, and change control borrowed from model risk management.26 Add the cost model, since a fifth of organizations already say operating costs have constrained their AI use.7 Then report to the board how many pilots reached production and, for those that did not, which row of the table was missing.
- Concrete marker: Every production agent is registered with owner, decisions touched, permitted actions, escalation rules, unit cost and review date.
- Concrete marker: Board reporting shows the pilot-to-production rate and each agent's override rate, not the number of pilots.
The leadership test
Leaders can locate their institution by asking six questions:
- Of the AI pilots we have run, how many are in production, and for those that are not, can we name the definition, policy, precedent or decision right that was missing?
- For the next agent we build, could we list today the decisions it touches, who owns each, and whether the agent prepares, proposes or acts?
- If two functions define the term the agent is using differently, which definition does it apply, and who decided that?
- Which cases must a person see before the agent acts, is that written down, and does the agent know?
- If a supervisor asked which policy and authority an agent applied last Tuesday, would we answer from the system or from an interview?
- If we replaced the model or the agent vendor next year, what would we have to rebuild, and who would own what remained?
Institutions that can answer these questions have the layer an agent needs. Those that cannot have a demo, and the next model release will not close the distance to the desk: a stronger model applies the same missing rules faster.
The institutions that capture value from AI will not be those with the most pilots. They will be those that wrote down what their terms mean, who may see what, what was decided before and who may decide now, and governed those answers as they govern models. The agent is the easy part. The institution's rules are the work, and in most firms it has barely begun.