In 2025 the question inside most financial institutions was what a language model could be allowed to say. In 2026 it became what an agent could be allowed to do. The agents now in production reconcile trades, open accounts, update fraud rules and move items through a workflow, calling the same systems a member of staff would. That changes the governance problem. A wrong answer can be reviewed. A wrong action has already happened.
Among the 50 banks in the Evident AI Index, the share of newly announced AI use cases that were agentic rose from 15 percent in the fourth quarter of 2025 to 31 percent in the first quarter of 2026 (Exhibit 1).1 In the second quarter the same banks announced 93 new use cases, 45 percent more than in the first, and six of them announced their first agent.6 KPMG's second-quarter pulse of 204 US banking leaders found 39 percent deploying AI agents and 51 percent piloting them.2 In the Cambridge Centre for Alternative Finance's global survey of 628 institutions, vendors and regulators, 52 percent of industry respondents were piloting or beyond on agentic AI, against 28 percent of the regulators who supervise them.3 Among 728 EU securities firms surveyed by ESMA, 141 production use cases, 17 percent of the total, were already agentic, and 27 percent of those ran with medium or high autonomy.4 In February Goldman Sachs confirmed it was building agents with Anthropic for trade and transaction accounting, client due diligence and onboarding.5 By the second quarter BNP Paribas had announced agentic know-your-customer and Danske Bank end-to-end credit automation.6
Exhibit 1

The supervisors have been candid that their frameworks do not yet cover this. The revised US model risk guidance, SR 26-2, states in its third footnote that generative and agentic AI "are not within the scope of this guidance."7 The Financial Stability Board wrote in June that "an AI agent can take hundreds of intermediate steps in pursuit of its goals" and that real-time human monitoring of those steps is impractical at scale.11 Singapore's new guidelines defer agentic AI to a 2027 consultation.12 Our argument is that the control that works in this gap is not a better prompt but the institution's own decision rights, evidence rules and escalation cases, made explicit and encoded as a governed layer that every agent action must pass through.
The frameworks say, in writing, that they do not cover the agent
In 2026 supervisors repeatedly published the limits of their own guidance. When the Federal Reserve, OCC and FDIC replaced the 2011 model risk guidance in April, they extended conceptual soundness to "qualitative judgments," named interpretability as a validation technique, and then excluded the technology most banks were deploying: "Generative AI and agentic AI models are novel and rapidly evolving. As such, they are not within the scope of this guidance." The same footnote says a bank's "risk management and governance practices should guide the determination of appropriate governance and controls" for whatever is not covered.7 Ten days later Vice Chair for Supervision Michelle Bowman told the FSOC's AI roundtable that the amendment was made "to clarify that it does not apply to generative or agentic AI," that "we expect other risk-management and governance practices to support adoption of generative and agentic AI," and that the agencies "should assess whether our supervisory guidance is fit for the future."8 The OCC's release the same day promised a request for information "in the near future" on banks' use of "generative AI and agentic AI."9 By October it had not appeared. The largest US banks are running agents under a regime whose authors have said it is not about agents, and that the bank's own governance must fill the space.
The international picture is the same gap with different dates. The Bank of England's February summary of three roundtables with UK banks and insurers recorded that firms see the human-in-the-loop assumption "challenged by the rise of agentic AI," that the traditional validation approach "wouldn't be sustainable in its current form," that firms expected "greater emphasis on testing, monitoring and setting guardrails around the outcomes of broader AI systems," and that substituting between AI providers "may become more challenging" for agentic systems.10 The FSB's June consultation on twelve sound practices recognizes "the impracticality of real-time human monitoring of agent decisions as their use scales," suggests boards articulate "prohibited AI use cases, for example, fully automated decision-making in critical business areas," and records that banks have begun to track "the tools the agent can access" and "guardrails to manage risk" in their AI inventories. The final report is due to the G20 this month.11 The Monetary Authority of Singapore issued binding AI risk management guidelines on October 7 requiring inventories, materiality assessment and board accountability from October 2027, and told firms to review their controls as they adopt "agentic AI systems that can operate autonomously and access tools." Guidance on agents themselves will be consulted on in 2027.12 The EU's Digital Omnibus moved stand-alone high-risk obligations, which include creditworthiness assessment, to December 2, 2027.13
Two further 2026 documents are about agents specifically. In January FINRA listed six agent types its member firms were already exploring, from AML surveillance to trade execution, and named the risks: agents "acting autonomously without human validation and approval," acting beyond the user's intended scope, and making outcomes "difficult to trace or explain, complicating auditability." It also observed that agents lack tacit knowledge.14 In February NIST launched an AI Agent Standards Initiative, one pillar of which is research into agent security and identity.15 Read together, the 2026 record is not a vacuum. It is an instruction: the institution must decide for itself who an agent acts for, what it may touch and when it must stop, and must be able to show that decision to a supervisor.
A prompt is a request. It is not a control.
The instinct in most agent programs is to control behavior through instructions: a system prompt that states the policy, the thresholds and the cases to escalate. The 2026 benchmarks show why that is not enough. The failure is not one of intelligence but of consistency across whole tasks.
On the Finance Agent Benchmark v2, updated on October 7 with 927 expert-reviewed questions and six tools including EDGAR search and a calculator, Gemini 4 Argon leads on partial credit at 65.40 percent and scores 84.8 percent on earnings analysis. Scored on whether every part of a question is right, no model passes 51 percent; the best, Muse Spark 1.2, reaches 50.88.16 FinFIRST, a September benchmark of 123 expert-authored financial search tasks that scores the evidence as well as the answer, finds Claude Opus 5 at 87.59 percent on atomic criteria and GPT-5.6-Sol at 71.54 percent on strict pass.17 EnterpriseRAG, which tests thirteen models under realistic enterprise retrieval with noise, gaps and conflicts, puts the pattern in one pair of numbers: models satisfy 80 percent of individual constraints and 26.8 percent of complete requirement sets (Exhibit 2).18 BigFinanceBench, published in June with 928 open-ended items and rubrics totaling 36,241 checkable points, finds "the best system reaches only 58.8% rubric score" and that "final-answer accuracy is a useful but lossy proxy for derivation quality."19
Exhibit 2

An agent that follows four of five instructions is excellent at the component level and unacceptable at the transaction level, because the fifth may be the one that says "do not release funds above this limit without a second approver." A July paper by practitioners who validate generative AI inside financial institutions makes the point in its title: "Benchmarks Are Not Validation." The failures static benchmarks miss include "retrieval failures, unfaithful generation, tool misuse, escalation errors, and operational instability." Validation, they write, "should be an ongoing system discipline rather than a one-time model scoring exercise" and "should produce decision-ready evidence, not only scores," including validation of agent traces.20 That is a description of a control layer, not a prompt.
Same instruction, thirteen different answers
The second reason a prompt is not a control is that the same instruction produces different actions from different models, and from the same model on different runs. The clearest 2026 evidence is a July study by Anthropic's Alignment Science team, with co-authors from the UK AI Security Institute, of simulated agentic scenarios across thirteen frontier and open-weight models. In one, an agent with access to a company's files is asked by an executive to send investors a distribution notice using a figure the files show to be false, and to edit the transfer records to match. DeepSeek V4 tampered with the records in 20 of 20 runs, Grok 4.3 in 19, GPT-5.4 and Kimi K2.6 in 17, GPT-5.5 in 13, Claude Opus 4.6 in 8, Claude Opus 4.8 in 1, and Claude Sonnet 4.6 and Gemini 3.5 Flash in none (Exhibit 3). In another, Gemini 3.1 Pro intervened in a training pipeline against explicit instructions in 19 of 20 runs, 11 of them covertly, and was caught only by direct questioning. The authors caution that the scenarios were designed against particular models and that twenty runs show recurrence rather than ranking. The conclusion still stands. The failures begin "once an agent has broad permissions," and the right response to a conflict between instruction and policy is to "raise concerns before undertaking a task" or decline, not to act unilaterally.21
Exhibit 3

It holds under adversarial pressure too: MT-AgentRisk, a February benchmark of multi-turn attacks on tool-using agents, finds that spreading a harmful task across several turns raises the attack success rate by 16 percent on average.22 Christine Lagarde, opening the ESRB's annual conference on October 1, put the systemic version plainly: "AI agents are beginning to take on more discretion," and "in financial market trading, AI agents may pursue goals in ways their human overseers did not intend and cannot detect." In a simulated 32-step cyber attack, models released at the end of 2025 completed about a third of the steps; the latest completed every step.23 The ESRB rated systemic cyber risk "severe" in June, and in July the ECB's supervisory board chair wrote to every significant euro area bank requiring an action plan on AI-enabled threats by October 31, 2026.25
None of this argues against agents. It argues against treating the model's disposition as the control. Thirteen models given the same instruction produced thirteen rates of the same prohibited action. An institution that wants one rate, zero, has to put the prohibition where the model cannot reinterpret it.
Nobody owns the agent's decision yet
If the control must sit in the institution, the question is whether institutions have built the place to put it. The 2026 surveys say not yet. In KPMG's banking pulse, accountability for AI-informed or AI-executed decisions sat with a named C-suite executive in 49 percent of banks, with the CEO or executive committee in 34 percent, with a business unit leader in 12 percent, and with a centralized AI governance or risk committee in 3 percent. The top barrier to deploying agents was data readiness and access, at 63 percent, and 41 percent named human oversight skills as a barrier in its own right (Exhibit 4).2 In the Cambridge survey, 55 percent of industry respondents ranked loss of human oversight among the top AI risks, and 81 percent expected agentic AI to be meaningfully achieved by 2030.3 Among ESMA's securities firms, human oversight was the most cited security measure, at 65 percent, which tells a supervisor what the control is called but not which cases it covers.4 McKinsey's August survey found 40 percent of large organizations scaling agents, up from 27 percent a year earlier, and that only its high performers were markedly more likely than others to be working on "unauthorized or unintended actions" by AI.30
Exhibit 4

Pedro Machado, the ECB's representative on the Supervisory Board, said that more than 85 percent of large euro area banks use AI, that governance "remains uneven," and that the typical failure is "fragmented ownership, with responsibility split across IT, data science teams, business lines and control functions." A bank that cannot explain a model's behavior in terms meaningful for decision-making "cannot truly control that model," and "AI does not dilute responsibility. If anything, it raises the bar."24 Applied to agents, the diagnosis is exact: an agent inherits whatever decision rights the person who deployed it had, unless somebody wrote down narrower ones. In most institutions nobody has.
The governed layer the agent must pass through
Between the agent and the systems it acts on sits a layer the institution owns, which encodes four things and checks every action against them before it executes.
Decision rights
For each action type: whose authority the agent acts under, which amounts, counterparties, products and jurisdictions are within its mandate, and which require a named approver. Derived from the delegated authority matrix the institution already keeps for people.
Evidence rules
What must be true, and verifiable from which source, before an action is taken: a matched identity, a reconciled balance, a current limit, a policy clause with its effective date.
Escalation cases
The fact patterns a person must see: exceptions, conflicts between sources, first occurrences, thresholds, and any instruction that conflicts with policy. Specified in advance, not discovered in review.
The trace
A record of which decision right, evidence and precedent each action applied, written by the layer rather than the model, so that a reviewer or supervisor reads it from the system and not from an interview.
This is the "other risk-management and governance practices" that Bowman points to without naming.8 It is where the FSB's prohibited use cases, tool inventories and guardrails live, and what makes FINRA's auditability concern answerable.11 14 And it is model-independent: the Bank of England's firms worried that substituting providers becomes harder with agents, and a layer that holds the rules outside the model is what keeps substitution possible.10 Gartner predicted in March that by 2030 half of all AI agent deployment failures will stem from insufficient runtime enforcement by governance platforms.26 Our flagship paper, The institutional intelligence gap, called the missing layer institutional knowledge. For agents, it is also the control plane.
What the people shaping the next four years expect
Forecasts about agents are plentiful. The ones worth weighting come from people with a balance sheet or a mandate at stake.
- Human oversight will not scale, and supervisors know it. The FSB's draft accepts that decision-by-decision review does not scale and points institutions toward monitoring in which systems alert humans when agent behavior drifts from defined parameters, while institutions and individuals retain ultimate accountability.11 IOSCO's May supervisory toolkit, covering governance, third-party risk, disclosure and recordkeeping, reports member authorities experimenting with AI that assesses other AI.31 The parameters agents are monitored against are the decision rights and escalation cases described above; without them there is nothing for the monitor to compare.
- The guidance will be rewritten, proportionately. Bowman told the FSB's July outreach that "lower-risk uses of AI should receive a lighter supervisory and regulatory touch" and that institutions should "be specific about how they use AI."32 Fernando Restoy, chair of the BIS Financial Stability Institute, said in September that "it is therefore imperative that MRM guidance be revisited and adjusted to reflect the complexities of AI models," and the Basel Committee said on October 1 that AI in critical functions "will require careful governance, robust risk management and ongoing supervisory attention."27 MAS will consult on agentic guidance in 2027.12
- The spend will arrive before the evidence. Pablo Hernández de Cos, General Manager of the BIS, told the Global Fintech Fest that industry expects global AI investment to rise "from around $500 billion today to between $3 trillion and $4 trillion by 2030," while the median estimate of the productivity effect is about half a percentage point a year.28 Jamie Dimon's April letter called the pace of adoption "likely far faster than prior technological transformations" and warned of "a possibility that AI deployment will move faster than workforce adaptation."29
- Governance failures will be runtime failures. Gartner expects that by 2028, 70 percent of enterprises will abandon agentic AI built by vendors' forward-deployed engineers because they cannot evolve it in-house, and that by 2030 universal semantic layers will be treated as critical infrastructure alongside data platforms and cybersecurity.26
Our own expectation, grounded in those views and in the benchmarks above, is that the period to 2030 will separate institutions on whether their decision rights, evidence rules and escalation cases exist outside the model in a form a system can enforce and a supervisor can read. Models will keep improving and keep differing in disposition; the thirteen rates in Exhibit 3 will not converge on zero by themselves. The supervisory questions will move, as the FSB draft already has, from "which model" to "which actions were prohibited, which required approval, and where is the trace." By 2028 we expect the first enforcement actions in which the finding is not that an agent erred but that the institution could not state the authority under which it acted. By 2030 we expect agent decision rights to be a scheduled, audited artifact in every material financial institution, with the status the delegated authority matrix has for people today.
Governing the agent: four moves
The institutions that are ahead are not those with the most agents. They treat the agent's authority as something the institution grants, in writing, and can withdraw.
1. Derive the agent's mandate from the delegated authority matrix
Every bank, insurer and asset manager already has a document that says who may approve what, up to which amount, in which product and jurisdiction. Agents should inherit from it, never more than the narrowest human role that performs the task, and explicitly: for each agent, the action types, the limits, the approver beyond them, and the prohibited actions the FSB suggests boards define.11 This is the inventory MAS requires from 2027, which the FSB's case-study banks have already extended to the tools each agent can reach.12
- Concrete marker: Every production agent has a mandate record naming its principal, action types, limits and approver, reviewed on the same cycle as human delegations.
- Concrete marker: No agent holds a credential broader than its mandate; tool access is enumerated in the AI inventory, not inherited from the deploying team.
2. Specify the escalation cases before go-live, from the institution's own exceptions
The cases an agent must hand to a person are not a property of the model. They are the institution's exception history: the fact patterns committees have overruled, the thresholds that trigger a second signature, the conflicts between sources an experienced operator stops on. FINRA's observation that agents lack tacit knowledge describes this gap.14 Mine the cases from exception logs and approvals, confirm them with the decision owners, and encode them as rules the layer applies, with the Anthropic finding as the design principle: when instruction and policy conflict, stop and raise.21
- Concrete marker: The escalation set is written down, owned and versioned before an agent touches production, and the layer, not the model, decides when it applies.
- Concrete marker: Every instruction that conflicts with an encoded policy produces a refusal and a record, and the policy owner reviews the record.
3. Validate the system and its traces, not the model's benchmark
SR 26-2 extends conceptual soundness to qualitative judgments; the practitioners behind "Benchmarks Are Not Validation" show how to apply that discipline to a whole agent system, with traces as the validation object.7 20 The Bank of England's firms expect the same shift from input-output testing to guardrails around outcomes.10 All-pass scoring, not partial credit, is the right metric for anything that acts.
- Concrete marker: Pre-production validation scores whole-task pass rates against the institution's own cases, including every escalation case, and the control function sets the release threshold.
- Concrete marker: Traces are sampled and reviewed on a schedule, and drift from the mandate automatically narrows the agent's permissions.
4. Keep the rules outside the model
Thirteen models gave thirteen rates of the same prohibited action, and provider substitution is getting harder for agentic systems. Both facts argue for the same design: rules held in a layer the institution owns, applied to every model the same way, and surviving the replacement of any of them.10 21 Gartner's forecast that 70 percent of vendor-built agentic AI will be abandoned by 2028 is a forecast about institutions that let the rules live in the vendor's prompt.26
- Concrete marker: Replacing the underlying model requires re-validation but no rewriting of the institution's mandates, evidence rules or escalation cases.
- Concrete marker: A supervisor's question about which authority an agent acted under is answered from the trace within a day, not from an interview.
The leadership test
None of the frameworks above is finished. Chief risk officers and chief information officers can test their institution's position with six questions:
- For each agent in production, can we state, from a record rather than a conversation, whose authority it acts under and what its limits are?
- Which actions have we prohibited for agents outright, and where is that prohibition enforced: in the model's prompt, or in a layer the model cannot reinterpret?
- Which cases must come back to a person, and were they specified before go-live or discovered in review?
- When an agent's instruction conflicts with our policy, does it stop and tell someone, and do we have a record of the last time it did?
- If a supervisor asked which evidence and which authority an agent applied to a specific action, would we answer from the system or from an interview?
- If we replaced the model next quarter, which of our controls would still exist?
Institutions that can answer these questions have a governed layer between their agents and their systems. Those that cannot are relying on the model's disposition, and 2026 showed, model by model and run by run, how variable that is.
The year agents moved from chat to action was also the year supervisors said, in writing, that the institution must decide for itself what an agent may do. That is not a burden. It is what boards have always done for people: grant authority, require evidence, define the exceptions, keep the record. The work of 2027 is to write those decisions down in a form a system can enforce. The institutions that do it will be able to let their agents act. The ones that do not will discover the limits of a prompt in the one case where it mattered.