The institutional intelligence gap
Financial institutions have wired AI into their data. Value is stalling because the definitions, policy, precedent and judgment that make an answer right for a particular institution were never written down in a form a machine can apply.
By Jared D. Yerian and Jennifer Kilian · 17 minute read
Download the PDFThree years into the enterprise AI cycle, the pattern in financial services is remarkably consistent. The model can find the credit memo, the policy manual, the board pack and the spreadsheet behind last quarter's numbers in seconds. Then a senior person rewrites the memo, corrects the definition, adds the exception the model missed, and explains why the committee decided differently last year. The work has moved. The judgment has not.
The 2026 numbers describe that gap with unusual precision. In the Cambridge Centre for Alternative Finance's global survey of 628 financial institutions, fintechs, vendors and regulators, published in April with the BIS, the IMF and the World Bank as partners, 81 percent of firms use AI and 40 percent say they are at the scaling or transforming stage, yet only 14 percent describe AI as transformational to their business, and 55 percent of the industry say its value is hard to measure.1 McKinsey's August survey tells the same story across all industries: nearly nine in ten organizations use AI, 44 percent are scaling it enterprise-wide, and still only 37 percent report any effect on EBIT, a figure that has not moved in a year (Exhibit 1).2 In February, the ECB's supervisory board member Pedro Machado put adoption among large euro area banks at more than 85 percent and named the governance problem directly: "fragmented ownership, with responsibility split across IT, data science teams, business lines and control functions."3

Our argument is that the gap is structural rather than technical, and that it will not close with the next model release. What makes an answer right inside a bank, a fund or an insurer is rarely contained in any single record. It sits in how the institution defines its terms, which policy applies to which case, what was decided before and why, and who has the authority to depart from the rule. That knowledge lives in the heads of experienced staff and partly in their documents. Almost nowhere has it been made explicit, governed and machine-usable. We call the result the institutional intelligence gap. In this paper we set out the 2026 evidence for it, what the people who will shape the next four years expect, and four moves that close it.
Access is no longer the bottleneck. Judgment is.
The models of October 2026 are very good at the parts of financial analysis that are written down, and they are converging. On the Finance Agent Benchmark, which gives agents search and filing tools and 927 questions reviewed by finance professionals, 76 models were scored this month. Gemini 4 Argon leads at 65 percent; Claude Fable 5.1, Claude Opus 5.5 and Claude Sonnet 5.5 sit at 58 to 59 percent; the GPT-5.6 and GPT-6 variants at 49 to 55 percent; and the open-weight models from Zhipu, Mistral, DeepSeek, Moonshot and Alibaba, including GLM 5.3, Mistral Large 4, DeepSeek V4.1, Kimi K3 and Qwen 3.8, at 51 to 58 percent, within a few points of the proprietary frontier and in some cases ahead of it. The spread between the best model and the twentieth is twelve points. The spread within any one model, across task types, is fifty. Every leader scores above 80 percent on earnings analysis, market analysis and general quantitative work. The scores fall away on the tasks that depend on how a particular house does things: adjustments, 60 percent; comparables, 52 percent; precedents, 50 percent; building the financial model an analyst would build, 35 percent. No model passes every part of a question more than 51 percent of the time (Exhibit 2).4

BigFinanceBench, published in June by a team from Rogo and OpenAI with 52 former investment-banking and private-equity professionals writing the questions, locates the failure more exactly. The best agents available at publication, Claude Opus 4.7, GPT-5.5 and Claude Sonnet 4.6, scored below 60 percent on expert rubrics and below 45 percent on final answers, and the three sat within 0.3 points of each other. The errors, the authors write, "mostly come before arithmetic": in source selection, metric definition and accounting adjustment. Once the set-up is right, the calculation scores barely differ between models.5 That is the institutional intelligence gap measured in a laboratory. The models have the mechanics. They do not have the house view.
The same pattern shows up outside finance. EnterpriseRAG, an August benchmark of 13 frontier models under realistic enterprise retrieval, found that models satisfy each individual instruction about 80 percent of the time but satisfy all of a task's requirements at once in only 27 percent of responses.6 And Anthropic's June Economic Index, drawing on roughly 9,700 surveyed users, found that people with fifteen or more years of experience rate the share of their work that AI can do about ten points lower than first-year workers do, a difference the authors attribute to "tacit or context-specific expertise." The reasons respondents most often gave for tasks AI will never take over were judgment, contextual awareness and situational reasoning.7
None of this is a reason to slow down. It is a reason to be precise about what is missing. The models are not short of data or language. They are short of the institution's own account of what its data means and how it decides.
Same data, different right answers
Two lenders can hold identical information about a borrower and reach different, defensible conclusions, because exposure, default, adjusted EBITDA, liquid asset and active customer each mean something slightly different in each house, and sometimes in each function of the same house. Risk, finance and the relationship team are often all correct within their own mandate. The problem is not the disagreement. It is that the distinctions were never represented clearly enough for a person outside the circle, or a system, to pick the right one.
Executives already know this is where their programs are stuck. In KPMG's second-quarter 2026 pulse of 204 US banking leaders, data readiness and access was the top barrier to deploying AI agents, cited by 63 percent, ahead of the complexity of the systems themselves; only 31 percent said they had full visibility of their AI operating costs, and just 3 percent had assigned accountability for AI decisions to a central governance committee.8 In the Cambridge survey, 46 percent of firms named legacy and siloed systems as a barrier and 40 percent named data quality; only 10 percent said their workforce was highly prepared.1 Among 728 EU securities firms surveyed by ESMA, 87 percent of the 847 AI use cases in production were internal-only and 77 percent had low or no autonomy. Firms expected modest cost savings and almost no revenue from generative AI, and 23 percent expected no savings at all.9
The ECB's reading is that the technology is neutral and the governance is not. A bank that cannot explain a model's output in terms its decision-makers, risk function and internal audit can challenge, Machado told supervisors in February, "cannot truly control that model."3 The Bank of England's February roundtables with UK firms reached a related conclusion from the other side: traditional input-output validation is "becoming unsustainable and less effective" for generative and agentic models, and the human-in-the-loop assumption that most control frameworks rest on is being "challenged by agentic AI."10 Both supervisors are describing the same missing object. Explaining an output requires the definitions it applied. Controlling an agent requires knowing which cases a person must see. Neither is a property of the model.
Five layers, one missing
It helps to describe the modern stack as five layers. Most institutions have invested in four of them.
When the fourth layer is missing, every model call improvises it, and every reviewer has to supply it again by hand. That is the work that has moved but not disappeared in so many AI programs, and it is why 80 percent of McKinsey's respondents report individual productivity gains while the enterprise numbers stand still.2
What governed meaning does to accuracy
The evidence that an explicit knowledge layer changes model behavior now comes from the current generation of models, and the effect is large and consistent. In an April 2026 study, three frontier models were asked the same 99 analytical questions about a retail business twice: once with only the database schema, and once with a four-kilobyte document of business definitions. Claude Opus 4.7 went from 50.5 to 67.7 percent, Claude Sonnet 4.6 from 46.5 to 68.7, and GPT-5.4 from 45.5 to 68.7. The gains were statistically significant for every model, and the three models were statistically indistinguishable from one another in each condition (Exhibit 3).11 The definitions mattered more than the choice of model.
A second 2026 study separated definitions from governance. On a 100-question enterprise analytics benchmark, a model writing SQL directly produced at least one hallucination in 79 percent of answers. Adding schema retrieval cut that to 54 percent. Adding business definitions cut it to 40 percent but still leaked row-level access controls in 35 percent of cases. Only when the definitions were combined with access policy and validation before execution did the hallucination rate fall to zero, with strict accuracy rising from zero to 79 percent.12 dbt Labs' April update of its own benchmark found the same asymmetry on an insurance dataset: once the business logic was modeled, a governed semantic layer answered 98 to 100 percent of questions correctly against 84 to 90 percent for direct text-to-SQL with the same models, and when the governed layer failed it returned an error, whereas text-to-SQL returned "plausible but incorrect numbers."13

Three caveats belong here. Two of these studies are vendor-authored, all three use modest question sets, and none is a bank. But the direction is the same in every one, it matches what BigFinanceBench found about where errors begin, and it is consistent with Gartner's April observation that in the agentic era "control of enterprise context is economic power."14 Definitions alone are not enough. They have to travel with who may see what, what counts as a valid answer, and when the system must stop.
The knowledge is walking out of the building
The institutional knowledge layer has always existed. It has simply been stored in people. That storage is becoming less reliable for a demographic reason and a structural one.
The demographic reason is tenure. In US commercial banking, the share of employed people aged 55 and over rose from 18.7 percent in 2015 to 21.3 percent in 2025; across finance and insurance it rose from 21.9 to 23.1 percent, and the number of finance and insurance workers aged 65 and over grew by a third, from 350,000 to 469,000.15 Finance is not an unusually old industry. The point is narrower: the judgment that a credit committee, an underwriting desk or a private bank relies on is concentrated in senior people, and those people are closest to the exit. In Kelly's survey of executives and employees reported in January 2026, 92 percent of executives said retirements will worsen their skills shortages, 67 percent believed their organization was prepared, 17 percent of their employees agreed, and fewer than half of organizations had a formal knowledge-transfer program in place before people retire (Exhibit 4).16 Among credit unions, the NCUA's succession-planning rule took effect on January 1, 2026, after the agency concluded that the absence of a plan was "one of the most common causes for unplanned and unforced credit union mergers."17

The structural reason is that writing the knowledge down does not work the way institutions assume. Asked to document how they decide, experienced people produce a policy summary. The convention that actually drives the decision, the exception they would make and the one they would not, stays in their heads. The Anthropic finding that the most experienced users see the most tacit content in their own work is the same point from the user's side.7 The practical implication is that the knowledge layer has to be built from how senior people actually decide, case by case, from minutes, exception logs and approvals, and then confirmed by them, rather than from what they say in a workshop.
Regulators are already asking for this layer
Supervisors spent 2026 saying, with unusual candor, that their own frameworks do not yet cover what banks are deploying. In April the Federal Reserve, OCC and FDIC replaced the 2011 model risk guidance with SR 26-2, broadening conceptual soundness to cover qualitative judgments and naming interpretability as a validation approach, and in the same document explicitly excluded generative and agentic AI as "novel and rapidly evolving."18 Two weeks later Vice Chair Bowman told the FSOC's AI roundtable that the agencies should "assess whether our supervisory guidance is fit for the future"; the promised interagency request for information had not been issued by October.19 The largest US banks are therefore deploying generative AI in a model risk regime that, by its own terms, does not describe it.
The direction of travel elsewhere is the same. The Financial Stability Board consulted in June on twelve sound practices for the responsible adoption of AI, four of them on organization-wide governance, with the final report due this month.20 The Monetary Authority of Singapore issued binding AI risk management guidelines on October 7 that require every financial institution to keep an AI inventory, assess materiality, and remain accountable for vendor models, with agentic AI guidance to follow in 2027.21 In the EU, the Digital Omnibus agreed in May moved the AI Act's high-risk obligations, which include creditworthiness assessment, to December 2027.22 In September the chair of the BIS Financial Stability Institute said that "the limited explainability of large language models" challenges existing supervisory expectations and that it is "imperative that MRM guidance be revisited," and the Basel Committee said on October 1 that integrating AI into critical functions "will require careful governance, robust risk management and ongoing supervisory attention."23
Read together, these expectations converge on the fourth layer. Explainability requires the definitions the system applied. Accountability requires decision rights. Human-in-the-loop requires knowing which cases the institution has decided a person must see. An institution that has built its knowledge layer can answer a supervisor's questions from the system. One that has not will answer them, as most do today, from interviews.
What the people shaping the next four years expect
Forecasts about AI are cheap. The ones worth weighting are those made by people with capital, a balance sheet or a supervisory mandate at stake, and in 2026 they were unusually specific.
- The investment will not wait for the evidence. The BIS General Manager, Pablo Hernández de Cos, told the Global Fintech Fest in September that global AI-related investment is expected to rise from around $500 billion today to between $3 trillion and $4 trillion by 2030, while the median estimate of its effect on productivity is around half a percentage point a year. "Should the returns to AI disappoint," he said, "a pullback in investment could turn today's capital expenditure boom into a bust."24 Goldman Sachs economists, writing in March, still found "no meaningful relationship between productivity and AI adoption at the economy-wide level," and noted that only 10 percent of S&P 500 management teams had quantified AI's impact on their use cases and 1 percent on earnings.25
- Adoption will outrun the organization. Jamie Dimon's April letter to JPMorganChase shareholders called the pace of adoption "likely far faster than prior technological transformations" and warned that "AI deployment will move faster than workforce adaptation."26 In the Cambridge survey, 81 percent of financial institutions expect agentic AI to be meaningfully achieved by 2030, and 24 percent expect a net reduction in roles by then.1 McKinsey's respondents expecting headcount reductions next year rose from 32 to 39 percent.2
- Bought context will not stick. Gartner predicted in September that by 2028, 70 percent of enterprises will abandon agentic AI built for them by vendors' forward-deployed engineers, because they "fail to build internal capability," and in April that by 2028 more than half of enterprises will stop paying for assistive AI in favor of outcome-based workflows.14 The knowledge that makes an agent useful inside a specific institution is, in Gartner's phrase, "economic power," and it is not transferable by contract.
- The risk horizon lengthens. The World Economic Forum's 2026 survey of more than 1,300 experts ranked adverse outcomes of AI thirtieth among global risks over two years and fifth over ten, the largest riser in the report.27
Our own expectation, grounded in those views and in the benchmarks above, is that the period to 2030 will separate institutions on one variable: whether their meaning, precedent and decision rights exist in a form their systems can use. Models will keep improving and keep converging; twenty proprietary and open-weight models already sit within twelve points of each other on the Finance Agent Benchmark, and open-weight models from China and Europe now match the US frontier on most task types. Supervisors will keep moving the questions from "which model" to "which definition, whose authority, and where does a person decide." And the people who hold the unwritten answers will keep retiring. By 2028 we expect the first supervisory findings to turn on an institution's inability to state the definitions and decision rights an agent applied, and by 2030 we expect the institutional knowledge layer to be as ordinary a line item in a bank's technology estate as the data warehouse is today.
Building the institutional knowledge layer: four moves
The institutions that are closing the gap are not doing so with a single platform purchase. They are doing four things, in roughly this order, and treating each as a governed asset rather than a project.
1. Start from decisions, not documents
The common mistake is to begin by ingesting every policy manual. Begin instead with the twenty decisions the institution makes most often and defends most carefully: a credit exception, a suitability determination, a limit breach, a reserve estimate, a counterparty downgrade. For each, write down the concepts it uses, the evidence it requires, the thresholds that apply, who may decide and who must be told. That inventory is the first version of the knowledge layer, and it is small enough to govern. It is also exactly the inventory MAS now requires and the FSB is about to recommend.
- Concrete marker: Each priority decision has a named owner, an explicit definition set and an evidence standard that a system can check.
- Concrete marker: Definitions carry scope, jurisdiction, effective date and the function that owns them, not just a label.
2. Mine judgment from cases, not from workshops
Experts cannot fully write down how they decide, but their decisions are recorded: committee minutes, exception logs, approval chains, redlines. Use models to extract the conventions those records reveal, then have the owners confirm or correct them. BigFinanceBench's finding that errors begin in metric definition and accounting adjustment, not arithmetic, says where to look first.5
- Concrete marker: Precedent is captured as a fact pattern and a decision, with the reason recorded, so it can be matched to a new case.
- Concrete marker: Retiring senior staff spend their final months validating extracted conventions, not writing memoirs.
3. Attach authority to meaning
A definition without a decision right is a glossary. Each concept, rule and precedent should carry who may use it, who may change it, who must approve a departure, and which outputs require a human before they are published. This is where most semantic-layer projects stop, and it is the step the 2026 results show carries the most weight: definitions alone left 40 percent of answers hallucinated and leaked access controls; definitions plus policy and validation did not.12
- Concrete marker: Every system output records which definition, evidence and authority were applied, so a reviewer or a supervisor can trace it.
- Concrete marker: The cases that must go back to a human are specified in advance, not discovered in review.
4. Govern the layer like a model
Finance already has the discipline for this. Conceptual soundness, independent validation, ongoing monitoring, change control and an inventory with owners are the elements of model risk management, and SR 26-2 now explicitly extends conceptual soundness to qualitative judgments.18 Applied to the knowledge layer, they turn institutional meaning into an asset that survives staff turnover, model replacement and the next vendor cycle.
- Concrete marker: The knowledge layer has an inventory, owners, review cycles and a change log, like any material model.
- Concrete marker: Replacing the underlying language model does not require rebuilding the institution's definitions, precedent or decision rights.
The leadership test
No institution has finished this work, and none of the frameworks above are a substitute for judgment about where to begin. Leaders can test their own institution's position by asking six questions:
- Which twenty decisions do we defend most often, and could a system today state the definitions, evidence and authority each one uses?
- When two functions report different numbers for the same concept, is there a recorded owner of the definition, or a meeting?
- Where does our senior judgment actually live, and what happens to it when the three most experienced people in that area retire?
- If a supervisor asked which definition and which precedent an AI output applied, would we answer from the system or from an interview?
- Which cases have we decided a human must see before an action is taken, and does the system know?
- If we replaced our language model next year, what would we have to rebuild?
Institutions that can answer these questions have a knowledge layer, whether or not they call it that. Those that cannot have a gap, and it is widening as the models improve, because better models reach further into the institution's data and apply the same missing judgment faster.
The organizations that capture the most value from AI in financial services will not be those with the most agents or the largest model. They will be the ones that made their own meaning explicit, attached authority to it, and governed it with the discipline they already apply to models. That is the institution's intelligence. The technology to use it has arrived. The work of writing it down has not been done.
Read the full paper and get the PDF
The rest of “The institutional intelligence gap”, every exhibit and the full source notes. We will also email you the PDF. One form unlocks all Penomic Research.
About the authors

Jared D. Yerian, CFA, CIRA, CDBV, Senior Board Advisor, Penomic. Former Partner at McKinsey & Company, where he was one of five founders of the global Recovery & Transformation Services practice, and later Senior Partner and Co-Lead of Transformation at Oliver Wyman. He has served in CFO, CRO and board advisory roles on complex financial and operational transformations, restructurings and M&A. LinkedIn

Jennifer Kilian, Senior Board Advisor, Penomic. Former Partner at McKinsey & Company and Co-Founder and CEO of Cognition Capital. A transformation executive working where AI, digital product and experience-led growth meet, advising CXOs and boards. LinkedIn
Institutions building an institutional knowledge layer can request a confidential briefing with the authors.
Request a briefingMore from Penomic Research
Why AI pilots stall after the demo
2026 surveys, bank disclosures and agent benchmarks show that AI pilots stall because the institution's definitions, precedent and decision rights were never written down and governed.
Govern the agent through the institution's decision rights
2026 adoption data, agent benchmarks, misalignment studies and supervisory statements show why agent control belongs in the institution's decision rights, not the model's prompt.
One number, five definitions
2026 supervisory findings and enterprise benchmarks show why governed definitions, with owners and effective dates, now return more than any model choice in financial AI.