By the autumn of 2026 the first draft of almost everything an investment bank produces can be generated by a machine. The pitch book, the trading comparables, the precedent transactions table, the first cut of the model and the summary of the target's last eight quarters now arrive in minutes. Then a managing director opens the file and starts changing it. The comparables set is wrong for this client. The adjustment to EBITDA is not the one the firm has defended in the last three processes. The precedent that matters was never announced and the one that was announced is being cited for the wrong reason. The work has been automated. The judgment that makes it saleable has not.
The adoption numbers are no longer in doubt. Evident's 2026 index of 50 of the world's largest banks, published on October 6, found that AI capability advanced nearly three times faster over the past year than the average of the previous three years, the fastest pace since the index began; JPMorganChase held first place for the fifth consecutive year, with Morgan Stanley ninth and Citigroup eighth. Yet only 12 of the 50 banks reported a realized or projected return on AI, up from eight a year earlier.1 Goldman Sachs' own economists found in March that 70 percent of S&P 500 management teams discussed AI on their quarterly calls, that 10 percent had quantified its effect on a use case and 1 percent on earnings, and that "we still do not find a meaningful relationship between productivity and AI adoption at the economy-wide level" (Exhibit 1).2 At Citigroup, 70 percent of staff were using the bank's proprietary AI tools by the fourth quarter of 2025 and 175,000 had been put through AI training, and the chief executive's summary of the lesson was that "someone using AI is going to probably be better at your job than you are."3
Exhibit 1

Our argument is that in investment banking the gap between deployment and value has a precise location. The models have mastered the work that is written down: the filings, the transcripts, the public comparables and the arithmetic that connects them. They have not mastered the house view, which is the set of choices a particular firm would defend in front of a client, a fairness committee or a litigator: which comparables belong in the set and which do not, which adjustments the firm makes and which it refuses, which precedents it cites for what, and how it prices a risk for this client in this market. That judgment is the bank's edge. It is held by senior people, it is taught by apprenticeship, and almost nowhere has it been written down in a form a system can use. In this paper we set out the 2026 evidence, what the people shaping the next four years expect, and four moves that turn the house view into an asset.
The analyst's work is done. The banker's is not.
The clearest measurement comes from the Finance Agent Benchmark, whose second version was published on October 7 with 927 questions written and reviewed by finance professionals to the standard of a second- or third-year investment banking analyst, each requiring several sources and "domain-specific convention." On the parts of the job that live in public documents, the best current models are strong: Gemini 4 Argon scores 84.8 percent on earnings analysis and 79.7 percent on market analysis, and the best general quantitative score is 81.8 percent. On the parts that depend on how a house does things, the scores fall away. The best adjustments score is 60.5 percent, the best comparables score 52.0 percent, the best precedents score 49.8 percent, and the best financial modeling score, on building the model an analyst would build, 34.5 percent. No model passes every part of a question more than 51 percent of the time (Exhibit 2).4
Exhibit 2

That ordering is not an accident of the test. Earnings analysis has one right answer that sits in the filing. A comparables set has a defensible answer that sits in the firm's experience of which peers a buyer will accept and which an opposing adviser will attack. A precedent table depends on knowing which announced deals were really comparable and which were distressed or mispriced. A model depends on dozens of conventions about normalization, synergies, leases and minority interests that differ between houses. The benchmark measures the distance between what is public and what is proprietary.
BigFinanceBench, published in June by a team at Rogo and OpenAI with 928 items written by 52 practitioners, most of them current or former investment bankers and private equity investors, locates the failure inside the workflow. The three leading agents, Claude Opus 4.7, GPT-5.5 and Claude Sonnet 4.6, score 58.8, 58.8 and 58.5 percent on expert rubrics and 41.5, 44.3 and 38.4 percent on the final answer, a headline tie within 0.3 points. Among trajectories where retrieval and definition were already perfect, calculation scores fell within a narrow band of 84.1 to 87.6 percent across nine models, and the authors conclude that "many failures accumulate before the final arithmetic stage rather than arising from arithmetic alone."5 The machine can compute. What it cannot reliably do is choose the right source, define the metric the way the firm defines it and apply the adjustment the firm would apply, which is to say the work of the associate and the vice president rather than the analyst.
A third 2026 benchmark, FinFIRST, built with the support of the investment banking team at China International Capital Corporation and 123 expert-authored tasks, adds a finding that should concern anyone who signs a fairness opinion. It scores not only the answer but the evidence chain behind it. Claude Opus 5 achieved the highest atomic score, 87.59 percent, and GPT-5.6 the highest strict pass rate, 71.54 percent, but across all systems 17.48 percent of correct answers were not fully supported by the evidence the agent had actually inspected, and for one model the rate was 34.62 percent. "Final-answer accuracy alone," the authors write, "cannot distinguish a well-supported result from one reached without a complete evidence chain" (Exhibit 3).6 A number that is right for the wrong reason is, in a pitch, a number the client cannot rely on and the banker cannot defend.
Exhibit 3

What the house view is, and why it is not in the data room
Ask a senior banker why a comparables set is right and the answer is rarely a rule. It is a history: the buyer who walked when a peer was included, the committee that struck an adjustment, the process in which the firm's precedent table was taken apart by the other side. The house view is the residue of those episodes, and it governs how the firm prices risk for a client and what it will say in writing. It is why two banks with the same data access and the same model produce different valuation ranges, and why clients pay one of them more.
The 2026 evidence says this knowledge is unusually resistant to being written down. In Anthropic's June Economic Index, drawing on about 9,700 survey respondents, people with fifteen or more years of experience rated the share of their work AI can do roughly ten percentage points lower than first-year workers did, a gap the authors attribute to "tacit or context-specific expertise," and the reasons most often given for tasks AI will never take over were judgment, contextual awareness and situational reasoning.7 The same shape appears when models are asked to meet an institution's full set of requirements at once. In EnterpriseRAG, an August benchmark of 13 frontier models under realistic enterprise retrieval, individual constraints were satisfied about 80 percent of the time, but all of a task's requirements were met together in only 26.8 percent of responses.8 A pitch book is nothing but a set of simultaneous requirements.
The vendors understand where the boundary sits, and their 2026 products describe it with some honesty. OpenAI's ChatGPT for Financial Services, launched on September 10 with Morgan Stanley and Evercore as design partners, ships with Daloopa, PitchBook, LSEG News and Crunchbase data built in, more than 50 connectors, and the ability to turn a peer comparison into an editable model or a pitch book on the firm's own templates. Its product lead's description of the ambition was that "a new intern on day one has all of this pre-loaded and ready to go"; the head of ChatGPT added that "there's a difference between what looks good in a demo and what is actually a usable output."9 JPMorgan's Asia Pacific head of investment banking, describing the firm's global rollout in May, put the gain in the same place: "AI streamlines the preparation of content and materials."10 Goldman Sachs, which spent six months co-developing agents with Anthropic for onboarding, reconciliation and compliance, listed help with pitch books as a possible future use rather than a current one.11 And JPMorgan's co-chief executive of the Commercial and Investment Bank told investors in February that banker and sales enablement was producing real productivity gains while some of the work was "table stakes" and not measurable.12 Content and materials are the intern's job. The house view is the product.
The apprenticeship that carried the house view is being cut
Investment banking never wrote its judgment down because it had a better mechanism: the apprenticeship. Analysts built the comps, associates were corrected on them, and by the time a banker could run a process the house view was in their head. That mechanism is being removed at the bottom at the same time as the work at the bottom is automated. McKinsey's QuantumBlack reported in June that banks are cutting junior analyst classes by as much as two-thirds, and that roughly 62 percent of banks' AI talent is sourced from those same junior cohorts; its senior partner's caution was that "banking is an apprenticeship business. Today's junior analysts become tomorrow's managing directors."13 Jamie Dimon told Bloomberg in Shanghai in May that "we will be hiring more AI people and fewer bankers in certain categories."14 In July, Goldman Sachs' chief financial officer said AI and process improvements were making employees more productive and that the firm was not pursuing a structural rework of its headcount, while the chief executive insisted AI would not replace "what matters most in driving our business, our extraordinary people."15
The labor data show where the effect lands. The Stanford Digital Economy Lab's June 2026 indicators, drawn from payroll data on about 4.6 million workers, found employment of 22- to 25-year-olds in the most AI-exposed occupations contracting at 3.8 percent a year since late 2022, against 2.0 percent growth for the least exposed; for workers aged 31 to 34 the decline was 1.7 percent, and for those aged 35 to 40 employment was growing about 2 percent. The divergence is concentrated among early-career workers and fades with age (Exhibit 4).16 Banking's version of this is a knowledge-transfer problem. If the analyst class shrinks by two-thirds and the survivors spend their first years correcting machine output rather than building it, the path by which the house view was learned narrows at the moment the senior people who hold it are closest to retirement.
Exhibit 4

Morgan Stanley's chief executive called AI "our friend" on the first-quarter call and "additive to what we have."17 That is probably right for the next two years and wrong for the next ten, unless the firm makes the house view explicit while the people who hold it are still in the building.
Clients and supervisors will ask who decided
The people on the other side of the pitch are moving faster than the banks assume. In a survey of 1,000 senior dealmakers in 27 countries commissioned by Datasite from FT Longitude in March and published in July, 96 percent were using or exploring AI for sourcing and screening, 50 percent used it regularly in due diligence, 62 percent said human-only decision-making was no longer defensible in complex transactions, and 43 percent said AI already made better deal decisions than humans in some situations. Yet 45 percent said the decision to sign should always be exclusively human, 33 percent wanted a human decision informed by AI, 7 percent would proceed on an AI recommendation without review, and 58 percent relied on human review to build confidence in AI output; accuracy, cited by 71 percent, was the attribute they demanded most.18 A client that uses AI in diligence and still insists on a human signature will want to know which parts of the adviser's valuation were judged, and by whom.
Supervisors are asking the same question from the other side. FINRA's 2026 oversight report, and the January note from its chief regulatory operations officer on AI agents at member firms, list among the risks of agentic systems "lack of industry-specific domain knowledge," reduced auditability of multi-step reasoning and action beyond the user's authority, and remind firms that supervision, recordkeeping and fair dealing obligations "continue to apply."19 The SEC chair told the FSOC's AI roundtable in March that "the Commission's mandate to protect investors is technology neutral," that "misconduct remains misconduct, regardless of the medium," and that an algorithm "cannot serve as the sole basis of an SEC enforcement action" because it cannot weigh credibility or intent.20 Neither regulator has written a rule for a machine-generated comparables table. Both have said the firm, not the model, owns the answer.
Gartner's March predictions describe the infrastructure this implies: by 2030, it expects universal semantic layers, the governed definitions an organization's systems share, to be treated as critical infrastructure alongside data platforms and cybersecurity, and half of AI agent deployment failures to stem from insufficient runtime enforcement of governance.21 In September it predicted that by 2028, 70 percent of enterprises will abandon agentic AI built for them by vendors' forward-deployed engineers, because the buyer never develops "the capabilities, governance, and operational ownership required to independently sustain and evolve the solution."22 For a bank, the capability in question is not engineering. It is the house view, in a form the agent can read.
What the people shaping the next four years expect
The forward views worth weighting came from people with a balance sheet, a franchise or a mandate at stake.
- The capital will arrive before the proof. The BIS General Manager told the Global Fintech Fest in September that AI-related investment is expected to rise from about $500 billion today to between $3 trillion and $4 trillion by 2030, while the median estimate of its effect on productivity is around half a percentage point a year, and that productivity gains so far are larger for less experienced workers than for senior ones. "A pullback in investment," he warned, "could turn today's capital expenditure boom into a bust."23 For investment banks that is both the fee pool of the decade and a warning about what the pool depends on.
- The workforce will change faster than the apprenticeship. Dimon's April letter said the pace of adoption "will likely be far faster than prior technological transformations," that "AI will definitely eliminate some jobs, while it enhances others," and that "there is a possibility that AI deployment will move faster than workforce adaptation to new job creation."24 His May remark about hiring fewer bankers in certain categories made the application to his own firm explicit.14
- The decision stays human, and the client will check. Seventy-one percent of the dealmakers in the Datasite survey expected firms that fail to adopt AI to struggle to compete within five years; 45 percent would never let the signing decision leave human hands; 26 percent had already delayed or canceled a hire because AI could do the work.18
- Bought context will not stick. Gartner's two 2028 predictions, that 70 percent of vendor-built agentic AI will be abandoned and that semantic layers will become critical infrastructure by 2030, say that the firm's own definitions and ownership are the part of the stack it cannot outsource.21 22
Penomic's own expectation, grounded in these views and the benchmarks above, is that the period to 2030 will separate investment banks on one variable: whether the house view exists outside the heads of the people who hold it. Models will keep converging; BigFinanceBench's three leaders already sit within 0.3 points. The analyst class will keep shrinking, and with it the apprenticeship. Clients will arrive at the pitch with their own machine-generated comparables and ask the bank what it would change and why. By 2028 we expect the leading houses to be running agents that produce comps, precedents and first-draft models against an explicit, governed statement of the firm's conventions, with every departure recorded and owned; by 2030 we expect the ability to show a client or a supervisor which convention, which precedent and whose authority stood behind a number to be a condition of mandates, not a differentiator. The banks that cannot do that will still be selling judgment. They will simply be unable to prove it was theirs.
Writing down the house view: four moves
None of this requires a bank to stop deploying the tools it has. It requires treating the judgment those tools cannot supply as an asset with owners, a structure and a review cycle, in four moves.
1. Inventory the conventions the firm defends
Start not with documents but with the decisions a senior banker makes a hundred times a year and would defend under cross-examination: the inclusion rules for a comparables set in each sector, the adjustments the firm makes to reported earnings and the ones it refuses, the criteria by which a precedent counts, the risk factors it prices for a given client type. Written as rules with reasons, these are the first version of the house view, and they are small enough to govern. The Finance Agent Benchmark's own category list, adjustments, comparables, precedents and modeling, is a usable table of contents.4
- Concrete marker: Each sector group has a named owner of its comparables, adjustment and precedent conventions, with the reasons recorded, not just the rules.
- Concrete marker: A machine-generated comps set or precedent table can be checked against the firm's conventions automatically, and every departure is flagged before a banker sees it.
2. Mine judgment from the redlines, not from the memo
Senior bankers cannot fully say how they decide, but their decisions are recorded: the version history of every pitch book, the fairness committee minutes, the comments on the model, the comparables struck and the reason given. Use models to extract the conventions those records reveal, then have the owners confirm or correct them. BigFinanceBench's finding that failures accumulate in source selection, definition and adjustment before arithmetic says where to look first.5
- Concrete marker: Every senior correction to a machine draft is captured as a fact pattern, a decision and a reason, and feeds the convention set rather than disappearing into the next version.
- Concrete marker: Managing directors approaching retirement spend time validating extracted conventions for their sector, and their sign-off is recorded.
3. Make the evidence chain part of the deliverable
FinFIRST showed that one in six correct answers from current agents was not fully supported by the evidence inspected.6 A pitch, a fairness opinion or a valuation range must carry its sources, the convention applied to each number and the person who approved any departure, so that a client, a committee or a supervisor can trace it. This is the control FINRA and the SEC have already described without naming it, and the attribute the Datasite respondents ranked first.18 19 20
- Concrete marker: Every number in client material records its source, the firm convention it applied and the reviewer who approved it, in a form that survives into the data room.
- Concrete marker: Outputs that must be seen by a human before they reach a client are specified in advance by convention and client type, not discovered in review.
4. Rebuild the apprenticeship around the conventions
If the analyst class is a third of its former size and the first draft is machine-made, the apprenticeship must change what it teaches. Juniors should be trained to interrogate the machine's comparables against the firm's conventions and to argue for a departure, because that argument is the house view in action. The convention set then becomes the curriculum, the firm's AI talent pipeline, which McKinsey put at 62 percent sourced from junior cohorts, keeps its intake, and the knowledge that used to take six years to acquire is available on the first day in a form the firm can defend.13
- Concrete marker: Analyst and associate reviews assess the quality of challenges to machine output against firm conventions, not the volume of material produced.
- Concrete marker: Replacing the firm's language model or vendor does not require rebuilding its conventions, precedent set or decision rights.
The leadership test
No bank has finished this work. Leaders can test their own firm's position by asking six questions:
- For our three largest sector groups, could a system today state which comparables we would include, which adjustments we would make and which precedents we would cite, with the reasons?
- When a managing director changes a machine-generated comps set, where does the reason go?
- If a client arrived with their own AI valuation and asked what we would change and why, could we answer from a documented convention or only from the room?
- If FINRA or the SEC asked which sources, conventions and approvals stood behind a number in a fairness opinion, would we answer from the system or from interviews?
- Which of our senior bankers hold conventions that exist nowhere else, and what is the plan for the year before they retire?
- If our analyst class is cut by two-thirds, what exactly will the survivors be learning, and from whom?
Banks that can answer these questions have a house view that exists outside their people. Those that cannot have a franchise that depends on tenure, and the tenure is shortening while the machines get faster at producing everything except the judgment.
The banks that win the next four years will not be the ones with the most agents or the earliest access to the largest model. They will be the ones that took the thing the senior banker knows and the pitch book does not, wrote it down, attached an owner to it, and made every machine in the firm use it. The technology to do that arrived in 2026. The decision to do it has not been taken in most houses, and it is the only part of the program a competitor cannot buy.