The intake has been automated. The appetite has not.
The progress at the front of the funnel is real and measurable. AIG's chief executive Peter Zaffino told investors in February that Lexington had processed more than 370,000 submissions with its generative AI tooling "without additional human capital resources," and in May reported a 55 percent reduction in time to quote, 30 percent more quoted submissions and roughly 40 percent more bound submissions in Lexington middle-market property.7 In Convr's survey of 211 commercial insurance professionals, 53.6 percent had AI in a production underwriting workflow, yet only 20.4 percent of leaders were highly confident their organization had a clear underwriting AI strategy.8
Underwriters draw a sharp line between that progress and the judgment it serves. In hyperexponential's June survey of 350 senior commercial property and casualty underwriters in the United States and United Kingdom, 51 percent said AI's greatest contribution so far had been saving time on manual administration and 21 percent said it had improved the quality of decisions. The drags on decision quality they named were inconsistent submission data, 44 percent, pressure to bind quickly, 38 percent, and no context on similar prior risks, 35 percent (Exhibit 2).9 Both surveys are vendor-sponsored; brokers see the same pattern. Aon's global chief broking officer for commercial risk, Cynthia Beveridge, said in August that AI is making underwriting "more selective and informed," and set the limit: "AI has not fundamentally changed pricing patterns, yet."10
Exhibit 2

Capgemini's World Property and Casualty Insurance Report, published in May from a survey of 344 senior executives, puts the enterprise consequence in numbers: 10 percent of property and casualty insurers have scaled AI, 55 percent cite no clear return, and 55 percent say it is unclear who owns their AI initiatives.11 The two halves of the underwriting job have different economics. Intake automation removes cost in any market. Appetite judgment determines the loss ratio, and it matters most when prices are falling, because that is when the gap between the risks to keep and the risks to let a competitor have is widest.
What the benchmarks measure
Two 2026 studies test the appetite problem directly. In January a team at Snorkel AI published UNDERWRITE, 300 tasks built with underwriting experts around a fictional commercial insurer: a nine-table policy database including an appetite matrix, and free-text guidelines with proprietary business rules, exposed to agents through tools. Across 13 models, answer correctness ranged from 90.3 percent for Claude Sonnet 4.5 down to 30 percent, and fell by roughly 20 points across four repeated attempts. The characteristic failure was not arithmetic. The smaller OpenAI models recommended products the fictional insurer did not offer in 58 to 66 percent of product-recommendation tasks, drawing on pretrained industry knowledge rather than the guidelines in front of them, and accuracy declined sharply as the expert answers became more surprising: Claude Haiku 4.5 went from nearly 100 percent on the least surprising tasks to 66 percent on the most.12
InsClaimBench, published on October 7 with 3,780 cases and six current models, decomposes the problem along the decision chain. GPT-5.6 Sol judged individual policy rules correctly 95.48 percent of the time and reached the right payout decision 80.05 percent of the time, but got both decision and amount right in 69.72 percent of cases and reproduced the full vector of rule judgments an auditor would need in 36.90 percent. Gemini 3.8 Flash, DeepSeek V4.1 Flash, Qwen 3.8 Flash and Kimi K2.6 spread from 14.68 to 31.96 percent on the last measure. The authors describe "a progressive loss of reliability" along the chain: "correct payouts can conceal intermediate errors" (Exhibit 3).13
Exhibit 3

An underwriting decision is a chain of exactly this kind: classification, eligibility, hazard grading, limit, pricing adequacy, terms, referral and authority are each a rule judgment, and the appetite is their composition with the carrier's exceptions on top. A model that applies each rule at 95 percent and reproduces the whole chain at 37 percent is not an underwriter. It is an excellent reader of guidelines that has never been told which readings the company would defend. The newest models show the same shape on general finance work: on Vals AI's Finance Agent Benchmark v2, updated on October 7, Gemini 4 Argon leads 76 models at 65.40 percent and Claude Fable 5.1 scores 58.88 percent, and the best category scores fall from 84.8 percent on earnings analysis to 49.8 percent on precedents and 34.5 percent on financial modeling, the categories where a house's own conventions decide the answer.14
The counter-evidence is just as consistent: when the conventions are supplied, accuracy moves more than it does between models. In an April paired benchmark, three frontier models answered the same 100 analytical questions with only a database schema and then with a short document of business definitions. Accuracy rose from 45.5 to 50.5 percent to 67.7 to 68.7 percent, a gain of 17 to 23 points for every model.15 dbt Labs' April benchmark on an insurance dataset of 15 tables found a governed semantic layer answered 98.2 percent of questions correctly with Claude Sonnet 4.6 and 100 percent with GPT-5.3 Codex, against 90.0 and 84.1 percent for text-to-SQL: "With text-to-SQL, failure looks like a plausible but incorrect answer. With the Semantic Layer, failure looks like an error message."16 Both studies are vendor-authored and use modest question sets, but the direction is the same. The appetite has to be an artifact the company controls and governs, not an instruction someone types into a prompt.
Where the appetite lives today
Most carriers have invested in four layers of the underwriting stack. The fifth, where appetite sits, is rarely built.
Data
Submissions, loss runs, exposure schedules, third-party enrichment, policy administration and claims. Increasingly digitized by intake tooling.
Models
Pricing and catastrophe models, extraction models, general language models. Validated under existing model risk frameworks.
Agents
Triage, clearance, document summarization, quote preparation, portfolio reporting. Where most 2026 deployments sit.
Appetite layer
The carrier's definitions, appetite rules, exclusions, referral triggers, precedents and their reasons, and authority limits for underwriters and delegated partners, made explicit, governed and machine-usable. Rarely built.
Decisions
Quote, decline, refer, bind, with terms and price, defensible to the chief underwriting officer, the reinsurer, the auditor and the supervisor.
Today the fourth layer is stored in people, and the people are leaving. In the Bureau of Labor Statistics' 2025 Current Population Survey, 24.8 percent of the 2.97 million people employed in insurance carriers and related activities were aged 55 or over, against 23.2 percent across all US industries.17 The BLS projects employment of insurance underwriters, 125,600 in 2025, to fall 4 percent by 2035 because "automated underwriting software allows workers to process applications quickly," while still generating about 6,800 openings a year, all of them replacements.18 In the Jacobson Group and Aon's third-quarter 2026 labor study, 11 percent of insurers planned to reduce staff, up from 7 percent in January, with automation the most common reason, and Jacobson's Jeffrey Blair described the market as "hiring to backfill key positions and bring in new talent, rather than hiring for growth."19
Underwriters are more candid than their employers' investment plans. In the hyperexponential survey, the top concern, named by 44 percent, was senior judgment leaving without being passed on. Thirty-seven percent of US respondents said senior judgment mostly lives in people's heads, 40 percent said expertise was captured poorly or not at all, and coaching and knowledge transfer ranked last of ten investment areas at 8 percent, behind end-to-end workflow automation at 52 percent and submission ingestion at 48 percent.9 KPMG's September survey of insurance leaders in 20 countries found 72 percent expecting underwriting to run as a hybrid model by 2029 and 57 percent having built agents or decision-support engines for underwriting, while only 15 percent had AI governance fully integrated into strategic planning and 11 percent a strong data foundation.20 The industry is funding the layer it has and starving the layer it is losing.
Delegated authority makes the same gap a counterparty problem. Delegated business is roughly 45 percent of Lloyd's gross written premium, and from January 2026 every managing agent has a dedicated DA Oversight Manager, because "performance deterioration of delegated business tends to be less visible and harder to remediate."21 In the United States, Conning estimates MGA premium at about $128 billion in 2025, growing 12 percent against roughly 5 percent for the wider market.22 Every binding authority is an appetite handed to a third party and checked months later through bordereaux. Carl Day, deputy chief underwriting officer at Apollo Underwriting, said in August that the pressure to compromise concentrates in new business, delegated authority, new products and changes to terms, and that rigid rules such as sign-off for any reduction end up "pushing your underwriters almost to select against themselves."23 Evan Greenberg's March 16 letter to Chubb shareholders states the alternative discipline in a sentence: "We shrink whole businesses when necessary to preserve an underwriting profit." Chubb's 2025 combined ratio was 85.7 percent, and "underwriting," the letter adds, "is not something Chubb outsources."24
Regulators are asking for this layer
Insurance supervisors have moved faster than their banking counterparts on AI, and what they ask for is this layer. As of October 8, 2026, 29 US jurisdictions had adopted the NAIC's Model Bulletin on the Use of Artificial Intelligence Systems by Insurers, with California, Colorado, New York and Texas on their own regimes.25 Since March 2, twelve states have been piloting the NAIC's AI Systems Evaluation Tool for examinations, with exhibits that quantify AI usage, assess governance and detail high-risk systems; adoption is expected in November.26 An examiner using that tool will ask how an AI system that influences underwriting is governed, tested and documented. At most carriers the honest answer is that the system is governed and the appetite it applies is not.
In Europe, EIOPA's Opinion on AI governance and risk management, adopted in August 2025 and the reference point for 2026 supervision, expects undertakings to keep records that enable "reproducibility and traceability," to give supervisors and auditors "a global and comprehensive explanation" of an AI system, to define "roles and responsibilities" in policy documents "including escalation procedures," and to remain "ultimately responsible for the AI systems that they use" whether or not they built them.27 EIOPA's survey of 347 undertakings in 25 countries, published in February, found 65 percent already using generative AI, 49 percent with a dedicated AI policy, up from 25 percent in 2023, and 32 percent of use cases in production; hallucination was the most cited risk.28 EIOPA's chair, Petra Hielkema, said in April that assisted systems "ensure that human judgement remains central in decision-making processes," while agentic AI "represents a shift towards technologies that can make decisions without direct human intervention."29
Bermuda followed in August. The Bermuda Monetary Authority's consultation on a guidance note for the responsible use of AI, open until October 30, creates no new statutory obligations, but it scales expectations with materiality: higher-impact uses "such as underwriting, pricing, trading, or customer decisions" may need independent challenge, audit records, human approval points and fallback arrangements, and because "agentic AI can take a series of autonomous actions," those systems may need kill-switch requirements.30 The London market shows how far governance has run ahead of content. In the Lloyd's Market Association's survey of 39 firms representing more than 60 percent of stamp capacity, 72 percent had an AI governance framework in place and 21 percent had one in development, more than 60 percent required mandatory human review of AI outputs, and deployment in core underwriting decisions remained limited.31 A framework that mandates human review is only as good as the record of what the humans are checking against. Where the appetite is tacit, review means a senior underwriter's recollection, which is not what the NAIC tool, the EIOPA Opinion or the BMA note describe (Exhibit 4).
Exhibit 4

Read together, these expectations converge on a single object. Reproducibility requires the rules the system applied. Explainability requires the definitions and the exceptions. Escalation requires referral triggers precise enough to fire. Accountability for third parties requires an appetite exact enough to test against a coverholder's bordereaux. A carrier that has built its appetite layer can answer an examiner from the system. One that has not will answer, as most do today, from interviews.
What the people shaping the next four years expect
The forecasts worth weighting come from people with a balance sheet, a market or a mandate at stake, and in 2026 they were specific.
- The soft market will last longer than the losses suggest. Marsh's John Donnelly expects current conditions "likely to persist absent a severe northern hemisphere storm season."1 Swiss Re Institute's path to a 7.7 percent return on equity by 2028 assumes the softening continues, and Tiernan told Lloyd's that "the 2027 planning season will be different" and "we should not expect growth in core markets."2, 3
- Submission automation will keep compounding, and the organization will not keep up. AIG's ambition is 500,000 Lexington submissions a year by 2030, and Zaffino said in May of AIG's agents, which once ran for under an hour: "today, they can run autonomously for as long as 30 hours."7 Accenture's Michael Reilly, writing in February, cited the firm's survey of 430 senior underwriting executives expecting AI adoption in underwriting to rise from 14 percent to 70 percent within three years.32
- Oversight will move from the model to the decision, and insurers will be on both sides of it. Gartner's strategic predictions of September 15 hold that "by 2030, insurers, not regulators, will drive AI governance" through underwriting standards for AI liability cover, and that 80 percent of Global 500 companies will by then contractually make their CIO or chief AI officer the "Evidence Custodian" for AI accountability.33
- The people who hold the appetite will leave on schedule. The BLS projects a 4 percent decline in underwriter employment to 2035 with 6,800 openings a year, a quarter of the insurance workforce is already over 55, and the share of carriers planning reductions rose from 7 to 11 percent between January and the third quarter.17, 18, 19
Our own expectation, grounded in those views and in the benchmarks above, is that the period to 2030 will separate carriers on one variable: whether their appetite exists in a form their systems, their people and their delegated partners can apply consistently. Intake will be commoditized. Models will keep improving and keep failing on a particular carrier's proprietary logic. Supervisors will keep moving the question from "which model" to "which rule, which exception, whose authority." By 2028 we expect the first market conduct findings to turn on a carrier's inability to state which appetite rules and precedents an AI-assisted decision applied, and the first reinsurer and fronting-carrier terms to require a machine-readable appetite from MGAs. By 2030 we expect the appetite layer to be as ordinary a line in a carrier's underwriting architecture as the pricing model is today.
Building the appetite layer: four moves
The carriers closing the gap are doing four things, in roughly this order, each treated as a governed asset rather than a project.
1. Turn appetite statements into governed definitions
The appetite guide is written for people, once. Rewrite it, class by class, as definitions a system can evaluate: eligibility, hazard thresholds, limit bands, territory, exclusions, minimum terms and pricing floors, each with an owner, an effective date and a reason. UNDERWRITE's finding that models invent products the carrier does not offer is the direct argument, and the April paired benchmark shows the definitions do more for accuracy than the choice of model.12, 15
- Concrete marker: Every appetite rule has an owner, a scope, an effective date and a recorded rationale, and the current version is the one the intake and triage agents evaluate.
- Concrete marker: A submission outside appetite is declined or referred by rule, and the rule is cited in the decision record.
2. Capture referral precedent, not just referral outcomes
Referrals are where the real appetite is decided, and most carriers keep only the outcome. For each, record the fact pattern, the trigger, the decision, the conditions and the reason, so the next similar case can be matched against it. Use models to extract this from referral emails and workbench notes, then have the senior underwriters who made the calls confirm or correct it. That is the knowledge 44 percent of underwriters fear losing and the context on prior risks that 35 percent say they lack.9
- Concrete marker: Precedent is stored as a fact pattern, a decision and a reason, and an agent presenting a new submission shows the closest precedents and how the carrier decided them.
- Concrete marker: Senior underwriters approaching retirement spend their last year validating extracted precedents rather than writing a handover note.
3. Make authority explicit for underwriters and delegated partners alike
Authority matrices exist in every carrier and are almost never enforced by the systems that make decisions. Express each underwriter's, coverholder's and MGA's authority as limits a system can check: class, limit, attachment, territory, terms and deviation from technical price, with the referral path for each breach. Lloyd's DA Oversight Managers, the PRA's expectation of "strong oversight arrangements" for delegated business and the BMA's human approval points are asking for exactly this record.6, 21, 30
- Concrete marker: Every bound risk carries the authority under which it was bound, and breaches are detected at bind, not at the bordereau.
- Concrete marker: An MGA's binding authority is expressed as the same rule set the carrier's own agents evaluate, and the two are reconciled automatically.
4. Govern the appetite layer like a model
The NAIC bulletin, the EIOPA Opinion and the BMA note all extend existing governance to AI rather than inventing new regimes. Apply the same discipline to the appetite layer: an inventory with owners, independent review, monitoring of outcomes against the rules, change control and an audit trail. EIOPA's expectation of "reproducibility and traceability" is satisfied by this record and by nothing else.27
- Concrete marker: The appetite layer appears in the AI inventory with a named owner, a review cycle and a change log, and an examiner can be shown it.
- Concrete marker: Replacing the language model, the intake vendor or the workbench does not require rebuilding the carrier's definitions, precedents or authorities.
The leadership test
No carrier has finished this work. Leaders can locate their company by asking six questions:
- If our three most senior underwriters in a class left this quarter, what share of our appetite in that class would still exist in a form anyone could apply?
- When a referral is approved as an exception, do we record the fact pattern and the reason, or only the outcome?
- Could our intake and triage agents today state which appetite rule and which precedent they applied to a given submission?
- Is the authority we have delegated to an MGA expressed in the same terms our own systems check, and would we know of a breach before the bordereau?
- If an examiner used the NAIC evaluation tool on our underwriting AI, would we answer from the system or from interviews?
- As rates fall, do we know, by rule rather than by instinct, which business we should be losing to competitors?
Carriers that can answer these questions have an appetite layer, whether or not they call it that. Those that cannot are relying on judgment that is accurate, unwritten and retiring, in a market about to test it.
The carriers that come through this cycle with the best loss ratios will not be those with the most submissions processed or the largest model. They will be the ones that wrote their appetite down as governed definitions, kept their referral precedents as evidence, attached authority to both, and governed the result with the discipline they already apply to pricing models. They will quote faster, because fewer submissions need a referral to find out what the company thinks, and defend every decision, because the system can say which rule it applied. The technology has arrived. The work of making the appetite explicit has not been done.