Huzzle Labs

The underwriting AI benchmark

The underwriting track measures how language models do the core job of an underwriter: read the submission, weigh the risk, and decide whether to offer cover and on what terms. Every case is grounded in real application materials and scored pass@1 against the outcome the work resolved to.

Risk & terms
accept, decline, or refer — with limits and exclusions
Submissions
applications, schedules, and supporting files
pass@1
checked against the recorded decision

What the underwriting track tests

Underwriting is a judgement made under uncertainty, but it isn't a guess. An underwriter takes a submission, identifies the exposures that matter, checks them against appetite and guidelines, and produces a decision a file can be built on: accept, decline, or refer — and, if accepted, the limits, deductibles, exclusions, and pricing inputs that go with it. The track asks a model to do exactly that, from the same materials, and then checks the answer.

Because each case resolves to a recorded outcome, the benchmark rewards models that reach the right call for the right reasons, not models that produce a confident-sounding memo. That's what separates an underwriting AI benchmark from a general reasoning test: the answer is the decision, and the decision is checkable.

Why underwriting is hard for AI

The hard part isn't writing a plausible rationale — it's reaching the same call an experienced underwriter would, from messy inputs, under specific rules.

See how this work is done end to end in the underwriting workflow.

Where models slip
  • The signal is buried — the deciding fact is often one line in a long application.
  • Guidelines are specific — appetite and referral triggers are precise; a near-miss is a miss.
  • Documents disagree — submissions contain inconsistencies a model must resolve, not average.
  • Terms compound — a defensible accept is still wrong if the limits or exclusions are.

Example case types

Appetite

In or out of appetite

Decide whether a commercial property risk falls within appetite given construction, occupancy, and loss history.

Terms

Exclusions & conditions

Set the correct exclusions and conditions for a liability risk with a flagged prior claim.

Authority

Referral triggers

Identify the trigger that takes a case out of an underwriter's authority and into referral.

Pricing

Pricing inputs

Determine the pricing inputs that follow from the exposure data once the controlling guideline is applied.

How it's scored

Models run pass@1 — one attempt, no retries — and the response is compared to the recorded underwriting outcome. A decision is scored against the decision actually made; numeric terms against the recorded values within a defined tolerance. The full rules are in the methodology. Underwriting is one of three families in the InsureBench insurance AI benchmark, alongside claims and actuarial work.

Leaderboard opening 2026. Built by Huzzle Labs.
Get in touch about InsureBench