How InsureBench scores AI
InsureBench is an insurance AI benchmark built on one principle: a model should be judged on the same work a practitioner is. Every case is grounded in real documents and resolves to a single answer checked against a recorded outcome — not a rubric, not a vibe.
Why insurance needs its own AI benchmark
The stakes are large and the adoption is real. McKinsey estimates AI could add up to ~$1.1 trillion in annual value to global insurance, with generative AI alone potentially unlocking $50–70 billion in revenue.1 A 2024 Deloitte survey found 76% of insurers had deployed generative AI in at least one function,2 yet BCG reports only about 7% have scaled it to production.3 The gap between deploying and trusting is exactly what a benchmark should illuminate.
Most AI benchmarks measure general reasoning, coding, or exam recall. Insurance work is different in kind. It turns on reading long, inconsistent policy documents, locating the clauses that control a question, reconciling files that don't agree, and arriving at a decision or number an auditor could check. A model can be fluent and still get the coverage call wrong — and in insurance, the call is what matters.
InsureBench measures that specific competence: an AI benchmark for insurance that scores models on underwriting, claims, and actuarial tasks as they're actually performed.
How a case is built
Each case starts from a real piece of insurance work and is rebuilt into a self-contained, checkable task.
Source documents
The policy wording, application, schedules, endorsements, or claim file a practitioner would have.
A precise question
Accept or decline, covered or not, or a number such as a reserve or payable amount.
Recorded outcome
The answer the work actually resolved to, verified by a practitioner.
pass@1 grade
One model attempt, compared to the recorded outcome — right or wrong.
Cases are written and reviewed with practising underwriters, claims handlers, and actuaries so the documents are realistic and the recorded outcome is defensible. Sensitive material is removed or synthesised so a case can be released without exposing private data, while keeping the structure and difficulty intact.
How models are scored
Every case resolves to a single verifiable answer. Models run pass@1: one attempt per case, no retries, no best-of-N. The response is compared to the recorded outcome — an accept/decline decision, a coverage determination, or a number within a defined tolerance. Scoring reflects the outcome, not the style, length, or confidence of the writing. A well-argued wrong answer scores the same as a terse one: zero.
Reporting it this way keeps the leaderboard honest. Because each answer is checkable, scores aren't a matter of judgement, and a model can't earn credit for sounding authoritative. It either reached the recorded outcome or it didn't.
What the benchmark deliberately rewards
Document grounding
Finding and applying the controlling clause rather than answering from prior knowledge.
Faithful to the terms
Following the policy as written — including exclusions and conditions — not a reasonable-sounding approximation.
Numerical discipline
Carrying the right tables, assumptions, and arithmetic through to a defensible figure.
Restraint
Declining to invent facts the documents don't support.
A GDPval-style benchmark for insurance
InsureBench follows the spirit of GDPval, OpenAI's 2025 benchmark that evaluates models on real, economically valuable occupational tasks across 44 occupations, judged against expert deliverables.6 Where GDPval spans many occupations, InsureBench goes deep on one industry — built around the tasks underwriters, claims handlers, and actuaries are paid to get right.
win-or-tie
Related research
InsureBench builds on a small but growing body of work evaluating language models on insurance tasks. These are prior efforts we draw on, not competing products.
UNDERWRITE
The most directly related work: an expert-built, multi-turn agentic underwriting benchmark over 13 frontier models, documenting domain-knowledge hallucination despite tool access.
INS-MMBench
The first hierarchical multimodal insurance benchmark, spanning auto, property, health, and agricultural insurance across 22 fundamental tasks.
InsQABench
A Chinese insurance QA benchmark across commonsense knowledge, structured databases, and unstructured documents.
INSEva
A comprehensive Chinese insurance LLM benchmark of 38,704 examples scoring both faithfulness and completeness.
InsureBench's distinction is its GDPval-style, occupational framing for Western insurance practice: document-grounded cases that resolve to a single verifiable outcome, scored pass@1, across underwriting, claims, and actuarial work.
Sources
- McKinsey & Company — The future of AI in the insurance industry. mckinsey.com
- Deloitte — Scaling generative AI in insurance (2024). deloitte.com
- BCG — Insurance leads AI adoption; now is the time to scale (2025). bcg.com
- Stanford HAI — Hallucinating law: legal mistakes with LLMs are pervasive. hai.stanford.edu
- Dsouza et al., Snorkel AI — Benchmarking Agents in Insurance Underwriting Environments (arXiv 2602.00456, 2026). arxiv.org
- Patwardhan et al., OpenAI — GDPval (arXiv 2510.04374, 2025). arxiv.org
- OpenAI — Introducing GDPval. openai.com
- Jin et al., Fudan University — INS-MMBench (arXiv 2406.09105, ICCV 2025). arxiv.org
- Ding et al. — InsQABench (arXiv 2501.10943, 2025). arxiv.org
- INSEva — A comprehensive Chinese insurance LLM benchmark (arXiv 2509.04455, 2025). arxiv.org