Tokens are the unit of cost, and agents spend them like a fleet, not a person. Every call pays for the context going in and the text coming out; an agentic workflow makes many calls per task; a fan-out makes many calls per second. None of this is an argument against agents. It is an argument for engineering the cost the way you engineer the quality.
1 Β· Token economics: where the money actually goes
Four drivers set the cost of an agentic workflow, and none of them is the model's sticker price:
- Context per call. An agent that reads a 40-page submission packet per step spends those input tokens on every step. Context discipline (send what the step needs, not the file room) is the biggest lever most teams miss.
- Calls per task. A five-stage pipeline makes five-plus calls; a retry loop makes more. Orchestrated agent teams have been measured at roughly 15x the token cost of a single agent (Anthropic's multi-agent research write-up); fan-out multiplies spend roughly linearly.
- Model tier per step. Frontier models cost several times more than mid-tier ones. Extraction and classification rarely need the frontier; judgment steps sometimes do. Route by step, not by habit.
- Rework loops. Output that fails review gets regenerated, and regeneration is spend. Eval gates that catch slop before it reaches a human also stop the cost of cycling it.
Budget anchors. Enterprise LLM agreements at mid-size scale run roughly $250k to $1M+ a year, with 25β40% discounts typical at the $1M tier (benchmark guide; directional, verify at procurement). Carrier context: IT spend averages ~4.5% of GWP (Datos Insights), and two-thirds of insurance CEOs plan to allocate 10β20% of budget to AI (KPMG CEO Outlook, PDF). Vendor purchases succeed about 67% of the time versus 33% for internal builds (MIT, PDF), which is an argument for buying the commodity layers and reserving build spend for what differentiates you.
Measure cost per completed task and dollars per underwriter hour returned, not price per token. A use case that cannot clear both bars is a demo. The winning comparison is machine cost per submission versus loaded human cost per submission reviewed, and it usually is not close.
2 Β· The operational realities the demo never shows
| Consideration | What it looks like in production | The practical rule |
|---|---|---|
| Latency and broker SLAs | Multi-step agents take seconds to minutes; brokers notice turnaround, not your architecture | Put speed where the broker sees it (acknowledgment, triage, appetite answer) and depth behind it |
| Enterprise terms | Zero data retention and no-training clauses, audit logging, regional processing | No enterprise agreement, no company data. This is the first control in the field guide's governance level |
| Model deprecation and drift | Vendors retire and upgrade models; behavior shifts silently under a pipeline that passed its evals in March | Re-run the golden dataset on every version change, and pin versions where the provider allows it |
| Rate limits and surge | A CAT event is a volume spike; provider rate limits are a hard ceiling | Capacity-plan for surge events, and keep a queue-and-degrade path that keeps humans working when the cap is hit |
| PII and residency | Claimant and policyholder data in prompts and logs; state and partner rules on where it may be processed | Minimize and de-identify by default; log retention is a compliance surface, not an IT detail |
| Vendor viability | 95% of H1 2026 insurtech funding went to AI startups (funding data), which means consolidation is coming | Due diligence on funding and runway; an exit plan for any vendor whose output feeds a regulated decision |
| Lock-in and portability | A harness written to one provider's quirks is a migration project later; both core vendors shipped agentic frameworks in 2026 (Guidewire, Duck Creek) | Keep context specs, evals, and hooks model-agnostic; they are your portable assets |
3 Β· The most impactful use cases for insurance companies, ranked
Ranked by measured value divided by implementation risk, from the carrier and vendor results behind the roadmap (verbatim source passages for the top entries are on the evidence page).
| # | Use case | Why it pays (measured) | Where it sits |
|---|---|---|---|
| 1 | Submission intake and triage | 2β5x underwriting speed and 370k+ submissions a year at AIG; Markel's 113% productivity uplift; 50β97% faster processing and +15% hit ratios at Sixfold customers (evidence) | Underwriting; the proven first move |
| 2 | Document extraction and summarization | Loss runs, SOVs, claims files: routine 50β80% time cuts on document-heavy work; the foundation every other use case reads from | Everywhere; assistive, low scrutiny |
| 3 | Claims triage, severity and litigation prediction | Attorney-involved claims cost ~4.9x more (CLARA data); early triage moves both cycle time and indemnity | Claims; decision support with human authority |
| 4 | Bordereaux processing | 85β94% processing time savings (Verodat); the unglamorous pain point of program business | Program/delegated-authority operations |
| 5 | Fraud detection | 5x more fraud detected at Tokio Marine (case study, vendor-reported) | Claims; scoring with human review |
| 6 | Knowledge access (RAG over guidelines) | Appetite and procedure answers in seconds; multiplies every other use case by keeping context current | Enterprise-wide; internal only |
| 7 | Actuarial filing research and pricing workbench | Filing research from weeks to hours (Akur8); 13 pricing tools in 13 weeks at Allianz Commercial (hx) | Actuarial/pricing; assistive |
| 8 | Leakage and subrogation | AI pre-payment controls prevent 90β95% of detectable leakage (analysis); $15β20B of subrogation goes uncollected annually (industry estimate) | Claims finance; Phase 3 material |
| 9 | Bounded agentic quoting | 3 days to ~3 minutes at Hiscox London Market; CFC's agentic pilot (evidence). Real, but gated: human authority, governance gates, bounded segments only | Underwriting; last, under Phase 3 gates |
Sequencing rule: start assistive (ranks 1β2), move to decision support with human authority (3β5), and reach bounded automation last (9). That is the roadmap's phase logic applied to a single function; the harness section of the practice ladder is the build manual for each step.