Prompted LinesAI guidance for insurance

Strategy · Procurement checklist · ~8 min

Evaluating AI vendors

A neutral checklist for evaluating AI vendors for specialty and commercial insurance: evidence, data, controls, integration, and total cost.

This checklist is for a leader, procurement team, or risk function assessing an AI product for specialty or commercial insurance before committing budget. It does not review or rank products. The same questions apply to a new supplier, an AI feature added to software you already license, and an internal build that relies on a model provider.

Leadership decision

Before the first demonstration, decide which workflow the product would change, who owns its output, and what evidence would justify expanding it. Evaluate the product against that decision, not against the demonstration.

Start from the workflow and its baseline

Treat a demonstration as a sample of material the vendor chose. The procurement question is narrower: does it improve one named workflow, on your files, at a cost and risk you accept? Record current volume, handling time, quality standard, and cost first, using the pilot decision sheet. Without that baseline, a vendor's result cannot be compared with the work it would replace.

State what the output will do: inform a person, produce a draft for approval, or trigger an action. That determines how much evidence and control the rest of this checklist requires.

Evidence to request

Vendor case studies report results from particular deployments, as described by the vendor or its client. The integration phases page treats such figures as evidence about their stated scope, not as audited or transferable benchmarks. Before procurement, request:

  • Results on comparable files: the same line of business, document types, and data quality as your own book. A test on a sample of your historical files is stronger evidence than a case study.
  • A definition of each metric: what was counted, over what period, and against which baseline. A high average accuracy can conceal serious errors in a small number of important fields (see validation).
  • The cost of exceptions and human review: the share of outputs that need correction, and reviewer time per file.
  • Reference customers running a comparable workflow, whom you can ask about implementation effort and support.

Data handling

The Costs & value page lists the enterprise terms to review: training restrictions, approved retention settings, audit access, and processing locations. Confirm them for the specific product and account, reading the vendor's trust documentation alongside the proposed contract (Resources explains this step).

Ask whether your data is used to train or improve models, how long inputs and outputs are retained, where they are processed, and who at the vendor or its suppliers can access them. Treat outputs derived from confidential inputs as confidential.

Model and change management

The site's validation guidance recommends re-measuring performance on every model version change, prompt change, or vendor update, pinning versions where supported, and re-running evaluations before accepting a change.

If the product relies on an underlying model that the vendor or its provider updates, ask how customers are notified of model and prompt changes, whether you can hold a version while you re-test, and how the vendor monitors output quality for drift over time.

Controls

  • Review proportional to consequence. Set the depth of review by how sensitive the data is and what the output affects. Reviewers need the source evidence behind each output, not only the output.
  • Least privilege. Give the product only the access its task needs. As the Agentic work page advises, enforce permissions outside the prompt and test that the application rejects unauthorized actions.
  • Logging. Check that the product records source references, model and workflow versions, outputs, and reviewer decisions, and that you can export those records.
  • External documents. Broker submissions and claimant correspondence can carry prompt-injection payloads, as the technical deep dive explains in its data-handling rules. Ask how the product handles them, especially if it can take actions.
  • Escalation and pause. Confirm that a named person can pause the use, withhold affected outputs, and restart only after revalidation.

Regulatory and third-party oversight

As the Governance page notes, the insurer remains accountable when it relies on a vendor. The NAIC Model Bulletin says an insurer's written AI Systems Program should address how it acquires, uses, or relies on AI systems developed by third parties. The considerations it lists may include, as appropriate, due diligence on the third party and its systems; contract terms, where appropriate and available, that provide audit rights or audit reports and require cooperation with regulatory inquiries; and activities to confirm the third party's compliance. It also says that, where an investigation or examination concerns third-party data, models, or AI systems, an insurer should expect the department to request due diligence records, contracts, audit or confirmation work, and validation, testing, and auditing documentation, including evaluation of model drift.

Adoption of the model bulletin varies by state; check each state's bulletin and other guidance individually. Have Legal/Compliance map the requirements to the proposed use. The Governance page summarizes the wider landscape.

Integration and exit

A vendor product may need substantial integration, and an internal build needs ongoing ownership; the integration phases page recommends comparing both on total cost, fit, support, and results on representative cases.

Ask which interfaces connect the product to your policy, claims, and document systems, and who maintains them. Confirm that you can export outputs, logs, and reviewer decisions in a usable format, and how your data is deleted on termination.

Total cost, including review and rework

License and usage fees are one part of the cost. The Costs & value page recommends adding implementation, security review, evaluation, training, and ongoing ownership to current quotes, and measuring cost per accepted task and net staff time released, including review and rework. Failed output that is regenerated is additional spend.

Ask the vendor to price a downside case with lower volume and more exceptions, as the pilot decision sheet recommends before expansion.

A checklist to copy

Use one row per area in a shortlist meeting. The red flags are prompts for follow-up questions, not automatic disqualifiers.

Ask the vendorEvidence to request · red flag
WorkflowWhich workflow and decision does the product support, and what does its output do?A written scope matching your pilot decision sheet.Red flag: results shown only in a general demonstration.
ResultsHow does it perform on files like ours?Results on a sample of your historical files, with metric definitions.Red flag: only case studies or an undefined accuracy figure.
Review costWhat share of outputs need correction, and how long does review take?Exception rates and reviewer time from comparable deployments.Red flag: savings figures that omit human review.
ReferencesWho uses it for comparable work?Reference customers you can contact.Red flag: references only from unrelated lines or tasks.
DataTraining use, retention, processing location, audit access?Contract terms and trust documentation for this product and account.Red flag: terms that cannot be confirmed in the contract.
ChangeHow are model and prompt changes notified and tested?Change-notice process, version pinning, drift monitoring.Red flag: updates applied with no chance to re-test.
ControlsHow are permissions, logs, external documents, and pausing handled?Permission model, sample logs, injection testing, pause procedure.Red flag: broad default access or logs you cannot export.
OversightWill the vendor support due diligence, audits, and regulatory inquiries?Audit rights or audit reports; a cooperation clause.Red flag: no commitment to support regulatory inquiries.
ExitHow do we leave?Export formats, termination and deletion terms.Red flag: outputs or records held in a form you cannot take with you.
CostWhat is the full cost at our volume?Quotes for seats, usage, implementation, and support; a downside case.Red flag: pricing that excludes exceptions or usage growth.
Next steps

Take a shortlisted product into a bounded pilot with a named owner, a baseline, and stop conditions, then decide at the pilot gate whether to expand, revise, or stop. Record the vendor diligence and approvals in your AI inventory and governance program; the Governance page includes a policy template.