Start with a worked example you can check. Then choose a model class, adapt a prompt template, and build a habit of verifying the result against its source.
Your first task: extract, then check
Use an approved tool and the fictional notes below. You need no company data or model comparison to try this exercise.
Copy this prompt and its sample notes
Extract the action items from these fictional meeting notes into a table with columns: action, owner, due date, supporting text. Use only the notes. Write "not stated" for missing owners or dates. Separate decisions from proposals, and do not invent deadlines.
Notes: Priya will check whether the submission checklist includes roof age by Friday. We agreed to keep the current checklist until that review is complete. Someone suggested asking brokers for photographs, but no decision was made and no owner was assigned.
Check the result before accepting it
The assigned action is checking the checklist. Its owner is Priya and its due date is Friday; no calendar date is given.
Keeping the current checklist is a decision. Asking for photographs is an undecided proposal, with owner and due date not stated.
Every supporting passage must appear in the notes. An invented task, owner, date, or quotation is a failed check.
Save the prompt, output, and corrections. Try another fictional example with a missing owner or conflicting dates before adapting the templates below. For a team pilot, record time spent prompting and reviewing against doing the task manually.
For a team pilot: record task volume, accepted results, material errors, review time, and total cost. Bring that evidence to the integration decision before expanding access or automation.
Choose a model class before you build
Use three working categories to compare capability, speed, and cost: economical, balanced, and frontier. They are a selection aid, not an industry standard or a promise that providers offer equivalent models. Define the task's quality bar, latency, and volume first, as recommended in Anthropic's model-selection guidance.
Start with the task, not the model.
Define the required artifact, evidence, acceptable error rate, response-time target, volume, and human review point.
↓
What quality and response time does this step require?
Use the lowest-cost tier that clears the quality bar on representative examples.
Fast / economical
Routine · high volume · low latency
Evaluate for classification, routing, schema-constrained extraction, simple summaries, document indexing, and first-pass data checks. Escalate failed validation and inputs outside the tested pattern.
Balanced reasoning
Candidate for multi-step knowledge work
Use for underwriting briefs, grounded RAG answers, document comparison, code generation, multi-step analysis, and most tool-using workflows. Test these tasks against representative inputs and review criteria.
Frontier
Novel · complex · difficult exceptions
Reserve for difficult synthesis, complex coding or debugging, adversarial review, ambiguous exceptions, and hard actuarial or coverage analysis. It still needs evidence and human approval.
Then evaluate and route.
Start fast when the task is routine and the quality bar is measurable. Start capability-first when the task is novel or failure is expensive, then test whether a lower tier clears the same evaluation. Do not use one model tier for every step just because it worked once.
Tier
Anthropic
OpenAI
Google
Fast / economical
Claude Haiku 4.5
GPT-5.6 Luna
Gemini 3.5 Flash-Lite
Balanced reasoning
Claude Sonnet 5
GPT-5.6 Terra
Gemini 3.8 Flash
Frontier
Claude Opus 5
GPT-5.6 Sol
Gemini 2.5 Pro
Selected examples checked on 9 September 2026 against the providers' catalogs (Anthropic, OpenAI, Google). This table is illustrative, not a performance ranking or a complete shortlist. The catalogs also list Claude Fable 5.1 and GPT-6 Astra for more demanding work; Google lists Gemini 3.1 Pro as a preview. Confirm availability, price, and lifecycle status before choosing.
Route by step, not by workflow
A submission workflow might use a fast model to classify attachments, a balanced model to create the underwriter brief, and a frontier model only when documents conflict or the case needs deeper analysis. Define escalation using failed validation, missing evidence, or inputs outside the evaluated pattern. Do not use the model's own confidence statement as the sole routing signal.
Non-negotiable rule: model tier does not change authority. A frontier model may be more capable, but it does not get to quote, bind, set reserves, change a rate, or file a document without the same permissions, deterministic controls, and human gate. For the implementation pattern, see the capability ladders' harness section.
These categories are provider-neutral, and the named examples above are a dated snapshot rather than a recommendation. The capability, speed, and cost framework is reflected in Anthropic's model-selection guidance; use each provider's current documentation and your own evaluations to select the actual model.
Beyond the prompt
For recurring team work, specify the documents the system can use, the actions it may take, and how the result will be verified. The harness section explains how to make those controls repeatable.
Anatomy of a good prompt
Illustrative prompt · renewal preparation
Make each instruction do a specific job
Read the prompt fragments in order. The annotations show what each one makes possible to review.
What you writeWhat it establishes
1Prepare a draft for a commercial lines underwriter reviewing a renewal.
Role and audienceSets the perspective. The underwriter still owns the decision.
2Summarize the losses, open large claims, and missing information relevant to the review.
TaskNames the work and the omissions a reviewer should look for.
3Use only the attached approved loss run. Cite the source page or row for each point.
Source materialMakes the evidence traceable to a specific document.
4Return a short file brief with a claims table and a separate list of questions for the underwriter.
FormatDefines the artifact and separates recorded facts from open questions.
5Write “not stated” for missing values. Preserve the valuation date and currency. Flag totals that have not been reconciled.
ConstraintsSpecifies how to handle gaps without inventing an answer.
Acceptance checkOpen the cited passages, reconcile the numbers, and resolve material gaps before accepting the draft.
Ten techniques that do most of the work
Paste the source material. Never ask the model to recall a policy form, regulation, or account from memory; give it the document and ask it to work from that. Check that the answer actually follows the supplied source.
Assign a role. Name the intended perspective and audience, such as a reinsurance treaty reviewer. A role instruction does not confer professional expertise.
Specify the output format. Table vs. prose vs. bullet brief; word limits; audience ("explain for a board member" vs. "for an actuary").
Show an example. If you want a specific style (a triage note, a file summary format), paste one good example and say "match this format." Check that the example matches the task and uses data permitted in the tool.
Ask for a checkable explanation. Request the supporting passages, assumptions, calculations, and unresolved questions. Use an available reasoning mode when evaluation shows it helps; OpenAI's reasoning guidance advises against requiring step-by-step internal reasoning.
Iterate. The first output is a draft, not a verdict. "Shorter." "More formal." "You missed the 2023 claim; redo with that included." Check that each revision resolves the identified defect.
Use it as a critic, not just a drafter. Paste your own memo or analysis and ask for a skeptical peer review. Review the criticism against your evidence before accepting it.
Break big tasks into steps. Separate extraction, reconciliation, analysis, and drafting so you can review the result at each stage.
Tell it what to do when unsure. Add "If the information isn't in the document, say 'not stated'; do not guess." Then check missing fields against the source; the instruction is not a guarantee.
Start fresh chats for new topics. For a new task, start with a clear brief and only the relevant documents. Remove obsolete instructions when continuing an existing conversation.
Copy-paste templates for insurance work
Examples use P&C documents; adapt them to your line of business. Use only data permitted in your approved tool under your organization's data rules. Treat every output as a draft for review.
Summarize a claim file
Prepare a draft summary for the claims handler using only the file below: (1) a brief overview, (2) coverage questions raised in the file, (3) recorded reserves and exposure estimates, with their dates and sources, (4) missing or conflicting information, (5) questions for the handler. Cite the page or passage supporting each point. Use "not stated" for missing information. Do not decide coverage or recommend a reserve change.
[paste approved file]
Extract a loss run to a table
Extract the claims from the loss run below into a table with columns: claim number, loss date, line of business, cause of loss, status, paid, incurred, source page or row. Preserve currency and valuation date. Use "not stated" for missing values; do not guess or infer. List uncertain entries and any reconciliation differences. Do not claim completeness unless the extracted record count and totals have been checked against the source.
[paste loss run]
Compare policy wordings
Prepare a draft comparison of the two endorsement wordings below. List differences you identify in coverage grants, exclusions, conditions, and definitions, with the supporting clause from each wording. Separate textual differences from their possible implications for [describe the risk]. Flag uncertainty and missing context for a qualified reviewer; do not give a final coverage determination.
[paste wording A] / [paste wording B]
Draft a communication
Draft an email to a retail broker declining the [risk type] submission for [named insured] because [reasons]. Professional and warm, under 150 words, keep the door open for other business, and do not promise to reconsider this risk.
Red-team your own work
Below is my draft analysis. Act as a skeptical peer reviewer: list the strongest objections, the weakest assumptions, anything missing, and alternative interpretations of the data. Rank by importance. Do not compliment the work.
[paste analysis]
Understand or write code (actuaries/analysts)
Explain what this [R/Python/SQL/VBA] code does, section by section, then flag any bugs, silent assumptions, or edge cases that could produce wrong numbers.
[paste code]
Turn a meeting into actions
From the meeting notes below, produce: decisions made, action items (owner + due date if stated), and open questions. Don't invent owners or dates that weren't stated.
[paste notes]
Habits of effective users
Verify anything that leaves your hands: numbers, quotes, citations, policy-language references. LLMs fabricate plausible-looking citations; check every one against the source.
Numbers → code, not mental math. For quantitative work, request a formula or code that you can run and inspect. Verify inputs, units, aggregation rules, and results; generated code can contain errors too.
Review the whole draft. Check the final document and its decisions, including details that initially looked correct.
Check important answers against sources. Repeating a question can reveal variation; a repeated answer still needs evidence.
Build a prompt library. Save useful prompts with sample inputs, expected results, and known failure cases. Recheck them when the model or task changes.
Give it your standards. Provide your team's checklist, style guide, or approved example, then check the output against it.
Use it to learn. Ask for an explanation and a quiz, then check the explanation against an authoritative reference before applying it at work.
Respect the data rules. Sanctioned tools only; no client or confidential data in personal accounts. No exceptions. See the policy template in Governance.
Common mistakes and their fixes
Mistake
Fix
Vague ask ("thoughts on this?")
State the task, audience, and format you want
Asking from memory ("what does ISO CG 00 01 exclude?")
Paste the actual form and ask about that document
Accepting the first draft
Check against acceptance criteria; revise until defects are resolved or escalate
Treating it as a search engine
Provide the source or use an approved search-enabled tool, then verify the retrieved evidence
One giant prompt for a complex job
Break into stages; review between them
Trusting confident citations
Open each citation and check the claim; confident wording is not evidence of accuracy
Letting long chats drift
New topic, new chat
Blaming the model for a vague prompt
Re-read your prompt as if you were a new hire receiving it: could you deliver from that instruction?
Every fix above is still one person improving one prompt. Once a task recurs often enough that its wording should stop varying, it stops being a prompting problem: Agentic work covers delegating the job itself, and the capability ladders' harness section is the build manual for making the result repeatable, with context specs, hooks, golden datasets, and eval gates.