Skip to content
Book a free consult
Interface sounds
Professional Services Transformation concept 6 min read

Case study

AI Claim Triage From Photos: First Notice of Loss to a Reviewed Assessment

How a property claims operation turns a folder of unsorted phone photos and a 62-page policy into a structured damage assessment: every line citing the photo it came from, the coverage clause quoted, uncertain lines flagged, and an adjuster approving before anything moves.

No client engagement behind this piece: this is how we would transform this category of product, with benchmark-sourced targets.

Who this is for

A typical regional property claims operation

Industry

Insurance claims (property and casualty), 15-60 adjusters and claims support staff

Legacy stack

A claims management systemPolicy documents as PDFsPhotos emailed or uploaded by policyholdersEstimating software

Engagement

Concept

An unsorted grid of forty-one claim photos becoming a structured damage assessment where every line cites the photo it came from, with the coverage clause quoted and one uncertain line flagged.
Contents

An adjuster opens claim 20-8841 on Tuesday. The policyholder has uploaded forty-one photos from their phone: a flooded living room shot from four angles, a hallway ceiling, a close-up of a baseboard, three that are almost certainly the same wall. In another window sits a 62-page policy PDF. The job is to turn that into a defensible damage assessment with coverage checked. It takes hours, it cannot start until the claim reaches the top of a queue, and by the time it does, the policyholder has called twice.

This is the transformation we would run for that operation. It is presented as a concept, built on the same method we use in client engagements, so you can see exactly what the before, the after, and the path between them look like.

What does the manual claims desk actually cost?

The cost is not one dramatic number, it is the queue. J.D. Power’s 2025 U.S. Property Claims Satisfaction Study, released in March 2025 and based on 5,178 homeowner customers who had filed a claim in the previous nine months, put the average property claim at 32.4 days from filing to finished repairs and more than 44 days from first notice of loss to final payment: the longest since the study began in 2008. The same study found claims completed within 10 days scored 762 on its 1,000-point satisfaction scale, dropping 167 points to 595 when repairs ran past 31 days.

The honest reading of those numbers matters more than the numbers. Most of that elapsed time is physical: contractors, materials, and scheduling, none of which software touches. What software can touch is the desk time, the stretch where a claim waits in a queue and then consumes an adjuster’s afternoon in sorting, measuring, and policy lookup. That is a real slice of the calendar, and it is the only slice this work claims.

What does the claims desk look like before and after?

Drag the handle. The before is the recreated manual desk; the after is the triage copilot designed in its place.

The triage copilot
41 photos grouped by room · 4 damage lines drafted · 1 flagged

Claim 20-8841 · review and approve

Drywall, living room IMG_4471, 4473

Water line at 18 in · 3 walls

Laminate flooring IMG_4472

Buckling, full room

Baseboard trim IMG_4475

Swelling along north wall

Ceiling stain, hallway Confirm

Source unclear from photos

Coverage: water damage, sudden discharge · policy p. 14, clause 3(b)
Approve No payment moves without an adjuster
The claims desk today
Claim 20-8841 · FNOL Tuesday day 6, not yet assessed
Photos from the insured 41 files
IMG_4471.jpg
IMG_4472.jpg
IMG_4473.jpg
IMG_4474.jpg
IMG_4475.jpg
IMG_4476.jpg
No order, no labels. Which room is this one?
policy_HO3_2024.pdf p. 14 of 62
Estimate, typed by hand
Drywall, living rm ?? sq ft
Flooring measure again
Baseboard which photos?
Scroll back through 41 photos for each line
Coverage checked by flipping the PDF, from memory Next claim

The assessment is still an adjuster’s, checked against the same policy. What disappears is the sorting, the scrolling back through forty-one photos for each line, and the hours a claim spends waiting for someone to start.

What we built

The concept above is not a mockup exercise. It is the output of the same engagement steps we run on real companies:

  • We map one claim type first, usually the most common water or wind loss, because it repeats constantly and its desk time is measurable against a baseline the operation already tracks.
  • We build the extraction against the operation’s own closed claims: photos, adjuster notes, and final assessments become the evaluation set, so accuracy is scored on this book of business rather than on a vendor demo.
  • We require the model to cite its image: every drafted line names the photo it was read from, and the reviewer can open that crop. A line the copilot cannot tie to a specific image does not get drafted.
  • We ground coverage in the policy document rather than in the model’s memory, quoting the clause and page so the adjuster is reading their own policy language, not a paraphrase of it.
  • We deliberately left settlement and fraud referral un-automated. The copilot assembles evidence and quotes clauses; deciding coverage, valuing the loss, and referring suspected fraud stay with the adjuster, because those calls carry legal and regulatory weight that a confidence score does not discharge.
  • We designed the flag before the automation: anything ambiguous, like a ceiling stain whose source is not visible in any photo, is surfaced as a question rather than filled in with a plausible guess.

What that costs the operation is mostly one person’s attention, not a budget line. Expect roughly a day of a senior adjuster’s time to pick the closed claims that become the evaluation set and settle what a correct assessment looks like on each, a few hours from whoever administers the claims system to arrange read access to photos and policies, and about two hours a week from the same adjuster during the pilot to review flagged lines and correct drafts. Those corrections are not overhead, they are the eval data that decides whether the copilot earns wider scope.

How does it work?

The claim timeline split in two: the desk portion from first notice of loss through grouping photos, reading damage, checking coverage, and an adjuster decision is what this compresses. The physical repair that follows is untouched.

Photos and the policy arrive the way they already do. The copilot groups images by room and damage type, drafts damage observations with a confidence score on each field, and cites the specific photo behind every line. It checks the loss against the policy and quotes the governing clause with its page number. The adjuster opens a claim where the sorting is done, the obvious lines are ready to confirm, and their attention goes to the flagged ones: the ambiguous ceiling stain, the photo that does not match the reported loss.

The failure mode worth naming is the one vision has and text does not: a misread field looks exactly as confident as a correct one, and there is no sentence to quote back as proof. That is why the design pins every line to its source image and puts an adjuster in front of the assessment before it counts. When the copilot is wrong, the cost is a corrected line in a review queue, and that correction becomes an eval case that decides whether the copilot has earned wider scope.

The operation’s data stays in its environment, the claims management system remains the system of record, and nothing is used to train anyone else’s models.

Why does this pay back?

For a claims operation the return lands on the two things the business is judged on. Cycle time improves at the only point software can improve it, by removing the days a claim spends waiting for an adjuster to have an afternoon free. And the assessment that results is better evidenced than the one it replaces: every line traceable to an image, every coverage call quoting the clause it rests on, which is worth as much in a dispute as it is in a fast month. The first slice is also a wedge, because the same extraction engine extends from water losses to wind, fire, and contents claims, one measured claim type at a time.

The honest limit: this pays back where photo-heavy, high-volume, relatively standard losses dominate the book. An operation whose claims are mostly large, complex, or litigated will find the copilot organizing evidence and little more, because the expensive judgment in those files was never the sorting. If policy documents are inconsistent across a legacy book, the first weeks are document work before they are AI. And an operation unwilling to keep an adjuster in the approval loop should not run this design at all: the whole safety argument depends on that gate.

The outcomes

Average property claim, filing to finished repairs

32.4 days1

Industry benchmark

Satisfaction drop when repairs pass 31 days

-167 pts2

Industry benchmark

Desk time from photos in to assessment ready

Days in a queue, then hours per claim Drafted on arrival, adjuster-reviewed

Assessment waiting, not started3

Design target

Time to the first shippable slice

4-6 wk4

Design target

1 J.D. Power 2025 U.S. Property Claims Satisfaction Study, released March 18, 2025, based on 5,178 homeowner insurance customers who filed a claim in the previous nine months, fielded January through December 2024. The same study puts first notice of loss to final payment at more than 44 days, both the longest since the study began in 2008.

2 J.D. Power 2025 U.S. Property Claims Satisfaction Study: claims completed within 10 days score 762 on a 1,000-point scale, falling 167 points to 595 when repairs take more than 31 days.

3 Design target for the triage copilot, measured against the recreated manual desk shown in the before/after above. Covers the desk portion of the claim only, not physical repair time.

4 Standard first-slice scope: one claim type, one metric, evals and an approval gate included.

Frequently asked questions

Can AI actually assess property damage from a policyholder's phone photos?

It can draft an assessment, not settle a claim. The copilot groups the photos by room, extracts damage observations with a confidence score per field, and cites the specific image behind every line. Anything it cannot read confidently, like a ceiling stain whose source is not visible, is flagged rather than estimated. An adjuster reviews and approves before the assessment counts for anything, which is the design, not a disclaimer.

Does this speed up claim settlement, or just the paperwork?

It compresses the desk time, which is one part of the claim and not the whole thing. J.D. Power's 2025 study put the average property claim at 32.4 days from filing to finished repairs and more than 44 days from first notice of loss to final payment. A large share of that is contractors, materials, and scheduling, which no software touches. What this addresses is the stretch where photos sit in a queue and an adjuster then spends hours sorting them, measuring, and flipping through a policy PDF.

Our claims decisions have to be defensible in a dispute. Does an AI assessment survive that?

That requirement shapes the whole design. Every drafted line carries the image it was read from and the policy clause it was checked against, so the audit trail is stronger than a typed estimate with no provenance. The adjuster's approval, edits, and overrides are all recorded. The copilot never makes the coverage decision; it assembles the evidence and quotes the clause so the person deciding can see exactly what they are deciding from.

What about fraud, staged damage, or photos that are not from this loss?

This design does not claim to detect fraud, and any vendor promising that from photos alone deserves scrutiny. What it does is make inconsistencies visible: images whose content does not match the reported loss type, duplicates across claims, and timeline gaps get surfaced for a human to judge. Fraud referral stays a trained adjuster's call, made with better-organized evidence in front of them.

What would trying this cost?

A scoped proof of value on a single claim type, usually the most common water or wind loss, priced as a fixed first slice. It ships in weeks with its own desk-time metric measured against today's baseline, so the decision to go further rests on a measured result rather than a promise.

Go deeper on the method

A spectrum from suggest to autonomous, with a human and AI handoff.
Human+AIWorkflow

Human-in-the-Loop, by Design, Building AI People Actually Trust

Autonomy is easy to demo and hard to ship. We design Human+AI workflows where the handoffs, guardrails, and overrides are first-class, so teams adopt them.

Read article
A chat bubble splits into two paths: one grounded and checked, one an AI hallucination marked with a warning.
AI ProductTrust

Spot the AI Hallucination: Can You Tell When It's Making It Up?

An interactive quiz. Judge 5 AI answers as grounded or hallucinated, learn why you often can't tell from the text, and what actually catches it.

Read article
A phone photo of a cracked machine bracket on the left, and on the right the structured facts read out of it: part, condition, severity, and one low-confidence field flagged for a person to confirm.
AI ProductWorkflow

When Your AI Should Look, Not Just Read: A Practical Guide to Multimodal at Work

Most business AI reads text, but a lot of what your business actually knows arrives as pixels: damage photos, whiteboards, scanned forms, screenshots. A short quiz sorts your workflow into the three honest tiers, and the failure modes of vision are not the ones you are watching for.

Read article

More transformations

A crowded RFQ inbox queueing on the left transforming into a priced, stock-checked quote with one flagged line and a rep approval button.
Wholesale & Distribution Concept

AI Quote Desk for a Wholesale Distributor: RFQ Inbox to Priced Quote in Minutes

How a parts distributor replaces the morning RFQ pile and the three-system price hunt with a copilot that extracts every line item, prices it from the company's own item master and contracts, checks stock, and drafts the reply, with a rep approving every quote before it leaves.

60-70%

Share of routine admin work generative AI can absorb

A pile of 300 applications screened by a keyword filter that drops good candidates with no reason recorded, next to the same pile summarized and ranked against the real criteria with a recruiter approving each advance and nobody auto-rejected.
Staffing & Recruiting Concept

AI Screening for a Recruiting Team: Read Every Application, Auto-Reject No One

How a staffing team buried in applications gets an AI screen that reads and ranks every candidate against the real job criteria, surfaces the people a keyword filter would miss, and logs every decision, while a recruiter advances or declines each one and nothing is rejected automatically.

~11,000/min

Get the next transformation in your inbox.

When we publish something worth your time, you will be first to know. No spam, unsubscribe anytime.

Your version of this

Run a document-heavy practice where errors are expensive? We will mapwhich workflow pays back first in a free consult.

We design and build every engagement ourselves. No juniors, no handoffs.

Free, no obligation Your data stays yours Reply within 1 business day