Skip to main content
Speculative Design / Agentic AI · 2025

When AI Acts Autonomously

What an adjuster does when the AI already decided 500 times today

  • Speculative Design / Agentic AI
  • Concept Designer & Researcher
At a glance
Role
Concept Designer & Researcher
Year
2025
Category
Speculative Design / Agentic AI

The findings are the argument. Everything below them is evidence.

Verdict first — for people with forty portfolios open

AI in insurance today mostly advises. It puts a recommendation on a screen and a person acts on it. This is a speculative look at the step after that — where the system approves the routine payout itself, and a human only sees the ones it wasn't sure about. The interesting problem is not the AI. It is what oversight has to look like when nobody reviewed the decision before the money went out, and a regulator asks you to account for it eighteen months later.

oversight interfaces

3

oversight interfaces

auto-resolve target

85%

auto-resolve target

review time, targeted

< 4min

review time, targeted

The brief — what this solves, and for whom

What an adjuster does when the AI already decided 500 times today

The design problem

What made it hard.

A mid-size insurer processes roughly 50,000 claims a month. 70% are straightforward — water damage, minor collisions, routine medical — yet adjusters spend most of their time on exactly these cases, leaving too little attention for the complex claims that actually need expert judgment. Mapping the current workflow (7 manual steps, 7-10 days to resolution) showed three natural zones of AI autonomy: Autonomous (~70% of claims, AI handles it end-to-end), Assisted (~20%, AI flags one decision point for a human), Escalated (~10%, AI organizes the data, human decides). The harder design question wasn't the AI — it was how an adjuster monitors 500 autonomous decisions a day without drifting into rubber-stamping, and what happens when the AI gets one wrong. > Users don't need to watch every AI decision. They need confidence the system is functioning correctly, and immediate access when something goes wrong.

The approach

What I did about it.

Three interconnected interfaces, each for a different cognitive mode: The Agent Dashboard — a control-room view of what the AI is doing in real time, closer to air traffic control than a claims spreadsheet. The Exception Queue — when the AI isn't confident, it routes to a human with its analysis and the specific source of uncertainty already highlighted, not raw data to investigate from scratch. Cuts average review time from 45 minutes to under 4. The Audit Trail — every autonomous decision logged with full reasoning chains, so any payout traces back to its inputs and logic. Auditability isn't a feature added on top; it's the architecture. The hardest problem wasn't any of these interfaces — it was automation complacency: once AI handles 85% of claims, humans reviewing the rest start rubber-stamping. The fix was 'attention anchors' — a source-document confirmation checkbox plus randomized verification prompts, just enough friction to interrupt autopilot without hurting throughput.

The findings

Finding 01

I designed the failure mode first

What I found

Before working out how a claim gets processed correctly, I worked out what happens when one goes wrong: how the error surfaces, how the policyholder is made whole, what the audit trail has to preserve for the accountability to survive. Starting there produced a different architecture than starting from the happy path and patching would have.

Finding 02

The on-off framing was the thing to reject

What I found

Stakeholders arrived wanting to know whether the AI was autonomous or not. It has to flex with claim complexity, adjuster experience, regulatory context and the model's own confidence — so the three zones exist to make that flex structural instead of something negotiated case by case.

The log

How it actually went, in order.

Loops, not phases. The order is the one it happened in, not the one it tidies into.

Seven manual steps, seven to ten days

I mapped the current claims workflow end to end: seven manual steps, multiple specialist handoffs, averaging 7 to 10 business days from first notice to resolution. The inefficiency was structural, not individual.

From this mapping, I identified three distinct zones where AI could operate at graduated levels of autonomy:

Zone 1: Autonomous. AI handles the claim end-to-end. The human sees a summary dashboard but does not intervene. Approximately 70% of claims.

Zone 2: Assisted. AI performs 80% of the analysis, but flags a specific decision point that requires human judgment. Approximately 20% of claims.

Zone 3: Escalated. AI gathers and organizes all relevant data, presents its findings, but the human makes the final determination. Approximately 10% of claims.

Three interfaces, because they're three different jobs

Three interfaces, because monitoring, deciding and answering to a regulator are three different mental jobs:

1. The Agent Dashboard. A control-room view that shows what the AI is doing in real time. Claims flowing through processing pipelines, auto-resolved cases accumulating, exceptions being flagged for review. The mental model is closer to air traffic control than to a traditional claims management spreadsheet.

2. The Exception Queue. When the AI is not sufficiently confident, it routes the claim to a human reviewer. But rather than presenting raw data and asking the adjuster to investigate from scratch, it presents its analysis, highlights the specific source of uncertainty, and suggests resolution options. This reduces average review time from 45 minutes to under 4 minutes.

3. The Audit Trail. Every autonomous decision is logged with full reasoning chains. Regulators can trace any payout back to the specific data inputs, confidence calculations, and decision logic that produced it. None of that can be added afterwards — the logging shape has to be decided before the first decision gets made.

The hard part was people going on autopilot

The most difficult design problem in this concept was not the AI interface itself. It was automation complacency.

When AI handles 85% of claims autonomously, the humans reviewing the remaining 15% begin to rubber-stamp. They trust the AI's pre-analysis too readily. Attention degrades precisely when it matters most, on the cases the AI flagged because it was uncertain.

I designed what I am calling 'attention anchors,' deliberate friction points that require genuine engagement with flagged cases. A confirmation checkbox stating 'I reviewed the source documents' paired with randomized verification prompts. The friction is minimal enough to preserve workflow efficiency, but sufficient to interrupt the autopilot pattern that high automation rates inevitably produce.

Two things I took away

- The regulator is a user. I put regulatory auditors in the persona map next to adjusters and policyholders, and their requirements — full auditability, a traceable decision chain, override authority — shaped the information architecture about as much as the adjusters' workflow did. Treat compliance as a constraint to be satisfied at the end and you get a weaker system with the same features. - This is recombination, not invention. Every pattern here came from something I had already watched fail somewhere else: trust calibration from the underwriting work, progressive disclosure from the risk platform, data density from an analyst-tools project at a bank. Speculative work is worth doing when you have enough production scar tissue to recombine honestly.

Open findings — filed against me

What I'd flag if this were someone else's project.

Open · What I'd change

What I'd do differently.

Every pattern here traces back to something I hit in production: trust calibration from the underwriting work, progressive disclosure from the risk platform, data density from a bank analyst-tools project. That's the honest description of what speculative work is — recombining things you've already seen fail. What I'd change is who validated it. I took this to operations leaders and not to adjusters, and the people who process claims all day would have found the holes in the three-zone model much faster than the people who manage them.

Open · On the numbers

How much the metrics are worth.

Nothing on this page was built or measured. This is speculative work: the numbers above are the assumptions the design rests on and the targets it aims at, which is why they sit under a different heading from the ones on my other studies. The one figure here that came from looking at how claims are handled today, rather than from the design, is the 45-minute review time. I removed a "100% auditability" metric that used to sit alongside these — every decision being traceable is an architectural property, not a result.