Skip to main content
AI product · one week · strategy to shipped code, alone

Drop in any screen. Get a senior designer's critique.

Deskcrit is a tool: upload any screen — yours, your team's, a checkout page you're arguing about — and it reviews it the way a senior designer would. Ranked findings, evidence from the actual pixels, a confidence score, and if you disagree, it defends the call or takes it back. This study is written in the same format the tool produces: decisions as findings, screenshots as captures.

Get your screen critiqued (opens in a new tab)Live, no signup. One critique on the house.
  • AI product
  • Prompt system design
  • Trust & calibration
  • Solo build
At a glance
Role
Strategy, interaction, prompt system, build — solo
Timeline
7 days, brief to live
Stack
Next.js · Anthropic vision API · Claude Code
Status
Live — public demo, July 2026

Four minutes end to end. In a hurry? The three findings are the argument — everything else is captures.

Verdict first — for people with forty portfolios open

One designer took an AI critique tool from brief to live public product in a week — and the interesting part is not the speed. It's the three findings below.

Days, brief to live

7

Days, brief to live

Strategy, interaction, prompt system and build, alone

Prototypes killed

3

Prototypes killed

All three chat-shaped, all three demoed well, all three dead on day 3

Parts to every finding

4

Parts to every finding

Severity, evidence from the pixels, confidence, the smallest fix

Accounts to check the claims

0

Accounts to check the claims

Public demo, one critique on the house, no signup

The brief — what this solves, and for whom

Senior critique is scarcest at the exact moment work ships.

The use case

A designer with a screen and no senior in the room.

Solo designers, small teams, anyone shipping without a design director on call. The moment is always the same: the work is about to go out, someone disagrees about it, and the feedback available is Slack emoji, a mentor slot on Thursday, or a chatbot that compliments everything.

The design problem

AI feedback exists. Trustworthy AI feedback doesn't.

Any chatbot will comment on a screen. None of it is usable as a review: nothing ranked, nothing tied to a pixel, everything stated with the same confidence — so you can't tell the finding that costs conversions from the nitpick. The problem was trust mechanics, not fluency.

The business problem

Review is a bottleneck priced like a luxury.

Professional audits cost thousands and land after the ship date. Senior reviewers are expensive, overbooked, and don't scale past their calendar. The bet: package senior judgment as software, and the free critique becomes the funnel.

Three findings

Finding 01Severity: high
Confidence 96% — I lived it

The first three versions worked, demoed well, and were the wrong product.

Why I flagged this

All three were chat-shaped: paste a screen, receive a fluent paragraph. Fluent paragraphs read as opinion. A $10k-a-month conversion leak and a nitpick about corner radius sat in the same prose with the same weight.

Evidence

I ran each version on screens I already knew the problems with. Every review sounded right. Under a minute of scrutiny, none survived the question a design lead actually asks: which of these do I fix first, and how sure are you.

What I did

Killed all three on day 3 and rebuilt around a different unit: the finding. Severity, evidence measured from the pixels, a confidence score, the smallest fix that resolves it. Chat survives only as the channel you argue through.

What it cost

Two days of a seven-day timeline, thrown away in public. Worth it: the three dead versions are the clearest proof on this page that the final shape was chosen, not defaulted into.
Finding 02Severity: high
Confidence 91% — the revision log agrees

A reviewer earns belief by admitting what it can't see.

Why I flagged this

A vision model states everything with the same confidence, including things a still image cannot prove — hover states, animation, what happens after the tap. Uniform certainty demos better. It's also the fastest way to train a user to ignore you.

What I did

Every finding carries a calibrated confidence score, and the prompt system bans unearned certainty outright. Push back with real evidence and it re-examines, then defends, concedes, or revises the severity — on the record.

What it cost

A 62% next to a finding invites dismissal in a demo. I took that trade. A reviewer that can say “downgrading, pending evidence” is the only kind whose High severity means anything.

Capture A — the signature move

A finding, a challenge, and the reviewer re-examining: bottom tap targets flagged at Medium and 74% confidence, challenged on the real hit area, re-examined as fair, and revised from Medium to Low
Challenged on a hit-target call, it re-measures and revises Medium → Low. Concession as a feature. This is the exchange as the live product page demonstrates it, rather than a capture from inside a session.
Finding 03Severity: medium
Confidence 88% — hardest to prove, easiest to feel

A reviewer that always finds five problems is lying some of the time.

Why I flagged this

Demo pressure says pad the list — an empty review looks broken at the exact moment someone decides whether to trust you. But every padded finding teaches the user to skim, and a reviewer you skim is a reviewer you've stopped using. The failure is silent.

What I did

Zero findings is a designed, first-class state. It reports what was checked, at what confidence, and where coverage ends — a clean bill with the inspection log attached, not praise. The empty state is the trust budget every full review spends from.
Captures B — C · the shipped product

Capture B

Deskcrit's empty state: a dark canvas with a dashed drop zone reading “Put your work on the desk”, and the reviewer's opening message in the panel on the right
The desk. One surface, one job: put your work down and it starts reading. No account, no onboarding tour.

Capture C

A live Deskcrit session: a sample checkout screen on the canvas, focus chips for heuristics, accessibility and visual craft selected, and the password prompt that follows the free critique
A live session. You aim it — heuristics, accessibility, visual craft — and it reads the actual pixels, not your description of them. This one has spent its free critique, so the demo tier's password gate is showing.
Session log

Seven days, unedited.

The process, kept honest: loops, not phases. Days 5 and 6 — the invisible prompt-system work — took more of the week than anything you can see.

Day 1
Mapped where designers actually get critique — Slack emoji, booked mentors, flattering chatbots. Wrote the brief: ranked, evidenced, challengeable.
Days 2–3
Built three chat-shaped prototypes. Tested them on screens whose problems I already knew. All three failed scrutiny. Killed.
Days 3–4
Rebuilt around the finding object: severity, evidence, confidence, fix. Designed and built the same day.
Days 5–6
The prompt system: banned-certainty rules, the challenge protocol, the zero-finding state. Most of the week lived here.
Day 7
Shipped public: demo mode, no accounts, one free critique — so anyone reading this can check the claims themselves.
Open findings — filed against me

What I'd flag if this were someone else's project.

Open · Medium

Calibration measures honesty, not accuracy.

I can measure the revision rate. I can't yet measure what nobody challenged. A model can be calibrated and still wrong in the quiet places.

Open · Medium

It reviews stills. Interaction is invisible.

Flow logic, latency, motion — the slow killers — are outside the frame. The banned-certainty rules stop it pretending otherwise. A boundary honestly stated is still a boundary.

Open · Won't fix

It's one person's judgment, encoded.

Deskcrit argues the way I argue. That's the product and the limit — it inherits my blind spots at scale, which is exactly what I'd flag in someone else's agent.

Every claim on this page has a live rebuttal channel.

Drop in a screen you actually care about. Challenge the finding you disagree with. How it argues is my work sample.