Drop in any screen. Get a senior designer's critique.
Deskcrit is a tool: upload any screen — yours, your team's, a checkout page you're arguing about — and it reviews it the way a senior designer would. Ranked findings, evidence from the actual pixels, a confidence score, and if you disagree, it defends the call or takes it back. This study is written in the same format the tool produces: decisions as findings, screenshots as captures.
- AI product
- Prompt system design
- Trust & calibration
- Solo build
- Role
- Strategy, interaction, prompt system, build — solo
- Timeline
- 7 days, brief to live
- Stack
- Next.js · Anthropic vision API · Claude Code
- Status
- Live — public demo, July 2026
Four minutes end to end. In a hurry? The three findings are the argument — everything else is captures.
One designer took an AI critique tool from brief to live public product in a week — and the interesting part is not the speed. It's the three findings below.
- Days, brief to live
7
Days, brief to live
Strategy, interaction, prompt system and build, alone
- Prototypes killed
3
Prototypes killed
All three chat-shaped, all three demoed well, all three dead on day 3
- Parts to every finding
4
Parts to every finding
Severity, evidence from the pixels, confidence, the smallest fix
- Accounts to check the claims
0
Accounts to check the claims
Public demo, one critique on the house, no signup
Senior critique is scarcest at the exact moment work ships.
The use case
A designer with a screen and no senior in the room.
Solo designers, small teams, anyone shipping without a design director on call. The moment is always the same: the work is about to go out, someone disagrees about it, and the feedback available is Slack emoji, a mentor slot on Thursday, or a chatbot that compliments everything.
The design problem
AI feedback exists. Trustworthy AI feedback doesn't.
Any chatbot will comment on a screen. None of it is usable as a review: nothing ranked, nothing tied to a pixel, everything stated with the same confidence — so you can't tell the finding that costs conversions from the nitpick. The problem was trust mechanics, not fluency.
The business problem
Review is a bottleneck priced like a luxury.
Professional audits cost thousands and land after the ship date. Senior reviewers are expensive, overbooked, and don't scale past their calendar. The bet: package senior judgment as software, and the free critique becomes the funnel.
Three findings
The first three versions worked, demoed well, and were the wrong product.
Why I flagged this
Evidence
What I did
What it cost
A reviewer earns belief by admitting what it can't see.
Why I flagged this
What I did
What it cost
Capture A — the signature move

A reviewer that always finds five problems is lying some of the time.
Why I flagged this
What I did
Capture B

Capture C

Seven days, unedited.
The process, kept honest: loops, not phases. Days 5 and 6 — the invisible prompt-system work — took more of the week than anything you can see.
- Day 1
- Mapped where designers actually get critique — Slack emoji, booked mentors, flattering chatbots. Wrote the brief: ranked, evidenced, challengeable.
- Days 2–3
- Built three chat-shaped prototypes. Tested them on screens whose problems I already knew. All three failed scrutiny. Killed.
- Days 3–4
- Rebuilt around the finding object: severity, evidence, confidence, fix. Designed and built the same day.
- Days 5–6
- The prompt system: banned-certainty rules, the challenge protocol, the zero-finding state. Most of the week lived here.
- Day 7
- Shipped public: demo mode, no accounts, one free critique — so anyone reading this can check the claims themselves.
What I'd flag if this were someone else's project.
Calibration measures honesty, not accuracy.
I can measure the revision rate. I can't yet measure what nobody challenged. A model can be calibrated and still wrong in the quiet places.
It reviews stills. Interaction is invisible.
Flow logic, latency, motion — the slow killers — are outside the frame. The banned-certainty rules stop it pretending otherwise. A boundary honestly stated is still a boundary.
It's one person's judgment, encoded.
Deskcrit argues the way I argue. That's the product and the limit — it inherits my blind spots at scale, which is exactly what I'd flag in someone else's agent.
Every claim on this page has a live rebuttal channel.
Drop in a screen you actually care about. Challenge the finding you disagree with. How it argues is my work sample.