A design critique you can argue with — and watch change its mind
Post a screen in Slack and you get emoji. Book a mentor and you're waiting until Thursday. Pay for an audit and it lands after you've already shipped. Paste it into a chatbot and you get flattery, which is worse than nothing. So I built the reviewer I actually wanted: it ranks findings by what they'll cost you, points at the pixels it's talking about, says so when it can't be sure — and if you push back, it has to defend the call or take it back.
- AI product
- Agent design
- Trust and calibration
- Prompt systems
- Designer who ships
- Product
- Deskcrit — AI design critique
- Role
- Strategy, interaction, prompt system, build — solo
- Timeline
- Brief to working product in one week
- Stack
- Next.js · Anthropic vision API
- Process
- Designed in Claude Design, built with Claude Code
- Status
- Live — public demo mode, July 2026
Five minutes end to end. If you only have one, read the exchange below — that's the product.
The primary action sits inside a scrolling region, so on a short viewport it is below the fold at rest.
- Evidence
- Points straight at the region and the button sitting inside it. It can't make a claim without showing you where.
- Capped
- It measured this off the picture, so it's an estimate and says so. And it's one screen — it has no idea whether the next step offers you the same action again.
The designer
This is a kiosk build. It runs on one fixed 1080p screen, mounted portrait, and nothing about it scrolls.
A fixed screen kills the fold problem, and the fold problem was the whole reason this was a Medium. It stays on the list at Low, though — that button is still the quietest thing in its own region, and that part had nothing to do with the viewport.
- week, brief to deployed
1
week, brief to deployed
nobody else on it — strategy, interaction, prompts, build
- iterations to get it right
4
iterations to get it right
the first three were wizards; the fourth threw the stepper out
- ways an argument can end
3
ways an argument can end
it holds, it folds, or it moves — and you see which
- steps before your first critique
0
steps before your first critique
no signup, no waitlist, no sample data to click through
All three made it worse to demo.
- 01
Threw out three finished versions with two days left
The first three were a wizard: fill in a form, watch it review, read a report. Nothing broken about it. But the whole pitch is a reviewer you can talk back to, and I'd buried the talking back in a textarea on the third screen.
I'd made arguing with it a feature you have to go find.
Read the reasoning - 02
Stopped it from sounding sure when it isn't
It can only see one screen, so it isn't allowed to score the flow. It measures tap targets off pixels, so it calls those estimates. And when the answer depends on something it can't see, it says so and hands the call back to you.
A cockier version would demo better. It would also be lying to you.
Read the reasoning - 03
Let it come back with almost nothing
It never pads a list to look thorough, so some runs turn up one small thing. Same reason it can't open with a compliment, can't tell you to “consider improving the hierarchy,” and can't say “as an AI.”
Bad for the demo. It's the only reason five findings means anything.
Read the reasoning
Three versions in, I realized I'd built the wrong shape.
Nothing changed about the judgment — same model, same findings, same way of ranking them. What changed is where you argue with it.
Wizard · stepper across the top
Three screens, one direction, and the only thing that makes this different sitting at the bottom of the last one.
One conversation
One surface, no stepper, and arguing is the obvious thing to do rather than a button you have to hunt for.
The fork
Three versions in, I had a wizard that worked: intake form, review screen, report, stepper across the top. Spend the last two days polishing it, or pull the whole architecture apart with most of the week gone.
The tension
Nothing was broken. That's what made it a hard call. But I'd taken “a reviewer you can talk back to” and built it as forms and panels, so talking back meant hunting for a textarea inside a card. A wizard suits a job that finishes. Arguing about a design doesn't finish.
The call
I pulled it apart. Version four is one conversation: you drop a screen in, findings come back as cards in the thread, and every one of them has Agree and Challenge sitting right there. If the thing that makes your product different is an interaction, that interaction has to be the interface.
The consequence
Arguing stopped being a feature and turned into just talking. It also cost me three finished versions in a one-week build, and left me with a product that's genuinely hard to screenshot — which matters more than it sounds when the way people find it is a portfolio page.
It admits what it can't see. That's the only reason to believe the rest.
I've spent twenty years designing AI platforms for banks and insurers, where a bad design decision doesn't cost you churn, it costs you a compliance finding. That work left me with one conviction: the hard part of human-AI collaboration was never making the machine smarter. It's whether a person can see how it got there and push back on it.
So the hedging here is engineered, not decoration. The model could always see the screen — that was never the hard part. What you get out of a general chatbot is commentary: nothing ranked, nothing you can go check, and it agrees with you. Getting from that to critique is all design work. None of it arrives with the model, which is why the system is the product and not the model.
The fork
Let it answer everything in the same confident voice, or draw a line around what it genuinely can't know from one screen and make it stay behind that line.
The tension
Every cap is the product admitting something, and admissions don't impress anybody. You get about ninety seconds from a first-time visitor. Spending part of that on “I can't tell from here” looks weak next to something that just answers.
The call
Three caps. Each one written as the sentence you'll actually read, not as a score — it can't grade the flow when it's only seen one screen, it calls pixel measurements estimates, and when the answer depends on context it doesn't have, it says so and gives you the call.
The consequence
Now when it doesn't hedge, you can believe it. That's the whole return. The cost is that the product is at its least impressive in the first ninety seconds, and the caps are the first thing a skeptic runs into.
Three things it isn't allowed to sound sure about.
- In the build
- 3
- Shown as a %
- 0
Anything about the flow
Where this step sits, what it assumes you already did, whether you can recover from an error here — all of that lives on screens it hasn't seen. It can raise the question. It can't put a number on it.
I can only see this step
Anything it measured off the picture
Tap targets and contrast ratios get read off the image, not out of your code or your tokens. Close enough to be worth telling you about, not close enough to state as fact — so it says estimate, and the severity gets held down to match.
Estimate
Calls that depend on what it can't see
Is this pattern wrong, or is it a kiosk, a disclosure your legal team wrote, a deadline somebody already fought about? You can't tell from a screen. So it names the tension and hands you the decision — which is half the reason the challenge flow exists at all.
Yours to call
The phrase on the right is what it actually says to you. I wrote the caps as sentences instead of percentages on purpose — “I can only see this step” tells you what to do next. A confidence score of 62% tells you nothing you can act on.
If it can't tell you there's nothing wrong, it can't tell you anything.
The fork
Make sure every run comes back with a decent list of findings, or let it come back nearly empty when there's nearly nothing to say.
The tension
Padding is invisible, which is exactly what makes it the disease. Three real findings and two made-up ones look the same to you until you act on one of the made-up ones. After that, nothing it ever tells you is worth much.
The call
“Nothing significant here” is a legitimate answer, and it never stretches a list to look thorough. Same reason I wrote the voice rules as bans, and wrote the ways it could embarrass itself before I wrote the happy path.
The consequence
Worse demo, better product. Somebody can run it and watch it barely do anything — and that's the only reason a list of five findings is worth reading when it does show up.
Three things it isn't allowed to say to you.
- Opening with a compliment
- “Great start!” before it gets to the findings. It's buying goodwill with the one thing a critique actually owns — that when it says something is good, you believe it.
- Advice you can't act on
- “Consider improving the hierarchy.” You can't do anything with that, you can't prove it wrong, and it reads exactly the same as if it never looked at your screen.
- “As an AI”
- A disclaimer about itself where a claim about your screen should be. If it's going to hedge, it should hedge on the evidence, not on what it is.
You find out what a voice is really made of in the weird cases, so I wrote those first. Upload a photo of your dog and it says “This looks like a very good dog, but I critique interfaces.” Not an error message, and not pretending it can review a dog.
Five trades, and what I paid for each one.
Every one of these made the thing harder to build or worse to demo. Take the cost column off a list like this and you're just showing off.
The trade
What I chose
What it costs
- Ship the wizard that worked, or throw out three versions and start over
- What I choseStarted over. Version four folded intake, review and report into one conversation, and arguing with a finding turned into just talking.
- What it costsThree finished versions gone, in a build that only had a week in it. And a conversation is much harder to screenshot than a report — which matters when people find the thing through a portfolio page.
- Let it answer everything, or cap it where it can't actually know
- What I choseCapped it, and wrote each cap as the sentence you read instead of a score.
- What it costsA cockier version demos better. Hedging looks like weakness the first time you meet it, and I'm betting on credibility, which pays a lot slower than confidence does.
- Guarantee a full list of findings, or let it say “nothing significant here”
- What I choseLet it. It never stretches a list to look thorough, so some runs come back nearly empty.
- What it costsAbout the worst demo you could design. Somebody shows up, runs it, and watches it do almost nothing. That's what I'm paying for five findings meaning five findings.
- Get the signup first, or lead with the critique
- What I choseLead with the critique. No signup, no waitlist, nothing to click through first.
- What it costsNo email list and no idea who's using it — part of why the number I care about is still a number I don't have.
- Describe the agent's judgment, or specify it tightly enough to ship
- What I choseSpecified it. Severity model, what happens when you challenge it, the confidence caps, the banned phrases, and the embarrassing cases written before the happy path.
- What it costsMost of the week went into a document nobody will ever see, for a product whose entire visible surface is a text field and a card.
The part nobody sees, which took most of the week.
What you see is a text field and a card. What's underneath it is a set of rules about judgment, and that's the entire difference between this and a chatbot with opinions.
Four principles, and what each one forbidsThe rules every decision was checked against
- Show the reasoning or don't ship the finding
- Nothing goes out without why it's a problem and where it is on your screen. An opinion with no reasoning behind it is just noise delivered confidently.
- You can argue with all of it
- Challenge anything. It either comes back with specifics or it backs down and changes the record — no third option where it repeats itself at you.
- It says when it can't tell
- This is the one people push back on, and it's the one I'd defend hardest. You can't trust the confident answers from something that's never uncertain.
- Nothing between you and the first critique
- No signup, no waitlist, no sample data to click through. You'll give a new product about ninety seconds, and I'd rather spend all of it on the actual thing.
The challenge protocol, and its three endingsDefended, conceded, revised — and what each one commits to
When you challenge something, the model gets your argument with one instruction attached: answer what they actually said, don't just say the finding again. Restating is the thing that makes an agent feel like a wall while it's technically replying to you, and it's the specific behavior I was trying to design out.
- Defended
- It's sticking with the call, and it has to tell you something new about your screen to do that. Repeating itself doesn't count as defending anything.
- Conceded
- You knew something it didn't, and the finding comes off the list. It comes off on the record though — the exchange stays where you can see it.
- Revised
- You were half right and the severity moves. It says what moved and why, so you can tell the difference between it reconsidering and it caving.
Nowhere does this product tell you to trust it. You either come to trust it or you don't, one argument at a time, and those three are the only ways an argument is allowed to finish.
Commentary and critique, side by sideThe four differences the system exists to produce
Commentary
What you get from a chatbot
- Nothing ranked, and it wants to agree with you — leads with a compliment
- Claims you can't go and check for yourself
- Sounds equally sure about all of it
- You read it and that's the end. There's nobody to argue with
Critique
What I specified this to do
- Ranked by what it'll cost you, and stingy with praise — no compliment opener
- Points at the pixels. Every claim lands somewhere specific
- Tells you where it's guessing, and holds itself down when it can't know
- You can argue with any of it, and the argument stays on the record
Where this could still fall over.
You'll notice there are no results on this page. It went live in July 2026 and I have no usage data yet. Putting a number here anyway would be the exact thing the product exists to argue against, so instead here's the number I'm waiting on: how often somebody pushes back at least once. If people argue with it, the trust design worked. If they just read it and close the tab, I've built a nicer report generator and I was wrong about the whole premise.
The hedging might just read as weak. Something that tells you “I can't see enough to call this one” is more useful than something that answers everything, and it is definitely less impressive. I'm betting credibility beats confidence over time. That's a bet, not a result, and I'll know from the pushback numbers.
I also don't know if people will argue with software. Everything they'd need is right there — Agree, Challenge, and a box that takes anything you want to type. But nobody's in the habit of talking back to a review tool, and habits usually have to be taught. If I have to teach it, the interface isn't done.
And no designer but me has taken a run at the voice. I wrote the bans, I wrote the awkward cases, and I'm also the person least likely to be annoyed by how it talks. The dog test tells me the tone holds. It tells me nothing about whether a room full of design directors would rank these findings in the same order it does.
This is a portfolio piece and I'm not going to pretend otherwise — it's the first of a few things I want to build on the same idea: take expert judgment, hand it to an agent, and leave the reasoning where you can check it. Critique first, research synthesis next, design system audits after that. It's also me making a point about what a designer can do now. I'd rather specify an agent's judgment tightly enough to ship it than mock up another screen of an AI product that doesn't exist.
Everything above is a claim. The product is the evidence.
Drop a screen in and push back on the first thing it tells you. It defends the call, softens it, or takes it back — that exchange is the entire argument of this page, and it takes about a minute to check.