OffSpec

Design QA · Figma frame vs. simulator screenshot

Design QA for screens that don't have a URL.

Two screenshots in, one assignable ticket out.

Drop in the Figma frame and a simulator screenshot. Get back the spacing, colour, type and alignment mismatches — scored against a published rubric, ranked Blocker / Should-fix / Nit, and written up as a ticket you can paste straight into Jira. About thirty seconds, no plugin, no CI, no card.

No account. No signup. Three free comparisons — enough to check one real screen and decide for yourself.

80% of what they build is inaccurate, but close enough that they push back when I tell them it's not right.
a designer, r/UXDesign

Run a comparison

3 free comparisons
Drop an image, or click to chooseFigma frame export. PNG or JPEG.
Drop an image, or click to chooseSimulator or device screenshot.
About 30 seconds. No account, no card.
01

Built for screens, not pages

Every AI design-QA tool on the market needs a live URL to crawl — it reads CSS off a running web page. Your iOS screen has no CSS to read. This compares two images, so it works on a simulator capture, a device screenshot, a TestFlight build, or a still from a screen recording.

02

You get a ticket, not a diff

Overlay tools show you a blurry red halo and leave you to work out what it means and go type it up. This names the element, says what the design shows and what shipped, and hands you a formatted, severity-tagged ticket ready to paste into Jira or Linear.

03

Severity, so you stop arguing

Every finding is tagged Blocker, Should-fix, or Nit — and every acceptance criterion carries the evidence behind it, not just the instruction. The 1px icon shift gets labelled a nit and stops eating your standup. The wrong brand colour gets labelled a blocker.

This is not a raw pixel diff. Two screenshots at different scale factors will never subtract to zero, and a tool that reports that noise is useless. Findings are triaged: a 1px icon shift is labelled a nit, a wrong brand colour is labelled a blocker.

How it works

  1. 01

    Export the frame

    Any PNG from Figma. No plugin to install, no edit access to the file, no Figma account required.

  2. 02

    Grab the screenshot

    ⌘S in the iOS Simulator, or the volume+power screenshot from a real device. Android Studio works the same way.

  3. 03

    Get the ticket

    Scale is normalised automatically. You get a 0–100 match score, every mismatch listed with what the design shows versus what shipped, and a copy-ready ticket at the bottom.

What it checks

Ten weighted categories, 61 sub-checks, weights summing to 100. The weight is also the cap: no category can ever deduct more than its share, however many findings land in it.

Pay for what you check

Free
3 comparisons

No card, no trial timer, no account nag. Enough to check one real screen and decide for yourself.

$9
25 comparisons

One-time. They don't expire. About the size of a full app's screen inventory, which is exactly the point: audit the whole build before you invoice.

Unlock 25 comparisons
$19
per month · coming

Unlimited comparisons and a direct Jira write, for teams shipping every week. Not here today; today you copy and paste, which honestly takes four seconds.

For context on the price: the cheapest competing tool with AI analysis starts at $39/mo with no free tier, and its full visual-analysis tier is $249/mo. Those are good tools, built for web teams with a procurement process. This is built for the person who has one screen to check right now.

Questions

Can't I just paste both screenshots into ChatGPT?
You can — once. You'll get different findings every run, a vibes-based score (models rate themselves ~9/10 unprompted), and whatever format it feels like today. Here the model only reports findings; the score is computed server-side against a 61-check weighted rubric, so an 87 means the same thing on every screen, every run. That consistency is what lets a team use it as a handoff gate — and the output is already a severity-tagged ticket, not a chat reply you reformat by hand.
Does this work for web too?
Yes — feed it a browser screenshot and it works fine. Mobile is what it's tuned for and honestly where it's most useful, because the web already has half a dozen tools that plug into a live URL and read your CSS directly. Those are more precise than an image comparison when a URL exists. When one doesn't, they can't help you at all.
Do I need a Figma plugin?
No, deliberately. Most alternatives run inside Figma, which means you need edit access to the client's file — something contract devs frequently don't have. This takes an exported PNG, so anyone who can see the design can run the check.
How accurate is it?
Strongest on the things that change a screen's meaning: a missing or extra element, a wrong brand colour, a type weight that shifted, a hierarchy that inverted. Weakest on fine geometry across a full-length screen — a corner radius that dropped a few points, or spacing drift under about 5% of the screen's width, can slip past it, because at that scale the difference is smaller than what the vision model resolves in one pass. Crop to the region you care about and it sees far more. It also can't see what a still image doesn't contain: interaction states, scroll behaviour, dynamic type, anything below the fold. Treat it as a well-prepared reviewer with good instincts and imperfect eyes, not a compiler.
What if the screenshot is a different size than the frame?
That's the normal case, not an error case. Simulator captures come out at 2x or 3x; frames are usually 1x. The engine establishes a shared reference element and reasons in proportions, so a 3x iPhone capture against a 1x frame is fine.
Does it send my designs anywhere?
The two images go to Anthropic's Claude via OpenRouter to generate your report, and are not used to train anything. Nothing is written to a database — there isn't one. If you're under an NDA that forbids third-party processing of client assets, don't use this; use a local overlay tool instead. Straight answer, since you were going to ask.
Where does the score come from?
A published rubric: ten weighted categories summing to 100, and a fixed points-per-severity formula applied by the server. The model finds and classifies; it never picks the number. That's why the same pair of screenshots scores the same twice, and why the number survives an argument with an engineering manager.

Launched today. Here's the honest version.

There are no customer logos on this page because there are no customers yet — this shipped on 13 August 2026. Rather than borrow credibility, here's what's actually verifiable.

Three free comparisons, no card

Run it on a screen you already know is wrong and see whether it finds what you found.

What it doesn't do yet

No direct Jira write (copy-paste only), no Android Studio plugin, no batch upload, no team accounts. Those are next. They aren't here today, and pretending otherwise would waste your time.

Why it exists

Every tool that does AI design QA needs a live URL, because it reads CSS off a running page. Native app screens don't have one, so mobile devs got left with overlay tools and their own eyeballs. That gap is the entire product.

Found a bug or a bad call? I'm one person and I'll actually reply. jordan@jwcholding.com