Which screenshot makes you want to try the product?
The one people see in search results. Everything rides on it.
Human taste > another AI opinion
Drop two versions. Get fast human decisions on your landing page, logo, screenshot, pricing, or anything else you’re about to ship.
10 free credits to startNo cardVote without an account
Which one would you ship?
“Which homepage would make you try the product?”
Pick one. This is the whole interaction.
Voters can also useAB
Landing page, logo, screenshot, copy, pricing — or anything else.
Drag two images in, or type two lines of copy. Nothing else to configure.
Take a suggested one or write your own. This is the only part that needs thought.
10, 25, 50, 100 or a number you pick. One credit per response.
The loop runs both ways. Give boops to earn credits, spend credits to get boops. That’s the entire economy — no seats, no contracts, no waiting for a research panel to schedule you in.
Two versions of the thing you can't decide on. Images or copy.
One question, in plain language. “Which one would you trust enough to buy?”
Real people compare and pick. One tap, one reason, thirty seconds.
You get counts, a confidence read and the reasons behind the split.
DemoTests marked Demo are fictional examples seeded with this deployment, not real human responses.
The one people see in search results. Everything rides on it.
Goes out with the announcement post. 1200x675.
Three plans either way. I keep going back and forth and I need to stop.
Showing the split hero for five seconds. I want to know if the product lands without a scroll.
Same copy, same product. A is centered and quiet, B is split with the dashboard visible.
Start free — no card
See it work on your data
You can produce forty homepages before lunch now. That was the hard part, and it stopped being hard. The bottleneck moved: you have options and no idea which one a stranger would trust.
Ask a model which one is better and you get a confident paragraph assembled from everything that has ever been written about design. It is articulate, it is instant, and it is not a preference. Nobody chose anything.
Boop measures the choice. Someone looked at both, picked one, and told you why. That’s a different kind of number — and it is the only one that predicts what happens when you ship.
“Variant B demonstrates stronger visual hierarchy and a more conventional information architecture, which may improve comprehension for first-time visitors.”
Plausible. Unfalsifiable. Zero humans involved.
67% picked B
82 people. Top reason: “Easier to understand.” Eleven of them wrote a sentence explaining it.
A count, an interval, and the reasons underneath.
Boop does use a model for one job: reading the written comments and grouping them into themes, so you don’t have to skim ninety sentences. It never casts a vote and never invents feedback. Every number on a results page comes from a person who clicked.
Nobody studies your homepage. They glance, they form an impression, and they decide whether to keep reading. A comparison tells you which version wins. A 5 Second Test tells you whether anything landed at all.
Show one design for exactly five seconds.
No scrubbing, no second look.
Hide it, then ask one question.
What do you remember? What does it do?
Read the answers side by side.
The gap between intent and recall is the finding.
You get five seconds with the design. Then it goes away and we ask what stuck.
Hero, above the fold, the whole page.
“Which one would make you try it?”
Marks, wordmarks, the small-size test.
“Which would you remember tomorrow?”
Screens, components, navigation shape.
“Which one feels easier to use?”
Tables, tiers, packaging.
“Which is easier to understand?”
App store, listings, docs.
“Which makes you want to try it?”
Static creative, social graphics.
“Which would you stop scrolling for?”
Two finalists, said out loud.
“Which is easier to say?”
Taglines, CTAs, empty states.
“Which line would make you click?”
“Which homepage would make you try the product?”
B won on comprehension, not aesthetics — people said A looked better and still picked B because it showed the product. The recurring complaint about A is that nothing on screen says what the company does.
Strongest signal
Seeing a real payout row made the product feel legitimate — six people mentioned it unprompted.
Biggest concern
Three voters found B's right-hand panel busy. The win is the screenshot, not the density.
Next action
Ship B, then test B against a version with one card removed from the panel.
The model reads the comments. It never casts a vote, never writes one, and never changes a count.
Illustrative results page built from the fictional sample test above.
Everything on this page above this line exists today. Everything below is planned work — no dates, no promises.
Ask only founders, only designers, or only people who have shipped something this year. Narrow the room before you ask the question.
Invite your own list — customers, beta testers, a Discord — and run the same test against people who actually use your product.
Shared credit pools, shared history, and a record of every call your team made before shipping.
Open a test from CI, block a deploy on a taste test, or pipe results back into your own dashboard.
Boop is a small, independent product. It exists because the loop it describes is one we wanted and couldn’t buy: not a research platform with a sales call, not a Slack poll of four colleagues who already know what you’re hoping to hear.
The design of the product is the argument. It is quiet, it is mostly white, and it gets out of the way of the thing you uploaded — because the thing you uploaded is what people are supposed to be looking at.
Read why human taste still mattersNo seats, no subscription required. Earn credits by helping other builders, or buy a pack when you’re in a hurry.
$0
Free forever
Earn credits by helping others. One small test included.
$9
50 Boops · one-off
One decision, settled today.
$29
250 Boops · one-off
A whole launch worth of calls.
$79
1,000 Boops · one-off
For teams shipping every week.
Checkout isn’t enabled on this deployment. The pricing above is real configuration, not a mockup — add your Stripe keys and these buttons open a live Checkout session. Until then, every credit is earned: 1 credit per 3 boops you give.
One human decision. A person looks at your two versions, picks one, and optionally tags why. That single act is the unit Boop is built around — it's the vote, the credit and the name.
Other people using Boop: founders, designers, developers, marketers and anyone who wandered into the feed. It's a general audience of internet builders, not a screened research panel. That's the right sample for “does this read as trustworthy” and the wrong one for “would a hospital procurement officer buy this.” Audience filters and private panels are on the roadmap, and are marked as not built yet.
Yes. Unlisted tests stay out of the public feed and are reachable only by link. Private tests are visible to you alone and are how you'd run a test against an audience you bring yourself. Public tests are the default because they're what earns you responses from the feed.
New accounts get 10 credits. You earn 1 more for every 3 boops you give, plus a bonus for leaving written feedback. Requesting one human response costs 1 credit. Buying a pack is a shortcut, not a requirement — the loop works without spending anything.
Yes, and it's one of the most useful things to test. Type two names, taglines or CTA lines and Boop typesets them properly rather than showing them as raw text, because how a line reads at size is part of what you're judging.
You show one design for exactly five seconds, then hide it and ask a single recall question — what do you remember, what does this do, what would you click. It measures whether anything landed at all, which a head-to-head comparison can't tell you.
No. Every vote and every comment comes from a person. AI is used for exactly one job: reading the written feedback and grouping it into themes so you don't have to skim ninety comments. It is labelled as such everywhere it appears, it never casts a vote, and it never changes a count. Sample tests seeded with the app are clearly marked Demo.
With a 95% Wilson score interval on the leading variant's share. A winner is only declared once the whole interval sits above 50% — meaning the observed preference is consistent enough that a coin flip is excluded at that sample size. Below that you'll see “Early signal” or “Moderate signal” instead of a result. It describes consistency at this sample size, not a guarantee about everyone on the internet.
Know what people pick before you ship. Two versions, a question, and a room full of people who have no reason to be polite about it.
Running in Demo Mode — everything works, nothing is stored permanently.