0
Table of Contents
How to test an MVP before you build it
Here's how to check your assumptions behind a product before you commit to a production build.
So, you have an MVP idea and a budget to build it. Before that budget turns into code, figure out first which of your assumptions would make the product fail if they turned out to be wrong, and what is the cheapest way to find out?
MVP testing answers that, and it's also the scope of this guide: checking the assumptions behind a product before you commit to a production build. Analytics on a shipped product and A/B tests in a live app are different jobs, and this guide doesn't cover them.
Below: the four classes of evidence a pre-build test can produce, a six-method matrix showing what evidence each test gives you and what it can’t, a rule for choosing the lightest test that still counts, and the ways founders misread the result. Three Merge client projects near the end show what feasibility-first and design-first tests look like in practice.
What "testing an MVP" means before you write code
The learning instrument, not a smaller product
Eric Ries, author of The Lean Startup, put the purpose this way in a 2011 guest article for TechCrunch:
The goal of the MVP is to begin the process of learning, not end it... Its goal is to test fundamental business hypotheses.
The Dropbox story from the same article shows what he meant. Drew Houston could not demo working software, so he recorded a three-minute video of the product as it was meant to work. He later recalled that the beta waiting list "went from 5,000 people to 75,000 people literally overnight." In Ries's vocabulary, the video was the MVP. The engineers were building the product at the same time; the video answered the demand question before the code could.
Steve Blank, author of The Four Steps to the Epiphany and creator of Customer Development, wrote in a 2013 post on his own site that:
A minimum viable product is not always a smaller/cheaper version of your final product... Think about cheap hacks to test the goal.
Marty Cagan, founder of Silicon Valley Product Group, in his SVPG article on the minimum viable product, defines an MVP as the smallest possible product that people choose to use or buy, can figure out how to use, and that the team can deliver with the resources available. The "smallest possible experiment to test a specific hypothesis" he calls an "MVP Test," so that nobody confuses an experiment with a product.
While Ries calls the Dropbox video an MVP, in Cagan's terms it is an MVP Test. In this guide, we follow Cagan's split because it keeps the experiment and the product apart: the MVP is the smallest product you would ship and stand behind, and MVP testing is the set of experiments you run before or alongside building it. The Dropbox video is one of those experiments.
What this article does not cover
Three neighboring topics get filed under "MVP testing" and are outside the scope of this guide:
- Post-build analytics. Session recordings, funnels, and retention on a product you have already instrumented. Useful, but it happens after the code exists.
- In-product A/B testing. Comparing two variants of a feature inside a live product. That is optimization of something that already works.
- The definition of an MVP and how to scope one. Covered elsewhere; this article assumes you have an idea and are deciding what to check first.
Those are worth their own guides. We want to answer a narrower question: what do I have to be right about, and how do I check it before I pay for a build?
The four evidence classes a pre-build test can produce
Every MVP validation method produces one or two kinds of evidence. Sorting methods by the evidence they produce, rather than by how popular they are, is what makes the choice manageable. There are four classes.
Demand: will anyone want this?
Evidence that a specific group of people reacts to the offer when they meet it in a real channel: they click, sign up, ask for access, or come back. Demand evidence is usually the easiest to collect and the easiest to overrate. A sign-up is a small act with no cost to the person signing up.
Usability: can they use it?
Evidence that the people in your target group can complete the core task with the interface you have in mind, without you explaining it. A product can have demand and still fail here, because the first session is often where a trial user decides whether to come back.
Feasibility: can we build it?
Evidence that the hard part of the product works at all with the data, integrations, models, or performance you need. For AI products, fintech, and anything that depends on a third-party data source or regulation, feasibility is often the riskiest assumption.
Willingness to pay: will they pay for it?
Evidence that someone commits something with a cost attached: a payment, a deposit, a signed pilot agreement, a purchase order. The commitment is what counts, not the channel it came through. A free waitlist form collects interest; the same landing page with a pre-order or deposit step collects willingness-to-pay evidence. A non-binding letter of intent is a different case. It shows that someone with authority took the offer seriously, but nothing was spent, so treat it as a weaker signal than money received.
We organize everything below around these four classes, because if you can't name which one a test produces, you won't know what the result means.
The six-method matrix: what each test shows and what it doesn't
The six MVP testing methods below are the ones we recommend when you test an MVP before building it. Each row states which evidence class the test produces, what it cannot show, and what it costs to run. You can read this matrix as a set of MVP tests to choose from, and the third column is there so you don't mistake a good result in one class for proof in another.
Method | Evidence class it produces | What it does not show | Relative effort |
Customer interviews | Sharpens the assumption; weak demand signal | Demand, willingness to pay, feasibility | Lowest |
Clickable prototype | Usability | Demand, willingness to pay, engineering feasibility | Medium (needs design) |
Landing page (with fake-door variant) | Demand; willingness to pay if the action is a pre-order or deposit | Usability, feasibility; willingness to pay when the action is a free sign-up | Low |
Concierge MVP / Wizard of Oz | Demand; willingness to pay if you charge | Feasibility of the automated version; scale | Medium to high (manual labor) |
Pre-sale / paid pilot | Willingness to pay; demand | Usability at scale, feasibility | Low to medium (sales time) |
Proof of concept (PoC) | Feasibility | Demand, willingness to pay | High (needs engineering) |
Effort is relative and reflects what each method needs from the team: a page builder, a designer, an engineer, or hours of manual service. Duration depends on the product. For example, Merge's PoC for Waffly, described below, took one month.
Customer interviews: clarify the assumption, not proof of demand
Interviews come first because they tell you what to test. A round of conversations with people who have the problem will turn a vague assumption into a specific, testable one. An interview round can also count as a test on its own, as long as you decide beforehand what you're listening for and how many people need to describe the problem unprompted before you call it confirmed. Without that written down, it's research, not evidence.
What interviews do not show: that anyone will act. People are polite, and they tend to describe the product they imagine rather than the one you will build. Treat what someone says they would do as a hypothesis, and go collect evidence of what they actually do.
Clickable prototype: usability signal, not a feasibility test
A clickable prototype is a set of designed screens wired together so a person can move through the core task. Put it in front of a small group of people from the target audience and watch where they stop. You learn whether the flow makes sense, whether the vocabulary matches how they think, and which step they skip.
What it does not show: demand, willingness to pay, or feasibility. A prototype session is a lab setting. The participant has agreed to spend time with you, so their engagement tells you nothing about whether a stranger would choose your product unprompted. And the screens are static designs wired together: they can show a result that no back-end, integration, or model has produced yet. If the risky question is whether the engineering works, that needs its own instrument, an engineering prototype, a research spike, or the proof of concept described below, not a more polished set of screens.
Landing page and the fake-door variant: demand
A landing page states the offer, names the audience, and asks for one action, usually an email address or a request for early access. Send a defined amount of traffic from the channel you would use in real life and measure the conversion rate. Which evidence class you get depends on the action you ask for: a free sign-up collects demand evidence, while a pre-order, deposit, or paid founder plan on the same page collects willingness-to-pay evidence, usually at a lower conversion rate. An MVP landing page is a fast demand test for a SaaS or fintech idea; Dropbox ran the same test with a video instead of a page. If you need the page designed and built rather than assembled in a page builder, Merge's product landing page design and development service exists for exactly this stage.
The fake door test is the variant for a single feature rather than a product. You add the button, menu item, or plan tier as if it existed, and when someone clicks it, you tell them it is not available yet and, ideally, ask why they wanted it. The click rate is the demand signal. Be direct with the person on the other side of the door: an honest "not built yet, tell us what you expected" keeps trust and gives you an interview lead.
What a landing page with a free form does not show: willingness to pay, usability, or feasibility. An email address costs nothing. A high sign-up rate on a fake door doesn't tell you the feature is buildable or that anyone would pay for it separately.
Concierge MVP and Wizard of Oz: demand and willingness to pay, manual back-end
Both methods deliver the outcome the product promises without the software that would automate it. In a concierge MVP, the customer knows a human is doing the work: you personally onboard, configure, and deliver, and you charge for it if the model allows. In a Wizard of Oz setup, the customer sees what looks like a working product while people behind the interface do the processing by hand.
Take the Zappos origin story, for example. According to a 2012 Fortune profile, founder Nick Swinmurn photographed shoes in local stores in 1999 and posted them online before holding any inventory, buying a pair from the store only when someone ordered it. Every order was a purchase decision, so the test produced willingness-to-pay evidence, not just interest.
What these methods do not show: that the automated version is feasible, or that unit economics hold once the humans are removed. A concierge service a handful of customers love may require an engineering effort its revenue cannot fund. That is about feasibility, and the concierge test does not answer it.
Pre-sale and paid pilot: testing willingness to pay
A pre-sale asks for a commitment before the product exists, and the commitments are not equal. Money received — a deposit or an annual plan at a founder price — is the strongest signal, because someone spent money. A paid pilot with a fixed scope and a start date is nearly as strong, because a budget owner had to approve the spend. A letter of intent is weaker: if it is non-binding, it records interest from a decision-maker, not a purchase, and it counts as demand evidence until a payment or a signed pilot follows. For B2B SaaS and fintech, a paid pilot with a few design partners is often the clearest willingness-to-pay evidence you can get. A concierge pilot that customers pay for produces the same class of evidence, but requires more manual work.
What it does not show: that the product will be usable by people who never had your onboarding call, or that the pilot scope is feasible at the price. Pilots also tend to attract customers who are unusually motivated. A pilot conversion is a strong signal for the segment that converted and a weak signal for anyone else.
Proof of concept: feasibility
A proof of concept builds the risky part and only the risky part. If the product depends on extracting structured data from messy documents, the PoC is the extraction pipeline on real samples, measured by accuracy, latency, and cost per document. If it depends on an integration or a regulatory workflow, the PoC is that workflow end to end, ugly and unpolished.
A PoC also doubles as investor evidence, because showing the hard part working is more convincing than describing it.
Merge's POC design service is a design-led concept-validation package: a product concept and architecture with a clickable prototype, fast branding, and an investor-ready pitch, delivered through a discovery, ideation, research, and deliverables sequence. It fits when you need to validate the concept and the pitch. It is not the engineering proof of concept described above, because that is an engineering task with its own scope.
What a PoC does not show: demand or willingness to pay. A working pipeline that nobody wants is an expensive pre-build failure, which is why the selection rule below says feasibility should not be the first test unless it is the riskiest assumption.
The pretend-first tests in this matrix (the landing page or video, the fake door, and the Wizard of Oz setup) are what Alberto Savoia, formerly Google's "Innovation Agitator," calls pretotyping, and his one-line summary applies to every row above:
Make sure you are building the right it before you build it right.
Mapping your riskiest assumption to the lightest valid test
Picking the right method matters more than knowing all of them, and the choice starts with the assumptions.
How to find your riskiest assumption
Write down every assumption the product depends on, as plainly as you can. The procedure is the same whether you are validating a SaaS, fintech, or app idea before building it. A working list for a pre-seed SaaS is usually longer than it first looks: who the user is, what they do today, why they would switch, what the product must do in the first session, what you can build with the team you have, what you can charge, and how you reach the first customers.
Teresa Torres, a product discovery coach and author of Continuous Discovery Habits, offers a taxonomy that helps with the listing step.
In an October 2023 Product Talk article, she names five types of assumptions: desirability ("why we think someone will want our solutions"), viability ("why a particular solution will be good for our business"), feasibility ("why we think we can build our solutions"), usability ("what our customer is able to do"), and ethical ("why we think there's no potential harm in offering our proposed solutions").
Use the five types to check that your list is complete. They describe what you assume; the four evidence classes describe what a test returns. The two lists overlap but do different jobs, which is why the matrix and the map use the evidence classes.
Torres also argues for cadence in a separate piece on assumption testing:
A regular cadence of assumption testing helps product teams quickly determine which ideas will work.
In practice, for a pre-build team, we read that as testing assumptions one at a time on a steady schedule rather than in one big validation phase.
Then rank. David J. Bland, co-author of Testing Business Ideas with Alexander Osterwalder and founder of Precoil, describes assumptions mapping in a 2020 Strategyzer article as:
A team exercise where desirability, viability, and feasibility hypotheses are made explicit and prioritized in terms of importance and evidence.
The two axes are the whole method: how important is the assumption to the business surviving, and how much evidence do you already have for it? The assumption that scores high on importance and low on evidence is the one you test first.
Rik Higham, a senior product manager, made the same argument in 2016 in a Hackernoon piece titled "The MVP is dead. Long live the RAT.":
There is a flaw at the heart of the term Minimum Viable Product: it's not a product... Instead of building an MVP identify your Riskiest Assumption and Test it.
RAT stands for riskiest assumption test, and the acronym is a reminder that the unit of work is one assumption, not one product.
If the list-and-rank step is where your team stalls, or if the founders disagree about the riskiest assumption, that disagreement is the first thing to resolve. Merge runs it as product UX discovery, which starts with user and competitor research and ends in a prioritized scope, but a whiteboard session is usually enough to get to a defensible shortlist.
The "lightest valid test" selection rule
Once you know the riskiest assumption, pick the test with three checks, in order:
- Which evidence class would settle the assumption? "Clinic managers will switch from their current tool" is a demand question. "We can pull bookings from both legacy systems" is a feasibility question. "Managers will pay the monthly price we have in mind" is willingness to pay.
- Which is the cheapest method that produces that class? For demand, a landing page beats a concierge MVP. For willingness to pay, a pre-sale beats building billing. For feasibility, a PoC on real data beats a full build.
- What will you observe, and what result did you decide in advance would count as a pass? For a landing page, that is a conversion rate against a threshold. For an interview round, it can be a count: how many of the people you talk to describe the problem unprompted. Either way, the observation, the decision rule, and what the sample cannot tell you are written down before the test starts. Without them, the result will be read in whatever direction you were already leaning.
"Lightest" means the cheapest method that still produces the right evidence class. It does not mean the cheapest method available. A landing page is the lightest demand test, and it is the wrong test for a feasibility assumption no matter how cheap it is. This is also the answer to "how do I validate an MVP idea fast": speed comes from testing one assumption with the smallest matching instrument, not from running every method at once.
Assumption → evidence class → test → success signal
Here's an example for a pre-seed fintech product that reconciles payouts for marketplace sellers. Swap in your own assumptions; the columns stay the same.
Riskiest assumption | Evidence class | Lightest valid test | Success signal (set before the test) |
Marketplace sellers are interested enough in payout reconciliation to request access | Demand | Landing page with a "get early access" form, traffic from the seller communities you would sell through | Conversion rate on a fixed visitor count, threshold set in advance |
Sellers will pay for reconciliation rather than use a spreadsheet | Willingness to pay | Pre-sale of a founder-priced annual plan to sign-ups from the landing page | A fixed number of paid commitments within a set window |
Sellers can categorize a payout in the interface without help | Usability | Clickable prototype tested with a small group of sellers | A fixed share complete the core task unaided |
We can pull payout data from the two largest marketplace APIs at the required frequency | Feasibility | PoC against both APIs with real seller accounts | Data pulled at the target interval with an error rate below the threshold |
Sellers will keep using it after month one | Demand (behavioral) | Concierge: reconcile manually for a small group of sellers for a fixed period | A fixed share ask to continue after the pilot ends |
Read the map top to bottom and notice that no single method covers more than two rows. That's why the matrix exists: a founder who runs one landing page and calls the idea validated has evidence for row one and nothing else.
Define your success signal before you run the test
What a pass looks like for each evidence class
- Demand: a conversion rate on a fixed amount of traffic from the channel you will actually use, or a fixed number of qualified sign-ups within a time window. Qualified means they match the segment, not just that they entered an email.
- Usability: the share of participants who complete the core task without help, plus the specific step where the rest stopped.
- Feasibility: a measurable property of the hard part on real inputs: accuracy, latency, cost per unit, or a clean end-to-end run through the regulated workflow.
- Willingness to pay: money received, a signed and dated pilot agreement, or a purchase order. A non-binding letter of intent is a weaker signal and should be labeled as one. Verbal enthusiasm does not count in this class.
Set the threshold before you see the number, not after
Write the pass threshold and the sample size down before the test starts, and tell someone. A result you see first and interpret second tends to get interpreted in your own favor. If you set a threshold for the share of visitors who request access and the result came in well under it, the test missed its threshold. That is not the same as the assumption being wrong. Before you decide whether to repeat, change, or stop, check the things that can produce a miss on their own: whether the traffic came from the segment you meant, whether the page said what you meant it to say, and whether the sample was large enough to mean anything. If those hold up, the assumption is probably wrong. If you never wrote a number down, a weak result will look good enough to keep going.
Thresholds should also tie to the business, not a benchmark. For a B2B product at a high price point, a small number of paid pilots from a short list of conversations may be a strong pass. For a consumer app that needs volume, a landing page conversion rate that would satisfy a B2B team may be a fail. The number depends on what must be true for the business to work, which is why the assumption list from the previous section comes first.
Common mistakes when interpreting a pre-build test
Mistaking sign-ups for willingness to pay
This is the misreading the four-class model exists to prevent. A waitlist shows that people were curious enough to type an email. It does not show that any of them will pay, and a large waitlist can still mean zero paying customers. If price is one of your riskiest assumptions, a demand test won't answer it. Run a pre-sale or a paid pilot.
Treating a small sample as final
A handful of usability sessions will usually surface the biggest flow problems. They are not enough to say "users can use it." Sign-ups from a founder's own network are not a market demand signal. Report the result with its sample size, and if the sample came from a channel you will not use at scale, say so. Small samples are fine for deciding what to fix next and poor for deciding to build.
Skipping the feasibility check after a strong demand signal
A strong demand result creates pressure to start building. For products whose core depends on an external data source, a model's accuracy, a regulated workflow, or an integration nobody on the team has done before, that pressure is where budgets go wrong. Not every feasibility assumption needs its own PoC. The ones that do are the ones that rank high on importance and low on evidence, the same rule as everywhere else in this guide. For those, run the PoC before the full build, even if the demand test passed with room to spare. The cost of a PoC is small compared with discovering the hard part doesn't work after the full build is underway.
When it's time to bring in a partner
You have a demand signal and need a clickable prototype or PoC fast
Interviews, a landing page, and a pre-sale can all be run by a founding team with no design or engineering help. Two of the six methods usually cannot, at least not quickly: the clickable prototype needs a designer who has done onboarding flows before, and the PoC needs an engineer who can make the risky part work on real data without building everything around it. If the assumption you are testing needs one of those, and your team cannot produce it quickly, that's when an outside team earns its fee. The mistake is bringing one in earlier, to build the whole product before the assumptions are tested.
What Merge has helped founders validate before building further
Waffly came to Merge with an idea for AI semantic search inside an audio player. In the matrix's terms, the first deliverable, a proof of concept delivered in one month, was a feasibility test.
Val Wikstrem, CEO at Waffly:
Achieving the delivery of a POC within just one month helped us to fundraise our angel round and begin to work on the complete MVP. Within three months through the help of Merge was a remarkable success in my view.
The PoC served two purposes at once: a feasibility test and an investor test, before the full MVP was built.
Agentless, a platform for U.S. home buyers who want to complete a purchase without an agent, started with technical feasibility research and market validation before any UX or development work. The open feasibility questions were whether property data could be obtained automatically and whether per-state legal rules could be handled, a clear case of feasibility being the riskiest assumption.
NFTBull needed a design-first MVP for hypothesis testing and concept validation, delivered as an MVP design and a Webflow site within a two-month window. The deliverables were design-led: an interface and a site the team could put in front of people to test its hypotheses and validate the concept before committing to a full platform. The case study records what was delivered and why, not usability or demand measurements.
If you have a tested assumption and the next step is a prototype, a PoC, or the first real release, Merge's MVP development services start from the evidence you already have and scope the smallest useful release around it.
What to test next after your MVP testing passes
What to build first
A passed test narrows the build. With demand and willingness-to-pay evidence in hand, the next question is which features belong in the first release and which can wait. Merge's guide on how to prioritize features in MVP design covers five prioritization techniques and the mistakes to avoid at that stage.
The full build-order guide
Once you've tested the assumption list and scoped the first release, you move to sequencing: what to build first, what to build next, and what to leave out entirely. That is a separate article: MVP development for startups: what to build first and why. It covers build order, what the first release has to include, and how that changes by startup type.
Usability testing vs MVP testing: where each fits
When your test is about an existing product or working prototype
Usability testing appears in the matrix above as one row, because it produces one evidence class, and it can run before any production code exists, on a clickable prototype. A related topic is user testing on a product or working prototype you already have: how people use it, which steps confuse them, which of two flows performs better. Merge's article on user testing methods and how to choose the right approach for your product covers moderated and unmoderated sessions and the rest of that toolkit. The difference between the two isn't whether a prototype exists – it's the question you're asking: whether an assumption holds before you commit to a build, or how something you have already built performs.
FAQ
What is MVP testing?
MVP testing is the set of experiments a team runs to check the assumptions behind a product idea before building it. Each experiment produces one of four kinds of evidence: demand, usability, feasibility, or willingness to pay. Eric Ries uses "MVP" for the experiment itself; Marty Cagan of SVPG calls it an "MVP Test" to keep it distinct from the MVP as a product. This guide follows Cagan's split.
What is the purpose of testing your MVP?
To find out, at the lowest possible cost, whether the assumption the product most depends on is true. In Eric Ries's words from 2011, the goal is "to test fundamental business hypotheses." A secondary purpose is to reduce the uncertainty around scope: a passed demand test and a passed feasibility PoC tell you which assumptions the first release no longer has to hedge against, which narrows the build without defining it.
How do you validate an MVP?
List the assumptions the product depends on, rank them by importance and by how little evidence you have, and run the lightest test that produces the evidence class the top assumption needs. Set the pass threshold before you run it. Then move to the next assumption. Validation is a sequence of small tests, not a single event.
What's the difference between an MVP and a PoC?
A proof of concept shows that the hard part can be built and answers only the feasibility row of the matrix above; an MVP is the smallest product real users can use, and the team can support.
