15 Assumption Testing Methods For Product Management Teams
“Assumption testing means checking the individual beliefs your product idea rests on, instead of building the whole idea and hoping it works. It beats idea testing because you don’t fall in love with a single assumption the way you do with a whole idea, so you’ll actually kill a bad one. The core tool is the Riskiest Assumption Test (RAT): list every assumption your idea needs to be true, pick the one most likely to sink you, and design the cheapest possible test for it. DoorDash tested “will people trust a third-party website for food delivery?” with a one-afternoon website and PDF menus before building anything. This guide covers how to plan assumption tests (assumption boards, opportunity solution trees, the four product risks) and 15 methods to run them, ordered from lowest-fidelity “asking” methods like user interviews up to “building” methods like prototypes and A/B tests.”
Why the best teams test assumptions instead of ideas
Assumption testing forces us to name what we believe to be true about our idea and find ways to prove it. It's not the same as idea testing. We aren't checking whether the whole idea is viable, we're checking whether the pillars it will stand on are as solid as we assume.
That distinction matters. The more time you spend on an idea, the more likely you are to fall in love with it. We don't get the same attachment to an assumption, because it's one small piece of the whole. Disproving a single assumption is far easier than killing an entire idea in one go.
Testing a whole idea means building a Minimum Viable Product (MVP). The problem with MVPs is right there in the name: "product". An MVP is the perfect excuse to skip assumption testing and start building. You may intend for it to stay small, but the definition tends to drift toward "Minimum Loveable Product" as you pile on more risky, untested features. The more you build, the harder a bad idea becomes to kill.
Move away from testing the idea and onto the assumptions, and you replace the MVP with the RAT, the Riskiest Assumption Test, a three-step process:
List the assumptions that must be true for your idea to succeed.
Identify which individual assumption is most important for that success.
Design and conduct a test to prove whether this assumption is true or not.
The best part about the RAT is that it doesn't need a full prototype or launch-ready product. It explicitly discourages one.
Riskiest Assumption Test: a real example (DoorDash)
In late 2012, the DoorDash founders were interviewing the owner of a macaroon shop in Palo Alto named Chloe. At the very last moment of the interview, Chloe received a call from a customer requesting a delivery and turned the order down. The founders were mystified: why would she turn business away?
It turned out Chloe had "pages and pages of delivery orders" and "no drivers to fulfil them". Over the following months, they heard "deliveries are painful" over and over while interviewing 200+ other small business owners. The supply side had clear potential: restaurant owners needed someone to handle deliveries for them.
But the DoorDash founders decided not to build anything for restaurant owners at the start. When they first launched, they didn't even tell the featured restaurant owners. In January 2013, the team created paloaltodelivery.com in a single afternoon: a basic website with PDF menus from a handful of local restaurants and a phone number for placing orders. They charged a flat $6 delivery fee, had no minimum order size, and collected payments from customers manually in person.
The point of paloaltodelivery.com wasn't to prove the idea to restaurant owners. The founders knew their riskiest assumption was whether people would trust a third-party website for their food deliveries. This approach was unheard of at the time. If the assumption turned out to be wrong, they'd have to find a completely different way to solve restaurant owners' delivery problem.
That isn't how it went. The same day they created the website, they received their first order: someone searching for "Palo Alto Food Delivery" found their site through a cheap AdWords campaign. Within weeks, the four founders were fielding so many orders, mostly from Stanford students, that they struggled to keep up. The experiment had proven their riskiest assumption true.
paloaltodelivery.com in May 2013, four months after its initial launch (via Wayback Machine)
How the best product teams plan their assumption tests
Assumption Boards
If you've ever taken part in a hackathon or Startup Weekend, you've likely been handed a Lean Canvas to map your assumptions. For teams just starting on a new idea, the Lean Canvas often pushes people to invent assumptions, giving them boxes to fill that they hadn't even considered. I prefer the more basic Assumption Board for early-stage teams: four boxes ("Hypothesis, Assumptions, Invalidated, and Confirmed") to capture their assumptions, identify which are riskiest, and track learnings from their tests as they de-risk the idea.
Back when I was a Product Innovation Lead at Unilever, we had sprint weeks where we'd complete a full iteration of the assumptions board every day for five days straight. We'd start each morning by identifying our current riskiest assumption, by lunchtime we'd started an experiment to test it, and before 5pm we'd update the board with the result. And we were inventing new types of ice cream, so you can't claim this pace is too fast for software. Assumption boards rely on your ability to identify your research hypothesis, and there's a good guide by Teresa Torres to help with that.
(Sorry for the blurry quality, these are screenshots of a video)
Opportunity Solution Trees
The Opportunity Solution Tree (OST) is a visual framework that helps product teams map their research space and figure out what to prioritise next. Instead of fixating on specific customer problems or feature ideas in isolation, the OST helps teams brainstorm the customer opportunities (unmet needs, pains, desires), the solutions that could address each one, and the experiments that test those solutions.
There are two less obvious tools built into Opportunity Solution Trees. The first is reframing: take what you're fixated on (say, a new feature idea), identify what it will accomplish (the customer problem your idea solves), then brainstorm other ways to accomplish that same objective (other feature ideas). There's more on reframing in my problem-brainstorming guide.
The second, compare-and-contrast decisions, matters more to assumption testing. Research shows that the more ideas you generate, the better your ideas tend to be. The same goes for product decisions: we're better at picking which customer problem to focus on once we've identified a wider range to consider. Teresa Torres, creator of the Opportunity Solution Tree, calls these "compare and contrast" decisions and "whether or not" questions.
Using an OST tool like Vistaly or Olta, teams can pick an opportunity or solution and add the underlying assumptions as a note on that node. Then, once they've picked which one to test, they have a paper trail of the assumptions made when that node was first added.
Problem Assumptions vs Solution Assumptions
When most people think about assumption testing, they think about idea validation and proving they're solving a customer problem. But as product management leader Marty Cagan says, "People don't buy the problem, they buy your solution, so don't spend a lot of time on the problem because you need as much time as possible to come up with the winning solution!"
This sounds like the opposite of most customer discovery advice, but there's an important nuance in Marty's words. When people validate a high-priority customer problem, they often assume this also validates their overall idea, when really all they're doing is skipping every assumption they've made about their intended solution. This is another reason the Opportunity Solution Tree is so useful. Splitting opportunities (customer needs and pains) from solutions pushes us to map our assumptions for each separately. If you're not using an OST, you can still add a step to your workflow to identify problem and solution assumptions separately, so you don't skip it.
The Four Product Risks
Marty doesn't just tell us to consider solution assumptions, he gives us the framework to do it, the Four Big Product Risks:
A traditional product trio tends to divide ownership of each risk like this:
Product Manager → Valuable + Viable
Product Designer → Usable
Product Engineer → Feasible
As Cesar Tapia points out, risks and assumptions are opposite sides of the same coin: risks hurt when they're true, assumptions hurt when they're false.
Planning Your Assumption Tests
I like these frameworks because they make it 10x easier to plan good assumption tests. The Opportunity Solution Tree, for one, forces us to ask questions about the links between the company's desired outcome, the problems customers face, and the solutions we think might solve them. That raises questions like "Is this the only way a customer would try to solve this problem?", "Is this solution too complex for the type of customer experiencing this need?", and "Will this experiment give me a new or clearer perspective on the proposed solution?".
The Four Big Risks gives even more targeted questions about our assumptions, especially the ones tied to solution ideas: "Do we have the time and resources to accomplish this proposed build?", "Does this feature solve a problem that's high enough on the customer's list of priorities?", and "Does this solution actually fit into the broader picture of how our company provides value to customers?". Just like in idea brainstorming, the more assumptions we can identify, the more likely we are to correctly diagnose the riskiest one in most need of testing.
How the best product teams test their assumptions
The best product teams run 10-20 experiments every week. If that number sounds absurd, it's because you're still thinking about assumption testing the wrong way. The goal is to avoid unnecessarily expensive experiments: the ones that take the most time, risk your credibility, or involve one-way decisions that can't be undone. My guiding principle is that the earlier you are in a project, the lower-fidelity your experiment should be. As certainty and conviction grow, the experiments move from "asking" methods to "building" methods.
Qualitative Research
1. User Interviews. The lowest-fidelity method of all is simply to ask customers questions that test your assumptions, the heart of discovery research. Recommended resources: The Mom Test for founders, or Deploy Empathy for product managers.
2. User Feedback. Customer support, success and account management teams sit on a treasure trove of complaints, suggestions and feedback. A few keyword searches can point you where to start digging.
3. Ethnography. Ethnography means spending time with or observing people to learn about their lives, feelings and habits. Watch customers attempt a routine workflow or task using their existing solution (or your own product). The more natural the situation, the better.
Quantitative Research
4. Scenario Testing. Scenario testing is a survey method that presents people with hypothetical choices and measures their preferences from the decisions they make. Unlike rating-scale questions (say 1-5 stars), scenario tests force people to compare options directly, giving you better insight into how they might act in a real situation. Common formats:
Pairwise Comparison: breaks a list into a series of head-to-head "pair votes", measures how often each option is picked, and ranks the list by relative importance.
Ranked Choice Voting: presents the full list to be ranked in order of preference.
Points Allocation: gives people a pool of credits to allocate across the options however they see fit, measuring the magnitude of their preferences.
All of the above Scenario Testing formats are available for free on OpinionX
OpinionX is a free research tool for creating scenario-based surveys for assumption testing using any of the formats above. It's free, and includes analysis features like comparing results by customer segment.
5. Paraphrase Testing. Paraphrase testing helps you work out whether people interpret a statement the way you expected. No fancy equipment needed: show the statement alongside an open-response text box and ask people to write what it means to them. You can count keywords or use thematic analysis to interpret the answers.
6. Quant Survey. I tend to opt for scenario testing on a survey-based assumption test, but if your assumption relates to an objective fact (how much time someone spends on a task per week, or the financial cost a problem represents to their team) then a traditional quantitative survey can work well. There's a separate guide on best practice and common pitfalls for quant surveys.
Demand Testing
7. Ad Testing. Online ads can source participants for your assumption test. DoorDash used AdWords to test whether anyone was searching for food delivery in Palo Alto. You can run Instagram ads with variations of the same image and text to see which gets the highest click-through rate. At Unilever, we used Facebook ads with concept product images that directed people to short surveys, telling us click-through rates alongside qualitative impressions and how well people understood the concepts.
8. Dry Wallet. Dry Wallet tests lead customers through a checkout experience to a dead end, like an "Out of Stock" message. These let customers prove their intention to buy without you investing in the product upfront, and they're particularly common among ecommerce companies.
Groupon: the founders created a blog with fake deals and offers to see if people would sign up. They didn't build the product until they'd proven the demand.
Buffer: the social media scheduling tool created a landing page and set of pricing plans before ever building the product. Click through a pricing plan and you were told the product wasn't ready yet, then shown an email sign-up form to be notified about launch.
9. Fake Door. Fake Door tests are like Dry Wallet tests but focused on a specific feature rather than a mock purchase of a whole product. Dropbox created a landing page video using paper cutouts to explain the original concept. Founder Drew Houston later said that "[the video] drove hundreds of thousands of people to the website. Our beta waiting list went from 5,000 people to 75,000 people literally overnight. It totally blew us away."
Prototype Testing
10. Impersonator. The Product Impersonator test uses competitor components to deliver the intended product or service without the customer knowing. Product Loops offers two good examples:
Zappos: initially bought shoes from local retailers as orders came in, instead of buying inventory upfront.
Tesla: in 2003 (pre-Elon Musk), Tesla built a prototype fully-electric roadster using a heavily modified, non-functional Lotus Elise to show prospective investors and buyers what the final design might look like.
11. Wizard of Oz. A quick way to test an assumption is to offer functionality or services to customers without actually building the process that delivers them. If the customer requests the service or pays upfront, you scramble to fulfil it manually behind the scenes.
DoorDash: the team had no delivery drivers or order-processing system when they first launched; they just sent whoever on the founding team was closest to the restaurant to pick the food up when an order came in.
Anchor: "We hired a couple of college interns, we said to them that people are going to push this magical button and say 'I want to distribute my podcast' and your job is to do all that manually but to them it's going to feel magical like it happened automatically. We just had college students submitting hundreds of thousands of podcasts." (Maya Prohovnik, VP of Product)
The video below talks about building "Wizard of Oz" features to close enterprise deals, like a "Generate Report" button that shows a "Report will take 48-72 hours to compile" notification while an employee creates the report manually.
In episode 679, join @RobWalling for a solo adventure where he answers listener questions about enabling “mock features” for closing big sales, phased launches and recovering from failed launches, content marketing for SaaS apps, consulting, and more.https://t.co/o4FYg0hhGE pic.twitter.com/Wpq64rACIE
— Startups For the Rest of Us (@startupspod) September 19, 2023
12. Wireframe Testing. The most common wireframe testing methods are usability tests (asking users to complete tasks with interactive prototypes) and the five-second test (showing an image or video of the prototype and asking users to explain what they see), a first-click test, or image-based pair voting (showing two images at a time to measure preferences).
Vanta: the first version of Vanta was just a spreadsheet they shared with new customers, testing whether a templated approach to SOC 2 applications could work. As founder Christina Cacioppo explains, "We started with really open-ended questions, then moved to spreadsheet prototypes, and then moved to prototypes generated with code. At the end of the six months, we started coding."
Product Analytics
13. A/B Testing. A/B testing is a common variable test that splits your users into two groups shown two different versions of the product, then tracks the difference in their usage. It's most commonly used for optimisation work (improving existing functionality), though some companies (like Spotify) use it to test the impact of new functionality too.
A/B testing needs large sample sizes to be statistically significant, which limits its use mostly to large companies.
“Unless you have at least tens of thousands of users, the [A/B test] statistics just don’t work out for most of the metrics that you’re interested in. A retail site trying to detect changes that are at least 5% beneficial, you need something like 200,000 users.”
14. Feature Adoption Rate. Most product analytics tools can track feature adoption rate. You can ship a V1 feature and measure how many users try it as a gauge of overall interest. This is how we launched our first needs-based segmentation feature for OpinionX: we shipped a really simple segmentation filter that only worked on the answers to one individual question. Customers quickly asked us to expand it to segment more types of data at once and to compare and correlate different customer segments.
15. "Ship It And See". This isn't an assumption test, it's an "everything" test. Be ready for a hard fall if you've jumped straight to shipping your idea with no experimentation, research or assumption testing along the way.
Source: @adhamdannaway on Twitter/X
Defining Success and Failure Upfront
Defining what success looks like before you run an assumption test reduces confirmation bias and, more importantly, stops teams disagreeing over what the results meant afterwards.
Clear purpose. The only way to define success for a test is to first understand why you're running it. The Opportunity Solution Tree helps here because its tree structure links everything together, and the Four Big Risks framework helps you articulate which aspect is riskiest. If your test is for a new feature, you know what customer need or pain the solution must address (OST) and which aspect of the solution you're most uncertain of (4BR).
Example 1: Vanta. Here's a clear example of assumption testing in action from Vanta's founder Christina Cacioppo, taken from a profile by First Round Review. In their first experiment, they went to Segment, a customer data platform, and interviewed its team to determine what the company's SOC 2 should look like and how far away it was from getting it. "We made them a gap assessment in a spreadsheet that was very custom to them and they could plan a roadmap against it if they wanted," she says.
Cacioppo was running a test to answer two questions: could her team deliver something credible, and would Segment think it was credible? The answer to both was "yes." And so the first low-tech version of Vanta was born, as a spreadsheet. "It actually went quite well, so we moved on to a second company, a customer operations platform called Front," she says.
For this experiment, Cacioppo wanted to test a new hypothesis: could she give Front Segment's gap assessment without telling them it was Segment's, and would they notice? "We used the same controls, the same rules and best practices, and still interviewed the Front team to see where Front was in their SOC 2 journey, so it was customised in that sense," she says. "But this test was pushing on the 'Can we productise it? Can we standardise this set of things?' And most importantly, 'Can they tell this spreadsheet was initially made for another company?'" They couldn't. And then Cacioppo got an email that sealed the deal in terms of validating her idea.
Christina outlines an initial assumption (can Vanta credibly provide SOC 2 compliance as a service), tests it, updates her riskiest assumption (the SOC 2 application process could be standardised), and tests that too. The assumptions are so clearly articulated that the outcome was an obvious success or failure.
Example 2: OpinionX. Six months after launching OpinionX, we lost our one and only paying customer. Having run 150+ interviews up to that point, we were sure our key problem statement, "it's hard to discover users' unmet needs", was the right pain point to focus on. Losing our only customer made us question whether we'd really tested that assumption. So we ran a quick test: we compiled a list of 45 problem statements and shared them with 600 target customers, of whom 150 agreed to help. Each person was shown 10 pairs of problem statements and asked to pick the bigger pain. Using that data, we ranked the statements from most to least important. And our key problem statement came DEAD LAST.
We'd spent over a year building a product that solved the wrong problem, all because we never tested whether the problem was a high priority for customers. When we interviewed participants from that ranking test, we learned they didn't struggle to discover unmet customer needs at all. On the contrary, many teams felt like they were drowning in unmet needs and couldn't work out which to prioritise.
We pivoted the product from problem discovery to problem ranking. One week later, we had four paying customers. There's more on this story here, or in the video below.
Comparing segments. Teams often fail to account for how customer segments influence the results of their assumption tests. Here's a 30-second example I made that shows how a test can look like a failure in aggregate while showing the opposite once you isolate a specific customer segment.
To account for segment differences, ask whether you've baked any assumptions about customer segments into the test you're designing. If so, you need a way to filter or compare results by segment. If that's not possible, isolate the segments (say, by creating identical but separate tests for each key segment) or test your segment hypothesis first.
One last tip: the difference between top-down and bottom-up segmentation. If you're building a test to measure feature importance, you can assume the results will vary by the pricing plan each customer is on. In top-down segmentation, you define your key segments upfront.
Imagine a research project where customers rank problem statements to show you their highest priority. Plenty of data points could shape the results, from company size and industry to seniority and job title. In cases like these, you want to collect enough data across different categories (or enrich results with existing customer data) to identify which segments move the results most afterwards.
Taken from my guide to needs-based segmentation
That's bottom-up segmentation: you let the data tell you which segments move your results most. There's more on bottom-up segmentation here.
When to test your assumptions, and what it saves you
The teams that win aren't the ones with the best ideas, they're the ones who find out fastest which of their beliefs are wrong. Assumption testing is how you do that cheaply: name the belief your idea rests on, pick the riskiest one, and run the lowest-fidelity test that can prove it false. Start with asking methods, move to building methods only as your certainty grows, and decide what success looks like before you run the test.
Frequently asked questions
What is assumption testing? Checking whether the individual beliefs a product idea depends on are actually true, rather than building the whole idea to find out. You list the assumptions your idea needs, identify the riskiest, and run a cheap test to prove or disprove it.
What is the Riskiest Assumption Test (RAT)? A three-step alternative to the MVP: list the assumptions your idea needs to be true, pick the single one most likely to sink it, and design the cheapest test that can prove whether it holds. It deliberately avoids building a full product.
What's the difference between assumption testing and idea testing? Idea testing evaluates a whole idea at once, usually by building an MVP, which is expensive and easy to fall in love with. Assumption testing breaks the idea into its underlying beliefs and tests them one at a time, so a wrong belief is cheap to catch and easy to abandon.
What are the main assumption testing methods? They range from low-fidelity asking methods (user interviews, user feedback, ethnography, scenario testing and quant surveys) to demand tests (ad tests, dry wallet, fake door), prototype tests (impersonator, Wizard of Oz, wireframe testing) and product analytics (A/B testing, feature adoption rate). Start low-fidelity and move up as your certainty grows.
How do product teams decide which assumption to test first? By identifying the riskiest one: the assumption that, if false, would do the most damage to the idea. Frameworks like the opportunity solution tree and the Four Big Risks (value, usability, feasibility, viability) help teams surface and prioritise their assumptions.
Every team on this list found out they were wrong about something before it got expensive to be wrong. That's the whole trick, and it's cheaper than the alternative every single time.
Want more like this? Over 42,000 researchers and product people get The Full-Stack Researcher in their inbox. Subscribe for the next one.
You can run scenario-based assumption tests on OpinionX for free: $0, unlimited surveys, unlimited researcher seats, capped at 25 participants per survey, then $900 a year to lift the cap (full pricing). Pairwise comparison, points allocation and ranked choice voting are all included, with results you can compare by customer segment.