A product-market fit score can give a team a shared language for a difficult question. It can also create false certainty if the number becomes detached from the people who answered, the experience they had and the business the company is trying to build.
Superhuman’s published survey account is worth reading because it makes an early product-learning process concrete. It is worth examining carefully because a memorable threshold is easier to repeat than the conditions under which a measurement is useful.
This article separates the historical report from Foundshore’s proposed evaluation method. It is not an independent audit of Superhuman, a statement about the company’s current performance, or a claim that the same score certifies every business.
What the historical account reports
In a November 12, 2018 First Round Review article, founder Rahul Vohra described work that began in summer 2017. The team surveyed people who had experienced the product, examined which users would be strongly disappointed to lose it, segmented responses and used feedback to inform product work. Vohra reported a starting score of 22%, a move to 33% after segmentation, and 58% within three quarters of subsequent product improvement. 1
These are reported survey results from a founder-authored case. They are not a trial-to-paid conversion rate, a retention curve or a controlled estimate of the process’s effect. Segmentation changes the interpretation of the denominator; product improvement and the passage of time also matter.
The useful starting point is that the team connected customer descriptions to decisions rather than treating every request as equally representative.
A diagnostic is different from a verdict
A diagnostic organizes inquiry. It may help identify a group receiving strong value or expose a gap between the founder’s story and the user’s experience. A verdict declares that the business has passed a test and can proceed without further examination.
Treat the score as the first kind of tool. A favorable response can coexist with weak purchasing authority, expensive onboarding or a small reachable market. An unfavorable response may reflect early users who never reached the product’s core task.
The next question is therefore not just “Is the number high enough?” It is “What does this number tell us about this defined group, and what decision can we responsibly make from it?”
Define who can answer meaningfully
Before distributing a survey, describe the experience required to answer it. Have participants completed the relevant task? Did they use a production version or a demonstration? Did a founder perform substantial work on their behalf?
For a B2B product, the daily user, purchasing owner and technical approver may have different perspectives. Record those roles rather than combining them into one undifferentiated user population.
Maintain the invitation and response counts. A result from respondents alone does not tell you what nonrespondents think. You do not need to invent a statistical correction, but you do need to keep the missing information visible.
These are proposed measurement practices, not claims about every detail of Superhuman’s historical survey operation.
Preserve comparable groups over time
When you narrow the target user, the change may be strategically sensible. It can also raise a score without the experience improving for any particular person. That does not make segmentation wrong; it changes what the result establishes.
Track both the chosen segment and the broader group where possible. Record product version, eligibility conditions, assistance and observation period. Avoid presenting a sequence of differently selected samples as if it were a controlled panel of the same users.
A comparison note can be short: “This wave focuses on users performing the weekly reporting task; the earlier wave included exploratory signups.” Such a sentence can prevent a graph from making a stronger claim than the evidence supports.
Turn responses into product questions
Ask what benefit the respondent values, how they currently obtain it and what prevents further use. Look for common tasks and constraints rather than only common job titles.
For a fictional workflow tool, several enthusiastic users might value reliable exception review more than faster first-pass output. That finding could change positioning, product priorities and which customers the company approaches.
Do not automatically build every requested feature. Determine whether the request belongs to the intended use case, whether the problem is repeated and what an actual test would establish. A request is evidence of an expressed preference; it is not a guaranteed purchase commitment.
Pair attitudes with behavior and economics
Dimension | Question it helps answer |
|---|---|
Reported value | What do users say they would miss? |
Task completion | Can they obtain the intended result? |
Repeat use | Do appropriate users return when the task recurs? |
Commercial progress | Does the buying organization proceed under real terms? |
Delivery effort | What assistance is needed to sustain the result? |
Retention | Do mature customer groups continue, and why do others stop? |
No single row substitutes for the others. The relevant evidence depends on the business and stage. A product used for an infrequent but important task should not be judged by a daily-use expectation simply because daily metrics are easy to collect.
Run one decision cycle instead of creating a dashboard ritual
Choose a question for the research. Collect responses from a defined group, examine the reasons behind them, and identify a limited change to test. Write what you expect to happen and what would challenge the proposal.
After the change, review behavioral and qualitative evidence alongside the survey. Did the task become easier? Did the target group understand the offer better? Did a commercial obstacle remain despite improved satisfaction?
Keep an explicit record of decisions not taken. A recurring request might be deferred because it belongs to a different segment or would make delivery unsustainable. That reasoning is part of product strategy, not a failure to listen.
What founders should avoid borrowing
Do not borrow a threshold while ignoring sample selection. Do not describe a percentage as scientific proof of company-wide fit. Do not use an attractive survey result to conceal poor retention or manual delivery costs.
Also avoid the opposite mistake: dismissing qualitative and attitudinal evidence because it is imperfect. Early-stage decisions rarely arrive with complete information. A well-defined diagnostic can be valuable precisely because it improves the next question.
A mentor or investor can challenge the interpretation, but the team should be able to show the source data, limits and reasoning without relying on a famous case as authority.
The case’s lasting usefulness
Superhuman’s account offers a concrete example of customer feedback entering a product decision process. It does not remove the need to understand the company’s own users, task, business model and stage.
Use the story to make your learning more explicit. Define who is speaking, understand what they value, distinguish selection from improvement, and test the decisions that follow. The number becomes useful when it supports that work—not when it allows the team to stop asking what customers actually need.
Sources and research scope
[1] How Superhuman Built an Engine to Find Product Market Fit — First Round Review / Rahul Vohra. Founder-authored historical case. Published 2018-11-12. Reviewed 2026-10-04. 2017–2018 survey process and reported movement from 22% to 58%. Not retention, conversion or a universal PMF threshold.






