Which Customer Problems Are Actually Worth Solving?
A customer problem is worth solving when evidence shows a consequential and persistent outcome gap in a coherent segment, users already pay to work around it, the need has credible future pull, the team can plausibly change it, and a bounded test could prove the opportunity wrong. Summary
At 10:14 on Tuesday morning, a large prospect asks for an executive dashboard.
At 10:19, sales marks the request urgent. It has a sponsor, a familiar solution, and a pleasingly crisp noun. By lunch, the dashboard has a row on the roadmap.
Seven current customers are doing something less presentable. Every Friday, they export transactions into spreadsheets, reconcile failed records by hand, and send screenshots through private channels because the product cannot explain which system is correct. Nobody has requested a “reconciliation confidence layer.” They have simply rearranged their week around its absence.
Which problem is worth solving?
The dashboard may be valuable. The reconciliation problem may be narrow, temporary, or outside the company’s ability to fix. The point is not that quiet problems always beat loud ones. It is that roadmaps reward legibility before they establish worth.
A named feature is easy to estimate. A visible buyer is easy to quote. A crowded voting board is easy to count. The more consequential question is harder: What outcome is failing, for whom, under which conditions, and what evidence shows that changing it is both valuable and possible?
A loud request is a witness. It is not the verdict.
The wrong question is “How painful is it?”
Many opportunity methods begin by rating pain, frequency, reach, strategic fit, willingness to pay, or confidence. These dimensions can improve a conversation. They become dangerous when the arithmetic creates more certainty than the evidence contains.
A founder may give one problem a pain score of nine and another a seven. A product manager may multiply each score by reach. A sales leader may raise the enterprise request because a contract is near. The spreadsheet returns a decimal. The decimal looks impartial. Its inputs are still judgments measured on scales that may not mean the same thing.
Customer problems are not lottery balls. One problem can have high frequency and low consequence. Another can be rare but catastrophic. A third can be important and widespread while remaining structurally impossible for this company to change. Adding unlike claims into one score lets strength in one category conceal absence in another.
The stronger unit is a problem warrant: a short argument that states why a defined outcome gap deserves the next level of investment. It does not promise that a product will succeed. It shows that the problem has enough independent evidence to justify reducing the next uncertainty.
The warrant has six questions:
- Consequence: What becomes materially worse if the problem persists?
- Behavior: What do customers already spend, build, delay, avoid, or risk because of it?
- Coherence: Which customers share the same cause and consequence?
- Future pull: Why should the need endure or spread rather than disappear?
- Solvability: Why can this team plausibly change the outcome?
- Disconfirmability: What result would make the team stop, narrow, or reframe the problem?
These are not six points on a scorecard. They are six witnesses. A glowing market forecast cannot repair an absent customer consequence. Ten interviews cannot compensate for a problem the company cannot affect. A strong workaround signal can justify another test, but not a full build when the intended segment is still incoherent.
Six witnesses must support the same problem claim
Consequence, behavior, segment coherence, future pull, solvability, and disconfirmability remain separate so one strong signal cannot hide a missing one.Consequence
Evidence requiredA costly outcome is visible
Reason to stop
No material consequence
Behavior
Evidence requiredPeople already spend effort
Reason to stop
Interest without action
Coherence
Evidence requiredOne segment shares the pattern
Reason to stop
Unrelated complaints
Future pull
Evidence requiredThe need is becoming more acute
Reason to stop
A fading edge case
Solvability
Evidence requiredA credible intervention exists
Reason to stop
No useful leverage
Disconfirmability
Evidence requiredFailure can be recognized
Reason to stop
No result could change the claim
Decision rule. A missing witness narrows or stops the claim; another witness cannot compensate for it.
The warrant is not an additive score. Each witness answers a different failure mode and can independently narrow, delay, or stop the investment.
Reconstruct the outcome beneath the request
A customer usually speaks in the language available to them.
“Add bulk editing.” “Give us an AI assistant.” “Build a mobile app.” “Make the dashboard simpler.” These statements are useful because they reveal where attention has landed. They are incomplete because they combine a situation, an obstacle, and a proposed remedy in one sentence.
Abbie Griffin and John Hauser’s voice-of-customer research treats customer needs as a hierarchy of primary, secondary, and tertiary needs. A primary need describes a broad outcome. Secondary and tertiary needs make that outcome more concrete. The distinction prevents a product team from confusing one interface idea with the progress it is meant to create.
Their work also tests a common shortcut: using frequency of mention as a substitute for importance. It is not necessarily valid. People repeat what is easy to notice and easy to say. A serious consequence may be rare, normalized, embarrassing, or difficult to explain. Satisfaction measures can also contain self-selection bias because dissatisfied customers may leave before the survey arrives.
John Narver, Stanley Slater, and Douglas MacLachlan make a related distinction between responsive and proactive market orientation. Responsive work addresses needs customers can express. Proactive work investigates latent needs: opportunities for value that customers do not yet recognize or cannot yet describe.
The practical move is simple. Rewrite a request without its proposed feature:
When this customer is in this situation, they cannot achieve this outcome, which creates this consequence.
“Add bulk editing” might become: “When an operations manager corrects a high-volume import, the system forces one-record-at-a-time repair, which extends a ten-minute exception into a two-hour interruption and increases the chance of inconsistent changes.”
That statement may still be wrong. It is now testable. The team can observe the situation, measure the interruption, compare workarounds, and discover that the actual cause is poor error classification rather than the absence of bulk editing.
Good discovery does not make customers speak product strategy. It stops the product team from treating the first proposed remedy as the final problem.
Costly behavior is stronger than stated preference
People are generous with hypothetical enthusiasm. Behavior is less polite.
A customer who says a problem is important has supplied one signal. A customer who spends four hours every week on a workaround has supplied another. A customer who buys a second tool, hires an operator, delays a launch, accepts legal exposure, or builds an internal system has revealed a cost that already exists.
This is why lead users matter. Eric von Hippel defines them as users who experience strong needs months or years before most of a market and expect substantial benefit from a solution. Because they already live in the future condition, they often attempt to solve the problem themselves.
Glen Urban and von Hippel’s 1988 PC-CAD study gives the idea behavioral teeth. The researchers contacted approximately 178 qualified users and received 136 screening questionnaires. A cluster of 38 lead-user firms differed sharply from 98 routine-user firms. Eighty-seven percent of the lead-user cluster had built its own PC-CAD system. Only 1% of the routine cluster had.
That is not a difference in survey adjectives. It is a difference in action.
The researchers then used lead-user information to develop a product concept. In a later evaluation, 78.6% of respondents chose that concept first among four alternatives. This was an early test in one 1985 industrial software market, with 71 complete or near-complete concept evaluations. It is not a universal conversion rate. It demonstrates a more durable pattern: customers who incur real costs to escape a constraint can reveal both the problem and the direction of future value.
The severe need appeared in what users had already built
Lead-user firms were far more likely to build their own PC-CAD systems, and a later sample strongly preferred the concept developed from their evidence.Behavior and preference are separate outcomes. The figure does not combine them into one effect.
Urban and von Hippel, 1988. Self-building is a costly behavioral signal; first choice is a separate concept preference. Neither percentage is a universal opportunity threshold.
The 3M evidence is larger in commercial ambition but needs equal restraint. Lilien and colleagues compared lead-user projects with conventional projects inside several 3M divisions. Five funded lead-user major-product-line projects carried an average management forecast of $146 million in annual sales after five years, compared with $62 million for 16 non-lead-user major product lines. The five lead-user projects represented $730 million in projected annual sales.
Those are forecasts, not realized causal revenue. The study took place in one company and used a natural experiment. Still, the intervention period produced seven major product lines in the studied divisions, five from lead-user work, against a historical average of 1.64 per four-year period.
The lesson is not “only interview extreme users.” It is “look for evidence that the problem has already forced adaptation.” A workaround, a shadow system, a repeated exception, a refusal to switch, or a home-built tool can reveal consequence and persistence more clearly than a preference rating.
Find the smallest coherent segment
“Small businesses need better reporting” is not a segment. It is a fog bank.
Two companies can share employee count and still experience different problems for different reasons. A retailer may need reporting because inventory and payments disagree. A consultancy may need it because project margin appears too late. A regulated clinic may need it because every adjustment requires an audit trail. One dashboard could touch all three. The causal problem is not the same.
A coherent segment shares four things:
- a recognizable situation;
- a similar causal obstacle;
- a consequence that matters in comparable ways; and
- a plausible route to adoption and value.
This definition prevents a team from averaging incompatible needs into a fictional customer. It also keeps a lead-user signal honest. A severe workaround in one advanced context does not prove that a broader market is ready. Urban and von Hippel did not stop after finding engineers who built their own tools. They tested the resulting concept with a broader sample of PC-CAD users.
Segment coherence is therefore not a demographic filter placed before research. It is a claim that improves as the team learns which conditions produce the same outcome gap.
Look for future pull without inventing a prophecy
Present demand can be a lagging indicator.
Most people cannot report a need created by conditions they have never experienced. A team researching AI-assisted operations in 2019, mobile payment behavior in 2005, or remote collaboration in 2018 would have found many mainstream users reasonably satisfied with familiar methods. Satisfaction with the present does not show how a need will change when technology, regulation, cost, or behavior changes.
Lead-user research offers one route: find people already living with the future constraint. Proactive market orientation offers another: study latent value rather than only expressed requests. Neither method gives permission to forecast whatever supports the pitch. Future pull needs a traceable mechanism.
Ask:
- Which external condition is changing?
- Who experiences that condition already?
- What new cost or opportunity does it create?
- Which behavior shows that the need is real now?
- What would prevent the condition from spreading?
The difference between “AI will change everything” and “support teams with this ticket mix already spend this much verifying generated answers” is the difference between a trend and a warrant.
Let solution work interrogate the problem
The phrase “fall in love with the problem, not the solution” contains a useful warning and a subtle trap.
The warning is clear: do not protect an artifact after evidence changes. The trap is the suggestion that a correct problem can be fully known in advance and then handed to a separate solution phase.
Eric von Hippel’s sticky-information theory explains why that handoff often fails. Customer context and technical knowledge can be costly to acquire, transfer, and use in a new location. When several sites hold necessary information, problem solving may move between them or be partitioned into subproblems.
Kees Dorst and Nigel Cross observed nine experienced industrial designers during conceptual design sessions. Their protocol analysis supports a co-evolution model: designers developed the problem space and solution space together. A proposed solution exposed a new constraint; the revised problem enabled a different solution.
This is not an excuse to build first and ask questions later. It is a reason to use small solution probes as research instruments.
A paper workflow, concierge service, clickable prototype, technical spike, or policy simulation can reveal whether:
- the assumed cause is real;
- the customer can recognize the proposed outcome;
- necessary data exists;
- the organization can change the relevant system; and
- solving one failure creates another.
Tian Heong Chan and Shi-Ying Lim’s study of IKEA hacks gives this starting-point question a useful data portrait. They scraped 3,356 posts from 2006 to 2019 and examined whether hackers began with a problem or with a product. Problem-first hacks had a 55% likelihood of becoming an exaptation—a novel use—versus 42% for product-first hacks.
Starting from a problem was associated with more novel uses
Across the coded IKEA-hack archive, problem-first searches had a higher reported probability of exaptation than product-first searches.The archive was observational. The difference is an association, not a causal estimate.
Chan and Lim, 2023. The comparison supports attention to search starting points; it does not make problem-first work universally superior or establish causation.
The study is observational, so it cannot prove that problem-first framing caused the 13-point difference. It also finds nuance: experience and product modularity helped product-first search catch up. Solutions can reveal valuable problems. The stronger rule is to know which direction the search is moving and what fixation it may create.
Invite customers into the right rooms
“Co-create with customers” sounds generous and modern. It is not a complete operating instruction.
Woojung Chang and Steven Taylor’s meta-analysis synthesized 123 correlations from 39 samples. Customer participation during ideation and launch improved new-product financial performance directly and indirectly through faster time to market. Participation during development slowed time to market and harmed financial performance. Effects also varied with technical turbulence, country, industry, customer type, and firm size.
The finding does not say that engineers should disappear into a locked room. It says that customer participation has different jobs at different phases.
Customers are strong authorities on their situation, constraints, workarounds, and perceived outcomes. They can generate novel ideas. Marion Poetz and Martin Schreier had company executives evaluate user and professional ideas blindly. User ideas scored higher on novelty and customer benefit and somewhat lower on feasibility, although feasibility remained high overall.
That division is productive. Customers can expand what the team notices. Product, design, engineering, operations, and legal specialists remain accountable for deciding what can be built, supported, secured, and sustained.
Confusion begins when one participant’s evidence is treated as authority over every question.
A crowd is not a verdict
Idea portals, support queues, interview repositories, and sales-request boards create valuable memory. They also create a visual illusion: the biggest pile must represent the biggest opportunity.
Barry Bayus studied 8,801 ideas from 4,285 contributors during Dell IdeaStorm’s first two years. Serial ideators were more likely than one-time contributors to have an idea implemented. Once an idea succeeded, however, the same contributor was unlikely to repeat that success, and later ideas became less diverse.
Implementation by Dell is not a clean measure of customer value. One community cannot stand for every market. The result still breaks a comfortable assumption: a productive source does not remain productive simply because it has a track record.
Selection needs diversity and discernment. Karan Girotra, Christian Terwiesch, and Karl Ulrich define four determinants of the best idea a group identifies: average idea quality, the number of ideas, variance in idea quality, and the group’s ability to recognize quality. In their experiment, individuals who worked independently before collaborating generated more and better ideas and selected them more effectively than groups that worked together throughout.
Applied to problem discovery, this suggests a useful rhythm:
- collect candidate problems from several independent channels;
- preserve minority and contradictory observations;
- write a warrant for each candidate before group ranking;
- compare missing evidence, not only confident opinions; and
- fund the next test that reduces the most consequential uncertainty.
The process is slower than moving the loudest request into design. It is much faster than discovering after launch that the team solved a sentence instead of a problem.
Write the warrant before the roadmap item
A problem warrant can fit on one page. It should contain enough detail to be challenged.
Customer and situation. Name the smallest coherent group and the moment when the outcome fails.
Consequence. State the lost time, money, quality, trust, access, safety, or strategic option. Separate inconvenience from material harm.
Behavioral evidence. Record what people do now: workaround, switch, delay, hire, build, abandon, escalate, or accept.
Future mechanism. Name the condition that should preserve or spread the need, plus the evidence that could show the condition is not developing.
Causal and capability hypothesis. Explain why a change within the company’s reach could improve the outcome. Identify sticky knowledge, technical dependencies, and operating limits.
Disconfirming test. Define the smallest credible observation that would make the team stop, narrow the segment, revise the cause, or abandon the opportunity.
Then make a bounded decision. Do not ask, “Should we build the product?” Ask, “Has this problem earned the next uncertainty-reducing investment?”
That investment may be ten interviews with a tighter segment, an event-log analysis, a prototype, a technical spike, a pricing test, or a deliberate no. Every step should make the warrant stronger, narrower, or unnecessary.
The dashboard request from Tuesday morning may survive this process. Perhaps executives cannot act because the product hides a material exception pattern, and a shared view is the lowest-cost remedy. Or perhaps the dashboard is a visible treatment for the same reconciliation failure customers already solve in spreadsheets.
The research does not tell the team which feature wins. It tells the team what a serious answer must contain.
The best customer problem is not the one with the most votes, the most eloquent buyer, or the cleanest feature attached. It is the one whose consequence, behavior, segment, future, solvability, and possible failure form a coherent reason to learn more.
References
- Bayus, B. L. (2013). Crowdsourcing New Product Ideas over Time: An Analysis of the Dell IdeaStorm Community. Management Science, 59(1), 226-244.
- Chan, T. H., & Lim, S.-Y. (2023). The Emergence of Novel Product Uses: An Investigation of Exaptations in IKEA Hacks. Management Science, 69(5), 2870-2892.
- Chang, W., & Taylor, S. A. (2016). The Effectiveness of Customer Participation in New Product Development: A Meta-Analysis. Journal of Marketing, 80(1), 47-64.
- Dorst, K., & Cross, N. (2001). Creativity in the Design Process: Co-evolution of Problem-Solution. Design Studies, 22(5), 425-437.
- Girotra, K., Terwiesch, C., & Ulrich, K. T. (2010). Idea Generation and the Quality of the Best Idea. Management Science, 56(4), 591-605.
- Griffin, A., & Hauser, J. R. (1993). The Voice of the Customer. Marketing Science, 12(1), 1-27.
- Lilien, G. L., Morrison, P. D., Searls, K., Sonnack, M., & von Hippel, E. (2002). Performance Assessment of the Lead User Idea-Generation Process for New Product Development. Management Science, 48(8), 1042-1059.
- Narver, J. C., Slater, S. F., & MacLachlan, D. L. (2004). Responsive and Proactive Market Orientation and New-Product Success. Journal of Product Innovation Management, 21(5), 334-347.
- Poetz, M. K., & Schreier, M. (2012). The Value of Crowdsourcing: Can Users Really Compete with Professionals in Generating New Product Ideas?. Journal of Product Innovation Management, 29(2), 245-256.
- Urban, G. L., & von Hippel, E. (1988). Lead User Analyses for the Development of New Industrial Products. Management Science, 34(5), 569-582.
- von Hippel, E. (1986). Lead Users: A Source of Novel Product Concepts. Management Science, 32(7), 791-805.
- von Hippel, E. (1994). “Sticky Information” and the Locus of Problem Solving: Implications for Innovation. Management Science, 40(4), 429-439.
Summary
Write a problem warrant, not a popularity score: name the customer and consequence, verify costly behavior, test segment coherence and future pull, show why your team can change the outcome, and state what evidence would make you stop.
- Translate the requested feature into the customer outcome, consequence, and situation beneath it.
- Look for costly behavior: repeated workarounds, switching, delay, self-built tools, or accepted risk.
- Define the smallest coherent segment that shares the cause and consequence, not only a demographic label.
- Check future pull with lead users, changing constraints, and latent needs instead of demand forecasts alone.
- Test whether your capabilities can change the outcome without depending on hidden customer or technical knowledge.
- Write the fastest credible disconfirming test and a stop condition before selecting a solution.
- Compare complete warrants, preserve minority evidence, and fund only the next uncertainty-reducing step.