Pricing Research Methods for Product Teams
Match your pricing research method to your decision stage, not your favorite toolkit.

Pricing decisions used to sit downstream of product and marketing. Not anymore. Pricing now touches finance, go-to-market, and the board, because the cost of getting it wrong has gotten steeper and the tools for getting it right have multiplied. Metronome's field report on SaaS teams found that AI product costs have turned pricing into a C-level call, with CFOs and revenue leads making real-time decisions about cost exposure rather than setting margin targets once a year and moving on.
The structural problem hasn't changed much, even as the tools have. Most teams still pick a pricing model before they have real evidence of what customers value or what they'll actually pay for it. A large share of new consumer products miss their launch targets, and a substantial portion of startup failures get traced back to no market need. Both numbers point at the same gap: assumptions about price get baked in early, before anyone tests them. The good news is that the research toolkit has expanded a lot in the past few years, from survey instruments that have been around for decades to AI-moderated interviews and synthetic consumer panels. The hard part now is knowing which tool fits which moment.
The core decision a team must make before choosing any method: what stage are you in?
"What should we charge at launch?" is not the same question as "why is our top tier underperforming?" and neither of those is the same as "how do we price a brand-new AI feature nobody's used before?" Yet teams often run the same research playbook for all three, and that's where things go sideways.
Three questions should come before any method gets picked.
First, what stage is the decision in: early concept, pre-launch calibration, or something already live that needs tuning? Second, how much speed does the team actually have: days or weeks? That gap used to be a scheduling detail. Now it's a real strategic constraint, because competitors move fast and a six-week research cycle can mean shipping a price nobody asked for. Third, what's the confidence threshold: is this a reversible experiment, or a contract structure that, once signed, is nearly impossible to unwind?
Teams tend to default to whatever method they already know how to run. That means commissioning a full conjoint study when a quick willingness-to-pay screen would have answered the question in two days, or leaning on a survey when behavioral testing would give a cleaner signal because real users behave differently than survey respondents.
For AI-native products, there's a fourth axis to add: cost structure. For AI-native products, pricing isn't just a question of what customers will pay, it's a question of what the unit economics can absorb. That distinction runs through the rest of this piece.
Van Westendorp Price Sensitivity Meter: fast willingness-to-pay framing in early concept stages
The Van Westendorp method asks four questions. At what price does this feel too cheap to trust? At what price is it a bargain? At what price does it start to feel expensive but still worth it? And at what price does it become too expensive to consider? Stack the answers across a sample and an acceptable price range falls out, not a single magic number, but a band.
It's a good fit early, before a team has committed to a business model, because it's fast to field and cheap per respondent. Deployed through modern survey platforms, a Price Sensitivity Meter can move faster than most quantitative studies, making it a better fit for near-term decisions than methods that require longer fieldwork windows.
What it won't do is tell a team how price interacts with features or competing offers, since it asks about price in isolation. That's a job for conjoint analysis, covered next.
A common mistake: treating the midpoint of the acceptable range as the launch price. The range shows where resistance starts to build on either end, not where the optimal price sits. Pairing the survey with a handful of follow-up interviews, even five or six, tends to reveal the reasoning behind the numbers in a way the scores alone can't.
Conjoint analysis: measuring trade-offs when features and price compete for attention
Conjoint analysis shows respondents different product configurations, feature sets, tiers, prices bundled together in varying combinations, and asks them to choose. From those choices, the method backs out part-worth utilities: how much value a customer assigns to each attribute relative to the others.
That output functions almost like a pricing simulator. A team can see what happens if it adds a feature, drops a price point, or restructures a bundle, all before building anything. That makes it a strong fit for pre-launch tier design, feature bundling, and competitive positioning, especially in situations where the real question isn't what customers will pay but what they'll give up to get something else.
Designing a conjoint study well takes real expertise. Picking the wrong attributes or under-sizing the sample produces utilities that look precise and mean nothing. For AI-native products, conjoint can help settle which value metric customers actually respond to, tokens, seats, or outcomes, before a team locks into a billing architecture that's painful to reverse later. With traditional panels, turnaround runs in weeks. Modern panel platforms can compress that significantly, which matters for teams working sprint to sprint rather than quarter to quarter.
Gabor-Granger and direct willingness-to-pay methods: calibrating a price point against a known demand curve
Gabor-Granger works sequentially. Respondents get asked whether they'd buy at a given price, then the price moves up or down and the question repeats. Chain enough of those answers together and a demand curve emerges, showing the price that maximizes revenue.
The method is narrower than Van Westendorp or conjoint by design. It doesn't map a range and it doesn't model trade-offs between features. It answers one question: at this price, how many people say yes? That makes it useful for validating a shortlisted price before launch, or testing a proposed price change on a product that's already live.
The limitation is baked into the design. Gabor-Granger assumes the respondent already understands the product and that the purchase decision comes down to price alone, with no competitive context and no feature trade-offs in the mix. In practice, it works best as a confirmation step after Van Westendorp or conjoint has already narrowed the field.
Qualitative pricing interviews: understanding the reasoning that survey instruments cannot capture
Surveys tell a team what customers say they'll pay. Interviews tell a team why, revealing the reference points, the budget context, and the comparisons a customer is making in their head against a competitor or a prior tool.
Qualitative work has historically been rationed. A typical research project might squeeze in a handful of interviews to add color to a quantitative dataset, not nearly enough to spot real patterns across different customer segments.
AI-moderated interviews have started to change that math. Running at a fraction of the cost of human-moderated sessions (a figure that hasn't been independently verified but occurs consistently in vendor claims), they let a team run pricing interviews across several segments and geographies inside a single sprint instead of sampling just one group and hoping it generalizes.
What they tend to do well in pricing work specifically: probe where a respondent's price anchor actually came from, dig into whether a purchase is an individual decision or something that has to clear procurement, and surface objections to specific pricing models like credits or usage-based billing that a fixed-scale survey question would never catch. They can also follow an unexpected answer down a rabbit hole instead of marching to the next scripted question. A lot of the real insight tends to live there.
For findings to actually get used, they need to trace back to evidence, meaning quotes and recordings linked directly to conclusions, so product and finance stakeholders can act on them with confidence instead of taking a summary on faith. For AI products specifically, Metronome's field report identified anxiety over unpredictable costs as the single biggest barrier to adoption, and qualitative interviews are the method best suited to figuring out what "unpredictable" actually means to a given buyer.
Behavioral price testing: what customers do when price changes
Every survey-based method shares the same weakness: it measures what people say, not what they do. Customers often say they'd pay X and then behave completely differently at the actual moment of purchase.
A/B price testing closes that gap. Real users see different prices, and the team measures conversion, upgrades, and churn. For a live product, this is the most reliable signal available, because it skips the risk of users stating preferences that don't match what they'd actually pay for.
It's not a fit for everything, though. Early-stage concepts have no product and no user base to test against. Some markets restrict price testing legally. And in some cases, the wrong price triggers a competitive response or customer backlash before the test even finishes running. There's also a speed cost: meaningful A/B tests on pricing often need weeks of traffic to reach a clear result, which makes them too slow for a lot of sprint-level decisions.
The strongest setup usually combines methods. Run a Price Sensitivity Meter to shortlist two or three candidate prices, A/B test the top two against live traffic, then run qualitative interviews to understand why the winner won.
Synthetic consumer panels: running pricing scenarios before a product or price exists
Synthetic panels are AI-generated respondents built from real-world data, historical surveys, behavioral logs, reviews, public opinion trends.
Construction usually involves established personality frameworks and affective modeling, a retrieval layer that pulls in domain-specific knowledge, and models fine-tuned against large behavioral datasets. A study from Stanford and Google DeepMind, run with 1,052 participants, found that AI digital twins replicated human survey answers with 85% accuracy, matched personality assessments at 80% correlation, and matched economic game behavior at 66% correlation. Properly calibrated synthetic panels reach 85 to 95% parity with real panels on concept, pricing, and positioning tests. Generic prompts run through an off-the-shelf model, without that calibration step, land closer to 55%, which is a big gap and worth taking seriously before trusting any synthetic output.
The speed difference is dramatic. According to the GreenBook GRIT report, concept-to-signal cycles that take four to eight weeks with traditional panels can take hours with a calibrated synthetic audience.
That speed opens up research that wasn't really possible before. A team can test willingness to pay for a product that doesn't exist yet, since there's no live user base to recruit from. It can simulate how different buyer personas respond to credit-based versus usage-based versus outcome-based pricing before a billing system gets built. It can run scenario analysis on a price increase or a tier restructure without exposing real customers to the change. And it can simulate global segments without the lead time international recruitment usually requires.
Adoption is already underway. Qualtrics' 2025 Market Research Trends Report found 73% of market researchers have used synthetic responses at least once, and the 2025 GRIT findings put satisfaction among active users at 87%.
None of this replaces live behavioral testing or human usability sessions, especially when a question depends on real hesitation or emotional reaction that a model trained on past data can't fabricate from scratch. And the 85 to 95% accuracy ceiling only holds when the panel is properly calibrated against real human benchmarks. Skip that step, and the results drift back toward the 55% floor.
Pricing research for AI-native products: why the standard methods need adaptation
AI products carry a cost of goods sold that moves with usage, inference, tokens, GPU compute. Pricing research has to answer two questions at once: what will customers pay, and what can the business actually afford to charge.
GitHub Copilot launched at $10 a month and reportedly cost Microsoft more than $80 a month in compute for its heaviest users, a gap that no willingness-to-pay study alone would have caught, because the study only answers half the question. Foundation model costs have also been falling fast, dropping 50 to 90% a year in some cases. A capability that cost six cents per call in early 2023 might run two-tenths of a cent by late 2024. Pricing research for these products has to account for a cost floor that keeps moving, not one that sits still.
That changes what research needs to answer. Which value metric, tokens, seats, documents, outcomes, feels fair and predictable to a given segment is a job for qualitative interviews. How anxiety about unpredictable costs affects willingness to adopt usage-based pricing calls for qualitative work paired with a Van Westendorp study framed around spend caps rather than flat prices. And what committed-use structures enterprise buyers will actually accept is a conjoint question, run across different contract configurations.
Industry reporting has found that most teams defaulted to cost-plus credit models with a 30 to 50% standard markup, largely because cost data was on hand and value data wasn't, reflecting a research gap rather than a strategy failure. That's a research gap. Synthetic panels can help close it by simulating how an individual developer versus an enterprise procurement team responds to usage-based versus hybrid pricing before a billing system gets locked in. Even so, adoption of AI in pricing work remains early. A poll at ISPOR Europe 2025, with 246 participants, found 64% of organizations aren't using AI in pricing work at all, and only 15% have adopted the tools, even in an industry where pricing complexity runs about as high as it gets.
Matching method to moment: a working framework for product teams
Break it down by stage, and the method choice gets a lot more obvious.
At the early concept stage, a team has a hypothesis but no product and no users yet: Van Westendorp brackets the price space fast, and synthetic panels simulate how different segments respond to different pricing models without needing a single live user. The goal here isn't precision, it's ruling out the price zones and pricing architectures that are clearly wrong before anyone commits engineering time to building them.
At the pre-launch calibration stage, the product exists and the team is choosing between a shortlist of pricing models: conjoint analysis models the feature-price trade-offs, and AI-moderated interviews dig into the reasoning behind those preferences and surface objections to specific billing structures before they become support tickets. The goal is picking the pricing model with the best odds of adoption among the segments that actually matter.
In-market optimization, when the product is live and pricing is already set, calls for a different toolkit. Behavioral A/B testing measures what customers do, not what they say, and Gabor-Granger validates specific price-change hypotheses against a known base of paying users. Qualitative interviews, run on customers who've already churned or downgraded, tend to explain the number in a way the dashboard never will.
None of these methods replaces sound judgment about the business itself. But running the wrong one, or running one when the moment calls for none at all, is how teams end up with a pricing page built on a guess dressed up as data.


