B2B SAAS CUSTOMER SEGMENTATION CASE STUDY

How Amplitude Used Behavioral Segmentation
to Build a Benchmark That Generated 1,120 Qualified Leads in Three Weeks

Amplitude had the product data. The challenge was making its benchmarks specific enough to answer a product leader’s real question: How does our product compare with similar products?

Amplitude was building a public benchmarking experience for product managers and product leaders. A single blended average would have combined unlike companies and product motions, making the result descriptive but not decision-ready.

HPI worked backward from the questions the audience would ask, cleaned and organized the product data, and created peer groups using company stage, product motion, and engagement depth. Aggregation and
k-anonymity thresholds were applied to protect customer privacy.

The final benchmark covered activation, retention, depth of use, and time to value. HPI also delivered the segmented dataset, reusable query patterns, and UX and messaging guidance needed for launch. In the first three weeks, the experience generated 1,120 qualified leads.

Background

Amplitude builds product analytics software that helps teams understand how people use digital products. The company planned a public benchmarking experience as a flagship lead-generation asset for two audiences: product managers and product leaders.

The concept sounded simple: show teams how activation, retention, depth of use, and time to value compared with the market. But a benchmark is only useful if the comparison is relevant. A blended average can be mathematically correct and still mislead when it combines products with different stages, motions, and engagement patterns.

That made the core problem one of segmentation, not visualization. Amplitude needed to decide which products belonged in the same peer group, which measures would answer useful product questions, and how to make those comparisons public without exposing customer data.

The Challenge

Turn a large dataset into a comparison a product leader could trust

One average hid the context

The default cuts combined unlike companies and product motions. They could describe the dataset, but they could not reliably answer the question a visitor cared about: what does good performance look like for a product like ours? Product leaders needed comparisons tied to activation, retention, depth of use, and time to value within a relevant peer group.

More specificity created a privacy constraint

The peer groups had to be narrow enough to feel relevant but large enough to protect individual customers. That meant privacy rules had to be part of the segmentation logic rather than a final publishing check.

Marketing needed something launch-ready

The team needed a segmented dataset, benchmark logic, and a clear launch narrative without replacing the data platform or asking engineering to rebuild the experience.

The Approach

Start with the decision, then build the segmentation around it

HPI treated the benchmark as a decision system rather than a reporting exercise. Instead of starting with every field in the dataset, the work started with what product managers and product leaders needed to understand and worked backward to the peer groups, measures, and privacy rules required to answer those questions.

  1. Define the audience and decision questions.
    HPI centered the work on two priority audiences – product managers and product leaders – and the questions they would bring to the benchmark: What does good activation look like for products like mine? Where do power users appear? How quickly should users reach value at our stage?
  2. Prepare the product data for comparison.
    Using SQL and Python, HPI cleaned and organized the source tables for product usage segmentation and customer cohort analysis so the same measures could be compared consistently across groups.
  3. Build peer-group logic around meaningful differences.
    Instead of one broad average, HPI created comparison groups using factors such as company stage, product motion, and engagement depth – the dimensions in the available data that made the benchmark more relevant to the visitor.
  4. Set the privacy rules before publishing the cuts.
    Strict aggregation rules and k-anonymity thresholds were applied so the benchmark could remain specific without creating a path to identify an individual customer.
  5. Turn the segments into four benchmark dimensions.
    The final views covered activation rate by peer group, retention curve shapes, depth-of-use quartiles, and time-to-value distributions. Each view was tied to a product decision rather than presented as an isolated chart.
  6. Package the analysis for launch.
    HPI delivered the segmented dataset, reusable query patterns, interface guidance, messaging direction, and calls to action needed to fit the existing launch plan. No engineering rebuild was required.

What Changed

The central decision was to stop treating the entire dataset as one population. Product leaders did not need the overall average; they needed a reference point that reflected the context of a product more like their own.

  • Peer groups organized the comparison around company stage, product motion, and engagement depth.
  • Activation, retention, depth of use, and time to value were calculated and presented within those peer groups rather than as one blended benchmark.
  • Aggregation and k-anonymity thresholds set a clear boundary between useful specificity and customer confidentiality.
  • The same segmentation logic could be reused across the benchmark rather than rebuilt separately for every chart.

The result was a fundamental shift in the question the benchmark could answer: from “Are we above or below the overall average?” to “How are we performing relative to products with a relevant context?”

The benchmark became a decision aid rather than a collection of descriptive charts.

The Result

Amplitude launched the benchmarking experience with audience-ready segments and no engineering rebuild. In the first three weeks, the benchmark generated 1,120 qualified leads across its two priority audiences: product managers and product leaders.

 

Measure Result Why it mattered
Qualified demand 1,120 qualified leads in three weeks Turned a data asset into a pipeline-generating experience.
Audience focus 2 priority audiences Kept the benchmark and launch narrative centered on product managers and product leaders.
Engineering impact 0 rebuilds required Allowed the team to launch using a segmented dataset, query patterns, and implementation guidance.
Benchmark design 4 decision-ready measure groups Covered activation, retention, depth of use, and time to value.
Privacy design Aggregation rules plus k-anonymity thresholds Protected customer confidentiality while preserving useful peer comparisons.

The benchmark did not ask visitors to interpret one generic average. It helped them compare their product with a relevant peer group and answer questions tied to real product decisions. That made the data more useful to the audience and gave Marketing a stronger reason for visitors to engage.

The same principle applies to a broader Customer Segmentation Study: segments become valuable when they change the decision, the experience, or the action that follows.

Frequently Asked Questions

Turn product data into decisions your audience can use.

HPI helps data-rich companies find the customer and product patterns that should shape benchmarks, product strategy, messaging, retention, and growth.
Start a Conversation

Customer evidence. Market clarity. Revenue decisions you can defend.