Pattern

Segment-Level Analysis (Simpson's Paradox)

Always check experiment results by segment, since an aggregate A/B-test winner can be a severe loser in the segment that matters most for revenue.

An aggregate winner that was quietly losing

Inside the growth organization Eric Glyman built at Ramp, the most tactical guardrail on experimentation is to check results by segment before propagating them. George Bonaci, Ramp's VP of Growth, illustrates it with a homepage test in which a red button beat every alternative "by far" on a pure aggregate basis.1 The result was counterintuitive, since red usually carries a "something is wrong, do not press it" connotation, but nothing outperformed it in the pooled data. The catch surfaced only under segmentation: in the enterprise segment, the same red button severely underperformed. The pattern is Simpson's Paradox, where a trend that holds in aggregate reverses inside subgroups.

The damage was downstream of the measurement. The web team codified "red is best practice" and it propagated to other teams, even though "most of Ramp's content, webinars, and direct mailers were aimed at enterprise." A tactic that won on the aggregate was quietly degrading the performance of everything built for the most important segment. Catching the paradox, then removing red from the enterprise-facing work, is described as "incredibly valuable."1

Generalizable learnings are powerful and dangerous

Bonaci embeds the story in a broader claim about post-mortems, part of the same experimentation practice as Growth as Experimentation. The highest form of output from an experiment, in this account, is not a single result but a learning that generalizes across the business. The DRI, the directly responsible individual who scoped the experiment, writes it up and asks not only whether it was run well but whether it taught something that applies to other parts of the company, with the cross-functional stakeholders present who can both learn from it and prevent the failure from recurring.

The red button is the cautionary twin of that ambition. A generalized learning is powerful precisely because it spreads, which is what makes it dangerous: propagating an un-segmented "best practice" spreads the error faster than a one-off mistake ever could. The instruction that follows is to segment a result before it becomes a company-wide rule, especially along the dimension that maps to revenue concentration, which at Ramp was enterprise.

Why the guardrail matters

The framing treats aggregate metrics as noisy proxies for what is true inside the segments a business actually cares about, connecting the pattern to Signal vs. Noise: a pooled win can be misleading signal. It also sharpens the prioritization discipline in Prioritize on Confidence and Time-to-Results, where "confidence" is meant to be confidence in the relevant segment rather than in the pooled average. Read together, the point is that the decision to trust a result should be indexed to the population that carries the revenue, not to the arithmetic mean across all of them.

The tradeoffs the pattern admits

The framework names its own limits. Segmenting too finely re-introduces the very problem it solves, because small per-segment samples become noisy again; there is a real tradeoff between protecting against Simpson's Paradox and preserving statistical power. And the story assumes you already know which segment matters before you slice the data. If revenue concentration shifts, the segmentation that was correct yesterday can mislead tomorrow, so the guardrail depends on a judgment about what matters that the analysis itself does not supply.

Practiced by

Connections

Loading connections…

References

  1. 01

    George Bonaci, VP of Growth at Ramp (20VC)

    George Bonaci, interviewed by Harry Stebbings · podcast

Related