Prerequisites
Read these first if A/B tests, exposure, or feature flags vs experiments are new.
- Statsig — Layers — mutual exclusion domains
- Embracing overlapping A/B tests — why isolation can hurt power
- Holdouts — measuring cumulative impact with a held-back floor
- Statsig — Feature flags — kill switch vs rollout vs experiment
This A/B ate that A/B
Two “winning” experiments ship the same week: a denser message list and a louder compose CTA. Metrics look great in isolation. Together, the inbox feels noisy and send drops. Neither dashboard was wrong; both were reading a population that quietly overlapped. Traffic topology is the thing nobody put in the readout.
This is not a rewrite of kill switch / rollout / experiment jobs. It is the failure mode when those jobs share a screen without a contract.
The wrong fight: “overlap is modern, layers are legacy”
Platforms that argue for overlapping assignment (e.g. Statsig’s public writing) are right about velocity and statistical power: isolating everything destroys both. The same docs still ship layers (mutual exclusion domains) and interaction detection — because some treatments are not independent parameters.
flowchart TB
subgraph overlap ["Overlapping — combinations exist"]
direction TB
uO[User U]
uO --> aO["Exp A: denser list"]
uO --> bO["Exp B: louder CTA"]
aO --> combo["U's inbox = denser + louder"]
bO --> combo
end
subgraph layer ["Layer — mutual exclusion"]
direction TB
pop[Users in one layer]
pop --> xor{"Hash once — A XOR B"}
xor -->|"slot A"| aL["Exp A only<br/>never enters Exp B"]
xor -->|"slot B"| bL["Exp B only<br/>never enters Exp A"]
end
Figure 1. Overlap: the same user can sit in A and B at once (combination cells). Layer: traffic is partitioned so a user is in at most one experiment in that domain.
The interesting decision is not “which religion?” It is: can these two treatments rewrite the same pixel or the same write path at once?
| Symptom | Likely topology mistake |
|---|---|
| Two “wins,” one ugly product | Overlap on colliding surfaces (list chrome + CTA + ranking) |
| Every test underpowered; teams fighting for users | Everything forced into layers |
| Cumulative product drift nobody can measure | No holdout floor — everyone is always in something |
| Ghost variants on old APKs | Overlapping remote config on clients that only understand half the matrix |
Where layers still earn rent on mobile
Colliding surfaces. List density, ranking, and compose CTAs are not independent knobs. Exclusive layers (or hard gates) keep ownership honest when two teams ship into the same screen.
Binary skew. Old APKs linger. Overlapping configs on clients that only know half the variants produce “impossible” combinations in telemetry. Layers plus version targeting cut that class of ghost.
Holdout floors. A small global holdout (often a few percent in org folklore) never enters a set of experiments so you can measure cumulative drift vs “everyone in everything.” That answers a different question than any single A/B — and it is easy to forget when overlap is the default pitch.
Overlapping remains the right default when interactions are rare, metrics are robust, and you can detect collisions. Layers are the circuit breaker when the product is one shared UI.
Topology does not fix missing exposure
Neither overlapping nor layers save the science if you never fire “user saw treatment T.” Sticky assignment, offline defaults, and no config fetch on the critical path still rule — then choose topology. Without exposure, you will ship another pair of “wins” that only look good because the dashboards never saw the combination users actually got.
For example, in a mail-shaped client, the dense-list + loud-CTA week is not a stats pedantry problem. It is a product incident caused by assignment topology. Fix the exclusion (or gate the ship) before you A/B a third chrome change on top.
References
- Interaction detection
- Tang et al. — Overlapping Experiment Infrastructure (KDD 2010) — layers as mutually exclusive domains so more tests share traffic
- Kohavi, Tang, Xu — Trustworthy Online Controlled Experiments — interactions, SRM, and why isolation is not free