Analytics

Sample Ratio Mismatch: Build an Assignment-to-Analysis Evidence Ledger

Before interpreting conversion lift, reconcile who was assigned, exposed and retained in analysis. A technical two-arm SRM guide with a tested stage-ledger implementation, numerical example and investigation contract.

A checkout experiment reports a better conversion rate, but one variant has fewer users than the configured split implies. Is the treatment winning, or are the people who disappear from the data changing the answer? Before defending the uplift, make the population reproducible. Sample ratio mismatch is a useful diagnostic, but the more valuable deliverable is the evidence ledger that explains where the population changed.

Editorial disclosure: Original, AI-assisted Growthcraft Editorial synthesis. All examples are synthetic, not Akshay's client results. The JavaScript example is tested in Node.js 24. The framework is an operational synthesis, not a new statistical method. This article does not certify an experiment, estimate causal lift or replace a qualified statistical review.

Four takeaways before you open the results

  • Check unique randomized units against the allocation configured before the observations—not against an idealized 50/50 split.
  • Keep assignment, exposure and analysis eligibility as separate populations with explicit evidence references.
  • Use a single-look chi-square diagnostic only within its assumptions; no flag does not mean the experiment is valid.
  • Investigate and document a cause before editing eligibility, restarting or interpreting a conversion winner.

Why faster experimentation still needs a population contract

Statsig's 19 August 2026 article on running faster tests discusses demand for shorter experiments and the risks of changing metrics or removing observations. That is a timely qualitative signal from a platform publisher, not a measured keyword trend or evidence that every company has this problem. This guide addresses one narrow operational consequence: how to preserve the population behind a test when growth teams are under pressure to reach a decision.

The task applies to consumer checkout, subscription onboarding and B2B product activation. The important difference is the randomized unit. A B2B product may allocate whole accounts while a consumer application allocates users. Counting every account member as an independent assignment inflates the apparent sample without changing the randomization. A website that repeatedly assigns by session has another contract again. Write the actual design down rather than inheriting the default denominator in an analytics dashboard.

This is not another article about calculating conversion significance. Outcome analysis asks whether a measured response differs under a valid design. Allocation diagnosis asks whether the observed group sizes are compatible with the configured assignment proportions and whether the population survived processing as intended. The second question is a prerequisite for understanding the first; it does not answer it.

Define the null allocation precisely

For two variants, let A and B be counts of unique independent randomized units at one named stage. Let N = A + B, and let q be the pre-specified probability of assignment to control. The expected counts are EA = Nq and EB = N(1 − q). Pearson's goodness-of-fit statistic is X² = (A − EA)²/EA + (B − EB)²/EB. With two fixed categories and no parameters fitted to these observations, the reference distribution has one degree of freedom.

The companion sample ratio mismatch calculator reports the upper-tail probability under that null. It is not the probability that a bug exists, the chance that treatment wins, or the percentage of missing users. Current Statsig documentation describes SRM checks using observed unique-user allocation and the configured split; our tool is independent and does not reproduce a vendor's complete health-check system.

The calculator requires each expected count to be at least five, using the conventional approximation caution described by NIST. This is a practical minimum, not a guarantee of accurate inference for every design. Below it, the interface withholds the approximate p-value. An appropriate exact-binomial procedure may suit a two-arm independent design, but its tail convention and decision policy must be specified. The tool does not silently swap methods.

A configured 60/40 split should not be tested against 50/50. A ramp from 10/90 to 50/50 should not be tested as though the final setting applied throughout. Separate fixed-allocation periods and review the design before combining evidence. Adaptive assignment, clustered observations and switchback tests need methods that respect those structures. Two numeric boxes cannot erase those dependencies.

Work a diagnostic without inventing a cause

Consider 4,800 control and 5,200 treatment users under a fixed 50/50 configuration. The expected counts are 5,000 each. Each squared deviation contributes 200²/5,000 = 8, so X² = 16. The one-degree-of-freedom p-value is approximately 0.0000633425. With a pre-specified single-look threshold of 0.001, this snapshot is flagged for investigation. The observed control share is 48%, two percentage points below configuration.

The threshold here is an illustrative review policy, not a universal industry standard. Choose and record it before seeing the result. A stricter threshold trades sensitivity against false alarms; it cannot repair bad identity mapping. The calculator applies strict p < threshold using the unrounded value. Its display uses scientific notation so a small probability is not casually rounded to 0.00.

Now imagine a second experiment with 6,000 control and 4,000 treatment users under a configured 60/40 split. Its statistic is zero and its p-value is one. Both groups could still lose the same proportion of exposure records, or both could have broken revenue events. A balanced ratio is compatible with some serious data failures. The correct label is no flag at this snapshot, not experiment passed.

For a third example, eight control and two treatment users under 90/10 allocation have expected counts nine and one. This checker refuses an approximate tail. That refusal is useful: a founder with a tiny sample should not mistake a precise-looking probability for a reliable decision. Separately, even a suitable allocation check would not establish enough outcome power to measure a modest conversion effect.

Build three ledgers, not one mutable denominator

The assignment ledger records the variant selected for each randomized unit and the immutable allocation version. The exposure ledger records delivery under a defined condition. The analysis ledger records the pre-specified eligibility rule applied to the data. These stages can legitimately differ, but each difference needs a definition and a reproducible query. Keep experiment ID, allocation version, unit type, event time, ingestion time and extraction cutoff in the data contract.

Use a stable internal unit key in your controlled warehouse, not a raw email address in a shared worksheet. Aggregate before exporting to the local tool or an LLM. Distinguish duplicate event delivery from conflicting assignment: five copies of one user's control event are not five users, while control and treatment records for the same unit require investigation. Choosing whichever variant appeared first may hide a crossover problem; an explicit policy and evidence are needed.

Start from an assignment spine and left-join downstream evidence. An inner join drops units without outcomes and can silently redefine the population. Capture row counts and unique-unit counts before and after every join. A many-to-many join may multiply events while leaving a plausible unique-user total, so reconcile both grains. Preserve unknown variants and downstream-only units in a separate rejection ledger rather than discarding them without a trace.

The reference code below is intentionally strict. It accepts already-scoped, anonymous synthetic rows and rejects conflicts or downstream rows absent from assignment. It does not implement your identity system, time-window filtering or warehouse connector. Those remain prerequisites. Its nested-stage contract requires analysis units to have exposure evidence; if your intended analysis is intention-to-treat on all assigned units, use a different declared relationship instead of copying that assumption unchanged.

Run a tested stage-ledger reference

The function de-duplicates repeated same-variant records, reports duplicate-row counts and returns per-stage unique-unit totals. It rejects variant conflicts within a stage, variant changes across stages, unknown variant labels and missing unit keys. The caller must select one fixed allocation version and cutoff before passing rows. Input rows are not mutated.

function stageLedger(stages) {
  const names = ['assignment', 'exposure', 'analysis'];
  const maps = {}, result = {};
  for (const name of names) {
    if (!Array.isArray(stages?.[name])) throw new Error('Missing stage');
    const units = new Map(); let duplicates = 0;
    for (const row of stages[name]) {
      if (!row || typeof row.unit !== 'string' || !row.unit.trim() ||
          !['control', 'treatment'].includes(row.variant))
        throw new Error('Invalid unit or variant');
      if (units.has(row.unit)) {
        if (units.get(row.unit) !== row.variant)
          throw new Error('Conflicting variant');
        duplicates++;
      }
      units.set(row.unit, row.variant);
    }
    maps[name] = units;
    result[name] = { control: 0, treatment: 0, duplicateRows: duplicates };
    for (const variant of units.values()) result[name][variant]++;
  }
  for (const [later, earlier] of [['exposure', 'assignment'], ['analysis', 'exposure']]) {
    for (const [unit, variant] of maps[later]) {
      if (maps[earlier].get(unit) !== variant)
        throw new Error('Downstream unit missing or changed variant');
    }
  }
  return result;
}
const fixture = {
  assignment: [{unit:'u1',variant:'control'}, {unit:'u1',variant:'control'},
    {unit:'u2',variant:'treatment'}, {unit:'u3',variant:'treatment'}],
  exposure: [{unit:'u1',variant:'control'}, {unit:'u2',variant:'treatment'}],
  analysis: [{unit:'u1',variant:'control'}]
};
console.log(stageLedger(fixture));
// assignment: control=1, treatment=2, duplicateRows=1
// exposure: control=1, treatment=1, duplicateRows=0
// analysis: control=1, treatment=0, duplicateRows=0

This tiny fixture verifies ledger behavior, not statistical significance; its expected counts would be too small for the calculator. Test a warehouse implementation with the same fixtures, plus duplicates arriving after the cutoff, null IDs, multiple allocation versions and mismatched timezones. Compare aggregates with an independently written extraction before treating a clean result as release evidence. In production, retain error counts and approved diagnostic samples privately rather than only throwing an exception.

Find the first unexplained divergence

Build a table with stage, variant counts, expected split, query version and cutoff. If assignment is already inconsistent, begin with allocation history, unit identity and extraction completeness. If assignment looks compatible but exposure does not, inspect delivery and exposure logging. If the discrepancy appears only after analytical filters, review each filter and join. This is an investigation order, not a classifier that proves root cause.

Microsoft Research's SRM discussion distinguishes detection from diagnosis and describes problems with conditions affected by treatment. For a practical checkout example, keeping only people who clicked a newly introduced widget cannot represent the same pre-treatment population in a control experience without that widget. Do not solve this by manufacturing matching control rows; revisit eligibility and the question being estimated.

Rank hypotheses with supporting and contradicting evidence. “Mobile exposure logging may be incomplete” needs a release timeline, a specific client version, aggregate event reconciliation and a falsifiable check. It is not established by a low global p-value. A treatment-driven increase in activity may also change a bot filter or outcome-dependent exclusion. Keep plausible explanations separate from confirmed causes.

The allocation evidence framework assigns each check to an owner and preserves the original contract. Its final decision can be hold interpretation, obtain missing evidence, fix and rerun, or proceed to a separate outcome review. None of those states automatically starts or stops a customer-facing rollout. Operational safety remains with the authorized experiment owner.

Separate diagnostics from repeated monitoring

Running a fixed-look test after every new event does not preserve the original single-look false-alarm interpretation. Searching many segments creates another multiplicity problem. A daily p-value chart can support engineering investigation, but repeated opportunities to flag require a monitoring policy. Use a supported sequential procedure or a planned family-wise review with a statistician when that is the actual need; this calculator implements neither.

Do not keep slicing until a clean subgroup appears and then quietly call it the experiment. A restricted population may answer a different question, and selecting it after observing outcomes adds further bias. Preserve the original result, document the reason for any new analysis and determine whether a fresh experiment is required. Likewise, increasing the threshold after seeing an inconvenient result changes the policy rather than explaining the data.

The numerical implementation computes the chi-square survival function through the regularized incomplete gamma function. NIST DLMF defines that function, and its continued-fraction representation supports a stable upper-tail computation. The implementation is tested against known one-degree-of-freedom tail values, symmetry under variant relabeling and scale behavior. At extreme statistics, floating-point underflow is disclosed instead of presenting zero as a mathematical impossibility.

Produce a review another team can reproduce

Save the calculator's JSON with the scope, counts, configured share, threshold, method version and limitations. Attach the stage ledger's aggregate evidence IDs and versioned queries in your own controlled system. Record unresolved assumptions explicitly. Avoid placing identifiers, credentials or confidential incident details in a public prompt. The free local checker does not need an AI API, sign-in or an audit allowance.

Use the SRM investigation prompt only to organize approved anonymous evidence. It requires facts and hypotheses to be separated, preserves missing-data states and rejects invented queries or launch verdicts. Its synthetic reference output is checked against the deterministic calculator; it is not a claim that a particular model performed the investigation. A human reviewer still owns the decision and verifies every evidence reference.

Once the population issue is resolved, return to the distinct questions of outcome measurement, uncertainty, guardrail effects and commercial value. A technically clean allocation may support further analysis, but cannot tell a merchant whether a discount preserves margin or tell a SaaS team whether activation leads to retained value. The useful end state is not a green badge. It is a bounded, reproducible explanation of which population was compared, what remains uncertain and what decision that evidence can support.

View all growth marketing articles