Measurement / calculator / Free to use
Two-Arm Sample Ratio Mismatch Calculator
Check observed unique-unit counts against a pre-specified traffic split with a transparent chi-square diagnostic, small-count guardrails and a downloadable evidence record.
Growthcraft Editorial · 2026-09-27. AI-assisted research and implementation. Examples are synthetic; Akshay's personal review is not claimed.
Enable JavaScript to change inputs in the interactive calculator. The complete formulas, default example and limitations are available below.
Use counts from one allocation contract
Supply the observed control and treatment counts at one explicitly named stage. Use independent unique randomized units, with one variant per unit, the same eligibility rule and a common extraction cutoff. Enter the configured control percentage; treatment receives the remainder. Do not infer the expected split from the observed data, or compare conversions when assignment counts are required.
Each count accepts integers from zero to one billion. The control share must be strictly between 0 and 100%. The review threshold accepts probabilities from 0.000001 to 0.1; enter 0.001, not 0.1, for the example's 0.1% threshold. These input bounds are implementation safeguards, not evidence-quality guarantees. The scope field is included in the downloadable record, so keep it anonymous.
The calculation and its assumptions
Let N = A + B and q be the configured control share as a fraction. Expected counts are EA = Nq and EB = N(1 − q). Pearson's statistic is X² = (A − EA)²/EA + (B − EB)²/EB. Two fixed categories and no fitted parameters give one degree of freedom. The p-value is the upper-tail probability under that null allocation, not the probability that the implementation is broken.
The local implementation evaluates Q(1/2, X²/2), the regularized upper incomplete gamma function, with a convergent series or continued fraction. See NIST DLMF definitions and continued fractions. It uses JavaScript double precision, not an exact-binomial test. Extremely small tails can underflow; the interface says below floating-point range rather than claiming mathematical probability zero.
A reproducible synthetic answer
At A = 4,800, B = 5,200 and q = 0.5, expected counts are 5,000 each and X² = 8 + 8 = 16. The p-value is about 0.0000633425, below 0.001. Investigate allocation and data processing before interpreting the outcome. The control share is 48%, a −2 percentage-point difference from the expected share—not a −2% relative difference.
At A = 6,000, B = 4,000 and q = 0.6, X² = 0 and p = 1. This says only that this snapshot exactly matches the configured ratio. It does not demonstrate that the test is unbiased. At zero total, there is no calculation. If either expected count is below five, the approximation is withheld; consider a suitable exact procedure or a pre-specified larger sample.
What to do with the result
An investigate state uses strict p < threshold on the unrounded value. Freeze the configuration, reconcile assignment/exposure/analysis populations and record falsifiable hypotheses. A no-flag state sends you to the rest of the data-quality checklist, not directly to a ship decision. Save the JSON alongside query versions and extraction time so another reviewer can reproduce the numbers.
Do not repeatedly refresh until the flag disappears. This is a single-look check without sequential or multiple-comparison adjustment. Segment analysis can suggest causes but is not an automatic repair. It cannot estimate outcome bias, treatment lift, sample-size sufficiency or financial value. More than two arms, changing allocations and dependent observations need a different method.
Method and source boundaries
Retrieved 26 September 2026. Statsig's current SRM documentation describes checking unique-unit allocation against the configured split. NIST's goodness-of-fit guidance supports the expected-count caution. Microsoft Research's 14 September 2020 article separates detection from diagnosis and discusses problems with treatment-dependent filtering. These sources do not endorse this original worksheet, prompt or policy threshold.