Measurement / framework / Free to use
The Experiment Allocation Evidence Framework
An editable assignment-to-analysis review that separates allocation faults, exposure loss and eligibility changes before a growth team interprets conversion lift.
Growthcraft Editorial · 2026-09-27. AI-assisted research and implementation. Examples are synthetic; Akshay's personal review is not claimed.
Copyable template
Select and copy the complete template below. With JavaScript enabled, you can edit, copy and download it in the interactive workspace.
EXPERIMENT ALLOCATION EVIDENCE REVIEW — original Growthcraft synthesis Decision and experiment version: Experiment owner / data owner / implementation owner / decision reviewer: Randomized unit (user, account or other) and independence rationale: Configured control share and configuration evidence: Fixed allocation window, timezone, extraction cutoff and late-event policy: Pre-specified look and flag threshold; repeated-monitoring policy: 1. FREEZE THE CONTRACT Eligibility before treatment; exclusion rules and version: Identity mapping / conflicting-assignment policy: Assignment, exposure and analysis stage definitions: Any allocation ramp, crossover or data gap? Evidence and owner: Status: usable contract / unresolved / incompatible with checker. 2. RECONCILE THE STAGES Stage | control unique units | treatment unique units | expected share | evidence ID Assignment: Exposure: Analysis eligibility: Units present downstream but absent upstream: Unknown variant / conflicting units / duplicates (report separately): Counts before and after every join, filter and deduplication: 3. CHECK, THEN INVESTIGATE Calculator export / expected counts / chi-square / p-value: State: no units / small expected counts / investigate / no flag. First stage with unexplained divergence: Hypothesis | supporting evidence | contradicting evidence | next falsifiable check | owner Do not delete a segment merely to remove the flag. 4. MAKE A BOUNDED HANDOFF Hold interpretation / fix and rerun / obtain missing evidence / proceed to separate outcome review: Reason and remaining uncertainties: Change version / replay plan / independent re-extraction: Reviewer, due date and evidence required to close: Outcome analysis, guardrails and commercial decision remain separate.
A health check before the conversion story
This original Growthcraft synthesis is for analysts, growth leads and engineers running two-arm checkout, onboarding or activation experiments. It turns an allocation warning into a bounded evidence request. It is not a new statistical test, a launch score or proof that a variant wins. Use it before interpreting lift, or when the exposure and outcome dashboards disagree.
Prerequisites are the saved allocation configuration, a stable randomized-unit definition, a fixed window and aggregate counts from each data stage. Do not upload customer identifiers. An account-randomized B2B experiment must count accounts here, not the people inside those accounts. A consumer experiment randomized by user must not count purchases as independent assignments.
Gate 1 — freeze a reproducible contract
The experiment owner supplies the allocation version and pre-treatment eligibility rule. The analyst records extraction cutoff, timezone, duplicate policy and scheduled review. If the split changed from 10/90 to 50/50, preserve separate allocation periods; one final dashboard setting does not describe the historical population. If units can move between variants or the assignment unit is unclear, stop and resolve that contract before calculating a reassuring number.
The worksheet's threshold is a documented review policy, not a benchmark. The worked example uses 0.001 for one planned look. Continuous monitoring or many segment checks requires another error-control design. A low p-value is a reason to investigate; a high one is not evidence that every part of the pipeline is sound.
Gate 2 — reconcile assignment, exposure and analysis
The data owner produces three rows with unique-unit counts and query/version references. Assignment means a decision was recorded. Exposure means the experience was delivered under the agreed logging definition. Analysis eligibility means the row survived the specified analytical filters. These are different populations. Check downstream-only units, conflicting variants and duplicate rows separately; never force the totals to balance by deleting inconvenient records.
For each join or filter, record counts on both sides and why the reduction is expected. An exposure condition affected by the treatment can alter the estimand. A stage-level discrepancy is diagnostic evidence, not automatic proof that the assignment algorithm is broken.
Gate 3 — work the synthetic incident
A checkout experiment has a fixed 50/50 split and 10,000 assigned users: 4,800 control, 5,200 treatment. Expected counts are 5,000 each. Pearson's statistic is 16; its one-degree-of-freedom upper-tail p-value is approximately 0.0000633425. At the pre-specified 0.001 threshold, the review state is investigate. Control's observed share is 48%, or −2 percentage points from configuration.
Suppose exposure counts are 4,700/5,100 and analysis counts 4,700/4,900. These additional synthetic rows do not explain the first discrepancy. Assignment already needs investigation; later losses need their own reconciliation. The implementation owner checks configuration version and identity conflicts. The data owner checks late arrivals, join keys and extraction boundaries. Neither owner may assume that a 200-user difference is caused by a particular browser without supporting evidence.
A second synthetic experiment with 6,000/4,000 units under a configured 60/40 split has statistic zero and p-value one. It has no allocation flag, but can still have broken revenue logging. A result of 8/2 at 90/10 has expected counts 9/1, so this calculator does not report an approximate p-value.
Gate 4 — close the evidence, not just the alert
The reviewer chooses hold interpretation, obtain evidence, fix and rerun, or proceed to a separate outcome review. A closure note names the cause supported by evidence, the affected scope, code/configuration change, independent re-extraction and remaining uncertainty. If restarting, fix the underlying fault first and use a new version; restarting alone can reproduce the same error.
Never salvage a desired uplift by repeatedly changing eligibility. A restricted analysis may answer a different question and needs explicit statistical review. Record it separately from the original test rather than quietly replacing the original population.
When not to use this shortcut
Adaptive allocation, switchback designs, clustered dependence and multivariate experiments do not satisfy this two-arm fixed-allocation contract. Small startups may have too few units for the approximation and too little outcome power even if allocation looks balanced. Use a suitable design and statistical method. The worksheet cannot establish causal identification, estimate lost revenue or authorize a rollout.
Method and source boundaries
Retrieved 26 September 2026. Statsig's current SRM documentation describes checking unique-unit allocation against the configured split. NIST's goodness-of-fit guidance supports the expected-count caution. Microsoft Research's 14 September 2020 article separates detection from diagnosis and discusses problems with treatment-dependent filtering. These sources do not endorse this original worksheet, prompt or policy threshold.