Performance Marketing

Creative Testing: Measure Delivery Concentration Before Declaring a Winner

Five uploaded ads are not five equally tested ideas. Build a reconciled impression ledger, calculate effective creative count and separate delivery evidence from causal learning, fatigue and creative quality.

Your team uploads five creative concepts. One receives 70% of the impressions, three receive 10% each and one receives none. The reporting deck calls this a five-way creative test and proposes retiring the ad with no delivery. Neither conclusion follows from the export. You know where observed exposure went; you do not yet know whether every idea was eligible, whether audiences were comparable or which idea would perform best under a controlled comparison.

This guide gives ecommerce, consumer-app, subscription and B2B startup teams a small reproducible diagnostic for that gap. It is a delivery review, not a replacement for an experiment or a creative strategy. All counts and business situations below are synthetic. They are not Akshay's client results, and the workflow is an original operating synthesis rather than a vendor-certified method.

Four takeaways for the next creative review

  • Keep listed, eligible and delivered creatives separate; a missing export row is not automatically a zero.
  • Calculate shares only from an additive reporting grain and one reconciled window.
  • Use inverse squared-share concentration to describe balance, not to estimate statistical sample size or creative quality.
  • Investigate eligibility and allocation before retiring an ad; design a controlled follow-up when the question is causal.

Use the creative delivery review framework to organize evidence, the delivery concentration calculator to inspect a snapshot and the evidence review prompt to challenge the proposed next action.

More assets do not automatically produce more learning

Google's October 1, 2026 creative guidance discusses reusing suitable social assets for YouTube, making changes gradually and prioritizing asset quality over volume. That is a timely reason to examine the learning process behind a growing creative library. It is not evidence that a particular number of variants is optimal, or that this topic has a measured keyword-search volume. Source: Google, creative testing and asset guidance.

Separately, Google's ad-rotation documentation explains that optimized serving can favor ads expected to perform better. It also warns that its non-optimized option does not guarantee equal impression percentages. The page scopes those settings to specified campaign types; do not assume the same control exists in every campaign or that changing it creates a randomized experiment. Source: Google Ads ad rotation.

The practical question is therefore narrower than which creative won: how much of the observed delivery reached each defined creative, and what evidence would justify the next decision? A delivery system can be doing its commercial job while providing an uneven opportunity to learn about alternatives. Those objectives need not be identical. Forcing a system to look balanced can also change costs and delivery; balance itself is not a business outcome.

Start with a grain that can actually be added

Write the analysis contract before exporting. Specify campaign or ad-group scope, a settled start-inclusive and end-exclusive window, timezone, export timestamp and metric definition. Choose whether a creative means a complete ad, an asset, a rendered combination or a versioned concept. A concept with several placements may map to several ad IDs; document that mapping rather than changing definitions between reviews.

Impressions must be mutually exclusive across the rows being summed. If one rendered ad contains several assets and each asset is credited with that impression, summing asset rows can count one impression several times. In that situation use a mutually exclusive ad or combination grain, or explicitly define a different asset-exposure metric. Do not label a sum of overlapping asset credits total ad impressions.

Build a roster independently of the nonzero export. Keep anonymous stable IDs, version, approval time, active dates, eligible placements and any exclusion reason. Left-join reported counts onto this roster only after verifying export completeness. A genuinely eligible ad with no delivered impressions can receive zero; an absent record caused by an incomplete export should remain missing until resolved. The website calculator expects that preparation to have happened and cannot infer it from numbers.

Aggregate daily rows only when periods are disjoint and the creative version has not changed. A duplicate export file should not double the count. The calculator rejects repeated IDs, including repeated zero rows, to make the ambiguity visible. Its IDs are case-sensitive, so A and a are distinct unless you intentionally normalize them upstream. Keep a reconciliation between the row sum and the same-scope platform total, including documented rounding or reporting differences where applicable.

Describe concentration without a made-up quality score

Let n_i be the impression count for creative i and N the sum of all counts. For positive N, define p_i=n_i/N. Squared-share concentration is C=sum(p_i²). Its reciprocal E=1/C describes the number of equally delivered categories that would have the same concentration. With k positive-count creatives, E lies between 1 and k; equal shares produce k, and a single delivered creative produces 1.

The arithmetic is the inverse Simpson index. The scikit-bio reference documents the reciprocal squared-proportion definition and distinguishes an optional finite-sampling correction. Here we use the uncorrected plug-in calculation descriptively on impression categories. This is a transfer of a standard formula, not a claim that its ecological interpretation validates marketing decisions. Source: scikit-bio inverse Simpson documentation.

Call E the effective creative count, but always add that it is not effective statistical sample size. Impressions can repeat for the same person, and people can encounter several ads. The value does not count independent users, independent experiments or genuinely different ideas. It is a compact summary of a distribution with a clearly defined category boundary.

Report N, listed count, positive-count creatives, zero-delivery count, largest share, C and E together. E alone conceals scale: 7, 1, 1, 1 and 7,000, 1,000, 1,000, 1,000 have the same concentration. The larger snapshot supplies more observed impressions, but neither guarantees unbiased comparisons. Adding zero rows changes the roster and zero count, not E. Splitting one creative into artificial IDs can inflate E, so stable grouping is part of the measurement contract.

A five-creative review, worked end to end

Creative IDImpressionsShare
Hook-A7,00070%
Hook-B1,00010%
Hook-C1,00010%
Hook-D1,00010%
Hook-E00%

Total delivery is 10,000. Five creatives are listed, four have positive delivery and one has zero delivery. C=0.7²+0.1²+0.1²+0.1²+0²=0.52. E=1/0.52, approximately 1.9231. The exposure distribution is as concentrated as approximately 1.92 equally delivered categories. It does not mean that exactly 1.92 ideas were tested or that the remaining ideas failed.

Suppose E was pending approval during most of the window. That explains a plausible eligibility constraint, not creative quality. Suppose A also ran in more placements. Its larger share now has at least two plausible contributors: eligibility and optimized serving. The snapshot cannot separate those effects. The review should request approval and placement history before recommending retirement or a follow-up test.

A useful memo states what is known: exposure was concentrated, the row sum reconciled, and E's eligibility is unresolved. It states what is not known: relative causal response, fatigue and mature downstream value. It then names an owner and the smallest next evidence request. This is more actionable than a colored winner badge because it tells the team which uncertainty is blocking the decision.

A small implementation you can verify

The following JavaScript accepts integer counts, returns null for undefined zero-total metrics and never silently converts a missing value to zero. It runs in a modern browser or Node.js without dependencies. The website adds a strict ID/count parser, scope notes, accessible errors and copyable JSON. Input limits are engineering bounds, not recommended campaign sizes.

function deliverySummary(counts) {
  if (!Array.isArray(counts) || counts.length === 0 || counts.length > 100)
    throw new Error("Use 1 to 100 counts");
  if (counts.some(n => !Number.isSafeInteger(n) || n < 0 || n > 1000000000))
    throw new Error("Counts must be integers from 0 to one billion");
  const total = counts.reduce((sum, n) => sum + n, 0);
  const delivered = counts.filter(n => n > 0).length;
  const shares = total ? counts.map(n => n / total) : null;
  const concentration = shares ? shares.reduce((sum, p) => sum + p * p, 0) : null;
  return { total, delivered, zero: counts.length - delivered,
    concentration, effective: concentration === null ? null : 1 / concentration };
}
console.log(deliverySummary([7000, 1000, 1000, 1000, 0]));

The total remains within exact integer range under these limits. Shares and reciprocal calculations use ordinary floating-point arithmetic, so comparisons in tests need a small numerical tolerance. Round only when presenting the result; preserve the unrounded values and method version in exported evidence. Do not round each share before squaring it, because that changes the calculation.

Store the anonymous scope with the result, including window, grain and source-export reference. Do not place customer identities, access tokens or raw audiences into this tool. The calculation runs locally without a provider API or account connection. The exported result is a review artifact, not an instruction file to upload into an ad platform.

Test properties, not just the default example

A known-answer test checks the 0.52 concentration above. Four equal positive counts should return E=4; one positive row and any number of zeros should return E=1. All-zero rows must return undefined shares and concentration, represented as null in JSON. Empty data, negative values, fractions, non-finite values, duplicate IDs and out-of-range counts must fail visibly rather than fabricate a result.

Three invariants are particularly useful. Permuting rows must not change aggregate metrics. Multiplying all counts by the same positive factor must leave concentration unchanged, provided the scaled inputs remain within range. Appending a zero row must leave concentration unchanged while increasing listed and zero counts. Also verify 1≤E≤k within floating-point tolerance for every positive-total fixture.

The companion implementation tests these properties across deterministic fixtures and executes this actual article snippet, not a separately retyped approximation. That establishes consistency of the published arithmetic; it does not validate a customer's source export, randomization or causal interpretation. Those remain review responsibilities.

Choose the next action from the unanswered question

If the question is whether the data are complete, reconcile exports and roster first. If it is whether a concept was eligible to run, inspect approval, schedule and placement evidence. If it is whether the current allocation supports the business objective, review mature outcomes, costs and customer quality separately. A concentration statistic cannot substitute for any of those checks.

If the question is which creative causes better outcomes under comparable conditions, prepare an experiment brief. Define the assignment unit, eligibility population, randomization, primary outcome, detectable effect, sample-size plan, guardrails and stopping rule. Review interference and repeated exposure. The appropriate platform workflow depends on the actual campaign type; a checkbox that changes rotation is not by itself a complete experimental design.

A small startup may not have enough volume to resolve a narrow causal difference affordably. In that case, use structured qualitative feedback, explicit message hypotheses and a bounded exploratory launch while labeling the evidence exploratory. Do not convert a low-volume allocation report into statistical certainty merely because the team needs a decision this week.

For B2C, preserve purchase maturity, returns and contribution economics when evaluating outcomes. For a B2B campaign, preserve qualification and sales-progression windows. For an app, separate install volume from activation and retention. These outcome contracts can change the commercial verdict without changing the delivery concentration at all.

What this diagnostic deliberately does not tell you

It does not diagnose fatigue. A change over time can reflect audience composition, auction conditions, placements, approval dates or a changed roster. To investigate fatigue, define a time-series question and compare the relevant exposure and outcome evidence under explicit controls. A single E value contains no time ordering or individual frequency information.

It does not prescribe an optimal number of ads, an equal-share target or a universal concentration threshold. Efficient delivery and broad learning can require different operating choices. It also cannot detect whether ten IDs are ten distinct ideas or cosmetic variants of one concept. That requires a creative taxonomy and human review of the actual messages.

Use the worksheet to record the evidence boundary, inspect the calculator output, and challenge the memo with the review prompt. For the wider production workflow, read the existing creative operating-system guide. The next useful action is the one that answers the unresolved business question—not the one that makes a distribution look prettier.

Editorial disclosure: AI-assisted research, writing and implementation by Growthcraft Editorial. Sources checked October 8, 2026. Examples are synthetic; Akshay's personal review and model-generated prompt evaluation are not claimed.

Need help turning this method into an operating plan?

Scope measurement, ownership and a practical growth backlog with Akshay. Start with the decision you need to make and the evidence available.

Explore hands-on growth leadership

View all growth marketing articles