Analytics & Measurement
Email Bot Clicks: Reconcile Recipients Before Rewriting the Campaign
A click-rate drop can be a definition change, not weaker demand. Build a recipient-level ledger that preserves mixed activity, exposes unknown flags and separates reporting from business impact.
A lifecycle dashboard says 9% of a campaign's recipients clicked. A second report says 3%. The team wants to rewrite the offer, suppress the apparently inactive audience and resend. Before doing any of those things, ask a smaller question: do the reports count the same people under the same classification rule? A different filter can create a dramatic reporting change without a single additional customer action.
This guide is for ecommerce retention teams, consumer subscription operators and startup marketers reviewing email performance. It builds a reproducible recipient-level review, not a bot detector. Every numerical example is synthetic. No example represents Akshay's client performance, and the operating method is an original synthesis rather than a vendor-approved process.
Four takeaways for lifecycle decisions
- Count distinct recipients for one send, not click events divided by recipients.
- A person can have both bot-flagged and explicitly unflagged events; retain that overlap when applying an unflagged-click policy.
- Missing classification is not false. Show unresolved recipients separately.
- A reporting filter changes the measurement definition; clicks alone do not establish human intent, inbox placement or incremental sales.
Use the email evidence review framework to organize ownership, the recipient reconciliation calculator for aggregate arithmetic and the review prompt to challenge an unsupported campaign decision.
Why this belongs in a campaign review now
Current documentation from two email platforms treats automated engagement as a reporting concern, rather than a rare curiosity. Mailchimp describes account-wide bot filtering and historical coverage limits. Klaviyo documents event-level classification and differences between reporting surfaces and attribution settings. These are qualitative evidence that the operational problem exists, not measured search demand or evidence of a new October product launch. Mailchimp documentation; Klaviyo documentation, updated January 26, 2026. Both were checked October 11, 2026.
The practical opportunity is to make campaign preparation more reliable before volume rises. A team can freeze its metric definition before comparing subject lines, sequences or offers. It should not discover halfway through a review that one screenshot used all recorded clicks and another excluded a category. This method is useful whenever settings or exports change; it does not rely on a seasonal forecast.
Start with a single-send contract
Define the campaign send identifier, delivered-recipient roster, click window, timezone and extraction timestamp. Use an interval such as [start, end), with an inclusive start and exclusive end, to avoid double-counting boundary events. Name the report surface and record whether the export is already filtered. A filtered export cannot reconstruct the omitted population unless a separate complete source is available.
Delivered is the roster for this reporting exercise, not a guarantee that the email reached the primary inbox. A recipient who received two sends belongs to two send-level observations, but must appear once in each send's unique-recipient denominator. If you want a campaign-family or monthly person metric, build that separately and give it a different name. Adding per-send unique counts does not produce monthly unique people.
Use stable anonymous IDs in your analytical workspace. Do not paste personal email addresses, access tokens or customer records into a public calculator or an AI prompt. The public companion calculator needs only five aggregate counts and an anonymous scope note. Consent, identity resolution and retention policies remain the responsibility of the team controlling the source data.
Deduplicate event IDs before aggregation. When duplicate records disagree, preserve the conflict rather than selecting the version that improves the rate. Quarantine out-of-roster recipients and missing identifiers. Late-arriving events require either a new versioned snapshot or a declared maturity cutoff. Never quietly change yesterday's denominator while continuing to call the report the same baseline.
A partition that preserves mixed activity
Normalize the source classification explicitly to true, false or unknown. A provider's false flag is evidence of its classification, not a universal guarantee of human behavior. In particular, missing, null and an empty string must not become false through loose coercion. If an export contains text values, validate and map only documented spellings before aggregation.
For each recipient, derive three indicators: at least one true flag, at least one false flag and at least one unknown flag. The following ordered rules create four disjoint buckets. First, if there is a false flag, retain the recipient: classify as mixed if a true flag also exists, otherwise unflagged/no-flagged. Second, among recipients without any false flag, classify as unresolved if any unknown flag exists. Finally, if every click is true, classify as flagged-only.
This order matters. The sequence true, null is unresolved, not flagged-only. The sequence false, null is retained under the explicit-unflagged policy. The sequence true, false is mixed and retained. A recipient with no observed click belongs to none of the four clicked buckets; keep that person in the delivered denominator. Do not manufacture a no-click event to force the person into a classification category.
Why not subtract everyone with a flagged event from all clicked recipients? Because recipients overlap across event classes. Removing a whole person because a security scan occurred can erase a later event that meets your retained definition. The objective is to apply an explicit reporting policy to events and then count recipients, not to assign a permanent bot identity to customers.
Reconcile a 10,000-recipient send
| Disjoint bucket | Recipients | Policy treatment |
|---|---|---|
| Unflagged, no flagged click | 240 | Retain |
| Both unflagged and flagged | 60 | Retain |
| Only flagged clicks | 500 | Exclude from retained policy |
| No unflagged; some unknown | 100 | Separate policy sensitivity |
The sum is 900 observed clicked recipients. With 10,000 delivered, raw unique-click rate is 9%. Retained recipients are 240+60=300, giving 3%. A second, deliberately more inclusive reporting policy retains the 100 unresolved recipients too, giving 4%. The difference is one percentage point, not one percent relative change. There are 9,100 delivered recipients with no observed click in this window.
The interval from 3% to 4% is not a confidence interval and not a bound on the true human click rate. It varies the treatment of one known unresolved bucket; it does not model classifier error, missed events, forwarded links, identity loss or selection bias. Call it a policy sensitivity comparison. That name tells the reader what changed and what did not.
Suppose the original 9% and later 3% screenshots both describe this same send. A definition change alone explains the arithmetic. It does not establish that the campaign improved or deteriorated. The immediate next action is to record the filter history and reconcile downstream evidence, not to change the offer based on the screenshot difference.
Implement the recipient classifier
The small JavaScript function below runs in modern Node.js or a browser. Its inputs are an already deduplicated delivered roster and deduplicated, in-window click records. It rejects out-of-roster recipients, duplicate roster IDs and unsupported flag values. It intentionally does not fetch data, parse vendor exports or infer flags from timing. Those are separate adapters with their own tests and access controls.
function partitionClicks(deliveredIds, events) {
if (!Array.isArray(deliveredIds) || !Array.isArray(events)) throw new Error("Arrays required");
const roster = new Set(deliveredIds);
if (roster.size !== deliveredIds.length || deliveredIds.some(id => typeof id !== "string" || !id.trim())) throw new Error("Invalid roster");
const people = new Map();
for (const event of events) {
if (!event || !roster.has(event.recipient) || ![true, false, null].includes(event.bot)) throw new Error("Invalid event");
const flags = people.get(event.recipient) || new Set();
flags.add(event.bot); people.set(event.recipient, flags);
}
const result = { delivered: roster.size, unflagged: 0, mixed: 0, botOnly: 0, unresolved: 0 };
for (const flags of people.values()) {
if (flags.has(false)) result[flags.has(true) ? "mixed" : "unflagged"]++;
else if (flags.has(null)) result.unresolved++;
else result.botOnly++;
}
return result;
}A compact fixture uses six delivered IDs. A has false, B has true and false, C has true, D has null, E has true and null, and F has no click. The resulting buckets are one unflagged, one mixed, one flagged-only and two unresolved. Five recipients clicked; two meet the retained definition. Repeat clicks do not inflate the bucket count because aggregation uses one set per recipient.
For production scale, implement the same precedence in your warehouse with a grouped send/recipient key and explicit boolean indicators. Filter by send and window before grouping. Join to the delivered roster deliberately; an inner join that drops unmatched click rows can conceal a reconciliation failure, so count and investigate those rows first. Save a versioned aggregate partition and the transformation version with the source snapshot.
Validate three invariants: the four buckets are nonnegative integers; their sum cannot exceed delivered recipients; and retained equals unflagged plus mixed. Reordering events or repeating an identical event must not alter recipient counts. Adding a flagged event to a previously unflagged recipient moves them to mixed but leaves retained unchanged. Adding a false event to an unresolved recipient moves them into retained. These are stronger checks than verifying one attractive example.
Separate restatement from experiment decisions
A filter snapshot belongs with a result. If a metric definition changes, keep the originally reported value and add a clearly labelled restatement. Do not splice old-definition and new-definition points into a seamless trend. If comparable raw evidence exists, recompute the full comparison window with one rule and disclose the revision. If it does not, mark a break and avoid numerical claims across it.
Klaviyo's documentation specifically cautions against changing the bot setting during an active campaign or fixed-date flow experiment. It also describes historical and current result surfaces that can reflect different definition states. Treat that as a vendor-specific warning, not a claim every platform behaves identically. Source: Klaviyo, experiment and reporting behavior.
The operating rule here is broader: the experiment owner signs off measurement changes before they affect a decision. Record the setting, timestamp, approver and affected reports. A reporting cleanup is not a reason to rerun a winner-selection process without documenting what changed. Historical declarations should remain auditable even if the team later decides a corrected analysis is more useful.
Connect the review to business evidence
Review mature orders, qualified enquiries or product actions with separately documented attribution and observation windows. A lower retained click rate can coexist with unchanged orders. That does not prove the campaign has no effect; it means the engagement definition and the business outcome are different measurements. Use an appropriate controlled design for a causal question, rather than translating a click-count correction into revenue loss.
Onsite evidence can help investigate, but missing web events are not proof of bots. Consent choices, blockers, navigation failure, cross-device behavior and identity limits can prevent matching. Report match coverage and unresolved cases instead of demanding one-to-one equality between an email platform and web analytics. This prevents a second measurement problem from being used as a false ground truth for the first.
For a B2C team, the next decision might be whether to investigate a landing-page promise, review send pressure or leave the offer unchanged while resolving classification. For a B2B startup it might be whether a nurture sequence generated qualified conversations rather than security-system activity. Neither requires a universal acceptable bot percentage. What matters is whether the remaining uncertainty could change the specific decision and who owns the missing evidence.
A defensible next step
Complete the framework worksheet, calculate both policies and ask the review prompt to name alternative explanations. Keep the analyst responsible for reconciliation and the lifecycle owner responsible for customer-facing changes. Stop if the roster does not reconcile, historical classification is incomplete or the comparison mixes definitions. A clean review should produce a bounded evidence request or an approved test proposal, not an automatic suppression list.
The result is less exciting than declaring a hidden campaign winner, but more useful: a report someone else can reproduce, an explicit account of what remains unknown and a campaign decision that does not depend on an accidental filter change.