Growth Strategy
AI Feature Unit Economics: From Request Costs to a Defensible Usage Allowance
Build a duplicate-aware task-cost ledger, include failed runs, and convert reconciled AI costs into a transparent monthly allowance scenario without confusing cost feasibility with customer value.
A growth team can promise 500 AI reports per month long before it can explain what those reports cost. The missing link is usually not a token-price table. It is the relationship between billable requests, useful completed tasks and the revenue allocated to the package. Build that relationship before turning an average API bill into a customer entitlement.
Editorial note: AI-assisted research, writing and implementation by Growthcraft Editorial. The operating method is an original synthesis and all examples are synthetic. No Akshay client result, personal review, vendor endorsement or production API execution is claimed. The executable example uses JavaScript with BigInt support and was tested with Node.js 24.
Four practical takeaways
- Chargeable model requests and valuable customer tasks have different identities. Preserve both; a failed task can still have substantial cost.
- Reconcile the entire task cohort before estimating cost per completion. Excluding abandoned work creates a flattering but unusable denominator.
- A monthly allowance can be tested against a contribution budget, but the resulting cap says nothing about willingness to pay or demand.
- Use mature cohorts, explicit cost categories and immutable charge IDs. Keep unknown usage and unapproved assumptions visible before a rollout decision.
Why this belongs in the growth planning process now
Stripe's 20 August 2026 report describes pricing leaders revisiting AI monetisation and the operating processes around it. It is a qualitative industry signal, not a representative demand survey or evidence of search volume. For a growth team, the useful implication is that packaging experiments need reproducible cost evidence, not only a conversion hypothesis. Stripe's pricing-leader report.
Provider usage is also more granular than a single “input tokens” number. Anthropic documents separate input, cache-creation and cache-read counts, with detailed cache-write categories. Its pricing documentation additionally describes tool charges. An integration that prices only one token field can miss part of the workload. Consult the current provider contract for your actual model and platform; no live prices are reproduced here. Prompt-caching usage fields and pricing documentation, accessed 20 September 2026.
This guide addresses a narrower job than choosing an AI marketing stack: producing an auditable cost-per-completion estimate and reviewing an included allowance. It does not recommend a vendor, set a market price, or establish whether an AI feature improves acquisition or retention. Those remain separate product and growth questions.
Define a task before counting a request
A task is the customer job whose completion you promise: an accepted report, a usable enrichment result, or an analysis that passes documented checks. A request is one provider interaction. One task might need several requests, while a request can fail after consuming billable work. An HTTP success, generated token count or returned JSON object is not automatically a useful task completion.
Write a versioned quality rule before collecting the denominator. For a report feature, it might require all mandatory sections, a valid source map and no unresolved validation error. The exact rule is product-specific; this article supplies no universal quality threshold. Keep retries, fallbacks and manual review connected to the originating task. If the definition changes, segment or restate the cohort instead of blending incomparable outputs.
The calculator assumes each completed task includes at least one observed model request. It is not appropriate for a mixed cohort containing cache-only results with no request, unless those are separated. The example also assumes one currency. A EUR package cannot be compared with an unconverted USD cost total merely because the two numbers look similar.
Use three records with different responsibilities
| Record | Identity and grain | What it establishes |
|---|---|---|
| Task | One task ID and definition version | Whether the customer's job met the completion rule |
| Request | One provider request ID belonging to one task | Which provider interactions supported or attempted the job |
| Charge | One immutable charge-line ID, request, category and rate version | The amount assigned to a billable cost component |
Do not use a task ID as a charge deduplication key: that would erase legitimate retries. Do not use a request ID alone either when the same request has separate model and tool costs. A stable charge-line ID identifies an immutable fact. Repeated delivery of that fact should have no effect; a conflicting amount under the same ID should stop reconciliation rather than silently choose the latest row.
In production, retain provider, account, model, service tier, rate-card version, currency, event time, observation time and the source of each amount. The example below deliberately consumes already-normalised charges. It is not an SDK adapter or an invoice parser. Resolve mixed currencies, discounts, cache-write pricing, rounding, adjustments and billing corrections before that boundary. Never place customer prompts, documents or API keys in this cost ledger simply to make reconciliation easier.
Close the cohort before dividing
Select tasks by their start time in a half-open window and record a later observation cutoff. If tasks are still pending, the costs and completions are not mature. Either wait or use a separately documented conservative reserve and a model designed for censoring. Do not move slow or failed tasks outside the cohort to improve its apparent efficiency.
All in-scope task-variable costs belong in the numerator, including failed and abandoned tasks. Only quality-qualified completed tasks belong in the denominator. If there are no completions, unit cost is undefined; the feature is not free. If the numerator is zero because a free allowance or credit covered the observed period, that is an observed accounting condition, not proof of permanent zero economic cost. Run a justified post-credit scenario separately.
Reconcile the ledger total to provider billing or the approved cost source. An unexplained difference is a review item with an owner, not automatically “rounding.” Determine whether its plausible impact could reverse the allowance decision. The method does not prescribe an arbitrary percentage tolerance; materiality depends on the available headroom and the decision being made.
Executable JavaScript — deduplicate charge evidence
Save this block as a local .mjs file and run it with Node.js 24. It uses no dependencies, network, credentials or paid API. Amounts enter as integer millionths of one currency unit to preserve the supplied ledger precision during summation. Conversion to floating-point currency happens only at the reporting boundary; this is an analytical example, not a billing engine.
function summariseTaskCosts(tasks, charges) {
if (!Array.isArray(tasks) || !Array.isArray(charges)) throw new Error("Arrays required");
const validId = x => typeof x === "string" && x.length > 0 && x.length <= 200 && x.trim() === x;
const taskMap = new Map();
for (const task of tasks) {
if (!task || !validId(task.id) || !["completed", "failed"].includes(task.state)) throw new Error("Invalid or immature task");
if (taskMap.has(task.id) && taskMap.get(task.id) !== task.state) throw new Error("Conflicting task state");
taskMap.set(task.id, task.state);
}
const seen = new Map(), requests = new Map(), coveredTasks = new Set();
const totals = { model: 0n, tool: 0n };
for (const row of charges) {
if (!row || !validId(row.id) || !validId(row.requestId) || !taskMap.has(row.taskId) ||
!["model", "tool"].includes(row.category) || !Number.isSafeInteger(row.micros) || row.micros < 0)
throw new Error("Invalid or out-of-scope charge");
const signature = JSON.stringify([row.requestId, row.taskId, row.category, row.micros]);
if (seen.has(row.id)) {
if (seen.get(row.id) !== signature) throw new Error("Conflicting charge");
continue;
}
if (requests.has(row.requestId) && requests.get(row.requestId) !== row.taskId) throw new Error("Request assigned to two tasks");
seen.set(row.id, signature);
requests.set(row.requestId, row.taskId);
coveredTasks.add(row.taskId);
totals[row.category] += BigInt(row.micros);
}
// Explicit zero-cost evidence is required too. Missing rows are not free work.
for (const id of taskMap.keys()) if (!coveredTasks.has(id)) throw new Error("Task missing cost evidence");
const total = totals.model + totals.tool;
if (total > BigInt(Number.MAX_SAFE_INTEGER)) throw new Error("Aggregate exceeds reporting precision");
const completed = [...taskMap.values()].filter(x => x === "completed").length;
return {
requests: requests.size, completed,
modelCost: Number(totals.model) / 1e6,
otherTaskCost: Number(totals.tool) / 1e6,
costPerCompletion: completed === 0 ? null : Number(total) / 1e6 / completed
};
}
// Synthetic: 1,000 completed tasks and 200 failed tasks; one request each.
const tasks = Array.from({ length: 1200 }, (_, i) => ({ id: "t" + i, state: i < 1000 ? "completed" : "failed" }));
const charges = tasks.flatMap((task, i) => [
{ id: "m" + i, requestId: "r" + i, taskId: task.id, category: "model", micros: 19500 },
{ id: "v" + i, requestId: "r" + i, taskId: task.id, category: "tool", micros: 3900 }
]);
const example = summariseTaskCosts(tasks, charges);
console.log(example);
// requests:1200, completed:1000, modelCost:23.4,
// otherTaskCost:4.68, costPerCompletion:0.02808
The two charge categories are normalised buckets, not a claim about provider response fields. A production adapter must check that all required components exist for each request, including explicitly zero-cost components. This reducer checks task coverage and identity consistency, but it cannot discover a missing tool fee if another charge already covers that task. That completeness check belongs upstream and in billing reconciliation.
Turn unit cost into a transparent allowance scenario
The synthetic cohort costs EUR 28.08 in total: EUR 23.40 model cost plus EUR 4.68 other task-variable cost. Dividing by 1,000 completions gives EUR 0.02808 per completed task. Dividing 1,200 requests by those completions gives 1.2 requests per completion. Here that ratio reflects failed tasks, not retries; calling it a 20% retry rate would be wrong.
Now define one account-month. Let net revenue allocated to this scope be EUR 49, other in-scope account-variable cost EUR 5, and planned completed tasks 500. Task cost is 500 × 0.02808 = EUR 14.04. Contribution after these listed costs is 49 − 5 − 14.04 = EUR 29.96, or approximately 61.14% of net revenue. This excludes any costs not entered; do not label it company profit or accounting gross margin.
For a chosen 60% target, the task-cost budget is 49 × (1 − 0.60) − 5 = EUR 14.60. The whole-task cap is floor(14.60 / 0.02808) = 519. At 520 tasks the cost becomes EUR 14.6016 and exceeds the budget. A displayed EUR 0.03 unit cost is useful for casual reading but too coarse for this boundary calculation.
Use the AI-feature allowance calculator to reproduce and export that scenario. It distinguishes unknown unit cost, zero revenue, a negative task budget and zero observed task cost. Those cases must not collapse into one cheerful green result.
Test failure modes before trusting the memo
Duplicate every charge and reverse the order: the result must remain unchanged. Change the amount on a duplicated ID: reconciliation must fail. Assign one request to two tasks: fail. Introduce a pending task, unknown category, negative amount, missing evidence row or unsafe integer: fail. An empty cohort should return null unit cost, not a divide-by-zero result.
At the calculator layer, validate blank, negative, fractional-count and non-finite inputs. Check the cost-based cap against both itself and the next integer task. Scaling a cohort's requests, completions and cost totals by the same factor should preserve unit cost and the resulting allowance. Increasing task cost while holding revenue and the target fixed must not increase the cap. These invariants catch errors that one attractive default example misses.
The implementation uses browser numbers for scenarios, so it does not replace an approved financial rounding policy. Keep billing calculations in an appropriate decimal or integer-unit system, retain source precision, and reconcile credits and adjustments. Treat an allowance sitting exactly on a computed boundary as a review issue, not spare operational capacity.
Stress the customer job, not just the average
A single average can conceal expensive segments. Separate long-context reports, short summaries, high-failure jobs and cache-cold sessions when their cost structures differ. Sample the relevant cohort rather than multiplying by a plausible-sounding “AI overhead factor.” If you use a hypothetical doubled unit cost of EUR 0.05616, label it as a scenario: the same budget now supports 259 tasks, not 519. No probability or confidence interval follows from that arithmetic.
Then ask whether the lower allowance still lets the customer finish the promised job. If it does not, changing the limit may reduce perceived value even when the spreadsheet improves. Alternatives include changing workflow scope, reducing unnecessary requests, improving first-pass quality, revising the package, or running a narrower experiment. Each needs its own value and operational evidence; this article does not claim any will improve conversion.
Separate a commercial monthly entitlement from concurrency and safety controls. A customer might have unused monthly tasks while a temporary operational limit prevents a burst. Explain that distinction in product copy and support procedures. Do not describe “unlimited” usage based on a zero-cost pilot or hide an undisclosed operational restriction inside the pricing story.
Make the rollout decision inspectable
Use the five-gate allowance framework to record scope, reconciliation, cost feasibility, customer value and rollout authority. The evidence-review prompt can organise the supplied facts, but a model must not invent an approver, price, user preference or missing source.
For the synthetic example, the base cost scenario fits a 500-task proposal. Missing value evidence and a doubled-cost stress case still justify a HOLD while the team investigates. Record the experiment segment, monitoring owner, review date, communication plan and rollback trigger before changing customer entitlements. Reopen the decision when the model, rate card, task definition, quality rule or packaging changes. The useful outcome is not a permanent number; it is a decision that another person can reproduce and challenge.