stripetoken billingai pricingusage-based pricingchurn prevention

Stripe's Own Pricing Lead Tried to Kill Token Billing. The Reason Was Churn, Not Margin.

Metronome's CEO says a token-markup invoice teaches AI customers to shop your margin. Here's the churn mechanism, and the fix Stripe just shipped.

XY
10 October 2026 · 8 min read

Metronome's CEO Scott Woody spent part of a blog post on Stripe's own site making a case against the pricing model his own company built its name on. Token billing — charging customers for the model tokens their usage consumes, plus a markup — is how most AI products meter cost today. Woody's argument, published October 1 on stripe.com/blog, isn't that token billing is broken as infrastructure. It's that showing it to the customer, line by line, is a slow-motion churn mechanism most AI pricing teams haven't connected to their renewal numbers yet.

Key stat
1 in 6
Stripe users past key revenue milestones now running hybrid pricing as of August 2026 — a model Metronome built support for 3 years earlier and watched sit unused for 18 months
Source: Scott Woody (Metronome CEO), "Why I tried to kill token billing (and why we kept it)," Stripe Blog, Oct 1, 2026

The mechanism is worth taking seriously precisely because it comes from the company Stripe paid roughly $1 billion to acquire specifically for its usage-metering expertise. If the people running the pricing infrastructure underneath a meaningful share of AI billing on Stripe are telling customers not to expose token math on an invoice, that's not a minor UX opinion. It's a warning about what that invoice does to a renewal.

What a token-markup invoice actually tells a customer

Token billing exists for a real reason: it's a safeguard against runaway compute costs, and it's easy to deploy and explain. For a model provider itself, token usage is the product. For almost every other AI company building on top of model APIs, tokens are just a convenient, countable basis to start metering from — not a reflection of what the customer is actually paying for.

The problem shows up the moment that metering becomes the customer-facing price. Woody's example is specific: an invoice itemizing 10 different models, each with its own markup — 10% on one, 5% on another, 25% on a third. Every one of those numbers is a fact the customer can now argue with. You've effectively told them your value is the spread between your price and someone else's model cost, at a moment when that underlying model cost is falling and getting more interchangeable by the month. A customer staring at a 25% line-item markup doesn't see sophistication. They see a number they can negotiate down, or a reason to find a provider who skips the markup entirely.

Pricing surfaceWhat the customer seesWhat it invites them to do
Raw token billingWhich models ran, tokens consumed per model, markup per modelContest the specific markup, or shop a provider without one
Unified creditsA single credit balance that burns down per action (an enrichment, a generated image)Evaluate the work delivered, not your model-routing decisions
Output-based pricingA price per completed, countable unit of workCompare you on results per dollar, not on margin structure
Outcome-based pricing (rare)A price tied to a business result like revenue or retentionDispute attribution — did your product cause the result, or something else

Notice that the failure mode isn't dissatisfaction with the product. It's dissatisfaction with the pricing mechanics sitting on top of a product the customer may still like. That distinction matters for anyone building a cancellation flow or a renewal process, because a customer leaving over markup math needs a completely different save conversation than one leaving over a missing feature — and a cancel-reason survey that only offers "too expensive" as a bucket will never tell you which one you're actually looking at.

Why this is a margin story that becomes a churn story

Woody frames the immediate risk as margin compression: as models get cheaper and more interchangeable, exposing your markup invites customers to challenge it or route around you, and competition pushes that margin toward zero. That's the mechanism from the vendor's side. From the customer's side, the same dynamic reads as a trust problem, and trust problems in billing are a pattern this blog has tracked before — we've written about how a mathematically correct but unexplainable invoice line erodes confidence in a billing relationship well before a customer files a complaint about it, and a markup-itemized AI invoice is the same failure mode wearing a different outfit.

The difference with token billing is that the number isn't a one-time mid-cycle oddity like a proration charge. It's recurring, it's visible every billing cycle, and it's attached to a cost basis — model pricing — that's publicly falling. A customer who sees your 25% markup on GPT-4-class tokens this month and reads a headline next month about that same model class getting cheaper has a very specific, very legible reason to open a renewal conversation you didn't initiate. That's a slower, quieter churn path than a bill-shock cancellation, and it tends to surface at renewal rather than mid-cycle — closer in shape to the committed-spend floor renewal fights we've covered before than to an immediate cancel-page event.

What each pricing metric actually measures
Output (countable, verifiable work)90
Input (raw cost, e.g. tokens)55
Outcome (attributed business result)20

Illustrative framing of Woody's three-metric taxonomy (input, outcome, output), not a measured index. Source: Stripe Blog, Oct 1, 2026.

The fix Metronome actually shipped: unified credits

Woody's answer isn't to rip out token metering — Metronome still runs token-level tracking on the backend so companies can keep routing models and tuning margin as workloads shift. The change is what sits on top of it. A unified credit model gives the customer one credit balance that draws down per product action, with each action consuming credits at a rate the vendor sets based on the underlying token cost. The invoice shows "1,200 records enriched" or "340 images generated," not a 10-model breakdown with a markup column. The complexity doesn't go away. It just stops being the customer's problem to audit.

That's a meaningful, specific difference from the usage-flatlining problem we've covered in usage-based billing churn generally. That earlier piece is about losing the signal that an account is quietly disengaging. This is a different failure: the signal is fine, the invoice lands exactly on schedule, and the customer leaves anyway because the invoice itself reads as an argument against staying. A unified credit layer doesn't fix usage decay. It fixes the specific thing that turns an engaged, paying customer into a renewal risk over pricing mechanics alone.

Output-based pricing, and why "outcome-based" is mostly marketing

Woody draws a three-way distinction worth keeping straight, because the terms get used interchangeably in a lot of AI pricing commentary. An input is a cost, like a token. An outcome is a business result — more revenue, more pipeline, less churn — that's valuable but genuinely hard to attribute to your product alone. An output is an objective, countable unit of completed work: an image generated, an email sent, a support ticket resolved.

His test case is Fin, the AI customer-support product that priced on resolved conversations and got widely described as the poster child for outcome-based pricing. Woody's read is that Fin's metric is actually an output — a resolved ticket is something you can count — not a true outcome, which would require proving the resolution itself (versus a human rep, or the customer giving up) caused whatever business result followed. Within a quarter of Fin introducing that pricing, all five leading companies in its category had committed to some version of outcome-based pricing, which is the fast-cascade pattern Woody says pricing models follow once a category matures: slow for years, then sudden once one credible player proves it out.

Genuine outcome-based pricing, in his framing, only works with monopoly-like market power or contract sizes large enough to justify instrumenting an entire workflow to prove causation every billing period. For nearly everyone else, what gets marketed as outcome-based pricing is an output metric wearing better branding. That's not a criticism of the tactic — Woody says explicitly that companies should "define your output metric, call it an outcome, and convince the market that you're right" — but it's a useful gut-check before you build pricing infrastructure around proving something you can't actually attribute cleanly.

What to check before your next pricing review

  • Audit what your invoice actually itemizes today. If a customer can see a per-model markup percentage anywhere on their bill, that's a specific number sitting in their inbox every cycle, available to contest at the next renewal whether or not they've said anything yet.
  • Separate your backend metering from your customer-facing metric. Woody's point isn't to stop tracking tokens — it's to stop pricing on them directly. Keep the token-level data for margin management; present the customer a credit balance or a completed-work count instead.
  • Pressure-test any "outcome-based" pricing you're running or considering. If you can't instrument the full workflow and prove causation every period, you're very likely running output-based pricing with outcome language on top of it — which is fine, as long as the metric itself is still objective and countable.
  • Watch markup-sensitive renewals as their own churn category. A customer who cancels right after a model-cost headline, or right after you've added a new model to the mix at a higher markup, is giving you a cleaner signal than a generic "too expensive" reason code. Capture it separately.

None of this is a reason to panic about existing token billing — Woody's own company still runs it as infrastructure, and it's a reasonable place to start metering a new AI product before you know what customers actually value. It's a reason to treat the presentation layer on top of it as a pricing decision with real churn consequences, not a cosmetic invoice-formatting choice. If your product already tracks usage-based or hybrid billing on Stripe, run your blended markup exposure through our gross margin calculator before your next renewal cycle — it's a fast way to see how much of your margin is sitting exposed on an invoice line a customer could contest, versus how much is protected behind a credit or output layer they never have to parse. And if a renewal conversation is already heading toward a markup dispute, that's exactly the context a cancellation or downgrade page from CancelFlow needs before the customer gets there — the save that works is an honest explanation of what they're actually paying for, not a discount on a markup they were right to question in the first place.

Frequently asked questions

What is token billing in AI SaaS pricing?+

Token billing, also called passthrough pricing, charges a customer for the underlying model tokens their usage consumes, plus a markup on top. It's the default starting point for most AI products because tokens are an easy, countable basis to meter — but it means your price is defined by your costs, not by the value the customer actually got.

What's the difference between token billing, unified credits, and output-based pricing?+

Token billing shows the customer the raw mechanics: which models ran, how many tokens each consumed, and the markup applied to each. A unified credit model keeps that metering on the backend but presents the customer with a single credit balance that burns down per product action — an enriched record, a generated image — regardless of which model produced it. Output-based pricing goes a step further and charges per completed unit of work directly, with no credit abstraction in between. Metronome's CEO, Scott Woody, describes unified credits as a stepping stone toward output-based pricing, which he calls the best current way to align AI pricing with value.

Why does showing customers a token-level invoice create churn risk?+

Because it defines your value as the gap between your price and your model provider's cost, and that gap is exactly what a customer can contest. Per Woody's framing, if an invoice itemizes 10 models at markups of 10%, 5%, and 25%, a customer now has a specific number to push back on, or a specific reason to route around you as cheaper model access becomes available. Competition drives that markup, and the vendor's margin, toward zero — and the first visible symptom is usually a renewal conversation about price, not a quiet product complaint.

Is outcome-based pricing a realistic alternative to token billing?+

Rarely, according to Woody. Outcome-based pricing — billing on a result like revenue generated or churn reduced — only holds up with monopoly-like market power or contract sizes large enough to justify instrumenting a workflow end-to-end to prove causation. For almost everyone else, what gets marketed as "outcome-based" is actually output-based: a countable, objectively verifiable unit of completed work, like a resolved support ticket or a generated image, dressed up with outcome language because it tests better with buyers.

Try CancelFlow

Stop losing subscribers today

One script tag. One function call. A live cancellation flow in under 10 minutes.

Start free trial →
← All postsHome