A token price is a very tempting threshold because it gives you one number to watch. If models keep getting cheaper, there should be some point where charging for AI usage stops making sense and the feature can simply disappear into the subscription.
I don't think that is the useful level of the problem.
A product is not packaging tokens. It is packaging a workload with a distribution: ordinary requests, strange requests, heavy users, long contexts, retries, tool calls, runtime, and sometimes a choice between very different models. A lower model price can move that distribution. It does not tell you whether the workload has become predictable enough to promise inside a flat subscription.
The useful threshold is the packaging boundary, not the token price.
That changes the question I would want answered. Instead of asking whether AI is broadly cheap enough to bundle, I would ask which specific AI workload is bounded enough, margin-compatible enough, and valuable enough to let the subscription carry it without usage pricing doing essential economic or control work.
Zoom shows compatibility, not a cutoff
Zoom is the strongest current example in this source set because several useful facts coexist without giving us a private unit-economics number.
AI Companion was introduced as an included capability for eligible paid users. Zoom has also described a federated approach that combines its own and third-party models as part of controlling AI cost. In its FY2026 reporting, the company said GAAP gross margin improved by 120 basis points to 77.0% while discussing cost optimization as AI usage scaled.
By 2026, the packaging had also become more visibly segmented. Core AI Companion capabilities remained included in eligible paid Workplace plans while ZoomMate introduced a separately paid agentic layer with AI credits.
I read that as compatibility evidence, not as proof of a magic cutoff.
It shows that useful core AI can be bundled while usage grows and the wider company maintains strong margins. It does not tell us Zoom's current inference cost per paid seat, the cost of a meeting summary, the shape of its heaviest accounts, or the contribution margin of AI Companion itself. Company gross margin is not feature-level AI margin, and the public evidence does not establish that cheaper inference caused the bundling decision.
That limit matters because it is exactly where the attractive story becomes too neat. We can see a viable packaging pattern from the outside. We cannot see the private threshold that made it viable.
The same product can rationally bundle and meter AI
The clearer pattern appears when the comparison moves from companies to workload classes.
Linear includes some AI capability while using AI Credits for coding sessions and Loops. Its documentation makes the source of variability unusually visible: model tokens and sandbox runtime can both affect the work.
GitHub Copilot uses another hybrid. Paid plans can include effectively unlimited use of narrower features such as code completions while broader AI usage is governed through premium requests or allowances. GitHub's billing documentation also makes the workload problem legible: a long agent session that touches more context and performs more work can consume more than a quick request.
Notion's Custom Agents use credits for a workload that can vary with how much an agent reads, decides, does, and how often it runs.
v0 is a useful counterexample to any simple "AI is becoming too cheap to meter" story. Its generations remain credit-governed, with consumption tied to model and token use.
None of these products establishes a universal rule that core AI should be bundled and agentic AI should be metered. The useful pattern is narrower: products keep more explicit controls around workloads whose execution depth, frequency, context, or cost distribution can vary enough to matter.
That is a packaging problem.
The five things I would want modeled before bundling
If I had to make the packaging call, I would want five things modeled before arguing about a provider's token price.
-
How bounded is the workload? A meeting summary, one suggested edit, or another tightly scoped job is easier to model than an open-ended agent that can gather context, call tools, retry, search, execute code, and keep going. The point is not that one class is better. It is that one makes an unlimited promise easier to price.
-
What actually creates cost variance? Tokens may matter, but so can model choice, runtime, retrieval, storage, search, third-party APIs, and tool execution. The product needs the distribution of the real workload, not the cheapest advertised model rate.
-
What does the heavy-use tail do to the plan? A low average can still be uncomfortable if a small set of accounts creates a very wide cost tail. Internally, I would want the expected cost per subscriber or accepted unit of value, plus the high-percentile cases that can redefine the economics of an "unlimited" promise.
-
What does inclusion do for the product? AI can be worth bundling even when it is not negligible in cost if inclusion makes the core subscription materially more useful. The inverse is also possible: a company may separately price an advanced capability because it is a distinct value surface, not simply because its inference bill is higher.
-
Does the workload need a budget-control surface? Credits, limits, and allowances can make variable consumption visible and governable. They can provide a spend ceiling or contain the tail without forcing every ordinary interaction into pay-per-use pricing.
This is why I would not read a credit system backwards into vendor COGS. Customer-facing credits are part of the product contract. They are not a receipt for the provider's internal inference bill.
Packaging is a spectrum
Once the workload is the unit of analysis, "bundled or metered" stops looking like a binary choice.
Included core AI makes sense when the workload is bounded enough, the cost distribution is understood well enough, and inclusion strengthens the main subscription enough that per-use pricing is unnecessary.
Subscription plus an included allowance fits work that is valuable enough to belong in the plan but variable enough that the product still needs a quota, credit pool, or overage mechanism for the heavy-use tail.
Separately metered or separately priced AI remains reasonable when the work is highly variable, compute-intensive, agentic, independently valuable, or difficult to absorb without making a small number of workloads dominate the plan economics.
The important part is that all three positions can coexist inside one product. That is not necessarily an awkward transition state. It can be the actual design.
The product decision is which workload belongs in which economic contract.
What the public evidence cannot answer
There is still a private calculation underneath all of this.
Public packaging pages can tell us what a vendor includes, meters, limits, or sells separately. Company reporting can sometimes tell us whether overall margins are healthy and whether management says AI cost optimization matters. Product documentation can expose which workload dimensions affect consumption.
None of that gives an outside observer the private feature-level cutoff.
It does not tell us the exact inference cost per seat, the true internal COGS of a customer credit, the distribution of expensive accounts, or the contribution margin of one AI surface. It also does not establish vendor motive. A separately priced agent may reflect cost variance, product segmentation, willingness to pay, budget control, or several of those at once.
So I would be careful with both directions of the argument.
Bundling does not prove that inference became negligible. Metering does not prove that inference is expensive.
The public evidence is much better at showing the shape of the packaging decision than the private number underneath it.
The boundary should be allowed to move
The useful threshold is not permanent because none of its inputs are permanent.
Model prices change. Routing and caching improve. Agent workloads expand. Customers use features differently from how product teams expect. A bounded feature can become more open-ended after one product update, while a workload that needed credits last year can become predictable enough to include later.
That means the packaging decision should be revisited as the workload changes rather than attached to one industry-wide token price.
I would treat "cheap enough to bundle" as a product-specific decision rule, not a market milestone. The question is whether this workload, with this distribution and this value, still needs usage pricing to do essential economic or control work.
For another workload in the same product, the answer can be different.
That is the boundary worth measuring.
// End of transmission. Draw the boundary deliberately. — ZYANE
