Insights / Commerce & SaaS · · 11 min read
Subscription plus usage pricing for AI SaaS
Why AI software products often need a subscription plus a usage component: the subscription pays for the service and predictable value, usage covers variable AI costs. How we design plans, set fair limits, choose usage units customers understand and avoid surprise bills.
Traditional software had a convenient property: once it was built, serving one more customer cost almost nothing. A flat monthly subscription worked because the heaviest user and the lightest user cost roughly the same to serve.
AI products break that assumption. Every briefing generated, every report written and every question answered calls a model, and every call has a cost. A customer who uses the product ten times as much costs roughly ten times as much to serve. A flat price that is comfortable for a typical customer can become a loss on a heavy one.
That is why many AI software products, including several in the Oryvelon group, use subscription plus usage pricing: a recurring fee with an included allowance, and a clear way to pay for more. This article explains how we think about it — what the subscription should cover, how to pick a usage unit, how to set fair limits, and the mistakes that make usage pricing feel like a trap.
We do not publish specific prices here. Pricing belongs to each company and changes with its market. The principles are what carry across.
Why flat pricing breaks for AI products
Consider a simplified example. Imagine an AI product priced at a flat monthly fee. Most customers use it moderately and cost a comfortable fraction of that fee in model calls, hosting and support. A small group of customers use it constantly. Their model costs alone exceed what they pay.
Three things tend to happen:
- Margins depend on light users subsidising heavy ones. That works until the mix shifts, which it often does as the product improves and people use it more.
- The company quietly degrades service — cheaper models, shorter outputs, hidden throttling — to protect margins. Customers notice and trust erodes.
- Pricing becomes a guessing game in which every new feature that calls a model makes the flat price less sustainable.
None of this means flat pricing is always wrong. For a product with low, predictable AI usage per customer, a flat price with a sensible fair-use policy is often best. The point is that the choice should be made knowing the cost structure, not by copying traditional SaaS. We cover the cost side in cost discipline for AI products.
The structure: base, allowance, beyond
Subscription plus usage has three parts.
The base subscription pays for access to the service and the value that does not scale with use: the product itself, integrations, data storage, the dashboard, support, security and ongoing improvement.
The included allowance is a quantity of the usage unit bundled into each plan. It should be generous enough that most customers on that plan never think about it.
Beyond the allowance, customers have a clear option: upgrade to the next plan, buy a top-up, or pay a published per-unit rate. Which option fits depends on the product.
Plans usually step up in both base price and allowance. The goal is that each customer lands on the plan that matches their normal usage, pays a predictable amount most months, and has a fair path when a month is unusual.
Choosing the right usage unit
The most important decision in usage pricing is the unit. Get it right and customers understand their bill without thinking. Get it wrong and every invoice is a surprise.
A good usage unit is:
- Meaningful to the customer. It corresponds to something they value and can see.
- Predictable. Customers can estimate how many they will use.
- Roughly proportional to cost, so the company is not exposed to heavy losses on one type of use.
- Easy to count and display in the product itself.
Tokens fail most of these tests. They are an internal measure of model input and output. Customers cannot see them, cannot predict them and should not have to learn what they are. A bill that says "4.2 million tokens" tells a store owner nothing.
Better units usually map to the product's core output. Across Oryvelon's companies, the natural units differ:
| Company | Core value | Natural usage unit |
|---|---|---|
| MerchNivo | Daily operational help for a Shopify store | Connected stores, plus briefings, analyses or questions within an allowance |
| ZodiVela | Personal discovery readings | Credits per reading, or readings included in a subscription |
| KeşifAtlası | Visa and relocation eligibility | Free test, then a paid report per case |
| EduRelia | AI learning for students and schools | Seats or learners per school licence |
Each unit is something the customer can picture. A store owner knows how many stores they run. A ZodiVela user understands that a reading costs a credit. A family understands that a KeşifAtlası report covers one relocation case. The internal cost of each unit is modelled by the company; the customer sees the unit.
Sizing the allowance
The included allowance is where fairness is won or lost.
Start from real or carefully estimated usage. For each plan, look at what a typical customer on that plan uses in a normal month, and set the allowance comfortably above it. The aim is that the large majority of customers never reach the allowance in an ordinary month.
Then check the economics at the allowance, not just at typical use. If a customer uses exactly their full allowance, does the plan still cover its costs with a reasonable margin? If not, either the allowance is too high or the base price is too low. This calculation belongs in each product's unit economics.
Finally, leave room for model costs to change. Model prices move, sometimes downwards, sometimes not. Routing some work to cheaper models, caching repeated work and trimming prompts can all lower cost per unit over time. We describe the first of those levers in model routing. When costs fall, a good company can raise allowances rather than cut prices, which customers generally appreciate more.
Fair limits, not traps
The difference between usage pricing that customers accept and usage pricing they resent is almost entirely about how limits behave. We hold ourselves to a few rules.
Limits are visible. Customers can see their current usage and remaining allowance at any time, in the product, in the unit they are billed in.
Warnings arrive early. A notice at a sensible point before the allowance runs out, and another close to it, with the options laid out. Nobody should discover a limit by hitting it.
No silent overages. If going beyond the allowance costs money, the customer chooses that — by setting a spending cap, approving a top-up or upgrading. A bill should never contain a charge the customer did not see coming.
The core service keeps working. When an allowance runs out, the product degrades gracefully rather than switching off. For MerchNivo, that might mean the daily briefing continues in a simpler form built directly from calculated store data, while additional on-demand analyses wait for a top-up or the next cycle. The owner still sees what needs attention today. See AI fallbacks and degraded modes.
Spending caps are available. Customers who want certainty can set a maximum monthly spend and know it will not be exceeded.
Limits reset predictably, on a stated date, with no confusing rolling windows.
Budgets on our side too
Usage pricing protects the customer from surprise bills. Budgets protect the company from surprise costs.
Every Oryvelon product that calls models does so through the group's shared AI gateway, which enforces budgets per product and, where needed, per customer. If a bug or a misuse pattern starts generating unexpected calls — a loop that regenerates the same briefing, a script hammering an endpoint — the budget catches it before the invoice does.
This matters for pricing because a usage allowance only makes sense if the company knows its cost per unit. Observability at the gateway gives each company that number, broken down by feature and model, so plans can be designed from facts rather than guesses.
Choosing between upgrade, top-up and pay-as-you-go
When a customer exceeds their allowance, there are three common paths. Each suits different products.
Upgrade to the next plan. Best when exceeding the allowance signals that the customer's normal usage has grown. A store that has added a second shop or doubled its order volume probably belongs on a higher plan permanently.
Top-up packs. Best when usage is lumpy. A ZodiVela user might want extra readings around a particular moment without changing their subscription. Credits bought in advance are clear and self-limiting.
Metered overage at a published rate. Best for business customers who prefer continuity over caps, and who have agreed to it with a cap in place. It needs the clearest communication of all.
Many products offer more than one. The important thing is that the customer chooses, with the cost visible, before it is incurred.
Worked example: an AI briefing product
Imagine a hypothetical AI product that sends store owners a daily briefing and answers operational questions on demand. Here is how we would reason about its pricing, without real numbers.
- Identify the cost drivers. The daily briefing runs once per store per day, a predictable cost. On-demand questions vary widely between customers. Deeper analyses — a monthly product review, say — cost more per request.
- Separate predictable from variable. The daily briefing is part of the base subscription for each connected store. It is the core value and its cost is stable.
- Choose the usage unit for the variable part. On-demand questions and analyses are counted as a single understandable unit, with deeper analyses counting as more than one, clearly labelled.
- Size allowances per plan so that a typical store on each plan stays within them.
- Design the limit behaviour. When the allowance is used, the daily briefing continues; on-demand requests show remaining allowance, then offer a top-up or upgrade.
- Check the worst case. A store at full allowance every month should still be a profitable customer on its plan.
The result is a price the customer can predict, with the variable cost controlled and the core service protected.
Consumer and business customers need different shapes
The same model does not fit every audience. Consumers and businesses think about spending differently, and the pricing shape should respect that.
Consumers want certainty and small decisions. A person using ZodiVela for a reflective reading does not want to monitor a meter. Credits bought in advance, or a subscription with a clear number of included readings, give them a fixed, visible cost. Metered overage is almost always wrong for consumers: it turns a light, personal product into something that feels like a utility bill. Consumer products should also make it easy to stop, pause or cancel, because trust at the exit is part of trust at the start.
Businesses want predictability and room to grow. A store owner using MerchNivo is making an operating decision. They care that the cost is stable month to month, that it scales sensibly as the business grows, and that nothing unexpected appears on the invoice. Plans tied to connected stores, with generous allowances and optional caps, fit that mindset.
Institutions want simple annual terms. For school licensing at EduRelia, usage is usually expressed as seats or learners for a period, agreed in advance. Schools budget annually; a usage meter would make purchasing harder rather than fairer. The variable AI cost still exists, so it is modelled inside the per-seat price with fair-use protections rather than exposed as a meter. We cover that model in B2B2C licensing for schools.
What to measure after launch
A pricing model is a hypothesis until customers use it. After launch, a handful of measures tell us whether it is working:
- Share of customers reaching their allowance each month, per plan. If it is high, the allowance is too low or customers are on the wrong plan.
- Cost per usage unit, tracked over time and by feature, so we notice when a new feature makes a unit more expensive.
- Gross margin per plan at typical and full usage, to confirm that no plan loses money on legitimate heavy users.
- Upgrade and top-up behaviour after limit warnings: do customers upgrade, buy more, or leave?
- Support contacts about billing, which are the clearest signal that something is confusing.
- Cancellations that mention price, read individually rather than counted only.
These numbers feed each company's regular review. They let us adjust allowances, add a plan or change the unit before a pricing problem becomes a retention problem.
Communicating price honestly
Usage pricing has a reputation problem, mostly earned by products that hide the details. Clear communication fixes most of it:
- Show the allowance and unit on the pricing page, in plain language.
- Give examples: "a typical store with daily briefings and occasional questions fits comfortably in this plan."
- Explain what happens at the limit before the customer signs up.
- Keep the invoice readable, with usage in the same unit as the pricing page.
- Announce changes in advance, and honour existing terms until the next renewal.
For consumer products especially, trust at the moment of payment carries over to trust in the product. We discuss the free-to-paid step in from free test to paid report and the recurring side in recurring revenue loops.
Common mistakes
- Pricing by tokens. Accurate for the company, meaningless to the customer.
- Allowances set at average usage. Half of customers would hit the limit every month.
- Hard shutdowns at the limit. The product disappears exactly when the customer is using it most.
- Silent overages. The fastest way to lose a customer and earn a dispute.
- Degrading quality instead of pricing honestly. Quietly swapping to a weaker model to protect margin erodes trust faster than a clear price change.
- Ignoring the worst case. A plan that loses money on its heaviest legitimate users is not a plan.
- Too many units. One or two usage measures per product; more than that and nobody can estimate a bill.
How this fits the group
Each Oryvelon company owns its pricing. There is no group-wide price list and no bundle that ties one company's plan to another's, because every company needs to stand alone and customers of one should never be pushed towards another. What the group shares is the method: understand cost per unit, choose a customer-meaningful unit, set generous allowances, keep limits visible and protect the core service.
Pricing decisions are revisited at each company's continue, stop or scale checkpoints, alongside unit economics, so a plan that looked right at launch is checked against what customers actually do.
Summary
AI software products carry a real variable cost per action, so a flat subscription alone can make heavy users unprofitable or push companies into quietly degrading service. Subscription plus usage pricing solves this with a base fee for the service and its stable value, an included allowance sized so most customers never reach it, and a clear path beyond it through upgrades, top-ups or capped overage. The usage unit should be something customers understand — stores, briefings, readings, reports or seats — never tokens. Limits must be visible, warned early and never produce silent charges, and when an allowance runs out the core service should keep working in a simpler form. On our side, gateway budgets and per-feature cost data make the economics knowable, so plans are built from facts and revisited at every checkpoint.
Questions and answers
What is subscription plus usage pricing?
It is a pricing model where customers pay a recurring fee that includes a set allowance, plus an additional charge or upgrade path for usage beyond that allowance. It suits products whose costs rise with use.
Why do AI SaaS products need usage-based pricing?
Because every AI request has a real cost that varies with how much a customer uses the product. A usage component keeps the heaviest users from making the whole plan unprofitable while keeping light users' bills simple.
Should AI products charge customers per token?
Usually not. Tokens are an internal cost measure most customers cannot predict. It is clearer to charge for units customers understand, such as reports, briefings, readings or connected stores.