WP-002 · Working Paper · October 2026
From SaaS to AI-Native Software
When Software Becomes a Producer: Pricing, Incentives and Scale under Variable Inference Costs
Abstract
Generative AI changes the production function of software: every answer, document or completed task consumes inference compute, so software acquires a positive marginal cost that scales with how intensively each customer uses it. I derive the consequences in three steps. First, flat access breaks. Heavy users become the least profitable, flat pricing’s share of attainable profit falls monotonically with inference cost, and it collapses once a unit of inference costs half the value of the first unit of use. Second, and centrally, the unit of sale follows control of compute. When the customer controls usage, as with copilots, usage pricing disciplines consumption and pass-through of inference cost exceeds Borch’s risk-sharing benchmark. When the provider controls the process, as with agents, usage pricing is cost-plus, and the efficient contract moves towards paying for outcomes. Outcome pricing is worth its verification cost when the compute bill per task is large. Third, with constant inference cost scale economies vanish; they survive only through firm-specific learning in inference. Common falls in compute prices raise margins without raising gross profit, and because AI also lowers the cost of building software, free entry pushes each firm’s gross profit towards its entry cost. Financial statements of 307 US-listed software firms are consistent with these predictions. Early generative-AI adopters saw gross margins fall four to five points relative to other firms after 2023 while growing faster, and market value tracks gross profit more closely than revenue. The learning rate in inference cannot be identified from public accounts.
Keywords software economics, generative AI, inference costs, nonlinear pricing, risk sharing, economies of scale, free entry, firm valuation
01The question
Traditional software sells access to an asset that costs almost nothing to reproduce. AI-native software sells work produced by a computation that costs something every time it runs.
The paper asks how that single change, a positive marginal cost of inference, alters what software firms charge for, who bears production risk, whether firms still enjoy economies of scale, and what it takes to reach a billion-dollar valuation.
02The cost shock
Conventional software has \(C(Q)=F+cQ\) with \(c\approx 0\). Generative AI adds a variable term that depends on model choice, context length and, for agents, the number of computational steps:
Customer heterogeneity becomes simultaneously demand heterogeneity and cost heterogeneity. The highest-revenue customer can be the least profitable one.
03Pricing as risk allocation
Under a flat subscription the supplier carries all usage variance, \(\operatorname{Var}(\pi_i)=c^2\operatorname{Var}(q_i)\). A two-part tariff \(P=F_0+pq\) separates two jobs: \(F_0\) recovers fixed costs and extracts surplus, while \(p\) allocates the marginal cost of inference.
- Fixed-price fragility. Higher dispersion of usage makes flat pricing riskier.
- Hybrid dominance. With \(F>0\), \(c>0\) and heterogeneous demand, hybrid contracts can weakly dominate both pure forms.
- Usage screening. When \(\operatorname{Cov}(V,q)>0\), menus screen on expected resource use as well as willingness to pay.
- Outcome pricing requires joint predictability of value and cost. The provider becomes the insurer of its own production process.
04The unit of sale
| Regime | Cost structure | Unit sold | Pricing |
|---|---|---|---|
| I · Traditional SaaS | Fixed | Seat | Subscription |
| II · AI-SaaS | Variable | Usage | Hybrid \(P_0+pq\) |
| III · AI-native | Task-dependent | Task / output | Per task |
| IV · Agentic | Outcome-dependent | Outcome | Share of value |
The transition is governed by marginal cost \(c\), usage dispersion \(\sigma_q\), the observability of value \(\alpha\) and the degree of labour substitutability \(s\). When AI performs a task previously done by a worker, the firm's allocation condition becomes \(PF_A>c_A\) against \(PF_L>w\). Software starts to be priced against labour.
With uncertain usage, pricing is risk allocation. When the customer controls consumption, as with a copilot, the optimal pass-through of inference cost lies above Borch's risk-sharing benchmark. When the provider controls the process, as with an agent, it lies below: usage pricing would be cost-plus and remove the incentive to economise, so outcome pricing wins.
05Do AI-native firms scale like software?
The ray economies-of-scale index, average cost over marginal cost, is the paper's central object:
With \(c\approx 0\), scale economies are effectively unbounded. With \(c>0\) they fade and the firm converges towards constant returns, like a services firm. Gross margin no longer improves with scale, and operating margin is capped at \(1-c/p\).
Scale returns only if the firm's own inference cost falls with volume, through routing, caching, distillation or its own models: \(c(Q)=c_0Q^{-b}\), so \(MC=(1-b)\,c(Q)\).
Cost-structure rotation. AI also lowers \(F\), the cost of building software, while raising \(c\). Break-even \(Q^*=F/(p-c)\) can fall, but so do entry barriers and the durability of any moat. The result is a paradox: it is easier to grow, and each dollar of revenue is worth less.
06The unicorn threshold
Separate the economics of scale (technology) from valuation (expectations). If markets capitalise gross profit with a multiple that depends on growth \(g\) and moat durability \(\delta\):
With exponential growth \(Q_t=Q_0e^{gt}\), the time to a $1bn valuation is \(T_U=\tfrac{1}{g}\ln(Q_U/Q_0)\). AI-native firms reach unicorn status faster only if their growth advantage outruns their higher threshold:
This race between growth and margin explains, within one condition, why some AI-native firms reach very large revenue unusually fast while markets decline to value them on classic SaaS multiples.
07Where economies of scale migrate
- Upstream. Foundation-model labs have the SaaS structure in the extreme: enormous training \(F\) and inference \(c\) falling with scale. Their price is the application layer's \(c\), with a risk of double marginalisation.
- Learning in inference. The firm-specific rate \(b\).
- Demand-side scale. Usage data that improves quality, and network effects.
- Distribution and workflow lock-in. The classic SaaS advantage, which survives.
The exit route is to anchor price to labour rather than to compute. With \(p=\alpha w\), gross margin becomes \(1-c/(\alpha w)\): it recovers as \(c\) falls while \(w\) does not, and the addressable market becomes labour spend rather than IT spend.
Make-or-buy completes the picture. A firm self-hosts when usage exceeds \(q^*=F_M/(p-c_M)\), internalising part of the upstream scale.
08Results
Cheaper to build, costlier to run, and worth less in a free-entry market.
- Flat pricing breaks down. Heavy users become the least profitable, flat pricing's share of attainable profit falls monotonically with inference cost, and it collapses once a unit of inference costs half the value of the first unit of use.
- Scale survives only through learning. The long-run scale index is 1/(1 − b), where b is the firm's own learning rate in inference. Common falls in compute prices raise reported margins without raising gross profit.
- Free entry: each firm's gross profit equals what it cost to build. When AI lowers that cost, the typical firm is worth less. Unicorns are exceptions with proprietary cost advantages or differentiation.
- Outcome pricing is worth verifying when the compute bill per task is large, the AI version of the procurement choice between cost-plus and fixed-price contracts.
| Evidence, 307 US-listed software firms | Estimate |
|---|---|
| Gross margin of early GenAI adopters vs. others, after 2023 | −4.6 pp** |
| Revenue growth of early adopters vs. others | +9.2 pp/yr** |
| Market value elasticity: gross profit vs. revenue (joint) | 0.72*** vs 0.41* |
| Elasticity of cost of revenue to revenue, before 2023 | 1.00 |
09Status
| Level | Where it stands |
|---|---|
| Model | Nine propositions and a corollary, proved and verified symbolically. |
| Numerics | Robustness over 5,000 parameter combinations and lognormal heterogeneity. |
| Evidence | SEC XBRL panel and 10-K text; descriptive. Working paper v1. |
Citation
Menéndez-Pidal, J. (2026). “From SaaS to AI-Native Software: When Software Becomes a Producer: Pricing, Incentives and Scale under Variable Inference Costs.” Project Frontier Working Paper No. 002, Madrid.
BibTeX
@techreport{menendezpidal2026saas,
author = {Men{\'e}ndez-Pidal, Jorge},
title = {From SaaS to AI-Native Software: When Software Becomes a Producer: Pricing, Incentives and Scale under Variable Inference Costs},
institution = {Project Frontier},
type = {Working Paper},
number = {002},
address = {Madrid},
year = {2026},
month = {oct}
}
Preliminary draft. Comments welcome; please do not cite without permission.