
Today’s legal AI pricing creates a dangerous illusion: that of a stable, predictable budget that is almost independent of usage.
That illusion will not survive AI agents.
Harvey was consuming around 1 trillion tokens per month in January 2026. By May, that figure had already reached 12 trillion tokens—or 144 trillion per year. A twelvefold increase in just a few months. (Business Insider)
The question is no longer whether prices will increase.
The question is: what will the multiplier be?
The figures should worry every LegalTech buyer
A calculation recently shared on LinkedIn by the excellent Raymond Blyd compares Harvey’s and Legora’s revenues with their estimated token consumption:
Harvey revenue: $2.08 per million tokens
Legora revenue: $1.39 per million tokens
VS likely real cost: an average of $15 per million tokens
The calculation is based on approximately $300 million in ARR for Harvey and $100 million for Legora. Harvey’s consumption comes from figures shared by its CEO, while Legora’s consumption is estimated at around 72 trillion tokens per year. (Sourcery)
For reference, OpenAI’s GPT-5.6 Sol is priced at $5 per million input tokens and $30 per million output tokens. Claude Fable 5 is priced at $10 for input and $50 for output. The average cost of $15 used by Raymond Blyd therefore relies on an assumed mix of input and output tokens, rather than an official flat rate.
If they had to charge the full economic cost merely to break even, your bill would increase by at least 7x. 🧨
This calculation is not an estimate of their margins. It does not account for model routing, caching, negotiated discounts, their own infrastructure costs or other revenue streams. It deliberately measures the maximum risk faced by a buyer who assumes that today’s prices will remain stable.
Even as an estimate, this stress test highlights the colossal gap between the prices currently charged per unit consumed and the real—and currently hidden—cost of those same units.
Even after correcting the calculation, the problem remains
Let us account for the fact that the largest LegalTech companies do not use the most expensive model for every request. Instead, they route each task to the model they consider most appropriate.
Using optimistic assumptions, we arrive at a much cheaper model mix, estimated at $7.80 per million tokens instead of $15.
The gap would still represent:
3.75 times Harvey’s estimated revenue
5.6 times Legora’s estimated revenue
The objection changes the multiplier. It does not change the conclusion: current prices probably do not yet reflect the mature economics of these products.
⏳ Business models have already started changing—and you will pay
Investors are currently funding explosive usage growth while vendors embed their platforms at the heart of legal work.
But once users—you—have become dependent on them, unlimited usage disappears. Quotas, credits, premium agents, overage charges and consumption-based pricing take over.
And this is no longer a hypothetical scenario.
Legora has officially announced that its Agent Pro product is moving to consumption-based pricing.
The company is effectively acknowledging that per-user pricing no longer works when agents perform tasks with highly variable complexity and duration. (Legora)
The shift has begun.
💸 Your €100,000 contract tomorrow: €375,000, €560,000… or more than €1 million?
Imagine a legal department currently spending €100,000 per year on an AI platform.
If the provider simply passed on the gap from the corrected $7.80 scenario, the same level of consumption would economically represent:
- approximately €375,000 based on Harvey’s ratio;
- approximately €560,000 based on Legora’s ratio.
Under the initial $15 assumption, the equivalent bill would exceed:
- €720,000 for Harvey;
- €1.1 million for Legora.
We are using these two LegalTech companies as examples, but the same calculation can be applied to any LegalTech or CRM provider whose legal AI relies on APIs from the largest frontier models.
This is not a prediction of these companies’ future pricing. They may reduce their costs, improve their margins or absorb part of the gap.
But it illustrates the scale of the risk taken by a customer who builds its entire strategy around these extremely powerful tools and assumes their prices will remain stable.
When growth subsidies disappear, prepare for the hangover. 😵
And by then, changing platforms will be much harder: teams will have been trained, workflows integrated and data centralised with the provider.
The alternative: smaller, specialised, internal and controlled AI
Not every legal task requires the most powerful model on the market, accessible only through an API.
Finding a clause, classifying a contract, extracting a deadline or applying an internal rule can be handled by a smaller, specialised AI system.
This is known as the Small AI approach: recalibrating AI to match real needs and stopping the use of a bazooka to kill mosquitoes.
🧠 A smaller, specialised AI can sometimes perform better using only a fraction of the resources.
It understands its specific domain. It delivers more predictable results. It can run on open models within European or sovereign infrastructure. Most importantly, its cost does not explode every time a new generation of American models is released.
This approach makes it possible to rely on LegalTech providers that develop their own AI internally—for example, using open-source model foundations—and therefore retain genuine control over their pricing and independence from OpenAI and similar providers.
This is the approach chosen by Mirmi AI : an AI specialised in understanding the context behind legal approvals, capable of running on controlled infrastructure without depending on prohibitively expensive external APIs.
By adapting computing power to the task, Mirmi AI’s architecture allows it to support AI costs that are 10 to 25 times lower than traditional LegalTech solutions, depending on the request.
There is also an environmental bonus: Small AI is far more sustainable for the planet. 🌍
In Mirmi AI’s case, measured water consumption is 8 to 12 times lower per query, with 3 to 6 times fewer CO₂ emissions than a standard GPT-4o query used as the benchmark. 🌿
Ultimately, the goal is not to consume the largest number of tokens using the biggest models available.
It is to produce the right decision with the minimum amount of computing power required.
Questions to ask before signing
Before committing to your next tool, your legal department should ask:
- Which uses will still be included in two years?
- Will agents be charged separately?
- Are there hidden quotas or a fair-use policy?
- Can the provider significantly increase its prices?
- Does it systematically rely on frontier models?
- Is there a specialised or sovereign alternative?
- How can you recover your data and workflows if the bill gets out of control?
Buyers who ask these questions now will remain in control. 💰😎
Everyone else will discover the true price of their LegalTech once they can no longer operate without it. 🔐🤯
This publication is based on an analysis first shared by Mirmi on LinkedIn.
View on LinkedIn