Skip to main content
MirmiDecision memory
Back to blog

Mirmi analysis

The next LegalTech moat? Data.

Alex HAYEM

For the past two years, everyone has been talking about models: GPT, Claude, Mistral, proprietary models… But the battle is starting to shift.

Legora now wants to compete directly with Westlaw and LexisNexis in legal research. To fill gaps in its database, the company is even scanning printed case law collections when some decisions are not available in a usable digital format.

Thomson Reuters has made an even clearer bet: $40 million invested in training its own model, Thomson, using an open-source base enriched with decades of Westlaw, Practical Law, Reuters content and legal expertise.

Why invest so much in data when frontier models keep getting better every month? Because the model is becoming less scarce. Good data is still hard to copy.

Tomorrow, several LegalTech companies may use the same Claude model or the same open-source foundation model. What will differentiate them is what sits behind it: case law, commentary, taxonomies, precedents… and increasingly, their customers’ own data. That is probably where the market will start to split.

Incumbents have decades of enriched legal content. New entrants are building their own datasets. And a third battle is already emerging: who will best understand how each organisation actually applies the law?

For legal departments, this is not yet a daily issue. But it is a useful way to read where the market is going. In a few years, asking “Which LLM do you use?” may be much less interesting than asking: “What data do you have that others don’t?”. And more importantly: “What are you learning from my data while I use your product?”

The next competitive advantage in LegalTech may not be having a smarter AI. It may be having a data asset that competitors simply cannot rebuild.

This publication is based on an analysis first shared by Mirmi on LinkedIn.

View on LinkedIn