Thomson Reuters Built Its Own LLM for $40 Million. On Legal Work, It Beats GPT-5.4.
Thomson Reuters' $40 million legal model is a case study in how proprietary data, expert alignment, and an open-weight base can outperform general-purpose systems on a tightly defined professional task.
For two decades, the conventional wisdom about large language models held roughly constant: bigger wins. More parameters, more compute, more data. The biggest models, from the biggest labs, with the biggest funding rounds, were supposed to own the market.
Thomson Reuters just complicated that story.
On August 24, the legal publishing giant launched Thomson 1.0, its proprietary large language model trained on the company's accumulated knowledge: more than 40,000 databases from Westlaw, decades of Practical Law guidance, Checkpoint tax research, and Reuters news content. The company says the effort cost roughly $40 million over two years, with the final training run costing less than $450,000.
On the tasks that matter to legal professionals, Thomson Reuters says the model does not just hold its own. Its internal benchmarks show stronger completeness and factuality than GPT-5.4 and Claude Sonnet 5 when the systems are connected to Thomson Reuters content. Independent verification is still developing, and the company has promised a technical report, but its framing is notably careful: broadly competitive with frontier models on web-only access, roughly equal or slightly better when plugged into TR data, and meaningfully ahead on document tasks where citations and authority matter most.
"Thomson needs to set the frontier of intelligence for legal," said Joel Hron, Thomson Reuters' chief technology officer.
Built on Qwen, Powered by Decades of Law
Thomson 1.0's most recent foundation is Qwen 3.5, the open-weight model from Alibaba Cloud. TR did not build from scratch. Instead, the team used a layered specialization approach: pre-training on proprietary content, targeted post-training guided by hundreds of subject-matter expert lawyers and tax professionals, then reinforcement learning that taught the model to navigate Westlaw and Practical Law tool interfaces directly.
That last step may be the most important. Westlaw's more than 40,000 databases span generations of legal publishing, and knowing how to retrieve from them is a skill that general-purpose models trained on public internet data fundamentally lack. TR's model does not just know legal content; it has been trained to navigate the retrieval system that organizes that content for professional use.
In academic testing cited by Thomson Reuters, the model's responses included links to treatises, making them more transparent and useful. Less than 10% of TR's available proprietary content has been used for training so far, according to the company.
CoCounsel Goes Agentic — on Claude
The architecture of TR's CoCounsel Legal AI platform shows how enterprise AI is actually being built in 2026. CoCounsel — the agentic layer that handles planning, multi-step reasoning, and workflow execution — is built on Anthropic's Claude Agent SDK. Thomson will be the default model for specific tasks within that system, beginning with Tabular Analysis, the high-volume document-review function where structured extraction and factual accuracy matter most.
The layers work like this: Claude Agent SDK orchestrates the workflow, understands intent, and manages the back-and-forth with users. Thomson handles legal-domain work — lookups, citations, and synthesis of dense primary sources — where proprietary training data gives it an edge.
This is not a single-model story. It is a layered architecture in which one vendor's agentic framework works alongside another organization's domain intelligence. CoCounsel administrators can configure alternative models for specific tasks, but the default bet is on TR's own model where its content advantage matters most.
A New Business Model: API Licensing
TR has not announced standalone pricing or availability for Thomson. The company says it has begun early conversations with large law firms and corporate legal departments about direct access and fine-tuning capabilities.
A smaller open-weight version is also planned for Hugging Face under a noncommercial academic license, a nod to the research and law-school communities that help define what legal AI tools need to do well. Customer data is not used for model training; Thomson Reuters says the model was built on its own proprietary and licensed content with expert guidance throughout.
The Broader Pattern
Thomson Reuters is not the first enterprise to pursue a domain-specific model strategy. Bloomberg built BloombergGPT, and Adobe trained Firefly on licensed image content. What has changed is that open-weight base models have become capable enough for a well-resourced enterprise to start from a strong foundation and achieve frontier-level performance within its domain without building a base model from scratch.
The reported $450,000 final training run is the number enterprise technology leaders will circle. It is within the budget of a major organization sitting on a large, clean proprietary data lake: legal publishers, financial-data providers, medical-record systems, and scientific publishers. The three ingredients TR assembled — a strong open-weight base, deep proprietary domain data, and expert-guided alignment — are potentially replicable by organizations with comparable assets.
Whether that matters depends on what you are building for. For tasks where specialized knowledge, citation accuracy, and domain-specific tool navigation matter, the TR playbook is now well documented. For general-purpose tasks or work requiring broad world knowledge, frontier models still hold the advantage. TR's competitive-on-the-web framing acknowledges that distinction.
The conventional wisdom is not wrong: bigger still wins in many domains. But Thomson is a signal that bigger is increasingly measured not in total parameter count, but in depth of relevant expertise. For a company that has spent generations accumulating exactly that expertise, the reframing is a considerable advantage.
Sources
Thomson Reuters — Thomson: a purpose-built foundation model for professionals: https://www.thomsonreuters.com/en/thomson-llm
Thomson Reuters — Next-generation CoCounsel Legal launch: https://www.thomsonreuters.com/en/press-releases/2026/august/thomson-reuters-launches-next-generation-of-cocounsel-legal-the-ai-ecosystem-built-for-legal-professionals
LawSites — Thomson Reuters launches Thomson: https://www.lawnext.com/2026/08/thomson-reuters-launches-thomson-its-own-proprietary-llm-trained-on-westlaw-and-practical-law-content.html