On 24 August 2026, Thomson Reuters announced it had trained its own large language model. The model is called Thomson. It is built for law, tax, news and the kind of work that surrounds them, and it has started appearing inside CoCounsel Legal for document review.
Most of the coverage picked up the word sovereignty. A software vendor that your firm probably already pays has decided it wants to own its own model rather than rent someone else's, and that is a genuinely interesting decision for the industry.
There is a more useful thing in the announcement, and it is a pair of numbers.
The two numbers
Thomson Reuters put the total development investment at $40 million. That figure covers talent, compute, staff, research, data curation and infrastructure, accumulated over a long period rather than spent in one go.
The final training run cost under $450,000 in GPU compute, over about three weeks.
Both numbers come from Thomson Reuters. Nobody has audited them. Take them as a company describing its own project, which is what they are. Even so, the distance between them is the part worth sitting with, because the training run is the part everybody pictures when they imagine a company building a model, and it accounts for roughly one percent of the money.
What the other ninety-nine percent bought

Thomson Reuters describes assembling a mid-training dataset of 200 billion tokens, selected out of a candidate pool of 19 trillion.
That is the whole story in one line. They looked at everything they had, and they used about one percent of it.
The rest was set aside. Superseded guidance, material that was correct when written and is wrong now, duplicates, drafts, content whose licensing made it unusable, and an enormous volume of perfectly good writing that simply was not authoritative enough to teach a model how to answer a tax question.
Somebody had to make each of those calls. That is what the money was for.
The base model, incidentally, was free to start from. Thomson built on Qwen3.5-397B-A17B, an open-weight model published by Alibaba. The foundation everyone treats as the hard part was the commodity in this project.
Why this matters to a firm that will never train anything
Your firm is not going to train a model. That is fine. The transferable part of this story has nothing to do with GPUs.
Your firm has a corpus. Memos, workpapers, engagement letters, the research file somebody built in 2019, four versions of the same policy, and a shared drive that everyone describes as a mess in the same resigned tone. Somewhere in it is the answer to most questions the firm gets asked twice.
Most of that corpus should never be used to answer a question, for exactly the reasons Thomson Reuters set aside 99% of theirs. It is out of date, or client-specific, or a draft, or it was one person's view and never became the firm's.
When a firm points an assistant at all of it, the assistant does not know the difference. It will answer from the superseded memo as readily as the current one, in the same calm tone, with the same formatting. There is no error message for having read the wrong file.
Deciding which of your material is authoritative is the work. It is unglamorous, it cannot be bought, and it is the same work Thomson Reuters paid tens of millions of dollars to do at a larger scale.
Owning the model and owning the advantage
The other useful thing in the announcement is what Thomson Reuters said about using it.
Joel Hron, the company's chief technology officer, put the strategy this way: "AI sovereignty is about owning the layers of the stack that matter to you, but ownership does not mean exclusivity."
On where Thomson actually gets used: "Thomson will be applied where its professional specialization, trust, validation and efficiency provide the strongest advantage. Other leading models may be used for capabilities where they are better suited."
Read that again with the $40 million in mind. A company that has just built its own model is saying plainly that it will keep sending work to other people's models when those are better for the job.
That is the correct engineering answer, and it is worth remembering the next time a vendor tells your firm that its proprietary model is the reason to buy. The model is one layer. Which layer matters depends on what you are trying to do.
On the benchmark claim
Thomson Reuters reports that Thomson-1.0-Large narrowly outperformed GPT-5.4 and Claude Sonnet 5 on completeness and factuality for tax, legal and news content.
Treat that as the vendor's own measurement, on the vendor's own subject matter, with no published scores, no named evaluation set and no independent replication. It may well be true. It is the kind of claim that deserves a follow-up question rather than a nod, and the follow-up question is what was in the evaluation set.
The question this leaves you with
If a company with Thomson Reuters' resources concluded that 99% of its own content was not fit to answer a professional question, it is worth asking what share of your firm's shared drive would survive the same review.
You do not need a model to start answering that. You need somebody to walk the folders and say, out loud, which documents the firm would stand behind today. Most firms have never done it, and it is the cheapest useful AI project available to them.
