
The Model Is No Longer the Moat: Why the Agent Decides Legal AI Quality
For standard contract work, the question "which AI model does it use?" has stopped being the useful one. The leading models from the major providers now perform close enough to each other that the difference an in-house legal team feels comes from what is built around the model: how much of the contract and its history the tool can hold, whether it knows your positions, and whether it does the work or only suggests it.
That was the core argument Aku Pöllänen, CEO and co-founder of Bind, made in Bind's live session on 24 September 2026, where he demonstrated Bind's agent drafting, reviewing and negotiating contracts for in-house teams. This article sets out that argument and what it means for teams choosing legal AI now.
The models have converged
The broad evidence is public. Stanford's 2025 AI Index reports that the score difference between the top and the tenth-ranked model fell from 11.9% to 5.4% in a year, and that the top two models were separated by just 0.7%. Open-weight models closed most of their gap with closed models over the same period, from 8% to 1.7% on some benchmarks.
Aku Pöllänen argued in Bind's live session on 24 September 2026 that the same pattern shows up in legal work specifically. For semi-standard contract drafting and negotiation, he said, the improvement end users get from each new model version is smaller than it was one and a half to two years ago, and the major providers and the best open models now handle this kind of work at a similar level. Basic legal reasoning is no longer the scarce part.
That has a practical consequence he made explicit: a team can get started with a general tool such as ChatGPT or Claude, and for occasional questions that is often enough. The gap opens when the work is high-volume and standardized, which is exactly the work in-house teams most want to take off their desks.
What now separates legal AI tools
If the model is roughly a constant, the variable is the agent: the software that decides what the model sees, what it is allowed to do, and what happens to the document. In the live session, three things carried that difference.
A fourth difference follows from those three: whether the tool acts. An assistant hands back suggestions that someone copies into Word. An agent edits the contract itself as tracked changes, writes the comments for the other side and keeps the round on record, while the lawyer decides every position. We explain the distinction in more detail in what a legal AI agent is.
What it means if you are buying legal AI
The convergence argument changes how an in-house team should run an evaluation.
Stop weighting the model name. Knowing which provider sits underneath tells you little about how the product will behave on your contracts, and vendors can and do switch models. Ask instead what the tool does with the model.
Test on your worst real contract, not a demo NDA. Use a long agreement with annexes and at least one term buried in a schedule. The context problem only shows up at length.
Run a negotiation, not a single review. Put the tool through three rounds. In one round, change something without tracking it, and in another, reverse a point you already conceded. See whether it notices.
Check your positions are used, and kept private. Load a playbook with a few real positions and confirm the tool applies them, and that your fallback positions do not leak into the comments the counterparty will read.
Look at the output, not the chat. A useful agent hands back the document, marked up and ready for you to read, not a list you still have to type in.
Our checklist for running this kind of test is in how to evaluate a legal AI agent.
What this looks like in practice
In the live session the difference was easiest to see in a multi-round negotiation. The contract becomes a tracked negotiation, each counterparty round is read against the previous one, and Bind prepares a response for the lawyer to approve step by step.
Step by step, with the names you will see in Bind:
- Actions → Start negotiation turns the contract into a negotiation, with the counterparty named and the playbooks for this contract type attached for every round.
- When the other side's redline arrives, Upload their redline adds it to the Timeline next to the original draft and every earlier round.
- Bind reads the round: what changed, including edits made without tracked changes, and whether the other side undid anything you changed before.
- A Review plan takes the lawyer through each change with a recommendation, its Severity and the reasoning. Nothing is edited until the lawyer decides.
- After Submit, Bind writes the response into the document as tracked changes, with comments for the other side that never reveal the playbook.
In-house legal teams at Atria, listed on Nasdaq Helsinki, and Outdoor Holding, listed on Nasdaq in the US, use Bind for this kind of work.
Sources
- Stanford HAI, The 2025 AI Index Report: model score gaps (11.9% to 5.4%, top two 0.7%, open vs closed 8% to 1.7%).
- Bind live session, 24 September 2026, presented by Aku Pöllänen. Statements attributed to the session are paraphrased, not quoted.
Ready to simplify your contracts?
See how Bind helps teams manage contracts from draft to signature in one platform.
Frequently asked questions
- Do different AI models give different results on contract work?
- Less than they used to. Stanford's 2025 AI Index found that the score gap between the top and the tenth-ranked model fell from 11.9% to 5.4% in a year, with the top two just 0.7% apart. For standard contract drafting and negotiation, Aku Pöllänen argued in Bind's live session on 24 September 2026 that newer model versions now bring end users smaller gains than they did one and a half to two years ago. Differences in output come increasingly from the software around the model.
- What is agentic tooling in legal AI?
- Agentic tooling is the software layer that turns a language model into a legal AI agent: it loads the whole contract package and the negotiation history, applies your templates and playbooks, plans the changes before making them, edits the document itself as tracked changes, and keeps the record across rounds. Two products running the same model can behave very differently because of this layer, which is why it is now the main thing to evaluate.
- Can in-house teams just use ChatGPT or Claude for contracts?
- For getting started and for one-off questions, yes: most general models handle basic legal reasoning well. The argument made in Bind's live session is that high-volume, standardized contract work is where general chat tools run out, because they do not hold your playbooks, your past contracts or a negotiation's full history, and they hand back suggestions rather than an edited document.
- What should in-house legal teams test when buying legal AI in 2026?
- Test the agent, not the model. Use a long real contract package with annexes, run a negotiation across at least three rounds including an edit the counterparty made without tracked changes, check whether your playbook positions are applied and kept out of the comments the other side sees, and check whether the tool edits the document itself or only suggests. Model benchmarks tell you little about any of these.
Bind is trusted by legal teams across Europe and the US

