Where your client file actually goes when you use an AI tool
Data security is the barrier firms name most often, and the answer is usually buried in a terms page. Here is what to look for, what the words mean, and the four questions that settle it.
Ask a room of attorneys what is stopping them adopting AI and data security comes back first. In the 8am 2026 legal industry survey, 46% named it as a significant barrier, ahead of ethics at 42% and trust in results at 39%.
It is the right worry. It is also a worry that almost nobody resolves, because resolving it means reading a terms page written to be skimmed, and the answer is usually three clicks further in.
Here is what is actually happening to a file, and the four questions that settle it.
Two companies, not one
Most legal AI products do not have their own model. They are built on one from OpenAI, Anthropic or Google.
That means your client file may pass through two organizations, and both sets of terms apply. Firms routinely read the vendor’s privacy policy with care and never ask whose model sits underneath it, which is the half where the file actually gets processed.
Ask. It is a fair question and any vendor who is uncomfortable answering it has told you something.
What the words mean
Three terms do most of the work, and they are not interchangeable.
Training. Whether your input is used to improve the model itself. On consumer tiers this is frequently the default. On business and API tiers it is usually not, and the major providers say so contractually. The word that matters is contractual: a reassurance in a marketing page is not a term you could enforce.
Retention. Whether a copy is kept after processing, and for how long. This is separate from training and is the one people miss. A provider may not train on your data and still hold it for a period for abuse monitoring. That is still a copy of your client’s medical history on a system you do not control.
Zero data retention. An arrangement where nothing is kept after the request completes. Several major providers offer it on business tiers, often only on request, which means the default is not it.
A firm that has confirmed “they do not train on our data” has answered one of three questions and may believe it has answered all three.
Does any of this waive privilege
The honest answer is that it depends on the arrangement, and that this is a question for your own malpractice carrier and your state’s rules rather than for a software vendor.
What can be said generally: disclosing privileged material to a third party can waive privilege, and vendors handling material under a confidentiality agreement as agents of the firm have long been treated differently from disclosure to an outside party. Firms have used outside copy services, records retrieval vendors and contract reviewers for decades on that basis.
Nothing about a tool being described as AI changes the analysis. What changes it is what the contract says and what the vendor actually does. A vendor that keeps your data indefinitely on shared infrastructure is a different proposition from one that never holds it at all, regardless of what either calls itself.
The question is not whether the tool is AI. It is how many organizations end up holding the file, and what each of them is contractually permitted to do with it.
The arrangement that avoids most of this
There is a structural answer that makes several of these questions disappear rather than answering them.
Run the software inside your own cloud account, on your own API keys, with the vendor given scoped access during the build and revoked at handover.
Then your file never sits in a vendor’s system. Your contract for model access is directly with the model provider, so you control the retention setting rather than inheriting somebody else’s. When the engagement ends, nothing needs deleting from anyone’s servers, because nothing was ever there.
This is not exotic and it is not more expensive in any way that matters. It is mostly a decision about whose account the bill lands in. The reason it is uncommon is that it is worse for the vendor: it removes the lock-in, and a customer whose data lives on your platform is harder to lose.
It also happens to be the arrangement that survives a client’s own security questionnaire, which matters more every year in commercial matters. It is the same discipline as never filing what you have not verified: decide the rule once, then the individual judgment calls stop mattering.
The four questions
Enough to settle it, and short enough to send in an email before a demo.
- Whose model is underneath, and what are their terms? Two companies, two contracts.
- Is our data used to train anything? In the contract, not on the website.
- How long is it retained after processing, and can that be set to zero? Retention is the one that gets missed.
- Where does it run, and whose account is it billed to? This decides how much of the rest matters.
A vendor who answers all four in writing is workable even if some answers are imperfect, because you can price the risk. A vendor who answers with a security badge and a paragraph about enterprise-grade encryption has not answered any of them.
What we do
Ours is the structural answer, because we think the others are harder to keep. Everything runs inside your AWS, Azure or Google tenant, billed to your accounts, on your own model provider contract configured for zero retention where the provider offers it. We get scoped access during the build and it is revoked at handover. Nothing you give us is used to train anything, and that is contractual rather than a promise on a page.
The part worth stating plainly: this arrangement is not us being generous. It is easier for us too, because it means we are never the custodian of a client’s medical history, and never the party who has to answer for it.
Questions we get asked
- Does using an AI tool waive attorney-client privilege?
- Disclosing privileged material to a third party can waive privilege, and whether a vendor counts depends on the arrangement. Vendors handling material under a confidentiality agreement, as agents of the firm, are generally treated like any other outside service. What matters is what the contract says and what the vendor actually does with the data, not whether the tool is described as AI.
- What is zero data retention?
- An arrangement where the model provider processes your input and keeps no copy afterward. Several major providers offer it on business tiers, often only on request. Without it, inputs may be stored for a period for abuse monitoring, which is a different thing from training but is still a copy of your file on somebody else's system.
- Is my data used to train the model?
- On consumer tiers, frequently yes by default. On business and API tiers, usually no, and the major providers state so contractually. The word to look for is contractual. A promise in a marketing page is not the same as a term you could enforce.
- What is the difference between the AI vendor and the model provider?
- Most legal AI products are built on a model from a company like OpenAI, Anthropic or Google. So your file may pass through two organizations, and both sets of terms apply. Firms often review the vendor's terms carefully and never ask whose model sits underneath.
- Where should the software actually run?
- The strongest arrangement is inside your own cloud account, on your own API keys, with the vendor given scoped access during the build and revoked at handover. Then your data never sits in a vendor's system at all, and ending the relationship does not require anyone to delete anything.
- What should be in the contract?
- Where data is processed and stored, how long it is kept, that it is not used to train any model, who may access it, what happens at termination, and notification obligations if there is a breach. If a vendor will not put those in writing, that is the answer.
Not ready to book a call
Send us five pages of a record set. We will send back what we found in it.
Five pages is enough, redacted however you like. You get a short video back within two business days showing what a chronology would surface from it: the dates, the gaps, the things worth knowing before the other side finds them.
- Send five pages of a real file. Redact whatever you like first.
- We run them and record what comes out, including what it misses.
- You get the video within two business days. If there is nothing worth showing, we say so.