---
title: "What Actually Happens to Your Data When You Use an AI Tool"
description: "The question is rarely whether AI is safe. It is which account the data went through, what the vendor's terms actually say, and whether your own permissions were ever correct. A plain guide for operators, with the policies quoted."
slug: "what-happens-to-your-data-in-ai-tools"
publishedDate: "2026-07-31"
readingTime: "12"
category: "Guides"
tags: ["ai-policy", "data-privacy", "vendor-management", "operations", "security"]
---

# What Actually Happens to Your Data When You Use an AI Tool

Somebody in your company pasted a client contract into a chatbot last week. Probably several somebodies, probably several documents. This is not a hypothetical, and it is not a discipline problem. The tools are useful, they are one browser tab away, and nobody told anyone where the line was.

The usual response to this is a memo banning AI tools, which reliably produces the same result as banning personal phones did: the behavior moves out of sight. The better response starts with a question most operators have never had answered plainly.

Where does the data actually go?

The honest answer is that it depends almost entirely on which account it went through, and barely at all on which model answered. That distinction is the whole subject. Once you understand it, a workable policy takes an afternoon to write instead of a quarter to argue about.

## The account matters more than the model

The same underlying model can sit behind a free consumer chatbot and behind a paid business tool, and the terms governing your data are completely different in each case. Same model, same answers, different contract.

Here is what the major vendors say, as published on their own pages, quoted rather than summarized.

On the business side, Anthropic's commercial terms state that "Anthropic may not train models on Customer Content from Services," and that the customer "(a) retains all rights to its Inputs, and (b) owns its Outputs" ([Commercial Terms of Service](https://www.anthropic.com/legal/commercial-terms)). OpenAI's developer documentation says that "data sent to the OpenAI API is not used to train or improve OpenAI models (unless you explicitly opt in to share data with us)" ([API data controls](https://developers.openai.com/api/docs/guides/your-data)). Microsoft states that with Microsoft 365 Copilot, "prompts, responses, and data accessed through Microsoft Graph aren't used to train foundation LLMs" ([Data, Privacy, and Security for Microsoft 365 Copilot](https://learn.microsoft.com/en-us/copilot/microsoft-365/microsoft-365-copilot-privacy)). Google says of Gemini in Workspace that "Workspace does not use customer data for training models without customer's prior permission or instruction," and that content "is not human reviewed or otherwise used for Generative AI model training outside your domain without permission" ([Generative AI in Google Workspace Privacy Hub](https://knowledge.workspace.google.com/admin/gemini/generative-ai-in-google-workspace-privacy-hub)).

Now the consumer side. Anthropic's consumer privacy page describes training on chats as something that happens when "you choose to allow us to use your chats and coding sessions to improve Claude," a setting on personal plans that does not exist in the same form on the commercial products ([consumer data and model training](https://privacy.claude.com/en/articles/10023580-is-my-data-used-for-model-training)). Consumer tiers across the industry generally carry some version of this switch, sometimes on by default, sometimes off, changeable by the vendor with notice.

So the practical rule is not "do not use AI with company data." It is closer to: company data goes through company accounts. An employee using a personal free account is operating under a different contract than the one your business signed, and no policy language you write changes that.

This is also the cheapest problem on the list to fix. Buying business seats is a line item. Explaining after the fact why a client's numbers went through somebody's personal login is not.

## What "we do not train on your data" does not cover

Training is the headline commitment and the narrowest one. Four other things happen to your data that the training sentence says nothing about, and each is worth checking rather than assuming.

### Retention

Content that is never used for training may still be stored for a while. OpenAI's documentation notes that "abuse monitoring logs are generated for all API feature usage and retained for up to 30 days," with a zero-retention arrangement available to approved organizations that "excludes customer content from abuse monitoring logs" ([API data controls](https://developers.openai.com/api/docs/guides/your-data)). Microsoft stores Copilot interaction history and states that admins can apply retention policies to it through Purview, while noting that the stored data "is encrypted while it's stored and isn't used to train foundation LLMs" ([Microsoft 365 Copilot privacy](https://learn.microsoft.com/en-us/copilot/microsoft-365/microsoft-365-copilot-privacy)).

Neither of these is alarming. Both are facts you would want to know before telling a client their information is never stored anywhere, because that sentence is usually wrong and is easy to avoid saying.

### Human review

Some vendors reserve the right to have humans look at flagged content for safety purposes. Some have turned that off for specific products. Microsoft states that "while abuse monitoring, which includes human review of content, is available in Azure OpenAI, Microsoft 365 Copilot services have opted out of it." Google states that Workspace content "is not human reviewed" for training outside your domain without permission. These are product-level decisions, not company-level ones, so the answer for one tool from a vendor does not carry to another tool from the same vendor.

### Subprocessors

The company on your invoice is often not the only company touching the request. Microsoft's own documentation points to separate pages for "Anthropic models in Microsoft Online Services" and "OpenAI as a subprocessor," and notes that models provided by Anthropic as a subprocessor "are currently excluded from the EU Data Boundary." If you have data residency obligations, that sentence is the kind of detail that decides whether a configuration is compliant, and it lives three clicks deep in a vendor doc rather than in the sales deck.

### The feedback button

Thumbs up and thumbs down are usually a separate data path with separate terms. Microsoft says it may use customer feedback to improve the product, that this feedback is not used to train the foundation models, and that admins can manage it. That is a reasonable arrangement, and it is still worth knowing that clicking thumbs-down on a response can send that response somewhere the response itself would not have gone.

None of this argues against using these tools. It argues for reading four paragraphs before signing, and for describing your own posture to clients accurately rather than generously.

## The larger risk is usually inside your own building

Here is the part that surprises people. In most companies the vendor is not the weak point. The weak point is that the AI tool works exactly as designed and surfaces things nobody realized were reachable.

Both Microsoft and Google are explicit that their tools inherit your existing access controls. Microsoft: "Microsoft 365 Copilot only surfaces organizational data to which individual users have at least view permissions," followed immediately by the warning that "it's important that you're using the permission models available in Microsoft 365 services, such as SharePoint, to help ensure the right users or groups have the right access to the right content." Google puts it the same way: Gemini "abides by your organization's existing controls and data handling practices."

Read that as what it is. The tool is not going to leak anything. It is going to be extremely good at finding whatever your permissions already allow.

Ten years of accumulated file sharing sits in most organizations. A folder shared with everyone in 2019 for one meeting. A drive where the salary review spreadsheet lives one level above the folder that was opened to the whole team. Under the old regime this was theoretically exposed and practically invisible, because finding it required knowing it existed. A good retrieval system removes that protection completely, and it does it on day one, in front of whoever asked an innocent question.

This is the single most common unpleasant surprise in a first-week rollout, and it is not really an AI problem. It is a permissions problem that AI made visible. Which means the fix is old, boring, and worth doing regardless: audit what is shared with everyone, tighten the obvious cases, and stage the rollout so a small group finds the surprises before the whole company does.

We have written before about [running a pilot that survives contact with real operations](https://elorati.com/blog/running-an-ai-pilot/). This is one of the specific reasons the staged version wins.

## A policy that fits on one page

Most AI policies fail because they are written as prohibitions and read as suggestions. The version that works sorts information into a small number of buckets and names the tool for each. Three tiers is usually enough.

**Open.** Information you would put in a public post or hand to a stranger: marketing copy, public pricing, general questions, anything already published. Any tool, any account. There is no reason to make people think hard about this tier, and pretending otherwise trains them to ignore the whole policy.

**Internal.** The daily material of the business: client work, internal documents, financials, drafts, operational data. Company accounts only, on the business tier, in tools you have actually reviewed. This is where most work happens and where most policies are silent, which is exactly why people improvise.

**Restricted.** Regulated or contractually protected information: health records, payment card data, anything covered by a client confidentiality clause, anything a regulator has an opinion about. Named tools with the paperwork in place, or not at all. Nothing goes into this tier by default, and no one gets to move something into it informally.

Two rules make the tiers hold. First, the list of approved tools is short, named, and current, because "approved tools" without a list means every employee is guessing. Second, there is a fast, blameless way to ask about something that does not obviously fit, and a fast way to add a tool. A policy with no intake path becomes a policy people route around, and the routing around is the actual risk.

Write it on one page. Longer documents do not get read, and an unread policy provides exactly as much protection as no policy.

## Four questions for any vendor

The vendor conversation is short if you know what to ask. Most of the surface area is covered by four questions, and the manner of the answer tells you as much as the content.

**Is our content used to train your models, and where is that written?** You want a link, not an assurance. If the answer is a sentence in an email rather than a clause in the terms, it is not a commitment, it is a hope.

**How long do you keep it, and can we shorten that?** Ask about both the product data and any abuse or safety logs. Retention windows are often configurable at the business tier and almost never at the consumer tier.

**Who else touches the request?** Which model providers, which cloud regions, which subprocessors. This is the question that catches the small tool built on top of a big model, where the wrapper's terms and the underlying provider's terms are two different documents and only one of them was shown to you.

**What happens when we leave?** Export path, deletion timeline, what remains. Ownership of your own material should be explicit rather than implied, which is the same principle that applies to [everything else built for you](https://elorati.com/blog/who-owns-your-custom-software/).

A vendor who has done this before answers all four in a few minutes and sends links. A vendor who has not will answer with reassurance, and the reassurance itself is the finding.

## When regulation changes the answer

If you handle health information, payment data, or anything a client contract specifically restricts, the analysis is different in one important way: the general terms are not sufficient, and the specific agreement is what matters. That usually means a signed data processing agreement, and in health care a business associate agreement with each vendor that will touch the data. Some vendors sign these for specific products only.

This is also where the shape of the deployment starts to matter. Running a model inside your own cloud tenant, or on hardware you control, changes the compliance conversation more than any setting in a consumer app will. That is a real option, and it costs real money, and for most operators it is not necessary. But it exists, and knowing it exists keeps the choice from being framed as "use the public tool or do nothing."

Get an opinion from someone qualified for your specific obligations rather than from a blog post. The useful thing we can say here is narrow: the tier of tool and the paperwork behind it are the variables that move, and both are decidable before anyone starts using anything.

## How we handle it

Elorati / Advisory work regularly starts here, because a company that is unsure where its data goes cannot make a confident decision about anything downstream. The sequence we tend to follow is unglamorous and short. Find out what people are already using, without treating it as an investigation, because the honest inventory is worth more than the tidy one. Get the business-tier accounts in place so the terms match the work. Look at what the retrieval tools can actually reach before rolling them out widely. Write the one-page tiering. Then build.

Custom systems we build get the same treatment, with the model provider accounts in your name so you can see the usage and the terms apply to you directly. That is the same principle as the rest of the [Elorati / Studio](https://elorati.com/#studio) handoff standard: the thing should be yours, and legible, including the parts that live at a vendor.

The reason to do this early is not compliance theater. It is that teams use these tools far more freely once someone has told them plainly what is fine. Ambiguity does not produce caution. It produces quiet improvisation, which is the outcome the memo was supposed to prevent.

## Frequently Asked Questions

### Is it safe to put client information into ChatGPT or Claude?

It depends on the account, not the product name. Business and API tiers from the major vendors state in their terms that customer content is not used for model training, while consumer tiers may include a training setting on personal plans. If the work is going through a company account on a business plan whose terms you have read, client information is generally in scope for ordinary internal work. If it is going through someone's personal login, a different contract governs it and your policy has no bearing on the outcome. Check the terms for the specific product and plan you are on, since these documents change.

### Will an AI assistant expose files people should not see?

Not by itself. Microsoft and Google both state that their assistants only surface content the individual user already has permission to access. The practical problem is that most organizations have years of over-broad sharing that was never noticeable because nobody could search across it. A capable retrieval tool makes all of that findable at once. The fix is a permissions audit before rollout and a staged launch, so the surprises turn up with a small group rather than with the entire company.

### Do we need a formal AI policy?

You need one page, and it is worth writing before rather than after. Sort information into open, internal, and restricted, name the approved tools for each tier, and provide a fast way to ask about anything ambiguous or to get a new tool reviewed. Longer policies do not get read, and prohibition-shaped policies push usage somewhere you cannot see it. The goal is to make the safe path the obvious one.

### What should we ask a small AI vendor built on top of a larger model?

Ask who else touches the request, and ask for both sets of terms. A wrapper product has its own agreement with you and its own agreement with the model provider underneath, and those can differ on training, retention, and data residency. Also ask where the data is processed, how long it is kept, whether human review of content is possible, and what the export and deletion path looks like if you stop working together. A vendor who has been through this answers with links rather than assurances.
