What leaves your machine when a coding tool runs?
The main data that leaves the machine is model input, the text an AI coding tool sends to a model provider at each step. That text includes the prompt, the files the agent read, and the output of the commands it ran. The tool can send usage data to its maker too, and the agent's commands can reach other hosts.
A coding agent is driven by a large language model (LLM) that usually runs on the provider's servers, not on the developer's machine. The model receives only what the tool sends it, so the tool sends the code the task needs. The maker and provider can be one company or two.
Data retention is how long a service keeps the data it receives, e.g. code sent to a model provider, before it deletes that data. Training terms say whether the service may use that data to improve its models. Both follow the terms for the account that sent the data.
A sandbox limits which files and hosts the agent's commands can reach. Model input still leaves, because the tool needs the model's API. The same questions apply to other tools that read code, e.g. hosted AI code review.
How does a coding tool send data?
A coding tool sends model input in a loop that runs in this order:
- The developer types a prompt, e.g. a request to change the checkout.
- The tool builds a request from its instructions, the conversation so far, and the prompt. Choosing that content is called context engineering.
- The tool sends the request over an encrypted connection to the model provider.
- The model replies, often with a request to read a file or run a command.
- The tool runs it on the machine and adds the result to the conversation.
- The next request carries the conversation again, or a summary once it grows long.
Data leaves the machine by these routes:
- Model input. It goes to the model provider. Some tools route it through the maker's backend, e.g. Cursor, whose data use page says this holds even with a user's own API key.
- Usage telemetry. The tool reports how it runs to its maker and, in some tools, to services the maker uses. Metrics can leave out content, e.g. Claude Code's metrics, which its data usage page says never include code, prompts, or file paths. Code and prompts still go to the provider as model input.
- Reports. A crash report carries error messages. A bug report can carry the whole conversation, including code, e.g. one sent with Claude Code's
/feedbackcommand. - Tool requests. A call to a remote Model Context Protocol (MCP) server or a web fetch goes to another host, under its terms.
- Commands. A command can reach any host the network allows, e.g. the code host that receives a
git push. - Cloud sessions. A background coding agent usually works on a clone of the whole repository, on machines its maker runs.
What is an example of what a coding session sends?
Here is an illustrative example. A developer at Acme Co. asks a coding agent to "Let customers edit their delivery address during checkout." The session sends data in seven steps:
- The first request carries the tool's instructions, Acme's rules file for agents, and the prompt.
- The model asks to read
src/checkout/address.ts. The tool sends all 180 lines of the file in the next request. - The agent runs
git log -5. Its output, with each author's name and email address, goes to the provider. - The agent runs
npm test, and one test fails. The failure line,expected 2 items, received 0, goes into the next request. - The agent asks a remote documentation server about the cart API, and that server's host receives the question.
- While the session runs, the tool sends usage metrics to its maker, e.g. how long each request took.
- When the session ends, the transcript stays on the laptop, and the provider keeps its copy for as long as Acme's account terms allow.
The tool never uploaded the whole repository. It sent two files, some history, and a test failure, plus the conversation around them. This example is simplified. A real session often reads many more files, and each one leaves the same way.
Which settings limit what leaves the machine?
Settings change who receives the data, what they may do with it, or how much the tool sends:
- Account type. Terms often differ by account type, e.g. Claude Code's commercial terms, under which Anthropic does not train on code or prompts unless the customer opts in. Consumer accounts choose in a setting. Turning off training does not by itself stop sending or storing.
- Privacy mode. Some tools put the training choice behind one switch, e.g. Cursor, whose security page says Privacy Mode keeps a user's data out of training.
- Zero data retention. Zero data retention (ZDR) is an agreement under which the provider processes each request and does not store it after responding. Access can depend on the plan, e.g. Claude Code, whose ZDR page lists it for qualified Claude for Enterprise accounts, enabled one organization at a time.
- Telemetry switches. Environment variables turn off reports to the maker, e.g.
DISABLE_TELEMETRY=1for Claude Code's usage metrics. - Deny rules. A rule in the tool's settings file, e.g.
Read(./.env), stops its file tools from reading that file. A script can still open it, so secrets belong outside the agent's reach. - Model route. Model input can go to a cloud platform the team already uses, e.g. Amazon Bedrock, where Claude Code turns off its metrics by default. A model on the team's own server keeps model input inside its network.
- Network limits. Egress control limits which hosts commands can reach, and least privilege limits which files and keys the agent holds.
In Claude Code, the switches and deny rules can go in the user settings file, ~/.claude/settings.json:
{
"env": {
"DISABLE_TELEMETRY": "1",
"DISABLE_ERROR_REPORTING": "1",
"DISABLE_FEEDBACK_COMMAND": "1"
},
"permissions": {
"deny": ["Read(./.env)", "Read(./.env.*)"]
}
}
What are the limits of vendor settings?
Vendor terms cover the vendor's own servers, and settings cover what the tool sends. They leave these gaps:
- Other recipients. Terms do not cover other hosts the agent contacts, e.g. a remote MCP server, which Claude Code's ZDR page puts outside its scope.
- Exceptions. Retention terms usually let the provider keep data for a legal duty or to investigate misuse.
- The account that signed in. Terms follow the account on each request, so a personal login on a work laptop gets personal terms.
- Local copies. Many tools keep transcripts on disk, e.g. Claude Code, which keeps them in plain text for a period that
cleanupPeriodDayssets. That session history holds each file the agent read. - Redaction. Some reports remove known key formats, the patterns secret scanning looks for. Model input usually goes as the agent read it.
- Injected instructions. Prompt injection can make the agent send data through a route the settings allow.
Terms and defaults also change between releases. A team can observe what leaves the machine, not what a service does with the data.
How is model input different from telemetry?
Model input is the content the model needs for its next step, so it holds prompts and code by design. It goes to the model provider. Usage telemetry is data about how the tool itself runs, e.g. how often requests fail. It goes to the tool's maker and, in some tools, to services the maker uses. Turning off telemetry does not change model input.
Review tools that read runtime context use the word in another sense. There, telemetry means production telemetry, records of how deployed software behaved for its users.
FAQs
Does a coding agent send the whole repository?
A coding agent usually sends the files and command output it reads during a task, not the whole repository at once. Over a long session, that can still add up to much of the codebase. A cloud session differs, because the agent usually works on a copy of the whole repository on the maker's machines.
What is zero data retention?
Zero data retention is an agreement under which a model provider processes each request and does not store it after returning the response. Providers usually keep exceptions for legal duties and misuse. The agreement covers only the provider's own servers, not the other hosts an agent contacts.
Can a coding tool run without sending code anywhere?
A coding tool can run without sending code to a model provider when its model runs on the team's own server. The tool's telemetry and the agent's own commands can still reach the network, so telemetry switches and egress control apply there too.
Can a model provider train on code it receives?
A model provider can train on code it receives only when its terms for that account allow it. Terms often differ by account type, and some tools offer a privacy mode. A setting that turns off training does not by itself stop sending or storing.