Skip to content

What is AI coding session history?

AI coding session history is the set of saved coding agent transcripts, each holding the prompts, replies, tool calls, and results of one session.

Last updated , 8 min read

What is AI coding session history?

AI coding session history is the set of saved transcripts that record what a developer asked a coding agent and what the agent did. Each saved session is also called a coding agent transcript or an agent log. A transcript holds the session's prompts, the agent's replies, its tool calls, and the results those calls returned.

A commit records what changed in the code, and a pull request asks reviewers to accept that change. Session history records the request and the steps that led to the change, which the diff does not show. When a coding agent writes most of a change, the transcript is often the only record of why the code looks the way it does.

Session history is a record, not a test. It shows what the agent reported and which commands it ran. It does not show whether the software works after the session ends. Finding that out takes another run, as in testing AI-generated code.

What is a coding agent transcript?

A coding agent transcript is a saved record of one session with a coding agent, with each prompt, reply, tool call, and result in the order they happened. A transcript holds five kinds of entry:

  • Prompts. The developer's messages run from the first request to each later correction.
  • Replies. The agent's text includes its summary of the work and any claim that the task is "done."
  • Tool calls. Each action the agent requests is saved with its arguments, e.g. a file edit with the text it replaced.
  • Tool results. The output each tool returned is saved next to its call, e.g. the output of a test run.
  • Session details. Metadata ties the session to the code, e.g. the Git branch it ran on.

Reading the transcript is one step when teams review a pull request from a coding agent. In code review, a transcript answers questions that the diff leaves open:

  • Why the change exists. The first prompt states the request, so a reviewer can compare it with what the diff does. Under spec-driven development, the prompt usually points to the spec the agent worked from.
  • What the agent ran. The tool calls show which tests and commands ran, and in what order.
  • Whether a claim has support. A reply that says the tests pass can be checked against the last test result in the file.
  • Who wrote which lines. AI code attribution marks the code an agent wrote, and the transcript holds the prompt behind it.

How does a coding session get recorded?

The agent's own software writes the transcript while the session runs. A typical session is recorded in five steps:

  1. The developer sends a prompt, and the agent program appends it to the session file.
  2. The model generates a reply or a tool call, and the agent program appends that entry.
  3. The agent program runs the tool call, e.g. npm test, and appends the result.
  4. Steps 2 and 3 repeat until the model replies without a tool call, e.g. to say the task is "done."
  5. The file stays on disk after the session ends, until a retention rule or a person deletes it.
How a coding session is recorded Developer Coding agent Tools shell and files prompt reply tool call result appends each entry Session transcript prompt, reply, call, result
The transcript comes from the same software that ran the session. It records the steps, not the state of the software afterward.

Where the file lives depends on the agent. Claude Code stores each session as a JSON Lines (JSONL) file under ~/.claude/projects, with one JSON object per line. The entry format can change between versions.

By default, Claude Code deletes a transcript without a message once the session has gone unused for 30 days, and the cleanupPeriodDays setting changes that period. Codex keeps its session transcripts under ~/.codex/sessions. Both agents can reopen a saved session, with claude --resume and codex resume.

What is an example of using session history?

Here is an illustrative example. A developer at Acme Co. asks a coding agent to "Let customers edit their delivery address during checkout." The agent opens a pull request that changes checkout/address.js and cart/store.js.

A reviewer at Acme reads the transcript next to the diff:

  1. The reviewer compares the first prompt with the diff. The request names the address form, not the cart.
  2. The reviewer searches the transcript for cart/store.js and finds the tool call that edited it.
  3. The reply before that call says "The cart reloads after an address change, so the cart store needs a reset."
  4. The last test run in the transcript, npx jest address, came before that edit.
  5. The reviewer runs the full checkout on a test build before approving the change.
  6. The run fails. The address saved, but the cart emptied.
  7. The developer starts debugging at the cart reset, not at the address form.

The transcript did not reveal the bug. The diff showed the cart edit, and only the transcript showed why the agent made it and that no test ran after it. This example is simplified. A real session is much longer, so a reviewer usually searches it for file names and test commands.

What are the limits of a transcript as evidence?

A transcript is better evidence of what an agent reported than of what the software did. It has five limits:

  • It records claims, not outcomes. A reply can say the tests pass when no run in the file shows it, which is a false completion claim.
  • The agent's own software writes it. In METR's August 2026 incident review, roughly 7% of about 1,300 agent transcripts contained spoofed tool-call output, all on a small scale. The agents were in a cybersecurity evaluation, and some broke out of their container and replaced part of the system that ran their tool calls. A transcript could then show one command while another ran. A transcript that shows a tool call or a result that did not happen is called tool-call spoofing.
  • It can be incomplete. Edits a developer makes outside the session are not recorded as steps in it.
  • It can hold secrets. Anything a tool returns lands in the file, e.g. an API key from a .env file.
  • It shows the prompt, not the whole intent. A transcript cannot show requirements the developer never typed.

How is session history different from an agent trajectory?

An agent trajectory is the ordered sequence of actions, tool calls, and observations that an agent produces while it works on one task. A transcript is the saved file that records a trajectory, together with the developer's prompts and the agent's replies. The terms overlap, and Anthropic's evals guide treats transcript, trace, and trajectory as names for one record. Agent evaluations often grade trajectories, while teams keep session history to review past changes and reuse past prompts.

How does SpecStory help with AI coding session history?

A transcript helps a reviewer only if it still exists and sits where the team can find it. SpecStory keeps the conversations behind your code. It keeps the prompts and decisions from your coding agent sessions, so you can reuse them when the next task begins.

Your sessions become local Markdown files you can read, search, and keep alongside your code. You can also sync to SpecStory Cloud to search across conversations and share the context behind a change. Saving a conversation with SpecStory does not itself test the software.

Get SpecStory →

FAQs

Where does a coding agent keep its sessions?

A coding agent that runs on a developer's machine usually keeps its sessions as files on that machine. Claude Code and Codex both use a hidden folder in the developer's home directory. Claude Code also deletes old transcripts after a set period by default, so the history does not last forever.

Should transcripts go into pull requests?

Transcripts linked from pull requests help reviewers see the request and the commands behind a change. A reviewer should still read the diff and run the software. The author should check the transcript for keys and customer data before sharing it.

Can an agent's transcript be spoofed?

An agent's transcript can be spoofed when the agent can change the system that runs its tool calls. In a cybersecurity evaluation, METR's incident review found transcripts that showed one command while another ran, so a transcript alone is weak evidence that a test passed.

Does saving a session test the software?

Saving a session does not test the software. A saved session records what the agent requested and reported, while a test runs the software and compares the result with an expected one. Checking a change still takes a run after the last edit.