← Blog

AI code review4 Oct 2026

The conversation spine: turning an agent session into useful code review context

Reviewers need more than a diff. The spine keeps the decisions, failures, edits, and verification that explain how the code got there.

A conversation thread connecting a failed test, code edit, decision, verification result, and source into a reviewer dossier
The spine preserves useful evidence and carries its source into one reviewer-ready view.

A pull request shows what changed. It rarely shows how the change got there.

That missing path matters more when an agent wrote the code. For AI code review, the useful context is sitting in the coding session: the user’s request, the agent’s proposal, a rejected approach, a failed command, a scope decision, the test that passed, and the thing nobody verified. By the time the pull request opens, most of that has disappeared.

Review Assist calls the useful path through that session the conversation spine. It is not a summary produced by another model. It is a deterministic view of the session that keeps the conversation and marks the important work around it.

Why searching the transcript did not work

Agent transcripts are mostly machinery. Across 41 measured sessions in one workspace, the raw transcripts were 85 MB. The actual conversation was a median 4.8% of that data. The rest was tool input, tool output, metadata, and repeated state.

The first approach was retrieval. The author agent paged through windows and searched for likely terms. On one session it pulled 388 KB across 34 calls and still built an incomplete picture of a conversation that was only 344 KB. Search cost more than reading the conversation, and it could never prove that the missing decision had not been overlooked.

The spine takes the opposite approach. It reads the session once, keeps the useful sequence, and pages the result without dropping material. In the measured set, 40 of 41 session spines fit in one call. The median was 11,886 tokens.

What the spine keeps

The spine keeps both sides of the conversation. That sounds obvious, but it is easy to get wrong. Roughly two thirds of the prose in the measured sessions came from the agent. A user reply such as “yes, do that” is useless without the proposal it accepts. A reported finding is often on the agent’s side too. Of 23 abandoned approaches found in seven real Intent Documents, 17 came from the agent’s messages.

It also harvests compact evidence from the work around the conversation:

  • •Structured questions and selected answers, because scope decisions often happen in a picker rather than in prose.
  • •Commands and edits as one-line events, so the author can see what ran and which files changed without carrying pages of output.
  • •Plan revisions as deltas, which expose work that was added or deliberately dropped.
  • •Failures and long gaps, which identify phases worth inspecting more closely.
  • •Stable transcript indices, so the author can open the raw entries around a claim when it needs the exact output or measurement.

How that helps review the written code

The spine is useful because it lets Review Assist ask better questions about the diff. It can recover the original problem and compare it with what was implemented. It can identify requirements that appeared halfway through the session. It can tell the difference between code that was intentionally excluded and code that the agent simply forgot.

Plan revisions expose another useful signal. If the plan dropped a migration but the diff changes a schema, the reviewer has something specific to challenge. If the session says a fallback was abandoned after a failed test but the same pattern appears in the final code, that is worth checking. If the agent claims a behavior was verified, the transcript index provides a path back to the command output behind that claim.

This is more useful than dumping the transcript into a prompt. The reviewer does not need every compiler line. It needs the pieces that change how the code should be judged: intent, constraints, assumptions, alternatives, failures, edits, and verification.

The raw session never goes to the reviewer

The boundary is important. The Author role can read the local session through the spine. The cold Intent Reviewer cannot read the transcript or the local filesystem. It reads the diff and asks questions. The Author answers from the session, and the server records which answers actually came from that role.

The coding transcript stays on the developer’s machine. What leaves the process is a curated Intent Document: the problem, requirements, assumptions, adopted approach, rejected alternatives, guided tour, and verification. The human reviewer gets the useful context without receiving a raw log of the developer’s conversation.

Completeness comes from paging, not deletion

An early version capped the spine at 600 KB and deleted assistant turns when it crossed the limit. That removed the side carrying most of the reasoning and most of the recorded alternatives. It made large sessions look complete when they were not.

The current design pages the spine on item boundaries. A large session costs more calls, but no category of evidence disappears to make it fit. Oversized individual turns are sliced with their character ranges attached. Every item still points back to its location in the raw transcript.

What the reviewer gets in practice

A reviewer does not see a conversation archive. They see answers to the questions that matter for this change: What was the actual problem? Which constraints shaped the code? What was tried and rejected? Which assumptions could still be wrong? What was verified, and what was not?

That is the job of the conversation spine. It turns a noisy agent session into traceable review context, while keeping the final review focused on the code.

Review AI-written code with its intent intact

Review Assist is free, open source, and stores none of your code.

Install Review Assist