Agent Sessions
Session-Bench v1
After a coding session ends, how much of the agent's work can another tool recover from the files the harness wrote? Session-Bench v1 runs the same two-turn task in each harness, records what really happened, and then checks 31 facts against the session files alone: every turn, response, tool call, result and exit code, the edited file before and after, the model, the token counts and the timestamps.
Ranking
Scores are out of 100: record fidelity (30), causality and context (20), usage and attribution (15), portability and openness (20), durability and signal (15). Each score is the mean of 3 runs. A harness name links to the document that says how its row was decided. Codex Desktop: shares the Codex CLI row.Cursor Desktop: not scored yet. Charts and a row-by-row reading are in the post.
The scores measure the files, not the agent. They are not scores of coding quality, price, speed or vendor reliability. The harnesses did not run the same model, and the benchmark does not grade the model.
What the numbers say
Nobody loses the conversation. Eleven of twelve harnesses keep every turn, response, tool call and result, and the link between a call and its result. Claude Desktop is the exception, and the cause is the run rather than the format: its model wrote the file from inside a shell command, so the edit has no result of its own and one result in four goes missing.
The spread comes from three other places.
- Saying things once. A reader that walks the session file from top to bottom should meet each event one time. In OpenCode, Kimi, OpenClaw and Antigravity every prompt, call and result is stored more than once, with no field that marks the later copies. Pi stores each event once.
- Counting tokens. Nine harnesses store input, output and cache token counts beside each response. Cursor CLI stores no usage in the session record. Only four keep enough to add the responses up and check them against a session total.
- Being one thing you can copy. Five harnesses keep parts of a session in stores that all sessions share, so there is no folder you can lift out and call “this session”. Claude Code, Claude Desktop and Codex write no format version.
The leaders do nothing clever: one append-only file or one small store per session, a version field, usage on each response, few repeats.
Read this before you quote a number
This is a first release. One small synthetic task, three runs per harness. It says nothing about long sessions, sub-agents, crashes or compaction. Some rules were waived or corrected after results were known. The report states each one:
- Replaced attempts. The rubric says: do not replace a scheduled run with a calibration or corrective attempt. This rule is waived for v1 for OpenClaw (two attempts replaced: one controller fault, one turn without a visible response), Claude Desktop (two corrective runs after invalid captures) and Kimi (five attempts stopped by a controller setting, the provider rate limit, or a reply without the required marker). No replaced attempt had a score. Each one is listed with its reason in the adapter document of the row.
- OpenCode native record. The first scoring read the
eventtable, in which every part has one to five update rows; duplicate safety is then zero. On 2026-10-05 a read without that table was tried and gave 97.4. It was withdrawn on 2026-10-08, because the rubric fixes the read before scoring. The score shown uses the first read. - OpenClaw reconciliation. The store holds the run token total twice and no usage per request. An earlier candidate gave 3 points for reconciliation; they are removed.
- Codex packets keep the address of this repository, which names the operator’s GitHub account, by owner choice. The home name, home paths, e-mail, account ids and tokens are aliased, blanked or zeroed in every packet; the limits below list what stays readable.
Known limits of the method (12)
- Rules changed after captures. The tiered evidence rule (2026-10-03), the single duplicate rule (2026-10-05) and the privacy rules were written after the first captures and after some first scores. The ranking is a careful comparison, not a pre-registered experiment.
- Duplicate safety depends on the native record that the decoder reads. A store that the decoder must open for one fact brings all its repeats into the count. Rows whose facts sit in one lean file do better than rows that spread facts over several stores.
- Reconciliation. The comparator accepts the decoder’s statement that per-response usage records sum to the declared total; it does not add them up itself.
- Response text. A native response matches the observed one by turn, role and response marker. The full response text is not compared.
- Usage attribution. A usage record can be joined to its response by turn when the format gives no response id.
- Action arguments. The comparator matches an observed action to a native one by id, turn, kind, path and shell command line. It does not compare the text of an edit. Kimi has its own check of the edit text; the other rows have none.
- Rules written for these packets. The Kimi edit-text check and the Copilot rule that finds a message text inside a store cell were written for these packets. The Copilot rule has no minimum text length; the shortest message here has more than 200 characters.
- Reviews. Each public packet set was checked by a separate agent session on the same host and account, not by a second operator. Three outside model reviews were run, of candidates v38, v39 and v40. Each returned NO-SHIP with findings. Their findings are answered in the notes above and in these limits. A fourth review, of candidate v41, found no defect that makes a score, a rank or a stated claim wrong; it was a static reading, without replays. This release differs from candidate v41 only in this sentence and in the replay command of REPRODUCE.md.
- Machine details that stay readable. Some packets keep device and inode numbers, file counts and capture times of the capture host in their inventories and receipts (in Codex also a
device:inodestring and three file counts of the operator’s Codex folder that the replay reads). The values are stable for the capture host, so a reader can tell that packets came from the same machine and folder. They hold no account id, name or content. - Vendor instruction text (system prompts, tool descriptions) is blanked at equal length in the public packets, with these exceptions: Pi keeps four short tool descriptions; Kimi keeps the first sentence of its system prompt, from which the harness name is read; Copilot keeps the operating-system line of its system prompt; Claude Code keeps the platform field of one environment record.
- Privacy gate. Every packet passes an automatic check of digests and hex values from 12 characters, and of shorter values under digest-named keys. A short hex value of 6 to 11 characters under an ordinary key is not checked. The check is a filter, not a proof that no private value remains.
- What ‘verified’ means here. A verified row has three replays that recompute all 31 metric states from the public packet, and an approving review by a separate agent session. It does not mean that every judgment in the adapter document is beyond dispute; each adapter document lists its judgment calls and their point effect.
Check it yourself
Every score can be recomputed. The release holds the sanitized session files of all 36 runs, the scoring source, and one command per run that replays the 31 metrics from those files.
- Report: ranking, notes and limits
- Replay instructions
- Rubric: the 31 metrics and how each is matched
If you think a row is wrong, open its document from the table. It lists every judgment call and what each one is worth.
This page shows a byte-identical copy of
data/leaderboard-v1.yml from the Session-Bench repository; scores
and evidence are changed there first. The earlier gate-based report card is at
Session-Bench v0.4.