Agent Sessions

Session-Bench v1

After a coding session ends, how much of the agent's work can another tool recover from the files the harness wrote? Session-Bench v1 runs the same two-turn task in each harness, records what really happened, and then checks 31 facts against the session files alone: every turn, response, tool call, result and exit code, the edited file before and after, the model, the token counts and the timestamps.

Bench v1.0 · released 2026-10-08 · 3 runs per harness · full report · replay every score · github.com/jazzyalex/session-bench

Ranking

HarnessScoreRecord30Causality20Usage15Portability20Durability15 1DeepSeek Harness CLI0.2.0-rc.2 · deepseek-flash96.930Record20Causality15Usage20Portability11.9Durability 2Pipi 1.0.0 · gpt-5.596.430Record20Causality12Usage20Portability14.4Durability 3Copilot CLI1.0.91 · gpt-6-luna96.030Record20Causality15Usage20Portability11Durability 4OpenCode CLI1.18.31 · opencode/muse-spark-1.3-contributor-free94.430Record20Causality15Usage20Portability9.4Durability 5Kimi Code2.1.1 · kimi-k2.7-code91.230Record20Causality12Usage20Portability9.2Durability 6Claude Code CLI2.1.272 · claude-sonnet-5[1m]89.330Record20Causality12Usage18Portability9.3Durability 7OpenClawOpenClaw 2026.9.8 (fc23bc8) · gpt-5.6-terra88.130Record20Causality12Usage17Portability9.1Durability 8Codex CLI0.154.0 · gpt-5.6-sol87.530Record20Causality15Usage15Portability7.5Durability 9HermesHermes Agent v0.21.5+4668.gdccb84b.dirty (2026.9.24) · upstream dccb84b9 · gpt-5.586.430Record20Causality5.5Usage17Portability13.9Durability 10Claude Desktop Code (Local)2.1.270 · claude-opus-585.628.8Record18.3Causality12Usage18Portability8.6Durability 11Antigravity1.2.17 · claude-sonnet-4-682.730Record20Causality8Usage15Portability9.7Durability 12Cursor CLI2026.10.01-e373342 · cursor-grok-4.5-high78.130Record20Causality3Usage15Portability10.1Durability
full marks70% or more of the categorybelow 70%

Scores are out of 100: record fidelity (30), causality and context (20), usage and attribution (15), portability and openness (20), durability and signal (15). Each score is the mean of 3 runs. A harness name links to the document that says how its row was decided. Codex Desktop: shares the Codex CLI row.Cursor Desktop: not scored yet. Charts and a row-by-row reading are in the post.

The scores measure the files, not the agent. They are not scores of coding quality, price, speed or vendor reliability. The harnesses did not run the same model, and the benchmark does not grade the model.

What the numbers say

Nobody loses the conversation. Eleven of twelve harnesses keep every turn, response, tool call and result, and the link between a call and its result. Claude Desktop is the exception, and the cause is the run rather than the format: its model wrote the file from inside a shell command, so the edit has no result of its own and one result in four goes missing.

The spread comes from three other places.

The leaders do nothing clever: one append-only file or one small store per session, a version field, usage on each response, few repeats.

Read this before you quote a number

This is a first release. One small synthetic task, three runs per harness. It says nothing about long sessions, sub-agents, crashes or compaction. Some rules were waived or corrected after results were known. The report states each one:

Known limits of the method (12)

Check it yourself

Every score can be recomputed. The release holds the sanitized session files of all 36 runs, the scoring source, and one command per run that replays the 31 metrics from those files.

If you think a row is wrong, open its document from the table. It lists every judgment call and what each one is worth.

This page shows a byte-identical copy of data/leaderboard-v1.yml from the Session-Bench repository; scores and evidence are changed there first. The earlier gate-based report card is at Session-Bench v0.4.