Claude Code对比评测长帖约 5 分钟

同一模型看套件花费差

Harness choice beats model choice

要点

  1. 同一模型同一任务,九套件通过率50%到67%,单次花费从1到18美元
  2. 一道难题Claude Code花381轮64美元,Pi用90轮2.5美元同样修好
  3. 求稳用Codex;任务会跑上千次选Pi;宁可早停别死磕选Exo
  4. 前缀缓存会让调试过的任务看起来更便宜,正式评测必须冷启动

原帖开头

A year ago the question was which model. Now it's which harness. Pi, Exo, Claude Code, Codex, DeepSeek Harness and 4 others. Same model, same tasks, same runtime. 360 runs, 2 billion tokens. Pass rates: 50% to 67%. Cost per pass: $1.05 to $18.34. Introducing FrontierHarness Ev