CodyBench Results
Benchmarks updated weekly by GitHub Actions.
How It Works
CodyBench runs A/B comparisons: same task with CodyMaster vs without. Each suite runs 3 times per platform. Scores are 0–100.
Latest Results
| Suite | With CM | Without CM | Delta | Date |
|---|---|---|---|---|
| TDD Regression Catch Rate | — | — | — | pending |
| Token Efficiency | — | — | — | pending |
| Memory Retention | — | — | — | pending |
Run cm bench locally to generate results.
Research Basis
- SkillsBench (peer-reviewed): curated skills +16.2pp, 2-3 focused skills +18.6pp
- Goose benchmark methodology: transparent runs, 3× repeat, publish even early/small results