| Model | Benchmark | Score | Source | Checked |
|---|---|---|---|---|
| GPT-6 Astra | OSWorld | 72.6% | Source ↗ | |
| GPT-6 Astra | SWE-bench Verified | 80.4% | Source ↗ | |
| GPT-6 Astra | AutomationBench | 41.4% | Source ↗ | |
| Claude Sonnet 4.6 | SWE-bench Verified | 77.2% | Source ↗ |
Methodology
Scores are copied from the vendor or harness page linked per row. They are point-in-time snapshots for comparing local models against hosted references, not targets to match on-device. Rows are re-checked weekly; any row older than ten days is flagged for refresh.