GPT‑6.1 Sol in an external test: strong scores, limited conclusions
Why perfect scores do not necessarily produce a clear winner.

The short answer
We reviewed joonlab’s published documentation. The measurements were made by the project author, not Findbest.
Announcement date: 30 Sept 2026
What changed?
The project dated 30 September reports 3,527 calls across seven tracks in a ChatGPT Pro Codex environment. Its coding table lists 58 out of 58 for Sol, with several comparison models tied or very close.
Who is this relevant to?
For teams considering a switch, this illustrates an essential benchmark question: are the tasks difficult enough to reveal meaningful differences?
What this means for your work
Our analysis: Record runtime, corrections and acceptance criteria alongside success rates. Do not automatically adopt an external winner. Choose tasks where your existing workflow struggles, and score all responses using the same criteria.
Limits of this report
The author notes few repetitions and ceiling effects. Full responses, the harness and hidden tests are absent from the repository. We did not reproduce the measurements; they establish neither general API performance nor guaranteed reliability.

