External test

GPT‑6.1 Sol in an external test: strong scores, limited conclusions

Why perfect scores do not necessarily produce a clear winner.

Illustration for Findbest.si
Findbest perspective

The short answer

We reviewed joonlab’s published documentation. The measurements were made by the project author, not Findbest.

Announcement date: 30 Sept 2026

What changed?

The project dated 30 September reports 3,527 calls across seven tracks in a ChatGPT Pro Codex environment. Its coding table lists 58 out of 58 for Sol, with several comparison models tied or very close.

Who is this relevant to?

For teams considering a switch, this illustrates an essential benchmark question: are the tasks difficult enough to reveal meaningful differences?

What this means for your work

Our analysis: Record runtime, corrections and acceptance criteria alongside success rates. Do not automatically adopt an external winner. Choose tasks where your existing workflow struggles, and score all responses using the same criteria.

Limits of this report

The author notes few repetitions and ceiling effects. Full responses, the harness and hidden tests are absent from the repository. We did not reproduce the measurements; they establish neither general API performance nor guaranteed reliability.

Read next

Ideas to put into practice →