Codex Reset
AI News
GitHub

QQ·微信群

CODEX / SIGNAL STUDIO

Your Codex, in focus.

Hassan releases GeoGuess Bench to evaluate AI models on GeoGuessr

Hassan

Hassan has released GeoGuess Bench, a benchmark designed to evaluate how well AI models play the geolocation game GeoGuessr. The benchmark provides each model with 210 photos from across the globe, asks it to predict where each photo was taken, and scores performance based on distance from the true location.

In the benchmark results, Claude Opus 5.5 took the top spot, scoring above Fable 5.1. Open models also demonstrated strong cost efficiency: Muse Glimmer 30B outperformed GPT 6 Astra at roughly 45 times lower cost, while GLM 5.3 Flash matched GPT 6.1 Sol at approximately one-fifth the cost. The full leaderboard and guesses are available on [GeoGuess Bench](http://geoguessbench.com), with code hosted on [GitHub](http://github.com/Nutlope/geoguessbench).