Starburst Leaderboard
Will be updated with further LLMs and people
If you wish to try Starburst, email me at chapinalc@gmail.com.
Era 1 (1-20)
Era 2 (21-50)
Era 3 (51-60)
Era 4 (61-100)
----====ASI Threshold?====----
Ethan Era (5) (101-130)
1 Human-day 115, very roughly 3.5σ above median
Cole Era (6) (131-160)
1 Human-day 156, very roughly 3σ above median
Era 7 (161-190)
Era 8 (191-220)
Era 9 (221-250)
----====Crisis Threshold====----
Era 10 (251-275)
Gemini 3.1 Pro Preview (Days 260, 260, 260, 260, and 265)
Gemini 3.5 Flash Extended (Days 260, 260, 260, 260, and 265)
GPT-5.5 Extended Thinking (Days 260, 265, and 265)
Claude Opus 4.6 Extended Thinking (Days 270, 295, and 345)
Era 11 (276-300)
Claude Opus 4.7 Adaptive Thinking (Days 280, 305, and 315)
Ryan Era (12) (301-350)
1 Human-day 319, very roughly 1σ above median
GPT-5.4 Extended Thinking (Days 310, 320, and 380)
GPT-5.2 Extended Thinking (5/10)
Gemini 3 Pro (1/10)
Era 13 (351-400)
Grok 4.20 (7/10)
Grok 4.1 Thinking (4/10)
Grok 4 Fast (4/10)
Grok 4 (2/10)1
GPT-5.1 Extended Thinking (1/10)
----====Weak AGI Threshold (If Above 50%)====----
Era 14 (401-450)
Estimated average adult American performance
Claude Opus 4.5 Extended Thinking (7/10)
GPT-5 Thinking (3/10)
o4-mini-high (3/10)
GPT-5 Extended Thinking (2/10)
o3-pro (1/5)2
Gemini-2.5-Pro-6-5 (1/5)*
o3 (1/10)
Era 15 (451-500)
Gemini-2.5-Pro-3-25 (2/5)*
Deepseek v3.2 Thinking (1/5)
Gemini-2.5-Pro (full release) (1/10)3
----====Genuine Reasoning Threshold====----
Era 16 (501-525)
Claude Sonnet 4.5 Extended Thinking (10/10)
Grok 3 (Thinking) (5/5)
Claude Opus 4.1 Extended Thinking (5/5)
o3-mini-high (2/3)*
Claude Opus 4 Extended Thinking (2/5)*
Claude Sonnet 4 Extended Thinking (2/5)*
o1 (1/3)*
Claude Sonnet 3.7 Extended Thinking (1/5)*
Era 17 (526-550)
GPT-4.1 (5/5)
GPT-4*
GPT-4o*
Deepseek R1*
Era 18 (551-575)
GPT-3.5*
Era 19 (576-600)
Era 20 (601-650)
.
For models released in 2026, we’ve switched to a more human-comparable benchmarking scheme in which we report the day the model solved the game.
Previously, we gave models only a single era of data and reported the reliability for the earliest era in which they could solve it. We did this to remove agency concerns and the possibility of models being overwhelmed by irrelevant data. This handholding to focus only upon core reasoning abilities is no longer necessary.
We expect a smart (by our sense of the term) person to solve Starburst in era 10, an average (adult American) person to solve it in era 14, and a dumb person to solve it in era 18. The smartest geniuses could probably solve it early in era 4, and an arbitrary superintelligence could probably do better than them. If adapted to not require language, some animals could probably solve it in the last era. These estimates could easily be off by an era, but more than that would be quite surprising.
Models are listed by the earliest era in which they are capable of solving Starburst, along with a fraction of solutions correct at that era for eras 16 and earlier. Human and LLM scores in eras 16 and especially later should not be treated as directly comparable, since LLMs may be able to gain a substantial advantage by throwing knowledge at the problem in those eras.
Note on calling it an AGI threshold: a model that reaches here probably has general intelligence roughly on par with the average adult American. It might (for a time, likely will) not be able to do everything the average human can do on a computer due to non-intelligence limitations such as time horizons, tool use, hallucinations, or perceptual abilities. I’d call a model with every cognitive ability at or above the average adult American strong AGI, even if it’s only average-intelligence-strong-AGI.
More information about Starburst: https://pennheretic.substack.com/p/you-can-do-fictional-theoretical
Models marked with an asterisk were tested on an old version of the prompt with a subtle ambiguity. This had little to no effect on results, but is noted for completeness.
At this era, Grok 4 gives a message like “Uh-oh, too much information for me to digest all at once. You know, sometimes less is more!“ in roughly 20% of attempts. These are treated as errors and therefore not counted towards the denominator.
Unlike most other models in this list, o3-pro is only accessible to users paying $200/month and enterprise customers. This essentially gives it an unfair advantage over the $20/month models on the leaderboard.
According to Google’s documentation, this is the same model as the 6-5 snapshot, but it performs markedly worse. This is likely due to the full release usually expending roughly half the thinking tokens of the snapshot. Even setting a thinking budget in AI Studio doesn’t fix this issue.

