Gen AI’s Big Problem is that it can’t solve Big Problems

Gen AI’s Big Problem is that it can’t solve Big Problems

There were red faces at OpenAI when staffers claimed an apparent “breakthrough” for flagship model GPT-5 that simply didn’t add up. It is a humiliating tale of exuberant over-claim that offers an important lesson for us all.

t all began with maths whiz Mark Sellke— who had just joined the AI start-up from Harvard University—posting on X that through using GPT-5 he and another academic had “found solutions” to several fiendishly complex Erdős problems.

Named after Hungarian mathematician Paul Erdős, they are a series of over one thousand mathematical challenges which mostly remain unresolved. While Sellke said GPT-5 had “found solutions”, Kevin Weil, OpenAI’s former product chief who now heads its science arm, said GPT-5 had also “made progress on 11 others”:

The phrase “made progress” implied GPT-5 was actually solving the problems.

OpenAI technical staffer Sebastien Bubeck posted: “Science acceleration via AI has officially begun.” Colleague Boris Power, head of applied research, joined the victory lap saying: “Wow, finally large breakthroughs at previously unsolved problems!!” Again, the implication was that GPT-5 had cracked the challenges.

It didn’t take long before maths experts spotted a eureka moment that wasn’t.

Thomas Bloom, a maths research fellow at the University of Oxford who keeps a tally of solved Erdős problems and those that remain ‘open’, accused Weil of a “dramatic misrepresentation”, saying GPT-5 had merely “found references” which solved problems. Notable figures joining the social media backlash included Yann LeCun, Meta’s chief AI scientist, who said the OpenAI staffers had been “hoisted by their own GPTards”, and Sir Demis Hassabis, CEO at Google DeepMind, who lamented: “This is embarrassing.”

Bubeck deleted his post and said on X he hadn’t meant to mislead anyone, adding: “Only solutions in the literature were found that’s it, and I find this very accelerating [sic] because I know how hard it is to search the literature.” Gary Marcus, scientist, author, entrepreneur and LLM critic, blogged: “Yeah, right. I don’t know anybody who believes his retrenchment.”

Marcus hoped the episode would be seen as a “teachable moment”:

“Some people (I won’t name the guilty) were extremely quick to take Bubeck at his word. But why? The claim would have been extraordinary and should have been vetted closely. I smell a really big dose of people believing what they want to believe.”

And what is it that they—desperately—want
to believe?

That LLMs such as GPT-5 are displaying signs of artificial general intelligence (AGI)—AI as smart, if not smarter, than you or I, the holy grail of the major AI developers.

GPT-5 wasn’t solving anything, it was merely looking for patterns in its training data and generating plausibly correct outputs. That’s all LLMs can do. They don’t solve problems or come up with novel solutions any more than they can genuinely create. But thank you Team OpenAI for providing us with another proof-point in the fightback narrative against Big Tech’s relentless AI hype.

It’s accelerating.

This article first appeared in Charting Gen AI, the global newsletter charting the impacts of generative AI on creators and human-made media, the ethics and behaviour of the AI companies, the future of copyright in the AI era, and the evolving AI policy landscape.

Subscribe: https://grahamlovelace.substack.com