OpenAI’s newest chatbot may be a whiz at math, but it seems to be lagging far behind humans in its academic rigor.
Last week the company announced 10 more artificial-intelligence-generated math advances that were found during internal development and testing of its next major large language model. This batch of results came from that LLM, Astra, and each one resolves or progresses a different “long-standing open problem” of “substantial interest” to the mathematical community. The company said that the total token cost was a mere $2,000.
The news quickly spread as yet another harbinger of AI’s promise—or threat—of outpacing humans to become a dominant, disruptive force in math and computer science. But over the days following the announcement, as flesh-and-blood experts poured over the nearly 250-page paper in detail, many grew frustrated.
On supporting science journalism
If you’re enjoying this article, consider supporting our award-winning journalism by subscribing. By purchasing a subscription you are helping to ensure the future of impactful stories about the discoveries and ideas shaping our world today.
Two of the most exciting results, the experts say, incorporate preexisting ideas from the recent mathematical literature without properly citing them. This contradicts OpenAI’s initial press release, which said that the problems Astra addressed “have been open and seen no progress on the main result for at least a decade.” (OpenAI has since updated the language to be more accurate).
“They are running roughshod over the work of others who came before them in a deliberate way,” says Steven Miller, a mathematician at Yeshiva University, who argues that OpenAI has effectively plagiarized his own research. “It seems completely systematic to me, and it points to research misconduct.”
The result Miller refers to concerns how many balls you can fit in a box—a seemingly simple problem, except these balls and boxes exist in a mathematical space of 1,000 dimensions—or even more. OpenAI’s paper improves the best estimate for how tightly these balls can possibly be packed. The LLM-generated proof hinges on a particular mathematical argument that it presented as its own but that actually first appeared in a 2016 paper by Miller and a collaborator.
Another of the 10 results resolves a long-standing question in group theory, which studies sets of mathematical objects called “groups” that interact in an organized way. Mathematicians have long wondered whether all groups have a property called “soficity,” the capacity to be faithfully approximated in a particularly way by other, simpler groups. The OpenAI paper establishes at least one group that lacks this property.
The discovery stunned Francesco Fournier-Facio, a mathematician at the University of Cambridge, who studies group theory—at least until he “engaged with this breakthrough as I would if a human had written it,” he says. The result, he and some of his colleagues found, wasn’t as novel as it first appeared. Like a number of recent AI breakthroughs, it pasted together ideas from the mathematical literature to build a new theorem. Once again, the LLM’s trick is its superhuman patience for assembling puzzle pieces, not the ability to make some profound intellectual leap.
In particular, Astra’s key mathematical step combined ideas first found in two papers from 2016 and 2019. Andreas Thom, a mathematician at the Dresden University of Technology, who co-authored both papers, summarized the result on MathOverflow.com, calling it “creative and at the same time elementary.”
OpenAI’s initial press release seemed to ignore—or be completely unaware of—these crucial, recent developments. Fournier-Facio argues that the two preceding papers show humans had not hit a stalemate with the soficity problem. OpenAI’s mathematicians did their best to attribute these ideas correctly in their paper, he says. But, in spite of their good intentions, “there is the big PR machine that wants to sound as impressive as possible and does not care about being 100 percent accurate,” he says.
“We take responsibility for the correctness of these results and are meeting the same standards generally expected of human mathematicians,” an OpenAI spokesperson said in a statement to Scientific American. “We plan to make small updates [to the paper] this week, consistent with standard academic practice.”
But as AI continues its campaign to conquer math without any built-in fealty to the field’s academic norms, some in the community are clearly losing patience. “OpenAI is now fully participating in high-level research,” Fournier-Facio says. “So they should be held to the same academic standards that we are.”
It’s Time to Stand Up for Science
If you enjoyed this article, I’d like to ask for your support. Scientific American has served as an advocate for science and industry for 180 years, and right now may be the most critical moment in that two-century history.
I’ve been a Scientific American subscriber since I was 12 years old, and it helped shape the way I look at the world. SciAm always educates and delights me, and inspires a sense of awe for our vast, beautiful universe. I hope it does that for you, too.
If you subscribe to Scientific American, you help ensure that our coverage is centered on meaningful research and discovery; that we have the resources to report on the decisions that threaten labs across the U.S.; and that we support both budding and working scientists at a time when the value of science itself too often goes unrecognized.
In return, you get essential news, captivating podcasts, brilliant infographics, can’t-miss newsletters, must-watch videos, challenging games, and the science world’s best writing and reporting. You can even gift someone a subscription.
There has never been a more important time for us to stand up and show why science matters. I hope you’ll support us in that mission.