Mathematician Andreas Thom publicly demanded proof that OpenAI did not train its recent non‑sofic groups breakthrough on his unpublished discussions, and the company’s response offered no verifiable evidence. The request, echoed by fellow researcher Gábor Kun, has sharpened scrutiny of how large‑scale AI models ingest academic material.
What the mathematicians asked
Thom wrote to OpenAI researchers Sébastien Bubeck and Mark Sellke, asking whether his conversations with ChatGPT were "part of the training data or accessible to the reasoning process." In the email chain he noted that no qualification, explanation, or evidence was provided in OpenAI’s reply. The Verge article records Thom’s exact wording: “No such qualification, explanation, or evidence was given.”
OpenAI’s public stance
OpenAI issued a statement that it "did not see any of their work through any means until they released it publicly" and that "no specific user data was accessed in order to solve this problem." The same statement added a caveat: "While unlikely, we cannot rule out that de‑identified data derived from their usage of our products helped improve our models." This language, quoted in the Verge piece, acknowledges a theoretical indirect influence but stops short of confirming or denying any direct use of Thom’s or Kun’s unpublished material.
Why the dispute matters now
The timing coincides with OpenAI’s own admission that rumors about its Navier‑Stokes work prompted a compute shift, and that de‑identified user data "may" have contributed to model improvements. That admission, referenced in the commission brief, has already raised concerns about academic data misuse. Thom’s demand for proof adds a concrete request from the research community, moving the conversation from speculation to a request for documentation.
Company background
OpenAI, founded on 11 December 2015, is headquartered in San Francisco and operates in the United States. The firm’s chief executive is Sam Altman, and it reports roughly 4,500 employees, according to Wikidata. The research packet flags the Wikidata figures as background only and advises confirmation against the company’s own filings before publication.
Implications for the AI sector
If OpenAI were to provide verifiable evidence that its models were trained without the unpublished discussions, it could set a precedent for transparency in AI training data practices. Conversely, the lack of such evidence leaves open the possibility that proprietary academic work could be incorporated into large‑scale models without explicit consent, a scenario that could provoke regulatory attention.
What remains unknown
- The specific datasets used to train the models behind the non‑sofic groups breakthrough have not been disclosed.
- OpenAI has not quantified how much, if any, de‑identified user interactions contributed to the model’s performance on the mathematical problem.
- Neither OpenAI nor the mathematicians have provided a timeline for any further exchange beyond the initial email and public statement.
Next steps for stakeholders
Researchers seeking assurance that their unpublished work remains private may request formal data‑usage audits from AI firms. Investors and policymakers, watching the growing intersection of AI capabilities and academic research, may look for clearer guidelines on training‑data provenance. For now, the dispute rests on the absence of documented proof, a gap that the mathematicians argue must be filled.
OpenAI’s non‑committal reply, while denying direct use of specific user data, leaves the door open for indirect influence via de‑identified interactions. The mathematicians’ demand for proof underscores a broader call for transparency that could shape future data‑governance standards in the AI industry.