Research
Mathematicians Demand Transparency from OpenAI on Training Data Sources

Mathematicians Demand Transparency from OpenAI on Training Data Sources

Updated September 10, 2026

Mathematicians are raising concerns about OpenAI's use of unpublished work in training its AI models, particularly in the field of mathematics. Researcher Andreas Thom has accused the company of unethical practices and a lack of transparency regarding the origins of its training data, following a recent controversy over similar allegations. This situation highlights ongoing debates about intellectual property and data usage in AI development.

Reporting notesBrief

Sources reviewed

1

Linked below for direct verification.

Official sources

0

Preferred when available.

Review status

Human reviewed

AI-assisted draft, editor-approved publish.

Confidence

High confidence

85/100 from the draft pipeline.

This AI Signal brief is meant to save busy builders time: what changed, why it matters, and where the reporting comes from.

This story appears to rely mostly on secondary or mixed-source reporting, so readers should treat it as a developing summary rather than a final word. If you spot an issue, email [email protected] or read our editorial standards.

Share this story

0 people like this

Why it matters

  • Developers and product teams must be aware of the ethical implications of using potentially unpublished or proprietary data in AI training, as this could lead to legal challenges.
  • The controversy may prompt stricter regulations around data sourcing for AI models, impacting how teams gather and utilize training data.
  • Transparency in AI training data can enhance trust among users and stakeholders, which is crucial for the adoption of AI technologies in various sectors.

Opening

Recent allegations from mathematicians regarding OpenAI's data sourcing practices have sparked a significant discussion about ethics and transparency in AI development. Researcher Andreas Thom has publicly questioned whether OpenAI's models, particularly in mathematics, have benefited from unpublished work, raising concerns about the ethical implications of using such data without proper acknowledgment. This situation underscores the need for clarity in how AI companies source their training data, especially when it involves intellectual property.

What happened

The controversy began when Andreas Thom, a mathematician, took to social media platform Mastodon to express his concerns about OpenAI's practices. He suggested that interactions he and his colleagues had with the ChatGPT chatbot prior to OpenAI's announcements could have influenced the AI's mathematical capabilities. Thom's accusations come on the heels of another mathematician's claims that OpenAI's models may have utilized unpublished research without consent, leading to accusations of unethical behavior and a lack of transparency.

This situation has raised alarms within the academic community, as it touches on the sensitive issue of intellectual property rights and the ethical use of research in AI training. The mathematicians involved are calling for OpenAI to provide proof that their work was not used inappropriately, highlighting the need for accountability in AI development.

Why it matters

The implications of this controversy extend beyond the immediate concerns of the mathematicians involved. Here are several key reasons why this issue is significant for developers, builders, operators, and product teams:

  • Ethical Data Usage: Developers must consider the ethical ramifications of using data that may be proprietary or unpublished. This situation serves as a reminder to ensure that all training data is sourced responsibly and with proper permissions.
  • Potential Legal Challenges: As the debate over data sourcing intensifies, companies may face legal repercussions if they are found to be using unpublished work without consent. This could lead to costly litigation and damage to reputation.
  • Regulatory Changes: The controversy may prompt regulatory bodies to impose stricter guidelines on how AI companies can source their training data. This could necessitate changes in data collection practices for developers and product teams.
  • Building Trust: Transparency in data sourcing can enhance trust among users and stakeholders. As AI technologies become more integrated into various sectors, maintaining a reputation for ethical practices will be crucial for long-term success.

Context and caveats

While the allegations made by Thom and others are serious, it is important to note that the sourcing of training data for AI models is a complex issue. OpenAI has not publicly responded to these specific claims, and the full extent of the data used in training its models remains unclear. The lack of transparency in AI development is a broader issue that many companies face, and it is not unique to OpenAI.

Furthermore, the academic community is grappling with the challenges posed by AI technologies, including how to protect intellectual property while still fostering innovation. As AI continues to evolve, these discussions will likely become more prominent.

What to watch next

As this situation develops, it will be important to monitor how OpenAI responds to the allegations and whether it takes steps to increase transparency regarding its training data. Additionally, the reactions from the broader academic community and regulatory bodies will be crucial in shaping the future landscape of AI development. Developers and product teams should stay informed about potential changes in regulations and best practices for data sourcing to ensure they remain compliant and ethical in their work.

In conclusion, the ongoing debate surrounding OpenAI's data sourcing practices serves as a critical reminder of the importance of ethics and transparency in AI development. As the industry continues to grow, addressing these concerns will be essential for fostering trust and ensuring responsible innovation.

OpenAImathematicstraining dataethicstransparency
AI Signal articles are AI-assisted, human-reviewed, and expected to link back to source material. Read our editorial standards or contact us with corrections at [email protected].

Comments

Log in with

Loading comments…

Ads and cookie choice

AI Signal uses Google AdSense and similar technologies to understand usage and, if you allow it, request ads. If you decline, we will not request display ads from this browser. See our Privacy Policy for details.