What Happens When the AI Slop Is Correct?
As AI produces correct work faster than people can understand it, what keeps humans meaningfully involved in research, software and shared knowledge?
What happens when the AI slop is correct? As models get better, their outputs are becoming increasingly ambitious. At the same time, they are not necessarily becoming more understandable. Math proofs resolving longstanding unsolved problems are within reach. The range of software you can generate with frontier models is growing. Yet, as generation accelerates, we risk producing more than we can understand and build upon. How can people retain a meaningful role in this process?
The researcher’s dilemma
In August of 2026, Levent Alpöge and Tristan Buckmaster were collaborating on novel math research. Then, around September 3rd-6th, they realized that OpenAI was concurrently working on a related Navier–Stokes result and formalization. According to Tristan’s account, they felt pressured to publish due to the imminent OpenAI result, despite missing important work to improve the quality of the papers. In Tristan’s statement, he says:
I am not happy about the presentation quality in these papers … The Euler writeup, in particular, can only be described as AI slop. I am sorry for this. The reasons are below, and they involve our being pressured by outside factors.
Although their results included Lean verification and public formalization, they were not easily digestible and left a significant burden on the reader. Even in the best circumstances where researchers and labs are acting in good faith, the incentives are aligned towards faster generation and sharing of frontier results. With sufficiently capable models, maintaining recognition for novel work incentivizes secrecy and rushing to publication (lest a much smarter AI get to it first).
Only two weeks later, it looks like OpenAI is sitting on many unpublished math proofs. Suppose, for the sake of argument, that OpenAI does publish 100 results which are largely terse Lean proofs somewhere in the AI slop spectrum. How much would they advance math? In theory, the number of math results would continue growing rapidly. Some of the results could be integrated into other research and applications without deep understanding. However, it feels like many of these proofs would just sit in the ether, causing substantial cognitive burden rather than empowering mathematicians.
It’s also very difficult to ignore the results: researchers working on related areas would likely be expected to reconcile the AI results with their own work, plus ignoring them could mean falling behind colleagues and AI itself. If the ability for AI models to generate novel proofs outpaces their ability to explain them and make them easily accessible, we are left with the impossible task of integrating the AI firehose into our collective knowledge. This could create a feedback loop: AI produces results faster than people can absorb them, then continues building on those results to push further ahead. Even as human understanding grows, more of the frontier could become accessible only through AI, with people increasingly dependent on machines to interpret and extend it. As Henry Cohn put it:
When you make a discovery, the knowledge doesn’t transfer seamlessly to everyone else … It can take a shockingly large amount of work to document your discovery in a form that is truly useful for others, and even after all that work, the collective effort required for the community to master and internalize it is enormous …
It’s clear by now that change is coming to the mathematics field. As the cost of AI proofs goes down, many will question the viability of a career in math. If recognition and viable careers remain concentrated on producing original results, faster machine discovery could make human contributions harder to sustain professionally. At the same time, AI can give many outside the field the ability to meaningfully contribute results and digest others’ work. In my opinion, math leaders such as Terence Tao and Henry Cohn are right to be concerned about the negative impacts of unstructured dumps of AI math generations on the community. To me, a world where vast stacks of proofs are sitting in some AI data dump while many mathematicians move on to different fields should actively be avoided through deliberate action.
Correct software slop
As a software engineer, I have way more experience dealing with a similar problem. As the volume of AI code grew, the burden on me and other PR reviewers became untenable. At some point, I could no longer review every single line of code, so I started relying on AI workflows. Pretty soon, it felt like several pieces of the architecture were no longer well understood by engineers. The code could be correct, and AI agents perfectly capable of making edits to keep things going, but keeping up rapidly became painful. One effect was that intentional architecture design also tended to decrease, usually because engineers were getting burnt out and were spending most of their time accelerating code changes and reviews with AI, rather than taking a step back and intentionally designing the high-level pieces of the system.
Another experience I’ve had is the difference between onboarding onto a new codebase with the help of mentors, versus onboarding exclusively with AI. When I onboarded onto Edge/Chromium, I had the privilege of having a few amazing coworkers who, through informal conversations and PR comments, taught me a lot of tribal knowledge and practices. They helped me become a strong contributor in a short period of time. The experience was completely different onboarding to a Rust codebase with AI, when no experts were available to help. The kind of comments agents made on my PRs were focused on fixes rather than teaching. I learned a lot about paranoid defensive coding, but not so much about good Rust practices and ways of maintaining a healthy codebase. As we prioritized code velocity, I had less time and fewer good examples to learn from. While these were very different circumstances, they showcase the benefits of intentional mentorship and the fact that AI does not provide it automatically.
Here is how Simon Willison puts it in his Cognitive Debt piece:
… I’ve been experimenting with prompting entire new features into existence without reviewing their implementations … I no longer have a firm mental model of what they can do and how they work, which means each additional feature becomes harder to reason about, eventually leading me to lose the ability to make confident decisions about where to go next.
It’s not just that writing and reviewing code by hand becomes more difficult. Even prompting for new features with high levels of detail becomes harder. We are offloading the implementation without retaining or developing enough understanding to direct it.
So what? If we can accomplish our goals, does it matter if only AI can make sense of the codebase? One issue with this is that you cannot hold an LLM accountable. The responsibility for deploying software still rests with people and organizations. Even if AI becomes capable of providing strong guarantees about its outputs, those accepting responsibility need enough understanding to evaluate them. Thus, people need to be able to understand and make informed decisions on the major components, their interactions, design rationale, and constraints over time. I believe that deep understanding of the codebase makes it easier to architect its future and articulate how it should evolve. An analogy would be a tech company sub-contracting their core product to a vendor firm. They hand product requirements to the vendor, but their incomplete understanding of the system limits what they can request or dispute. Bridging the knowledge gap is possible, but not free.
There is another side of the coin though. AI is making participation in complex projects much easier, even for those who do not consider themselves technical. One great example is the Linux distribution Omarchy, a project by DHH which is designed to facilitate customization via agentic prompting. I think acceleration of progress with AI is great, and the spirit of expanding access to building technology to as many people as possible should be the goal. Readable code, inspectable architecture, and faithful explanations help achieve this in the broadest sense, letting people participate at various levels of abstraction.
The incentive problem and practical solutions
Math and software seem to be running into a similar incentive problem. Publishing a novel result or shipping a feature brings immediate, visible rewards. Helping people understand the work takes additional effort, and the benefits may only become apparent much later. AI makes it possible to accelerate the first part without necessarily investing in the second. Even people who care deeply about teaching and maintaining shared understanding can feel pressured to keep shipping when their peers and competitors are moving faster. My concern is that individually reasonable decisions could, at the limit, collectively lead us toward a world where people can no longer meaningfully participate in the pursuit of knowledge.
Improving capabilities of models does not necessarily improve human understanding. For example, the DeepSeek-R1 paper describes R1-Zero achieving substantial reasoning improvements while exhibiting poor readability and language mixing in its reasoning traces. For R1, the researchers deliberately introduced readable training examples and a reward for language consistency. They report that the latter slightly reduced task performance while improving readability and alignment with human preferences. Similarly, in its Opus 5.5 announcement, Anthropic explicitly identified communication as an improvement over Opus 5, responding to user feedback. We should continue explicitly training for readability, faithfulness, and human learning. An important measure of success is whether users can subsequently apply a concept, identify a mistake, or make an informed modification. Outputs should not only be correct, but also understandable and empowering to people.
Even with perfect communication between AI and people, the sheer volume of output will make it impossible to fully digest everything. As AI becomes capable of generating correct and novel outputs much faster than us, it could exceed even a specialist community’s ability to integrate them. We should preserve our ability to choose what matters, acquire relevant understanding, and meaningfully direct or challenge the work. We can do this in two complementary ways:
- Abstractions can let us architect and understand the high-level components of artifacts like code and proofs, without the need to verify or understand all the lower-level implementation details. While this is imperfect, it provides a scalable way for us to keep contributing. For example, the new programming language Bend allows stating requirements as formal laws, with machine-checked proofs that the implementation satisfies them. This lets programmers control the requirements of the codebase which constrain the kinds of changes AI agents can make. Deciding whether the specifications capture what actually matters remains a human responsibility.
- Institutions designed to reward integration of AI-generated content into human knowledge can strengthen our ability to keep up. The key is expanding what earns credit: turning an isolated result into shared, usable knowledge deserves recognition comparable to discovering it. From an outsider perspective, it seems like this is in line with the goals of the Advisory Group on Mathematics and Artificial Intelligence, for example. A funded fellowship could reward researchers for explaining, simplifying, and connecting AI-generated mathematical results, with success measured by whether their work helps others understand and build on those discoveries. For software, companies should consider budgeting time and resources for mentorship, knowledge sharing, and cleanup of code. Although these practices have immediate costs, their value accrues long-term. Neglected understanding makes future decisions and participation harder, accumulating as cognitive debt.
Abstractions can give people manageable decisions, and education and practice equip them to make those decisions. At the same time, AI can become increasingly capable of explaining, organizing, and teaching. With the combination of these approaches, we are in a much better position to not only understand frontier results on demand, but systematically integrate them into our collective knowledge. Through deliberate effort, we can build a world where AI not only accelerates human progress, but also empowers people to understand, question, and shape what AI creates.