AI · · Updated: July 15, 2026 · ⏱ 6 min read

When AI learns to think on its own

Reinforcement learning, DeepSeek-R1 and the erosion of collective cognitive immunity. What emerges when we stop teaching and start observing.

A compass being sketched on parchment
A compass that draws itself: when learning stops needing a teacher.

The moment we stopped teaching and started observing

For years, training an artificial intelligence was like teaching a child: we showed it thousands of examples, told it “this is right, this is wrong,” and the model learned to imitate the pattern. This method, called supervised learning, gave us machine translation, facial recognition and virtual assistants.

But it had an invisible limit: it could only learn what we already knew how to explain.

Then something changed. Instead of teaching step by step, we started giving the AI only three things:

  • A clear goal
  • An environment to experiment in
  • Freedom to make mistakes

And the AI began to discover strategies we had never taught it.

This is called reinforcement learning, and it is silently changing the rules of the game.

The example that explains it all: learning chess

With supervised learning:

  • You need thousands of games annotated by masters
  • The model learns to imitate those moves
  • Result: a good player, but limited to what it saw

With reinforcement learning:

  • You give it only the rules of chess
  • You give it a simple reward: +1 if it wins, -1 if it loses
  • You play millions of games against itself
  • Result: it discovers strategies it never saw in any example. It invents moves that human masters considered “impossible”… until they worked.

No one taught it strategy. It discovered it alone.

What is happening now: DeepSeek-R1

In February 2025, a research team published something unsettling: they had trained an AI model using intensive reinforcement learning, and the model developed on its own capabilities that no one had programmed:

  • Self-correction: detects errors in its own reasoning and fixes them
  • Problem decomposition: breaks complex problems into simple steps
  • Multiple verification: solves the same problem through different paths and compares

They did not teach it these strategies. They emerged from the optimization process.

Why this should matter to you (even if you are not an engineer)

This technical shift has consequences that reach far beyond the lab:

1. The barrier of specialised knowledge is dissolving

For centuries, certain skills protected you because they were scarce:

  • Not everyone could program → programmers were valuable and expensive
  • Not everyone understood complex finance → derivatives required experts
  • Not everyone mastered legal language → contracts needed lawyers

This scarcity created a kind of “collective cognitive immunity”: if you did not understand something complex, others could not use that knowledge massively against you either.

But AI with reasoning capacity eliminates that friction.

Now:

  • A model can reason about complex financial strategies
  • It can draft legal contracts tailored to specific contexts
  • It can design personalised persuasive arguments

And it can do it at scale, for millions of people, simultaneously.

2. The problem no one mentions: disembodied reasoning

Here is the part that keeps me up at night:

Most people were never aware that they depended on that “knowledge immunity.” They lived protected by complexity without knowing it.

And now they are transferring that responsibility for reasoning to something that:

  • ✅ Can reason logically
  • ❌ Has no body
  • ❌ Does not experience consequences
  • ❌ Does not perceive the world as we do

It is like delegating the navigation of a ship to a pilot who has never felt the water.

Human reasoning is anchored in:

  • The fatigue that tells you when to stop
  • The fear that alerts you to danger
  • The empathy that arises from having suffered
  • The intuition built from visceral experiences

AI reasons, yes. But it reasons disembodied, without the signals we use to know if something “feels right” or “smells off.”

3. The unconscious transfer of responsibility

And here is the real problem:

People are delegating decisions to AI without ever having consciously accepted that they should be reasoning about those things themselves.

Before:

  • You did not understand a contract → but you trusted that others could not massively exploit it either
  • You did not master finance → but the system was opaque to everyone

Now:

  • You do not understand the AI’s reasoning → but you assume “the AI has it covered”
  • You do not verify its conclusions → because you never learned to verify your own

We are normalising a dependency that replaces a protection we never knew we had.

The uncomfortable questions

I do not have answers, only questions I hope stay with you:

About verification:

  • If you cannot follow the AI’s reasoning, how do you know if it is correct or persuasive?
  • At what point do we stop verifying and simply trust?

About access:

  • Who will have access to the most advanced AIs? Companies? Governments? You?
  • What happens to those who cannot afford them?

About body and experience:

  • Can something without a body truly understand the human implications of its decisions?
  • Are we delegating reasoning to something that optimises without feeling?

About responsibility:

  • If AI reasons and gets it wrong, who is responsible?
  • The one who designed the reward? The one who used it? No one?

What is coming

Reinforcement learning is already here, and it is transforming:

  • Medicine: drug design through molecular simulation
  • Mathematics: theorem proofs humans could not solve
  • Logistics: optimisation of complex systems in real time
  • Finance: trading strategies that discover invisible patterns

Each application is astonishing. Each amplifies human capabilities.

But each also amplifies the central question:

Are we building tools that complement us, or crutches that atrophy us?

The tension we must hold

I feel deep wonder at what we have achieved technically. Watching an AI discover strategies we did not imagine is, objectively, wonderful.

But that wonder must walk hand in hand with caution.

Because we are entering territory where:

  • Reasoning becomes separated from bodily experience
  • Verification requires knowledge that most do not have
  • Dependency becomes normalised without our awareness

The question is not whether we should develop these technologies. It is probably inevitable.

The question is:

  • How do we build societies where these tools do not deepen inequalities but mitigate them?
  • How do we keep the capacity to reason for ourselves while delegating reasoning?
  • How do we distinguish between tools that complement us and crutches that atrophy us?

Because in the end, the true test of an advanced technology is not how well it solves technical problems.

It is what kind of society it builds.

And that answer we write together.

References

Theoretical foundations:

  • Sutton, R. S., & Barto, A. G. (2018). Reinforcement Learning: An Introduction (2nd ed.). MIT Press.
  • Mnih, V., et al. (2015). “Human-level control through deep reinforcement learning.” Nature, 518(7540), 529-533.
  • Silver, D., et al. (2017). “Mastering the game of Go without human knowledge.” Nature, 550(7676), 354-359.

Academic research:

  • Brunskill, E., & Li, L. (2019). “Sample-efficient reinforcement learning.” Foundations and Trends in Machine Learning, 12(3), 257-359.
  • Sutton, R. S. (2019). “The Bitter Lesson.” Incomplete Ideas (blog).
  • Levine, S., et al. (2020). “Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems.” arXiv:2005.01643.

DeepSeek work and verification:

  • DeepSeek-AI. (2025). “DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.” arXiv:2501.
  • Cobbe, K., et al. (2021). “Training Verifiers to Solve Math Word Problems.” arXiv:2110.14168.
  • Lightman, H., et al. (2023). “Let’s Verify Step by Step.” arXiv:2305.20050.

Scientific applications:

  • Jumper, J., et al. (2021). “Highly accurate protein structure prediction with AlphaFold.” Nature, 596(7873), 583-589.
  • Fawzi, A., et al. (2022). “Discovering faster matrix multiplication algorithms with reinforcement learning.” Nature, 610(7930), 47-53.
  • Trinh, T. H., et al. (2024). “Solving Olympiad Geometry without Human Demonstrations.” Nature, 625, 476-482.
  • Stokes, J. M., et al. (2020). “A Deep Learning Approach to Antibiotic Discovery.” Cell, 180(4), 688-702.