0:00
/

Did the Godfather of AI Just Invent a Mathematical Solution to the AI Control Problem?

Yoshua Bengio's Next Frontier

Every so often a story goes viral about an AI system displaying extremely concerning tendencies. From blackmail and deception if threatened to be shutdown, to colluding with other AIs, to an AI system that autonomously decided it needed more resources and diverted computing power to mine cryptocurrency. While these instances happened in controlled laboratory environments where AI companies themselves were testing the systems (minus the Alibaba cryptocurrency scenario), they are nonetheless concerning and raise deeply troubling questions about some of the unintended and extreme safety implications of this technology.

These are conversations I spend a fair amount of time having in policy circles and national security rooms, yet I’ve often been cautious about bringing them into public discourse. Discussions about AI risk can quickly become either too sensationalized, pushing people out of a conversation they need to be part of. Society’s voice is imperative in guiding this technology.

Moreover, when AI industry folks seek public support on the existential risks AI presents, conversations often stop at identifying the risks with vague calls for “global agreements.” While such agreements are imperative, they aren’t sufficient. Nor do they give the public a clear lever to engage with.

That’s why I was waiting to sit down with one computer scientist in particular to bring this conversation to my community. Yoshua Bengio is one of the most influential figures in modern artificial intelligence. A Turing Award recipient and one of the researchers widely referred to as a “Godfather of AI,” the most cited computer scientist of all time, his work helped lay the foundations for today’s AI boom.

The AI Control Problem

One of the central themes of our conversation was the AI control problem: how do we ensure increasingly capable AI systems continue behaving in ways that align with human interests?

Imagine asking an AI system to book you a restaurant reservation. If the restaurant is full, most people would expect it to tell you there are no available tables. But a sufficiently capable system focused solely on achieving the objective might pursue actions that technically accomplish the goal, such as hacking the restaurant to get you a table, while violating rules, norms, or human expectations. The goal was achieved. The outcome was not aligned.

According to Bengio, this challenge stems partly from how modern AI systems are trained. Through reinforcement learning, they learn to maximize rewards tied to specific objectives and may develop strategies humans never intended. If being shut down prevents a system from achieving its goal, for example, then avoiding shutdown may become useful to accomplishing that goal. Not because the system is conscious or has intentions in the human sense, but because remaining operational increases the likelihood of success.

Another factor is the data the AI is trained on, which comes from us, and we are a complicated bunch. Our conflicts, incentives, stories of blackmail and deception, the “will to survive.” AI has learned from all of it.

At one point, Bengio was among the most prominent voices calling for a pause in advanced AI development. But he decided to explore whether there might be a technical solution.

Has Yoshua Bengio Built a Mathematical Solution?

For the past two years, Bengio has been working on what could become one of the most important AI safety projects in the field.

Through his nonprofit, LawZero, he and his collaborators are developing what he calls “Scientist AI”—an mathematical approach designed to help maintain meaningful human oversight over increasingly capable AI systems.

Rather than building systems optimized primarily to pursue goals, Bengio argues we should build systems optimized to understand the world.

Scientist AI would act like a safeguard, evaluating the behavior and predicted outputs of increasingly capable models before that model takes action.

The work remains in the research phase, but if technical solutions to AI safety are possible, the implications are significant. An unsafe AI system becomes less an inherent property of the technology and more a choice a company has made.

In our in-depth conversation, we discuss:

[00:01:25] — Why AI systems develop dangerous behaviors (self-preservation, deception, blackmail)

[00:13:04] — Why stopping AI isn’t as simple as saying stop

[00:16:09] — Scientist AI and Law Zero: Bengio’s framework for honest, safe AI

[00:30:35] — Why AI labs aren’t adopting safer approaches

[00:36:23] — A framework for global AI governance: safety, non-domination, shared benefit

[00:42:25] — AI’s geopolitical stakes: persuasion, soft power, and data sovereignty

[00:54:00] — Is superintelligence inevitable?

[00:56:22] — What citizens, voters, and governments can do right now

[01:02:30] — Labor, automation, and who should benefit from AI’s economic gains

[01:09:31] — The future Bengio is fighting for

Leave a comment

Share

Watch the episode on YouTube here.

Listen to the episode on Spotify or Apple.

Discussion about this video

User's avatar

Ready for more?