Nouveau : AI international speaker


A world model is an AI that learns how the physical world behaves, just by watching video. It does not learn language. It learns things like: if you drop a glass, it breaks. If you push a box, it slides. Once it has learned enough of these rules, it can predict what happens next in a situation it has never seen before.

This is different from ChatGPT, Claude, or Gemini. Those tools predict the next word in a sentence. A world model predicts the next moment in a scene, like a simulation running in a video game.

In 2026, this became one of the biggest bets in AI.

Yann LeCun, one of the most respected AI scientists alive, left Meta to raise $1.03 billion for a lab built entirely around this idea.
Fei-Fei Li, the Stanford professor often called the « godmother of AI, » raised a billion dollars for the same bet. Investors put over $3 billion into world model startups in the first six months of 2026 alone.

Here is the blunt version: chatbots are great at sounding smart. World models are being built because sounding smart and being physically right are two different skills, and robots, self-driving cars, and factories need the second one.

Why the DragonBall Comparison Fits

In DragonBall, Goku never stops training after his first big transformation. Each new form that once looked like the final upgrade turns out to be a stepping stone. The next fight always needs a power nobody trained for yet.

Chatbots are that first transformation. They got AI very good at language.

But language alone hit a wall: a chatbot can describe gravity perfectly in a sentence and still have zero idea what actually happens when something falls. World models are the next form, built on watching and predicting the real world instead of reading about it.

What A World Model Actually Does

A world model watches something, like a video of a room or a robot’s camera feed, and builds an internal picture of how that world works. Not a description in words. A working simulation it can run forward in its « head » to guess what happens next.

Three things separate a real world model from a fancy video generator.

First, it stays consistent:
if you look away from an object and back, it is still there, in the same place.
Second, it reacts: an action you take inside it leads to a believable result.
Third, what it learns can transfer to a real robot or vehicle, not just stay inside the simulation.

Most products calling themselves « world models » in 2026 fail at least one of these three tests. If a system cannot keep a scene stable for more than a few seconds, it is a video generator with better marketing, not a world model.

Why LeCun and Fei-Fei Li Are Betting Billions on This

Yann LeCun’s argument is simple: language alone will never produce a truly intelligent AI, because reading about the world is not the same as experiencing it. He made this argument for years while running Meta’s AI lab. Now he is backing it with $1.03 billion of funding for his new company, Advanced Machine Intelligence Labs, per Fortune, May 2026, the largest first funding round for an AI company in European history.

LLMs are « not the path to real intelligence. They’re a detour, » LeCun told Bloomberg’s « The Close » on May 21, 2026, predicting they will become « largely obsolete » across most applications within five years. Yann LeCun, per Crypto Briefing

Fei-Fei Li’s version of the same bet is called « spatial intelligence, » the idea that AI needs a real sense of 3D space. Her company, World Labs, launched a product called Marble in late 2025. It builds downloadable 3D worlds and costs between $35 and $95 a month.

« They remain wordsmiths in the dark; eloquent but inexperienced, knowledgeable but ungrounded. » Fei-Fei Li, on large language models, in « From Words to Worlds »

Neither of them says chatbots are useless. Both say chatbots hit a ceiling: they are excellent with words, but they have never touched, dropped, or bumped into anything. More text and more computing power will not fix that gap. It needs a different kind of training entirely.

Why A Four-Year-Old Beats The Biggest LLM

At an ETH Zürich lecture in May 2026, LeCun made his argument concrete with one comparison. A four-year-old child, just from ordinary seeing, touching, and moving around, has absorbed more raw sensory data than the biggest language models trained on the entire internet’s text. He calls this Moravec’s paradox: tasks a toddler finds trivial, like understanding gravity or walking across a room, remain some of the hardest problems in AI, while tasks that feel advanced to us, like passing a bar exam, turn out to be comparatively easy for a language model, per his talk covered by StartupHub.ai, June 2026.

His fix is what he calls an objective-driven architecture: separate modules for perception, memory, a world model, and an action planner, all working toward explicit goals instead of just predicting the next word. Inside it sits JEPA, Joint Embedding Predictive Architecture, which predicts an abstract representation of what happens next rather than generating it pixel by pixel, the same reason text prediction stays sharp while video prediction turns blurry. LeCun compares the result to Model Predictive Control, a decades-old robotics technique, run at a scale nobody has attempted before.

The Three Companies Racing to Build This

Google DeepMind, World Labs, and NVIDIA are not building the same thing, even though headlines lump them together. Knowing the difference matters if you ever plan to use one of these tools.

Genie 3, from Google DeepMind, generates a video world in real time that responds to your actions as you explore it, like walking through a video game that AI is inventing as you go. You cannot download the world it makes; you can only visit it through Google’s system.

Marble, from World Labs, makes a 3D world you can actually download and reuse, built with a technique called Gaussian Splatting (a way of representing 3D space using millions of tiny colored dots instead of solid shapes).

NVIDIA Cosmos takes a third approach: it is free and open for companies to run on their own computers, and it was trained on 20 million hours of real footage from cars, factories, and robots.

The simple way to remember the split: Genie 3 and Marble are closed tools for creators. Cosmos is open infrastructure for robotics and self-driving car teams who need to train at scale. NVIDIA says Cosmos has been downloaded 2 million times, per MIT Technology Review, April 2026, and reportedly rattled OpenAI internally over how far ahead this gives NVIDIA on physical AI.

Where This Already Works Today, And Where It Doesn’t Yet

A world model earns its money the moment it can safely simulate something too dangerous or too rare to test in real life. Self-driving car teams use this constantly: black ice, a child running into the street, a strange intersection, all created inside a simulation instead of waited for on real roads.

Robotics is the second place this already works. NVIDIA trains robots by showing them millions of hours of human video, then transferring what the AI learned onto the robot’s own body through extra training. This turns a training process that used to take months of real-world trial and error into weeks of simulation.

But there is a real limit nobody puts in the demo video. Small errors in the simulation build up over time, called « drift, » until the predicted world stops matching anything physically possible. These models also struggle outside the specific environments they were trained on, and building a separate world model for every new use case costs real money.

Anyone selling you a world model as a total replacement for your entire AI setup is overselling it, the same way « AGI by 2025 » was oversold two years ago. The honest version of this technology is narrower than the pitch, and it is still worth billions.

The Money Behind the Bet

The amount of money moving into this space is the real proof that this is not just researchers being excited.

Per Forbes, June 2026, investors put over $3 billion into world model startups in just the first half of 2026. The wider market these tools support, sometimes called « physical AI, » is expected to grow from around $110.8 billion in 2026 to $960.4 billion by 2033, per MarketsandMarkets, 2026.

There is also a smaller, fast-growing market just for the synthetic training data these models produce. That market was worth $2.03 billion in 2025 and is projected to reach $63.95 billion by 2035, per Kaiso Research, 2026. In plain terms: a huge amount of money is betting that fake, AI-generated experience can replace real-world testing at a fraction of the cost.

To put LeCun’s funding round in perspective, $1.03 billion is more early-stage money than many entire countries raise for AI in a full year. This is no longer a research curiosity. It is one of the biggest destinations for AI capital in 2026, and most business leaders are not even tracking it yet.

World Model Products Compared

Here is the simple version: these four efforts solve different problems, so don’t compare them like they’re competing on the same checklist.

ProductCompanyWhat It MakesCan You Use It Yourself?Best For
Genie 3Google DeepMindA real-time video world you exploreNo, closed to the publicImmersive demos and creative experiences
MarbleWorld Labs (Fei-Fei Li)A downloadable 3D worldLimited, $35 to $95 a monthBuilding and reusing 3D environments
CosmosNVIDIATools to predict, test, and simulate physicsYes, free and open, run on your own serversTraining robots and self-driving cars
AMI Labs (LeCun)Advanced Machine Intelligence LabsFoundational research, no product yetNo, not commercial yetLong-term research into general intelligence

The real story in this table: the two tools people can actually buy today, Genie 3 and Marble, are closed and built for creators. The one with the clearest path to serious enterprise revenue, Cosmos, is the one that’s open. Calling any single one of these « the winner » this early misses that they are three different businesses wearing the same label.

I believe in this shift more than almost anything else happening in AI right now. It is the first change since the early GPT days that makes me want to build hands-on again, not just write about it. If you run a technical team and want to get ahead of this before it gets priced into everything, that is exactly the kind of work I run through the Claude Sprint: map the real use case first, then commit to a stack.

FAQ

Q: What is a world model in AI, in simple terms?

A: It’s an AI that learns how the physical world behaves by watching video, then predicts what happens next when something moves or changes. It simulates reality instead of just describing it in words.

Q: How is a world model different from ChatGPT?

A: ChatGPT predicts the next word in a sentence. A world model predicts the next moment in a physical scene, the way a video game engine predicts what happens when you move a character.

Q: Is this actually worth investing in for a normal business right now?

A: Only if you work in robotics, self-driving vehicles, or heavy simulation. For most everyday business AI use, like writing, customer service, or data analysis, this technology is not relevant yet.

Q: Why do companies get world models wrong?

A: They treat every « world model » as the same thing and pick based on hype instead of what it actually does. A closed creative tool like Marble solves a completely different problem than open infrastructure like Cosmos.

Q: Will world models replace chatbots like ChatGPT or Claude?

A: No. In 2026, most researchers agree these are two different specialties, not a replacement. Chatbots stay best at language and reasoning with words. World models matter when you need to predict a physical outcome.

Q: What is Yann LeCun building now?

A: Advanced Machine Intelligence Labs, started after he left Meta in late 2025. It focuses entirely on world models as a path to smarter AI, backed by a $1.03 billion first funding round.

Q: What is the biggest weakness of world models today?

A: Errors build up over time. A small mistake early in a simulation snowballs until the predicted world stops making physical sense. That, plus struggling outside familiar environments, remains unsolved as of mid-2026.

The Verdict

World models are not a chatbot upgrade. They are a completely separate race, with more money poured in over six months than most AI research areas see in five years, led by two of the most respected scientists in the field betting against the mainstream. If your AI plan in 2026 still begins and ends with « which chatbot should we use, » you are optimizing for the wrong cycle.

Pick a lane, physical simulation or language tools, and stop assuming one will quietly swallow the other.