CoRL 2026 Workshop

Do Robots Need World Models?

A Debate on Dynamics Learning, Planning, and End-to-End Robot Learning

CoRL 2026 · Half-day workshop · Date & location TBD

Overview

World models are rapidly becoming a central paradigm in robot learning. Broadly, they are predictive systems that let robots anticipate how the world may change under their actions — spanning dynamics, contact, geometry, and physics-based models, as well as latent predictors, video models, learned simulators, and world–action models. Yet their growing impact has exposed an unresolved tension: should predictive structure be made explicit to support planning, counterfactual reasoning, data efficiency, safety, and generalization — or can it be absorbed implicitly into end-to-end policies that map observations and goals directly to actions?

World models offer reusable structure but may introduce brittleness, hallucinated predictions, modeling errors, slow inference, or unnecessary complexity, while imitation, diffusion, and vision–language–action (VLA) policies suggest some of this structure can be learned from data without exposing a separate model for prediction or planning. This workshop turns that tension into a structured scientific debate — asking not whether world models are necessary in the abstract, but when, why, and in what form they help, where they introduce brittleness, and when end-to-end policies are the better choice. Our goal is to turn the debate into concrete design principles for choosing among explicit world models, implicit predictive policies, and hybrid systems.

The question cuts across reinforcement learning, control, computer vision, simulation, and embodied AI. The workshop is built for a broad audience: newcomers who want a grounded introduction to world models, and experts from academia and industry who want to debate when and why to use them — and confront their limits.

The Six Motions

Rather than surveying world-model approaches broadly, the workshop stress-tests six concrete claims shaping the field today. Each is used as a debate motion: invited speakers defend assigned sides, the audience votes before and after, and a community white paper consolidates the outcome.

01

End-to-end policies can internalize the predictive structure, with learned abstractions, needed for robot behavior.

Visuomotor, diffusion, and VLA models may already embed action-conditioned knowledge about dynamics without exposing it for planning.

Open problem: when is implicit structure sufficient, and when does its absence cause brittleness under unseen objectives, novel objects, longer horizons, distribution shift, or failure recovery?

02

Explicit prediction of dynamics benefits generalization to diverse goals.

Learning world models provides an independence between dynamics and task objectives, supporting planning for novel goals and test-time scaling of compute.

Open problem: what characteristics of a downstream task demand an explicit model or make it better suited to be learned from end-to-end training — and how should this be quantified?

03

General-purpose policies are easier to learn than general-purpose world models.

A policy need only retain information relevant for action selection, whereas a reusable model demands broader counterfactual coverage and faces challenges in defining a state abstraction.

Open problem: is it easier to scale learning for general-purpose policies, or can one tractable model abstraction support broad reuse and learning from broader data?

04

Explicit models enable reproducible policy evaluation, supporting policy improvement and safe deployment at scale.

Policy evaluation drives policy improvement and is critical for assessing behavior at deployment time, yet real-world evaluation may be costly or dangerous.

Open problem: when does reproducible model-based evaluation preserve real-world policy rankings and improvement directions — or introduce a new failure mode: trusting a model of the world that is itself wrong?

05

The real question is not whether to use world models, but where the model should live.

Predictive dynamics can appear as a deploy-time planner, a learned simulator, a training-time data generator, a pre-trained policy representation, a safety monitor, or an evaluation tool.

Open problem: which role fits which regime, given that a model useful for pretraining may fail at online control, and image prediction may not support contact-rich manipulation?

06

Current benchmarks cannot tell us whether world models are necessary.

Visual prediction quality is insufficient, and task success alone hides differences in sample efficiency, robustness, and generalization.

Open problem: how do we design comparisons under matched data, compute, and deployment constraints — covering distribution shift, long horizons, contact-rich interaction, failure recovery, and real-world transfer?

Invited Speakers

The speakers are balanced across the debate. Following the workshop's format, each speaker is assigned a motion to defend — steelmanning the opposing view first, and closing with one concrete failure case of their own paradigm — so both sides are argued vigorously, independent of personal lean.

Chelsea Finn

Confirmed

Stanford University / Physical Intelligence

Against world models · implicit / end-to-end

Speaking to Motions 1 & 3: work on imitation learning and generalist robot policies argues that robust behavior emerges from large-scale learning, with predictive structure embedded implicitly and increasingly absorbed by scale.

Vincent Sitzmann

Confirmed Possibly online

Massachusetts Institute of Technology / Rhoda AI

Against explicit world models · learned / generative

Speaking to Motions 1 & 3: work on neural scene representations and generative models argues the structure needed for embodied behavior can be learned directly from data at scale, rather than imposed by hand-structured models.

Yunzhu Li

Confirmed

Columbia University

For world models · explicit

Speaking to Motions 2, 4 & 5: research on structured world models, physical scene understanding, and deformable-object manipulation offers a perspective on when explicit, queryable predictive structure is needed for planning, evaluation, and generalization.

Georgia Chalvatzaki

Confirmed

Technical University of Darmstadt

For structured world models · hybrid

Speaking to Motions 5 & 6: work bridging learned representations and classical planning argues for structured, model-based components over monolithic end-to-end policies — while clarifying where a model should live and how we would measure whether it helped.

Schedule

Half-day program. Invited talks are 20 min + 5 min Q&A. Times are tentative and will be finalized once the CoRL 2026 program is set.

TimeDurationSession
08:50–09:0010 minWelcome & opening vote on the six motions
09:00–09:2525 minInvited talk: Chelsea Finn implicit / end-to-end
09:25–09:5025 minInvited talk: Yunzhu Li explicit world models
09:50–10:2030 minContributed spotlights
10:20–11:0040 minCoffee break & poster session
11:00–11:2525 minInvited talk: Georgia Chalvatzaki structured / hybrid
11:25–11:5025 minInvited talk: Vincent Sitzmann implicit
11:50–12:4050 minStructured debate & closing vote
12:40–12:5010 minAwards & closing remarks

Call for Papers

We solicit two kinds of contributions via OpenReview — standard research papers and short position papers arguing for the relevant topics. We explicitly invite in-progress work, head-to-head comparisons, and negative or surprising results (a paradigm that failed where it was expected to win) — exactly the evidence the debate needs. We only accept submissions that have not been published elsewhere or accepted to the main CoRL 2026 conference.

Reviewing follows the CoRL main-track reciprocal model: each submission nominates one author to review another, ensuring at least two reviews per paper, with an organizer serving as editor. Accepted work appears as spotlights and posters, with an optional short video. Two targeted awards recognize both empirical and argumentative contributions: Best Research Paper and Best Position Paper.

Topics include: explicit and latent world models for control and planning; end-to-end, diffusion, and VLA policies; hybrid model-based / model-free systems; counterfactual and long-horizon reasoning; uncertainty and safe deployment; and protocols benchmarking explicit vs. implicit approaches under data and compute.

Important Dates

All dates are placeholders and will be confirmed alongside the CoRL 2026 schedule.

Paper submission deadlineTBD
Notification to authorsTBD
Camera-ready deadlineTBD
Workshop dayCoRL 2026 · TBD

Organizers

Alberta Longhini
Alberta Longhini Stanford University alberta@stanford.edu
Wenlong Huang
Wenlong Huang Stanford University wenlongh@stanford.edu
Bardienus P. Duisterhof
Bardienus P. Duisterhof Carnegie Mellon University bduister@cmu.edu
Lasse Peters
Lasse Peters UC Berkeley lasse.peters@berkeley.edu
Holly Dinkel
Holly Dinkel Samsung Research America holly.dinkel@gmail.com
Jeffrey Ichnowski
Jeffrey Ichnowski Carnegie Mellon University jeffi@cmu.edu
Fei-Fei Li
Fei-Fei Li Stanford University — World Labs feifeili@stanford.edu