About Yilun Du
I develop generative reasoning methods that enable agents to solve tasks beyond those encountered during training.
My work treats sampling, optimization, and search as general-purpose mechanisms for reasoning. I study how energy-based, diffusion, and world models can be composed to meet new objectives and constraints across visual generation, planning, and robotics.
My long-term goal is to build compositional world models that can reason, plan, and act reliably in unfamiliar physical environments.
I am an Assistant Professor at Harvard's Kempner Institute and in Computer Science, where I lead the Embodied Minds Lab. I am also a part-time research scientist at NVIDIA Research. I received my PhD from MIT EECS, advised by Leslie Kaelbling, Tomas Lozano-Pérez, and Joshua B. Tenenbaum. Before joining Harvard, I was a research fellow at OpenAI and a senior research scientist at Google DeepMind.
Research
Generative inference for reasoning and action
My research studies how generative models can solve problems beyond their training distribution by composing learned models with new goals and constraints, then using inference-time computation to find solutions. I develop this idea across generative modeling, reasoning and planning, embodied world models, and multi-agent intelligence.
I develop generative models that represent spaces of possible solutions and use iterative inference to find them. My work spans energy-based, diffusion, and language models across visual generation, sequence prediction, and reasoning.
Representative work: deep EBMs, equilibrium matching, and reasoning with sampling.
I study how learned models, objectives, and constraints can be composed at inference time to solve unfamiliar reasoning and planning problems. I apply this idea to abstract reasoning, long-horizon trajectory generation, and decision-making.
Representative work: compositional energy minimization, Diffuser, and generative trajectory stitching.
I build generative world models that predict possible futures, then use inference and planning over those futures to extract actions. My work uses large video and trajectory models for long-horizon planning, robot control, and generalization across tasks and environments.
Representative work: Diffusion Policy, UniPi, Video Language Planning, and Large Video Planner.
I study how multiple language models can be composed into an intelligent system. Their interactions provide inference-time computation for proposing, critiquing, verifying, and coordinating solutions.
Representative work: multi-agent debate, multi-agent verification, and Economy of Minds.
Common thread. By composing learned models and searching for solutions at inference time, we can build intelligent systems that generalize reliably to new tasks, goals, and environments.