About Yilun Du
I develop compositional generative reasoning methods that enable physical agents to solve tasks beyond those encountered during training.
My work studies how generative models, composition, and search can serve as general-purpose mechanisms for solving complex unseen tasks. We study these techniques across scene understanding, natural language reasoning, and robotics.
My long-term goal is to build compositional world models that enable agents to reason, plan, and act reliably in unfamiliar physical environments.
I am an Assistant Professor at Harvard's Kempner Institute and in Computer Science, where I lead the Embodied Minds Lab. I am also a part-time research scientist at NVIDIA Research. I received my PhD from MIT EECS, advised by Leslie Kaelbling, Tomas Lozano-Pérez, and Joshua B. Tenenbaum. Before joining Harvard, I was a research fellow at OpenAI and a senior research scientist at Google DeepMind.
I explore how iterative inference and search can be combined with generative models to solve complex tasks. My work spans energy-based, diffusion, and language models across visual generation, sequence prediction, and reasoning.
Representative work: deep EBMs, equilibrium matching, and reasoning with sampling.
I study how learned models, objectives, and constraints can be composed at inference time to solve unfamiliar reasoning and planning problems. I apply this idea to abstract reasoning, long-horizon trajectory generation, and decision-making.
Representative work: compositional energy minimization, Diffuser, and generative trajectory stitching.
I build generative world models that predict possible futures, then use inference and planning over those futures to extract actions. My work uses large video and trajectory models for long-horizon planning, robot control, and generalization across tasks and environments.
Representative work: Diffusion Policy, UniPi, Video Language Planning, and Large Video Planner.
I study how multiple language models can be composed into an intelligent system. Their interactions enable inference-time reasoning through proposing, verifying, recombining, and refining solutions.
Representative work: multi-agent debate, multi-agent verification, and Economy of Minds.
Common thread. By composing learned models and searching for solutions at inference time, we can build intelligent systems that generalize reliably to new tasks, goals, and environments.