In this interview, Tan Jie, Senior Research Scientist at Google DeepMind Robotics, traces his unlikely path from computer graphics to cutting-edge robotics and offers an insider’s view of two revolutionary paradigm shifts that have reshaped the field.
Tan started his career in animation and graphics, even interning at Pixar, where he simulated lifelike character motion through physics-based AI. He soon realized that controlling a virtual character in simulation is deeply similar to controlling a real robot. By 2018, while at Google Brain, he published a landmark paper on deep reinforcement learning for quadruped robots, pioneering the “sim-to-real” transfer that now powers the agile movements of countless robots. This was the first shift: reinforcement learning replaced intricate manual control models, giving robots a robust “cerebellum” for motion. Suddenly, the acrobatics once exclusive to Boston Dynamics could be replicated by open-source humanoids.
The second shift arrived with large language models (LLMs). Before, robots lacked common sense and couldn’t understand natural language; you had to program each movement. Now, multimodal models like Gemini and GPT can both comprehend instructions and reason about tasks, acting as a robot’s “cerebrum.” Tan’s team recently demonstrated Gemini Robotics 1.5, where the model “thinks” through steps, enabling complex manipulation via simple voice commands. However, these models must still be augmented with robot action data, so they remain extensions of LLMs rather than wholly independent “embodied foundation models.” The thin gap between planning and execution is where most research currently focuses.
Despite rapid progress—fueled by intense competition, massive investments, and Silicon Valley’s new 996 work culture—Tan is sober about the road ahead. The critical bottleneck is data: text is abundant on the web, but diverse, high-quality robot interaction data is extremely scarce. Startups are building data foundries using teleoperation and simulation, but it’s far from enough. Performance is also far from reliable; while simple pick-and-place tasks can reach 90%+ success, delicate actions like zipping a zipper hover around 30–40%, rendering them impractical for real-world deployment. Drawing a parallel with autonomous driving, Tan estimates it will take 2–3 more years to reach a “GPT-3 moment” for robotics, and another 5–10 years for robust commercial products.
Finally, Tan reflects on the philosophical divide between generalist humanoid robots and specialized domain-by-domain solutions. He believes both paths are valid, but predicts that visionary leaders and concentrated investment will likely push the generalist route to dominance, much as general language models supplanted specialized ones. While challenges loom large, Tan remains optimistic: with steady breakthroughs in data, compute, and algorithms, truly capable home and workplace robots may be closer than we think.



![[CVPR'23 WAD] Keynote - Jiyang Gao, Momenta](https://i.ytimg.com/vi/hFQLJIvdQNU/maxresdefault.jpg)