This video is from the podcast 'Zhang Xiaojun Business Interview', where hosts Zhang Xiaojun and Guangmi talk with Yao Shunyu, a researcher at OpenAI. Yao graduated from Tsinghua's Yao Class and Princeton, and started researching LLM-based agents early in his PhD, now six years in. In April 2025, he published an influential blog post 'The Second Half', declaring that the main thread of AI has entered its second half. The interview explores his personal journey, the history and future of AI agents, and the central role of language.
Yao's core argument is that language is the most essential tool for achieving AGI, because almost everything can be represented through language. He reviews three waves of AI agents: first, symbolic AI, which built agents by hand-crafting rules but hit bottlenecks—rules couldn't cover all cases and couldn't generalize across tasks, leading to the first AI winter. Second, deep reinforcement learning, exemplified by AlphaGo, which learned through massive trial and error in specific environments but couldn't generalize to new ones and required extensive environment-specific engineering. Third, the LLM era, where models gain rich linguistic priors through pretraining, enabling reasoning and thus generalization to new tasks. He believes reasoning is the key differentiator that allows agents to generalize, unlike previous waves.
Yao shares his research journey, emphasizing the importance of 'non-consensus' and 'simple and general' approaches. He realized early that the environment matters, so he created WebShop, a simulated e-commerce environment for agents. He then proposed the ReAct framework, combining reasoning and action, which became one of the most general agent architectures. He recalls that the academic community at the time focused on training models (like BERT), but he believed studying how to use models (e.g., via prompting) was more valuable, because training models often lagged behind companies like OpenAI by years. His favorite work is ReAct because it's simple, general, and marked a paradigm shift from 'training models' to 'using models'. He also notes that agent research has two lines: methods (like ReAct) and environments (from games to internet, code, etc.), which reinforce each other.
Regarding OpenAI's five levels (chatbot, reasoner, agent, innovator, organizer), Yao explains the logic: first, linguistic priors enable chatbots; reasoning enables agents; next steps are self-reward and self-exploration, and multi-agent organization. He believes the most critical capability for improving agents is context handling (memory and online learning). For environments, he highlights code as particularly important, because it provides closed-loop feedback and verification, like human hands, making it a key path to AGI.
The interview also touches on counterpoints and limitations. Yao acknowledges current agents still face challenges like environment dependence, limited generalization, and technological immaturity. He recalls that when he started agent research, technology wasn't ready, and most people worked on QA or translation—agent research wasn't consensus. But he stuck with his non-consensus judgment, believing that immaturity was exactly the right time to start. He warns that the future of agents requires not only method innovation but also more realistic and challenging environments.
Overall, Yao's narrative offers a glimpse into a frontier researcher's thinking: from personal experience, he maps the evolution of AI agents from symbolic AI to deep RL to LLMs, emphasizing the centrality of language and reasoning, and the value of simple, general methods. His vision for the future—self-reward, multi-agent organization, code environments—provides listeners with key perspectives for understanding the second half of AI.



![[CVPR'23 WAD] Keynote - Jiyang Gao, Momenta](https://i.ytimg.com/vi/hFQLJIvdQNU/maxresdefault.jpg)