Ilya Sutskever, co-founder and former chief scientist of OpenAI, knows better than anyone how large language models are built. Yet in a recent conversation, he struck a more sober tone, pointing to a curious disconnect: today’s models ace benchmarks and evaluations, but stumble on real tasks—for instance, fixing a bug by politely acknowledging the error, introducing a new one, then reverting to the original bug in a loop when prompted again. For Sutskever, this “high-score, low-work” mismatch signals that the era of pure scaling—piling on more data and compute—is giving way to an era driven by genuine research.
His central claim is simple: strong eval results don’t equal real competence, and AI’s economic impact is lagging far behind for deep reasons. Global AI investment has already reached about 1% of GDP, yet most people feel no change in daily life. Some call this a “slow takeoff,” speculating the apathy might even survive the singularity. Sutskever disagrees, insisting that economic forces will eventually push AI into every sector; we’re just in the very early stages of diffusion, and the disruption will be sharp when it hits.
The bug-fixing example perfectly captures the gap. An advanced model passes tough coding tests, but when asked to correct a piece of code, it politely agrees, offers a fix that introduces a new error, then, when that is pointed out, apologizes again and reverts to the original bug, cycling endlessly. Sutskever and his interlocutor offer two explanations: first, fine-tuning techniques like RLHF might overly narrow the model’s attention, making it sharp on certain dimensions but oblivious to basic consistency; second, pre-training on indiscriminate internet data may produce broad but shallow knowledge, causing breakdowns in novel contexts. In short, today’s models excel at mimicry, not reliable reasoning.
Counterarguments exist. Technology diffusion is historically slow—electricity and the internet took decades to rewire economies—so the lag may be normal. Benchmarks might simply be poor measures, and human programmers also create bug loops. But Sutskever focuses on the pattern of failure: if models are statistical matchers, they may never learn to robustly solve new problems, puncturing hopes of an imminent economic revolution.
Ultimately, Sutskever is urging a shift from scaling to tackling foundational problems that make models not just know the answer but understand what they’re doing. Only when AI can perform consistently in the open world will those shiny eval scores become trustworthy, and the real economic show will begin. It’s a reminder to look past the milestone hype and ask: is this intelligence, or just a high-resolution mirror?
