HOME>WHAT'S NEW_LIST

Beyond ‘text world’: second phase of intelligence revolution from 2026 WAIC

Source:Chinese Social Sciences Today 2026-08-01

On July 17, a large screen in the Golden Hall of the Shanghai World Expo Center showed an infant repeatedly handling a toy: reaching out, making contact, missing, and trying again. No one had supplied a pre-labeled “correct answer,” yet through repeated action and feedback, the child was quietly learning the rules of the world. Richard Sutton, 2024 Turing Award laureate and a professor of computer science at the University of Alberta in Canada, used the scene to frame his remarks on AI. Rather than dwelling on the dazzling capabilities of large language models, he candidly described today’s AI as weak and unreliable. Existing models, he explained, can organize knowledge that humanity has already accumulated, but cannot yet discover genuinely new knowledge on their own. They produce fluent answers but cannot yet distinguish truth from falsehood because they have never tested their judgments against the physical world.

At parallel forums held during the 2026 World Artificial Intelligence Conference (WAIC), scientists returned repeatedly to a simple but fundamental question: How do machines truly comprehend the world? If AI can “talk” but cannot “act,” how real is its intelligence? As it moves beyond generating text and begins to interact with physical reality, on what grounds should people trust it? Once AI starts acting on humanity’s behalf, how can human beings preserve direction, accountability, and clear boundaries?

Wang Jian, an academician of the Chinese Academy of Engineering, director of Zhijiang Lab, and founder of Alibaba Cloud, reminded the audience that when people speak of “large models” today, they generally mean large language models. Yet text records only the knowledge that humans have written down and published. Vast quantities of scientific information—including spectra, seismic waves, sound waves, and observational signals—do not exist in linguistic form. Citing the field of geoscience, Wang noted that more than 70% of geoscientific information is not contained in text, adding that “a single spectrum is worth more than thousands of images.”

Su Hao, a distinguished professor from Fudan University (FDU) and founding director of the Institute of General Physical Intelligence at FDU, compared language to the “shadow” of the world. Humans first navigate the physical world, then condense their experiences into language; AI models, by contrast, have studied only the shadow, never the tangible reality that casts it. To understand the world in any meaningful sense, AI must leave the “text world” and engage directly with physical reality.

Once AI transcends text and enters the physical world, it faces the dual imperative of advancing development and strengthening governance. As industrial applications rapidly expand and AI begins acting on humans’ behalf, critical questions arise: From where does its authority originate? Who is responsible for the consequences of its actions? How can its safety be verified?

Yin Qi, chairman of the Shanghai-based AI company StepFun, argued that the third wave of AI would likely see intelligent agents operating in the physical world. Yet he cautioned that this transition would represent not merely a leap in capability, but a restructuring of social order. Society must determine whom intelligent agents act for, who bears responsibility for the consequences of their actions, and how their identities and permissions can remain verifiable and controllable.

In a video address, Yoshua Bengio, 2018 Turing Award laureate and a professor of computer science at the University of Montreal in Canada, warned that frontier AI capabilities are advancing at breakneck speed, with the complexity of tasks that AI systems can plan nearly doubling every few months. He recommended that all high-risk systems must be required to demonstrate their safety before release, that individual companies not be permitted to set their own risk thresholds, and that international verification mechanisms be established.

Xue Lan, director of the Institute for AI International Governance at Tsinghua University, proposed that existing principal-agent arrangements must be adjusted to accommodate machine agents. Because machines cannot bear social responsibility themselves, accountability must remain traceable across every stage of development, deployment, and use.

Mark Nitzberg, executive director of the Center for Human-Compatible AI at UC Berkeley in the United States, maintained that AI safety cannot be treated as an afterthought, with testing and remediation beginning only after a system has been built. Instead, developers must assume the responsibility of proving safety from the design stage onward.

Nicholas Dirks, president and chief executive officer of the New York Academy of Sciences, similarly cautioned that regulation should not begin only after a crisis has occurred. Because AI’s impact extends across fields including finance, medicine, and transportation, forward-looking forms of accountability must be established before risks materialize.

Concluding a roundtable discussion, Zhou Bowen, director of the Shanghai Artificial Intelligence Laboratory and a professor at Tsinghua University, distilled the issue into a clear boundary for the intelligent age: “AI may act on humans’ behalf, but it cannot take responsibility for them.”

Editor:Yu Hui

Copyright©2023 CSSN All Rights Reserved

Copyright©2023 CSSN All Rights Reserved