For the past two years, generative AI has largely competed on what it can create. At the 2026 World Artificial Intelligence Conference (WAIC), Kunlun Tech Chairman and CEO Fang Han argued that the next phase will be defined by what AI can understand and do in simulated and physical environments. Calling 2026 the “first year of world models,” Kunlun Tech is positioning its Skywork AI research around a shift from generating media to modeling space, memory, motion and interaction in real time.
The generative AI market has spent much of its recent history chasing better outputs: more realistic images, longer videos, more coherent text and increasingly convincing synthetic voices.
World models point toward a different objective.
Instead of asking AI to produce another piece of content, researchers are trying to build systems that understand how objects, environments and actions behave over time. That distinction could eventually make world models an important layer for gaming, robotics, autonomous systems, simulation and interactive entertainment.
At WAIC on July 19, Fang Han, chairman and CEO of Kunlun Tech, described 2026 as the first year of world models. The company is the parent of Skywork AI, whose research has increasingly focused on interactive video and simulated environments.
The announcement reflects a broader transition in AI infrastructure: from models that generate representations of the world to models that can maintain an internal representation of how that world works.
From video generation to interactive environments
The technical challenge is considerably harder than producing a convincing video clip.
A conventional video model can generate a sequence that looks plausible from frame to frame. A world model must maintain consistency when the user changes the environment, moves the camera or interacts with objects.
Skywork AI’s Matrix-Game 3.5 attempts to address that problem through a mechanism called Patch Memory.
Rather than treating previous frames as complete images, the system divides them into smaller patches and associates those patches with three-dimensional spatial information. The objective is to help the model remember where objects exist even when their visual appearance changes because the camera moves.
The distinction is important. A frame-based system essentially remembers what an image looked like. A spatial memory system attempts to remember what exists in a particular location.
Skywork AI says Matrix-Game 3.5 can generate 720p streaming video at roughly 20 frames per second using a 5-billion-parameter model, with memory extending to approximately one minute.
If those performance characteristics hold across broader testing, the implications extend beyond video generation. Real-time interaction is a prerequisite for applications in which AI must respond continuously rather than generate a predetermined clip.
That could include virtual environments, game development and eventually physical AI applications.
Geometry becomes part of the model
Matrix-Game 3.5 also incorporates PRoPE geometric position encoding and Warped RoPE, moving positional information toward a representation based on physical geometry.
This matters because conventional positional encoding is primarily concerned with relationships inside sequences or image grids. A world model needs to understand relationships in three-dimensional space.
The architecture therefore attempts to connect visual information with camera position, depth and movement.
That approach reflects an increasingly important direction in AI research: improving capability through architecture and data representation rather than simply increasing parameter counts.
Skywork AI says most of Matrix-Game 3.5’s interactive functionality does not require adding large numbers of parameters to the underlying model. Instead, components such as geometric encoding, memory, camera control and dynamic-static separation are designed as modular additions.
That could make the technology easier to transfer between video foundation models.
Data may be the bigger competitive advantage
The underlying data pipeline could ultimately matter as much as the model architecture.
World models need training data that contains information about physical relationships, not simply attractive images or videos. Skywork AI says it built an automated pipeline capable of reconstructing 3D information from video, including camera poses, metric depth and camera characteristics.
The resulting dataset reportedly contains more than 5 million video segments and 10,000 hours of training material across more than 1,200 game scenes.
That points to an emerging infrastructure layer for physical AI: systems capable of converting ordinary video into structured information that models can use to learn about environments.
Gaming is particularly useful for this research because games offer controllable environments, repeatable interactions and extensive visual diversity.
It is also why the gaming industry could become one of the earliest commercial beneficiaries of world models.
Open source could create an ecosystem effect
Skywork AI is also pursuing an open-source strategy that could prove strategically important.
The company says researchers have used earlier Matrix-Game models as foundations or benchmarks for projects involving NVIDIA, Zhejiang University, NYU and Adobe.
That matters because the influence of a model architecture is not necessarily measured by direct commercial revenue. A research foundation that becomes widely adopted can become an interoperability layer for an entire ecosystem.
The strategy resembles the broader open-source dynamic seen elsewhere in AI. Meta, Google, NVIDIA and numerous academic laboratories have increasingly used open models, frameworks and benchmarks to accelerate research while competing on infrastructure and applications.
For Kunlun Tech, establishing Matrix-Game as a common foundation for world-model research could therefore be more valuable than simply releasing another standalone video generator.
World models need perception, reasoning and decisions
The technology discussion at WAIC also highlighted that a complete world model requires more than visual perception.
Academician Zhou Zhihua of the Chinese Academy of Sciences described a three-layer structure involving perception, understanding and decision-making.
The decision layer is especially difficult because errors can compound over multiple steps. A small mistake in one action can change the state of the environment and make subsequent predictions increasingly unreliable.
Zhou also introduced the concept of “learnware,” in which models are paired with automatically generated specifications that allow heterogeneous AI systems to be reused and combined without exposing their underlying data.
The idea is notable because it points toward a more modular AI ecosystem. Instead of building increasingly large monolithic models, developers could potentially assemble specialized models and capabilities into reusable systems.
That architecture could become important as world models move from research demonstrations toward production environments.
The next battleground may be physical AI
The broader industry trajectory is increasingly visible.
Generative AI reduced the cost of producing digital content. World models aim to reduce the cost of understanding and simulating environments.
That could eventually affect robotics, autonomous vehicles, gaming, industrial simulation and embodied AI.
The distinction between digital and physical AI is also becoming less clear. A model trained inside a game environment can learn useful representations of movement, objects and spatial relationships that may eventually inform robotics or other physical systems—although transferring knowledge from simulated environments to the real world remains a major technical challenge.
The strongest near-term opportunity is likely gaming and interactive media, where the environment is already digital and controllable.
For enterprise AI teams, the development is worth watching because world models could introduce a new category of infrastructure alongside large language models, vision models and generative video systems.
The question is no longer simply whether AI can generate a convincing representation of the world.
It is whether AI can maintain a persistent model of that world, predict what happens next and respond to changes in real time.
If that becomes reliable, generative AI’s next chapter could look less like content production and more like simulation.
Market Landscape
World models sit at the intersection of several rapidly developing markets.
Generative video companies are pursuing longer clips, consistent characters and controllable camera movement, while researchers in robotics and autonomous driving are working toward models that understand physical environments.
Companies such as NVIDIA, Google DeepMind, Meta and Adobe are investing across adjacent areas including generative media, simulation, computer vision and physical AI.
The competitive distinction is increasingly shifting from parameter count to data quality, spatial representation, memory, inference efficiency and real-time interaction.
For enterprises, that could create a new infrastructure stack: foundation models at the bottom, world-model or simulation layers above them, and applications for gaming, robotics, industrial digital twins and autonomous systems at the top.
The economics remain uncertain. Training spatially grounded models requires large-scale video and structured data pipelines, while real-time inference requires significantly greater efficiency than offline content generation.
Still, if models can achieve persistent memory and reliable physical reasoning at manageable compute costs, world models could become an important complement to today’s language-model infrastructure.
Top Insights
- Kunlun Tech is positioning 2026 as a turning point for world models, shifting AI competition from content generation toward spatial understanding, memory and real-time interaction.
- Skywork AI’s Matrix-Game 3.5 uses Patch Memory and geometric encoding to improve object persistence, potentially enabling more consistent interactive video environments for gaming and simulation.
- A large-scale annotated video pipeline gives Skywork AI structured spatial training data, highlighting data infrastructure as an increasingly important competitive advantage in world-model development.
- The open-source Matrix-Game ecosystem could extend Skywork AI’s influence beyond its own products, giving researchers and technology companies a foundation for experimenting with interactive world models.
- Gaming may become an early commercial proving ground for world models before the technology expands into robotics, autonomous systems, simulation and other physical-AI applications.
Power Tomorrow’s Intelligence — Build It with TechEdgeAI




