广告
加载中

重磅发布|极佳视界首次提出通用世界模型L1-L5分级体系 GWM正在成为AGI的最新前沿

龚作仁 2026-08-27 11:05
龚作仁 2026/08/27 11:05

邦小白快读

EN
全文速览

本文核心干货是极佳视界在2026世界机器人大会上首次提出通用世界模型(GWM)的L1-L5分级体系,为AGI领域的世界模型发展划定了清晰的能力边界,指明了技术方向。

1. GWM从L1到L5依次为视频级、动作级、几何级、物理级、运营级,对世界的约束逐层增强,从仅满足外观真实逐步升级到可长程稳定运营,高级别不会替代低级别,是叠加约束的关系。

2. 分级框架对应表征、智能、时间三条互相咬合的技术进阶轴,目前该体系已经在具身智能、自动驾驶、内容创作三个场景落地验证,同时这套分级也重新定义了不同层级模型的评测标准,避免将演示效果误判为技术实力。

本次极佳视界发布通用世界模型分级体系,为AGI领域的品牌建设和技术布局提供了清晰方向,也展现了AGI领域最新的产业发展趋势。

1. 品牌营销层面,极佳视界通过行业顶级会议发布原创分级体系,建立了行业标准制定者的品牌定位,依托通用世界模型北京市重点实验室的合作背书,强化了技术品牌影响力,这类路径值得同类科技品牌参考。

2. 产业趋势层面,当前AGI已经从单纯的语言智能转向需要理解物理世界的世界智能,通用世界模型成为AGI新的发展前沿,具身智能、自动驾驶、内容创作三个领域都有明确的市场需求,为品牌布局新赛道提供了清晰方向。

3. 产品研发层面,分级框架给出了清晰的技术路径,品牌可以按照自身定位选择对应层级的技术研发方向,避免方向错配浪费资源。

通用世界模型的兴起给AI领域相关卖家带来了明确的增长机会,同时也给出了清晰的风险提示,核心干货整理如下:

1. 增长机会层面,当前AGI发展进入新的阶段,从语言智能转向世界智能,具身智能、自动驾驶、内容创作三个核心场景都对更高能力的世界模型有明确需求,不同能力层级的模型都有对应落地场景,卖家可根据自身技术优势匹配对应赛道。

2. 风险提示层面,行业此前存在世界模型概念混淆的问题,很多从业者把低层级演示效果等同于高层级技术能力,既容易误导市场也会带来自身品牌信任危机,需要按照分级框架明确自身技术定位。

3. 可学习点,极佳视界通过建立行业标准坐标系的方式建立技术壁垒,这种路径值得中小AI卖家参考,先明确能力边界再逐步进阶升级。

通用世界模型的发展给工厂带来了新的商业机会,也为工厂推进智能化数字化转型提供了新的启示,干货整理如下:

1. 产品生产和设计需求层面,随着通用世界模型在具身智能、自动驾驶领域的落地,市场对工业机器人、自动驾驶相关零部件的精度要求大幅提升,几何一致性、物理规律准确性、长程运行稳定性成为核心要求,工厂可针对这些需求调整产品设计和生产标准。

2. 商业机会层面,三个核心落地场景带动了相关硬件产品的需求增长,从服务机器人到自动驾驶核心部件,都有持续的增量市场空间,工厂可针对性布局相关产能,对接市场新需求。

3. 转型启示层面,工厂推进智能化数字化可以对接通用世界模型的落地需求,围绕不同层级模型的能力要求开发适配的硬件产品,逐步升级自身产品体系,抓住AGI发展的新行业红利。

通用世界模型行业当前已经出现了明确的痛点和新的发展趋势,为AI领域服务商提供了新的业务方向,干货整理如下:

1. 行业发展趋势:AGI已经从语言智能进入世界智能的发展新阶段,通用世界模型成为AGI的最新前沿,技术发展沿着表征、智能、时间三条进阶轴推进,在具身智能、自动驾驶、内容创作三个领域都有大规模落地需求,行业整体增长空间十分广阔。

2. 客户核心痛点:此前行业对世界模型的概念边界模糊,不同能力层级的模型没有统一区分标准,评测体系错位,导致客户无法准确判断技术能力,研发方向也容易出现错配,造成资源浪费。

3. 可落地的解决方案:极佳视界提出的L1-L5分级体系,为不同能力模型提供了统一的坐标系,也给出了对应不同层级的评测标准,服务商可以依托这套体系为客户提供技术咨询、评测、研发对接等相关服务,匹配不同场景客户的需求。

通用世界模型的发展给AI平台带来了新的需求和发展方向,也明确了需要规避的风险,干货整理如下:

1. 市场对平台的核心需求:当前通用世界模型行业处于发展初期,需要平台搭建连接研发、产业、资本的协同生态,统一行业话语体系,推动不同层级技术的研发和落地,为行业参与者对接应用场景。

2. 平台运营和招商方向:平台可以围绕L1-L5五个层级的模型,对接不同技术能力的企业,覆盖从内容创作(L1为主)到长程运营智能体(L5为主)的全赛道,同时可以针对三个核心落地场景打造专属招商和运营板块,吸引不同定位的企业入驻。

3. 风险规避:平台需要引导企业按照分级框架明确自身技术定位,区分演示效果和实际技术能力,避免行业出现概念炒作、虚假宣传等问题,维护行业健康发展,同时建立对应分级的评测和准入标准,筛选真正符合能力要求的项目。

本文提出了通用世界模型领域全新的分级体系,是AGI领域重要的产业新动向,为相关研究提供了清晰的研究框架,干货整理如下:

1. 产业新动向:当前AGI的发展已经从大语言模型的语言智能,拓展到需要理解、预测物理世界的世界智能阶段,通用世界模型(GWM)正式成为AGI发展的最新前沿,行业已经出现大量技术落地尝试,概念扩张后急需统一的标准梳理混乱。

2. 新的理论贡献:极佳视界基于自身技术积累和落地实践,提出了L1-L5五级能力分级框架,明确了每个层级的能力定义、核心问题和检验标准,同时指出分级不是五代替代,而是对应表征、智能、时间三条互相咬合的技术进阶轴,高级是在低级基础上叠加约束。

3. 研究启示:这套分级体系不仅明确了技术发展路径,还改变了原有的研发和评测方式,为产业研究提供了统一的坐标系,后续研究可以围绕不同层级的技术突破、跨层级技术整合、不同场景的落地优化等方向展开,推动通用世界模型的技术发展。

返回默认

声明:快读内容全程由AI生成,请注意甄别信息。如您发现问题,请发送邮件至 run@ebrun.com 。

我是 品牌商 卖家 工厂 服务商 平台商 研究者 帮我再读一遍。

Quick Summary

This article highlights the core insight that ExViewer first introduced the L1-L5 classification framework for General World Models (GWM) at the 2026 World Robot Conference, which draws clear capability boundaries and charts a clear technical direction for world model development in the AGI space.

1. GWM levels L1 to L5 are defined sequentially as Video-level, Motion-level, Geometry-level, Physics-level, and Operation-level, with increasingly strict constraints on world representation. The framework progresses from only meeting visual realism requirements to enabling long-term stable operation, and higher levels do not replace lower levels — they simply add additional layers of constraints.

2. The classification framework corresponds to three interlocking technical advancement axes: representation, intelligence, and time. To date, the framework has been validated in three real-world scenarios: embodied intelligence, autonomous driving, and content creation. The classification also redefines evaluation standards for models at different levels, preventing misjudgment of demo performance as full technical capability.

ExViewer’s release of the General World Model classification framework provides clear guidance for brand building and technical positioning in the AGI industry, while highlighting the latest industrial development trends in the space.

1. From a brand marketing perspective, releasing an original classification framework at a top-tier industry conference has established ExViewer’s positioning as a standard-setter in the space. Backed by its partnership with the Beijing Key Laboratory of General World Models, the company has strengthened its technology-focused brand influence — a path that other comparable technology brands can reference.

2. In terms of industrial trends, AGI has shifted from pure language intelligence to world intelligence that requires understanding of the physical world, making General World Models the new frontier of AGI development. Clear market demand already exists in three fields: embodied intelligence, autonomous driving, and content creation, giving brands a clear direction for new track positioning.

3. For product R&D, the classification framework outlines a clear technical path, allowing brands to select a R&D direction aligned with their own positioning to avoid resource waste from misaligned strategy.

The rise of General World Models brings clear growth opportunities for AI industry sellers, while also highlighting key risks. Key takeaways are as follows:

1. On the opportunity side, AGI has entered a new phase of development, shifting from language intelligence to world intelligence. Three core scenarios — embodied intelligence, autonomous driving, and content creation — all have clear demand for higher-capability world models, and models of every capability level have matching application scenarios. Sellers can align their own technical advantages with the most suitable track.

2. For risk mitigation, the industry has long suffered from conceptual ambiguity around world models, with many practitioners equating low-level demo performance with high-level technical capability. This misrepresentation misleads the market and risks eroding brand trust, so sellers should clarify their technical positioning based on this classification framework.

3. A key takeaway is that ExViewer built technical barriers by establishing an industry standard coordinate system, a strategy that small and medium-sized AI sellers can reference: clarify capability boundaries first, then advance step-by-step.

The development of General World Models brings new business opportunities for factories, and new insights for advancing intelligent and digital transformation. Key takeaways are as follows:

1. For product design and manufacturing requirements, as General World Models are deployed in embodied intelligence and autonomous driving, market requirements for the precision of industrial robots and autonomous driving-related components have increased significantly. Geometric consistency, physical law accuracy, and long-term operational stability are now core requirements, and factories can adjust their product design and manufacturing standards to meet these needs.

2. In terms of business opportunities, the three core deployment scenarios have driven growing demand for related hardware products, from service robots to core autonomous driving components, creating sustained incremental market space. Factories can adjust capacity布局 to align with these new market demands.

3. For transformation, factories advancing intelligent and digital upgrades can align with the deployment requirements of General World Models, develop compatible hardware products that match the capability requirements of different model levels, gradually upgrade their product portfolios, and capture new industry dividends from AGI growth.

The General World Model industry now has clear pain points and new development trends, opening up new business directions for AI-focused service providers. Key takeaways are as follows:

1. Industry trends: AGI has moved beyond the language intelligence phase to the new era of world intelligence, with General World Models becoming the latest frontier of AGI development. Technology advances along three interlocking axes: representation, intelligence, and time, with large-scale deployment demand in three fields: embodied intelligence, autonomous driving, and content creation, creating very broad overall growth space for the industry.

2. Core customer pain points: The industry has long had vague conceptual boundaries for world models, with no unified classification standard for models of different capability levels and misaligned evaluation systems. This leaves customers unable to accurately assess technical capability, and often leads to misaligned R&D directions and wasted resources.

3. Deployable solutions: ExViewer’s L1-L5 classification framework provides a unified coordinate system for models of different capabilities, alongside matching evaluation standards for each level. Service providers can leverage this framework to offer customers technical consulting, evaluation, R&D matchmaking and other related services, to match the needs of customers across different scenarios.

The development of General World Models brings new demands and development directions for AI platforms, while also clarifying risks to avoid. Key takeaways are as follows:

1. Core market demand for platforms: The General World Model industry is still in an early development stage, and requires platforms to build a collaborative ecosystem connecting R&D, industry, and capital, unify industry terminology, promote R&D and deployment of technologies at all levels, and connect industry participants with application scenarios.

2. Platform operation and investment promotion direction: Platforms can organize businesses around the five L1-L5 model levels, to onboard enterprises with different technical capabilities, covering the full track from content creation (dominated by L1) to long-term operating agents (dominated by L5). Platforms can also build dedicated investment promotion and operation segments for the three core deployment scenarios to attract enterprises with different positioning.

3. Risk mitigation: Platforms should guide enterprises to clarify their technical positioning based on the classification framework, distinguish between demo performance and actual technical capability, prevent concept hype and false advertising in the industry, maintain healthy industry development, and establish classification-aligned evaluation and access standards to screen projects that actually meet capability requirements.

This article introduces a brand new classification framework for the General World Model field, representing an important new industrial development in AGI that provides a clear research framework for related studies. Key takeaways are as follows:

1. New industrial development: AGI development has expanded beyond the language intelligence of large language models, and entered the world intelligence phase that requires understanding and predicting the physical world. General World Models (GWM) have officially become the latest frontier of AGI development, with a large number of industry deployment attempts already underway. Following rapid conceptual expansion, the industry urgently needs unified standards to resolve existing confusion.

2. New theoretical contribution: Drawing on its own technical accumulation and deployment practice, ExViewer proposes the L1-L5 five-level capability classification framework, defining clear capability definitions, core problems and testing standards for each level. The framework also clarifies that the classification does not represent sequential replacement of generations, but corresponds to three interlocking technical advancement axes of representation, intelligence, and time, with higher levels adding constraints on top of lower levels.

3. Research implications: This classification framework not only clarifies the technical development path, but also reshapes existing R&D and evaluation methods, providing a unified coordinate system for industrial research. Future research can focus on technical breakthroughs at different levels, cross-level technical integration, and deployment optimization across different scenarios to drive technological progress in General World Models.

Disclaimer: The "Quick Summary" content is entirely generated by AI. Please exercise discretion when interpreting the information. For issues or corrections, please email run@ebrun.com .

I am a Brand Seller Factory Service Provider Marketplace Seller Researcher Read it again.

当行业里越来越多的视频生成模型、交互环境模型、机器人策略模型都被称为“世界模型”,一个更根本的问题随之浮现:它们建模的,究竟是哪一层世界?

在北京极佳视界科技有限公司与通用世界模型北京市重点实验室联合主办的2026世界机器人大会“通用世界模型前沿论坛”上,极佳视界创始人、CEO黄冠以《通用世界模型引领AGI的最新前沿》为题发表主旨演讲。他提出,AGI不能只有语言智能,更需要能够在物理空间中理解、生成、预测和行动的“世界智能”。而要判断世界智能走到了哪里,行业首先需要一套能够区分能力边界的坐标系。

基于极佳视界在通用世界模型领域的技术积累,以及在具身智能、自动驾驶、内容创作等多个领域的落地实践,黄冠表示:“通用世界模型(GWM)可以分为L1-L5分级框架——从生成视觉真实的视频世界,逐级走向动作、几何、物理,最终抵达可长程运营的通用无限世界。”

图片

图片

01 当所有模型都叫“世界模型”,行业更需要一把尺子

世界模型正在快速走出单一技术范畴:它可以是生成未来画面的视频模型,可以是接受动作输入、预测环境变化的交互式模拟器,也可以成为机器人规划与决策的内部模型。概念边界的扩张说明技术正在成熟,但也带来新的混淆——画面更逼真,并不等于空间更准确;能够预测下一帧,并不等于能够预测一次接触、碰撞或长程操作的真实后果。

L1-L5描述的是模型对世界施加的约束越来越强:从外观真实,到动作可用;从三维结构稳定,到物理规律可信;再到状态、记忆与反馈可以在更长时间尺度上持续运转。越往上,模型越不能只依赖一段看起来合理的视频来证明自己。

通用世界模型的“通用”,至少要经得住三重检验:能否跨场景迁移,能否被行动与结果验证,能否在时间推移中保持一致。

02 L1-L5,模型究竟学会了哪一层“世界”?

图片

▲ 极佳视界提出的通用世界模型L1-L5分级框架

▍L1视频:先让世界“看起来真实”

能力定义:根据各种输入生成视觉真实的通用视频世界。这一层解决的是世界的视觉表达:画面质量、时间连续性、语义一致性与可控性。它是世界模型的重要入口,但视觉可信仍不等于结构和规律可信。一个模型能生成逼真的机械臂,并不代表它知道夹爪与物体发生接触后会怎样。

回答的问题:模型能否稳定生成一个可信、连贯、可控的世界?

▍L2动作:从“看见未来”到“做出选择”

能力定义:生成高效、合理的通用执行动作。在这一层,模型不再只是呈现世界,而要把环境理解转化为可执行的动作序列。对于机器人而言,动作必须在约束下完成任务;对于交互环境而言,用户输入必须引起合理反馈。世界模型开始从内容生成器转向决策与行动的基础设施。

回答的问题:动作是否可执行、有效率,并能真正推进任务?

▍L3几何:让世界从像素表面获得几何信息

能力定义:生成几何准确的通用空间世界。空间关系、尺度、遮挡、多视角一致性与三维结构在这里成为硬约束。只要视角变化,物体的位置、形状和相互关系就必须保持一致。世界不再是一串逐帧“猜出来”的画面,而是一个可以被定位、测量和导航的空间。

回答的问题:跨视角、跨传感器、跨时间观察时,空间结构是否仍然成立?

▍L4物理:从空间正确走向因果正确

能力定义:生成物理精确的通用物理世界。世界除了有几何结构,还要服从接触、摩擦、碰撞、力与运动等规律。模型不仅要预测“接下来像什么”,还要回答“采取这个动作会发生什么”,并在更长的滚动预测中控制误差累积。物理精确性决定世界模型能否服务真实机器,而不只是服务视觉体验。

回答的问题:同样的状态与动作,能否稳定地产生符合物理和因果的结果?

▍L5运营:让一个世界长期存在、持续反馈

能力定义:生成可长程运营的通用无限世界。这里的“运营”不是商业运营,而是让世界在长时间尺度上维持状态、规则、记忆和交互,并在持续反馈中演化。模型需要处理长程任务、多主体协同、异常恢复和开放式变化,使世界从一次性生成结果变成可持续运行的智能环境。

回答的问题:时间不断向前时,世界能否记得、响应、恢复并保持内部一致?

“视频决定世界看起来是否真实,几何与物理决定它是否真正可以被进入、被预测、被行动;长程运营则决定这个世界能否持续存在。”

03 五个层级,实际对应三条同时推进的进阶轴

把L1-L5简单理解成五代模型,会低估这套框架的意义。它真正刻画的是三条彼此咬合的技术进阶轴。只有三条轴同时向前,世界模型才可能从演示走向基础能力。

▍第一条:表征轴——从像素到几何,再到物理

像素让模型学会世界的外观,几何让模型获得稳定的空间结构,物理让模型把结构与因果规律连接起来。越往后,模型越需要显式或隐式地保存对象、状态、关系与动力学,而不是仅靠局部纹理延续下一帧。

▍第二条:智能轴——从生成到行动,再到闭环反馈

生成是提出一种可能的未来,行动是在诸多未来中选择一条可执行路径,闭环则要求模型根据结果不断修正策略。世界模型只有进入“预测—行动—反馈—再预测”的循环,才能真正成为智能体的内部环境。

▍第三条:时间轴——从短片段到长程任务,再到持续世界

短时生成主要考验局部连贯,长程任务会暴露记忆丢失、状态漂移与误差累积,而持续世界还要面对开放变化和多主体交互。L5之难,并不只是“生成得更久”,而是让因果、状态和目标在更长时间尺度上仍然成立。

需要强调的是:更高层级不是替代更低层级,而是在其上叠加更强约束。没有稳定的视频表征,动作无从观察;没有准确的几何与物理,长程运行也只会放大错误。

04 三个场景,校准同一个世界

L1-L5并非从单一场景倒推出的抽象概念。极佳视界自2023年成立以来,持续布局「通用世界模型(GWM)」的全栈技术体系,并应用在具身智能、自动驾驶、内容创作等不同的场景。不同场景对视觉、空间、物理、动作和长程稳定性的要求不同,却在共同校准同一件事:模型是否真的理解世界,而不仅是拟合画面。

▍具身智能:用行动结果检验预测

在具身智能场景中,黄冠重点介绍了GigaBrain-0.7。其System 1以VLA生成连续未来动作,System 2以VLM理解物理常识、因果关系并完成高层规划,System 3则引入通用世界价值模型,对未来世界及动作价值进行预测,再结合在线强化学习形成持续自进化闭环。现场展示的长程任务包含20余个原子动作,并实现超过20分钟的一镜到底执行。

这里对应的不只是L2的动作生成。高层规划依赖对环境状态的理解,世界价值模型要判断动作的未来后果,长程执行又会检验记忆、状态保持与异常恢复。机器人因此成为高等级世界模型最严格的“真机考场”之一。

图片

图片

▍自动驾驶:把视觉世界压进几何与物理约束

在自动驾驶场景中,DriveDreamer Engine由World Model、Agentic Model与Harness构成,支持结构化多输入视频生成、HD Map与3D Box控制、多视角生成,以及RGB与LiDAR之间的跨传感器一致性。极佳视界还将世界重建扩展到10公里以上路网,并在部分场景中把重建时间从约3小时压缩到约3分钟。

驾驶场景对“看起来合理”的容忍度极低:车道结构、车辆位置、遮挡关系与传感器观测必须一致,智能体的动作还要引起正确的交通演化。这正是从L1视频走向L3几何、L4物理的典型压力测试。

图片

图片

▍内容生成:在开放创作中检验可控与一致

在内容生成场景中,极佳视界以YiSu 2.0、YiSu Agentic与YiSu Engine构成产品矩阵,并通过WonderTurbo、HumanDreamer、Motion-R1、WorldDreamer等模型持续推进长镜头、三维一致性、人体运动和空间交互。

内容创作看似更接近L1,却也会迅速触碰更高层级:镜头一旦拉长、视角一旦切换、角色一旦与环境互动,几何、遮挡与物理稳定性就会成为决定可用性的瓶颈。

图片

05 L1-L5真正改变的是研发与评测方式

分级框架的价值,不只在于告诉行业“下一步做什么”,还在于提醒行业“不同层级不能用同一套指标”。如果用视频美学指标衡量物理世界模型,或只用短时成功率衡量长程智能体,就会把演示质量误当成技术等级。

图片

▲ 从L1到L5,评测对象也应随能力边界升级

当指标随层级升级,行业才能把“演示得好”与“能力走得远”区分开:L1可以回答生成质量,L2要看动作是否完成任务,L3与L4要验证空间和因果是否正确,L5则必须接受长时间、开放环境与持续交互的考验。

L1-L5由此不仅是一张技术地图,也可能成为连接研发目标、数据建设与评测标准的共同语言。

结语 从语言到世界,GWM正在成为AGI的最新前沿

大语言模型证明了规模化学习可以压缩人类语言与知识,但物理世界不会仅以文本形式存在。它有空间,有动作,有摩擦与碰撞,也有不断延展的时间。世界模型要承接AGI的下一段路,就必须把这些约束逐层纳入同一个可学习、可预测、可行动的系统。

黄冠在演讲中提出的L1-L5没有把未来简化成一张年份表,而是给出了一条能力累积路径:先让世界可见,再让行动发生;再让空间成立、规律可信,最终让智能在一个持续运行的世界中形成闭环。它的真正价值,不是为今天的模型贴标签,而是让行业更清楚地看到,通用世界模型的巨大价值和坚实路径。

图片

从语言到世界,极佳视界正在以「通用世界模型」的全栈技术体系,以及具身智能、自动驾驶、内容生成三类场景的落地实践,持续验证这条进阶路径。真正的通用世界模型,也将带来一个完全不一样的世界。

注:文/龚作仁,文章来源:Laborer,本文为作者独立观点,不代表亿邦动力立场。

文章来源:Laborer

广告
微信
朋友圈

FAQ回顾

通用世界模型L1-L5分级分别对应什么能力?

通用世界模型L1为生成视觉真实的通用视频世界,L2为生成高效合理的通用执行动作,L3为生成几何准确的通用空间世界,L4为生成物理精确的通用物理世界,L5为生成可长程运营的通用无限世界,层级越高对世界的约束越强,需叠加前序层级能力。

通用世界模型L1-L5分级体系有什么行业价值?

该分级体系可作为通用世界模型能力边界的判断坐标系,区分不同层级模型的评测标准,避免将演示质量误判为技术等级,还可成为AGI领域研发目标、数据建设与评测标准的通用行业语言。

通用世界模型目前可应用在哪些场景?

通用世界模型已可落地三类场景,分别是用行动结果检验预测的具身智能场景、对几何与物理约束要求极高的自动驾驶场景,以及考验可控性与一致性的内容创作场景,极佳视界已在三类场景布局相关技术验证路径。

通用世界模型和AGI发展有什么关联?

AGI不能仅具备语言智能,还需拥有可在物理空间理解、生成、预测和行动的世界智能,通用世界模型是AGI发展的最新前沿,L1-L5分级体系明确了世界智能的能力累积进阶路径。

这么好看,分享一下?

朋友圈 分享

APP内打开

+1
+1
微信好友 朋友圈 新浪微博 QQ空间
关闭
收藏成功
发送
/140 0