广告
加载中

GigaWorld-Policy-0.5 发布:AutoResearch加持的世界动作模型 速度更快 效果更强

龚作仁 2026-07-17 10:38
龚作仁 2026/07/17 10:38

邦小白快读

EN
全文速览

本文核心内容是极佳视界联合清华大学发布了新一代通用机器人领域的世界动作模型GigaWorld-Policy-0.5,该模型针对性解决了传统模型存在的三大行业痛点,核心干货信息如下

1. 传统世界动作模型存在三大问题:人工调参耗时耗力、推理延迟高算力开销大、真实场景泛化性不足,严重限制了行业落地发展

2. 新模型有三大核心升级:搭载AutoResearch智能训练管线实现全自动超参搜索,可减少超50%人工实验成本;采用MoT混合专家架构重构推理链路,RTX4090环境下推理延迟低至85ms,速度提升超60%;用超10000小时全域视频预训练沉淀通用物理先验,真机任务性能提升6%

3. 目前该模型已经完全开源,项目主页、技术报告、代码仓库均已公开,感兴趣的用户可自行访问获取资源测试体验

本次发布的新模型为通用机器人产业落地扫清了核心技术障碍,对机器人相关品牌的产品研发、市场布局有较高参考价值,核心干货如下

1. 消费趋势层面:当前通用机器人技术已经突破落地瓶颈,可低成本应用在家居服务、柔性仓储、餐饮加工等多个真实场景,相关品牌可提前针对这些大众需求场景布局产品,抢占市场先机

2. 产品研发层面:新模型给出了解决传统模型调参难、泛化弱、落地慢三大痛点的可行路径,品牌可参考AutoResearch自动调参、MoT架构解耦、大规模预训练这套技术方案,优化自身机器人产品性能,降低研发成本

3. 成本控制层面:新模型可在消费级RTX4090显卡上稳定运行,大幅降低了部署的硬件成本,品牌可以推出性价比更高的终端产品,提升自身产品的市场竞争力

本次通用机器人的技术升级给相关领域卖家带来了新的增长机会,核心干货整理如下

1. 市场机会层面:新模型突破技术瓶颈后,通用机器人可低成本落地多个To B和To C场景,To B卖家可针对性开发家居、仓储、餐饮领域的机器人解决方案,拓展自身业务边界;To C卖家可布局消费级家用服务机器人产品,抓住新一波消费需求风口

2. 可学习经验:新模型的AutoResearch自动调参管线可降低研发门槛,减少超50%的人力研发成本,卖家可参考这套自动化研发模式降低自身研发投入,提升研发效率

3. 风险提示:当前技术还处于持续迭代阶段,后续会往端侧推理、双臂机器人、移动操作机器人方向升级,卖家需要跟进技术迭代速度,避免产品技术落后被市场淘汰

本次新技术发布给生产工厂尤其是机器人相关制造工厂带来了新的商业机会和数字化研发启示,核心干货整理如下

1. 商业机会层面:新模型大幅降低了通用机器人的落地成本,家居服务、柔性仓储、餐饮加工等多个场景对机器人的需求会快速增长,相关工厂可提前拓展这些场景对应机器人产品的生产线,抢占市场订单增量

2. 产品生产设计层面:新模型满足机械臂实时闭环控制的硬性要求,且可在消费级显卡部署,工厂开发机械臂等机器人产品时,可参考这套技术方案降低对高端硬件的要求,压缩生产和研发成本,推出更有价格竞争力的产品

3. 数字化研发启示:自动化训练调参管线可实现无人化研发,工厂推进自身研发数字化转型时,可参考这类自动化研发流程,减少人力投入,缩短研发周期,提升整体研发效率

通用机器人领域正从实验室走向产业落地,行业痛点清晰,新技术的发展给服务商带来了新的业务机会,核心干货如下

1. 行业发展趋势:当前通用具身智能是机器人领域的核心发展方向,世界动作模型技术成熟后,产业落地速度会明显加快,市场对相关技术服务的需求会持续上涨,赛道增长空间广阔

2. 客户核心痛点:目前机器人研发企业普遍面临传统世界动作模型调参成本高、推理延迟无法满足落地要求、泛化性差场景适配难三大痛点,这是服务商可以切入的核心需求点

3. 可落地解决方案参考:本次发布的新模型给出了成熟的解决方案,AutoResearch自动调参、MoT混合专家架构解耦训练推理、万小时大规模预训练这套技术路径,服务商可以将其优化打包,输出给中小机器人研发企业,满足客户降本增效的需求,提前布局赛道抢占份额

本次通用机器人技术突破,推动产业落地加速,对机器人产业相关平台的运营和布局有较多启发,核心干货整理如下

1. 市场需求层面:当前大量中小机器人研发团队有降低研发成本、降低部署门槛的需求,平台可针对性推出相关配套服务,比如研发工具共享、算力补贴、技术对接等,吸引更多开发者入驻平台

2. 招商布局层面:技术突破后,通用机器人领域的创业项目会明显增多,且已经具备商业落地可能性,平台可将家居、仓储、餐饮场景的通用机器人项目作为重点招商方向,丰富平台生态,抢占新赛道的生态优势

3. 风险规避:当前通用机器人技术仍在迭代过程中,部分创业项目技术落地能力不足,平台在招商引入项目时,需要重点考察项目的技术落地能力,筛选真正掌握核心技术的项目,降低平台自身的运营风险

本文披露了世界动作模型领域的最新研究成果,对通用机器人领域的学术研究和产业研究都有较高的参考价值,核心干货如下

1. 产业新动向:当前世界动作模型已经补齐了调参难、泛化弱、落地慢三大短板,首次实现训练便捷性、通用操作性能、实时部署效率三者兼顾,通用机器人技术正式走出实验室,开始走向多元场景落地,通用具身智能的产业价值开始逐步释放

2. 技术创新方向:本文针对传统模型的痛点提出了三个创新性解决方案,分别是AutoResearch全自动化超参搜索训练管线、MoT混合专家架构解耦训练与推理链路、万小时全域视频预训练底座,为后续该领域的研究提供了新的清晰路径

3. 未来研究方向明确,文中指出后续行业会围绕端侧推理性能优化、双臂机器人适配、移动操作机器人适配方向迭代,研究者可围绕这些方向开展后续相关研究,探索产业更大的价值空间

返回默认

声明:快读内容全程由AI生成,请注意甄别信息。如您发现问题,请发送邮件至 run@ebrun.com 。

我是 品牌商 卖家 工厂 服务商 平台商 研究者 帮我再读一遍。

Quick Summary

This article covers the launch of GigaWorld-Policy-0.5, a new-generation generalist world action model for general-purpose robots, jointly released by ExViewer and Tsinghua University. The model addresses three long-standing pain points of conventional models in the industry, with key takeaways as follows:

1. Traditional world action models face three major issues that severely hinder industrial adoption: time- and resource-intensive manual parameter tuning, high inference latency that drives up computing costs, and poor generalization to real-world scenarios.

2. The new model delivers three core upgrades: it leverages the AutoResearch automatic training pipeline to enable fully automated hyperparameter search, cutting manual experiment costs by over 50%; it reconstructs the inference pipeline with the Mixture of Experts (MoT) architecture, delivering inference latency as low as 85ms on the RTX 4090, a speed improvement of over 60%; and it incorporates over 10,000 hours of full-domain video pre-training to build general physical priors, boosting real robot task performance by 6%.

3. The model is now fully open-source, with its project homepage, technical report, and code repository all publicly available. Interested users can access the resources and test the model on their own.

The newly released model removes core technical barriers to the industrial adoption of general-purpose robots, offering high reference value for product R&D and market positioning for robot-related brands. Key takeaways are as follows:

1. Consumer trend perspective: General-purpose robot technology has now broken through long-standing adoption bottlenecks, enabling low-cost deployment across multiple real-world scenarios including home services, flexible warehousing, and food processing. Relevant brands can get ahead of the curve by developing products tailored to these high-demand consumer scenarios to capture early market share.

2. Product R&D perspective: The new model provides a viable solution to the three core pain points of traditional models: difficult parameter tuning, weak generalization, and slow deployment. Brands can adopt its technical framework—AutoResearch automatic parameter tuning, MoT architecture decoupling, and large-scale pre-training—to optimize their robot product performance and cut R&D costs.

3. Cost control perspective: The new model runs stably on consumer-grade RTX 4090 graphics cards, significantly reducing hardware deployment costs. This allows brands to launch more cost-competitive end products and improve their market competitiveness.

This technical upgrade for general-purpose robots opens new growth opportunities for sellers in related fields, with key takeaways as follows:

1. Market opportunity perspective: After breaking through core technical bottlenecks, general-purpose robots can now be deployed at low cost across multiple B2B and B2C scenarios. B2B sellers can develop customized robot solutions for the home, warehousing, and catering sectors to expand their business scope, while B2C sellers can position themselves for consumer-grade home service robot products to capitalize on the coming wave of consumer demand.

2. Actionable insights: The AutoResearch automatic parameter tuning pipeline featured in the new model lowers R&D barriers and cuts manual R&D costs by more than 50%. Sellers can adopt this automated R&D framework to reduce their own R&D investment and improve development efficiency.

3. Risk note: The technology is still in a period of continuous iteration, with future development focused on edge inference, dual-arm robots, and mobile manipulation robots. Sellers need to keep up with technical iteration to avoid being sidelined by outdated product technology.

This new technology release brings new business opportunities and digital R&D insights for manufacturing factories, especially robot manufacturing facilities, with key takeaways as follows:

1. Business opportunity perspective: The new model drastically reduces the deployment cost of general-purpose robots, and demand for robots across scenarios including home services, flexible warehousing, and food processing is set to grow rapidly. Relevant factories can expand production lines for robot products targeting these scenarios in advance to capture growing order volume.

2. Product design and manufacturing perspective: The new model meets the strict requirements for real-time closed-loop control of robotic arms and can be deployed on consumer-grade graphics cards. When developing robotic arms and other robot products, factories can adopt this technical solution to reduce reliance on high-end hardware, cut R&D and manufacturing costs, and launch more price-competitive products.

3. Digital R&D insights: Automated training and parameter tuning pipelines enable lights-out R&D. When advancing digital R&D transformation, factories can reference this automated development workflow to reduce labor input, shorten R&D cycles, and improve overall development efficiency.

The general-purpose robot industry is moving from lab research to industrial deployment, with clearly defined industry pain points, and new technological advancements open up new business opportunities for service providers. Key takeaways are as follows:

1. Industry development trend: General embodied intelligence is currently the core development direction of the robot industry. As world action model technology matures, industrial adoption will accelerate significantly, driving sustained growth in demand for related technical services and creating broad room for market expansion.

2. Core customer pain points: Robot R&D companies universally face three pain points with traditional world action models: high parameter tuning costs, inference latency that fails to meet deployment requirements, and poor generalization that makes scenario adaptation difficult. These are the core demand points that service providers can target.

3. Reference for deployable solutions: The newly released model provides a mature solution framework that includes AutoResearch automatic parameter tuning, MoT mixture-of-experts architecture to decouple training and inference, and large-scale pre-training based on 10,000 hours of video data. Service providers can optimize and package this technical approach to offer to small and medium-sized robot R&D companies, helping customers cut costs and improve efficiency while allowing service providers to position themselves early to capture market share.

This technical breakthrough in general-purpose robots accelerates industrial adoption, offering valuable insights for the operation and strategic positioning of robot industry platforms. Key takeaways are as follows:

1. Market demand perspective: A large number of small and medium-sized robot R&D teams currently have strong demand for lower R&D costs and lower deployment barriers. Platforms can launch targeted supporting services such as R&D tool sharing, computing subsidies, and technical matchmaking to attract more developers to join the platform.

2. Investment and recruitment perspective: Following this technical breakthrough, the number of general-purpose robot startups will grow significantly, and the sector now has clear commercial potential. Platforms can prioritize recruiting general-purpose robot projects targeting home, warehousing, and catering scenarios to enrich platform ecology and secure an early ecological advantage in the new track.

3. Risk mitigation: General-purpose robot technology is still in the iteration phase, and many startups lack solid commercial deployment capabilities. When onboarding new projects, platforms should prioritize evaluating projects' real-world deployment capabilities and screen for teams with core technical expertise to reduce the platform's own operational risk.

This article discloses the latest research progress in the field of world action models, offering high reference value for both academic and industrial research in the general-purpose robot sector. Key takeaways are as follows:

1. New industry trends: World action models have now addressed the three core shortcomings of difficult parameter tuning, weak generalization, and slow deployment, achieving a balance between training convenience, general operation performance, and real-time deployment efficiency for the first time. General-purpose robot technology has officially moved out of the lab toward deployment across diverse scenarios, and the industrial value of general embodied intelligence is beginning to be unlocked.

2. Technical innovation direction: This work proposes three innovative solutions to the pain points of traditional models: the AutoResearch fully automated hyperparameter search training pipeline, the MoT mixture-of-experts architecture to decouple training and inference pipelines, and a pre-training base built from 10,000 hours of full-domain video. This provides a clear new path for future research in the field.

3. Clear future research directions: The paper points out that the industry will next iterate around optimizations for edge inference performance, adaptation for dual-arm robots, and adaptation for mobile manipulation robots. Researchers can focus their follow-up work on these directions to unlock greater industrial value in the sector.

Disclaimer: The "Quick Summary" content is entirely generated by AI. Please exercise discretion when interpreting the information. For issues or corrections, please email run@ebrun.com .

I am a Brand Seller Factory Service Provider Marketplace Seller Researcher Read it again.

通用机器人领域的世界动作模型(WAM)凭借 “预判未来画面” 的物理先验,显著提升机械臂抓取、多步骤操作泛化能力,但长期存在三大行业卡点:

模型同时建模视觉动态与动作序列,超参组合复杂,人工调参耗时耗力;

训练推理同流程,必须生成未来视频,算力开销巨大,推理延迟过高;

缺少大规模视觉、机器人动作联合预训练,真实场景泛化性不足。

极佳视界联合清华大学正式发布GigaWorld-Policy-0.5,针对性推出三大颠覆性突破:

搭载AutoResearch智能训练管线,全流程自动化超参搜索,摆脱人工反复调参,锁定最优训练配置;

以超10000小时全域视频数据完成底层预训练,海量学习自然场景物体运动、交互逻辑与空间变换规律,沉淀扎实通用世界动态先验;

基于MoT混合专家架构重构推理链路,配套算子深度优化,大幅削减推理计算量,RTX4090环境下推理延迟低至85ms。

图片

01 AutoResearch智能训练管线

全自动超参搜索,告别低效人工调参

传统WAM需要同时优化视觉动态预测、机器人动作生成两大相互牵制的训练目标,学习率、批次大小、预热步数、损失权重等数十项超参互相影响,单一参数改动就会同步改变视觉、动作两大分支的收敛效果。研发人员往往需要数十轮手动跑实验、反复对比验证,耗费大量算力资源与人力时间,很难快速找到兼顾建模效果与训练稳定性的最优配置。

GigaWorld-Policy-0.5配套自研Agent驱动的AutoResearch自动化研究管线,搭建完整无人化调参闭环:先通过1K步小规模试点训练批量遍历所有候选超参组合,快速筛除效果较差的配置;再以动作预测MSE、任务成功率为核心量化指标自动择优;最后基于筛选出的最优参数开展长周期完整训练,输出高性能模型权重。

整套流程无需人工盯守干预,可适配分拣、摆盘、长时序料理等各类机器人任务场景,大幅降低WAM算法研发门槛,减少超50%人工实验成本。

图片

02 MoT混合专家架构 + 算子深度优化

RTX4090推理延迟低至85ms

传统WAM训练与推理采用完全相同的计算链路,真机部署阶段仍要完整生成海量未来视频Token,带来巨额算力开销,主流基线推理延迟动辄200ms以上,部分模型甚至达到数秒,无法满足机械臂实时闭环控制的硬性要求。

GigaWorld-Policy-0.5创新采用MoT混合专家架构,将视觉动态建模、动作生成拆分为两套独立专家模块,采用动作中心化因果掩码实现训练、推理链路解耦:训练阶段双专家协同工作,利用未来视觉预测提供稠密监督;真机推理时直接跳过视觉专家、仅启动轻量化动作分支解码动作,省去视频生成全部计算量。同时团队配套完成全链路工程优化,基于KV缓存复用、算子编译加速,最终落地轻量化运行时,剥离Python运行冗余开销。

实测在RTX4090消费级显卡环境下,单轮推理延迟低至85ms,相比前代WAM模型速度提升超60%,无需高端算力即可支撑机械臂高速连续作业,大幅降低企业真机部署硬件成本。

图片

03 10000小时全域视频预训练底座

沉淀通用物理动态先验

模型泛化能力的核心根基,来源于大规模、多元化的预训练数据支撑。过往多数WAM仅依靠少量机器人仿真或真机轨迹训练,缺少对真实世界物体运动、空间交互的底层认知,面对陌生物体、全新工作台布局时极易出现操作失误。

GigaWorld-Policy-0.5视觉专家模块基于超10,000小时视频数据完成底层视频专家模型预训练,在此基础上,模型进一步引入超过2,000小时的动作—视频数据及数千小时机器人动作轨迹,开展AC-WM与WAM混合二次预训练。

训练数据广泛覆盖日常家居、仓储、餐饮等多类真实场景,让模型充分学习物体位移、碰撞、形变、空间位置变换等通用物理规律,通过联合建模世界动态与动作生成,深度耦合动作指令、机器人行为和视觉场景变化,使模型能够精准理解“何种动作会引发何种场景变化”,并据此生成有效的操作动作。

实验结果显示,依托这套万小时视频预训练体系,GigaWorld-Policy-0.5在真机任务表现上,相比上一代性能提升6%,运行速度提升62%。

图片

综合AutoResearch自动化训练管线、万小时级大规模视觉预训练、MoT架构 + 算子深度极速推理三大核心升级,GigaWorld-Policy-0.5补齐了传统世界动作模型调参难、泛化弱、落地慢的三大短板,首次实现训练便捷性、通用操作性能、实时部署效率三者兼顾。

这款全新迭代的WAM框架,让具备物理预判能力的通用机器人技术真正走出实验室,能够低成本落地家居服务、柔性仓储、餐饮加工等多元实操场景。后续极佳视界将持续迭代端侧推理性能,拓展双臂机器人、移动操作机器人适配方案,进一步释放世界动作模型在通用具身智能领域的产业价值。

注:文/龚作仁,文章来源:Laborer,本文为作者独立观点,不代表亿邦动力立场。

文章来源:Laborer

广告
微信
朋友圈

FAQ回顾

GigaWorld-Policy-0.5是什么?

它是极佳视界联合清华大学发布的通用机器人领域世界动作模型,搭载AutoResearch智能训练管线、MoT混合专家架构,基于超10000小时全域视频预训练,解决了传统模型调参难、泛化弱、落地慢的痛点,可低成本落地家居服务、柔性仓储等场景。

传统世界动作模型存在哪些行业痛点?

传统世界动作模型存在三大行业卡点:一是超参组合复杂,人工调参耗时耗力;二是训练推理同流程,必须生成未来视频,算力开销大、推理延迟高;三是缺少大规模视觉与机器人动作联合预训练,真实场景泛化性不足。

GigaWorld-Policy-0.5相比传统世界动作模型有什么优势?

它有三大核心升级:一是搭载AutoResearch智能训练管线,全自动超参搜索可减少超50%人工实验成本;二是采用MoT混合专家架构加算子优化,RTX4090下推理延迟低至85ms,速度提升超60%;三是经万小时视频预训练,真机任务性能较上一代提升6%,泛化性更强。

这么好看,分享一下?

朋友圈 分享

APP内打开

+1
+1
微信好友 朋友圈 新浪微博 QQ空间
关闭
收藏成功
发送
/140 0