广告
加载中

中国智能体登顶OSWorld 90.2%成功率反超海外巨头

龚作仁 2026-08-05 13:02
龚作仁 2026/08/05 13:02

邦小白快读

EN
全文速览

本文核心重点是中国实在智能自研的实在Agent拿下全球AI智能体评测OSWorld冠军,干货内容如下:

1. 核心成绩:实在Agent以90.2%的任务成功率登顶,是该榜单发布以来首个突破90%成功率的智能体,成功率反超OpenAI、Meta、Anthropic等海外硅谷巨头,90%是行业公认从“勉强能跑”的实验状态迈入稳定可用、可靠落地产业阶段的临界点。

2. 行业赛道变化:当前AI竞赛已经完成切换,从早期拼模型参数大小、拼单步推理能力,转向拼工程落地能力,核心比拼Harness能力,也就是让AI能稳定把事做成的能力。

3. 实在Agent的实际价值:它深耕GUI屏幕操作八年,不管系统有没有现成接口都能操作,能帮用户快速完成复杂任务,比如十几分钟就能搞定过去需要几个月才能完成的数据整理工作。

本文给AI赛道品牌商梳理了产业趋势、产品研发方向和竞争机会,干货内容如下:

1. 产业与消费趋势:当前AI已经从概念转向落地,企业端用户核心需求从“能聊天的AI”变成“能干活的AI”,用户愿意为能稳定完成复杂任务、适配各类复杂IT环境的产品付费,国产智能体已经从追赶者变成定义者,赛道机会广阔。

2. 产品研发方向:要重视Harness工程层建设,不能只一味堆大模型参数,Harness作为连接大模型和场景的中间层,能把大模型能力转化为稳定交付,还能把可服务的企业场景从有API的系统拓宽到无接口的老旧系统,整体市场空间从约2000亿美元扩展至1.5万亿美元。

3. 竞争策略:可以深耕工程落地赛道,避开海外巨头在通用大模型基座上的优势,在工程化、GUI多模态操控上建立自身壁垒。

本文给AI相关卖家、To B服务卖家梳理了智能体赛道的机会、方向和风险,干货内容如下:

1. 增长市场机会:当前AI智能体已经正式迈入产业落地阶段,企业需求旺盛,数据显示国内桌面端办公Agent月访问量从2026年3月的超2000万次增长到6月的超6000万次,三个月增长两倍,赛道增长速度极快。

2. 需求痛点与新商业模式:传统纯API对接模式存在成本高、稳定性差的痛点,老旧系统无接口开发费用高,API调整就容易瘫痪,Harness加GUI结合的模式解决了这个痛点,有接口走接口无接口做屏幕操作,大幅降低后期维护成本,适配绝大多数企业的IT环境,可落地性更强。

3. 风险提示:如果依旧只拼模型参数不重视工程落地能力,会逐步被市场淘汰,当前行业已经进入“比谁能干活”的阶段,需要尽早布局工程体系。

本文给制造类工厂带来了数字化转型的启示和商业机会,干货内容如下:

1. 适配工厂的真实需求:大多数工厂都存在大量运行多年的老旧产线内部系统、不同品牌的工业系统,开发接口成本极高,一个工业系统接口开发费用就可能高达20万,实在Agent这类Harness架构的智能体不用开发接口,就能像人一样看懂屏幕操作软件,获取需要的产线数据,完美匹配工厂的痛点需求。

2. 实际效率提升非常明显:原来工厂整理产线数据、上新系统,需要等待数月开发接口,现在只需要给智能体下达指令,十几分钟就能生成整理好的报表,大幅缩短流程,降低时间和人力成本。

3. 数字化转型启示:工厂不需要花大价钱全部替换现有老旧系统,也不用投入高额成本开发大量接口,借助成熟的AI智能体能力,就能低成本快速推进数字化落地,降低转型的成本和风险。

本文给AI服务商、数字化服务商梳理了行业趋势、客户痛点和解决方案参考,干货内容如下:

1. 行业发展趋势:AI行业对落地的认知已经完成系统性升级,从早期的Prompt Engineering到Context Engineering,现在已经进入Harness Engineering的核心范式,Harness作为将大模型能力转化为稳定交付的中间层,市场空间从原来的约2000亿美元扩展到1.5万亿美元,是接下来服务商的核心赛道。

2. 客户核心痛点:企业IT环境复杂,大量老旧系统没有公开接口,纯API对接不仅开发成本高,而且稳定性差,API只要微调就可能导致系统瘫痪,后期维护成本高到难以承受,传统解决方案无法满足客户需求。

3. 可参考的解决方案:走API加GUI融合的路线,不站队不设限,有接口走接口保证效率,没有接口就用屏幕语义理解技术实现拟人操作,同时积累真实场景的异常数据,迭代优化自身的Harness工程体系,提升复杂场景下的稳定性。

本文给AI平台商梳理了产业需求、布局方向和风险规避要点,干货内容如下:

1. 产业对平台的核心需求:当前通用大模型的能力已经逐步成熟,产业端核心需求是能连接大模型和各类企业复杂场景的中间层平台,解决大模型“能说不能干”落不了地的问题,平台的核心价值转向赋能落地。

2. 行业最新布局风向:国内头部科技企业阿里、腾讯都已经完成团队调整整合,核心方向都是强化模型与用户场景之间的Harness层,Harness已经成为行业共识的核心布局方向。

3. 风向规避和发展方向:不要盲目跟风和海外巨头比拼通用大模型参数,要发挥国内产业落地积累的优势,聚焦工程化能力建设,同时智能体是全产业的系统工程,平台可以牵头共建产业生态,争夺下一代AI基础设施的定义权,建立自身的竞争壁垒。

本文给人工智能和产业方向研究者提供了智能体赛道的最新产业动向和研究方向,干货内容如下:

1. 产业最新动向:全球AI智能体竞赛已经完成赛道切换,从原来比拼模型参数大小、单步推理能力,转向比拼工程化落地能力也就是Harness能力,中国实在智能的实在Agent以90.2%的成功率登顶全球权威评测基准OSWorld,首次突破90%临界点,成功反超硅谷巨头,国产智能体已经从追赶者变成赛道定义者。

2. 新的研究方向:现有实证研究显示,同一底层模型因为Harness设计不同,智能体表现差距可达6%到17%,编码任务解决率可从约5%跃升到30%以上,Harness工程、GUI多模态交互、真实场景任务落地都是接下来值得深入研究的新方向。

3. 商业模式与产业研究启示:智能体的核心竞争力不在模型参数,而在工程落地能力,未来Harness会成为AI产业的新底座,掌握Harness工程范式就能掌握下一代企业级软件的定价权,为产业研究提供了新的核心课题。

返回默认

声明:快读内容全程由AI生成,请注意甄别信息。如您发现问题,请发送邮件至 run@ebrun.com 。

我是 品牌商 卖家 工厂 服务商 平台商 研究者 帮我再读一遍。

Quick Summary

This article highlights that Chinese AI firm RealAI's independently developed Shizhen Agent has won first place in OSWorld, the global AI agent benchmark evaluation. Key takeaways are as follows:

1. Core results: Shizhen Agent claimed the top spot with a 90.2% task success rate, becoming the first AI agent ever to break the 90% threshold since the benchmark launched. It outperformed Silicon Valley giants including OpenAI, Meta and Anthropic. The 90% mark is widely recognized as the critical threshold that moves AI agents from a "barely functional" experimental state to a stable, reliable stage ready for industrial deployment.

2. Shift in the AI race: The focus of AI competition has shifted from competing on model size and single-step reasoning capability to engineering and implementation capabilities. The core competition now centers on Harness capability, which refers to an AI's ability to complete tasks reliably.

3. Practical value: Built on eight years of expertise in GUI screen operation, Shizhen Agent can operate across systems regardless of whether open APIs are available, enabling users to complete complex tasks quickly. For example, it can finish in 10-odd minutes data sorting work that previously took months to complete.

This article outlines industry trends, product R&D directions and competitive opportunities for AI brands, with key insights below:

1. Industry and consumer trends: AI has now shifted from concept to real-world deployment. Enterprise users' core demand has changed from "AI that can chat" to "AI that can work". Users are willing to pay for products that can reliably complete complex tasks and adapt to diverse complex IT environments. Domestic Chinese AI agents have evolved from followers to trailblazers that define the new industry standard, creating huge room for growth in the sector.

2. Product R&D direction: Brands should prioritize building out the Harness engineering layer rather than just scaling up large model parameters. As an intermediate layer connecting large models to application scenarios, Harness converts large model capabilities into stable, deliverable solutions. It also expands the addressable enterprise market from systems with available APIs to include legacy systems without open interfaces, growing the overall market size from roughly $200 billion to $1.5 trillion.

3. Competitive strategy: Brands can focus on the engineering implementation track to avoid competing head-on with overseas giants' advantages in general-purpose large model bases, and build their own competitive moats in engineering and multimodal GUI manipulation.

This article breaks down opportunities, directions and risks in the AI agent track for AI and B2B service sellers, with key takeaways below:

1. High-growth market opportunity: AI agents have officially entered the industrial implementation phase, with robust enterprise demand. Data shows that monthly visits to domestic desktop office agents grew from over 20 million in March 2026 to over 60 million in June 2026, a three-fold increase in just three months, indicating extremely rapid sector growth.

2. Pain points and new business models: The traditional pure API integration model suffers from high costs and poor stability. Developing interfaces for legacy systems is very expensive, and even minor API adjustments can lead to system failure. The combined Harness+GUI model solves this pain point: it uses APIs when available, and operates directly via the screen when interfaces are unavailable, drastically cutting downstream maintenance costs, adapting to the vast majority of enterprise IT environments, and offering far stronger deployability.

3. Risk warning: Sellers who still focus solely on model parameters and neglect implementation capability will gradually be phased out by the market. The industry has now entered an era where capability to deliver real work is the core differentiator, so players need to build out their engineering systems as soon as possible.

This article outlines insights and business opportunities from the AI agent track for manufacturing factories pursuing digital transformation, with key takeaways below:

1. Alignment with factories' real needs: Most factories operate many long-running legacy internal production line systems and industrial systems from different vendors, and developing custom interfaces for these systems is extremely expensive—developing a single industrial system interface can cost up to 200,000 RMB. Harness-architecture agents like Shizhen Agent do not require custom interface development; instead, they can read screens and operate software just like a human worker to extract required production line data, perfectly matching the core pain points of factories.

2. Dramatic efficiency gains: Previously, factories had to wait months for interface development to organize production line data and integrate new systems. Now, users can simply issue a command to the AI agent, which generates a well-organized report in 10-odd minutes, drastically shortening workflows and cutting both time and labor costs.

3. Insights for digital transformation: Factories do not need to spend heavily to fully replace their existing legacy systems, nor invest huge sums in developing large numbers of custom interfaces. Leveraging capable mature AI agents allows them to advance digital implementation quickly at low cost, reducing both the cost and risk of transformation.

This article sorts out industry trends, customer pain points and solution references for AI and digital service providers, with key insights below:

1. Industry development trends: The AI industry has completed a systematic upgrade in its understanding of implementation. It has evolved from early Prompt Engineering to Context Engineering, and now entered the core paradigm of Harness Engineering. As the intermediate layer that converts large model capabilities into stable, deliverable solutions, Harness has grown its total addressable market from roughly $200 billion to $1.5 trillion, making it the core growth track for service providers going forward.

2. Core customer pain points: Enterprise IT environments are highly complex, with a large number of legacy systems lacking open interfaces. Pure API integration not only comes with high development costs, but also poor stability: even minor API adjustments can cause system outages, and downstream maintenance costs become unsustainable. Traditional solutions cannot meet customer requirements.

3. Reference solution: Providers should adopt an integrated API+GUI approach that avoids vendor lock-in and remains flexible: use APIs when available to ensure efficiency, and use semantic screen understanding technology for human-like operation when no interface exists. At the same time, providers should accumulate anomaly data from real-world scenarios to iterate and optimize their own Harness engineering systems, and improve stability in complex scenarios.

This article sorts out industrial demand, layout directions and risk mitigation points for AI platform providers, with key insights below:

1. Core industrial demand for platforms: Capabilities of general-purpose large models have gradually matured, so the core demand from the industrial side is for intermediate layer platforms that can connect large models to complex enterprise scenarios, solving the problem of large models being "able to talk but unable to act" for real-world implementation. The core value of platforms has now shifted to enabling implementation.

2. Latest industry layout trends: Leading Chinese tech firms including Alibaba and Tencent have already completed internal team adjustments and integrations, with a core focus on strengthening the Harness layer between models and user scenarios. Harness is already a widely agreed core layout direction across the industry.

3. Risk mitigation and development direction: Platforms should not blindly compete with overseas giants on general-purpose large model parameters. Instead, they should leverage China's advantages in industrial implementation experience, and focus on building engineering capabilities. At the same time, as agents are a cross-industry systems engineering project, platforms can take the lead in building a joint industrial ecosystem, compete for the right to define the next generation of AI infrastructure, and build their own competitive moats.

This article shares the latest industry developments and research directions in the AI agent track for artificial intelligence and industry researchers, with key insights below:

1. Latest industry developments: The global AI agent competition has completed a track shift. It has moved from competing on model size and single-step reasoning capability to competing on engineering implementation capability, or Harness capability. China's RealAI has claimed the top spot on the globally authoritative OSWorld benchmark with a 90.2% success rate, marking the first time any agent has broken the 90% threshold and outperformed Silicon Valley giants. Chinese domestic AI agents have evolved from industry followers to track definers.

2. New research directions: Existing empirical research shows that for the same base model, different Harness designs can lead to a 6% to 17% gap in agent performance, and coding task resolution can jump from roughly 5% to over 30%. Harness engineering, multimodal GUI interaction, and real-world task implementation are all promising new directions for in-depth research going forward.

3. Insights for business model and industrial research: The core competitiveness of AI agents lies not in model parameters, but in engineering implementation capability. Harness will become the new foundational layer of the AI industry in the future, and mastery of the Harness engineering paradigm will bring pricing power over the next generation of enterprise software, opening up a new core research topic for industrial research.

Disclaimer: The "Quick Summary" content is entirely generated by AI. Please exercise discretion when interpreting the information. For issues or corrections, please email run@ebrun.com .

I am a Brand Seller Factory Service Provider Marketplace Seller Researcher Read it again.

2026年7月27日,中国智能体“实在Agent”以90.2%的任务成功率登顶OSWorld全球总榜,同时拿下智能体分榜双料冠军。Meta、Anthropic、OpenAI等硅谷巨头的公开最高纪录,集体被一家中国公司反超。

据了解,这家公司是来自浙江的实在智能。其自研的“实在Agent”成为OSWorld自2024年发布以来首个突破90%成功率的Computer-Use Agent。90%是行业公认的关键临界点——在此之下,智能体还停留在“勉强能跑”的实验状态;跨过这条线,才意味着迈入稳定可用、可靠落地的产业阶段。

这场暗流之下的角力,终于有了一个清晰的结果。

01从“堆模型”到“拼工程”:AI竞赛的赛道切换

2023年,所有人都在问“哪个模型参数最多”;2024年,问题变成“哪个模型推理最强”;到了2026年,产业界的追问已经变成——“哪个智能体最能干活”。

OSWorld的出现,正是为了回答这个问题。

该榜单由香港大学、卡内基梅隆大学、滑铁卢大学等机构于2024年发表于人工智能顶会NeurIPS,是目前全球公认的Computer-Use Agent核心评测基准。

与仅支持网页交互或API调用的传统基准不同,OSWorld在真实Ubuntu虚拟机中部署了大量应用程序,设置了361个实战任务,覆盖办公、编程、UI设计、系统运维等高频场景。任务是否完成由机器自动判定,无任何人工主观干预。

正因如此,OSWorld成了全球科技巨头验证智能体能力的“必争之地”。OpenAI发布GPT-5.4时强调“首次超越人类”,Google发布Gemini 3.6 Flash时标注“83%全场最高分”,Anthropic发布Claude Sonnet 5时突出“81.2%”。

复盘OSWorld的迭代曲线,整个行业的进步一目了然:

●2024年,最强智能体成功率仅12.2%

●2025年先后攀升至22.0%、72.6%,首次超越人类操作基线

●2026年5月推高至83.6%

● 两个月后,实在Agent以90.2%登顶

两年时间,从12%到90%。但真正值得追问的不是“谁跑得更快”,而是“为什么跑得更快”。拆解实在Agent的得分,答案逐渐清晰:

跨应用协同:该任务模拟的是真实办公中最繁琐的跨系统流程,要求智能体在多个软件间连续穿梭。实在Agent在93个任务中成功率84.7%,领先第二名近10个百分点。

非标界面图像处理:GIMP界面充满浮动面板和非标准工具栏,堪称视觉识别的“噩梦级”场景。实在Agent在26个任务中拿下24个,成功率92.3%。

系统底层操作:24个任务实现100%满分零失误,覆盖进程管理、权限修改、命令行操作。

从跨应用到复杂界面,再到底层系统——这些“硬骨头”场景拼的不是模型“能想多深”,而是工程“能做多稳”。行业给这种工程能力起了一个名字:Harness,它决定了AI能不能把“想做的事”真正“做成”。

这一判断有扎实的实证支撑。斯坦福大学与清华大学的研究表明,在同一底层模型的前提下,仅因Harness设计不同,智能体的表现差距可达6%至17%。在编码基准测试中,同一模型仅改变Harness设计,解决率可从约5%跃升至30%以上。

02 Harness:AI的操作系统,实在智能的先行者之路

Harness可以通俗地理解为AI的“操作系统”——就像电脑有了Windows才能稳定运行各种软件一样,AI有了Harness,才能可靠地完成一系列复杂操作。任务怎么拆、工具怎么调、出错了怎么补救——所有这些让AI“不掉链子”的机制,都属于Harness的范畴。

Harness并非一个新概念,但它在2026年成为AI工程领域的核心范式:从Prompt Engineering到Context Engineering,再到Harness Engineering,行业对“如何让模型稳定交付”的认知正在完成一次系统性升级。

中信证券在近期研报中指出,Agent走向长任务执行与多智能体协同后,状态维护、错误传播与Token成本加速上升,模型单步能力强并不等于能稳定交付。Harness是将底层智能转化为稳定交付的中间层,可触达空间由约2000亿美元扩展至1.5万亿美元。换言之,过去AI只能服务有标准接口的企业,而Harness让那些‘没有API’的老旧系统也能被AI接管,企业场景的可触达范围被大幅拓宽

就如同DeepSeek除了模型核心技术,工程化能力是其大幅降低成本的关键一样,Harness对智能体的意义同样如此——它决定了AI能不能稳定、可控、低成本地“把事办成”。

在Harness这个赛道上,成立于2018年的实在智能是先行者。当行业还在追捧纯API对接时,公司创始人兼CEO孙林君做出了一个在今天看来极具前瞻性的判断。

“我们做智能体,就围绕第一性原理——你不能干的,我能不能干?谁的能力边界更广,谁的成本更低、效率更高?”

这个判断背后是对企业IT环境的深刻理解。孙林君在今年的世界人工智能大会上算过一笔账:像西门子这样的工业系统,开发一个接口费用可能高达20万。更不用说那些运行了几十年的内部系统,外面根本看不到接口。而电商平台的API策略三天两头微调,哪怕只是一个3%概率的A/B测试,纯API对接的系统就可能当场瘫痪。

“纯靠API对接的产品,后期维护成本高到无法想象,最后大多只能给客户退款。”

所以实在智能做实在Agent的逻辑很简单——不站队,不设限。有现成接口的就走接口,高效顺畅;没有接口的,AI就切换成“人肉模式”——像真人一样看屏幕、动鼠标、敲键盘,把数据拿回来,把事办成。

用户感知层面,变化是颠覆性的——过去需要IT部门花数月开发接口才能连通的系统,现在智能体可以直接“看屏幕”操作。一位制造业客户曾对孙林君说:“以前上个新系统要等半年,现在告诉Agent‘把上个月的产线数据整理成报表’,十分钟就回来了。”

从用户角度看,Harness的价值很简单:AI不再是一个需要你维护的“工具”,而是一个能自己动手干活的“同事”。

03八年坚守:从GUI到AGI的终局判断

2025年底至2026年初,Harness完成从概念到共识的跃迁。但实在智能在这条路上的探索,远早于行业的集体转向。

“技术圈有一种流行的看法:只有通过API和MCP对接才是先进的,像人一样看屏幕、点鼠标的GUI路线,只是过渡期的‘妥协方案’。” 孙林君直言。

但实在智能从一开始就把GUI作为核心能力来建设。“这不是技术路线的站队问题,而是对AI终局的判断问题。”

这个判断的核心逻辑是,AGI要实现真正的通用智能,一定需要在多模态能力上有所突破——因为只有像人一样感知和操作这个世界,AI才能真正理解这个世界。

在孙林君看来,AI的终局不在参数里,而在对真实世界的感知与交互中。“未来要想实现AGI,必须先让AI认识世界、接触世界、操作世界。” 这里的“世界”,首先就是人类每天都在面对的数字世界——操作系统、软件界面、业务流程。

正是基于这一判断,实在Agent在API之外深耕GUI八年。实在Agent通过自研的ISSUT(智能屏幕语义理解)技术和融合拾取能力,让AI能像人一样“看懂”屏幕上的UI界面,在没有数据接口的情况下也能精准获取数据、操作软件。

八年时间,实在Agent沉淀下来的是一整套“会生长的”Harness工程体系——

真实场景的数据飞轮:实在Agent已服务超6000家企业客户,覆盖制造、电商、能源、跨境、医药、物流运输等核心行业。每一次真实场景的运行,都是一次Harness的练兵——异常报错如何处理,流程卡顿如何恢复,边界案例如何兜底。千万级异常场景的迭代优化,让实在Agent具备了实验室模型难以企及的“实战免疫力”。

从自动化到Agent的工程跃迁:实在智能的八年,恰好也是企业自动化向AI Agent演进的关键周期。自动化时代的流程编排、异常处理、系统集成经验,被完整地继承并升级为Agent时代的Harness能力。这种“工程基因”的延续,是纯粹的模型公司短期内难以复制的。

“模型是可插拔的‘大脑’,但光有大脑远远不够。真正决定AI能不能‘动手干活’的,是Harness工程能力——如何拆解复杂任务、如何调度各类软件工具、如何在执行出错时自我纠错、如何组织多个智能体协同作业。”

实在智能的八年坚守,本质上是对一个判断的反复验证:AI的终局不在模型参数里,而在工程落地的细节里。

04国产智能体:从单点突破到产业底座

实在Agent登顶OSWorld,不仅是单个企业的技术突破,更是国产智能体产业的一个标志性事件。

格局正在变化。近期阿里将QoderWork、悟空与MuleRun整合至千问办公,腾讯推动WorkBuddy与QClaw团队收敛。上述动作形式各异,但本质上均在强化位于模型与用户场景之间的Harness层。易观分析发布的《中国办公智能体平台市场研究报告2026》显示,中国桌面端办公Agent月访问量由2026年3月超2000万次增至6月超6000万次。

国产智能体,正在从“追赶者”变成“定义者”。

但压力同样存在。在通用大模型基座上,硅谷巨头仍凭借算力与资本优势保持领跑。中国企业要在智能体赛道实现超越,必须在工程化落地、GUI多模态操控与真实工作流打通上建立不可替代的壁垒。

实在智能的突破,恰恰证明了这条路径的可行性。

从2018年成立,到2023年推出行业首个智能体,再到2026年登顶OSWorld。实在智能的每一步,都在回答同一个问题:如何让AI从“能聊天”变成“能干活”。而OSWorld 90.2%的成绩,是对这个问题最有力的回答。

AI智能体的竞争不是一家公司对另一家公司的较量,而是整个产业生态的系统工程。中国在智能体工程化落地上的积累,是参与这场全球竞赛的核心筹码,行业需要形成合力、共建生态,才能在下一代AI基础设施的定义权上占据主动。

一位人工智能领域资深分析人士向记者表示:“实在Agent登顶OSWorld的意义,不只是一次刷榜。它证明了在Harness工程这条路上,中国团队有能力定义标准。当AI智能体网络成为继电网、互联网之后人类历史上的第三张全球网络,成为所有产业之下的新底座时,谁掌握了Harness的工程范式,谁就掌握了下一代企业级软件的定价权。从这个角度看,实在智能的突破,是国产智能体产业的一个起点,而不是终点。”

注:文/龚作仁,文章来源:Laborer,本文为作者独立观点,不代表亿邦动力立场。

文章来源:Laborer

广告
微信
朋友圈

FAQ回顾

OSWorld是什么类型的评测基准?

OSWorld是香港大学、卡内基梅隆大学等机构2024年发表于人工智能顶会NeurIPS的Computer-Use Agent核心评测基准,在真实Ubuntu虚拟机中设置361个覆盖办公、编程、运维等场景的实战任务,结果由机器自动判定,是全球公认的智能体能力验证标准。

AI领域的Harness是什么?有什么作用?

Harness俗称AI的操作系统,是衔接底层模型与用户场景的中间层,负责任务拆解、工具调度、错误补救等保障AI稳定交付的机制,可让智能体在无API接口的老旧系统也能稳定作业,能将AI企业场景可触达空间从2000亿美元扩展至1.5万亿美元。

实在Agent登顶OSWorld有什么行业意义?

实在Agent是OSWorld发布以来首个突破90%成功率的Computer-Use Agent,90%是行业公认的智能体从实验状态迈入稳定可用产业阶段的临界点,这一突破也证明中国团队在Harness工程领域已具备定义行业标准的能力。

当前AI智能体赛道的核心竞争壁垒是什么?

当前AI智能体赛道的竞争已从模型参数比拼转向Harness工程能力竞争,同一底层模型下仅Harness设计差异就能带来6%-17%的表现差距,工程化落地能力、GUI多模态操控能力、真实工作流打通能力是核心竞争壁垒。

这么好看,分享一下?

朋友圈 分享

APP内打开

+1
+1
微信好友 朋友圈 新浪微博 QQ空间
关闭
收藏成功
发送
/140 0