广告
加载中

腾讯清华联合测试 主流AI搜索准确率最高不足五成

亿邦AI 2026-07-07 10:08
亿邦AI 2026/07/07 10:08

邦小白快读

EN
全文速览

本次腾讯混元联合清华大学公开了主流AI搜索代理的最新测试结果,核心干货如下:

1. 当前AI搜索的核心问题不是搜索环节出错,而是用户查询存在歧义时,极少主动向用户申请澄清,单个未解决的歧义就会导致整个推理链失效,反复检索不澄清的效果还不如直接猜测。

2. 无歧义提示前提下,所有参与测试的主流大模型AI搜索准确率都不足五成,最高的豆包Seed 2.0 Pro仅为43.1%,表现较弱的模型准确率仅10%以上。仅靠大模型内置知识库根本无法完成这类带歧义的搜索任务,切断搜索权限后准确率会暴跌。

3. 先搜索再追问澄清歧义的模型成功率高达93.4%,远高于不追问直接猜测的56.5%,未来AI搜索一定会增加用户交互澄清环节,使用AI搜索时遇到模糊问题可以主动给模型补充信息提升准确率。

本次测试结果对品牌布局AI搜索流量、适配新搜索环境有重要参考价值,核心干货如下:

1. 当前AI搜索对信息清晰度要求极高,移除查询中的歧义后,各模型准确率可提升26.8至40.2个百分点,说明品牌信息的明确性直接决定AI搜索能否正确输出品牌相关结果,如果品牌产品参数、场景标注模糊,会大幅降低被AI检索推荐的概率。

2. AI搜索高度依赖外部公开搜索信息,切断搜索工具权限后头部模型准确率也会暴跌近一半,说明品牌的公开网络信息布局仍然是AI搜索时代获客的核心,不需要过度依赖大模型内置知识库的训练。

3. 未来AI搜索会新增歧义澄清交互环节,品牌需要提前适配新规则,梳理不同场景下用户对品牌产品的模糊搜索需求,标准化自身公开信息,提前抢占AI搜索的流量先机。

本次测试结果给布局AI搜索流量渠道的卖家明确了风险与机会,核心干货如下:

1. 当前AI搜索处理用户模糊查询的能力极差,单个未解决歧义就会导致整个搜索结果出错,如果卖家的产品信息标注模糊、关键词不统一、参数不明确,会大幅降低被AI搜索正确推荐的概率,这是当前布局AI搜索流量的核心风险。

2. 机会层面,现有AI搜索整体准确率不足五成,只要卖家明确标准化自身产品信息,消除信息层面的歧义,就能让产品被AI检索推荐的准确率提升20到40个百分点,远高于同行,可提前抢占AI搜索的早期流量红利。

3. 目前头部大模型厂商已经开始针对歧义处理短板做优化,未来AI搜索一定会成为新的重要流量渠道,卖家需要跟进技术变化,及时调整自身线上信息的发布规则,适配新的搜索逻辑。

本次测试结论给推进数字化和电商转型的工厂带来多方面启示,核心干货如下:

1. 线上获客层面,AI搜索已经成为新的用户流量入口,当前AI搜索对信息清晰度要求极高,工厂在发布产品信息时,需要明确标注产品参数、适用场景、版本型号等信息,避免因信息模糊导致AI检索错误,降低获客效果。

2. 数字化转型层面,AI搜索技术可帮助工厂提升供应链信息检索、市场需求调研的效率,但目前AI搜索对非标准化信息的处理能力很差,工厂内部推进信息标准化改造,能大幅提升AI工具的使用效率,充分发挥数字化工具的价值。

3. 商业机会层面,当前通用AI搜索在专业生产场景的歧义处理能力极弱,有技术能力的工厂可以探索开发适配自身生产领域的专业AI搜索工具,满足行业内模糊信息检索的需求,打造新的增长曲线。

本次测试明确了AI搜索领域的核心痛点和发展方向,给AI技术服务商指明了研发方向,核心干货如下:

1. 当前行业核心痛点:主流大模型的AI搜索在处理真实场景的歧义查询时能力极差,无歧义提示下平均端到端准确率仅28.6%,即便是头部模型准确率也不到五成,单个未解决歧义就会导致整个结果失效,用户体验极差,市场对成熟的优化方案有强烈需求。

2. 技术研发方向:仅靠提示引导模型识别歧义只能提升歧义识别率,无法有效提升最终检索准确率,未来研发的核心是建立将搜索不确定性转化为有效用户交互的机制,提升模型主动澄清歧义的能力。

3. 目前头部厂商已经探索出可行的优化方向,比如Anthropic提升不确定性标记频率、Perplexity重构搜索工作流,服务商可参考这些方向结合自身优势开发针对性解决方案,抢占市场先机。

本次测试揭露了AI搜索平台当前的核心问题,指明了平台优化运营的方向,核心干货如下:

1. 用户核心需求:真实场景中用户的搜索查询大多带有模糊、歧义甚至错误,现有AI搜索无法满足该核心需求,平台需要把歧义交互能力作为AI搜索的核心优化方向,才能提升用户满意度和留存率。

2. 风险规避:当前所有模型都高度依赖外部搜索工具,仅靠内置知识库无法完成搜索任务,切断搜索权限后头部模型准确率也会暴跌,平台要规避过度依赖内置知识库、放弃外部搜索对接的错误方向。

3. 生态与运营:当前AI搜索的歧义处理技术还不成熟,平台可以引入相关技术服务商共建生态,填补技术短板;在模型测试环节,可以采用本次推出的DiscoBench基准,该基准贴合中文网络搜索的典型特征,能更准确反映模型的真实能力,帮助平台筛选优质模型。

本次研究为AI搜索领域带来了新的研究成果与方向,核心干货如下:

1. 产业新动向:当前AI搜索代理的核心瓶颈不是检索和推理能力,而是歧义处理与用户交互能力,参与测试的所有主流模型端到端准确率最高不足五成,技术成熟度较低,还有很大的研发创新空间。

2. 研究领域新进展:此前业内常用的测试基准都假设用户查询无歧义,不符合真实搜索场景,本次推出的DiscoBench测试基准填补了该领域空白,覆盖11个知识领域,包含211项任务共463个中文歧义点,还将歧义分为四大类,更贴合中文搜索的实际特征。

3. 未来研究方向:研究发现歧义识别能力与追问质量不存在强关联,仅提升歧义识别率无法有效提升最终搜索准确率,未来需要重点研究如何让模型将识别到的不确定性转化为有效的用户澄清交互,这是该领域新的核心研究问题。

返回默认

声明:快读内容全程由AI生成,请注意甄别信息。如您发现问题,请发送邮件至 run@ebrun.com 。

我是 品牌商 卖家 工厂 服务商 平台商 研究者 帮我再读一遍。

Quick Summary

Tencent Hunyuan and Tsinghua University have jointly released the latest test results of mainstream AI search agents. Key findings are as follows:

1. The core problem of current AI search does not lie in errors during the retrieval phase. When user queries contain ambiguity, very few AI search systems proactively ask users for clarification. A single unresolved ambiguity can break the entire reasoning chain, and repeated retrieval without clarification performs worse than simply guessing an answer.

2. Even when given unambiguous prompts, all tested mainstream large model-powered AI search tools have an accuracy rate below 50%. Doubao Seed 2.0 Pro, the top performer, only scored 43.1%, while weaker models only reached just above 10%. Large models' built-in knowledge bases alone are completely insufficient for ambiguous search tasks, and accuracy plummets when search access is disabled.

3. Models that search first and then ask follow-up questions to clarify ambiguity achieve a 93.4% success rate, far exceeding the 56.5% success rate of guessing without follow-ups. AI search will definitely add interactive user clarification steps in the future, and users can proactively provide additional context to AI search when asking vague questions to improve accuracy.

These test results offer key insights for brands positioning themselves for AI search traffic and adapting to the new search environment. Key takeaways are as follows:

1. Current AI search requires extremely clear information. Removing ambiguity from queries can improve model accuracy by 26.8 to 40.2 percentage points, meaning the clarity of a brand's information directly determines whether AI search can correctly output brand-related results. Vague product parameters and scenario labels will sharply reduce the probability of a brand being retrieved and recommended by AI.

2. AI search relies heavily on public external search information. Even top-tier models see their accuracy drop by nearly half when access to search tools is cut off. This shows that building out public online brand information remains the core of customer acquisition in the AI search era, and brands do not need to over-rely on fine-tuning large models' built-in knowledge bases.

3. AI search will add interactive ambiguity clarification steps in the future. Brands should adapt to this new rule in advance by mapping users' vague search queries for their products across different scenarios, standardizing their public information, and getting a head start in capturing AI search traffic.

These test results clarify the risks and opportunities for sellers looking to tap into AI search as a new traffic channel. Key insights are as follows:

1. Current AI search performs extremely poorly when handling ambiguous user queries, and a single unresolved ambiguity can lead to completely incorrect search results. If a seller's product information is vaguely labeled, uses inconsistent keywords or unclear parameters, their chances of being correctly recommended by AI search will drop sharply, which is the core risk of entering AI search traffic today.

2. On the opportunity side, the overall accuracy of existing AI search is below 50%. As long as sellers standardize their product information and eliminate informational ambiguity, they can boost the accuracy of their products being retrieved and recommended by 20 to 40 percentage points, outperforming most peers and capturing early AI search traffic dividends ahead of competitors.

3. Leading large model developers are already optimizing their ambiguity handling capabilities, and AI search will certainly become an important new traffic channel. Sellers should keep up with technological changes, adjust their online information publishing rules in a timely manner, and adapt to the new search logic.

These test conclusions bring multiple insights for factories advancing digital transformation and e-commerce expansion. Key takeaways are as follows:

1. For online customer acquisition: AI search has become a new user traffic entry point, and current AI search requires extremely clear information. When publishing product information, factories should clearly label product parameters, applicable scenarios, versions and models to avoid AI retrieval errors caused by vague information that hurt customer acquisition performance.

2. For digital transformation: AI search technology can help factories improve the efficiency of supply chain information retrieval and market demand research, but current AI search performs poorly when handling unstandardized information. Factories that standardize internal information will see greatly improved efficiency when using AI tools, and can fully unlock the value of digital tools.

3. For new business opportunities: General-purpose AI search performs extremely poorly at handling ambiguity in professional production scenarios. Factories with technical capabilities can explore developing industry-specific AI search tools tailored to their production field to meet the demand for ambiguous information retrieval within the sector, and build a new growth curve.

This test clarifies the core pain points and development direction of the AI search field, and points out clear R&D directions for AI technology service providers. Key findings are as follows:

1. Core industry pain point: Mainstream large model-powered AI search performs extremely poorly when handling ambiguous queries in real-world scenarios. The average end-to-end accuracy is only 28.6% with unambiguous prompts, and even top models have accuracy below 50%. A single unresolved ambiguity can invalidate the entire result, leading to poor user experience, and the market has strong demand for mature optimization solutions.

2. Technology R&D direction: Prompting models to identify ambiguity only improves ambiguity recognition rates, and cannot effectively boost final retrieval accuracy. The core of future R&D is building a mechanism that converts search uncertainty into effective user interaction, to improve models' ability to proactively clarify ambiguity.

3. Leading players have already explored viable optimization directions: for example, Anthropic has increased the frequency of uncertainty tagging, and Perplexity has restructured its search workflow. Service providers can reference these directions to develop targeted solutions that match their own advantages, and capture first-mover advantage in the market.

This test reveals the core problems of current AI search platforms, and points out the direction for platform operation optimization. Key insights are as follows:

1. Core user demand: Most real-world user search queries are vague, ambiguous, or even incorrect, and existing AI search cannot meet this core user need. Platforms must prioritize improving interactive ambiguity clarification as a core optimization direction to boost user satisfaction and retention.

2. Risk mitigation: All current models rely heavily on external search tools, and built-in knowledge bases alone cannot complete search tasks. Even top models see accuracy plummet when search access is removed. Platforms should avoid the wrong strategic direction of over-relying on built-in knowledge bases and abandoning external search integration.

3. Ecosystem and operations: Ambiguity handling technology for AI search is still immature. Platforms can partner with relevant technology service providers to co-build an ecosystem and fill technical gaps. For model testing, platforms can adopt the newly released DiscoBench benchmark, which aligns with the typical characteristics of Chinese web search, can more accurately reflect models' real-world capabilities, and helps platforms screen for high-performing models.

This study delivers new research findings and directions for the AI search field. Key contributions are as follows:

1. New industry insight: The core bottleneck of current AI search agents is not retrieval or reasoning capability, but ambiguity handling and user interaction. The highest end-to-end accuracy among all tested mainstream models is below 50%, meaning technical maturity remains low and there is still large room for R&D and innovation.

2. New progress in research: All previously commonly used industry test benchmarks assume that user queries are unambiguous, which does not match real search scenarios. The newly launched DiscoBench benchmark fills this gap. It covers 11 knowledge domains, includes 211 tasks with 463 Chinese ambiguity instances, categorizes ambiguities into four main types, and is far more aligned with the actual characteristics of Chinese search.

3. Future research direction: This study finds that there is no strong correlation between ambiguity recognition ability and the quality of clarification follow-ups. Only improving ambiguity recognition rates cannot effectively boost final search accuracy. Future research should focus on how models can convert recognized uncertainty into effective user clarification interactions, which is the new core research question in this field.

Disclaimer: The "Quick Summary" content is entirely generated by AI. Please exercise discretion when interpreting the information. For issues or corrections, please email run@ebrun.com .

I am a Brand Seller Factory Service Provider Marketplace Seller Researcher Read it again.

2026年7月,腾讯混元联合清华大学团队推出全新AI搜索代理测试基准DiscoBench,相关测试结果于近期公开。测试结果显示AI搜索代理极少在多步检索任务的搜索环节出错,核心问题集中在用户查询存在歧义时,未主动向用户申请澄清。反复检索的表现往往低于直接猜测。

此前业内常用的GAIA BrowseComp等基准均假设用户查询完整无歧义,而现实场景中用户查询常存在模糊、不完整甚至事实错误的情况,长推理链中未解决的歧义会不断累积,引导搜索代理偏离正确路径。如果模型在早期节点选择错误实体,后续检索语法再规范也无法匹配真实需求。

该基准覆盖11个知识领域,包含211项任务共463个歧义点,研究团队将歧义分为四类,分别是描述匹配多个实体、适用不同时段或版本、存在多个有效排序评估标准、包含事实错误。数据集以中文为主,匹配中文网络搜索典型特征。测试过程中,代理在每个检查点可选择继续搜索、向用户申请澄清、直接给出答案三个动作,当代理提出有效追问时,由大语言模型驱动的用户模拟器会释放预设线索缩小搜索范围,所有搜索请求通过Tavily搜索引擎运行,Gemini 3 Flash承担模拟器角色。

团队测试了过去半年发布的11款大模型,无歧义提示的前提下,豆包Seed 2.0 Pro端到端准确率最高为43.1%,Gemini 3.1 Pro以40.8%紧随其后,Claude Opus 4.7为39.8%,表现较弱的MiniMax M2.7和Qwen3.6 Max准确率仅分别为16.1%和12.3%。部分模型单步骤得分与整体结果存在差距,Claude Opus 4.7单检查点正确率达57%,但端到端准确率仅39.8%,单个未解决的歧义就足以导致整个推理链失效。

团队还增设引导模式测试,系统提示明确要求代理关注歧义,存疑时发起追问。10款测试模型平均端到端准确率从28.6%上升至33.7%,歧义检测F1值从45.3%大幅提升至64.9%,提示仅帮助模型识别歧义,未有效提升最终检索成功率。Claude Opus 4.7在引导模式下,尽管单检查点通过率上升,端到端准确率甚至出现小幅下降。

行为特征分析拆解了代理在歧义检查点的选择差异,先搜索再追问的模型平均成功率达93.4%,不追问直接猜测的成功率为56.5%,反复搜索仍不追问直接猜测的模型成功率仅51.9%。反复搜索行为说明模型已识别到歧义,但未转化为用户交互动作。

歧义识别能力与追问质量不存在强关联,Qwen3.6 Max歧义检测F1值仅16%,平均每任务仅发起0.07次追问,但发起的追问中94.7%符合事实,89.5%可推动检索进展。MiniMax M2.7追问频率高,但追问有效率仅60.7%至66.5%。四类歧义中,事实错误最易检测,实体歧义与标准歧义难度更高。

测试同时显示,切断搜索工具访问权限后,豆包Seed 2.0 Pro准确率从43.1%降至2.4%,Gemini 3.1 Pro从40.8%降至19.9%,DiscoBench任务无法仅通过模型内置知识库完成。移除查询中的歧义后,各模型准确率可提升26.8至40.2个百分点,研究团队认为未来搜索代理除检索和推理能力外,还需建立将搜索不确定性转化为用户交互的机制。

目前已有厂商针对该类短板推出优化方案,Anthropic最新更新的Claude Opus 4.8已提升不确定性标记频率,代码遗留bug概率较前代下降约四分之三。Perplexity推出Search as Code功能,允许模型将搜索工作流编写为Python程序,而非调用预设API。

文章来源:亿邦动力

广告
微信
朋友圈

FAQ回顾

当前主流AI搜索的准确率大概是多少?

2026年腾讯混元联合清华大学的DiscoBench测试显示,参与测试的11款主流大模型在无歧义提示前提下,AI搜索端到端准确率最高为豆包Seed 2.0 Pro的43.1%,不足五成,表现较弱的模型准确率仅12%左右。

导致AI搜索准确率低的核心原因是什么?

当前AI搜索的核心问题集中在用户查询存在歧义时,未主动向用户申请澄清。长推理链中未解决的歧义会不断累积,引导搜索代理偏离正确路径,单个未解决的歧义就足以导致整个推理链失效。

DiscoBench测试基准有什么特点?

DiscoBench是腾讯混元联合清华大学推出的AI搜索代理测试基准,覆盖11个知识领域,包含211项任务共463个歧义点,数据集以中文为主,匹配中文网络搜索典型特征,更贴合用户查询真实场景。

如何有效提升AI搜索的准确率?

AI搜索代理在遇到歧义时先搜索再追问的平均成功率达93.4%,远高于不追问直接猜测的表现;移除查询中的歧义后,各模型准确率可提升26.8至40.2个百分点,厂商也可针对性优化歧义交互短板。

这么好看,分享一下?

朋友圈 分享

APP内打开

赞 +1
+1
微信好友 朋友圈 新浪微博 QQ空间
关闭
收藏成功
发送
/140 0