广告
加载中

腾讯推出AI模型Gander 可实时对话同步处理后台任务

亿邦AI 2026-09-21 14:43
亿邦AI 2026/09/21 14:43

邦小白快读

EN
全文速览

这篇文章核心介绍了腾讯团队2026年9月推出的新型AI研究模型Gander的功能特点、当前进展,以及对话式AI领域的行业现状,能帮大家快速了解下一代语音智能助手的发展方向。

1. 和大家现在常用的需要轮流对话的语音助手不同,Gander可以在和用户实时对话的同时在后台处理复杂任务,支持用户随时打断对话,还能主动追问或同步任务进度,在公开测试里它对对话时机的把控表现最优,打断用户说话的比例仅为8%,远低于多数竞品。

2. 目前该项目的代码已经在GitHub上线并提供演示页面,感兴趣的用户可以直接在线体验,后续走完开源流程后会公开模型权重和训练数据,腾讯此前推出的同类AI模型已经落地微信、元宝等大家日常使用的产品。

3. 目前这类技术还处在早期阶段,存在任务准确率略低、图像计数和定位能力不足的短板,响应延迟、交互打断体验差仍是全行业待解决的问题,距离大规模成熟落地还有一段时间。

这篇文章披露的全双工AI技术进展与用户行为特征,能够为品牌布局智能交互服务、优化用户体验提供明确的参考方向。

1. 从用户行为与消费趋势看,用户使用对话式智能服务时随时打断是非常普遍的操作,资深用户打断比例达9%,新用户打断比例约5%,传统轮流交互的语音助手已经无法适配真实对话场景,低打扰、支持即时打断、对话同时同步处理任务的流畅交互体验,是用户对智能服务的核心需求。

2. 从产品研发方向看,将对话流程管控和复杂任务推理拆分的双模块设计已经成为行业主流技术路径,既可以保证秒级响应速度,又能灵活替换后台大模型实现能力升级,目前行业普遍存在响应延迟、打断体验差的痛点,音视频精准感知、任务准确率是需要重点补齐的短板。

3. 从落地参考看,腾讯的同类型AI技术已经落地微信、元宝等C端产品,品牌可以依托这类成熟技术搭建智能客服、智能导购系统,降低用户咨询等待成本。

这篇文章提到的全双工AI技术发展动向,能够为各平台卖家优化经营效率、挖掘新的经营机会提供实用参考。

1. 从需求变化与机会提示看,当前商家使用的传统智能客服、智能接待工具普遍存在响应延迟、无法承接用户随时打断需求的问题,能够边和用户实时对话、边后台处理订单查询、商品推荐、售后登记等任务的新型智能接待工具存在明确的市场缺口,尤其是微信生态内的卖家,可以重点关注腾讯智能体能力的开放动态,率先接入获得体验优势。

2. 从可复用的经验看,卖家搭建自有智能接待体系时,可以借鉴双模块拆分思路,前端用轻量模块保证秒级响应、承接用户打断需求,后台接入不同能力的模型处理复杂经营任务,无需重复训练前端模块就能灵活升级服务能力。

3. 从风险提示看,目前这类双模块AI系统仍处在早期发展阶段,任务准确率、音视频识别能力存在短板,行业也没有统一评估标准,卖家选型时不要盲目追新,要重点考察实际经营场景下的任务完成准确率。

这篇文章披露的新一代全双工AI技术发展趋势,能够为工厂研发智能硬件、推进数字化转型、挖掘新的商业机会提供方向指引。

1. 从产品设计生产需求看,传统搭载轮流交互语音助手的智能音箱、智能家居、车载语音设备已经无法适配用户边对话边发指令、随时打断的真实使用场景,搭载支持全双工交互、双模块运行的AI系统将成为下一代智能硬件的核心升级方向,相关硬件厂商可提前布局适配接口。

2. 从商业机会看,当前全双工AI系统仍在技术快速迭代期,Gander这类模型的后台推理模块可灵活替换接入不同智能体系统,硬件厂商可提前对接技术团队,推出适配该类技术的新一代智能交互硬件抢占市场;模型训练需要覆盖不同背景噪音场景的海量样本,相关场景测试、试点落地服务也有明确合作空间。

3. 从数字化转型启示看,工厂可以借鉴双模块拆分思路搭建智能生产管控系统,前端轻量模块做实时生产流程响应,后台模块做复杂产能调度、质量检测推理,兼顾响应速度和复杂任务处理能力。

这篇文章梳理的全双工对话AI技术路径与行业共性痛点,能够为AI服务商、企业服务厂商优化产品能力、捕捉市场需求提供清晰参考。

1. 从行业发展趋势看,拆分对话管控与复杂任务处理的双模块架构已经成为腾讯、OpenAI、Sakana AI等全球头部厂商的共同技术选择,后台推理模块可灵活替换不同大模型、无需重新训练对话模块的设计,能够大幅降低技术迭代成本,是下一代对话式智能体的主流技术方向。

2. 从客户痛点看,目前所有对话式语音、聊天智能体开发团队都普遍面临响应延迟、用户打断处理体验差的问题,现有产品普遍存在打断用户比例过高、对话时机把控不准的短板,同时双模块系统目前还缺乏统一的行业评估标准,客户选型缺乏参考依据。

3. 从解决方案参考看,服务商可以借鉴双模块设计思路,用秒级响应的轻量前端模块承接对话流管控、打断识别处理,搭配可灵活替换的后台推理模块处理复杂任务,同时针对现有模型在音视频理解、任务准确率上的短板做定向优化,形成差异化竞争力。

这篇文章披露的头部企业智能体布局动作、行业共性问题,能够为各类型平台布局内置智能功能、完善平台生态提供实用参考。

1. 从行业最新做法看,腾讯持续在开源大模型、智能体领域加码,此前发布的Hy3开源模型已经落地微信、元宝、WorkBuddy等自有平台产品,同时正谈判拟拿下智能体初创公司Manus的最大股权,匹配微信内置智能体的业务规划;新推出的Gander模型已上线开源代码,后续将开放权重与训练数据。

2. 从平台用户的真实需求看,当前平台内置的对话式智能助手普遍存在响应延迟、打断交互体验差的问题,数据显示资深用户使用智能助手时打断比例达9%,传统轮流交互的模式已经无法满足用户自然对话的需求。

3. 从运营与风险规避提示看,目前双模块全双工智能体仍处在发展早期,尚无明确的规模化落地方案,也缺乏统一评估标准,平台布局相关功能时不要盲目追求参数规模,要重点打磨交互流畅度与任务准确率,可通过开源合作、投资初创团队的方式快速补全技术能力,降低自研成本。

这篇文章系统梳理了全双工对话AI领域的最新技术进展、产业布局与现存问题,能够为相关领域研究者提供前沿研究方向与产业观察素材。

1. 从产业新动向看,拆分对话流管控与复杂任务推理的双模块架构已经成为全球头部科技企业的共同技术选择,腾讯Gander、OpenAI GPT-Live、Sakana AI Fugu等模型均采用类似思路,后台推理模块可灵活插拔不同大模型的设计大幅降低了系统迭代成本,全双工交互、支持随时打断、边对话边处理后台任务成为下一代智能体的核心发展方向。

2. 从领域新问题看,现有双模块系统仍存在任务准确率易受语音识别环节拖累、音视频精准感知能力不足、规模化落地方案缺失的问题,行业暂未形成针对双模块系统的统一评估标准,交互打断、响应延迟仍是全行业待解的共性难题,测试数据显示主流产品打断用户比例从8%到48%不等,体验差距较大。

3. 从研究参考价值看,Gander模型通过约270万个样本完成训练,专门加入噪音场景、非面向模型发言场景的训练样本实现恰当静默的效果,相关训练思路、小脑加大脑的模块设计具备较高研究价值,目前项目代码已上线GitHub,后续将开放权重与训练数据供行业研究使用。

返回默认

声明:快读内容全程由AI生成,请注意甄别信息。如您发现问题,请发送邮件至 run@ebrun.com 。

我是 品牌商 卖家 工厂 服务商 平台商 研究者 帮我再读一遍。

Quick Summary

This article outlines the core features and current development status of Gander, a new AI research model launched by Tencent’s team in September 2026, alongside the broader state of the conversational AI sector, offering readers a quick overview of where next-generation voice assistants are headed.

1. Unlike the turn-taking voice assistants most people use today, Gander can process complex tasks in the background while holding real-time conversations with users. It supports seamless user interruption, and can proactively ask follow-up questions or share task progress updates. In public testing, it delivered the industry’s strongest performance in conversation timing: it interrupted users only 8% of the time, a rate far lower than most competing products.

2. The project’s code is already available on GitHub alongside a live demo, so interested users can test the model directly online. Tencent plans to release the model weights and training data once the full open-source process is complete. The company’s previous generation of comparable AI models has already been integrated into widely used consumer products including WeChat and Yuanbao.

3. This category of technology remains in its early stages. Current limitations include slightly lower task accuracy and weak performance in image counting and localization. Response latency and poor interruption handling remain industry-wide pain points, meaning the technology is still some distance away from large-scale, mature commercial deployment.

The advances in full-duplex AI technology and user behavior insights outlined in this article provide clear guidance for brands planning intelligent interactive services and optimizing user experience.

1. From a user behavior and consumer trend perspective, interrupting mid-interaction is a very common action when users engage with conversational intelligent services: power users interrupt 9% of the time, while new users do so around 5% of the time. Traditional turn-taking voice assistants can no longer fit real-world conversation scenarios. A low-interruption, smooth interactive experience that supports instant interruption and processes tasks concurrently during dialogue is the core user demand for intelligent services.

2. From a product R&D perspective, the dual-module design that separates conversation flow control from complex task reasoning has become the dominant industry technical path. This structure delivers sub-second response speeds while allowing flexible upgrades to backend large language models. At present, widespread industry pain points include response latency and poor interruption experience, with precise audio and video perception and task accuracy as key capabilities to strengthen.

3. For deployment reference, Tencent’s comparable AI technology has already been rolled out in consumer-facing products including WeChat and Yuanbao. Brands can leverage this mature technology to build intelligent customer service and smart shopping guide systems, cutting user wait times for inquiry support.

The development trends in full-duplex AI technology covered in this article offer practical guidance for sellers across platforms to improve operational efficiency and identify new business opportunities.

1. From a demand shift and opportunity perspective, the traditional intelligent customer service and reception tools most merchants use today generally suffer from slow response times and an inability to support users’ need to interrupt mid-conversation. There is clear market demand for a new generation of intelligent reception tools that can handle real-time dialogue while processing backend tasks including order queries, product recommendations and after-sales registration in parallel. Sellers operating in the WeChat ecosystem in particular should closely follow updates on Tencent’s open intelligent agent capabilities, as early integration can deliver a tangible experience advantage.

2. From a reusable experience perspective, when building proprietary intelligent reception systems, sellers can adopt the dual-module split design: use a lightweight frontend module to guarantee sub-second responses and handle user interruptions, and connect backend models with different capabilities to process complex operational tasks. This structure allows flexible service capability upgrades without retraining the frontend module.

3. As a risk note, this category of dual-module AI systems remains in an early development stage, with shortcomings in task accuracy and audio/video recognition, and no unified industry evaluation standard. Sellers should not blindly chase novelty when selecting tools, and instead prioritize testing task completion accuracy in real operating scenarios.

The development trends of next-generation full-duplex AI technology covered in this article provide direction for factories developing smart hardware, advancing digital transformation, and identifying new commercial opportunities.

1. From a product design and manufacturing demand perspective, traditional smart speakers, smart home devices and in-car voice systems equipped with turn-taking voice assistants can no longer adapt to real usage scenarios where users give commands while talking and interrupt interactions freely. Integrating AI systems that support full-duplex interaction and dual-module operation will become the core upgrade direction for next-generation smart hardware, and related hardware manufacturers can pre-develop compatible interfaces to prepare.

2. From a commercial opportunity perspective, full-duplex AI systems are still in a period of rapid technological iteration. The backend reasoning modules of models such as Gander can be flexibly replaced and connected to different intelligent agent systems, so hardware manufacturers can engage with technical teams early to launch new interactive smart hardware adapted to this technology and capture market share. There is also clear cooperation potential for scenario testing and pilot deployment services, as model training requires massive volumes of samples covering diverse background noise scenarios.

3. For digital transformation insights, factories can reference the dual-module split structure to build intelligent production management and control systems: use a lightweight frontend module for real-time production flow responses, and a backend module for complex capacity scheduling and quality inspection reasoning, balancing fast response speeds and complex task processing capabilities.

The technical paths for full-duplex conversational AI and common industry pain points outlined in this article offer clear reference for AI service providers and enterprise service vendors to optimize product capabilities and capture market demand.

1. From an industry development trend perspective, the dual-module architecture that separates conversation control from complex task processing has become a shared technical choice for leading global players including Tencent, OpenAI and Sakana AI. The design, which allows flexible replacement of backend reasoning modules with different large models without retraining the conversation module, can significantly reduce technology iteration costs, and is set to become the mainstream technical direction for next-generation conversational agents.

2. From a customer pain point perspective, all development teams building conversational voice and chat agents currently face widespread issues with response latency and poor user interruption handling. Existing products generally suffer from overly high user interruption rates and inaccurate conversation timing judgment. Meanwhile, there is no unified industry evaluation standard for dual-module systems, leaving customers with no clear reference for vendor selection.

3. For solution reference, service providers can adopt the dual-module design: use a lightweight frontend module with sub-second response to manage conversation flow and process interruption detection, paired with a flexibly replaceable backend reasoning module to handle complex tasks. Providers can also build differentiated competitiveness by making targeted improvements to existing models’ shortcomings in audio/video understanding and task accuracy.

The moves by leading enterprises in intelligent agent deployment and common industry problems covered in this article offer practical reference for platforms of all types to build built-in intelligent features and improve their platform ecosystems.

1. From recent industry practices, Tencent is continuing to scale up investment in open-source large models and intelligent agents. Its previously released Hy3 open-source model has already been deployed across its own platform products including WeChat, Yuanbao and WorkBuddy, and the company is in talks to acquire a majority stake in intelligent agent startup Manus, aligned with its business plan to build built-in agents for WeChat. Its newly launched Gander model has already released open-source code, with model weights and training data set to be opened in the future.

2. From the perspective of real platform user needs, the conversational intelligent assistants currently built into most platforms generally suffer from response latency and poor interruption interaction experiences. Data shows power users interrupt smart assistants 9% of the time, and the traditional turn-taking interaction model can no longer meet users’ demand for natural conversation.

3. For operation and risk mitigation guidance, dual-module full-duplex intelligent agents remain in an early development stage, with no clear large-scale deployment plans and no unified evaluation standard. When rolling out related features, platforms should not blindly pursue larger parameter scales, but instead prioritize polishing interaction fluency and task accuracy. They can quickly fill technical capability gaps and reduce in-house R&D costs through open-source collaboration and investments in startup teams.

This article systematically reviews the latest technical advances, industry deployment landscape and remaining challenges in the full-duplex conversational AI field, providing cutting-edge research directions and industry observation material for researchers in related domains.

1. From new industry trends, the dual-module architecture that splits conversation flow control from complex task reasoning has become a shared technical choice for leading global technology firms. Models including Tencent’s Gander, OpenAI’s GPT-Live and Sakana AI’s Fugu all adopt a similar design, where the flexibly swappable backend reasoning module that supports integration with different large models sharply reduces system iteration costs. Full-duplex interaction, support for seamless interruption, and parallel backend task processing during conversation have emerged as core development directions for next-generation intelligent agents.

2. From unresolved field challenges, existing dual-module systems still face issues including task accuracy that is easily dragged down by errors in the speech recognition link, insufficient precise audio and video perception capabilities, and a lack of scalable large-scale deployment solutions. The industry has not yet formed a unified evaluation standard for dual-module systems, while interaction interruption handling and response latency remain common, unresolved industry-wide problems. Test data shows the rate at which mainstream products interrupt users ranges from 8% to 48%, indicating large experience gaps across offerings.

3. For research reference value, the Gander model was trained on approximately 2.7 million samples, with targeted training samples covering noisy scenarios and utterances not directed at the model to achieve appropriate silent response behavior. Its training approach and “cerebellum + cerebrum” modular design hold notable research value. The project code is now available on GitHub, and its weights and training data will be released for industry research use in the future.

Disclaimer: The "Quick Summary" content is entirely generated by AI. Please exercise discretion when interpreting the information. For issues or corrections, please email run@ebrun.com .

I am a Brand Seller Factory Service Provider Marketplace Seller Researcher Read it again.

2026年9月,腾讯混元语音团队联合多家高校研究人员推出AI研究模型Gander,可在与用户保持实时对话的同时,于后台处理复杂智能体任务。

当前主流语音助手多采用与用户轮流交互的模式,无法适配真实对话中双方随时打断、边听边说的交互场景。Gander可同步接收处理语音、图像、文本输入,即便是在自身输出语音的阶段,也能持续处理多模态信息,支持用户随时中断对话,模型无需用户提示即可主动追问问题或同步任务进度。

为平衡对话响应速度与复杂任务推理时长,Gander采用双模块拆分设计,两个模块参照人体结构被命名为“小脑”与“大脑”。其中“小脑”按秒级管控对话流程,决定倾听、发言或在用户打断时停止输出,依托近两分钟的对话内容作为记忆,无需单独的语音起止检测模块即可完成决策。“大脑”则负责在后台完成推理及复杂智能体任务,该模块可直接替换为Codex、Claude Code等不同智能体系统,无需重新训练对话模块,测试中该位置曾接入OpenAI GPT-5.6系列的某款未披露具体名称的模型,底层模型升级即可带动整套系统能力提升。

在现有公开基准测试中,Gander在Full-Duplex-Bench v3语音助手场景测试中取得对话时机表现最优成绩,100个测试场景下均能在恰当时间点开口发言,打断用户的比例为8%,低于GPT-Realtime的13.5%,对比表现最弱的竞品近48%的打断比例优势明显。该模型以相对较小的参数规模,即可对标GPT-Realtime、Gemini Live、Grok等商用系统。

测试同时暴露了模型的短板。Gander的任务准确率略落后于竞品,相关技术报告内容显示,现有测试对整套系统统一打分,语音识别和输出环节的误差也会拉低最终得分,若直接给“大脑”模块输入文本,其得分会有明显提升。受训练方向侧重流畅对话而非精准感知影响,Gander的音视频理解能力弱于其基础模型,在物体计数、图像定位类任务上表现存在不足。

据技术报告披露,Gander共使用约270万个样本完成训练,部分训练样本用于教会模型在存在背景噪音、对话场景中无人面向模型发言时保持安静。目前该项目的代码已在GitHub上线并提供演示页面,研究团队待开源流程完成后,将对外公开模型权重与训练数据。目前该项目仍处于早期阶段,如何完成规模化落地尚无明确方案,行业内也暂未形成针对此类双模块系统的统一评估标准。

今年7月腾讯曾发布开源语言模型Hy3,缩小了与头部竞品在智能体任务上的能力差距,该模型已落地WorkBuddy、元宝、微信等产品。目前腾讯正谈判拟拿下智能体初创公司Manus的最大股权,此前北京方面叫停了Meta对该公司的收购,该交易被腾讯视为与自身微信内置智能体等业务规划相匹配。

行业层面,拆分对话与任务处理模块已成为多家厂商的技术选择。OpenAI的GPT-Live将对话流程与推理环节分离,在对话持续进行的同时,将网页搜索、智能体任务交由后台模型处理。Sakana AI推出的Fugu模型作为独立语言模型,可从可扩展的模型池中调用其他模型完成任务。OpenAI同时在测试主动式智能体,可自动创建后续任务,无需用户主动触发即可联系用户。

实际落地中,交互打断与响应延迟仍是普遍待解的问题。Anthropic的相关分析数据显示,资深用户在使用Claude Code时,约9%的工作步骤会出现打断操作,新用户的打断比例约为5%,相关行业调研结果显示,对话式语音及聊天智能体开发团队普遍反馈延迟问题较为突出。

本文首发于 亿邦动力 官方网站

文章来源:亿邦动力

广告
微信
朋友圈

FAQ回顾

腾讯Gander是什么?

Gander是2026年9月腾讯混元语音团队联合多家高校研发的AI研究模型,采用“小脑”“大脑”双模块拆分设计,可在与用户实时对话的同时,于后台处理复杂多模态智能体任务,支持用户随时中断对话。

Gander和主流语音助手相比有什么优势?

主流语音助手多采用轮流交互模式,无法适配真实对话中随时打断、边听边说的场景。Gander可同步处理语音、图像、文本输入,输出语音阶段也能持续处理信息,可主动追问或同步任务进度,在Full-Duplex-Bench v3测试中打断用户比例仅8%,低于GPT-Realtime的13.5%,表现优于多数竞品。

Gander AI模型目前存在哪些短板?

Gander当前任务准确率略落后于竞品;受训练方向侧重流畅对话影响,音视频理解能力弱于其基础模型,在物体计数、图像定位类任务上表现不足;项目仍处于早期阶段,暂无明确规模化落地方案。

当前对话式智能体技术的主流发展方向是什么?

拆分对话流程与任务处理模块已成为行业主流技术选择,OpenAI GPT-Live、Sakana AI Fugu等模型均采用类似架构,可在对话进行时将搜索、复杂任务交由后台模型处理。据Anthropic统计,用户使用智能体时普遍存在打断操作,交互打断、响应延迟仍是行业待解问题。

这么好看,分享一下?

朋友圈 分享

APP内打开

+1
+1
微信好友 朋友圈 新浪微博 QQ空间
关闭
收藏成功
发送
/140 0