广告
加载中

传音TEX AI语音算法团队斩获MLC-SLM 2026国际挑战赛Task 1亚军

龚作仁 2026-08-04 10:22
龚作仁 2026/08/04 10:22

邦小白快读

EN
全文速览

本文核心信息是传音TEX AI中心语音算法团队获得MLC-SLM 2026国际挑战赛Task 1赛道亚军,该赛事是国际语音领域顶级会议的权威赛事,本次总结分享了参赛的技术思路和成果干货。

1.本次赛事共有全球110支队伍参赛,任务要求处理未切分、未标注说话人的真实多人长对话音频,输出带说话人信息的转写结果,评价指标同时考察准确率、位置和归属,难度很高,传音最终以tcpMER 15.41%的成绩拿到全球第二。

2.针对赛事难点,团队没有只优化单一模型,而是搭建了三级级联系统,两条链路分别处理说话人区分和高精度转写,最终融合生成结果,形成了全流程闭环优化能力。

3.目前该技术已经应用到传音智能终端的多个场景,未来还会拓展更多应用方向。

本文对布局出海新兴市场的消费电子品牌,在产品研发和技术竞争力打造上有较多参考干货。

1.产品研发方向参考:新兴市场语言环境复杂,多语种语音交互是用户体验的核心痛点,提前布局多语种语音识别、交互技术,能帮助品牌打造差异化体验优势。传音已经构建了覆盖多种语言、本地口音和方言变体的技术能力,落地到多个终端场景验证了价值。

2.技术研发路径参考:品牌研发技术不用局限单一技术路线,传音采用级联系统+端到端方案并行探索的路线,通过增量式实验逐步迭代优化指标,从22.78%降到15.41%,这种稳扎稳打的研发思路可借鉴。

3.品牌营销参考:拿下国际顶级赛事的好成绩,可以作为品牌技术实力的背书,用于品牌营销强化专业认知。

本文对出海布局消费电子业务的卖家,在机会判断、能力布局上有较多参考干货。

1.机会提示:出海新兴市场的差异化竞争突破口已经显现,复杂多语言环境下的优质语音交互体验是用户未被充分满足的需求,提前布局相关能力能获得竞争优势。

2.可学习的经验:针对多语种语音技术研发,不要局限在单一模型优化,可参考传音双路线并行探索、全链路闭环打磨、增量实验逐步优化的思路,降低研发试错成本,逐步提升效果。

3.风险提示:随着智能终端的发展,语音交互能力已经成为核心竞争力之一,缺乏相关技术布局的卖家很容易在后续产品竞争中陷入劣势,要么提前自研,要么对接具备成熟技术的供应链补齐能力。

本文对出海相关的消费电子代工厂、布局自有终端业务的工厂,有不少参考干货。

1.产品生产设计需求参考:当前面向新兴市场出海的智能终端,已经明确要求配套成熟的多语种语音交互能力,工厂在产品设计和生产配套阶段,需要把多语种语音识别、交互能力纳入硬性需求,才能匹配品牌客户的要求。

2.商业机会参考:多语种语音技术落地需要硬件端的适配整合,具备硬件整合能力的工厂,可以对接有技术输出需求的AI技术企业、品牌方,拿到更多出海相关的订单,拓展自身业务边界。

3.数字化研发启示:工厂推进技术升级和数字化研发,可以参考传音从基线出发、增量式实验逐步优化的思路,稳扎稳打打磨能力,降低研发试错成本,逐步提升技术水平。

本文对做AI语音技术服务、智能终端技术配套的服务商,明确了行业方向和落地参考干货。

1.行业发展趋势:当前多语种真实场景语音处理已经成为行业核心研发方向,国际领域也将真实多人对话场景的全链路语音技术作为重点考察方向,相关需求会持续增长。

2.客户核心痛点:出海品牌面向新兴市场,需要能处理未标注长音频、说话人区分、多语种转写的全链路语音解决方案,单一环节的技术优势无法满足落地需求,全链路协同能力才是客户真正需要的。

3.可参考的解决方案:技术路线可参考传音的级联系统方案,通过两条链路分别处理说话人结构区分和高精度转写,最终融合输出结果;同时可布局级联+端到端并行研发路线,为后续技术迭代预留空间,适配不同业务场景需求。

本文对做AI技术服务平台、出海消费电子招商平台,在平台运营和布局上有不少参考干货。

1.商家需求洞察:当前入驻平台的出海品牌,普遍有多语种全链路语音技术的需求,平台需要提前布局相关的技术商家供给,才能满足客户需求,提升平台竞争力。

2.招商方向参考:平台可以重点引入类似传音这类具备成熟多语种语音技术全链路落地能力的企业入驻,丰富平台的技术服务品类,吸引更多有出海需求的品牌客户入驻,扩大平台规模。

3.风险规避提示:多语种语音技术落地非常依赖全链路协同能力,单一技术点的优势无法满足客户落地需求,平台在引入相关技术商家时,需要重点考察商家的全链路工程落地能力,避免引入不成熟技术影响平台口碑,带来运营风险。

本文对AI语音领域的产业和学术研究者,呈现了多语种语音领域的最新产业研动向,有较高参考价值。

1.领域新动向:当前多语种对话语音模型领域,核心研究挑战已经转向真实场景的落地问题,本次国际赛事将未预切分、未标注说话人的真实长对话处理作为核心赛道,全链路算法协同能力是领域考察的核心方向,吸引了全球110支队伍参与,足见热度。

2.技术路线参考:传音提出的三级级联系统技术路线,通过两条链路分别处理说话人结构恢复和多语种转写,最终融合输出结果,实现了优异的指标,同时采用级联+端到端并行探索的路线,为后续技术演进预留空间,这种研发思路值得研究者参考。

3.产业化方向参考:多语种语音技术已经在新兴市场智能终端场景实现大规模落地,能解决真实用户痛点,产生明确商业价值,明确了相关技术的产业化落地方向。

返回默认

声明:快读内容全程由AI生成,请注意甄别信息。如您发现问题,请发送邮件至 run@ebrun.com 。

我是 品牌商 卖家 工厂 服务商 平台商 研究者 帮我再读一遍。

Quick Summary

This article shares the key takeaways and technical insights from Transsion's TEX AI Center speech algorithm team, which won second place in Task 1 of the MLC-SLM 2026 International Challenge, a top authoritative competition in the global speech technology field.

1. A total of 110 teams from around the world participated in the challenge. The task required participants to process unsegmented, unannotated long-form real multi-speaker conversation audio, and output transcription results with speaker information. Evaluation covers accuracy, timestamp alignment and speaker attribution, making it an extremely challenging task. Transsion secured second place globally with a tcpMER of 15.41%.

2. To address the challenge's difficulties, the team did not just optimize a single model. Instead, it built a three-stage cascaded system with two separate pipelines to handle speaker diarization and high-accuracy transcription respectively, before fusing the outputs to generate the final result. This approach enabled full closed-loop optimization across the entire processing workflow.

3. This technology is already deployed in multiple scenarios across Transsion's smart devices, and will be expanded to more application areas in the future.

This article offers valuable actionable insights for consumer electronics brands targeting emerging overseas markets, particularly for product R&D and building technological competitiveness.

1. Product R&D guidance: The language environment in emerging markets is highly complex, and multi-lingual voice interaction is a core pain point for user experience. Proactive investment in multi-lingual speech recognition and interaction technology allows brands to build a differentiated experience advantage. Transsion has already built technical capabilities covering multiple languages, local accents and dialect variants, and has validated its value through deployment across multiple device scenarios.

2. R&D strategy guidance: Brands do not need to limit themselves to a single technical route. Transsion adopted a parallel exploration approach combining cascaded systems and end-to-end solutions, iteratively optimizing performance through incremental experiments, reducing tcpMER from 22.78% to 15.41%. This steady, methodical R&D approach is highly replicable.

3. Brand marketing guidance: Top results at leading international competitions serve as strong third-party validation of a brand's technical strength, and can be leveraged in brand marketing to reinforce professional positioning.

This article provides useful reference insights for sellers building consumer electronics businesses in overseas markets, covering opportunity assessment and capability building.

1. Opportunity outlook: A clear breakout point for differentiated competition in emerging overseas markets has emerged: high-quality voice interaction experiences in complex multi-lingual environments remains an unmet user need. Proactive investment in related capabilities can deliver significant competitive advantages.

2. Actionable lessons: For multi-lingual speech technology R&D, avoid limiting work to single-model optimization. You can follow Transsion's approach of parallel exploration of dual technical routes, full closed-loop workflow refinement, and incremental iterative optimization to reduce R&D trial-and-error costs while steadily improving performance.

3. Risk warning: As smart devices evolve, voice interaction capability has become a core competitive differentiator. Sellers without relevant technical investment will likely fall behind in future product competition. Brands should either invest in independent R&D proactively, or partner with supply chain players with mature technology to fill this capability gap.

This article offers valuable insights for consumer electronics contract manufacturers serving overseas markets, and factories building their own end-device businesses.

1. Product design and manufacturing guidance: For smart devices exported to emerging markets today, mature multi-lingual voice interaction capability has become a clear required specification. Factories need to include multi-lingual speech recognition and interaction capabilities as hard requirements during the product design and manufacturing preparation stage to meet brand clients' demands.

2. Business opportunity outlook: Deploying multi-lingual speech technology requires hardware integration and adaptation. Factories with existing hardware integration capabilities can partner with AI technology companies and brands looking for technology output partnerships to secure more overseas-related orders and expand their business scope.

3. Digital R&D inspiration: For factories advancing technological upgrading and digital R&D, Transsion's approach of starting from a solid baseline and incremental iterative optimization is a useful reference. This method allows teams to steadily refine capabilities, reduce R&D trial-and-error costs, and steadily improve technical standards.

This article clarifies industry direction and provides deployment reference for AI voice technology service providers and smart device technology supporting service providers.

1. Industry development trend: Multi-lingual real-scenario speech processing has become the core R&D direction for the industry. The international community also prioritizes full-pipeline speech technology for real multi-speaker conversation scenarios, and related demand will continue to grow.

2. Core client pain points: Overseas brands targeting emerging markets need full-pipeline voice solutions that can handle unannotated long-form audio, speaker diarization, and multi-lingual transcription. Technical advantages in a single link cannot meet real-world deployment needs, and full-pipeline collaborative capability is what clients actually require.

3. Reference solution: You can adopt Transsion's cascaded system approach, which uses two separate pipelines to handle speaker diarization and high-accuracy transcription respectively before fusing outputs; meanwhile, pursuing parallel R&D for both cascaded and end-to-end architectures reserves space for future technical iteration, and allows adaptation to different business scenario requirements.

This article offers useful insights for AI technology service platforms and overseas consumer electronics recruitment platforms on platform operation and strategic layout.

1. Merchant demand insight: Currently, the majority of overseas brands on these platforms have demand for full-pipeline multi-lingual speech technology. Platforms need to proactively build a supply of relevant technology merchants to meet client demand and improve platform competitiveness.

2. Recruitment direction guidance: Platforms can prioritize onboarding enterprises like Transsion that have mature full-pipeline deployment capability for multi-lingual speech technology. This enriches the platform's technology service categories, attracts more overseas-focused brand clients, and expands platform scale.

3. Risk mitigation guidance: Deploying multi-lingual speech technology relies heavily on full-pipeline collaborative capability, and advantages in a single technical point cannot meet client deployment needs. When onboarding technology providers, platforms should prioritize evaluating vendors' full-pipeline engineering deployment capability to avoid onboarding immature technology that harms platform reputation and creates operational risks.

This article presents the latest industry and academic developments in multi-lingual speech technology, and offers high reference value for both industrial and academic researchers in the AI speech field.

1. New field trends: In the multi-lingual conversational speech model space, the core research challenge has shifted to real-world deployment problems. This international challenge took processing unsegmented, unannotated real long-form conversations as a core track, with full-pipeline algorithm collaboration as the core evaluation dimension. The track drew 110 teams globally, reflecting its high research momentum.

2. Technical route reference: Transsion's three-stage cascaded system architecture uses separate pipelines to recover speaker structure and complete multi-lingual transcription, before fusing outputs to generate final results. This approach delivered strong performance, while parallel exploration of both cascaded and end-to-end routes reserves room for future technical evolution. This R&D approach is valuable for researchers to reference.

3. Industrialization direction reference: Multi-lingual speech technology has already achieved large-scale deployment in smart device scenarios targeting emerging markets, where it solves real user pain points and delivers clear commercial value. This clarifies the industrialization direction for related technologies.

Disclaimer: The "Quick Summary" content is entirely generated by AI. Please exercise discretion when interpreting the information. For issues or corrections, please email run@ebrun.com .

I am a Brand Seller Factory Service Provider Marketplace Seller Researcher Read it again.

近日,在国际语音领域顶级会议INTERSPEECH 2026卫星活动——第二届多语种对话语音语言模型挑战赛 (Multilingual Conversational Speech Language Model Challenge,MLC-SLM)Task 1赛道中,传音TEX AI中心语音算法团队以自主研发的说话人归属多语种自动语音识别系统取得tcpMER 15.41%的成绩,在全球110支参赛队伍中位列第二。此次突破展现了传音在多语种语音识别、说话人日志、时间对齐与长音频工程化系统等关键方向的系统性技术实力。

本届赛事聚焦多语言对话语音模型面临的关键技术挑战,共吸引来自全球工业界与学术界的110支队伍报名参赛。官方训练数据约2100小时,覆盖14种语言及21个语言变体。MLC-SLM 2026 Task 1面向真实多人对话场景,要求系统直接处理未预切分、未标注说话人的长音频,同时输出“谁在何时说了什么”的带说话人转写结果。挑战赛评价指标tcpMER同时考察文本准确率、时间位置与说话人归属,任何单一环节的偏差都会传导到最终结果,是对语音算法全链路协同能力的高强度检验。

面对多语种、快速轮次切换、短暂停顿、说话人交叠以及长音频处理等挑战,传音TEX AI中心语音算法团队没有局限于单一模型优化,而是构建了由Speaker Diarization(说话人分离)、Long-form Multilingual ASR(长音频多语种自动语音识别)和Speaker-Transcription Fusion(说话人与转写融合)组成的级联系统:一条链路恢复跨片段一致的说话人结构,对不同说话人进行准确区分;另一条链路负责对多语种长音频进行高精度转写,并为识别结果标注精确的时间戳,最终融合生成可评测、可追溯的说话人归属转写结果。系统形成了从数据清洗、模型训练、推理解码、时间对齐到结果融合的闭环优化能力。

持续、可复现的实验验证是团队取得突破的重要基础。团队从基线出发,通过Local EEND任务适配、融合策略优化、说话人表征与聚类后端升级、SD条件化切分等增量式实验,将赛事评价指标tcpMER从22.78%逐步降低至15.41%,绝对下降7.37个百分点,相对下降约32.4%。此外,团队还完成了端到端方案VibeVoice-ASR的适配验证,形成了级联系统与端到端系统并行探索的技术路线,为后续算法演进储备了更多可能。

此次赛事成绩突破,验证了传音TEX AI中心语音算法团队在多语种语音理解领域的深厚积累。目前,团队已具备将说话人日志、多语种自动语音识别和精细化时间对齐能力灵活组合的工程能力,可根据不同业务场景灵活替换和升级识别模型,并保留可检查、可定位、可持续迭代的中间结果链路,为相关技术在实际业务中的快速落地和持续迭代奠定了基础。

多年来,传音持续加大在多语种语音识别、语音合成及降噪等领域的研发投入,技术能力已覆盖多种语言及本地口音、方言变体,并广泛应用于智能终端的通话、录音、翻译等多个场景,为传音在新兴市场复杂语言环境下的语音体验优势提供了坚实支撑。

未来,传音将持续深化多语种语音交互与语音理解技术研究,推动相关能力在全场景录音、智能终端及更多真实语音交互场景中落地,为全球用户提供更自然、更精准、更具理解力的智能语音体验。

注:文/龚作仁,文章来源:Laborer,本文为作者独立观点,不代表亿邦动力立场。

文章来源:Laborer

广告
微信
朋友圈

FAQ回顾

MLC-SLM 2026挑战赛Task 1的考核要求是什么?

MLC-SLM 2026 Task 1面向真实多人对话场景,要求系统直接处理未预切分、未标注说话人的长音频,同时输出“谁在何时说了什么”的带说话人转写结果,评价指标tcpMER同时考察文本准确率、时间位置与说话人归属,检验语音算法全链路协同能力。

传音在MLC-SLM 2026挑战赛中取得了什么成绩?

传音TEX AI中心语音算法团队以自主研发的说话人归属多语种自动语音识别系统取得tcpMER 15.41%的成绩,在全球110支参赛队伍中位列第二,赛事指标tcpMER较基线绝对下降7.37个百分点,相对下降约32.4%。

传音的多语种语音技术有哪些应用场景?

传音多语种语音技术已覆盖多种语言、本地口音及方言变体,目前广泛应用于智能终端的通话、录音、翻译等场景,未来还将落地全场景录音、智能终端及更多真实语音交互场景,为用户提供更优质的智能语音体验。

这么好看,分享一下?

朋友圈 分享

APP内打开

+1
+1
微信好友 朋友圈 新浪微博 QQ空间
关闭
收藏成功
发送
/140 0