广告
加载中

Meta推出Muse语音转写 支持20+说话人每小时定价0.18美元

亿邦AI 2026-09-04 09:20
亿邦AI 2026/09/04 09:20

邦小白快读

EN
全文速览

本文核心介绍Meta最新推出的实时语音转写产品Muse的基础信息、核心能力和优缺点,普通读者可以快速了解这款新产品的实用价值。

1. 核心能力:Muse支持流式转写,无需等录音结束就能同步出结果,支持识别20名以上说话人分角色,覆盖超70种语言,其中25种经过充分验证,支持多语言无缝切换,技术上将分角色识别集成到模型训练中,不需要单独后处理。第三方测评显示它的单词错误率仅3.1%,位列参测产品第一,分角色错误率也低于竞品。

2. 价格与不足:定价为每小时0.18美元,按实际处理时长计费,零数据留存服务和标准服务同价,目前不支持单词级时间戳、情绪检测等功能,单个租户最多8条并发流,单会话最长60分钟。

Muse的入局为语音相关领域的品牌商带来了产品研发、市场竞争和消费趋势的多维度参考。

1. 消费趋势观察:当前市场对多说话人实时转写的需求持续上升,这类技术广泛应用在会议系统、通话分析、智能助手等场景,用户既要求转写和分角色的准确率,也对服务成本有更高要求,低价高准确率的产品会更受市场欢迎。

2. 竞争参考:当前实时转写赛道已有多个玩家,不同产品的说话人识别上限、定价差异较大,Muse以低价高准确率切入,会倒逼全行业在准确率和运营成本层面升级,品牌布局相关业务需要找准差异化定位。

3. 产品研发参考:Muse采用语音识别、分角色联合训练的技术路径,省去单独后处理流程,这种架构可以供自研语音产品的品牌参考。

Muse的推出为做AI语音相关业务的卖家带来了新的机会,同时也提示了需要警惕的竞争风险。

1. 新机会:卖家不需要完全自研语音转写能力,可以直接集成Muse的公开API开发会议记录整理、直播转写、内容创作辅助等面向企业或个人用户的产品,大幅降低开发门槛和成本。Muse本身每小时0.18美元的定价很低,卖家既可以预留充足的利润空间,也可以通过价格优势快速抢占市场。

2. 风险提示:当前赛道竞争已经十分激烈,巨头入场后会进一步加剧价格和技术竞争,中小卖家尽量避开通用场景的红海竞争,选择细分场景切入更易存活。

3. 注意事项:要提前注意Muse的现有功能限制,比如并发量、会话时长、缺少部分高级功能,在给客户提供服务时提前说明,做好功能补足方案。

对于布局智能硬件、音频相关产品的工厂,以及想要推进数字化升级的工厂,本文提供了不少有价值的参考信息和商业机会。

1. 产品生产需求:当前消费端和企业端对带语音转写功能的智能录音笔、会议硬件等产品需求持续增长,多说话人分角色转写是用户核心痛点之一,工厂可以对接品牌客户需求,开发适配这类功能的终端硬件,打开新的增长空间。

2. 数字化升级启示:工厂推进内部数字化,比如生产会议记录整理、沟通内容存档都需要语音转写能力,可以通过集成这类低成本API实现,不需要投入高额成本自研,能降低工厂数字化改造的门槛和成本。

3. 商业合作机会:工厂可以和需要集成语音转写能力的品牌方合作,依托Muse的低成本高准确率能力,推出高性价比的终端产品,共同开拓市场。

对于AI语音服务商、ToB技术服务商来说,本文清晰梳理了实时语音转写赛道的当前格局,明确了行业发展方向和客户痛点,可参考价值很高。

1. 行业发展趋势:实时语音转写赛道已经进入精细化成本和技术竞争阶段,核心竞争点聚焦在转写准确率、多说话人分角色准确率、服务成本三个方面,Meta这类互联网巨头入场,会加速行业洗牌,推动全行业降本提效,小玩家如果没有差异化优势会被逐步淘汰。

2. 客户痛点:客户越来越看重多说话人识别、长音频处理能力,同时对成本敏感,传统服务商高定价还额外收取分角色功能费用的模式,已经不再适配市场需求。

3. 解决方案方向:中小服务商可以参考Muse的联合训练技术架构,减少后处理流程降低成本,也可以走差异化路线,针对需要更多说话人识别、情绪检测、声音事件检测等高级功能的细分客户提供服务,避开和Muse的直接价格竞争。

对于做AI PaaS平台、开放技术平台的平台商来说,Muse的推出改变了赛道格局,需要调整运营布局应对新变化。

1. 市场需求变化:当前平台的入驻开发者和企业客户,对低成本、高准确率的多说话人语音转写API需求非常旺盛,平台尽快引入这类产品,可以有效提升平台对客户的吸引力,拉动招商增长。

2. 运营调整方向:Muse的低定价会冲击平台现有同类产品的价格体系,平台需要调整产品组合,既可以引入Muse这类高性价比产品满足价格敏感型客户,也可以加大力度扶持差异化的竞品,满足不同层级客户的需求。

3. 风险规避要点:要明确告知客户Muse现有功能限制,比如并发量、会话时长、缺失高级功能等,避免引发客户纠纷,同时要提前布局细分场景的产品,规避巨头入场带来的集中度提升风险。

本文为AI语音识别领域的研究者提供了实时语音转写赛道最新的产业动向,也给出了新的研究方向和案例参考。

1. 产业新动向:原本相对分散的实时语音转写赛道迎来巨头入场,Meta推出的Muse在技术上采用了新架构,将说话人归属识别直接集成到自回归多模态架构中,和语音识别、端点检测联合训练,打破了原来把说话人聚类作为独立下游流程的常规路径,技术路径的创新值得深入研究。

2. 新研究问题:当前Muse这类产品在说话人识别上限、附加功能开发上还有明显不足,如何在保持低错误率、低成本的前提下,支持更多说话人识别,补充单词级时间戳、情绪检测等功能,是接下来值得研究的方向。

3. 商业模式研究参考:目前行业主流采用按处理时长计费的模式,且零数据留存服务和标准服务同价,反映了市场对数据隐私的重视,为研究者研究用户需求偏好和产业商业模式提供了新的典型案例。

返回默认

声明:快读内容全程由AI生成,请注意甄别信息。如您发现问题,请发送邮件至 run@ebrun.com 。

我是 品牌商 卖家 工厂 服务商 平台商 研究者 帮我再读一遍。

Quick Summary

This article introduces the core basics, key capabilities, pros and cons of Meta's newly launched real-time speech transcription product, Muse. General readers can quickly understand the practical value of this new product.

1. Core capabilities: Muse supports streaming transcription, which delivers results synchronously without waiting for the entire recording to finish. It can identify and separate speakers for more than 20 participants, covers over 70 languages (25 of which are fully validated), and enables seamless switching between multiple languages. Technically, it integrates speaker diarization directly into model training, eliminating the need for separate post-processing. Third-party testing shows it achieves a word error rate of just 3.1%, ranking first among all tested products, with a lower diarization error rate than competitors.

2. Pricing and limitations: It is priced at $0.18 per hour, billed based on actual processing time, with the same rate for both zero-data-retention and standard services. Currently, it does not support features such as word-level timestamps or emotion detection, and is capped at 8 concurrent streams per tenant with a maximum 60-minute session length.

Muse's market entry provides multi-dimensional insights for brands in speech-related sectors across product R&D, market competition and consumer trends.

1. Consumer trend observation: Demand for real-time multi-speaker transcription is growing steadily. This technology is widely used in meeting systems, call analytics, intelligent assistants and other scenarios. Users demand high accuracy for both transcription and speaker separation, while being more cost-sensitive, making low-price, high-accuracy products particularly competitive in the market.

2. Competitive insights: The real-time transcription space currently has multiple players, with wide variations in speaker recognition limits and pricing. By entering the market with low pricing and high accuracy, Muse will force the entire industry to improve accuracy and cut operational costs, requiring brands entering this space to establish clear differentiated positioning.

3. Product R&D reference: Muse uses a joint training architecture for speech recognition and speaker diarization that eliminates separate post-processing steps, a framework that can serve as a reference for brands developing in-house speech products.

The launch of Muse opens new opportunities for sellers offering AI speech-related services, while also signaling competitive risks to watch for.

1. New opportunities: Sellers do not need to fully develop speech transcription capabilities in-house. They can directly integrate Muse's public API to build products for enterprise or consumer users such as meeting note organization, live stream transcription, and content creation assistance, drastically lowering development barriers and costs. With Muse's low price of $0.18 per hour, sellers can retain healthy profit margins while gaining price advantages to capture market share quickly.

2. Risk warning: Competition in this segment is already intense, and the entry of a tech giant will further intensify price and technology competition. Small and medium-sized sellers are advised to avoid the red sea of general use cases, as targeting niche segments offers a much higher chance of survival.

3. Key notes: Sellers should account for Muse's current functional limitations, such as its concurrency cap, session length limit, and lack of certain advanced features, disclose these limits to customers in advance, and prepare complementary solutions.

This article provides valuable insights and business opportunities for factories developing smart hardware and audio products, as well as factories pursuing digital transformation.

1. Product development demand: Consumer and enterprise demand for smart voice recorders and meeting hardware with built-in transcription capabilities is growing steadily, and multi-speaker diarized transcription is one of the core user pain points. Factories can align with brand clients' needs to develop end hardware that supports this functionality, unlocking new growth opportunities.

2. Insights for digital transformation: Factories need speech transcription for internal digitalization use cases such as meeting minute organization and communication content archiving. Integrating a low-cost API like Muse eliminates the need for costly in-house development, lowering the barrier and cost of digital transformation for factories.

3. Business partnership opportunities: Factories can partner with brands looking to integrate speech transcription capabilities. Leveraging Muse's low cost and high accuracy, the two parties can launch cost-effective end products and co-develop the market together.

For AI speech service providers and B2B technology service providers, this article clearly maps the current landscape of the real-time speech transcription segment, clarifies industry development direction and core customer pain points, offering high reference value.

1. Industry development trend: The real-time transcription sector has entered a phase of refined competition around cost and technology, with core competition focused on three dimensions: transcription accuracy, multi-speaker diarization accuracy, and service cost. The entry of a major internet player like Meta will accelerate industry consolidation and push the entire sector to cut costs and improve efficiency. Small players without clear differentiated advantages will be gradually phased out.

2. Customer pain points: Customers increasingly prioritize multi-speaker recognition and long audio processing capabilities, while remaining highly cost-sensitive. The traditional model used by established service providers, which charges high base prices plus additional fees for speaker diarization functionality, no longer aligns with market demand.

3. Solution direction: Small and medium-sized service providers can reference Muse's joint training architecture to reduce post-processing steps and cut costs. They can also pursue a differentiation strategy by serving niche customers that require advanced features such as support for more speakers, emotion detection, and sound event detection, avoiding direct price competition with Muse.

For operators of AI PaaS platforms and open technology platforms, Muse's entry has reshaped the segment's competitive landscape, requiring adjustments to operational strategy to adapt to new market conditions.

1. Shifting market demand: Registered developers and enterprise clients on these platforms now have strong demand for low-cost, high-accuracy multi-speaker transcription APIs. Adding this type of product to the platform portfolio can effectively improve the platform's attractiveness to customers and drive business growth.

2. Operational adjustment direction: Muse's low pricing will disrupt existing pricing structures for comparable products on the platform. Platforms need to adjust their product portfolios: they can add high-performance, low-cost products like Muse to serve price-sensitive customers, while also stepping up support for differentiated competing products to meet the needs of customers across different tiers.

3. Risk mitigation: Platforms must clearly disclose Muse's existing functional limitations to customers, including its concurrency cap, session length limit, and lack of advanced features, to avoid customer disputes. They should also proactively build out products for niche use cases to mitigate the risk of rising market concentration brought by the giant's entry.

This article provides researchers in AI speech recognition with the latest industry developments in the real-time speech transcription space, along with new research directions and case references.

1. New industry developments: The previously fragmented real-time transcription segment has now attracted a major tech giant entry. Meta's Muse adopts an innovative technical architecture that integrates speaker diarization directly into an autoregressive multimodal architecture, enabling joint training with speech recognition and voice activity detection. This breaks from the conventional approach that treats speaker clustering as a separate downstream process, making this technical innovation worthy of in-depth study.

2. New research directions: Existing products like Muse still have clear limitations in maximum speaker count and additional functionality development. Research on how to support more speakers and add features such as word-level timestamps and emotion detection while retaining low error rates and low costs is a promising direction for future work.

3. Reference for business model research: The industry's mainstream per-processing-hour pricing model, which offers zero-data-retention services at the same rate as standard services, reflects the market's strong emphasis on data privacy, providing a new typical case for researchers studying user demand preferences and industry business models.

Disclaimer: The "Quick Summary" content is entirely generated by AI. Please exercise discretion when interpreting the information. For issues or corrections, please email run@ebrun.com .

I am a Brand Seller Factory Service Provider Marketplace Seller Researcher Read it again.

2026年9月2日,Meta超智能实验室推出音频感知模型Muse Voice Transcribe,正式进入实时语音转写赛道。该产品公开API定价为每处理1小时音频0.18美元,集成语音流转写、端点检测与说话人分角色功能,最多可同时识别20名以上说话人。

官方公开资料显示,Muse采用流式处理架构,无需等待录音完成即可同步转写内容。模型训练覆盖超70种语言,初始版本有25种语言经过充分验证,支持超1小时长音频处理、无缝多语言切换、语言和关键词偏置,无需单独后处理流程即可完成说话人分角色。

Muse的20+说话人识别能力处于行业较高水平但并非行业最高。现有公开产品文档显示,Speechmatics实时转写服务默认支持50名说话人识别,上限可上调至100名,亚马逊Transcribe流式转写最多支持识别30名独特说话人,Soniox单会话最多支持15名说话人识别,AssemblyAI流式分角色最多支持10名,xAI的实时语音转写API支持分角色功能但未公开最高说话人数量上限。Meta公开演示中未展示20人以上同时参与的场景,主力实时演示使用8名说话人,长录音演示包含11名标注参与者,20+为公开标称的模型能力。

技术架构上,Muse将说话人归属直接集成到自回归多模态架构中,音频以80毫秒为单位分块输入,模型通过自适应延迟机制判断是否需要更多音频信息再输出文本,强化学习训练同时兼顾单词错误率和延迟指标。语音识别、分角色和端点检测联合训练,而非将说话人聚类作为独立下游流程运行。API将分角色设为与一键通、端点检测并列的一等运行模式,说话人标签以A B等形式在单会话内生效,提供轮次级而非单词级时间戳。

定价层面,Muse定价为每1000分钟3美元,即每小时0.18美元,流式与非流式转写价格一致,零数据留存处理与标准处理定价相同,按实际处理的音频时长计费,向下取整到秒。同类服务公开报价对比显示,Soniox同类服务定价更低,约每小时0.12美元,支持最多15名说话人分角色。Speechmatics定价约每小时0.24美元,默认支持50名说话人上限可调至100名。亚马逊Transcribe约每小时0.6美元,支持最多30名说话人。其余服务商定价多在0.32美元至1.02美元区间,部分服务商需单独购买分角色功能,附加费用约每小时0.12美元。

第三方独立AI测评机构Artificial Analysis 9月1日的流式语音转写评估结果显示,Muse最终转写单词错误率为3.1%,位列测评首位,优于Cartesia ElevenLabs 谷歌Gemini等同类产品。其分角色错误率在AMI-IHM AMI-SDM和VoxConverse数据集上平均为17.5%,低于参测的竞品系统。

目前Muse仍存在部分功能限制。Meta API目前仅提供轮次级时间戳,不支持单词级时间戳单词级置信度评分声音事件检测及情绪检测,默认每个租户最多同时使用8条并发流,实时会话最长持续60分钟,到期后应用需重新连接。Muse的推出为企业开发会议系统通话分析实时助手等场景提供新选择,也将迫使同行在具备说话人识别的准确率和总运营成本层面展开竞争。

本文首发于 亿邦动力 官方网站

文章来源:亿邦动力

广告
微信
朋友圈

FAQ回顾

Muse语音转写是什么?

Muse是Meta超智能实验室2026年9月推出的实时语音转写产品,定价每小时0.18美元,集成语音流转写、端点检测与说话人分角色功能,最多可同时识别20名以上说话人,覆盖超70种语言,无需等待录音完成即可同步转写内容。

主流实时语音转写产品最多支持识别多少名说话人?

目前主流实时语音转写产品中,Speechmatics默认支持50名说话人上限可调至100名,亚马逊Transcribe最多支持30名,Meta Muse支持20名以上,Soniox支持15名,AssemblyAI支持10名,xAI相关产品未公开最高说话人数量上限。

Muse语音转写的准确率怎么样?

第三方独立AI测评机构Artificial Analysis的评估结果显示,Muse转写单词错误率为3.1%,位列测评首位,优于Cartesia、ElevenLabs、谷歌Gemini等同类产品,分角色错误率在相关数据集上平均为17.5%,低于参测竞品系统。

Muse语音转写的定价在行业里是什么水平?

Muse语音转写定价为每小时0.18美元,流式与非流式、零数据留存与标准处理价格均一致。行业内仅Soniox定价更低约每小时0.12美元,Speechmatics约0.24美元,亚马逊Transcribe约0.6美元,其余服务商多在0.32至1.02美元区间。

企业选购实时语音转写服务要考虑哪些因素?

企业选购实时语音转写服务可结合会议系统、通话分析、实时助手等使用场景,重点考量可识别说话人上限、单词错误率、分角色准确率、定价、支持语言数量、并发流额度、会话时长上限、数据留存规则等核心指标。

这么好看,分享一下?

朋友圈 分享

APP内打开

+1
+1
微信好友 朋友圈 新浪微博 QQ空间
关闭
收藏成功
发送
/140 0