广告
加载中

微软发布MAI-Transcribe-2 语音识别定价低至每小时10美分

亿邦AI 2026-09-04 09:30
亿邦AI 2026/09/04 09:30

邦小白快读

EN
全文速览

本文核心信息是微软在2026年9月推出全新自研语音识别模型MAI-Transcribe-2,主打高性价比、多场景适配,普通用户和中小开发者都可以直接测试使用,核心干货如下

1. 价格优势突出,定价低至每小时音频10美分,较初代产品降价约72%,年处理10万小时音频仅需1万美元,成本远低于同类产品。

2. 功能齐全适配多场景,支持60种语言识别,覆盖混合语言场景,原本专业服务商的付费增值功能全部包含在基础定价中,还提供逐字、清洁两种输出模式,适配合规记录、字幕笔记创作等不同需求。

3. 目前已经正式上线,可直接在微软Foundry开发者模型市场及MAI Playground测试环境体验,模型精度和速度的平衡表现居行业前列,比OpenAI、谷歌等头部竞品速度快数倍到十倍不等。

本次微软推出MAI-Transcribe-2,给AI语音赛道品牌商带来了产品研发、定价竞争、品牌发展多维度的参考,核心干货如下

1. 产品研发方向明确,瞄准真实商业场景痛点,针对常见的背景噪音、低质量录音、重叠语音做优化,而非追求理想工作室环境的表现,同时贴合不同场景需求设计功能,适配多语言混合使用,研发方向贴合真实用户需求。

2. 定价竞争策略清晰,首发定价低至每小时10美分,价格已经低于多数专业转录服务商的大订单采购价,同时把竞品放在付费层的功能全部开放给基础定价,靠功能普惠+低价抢占市场份额。

3. 发展模式可参考,微软通过自研模型替代外部采购,大幅降低自身多业务线的成本,初代模型由10人小团队开发,不受内部官僚约束,GPU成本仅为同类模型的一半,降本效果突出。

对于AI语音转录相关卖家来说,本次微软新品发布带来了新的市场机会,同时也提示了需要警惕的风险,核心干货如下

1. 市场机会方面,卖家可以依托微软这款低成本高性能的模型,开发针对呼叫中心转录、会议字幕、内容创作笔记等垂直场景的服务,依托微软模型的能力降低自身开发和算力成本,快速切入市场。目前专业服务商仅在医疗法律等垂直领域有优势,中小卖家可以在通用场景的垂直细分方向挖掘机会。

2. 风险提示方面,微软的低价策略已经大幅拉低通用语音转录的价格区间,做通用转录业务的卖家会面临极强的价格竞争,利润空间会被大幅压缩,需要尽快转型。

3. 落地注意事项,本次10美分是首发优惠价格,官方未公布优惠截止时间、常规定价以及数据处理规则等核心信息,卖家合作落地前需要完成对应验证,避免后续出现成本不可控的问题。

本次微软发布新款AI语音模型,给工厂推进数字化转型、开发AI应用带来了多方面的启示,核心干货如下

1. 产品开发思路启示,微软这款模型没有追求实验室理想环境的完美表现,而是针对真实商业场景的常见痛点,比如背景噪音、低质量音频做优化,贴合实际使用需求。工厂开发数字化AI应用时,也可以参考这个思路,不用一味追求实验室级的指标,优先贴合生产一线的真实场景,比如车间语音控制、巡检记录转录等场景,优先解决实际痛点。

2. 降本思路启示,微软通过自研模型替代外部采购的OpenAI模型,大幅降低了自身业务成本,工厂推进数字化也可以参考,依托成熟的公开基础模型,自研适配自身需求的应用,减少高额的外部采购支出。

3. 创新模式启示,初代模型由10人核心小团队开发,不受内部官僚体系约束,研发成本仅为同类模型的一半,工厂做数字化创新可以采用小团队灵活试错的模式,降低创新成本,提升创新效率。

对于AI语音转录相关服务商来说,本次发布明确了行业发展趋势,也指明了差异化发展的方向,核心干货如下

1. 行业发展趋势,通用语音识别领域正在快速走向功能普惠化、价格下探,原本专业服务商的付费增值功能,比如说话人分镜、多语言识别等,现在已经被微软纳入基础定价,价格也远低于行业平均水平,通用转录赛道的利润空间被大幅压缩,行业转型迫在眉睫。

2. 客户痛点清晰,当前企业用户对语音转录的成本敏感度持续提升,同时对模型的精度、速度平衡,多场景适配能力的要求越来越高,低成本高性能的通用模型是市场的核心需求。

3. 解决方案与机会,目前微软仅在通用语音识别领域具备优势,专业服务商在医疗、法律等垂直领域的深度适配能力仍然具备差异化优势,服务商可以聚焦垂直领域做深度开发,打造专属适配能力,避开通用市场的价格战,巩固自身优势。

本次微软新品发布,给AI模型平台商的运营、生态建设带来了不少参考方向,核心干货如下

1. 用户需求方向,当前企业和开发者用户对AI模型的降本需求强烈,偏好高性价比、功能齐全的通用模型,平台在引入模型和生态建设时,需要侧重引入这类能够满足用户降本需求的模型,提升平台对用户的吸引力。

2. 运营模式参考,微软将本次发布的新模型放在自身的Foundry开发者模型市场和MAI Playground测试环境上线,既方便用户测试体验,也能够为自身平台引流,带动平台生态发展,这种产品发布+平台引流的模式值得同类平台参考。

3. 风险规避方向,本次发布中微软未明确优惠截止时间、常规价格、用户数据存储规则等核心信息,不利于企业用户落地,平台在引入模型时,需要要求提供商明确核心信息,同时提示用户落地前做好验证,规避不确定风险,维护平台信誉。

本次微软推出MAI-Transcribe-2,反映了当前全球AI大模型产业的多个新动向,值得研究者关注,核心干货如下

1. 产业竞争新动向,微软和OpenAI修订合作协议后,推出自研模型直接对标OpenAI、谷歌等头部竞品,依靠精度速度优势和低价策略抢占通用语音识别市场,说明AI大模型赛道的竞争已经从技术研发竞争延伸到市场价格、生态布局的竞争,行业整合速度加快。

2. 研发模式新变化,该系列初代模型由10人核心小团队开发,不受企业内部官僚体系约束,GPU研发成本仅为其他同类先进模型的一半,这种轻量化、扁平化的AI研发模式已经成为新的研发方向,能够有效控制研发成本,提升研发效率。

3. 商业模式新变化,越来越多科技巨头开始采用自研模型替代外部采购,既降低自身现有业务的成本,又对外开放模型获取新的营收,这种自研内用+对外输出的商业模式成为大模型产业的新主流方向,值得深入研究。

返回默认

声明:快读内容全程由AI生成,请注意甄别信息。如您发现问题,请发送邮件至 run@ebrun.com 。

我是 品牌商 卖家 工厂 服务商 平台商 研究者 帮我再读一遍。

Quick Summary

This article covers Microsoft’s newly launched in-house automatic speech recognition model MAI-Transcribe-2, set for release in September 2026. Positioned for high cost-performance and multi-scenario adaptability, it is available for direct testing by general users and small-to-medium developers. Key details are as follows:

1. It features a dramatic price cut: starting at just $0.10 per audio hour, a 72% reduction from the first-generation model. Processing 100,000 hours of audio per year costs only $10,000, far lower than competing products.

2. It offers full-featured support for diverse use cases: it recognizes 60 languages, including mixed-language input. All paid premium features previously exclusive to professional service providers are included in the base pricing. It also offers two output modes—verbatim and cleaned—to meet different needs, from compliance recording to subtitle and note creation.

3. The model is already live and available for testing on Microsoft’s Foundry developer model marketplace and MAI Playground. It strikes an industry-leading balance between accuracy and speed, outpacing top competitors including OpenAI and Google by a factor of several times to 10x.

Microsoft’s launch of MAI-Transcribe-2 offers multi-dimensional insights for brands in the AI speech industry across product development, pricing competition, and growth strategy. Key takeaways are as follows:

1. Clear product development direction: The model is optimized for common pain points in real-world business scenarios, including background noise, low-quality recordings, and overlapping speech, rather than chasing peak performance in controlled studio environments. It also incorporates features tailored to diverse use cases, including mixed-language support, aligning development efforts closely with actual user needs.

2. Clear competitive pricing strategy: Priced from just $0.10 per hour at launch, the model undercuts the bulk order rates of most professional transcription service providers. Additionally, it includes all features that competitors lock behind paid tiers in its base pricing, using accessible features and aggressive low pricing to capture market share.

3. A replicable growth model: By replacing third-party purchased models with in-house development, Microsoft drastically cuts costs across its multiple business lines. The first-generation model was built by a 10-person small team free of internal bureaucratic constraints, with GPU costs totaling just half that of comparable models, delivering significant cost reduction.

For sellers offering AI transcription-related services, Microsoft’s new model launch brings new market opportunities alongside notable risks to watch for. Key takeaways are as follows:

1. Market opportunities: Sellers can leverage this low-cost, high-performance model to build vertical services for use cases including call center transcription, meeting subtitles, and content creation note-taking. The model reduces sellers’ own development and computing costs, enabling rapid market entry. Since professional providers currently only hold advantages in niche verticals such as medical and legal transcription, small and medium sellers can tap opportunities in segmenting general-purpose use cases.

2. Risk warnings: Microsoft’s low pricing has pulled down the overall price range for general speech transcription drastically. Sellers focused on general transcription will face intense price competition that will significantly compress profit margins, making a quick strategic transition necessary.

3. Implementation notes: The $0.10 per hour rate is an introductory launch discount, and Microsoft has not yet announced key details including when the promotion will end, the regular pricing, and data processing rules. Sellers must complete due verification before partnering on implementation to avoid uncontrolled future costs.

Microsoft’s launch of this new AI speech model offers multiple insights for manufacturers advancing digital transformation and building in-house AI applications. Key takeaways are as follows:

1. Insights for product development: Rather than chasing perfect performance in ideal laboratory settings, Microsoft optimized the model for common pain points of real-world business scenarios, such as background noise and low-quality audio, to align with actual usage needs. Manufacturers can follow this example when building digital AI applications: instead of blindly pursuing lab-level benchmark metrics, they should prioritize alignment with real frontline production scenarios, such as workshop voice control and inspection note transcription, and focus on solving actual pain points first.

2. Insights for cost reduction: By replacing externally purchased OpenAI models with in-house development, Microsoft drastically cut its own operating costs. Manufacturers can adopt the same approach for digital transformation: build in-house applications tailored to their own needs on top of mature, open foundational models, to reduce high third-party procurement spending.

3. Insights for innovation models: The first-generation MAI-Transcribe model was developed by a 10-person core small team unconstrained by internal bureaucracy, with total R&D costs reaching just half that of comparable models. Manufacturers can adopt this flexible small-team trial-and-error model for digital innovation, cutting innovation costs while boosting efficiency.

For AI transcription-related service providers, this launch clarifies industry development trends and points the way to differentiated growth. Key takeaways are as follows:

1. Industry trends: The general speech recognition space is rapidly moving toward accessible features and falling prices. Premium paid features once exclusive to professional providers, such as speaker diarization and multi-language recognition, are now included by Microsoft in its base pricing at a far lower rate than the industry average. This has drastically compressed profit margins in the general transcription segment, making industry transformation urgent.

2. Clear customer pain points: Enterprise users are becoming increasingly cost-sensitive for speech transcription services, while demanding a better balance between accuracy and speed and stronger multi-scenario adaptability. Low-cost, high-performance general-purpose models have become the core market demand.

3. Solutions and opportunities: Microsoft only holds a competitive advantage in general speech recognition. Professional service providers still retain a differentiated edge in their deep vertical adaptability for sectors such as healthcare and legal services. Providers should focus on deep development in these vertical segments, building tailored industry-specific capabilities, avoiding the price war in the general market and consolidating their existing advantages.

Microsoft’s new model launch offers multiple useful references for AI model platform operators in operations and ecosystem building. Key takeaways are as follows:

1. User demand direction: Enterprises and developers currently have strong demand for AI model cost reduction, and prefer high-cost-performance, full-featured general-purpose models. When curating model catalogs and building ecosystems, platforms should prioritize models that meet users’ cost-reduction needs to boost platform attractiveness.

2. Operational model reference: Microsoft launched the new model on its own Foundry developer model marketplace and MAI Playground testing environment, making it easy for users to test the product while driving traffic to its own platform to fuel ecosystem growth. This product launch + platform traffic model is a valuable reference for peer platforms.

3. Risk mitigation: Microsoft did not clarify key details including the discount end date, regular pricing, and user data storage rules for this launch, which creates barriers for enterprise implementation. When onboarding new models, platforms should require providers to disclose all core information, and prompt users to complete verification before implementation to avoid uncertain risks and protect platform reputation.

Microsoft’s launch of MAI-Transcribe-2 reveals several new trends in the global large AI model industry that warrant researcher attention. Key observations are as follows:

1. New industry competition dynamics: After revising its partnership agreement with OpenAI, Microsoft launched this in-house model to directly compete with head-on competitors including OpenAI and Google, capturing share in the general speech recognition market through its advantages in accuracy, speed, and aggressive pricing. This shows that competition in the large AI model industry has expanded beyond pure R&D competition to include pricing and ecosystem positioning, accelerating industry consolidation.

2. New shifts in R&D models: The first-generation model of this series was developed by a 10-person core small team unconstrained by internal corporate bureaucracy, with GPU R&D costs totaling just half that of other comparable advanced models. This lightweight, flat AI R&D model has emerged as a new development direction, enabling effective R&D cost control and improved efficiency.

3. New shifts in business models: More and more large technology companies are replacing external procurement with in-house model development. This cuts costs for their existing business operations while allowing them to generate new revenue by offering models to external users. This "in-house development for internal use + external commercialization" business model has become a new mainstream direction in the large model industry, and merits further in-depth research.

Disclaimer: The "Quick Summary" content is entirely generated by AI. Please exercise discretion when interpreting the information. For issues or corrections, please email run@ebrun.com .

I am a Brand Seller Factory Service Provider Marketplace Seller Researcher Read it again.

2026年9月3日,微软AI推出MAI-Transcribe-2语音识别模型。该产品定价为每小时音频10美分,较五个月前发布的同系列初代产品36美分的价格下降约72%。年处理10万小时呼叫中心音频的企业,使用该模型的年支出可从3.6万美元降至1万美元。

MAI-Transcribe-2支持60种语言识别,覆盖数量较今年6月推出的1.5版本43种、4月初代版本25种持续提升。产品可在微软Foundry开发者模型市场及MAI Playground测试环境使用,针对商业场景常见的背景噪音、低质量录音、重叠语音优化,而非适配纯净的工作室录音环境。

该模型基础功能包含说话人分镜、词级时间戳、关键词偏置、自动语言识别,还可切换两种输出格式。逐字模式保留所有语气词、停顿内容适配合规及法律场景,清洁模式剔除冗余内容生成易读的字幕及笔记。产品还支持句中多语言切换识别,覆盖印式英语、西班牙式英语等混合语言使用场景。上述功能此前多为专业转录服务商的付费增值功能,此次全部包含在基础定价中。

公开测试数据显示,MAI-Transcribe-2在FLEURS基准测试中,60种语言平均词错率为5.2%,位列同测试第一。在独立机构Artificial Analysis的词错率排行榜中位列第二,同时占据精度延迟帕累托前沿,即无竞品可在精度超过该模型的同时速度更快,也无竞品可在速度超过该模型的同时精度更高。速度层面,该模型较OpenAI GPT-Transcribe快10倍,较ElevenLabs Scribe v2快7倍,较谷歌Gemini 3.5 Transcribe快5倍。

五个月内,微软AI已连续发布三款同系列语音识别模型,每次语言覆盖提升约40%,同步新增竞争对手放置在付费层级的功能。相关信息披露,初代同系列模型由10人核心小团队开发,不受内部官僚体系约束,GPU成本仅为其他同类最先进模型的一半。

微软此前已向OpenAI投入超130亿美元资金,双方合作关系在2025年10月完成重组,微软首次获得独立或联合第三方研发AGI的权限。2026年4月双方再次修订合作协议,微软结束对OpenAI模型的独家访问权,也不再支付对应收入分成,此后微软陆续推出多款MAI系列自研模型。

自研模型替代外部模型可直接降低微软的业务成本。目前微软已在Word、Excel等产品中使用自研MAI模型回应部分用户提示,替换此前使用的外部厂商模型。Teams会议音频、Nuance临床文档业务、Azure语音服务的相关语音工作负载,都可迁移至MAI-Transcribe-2运行,进一步减少对外采购支出。

此次发布的对标对象为OpenAI、谷歌、ElevenLabs等前沿实验室的同类产品,0.1美元每小时的定价已低于多数专业转录服务商的大订单采购价格,专业服务商仅在医疗、法律等垂直领域的深度适配能力上暂存优势。微软未提及在精度上超过阿里巴巴的同类语音识别模型,其面向有中国模型采购意向的美国企业的核心优势,为精度相当、推理速度更快、价格更低,且供应商符合合规要求。

此次发布仍有部分信息待明确。10美分每小时为首发优惠价格,官方未公布优惠截止时间及常规定价,也未提及实时流转录的相关能力,各语言单独识别精度、说话人分镜准确率、用户数据存储处理规则等内容也未同步披露,企业用户落地使用前还需完成对应验证。

目前MAI-Transcribe-2已正式在微软Foundry及MAI Playground上线。

本文首发于 亿邦动力 官方网站

文章来源:亿邦动力

广告
微信
朋友圈

FAQ回顾

MAI-Transcribe-2是什么?

是微软AI于2026年9月3日推出的语音识别模型,定价低至每小时音频10美分,支持60种语言识别,针对商业场景的背景噪音、低质量录音、重叠语音优化,可在微软Foundry开发者模型市场及MAI Playground使用。

MAI-Transcribe-2相比同类语音识别产品有哪些优势?

该模型在FLEURS基准测试中60种语言平均词错率为5.2%位列第一,推理速度比OpenAI GPT-Transcribe快10倍、比谷歌Gemini 3.5 Transcribe快5倍,定价更低,还将原专业转录服务商的付费增值功能纳入基础定价。

MAI-Transcribe-2可应用在哪些场景?

它针对商业场景优化,逐字输出模式适配合规、法律场景,清洁模式可生成易读的字幕及笔记,还覆盖印式英语等混合语言使用场景,微软Teams会议、Azure语音服务等工作负载都可迁移运行。

使用MAI-Transcribe-2前需要注意哪些待明确的信息?

目前其10美分每小时为首发优惠价,官方未公布优惠截止时间及常规定价,也未提及实时流转录能力、各语言单独识别精度、用户数据存储处理规则等内容,企业落地前需完成对应验证。

这么好看,分享一下?

朋友圈 分享

APP内打开

+1
+1
微信好友 朋友圈 新浪微博 QQ空间
关闭
收藏成功
发送
/140 0