广告
加载中

微信AI团队披露617B参数WeLM大模型训练方案

亿邦AI 2026-08-14 17:01
亿邦AI 2026/08/14 17:01

邦小白快读

EN
全文速览

这篇文章核心信息是腾讯微信AI团队公开了WeLM大模型家族的全新训练缩放方案,带来了大模型训练领域的新突破,核心干货如下:

1. 本次推出了名为Hidden Decoding的全新训练方法,不需要修改原有的Transformer主干架构,只需要把单个token拓展为多条内部计算流,就可以训练出更大参数的大模型,一共产出了WeLM-HD4-80B与WeLM-HD4-617B两款新模型。

2. 两款模型的参数和性能都有明确数据,80B版本激活参数为30亿,617B版本激活参数为230亿,在九项通用基准测试中的表现都优于同等匹配度的自回归基线模型。

3. 成本方面,80B版本训练成本是基线模型的5.1倍,617B版本是基线模型的4.4倍,大参数模型的成本效益更高。

本次微信AI公开大模型新训练方案,给品牌布局智能化运营、把握消费趋势提供了新的参考方向,核心干货如下:

1. 技术路线降低了大模型升级改造门槛,不需要重构原有技术架构就能实现参数和性能提升,后续品牌接入大模型能力改造自身现有系统的成本和难度都会下降,可用于优化品牌营销内容生成、智能客服交互、个性化推荐等业务。

2. 公开数据显示617B大参数版本的训练成本仅为基线模型的4.4倍,低于80B版本的5.1倍,说明大模型规模化训练的成本控制已经取得新进展,未来品牌接入大模型服务的成本有望逐步降低。

3. 新模型性能优于同类基线模型,能支撑更复杂的智能化需求,可帮助品牌更好地适应当下消费者对智能化交互、个性化服务的消费趋势,提升品牌用户体验。

本次微信AI公开大模型新训练方案,给微信生态及全渠道卖家带来了新的技术赋能机会和方向提示,核心干货如下:

1. 技术进展说明大模型落地的成本和门槛正在下降,性能不断提升,后续微信及腾讯生态的各类商家运营工具大概率会完成智能化升级,卖家可以期待更好用的AI工具,覆盖智能文案生成、智能客服用、个性化选品预测等多个运营场景。

2. 大模型技术升级会带动生态内用户交互体验升级,提前布局智能化运营的卖家可以率先享受技术红利,抓住流量和转化的新增长机会,挖掘新的消费需求。

3. 腾讯微信公开技术方案,预示着AI能力开放会成为接下来的重要方向,卖家可以持续关注相关的开放政策和扶持措施,及时接入新的AI能力提升自身运营效率,跟上行业变化节奏。

本次微信AI公布的大模型训练新方案,给工厂推进数字化、智能化升级带来了新的启示和商业机会,核心干货如下:

1. 技术路线给工厂数字化升级提供了新思路,不需要推翻重构现有系统,就能实现能力的扩容升级,工厂可以参考这种渐进式升级思路,推进自身生产、管理系统的智能化改造,降低改造的难度和成本,减少对现有生产的影响。

2. 大模型训练的成本控制取得新进展,未来面向制造业的AI产品,比如产品AI设计、生产需求预测、供应链管理工具的成本会逐步下降,工厂可以用更低的成本引入AI能力优化生产和设计,匹配市场需求。

3. 腾讯AI技术的不断成熟,未来大概率会推出面向制造业的落地服务,工厂可以提前关注相关的合作机会,借助大模型能力打通消费端需求数据,优化产品生产和设计,提升自身市场竞争力。

本次微信AI公开的Hidden Decoding大模型训练方案,给AI及相关行业服务商提供了行业新趋势和技术参考,核心干货如下:

1. 带来了新的技术方向,解决了原有大模型升级需要重构Transformer主干架构的痛点,只需要做内部拓展就能提升参数规模和性能,这个方案可以给服务商开发大模型服务提供新的技术路线参考,降低开发升级的成本。

2. 公开的成本数据显示,大参数模型的训练成本倍率反而低于小参数模型,说明规模化大模型的成本效益更高,服务商可以参考这个规律调整自身的产品路线,布局更大参数的大模型服务,提升性价比。

3. 新方案训练出的模型性能优于同类基线模型,服务商可以借鉴这个技术思路优化自身给客户提供的大模型解决方案,在控制成本的同时提升模型效果,更好地满足客户对高性能AI的需求,提升自身方案竞争力。

本次微信AI团队公开大模型新训练方案,反映了大模型时代平台技术建设的新方向,给各类平台的运营和发展带来了参考,核心干货如下:

1. 平台升级AI能力不需要全盘重构原有技术架构,采用渐进式拓展的方案就能实现大模型参数和性能的提升,大大降低了平台升级智能化服务的技术门槛和成本投入,适合各类平台参考这个思路升级自身的AI能力。

2. 该训练方案实现了性能提升和成本可控的平衡,平台可以依托这类技术升级,给平台内的入驻商家、C端用户提供更优质的AI工具和服务,提升平台整体的用户体验和竞争力。

3. 微信作为头部生态平台公开技术方案,也预示着开放AI能力会成为平台生态建设的重要方向,平台可以依托领先的AI能力吸引更多商家入驻,优化招商效果,同时需要提前做好技术合规相关的准备,规避技术落地带来的潜在风险。

本次微信AI公开的617B参数WeLM大模型缩放方案,给大模型领域的研究提供了新的产业动向和研究素材,核心干货如下:

1. 本次提出了全新的Hidden Decoding训练方法,打破了原有大模型缩放需要修改Transformer主干架构的限制,探索出了大模型扩容的新路线,给大模型领域的学术研究和产业研究都提供了新的方向。

2. 官方公开了完整的实证数据,包括模型参数、测试表现、训练成本,明确两款模型性能优于同等匹配度的自回归基线模型,80B和617B版本训练成本分别为基线的5.1倍和4.4倍,大参数模型成本效益更高,这些数据可以给后续大模型训练成本、缩放规律的研究提供扎实的实证参考。

3. 本次产业端的探索也为研究大模型商业化落地提供了新的案例,渐进式扩容的路线对于降低落地成本、推进大模型商业化应用有重要参考价值,有助于研究者探索更适合产业落地的大模型商业模式和技术路径。

返回默认

声明:快读内容全程由AI生成,请注意甄别信息。如您发现问题,请发送邮件至 run@ebrun.com 。

我是 品牌商 卖家 工厂 服务商 平台商 研究者 帮我再读一遍。

Quick Summary

This article summarizes a key breakthrough in large language model (LLM) training from Tencent WeChat AI, which publicly released a novel scaling approach for its WeLM model family:

1. The team introduced a new training method called Hidden Decoding. Instead of modifying the original Transformer backbone, it scales up models by expanding single tokens into multiple internal computation streams. Two new models were released in this way: WeLM-HD4-80B and WeLM-HD4-617B.

2. Performance and parameter details are fully disclosed: the 80B version has 3 billion active parameters, while the 617B version has 23 billion active parameters. Both outperformed size-matched autoregressive baseline models across 9 common benchmark tests.

3. In terms of training cost, the 80B model costs 5.1 times that of the baseline, while the 617B costs 4.4 times, demonstrating higher cost efficiency for larger-parameter models under this approach.

Tencent WeChat AI’s newly released open LLM training scheme offers new insights for brands looking to build out intelligent operations and adapt to evolving consumer trends:

1. This technical approach lowers the barrier to upgrading large model capabilities. It enables parameter and performance improvements without requiring a full reconstruction of existing technical architectures, which will reduce the cost and complexity for brands integrating large models into their current systems. The technology can be leveraged to optimize core brand use cases including marketing content generation, intelligent customer service, and personalized recommendation.

2. Publicly released data shows the 617B large-parameter model only costs 4.4 times as much as the baseline model, lower than the 5.1x cost of the 80B model. This marks new progress in cost control for large-scale model training, and suggests the cost of accessing large model services for brands will likely fall gradually over time.

3. The new models outperform comparable baseline models, making them capable of supporting more complex intelligent demands. This can help brands better adapt to current consumer expectations for intelligent interaction and personalized service, ultimately improving end-user experience.

WeChat AI’s open-source new large model training approach points to new opportunities for technological empowerment for sellers operating in the WeChat ecosystem and across all channels:

1. This technical advance confirms that the cost and barriers to LLM deployment are falling even as performance improves. It is very likely that various merchant operation tools across WeChat and the broader Tencent ecosystem will receive AI-powered upgrades soon. Sellers can expect more capable AI tools covering core operation scenarios including intelligent copywriting, customer service, and personalized product demand forecasting.

2. Upgrades to large model technology will drive overall improvements to user experience across the ecosystem. Sellers that proactively build out intelligent operations will be first to capture technological dividends, unlock new growth opportunities for traffic and conversion, and tap into unmet consumer demand.

3. Tencent WeChat’s public release of this technical solution signals that open AI capabilities will be a key priority going forward. Sellers should keep an eye out for related open policies and support programs, and integrate new AI capabilities in a timely manner to boost operational efficiency and keep pace with industry changes.

The new large model training scheme released by WeChat AI offers new insights and business opportunities for factories advancing digital and intelligent transformation:

1. This technical route provides a new framework for factories’ digital upgrades. It enables capability expansion without requiring factories to completely discard and rebuild their existing systems. Factories can adopt this incremental upgrade approach to advance intelligent transformation of production and management systems, cutting the difficulty and cost of transformation while minimizing disruption to ongoing operations.

2. New progress has been made in controlling LLM training costs, which means the cost of AI products built for the manufacturing sector—including AI-aided product design, production demand forecasting, and supply chain management tools—will gradually decline. This will allow factories to adopt AI capabilities to optimize production and design to match market demand at a lower cost.

3. As Tencent’s AI technology matures, the company is very likely to launch end-to-end services for the manufacturing industry. Factories can monitor for upcoming cooperation opportunities in advance, and leverage large model capabilities to connect consumer demand data with in-house operations, optimizing product production and design to improve overall market competitiveness.

WeChat AI’s newly open Hidden Decoding LLM training approach offers valuable insights into emerging industry trends and a new technical reference for AI and related service providers:

1. The method establishes a new technical direction that solves the core pain point of traditional LLM scaling, which previously required full reconstruction of the Transformer backbone. This approach delivers larger parameter scales and better performance through internal expansion only, providing service providers with a new technical route that lowers the cost of development and upgrades.

2. Public cost data shows that larger parameter models actually have a lower training cost multiple than smaller models, proving that scaled-up LLMs deliver better cost efficiency. Service providers can use this finding to adjust product strategy and prioritize building larger-parameter LLM services to improve price-performance.

3. The models trained via this new approach outperform comparable baseline models. Service providers can adopt this technical idea to optimize the LLM solutions they deliver to clients, boosting model performance while controlling costs to better meet client demand for high-performance AI and improve the competitiveness of their offerings.

WeChat AI’s public release of a new LLM training scheme reflects emerging directions for platform technology development in the age of large models, offering valuable reference for the operation and growth of all types of platforms:

1. This approach shows that platforms can upgrade their AI capabilities without a full reconstruction of existing technical architectures. A gradual expansion strategy can deliver gains in LLM parameter size and performance, greatly lowering the technical barrier and capital investment required for platforms to upgrade intelligent services, making it a suitable framework for any platform looking to boost its AI capabilities.

2. This training scheme strikes a balance between performance improvements and cost control. Platforms can leverage this type of technical upgrade to deliver higher-quality AI tools and services to on-platform merchants and end consumers, improving overall user experience and platform competitiveness.

3. As a leading ecosystem platform, WeChat’s open release of this technical solution also signals that opening up AI capabilities will become a core focus for platform ecosystem building. Platforms can leverage cutting-edge AI capabilities to attract more merchants and improve recruitment outcomes, while also needing to prepare in advance for technical compliance to mitigate potential risks from new technology deployment.

Tencent WeChat AI’s open release of its 617B-parameter WeLM scaling scheme provides new industry insights and empirical material for LLM research:

1. The work introduces a novel Hidden Decoding training method that removes the constraint that LLM scaling requires modifications to the Transformer backbone. It explores a new route for LLM capacity expansion, opening up new research directions for both academic and industrial LLM research.

2. The team has publicly released full empirical data, including model parameters, test performance, and training costs. The work confirms that both new models outperform size-matched autoregressive baseline models, with training costs of 5.1x and 4.4x the baseline for the 80B and 617B versions respectively, and demonstrates larger models deliver higher cost efficiency. This full dataset provides solid empirical reference for future research on LLM training costs and scaling laws.

3. This industry-led exploration also provides a new case study for research on LLM commercialization. The incremental expansion route offers important reference for reducing deployment costs and advancing commercial LLM adoption, helping researchers explore more commercially viable business models and technical paths for industry-ready large models.

Disclaimer: The "Quick Summary" content is entirely generated by AI. Please exercise discretion when interpreting the information. For issues or corrections, please email run@ebrun.com .

I am a Brand Seller Factory Service Provider Marketplace Seller Researcher Read it again.

2026年8月13日,腾讯微信AI团队公开WeLM模型家族的全新缩放方案。团队采用名为Hidden Decoding的训练方法,在不增大Transformer主干架构的基础上,将单个token拓展为多条内部计算流,完成WeLM-HD4-80B与WeLM-HD4-617B两款模型的训练。

两款模型中,80B版本激活参数为30亿,617B版本激活参数为230亿。在团队开展的测试中,两款模型在九项通用基准测试中的表现,均优于同等匹配度的自回归基线模型。

官方公开数据显示,80B版本模型训练成本为基线模型的5.1倍,617B版本模型训练成本为基线模型的4.4倍。

文章来源:亿邦动力

广告
微信
朋友圈

FAQ回顾

617B参数的WeLM大模型采用了什么训练方法?

617B参数WeLM大模型由腾讯微信AI团队研发,采用Hidden Decoding训练方法,在不增大Transformer主干架构的基础上,将单个token拓展为多条内部计算流,激活参数为230亿。

WeLM大模型的性能表现怎么样?

WeLM-HD4-80B与WeLM-HD4-617B两款模型在九项通用基准测试中的表现,均优于同等匹配度的自回归基线模型,其中80B版本激活参数为30亿。

WeLM大模型的训练成本比基线模型高多少?

官方公开数据显示,WeLM-HD4-80B版本模型训练成本为基线模型的5.1倍,WeLM-HD4-617B版本模型训练成本为基线模型的4.4倍。

这么好看,分享一下?

朋友圈 分享

APP内打开

+1
+1
微信好友 朋友圈 新浪微博 QQ空间
关闭
收藏成功
发送
/140 0