广告
加载中

对话徐直军:被低估的灵衢 重新理解AI算力

张帅 2026-09-21 08:36
张帅 2026/09/21 08:36

邦小白快读

EN
全文速览

这篇文章通过华为轮值董事长徐直军、半导体首席科学家廖恒的公开解读,披露了华为此前被公众忽略的AI算力关键布局,帮大家重新理解AI算力的真实价值逻辑。

1. 先纠正大众普遍的认知偏差:AI算力的核心瓶颈早已不只是单颗AI芯片的性能,当数千、数万芯片协同工作时,芯片间的数据等待、通信开销、链路故障会大幅抵消单芯片性能,买到的峰值算力和实际能用的有效算力存在巨大差距。

2. 华为此次重点披露的灵衢互联技术是解决这一问题的核心,搭配Peerium计算架构可支撑百万级处理器像单台计算机一样协作,能减少跨设备的协议转换开销;华为选择NPO近封装光学路线做硬件支撑,相比行业其他路线成本降低40%还更便于维护。

3. 目前相关产品落地节奏明确,昇腾950已上市,25.6万卡集群正在部署,960系统进入测试阶段,明年会有大量基于昇腾原生训练的大模型集中出现,灵衢协议也将开放给全产业共建生态。

华为轮值董事长徐直军在2026年全联接大会上披露的AI算力布局思路,在技术传播、产品研发、生态建设层面可为科技品牌的经营决策提供高价值参考。

1. 品牌传播层面,华为去年对外公布芯片路线图回应供应链质疑、建立市场信心,今年转向解读系统级互联技术灵衢,逐步释放全栈技术能力,避免品牌价值被单一芯片产品绑定遮蔽的传播节奏,值得硬科技品牌借鉴。

2. 产品研发层面,当前算力客户的决策逻辑已经发生变化,头部模型厂商不再只关注单芯片峰值参数,更看重实际的模型浮点算力利用率、训练任务稳定性、综合成本,品牌研发要围绕用户真实使用价值做设计,而非单纯堆砌硬件参数。

3. 生态建设层面,华为选择开放核心的灵衢互联协议,联合全产业链共建技术标准,在技术路线选择上兼顾性能、40%的成本优化空间、维护便利性、供应链成熟度选择NPO路线,更易实现规模化落地,巩固品牌的行业引领地位。

华为轮值董事长徐直军披露的AI算力产品线落地节奏、技术优势和生态开放策略,为算力领域从业者明确了市场机会、需求变化和经营注意事项。

1. 增量市场机会清晰,面向推理场景的昇腾950PR已经上市,面向训练场景的950DT预计今年底到明年规模供货,明年大量大模型训练需求将转移到昇腾950超节点及集群上,当前相关产品产能仍紧张,提前布局相关产品销售、配套服务的卖家将率先吃到增量红利。

2. 客户需求变化明确,采购方不再单纯比价、比拼峰值算力规模,核心关注模型浮点算力利用率、任务运行稳定性、训练迭代速度、综合使用成本,基于灵衢技术的昇腾超节点已经过头部客户验证,是核心销售卖点。

3. 生态合作机会明确,华为已开放灵衢UB协议欢迎全产业链适配,卖家可提前布局相关适配产品、增值服务,同时要规避单纯炒作单芯片性能的营销误区,重点突出系统级的实际使用价值。

华为明确的AI算力技术路线和规模化部署节奏,为上下游硬件制造工厂指明了产品设计方向、订单机会和技术储备思路,相关信息来自华为轮值董事长徐直军、半导体首席科学家廖恒的公开发布。

1. 产品设计与生产需求明确,灵衢统一互联协议将覆盖柜内、柜间、数据中心全场景互联,相关交换芯片、NPU、CPU、光互联器件都需匹配统一协议标准;华为已确定选择NPO近封装光学路线而非CPO路线,相关Hi-ONE光引擎模组、板上短距铜连接、配套光连接部件将迎来大规模需求,该路线可降低约40%成本,对产业链分工、后期维护更友好,适配规模化量产要求。

2. 批量订单机会确定性强,当前昇腾950已上市,25.6万卡集群正在部署,采用NPO技术的昇腾960系统进入测试阶段,明年训练类算力产品将进入规模供货期,配套硬件生产订单将持续释放。

3. 工厂可提前围绕统一互联协议、NPO光互联相关生产工艺做技术储备,匹配华为UB技术从1.0到2.0、2.1的迭代节奏,抓住算力基础设施建设的批量订单机会。

华为轮值董事长徐直军、半导体首席科学家廖恒在2026年全联接大会上分享的AI算力行业发展判断,为服务商明确了技术趋势、客户痛点和服务布局方向。

1. 行业发展趋势上,AI算力竞争已经从单芯片性能比拼进入系统级算力效率竞争阶段,互联技术、系统架构对有效算力的影响权重持续提升,统一开放的互联协议、全光互联是未来AI计算的确定发展方向。

2. 当前客户核心痛点突出,大模型训练中多芯片协同的通信开销大、跨层级协议转换成本高,单卡故障、链路短暂异常容易造成运行许久的训练任务出现巨大损失,实际采购的峰值算力和可用的有效算力差距大,模型浮点算力利用率偏低,直接推高训练成本。

3. 可落地的服务方向清晰,服务商可依托华为Peerium计算架构、灵衢统一互联技术、NPO全光超节点产品,为客户提供高利用率、高稳定性的算力支撑服务,结合华为原厂的效率优化支持帮助客户降本提效,还可提前布局灵衢协议开放后的适配服务、运维服务,挖掘生态红利。

华为轮值董事长徐直军披露的AI算力产业发展路径与产品落地节奏,为算力服务平台、产业生态平台的资源搭建、招商运营、风险规避提供了明确指引。

1. 平台要精准匹配产业端核心需求,当前大模型企业、算力客户需要的不是单纯的峰值算力堆砌,而是高模型浮点算力利用率、高运行稳定性、低综合训练成本的系统级算力服务,平台在算力资源搭建上要优先选择经过验证的、能支撑多芯片高效协同的架构与产品,减少通信开销和故障带来的损失。

2. 平台招商与生态建设方向明确,华为正在开放灵衢UB互联协议,推动全产业链共建NPO技术、统一互联的产业标准,平台可针对性招纳适配灵衢协议的硬件厂商、软件服务商、模型开发企业入驻,构建统一协议下的协同算力生态。

3. 平台运营要做好风向规避,避免陷入单纯比拼单芯片参数、峰值算力规模的低维度竞争,不要盲目押注成本、可维护性、产业链成熟度未达规模化要求的技术路线,同时要提前储备昇腾950相关算力资源,匹配明年即将到来的大模型训练需求高峰。

华为轮值董事长徐直军、半导体首席科学家廖恒分享的AI算力架构创新、技术路线选择、生态建设实践,为AI算力产业研究提供了丰富的新观察样本与研究方向。

1. 产业新动向值得重点关注,AI算力的竞争焦点正在从单芯片性能转向系统级有效算力,芯片间互联技术成为影响算力效率的核心变量;华为提出的Peerium计算架构、灵衢统一总线协议,试图打破传统计算机内部总线与外部网络的边界,实现大范围设备的统一协议互联,支撑百万级处理器像单台计算机协同,是计算机体系结构设计的重要创新方向。

2. 产业共性问题的务实解法具备研究价值,超大规模集群下的故障容错、通信开销对冲、算力利用率提升是行业共性难题,华为通过近故障点检错重传、路径冗余、光源冗余设计降低故障影响,选择NPO路线平衡性能、40%的成本优势、产业链分工、可维护性,为行业技术路线选择提供了务实参考。

3. 商业模式与政策启示明确,华为先在自有昇腾产品迭代中验证灵衢技术,再开放核心协议推动全产业共建标准的生态模式,以及从2019年就开始布局互联技术的长周期研发逻辑,为硬科技底层技术推广、产业政策支持长期基础研发提供了鲜活样本。

返回默认

声明:快读内容全程由AI生成,请注意甄别信息。如您发现问题,请发送邮件至 run@ebrun.com 。

我是 品牌商 卖家 工厂 服务商 平台商 研究者 帮我再读一遍。

Quick Summary

Drawing on public remarks from Huawei Rotating Chairman Xu Zhijun and Chief Semiconductor Scientist Liao Heng, this article sheds light on Huawei’s previously underdiscussed core AI computing power strategy, offering readers a clearer understanding of the real value dynamics of AI compute.

1. First, it corrects a widespread public misconception: the core bottleneck for AI compute is no longer the raw performance of individual AI chips. When thousands or tens of thousands of chips work in tandem, cross-chip data latency, communication overhead, and link failures can significantly erode single-chip performance, creating a massive gap between purchased peak compute capacity and actually usable, effective compute.

2. Huawei’s newly highlighted Lingqu interconnect technology is the core solution to this problem. Paired with the Peerium computing architecture, it can enable millions of processors to operate as seamlessly as a single computer, cutting cross-device protocol conversion overhead. For hardware enablement, Huawei has adopted the near-packet optics (NPO) route, which delivers 40% lower costs than alternative industry approaches while simplifying maintenance.

3. The rollout timeline for related products is clearly defined: the Ascend 950 is already on the market, a 256,000-card cluster is being deployed, and the 960 system has entered testing. Next year will see a wave of large models natively trained on Ascend hardware, and the Lingqu protocol will be opened to the entire industry to support joint ecosystem development.

The AI computing strategy outlined by Huawei Rotating Chairman Xu Zhijun at the 2026 Huawei Connect conference offers high-value reference for tech brands’ decision-making across technology communications, product R&D, and ecosystem building.

1. For brand communications: Last year, Huawei released its chip roadmap to address supply chain skepticism and build market confidence; this year, it is shifting focus to explaining its system-level Lingqu interconnect technology, gradually revealing its full-stack technical capabilities to avoid brand value being overshadowed by association with a single chip product. This carefully paced communications approach is a valuable model for deep tech brands.

2. For product R&D: The decision-making logic of compute customers has evolved. Leading large model developers no longer focus solely on single-chip peak specifications, but instead prioritize actual model floating-point compute utilization, training task stability, and total cost of ownership. Brands should design R&D around real user value rather than simply stacking hardware specs.

3. For ecosystem building: Huawei’s decision to open its core Lingqu interconnect protocol to co-build industry standards across the supply chain, and its selection of the NPO route—which balances performance, 40% cost reduction, maintenance convenience, and supply chain maturity—will enable easier large-scale deployment and reinforce the brand’s industry leadership position.

Xu Zhijun’s disclosures on the rollout timeline, technical advantages, and ecosystem opening strategy of Huawei’s AI compute product line clarify market opportunities, shifting demand, and operational considerations for players across the compute sector.

1. Clear incremental market opportunities: The Ascend 950PR for inference scenarios is already commercially available, while the 950DT for training scenarios is expected to enter mass supply from late this year to next year. A large share of large model training demand will shift to Ascend 950 supernodes and clusters next year. As current production capacity for related products remains tight, sellers that lay out early for related product sales and supporting services will be first to capture incremental gains.

2. Clear shifts in customer demand: Buyers no longer make decisions purely based on price comparisons or peak compute scale. Their core priorities are model floating-point compute utilization, task operation stability, training iteration speed, and total cost of ownership. Ascend supernodes powered by Lingqu technology, already validated by top-tier customers, represent a core sales differentiator.

3. Clear ecosystem partnership opportunities: Huawei has opened the Lingqu UB protocol to invite cross-industry adaptation. Sellers can pre-position for related adapted products and value-added services, while avoiding the marketing pitfall of hyping single-chip performance alone, and instead highlighting system-level practical use value.

Huawei’s clearly defined AI compute technology roadmap and large-scale deployment timeline provide clear guidance on product design direction, order opportunities, and technology reserve planning for upstream and downstream hardware manufacturing factories, based on public announcements from Huawei Rotating Chairman Xu Zhijun and Chief Semiconductor Scientist Liao Heng.

1. Clear product design and production demand: The unified Lingqu interconnect protocol will cover all interconnection scenarios within cabinets, between cabinets, and across data centers, requiring related switch chips, NPUs, CPUs, and optical interconnect components to align with unified protocol standards. Huawei has confirmed it will adopt the NPO (near-packet optics) route rather than the CPO route, which will drive large-scale demand for related Hi-ONE optical engine modules, short-reach on-board copper connections, and supporting optical connection components. The NPO route reduces costs by roughly 40%, is more friendly to supply chain division of labor and post-deployment maintenance, and is well suited for mass production.

2. High certainty for bulk order opportunities: The Ascend 950 is already on the market, a 256,000-card cluster is being deployed, and the NPO-powered Ascend 960 system has entered testing. Training-focused compute products will enter mass supply next year, with supporting hardware production orders to be released on an ongoing basis.

3. Factories can build technology reserves in advance around production processes related to the unified interconnect protocol and NPO optical interconnects, aligning with the iteration rhythm of Huawei’s UB technology from version 1.0 to 2.0 and 2.1, to capture bulk order opportunities from computing infrastructure buildout.

Insights on AI compute industry development shared by Huawei Rotating Chairman Xu Zhijun and Chief Semiconductor Scientist Liao Heng at the 2026 Huawei Connect conference clarify technology trends, customer pain points, and service layout directions for service providers.

1. For industry development trends: AI compute competition has moved beyond single-chip performance rivalry to enter the stage of system-level compute efficiency competition. Interconnect technology and system architecture carry growing weight in determining effective compute output. Unified, open interconnect protocols and all-optical interconnect are confirmed future directions for AI computing.

2. For current core customer pain points: In large model training, multi-chip collaboration incurs high communication overhead, cross-layer protocol conversion costs are elevated, and single-card failures or transient link anomalies can cause massive losses for long-running training tasks. There is a wide gap between procured peak compute and available effective compute, with low model floating-point compute utilization directly driving up training costs.

3. Clear actionable service directions: Service providers can leverage Huawei’s Peerium computing architecture, unified Lingqu interconnect technology, and NPO all-optical supernode products to deliver high-utilization, high-stability compute support services for customers, helping them reduce costs and improve efficiency in coordination with Huawei’s native efficiency optimization support. They can also pre-position for adaptation and O&M services following the opening of the Lingqu protocol to capture ecosystem dividends.

The AI compute industry development path and product rollout timeline disclosed by Huawei Rotating Chairman Xu Zhijun provide clear guidance for compute service platforms and industry ecosystem platforms on resource setup, investment attraction and operation, and risk mitigation.

1. Platforms should accurately match core industry demand: What large model enterprises and compute customers need is not simply stacked peak compute capacity, but system-level compute services with high model floating-point compute utilization, high operational stability, and low comprehensive training costs. When building compute resources, platforms should prioritize validated architectures and products that support efficient multi-chip collaboration, to reduce losses from communication overhead and failures.

2. Clear direction for platform investment attraction and ecosystem building: Huawei is opening the Lingqu UB interconnect protocol to drive cross-industry co-construction of industry standards around NPO technology and unified interconnects. Platforms can target recruitment of hardware vendors, software service providers, and model developers adapted to the Lingqu protocol to build a collaborative compute ecosystem under a unified protocol.

3. Platform operators should proactively avoid market pitfalls: Steer clear of low-dimensional competition focused solely on comparing single-chip parameters and peak compute scale, and avoid making blind bets on technology routes that have not met mass deployment requirements in terms of cost, maintainability, and supply chain maturity. At the same time, reserve Ascend 950-related compute resources in advance to prepare for the coming surge in large model training demand next year.

The AI computing architecture innovations, technology route choices, and ecosystem building practices shared by Huawei Rotating Chairman Xu Zhijun and Chief Semiconductor Scientist Liao Heng provide rich new observation samples and research directions for AI compute industry research.

1. Key new industry dynamics merit close attention: The competitive focus of AI compute is shifting from single-chip performance to system-level effective compute, with cross-chip interconnect technology emerging as a core variable affecting compute efficiency. Huawei’s proposed Peerium computing architecture and Lingqu unified bus protocol aim to break down the traditional boundary between internal computer buses and external networks, enabling unified protocol interconnection for large-scale device groups and supporting collaboration across millions of processors as if they were a single computer—representing a key innovation direction for computer architecture design.

2. Practical solutions to common industry challenges hold significant research value: Fault tolerance for ultra-large-scale clusters, mitigation of communication overhead, and improvement of compute utilization are widespread industry pain points. Huawei reduces failure impact through near-fault-point error detection and retransmission, path redundancy, and light source redundancy designs, and has selected the NPO route to balance performance, a 40% cost advantage, supply chain division of labor, and maintainability, providing a pragmatic reference for industry technology route selection.

3. Clear business model and policy implications: Huawei’s ecosystem model—first validating Lingqu technology through iterations of its own Ascend products, then opening the core protocol to drive cross-industry standard co-construction—along with its long-cycle R&D logic of laying out interconnect technology as early as 2019, offers a vivid case study for commercialization of underlying deep tech and for industrial policies supporting long-term basic R&D.

Disclaimer: The "Quick Summary" content is entirely generated by AI. Please exercise discretion when interpreting the information. For issues or corrections, please email run@ebrun.com .

I am a Brand Seller Factory Service Provider Marketplace Seller Researcher Read it again.

灵衢的故事,就藏在这些连接里。

芯片终于不再是唯一的悬念。

一年前,华为轮值董事长徐直军公布昇腾芯片路线图时,人们首先想知道,华为能不能做出来AI芯片,什么时候能供货。从950到960、970,未来数年的产品排在一起,对于一家经历过供应链断裂的公司而言,这张路线图本身就是一种信心。

但这些问题太迫切,以至于同场公布的另一项技术,没有得到同样的关注。“去年我讲了互联协议,但大家都关注芯片去了,没人关注灵衢,但灵衢是关键。”徐直军说。

一年过去,AI芯片的路线图开始变成产品,昇腾950已经上市,Atlas 950超节点开始交付,基于它构建的25.6万卡集群正在部署,下一代Atlas 960系统也进入测试阶段。产能依然紧张,但随着研发与产品迭代持续推进,外界对华为的期待也在变化。

能不能做出下一颗芯片,仍然重要。只是到了今天,这个问题已经不足以解释华为在AI算力上的全部布局。

有了AI芯片,然后呢?

2026年全联接大会期间,徐直军与华为半导体首席科学家廖恒,试图把讨论推进到下一个层级:当芯片从几颗增加到数千、数万乃至更多,怎样让它们像一台计算机一样工作?

芯片的数量容易看见,芯片之间的等待却不容易看见。大模型训练需要大量处理器频繁交换数据、同步结果,一颗芯片算得再快,也要等其他芯片完成工作。一条链路短暂异常,还可能影响一个运行了很久的任务,买到了多少峰值算力,与最终能用上多少有效算力,中间隔着整个系统。

本次大会上,华为发布了Peerium计算架构,按照华为的定义,它通过嵌套并行、统一内存寻址和平等互联,可支撑百万级处理器像一台计算机一样协作。灵衢(UnifiedBus:UB)是实现这套架构的关键互联技术,全球首个采用NPO技术的超节点——昇腾960超节点正式发布。

没有人比华为更懂连接,比如通信产业的有线、无线技术,华为的发展史就是一部不断连接一切的过程,到了AI时代,连接和计算又默契地走到一起。去年,徐直军用芯片路线图回答了外界对华为未来的追问;今年,他想解释,为什么那些芯片之间的连接,同样值得被看见。

灵衢的故事,就藏在这些连接里。

AI时代,计算机也要换一种设计方式

过去一年,超节点成为AI产业的热门词汇。但在徐直军看来,同一个概念下面,产品之间存在很大区别。

“‘超节点’这个词,人人都在讲我做了超节点,但是不同厂家做的超节点是有很大的区别。”

区别首先在于,越来越多的处理器连接起来以后,能不能高效地协同工作。大模型训练需要把任务分给大量处理器,每颗芯片完成自己的计算,还要和其他芯片交换结果,再进入下一阶段。数据没有到位,计算就无法继续,一个短板会制约整个系统。

企业购买的芯片越来越多,并不意味着这些等待会自行消失,处理器增加带来的收益,绝大部分被通信和协议转换的开销抵消。AI仍然需要更快的芯片,也越来越需要能够把这些芯片组织起来的计算架构。

廖恒用两组层层嵌套的套娃解释华为的思路:软件把任务分成不同层级的并行工作,硬件也从芯片、板卡、机柜,一直到超节点和更大的集群。Peerium希望让两组结构相互配合,软件侧的组织方式被概括为Nested BSP,硬件侧也通过灵衢互联层层折叠时间,来与软件侧对应。

Peerium架构能否立住,首先要看互联。过去,总线管近处,网络管远处,各有各的分工,AI把更多处理器拉进同一个任务,也让这条分界线上的代价越来越难以忽略。

近距离的设备互联追求低时延、高带宽,远距离的网络通信需要兼顾更大的覆盖范围、规模和复杂环境,各自发展出不同的技术体系。如今,联接通信不断跨越原有边界,不同通信机制之间的衔接成本开始被放大。

行业都在寻找解决办法,英伟达通过NVLink和集群网络优化不同层级的通信,AMD与开放标准阵营也在发展加速器互联。华为的答案,是把更大范围的设备放进同一套互联规则。

“谁都想柜内、柜间、在一个数据中心之间和一个zone之间最好速度都一样,那多牛,那就是一个巨大的空间,构建一台计算机。”徐直军说。

物理距离不会因协议统一而消失,不同位置的带宽和时延仍有差别,但华为希望尽量减小跨越这些边界时的性能损失。

“这也是灵衢的价值,它能做到柜间、柜内还有数据中心之间,甚至是在一个数据中心区域里面都能够一个协议,还能高速,因此我们叫总线协议。”他说。

总线原本是人们理解一台计算机内部连接的方式,华为希望把它的能力延伸到更大的空间,让CPU、NPU、内存、存储等资源可以更直接地交互,减少部分通信必须经过CPU协调和不同协议转换的环节。

“当时在决策做这个产品的时候更多希望放在计算机部门而不是传统的网络部门,这样才能真正做出一个高速、低时延、又统一协议的UB出来。”徐直军回忆道。

华为重做互联的底气,来自此前亲自走过的所有关于连接的路。2019年之前,华为团队已经经历了多处理器互联、PCIe、RDMA以及存储和内存接口等多种技术的研发与交付。它们服务于不同产品,也让华为有机会从多个环节理解系统之间的衔接。

接下来要做的,是让这些能力围绕同一个计算目标重新组织起来,为了让总线走出机箱,华为准备了多年。2019,互联被确立为计算系统架构的战略方向。2020年7月,Unified Bus技术项目正式立项。

统一协议只是开始。发起请求的计算芯片、传递数据的交换芯片、接收访问的设备,都需要理解并执行同一套规则,原本分属不同产品的研发工作,需要在这里衔接起来。

自立项至2023年,UB经历了技术规范定义、协议版本冻结,进入NPU、CPU和交换芯片的设计与投片阶段,随后,还有套片测试和系统验证。协议要走进硅片,硅片要组成系统,系统还要跑过真实的任务,灵衢就是这样一步步做出来的。

徐直军介绍,昇腾910C超节点采用UB 1.0,昇腾950超节点采用UB 2.0,昇腾960超节点将采用UB 2.1。灵衢的演进,已经和华为计算产品的节奏联系在一起。

但通信规则统一之后,信号仍然要穿过真实的线路。近处用铜,远处用光,变的是传输介质,延续的是同一套通信规则,这个原则要在更大规模上成立,还需要光互联器件的配合。

退后一步,罕见的妥协

当速率不断提高,传统电连接开始面对更严苛的距离和信号质量约束。

为了缩短连接距离,可以把芯片放得更近,但随之增加的功率密度,又会把压力转移到供电和散热。计算设备怎样摆放,已经不只是机房规划的问题,也开始影响系统能够做到多大。

光互联提供了更大的空间,数据转成光信号之后,可以在更远的距离上传输,设备不必全部挤在一起。

华为选择的NPO,即近封装光学,把光引擎放在主芯片附近。徐直军介绍,Hi-ONE模组与昇腾芯片之间保留约五厘米的板上铜连接,其余相关互联转向光连接。“基于NPO我们是一个全光的超节点,铜就很少了。铜只在PCB板上,没有铜缆了。”

这里的全光,指的是系统互联形态,芯片内部的计算和靠近芯片的短距离连接仍然依靠电子电路

另一条受到行业关注的路线是CPO,将光引擎与主芯片共同封装,进一步缩短电连接。博通、英伟达等厂商都在推进相关产品,两种路线追求更高效的光互联,也需要分别处理封装、散热、成本和维护问题。

徐直军表示,选择NPO而非CPO时,这个问题在华为没有过争论。“因为华为自己会做光器件,又做光模块、昇腾芯片,清楚哪种情况是最佳的选择。我们很早就达成共识做NPO,没做CPO。”

他同时也表示,“NPO是工程、成本、产业链分工与维护以及各方向上一个平衡的选择,没有对错。”NPO保留更长一点的电连接,换来了接近40%的成本下降,以及光引擎相对独立维护的空间。

廖恒回忆,约二十年前,他就考虑过让芯片I/O转向光连接,原理上的吸引力一直存在,但温度变化、光学特性和可靠性,让产品实现远比电路图复杂。今天的Hi-ONE,是经过一系列务实取舍才走到产品化阶段的结果。

华为此前在激光器、硅光和高速电子电路上的积累,也在这里汇合。一项技术能不能采用,需要考虑的不仅是性能上限,还包括能否制造、长期运行后怎样维护,以及供应链能否跟上。

“所以,我们坚定不移走NPO的道路,推动产业界形成标准,全产业链共同把NPO发展好,但也不是说CPO不行。”徐直军说。光向芯片靠近,机柜之间反而可以拉开距离。疏朗地站,紧密地算,背后是华为对性能、维护和规模的一次综合取舍。

基于昇腾原生训练的大模型,或在明年集中出现

几颗芯片组成的系统,和数万乃至更多芯片组成的系统,对故障的感受完全不同。规模足够大时,某个器件失效、某条链路短暂异常,都需要被纳入正常的设计考虑。

“在做大模型训练过程中,突然一个卡掉下去了,或者是突然一个卡的性能下降了,都会对训练造成巨大的损失。”徐直军说。

廖恒解释,光链路有时会出现短暂异常,虽然没有永久损坏,却可能干扰训练。华为通过更接近故障位置的检错与重传、替代连接路径,以及光源冗余等设计,尽量减少底层问题对应用的影响,它们不如峰值算力显眼,却决定了任务能否持续运行。

徐直军反复提到MFU,也就是模型浮点算力利用率,这个指标关注的是,系统的计算能力有多少真正用在模型计算上。“基于UB互联技术,基于Peerium计算架构做出的超节点和集群,可以大幅度提升MFU,这已经被验证了。”

徐直军表示,这也是头部模型厂商看重昇腾950的原因之一,华为还可以通过原厂服务,为持续提升效率提供支持,从而降低训练成本、加快训练迭代。

这给灵衢找到了一个更容易理解的位置,客户采购算力,是为了完成训练和推理任务。连接速度、系统稳定性和软件效率,最后都需要体现为任务完成得更快、综合成本更优。

眼下,华为仍处在把能力转化为大规模供给的过程中。950PR已经上市,主要面向推理;用于训练的950DT预计今年年底或明年开始规模供货。基于Atlas 950超节点的25.6万卡集群正在部署,基于NPO的Atlas 960系统正在测试。

“我相信从明年开始大量模型的训练会构筑在基于950DT的超节点上或集群上。相信未来不是我们去推动的力度有多大,而是我们供货的能力有多大。”徐直军说。

而在华为自己的产品之外,灵衢还有另一道题要答,其他企业是否愿意采用。

先在自己的产品里验证,再让更多企业参与,技术面对的环境也随之变化。华为可以协调自己的芯片、设备和软件,其他企业接入之后,还需要解决产品适配、兼容验证和长期演进的问题。

“为什么开放UB协议?就是相信整个产业界会朝着这条路线来,这样的话,为AI走向AGI,整个产业界共同努力。”徐直军说。

华为希望通过计算产品实现商业回报,而互联只有获得更多参与者,才能形成更大的生态,今天的开放,也是在为未来的产品和市场寻找更广泛的支撑。

回看过去几年的经历,人们容易把华为的创新都放进制裁后的突围故事里,但没有凭空而来的底层创新,华为的计算架构和互联研究在制裁之前已经开始,后来的持续投入,才形成今天的体系。

“UB是一个创新之路,也是我认为未来AI计算的必由之路。”一年前,芯片路线图回答了外界对未来供给的担忧。一年后,华为希望外界看到,计算背后还有影响更为深远的连接。

注:文/张帅,文章来源:钛媒体(公众号ID:taimeiti),本文为作者独立观点,不代表亿邦动力立场。

文章来源:钛媒体

广告
微信
朋友圈

FAQ回顾

华为灵衢(UB)互联技术是什么?

灵衢(UnifiedBus,简称UB)是华为研发的统一总线互联技术,可实现柜内、柜间、数据中心区间单协议高速互联,支撑Peerium计算架构下百万级处理器协同工作,减少协议转换开销,提升AI集群的模型浮点算力利用率。

华为AI超节点为什么选择NPO近封装光学路线?

NPO是华为综合平衡工程难度、成本、产业链分工与运维需求做出的选择,相较CPO路线可降低近40%的成本,预留光引擎独立维护空间,适配大规模AI集群的全光互联需求。

影响AI大模型训练集群有效算力的核心因素有哪些?

单芯片峰值算力并非唯一决定因素,芯片间互联的协议开销、通信时延、链路稳定性、故障冗余设计、软硬件协同能力都会影响最终可用算力,模型浮点算力利用率(MFU)是核心衡量指标。

华为昇腾AI算力产品当前的交付节奏是怎样的?

目前昇腾950PR已上市,主要面向推理场景;面向训练的昇腾950DT预计2026年底或2027年规模供货;基于Atlas 950超节点的25.6万卡集群正在部署,采用NPO技术的Atlas 960系统正处于测试阶段。

这么好看,分享一下?

朋友圈 分享

APP内打开

+1
+1
微信好友 朋友圈 新浪微博 QQ空间
关闭
收藏成功
发送
/140 0