广告
加载中

大模型下半场怎么打?智谱悄悄花几亿“买”下了国家队

吴梅梅 2026-07-23 11:19
吴梅梅 2026/07/23 11:19

邦小白快读

EN
全文速览

本文核心讲了大模型行业进入下半场后的竞争变化,核心干货可总结为以下几点

1. 当前头部大模型公司普遍面临算力瓶颈问题,月之暗面Kimi因为用户请求量超出集群承载极限,已经紧急暂停C端新用户订阅;智谱半年前也出现过算力不足,不得不限购新用户,高流量下还出现推理混乱、用户体验变差的问题。

2. 行业内已经出现两种截然不同的竞争路线:月之暗面选择聚焦模型能力,狂卷参数大小和榜单排名;智谱则选择悄悄布局底层基础设施,花数亿元收购了中科院孵化的编译器公司中科加禾,还通过旗下基金投资了AI基建全产业链,甚至买了独栋写字楼储备长期发展,为后续自研芯片提前布局。

3. 目前行业已经形成共识,大模型竞争已经从拼算法能力转向拼底层工程能力,稳定、低成本的输出能力才是长期竞争的核心。

本文揭示了大模型品牌行业的最新竞争趋势和用户需求变化,能给大模型品牌经营带来不少参考

1. 消费趋势与用户行为变化:大模型已经从前期的能力展示阶段进入规模化服务阶段,用户不再单纯追捧参数大小和榜单排名,更看重模型输出的稳定性、推理延迟的可控性和使用成本,一旦出现算力不足体验下降,就会引发用户不满,直接损伤品牌口碑。

2. 定价与竞争参考:智谱GLM-5.2在核心性能追平甚至超越美国顶级闭源模型的前提下,把运行成本降到对方的六到七分之一,靠着“便宜又能打”实现了首周日均调用量暴涨27倍,这种高性价比的竞争路线非常值得品牌商参考。

3. 风险提示:受美国出口管制影响,国产大模型的供应链存在不确定性,提前布局国产算力适配能力,才能从根本上保障业务长期稳定运营。

本文清晰展现了当前大模型行业的机会与风险,给AI相关领域的从业者提供了不少实用参考

1. 风险提示:大模型行业当前最大的卡点是底层算力基础设施供给不足,哪怕是头部公司拿到爆发式流量,也会因为算力瓶颈被迫限制新用户,还会损伤已有用户体验,中小从业者更需要提前做好算力规划,规避算力不足的风险。

2. 新增增长机会:大模型进入规模化落地阶段后,底层AI基础设施已经成为全行业的刚性需求,从算力集群搭建、推理性能优化、跨芯片适配到编译器开发,都有明确的市场缺口,是当前值得切入的新增增长赛道。

3. 可参考的发展模式:智谱通过自有产业基金布局上下游全链条,绑定产业资源的模式值得借鉴,同时国产算力替代的大趋势下,做国产芯片适配相关的业务,拥有明确的政策和市场红利。

本文关于大模型产业发展的路径,也给制造工厂推进数字化转型、挖掘商业机会带来不少启示

1. 数字化转型的启示:当前AI落地已经从概念验证走向规模化应用,底层基础设施能力是落地成功的核心保障,工厂推进AI数字化转型的时候,不能只关注算法模型的参数大小和榜单分数,也要提前关注底层算力适配、推理稳定这些工程能力,避免出现模型看起来很强,实际用起来不稳定、成本居高不下的问题。

2. 新增商业机会:当下国产AI芯片生态分散,不同芯片架构互不兼容,模型适配成本随芯片种类指数级增长,已经成为全行业的痛点,如果工厂有相关技术配套基础,可以切入AI底层基建相关的配套领域,抓住国产替代的行业风口。

3. 发展思路参考:智谱提前布局全链条核心能力的思路,也给工厂长期发展提供了参考,面对不确定的市场变化,提前布局核心能力才能接住后续的增长需求。

本文明确了当前大模型行业的发展趋势,也点出了服务商可以切入的客户痛点和方向,干货总结如下

1. 行业发展趋势:大模型下半场的竞争已经从算法能力比拼转向底层工程能力比拼,稳定、低成本、高产出的Token系统能力是市场的核心需求,AI基础设施相关服务会迎来快速增长的红利期。

2. 核心客户痛点:当前大模型厂商的核心痛点是算力供给不稳定,国产芯片生态分散,不同厂商芯片的底层指令集、架构互不兼容,模型适配成本随芯片种类指数级增长,同时高并发场景下推理稳定性差、运营成本居高不下。

3. 解决方案方向:跨芯片统一适配层、推理性能优化、算力集群管理都是市场迫切需要的解决方案,中科加禾的虚拟指令集统一方案已经验证了效果,能实现推理时延最高降低74倍、能效比提升1.46倍,这个技术方向值得服务商参考布局。

本文揭示了大模型行业对AI平台的核心需求,也给平台运营发展指明了方向,干货总结如下

1. 商家对平台的核心需求:当前大模型厂商对平台最核心的需求是稳定、低成本、可扩展的算力基础设施服务,国产大模型普遍需要适配多类国产芯片,原来分散的生态无法满足需求,平台需要搭建统一的适配层,降低厂商的适配成本。

2. 平台新的发展机会:大模型进入规模化落地阶段后,推理优化、算力调度相关的服务需求快速增长,平台可以围绕底层基建能力打造新的增值服务模块,吸引更多大模型客户入驻,拓展自身的营收边界。

3. 风险规避方向:平台运营要提前做好算力容量规划,避免出现流量暴涨后算力不足,影响全平台用户体验的问题,同时可以参考智谱的做法,通过投资布局上下游全链条,整合产业资源,提升自身的抗风险能力和服务能力。

本文呈现了大模型产业下半场的最新竞争动向,也提出了行业发展的新问题,给产业研究提供了新鲜的研究样本,核心内容总结如下

1. 产业最新动向:大模型行业进入下半场后,已经分化出两种清晰的竞争路线,一种是聚焦模型能力,通过增大参数比拼榜单成绩的“大力出奇迹”路线,另一种是提前布局底层AI基础设施,打造从芯片到应用全工程闭环的路线,目前前者已经遭遇明确的算力瓶颈,后者更适配当前国产供应链的发展环境。

2. 行业新问题:美国出口管制导致国内大模型企业无法自由采购高端英伟达GPU,而国内国产算力生态分散、不同芯片适配成本高,已经成为制约国产大模型规模化发展的核心问题,值得深入研究破解方向。

3. 商业模式研究:智谱通过自身业务加产业基金的方式布局全产业链,整合底层基建能力的模式,是大模型行业一种新的商业模式探索,为研究国产大模型的自主发展路径提供了典型案例。

返回默认

声明:快读内容全程由AI生成,请注意甄别信息。如您发现问题,请发送邮件至 run@ebrun.com 。

我是 品牌商 卖家 工厂 服务商 平台商 研究者 帮我再读一遍。

Quick Summary

This article analyzes the competitive shifts in the large language model (LLM) industry as it enters its second phase. Its key takeaways are as follows:

1. Leading Chinese LLM developers are now widely facing computing power bottlenecks. Moonshot AI, developer of the Kimi chatbot, has urgently suspended new user subscriptions for its C-end service after user demand exceeded its cluster capacity limit. Six months ago, Zhipu AI also encountered insufficient computing power, forcing it to restrict new user access and suffer from inconsistent inference outputs and degraded user experience during periods of high traffic.

2. Two distinct competitive strategies have emerged in the industry: Moonshot AI focuses on advancing model capabilities, prioritizing larger parameter sizes and higher benchmark rankings. In contrast, Zhipu AI has quietly built out its underlying infrastructure, spending hundreds of millions of yuan to acquire中科加禾 (Zhongke Jiahe), a compiler developer incubated by the Chinese Academy of Sciences. It has also invested across the entire AI infrastructure value chain via its affiliated funds, and even purchased an entire office building to support long-term development and prepare for in-house chip development.

3. It is now an industry consensus that LLM competition has shifted from a race for algorithm performance to a competition for underlying engineering capabilities. Stable, low-cost inference output is the core of long-term competitiveness.

This article outlines the latest competitive trends and shifting user demands in the LLM industry, offering valuable insights for LLM brand operators:

1. Shifting consumer trends and user behavior: LLMs have moved from an early-stage capability demonstration phase to large-scale commercial service. Users no longer chase large parameter sizes and benchmark rankings alone; they now prioritize stable output, predictable inference latency and cost-effectiveness. Insufficient computing power that degrades user experience directly triggers customer dissatisfaction and damages brand reputation.

2. Reference for pricing and competition: Zhipu AI’s GLM-5.2 matches and even outperforms top-tier closed-source LLMs from the U.S. while cutting operating costs to 1/6 to 1/7 of the U.S. model’s level. This "high performance at low cost" strategy drove a 27-fold surge in average daily API calls in its first week post-launch, making it a valuable competitive blueprint for LLM brands.

3. Risk alert: U.S. export controls have created supply chain uncertainty for domestic Chinese LLMs. Preparing domestic computing power adaptation capabilities in advance is essential to guarantee long-term stable business operations.

This article clearly outlines the current opportunities and risks in the LLM industry, offering practical insights for AI industry practitioners:

1. Risk alert: The biggest bottleneck facing the LLM industry today is insufficient supply of underlying computing infrastructure. Even leading companies have to restrict new user access and damage existing user experience when hit by explosive traffic growth due to computing power shortages. Small and medium-sized practitioners must plan computing capacity in advance to avoid this risk.

2. New growth opportunities: As LLMs enter the large-scale deployment phase, underlying AI infrastructure has become an industry-wide rigid demand. Clear market gaps exist across the entire space, from computing cluster construction, inference performance optimization and cross-chip adaptation to compiler development, making this a promising new growth track for entry.

3. Referential development model: Zhipu AI’s approach of investing across the upstream and downstream value chain via its industry fund to lock in industrial resources is worth learning from. Against the broader trend of domestic computing substitution, businesses focused on domestic chip adaptation stand to benefit from clear policy and market dividends.

This article’s analysis of LLM industry development paths offers key insights for manufacturing facilities pursuing digital transformation and new business opportunities:

1. Insights for digital transformation: AI deployment has now moved from proof-of-concept to large-scale application, and underlying infrastructure capability is the core guarantee of successful implementation. When advancing AI-driven digital transformation, factories should not focus solely on model parameter sizes and benchmark scores. They must also prioritize engineering capabilities such as domestic computing power adaptation and inference stability, to avoid the common pitfall of models that look impressive on paper but perform unreliably and cost too much in actual use.

2. New business opportunities: China’s domestic AI chip ecosystem is currently fragmented, with incompatible architectures across different chip vendors, leading adaptation costs to grow exponentially as the number of chip types increases. This has become an industry-wide pain point. Manufacturers with relevant technical foundations can enter the supporting field of AI underlying infrastructure to capitalize on the domestic substitution trend.

3. Reference for long-term development strategy: Zhipu AI’s approach of proactively building core end-to-end capabilities offers a useful blueprint for long-term factory growth. Preemptively building core capabilities allows businesses to capture future growth opportunities amid market uncertainty.

This article clarifies the current development trend of the LLM industry and identifies key customer pain points and entry opportunities for service providers, with key takeaways as follows:

1. Industry development trend: Competition in the LLM industry’s second phase has shifted from algorithm performance to underlying engineering capabilities. The market’s core demand is for stable, low-cost, high-throughput token processing systems, and AI infrastructure-related services are entering a period of rapid growth.

2. Core customer pain points: The top challenge for LLM developers today is unstable computing power supply. The fragmented domestic chip ecosystem, with incompatible instruction sets and architectures across vendors, pushes model adaptation costs to grow exponentially with the number of chip types. Additionally, poor inference stability under high concurrency and persistently high operating costs remain major problems.

3. Solution directions: Cross-chip unified adaptation layers, inference performance optimization, and computing cluster management are all urgently needed solutions. Zhongke Jiahe’s unified virtual instruction set solution has already proven effective, delivering up to a 74x reduction in inference latency and a 1.46x improvement in energy efficiency. This technical direction is a promising area for service providers to explore.

This article identifies the core demands of the LLM industry for AI platforms and outlines development directions for platform operators, with key takeaways as follows:

1. Core merchant demands for platforms: The top demand from LLM developers for platforms is stable, low-cost, scalable computing infrastructure services. Most domestic LLMs need to adapt to multiple types of domestic chips, and the existing fragmented ecosystem cannot meet this need. Platforms need to build a unified adaptation layer to cut developers’ adaptation costs.

2. New platform growth opportunities: As LLMs enter large-scale deployment, demand for inference optimization and computing scheduling services is growing rapidly. Platforms can build new value-added service modules centered on underlying infrastructure capabilities to attract more LLM customers and expand revenue streams.

3. Risk mitigation directions: Platform operators should complete computing capacity planning in advance to avoid insufficient computing power after sudden traffic surges that degrade user experience across the entire platform. They can also follow Zhipu AI’s example to invest across the upstream and downstream value chain, integrate industrial resources, and improve both risk resilience and service capabilities.

This article presents the latest competitive dynamics in the second phase of the LLM industry, raises new questions about industry development, and provides fresh research samples for industrial research. Key findings are summarized as follows:

1. Latest industry dynamics: After entering the second phase, the LLM industry has split into two clear competitive strategies. One is the "brute force" approach focused on model capability, competing on benchmark performance via larger parameter sizes. The other is preemptively building out underlying AI infrastructure to create a full closed-loop engineering chain from chips to applications. The former has already hit clear computing bottlenecks, while the latter is better aligned with the development environment of China’s domestic supply chain.

2. New industry challenges: U.S. export controls prevent Chinese LLM firms from freely purchasing high-end NVIDIA GPUs, while the fragmented domestic computing ecosystem and high cross-chip adaptation costs have become the core bottleneck restricting the large-scale development of Chinese domestic LLMs. This problem is worthy of in-depth research to identify solutions.

3. Business model research: Zhipu AI’s model of laying out the entire industrial chain through a combination of core business and an industry fund to integrate underlying infrastructure capabilities represents a new business model exploration in the LLM industry. It provides a representative case study for research on the independent development path of Chinese domestic LLMs.

Disclaimer: The "Quick Summary" content is entirely generated by AI. Please exercise discretion when interpreting the information. For issues or corrections, please email run@ebrun.com .

I am a Brand Seller Factory Service Provider Marketplace Seller Researcher Read it again.

作者 | 吴梅梅

来源 |IT桔子

7月19日深夜,一条消息在AI圈炸开了锅。月之暗面Kimi发布紧急公告:因K3模型上线后用户请求量逼近集群承载极限,即日起暂停C端新用户订阅,全部算力优先保障已订阅用户。

一个大模型独角兽因为算力瓶颈不得不主动限制用户订阅。

同样是大模型明星公司的智谱,就在前几天被爆斥资数亿元人民币全资收购了一家AI Infra公司中科加禾(XCore Sigma)。

原来,智谱早在半年前也经历了同样的算力紧缺和限购。更早的经历让智谱更早的采取了行动——

在大模型的下半场竞争格局中,算力争夺成为牌桌上显性的战争,但藏在背后的是,Kimi、智谱两家截然不同的战略选择——有人疯狂卷模型能力和参数,有人在悄悄布局底层基础设施与生态整合。

智谱数亿买下中科院孵化项目,

打造AI芯片的“翻译官”

中科加禾是谁?

这家公司的背景堪称国内AI Infra圈最硬核的那一档——由中科院计算所孵化,创始人崔慧敏博士曾是中科院计算所编译实验室团队负责人,核心团队曾深度参与过龙芯、神威、寒武纪、华为昇腾等多款国产芯片的编译器研发,可以说是“国家队级”的技术班底。

这家公司具体是做什么的呢?通俗地说,中科加禾是给国产芯片写“翻译软件”的公司。

当前国内算力生态是一幅相当“分裂”的画面:华为昇腾、寒武纪、海光信息等国产芯片与英伟达GPU的底层指令集、硬件架构、软件栈互不兼容。模型厂商想把一个训练好的模型部署到不同芯片上,适配工作量随芯片种类指数级增长。

中科加禾的看家本领,就是用一层精妙的虚拟指令集软件,把这些互不相通的“生态烟囱”统一起来——对上提供标准接口,对下适配各家芯片,让大模型像水一样在不同芯片间丝滑流动。

据悉,中科加禾 (南京)自主研发了一款跨平台、高性能的 异构原生推理引擎SigInfer,在实验室测试中交出了“推理时延最高降低74倍、能效比提升1.46倍”的成绩。

被列入美国实体清单后,智谱无法自由采购英伟达的高端GPU,只能转向国产算力解决方案。因此,中科加禾的看家本领正是智谱所需要的,能帮其解决国产芯片适配难题。

这笔数亿元投资源头可以追溯到智谱经历了几次“算力焦虑”:

今年1月,智谱罕见宣布对GLM Coding Plan实施“限购”,每日新用户配额被砍掉八成。理由极其直白:算力跟不上了。

3月,GLM-5高调发布,日调用量冲上数亿次。但很快,用户开始抱怨模型在复杂任务中“变笨”——乱码、复读、生僻字。

事后,智谱在一篇名为《Scaling Pain》(成长的烦恼)的复盘博客中揭开了谜底:高并发下推理架构的KV Cache(临时记忆)管理出现系统性工程冲突,“读题”和“写答案”的服务器在数据协调时出现混乱。

6月,GLM-5.2发布后成为Vercel平台增长最快的模型,首周日均Token调用量暴涨27倍;这一爆发式增长主要得益于GLM-5.2在保持极低运行成本(仅为美国前沿闭源模型的六到七分之一)的同时,其核心编程与智能体基准测试性能已追平甚至超越Anthropic和 OpenAI的顶级模型,实现了“便宜又能打”的竞争优势。

但这“有喜有忧”,用户量暴涨也意味着后台基础设施承受着“近乎崩溃边缘的压力测试”。

正如智谱在那篇技术博客中的感叹:“跑不稳、跑不快、跑不便宜,一切归零。”

所以,智谱这笔收购的逻辑是清晰的:短期解决高并发推理的工程冲突,中期统一国产异构算力生态降低适配成本,长期为评估中的自研AI推理芯片储备编译器能力。

一个值得关注的细节是,今年7月7日,外媒曝出智谱正在与国内芯片设计公司接洽,评估合作开发定制AI推理芯片的可能性。如果这一计划落地,中科加禾的编译器能力将成为最关键的技术拼图。

智谱、Z基金在上半年的

AI基础设施“扫货清单”

中科加禾其实只是智谱在AI基础设施层面最显性的一步棋。如果拉出智谱和Z基金近半年的投资组合,你会发现它在用资本编织一张远比想象中更密实的网。

智谱作为第一大创始LP,几年前发起设立了星连资本(即Z基金),首关规模15亿元。IT桔子数据显示,Z基金已投资超过50家人工智能企业。

仅在2026年上半年,智谱和Z基金在AI基础设施层面的投资组合便已经初具规模:

·基流科技:AI算力集群“包工头”,已向港交所递交招股书,有望成为智谱系第一个IPO。星连资本2023年12月参与天使轮,2026年3月跟投C投融资,IPO前以7.7%持股成为第一大外部股东。

·无问芯穹:大模型软硬件协同优化,核心团队出自清华。

·趋境科技:专注推理优化,今年5月由Z基金和高瓴创投等联合领投pre-A轮融资。

更早在几年前,Z基金投资了智谱与摩尔线程联合成立的算力公司数道智算,芯片设计公司清程极智、行云集成电路等。

从算力集群(基流科技)到推理优化(无问芯穹、趋境科技),从AI平台(硅基流动)到编译器(中科加禾)——智谱正在用资本画一条从芯片到Token的完整工程闭环。

3.6亿买下钻石大厦,

大模型公司的资产逻辑也在变

在商办资产价格处于历史低位、融资成本持续下降的背景下,今年4月,智谱做了一件让商业地产圈和AI圈同时侧目的事。

智谱在港交所发布公告:拟以不超过3.61亿元,收购北京海淀区钻石大厦相关资产,用于公司总部自用。

这是一栋位于中关村软件园核心位置、总建筑面积约2.27万平方米的甲级独栋写字楼,周边聚集着联想、百度、腾讯、新浪等科技巨头,毗邻清华大学、北京大学和中国科学院。

公告显示,收购将通过购买北京红钻科技发展有限公司100%股权完成,包括约8162万元股权对价,以及承接约2.79亿元相关债务。

据天眼查信息,工商变更近日已经完成——智谱正式成为钻石大厦的新主人。

为什么要买楼?智谱在公告里说得明白:满足日常行政及大模型业务营运需求,提升运营效率及稳定性,优化资产结构,增强抗风险能力。

当一家大模型领域的上市公司从“轻资产租赁”模式转向“重资产持有”时,这向外界传递出一种公司正趋于成熟壮大、稳健发展的积极信号。

大模型下半场,

“算法信仰”或让位于“工程信仰”

当Kimi因算力告急关停订阅,智谱花数亿元买下中科加禾。这两件事指向的是大模型竞争进入下半场后,两种战略路线的公开摊牌。

月之暗面的策略高度聚焦模型能力,把所有筹码押在模型参数的军备竞赛——K3以2.8万亿参数登顶全球开源模型,Arena评测榜首,ARR半年内从2亿冲到3亿美元。

但这种“大力出奇迹”的打法,在面对指数级增长的用户请求时,被算力瓶颈狠狠堵在了基础设施的墙上。

虽然智谱最新发布的GLM-5.2总参数量为7440亿 ,不及K3的三分之一。

但智谱选择悄悄在底层工程能力上布了一张大网:买编译器公司补齐推理优化短板、用Z基金串联整个AI Infra产业链、联合清华建孵化器、和摩尔线程合资算力公司、甚至准备自研AI推理芯片。

中国工程院院士郑纬民在最近的WAIC论坛上有一句话一语中的:“智能体时代真正稀缺的不只是芯片,而是稳定、低成本、高等级地产出Token的系统能力。”

这解释了当大模型从“能力秀”走向“规模化服务”,决定用户体验的不再是模型在榜单上的分数,而是Token产出的稳定性、推理延迟的可控性、以及单位成本的可压缩空间。

而智谱这几年“买买买”的目的就是在为一场更漫长、更深层的战争储备弹药。

注:文/吴梅梅,文章来源:IT桔子(公众号ID:itjuzi521),本文为作者独立观点,不代表亿邦动力立场。

文章来源:IT桔子

广告
微信
朋友圈

FAQ回顾

智谱为什么要收购中科加禾?

中科加禾是中科院计算所孵化的AI Infra公司,核心团队参与过龙芯、寒武纪、华为昇腾等多款国产芯片编译器研发,可统一异构算力生态,帮助智谱解决国产芯片适配难题,短期缓解高并发推理冲突,中长期降本并储备自研芯片技术。

大模型下半场竞争的核心是什么?

大模型下半场竞争已从算法参数军备竞赛转向底层工程能力比拼,核心是稳定、低成本、高等级产出Token的系统能力,包括推理时延可控性、单位成本压缩空间、高并发下服务稳定性,直接决定用户实际体验。

中科加禾主要做什么业务?

中科加禾主要为国产芯片研发编译器,通过虚拟指令集软件打通不同国产芯片互不兼容的生态壁垒,其自主研发的异构原生推理引擎SigInfer可最高降低74倍推理时延,提升1.46倍能效比,大幅降低大模型跨芯片适配成本。

智谱在AI基础设施领域有哪些布局?

智谱除斥资数亿元收购中科加禾外,还通过旗下首关规模15亿元的星连资本(Z基金)投资了基流科技、无问芯穹、趋境科技等超50家AI企业,覆盖算力集群、推理优化、芯片设计等环节,构建从芯片到Token的完整工程闭环。

这么好看,分享一下?

朋友圈 分享

APP内打开

+1
+1
微信好友 朋友圈 新浪微博 QQ空间
关闭
收藏成功
发送
/140 0