广告
加载中

马斯克带着Grok 4.5杀回AI第一梯队:性能逼近Claude 成本打到更低

林夕 2026-07-13 11:32
林夕 2026/07/13 11:32

邦小白快读

EN
全文速览

本文介绍了马斯克旗下xAI最新推出的大模型Grok 4.5的核心信息,以及当前AI大模型行业的最新变化,核心干货如下:

1. Grok 4.5核心信息:参数规模1.5万亿,基于MoE架构,支持50万上下文窗口,性能接近顶级大模型Claude Opus,核心优势是成本极低,GDPval任务成本仅0.49美元,较行业榜首降价近90%,完成同类代码任务Token消耗量仅为Claude Opus的四分之一。

2. 升级背景:xAI此前因团队调整淡出主流AI竞争,收购AI编程工具Cursor后,引入数万亿Token的真实开发交互数据训练模型,补上了代码场景短板,重点提升了编码和AI Agent能力。

3. 行业变化:大模型竞争已经从拼参数、拼测试分数,转向拼真实任务的成本效率,能覆盖大部分场景的高性价比模型更受市场欢迎。

本文揭示了当前大模型行业的竞争逻辑变化与企业端需求趋势,能给AI相关品牌商提供多方面参考,核心干货如下:

1. 消费需求趋势:企业端需求已经从“模型能不能完成任务”转向“完成任务需要多少成本”,企业不再盲目追求顶级模型性能,能够覆盖90%工作场景、成本仅为顶级模型十分之一的产品,更容易成为企业的主流选择。

2. 产品研发方向:单纯堆参数规模已经不是最优路径,积累真实场景的动态过程数据,才能建立真正的竞争壁垒,Grok 4.5依托Cursor的真实开发交互数据提升能力,证明了场景数据的价值。

3. 品牌竞争新逻辑:行业不再以Benchmark排名论高低,单任务成本、Token效率已经成为新的竞争指标,品牌可以主打真实场景性价比,建立差异化竞争优势。

本文梳理了大模型行业的最新变化,给AI工具类卖家提供了明确的趋势参考、机会提示与风险提醒,核心干货如下:

1. 市场机会:当前企业端对AI的需求已经转向低成本落地,主打低单任务成本、高Token效率的AI产品,在代码开发、金融、法律、日常办公等多个知识工作领域都有广阔市场空间。

2. 可学习的经验:可以参考xAI的路径,通过收购场景类产品补足数据短板,真实场景的过程性动态数据比公开数据更有竞争力,能帮助产品快速建立用户信任与粘性。

3. 商业模式参考:马斯克垂直整合全链路(算力、数据、应用)的模式,可通过系统优化降低整体成本,适合有资源的卖家参考复制。

4. 风险提示:当前AI代码、企业服务场景竞争已经非常激烈,头部玩家已经建立优势,没有差异化性价比优势的卖家很难突围。

本文关于大模型行业发展的分享,对工厂推进数字化转型、落地AI应用有不少启示,也带来了新的商业机会,核心干货如下:

1. 需求方向启示:工厂推进数字化、AI落地时,不需要盲目追求最顶级的AI模型能力,应该优先选择能覆盖自身大部分生产设计需求、成本更低的AI产品,降低数字化转型的整体投入。

2. 商业机会:大模型已经成熟落地到代码开发、办公生产等场景,工厂可以接入Grok 4.5这类高性价比大模型,改造自身的产品研发设计、生产流程优化、日常办公等环节,有效降低研发和运营成本。

3. 转型路径参考:工厂落地AI可以参考行业经验,先积累自身真实生产业务数据,用业务数据优化AI模型,比直接套用通用大模型效果更好,同时可以参考垂直整合降本的思路,整合生产、数据、应用各个环节,降低转型的整体成本。

本文梳理了大模型行业的最新发展动向,给AI相关服务商明确了行业趋势、客户痛点和解决方案方向,核心干货如下:

1. 行业发展趋势:AI行业竞争已经从拼参数规模、拼测试成绩的“智能竞赛”,转向拼运行效率、拼使用成本的“工程竞赛”,高性价比、低成本的AI服务会成为未来的主流需求方向。

2. 客户核心痛点:当前企业客户已经解决了“AI能不能做”的问题,核心痛点转变为AI使用综合成本过高,完成同一项任务Token消耗大、反复调用次数多、人工修正成本高,导致AI很难大规模落地。

3. 解决方案方向:服务商可以围绕“单任务成本”这个新指标打造产品,优化模型训练方式,引入真实场景的动态交互数据提升模型处理复杂任务的能力,降低Token消耗,还可以通过全链路系统优化进一步降低成本,满足企业大规模落地AI的需求。

本文披露了大模型行业的最新变化,给大模型平台商明确了市场需求方向和运营调整思路,核心干货如下:

1. 市场需求变化:平台引入模型不能只看参数规模和测试排名,企业用户现在更需要能低成本完成真实任务的高性价比模型,平台需要调整模型引入的评价标准,加入单任务成本、Token效率等新指标。

2. 平台运营可参考的新路径:平台可以打造从上游算力、中游数据到下游应用的完整闭环,控制底层核心资源,通过全链路系统优化降低整体成本,和纯生态合作模式形成差异化竞争。

3. 招商与风险规避方向:平台招商可以重点倾斜面向企业真实工作场景的模型产品,扶持主打性价比的项目,同时要重视开发者场景这个核心入口,布局开发者AI工具积累用户和数据;需要注意规避同质化竞争,当前顶级模型、通用AI工具赛道已经被头部玩家占据,新进入者很难突围。

本文梳理了大模型行业的最新产业动向,提出了行业竞争的新逻辑,为AI产业研究提供了新的方向和案例,核心干货如下:

1. 产业新动向:大模型的竞争逻辑已经发生根本性转变,过去两年行业比拼参数规模、Benchmark测试成绩,当前AI进入企业工作流后,比拼的核心转变为完成真实任务的综合成本,Token效率成为新的竞争维度,单任务成本Cost per Task成为新的行业评价指标。

2. 新商业模式案例:不同于OpenAI、Anthropic依靠模型能力加生态合作的扩张模式,马斯克的xAI走出了新模式,就是垂直整合从上游算力、中游数据到下游应用的全链路,通过系统优化降低整体成本,这种新模式的竞争力和可持续性都值得深入研究。

3. 新研究问题:xAI作为后来者,通过性价比路线重新回到AI第一梯队竞争,为研究头部格局下后来者破局提供了新案例;同时AI大规模进入企业工作流后,成本已经成为核心瓶颈,这个新问题也值得研究者深入探讨。

返回默认

声明:快读内容全程由AI生成,请注意甄别信息。如您发现问题,请发送邮件至 run@ebrun.com 。

我是 品牌商 卖家 工厂 服务商 平台商 研究者 帮我再读一遍。

Quick Summary

This article outlines core details about Grok 4.5, the latest large language model (LLM) from Elon Musk's xAI, and the latest shifts in the global LLM industry:

1. Core specifications of Grok 4.5: The 1.5-trillion-parameter Mixture-of-Experts (MoE) model supports a 500,000-token context window, with performance matching top-tier models like Claude Opus. Its key advantage is extremely low cost: the cost for GDPval benchmark tasks comes out to just $0.49, a nearly 90% drop from the current industry leader, and it consumes only one-quarter the tokens that Claude Opus uses for equivalent coding tasks.

2. Context for the upgrade: xAI had stepped back from mainstream AI competition amid team restructuring, but after acquiring AI coding tool Cursor, it gained access to trillions of tokens of real-world development interaction data to train its new model. This filled the gap in the company's coding capabilities, with the update focusing heavily on improved coding and AI agent performance.

3. Industry shift: LLM competition has shifted from a race for larger parameter counts and higher benchmark scores to a focus on cost efficiency for real-world tasks. Cost-effective models that can handle the majority of common use cases are now the most in-demand among customers.

This article outlines shifting competitive dynamics and evolving enterprise demand trends in the LLM industry, offering key takeaways for AI brand builders:

1. Shifting enterprise demand: Business customers have moved from asking "can this model complete the task" to asking "how much will this task cost". Companies no longer blindly pursue top-tier model performance; products that can handle 90% of common work scenarios at just one-tenth the cost of premium models are far more likely to become mainstream enterprise choices.

2. Product R&D priorities: Simply scaling parameter counts is no longer the optimal development path. Building competitive moats now requires accumulating dynamic, real-world scenario data, a dynamic proven by Grok 4.5, which boosted its capabilities by leveraging Cursor's real developer interaction data.

3. New rules for brand competition: Benchmark rankings are no longer the main measure of success. Per-task cost and token efficiency have emerged as the new core competitive metrics. Brands can build differentiated competitive advantages by positioning themselves around real-world cost-performance.

This article maps the latest shifts in the LLM industry, providing clear trend insights, opportunity alerts, and risk warnings for AI tool sellers:

1. Market opportunities: Enterprise demand for AI has shifted toward low-cost deployment. AI products focused on low per-task cost and high token efficiency have massive market opportunity across multiple knowledge work sectors, including software development, finance, legal services, and general office work.

2. Actionable takeaways: Sellers can follow xAI's playbook to fill data gaps by acquiring scenario-specific products. Dynamic process data from real use cases delivers far more competitive value than public data, helping products build user trust and loyalty quickly.

3. Business model reference: Musk's end-to-end vertical integration strategy, covering computing power, data, and applications, reduces overall costs through system-wide optimization, making it a replicable model for well-resourced sellers.

4. Risk warning: Competition in AI coding and enterprise AI services is already extremely intense, with incumbents holding strong advantages. Sellers without differentiated cost-performance advantages will struggle to break through.

This article's analysis of LLM industry trends offers key insights and new business opportunities for factories pursuing digital transformation and AI adoption:

1. Demand prioritization insight: Factories do not need to chase top-tier LLM performance blindly when rolling out AI for digital transformation. They should prioritize low-cost AI products that meet the majority of their production and design needs to reduce overall transformation investment.

2. New business opportunities: LLMs are now mature enough for widespread deployment in coding and office production workflows. Factories can integrate cost-effective models like Grok 4.5 to upgrade product R&D design, production process optimization, and daily administrative operations, effectively cutting R&D and operating costs.

3. Transformation pathway reference: Factories can improve AI outcomes by first accumulating real production business data to fine-tune models, rather than relying on off-the-shelf generic LLMs. They can also adopt the vertical integration cost-reduction framework to integrate production, data, and application workflows, cutting overall transformation costs.

This article breaks down the latest LLM industry developments, clarifying industry trends, customer pain points, and solution directions for AI service providers:

1. Industry trend: AI competition has shifted from an "intelligence race" focused on parameter scale and benchmark scores to an "engineering race" focused on operational efficiency and usage cost. Cost-effective, low-price AI services will become the dominant mainstream demand going forward.

2. Core customer pain point: Enterprise customers have already resolved the question of "can AI do this work". Their core pain point now is the high overall cost of AI adoption: high token consumption per task, repeated API calls, and high human correction costs all stand in the way of large-scale AI deployment.

3. Solution direction: Service providers can build products around the new core metric of "per-task cost", optimizing model training methods and integrating real-world dynamic interaction data to improve complex task performance and reduce token consumption. End-to-end system optimization can further cut costs to meet enterprises' needs for large-scale AI deployment.

This article outlines the latest shifts in the LLM industry, clarifying market demand and operational adjustment priorities for LLM platform operators:

1. Shifting market demand: Platforms should no longer select models based solely on parameter scale and benchmark rankings. Enterprise users now prioritize cost-effective models that deliver low-cost performance on real tasks. Platforms need to update their model evaluation criteria to add new metrics including per-task cost and token efficiency.

2. New operational pathway reference: Platforms can build a complete closed loop spanning upstream computing power, midstream data, and downstream applications, controlling core underlying infrastructure and reducing overall costs through system-wide optimization, creating a differentiated offering compared to pure open ecosystem partnership models.

3.招商和风险规避方向:招商可以重点倾斜面向企业真实工作场景的模型产品,扶持主打性价比的项目,同时要重视开发者场景这个核心入口,布局开发者AI工具积累用户和数据;需要注意规避同质化竞争,当前顶级模型、通用AI工具赛道已经被头部玩家占据,新进入者很难突围。

This article organizes the latest industry developments in the LLM space and introduces a new competitive logic for the sector, providing new directions and cases for AI industry research:

1. New industry dynamic: The core competitive logic of LLMs has fundamentally shifted. Over the past two years, the industry competed on parameter scale and benchmark test performance. Now that LLMs are integrated into enterprise workflows, competition centers on the total cost of completing real-world tasks. Token efficiency has emerged as a new competitive dimension, and cost per task has become a new industry evaluation metric.

2. New business model case: Unlike the expansion model used by OpenAI and Anthropic, which combines model capability with ecosystem partnerships, Elon Musk's xAI has pursued a new path: vertical integration of the full stack from upstream computing power, to midstream data, to downstream end applications, cutting overall costs through system-wide optimization. The competitiveness and long-term sustainability of this new model deserve in-depth research.

3. New research questions: As a late entrant that has reclaimed a spot in the top tier of AI competition through a cost-performance focused strategy, xAI offers a new case study for how late movers can break into an established market dominated by incumbents. Furthermore, as AI scales widely into enterprise workflows, cost has become the core bottleneck for growth — this new issue also calls for further in-depth exploration by researchers.

Disclaimer: The "Quick Summary" content is entirely generated by AI. Please exercise discretion when interpreting the information. For issues or corrections, please email run@ebrun.com .

I am a Brand Seller Factory Service Provider Marketplace Seller Researcher Read it again.

作者丨林夕

编辑丨小龙

在过去一年,大模型竞争越来越像一场没有终点的军备竞赛。

OpenAI、Anthropic、Google不断刷新模型能力边界,Claude凭借代码能力在开发者市场建立优势,GPT系列则持续向企业应用扩张。

相比之下,马斯克旗下xAI推出的Grok,逐渐淡出了主流模型竞争。

这并不是马斯克最初设想的结果。

2023年成立xAI时,马斯克曾希望打造一家能够挑战OpenAI的AI公司。但过去一年,xAI经历了团队调整、战略变化,Grok系列也没有持续保持此前的关注度。

7月9日,马斯克发布Grok 4.5。

这一次,他没有强调“全球最强模型”,也没有试图用一个更高的Benchmark成绩证明自己。

Grok 4.5给出的卖点是:接近Claude Opus级别的能力,更快的速度、更低的Token消耗,以及更低的使用成本。

Grok 4.5的GDPval任务成本仅$0.49,较榜首降价近90%

马斯克想证明的,不是Grok拥有最高的模型智商,而是在大量真实任务中,它可能是更划算的选择。

这背后反映的,是大模型竞争逻辑正在变化。

过去两年,行业关注的是模型能力上限。

谁的参数更多,谁的测试成绩更高,谁就更容易获得关注。

但随着AI开始进入企业工作流,企业面对的问题正在从“模型能不能做到”转向“完成任务需要多少成本”。

Grok 4.5试图回答的,正是这个问题。

收购Cursor后的第一张牌

Grok 4.5重新进入前沿模型竞争

对于xAI而言,Grok 4.5并不仅是一款新模型。

更重要的是,这是马斯克AI版图调整后的第一张牌。

此前,马斯克曾公开承认,xAI在编程能力方面落后于竞争对手,并推动团队进行调整。

这并不难理解。

在当前AI商业化进程中,软件开发已成为最成熟、最落地的应用场景之一。一方面,开发者乐于为切实提升编码效率的工具付费;另一方面,企业也愿意为能够有效降低研发成本的产品买单。正因如此,代码能力正逐渐演变为大模型竞争的重要入口,谁能在这一环节建立起优势,谁就更容易获得开发者的信任与黏性。而Cursor的出现,恰好弥补了xAI此前在代码场景侧的短板。

与传统模型主要依赖公开代码仓库进行训练不同,Cursor所拥有的,是大量来自真实软件开发过程的行为数据与上下文信息,这种过程性的、动态的数据积累,恰恰构成了更具竞争力的壁垒,也让xAI在通往AGI的道路上多了一块坚实的落脚点。

模型可以通过代码学习,某个函数应该如何编写。但真正复杂的软件工程,并不只是生成一段代码,而是包含开发者如何理解庞大的代码库、定位问题、调用工具、反复修改验证,以及最终与AI协作完成一个完整任务的全过程。

这些经验,并不会简单存在于公开代码中。

Cursor的数据价值,就在于它记录了真实开发过程。

据Cursor透露,Grok 4.5训练过程中引入了数万亿Token规模的Cursor数据。这些数据不仅包含代码,还包括开发者与代码库、工具以及AI Agent之间的交互记录。

换句话说,Grok 4.5学习的不只是“如何写代码”,而是在学习“工程师如何完成软件开发”。

这也是为什么Grok 4.5此次重点提升的是Coding和Agent能力。它不再只是一个代码补全工具,而是在尝试成为能够处理复杂软件任务的AI Agent。

从模型本身来看,Grok 4.5已经进入前沿模型竞争区间。

这是一款基于MoE架构的大模型,参数规模约1.5万亿,支持最高50万上下文窗口,训练过程中使用了xAI的大规模计算资源,包括数万块英伟达GB300 GPU。

但相比单纯扩大参数规模,Grok 4.5更大的变化来自训练方式

xAI和Cursor对训练数据进行了去重、质量评分和领域筛选,提高数据质量密度。

强化学习阶段,则重点围绕多步骤软件工程和技术任务展开。

这类任务与传统Benchmark不同。

模型需要理解目标、拆解问题、调用工具、执行任务,并根据结果不断调整,这更接近未来Agent真正进入工作场景的方式。因此,Grok 4.5的定位也正在发生变化:过去,Grok更多被看作一个带有社交属性的聊天机器人;而如今,它正逐渐向企业生产力工具靠近,除了软件工程场景外,还开始覆盖金融、法律等知识工作领域,并进入Excel、PowerPoint、Word等办公任务。

从这个角度看,收购Cursor并不是简单增加一个产品入口,而是帮助xAI获得了一套进入企业市场的能力。

不再争“最强模型”

马斯克押注AI时代的新指标

如果只看模型能力,Grok 4.5并不是一次彻底颠覆。

它没有宣布全面超过GPT或者Claude。

甚至马斯克自己也承认,在一些极高难度任务中,部分旗舰模型仍然具备优势。

但这一次,他关注的是另一个指标:成本。

发布Grok 4.5后,马斯克反复强调,能力、更快速度和更低成本三者结合,才是模型真正的竞争力。

这也是AI行业正在发生的一种变化。

过去,大模型公司喜欢展示Benchmark成绩。

但对于企业来说,排行榜第一并不一定意味着商业价值最高。

企业购买的不是模型分数,而是任务结果。

一个模型如果完成同样工作,需要更少Token、更少调用、更少等待时间,它创造的价值可能更高。

因此,一个新的评价指标正在受到关注:Cost per Task。

即:完成一个真实任务,需要多少成本。

Grok 4.5正是围绕这一指标展开竞争。

从API价格来看,Grok 4.5输入价格为每百万Token 2美元,输出价格为每百万Token 6美元,明显低于部分旗舰模型。

但更重要的是,它降低成本并不只是因为价格低。

而是因为完成同样任务时,需要消耗更少Token。

在Agent时代,模型成本并不只是一次调用的价格,而是完成一个任务的综合成本,包括输入和输出Token消耗、工具调用次数、失败后的重复执行,以及最终仍需要人工介入修正的成本。

一个模型如果需要不断生成长答案、反复修改,最终账单可能远高于表面价格。

Grok 4.5试图解决的,就是效率问题。

在SWE Bench Pro测试中,Grok 4.5平均输出约15954个Token,而Claude Opus 4.8 Max完成类似任务平均需要约67020个Token,前者消耗量约为后者四分之一。

这意味着,Token效率正在成为模型竞争的新维度。

Artificial Analysis的数据也显示,Grok 4.5在真实知识工作任务测试中进入前列,同时单任务成本明显低于多个排名靠前的模型。

其中,一个GDPval任务成本约0.49美元。

这正是马斯克想强调的:未来AI竞争,可能不只是比谁拥有最强的大脑。

而是比谁能让智能以更低成本被大量使用。

当模型能力差距缩小时,企业未必需要永远追求最高能力。

如果一个模型能够覆盖90%的工作场景,却只需要十分之一成本,它可能更容易成为企业默认选择。

Grok 4.5押注的,就是这个市场。

从模型到基础设施

马斯克正在重建AI版图

Grok 4.5真正值得关注的地方,并不只是模型本身。它背后体现的是马斯克一直以来的产业逻辑:通过垂直整合降低系统成本。

过去,无论是特斯拉还是SpaceX,马斯克都没有选择只做单一环节。

特斯拉同时掌握电池、电驱、软件和制造体系。

SpaceX同时掌握火箭研发、制造和发射能力。

进入AI领域后,他也在尝试建立类似体系。

Grok 4.5背后,已经形成了一条AI链路。

上游,是算力。

xAI建设了Colossus超级计算中心,并持续扩大GPU规模,为模型训练提供基础设施。

中间,是数据。

除了互联网数据,xAI拥有X平台的信息流,同时通过Cursor获得大量真实开发过程数据。

从算力、数据到应用的完整闭环,由Cursor作为开发者入口、Grok Build覆盖创作与办公场景、特斯拉与SpaceX内部工程团队提供真实使用环境,最终在下游应用中得以落地实现。

这种模式与OpenAI、Anthropic有所不同。后两者更多依靠模型能力和生态合作扩大影响力。

而马斯克希望控制更多底层资源,通过系统优化降低成本。

不过,Grok 4.5的发布,也并不意味着xAI已经赢得竞争。

事实上,这款模型背后有着强烈的压力。

过去一年,Grok系列曾因生成不当内容引发争议,xAI团队也经历调整。

与此同时,AI编程市场竞争越来越激烈。Anthropic凭借Claude系列建立优势,Cursor也面临来自其他AI编程工具的竞争。

因此,对于xAI来说,Grok 4.5不仅是一款产品升级,更是一次重新证明自己的机会。

它需要向市场证明:xAI不仅可以训练模型,也能够把模型变成真正有商业价值的产品。

目前来看,Grok 4.5释放出的信号很明确:AI竞争正在从“智能竞赛”进入“工程竞赛”。

未来的竞争,不只是看谁能训练出更大的模型。还要看谁能让模型运行得更快、成本更低,并真正进入企业工作流程。

Grok 4.5是否能够改变AI市场格局,现在还没有答案。

但至少这一次,马斯克重新把Grok带回了牌桌。

而新的竞争规则正在形成,未来几年,AI行业争夺的不一定只是“谁拥有最强的大脑”。

更重要的是,谁能让这个大脑,以最低成本服务更多真实需求。

注:文/林夕,文章来源:创业邦(公众号ID:ichuangyebang ),本文为作者独立观点,不代表亿邦动力立场。

文章来源:创业邦

广告
微信
朋友圈

FAQ回顾

Grok 4.5大模型有哪些核心优势?

Grok 4.5是xAI推出的大模型,参数规模约1.5万亿,支持最高50万上下文窗口,性能接近Claude Opus级别,运行速度更快,单GDPval任务成本仅0.49美元,较行业榜首降价近90%,Token消耗量仅为Claude Opus的四分之一。

当前大模型行业竞争出现了哪些新趋势?

当前大模型行业竞争逻辑已从比拼模型能力上限转向成本竞争,“Cost per Task(单任务完成成本)”成为新评价指标,企业更关注模型完成任务的综合成本,Token效率、适配企业工作流成为核心竞争维度。

Grok 4.5的代码能力为什么能得到提升?

Grok 4.5训练过程中引入了数万亿Token规模的Cursor真实开发数据,包含开发者与代码库、工具、AI Agent的交互记录,强化学习阶段重点围绕多步骤软件工程任务展开,可支撑复杂软件任务处理。

Grok 4.5主要面向哪些应用场景?

Grok 4.5当前定位为企业生产力工具,除核心软件工程场景外,还覆盖金融、法律等知识工作领域,可支持Excel、PowerPoint、Word等各类办公任务,适配企业日常工作流需求。

这么好看,分享一下?

朋友圈 分享

APP内打开

+1
+1
微信好友 朋友圈 新浪微博 QQ空间
关闭
收藏成功
发送
/140 0