广告
加载中

复旦女学霸 融了100亿

张楠 2026-07-23 06:35
张楠 2026/07/23 06:35

邦小白快读

EN
全文速览

本文核心介绍了复旦出身的AI创业者Lin Qiao创办的Token工厂Fireworks AI的发展情况和Token赛道的行业概况,核心干货如下:

1. 项目核心信息:Fireworks AI成立仅四年,最新完成102亿元人民币D轮融资,投后估值达到175亿美元,不到两年估值涨了30多倍,被黄仁勋称为AI工厂的台积电。核心业务不做前沿大模型,只做AI推理,帮企业微调、托管开源模型,按调用量收费。

2. 核心优势:自研Fire Attention CUDA内核栈,推理速度是同价位对手的5倍,同等质量下定价比闭源模型便宜5到10倍,目前ARR突破10亿美元,日处理token超40万亿,95%的token来自企业定制模型,服务了Uber、Notion、Quora等知名客户。

3. 行业背景:当前开源大模型质量逼近闭源,企业降本需求强烈,推理层成为AI赛道新的热门风口,中美都有大量玩家布局。

本文给品牌商带来AI落地趋势和业务布局相关干货,核心内容如下:

1. 行业消费趋势:当前AI产业需求已经从追求通用大模型转向定制化专用模型,开源模型质量与闭源模型差距不断缩小,推理成本仅为闭源的1/5-1/10,越来越多企业开始选择开源+第三方推理服务的AI落地方案。

2. AI落地路径参考:Fireworks这类Token工厂可以帮助品牌商基于自身的用户数据、业务流程、客户关系做专属模型微调,品牌商不需要自行采购GPU、搭建AI团队,就能低成本获得适配自身业务的AI能力,覆盖文本、图像、多模态等多种需求,还满足各类合规要求。

3. 布局提示:当前Agent类应用爆火带动单次AI请求的token量暴涨,用户对AI互动的需求越来越高,品牌布局AI必须提前控制推理成本,避免出现用户用得越多亏损越多的情况。

本文给AI相关卖家梳理了Token赛道的机会、增长数据和风险提示,核心干货如下:

1. 市场机会:当前AI产业的资本正在从大模型训练环节转向推理服务层,开源模型质量提升叠加企业降本需求,催生了极大的Token工厂市场需求;推理是企业持续投入的成本项,降低单次调用成本已经成为全行业的核心痛点,市场空间广阔。

2. 行业增长情况:头部玩家Fireworks成立不到两年,估值从5.52亿美元涨到175亿美元,涨幅超过30倍,ARR突破10亿美元,同比增长5倍,日处理token量从15万亿涨到40万亿以上,行业增长速度极快,中国也有玩家冲击AI Token工厂第一股。

3. 风险提示:该赛道毛利率受算力采购价、利用率、客户压价多重影响,竞争加剧后容易变成高收入低利润的转售生意,还会面临云厂商和大模型厂商的两头夹击,未来大概率像云计算一样走向头部集中,中小卖家需要提前警惕风险。

本文给工厂推进AI数字化转型带来新的机会和启示,核心干货如下:

1. 需求匹配度高:工厂都有专属的生产数据、工艺流程和设计需求,通用大模型很难适配工厂的实际生产场景,需要定制化的专用AI模型,Token工厂可以帮助工厂基于自身生产数据低成本微调部署专属模型,不需要工厂投入大量资金采购算力、挖专业AI团队,大幅降低了AI转型的门槛。

2. 成本优势明显:同等模型质量下,Token工厂的推理成本比闭源大模型低5-10倍,工厂可以以较低成本持续使用AI优化生产流程、辅助产品设计,不会因为推理成本过高限制AI的规模化使用,适合工厂长期迭代AI能力。

3. 转型路径启示:工厂做AI数字化不需要死磕自己训练大模型,可以借助第三方Token工厂的基础设施,快速把自身积累的行业经验、生产数据转化为可用的AI能力,加速AI落地,降低转型试错成本。

本文给AI相关服务商梳理了推理赛道的行业趋势、客户痛点和成熟解决方案,核心干货如下:

1. 行业发展趋势:当前AI产业已经从拼大模型参数、拼训练能力转向拼落地效率,推理层成为资本重新估值的热门赛道,中美都有大量玩家涌入,头部玩家的收入、token处理量都实现了数倍增长,整个赛道处于高速扩张期,市场需求远未被满足。

2. 核心客户痛点:企业的核心痛点是通用大模型不符合自身业务需求,闭源大模型推理成本过高,自身又缺部署AI的基础设施和专业人才,自己搭建推理架构成本高、周期长,很难快速落地。

3. 头部玩家解决方案:头部玩家Fireworks的路径是主打工程效率壁垒,自研CUDA内核优化推理速度,推出兼容OpenAI的标准化API,支持多种微调方式,靠软件技术提升算力利用率,把推理成本压到远低于闭源模型的水平,同时满足企业合规要求,覆盖多类开源模型,适配企业定制化需求。

本文给AI相关平台商梳理了Token赛道的需求变化、行业动态和风险规避方向,核心干货如下:

1. 当前客户需求变化:越来越多平台客户的AI需求已经从通用大模型推理转向定制化开源模型推理,客户希望能基于自身私有数据做模型微调,获得性价比更高的推理服务,传统通用推理服务已经不能完全满足客户的差异化需求。

2. 行业头部玩家的最新做法:头部Token平台Fireworks从20多家供应商采购GPU,和微软、英伟达深度绑定合作,靠自研软件技术提升算力吞吐和推理速度,主动优化客户结构,从早期依赖少数AI工具客户拓展到多领域企业客户,不断扩大规模提升服务可靠性。

3. 风向规避提示:该赛道存在被上下游夹击的风险,大云厂商、大模型厂商都可能直接切入推理服务,价格战会大幅挤压利润空间,平台商需要打造自身的工程效率壁垒,提升算力利用率,打造客户离不开的差异化服务能力,避免沦为单纯的算力转售商,同时需要关注毛利率健康,警惕估值泡沫风险。

本文给产业研究者提供了Token工厂这个新兴赛道的最新产业动向、商业模式和待研究问题,核心干货如下:

1. 产业最新动向:当前AI技术栈的推理层正在被资本重估,资本投入正在从大模型训练环节转向推理服务层,Token工厂成为中美AI赛道共同的热门方向;美国赛道资本更激进,估值更高,英伟达等产业巨头直接下场投资,客户以全球化企业软件公司为主;国内玩家主打异构芯片适配,满足本土芯片生态需求,也获得了资本和产业资本的支持,增长速度很快。

2. 核心商业模式:Token工厂的本质是把底层算力、开源大模型、推理引擎打包封装成标准化API,开发者只需写几行代码就能调用,按调用量按量收费,核心帮企业做定制模型的微调与托管,相当于AI产业的卖水人,定位类似AI工厂的台积电,赚推理环节的过路费。

3. 值得研究的新问题:该赛道未来是否会像云计算一样走向头部集中,独立推理层能否建立起客户离不开的垄断性效率壁垒,还是会沦为低毛利的算力转售生意,本土芯片生态下Token工厂的发展路径和美国模式有什么差异,都是值得深入研究的新产业问题。

返回默认

声明:快读内容全程由AI生成,请注意甄别信息。如您发现问题,请发送邮件至 run@ebrun.com 。

我是 品牌商 卖家 工厂 服务商 平台商 研究者 帮我再读一遍。

Quick Summary

This article introduces the development of Fireworks AI, a "token factory" founded by Fudan University-educated AI entrepreneur Lin Qiao, and provides an overview of the token sector:

1. Core project background: Founded just four years ago, Fireworks AI recently completed a Series D financing round of RMB 10.2 billion, reaching a post-investment valuation of $17.5 billion — a more than 30-fold valuation increase in less than two years. Jensen Huang has called it "the TSMC of AI factories." Instead of developing cutting-edge large foundation models, the company focuses exclusively on AI inference, helping enterprises fine-tune and host open-source models, with a pay-as-you-go pricing model based on API call volume.

2. Key competitive advantages: Its self-developed Fire Attention CUDA kernel stack delivers 5x faster inference speed than competitors at the same price point, and its pricing is 5 to 10 times lower than closed-source models for equivalent output quality. Fireworks currently has an annual recurring revenue (ARR) exceeding $1 billion, processes over 40 trillion tokens daily, 95% of which come from enterprise-customized models, and counts well-known clients including Uber, Notion and Quora.

3. Industry context: As the quality of open-source large models nears that of closed-source alternatives and enterprises face strong demand for cost reduction, the inference layer has emerged as a new hot spot in the AI sector, with a large number of players entering the market in both China and the U.S.

This article shares key insights on AI implementation trends and business positioning for brand owners:

1. Industry consumption trend: AI industry demand has shifted from pursuing general-purpose large models to customized, use-case-specific models. As the quality gap between open-source and closed-source models narrows continuously, and inference cost for open-source setups is only 1/10 to 1/5 of that for closed-source models, a growing number of enterprises are adopting the "open-source model + third-party inference service" AI implementation approach.

2. A reference path for AI deployment: Token factories like Fireworks can help brand owners fine-tune proprietary models based on their own user data, business processes and customer relationships. Brands do not need to purchase GPUs or build an in-house AI team to get low-cost AI capabilities adapted to their business, covering text, image, multi-modal and other use cases while meeting various compliance requirements.

3. Strategic reminder: The current boom in agent-based applications has driven a sharp surge in token volume per AI request, as user demand for AI interaction grows steadily. Brands must proactively control inference costs when rolling out AI capabilities to avoid the scenario where higher user engagement leads directly to larger losses.

This article outlines market opportunities, growth data and risk alerts for AI-related sellers in the token sector:

1. Market opportunity: AI industry capital is shifting from large model training to the inference service layer. Rising open-source model quality combined with enterprise cost-cutting demand has created massive market demand for token factories. Inference is an ongoing cost item for enterprises, and reducing per-call inference cost has become a core pain point across the industry, leaving enormous room for market growth.

2. Industry growth: Leading player Fireworks saw its valuation surge from $552 million to $17.5 billion in less than two years, a more than 30-fold increase. Its ARR exceeds $1 billion, representing 5x year-over-year growth, while its daily token processing volume has grown from 15 trillion to over 40 trillion, indicating extremely rapid industry expansion. Chinese players are also competing to become the first listed AI token factory.

3. Risk alert: Sector gross margins are affected by multiple factors including GPU procurement costs, utilization rates and customer pricing pressure. As competition intensifies, the sector can easily become a high-revenue, low-margin resale business, and players face pressure from both cloud providers and large model developers. The sector is likely to see concentration among top players similar to the cloud computing industry, so small and medium-sized sellers should proactively prepare for these risks.

This article shares new opportunities and insights for factories advancing AI-driven digital transformation:

1. High demand alignment: Every factory has proprietary production data, processes and design requirements, and general-purpose large models are difficult to adapt to actual factory production scenarios, requiring customized, specialized AI models. Token factories can help factories fine-tune and deploy proprietary models at low cost based on their own production data, eliminating the need for heavy investment in GPU procurement or hiring specialized AI teams, drastically lowering the barrier to AI transformation.

2. Clear cost advantages: For equivalent model quality, inference services from token factories cost 5 to 10 times less than closed-source large models. Factories can continuously use AI to optimize production processes and assist product design at low cost, without scaling being restricted by exorbitant inference costs, making this model suitable for long-term iterative AI capability building.

3. Insights for transformation paths: Factories do not need to force in-house large model training to complete AI digital transformation. They can leverage the infrastructure of third-party token factories to quickly convert accumulated industry expertise and production data into usable AI capabilities, accelerating AI implementation and cutting trial-and-error costs for transformation.

This article outlines industry trends, core customer pain points and mature solutions for AI-related service providers in the inference sector:

1. Industry development trend: The AI industry has shifted from competing on large model parameter counts and training capacity to competing on implementation efficiency. The inference layer has become a hot sector receiving revaluation from capital, with a flood of new players in both China and the U.S. Leading players have recorded multiple-fold growth in revenue and token processing volume, the entire sector is in a period of rapid expansion, and market demand is far from saturated.

2. Core customer pain points: The core pain point for enterprises is that general-purpose large models do not fit their business needs, closed-source large models carry prohibitive inference costs, and enterprises lack the in-house infrastructure and professional talent to deploy AI. Building an in-house inference architecture is costly and time-consuming, making rapid implementation very difficult.

3. Solution from leading players: Top player Fireworks builds its competitive moat around engineering efficiency. It develops proprietary CUDA kernels to optimize inference speed, offers standardized APIs compatible with OpenAI, supports multiple fine-tuning methods, and leverages software technology to improve computing utilization, pushing inference costs far below the level of closed-source models. It also meets enterprise compliance requirements, supports a wide range of open-source models, and adapts to enterprise customization needs.

This article outlines changing demand, industry dynamics and risk mitigation directions for AI-related platform operators:

1. Changing customer demand: A growing number of platform customers have shifted their AI demand from general large model inference to customized open-source model inference. Customers want to fine-tune models based on their own private data and get more cost-effective inference services, so traditional general inference services can no longer fully meet customers' differentiated needs.

2. Latest practices from leading sector players: Top token platform Fireworks procures GPUs from more than 20 suppliers and maintains deep partnerships with Microsoft and NVIDIA. It leverages proprietary software technology to increase computing throughput and inference speed, and proactively optimizes its customer base, expanding from early reliance on a small number of AI tool clients to enterprise clients across multiple sectors, continuously scaling up and improving service reliability.

3. Risk mitigation guidance: The sector faces the risk of being squeezed by upstream and downstream players: large cloud providers and large model developers can both enter the inference service market directly, and price wars will drastically compress profit margins. Platform operators need to build their own engineering efficiency moats, improve computing utilization, and develop differentiated service capabilities that customers cannot replace, to avoid becoming a pure computing reseller. They also need to maintain healthy gross margins and be alert to valuation bubble risks.

This article shares the latest industry developments, business models and open research questions for industry researchers focused on the emerging token factory sector:

1. Latest industry developments: The inference layer of the AI tech stack is currently being revalued by capital, with investment shifting from large model training to the inference service layer. Token factories have become a shared hot direction in the AI sectors of both China and the U.S. In the U.S., sector capital is more aggressive and valuations are higher, with industry giants such as Nvidia directly investing, and clients are mostly global enterprise software companies. Chinese players focus on heterogeneous chip adaptation to meet the needs of the domestic chip ecosystem, and have also won support from financial and industrial capital, with very rapid growth.

2. Core business model: Token factories essentially package underlying computing power, open-source large models and inference engines into standardized APIs. Developers can access the service with just a few lines of code, and pay only for the volume they use. Their core value is helping enterprises fine-tune and host customized models, making them the "water sellers" of the AI industry, positioned as the "TSMC of AI factories" that collects tolls for the inference环节.

3. Open research questions: Key unaddressed questions for further research include: Will the sector eventually consolidate around a small number of leading players, similar to cloud computing? Can independent inference layer players build monopolistic efficiency moats that keep customers locked in, or will they devolve into low-margin computing resellers? And what are the key differences between the development paths of token factories under the domestic chip ecosystem and the U.S. model?

Disclaimer: The "Quick Summary" content is entirely generated by AI. Please exercise discretion when interpreting the information. For issues or corrections, please email run@ebrun.com .

I am a Brand Seller Factory Service Provider Marketplace Seller Researcher Read it again.

全球都在建的Token工厂。

有位计算机系的“复旦学姐”,在美国办了一家公司,被黄仁勋叫作“AI工厂的台积电”,红杉、英伟达、Index、TCV、光速这些顶级VC轮番押注,成立才四年,估值已经达到了175亿美元(约合1185亿元人民币)。

这家公司是Fireworks AI,按现在流行的话说,Fireworks AI是一家Token工厂。既不训练前沿大模型,也不做面向消费者的AI应用,只做推理,帮企业把开源模型微调好、托管好,然后按调用量收钱。今年GTC大会上,黄仁勋和Lin Qiao有场对谈,老黄直言,“In a lot of ways,you're the TSMC of AI factories.”

这家公司刚完成了15.05亿美元(约合102亿元人民币)的D轮,投后估值175亿美元,由Atreides Management、Index Ventures和TCV三家领投,英伟达跟投,老股东Lightspeed、Bessemer、Menlo以及红杉、基准资本等一众机构继续加码。

2022年,复旦学姐Lin Qiao从Meta离开,在ChatGPT发布前一个月,于加州红木城拉起一支七人队伍创办了Fireworks AI。

2023年,红杉、英伟达、AMD投了A轮;2024年B轮估值约5.52亿美元;2025年10月C轮2.5亿美元、估值冲到40亿;再到眼下这轮D轮175亿美元估值。

不到两年时间,估值从5.52亿翻到175亿美元,涨了约30多倍。ARR突破10亿美元,是上一年同期的5倍;平台每天处理的token从15万亿涨到40万亿以上;更关键的一点是,平台上超过95%的token来自客户在自己数据上定制过的模型,而不是现成的通用模型。

随着Kimi K3开源,以及DeepSeek V4满血版即将推出,基础模型的估值和壁垒受到冲击,推理层是AI技术栈里正在被资本悄悄重估的那一层,钱正从“谁模型最强”挪向“谁能把模型稳定便宜地跑起来”。

创始人Lin Qiao也是个容易被忽略的角色,从底层框架PyTorch一路干到175亿美元,存在感却远低于许多前沿模型创始人。更妙的是这门生意全球都在做,美国有Together、Baseten,中国也有硅基流动、无问芯穹、九章云极、趋境科技等,把中美两拨人摆在一起,正好能看清“Token工厂”到底是不是一门好生意。

AI台积电

Fireworks的产品说穿了是一套推理云平台。开发者通过一套兼容OpenAI的API,就能调用平台上200多个开源模型,覆盖文本、图像、多模态。

企业客户可以在自己的专有数据上做微调,支持SFT(监督微调)、DPO(直接偏好优化)和RFT(强化微调),还能同时挂上最多100个LoRA适配器而不额外加价。合规层面它拿到了SOC2、HIPAA、GDPR这些企业入场券,2026年3月又接进了微软的Foundry,意味着Azure上的客户能直接走它的推理。

速度是Fireworks对外讲的核心故事。它自研了一套叫Fire Attention的CUDA内核栈,在DeepSeekV4 Pro上能做到每秒167个token,对外宣称是同价位对手的5倍;内部用了解耦服务、语义缓存和推测解码,把长尾延迟压住。

这套工程能力是它和“随便租几张卡跑开源模型”的本质区别,它卖的不止是算力,是把算力榨出更高吞吐的软件效率。

目前,Fireworks无服务器推理起步价大约每百万token 0.2美元(8B级别模型)到0.9美元(70B级别),另外也按小时出租GPU。一个信号性的变化是:今天平台上95%的token来自定制模型。

换句话说,想要把AI用起来,基础大模型远远满足不了企业需求,必须拥有属于自己的、贴着自己业务数据的模型。Fireworks在融资声明里讲得很直白,企业手里都有别人没有的东西,客户关系、工作流、数据、对质量的定义,Fireworks干的事就是把这些变成能拥有、能持续改进的专用智能。

Fireworks的客户包括Cursor、Notion、Quora、Sourcegraph、Uber、Shopify、Perplexity、等知名公司,还有一家叫Sentient的公司,靠Fireworks24小时扛住了180万人的候补排队。一个转折性细节是,Fireworks早先大约一半收入来自AI编程工具Cursor,现在客户结构已经散开到更广泛的企业,这与它从“尝鲜”走向“生产级渗透”的判断能对上。

Fireworks的投资人之一是Atreides,掌门人Gavin Baker早年重仓过Meta和英伟达,Gavin Baker认为,前沿模型和开源模型会越来越合用,而Fireworks是少数被企业信任、用来在生产里训练并托管开源模型的平台。

Index是连投数轮的老股东,TCV投过Netflix和Spotify这类基础设施型生意。再加上英伟达这个“既是金主又是供货商”的角色,整张股东表几乎全是基础设施和企业软件的长线钱,一边投基础模型,一边拿Fireworks在企业级调用做对冲。

这轮融资,Fireworks主要还是为了扩大算力,目前Fireworks从20多家供应商采购GPU,微软是其中之一。另外,Fireworks还打算把工程团队从约200人扩到年底600人、深化和微软与英伟达的合作。换句话说,它要把跑模型这件事的规模和可靠性,以及服务客户的规模上再往上推一个量级。

掌舵人Lin Qiao

Lin Qiao在中国长大,父亲是一家造船厂的资深机械工程师,从零开始造货船,她从小就会对着船舶蓝图读那些精确的角度和尺寸,迷上了把复杂东西拆成可计算零件的感觉;初中起她就偏在数理化上,高中时第一次用BASIC写了个贪吃蛇小游戏,那时就认定计算机科学是自己的未来。

1995年,Lin Qiao考上复旦大学计算机科学专业,本硕连读,又从这里第一次走进AI的世界。再之后她才赴美,在加州大学圣塔芭芭拉分校拿到计算机科学博士,方向是分布式系统和数据库管理。

博士毕业后她先去了IBM,在硅谷实验室和阿尔马登研究中心做研究员,主导过一套极速数据分析优化器的研发。接着是LinkedIn四年,做分布式数据服务,参与过Gobblin数据摄取框架。

Lin Qiao的职业生涯转折点出现在2015年7月,她加入Meta,一待七年,从数据基础设施的高级经理一路做到工程高级总监,手下团队从5个人长到300多人。

在Meta她干成了一件影响整个行业的事:共同创立并主导了PyTorch。

最早这被她当成“六个月就能搞定的项目”,结果变成一场历时五年的底层重建,把Meta整个AI工作负载从数据加载、分布式推理到跨全球数据中心和数十亿设备的训练全部重写一遍。等到她离开时,这套系统每天要撑起超过5万亿次推理。她后来有个比方:在Meta她建了引擎,现在她想建路。

Lin Qiao离开Meta的时点很妙。2022年10月,就在ChatGPT发布前一个月,她从Meta出走创立Fireworks。

当时她的判断是,企业都想上AI,但缺的是把模型真正部署进生产环境的基础设施、资源和人才,而她在Meta花五年解决的正是这个。Lin Qiao的想法是,把Meta要五年的事,压缩到五周甚至五天。七人创始团队里六人来自Meta的PyTorch组,一人来自谷歌的Vertex AI;其中三位是华人,除了Lin Qiao,还有Benny Chen,UCLA本科,Meta广告基础设施负责人和ChenyuZhao伯克利本科,谷歌Vertex AI技术经理。

一开始,Lin Qiao对外讲的理念叫“复合AI”(compound AI),主张用成百个小专家模型分别解决窄问题,而不是死磕单个大模型。落到Fireworks上,就是帮企业把开源模型调成自己的专用智能。

“AI有两条路。一条是智能属于少数几家大实验室,其他人全去租。另一条是全世界每家公司都建自己的专用智能,由只有它懂的那个领域来塑形。我们在建第二条路。”

据CNBC,同等质量下Fireworks比闭源模型便宜5到10倍。这恰恰是它增长的发动机:当最新模型的账单让老板和CFO们坐立难安,当开源模型能力与闭源模型差距缩小到某种程度,企业就开始认真考虑用开源替代了。

全球都在建的Token工厂

眼下,不少公司都在说自己是“Token工厂”,指的就是把底层算力、各种开源大模型、推理引擎打包封装成标准化API,开发者不用自己买显卡、搭环境,写几行代码、按量付费就能用上AI。谁掌握了服务这一层,谁就收这道过路费。

这门生意眼下同时在中美两边升温。

美国这边,Fireworks不是孤例,Together AI在7月1日刚关了8亿美元C轮,估值83亿美元,由沙特阿美旗下的Aramco Ventures领投,英伟达和Salesforce跟投,年预订额据称已破11.5亿美元,靠一套叫ATLAS的自适应推测解码引擎加速。

Baseten更猛,6月23日宣布15亿美元F轮、估值130亿美元,18个月融了四轮,收入同比涨约20倍,客户名单里甚至有Open Evidence、Harvey、Cursor、Notion,英伟达同样是它的股东。

再往外数还有Anyscale,贾扬清创立的Lepton、Modal、Replicate、Groq、Cerebras一众玩家。而且slogan大差不差:own your intelligence。

中国这边,故事换了个讲法,但内核不一样,流行的是异构架构。说白了,就是怎么把英伟达和其他国产芯片一起用起来,让模型推理更有效率。

比如硅基流动,创始人袁进辉是清华博士,早年创办过深度学习框架OneFlow,2023年8月在经历被并购又分拆的波折后重新出发。

这家公司今年已向港交所递表,想冲“AI Token工厂第一股”,成立不到三年完成七轮融资,投后估值77.4亿元人民币(约11亿美元),其中一笔超20亿元的B轮创下国内第三方MaaS赛道最大单笔纪录;阿里巴巴是最大机构股东,持股7.42%,华为旗下的哈勃科技持股4.07%。

硅基流动的跨芯片、多模型适配能力据称全球第一,平台累计支持170多个主流模型,2025年DeepSeek爆火、官网被打瘫时,它是业内第一家走通国产昇腾芯片部署满血版DeepSeek的公司。

再比如无问芯穹,由清华大学电子工程系教授汪玉发起、夏立雪任CEO,不久前拿到超7亿元融资,国资和产业资本入场,其Agentic MaaS平台上线160多种模型,日均token调用量较2025年底增长超20倍,公司内部提出一个“AI生产力=智能规模×Token生产效率×Token价值转化”的公式,把自己类比成石化产业链里的炼化厂,把能源转成数字石油

老牌的厂商比如九章云极,由清华大学电子工程系博士方磊创办、此前完成多轮战略融资,北京两大国有AI产业基金联合领投,国资与头部产业资本持续加码。前阵子刚开完“Token工厂”的发布会,其新一代Alaya NeW智算云AI工厂体系汇聚1000+大模型与行业专用模型,平台日均Token流转承载力目标直指10万亿,相较2025年底平台处理量增长超数十倍,首创DCU“一度算力” 计量单位,让算力像电力一样流通、计价、取用。

另外,中美两拨人最大的差别还在站位。美国那边的资本更猛、估值更高,客户是全球化的企业软件公司,很多公司英伟达都是直接下场当股东。国内这边,因为芯片生态碎片化,反而给这些中间层的Token工厂留了空间,但另一个问题是火山、阿里云、百度云这些大云厂就在前面,DeepSeek这类模型方也可能自己把推理做了。

未来如何就看故事要怎么讲了,但至少目前来看,需求很大。

自从年初agent的爆火以来,单次请求的token量暴增;开源模型的质量逼近闭源,企业有了真能用的替代;以及,企业终于开始认真算账,训练是一次性的,推理是持续流血的消耗,产品越受欢迎、调用越频繁、亏得可能越狠。

比如有投资人就说过,一些AI眼镜、AI玩具厂商巴不得用户少用点AI能力,拿回去当个摆设就好,因为“Token真烧不起”。

于是“能不能把每次调用的成本压下来”变成了生死线,而夹在模型、GPU、云厂和应用之间的Token工厂们,恰好是解题的人。

但要做Token,毛利率太重要了,算力采购价、利用率、客户压价,每一项都啃利润。如果未来竞争加剧、大家用低价抢客户,它就可能变成“高收入、低利润”的转售生意。如果未来训练的强度降低,也容易被云厂商和大模型公司两头夹击。

总之,不管Fireworks也好,硅基流动也罢,归根到底Token经济学的时代,Token相当于卖水人的角色,这和云计算早期那批卖服务器的公司有点像。可云计算最后收敛成了AWS、微软、谷歌三家,推理层会不会也走这条收敛路?独立推理层到底是台积电还是租车行,得看它能不能把“跑模型”变成客户离不开的、带垄断性的效率。

注:文/张楠,文章来源:投中网(公众号ID:China-Venture),本文为作者独立观点,不代表亿邦动力立场。

文章来源:投中网

广告
微信
朋友圈

FAQ回顾

什么是AI Token工厂?

AI Token工厂是将底层算力、开源大模型、推理引擎打包封装成标准化API的服务商,开发者无需自行采购显卡搭建环境,只需按调用量付费即可使用AI推理服务,还可支持企业基于自有数据微调定制专属模型。

Fireworks AI的核心优势有哪些?

Fireworks AI自研Fire Attention CUDA内核栈,同价位推理速度可达对手5倍;兼容OpenAI API,支持200多个开源模型调用及微调,可同时挂载100个LoRA适配器不加价;已获SOC2、HIPAA、GDPR合规资质,同等推理质量下比闭源模型便宜5到10倍。

AI推理服务赛道中美玩家有什么差异?

美国AI推理服务赛道资本热度更高、企业估值更高,客户多为全球化企业软件公司,英伟达多直接持股;国内更侧重异构架构适配,可同时兼容英伟达与国产芯片,同时面临云厂商、大模型公司的双重竞争压力。

Fireworks AI本轮D轮融资的资金将用于哪些方向?

Fireworks AI本次15.05亿美元D轮融资将主要用于扩大算力储备,从20多家供应商处采购GPU;同时计划将工程团队从200人扩张至年底的600人,深化与微软、英伟达的合作,提升服务规模与可靠性。

这么好看,分享一下?

朋友圈 分享

APP内打开

+1
+1
微信好友 朋友圈 新浪微博 QQ空间
关闭
收藏成功
发送
/140 0