广告
加载中

中国数据生成团队 首次登上Nature期刊

王露 2026-08-13 08:57
王露 2026/08/13 08:57

邦小白快读

EN
全文速览

本文核心介绍了中国首个登上Nature主要期刊的数据生成科创公司维纳智能的发展背景与核心业务,核心干货如下:

1. 核心成果:维纳智能由港科大教授柳崎峰创办,团队完成的AI辅助肾癌手术决策论文登上Nature通讯,提出的RDPM模型在外部多中心测试AUC达到0.788到0.873,可为肾癌手术决策提供量化支撑。

2. 核心行业观点:柳崎峰将大模型发展分为三个阶段,当前已经进入第三阶段“从Token到数据”,核心是让AI主动提问生成推理问答数据,形成“数据→模型→Token→数据”的自主学习闭环,解决当前大模型落地准确率低的痛点。

3. 商业化进展:维纳智能已经完成5000万港元种子轮融资,且验证了小团队、无行业专家、低成本跨行业复制的能力,落地了四个不同领域的头部客户。

本文披露了AI产业最新发展动向,可为各行业品牌的AI转型、业务布局提供参考,核心干货如下:

1. 产业消费趋势:当前大模型在千行百业落地难,核心痛点是专业领域准确率不足,根源是缺少高质量推理问答训练数据,市场对可规模化的自动数据生成技术需求迫切,这是品牌AI转型必须关注的核心趋势。

2. 产品研发方向:品牌布局AI相关产品时,不要只关注大模型本身,要重视训练数据的质量,可通过自动生成带推理链的行业专属数据,提升AI应用的准确率,解决答不准、优化难的问题。

3. 合作机会:维纳智能已经验证跨行业数据生成的可行性,医疗、金融、政务等多领域品牌,可通过对接这类专业服务商,低成本获得专业训练数据,加快AI产品落地。

本文披露了AI大模型领域最新的创业方向与商业机会,可为AI领域相关卖家提供方向参考,核心干货如下:

1. 新增蓝海机会:AI大模型发展已经从拼算力、拼大模型参数,转向拼高质量训练数据,传统人工标注成本高、难规模化、缺少推理过程,无法满足行业需求,自动推理数据生成是新的明确增长机会。

2. 可参考的成熟商业模式:维纳智能已经跑通“AI自动生成带完整思维链cQrA数据”模式,可实现小团队、无行业专家、低成本跨行业复制,已经落地四个高要求领域的头部客户,模式可复制性已经得到验证。

3. 风险提示:当前行业存在“大模型问答过时,只需要执行能力”的误区,实际上执行能力的基础是单智能体准确率,卖家布局AI业务要先解决数据质量问题,避开基础不牢的风险。

本文分享的AI发展逻辑,可为制造工厂推进数字化、AI转型提供多方面启示,核心干货如下:

1. AI生产设计的需求要点:工厂落地AI赋能产品设计、生产质检、流程优化等场景时,同样会遇到大模型准确率不足的问题,核心原因是缺少对应场景的高质量推理交互训练数据,补齐数据缺口才能提升AI落地效果。

2. 数字化转型的启示:工厂推进AI转型不要盲目追求参数更大的大模型,核心是要建立“场景数据生成-模型训练优化-数据迭代升级”的闭环,让AI在工厂实际场景中持续自我进化,适配工厂的个性化需求。

3. 商业机会:工厂可对接专业的自动数据生成服务商,低成本获得适配自身场景的高质量AI训练数据,比传统人工标注成本更低、效率更高,还能形成持续迭代的机制,加快AI转型的进度,降低转型成本。

本文明确了AI大模型行业最新的发展趋势、客户痛点与解决方案方向,可为AI相关服务商提供参考,核心干货如下:

1. 行业发展趋势:当前AI大模型已经进入第三发展阶段,业内已经形成共识:推理交互数据生成的质量决定了大模型的能力上限,数据生成领域将成为接下来AI行业的核心增长热点,市场需求会持续放大。

2. 核心客户痛点:当前各行各业落地大模型,普遍遇到准确率低、优化难度大、人工标注成本高、无法规模化复制的痛点,传统数据标注服务只给答案缺少推理过程,无法满足行业对高质量训练数据的需求。

3. 可参考的解决方案:服务商可参考维纳智能的路径,打造AI自动生成带完整思维链专业数据的服务,形成“数据生成反哺模型迭代”的闭环,替代传统人工标注,既降低客户成本,又能提升数据质量,可适配跨行业的客户需求。

本文披露了AI行业对平台的新需求,可为AI平台的招商、运营、方向调整提供参考,核心干货如下:

1. 行业对平台的新需求:大模型进入自主学习闭环的发展阶段后,行业需要适配新工作负载的算力平台,现有大多平台仅适配前两个阶段的大模型训练需求,需要提前升级能力适配新范式。

2. 平台招商新机会:AI三要素中算力、大模型都已经有较多布局,数据生成赛道还未被充分满足,维纳智能这类已经完成技术突破和商业化验证的头部数据生成公司,是优质的招商对象,引入这类企业可以完善平台AI赛道布局,抓住新的增长机会。

3. 风向规避:平台布局AI赛道不要盲目跟风只追具身智能、执行能力等概念,要避开“只谈执行不谈基础准确率”的误区,提前重视数据生成赛道,布局适配自主学习闭环范式的相关服务,提前卡位下一阶段的行业增长。

本文披露了AI大模型领域最新的产业动向与理论创新,对AI领域研究者有较高的研究参考价值,核心干货如下:

1. 最新产业新动向:当前大模型产业已经进入新的发展阶段,业内已经形成“推理交互数据生成决定大模型能力上限”的共识,中国已经出现头部创业公司,技术成果登上Nature主刊系列,完成了跨行业的商业化验证,数据生成已经成为AI产业新的核心赛道。

2. 新的理论与观点:柳崎峰提出大模型发展三阶段论,提出AI自主学习的核心是“数据→模型→Token→数据”的反馈闭环,原创性提出了cQrA、cTrA的数据定义,既为专业领域大模型训练提供了新燃料,也提供了新的评测标尺,该逻辑还延伸到了具身智能领域,拓展了研究方向。

3. 新商业模式验证:产业界已经验证了小团队跨行业复制自动数据生成的商业模式,突破了传统人工标注的人力瓶颈,为AI数据产业提供了新的研究样本。

返回默认

声明:快读内容全程由AI生成,请注意甄别信息。如您发现问题,请发送邮件至 run@ebrun.com 。

我是 品牌商 卖家 工厂 服务商 平台商 研究者 帮我再读一遍。

Quick Summary

This article introduces the background and core business of VINA AI, the first Chinese generative AI startup to be featured in a main Nature journal. Key takeaways are as follows:

1. Core achievement: Founded by HKUST professor Liu Qifeng, VINA AI published an AI-assisted kidney cancer surgery decision-making paper in Nature Communications. The proposed RDPM model achieved an AUC score of 0.788 to 0.873 in external multi-center trials, providing quantitative support for clinical kidney cancer surgery decisions.

2. Key industry insight: Professor Liu divides large model development into three stages, and argues the industry has now entered the third stage: "from tokens to data". The core of this stage is enabling AI to proactively generate reasoning question-and-answer data, forming a closed autonomous learning loop of "data → model → token → data" to solve the long-standing problem of low accuracy in real-world large model applications.

3. Commercial progress: VINA AI has completed a HKD 50 million seed round. It has already validated its ability to expand across industries with a small team, no dedicated industry experts, and low cost, and has secured four leading enterprise clients across different fields.

This article outlines the latest developments in the AI industry, providing actionable insights for brands across sectors pursuing AI transformation and business布局. Key takeaways are as follows:

1. Industry trend: Large models have struggled to deliver reliable results in most vertical industries, with the core pain point being insufficient accuracy in professional use cases. This shortfall stems from a lack of high-quality reasoning training data, creating urgent market demand for scalable automatic data generation technology — a core trend brands must watch in their AI transformation journeys.

2. Product development guidance: When building AI-powered products, brands should not only focus on large models themselves. They must prioritize the quality of training data: automatically generating industry-specific training data with complete reasoning chains can significantly improve AI application accuracy, and resolve common problems such as incorrect outputs and difficult model optimization.

3. Collaboration opportunities: VINA AI has already validated the feasibility of cross-industry automatic data generation. Brands in healthcare, finance, government affairs and other vertical sectors can partner with specialized service providers like VINA to obtain high-quality professional training data at low cost, accelerating the launch of their AI products.

This article introduces the latest startup directions and commercial opportunities in the large model space, providing strategic guidance for AI-focused sellers. Key takeaways are as follows:

1. New blue ocean opportunity: Large model competition has shifted from competing for computing power and parameter size to competing for high-quality training data. Traditional manual annotation cannot meet industry demand, as it is costly, hard to scale, and fails to capture complete reasoning processes. Automatic reasoning data generation has emerged as a clear new growth opportunity.

2. Proven replicable business model: VINA AI has built a working "automatic AI generation of cQrA data with complete thinking chains" model. This model enables cross-industry expansion with a small team, no industry experts, and low cost, and has already been deployed with leading clients in four high-standard sectors, proving its replicability.

3. Risk warning: There is a common industry misconception that "large model chat is obsolete, and only execution capability matters". In reality, reliable execution capability depends on accurate single-agent performance. Sellers building AI businesses must resolve data quality issues first to avoid the risk of building on an unstable foundation.

This article shares AI industry development insights that offer multiple lessons for manufacturing factories pursuing digital and AI transformation. Key takeaways are as follows:

1. Requirements for AI applications: When factories deploy AI for product design, production quality inspection, process optimization and other use cases, they also face the problem of insufficient large model accuracy. This is mainly caused by a lack of high-quality reasoning interaction training data for their specific scenarios, and closing this data gap is the key to improving AI deployment outcomes.

2. Lessons for digital transformation: Factories pursuing AI transformation should not blindly chase larger-parameter large models. The core priority is to build a closed loop of "scene data generation → model training and optimization → data iteration and upgrade", enabling AI to continuously self-improve in actual factory scenarios and adapt to factories' personalized needs.

3. Commercial opportunities: Factories can partner with professional automatic data generation service providers to obtain high-quality AI training data adapted to their specific scenarios at low cost. This approach delivers lower cost and higher efficiency than traditional manual annotation, while enabling continuous data iteration, accelerating AI transformation and cutting overall transformation costs.

This article clarifies the latest industry trends, customer pain points and solution directions for the large model industry, providing guidance for AI-focused service providers. Key takeaways are as follows:

1. Industry development trend: Large models have now entered the third stage of development, with a growing industry consensus that the quality of reasoning interaction data generation defines the upper limit of large model capability. The data generation field will become the core growth hotspot of the AI industry in the coming period, with market demand continuing to expand rapidly.

2. Core customer pain points: Enterprises across all sectors currently face common pain points when deploying large models: low accuracy, difficult optimization, high manual annotation costs, and inability to scale. Traditional data annotation services only provide final answers without reasoning processes, and cannot meet industry demand for high-quality training data.

3. Reference solution framework: Service providers can follow VINA AI's path to build services that automatically generate professional data with complete thinking chains, forming a closed loop where data generation feeds back into model iteration. This approach replaces traditional manual annotation, reduces customer costs, improves data quality, and can meet the needs of cross-industry customers.

This article outlines new industry demands on AI platforms, providing guidance for AI platforms on investment recruitment, operation and strategic adjustment. Key takeaways are as follows:

1. New industry demand for platforms: As large models enter the stage of closed-loop autonomous learning, the industry needs computing platforms adapted to this new workload. Most existing platforms are only built for the large model training requirements of the first two development stages, and need to upgrade their capabilities in advance to adapt to this new paradigm.

2. New investment opportunity for platforms: Among the three core elements of AI, computing power and large model development already have extensive market布局, while the data generation track remains underserved. Leading data generation companies like VINA AI, which have already achieved technical breakthroughs and commercial validation, are high-quality investment targets. Bringing in such companies can help platforms complete their AI industry布局 and capture new growth opportunities.

3. Risk mitigation: When布局 the AI track, platforms should not blindly chase trending concepts such as embodied intelligence or execution capability. They should avoid the misconception of "talking only about execution while ignoring foundational accuracy", prioritize the data generation track, build services adapted to the closed-loop autonomous learning paradigm, and position themselves early for the next phase of industry growth.

This article discloses the latest industrial developments and theoretical innovations in the large model field, offering high research value for AI researchers. Key takeaways are as follows:

1. Latest industrial development: The large model industry has now entered a new development stage, with growing consensus that "reasoning interaction data generation defines the upper limit of large model capability". China has already produced leading startups in this space, whose research has been accepted into Nature's main journal series and completed cross-industry commercial validation. Data generation has become a new core track in the AI industry.

2. New theoretical insight: Liu Qifeng proposed a three-stage framework for large model development, arguing that the core of autonomous AI learning is the closed feedback loop of "data → model → token → data". He also originally defined the cQrA and cTrA data categories, which provide both new training fuel and new evaluation metrics for vertical-domain large model training. This framework also extends to the embodied intelligence field, opening new research directions.

3. New validated business model: The industry has already validated the business model of cross-industry automatic data generation with small teams, breaking through the labor bottleneck of traditional manual annotation and providing new research samples for the AI data industry.

Disclaimer: The "Quick Summary" content is entirely generated by AI. Please exercise discretion when interpreting the information. For issues or corrections, please email run@ebrun.com .

I am a Brand Seller Factory Service Provider Marketplace Seller Researcher Read it again.

推理交互数据生成决定了大模型能力上限。

一个香港创业团队意外进入我们的视野。

前不久,一篇关于AI辅助肾癌手术决策的论文登上Nature通讯(https://www.nature.com/articles/s41467-026-73813-7)。凭借这篇论文,维纳智能成为中国首个、世界第四个登上Nature主要期刊(近三年IF>10)的数据生成科创公司——之前发文的中国大模型公司,是DeepSeek与面壁智能。

身后创始人柳崎峰教授,曾在香港科技大学建起全球首个千卡H800 SuperPod集群,预训练出中国第三家千亿参数大模型,管理项目研发资金超1亿美金。此后他盯上了如何提高AI的“提问”能力,来生成高质量推理问答数据,而这是即将爆发的AI自主学习的关键之一,于是在香港创办了维纳智能。

带着好奇,投资界与柳崎峰深谈近三个小时。话题从这篇论文说起,一路聊到大模型与具身智能,以及他对AI下一程的理解。

从一篇Nature通讯论文说起

别人是“从论文中来,到论文中去”,柳崎峰却是“从问题中来,到问题中去”。

2025年初,柳崎峰的亲属身患肾癌,主治医生是中山大学肿瘤医院张志凌主任。和其他所有肾癌手术一样,医生一直面临一个临床难题——部分肾切除和根治性肾切除之间,能否有更量化更智能的判断依据?

这个难题的本质是:AI能否预测真实世界里的复杂选择。

于是就在医院病房里,一场跨越医学与AI的合作就此展开:张志凌负责医学工作并联合多家医院完成数据收集,维纳智能负责AI和数据处理工作。论文共同第一作者王雅田,是港科大博士生、维纳智能实习生,由柳崎峰与罗文寒教授联合指导。

针对多源异构稀疏数据挑战,团队提出RDPM模型,将3D影像和临床变量/指标放进同一个预测框架,在1621例患者队列中完成训练与验证,外部多中心测试AUC达到0.788到0.873。论文预测了患者远期肾衰退风险,为高度依赖经验的手术决策提供了可量化支撑。

维纳智能在医疗这样最难容错的场景中,完成了一次AI预测的公开检验。这也指向了柳崎峰接下来要讲的AI预测的另一面——预测其实是大模型的底层机制,用预测下一个Token的方式生成答案,天生“善答”。而维纳智能的聚焦点,是更进一步——让AI不仅善答,更要“善问”。要让AI有“学问”,既要“学”得进去,更要“问”得出来。

港科大教授创业,联想创投领投第一轮

“别人是火什么,做什么。而他是做什么,火什么。”身边的朋友聊起对柳崎峰的过往印象。

这句话并非夸张。早在2001年,柳崎峰进入中国科学院自动化所模式识别国家重点实验室,师从谭铁牛院士——国际模式识别领域最高奖-傅京孙奖2022年得主。此后他先后任Samsung Lab研究员、Yahoo!Lab数据科学家、平安集团Gamma AI Lab主任、中国科学院香港创新研究院AI总监等职。2018年,他与杨强院士联合发起香港人工智能与机器人学会,2021年,他便极具前瞻性地为港府撰写了“香港云脑”和“香港基础大模型”建议书,成为香港AI超算建设与大模型训练的早期推动者。

看似分散的经历,都指向同一件事:让机器从复杂信息中找到模式并形成判断。

真正的转折发生在ChatGPT开始火爆的2023年。彼时柳崎峰在特区政府和学校领导的大力支持下,在香港科技大学联合6大高校,与郭毅可院士联合发起香港生成式人工智能研发中心,带队建成世界首个千卡H800 SuperPod AI超算集群,2024年,完成中国第三家千亿MoE大模型预/后训练。对于香港AI发展来说,这是一个关键节点。

也正是这段经历,他看见了下一道缺口:大模型越往深处走,越依赖高质量数据,永远是“数据为王”,尤其是千行百业的推理问答数据,于是让大模型能够高质量地“提问”,就成了首要关键。

柳崎峰把大模型发展拆成三个阶段:先是“从数据到模型”,用互联网大数据做预训练;再是“从模型到Token”,大模型开始输出Token生成内容或执行任务;接下来,是从“Token到数据”——让大模型系统主动提问、分步思考、校对回答,即推理问答数据生成。如此形成“数据→模型→Token→数据”的反馈大闭环,从而使AI具备自主学习能力。

AI自主学习的目的是获取“学问”,而“学问”就是“学”习训练 + 提“问”回答。清代学者刘开在《问说》中写道:“君子之学必好问。问与学,相辅而行者也。非学无以致疑,非问无以广识。”

2024年7月,香港维纳智能正式成立。公司名取自诺伯特·维纳——控制论的创始人。柳崎峰看重的,正是控制论里的反馈闭环。维纳智能的使命,就是让AI“问”得准、“答”得对,进而实现“数据→模型→Token→数据”的大闭环,从而让Agentic AI在专业领域自主演化。

维纳智能的任务,是要解决一个反直觉的问题,即一方面大模型发展一日千里,另一方面,大模型在企业落地依然非常难。原因很简单,就是准确度低。用学生备考来类比——光有教科书(专业文档),缺少习题集(推理问答数据),考试成绩不可能高(系统准确度低)。因为背教科书获取的是死的知识,而做习题集练的是活的解题能力。维纳智能在做的事,就是帮千行百业补上这本“习题集”,让AI不仅“学教材”,而且“做习题”,从而解决当下Agent泛滥所面临的测不准、优化难、答不准的瓶颈问题。

当下有一种流行说法:大模型问答已过时,执行任务才是关键。这未免流于表面。执行能力取决于两大支柱:单智能体在专业领域的准确度,与多智能体间的协同能力。可现实是,当前执行能力远未可靠。症结之一便在于单体问答准确率往往不足70%——连“可信”的门槛都迈不过,更遑论“协作”。

“习题”的具体定义是cQrA:context、Question、reasoning、Answer。context是任务现场,Question是生成的问题,reasoning是推理过程,Answer是校验后答案。换言之,维纳智能要让模型在一个具体行业语境里,同时生成问题、答案和推理过程。

这也拉开了与传统数据标注的距离。传统数据标注严重依赖人工甚至专家,成本高、难规模化,只给答案,缺少推理,专家经验被重复劳动消耗。维纳智能则让Agentic AI化身不知疲倦的智能专家团队,自动生成带完整思维链的cQrA数据,彻底突破人力瓶颈。比省钱更关键的是,闭环机制使每一轮生成的数据反哺生成和评价模型,驱动下一轮迭代在精度与逻辑上持续跃升——由此实现从“手工作坊”到“自我进化知识工厂”的质变。

维纳智能很快进入产业界和投资人视野。公司刚成立不久,便完成5000万港元种子轮融资,由联想创投领投。联想创投一直沿AI三要素赛道下注:算力投了沐曦、寒武纪等,模型投了智谱、阶跃星辰等,数据这一环,落到了维纳智能。同时,沐曦和维纳智能深度合作,在即将到来的“数据→模型→Token→数据”的大闭环时代,一个提前为未来范式设计了算力平台,一个提前为未来范式定义了工作负载。

AI下一程,“让我们生成这个世界”

商业化验证,先从两个灵魂拷问开始。

拷问一:在没有大规模专家标注的情况下,生成的数据是否有专业机构买单?

拷问二:是否可以做到跨行业、可复制?

为此,维纳智能顶住压力,打破了传统2B科创公司“深度优先”原则,即一定先要“击穿某个行业”,而是采用了“广度优先”,故意选了四个看似毫不相干、但准确度要求很高的行业:价值观安全、政务、保险、赛马,且每个行业都落地了头部客户。

“我们证明了可以用小团队、无行业专家、低成本实现跨行业复制。”柳崎峰表示,如今实现了“0到4”的验证,下一步就是“1到M x N”的推广(M个行业,每个行业N个头部客户)。

这背后,是对数据价值的长期判断。

在柳崎峰看来,中美AI的差距,很大程度上源于对数据的认知差异——数据长期被视为“脏活累活”,数据工程师的薪酬普遍低于算法与模型工程师。

但格局正在扭转。随着数据生产从人工标注走向推理、交互与闭环反馈,大模型公司在数据侧的投入持续加码,推理交互数据生成决定了大模型能力上限,已是业内共识。

而未来最核心的,不是模型,甚至不是数据本身,而是那个“大闭环”。正如进化的关键既非男人也非女人,而是交配与自然选择——染色体复制、交叉、变异和适者生存的机制。数据蒸馏只是借力外部的一种路径,真正的护城河,是建立起模型训练与数据生成相互驱动的自主学习闭环,用模型协同与反馈机制持续生成高质量数据。

这个判断也延伸到当下最热的具身智能。

传统方式训练具身智能,基于对人类的模仿。而真正智能应像婴儿学步“摸爬滚打”:在持续跌倒与尝试中自主生成动作数据,再通过闭环反馈迭代优化决策模型。

柳崎峰说,数字世界的闭环训练逻辑,已经延伸至物理世界。无论是Agent进入行业,还是机器人走向现场,都离不开提前自主生成的海量高质量推理交互数据。相应地,cQrA演进为cTrA——场景(context)、任务(Task)、推理(reasoning)、执行(Action),这既是训练的新燃料,也是评测的新标尺。

路刚刚开始,柳崎峰给出的答案,指向那个不远的未来:“让我们生成这个世界!”

注:文/王露,文章来源:投资界(公众号ID:pedaily2012),本文为作者独立观点,不代表亿邦动力立场。

文章来源:投资界

广告
微信
朋友圈

FAQ回顾

维纳智能是做什么的?

维纳智能是2024年7月在香港成立的数据生成科创公司,核心业务是通过AI自动生成带完整思维链的cQrA推理问答数据,构建“数据→模型→Token→数据”的自主学习闭环,帮助千行百业解决大模型落地准确度低的痛点问题。

AI辅助肾癌手术决策的RDPM模型效果如何?

针对多源异构稀疏数据挑战,维纳智能团队提出的RDPM模型将3D影像和临床变量纳入同一预测框架,在1621例患者队列中完成训练与验证,外部多中心测试AUC达到0.788到0.873,可预测患者远期肾衰退风险,为肾癌手术决策提供可量化支撑。

大模型在企业落地难的核心原因是什么?

当前大模型在企业落地难度大的核心原因是准确度低,类比学生备考仅靠教科书(专业文档),缺少习题集(推理问答数据),导致单体问答准确率往往不足70%,达不到可信门槛,也难以支撑多智能体间的协同任务。

维纳智能的cQrA数据相比传统数据标注有什么优势?

传统数据标注依赖人工甚至专家,成本高、难规模化,仅提供答案缺少推理过程。维纳智能的cQrA数据由AI自动生成,包含语境、问题、推理过程、校验后答案,可突破人力瓶颈,还能通过反馈闭环持续迭代提升精度与逻辑水平。

这么好看,分享一下?

朋友圈 分享

APP内打开

+1
+1
微信好友 朋友圈 新浪微博 QQ空间
关闭
收藏成功
发送
/140 0