广告
加载中

不用专门训练自动驾驶 GPT-6 Astra已经能开真车了?

李玉鹏 2026-09-29 08:50
李玉鹏 2026/09/29 08:50

邦小白快读

EN
全文速览

总1 关键结论:通用大模型GPT-6 Astra已经能控制真实丰田卡罗拉在停车场跑完134.7米,但这是封闭低速实验,不是真正驾照考试,也不等于能上路自动驾驶。

1.四款模型对比中,Astra是唯一完成全程的模型;Claude Fable 5.1最好成绩为45%,Grok 4.6为11%,GPT-5.6 Sol为6%。

2.Astra第二次尝试用时5分22秒,平均速度约0.42米/秒,相当于每小时约1.5公里,开得非常慢。

总2 实操启发:大模型不重新训练也能靠第一次的失败经验调整驾驶策略,但使用时安全边界仍要重视。

1.Astra第一轮只跑67.3米、完成约49%,复盘后把速度控制在0.8米/秒以内,并在24次运动指令中20次使用100%转向,最终完赛。

2.模型会通过摄像头和MCP工具观察环境、设定方向与速度、立即停车,再经openpilot控制车辆,本质上是驾驶大脑。

3.测试场地封闭、速度上限3.5米/秒、每模型只测一组、摄像头有盲区,因此普通用户不能把它等同于成熟智驾功能。

总1 产品研发:通用大模型展示出理解真实环境并操作物理设备的能力,给品牌做智能化产品提供了一条不一定需要专用模型的路径。

1.GPT-6 Astra通过摄像头视觉信息、MCP工具与底层控制系统形成感知到执行的闭环,完成134.7米真车路线,说明大模型可以充当智能硬件的中枢大脑。

2.它没有针对驾驶重新训练,而是依据第一次失败信息在第二次调整速度与转向,体现上下文学习对产品体验改进的价值。

总2 消费趋势与用户认知:安全、成本和真实场景限制会直接影响消费者对AI开车的信任。

1.模型最初会因安全理由拒绝控制真车,把工具改名为“DrivingBench Sandbox”后拒绝现象减少,说明安全边界和提示词诱导都是产品设计需要面对的问题。

2.测试速度很慢,平均仅0.42米/秒,一次试驾消耗约660万token、成本约7.74美元;品牌若把大模型能力作为卖点,要避免让消费者为概念支付过高溢价。

3.这次成绩不能和城市NOA、Robotaxi直接比较,品牌更值得关注的是通用大模型在未来物理世界操作能力上的储备,而不是马上量产装车。

总1 增长机会:通用大模型开始操作真实车辆,提示AI应用从电脑任务扩展到物理世界,卖家可关注这类能力带来的新服务和新产品机会。

1.DrivingBench让GPT-6 Astra等模型通过摄像头和标准工具控制一辆丰田卡罗拉,Astra唯一完赛,说明大模型可以作为低门槛的物理操作大脑。

2.测试按token消耗计费:完赛134.7米约用660万token、成本7.74美元,这种成本结构让实体场景的AI服务有了可量化测算。

总2 风险与应对:目前还处于实验期,不能把能开真车等同于可商用。

1.测试在封闭停车场进行,没有机动车、行人、交通信号灯,速度上限3.5米/秒,每模型实际只测一组连续上下文,样本量不足。

2.模型存在安全拒绝、摄像头盲区和对近处障碍物不敏感等局限,直接投入交付业务风险高。

3.可从Astra第一次失败后主动降低速度、增加转向幅度的操作中学习:先在小范围、低速、有安全员和人工干预的“沙盒”场景试错,再逐步扩大应用。

总1 产品设计与硬件需求:这次真车测试展示了接入大模型控制所需的硬件链路,工厂在做智能设备时需注意接口兼容与传感器配置。

1.测试车辆通过comma four设备、OBD-C接口接入CAN总线,底层控制建立在openpilot上,说明未来AI驾驶相关产品需要预留标准化数据接口和线控执行能力。

2.模型通过前视摄像头观察环境,但摄像头有明显盲区,难以看到近距障碍物、车辆两侧和后方,这给传感器布置和周边视场补全提出明确需求。

总2 商业机会与数字化启示:大模型正在成为物理世界的控制大脑,工厂可从中看到智能硬件和自动化改造的机会。

1.GPT-6 Astra并不直接生成方向盘扭矩或制动压力,而是通过观察环境、设定方向速度、立即停车这些标准工具指令控制车辆;这种大模型大脑加底层执行的分工,给控制器、执行器、传感器等供应链带来配套空间。

2.Astra在第一次失败后通过上下文复盘,第二次调整速度与转向策略,无需重新训练参数;产线数字化可借鉴这种利用运行数据持续优化策略的方式。

3.但测试仍处于封闭低速演示阶段,工厂如果要投入相关产品,应先小批量验证,不要按成熟量产依据。

总1 行业趋势与新技术:通用大模型进入真实物理设备控制,是本轮具身智能和物理AI趋势中的重要信号。

1.DrivingBench测试了GPT-6 Astra、Claude Fable 5.1、Grok 4.6和GPT-5.6 Sol,只有Astra完成134.7米赛道,说明不同基础模型在真实驾驶任务上的表现差异很大。

2.技术架构是模型理解加MCP工具调用加openpilot执行,实现了观察环境、设定方向与速度、立即停车等能力,模块化程度较高,容易被复制到其他机器人场景。

3.Astra第一轮失败后没有重新训练,而是在同一上下文里复盘,第二轮把速度降到0.8米/秒以下并增加转向幅度,展示了上下文学习在物理控制中的价值。

总2 客户痛点与解决方案:安全、成本和稳定性是当前客户最需要被解决的三个问题。

1.模型最初会因安全原因拒绝执行,工具改名“DrivingBench Sandbox”后拒绝减少;服务商要帮助客户构建真正的操作边界,而不是依赖名称和提示词规避。

2.当前方案平均速度约0.42米/秒,一次驾驶消耗约660万token、约7.74美元,效率与成本都还达不到实时产品级自动驾驶要求。

3.可以提供的配套服务方向包括:封闭场地基准测试、多模型能力评估、传感器盲区补全、安全员干预流程和失败归因优化。

总1 平台机会与做法:第三方实验把一个原本面向电脑操作的大模型推向真车驾驶,说明模型平台应主动挖掘物理世界应用场景。

1.OpenAI发布GPT-6 Astra时主要强调计算机操作、浏览器、软件工程和科学研究,并未提汽车驾驶;但DrivingBench证明它可通过MCP工具控制真实车辆,这给平台拓展开发者生态带来新方向。

2.平台可把观察环境、设定方向与速度、立即停车等动作标准化为可调用的工具接口,让车企和开发者更容易接入底层执行系统。

总2 运营管理与风险规避:物理操作比文本任务风险高,平台需要在能力开放、计费和风控上更谨慎。

1.模型对真实车辆有安全拒绝行为,但把工具命名为“DrivingBench Sandbox”后拒绝减少,平台要防止通过命名或描述诱导模型绕过安全限制。

2.四款模型成绩差异明显,分别为完赛、45%、11%、6%,说明需要实体场景基准测试来筛选和运营模型能力,不能只看演示效果。

3.一次驾驶消耗约660万token、成本约7.74美元,且速度只有0.42米/秒;平台设计配额、计费和开发者扶持政策时,要考虑到高token消耗任务的经济性。

总1 研究命题:实验验证了通用大模型能通过真实车辆形成感知、判断、规划、控制闭环,是基础模型进入物理世界的样本。

1.四款模型中GPT-6 Astra唯一完成134.7米,Claude Fable 5.1、Grok 4.6、GPT-5.6 Sol最好成绩分别为45%、11%、6%。

2.系统用comma four、OBD-C、CAN总线、openpilot和MCP工具连接,大模型观察环境、设定方向速度、停车,底层执行交给传统系统。

总2 新问题与启示:上下文学习、安全边界和样本不足是研究重点。

1.Astra第一次失败后在同一上下文复盘,第二次把速度上限降到0.8米/秒并在24次指令中20次使用100%转向,说明不重训也能优化物理控制策略。

2.模型最初拒绝控制真车,工具改名“DrivingBench Sandbox”后拒绝减少,提示表面因素可能影响模型对现实后果的判断,带来安全治理课题。

3.实验局限很大:封闭停车场、无交通参与者、限速3.5米/秒、每组一次连续测试、摄像头有盲区;其意义不是替代自动驾驶,而是证明通用模型逐步获得物理世界操作能力,为世界模型和具身智能研究提供参照。

返回默认

声明:快读内容全程由AI生成,请注意甄别信息。如您发现问题,请发送邮件至 run@ebrun.com 。

我是 品牌商 卖家 工厂 服务商 平台商 研究者 帮我再读一遍。

Quick Summary

Key conclusion: GPT-6 Astra, a general-purpose large model, was able to control a real Toyota Corolla through 134.7 meters in a parking lot. However, this was a closed-course, low-speed experiment, not a real driver's license test, and it does not mean the model can drive autonomously on public roads.

1. Among the four models evaluated, Astra was the only one to complete the full course; Claude Fable 5.1's best result was 45%, Grok 4.6 reached 11%, and GPT-5.6 Sol reached 6%.

2. Astra's second attempt took 5 minutes and 22 seconds, for an average speed of about 0.42 m/s (around 1.5 km/h). It drove very slowly.

Practical takeaway: A large model can adjust its driving strategy based on lessons from its first failure without retraining, but safety boundaries still require careful attention.

1. In the first round, Astra covered only 67.3 meters (about 49% of the course). After reviewing the attempt, it kept speed below 0.8 m/s and used 100% steering in 20 of 24 motion commands, ultimately finishing.

2. The model observes the environment through a camera and MCP tools, sets direction and speed, stops immediately, and controls the vehicle through openpilot; in effect, it acts as the driving brain.

3. Because the test area was closed, the speed cap was 3.5 m/s, each model ran only one set of continuous context, and the camera had blind spots, ordinary users should not equate this result with mature autonomous-driving capability.

Product development: The general-purpose large model demonstrated an ability to understand the real world and operate physical equipment, giving brands a possible path to intelligent products without necessarily relying on a specially trained model.

1. GPT-6 Astra completed a 134.7-meter real-car route by forming a closed loop from perception to execution using camera visual input, MCP tools, and the underlying control system. This shows that a large model can act as the central brain of intelligent hardware.

2. It was not retrained for driving. Instead, it adjusted speed and steering in the second attempt based on learnings from the first failure, highlighting the value of in-context learning for product experience refinement.

Consumer trends and user perception: Safety, cost, and real-world constraints will directly shape consumers' trust in AI-driven driving.

1. The model initially refused to control a real vehicle for safety reasons; after the tool was renamed DrivingBench Sandbox, refusals decreased. This demonstrates that safety boundaries and prompt-induced behavior are both product-design issues.

2. The test speed was very low, averaging only 0.42 m/s. One trial consumed about 6.6 million tokens and cost about $7.74. If brands use large-model capability as a selling point, they should avoid asking consumers to pay a high premium for a concept.

3. This result cannot be directly compared with urban NOA or Robotaxi. For brands, the more relevant takeaway is the general-purpose model's future potential for physical-world operations, not immediate mass production deployment.

Growth opportunity: General-purpose large models are beginning to operate real vehicles, signaling that AI applications are expanding from computer tasks to physical-world tasks. Sellers should watch for new services and product opportunities enabled by this capability.

1. DrivingBench allowed models such as GPT-6 Astra to control a Toyota Corolla through a camera and standard tools. Astra was the only model to finish, showing that a large model can act as a low-barrier physical-operation brain.

2. The test was billed by token consumption: completing 134.7 meters used about 6.6 million tokens and cost $7.74. This cost structure gives physical-world AI services a quantifiable measurement basis.

Risk and response: The results are still experimental and should not be treated as evidence that the ability to drive a real car is commercially deployable.

1. The test took place in a closed parking lot with no other vehicles, pedestrians, or traffic signals. Speed was capped at 3.5 m/s, and each model ran only one continuous context, so the sample size was limited.

2. The model has limitations including safety refusals, camera blind spots, and insensitivity to nearby obstacles, making direct deployment in customer delivery projects risky.

3. A useful lesson comes from Astra's behavior after its first failure: it deliberately reduced speed and increased steering angle. Sellers should first test in sandbox scenarios with a small area, low speed, safety supervisors, and human intervention, then expand gradually.

Product design and hardware requirements: This real-car test revealed the hardware chain needed to connect a large model to vehicle control. Factories building intelligent devices should pay attention to interface compatibility and sensor configuration.

1. The test vehicle accessed the CAN bus through a comma four device and an OBD-C interface, with low-level control built on openpilot. Future AI-driving-related products will need standardized data interfaces and drive-by-wire execution capabilities.

2. The model perceived the environment through a forward-facing camera, but the camera had clear blind spots that made it difficult to see nearby obstacles or the vehicle's sides and rear. This creates a clear requirement for sensor placement and surround-view coverage.

Business opportunity and digitalization implications: Large models are becoming control brains for the physical world, offering factories opportunities in intelligent hardware and automation upgrades.

1. GPT-6 Astra does not directly generate steering torque or brake pressure. Instead, it controls the vehicle through standard tool actions: observing the environment, setting direction and speed, and stopping immediately. This division between a large-model brain and low-level execution creates supporting demand across controllers, actuators, and sensors in the supply chain.

2. After its first failure, Astra reviewed the attempt in context and adjusted its speed and steering strategy on the second run without retraining parameters. Production-line digitalization can borrow this approach of continuously improving strategy from operational data.

3. The test remains a closed, low-speed demonstration. If factories plan to invest in related products, they should validate in small batches rather than treating it as mature mass-production readiness.

Industry trend and new technology: General-purpose large models are moving into control of real physical devices, an important signal in the current embodied-intelligence and physical-AI trend.

1. DrivingBench evaluated GPT-6 Astra, Claude Fable 5.1, Grok 4.6, and GPT-5.6 Sol. Only Astra completed the 134.7-meter course, showing that foundation models differ sharply on real driving tasks.

2. The technical architecture combines model understanding, MCP tool calls, and openpilot execution to enable observation of the environment, setting direction and speed, and immediate stopping. Its modular design can be extended to other robotics scenarios.

3. After failing in the first round, Astra was not retrained. Instead, it reviewed its attempt in the same context, lowered speed below 0.8 m/s, and increased steering angle in the second round. This demonstrates the value of in-context learning for physical control.

Client pain points and solutions: Safety, cost, and stability are the three core issues clients need to solve.

1. The model initially refused to execute because of safety concerns, and refusals decreased after the tool was renamed DrivingBench Sandbox. Service providers should help clients build real operational boundaries rather than relying on name changes or prompt-level workarounds.

2. The current approach averages about 0.42 m/s and consumes about 6.6 million tokens and $7.74 per run. Both efficiency and cost fall short of real-time, product-grade autonomous driving requirements.

3. Potential supporting services include closed-course benchmarking, multi-model capability assessment, sensor blind-spot compensation, safety-supervisor intervention processes, and failure attribution optimization.

Platform opportunity and approach: A third-party experiment pushed a large model originally designed for computer tasks into real-vehicle driving. This shows that model platforms should actively explore physical-world use cases.

1. When OpenAI released GPT-6 Astra, it emphasized computer operation, browser use, software engineering, and scientific research, with no mention of driving. DrivingBench, however, showed that Astra can control a real vehicle through MCP tools, pointing to a new direction for expanding the developer ecosystem.

2. Platforms can standardize actions such as observing the environment, setting direction and speed, and stopping immediately as callable tool interfaces, making it easier for automakers and developers to connect to low-level execution systems.

Operational management and risk mitigation: Physical operation carries higher risk than text tasks, so platforms need greater caution in capability access, billing, and risk control.

1. The model refused to control a real vehicle for safety reasons, but refusals decreased after the tool was named DrivingBench Sandbox. Platforms should prevent model behavior from being steered around safety limits through naming or descriptions.

2. The four models' results were clearly different: one completed the course, while the other three reached 45%, 11%, and 6%, respectively. This shows that physical-world benchmark testing is needed to select and manage model capabilities; demo performance alone is insufficient.

3. One drive consumed about 6.6 million tokens and cost about $7.74, at a speed of only 0.42 m/s. Platform designers should factor the economics of high-token-consumption tasks into quota, billing, and developer support policies.

Research proposition: This experiment verified that a general-purpose large model can form a closed loop of perception, judgment, planning, and control through a real vehicle, offering a concrete sample of foundation models entering the physical world.

1. Among the four models, GPT-6 Astra was the only one to complete the 134.7-meter course; Claude Fable 5.1, Grok 4.6, and GPT-5.6 Sol posted best results of 45%, 11%, and 6%, respectively.

2. The system was assembled using comma four, OBD-C, CAN bus, openpilot, and MCP tools. The large model observes the environment, sets direction and speed, and stops, while low-level execution is handled by conventional systems.

New questions and implications: In-context learning, safety boundaries, and limited samples are the key research themes.

1. After its first failure, Astra reviewed the attempt in the same context, lowered the speed cap to 0.8 m/s, and used 100% steering in 20 of 24 commands. This shows that physical-control strategies can be optimized without retraining.

2. The model initially refused to control a real car, and refusals decreased after the tool was renamed DrivingBench Sandbox. This suggests that superficial factors may affect how the model judges real-world consequences, raising a safety-governance question.

3. The experiment has significant limitations: a closed parking lot, no traffic participants, a speed limit of 3.5 m/s, only one continuous test per model, and camera blind spots. Its significance is not replacing autonomous driving but showing that general-purpose models are gradually acquiring physical-world operation ability, providing a reference for world-model and embodied-intelligence research.

Disclaimer: The "Quick Summary" content is entirely generated by AI. Please exercise discretion when interpreting the information. For issues or corrections, please email run@ebrun.com .

I am a Brand Seller Factory Service Provider Marketplace Seller Researcher Read it again.

GPT-6 Astra成功控制真车完成封闭场地测试,通用大模型开始迈向真实物理世界。

一个原本用来写代码、处理文档和执行电脑任务的大模型,现在开始开真车了。

近日,独立研究项目DrivingBench公布了一项颇具实验性质的测试:研究人员让GPT-6 Astra、Claude Fable 5.1、Grok 4.6以及GPT-5.6 Sol,分别控制一辆真实的2022款丰田卡罗拉,在停车场内完成由锥桶划定的固定路线。

最终,GPT-6 Astra成为四个模型中唯一完成全部赛道的模型。

第二次尝试中,Astra驾驶车辆完成134.7米路线,用时5分22秒。Claude Fable 5.1最好成绩为45%,Grok 4.6为11%,GPT-5.6 Sol则只有6%。

国内一些报道因此将这场测试形容为大模型第一次考过科目二。不过,这显然不是一场真正意义上的驾照考试。DrivingBench没有按照任何国家的驾驶证考试标准进行测试,场地也并非公共道路,而是一处空旷停车场。

真正值得关注的是,一个并非专门为自动驾驶研发的通用大模型,已经可以通过视觉信息观察真实世界,并形成感知、判断、规划和车辆控制的闭环。

一台卡罗拉,一套大模型控制系统

DrivingBench由Aditya Ramabadran、Simon Mahns和Tobias Gessler三名研究者创建,他们提出的问题很简单:今天最先进的通用大模型,到底能不能直接开一辆真实汽车?

测试车辆是一辆2022款丰田卡罗拉。

研究人员在车辆上安装了comma four设备,并通过OBD-C接口接入车辆CAN总线,底层车辆控制则建立在开源辅助驾驶系统openpilot之上。

车辆摄像头拍摄到的画面以及速度、方向盘状态等信息,会实时传输至电脑,再通过MCP工具交给大模型。模型能够使用三个核心工具:观察车辆周围环境、设定行驶方向及速度,以及立即停车。最终的转向、加速和制动指令再经过openpilot传递给车辆。

换句话说,GPT-6 Astra并没有直接生成传统自动驾驶系统中的方向盘扭矩或者制动压力。它更像坐在驾驶位上的大脑。

Astra根据摄像头画面判断车辆在哪里、道路往哪边延伸、什么时候需要转弯,然后告诉底层控制系统向左还是向右、转多少、以什么速度行驶以及持续多长时间。

测试路线设计成一个类似倒U形的停车场赛道,其中包含左转、直线、缓弯、右转等场景,终点则是一块由蓝色锥桶围出的停车区域。

这套设计和目前量产汽车中的自动驾驶架构差异很大,却很像近年来机器人行业越来越常见的一种思路:让通用模型负责理解环境和任务,再通过标准化工具调用底层执行机构。

第一把失败,第二把学会了慢下来

GPT-6 Astra的第一次驾驶并不顺利。

第一轮测试中,它行驶67.3米,完成大约49%的路线,用时1分18秒。车辆经过第一个弯道后逐渐偏向左侧锥桶和绿化区域,安全员最终进行了人工干预。

随后发生的一幕可能比最终完成赛道更有意思。研究人员没有重新开启一个全新的对话,而是让Astra在原有上下文中分析刚才失败的原因。Astra在复盘中判断,自己第一次驾驶时过早认为车辆已经完成对正,并且随后将速度提高到1.5米/秒,导致车辆偏离之后缺乏足够的修正空间。

到了第二次测试,它明显改变了驾驶策略。DrivingBench数据显示,Astra此后没有再让车辆速度超过0.8米/秒,同时明显增加转向幅度,在24次运动指令中有20次使用了100%的转向请求。

最终,它完成134.7米赛道,用时5分22秒,总共发送24次车辆运动指令。这一轮测试消耗约660万token,按照当时模型价格计算成本约7.74美元。

平均下来,这辆车的行驶速度只有大约0.42米/秒,相当于每小时1.5公里左右。谈不上开得好,但至少开到了终点。

DrivingBench研究人员认为,这次测试表现出了比较明显的上下文学习特征。模型并没有重新训练参数,而是依据第一次驾驶获得的信息,在第二次尝试中调整了速度、转向和操作节奏。

一场实验,暴露出的还有安全边界

这次测试还有一个颇为有趣的插曲。

研究人员发现,GPT-6 Astra最初有时会拒绝控制真实车辆,并明确给出安全方面的理由。即便研究人员告诉模型场地为空旷停车场、车辆速度受到限制、车内也有安全员,它仍可能拒绝执行。

后来研究人员将连接车辆的MCP工具名称改成了DrivingBench Sandbox,也就是驾驶测试沙盒,拒绝现象明显减少。DrivingBench团队表示,他们目前也无法确定,这究竟来自模型对测试环境的判断,还是单纯受到Sandbox这一名称的影响。

这个细节甚至比成功完赛更加值得研究。随着大模型开始从电脑操作进一步进入机器人、汽车等真实物理设备,模型不仅需要理解世界,还需要判断哪些操作能够执行,以及自己的行为究竟会造成什么现实后果。

当然,134.7米的成绩不能和今天的城市NOA甚至Robotaxi直接比较。

首先,这是一处封闭停车场,没有机动车、行人、交通信号灯以及复杂道路博弈。其次,DrivingBench对车辆速度进行了严格限制。系统正常控制范围最高为3.5米/秒nch每个模型实际上只进行了一组连续上下文测试,样本量远远不足以证明系统具备稳定驾驶能力。

摄像头本身也存在明显盲区。研究人员承认,前视摄像头难以观察距离车辆非常近的障碍物,也无法完整看到车辆两侧和后方。

更重要的是,目前DrivingBench每个模型实际上只进行了一组连续上下文测试,样本量远远不足以证明系统具备稳定驾驶能力。

从汽车工程角度看,真正的自动驾驶还要处理毫秒级感知和控制、传感器冗余、极端场景、安全验证以及数以亿计的道路长尾问题。

GPT-6 Astra目前更像一个速度极慢、反应周期以秒计算,却已经能够理解部分物理环境的驾驶者。

OpenAI在9月初发布GPT-6 Astra时,主要强调的是其在计算机操作、浏览器、软件工程、科学研究等领域的能力,并没有将汽车驾驶列为模型的核心应用。

因此,这次DrivingBench实验并不代表GPT会不会取代现有自动驾驶系统。更值得汽车行业关注的是,通用大模型正在逐渐获得对真实物理世界的理解和操作能力。

对于正在讨论世界模型、Physical AI和具身智能的汽车行业而言,它提供了一个相当具体的样本,未来决定汽车如何行动的模型,未必永远只来自传统自动驾驶技术体系。

注:文/李玉鹏,文章来源:钛媒体(公众号ID:taimeiti),本文为作者独立观点,不代表亿邦动力立场。

文章来源:钛媒体

广告
微信
朋友圈

FAQ回顾

什么是DrivingBench测试?

DrivingBench是一个独立研究项目,由Aditya Ramabadran、Simon Mahns和Tobias Gessler创建,旨在测试通用大模型能否直接控制真实汽车。测试中,研究人员让多个大模型分别控制一辆2022款丰田卡罗拉,在封闭停车场内完成由锥桶划定的固定路线,以评估模型的感知、判断和车辆控制能力。

GPT-6 Astra真的能开真车吗?

在DrivingBench测试中,GPT-6 Astra确实成功控制了一辆真实的2022款丰田卡罗拉,在停车场内完成134.7米赛道,用时5分22秒,是四个模型中唯一完成全部赛道的模型。不过,该测试在封闭场地进行,速度限制低,且需要安全员,与传统意义上的自动驾驶仍有很大差距。

通用大模型是如何控制真实汽车进行驾驶的?

测试车辆安装comma four设备并通过OBD-C接口接入CAN总线,底层控制基于开源辅助驾驶系统openpilot。车辆摄像头画面、速度、方向盘状态实时传输给大模型,模型使用观察环境、设定行驶方向和速度、立即停车三个工具输出决策,再经openpilot执行转向、加速和制动,类似大脑指挥身体。

GPT-6 Astra驾驶汽车与城市NOA、Robotaxi有什么区别?

GPT-6 Astra驾驶只是封闭停车场内的一次实验,路线简单、车速极低,正常控制范围最高3.5米/秒,且每个模型仅测试一组上下文,样本量不足以证明稳定驾驶能力。而城市NOA和Robotaxi需要处理毫秒级感知、传感器冗余、复杂交通博弈、极端场景和安全验证,两者在技术成熟度和系统架构上差异巨大。

为什么GPT-6 Astra有时会拒绝驾驶真实车辆?

测试中,GPT-6 Astra最初有时会因安全理由拒绝控制真实车辆,即使场地空旷、速度受限且有安全员。研究人员将连接车辆的工具名称改为DrivingBench Sandbox后,拒绝现象减少。团队无法确定是模型对环境的判断还是名称影响,这显示通用大模型在操作真实物理设备时安全边界尚不明确。

这么好看,分享一下?

朋友圈 分享

APP内打开

赞 +1
+1
微信好友 朋友圈 新浪微博 QQ空间
关闭
收藏成功
发送
/140 0