城河并非更好的模型)
大家都在押注更大的模型。但對于物理人工智能而言智能本身并非倍增器學(xué)習(xí)才是。彼得·路德維希2026年7月28日A billion machines will become autonomous or intelligent over the next ten years. Cars, trucks, tractors, mining haulers, defense systems, warehouse robots, humanoids—the physical economy will be rebuilt around software that perceives, decides, and acts.The prevailing assumption about how we get there goes something like this: models keep improving, world models mature, foundation models for robotics arrive, and autonomy falls out the other end. Intelligence is the whole game; scale the intelligence and the machines will follow.I’ve spent a decade helping Applied Intuition build the software infrastructure behind many of the world’s most ambitious physical AI programs, from software-defined vehicles and autonomous trucks to construction equipment, mining systems, defense platforms, and robotics. That vantage point has given me a front-row seat to where the industry is accelerating—and where it continues to slow itself down. I’m as bullish on the intelligence as anyone, but the prevailing assumption gets the math wrong. Deployed physical AI is a product of two variables: the capability of the models, and the capacity of the engineering system around them—i.e. how requirements become software, how software gets validated, and how validated systems get deployed, monitored, and improved. The industry has largely poured everything into the first variable while the second sits roughly where it was a decade ago, built for quarterly releases and hundred-person integration teams. Frontier intelligence running on a legacy engineering system doesn’t produce frontier outcomes, because the old engineering system is the limiting factor.The contrarian bet, then, isn’t against intelligence. It’s that the next order of magnitude in physical AI comes from making the engineering system as intelligent as the models it carries. The industry’s roadmap, however, rests on a handful of assumptions that once made sense but no longer match where physical AI is headed.Smarter Models Don’t Create Deployed MachinesThe gap between a capable model and a certified, operating machine is enormous, and model quality alone doesn’t close it. A model that’s 20% better in benchmark terms still has to be integrated with many other software components, tested across millions of scenario variations, traced against safety requirements, validated on hardware, rolled out to a fleet, and monitored in the field. In a typical program, that pipeline — not the model — sets the tempo. Teams take delivery of a meaningfully better model and then spend two quarters proving it’s safe to ship.This is why world models, as remarkable as they are, won’t get us to a fully autonomous future on their own. They are advancing faster than engineering organizations can keep up. Every leap in model capability relocates the bottleneck rather than eliminating it. The constraint moves downstream, from “can the machine perceive the world?” to “can we validate, integrate, and operate what the machine can now do?” A team whose validation cycle takes months is, in effect, throttling frontier AI down to the speed of its own process.Here’s the implication the industry hasn’t priced in: as models commoditize toward the frontier, two companies with access to the same intelligence will have wildly different outcomes. The difference will be determined by how fast their engineering systems can absorb what the models can do. Intelligence is becoming ubiquitous. The ability to operationalize it isn’t.Digital AI Doesn’t Transfer to Physical AIThe second assumption is subtler: that the agentic revolution happening in digital work will naturally extend to physical AI. Direct the coding agents and copilots at the autonomy stack, and the same productivity gains will follow.They won’t, because most digital AI stops at documents, conversations, and code. Physical AI work doesn’t live there. It lives in drive logs and sensor data, in simulation runs and hardware-in-the-loop test rigs, and in requirements databases and validation reports. It lives in fleet telemetry streaming from real vehicles on real roads and job sites. An agent that has never seen a disengagement, doesn’t know why a perception regression matters, and can’t trace a requirement to a test case, is not a productivity tool in this domain. It’s a liability with a friendly interface.Making agents genuinely capable in deploying physical AI itself is a frontier intelligence problem, one that is entirely different than training a bigger model. The agents need access to the actual data and tools of the trade—simulators, data pipelines, validation systems—through interfaces hardened enough to trust. They need embedded domain judgment and the accumulated knowledge of what “validated” actually means when the artifact ships into a multi-ton machine. And they need evaluation and governance built in, because in this domain a plausible answer and a correct answer can be separated by fatal consequences.When I say physical AI needs its own agentic platform, I’m talking about a platform that combines state-of-the-art models grounded in the data layer, tooling, and domain expertise of physical systems, with evaluation and governance native to the platform. It is not a chatbot bolted onto engineering tools, and it is not something you get by fine-tuning a general-purpose agent. It’s a different architecture, and it demands as much AI innovation as the models themselves.The future of physical AI wont be determined by the smartest models alone, but by the engineering systems that turn intelligence into autonomous machines at scale.Speed Doesn’t Compromise SafetyWhen discussing agentic capabilities in physical AI development, the reflexive objection is that agents have no place in safety-critical engineering. Automation means moving fast and perhaps getting a few things wrong. It’s a reasonable position when the thing in question weighs several tons.However, in safety-critical systems, the speed of your feedback loopisa safety mechanism. When validation takes weeks, teams test at irregular milestones. When it takes minutes, they test on every change. Problems surface earlier, when they’re cheap to fix. Coverage expands to orders of magnitude more scenarios. Requirements get implemented more accurately and verified more often. The slow, careful-looking process isn’t by its nature the safest one. It’s often the one where defects age quietly for months before anyone notices.What actually makes agents safe in this domain isn’t slowing them down. It’s drawing the line correctly. Automate development, validation, and operations workflows, so safety-critical issues get resolved faster, and refuse to automate certification, regulatory sign-off, and final engineering judgment. High-stakes agents propose; humans decide. Anything touching production systems runs behind approval gates. This isn’t a temporary concession while the models improve. It’s the correct permanent architecture for physical AI, in the same way that a well-designed autonomous vehicle has a defined operational domain rather than unlimited authority.When agents help build physical AI faster, that speed can make the end product safer.Models Don’t Compound. Systems Do.When you understand these assumptions are holding physical AI back, the logical next step is combining better intelligence with faster learning, creating an agentic flywheel where frontier models and frontier engineering systems feed each other.A continuously turning flywheel looks like this. A machine underperforms in the field. Agents mine the operational data to find out where and why, or synthetically generate the scenarios that expose the gap. Findings become requirements, and requirements become test cases. Test cases become validated software, which is deployed to the fleet. The fleet generates new data, which makes the models better, which makes the agents better, which speeds up the next turn of the loop. Every turn used to take months and a room full of specialists. Each stage that agents accelerate doesn’t just save time; it increases the number of turns, and each turn compounds both the speed and the intelligence. Smarter models turn the loop faster. A faster loop makes the models smarter. That’s the compounding effect the industry is leaving on the table when it treats intelligence as the whole game.We know this flywheel is real because we’ve been running it on ourselves. Applied Intuition has spent a decade at the frontier of physical intelligence — perception, simulation, validation, vehicle software across automotive, trucking, mining, agriculture, and defense. Over the past year we built an agentic platform grounded in that infrastructure. It’s called Dana, and our engineers have built more than a thousand internal apps and agents on it. With Dana, development cycles have become roughly 20x faster, with higher output quality. Deployments went from once every few weeks to multiple times a day. Applications that took months to build now take days or hours. And, tellingly, applications that would never have justified months of effort now get built regularly. When the cost of building drops by an order of magnitude, the set of things worth building expands by more than an order of magnitude.The intelligence made the engineering system possible; the engineering system made the intelligence matter. Neither alone gets you there.What This Means for the Next DecadeIf the industry keeps betting on intelligence alone, the next decade of physical AI looks like a slow one: dazzling demos, decade-long programs, and a widening gap between what machines can do in a lab and what’s actually operating in the world. The models will be extraordinary and the deployment curve will stay stubbornly flat, because every improvement will queue up behind engineering organizations that absorb change at last year’s speeds.Pair frontier intelligence with agentic engineering systems and the curve bends. Physical AI starts reaching its full potential. Farms that stabilize output through labor and climate shocks. Mines with continuous, safer extraction. Freight networks that self-route around disruption. Defense systems that hold under degraded conditions. Multi-hundred-billion-dollar markets converging on the same stack, with the learning loop at the center of all of them.Software ate the world by making it cheap to build applications for the digital economy. Physical AI will do the same for the physical one, but only if building intelligent machines becomes as fast and iterative as building software, without compromising the discipline safety-critical systems demand. That takes the best modelsanda reinvention of how we engineer, and the second half is the one almost nobody is building.Twenty years from now, we won’t remember which company had the best world model in 2027. We’ll remember which company figured out how to continuously turn intelligence into deployed systems. That’s the problem we’ve been working on.未來十年將有十億臺機(jī)器實(shí)現(xiàn)自主運(yùn)行或具備智能。汽車、卡車、拖拉機(jī)、礦用運(yùn)輸車、國防系統(tǒng)、倉庫機(jī)器人、人形機(jī)器人——實(shí)體經(jīng)濟(jì)將圍繞能夠感知、決策和行動的軟件進(jìn)行重建。目前普遍認(rèn)為實(shí)現(xiàn)這一目標(biāo)的過程大致如下模型不斷改進(jìn)世界模型日趨成熟機(jī)器人技術(shù)的基礎(chǔ)模型最終形成自主性也隨之而來。智能是關(guān)鍵所在提升智能規(guī)模機(jī)器自然會隨之發(fā)展。過去十年我一直幫助Applied Intuition公司構(gòu)建支撐眾多全球最具雄心的物理人工智能項(xiàng)目的軟件基礎(chǔ)設(shè)施涵蓋軟件定義車輛、自動駕駛卡車、建筑設(shè)備、采礦系統(tǒng)、國防平臺和機(jī)器人等領(lǐng)域。這段經(jīng)歷讓我得以近距離觀察行業(yè)的發(fā)展趨勢以及它自身發(fā)展停滯不前的原因。我對人工智能的未來充滿信心但目前普遍的假設(shè)存在邏輯錯誤。物理人工智能的部署取決于兩個變量模型的能力以及圍繞模型構(gòu)建的工程系統(tǒng)的容量——即需求如何轉(zhuǎn)化為軟件、軟件如何得到驗(yàn)證以及經(jīng)過驗(yàn)證的系統(tǒng)如何部署、監(jiān)控和改進(jìn)。行業(yè)目前幾乎將所有資源都投入到了第一個變量上而第二個變量卻仍然停留在十年前的水平仍然沿用著季度發(fā)布和百人集成團(tuán)隊(duì)的模式。在老舊的工程系統(tǒng)上運(yùn)行的前沿智能無法產(chǎn)生前沿成果因?yàn)槔吓f的工程系統(tǒng)本身就是制約因素。因此這種反主流觀點(diǎn)并非反對智能本身而是認(rèn)為物理人工智能的下一個飛躍階段將來自于使工程系統(tǒng)本身達(dá)到與其承載的模型相同的智能水平。然而行業(yè)路線圖卻建立在一些曾經(jīng)合理但如今已不再符合物理人工智能發(fā)展方向的假設(shè)之上。更智能的模型不會創(chuàng)建已部署的機(jī)器一個性能優(yōu)異的模型與一臺經(jīng)過認(rèn)證、可實(shí)際運(yùn)行的機(jī)器之間存在著巨大的差距單靠模型質(zhì)量的提升并不能彌合這一差距。即使一個模型在基準(zhǔn)測試中提升了20%它仍然需要與其他眾多軟件組件集成在數(shù)百萬種場景變化中進(jìn)行測試對照安全要求進(jìn)行驗(yàn)證在硬件上進(jìn)行驗(yàn)證部署到整個機(jī)隊(duì)并在現(xiàn)場進(jìn)行監(jiān)控。在一個典型的項(xiàng)目中決定項(xiàng)目進(jìn)度的是整個流程而不是模型本身。團(tuán)隊(duì)拿到一個性能顯著提升的模型后還需要花費(fèi)兩個季度的時間來證明其安全性才能最終交付使用。這就是為什么世界模型雖然卓越但僅靠它們本身無法帶我們走向完全自主的未來。它們的進(jìn)步速度遠(yuǎn)遠(yuǎn)超過了工程組織的響應(yīng)速度。模型能力的每一次飛躍都只是將瓶頸轉(zhuǎn)移到了其他地方而不是消除它。限制因素向下游轉(zhuǎn)移從“機(jī)器能否感知世界”變成了“我們能否驗(yàn)證、集成并運(yùn)行機(jī)器現(xiàn)在能夠做到的事情”一個驗(yàn)證周期需要數(shù)月之久的團(tuán)隊(duì)實(shí)際上是在將前沿人工智能的速度限制在自身流程的速度之內(nèi)。行業(yè)尚未充分考慮以下影響隨著模型向前沿領(lǐng)域商品化兩家擁有相同情報資源的公司最終會得出截然不同的結(jié)果。這種差異取決于它們的工程系統(tǒng)吸收模型功能的速度。情報正變得無處不在但將其轉(zhuǎn)化為實(shí)際應(yīng)用的能力卻遠(yuǎn)未普及。數(shù)字人工智能無法轉(zhuǎn)化為物理人工智能第二個假設(shè)更為微妙數(shù)字工作中正在發(fā)生的智能體革命自然會擴(kuò)展到物理人工智能領(lǐng)域。引導(dǎo)編碼智能體和副駕駛關(guān)注自主系統(tǒng)架構(gòu)同樣的生產(chǎn)力提升也將隨之而來。他們不會因?yàn)榇蠖鄶?shù)數(shù)字人工智能都止步于文檔、對話和代碼。而物理人工智能的工作并非如此。它存在于駕駛?cè)罩竞蛡鞲衅鲾?shù)據(jù)中存在于模擬運(yùn)行和硬件在環(huán)測試平臺中存在于需求數(shù)據(jù)庫和驗(yàn)證報告中。它存在于來自真實(shí)道路和作業(yè)現(xiàn)場真實(shí)車輛的車隊(duì)遙測數(shù)據(jù)流中。一個從未經(jīng)歷過用戶脫離、不了解感知回歸為何重要、也無法將需求追溯到測試用例的智能體在這個領(lǐng)域并非生產(chǎn)力工具。它只是一個界面友好的累贅。讓智能體真正具備部署物理人工智能的能力是前沿智能領(lǐng)域的一大難題這與訓(xùn)練一個更大的模型截然不同。智能體需要通過足夠可靠的接口訪問實(shí)際數(shù)據(jù)和工具——模擬器、數(shù)據(jù)管道、驗(yàn)證系統(tǒng)等等。它們需要具備嵌入式領(lǐng)域判斷能力以及在將產(chǎn)品交付給一臺重達(dá)數(shù)噸的機(jī)器時“驗(yàn)證”的真正含義。此外它們還需要內(nèi)置評估和治理機(jī)制因?yàn)樵谶@個領(lǐng)域一個看似合理的答案和一個正確的答案之間可能存在致命的后果。我所說的物理人工智能需要其自身的代理平臺指的是一個將基于物理系統(tǒng)數(shù)據(jù)層、工具和領(lǐng)域?qū)I(yè)知識的尖端模型與平臺原生評估和治理功能相結(jié)合的平臺。它并非簡單地將聊天機(jī)器人附加到工程工具上也不是通過微調(diào)通用代理就能實(shí)現(xiàn)的。它是一種不同的架構(gòu)并且對人工智能創(chuàng)新提出了與模型本身同等的要求。物理人工智能的未來不僅僅取決于最智能的模型還取決于將智能大規(guī)模轉(zhuǎn)化為自主機(jī)器的工程系統(tǒng)。速度并不影響安全性在討論物理人工智能開發(fā)中的智能體能力時人們往往會反駁說智能體在安全攸關(guān)的工程領(lǐng)域沒有立足之地。自動化意味著快速行動但也可能導(dǎo)致一些錯誤。當(dāng)目標(biāo)物體重達(dá)數(shù)噸時這種觀點(diǎn)不無道理。然而在安全關(guān)鍵型系統(tǒng)中反饋循環(huán)的速度本身就是一種安全機(jī)制。當(dāng)驗(yàn)證需要數(shù)周時間時團(tuán)隊(duì)只能在不規(guī)則的里程碑節(jié)點(diǎn)進(jìn)行測試。而當(dāng)驗(yàn)證只需幾分鐘時他們就能對每一次變更進(jìn)行測試。問題會更早地被發(fā)現(xiàn)從而降低修復(fù)成本。測試覆蓋范圍也因此擴(kuò)展到更多場景。需求能夠得到更準(zhǔn)確的實(shí)現(xiàn)并被更頻繁地驗(yàn)證。緩慢而謹(jǐn)慎的流程本質(zhì)上并非最安全的。在這種流程中缺陷往往會悄無聲息地存在數(shù)月之久直到有人發(fā)現(xiàn)。在這個領(lǐng)域真正確保智能體安全的并非降低其運(yùn)行速度而是正確劃定界限。自動化開發(fā)、驗(yàn)證和運(yùn)維工作流程以便更快地解決安全關(guān)鍵問題同時拒絕自動化認(rèn)證、監(jiān)管審批和最終工程判斷。高風(fēng)險智能體提出方案由人類做出決定。任何涉及生產(chǎn)系統(tǒng)的事項(xiàng)都必須經(jīng)過審批流程。這并非模型改進(jìn)期間的臨時妥協(xié)而是物理人工智能的正確永久架構(gòu)正如設(shè)計(jì)良好的自動駕駛汽車擁有明確的運(yùn)行范圍而非無限的權(quán)限一樣。當(dāng)智能體幫助更快地構(gòu)建物理人工智能時這種速度可以使最終產(chǎn)品更安全。模型不會產(chǎn)生復(fù)合效應(yīng)系統(tǒng)才會。當(dāng)你理解這些假設(shè)阻礙了物理人工智能的發(fā)展時合乎邏輯的下一步就是將更智能的技能與更快的學(xué)習(xí)速度結(jié)合起來創(chuàng)建一個智能飛輪使前沿模型和前沿工程系統(tǒng)相互促進(jìn)。一個持續(xù)運(yùn)轉(zhuǎn)的飛輪看起來是這樣的一臺機(jī)器在實(shí)際應(yīng)用中性能不佳。智能體會挖掘運(yùn)行數(shù)據(jù)找出問題所在及原因或者合成場景來暴露差距。發(fā)現(xiàn)的問題轉(zhuǎn)化為需求需求轉(zhuǎn)化為測試用例。測試用例轉(zhuǎn)化為經(jīng)過驗(yàn)證的軟件并部署到整個系統(tǒng)中。系統(tǒng)生成新的數(shù)據(jù)從而改進(jìn)模型改進(jìn)智能體進(jìn)而加快循環(huán)的下一輪。過去每一輪循環(huán)都需要數(shù)月時間并且需要一屋子的專家參與。智能體加速的每一個階段不僅僅節(jié)省了時間它增加了循環(huán)次數(shù)而每一次循環(huán)都會同時提升速度和智能水平。更智能的模型能夠更快地完成循環(huán)。更快的循環(huán)速度又使模型更加智能。這就是行業(yè)將智能視為全部時所忽略的復(fù)合效應(yīng)。我們深知這種飛輪效應(yīng)真實(shí)存在因?yàn)槲覀円恢痹谧陨韺?shí)踐中驗(yàn)證它。Applied Intuition 十年來一直致力于物理智能領(lǐng)域的前沿研究涵蓋感知、仿真、驗(yàn)證以及汽車、卡車、采礦、農(nóng)業(yè)和國防等行業(yè)的車輛軟件。過去一年我們基于這一基礎(chǔ)架構(gòu)構(gòu)建了一個智能體平臺名為 Dana。我們的工程師已在其上開發(fā)了超過一千個內(nèi)部應(yīng)用程序和智能體。借助 Dana開發(fā)周期縮短了約 20 倍同時輸出質(zhì)量也顯著提高。部署頻率從幾周一次提升到每天多次。過去需要數(shù)月才能構(gòu)建的應(yīng)用程序現(xiàn)在只需幾天甚至幾小時即可完成。更重要的是那些過去根本不值得花費(fèi)數(shù)月時間開發(fā)的應(yīng)用程序現(xiàn)在也開始定期構(gòu)建。當(dāng)構(gòu)建成本降低一個數(shù)量級時值得構(gòu)建的項(xiàng)目數(shù)量級也會相應(yīng)增加。智慧使工程系統(tǒng)成為可能工程系統(tǒng)使智慧發(fā)揮作用。兩者缺一不可。這對未來十年意味著什么如果業(yè)界繼續(xù)只依賴智能那么未來十年物理人工智能的發(fā)展將會十分緩慢令人眼花繚亂的演示、長達(dá)十年的項(xiàng)目以及機(jī)器在實(shí)驗(yàn)室中的表現(xiàn)與實(shí)際應(yīng)用之間日益擴(kuò)大的差距。模型固然會非常出色但部署曲線卻會始終保持平緩因?yàn)槊恳豁?xiàng)改進(jìn)都將排在那些以去年速度吸收變革的工程團(tuán)隊(duì)之后。將前沿智能與智能工程系統(tǒng)相結(jié)合曲線將發(fā)生轉(zhuǎn)變。物理人工智能開始充分發(fā)揮其潛力。農(nóng)場能夠應(yīng)對勞動力和氣候沖擊穩(wěn)定產(chǎn)量。礦山能夠持續(xù)、安全地開采。貨運(yùn)網(wǎng)絡(luò)能夠自動繞過中斷。防御系統(tǒng)能夠在惡劣條件下保持有效運(yùn)行。數(shù)千億美元的市場將匯聚到同一技術(shù)棧上而學(xué)習(xí)循環(huán)則是所有這些市場的核心。軟件通過降低構(gòu)建數(shù)字經(jīng)濟(jì)應(yīng)用的成本徹底改變了世界。物理人工智能也將對物理世界產(chǎn)生同樣的影響但這只有在構(gòu)建智能機(jī)器的速度和迭代性能夠與構(gòu)建軟件一樣快并且不損害安全關(guān)鍵系統(tǒng)所要求的嚴(yán)謹(jǐn)性時才能實(shí)現(xiàn)。這需要最佳模型和對工程方式的徹底革新而后半部分幾乎無人涉足。二十年后我們不會記得哪家公司在2027年擁有最佳的世界模型。我們會記住哪家公司找到了將情報持續(xù)轉(zhuǎn)化為可部署系統(tǒng)的方法。這正是我們一直在努力解決的問題。