AREX Feed Article
Nemotron 3 Ultra 零微调打穿前沿模型,Palantir CTO:差点觉得自己被 PUA
"我简直觉得自己被 PUA 了"
8 月 3 日,Palantir 第二季度财报电话会上,CTO Shyam Sankar 抛出一个数据点:团队将 NVIDIA Nemotron 3 Ultra 部署上线,零后训练,24 小时内,这款"原味"开源模型在客户生产任务上跑赢了前沿模型。Sankar 的发言片段由 X 用户 后三天内获得超过 32 万次查看,Sankar 本人。
两天后,:
Thanks @ssankar for putting NVIDIA Nemotron 3 Ultra to the test with no post-training. 24 hours later, it was outperforming frontier models on the tasks @PalantirTech customers needed to solve.
Sankar 在电话会上的原话——显示他说:
I literally almost felt gaslit when, within 24 hours of getting Nemotron up with no post-training, this is vanilla Nemotron Ultra, it did better than frontier.
他接着解释为什么这件事和所有人的直觉相反:
If you just looked at the numbers, you would say, 'It's nowhere near the Frontier. That shouldn't even be possible.' But of course, the benchmarks are wrong. I mean, the benchmarks are right for what the benchmark's measuring, but that's not my business. Those are not the tasks my customers had that they were trying to solve.
据 SeekingAlpha 发布的财报电话会记录,Palantir 在将 Nemotron Ultra 接入技术栈的 24 小时内,在 5 个生产任务上发现该模型未经后训练即超越了前沿模型。
Benchmark 合伙人:这是开源模型在应用开发者手里的力量
:
This is the power of open weight models in the hands of AI application developers. They're able to deliver significantly better results vs foundational models.
,只回复了一个词——"Yes!" 次日,,用"钢铁与汽车"的类比为开源权重模型辩护:模型权重是 AI 的钢铁——不应被监管,真正需要 crash-test 的是上层应用。
Palantir CEO Alex Karp 在同场财报电话会上给出了呼应,据 :
You own the weights, you own the alpha, you own everything.
NVIDIA 应用深度学习研究 VP Bryan Catanzaro 该模型实现了"前沿精度、5 倍速度、30% 成本降低"。
这场回声酝酿了两个月
Nemotron 3 Ultra 发布于 2026 年 6 月 4 日——NVIDIA 迄今最大的开源模型:5500 亿参数,混合 Mamba-Attention 的混合专家架构(MoE),每次推理仅激活约 550 亿参数,以 OpenMDW-1.1 许可协议在 Hugging Face 公开发布权重。
但在公开 benchmark 上,这款模型的成绩并不惊艳。在 Artificial Analysis Intelligence Index 综合榜单上,Nemotron 3 Ultra 得分 48.2,而同期闭源前沿模型 Claude Opus 4.8 得分 61.4,GPT-5.5 得分 60.2——差距超过 13 分。即便在同为开源阵营的 Kimi K2.6(53.9)面前也落后近 6 分。
Palantir 的测试打在了这道裂缝上:benchmark 测的是通用能力,客户要的是解决具体问题。当模型被嵌入企业已有的数据管道和业务逻辑中,分数低一截的模型反而赢了。
7 月 24 日,NVIDIA 与微软、Meta、Palantir 等公司联署了一封题为"开源权重与美国 AI 领导力"的公开信,反对对开源模型施加"过早限制";签署方在两天内从约 25 家增至约 50 家。
5500 亿参数,每次只激活十分之一
Nemotron 3 Ultra 的技术路线与主流 Transformer 模型有显著差异。它交替使用三种层级:Mamba-2 状态空间层以固定大小记忆状态处理序列、传统注意力层负责精确长距离查找、名为 LatentMoE 的混合专家层将每个 token 路由到少量专家子网络。
这套架构的直接收益体现在长上下文推理效率上。在长文本检索测试 RULER 中,Nemotron 3 Ultra 在 100 万 token 上下文窗口下得分 94.7%,是 NVIDIA 在技术报告中重点突出的数字。
模型同时内置多 token 预测(MTP)机制,可在单次前向传播中生成多个 token 并验证,减少生成响应所需的串行步骤。
在编程任务上,SWE-bench Verified 得分 71.9%,SWE-bench Multilingual 得分 67.7%。
排行榜失灵之后,什么才是标准?
Palantir 的测试结果触及了一个正在发酵的争议:公开 benchmark 到底在衡量什么,又漏掉了什么。
过去两年,前沿模型的竞争主要围绕少数几个综合榜单展开——AA Intelligence Index、Humanity's Last Exam、GPQA Diamond。闭源厂商用这些分数证明自己的模型"更聪明",开源阵营则用"逼近闭源"来争取开发者。
但 Sankar 的发言指向一个更朴素的真相:客户的生产任务不是 benchmark 题目。一个在 Humanity's Last Exam 上只得 26.7% 的模型,在特定企业的数据环境和任务定义下,可以击败综合得分高出 13 分的闭源竞品。
Puttagunta 的评论点出了关键变量——"在应用开发者手里"。Nemotron 3 Ultra 不是在裸跑 benchmark 时赢的,而是被嵌入 Palantir 的数据管道和业务逻辑之后赢的。模型权重可以下载、部署、嵌入企业环境,这是闭源 API 做不到的。Karp 所说的"你拥有权重,你就拥有 alpha",正是这个意思。
决胜不在排行榜上
Sankar 的"gaslit"不是一句修辞夸张。它揭示了一个正在成形的行业判断:模型好不好,通用 benchmark 说了不算,客户的生产任务说了才算。Nemotron 3 Ultra 在 Palantir 手中的表现不证明它比 GPT-5.5 或 Claude Opus 4.8"更聪明"。但它证明了一件更重要的事:当开源权重模型被掌握领域数据和业务逻辑的团队部署时,排行榜上 13 分的差距可能毫无意义。
Footnotes
-
,2026 年 8 月 4 日。摘要中提及"Within 24 hours of bringing Nemotron Ultra into our stacks, we found 5 production tasks where a standard Nemotron Ultra model without post-training beat frontier models."
-
,2026 年 8 月 5 日。该文综合引用了 NVIDIA 官方技术报告、Artificial Analysis 榜单数据及公开信相关报道。_