今日焦点 · 大模型

大模型能力提升路线图:从"堆参数"到训练全栈 + 外层程序

把 2026 年可核查的公开证据整理成一张六层能力路线图——预训练、后训练 RL、推理时计算、上下文与记忆、智能体与 Harness、世界模型。含 Meta ScaleRL 40 万 GPU 小时实验结论、RL 预算占比 10%–30% 口径、Chinchilla 对比、Meta-Harness 6x 差距等数据锚点,并给出优先级表与算法工程师/产品经理的行动建议。
来源:本站原创2026-09-10
阅读全文 →

最新文章

LATEST 共 222 条
没自带 Token,就不配上大学了?
这个月生活费是给外卖,还是给大模型? #欢迎关注爱范儿官方微信公众号:爱范儿(微信号:ifanr),更多精彩内容第一时间为您奉上。
大模型 爱范儿 5天前
TransNormal-2: Geometry-Grounded Rectified Flow with Edge-Aware Decoding for Precise Normal Estimation
Diffusion-based models enable monocular geometry estimation, yet their pixel-space precision is limited by a shared, und…
行业动态 HuggingFace Daily Papers 5天前
Multi-Grid Post-Training for Long-Form Multi-Shot Video Generation
Generating long-form multi-shot videos requires coherent within-shot motion and visually consistent narratives across sh…
行业动态 HuggingFace Daily Papers 5天前
Cadence: Error-Bounded Lossy Compression of Demand Time Series with a Time-Series Foundation Model
We present Cadence, an error-bounded lossy compressor for numeric time series pairing a 330M-parameter time-series found…
行业动态 HuggingFace Daily Papers 6天前
GPT-6 突然全量上线,额度重置再+1,全网实测效果太离谱
GPT-6 Astra 的真正野心,是接管人类的电脑 #欢迎关注爱范儿官方微信公众号:爱范儿(微信号:ifanr),更多精彩内容第一时间为您奉上。
大模型 爱范儿 6天前
Steering Geometry: Validating Human Value Geometry in LLM Steering Space
As large language models (LLMs) are increasingly deployed in alignment-sensitive contexts, activation steering has emerg…
大模型 HuggingFace Daily Papers 6天前
Counter-Swarm Doctrine: Containing Coordinated Agent Intrusions
Agents can turn shared infrastructure into a channel for coordinated intrusion. The HF Mirror incident and a separate pu…
智能体 HuggingFace Daily Papers 6天前
DriveZero: End-to-End Driving Beyond Human Demonstrations
Most end-to-end autonomous-driving systems learn by imitating human driving logs, leaving their learned behavior constra…
行业动态 HuggingFace Daily Papers 6天前
Ask Before You Optimize: Dynamic Pre-Formulation Clarification for Interactive Optimization
Large language models (LLMs) are increasingly used to formulate optimization models from natural-language problem descri…
大模型 HuggingFace Daily Papers 9-4
GE-Act 2.0: Pretraining and Scaling a World-Action Model for Robotic Manipulation
World-action models (WAM) predict future states to guide robot actions, enabling learning from both action-free video an…
智能体 HuggingFace Daily Papers 9-4
Beneath the Surface of Chains-of-Thought: A Mechanistic Interpretation of Reasoning Operations in LLMs
Reasoning in large language models unfolds through diverse functional operations, such as problem formulation, goal deco…
研究前沿 HuggingFace Daily Papers 9-4
WorldSculpt: Generating Compositional Worlds from Grounded Videos
We study the problem of generating a compositional 3D representation of a cluttered scene containing hundreds of objects…
行业动态 HuggingFace Daily Papers 9-4
← 上一页 第 12 / 19 页 · 共 222 条 下一页 →