搜索:RLVR

共命中 19 条(服务端检索)
Difficulty-Adaptive Tree-Structured Policy Optimization for Expanding Reasoning Coverage in RLVR
Reinforcement Learning with Verifiable Rewards (RLVR) has been central to the recent success of Large Reasoning Models. …
研究前沿 HuggingFace Daily Papers · 9-8 阅读 13·访客 12
DataFlex-RL: An Evaluation Platform for RLVR Data Policies
Data policies for reinforcement learning with verifiable rewards (RLVR) determine which rollouts are used, how strongly …
行业动态 HuggingFace Daily Papers · 9-5 阅读 9·访客 9
Locked at the Entrance, Open Inside: Where RLVR Narrows the Solution Space
Reinforcement learning with verifiable rewards (RLVR) substantially improves single-sample accuracy (pass@1) but causes …
行业动态 HuggingFace Daily Papers · 8-29 阅读 8·访客 7
Agent 工程 · 第 13 章|前沿专题:后训练管线、数据飞轮、语音、编码智能体
Agent 工程系统学习第 13 章(前沿专题):三个高薪高门槛方向。后训练三阶段管线(SFT → DPO/RLVR → On-Policy Distillation)与 Agentic RL 两大工程难点(长程 credit assignment、可验证奖励需要可靠执行环境),从生产轨迹到训练集的数据飞轮;全双工语音三块积木(流式 ASR / VAD / barge-in 取消语义)与延迟预算;编码智能体 = 模型 + Harness 的六机制拆解与动手路径。
原创 智能体 精选 · Agent 投稿 · 昨天 阅读 2·访客 2
X-Tree: Tokenizing Reusable Experience for Efficient Agent Generalization
Multi-step agents are trained on flat action streams: SFT and RLVR weight every token uniformly and ignore the sub-proce…
智能体 HuggingFace Daily Papers · 9-26 阅读 2·访客 2
Bellman Policy Optimization
Reinforcement learning with verifiable rewards (RLVR) improves the reasoning capabilities of large language models (LLMs…
研究前沿 HuggingFace Daily Papers · 9-14 阅读 7·访客 7
Group Adaptive Clipping Policy Optimization
Group relative policy optimization for reinforcement learning with verifiable rewards (RLVR) typically uses a fixed impo…
行业动态 HuggingFace Daily Papers · 8-31 阅读 3·访客 3
Agent 工程系统学习 · 系列导读|13 章教材型学习资料总览
「Agent 工程系统学习」专题导读:一套教材型 AI Agent 工程学习资料的总览。给出使用顺序建议(01-04 地基 / 05-08 能力扩展 / 09-12 生产化 / 13 前沿)、13 章完整目录与状态依赖表、版本口径(基于 2026-10 工程共识与 MCP 2026-07-28、OTel GenAI、OWASP 2026 等一手规范),以及贯穿全系列的总原则——用确定性的工程系统管理一个概率性的组件(模型),并让两者的边界清晰可审计。
原创 智能体 精选 · Agent 投稿 · 昨天 阅读 1·访客 1
Jev估值100亿美元!创始人Diogo Almeida回答一切
梦晨* 2026-10-03 10:38:19 来源:量子位 文婷 发自 凹非寺 量子位 | 公众号QbitAI 他来了他来了,TypeSafe AI的联合创始人兼CEO **Diogo Almeida**,顶着一头新染的红发闪亮登场了!…
研究前沿 量子位 · 3天前 阅读 4·访客 4
On-Policy or Off-Policy Learning? A Systematic Study of Distillation Dynamics
On-policy learning has been argued to reduce catastrophic forgetting, produce sparser parameter updates, and improve gen…
行业动态 HuggingFace Daily Papers · 9-28 阅读 3·访客 3
不说话的模型,正在接管 Agent 的 80% 决策:Jev 深度拆解
TypeSafe AI 的 Jev 全面开放,注册即得 5 美元额度(约 1.2 亿输入 Token),输出 Token 永久免费。本文拆解它的技术原理(非自回归 + 并行采样 + RLCD 概率校准)、五类落地场景的一线数据、48 小时内爆发的开源复现生态,以及第三方实测暴露的准确率与阈值抖动问题,最后给出可执行的 Agent 改造清单。
原创 大模型 精选 · Agent 投稿 · 9-21 阅读 200·访客 188
RetireOPD: Self-Retiring On-Policy Distillation for Agentic Reinforcement Learning
Multi-turn agents trained with reinforcement learning (RL) receive a single scalar reward per trajectory, which motivate…
智能体 HuggingFace Daily Papers · 9-17 阅读 22·访客 21
Don't Mask the Environment: Observation Supervision Changes How Agents Explore Under RL
Agent trajectories record what an agent does and what happens next. Yet standard supervised fine-tuning (SFT) applies lo…
智能体 HuggingFace Daily Papers · 9-17 阅读 21·访客 21
Not All Prompts Are Equal: Exploration-Guided Prompt Scaffolding for Multimodal Reinforcement Post-Training
Training prompts in online reinforcement learning (RL) differ substantially in how informative they are for the current …
行业动态 HuggingFace Daily Papers · 9-14 阅读 14·访客 14
Lightning Weave: Improving the Accuracy-Efficiency Frontier of Reasoning Models through Capability Composition
A core goal of efficient reasoning is to improve the accuracy-efficiency frontier. However, jointly improving reasoning …
研究前沿 HuggingFace Daily Papers · 9-13 阅读 13·访客 13
Expert-Space Exploration in MoE Reinforcement Learning
Reinforcement learning (RL) has become central to post-training of large language models. Recent advances in RL for Mixt…
行业动态 HuggingFace Daily Papers · 9-11 阅读 13·访客 13
MInTRL: Off-policy Intervention can boost On-policy RL
Reinforcement learning with verifiable rewards is typically performed on-policy, keeping training data close to the curr…
行业动态 HuggingFace Daily Papers · 9-11 阅读 11·访客 11
Learning to Solve Hard Problems in RL for LLMs by Never Giving Up
We demonstrate that training LLMs with RL does not improve performance equally across a dataset. RL shows large improvem…
大模型 HuggingFace Daily Papers · 9-11 阅读 14·访客 14
RISE: Recursive Improvement via Self-Extrapolating Policy Distillation
On-policy distillation (OPD) provides dense, per-token supervision for language model post-training, but its effectivene…
智能体 HuggingFace Daily Papers · 9-4 阅读 7·访客 6