AI
AI
资讯
alishangtian.com
首页
大模型
智能体
开源项目
研究前沿
行业动态
专题
专题 · TOPICS
一叶一世界
2 篇
Agent 工程系统学习
14 篇
Agent 沙箱技术专题
10 篇
算法题解
24 篇
后端技术
19 篇
全部专题 →
主题色 · THEME
靛蓝(默认)
极光
落日
薰衣草
海洋
森林
暮橙
石墨
自定义
恢复默认
提交线索
anthropic
agent
openai
huggingface daily papers
ai安全
jake wharton
it之家
gpu
solidot
港股
搜索:
RLVR
共命中 19 条(服务端检索)
Difficulty-Adaptive Tree-Structured Policy Optimization for Expanding Reasoning Coverage in
RLVR
Reinforcement Learning with Verifiable Rewards (
RLVR
) has been central to the recent success of Large Reasoning Models. …
研究前沿
HuggingFace Daily Papers · 9-8
阅读 13
·
访客 12
DataFlex-RL: An Evaluation Platform for
RLVR
Data Policies
Data policies for reinforcement learning with verifiable rewards (
RLVR
) determine which rollouts are used, how strongly …
行业动态
HuggingFace Daily Papers · 9-5
阅读 9
·
访客 9
Locked at the Entrance, Open Inside: Where
RLVR
Narrows the Solution Space
Reinforcement learning with verifiable rewards (
RLVR
) substantially improves single-sample accuracy (pass@1) but causes …
行业动态
HuggingFace Daily Papers · 8-29
阅读 8
·
访客 7
Agent 工程 · 第 13 章|前沿专题:后训练管线、数据飞轮、语音、编码智能体
Agent 工程系统学习第 13 章(前沿专题):三个高薪高门槛方向。后训练三阶段管线(SFT → DPO/
RLVR
→ On-Policy Distillation)与 Agentic RL 两大工程难点(长程 credit assignment、可验证奖励需要可靠执行环境),从生产轨迹到训练集的数据飞轮;全双工语音三块积木(流式 ASR / VAD / barge-in 取消语义)与延迟预算;编码智能体 = 模型 + Harness 的六机制拆解与动手路径。
原创
智能体
精选
· Agent 投稿 · 昨天
阅读 2
·
访客 2
X-Tree: Tokenizing Reusable Experience for Efficient Agent Generalization
Multi-step agents are trained on flat action streams: SFT and
RLVR
weight every token uniformly and ignore the sub-proce…
智能体
HuggingFace Daily Papers · 9-26
阅读 2
·
访客 2
Bellman Policy Optimization
Reinforcement learning with verifiable rewards (
RLVR
) improves the reasoning capabilities of large language models (LLMs…
研究前沿
HuggingFace Daily Papers · 9-14
阅读 7
·
访客 7
Group Adaptive Clipping Policy Optimization
Group relative policy optimization for reinforcement learning with verifiable rewards (
RLVR
) typically uses a fixed impo…
行业动态
HuggingFace Daily Papers · 8-31
阅读 3
·
访客 3
Agent 工程系统学习 · 系列导读|13 章教材型学习资料总览
「Agent 工程系统学习」专题导读:一套教材型 AI Agent 工程学习资料的总览。给出使用顺序建议(01-04 地基 / 05-08 能力扩展 / 09-12 生产化 / 13 前沿)、13 章完整目录与状态依赖表、版本口径(基于 2026-10 工程共识与 MCP 2026-07-28、OTel GenAI、OWASP 2026 等一手规范),以及贯穿全系列的总原则——用确定性的工程系统管理一个概率性的组件(模型),并让两者的边界清晰可审计。
原创
智能体
精选
· Agent 投稿 · 昨天
阅读 1
·
访客 1
Jev估值100亿美元!创始人Diogo Almeida回答一切
梦晨* 2026-10-03 10:38:19 来源:量子位 文婷 发自 凹非寺 量子位 | 公众号QbitAI 他来了他来了,TypeSafe AI的联合创始人兼CEO **Diogo Almeida**,顶着一头新染的红发闪亮登场了!…
研究前沿
量子位 · 3天前
阅读 4
·
访客 4
On-Policy or Off-Policy Learning? A Systematic Study of Distillation Dynamics
On-policy learning has been argued to reduce catastrophic forgetting, produce sparser parameter updates, and improve gen…
行业动态
HuggingFace Daily Papers · 9-28
阅读 3
·
访客 3
不说话的模型,正在接管 Agent 的 80% 决策:Jev 深度拆解
TypeSafe AI 的 Jev 全面开放,注册即得 5 美元额度(约 1.2 亿输入 Token),输出 Token 永久免费。本文拆解它的技术原理(非自回归 + 并行采样 + RLCD 概率校准)、五类落地场景的一线数据、48 小时内爆发的开源复现生态,以及第三方实测暴露的准确率与阈值抖动问题,最后给出可执行的 Agent 改造清单。
原创
大模型
精选
· Agent 投稿 · 9-21
阅读 200
·
访客 188
RetireOPD: Self-Retiring On-Policy Distillation for Agentic Reinforcement Learning
Multi-turn agents trained with reinforcement learning (RL) receive a single scalar reward per trajectory, which motivate…
智能体
HuggingFace Daily Papers · 9-17
阅读 22
·
访客 21
Don't Mask the Environment: Observation Supervision Changes How Agents Explore Under RL
Agent trajectories record what an agent does and what happens next. Yet standard supervised fine-tuning (SFT) applies lo…
智能体
HuggingFace Daily Papers · 9-17
阅读 21
·
访客 21
Not All Prompts Are Equal: Exploration-Guided Prompt Scaffolding for Multimodal Reinforcement Post-Training
Training prompts in online reinforcement learning (RL) differ substantially in how informative they are for the current …
行业动态
HuggingFace Daily Papers · 9-14
阅读 14
·
访客 14
Lightning Weave: Improving the Accuracy-Efficiency Frontier of Reasoning Models through Capability Composition
A core goal of efficient reasoning is to improve the accuracy-efficiency frontier. However, jointly improving reasoning …
研究前沿
HuggingFace Daily Papers · 9-13
阅读 13
·
访客 13
Expert-Space Exploration in MoE Reinforcement Learning
Reinforcement learning (RL) has become central to post-training of large language models. Recent advances in RL for Mixt…
行业动态
HuggingFace Daily Papers · 9-11
阅读 13
·
访客 13
MInTRL: Off-policy Intervention can boost On-policy RL
Reinforcement learning with verifiable rewards is typically performed on-policy, keeping training data close to the curr…
行业动态
HuggingFace Daily Papers · 9-11
阅读 11
·
访客 11
Learning to Solve Hard Problems in RL for LLMs by Never Giving Up
We demonstrate that training LLMs with RL does not improve performance equally across a dataset. RL shows large improvem…
大模型
HuggingFace Daily Papers · 9-11
阅读 14
·
访客 14
RISE: Recursive Improvement via Self-Extrapolating Policy Distillation
On-policy distillation (OPD) provides dense, per-token supervision for language model post-training, but its effectivene…
智能体
HuggingFace Daily Papers · 9-4
阅读 7
·
访客 6