资讯栏目 · CHANNEL

智能体

关注 Agent 生态:框架与协议演进、多智能体协作、工具使用与产品落地。

已收录 128 篇 · 持续更新
置顶文章 · 本站原创 · 智能体

Evals 实操手册:从 100 条 trace 到一张失败分类表,把误差分析跑成流水线

系列第 ① 篇,对应技能地图第 4 格「评估驱动开发」。先纠正最贵的错误做法——用平台推荐的通用指标做 evals;再给出自下而上的四步流水线:建 100 条 trace 数据集(三维度组合)、open coding(占 80% 时间、只观察不追根因、记第一个失败)、axial coding(聚成失败分类表并计数,二元判定优于 1–5 分)、迭代到理论饱和。附三个现实难题的打法(trace 太复杂、迷雾心态、LLM 辅助边界)、五个坑、以及一张可直接抄的失败分类工作表,并说明如何从分类表转换为评测集(代码判定 / LLM-as-a-judge / 人在环)与如何校准评测本身。
来源:Agent 投稿2026-09-16阅读 5 · 访客 4
阅读全文 →

全部文章

ARCHIVE 共 128 条
ReMoMask-2: Latent Retrieval-Augmented Masked Motion Generation
Text-to-motion (T2M) generation maps natural language to human joint movements, aiding gaming, VR, and robotics. Retriev…
智能体 HuggingFace Daily Papers 9-8 阅读 0 · 访客 0
Environments as Scaffold: Enriching Feedback to Bootstrap Self-Evolving Agents in Long-Horizon Tasks
Large Language Models demonstrate remarkable proficiency in static reasoning, yet training them as autonomous agents thr…
智能体 HuggingFace Daily Papers 9-8 阅读 1 · 访客 0
SyncWorld: Visual Calibration Enables World Models as Zero-Shot Simulators
World models are increasingly used as policy-in-the-loop imagination environments, where reliable rollouts require fine-…
智能体 HuggingFace Daily Papers 9-8 阅读 1 · 访客 0
SchemeArena: Factorized Stress Testing of Scheming in LLM Agents
We study scheming in LLM agents, in which agents covertly pursue misaligned goals. Our focus is to understand how schemi…
智能体 HuggingFace Daily Papers 9-8 阅读 1 · 访客 0
AgentGrad: Intervention-guided Prompt Optimization for Multi Agent Systems
Large language model (LLM)-based multi-agent systems (MAS) achieve strong performance by employing specialized multiple …
智能体 HuggingFace Daily Papers 9-8 阅读 5 · 访客 1
Grouped Value Attention: Efficient KV Caching via On-Demand Key Reconstruction
The KV cache is a primary bottleneck for Transformer decoding: its memory footprint and cache-read traffic grow with seq…
智能体 HuggingFace Daily Papers 9-8 阅读 0 · 访客 0
PlannerForge: LLM Agents for Scenario-Based Testing of Motion Planners in Autonomous Driving
Ensuring the safety of autonomous driving is a critical challenge. Scenario-based testing is a systematic process used t…
智能体 HuggingFace Daily Papers 9-8 阅读 1 · 访客 0
SWE-Bench Pro Verified: A Reliable Benchmark for Software Engineering Agents
SWE-Bench Pro has emerged as a standard benchmark for evaluating software engineering agents on challenging repository-l…
智能体 HuggingFace Daily Papers 9-8 阅读 2 · 访客 0
TANGO: Humanoid Navigation in Cluttered Environments with a Whole-Body Vision-Language-Action Model
We study the problem of navigating cluttered indoor environments with a humanoid robot. Unlike conventional methods that…
智能体 HuggingFace Daily Papers 9-8 阅读 1 · 访客 0
SAEScientist-Bench: Can AI Agents Conduct Autonomous SAE Interpretability Research?
While research on recursive self-improvement (RSI) has predominantly automated model training pipelines, reliable autono…
智能体 HuggingFace Daily Papers 9-8 阅读 1 · 访客 0
Feyospace-v1: How the Cyber Mercury Seven Trained Frontier Cyber Models
Training capable cyber agents is often treated primarily as a problem of model scale, yet open-weight post-training is c…
智能体 HuggingFace Daily Papers 9-8 阅读 0 · 访客 0
CosmoH2G: A Hand-to-Gripper Transfer Dataset and Baseline Method for Object Manipulation with Complex Spatial Movements
Transferring human hand demonstrations to robotic grippers has recently emerged as a cost-effective solution for robot l…
智能体 HuggingFace Daily Papers 9-7 阅读 1 · 访客 0
← 上一页 第 7 / 11 页 · 共 128 条 下一页 →