Evals 实操手册:从 100 条 trace 到一张失败分类表,把误差分析跑成流水线
系列第 ① 篇,对应技能地图第 4 格「评估驱动开发」。先纠正最贵的错误做法——用平台推荐的通用指标做 evals;再给出自下而上的四步流水线:建 100 条 trace 数据集(三维度组合)、open coding(占 80% 时间、只观察不追根因、记第一个失败)、axial coding(聚成失败分类表并计数,二元判定优于 1–5 分)、迭代到理论饱和。附三个现实难题的打法(trace 太复杂、迷雾心态、LLM 辅助边界)、五个坑、以及一张可直接抄的失败分类工作表,并说明如何从分类表转换为评测集(代码判定 / LLM-as-a-judge / 人在环)与如何校准评测本身。
阅读全文 →
全部文章
ARCHIVE 共 128 条办公 Agent 大乱战,新势力TeleAgent 凭什么坐上牌桌
谁能成为真正的「国民级AI 办公助理」 #欢迎关注爱范儿官方微信公众号:爱范儿(微信号:ifanr),更多精彩内容第一时间为您奉上。
OpenAI 宣布攻克千禧年难题,清华姚班传奇陈立杰:不可思议的时代
1 万个 AI Agent,挑战百年数学难题 #欢迎关注爱范儿官方微信公众号:爱范儿(微信号:ifanr),更多精彩内容第一时间为您奉上。 ]
首个走进联合国的中国教育 Agent,正在打开下一个 Token 入口
Coding 之后,教育 Agent 正在成为下一场 Token 战争 #欢迎关注爱范儿官方微信公众号:爱范儿(微信号:ifanr),更多精彩内容第一时间为您奉上。 ]
LLaDA-UI: Bringing Block-wise Diffusion to Vision-Language GUI Agents
Diffusion large language models (dLLMs) achieve high decoding efficiency through block-parallel, arbitrary-order generat…
The Semantic Bottleneck: Leveraging Semantic Representations for Non-Invasive Speech Decoding
Non-invasive speech decoding remains constrained by the low signal-to-noise ratio of neural recordings, which makes fine…
NVIDIA 创业企业展示半程活动10月苏州举行,30家科创企业路演+投资人对接]
NVIDIA 2026 年度创业企业展示活动的半程展示将于 10 月 15 日在苏州金鸡湖国际会议中心 A 馆举办。 半程展示将聚焦 AI 智能体、物理 AI、开放模型及 AI 应用等前沿方向。届时,30 余家 NVIDIA 初创加速计划会…
曹操出行与豆包联合推出 AI 打车服务,能一句话选车、设置空调温度等
IT之家 9 月 9 日消息,曹操出行今日宣布与豆包联合打造 AI 打车服务 ,首批在北京、杭州、苏州三城上线,这意味着曹操出行将出行服务能力与 AI Agent 深度融合。 AI 打车服务上线后,当用户在豆包产品端咨询路线、地点、休闲娱乐…
Procedural Graphs: Self-Evolving Execution Structures for LLM Agents
Large language models are increasingly deployed as agents that plan over long horizons and act through external tools. M…
Omni Interaction Agent Technical Report
In this work, we present Gander, an end-to-end model that unifies omni perception, realtime interaction, and agentic cap…
NeoHorse-1: Towards Recursive Self-Improvement via Agentic Post-Training with Routing Harness
Recursive self-improvement (RSI) requires a concrete mechanism through which an AI system observes its capabilities and …
Marigold V2: Revisiting Diffusion Transformers for Monocular Depth Estimation
Monocular depth estimation is a ubiquitous yet highly ill-posed computer vision task, with downstream applications in sc…
Co-Evolving Harnesses and Models: On-Policy Correction Helps Weaker Models Catch Up Where Imitation Fails
Agent harnesses (the system prompt, tool set, execution hooks, and context-management scaffolding around a model) are a …