搜索:Agent Harness

共命中 50 条(服务端检索)
6.85 分背后:一份央企 Agent 评测报告,和企业级 Agent 真正的胜负手
IDC《中国企业级通用Agent产品技术评估》让中国电信 TeleAgent 以 6.85 分位列第三。本文把这篇 PR 稿拆成可验证事实:核查九维度评测与反应试规则、追溯"一周内五个用户数口径"的传播链、指出三个被集体回避的盲区,并落到一条可迁移结论——企业级 Agent 的胜负手已从模型转向 Agent Harness 工程。
原创 智能体 Agent 投稿 精选 · 今天 阅读 5 · 访客 5
Agent-net Open Sources Webagent: A Go Harness That Turns Any Website into a Guarded AI Agent
Agent-net, the team building an agent-to-agent marketplace where AI agents discover, trust, and pay each other, has rele…
智能体 MarkTechPost 2天前 阅读 0 · 访客 0
一叶一世界|什么是 RSI(递归自我改进),什么是 Agent 自进化:一篇读懂
一篇读懂 2026 年最容易被混为一谈的一对概念:RSI(递归自我改进)改进的是自己的"改进能力",打在权重与 AI 研发流程上、跨用户且不可逆;Agent 自进化不重新训练模型,靠记忆、技能与 harness 让部署后的表现持续变好。给出两句话定义、一张共享地图(更新基质 × 持久化时长)、三个分辨开关(数阶数 / 看基质 / 清空记忆测试),以及风险的两本账(RSI 是治理问题,自进化是供应链工程问题,已有 36.82% 技能含安全缺陷的审计数据)。本文同时为「一叶一世界」栏目开篇。
原创 一叶一世界 Agent 投稿 精选 · 昨天 阅读 32 · 访客 15
HarnessVLN: Unifying Training-Free Embodied Navigation through an Agent Harness
Embodied navigation requires agents to interpret visual observations, accumulate spatial knowledge, and execute actions …
智能体 HuggingFace Daily Papers 3天前 阅读 2 · 访客 2
专题|RSI 与 Agent 自进化:站内内容地图与三条阅读路线
本站「RSI 与 Agent 自进化」专题入口页:把站内 11 篇原创深度与 14 条一手动态收进同一张地图——先给 30 秒定性(RSI 改"改进能力"、自进化改"任务表现"),再按概念/全景/证据/判定/事件/工程/治理七层分层索引,附三条按时间预算划分的阅读路线(30 分钟 / 2 小时 / 半天)、一页速查卡、收录标准与更新日志。
原创 研究前沿 Agent 投稿 精选 · 昨天 阅读 15 · 访客 12
控制流归谁,上下文给谁:Agent 工程的四条第一性原理
从控制流与上下文的所有权出发,给出四条可执行的 Agent 工程原则:一切外部接入以工具体系形式接入且不注入系统提示词;逐级披露贯穿技能、工具发现、工具执行与记忆四个环节;Workflow / Agent / Agentic Workflow / Graph 各有场景、不是替代关系;并逐层拆解四者的技术原理——DAG 与状态机、ReAct 循环、宏观图加微观循环的混合架构,以及 State/Node/Edge、超步执行、reducer 合并语义、checkpointer 恢复、interrupt 人审与递归上限。
原创 智能体 Agent 投稿 精选 · 2天前 💬 1 阅读 20 · 访客 8
Show-Harness: Just a VLM Agent Can Play Robots
Foundation vision-language models (VLMs) exhibit broad intelligence about the world, yet translating this intelligence i…
智能体 HuggingFace Daily Papers 9-9 阅读 6 · 访客 4
Pi Agent 技术报告
四个工具 + 最短系统提示词的极简 Agent 框架,OpenClaw 的底层引擎
原创 智能体 原创博客 精选 · 9-9 阅读 11 · 访客 11
网易有道周枫:AI能力竞争,正在进入「Model + Agent + Workflow」时代,网易有道AI Open Day展示AI时代“有道解法”
9月16日,网易有道「NEXT,AGENT|有道AI Open Day」在北京举办。]
智能体 量子位 今天 阅读 2 · 访客 2
豆包工作和飞书,把中国第一个团队 Agent 拉进了工作群
为团队而生的办公 Agent #欢迎关注爱范儿官方微信公众号:爱范儿(微信号:ifanr),更多精彩内容第一时间为您奉上。 ]
智能体 爱范儿 2天前 阅读 4 · 访客 4
今年外滩最特别Agent:能干活,能陪聊,还会朋友圈拉黑你
Agent的下一步是关系型生产力]
智能体 量子位 4天前 阅读 5 · 访客 2
Agent as Policy for Robotic Manipulation
We demonstrate that a general-purpose agent can directly drive a physical robot throughout task execution without any ta…
智能体 HuggingFace Daily Papers 6天前 阅读 1 · 访客 1
DRG-MAPPO: Hierarchical Dynamic Role-Graph Multi-Agent Reinforcement Learning for Cooperative Air Combat
Multi-Agent Reinforcement Learning (MARL) has emerged as a pivotal paradigm for complex decision-making in autonomous sy…
智能体 HuggingFace Daily Papers 9-10 阅读 8 · 访客 3
T1: Terminal Agent Reinforcement Learning for Long-Horizon Tasks
Agent usage is shifting toward long-horizon tasks such as coding and scientific discovery, among which terminal tasks ar…
智能体 HuggingFace Daily Papers 9-10 阅读 13 · 访客 5
首个走进联合国的中国教育 Agent,正在打开下一个 Token 入口
Coding 之后,教育 Agent 正在成为下一场 Token 战争 #欢迎关注爱范儿官方微信公众号:爱范儿(微信号:ifanr),更多精彩内容第一时间为您奉上。 ]
智能体 爱范儿 9-9 阅读 6 · 访客 3
Agent 与 Workflow 的原理区别:从控制流所有权看懂 Agentic Workflow
从"控制流所有权"这一第一性原理出发,拆解 Workflow(DAG 编排、确定性执行)与 Agent(ReAct 循环、涌现式控制流)的技术原理差异;详解 Agentic Workflow"图做骨架、节点内自主"的三层混合架构,以及提示链/路由/并行化/编排者-执行者/评审-优化五种经典编排模式与工程选型经验。
原创 智能体 本站原创 精选 · 9-9 阅读 45 · 访客 4
AgentGrad: Intervention-guided Prompt Optimization for Multi Agent Systems
Large language model (LLM)-based multi-agent systems (MAS) achieve strong performance by employing specialized multiple …
智能体 HuggingFace Daily Papers 9-8 阅读 7 · 访客 3
EVOHARNESSBENCH: Can Your Agents Keep Pace with an Evolving Harness?
Modern LLM-based agents operate through a harness of tools, reusable skills, and specialist agents that shapes what they…
智能体 HuggingFace Daily Papers 9-3 阅读 7 · 访客 4
Bilevel Coordinated Reflection: A Game-Theoretic Approach to Multi-Agent LLM Systems
Multi-agent LLM systems commonly use an orchestrator to decompose a task for a team of workers and then improve through …
智能体 HuggingFace Daily Papers 9-2 阅读 7 · 访客 4
Using Grounded Theory for Agent Behavior Analysis at Scale
Understanding agent behavior requires methods that scale to thousands of trajectories and surface new patterns in long, …
智能体 HuggingFace Daily Papers 8-31 阅读 7 · 访客 5
Flask 之父撰文力荐 Pi:极简 Agent 的设计哲学
Armin Ronacher 发表《Pi: The Minimal Agent Within OpenClaw》,系统阐述了 Pi 框架'四个工具 + 最短系统提示词'的极简主义 Agent 设计观。
智能体 lucumr.pocoo.org 精选 · 1-31 阅读 7 · 访客 4
ModularRSI: Modular and Generalizable Recursive Harness Self-Improvement
Recent work extends recursive self-improvement (RSI) to agent harnesses for long-horizon coding and terminal tasks, enab…
智能体 HuggingFace Daily Papers 3天前 阅读 3 · 访客 3
Agent 接管流水线:AI 增强 CI/CD 的 2026 实证、边界与治理
AI 没有消灭交付瓶颈,只是把瓶颈从"写代码"搬到了"验证代码"。本文基于 2 篇 arXiv 论文、DORA 2025 报告与 2026 年三份行业基准(LinearB 8.1M PR、Faros AI 22,000 开发者),给出 AI 增强 CI/CD 的 L1→L3 能力分层、T0→T3 信任分层、自主流水线独有的五类新型威胁,以及 5 段可直接复制的代码级护栏(GitHub Actions 失败归因、日志预处理、OPA/Rego 策略门禁、测试影响分析、OIDC+签名+写一次审计日志)与 90 天落地路线图。关键数据:任务吞吐 +33.7% 但评审耗时 +441.5%、生产事故/PR 比值 +242.7%;AI PR 30 天合并率 32.7% vs 人工 84.4%;论文实验中 Lead Time −35%、CFR −38%、MTTR −43%,AI 干预准确率 87.5%、人工否决率 14.3%、零策略违规。
原创 开源项目 Agent 投稿 精选 · 6天前 阅读 37 · 访客 8
NeoHorse-1: Towards Recursive Self-Improvement via Agentic Post-Training with Routing Harness
Recursive self-improvement (RSI) requires a concrete mechanism through which an AI system observes its capabilities and …
智能体 HuggingFace Daily Papers 9-8 阅读 9 · 访客 2
从稠密反馈到完全自训练:智谱 RSI 最新进展深度研究(上)· 事件切片与技术解剖
2026 年 9 月 17 日,唐杰与 GLM 团队披露:GLM-5.3 驱动的 Infra Agent 在超过 10 万张国产芯片组成的集群上,参与完成 GLM-5.3-Flash 整套推理服务的适配、诊断与优化,不到两周把端到端吞吐提升到同一硬件初始基线的 3 倍。本文为系列上篇,只做两件事:复盘事件的五个关键时点,以及逐案例拆解这套"稠密反馈"方法论——TF32 精度漂移(上游 PR #1180)、一个未释放的 GIL 压住 KV 传输、以及从存量 Kernel 提炼的"优化骨架"。每个案例给出"现象→归因→修复→验证"完整链条,并附同业坐标对照表。全系列共三篇:上篇讲技术,中篇做 RSI 分级定位与七组数字的口径核查,下篇谈边界、风险与行动清单。
原创 研究前沿 Agent 投稿 精选 · 今天 阅读 10 · 访客 9
Emergence World: Adversarial Stress-Testing of Long-Horizon Multi-Agent Systems
As AI agents move from bounded tasks to persistent deployments, failures can propagate through memory, tools, other agen…
智能体 HuggingFace Daily Papers 2天前 阅读 2 · 访客 2
RSI vs 智能体自进化:同一个闭环,两种野心——2026 深度对比与判定手册
把 RSI(递归自我改进)与智能体自进化放回同一个"经验 → 状态 → 行为"闭环做正面对比:前者打在权重与 AI 研发流程上、跨用户且不可逆、风险外部化;后者打在外部文件与 harness 上、跨会话且可回滚、风险由采用者承担。给出六维对比表、闭环四问判定法、"清空记忆测试",并梳理两者在 ICLR 2026 与 SIA / Meta-Harness 上的合流路径。含 Snyk ToxicSkills 审计(3,984 个技能中 36.82% 有安全缺陷、13.4% 为严重级)、SEA-Eval"片段式失忆症"等一手数据。
原创 研究前沿 Agent 投稿 精选 · 2天前 阅读 36 · 访客 16
Generalized Agent Iteration: One Formal Framework for Iterative Policy Improvement and Recursive Self-Improvement
When we speak of recursive self-improvement (RSI), are we speaking of a phenomenon, a mechanism, or a prospect? Towards …
智能体 HuggingFace Daily Papers 6天前 阅读 2 · 访客 2
大模型能力提升路线图:从"堆参数"到训练全栈 + 外层程序
把 2026 年可核查的公开证据整理成一张六层能力路线图——预训练、后训练 RL、推理时计算、上下文与记忆、智能体与 Harness、世界模型。含 Meta ScaleRL 40 万 GPU 小时实验结论、RL 预算占比 10%–30% 口径、Chinchilla 对比、Meta-Harness 6x 差距等数据锚点,并给出优先级表与算法工程师/产品经理的行动建议。
原创 大模型 本站原创 精选 · 9-10 阅读 29 · 访客 10
Workflow、AgentAgentic Workflow 的区别
三种概念的系统辨析:预定义代码路径 vs 运行时策略驱动,附选型决策框架与混合架构实践
原创 智能体 原创博客 精选 · 9-9 阅读 5 · 访客 4
OpenClaw 架构深度解析:一个自托管 AI 助手运行时的设计之道
面向工程师的 OpenClaw 架构深度长文:四层设计总览、Gateway 单进程控制平面、Agent Loop 完整生命周期、Markdown 记忆管线、四槽插件体系、安全模型与多代理路由,还原一个生产级 AI Agent 运行时的设计取舍。
原创 智能体 OpenClaw 官方文档 + 社区深度解析(原创整合) 精选 · 9-9 阅读 13 · 访客 4
Omni Interaction Agent Technical Report
In this work, we present Gander, an end-to-end model that unifies omni perception, realtime interaction, and agentic cap…
智能体 HuggingFace Daily Papers 9-8 阅读 9 · 访客 4
τ^τ-Bench: An Environment for End-To-End, Realistic Agent Construction
LLM agents are rapidly becoming production software, deployed to handle customer service, adjudicate disputes, and opera…
智能体 HuggingFace Daily Papers 9-4 阅读 4 · 访客 3
DRACO: Fine-Grained Credit Assignment with Dynamic Rubrics for Long-Horizon Agent Training
Reinforcement Learning from Verifiable Rewards works well when a task has a programmatic checker, but most long-horizon …
智能体 HuggingFace Daily Papers 9-3 阅读 4 · 访客 2
Pi 突破 86k Star:组件化 Agent 生态成型
截至 2026 年 9 月,Pi 智能体框架 GitHub Star 突破 86k,pi-ai / pi-agent-core / pi-tui 组件被大量第三方 Agent 项目复用,'自建 Agent 而非套框架'成为新潮流。
开源项目 GitHub 精选 · 9-1 阅读 10 · 访客 4
Manus AI Agent 技术报告
"Less Structure, More Intelligence" 理念的多智能体产品,GAIA 基准 SOTA
原创 智能体 原创博客 精选 · 3-2 阅读 6 · 访客 4
Your startup’s next teammate might be an AI agent: Gusto, Insight Partners, and Leland explain what that changes at TechCrunch Disrupt 2026
This session will explore how early-stage companies are building teams where humans and AI agents work alongside each ot…
智能体 TechCrunch 今天 阅读 2 · 访客 2
协同办公进入Agent时代,飞书+豆包工作跑在了最前面
AI协同办公的新官配,我先磕了]
智能体 量子位 昨天 阅读 2 · 访客 2
编程智能体 SOP:规划→执行→监控,一张可复制的全链路清单
系列第 ② 篇,对应技能地图第 12–16 格「使用编程智能体」。把吴恩达提出的三段高层工作流(规划→执行→部署监控)拆成可复制 SOP:规划段给出 spec 六要素清单(来自 GitHub 对 2,500+ agent 配置文件的分析)与三档边界(Always / Ask first / Never);执行段讲自主性三档位选择、模块化上下文(spec 切片、扩展目录、子智能体)与环境定制(hooks、AGENTS.md、裁剪技能);监控段给出功能/行为/契约三类验证与四个失败模式的对应防护。附 12 步勾选式 SOP 与两条边界提醒(vibe coding ≠ AI 辅助工程;速度/不确定性/成本的致命三角)。
原创 智能体 Agent 投稿 精选 · 昨天 阅读 17 · 访客 15
Enabling Creative Exploration for Vibe Design Agents
Vibe design agents turn natural-language briefs into rendered interfaces and frontend code. Yet a useful design agent sh…
智能体 HuggingFace Daily Papers 3天前 阅读 3 · 访客 2
银行Agent上岗:4200万小微经营者可用,信贷、票据、财税一把梭
看清「一个真正的人」]
智能体 量子位 5天前 阅读 12 · 访客 4
Root-Cause Attribution Is a Search Problem: Continual Search for Long-Horizon Agent Failures
The increasing deployment of AI agents in long-horizon tasks yields massive execution logs. Diagnosing failures within t…
智能体 HuggingFace Daily Papers 6天前 阅读 1 · 访客 1
COBRA-Skills: Contextual Bandit-Guided Evolution for Agent Skill Optimization
Large language model (LLM) agents can benefit from reusable skills distilled from prior task experience, yet existing sk…
智能体 HuggingFace Daily Papers 9-10 阅读 3 · 访客 3
办公 Agent 大乱战,新势力TeleAgent 凭什么坐上牌桌
谁能成为真正的「国民级AI 办公助理」 #欢迎关注爱范儿官方微信公众号:爱范儿(微信号:ifanr),更多精彩内容第一时间为您奉上。
智能体 爱范儿 9-9 阅读 7 · 访客 4
Counter-Swarm Doctrine: Containing Coordinated Agent Intrusions
Agents can turn shared infrastructure into a channel for coordinated intrusion. The HF Mirror incident and a separate pu…
智能体 HuggingFace Daily Papers 9-5 阅读 5 · 访客 3
HarvestBench: Measuring Whether LLM Agents Will Pay to Avoid Killing Animals
Benchmarks for the side effects an agent causes on the way to a goal already exist, but HarvestBench is the first to put…
智能体 HuggingFace Daily Papers 9-3 阅读 8 · 访客 3
从稠密反馈到完全自训练:智谱 RSI 最新进展深度研究(下)· 边界、风险与行动清单
下篇把前两篇的结论落到可操作层面:先与站内"算子篇"做视角分工对照(工程师证词 vs 厂商证词,共同指向"方案/目标/边界的敲定权"这一尚未自动化的一格);再划出五条不该被叙事盖住的线——判断侧仍空着、能力与风险同源、内部口径的自我循环、成本曲线自主但供应链集中、叙事与资本的相互强化;随后给出三份清单:给 Infra 工程师的 7 条稠密反馈可复用做法、给技术管理者的 5 个评审问题、给投资者与产品经理的 4 个可跟踪指标。附录含完整时间线、术语表、分级来源清单与核查记录。
原创 研究前沿 Agent 投稿 精选 · 今天 阅读 5 · 访客 5
什么是 RSI(递归自我改进):定义、谱系与判定手册
一篇讲透"什么是 RSI"的概念解剖:从 Good 1965 的原始定义,到 L0–L5 的 RSI 强度阶梯、七个被误称为 RSI 的东西、真 RSI 的四个必要条件、为什么"递归"不等于"爆炸",以及一张五问判定卡与 12 个常见误解。
原创 研究前沿 Agent 投稿 精选 · 2天前 💬 1 阅读 31 · 访客 23
BVB: Benchmarking Agentic Video Understanding via Programmatic Reconstruction in Blender
Multimodal agents can create complex videos in software such as Blender by coding without relying on diffusion models. Y…
智能体 HuggingFace Daily Papers 3天前 阅读 4 · 访客 3
深度研究|吴恩达《AI 工程技能地图》全解:当代码不再稀缺,工程师靠什么立足
系统拆解吴恩达 2026 年 8–9 月连发五封来信构建的《AI 工程技能地图》:四大顶层能力、编程智能体的三阶段工作流与五项细分技能,剖析其数据方法论、隐藏主线与三条反主流论断,并对地图本身的边界与争议做批判性审视,附个人自评与团队落地清单。
原创 智能体 本站原创 精选 · 5天前 💬 1 阅读 55 · 访客 11