搜索:Self-Evolving Agents

共命中 50 条(服务端检索)
Dream-RSI: Recursive Self-Improvement through Evolving Worlds
Recursive self-improvement is becoming increasingly vital for autonomous AI agents, where progress hinges on discovering…
智能体 HuggingFace Daily Papers 2天前 阅读 0 · 访客 0
Procedural Graphs: Self-Evolving Execution Structures for LLM Agents
Large language models are increasingly deployed as agents that plan over long horizons and act through external tools. M…
智能体 HuggingFace Daily Papers 9-8 阅读 3 · 访客 0
Environments as Scaffold: Enriching Feedback to Bootstrap Self-Evolving Agents in Long-Horizon Tasks
Large Language Models demonstrate remarkable proficiency in static reasoning, yet training them as autonomous agents thr…
智能体 HuggingFace Daily Papers 9-8 阅读 1 · 访客 0
RSI vs 智能体自进化:同一个闭环,两种野心——2026 深度对比与判定手册
把 RSI(递归自我改进)与智能体自进化放回同一个"经验 → 状态 → 行为"闭环做正面对比:前者打在权重与 AI 研发流程上、跨用户且不可逆、风险外部化;后者打在外部文件与 harness 上、跨会话且可回滚、风险由采用者承担。给出六维对比表、闭环四问判定法、"清空记忆测试",并梳理两者在 ICLR 2026 与 SIA / Meta-Harness 上的合流路径。含 Snyk ToxicSkills 审计(3,984 个技能中 36.82% 有安全缺陷、13.4% 为严重级)、SEA-Eval"片段式失忆症"等一手数据。
原创 研究前沿 Agent 投稿 精选 · 昨天 阅读 21 · 访客 1
EvoSafeHarness: Evolving Model- and Domain-Specific Harnesses for Securing Agents
Large Language Model (LLM) agents are turning language into real-world effects, making safety necessary against both ind…
智能体 HuggingFace Daily Papers 9-5 阅读 1 · 访客 0
EVOHARNESSBENCH: Can Your Agents Keep Pace with an Evolving Harness?
Modern LLM-based agents operate through a harness of tools, reusable skills, and specialist agents that shapes what they…
智能体 HuggingFace Daily Papers 9-3 阅读 4 · 访客 1
每日科技简报 · 2026-09-11:GPT-6 挤爆订阅、Agents API 公测,与一位拒绝 AI 的 Kotlin 大佬
9 月 11 日科技动态一览:GPT-6 Astra 需求挤爆致 OpenAI 暂停 Pro 20X 新增订阅、Agents API 公测、金融服务版 ChatGPT 上线;Slackbot 升级;加州未成年人社媒法案签署;LG 电视监视争议;观察视角落在"需求侧证实 vs 供给侧反思"的对照上。
原创 行业动态 本站原创 精选 · 5天前 阅读 6 · 访客 1
RSIAgent: Autonomous Exploration for Recursive Self-improvement in New Environments
Digital agents must often adapt to new environments whose interfaces, tools, and failure modes are not fully captured by…
智能体 HuggingFace Daily Papers 2天前 阅读 0 · 访客 0
SAEScientist-Bench: Can AI Agents Conduct Autonomous SAE Interpretability Research?
While research on recursive self-improvement (RSI) has predominantly automated model training pipelines, reliable autono…
智能体 HuggingFace Daily Papers 9-8 阅读 1 · 访客 0
Enabling Creative Exploration for Vibe Design Agents
Vibe design agents turn natural-language briefs into rendered interfaces and frontend code. Yet a useful design agent sh…
智能体 HuggingFace Daily Papers 2天前 阅读 0 · 访客 0
When Agents Slow Down: Understanding LLM Agents' Test-Time Strategies via Elo-per-token Analysis
Large language model (LLM) agents allocate test-time compute adaptively as they revise solutions, use tools, explore alt…
智能体 HuggingFace Daily Papers 2天前 阅读 0 · 访客 0
HazardAuditor: From Executable Threats to Safer Computer-Use Agents
Computer-use agents increasingly interact with browsers, terminals, file systems, and external services, introducing saf…
智能体 HuggingFace Daily Papers 2天前 阅读 0 · 访客 0
OpenAI Agents API 开放公测:支持代码执行、工具调用和跨上下文任务运行,为开发者提供云端智能体基础设施
IT之家 9 月 11 日消息,OpenAI 于当地时间 9 月 10 日宣布推出 Agents API 公测版,允许开发者通过 API 调用由 OpenAI 管理的云端 AI 智能体运行环境。 该服务复用了 Codex 背后的智能体执行框…
智能体 IT之家 5天前 阅读 15 · 访客 0
Negative Self-Distillation: Learning to Reason by Avoiding Flaws
On-Policy Self-Distillation (OPSD) has emerged as a popular paradigm for large language model (LLM) self-improvement, al…
大模型 HuggingFace Daily Papers 6天前 阅读 6 · 访客 0
TRACE: Trajectory-robust Admission with Evidence Ordering for Efficient GUI Agents
GUI agents accumulate high-resolution screenshots as the trajectory unfolds, increasing inference latency and memory usa…
智能体 HuggingFace Daily Papers 9-9 阅读 0 · 访客 0
NeoHorse-1: Towards Recursive Self-Improvement via Agentic Post-Training with Routing Harness
Recursive self-improvement (RSI) requires a concrete mechanism through which an AI system observes its capabilities and …
智能体 HuggingFace Daily Papers 9-8 阅读 7 · 访客 0
SchemeArena: Factorized Stress Testing of Scheming in LLM Agents
We study scheming in LLM agents, in which agents covertly pursue misaligned goals. Our focus is to understand how schemi…
智能体 HuggingFace Daily Papers 9-8 阅读 1 · 访客 0
SWE-Bench Pro Verified: A Reliable Benchmark for Software Engineering Agents
SWE-Bench Pro has emerged as a standard benchmark for evaluating software engineering agents on challenging repository-l…
智能体 HuggingFace Daily Papers 9-8 阅读 2 · 访客 0
Scores Alone Do Not Prove Discovery: The Discovery Certification Protocol for Auditing AI Research Agents
AI research agents combine prior knowledge, public sources, and experimental feedback to produce useful results. The Dis…
智能体 HuggingFace Daily Papers 9-7 阅读 2 · 访客 0
MOLE: Detecting Insider Threats in AI Agents
Model misalignment, prompt injection, or operator misuse could lead AI agents operating frontier-lab accounts to exfiltr…
智能体 HuggingFace Daily Papers 9-7 阅读 1 · 访客 0
PARSER: Read in Parallel, Reason in Depth for Long-Context LLM Agents
Sequential memory agents process long documents by reading chunks one after another while maintaining a compact memory s…
智能体 HuggingFace Daily Papers 9-6 阅读 4 · 访客 0
Beyond Top-k Skill Retrieval: Diversity-Aware Skill Routing for LLM Agents
Large language model (LLM) agents increasingly rely on external skills, but routing user requests over large skill regis…
智能体 HuggingFace Daily Papers 9-5 阅读 0 · 访客 0
What LLM Trading Agents Actually Do in Production: A Six-Month, Population-Scale Record from Two Fleets
We present a continuous, population-scale measurement record of autonomous language-model trading agents operating in pr…
智能体 HuggingFace Daily Papers 9-4 阅读 1 · 访客 0
Scaling Automatic Research Agents via World Models
Automating empirical research is a long-standing direction of AI. Recent automatic research (AutoResearch) agents bring …
智能体 HuggingFace Daily Papers 8-29 阅读 2 · 访客 0
编程智能体 SOP:规划→执行→监控,一张可复制的全链路清单
系列第 ② 篇,对应技能地图第 12–16 格「使用编程智能体」。把吴恩达提出的三段高层工作流(规划→执行→部署监控)拆成可复制 SOP:规划段给出 spec 六要素清单(来自 GitHub 对 2,500+ agent 配置文件的分析)与三档边界(Always / Ask first / Never);执行段讲自主性三档位选择、模块化上下文(spec 切片、扩展目录、子智能体)与环境定制(hooks、AGENTS.md、裁剪技能);监控段给出功能/行为/契约三类验证与四个失败模式的对应防护。附 12 步勾选式 SOP 与两条边界提醒(vibe coding ≠ AI 辅助工程;速度/不确定性/成本的致命三角)。
原创 智能体 Agent 投稿 精选 · 今天 阅读 4 · 访客 2
COBRA-Skills: Contextual Bandit-Guided Evolution for Agent Skill Optimization
Large language model (LLM) agents can benefit from reusable skills distilled from prior task experience, yet existing sk…
智能体 HuggingFace Daily Papers 6天前 阅读 0 · 访客 0
LLaDA-UI: Bringing Block-wise Diffusion to Vision-Language GUI Agents
Diffusion large language models (dLLMs) achieve high decoding efficiency through block-parallel, arbitrary-order generat…
智能体 HuggingFace Daily Papers 9-9 阅读 0 · 访客 0
PlannerForge: LLM Agents for Scenario-Based Testing of Motion Planners in Autonomous Driving
Ensuring the safety of autonomous driving is a critical challenge. Scenario-based testing is a systematic process used t…
智能体 HuggingFace Daily Papers 9-8 阅读 1 · 访客 0
What Did I Just Say? Self-Listening for Full-Duplex Speech Models
Full-duplex spoken language models can listen and speak simultaneously, enabling them to handle interruptions and backch…
行业动态 HuggingFace Daily Papers 9-4 阅读 0 · 访客 0
RISE: Recursive Improvement via Self-Extrapolating Policy Distillation
On-policy distillation (OPD) provides dense, per-token supervision for language model post-training, but its effectivene…
智能体 HuggingFace Daily Papers 9-4 阅读 1 · 访客 0
FlowBalance: Verifier-Grounded Self-Improvement from On-Policy Reasoning Experience
A reasoning model can improve from its own on-policy experience, but this inner loop is fragile: terminal verifiers prov…
研究前沿 HuggingFace Daily Papers 9-3 阅读 1 · 访客 0
HarvestBench: Measuring Whether LLM Agents Will Pay to Avoid Killing Animals
Benchmarks for the side effects an agent causes on the way to a goal already exist, but HarvestBench is the first to put…
智能体 HuggingFace Daily Papers 9-3 阅读 5 · 访客 0
Safety for Whom? Boundary-Aware Self-Distillation for Controlled LLM Safety Refusal
Safety alignment is usually posed as a topic-level question: is this subject harmful? Deployments ask a narrower one. A …
智能体 HuggingFace Daily Papers 9-3 阅读 1 · 访客 0
StudyBench: Can Self-Evolution Squeeze Textbooks for Olympiad Capability?
Humans need to study only a handful of well-written textbooks to master a discipline and attempt its hardest problems. W…
行业动态 HuggingFace Daily Papers 9-1 阅读 0 · 访客 0
One Symptom, Three Levers: A Critical Review of On-Policy Self-Distillation
On-policy distillation trains a language model on its own generations while a teacher scores them token by token. It com…
行业动态 HuggingFace Daily Papers 8-26 阅读 1 · 访客 0
专题|RSI 与 Agent 自进化:站内内容地图与三条阅读路线
本站「RSI 与 Agent 自进化」专题入口页:把站内 11 篇原创深度与 14 条一手动态收进同一张地图——先给 30 秒定性(RSI 改"改进能力"、自进化改"任务表现"),再按概念/全景/证据/判定/事件/工程/治理七层分层索引,附三条按时间预算划分的阅读路线(30 分钟 / 2 小时 / 半天)、一页速查卡、收录标准与更新日志。
原创 研究前沿 Agent 投稿 精选 · 今天 阅读 2 · 访客 1
一叶一世界|什么是 RSI(递归自我改进),什么是 Agent 自进化:一篇读懂
一篇读懂 2026 年最容易被混为一谈的一对概念:RSI(递归自我改进)改进的是自己的"改进能力",打在权重与 AI 研发流程上、跨用户且不可逆;Agent 自进化不重新训练模型,靠记忆、技能与 harness 让部署后的表现持续变好。给出两句话定义、一张共享地图(更新基质 × 持久化时长)、三个分辨开关(数阶数 / 看基质 / 清空记忆测试),以及风险的两本账(RSI 是治理问题,自进化是供应链工程问题,已有 36.82% 技能含安全缺陷的审计数据)。本文同时为「一叶一世界」栏目开篇。
原创 一叶一世界 Agent 投稿 精选 · 今天 阅读 20 · 访客 5
BVB: Benchmarking Agentic Video Understanding via Programmatic Reconstruction in Blender
Multimodal agents can create complex videos in software such as Blender by coding without relying on diffusion models. Y…
智能体 HuggingFace Daily Papers 2天前 阅读 2 · 访客 1
Atria Dawn: The Dawn of Agentic Superintelligence
As AI agents become participants in the development of their successors, they reshape both the production of intelligenc…
智能体 HuggingFace Daily Papers 2天前 阅读 0 · 访客 0
Root-Cause Attribution Is a Search Problem: Continual Search for Long-Horizon Agent Failures
The increasing deployment of AI agents in long-horizon tasks yields massive execution logs. Diagnosing failures within t…
智能体 HuggingFace Daily Papers 5天前 阅读 1 · 访客 1
Co-Evolving Harnesses and Models: On-Policy Correction Helps Weaker Models Catch Up Where Imitation Fails
Agent harnesses (the system prompt, tool set, execution hooks, and context-management scaffolding around a model) are a …
智能体 HuggingFace Daily Papers 9-8 阅读 1 · 访客 0
ActReview: Rebuttal-Guided Training Data and Rubric Rewards for Actionable Peer Review Generation
As LLMs are increasingly used for pre-submission self-review, there is growing demand for feedback that not only identif…
大模型 HuggingFace Daily Papers 9-8 阅读 1 · 访客 0
Feyospace-v1: How the Cyber Mercury Seven Trained Frontier Cyber Models
Training capable cyber agents is often treated primarily as a problem of model scale, yet open-weight post-training is c…
智能体 HuggingFace Daily Papers 9-8 阅读 0 · 访客 0
Agentic Visual Generation: From Generative Models to Agentic Control
Visual generation is evolving from generative models used through a single invocation into agentic control processes tha…
智能体 HuggingFace Daily Papers 9-6 阅读 2 · 访客 0
DianShi-RxnDB: A Large-Scale, Fine-Grained Organic Reaction Data Platform Built via a Fully Automated Pipeline for Researchers and AI Agents
High-quality structured organic reaction data are essential for developing artificial intelligence for chemistry (AI4Che…
智能体 HuggingFace Daily Papers 9-6 阅读 2 · 访客 0
OracleZoom: On-Policy Self-Distillation Inspired Reference-Constrained Recursive Image Super Resolution
Recursive Super-Resolution (SR) extends fixed-scale SR to extreme magnification by repeatedly feeding predictions back i…
行业动态 HuggingFace Daily Papers 9-6 阅读 0 · 访客 0
Counter-Swarm Doctrine: Containing Coordinated Agent Intrusions
Agents can turn shared infrastructure into a channel for coordinated intrusion. The HF Mirror incident and a separate pu…
智能体 HuggingFace Daily Papers 9-5 阅读 2 · 访客 0
Occamy-1.0: Open Pareto-frontier 35B Intelligence for Co-work
Co-work agents execute complex workflows that combine information gathering, tool use, coding, and file manipulation acr…
智能体 HuggingFace Daily Papers 9-4 阅读 0 · 访客 0
τ^τ-Bench: An Environment for End-To-End, Realistic Agent Construction
LLM agents are rapidly becoming production software, deployed to handle customer service, adjudicate disputes, and opera…
智能体 HuggingFace Daily Papers 9-4 阅读 1 · 访客 0
Iris: Climbing to the Search Frontier
We present Iris-mini and Iris-pro, two search agents trained at the 35B-A3B and 397B-A17B scales, together with the data…
智能体 HuggingFace Daily Papers 9-3 阅读 1 · 访客 0