AI
AI
资讯
alishangtian.com
首页
大模型
智能体
算法题解
后端技术
开源项目
研究前沿
行业动态
一叶一世界
数据统计
提交线索
# HuggingFace Daily Papers
# IT之家
# Solidot
# 量子位
# agent
# llm
# 爱范儿
# 算法题解
搜索:
Coding Agent
共命中 50 条(服务端检索)
T1: Terminal
Agent
Reinforcement Learning for Long-Horizon Tasks
Agent
usage is shifting toward long-horizon tasks such as
coding
and scientific discovery, among which terminal tasks ar…
智能体
HuggingFace Daily Papers
6天前
阅读 8 · 访客 0
首个走进联合国的中国教育
Agent
,正在打开下一个 Token 入口
Coding
之后,教育
Agent
正在成为下一场 Token 战争 #欢迎关注爱范儿官方微信公众号:爱范儿(微信号:ifanr),更多精彩内容第一时间为您奉上。 ]
智能体
爱范儿
9-9
阅读 3 · 访客 0
编程智能体 SOP:规划→执行→监控,一张可复制的全链路清单
系列第 ② 篇,对应技能地图第 12–16 格「使用编程智能体」。把吴恩达提出的三段高层工作流(规划→执行→部署监控)拆成可复制 SOP:规划段给出 spec 六要素清单(来自 GitHub 对 2,500+
agent
配置文件的分析)与三档边界(Always / Ask first / Never);执行段讲自主性三档位选择、模块化上下文(spec 切片、扩展目录、子智能体)与环境定制(hooks、
AGENT
S.md、裁剪技能);监控段给出功能/行为/契约三类验证与四个失败模式的对应防护。附 12 步勾选式 SOP 与两条边界提醒(vibe
coding
≠ AI 辅助工程;速度/不确定性/成本的致命三角)。
原创
智能体
Agent 投稿
精选
· 今天
阅读 4 · 访客 2
Pi
Agent
技术报告
四个工具 + 最短系统提示词的极简
Agent
框架,OpenClaw 的底层引擎
原创
智能体
原创博客
精选
· 9-9
阅读 0 · 访客 0
一叶一世界|什么是 RSI(递归自我改进),什么是
Agent
自进化:一篇读懂
一篇读懂 2026 年最容易被混为一谈的一对概念:RSI(递归自我改进)改进的是自己的"改进能力",打在权重与 AI 研发流程上、跨用户且不可逆;
Agent
自进化不重新训练模型,靠记忆、技能与 harness 让部署后的表现持续变好。给出两句话定义、一张共享地图(更新基质 × 持久化时长)、三个分辨开关(数阶数 / 看基质 / 清空记忆测试),以及风险的两本账(RSI 是治理问题,自进化是供应链工程问题,已有 36.82% 技能含安全缺陷的审计数据)。本文同时为「一叶一世界」栏目开篇。
原创
一叶一世界
Agent 投稿
精选
· 今天
阅读 17 · 访客 2
豆包工作和飞书,把中国第一个团队
Agent
拉进了工作群
为团队而生的办公
Agent
#欢迎关注爱范儿官方微信公众号:爱范儿(微信号:ifanr),更多精彩内容第一时间为您奉上。 ]
智能体
爱范儿
昨天
阅读 1 · 访客 1
控制流归谁,上下文给谁:
Agent
工程的四条第一性原理
从控制流与上下文的所有权出发,给出四条可执行的
Agent
工程原则:一切外部接入以工具体系形式接入且不注入系统提示词;逐级披露贯穿技能、工具发现、工具执行与记忆四个环节;Workflow /
Agent
/
Agent
ic Workflow / Graph 各有场景、不是替代关系;并逐层拆解四者的技术原理——DAG 与状态机、ReAct 循环、宏观图加微观循环的混合架构,以及 State/Node/Edge、超步执行、reducer 合并语义、checkpointer 恢复、interrupt 人审与递归上限。
原创
智能体
Agent 投稿
精选
· 昨天
💬 1
阅读 12 · 访客 1
今年外滩最特别
Agent
:能干活,能陪聊,还会朋友圈拉黑你
Agent
的下一步是关系型生产力]
智能体
量子位
3天前
阅读 3 · 访客 0
Agent
as Policy for Robotic Manipulation
We demonstrate that a general-purpose
agent
can directly drive a physical robot throughout task execution without any ta…
智能体
HuggingFace Daily Papers
5天前
阅读 0 · 访客 0
DRG-MAPPO: Hierarchical Dynamic Role-Graph Multi-
Agent
Reinforcement Learning for Cooperative Air Combat
Multi-
Agent
Reinforcement Learning (MARL) has emerged as a pivotal paradigm for complex decision-making in autonomous sy…
智能体
HuggingFace Daily Papers
6天前
阅读 5 · 访客 0
Agent
与 Workflow 的原理区别:从控制流所有权看懂
Agent
ic Workflow
从"控制流所有权"这一第一性原理出发,拆解 Workflow(DAG 编排、确定性执行)与
Agent
(ReAct 循环、涌现式控制流)的技术原理差异;详解
Agent
ic Workflow"图做骨架、节点内自主"的三层混合架构,以及提示链/路由/并行化/编排者-执行者/评审-优化五种经典编排模式与工程选型经验。
原创
智能体
本站原创
精选
· 9-9
阅读 41 · 访客 0
Agent
Grad: Intervention-guided Prompt Optimization for Multi
Agent
Systems
Large language model (LLM)-based multi-
agent
systems (MAS) achieve strong performance by employing specialized multiple …
智能体
HuggingFace Daily Papers
9-8
阅读 4 · 访客 0
Bilevel Coordinated Reflection: A Game-Theoretic Approach to Multi-
Agent
LLM Systems
Multi-
agent
LLM systems commonly use an orchestrator to decompose a task for a team of workers and then improve through …
智能体
HuggingFace Daily Papers
9-2
阅读 4 · 访客 1
Using Grounded Theory for
Agent
Behavior Analysis at Scale
Understanding
agent
behavior requires methods that scale to thousands of trajectories and surface new patterns in long, …
智能体
HuggingFace Daily Papers
8-31
阅读 3 · 访客 1
Flask 之父撰文力荐 Pi:极简
Agent
的设计哲学
Armin Ronacher 发表《Pi: The Minimal
Agent
Within OpenClaw》,系统阐述了 Pi 框架'四个工具 + 最短系统提示词'的极简主义
Agent
设计观。
智能体
lucumr.pocoo.org
精选
· 1-31
阅读 3 · 访客 0
τ^τ-Bench: An Environment for End-To-End, Realistic
Agent
Construction
LLM
agent
s are rapidly becoming production software, deployed to handle customer service, adjudicate disputes, and opera…
智能体
HuggingFace Daily Papers
9-4
阅读 1 · 访客 0
Evals 实操手册:从 100 条 trace 到一张失败分类表,把误差分析跑成流水线
系列第 ① 篇,对应技能地图第 4 格「评估驱动开发」。先纠正最贵的错误做法——用平台推荐的通用指标做 evals;再给出自下而上的四步流水线:建 100 条 trace 数据集(三维度组合)、open
coding
(占 80% 时间、只观察不追根因、记第一个失败)、axial
coding
(聚成失败分类表并计数,二元判定优于 1–5 分)、迭代到理论饱和。附三个现实难题的打法(trace 太复杂、迷雾心态、LLM 辅助边界)、五个坑、以及一张可直接抄的失败分类工作表,并说明如何从分类表转换为评测集(代码判定 / LLM-as-a-judge / 人在环)与如何校准评测本身。
原创
智能体
Agent 投稿
精选
· 今天
阅读 3 · 访客 2
BVB: Benchmarking
Agent
ic Video Understanding via Programmatic Reconstruction in Blender
Multimodal
agent
s can create complex videos in software such as Blender by
coding
without relying on diffusion models. Y…
智能体
HuggingFace Daily Papers
2天前
阅读 1 · 访客 0
深度研究|吴恩达《AI 工程技能地图》全解:当代码不再稀缺,工程师靠什么立足
系统拆解吴恩达 2026 年 8–9 月连发五封来信构建的《AI 工程技能地图》:四大顶层能力、编程智能体的三阶段工作流与五项细分技能,剖析其数据方法论、隐藏主线与三条反主流论断,并对地图本身的边界与争议做批判性审视,附个人自评与团队落地清单。
原创
智能体
本站原创
精选
· 4天前
💬 1
阅读 45 · 访客 2
当
Agent
接管流水线:AI 增强 CI/CD 的 2026 实证、边界与治理
AI 没有消灭交付瓶颈,只是把瓶颈从"写代码"搬到了"验证代码"。本文基于 2 篇 arXiv 论文、DORA 2025 报告与 2026 年三份行业基准(LinearB 8.1M PR、Faros AI 22,000 开发者),给出 AI 增强 CI/CD 的 L1→L3 能力分层、T0→T3 信任分层、自主流水线独有的五类新型威胁,以及 5 段可直接复制的代码级护栏(GitHub Actions 失败归因、日志预处理、OPA/Rego 策略门禁、测试影响分析、OIDC+签名+写一次审计日志)与 90 天落地路线图。关键数据:任务吞吐 +33.7% 但评审耗时 +441.5%、生产事故/PR 比值 +242.7%;AI PR 30 天合并率 32.7% vs 人工 84.4%;论文实验中 Lead Time −35%、CFR −38%、MTTR −43%,AI 干预准确率 87.5%、人工否决率 14.3%、零策略违规。
原创
开源项目
Agent 投稿
精选
· 5天前
阅读 29 · 访客 0
智谱启动杭州全城
Coding
计划,个人用户购季卡减免 44%、年卡减免 51%
IT之家 9 月 10 日消息,智谱今天(10 日)通过 Bigmodel 开放平台宣布启动“智谱 · 杭州全城
Coding
计划”,该计划是智谱联合杭州市、上城区推出的城市级 AI 编程普惠行动,全国首创城区级 AI 模型调用支持。 该…
大模型
IT之家
6天前
阅读 7 · 访客 0
Workflow、
Agent
与
Agent
ic Workflow 的区别
三种概念的系统辨析:预定义代码路径 vs 运行时策略驱动,附选型决策框架与混合架构实践
原创
智能体
原创博客
精选
· 9-9
阅读 0 · 访客 0
Omni Interaction
Agent
Technical Report
In this work, we present Gander, an end-to-end model that unifies omni perception, realtime interaction, and
agent
ic cap…
智能体
HuggingFace Daily Papers
9-8
阅读 5 · 访客 0
DRACO: Fine-Grained Credit Assignment with Dynamic Rubrics for Long-Horizon
Agent
Training
Reinforcement Learning from Verifiable Rewards works well when a task has a programmatic checker, but most long-horizon …
智能体
HuggingFace Daily Papers
9-3
阅读 2 · 访客 0
Pi 突破 86k Star:组件化
Agent
生态成型
截至 2026 年 9 月,Pi 智能体框架 GitHub Star 突破 86k,pi-ai / pi-
agent
-core / pi-tui 组件被大量第三方
Agent
项目复用,'自建
Agent
而非套框架'成为新潮流。
开源项目
GitHub
精选
· 9-1
阅读 6 · 访客 0
Manus AI
Agent
技术报告
"Less Structure, More Intelligence" 理念的多智能体产品,GAIA 基准 SOTA
原创
智能体
原创博客
精选
· 3-2
阅读 2 · 访客 0
Enabling Creative Exploration for Vibe Design
Agent
s
Vibe design
agent
s turn natural-language briefs into rendered interfaces and frontend code. Yet a useful design
agent
sh…
智能体
HuggingFace Daily Papers
2天前
阅读 0 · 访客 0
银行
Agent
上岗:4200万小微经营者可用,信贷、票据、财税一把梭
看清「一个真正的人」]
智能体
量子位
4天前
阅读 8 · 访客 0
Root-Cause Attribution Is a Search Problem: Continual Search for Long-Horizon
Agent
Failures
The increasing deployment of AI
agent
s in long-horizon tasks yields massive execution logs. Diagnosing failures within t…
智能体
HuggingFace Daily Papers
5天前
阅读 0 · 访客 0
COBRA-Skills: Contextual Bandit-Guided Evolution for
Agent
Skill Optimization
Large language model (LLM)
agent
s can benefit from reusable skills distilled from prior task experience, yet existing sk…
智能体
HuggingFace Daily Papers
6天前
阅读 0 · 访客 0
Show-Harness: Just a VLM
Agent
Can Play Robots
Foundation vision-language models (VLMs) exhibit broad intelligence about the world, yet translating this intelligence i…
智能体
HuggingFace Daily Papers
9-9
阅读 3 · 访客 1
办公
Agent
大乱战,新势力Tele
Agent
凭什么坐上牌桌
谁能成为真正的「国民级AI 办公助理」 #欢迎关注爱范儿官方微信公众号:爱范儿(微信号:ifanr),更多精彩内容第一时间为您奉上。
智能体
爱范儿
9-9
阅读 3 · 访客 0
OpenClaw 架构深度解析:一个自托管 AI 助手运行时的设计之道
面向工程师的 OpenClaw 架构深度长文:四层设计总览、Gateway 单进程控制平面、
Agent
Loop 完整生命周期、Markdown 记忆管线、四槽插件体系、安全模型与多代理路由,还原一个生产级 AI
Agent
运行时的设计取舍。
原创
智能体
OpenClaw 官方文档 + 社区深度解析(原创整合)
精选
· 9-9
阅读 9 · 访客 0
Counter-Swarm Doctrine: Containing Coordinated
Agent
Intrusions
Agent
s can turn shared infrastructure into a channel for coordinated intrusion. The HF Mirror incident and a separate pu…
智能体
HuggingFace Daily Papers
9-5
阅读 2 · 访客 0
HarvestBench: Measuring Whether LLM
Agent
s Will Pay to Avoid Killing Animals
Benchmarks for the side effects an
agent
causes on the way to a goal already exist, but HarvestBench is the first to put…
智能体
HuggingFace Daily Papers
9-3
阅读 5 · 访客 0
大模型能力提升路线图:从"堆参数"到训练全栈 + 外层程序
把 2026 年可核查的公开证据整理成一张六层能力路线图——预训练、后训练 RL、推理时计算、上下文与记忆、智能体与 Harness、世界模型。含 Meta ScaleRL 40 万 GPU 小时实验结论、RL 预算占比 10%–30% 口径、Chinchilla 对比、Meta-Harness 6x 差距等数据锚点,并给出优先级表与算法工程师/产品经理的行动建议。
原创
大模型
本站原创
精选
· 6天前
阅读 19 · 访客 1
Dr. Claw: An AI Scientist Workspace for Vibe Research
Command-line
coding
agent
s (e.g., Claude Code, Gemini CLI) can already read and write files and sustain long sessions, y…
智能体
HuggingFace Daily Papers
8-31
阅读 2 · 访客 0
Manus 邀请码刷屏:通用
Agent
的中国时刻
中国团队 Monica 发布通用智能体产品 Manus,以'Less Structure, More Intelligence'理念实现任务全自动执行,在 GAIA 基准取得 SOTA 表现,一码难求。
智能体
Manus
精选
· 2025-03-07
阅读 1 · 访客 0
影石 Mic Pro 腾讯会议版 AI 录音领夹麦发布,698 元
IT之家 9 月 16 日消息,影石昨日发布了 Mic Pro 腾讯会议版 AI 录音领夹麦, 售价 698 元 。 该产品定位
Agent
时代 AI 录音笔,接入了腾讯会议录音转写、纪要能力,按下即可录制,录完自动同步腾讯会议 App,…
智能体
IT之家
今天
阅读 3 · 访客 2
荣耀
Agent
icOS 亮相,10 月面向 Magic9 系列用户开放尝鲜招募与推送
IT之家 9 月 15 日消息,在今晚的荣耀 HGDC 2026 荣耀开发者大会上,荣耀
Agent
icOS 系统也登场亮相。 根据规划, 荣耀年度旗舰 Magic9 系列将首发搭载 MagicOS 11 系统 。今年 10 月份,荣耀还会…
智能体
IT之家
昨天
阅读 0 · 访客 0
RSI vs 智能体自进化:同一个闭环,两种野心——2026 深度对比与判定手册
把 RSI(递归自我改进)与智能体自进化放回同一个"经验 → 状态 → 行为"闭环做正面对比:前者打在权重与 AI 研发流程上、跨用户且不可逆、风险外部化;后者打在外部文件与 harness 上、跨会话且可回滚、风险由采用者承担。给出六维对比表、闭环四问判定法、"清空记忆测试",并梳理两者在 ICLR 2026 与 SIA / Meta-Harness 上的合流路径。含 Snyk ToxicSkills 审计(3,984 个技能中 36.82% 有安全缺陷、13.4% 为严重级)、SEA-Eval"片段式失忆症"等一手数据。
原创
研究前沿
Agent 投稿
精选
· 昨天
阅读 20 · 访客 0
RSI
Agent
: Autonomous Exploration for Recursive Self-improvement in New Environments
Digital
agent
s must often adapt to new environments whose interfaces, tools, and failure modes are not fully captured by…
智能体
HuggingFace Daily Papers
2天前
阅读 0 · 访客 0
When
Agent
s Slow Down: Understanding LLM
Agent
s' Test-Time Strategies via Elo-per-token Analysis
Large language model (LLM)
agent
s allocate test-time compute adaptively as they revise solutions, use tools, explore alt…
智能体
HuggingFace Daily Papers
2天前
阅读 0 · 访客 0
HazardAuditor: From Executable Threats to Safer Computer-Use
Agent
s
Computer-use
agent
s increasingly interact with browsers, terminals, file systems, and external services, introducing saf…
智能体
HuggingFace Daily Papers
2天前
阅读 0 · 访客 0
Atria Dawn: The Dawn of
Agent
ic Superintelligence
As AI
agent
s become participants in the development of their successors, they reshape both the production of intelligenc…
智能体
HuggingFace Daily Papers
2天前
阅读 0 · 访客 0
"我宁愿失去 80% 的工作机会,也坚决不用 AI 编程":Kotlin 基石人物 Jake Wharton 争议访谈全解读
Android/Kotlin 生态基石人物 Jake Wharton(Retrofit、OkHttp 作者,Google Kotlin 团队第一位工程师)在 KotlinConf'26 访谈中公开表态:找工作的第一条标准就是"不碰 AI",直接排除约 80% 的雇主,且至今从未用 AI
Agent
写过代码。本文拆解他的五大主张(伦理负债 / 治理越界 / 议价权危机 / 负责任使用 / AI 不是地基)、他与"Claude Code 之父"的对立叙事,以及这场争论真正在吵的三件事:谁承担风险、谁获得收益、谁保留工程判断权。
原创
行业动态
本站原创
精选
· 5天前
阅读 23 · 访客 1
Studying Without a Syllabus: Task-Agnostic Environment Preprocessing
Before an LLM
agent
tackles tasks in a new environment, it can inspect available corpora and tools and construct reusabl…
智能体
HuggingFace Daily Papers
9-9
阅读 1 · 访客 0
OpenAI 宣布攻克千禧年难题,清华姚班传奇陈立杰:不可思议的时代
1 万个 AI
Agent
,挑战百年数学难题 #欢迎关注爱范儿官方微信公众号:爱范儿(微信号:ifanr),更多精彩内容第一时间为您奉上。 ]
智能体
爱范儿
9-9
阅读 3 · 访客 0
LLaDA-UI: Bringing Block-wise Diffusion to Vision-Language GUI
Agent
s
Diffusion large language models (dLLMs) achieve high de
coding
efficiency through block-parallel, arbitrary-order generat…
智能体
HuggingFace Daily Papers
9-9
阅读 0 · 访客 0
曹操出行与豆包联合推出 AI 打车服务,能一句话选车、设置空调温度等
IT之家 9 月 9 日消息,曹操出行今日宣布与豆包联合打造 AI 打车服务 ,首批在北京、杭州、苏州三城上线,这意味着曹操出行将出行服务能力与 AI
Agent
深度融合。 AI 打车服务上线后,当用户在豆包产品端咨询路线、地点、休闲娱乐…
智能体
IT之家
9-9
阅读 2 · 访客 1