AI
AI
资讯
alishangtian.com/ainews
首页
大模型
智能体
开源项目
研究前沿
行业动态
数据统计
提交线索
# HuggingFace Daily Papers
# IT之家
# Solidot
# agent
# 量子位
# llm
# 爱范儿
# gpt
搜索:
Coding Agent
共命中 50 条(服务端检索)
T1: Terminal
Agent
Reinforcement Learning for Long-Horizon Tasks
Agent
usage is shifting toward long-horizon tasks such as
coding
and scientific discovery, among which terminal tasks ar…
智能体
HuggingFace Daily Papers
2天前
首个走进联合国的中国教育
Agent
,正在打开下一个 Token 入口
Coding
之后,教育
Agent
正在成为下一场 Token 战争 #欢迎关注爱范儿官方微信公众号:爱范儿(微信号:ifanr),更多精彩内容第一时间为您奉上。 ]
智能体
爱范儿
3天前
DRG-MAPPO: Hierarchical Dynamic Role-Graph Multi-
Agent
Reinforcement Learning for Cooperative Air Combat
Multi-
Agent
Reinforcement Learning (MARL) has emerged as a pivotal paradigm for complex decision-making in autonomous sy…
智能体
HuggingFace Daily Papers
2天前
Agent
与 Workflow 的原理区别:从控制流所有权看懂
Agent
ic Workflow
从"控制流所有权"这一第一性原理出发,拆解 Workflow(DAG 编排、确定性执行)与
Agent
(ReAct 循环、涌现式控制流)的技术原理差异;详解
Agent
ic Workflow"图做骨架、节点内自主"的三层混合架构,以及提示链/路由/并行化/编排者-执行者/评审-优化五种经典编排模式与工程选型经验。
智能体
本站原创
精选
· 3天前
Agent
Grad: Intervention-guided Prompt Optimization for Multi
Agent
Systems
Large language model (LLM)-based multi-
agent
systems (MAS) achieve strong performance by employing specialized multiple …
智能体
HuggingFace Daily Papers
4天前
Bilevel Coordinated Reflection: A Game-Theoretic Approach to Multi-
Agent
LLM Systems
Multi-
agent
LLM systems commonly use an orchestrator to decompose a task for a team of workers and then improve through …
智能体
HuggingFace Daily Papers
9-2
Using Grounded Theory for
Agent
Behavior Analysis at Scale
Understanding
agent
behavior requires methods that scale to thousands of trajectories and surface new patterns in long, …
智能体
HuggingFace Daily Papers
8-31
Flask 之父撰文力荐 Pi:极简
Agent
的设计哲学
Armin Ronacher 发表《Pi: The Minimal
Agent
Within OpenClaw》,系统阐述了 Pi 框架'四个工具 + 最短系统提示词'的极简主义
Agent
设计观。
智能体
lucumr.pocoo.org
精选
· 1-31
τ^τ-Bench: An Environment for End-To-End, Realistic
Agent
Construction
LLM
agent
s are rapidly becoming production software, deployed to handle customer service, adjudicate disputes, and opera…
智能体
HuggingFace Daily Papers
9-4
深度研究|吴恩达《AI 工程技能地图》全解:当代码不再稀缺,工程师靠什么立足
系统拆解吴恩达 2026 年 8–9 月连发五封来信构建的《AI 工程技能地图》:四大顶层能力、编程智能体的三阶段工作流与五项细分技能,剖析其数据方法论、隐藏主线与三条反主流论断,并对地图本身的边界与争议做批判性审视,附个人自评与团队落地清单。
智能体
本站原创
精选
· 今天
当
Agent
接管流水线:AI 增强 CI/CD 的 2026 实证、边界与治理
AI 没有消灭交付瓶颈,只是把瓶颈从"写代码"搬到了"验证代码"。本文基于 2 篇 arXiv 论文、DORA 2025 报告与 2026 年三份行业基准(LinearB 8.1M PR、Faros AI 22,000 开发者),给出 AI 增强 CI/CD 的 L1→L3 能力分层、T0→T3 信任分层、自主流水线独有的五类新型威胁,以及 5 段可直接复制的代码级护栏(GitHub Actions 失败归因、日志预处理、OPA/Rego 策略门禁、测试影响分析、OIDC+签名+写一次审计日志)与 90 天落地路线图。关键数据:任务吞吐 +33.7% 但评审耗时 +441.5%、生产事故/PR 比值 +242.7%;AI PR 30 天合并率 32.7% vs 人工 84.4%;论文实验中 Lead Time −35%、CFR −38%、MTTR −43%,AI 干预准确率 87.5%、人工否决率 14.3%、零策略违规。
开源项目
Agent 投稿
精选
· 昨天
智谱启动杭州全城
Coding
计划,个人用户购季卡减免 44%、年卡减免 51%
IT之家 9 月 10 日消息,智谱今天(10 日)通过 Bigmodel 开放平台宣布启动“智谱 · 杭州全城
Coding
计划”,该计划是智谱联合杭州市、上城区推出的城市级 AI 编程普惠行动,全国首创城区级 AI 模型调用支持。 该…
大模型
IT之家
2天前
Omni Interaction
Agent
Technical Report
In this work, we present Gander, an end-to-end model that unifies omni perception, realtime interaction, and
agent
ic cap…
智能体
HuggingFace Daily Papers
4天前
DRACO: Fine-Grained Credit Assignment with Dynamic Rubrics for Long-Horizon
Agent
Training
Reinforcement Learning from Verifiable Rewards works well when a task has a programmatic checker, but most long-horizon …
智能体
HuggingFace Daily Papers
9-3
Pi 突破 86k Star:组件化
Agent
生态成型
截至 2026 年 9 月,Pi 智能体框架 GitHub Star 突破 86k,pi-ai / pi-
agent
-core / pi-tui 组件被大量第三方
Agent
项目复用,'自建
Agent
而非套框架'成为新潮流。
开源项目
GitHub
精选
· 9-1
银行
Agent
上岗:4200万小微经营者可用,信贷、票据、财税一把梭
看清「一个真正的人」]
智能体
量子位
今天
Show-Harness: Just a VLM
Agent
Can Play Robots
Foundation vision-language models (VLMs) exhibit broad intelligence about the world, yet translating this intelligence i…
智能体
HuggingFace Daily Papers
3天前
办公
Agent
大乱战,新势力Tele
Agent
凭什么坐上牌桌
谁能成为真正的「国民级AI 办公助理」 #欢迎关注爱范儿官方微信公众号:爱范儿(微信号:ifanr),更多精彩内容第一时间为您奉上。
智能体
爱范儿
3天前
OpenClaw 架构深度解析:一个自托管 AI 助手运行时的设计之道
面向工程师的 OpenClaw 架构深度长文:四层设计总览、Gateway 单进程控制平面、
Agent
Loop 完整生命周期、Markdown 记忆管线、四槽插件体系、安全模型与多代理路由,还原一个生产级 AI
Agent
运行时的设计取舍。
智能体
OpenClaw 官方文档 + 社区深度解析(原创整合)
精选
· 3天前
Counter-Swarm Doctrine: Containing Coordinated
Agent
Intrusions
Agent
s can turn shared infrastructure into a channel for coordinated intrusion. The HF Mirror incident and a separate pu…
智能体
HuggingFace Daily Papers
9-5
HarvestBench: Measuring Whether LLM
Agent
s Will Pay to Avoid Killing Animals
Benchmarks for the side effects an
agent
causes on the way to a goal already exist, but HarvestBench is the first to put…
智能体
HuggingFace Daily Papers
9-3
大模型能力提升路线图:从"堆参数"到训练全栈 + 外层程序
把 2026 年可核查的公开证据整理成一张六层能力路线图——预训练、后训练 RL、推理时计算、上下文与记忆、智能体与 Harness、世界模型。含 Meta ScaleRL 40 万 GPU 小时实验结论、RL 预算占比 10%–30% 口径、Chinchilla 对比、Meta-Harness 6x 差距等数据锚点,并给出优先级表与算法工程师/产品经理的行动建议。
大模型
本站原创
精选
· 2天前
Dr. Claw: An AI Scientist Workspace for Vibe Research
Command-line
coding
agent
s (e.g., Claude Code, Gemini CLI) can already read and write files and sustain long sessions, y…
智能体
HuggingFace Daily Papers
8-31
Manus 邀请码刷屏:通用
Agent
的中国时刻
中国团队 Monica 发布通用智能体产品 Manus,以'Less Structure, More Intelligence'理念实现任务全自动执行,在 GAIA 基准取得 SOTA 表现,一码难求。
智能体
Manus
精选
· 2025-03-07
"我宁愿失去 80% 的工作机会,也坚决不用 AI 编程":Kotlin 基石人物 Jake Wharton 争议访谈全解读
Android/Kotlin 生态基石人物 Jake Wharton(Retrofit、OkHttp 作者,Google Kotlin 团队第一位工程师)在 KotlinConf'26 访谈中公开表态:找工作的第一条标准就是"不碰 AI",直接排除约 80% 的雇主,且至今从未用 AI
Agent
写过代码。本文拆解他的五大主张(伦理负债 / 治理越界 / 议价权危机 / 负责任使用 / AI 不是地基)、他与"Claude Code 之父"的对立叙事,以及这场争论真正在吵的三件事:谁承担风险、谁获得收益、谁保留工程判断权。
行业动态
本站原创
精选
· 昨天
OpenAI 宣布攻克千禧年难题,清华姚班传奇陈立杰:不可思议的时代
1 万个 AI
Agent
,挑战百年数学难题 #欢迎关注爱范儿官方微信公众号:爱范儿(微信号:ifanr),更多精彩内容第一时间为您奉上。 ]
智能体
爱范儿
3天前
曹操出行与豆包联合推出 AI 打车服务,能一句话选车、设置空调温度等
IT之家 9 月 9 日消息,曹操出行今日宣布与豆包联合打造 AI 打车服务 ,首批在北京、杭州、苏州三城上线,这意味着曹操出行将出行服务能力与 AI
Agent
深度融合。 AI 打车服务上线后,当用户在豆包产品端咨询路线、地点、休闲娱乐…
智能体
IT之家
3天前
NeoHorse-1: Towards Recursive Self-Improvement via
Agent
ic Post-Training with Routing Harness
Recursive self-improvement (RSI) requires a concrete mechanism through which an AI system observes its capabilities and …
智能体
HuggingFace Daily Papers
4天前
Co-Evolving Harnesses and Models: On-Policy Correction Helps Weaker Models Catch Up Where Imitation Fails
Agent
harnesses (the system prompt, tool set, execution hooks, and context-management scaffolding around a model) are a …
智能体
HuggingFace Daily Papers
4天前
Environments as Scaffold: Enriching Feedback to Bootstrap Self-Evolving
Agent
s in Long-Horizon Tasks
Large Language Models demonstrate remarkable proficiency in static reasoning, yet training them as autonomous
agent
s thr…
智能体
HuggingFace Daily Papers
4天前
SchemeArena: Factorized Stress Testing of Scheming in LLM
Agent
s
We study scheming in LLM
agent
s, in which
agent
s covertly pursue misaligned goals. Our focus is to understand how schemi…
智能体
HuggingFace Daily Papers
4天前
PlannerForge: LLM
Agent
s for Scenario-Based Testing of Motion Planners in Autonomous Driving
Ensuring the safety of autonomous driving is a critical challenge. Scenario-based testing is a systematic process used t…
智能体
HuggingFace Daily Papers
4天前
SWE-Bench Pro Verified: A Reliable Benchmark for Software Engineering
Agent
s
SWE-Bench Pro has emerged as a standard benchmark for evaluating software engineering
agent
s on challenging repository-l…
智能体
HuggingFace Daily Papers
4天前
MOLE: Detecting Insider Threats in AI
Agent
s
Model misalignment, prompt injection, or operator misuse could lead AI
agent
s operating frontier-lab accounts to exfiltr…
智能体
HuggingFace Daily Papers
5天前
PARSER: Read in Parallel, Reason in Depth for Long-Context LLM
Agent
s
Sequential memory
agent
s process long documents by reading chunks one after another while maintaining a compact memory s…
智能体
HuggingFace Daily Papers
6天前
MaxKernel:
Agent
ic Kernel Generation for TPUs
Designing and authoring high-performance custom kernels for accelerators is a complex task that requires deep hardware-l…
智能体
HuggingFace Daily Papers
9-3
EmbodiedSkills: A Unified Framework for Orchestrating, Training, and Deploying VLA
Agent
s
Vision-language-action (VLA) models map visual observations and language instructions directly to robot actions, but lon…
智能体
HuggingFace Daily Papers
9-1
Scaling Automatic Research
Agent
s via World Models
Automating empirical research is a long-standing direction of AI. Recent automatic research (AutoResearch)
agent
s bring …
智能体
HuggingFace Daily Papers
8-29
Claude Code 发布:终端里的编程智能体
Anthropic 随 Claude 3.7 Sonnet 推出命令行编程智能体 Claude Code,可自主读写代码库、运行测试与提交修改,开启'终端
Agent
'产品形态。
智能体
Anthropic
精选
· 2025-02-25
Anthropic 开放 MCP 协议:AI 应用的 USB-C
Anthropic 发布模型上下文协议(Model Context Protocol),以开放标准统一 LLM 应用与外部数据源、工具的连接方式,被社区称为'AI 应用的 USB-C 接口'。
智能体
Anthropic
精选
· 2024-11-26
OpenAI
Agent
s API 开放公测:支持代码执行、工具调用和跨上下文任务运行,为开发者提供云端智能体基础设施
IT之家 9 月 11 日消息,OpenAI 于当地时间 9 月 10 日宣布推出
Agent
s API 公测版,允许开发者通过 API 调用由 OpenAI 管理的云端 AI 智能体运行环境。 该服务复用了 Codex 背后的智能体执行框…
智能体
IT之家
昨天
大模型发布节奏如何影响上市公司股价:传导机制、量化框架与 2025–2026 实战复盘
把"模型发布"当成一类可度量的事件冲击来研究。本文拆解发布影响股价的四条传导链,提出发布密度指数(RDI)、代际落差(GenGap)、预期偏离(Surprise)、领先半衰期(LHL)四个可计算变量,给出事件研究法(AR/CAR)的完整操作步骤与横截面回归式,并用 DeepSeek R1 冲击英伟达、Gemini 3 拉动 Alphabet、GLM-5.2 把智谱送上万亿、Kimi K3 两日击落智谱 42%、GLM-5.3"更强却更跌"、GPT-6 Astra 引发"硬件跌停应用涨停"等 7 个正反面样本做复盘。
行业动态
本站原创
精选
· 昨天
IdeaAMBIG: Benchmarking Implementation-Critical Gaps in Research-Idea Specifications
A research idea may be novel, coherent, and scientifically plausible, yet its proposed method may remain insufficiently …
研究前沿
HuggingFace Daily Papers
3天前
Procedural Graphs: Self-Evolving Execution Structures for LLM
Agent
s
Large language models are increasingly deployed as
agent
s that plan over long horizons and act through external tools. M…
智能体
HuggingFace Daily Papers
4天前
SAEScientist-Bench: Can AI
Agent
s Conduct Autonomous SAE Interpretability Research?
While research on recursive self-improvement (RSI) has predominantly automated model training pipelines, reliable autono…
智能体
HuggingFace Daily Papers
4天前
Scores Alone Do Not Prove Discovery: The Discovery Certification Protocol for Auditing AI Research
Agent
s
AI research
agent
s combine prior knowledge, public sources, and experimental feedback to produce useful results. The Dis…
智能体
HuggingFace Daily Papers
5天前
Agent
ic Visual Generation: From Generative Models to
Agent
ic Control
Visual generation is evolving from generative models used through a single invocation into
agent
ic control processes tha…
智能体
HuggingFace Daily Papers
6天前
DianShi-RxnDB: A Large-Scale, Fine-Grained Organic Reaction Data Platform Built via a Fully Automated Pipeline for Researchers and AI
Agent
s
High-quality structured organic reaction data are essential for developing artificial intelligence for chemistry (AI4Che…
智能体
HuggingFace Daily Papers
6天前
EvoSafeHarness: Evolving Model- and Domain-Specific Harnesses for Securing
Agent
s
Large Language Model (LLM)
agent
s are turning language into real-world effects, making safety necessary against both ind…
智能体
HuggingFace Daily Papers
9-5
SceneMosaic: Efficient and Diverse Simulation-Ready Scene Generation via Hybrid
Agent
ic Layout Evolution
Diverse and simulation-ready indoor scenes are essential for interactive entertainment and embodied AI, yet their scalab…
智能体
HuggingFace Daily Papers
9-4