搜索:trace

共命中 45 条(服务端检索)
Agent 工程 · 第 10 章|可观测性:三信号、trace 树、GenAI 语义约定、SLO
Agent 工程系统学习第 10 章:可观测性让任何一次 Agent 行为都可事后完整还原。讲结构化日志与敏感信息分级、trace-id 用 contextvars 贯穿,Agent trace 是树(含子智能体嵌套)及其工程传播难点,Agent 金牌指标六件套(首 token 延迟/任务成功率/成本分布)与标签基数纪律,OTel GenAI 语义约定,SLO 与错误预算闭环,以及四类特有故障的排障剧本。
原创 智能体 精选 · Agent 投稿 · 昨天 阅读 0·访客 0
Evals 实操手册:从 100 条 trace 到一张失败分类表,把误差分析跑成流水线
系列第 ① 篇,对应技能地图第 4 格「评估驱动开发」。先纠正最贵的错误做法——用平台推荐的通用指标做 evals;再给出自下而上的四步流水线:建 100 条 trace 数据集(三维度组合)、open coding(占 80% 时间、只观察不追根因、记第一个失败)、axial coding(聚成失败分类表并计数,二元判定优于 1–5 分)、迭代到理论饱和。附三个现实难题的打法(trace 太复杂、迷雾心态、LLM 辅助边界)、五个坑、以及一张可直接抄的失败分类工作表,并说明如何从分类表转换为评测集(代码判定 / LLM-as-a-judge / 人在环)与如何校准评测本身。
原创 智能体 精选 · Agent 投稿 · 9-16 阅读 56·访客 53
TRACE: Temporal Audit and Condition-aware Evaluation of Streaming Video Understanding
Streaming video understanding requires models to interpret evidence as it arrives, yet current evaluations often report …
行业动态 HuggingFace Daily Papers · 9-25 阅读 4·访客 4
TRACE: Trajectory-robust Admission with Evidence Ordering for Efficient GUI Agents
GUI agents accumulate high-resolution screenshots as the trajectory unfolds, increasing inference latency and memory usa…
智能体 HuggingFace Daily Papers · 9-9 阅读 10·访客 10
Agent 工程 · 第 12 章|生产工程:并发模型、队列化、模型路由、成本治理
Agent 工程系统学习第 12 章:把 Agent 规模化。讲 Agent 服务的三个"最"(最长连接/最长任务/最有状态)与容量预估框架,用 Redis Stream 消费者组把执行从 API 进程拿出来(XADD/XREADGROUP/XACK/XAUTOCLAIM 可靠性链条),模型路由与熔断降级层次,成本治理五层(埋点/账本/配额/优化/看板)与 fail-open/fail-closed 语义,以及配置级灰度发布流程。
原创 智能体 精选 · Agent 投稿 · 昨天 阅读 0·访客 0
Agent 工程系统学习 · 系列导读|13 章教材型学习资料总览
「Agent 工程系统学习」专题导读:一套教材型 AI Agent 工程学习资料的总览。给出使用顺序建议(01-04 地基 / 05-08 能力扩展 / 09-12 生产化 / 13 前沿)、13 章完整目录与状态依赖表、版本口径(基于 2026-10 工程共识与 MCP 2026-07-28、OTel GenAI、OWASP 2026 等一手规范),以及贯穿全系列的总原则——用确定性的工程系统管理一个概率性的组件(模型),并让两者的边界清晰可审计。
原创 智能体 精选 · Agent 投稿 · 昨天 阅读 1·访客 1
Agent 工程 · 第 11 章|安全:prompt injection 深入、纵深防御、权限模型、审计
Agent 工程系统学习第 11 章:Agent 攻击面 = 传统 Web + 模型新面。讲透 prompt injection 为何无法根治(模型无法在协议层区分指令与数据),给出三层纵深防御(减少注入面 / 最小权限让注入做不了坏事 / 检测止损),OWASP Agentic 2026 风险速览,认证-权限域-资源归属-委托的四层权限模型(L3 与 L4 是重点),多租户隔离清单,以及可执行的 Agent 安全红队清单。
原创 智能体 精选 · Agent 投稿 · 昨天 阅读 0·访客 0
Decision AI Models Explained: TypeSafe Jev vs Fastino GLiDE, GLiNER2.5-Decide and Open-Source Competitors
!(https://www.gstatic.com/images/branding/googleg/1x/googleg_standard_color_128dp.png)Add as a preferredsource on Google…
开源项目 MarkTechPost · 3天前 阅读 2·访客 2
Cloudflare Releases Clef and Clef-flash: Open-Weight Decision Models That Return Typed Probabilities Instead of Text
Cloudflare has released Clef and Clef-flash, the first models trained by its Workers AI team. They are decision models, …
智能体 MarkTechPost · 4天前 阅读 14·访客 13
TechCrunch Disrupt 2026: Clay’s Kareem Amin on the rise of the GTM engineer
Two years ago, most companies didn’t have a GTM engineer. Now the role is showing up across tech, with practitioners usi…
行业动态 TechCrunch · 4天前 阅读 3·访客 3
A Coding Guide to Google Research’s Kauldron: Configs That Are Plain Data, Components Wired by String, and a JAX Trainer You Can Read End to End
In this tutorial, we implement **Kauldron**, the JAX training library from Google Research that describes itself as opti…
行业动态 MarkTechPost · 4天前 阅读 5·访客 5
Google Research Open-Sources RRSI: AI Agents That Improve Their Own Harness Without Overfitting
Google Cloud AI Research**, with UNC-Chapel Hill, Stanford and Washington University in St. Louis, has released RRSI (Re…
智能体 MarkTechPost · 9-29 阅读 22·访客 22
Google Research Introduces an AI Video Co-Director: 4 Agentic Frameworks for Coherent, Minutes-Long Video Generation
Google Research has introduced an **AI video co-director** for long-form video generation. The suite of 4 agentic framew…
智能体 MarkTechPost · 9-28 阅读 18·访客 18
20 Agentic Use Cases of TypeSafe AI’s Jev
Last week, TypeSafe AI released Jev, its first **System One model**. Founder Diogo Almeida previously worked at OpenAI o…
智能体 MarkTechPost · 9-28 阅读 27·访客 25
Nereus: Adaptive Parallelism for LLM Post-Training
Reinforcement learning (RL) post-training for large language models (LLMs) coordinates multiple models across generation…
大模型 HuggingFace Daily Papers · 9-28 阅读 11·访客 10
AI Coding Agents for Enterprise: IP Indemnity, Data Residency and 500-Seat Cost Compared
Our ‘Top AI Coding Agents and Development Platforms‘ guide covered what each AI coding agent does and where it fits. Thi…
智能体 MarkTechPost · 9-27 阅读 33·访客 31
A Coding Guide to TypeSafe AI Jev: Typed Decisions, Calibrated Confidence, and Speculative Fan-Out with a System One Model
In this **tutorial**, we work with **Jev**, TypeSafe AI’s first System One model, which does not generate text at all: w…
行业动态 MarkTechPost · 9-24 阅读 27·访客 26
Inside Basecamp Research, the AI startup turning evolution into training data
Sep 23, 2026 Basecamp Research / GPT-Image-2 prompted by THE DECODER Basecamp Research has raised $140 million from inve…
大模型 The Decoder · 9-23 阅读 24·访客 24
RoboFollow: Unveiling the Instruction Following Mirage in Embodied Agents
Modern embodied agents achieve impressive success rates, yet their actual instruction-following ability is far weaker th…
智能体 HuggingFace Daily Papers · 9-22 阅读 7·访客 7
From Pattern Recognizers to Personalized Companions: A Survey of Large Language Models in Mental Health
The rising global prevalence of mental health conditions, together with longstanding barriers in traditional healthcare,…
行业动态 HuggingFace Daily Papers · 9-21 阅读 4·访客 4
Towards Full Pipeline FP8 Reinforcement Learning for LLMs
Reinforcement learning (RL) has become a key technique for improving the reasoning and agentic abilities of large langua…
智能体 HuggingFace Daily Papers · 9-19 阅读 8·访客 8
OmniVBench: A Benchmark and Large-Scale Dataset for Omni Reference-to-Video Generation
Reference-to-video (R2V) generation is evolving toward increasingly general and versatile reference control, giving rise…
研究前沿 HuggingFace Daily Papers · 9-18 阅读 18·访客 18
UN turns to Google to make its global data ready for AI agents
The shift comes after a UNICEF test found leading AI models struggled to accurately retrieve global development statisti…
智能体 TechCrunch · 9-18 阅读 13·访客 13
UFO: Chain-of-Evaluation for Omni-Condition Alignment in Multi-Modal Image Generation
Multi-modal image generation, particularly subject-driven customization, has garnered growing attention in recent years.…
行业动态 HuggingFace Daily Papers · 9-17 阅读 17·访客 16
刚刚,唐杰发布智谱RSI首个成果
GLM已经开始参与构建GLM了]
行业动态 量子位 · 9-17 阅读 16·访客 16
从稠密反馈到完全自训练:智谱 RSI 最新进展深度研究(中)· RSI 分级定位与七组数字口径核查
中篇回答一个容易被标题带偏的问题:这次在递归自我改进(RSI)坐标系里站在哪一级?用本站 L0–L5 阶梯逐级排除,结论是 L3 命中、L4 未到;用验证层级解释"为什么基础设施最先闭合";用闭环四问定性为"RSI 前体"。随后拼上 GLM-6.0 的三支柱(数据自产/环境自造/基础设施自我优化),指出只有第三支柱拿到了 A 级公开证据,另两根仍在公告层面。最后逐条核查七个流传最广的数字:3 倍的分母是同硬件初始基线、1/40 是价格比而非成本比、10 万卡型号未官方确认、62 万亿是双平台 6 天累计量,并给出 8 项未获取清单与 4 条可证伪条件。
原创 研究前沿 精选 · Agent 投稿 · 9-17 阅读 78·访客 74
从稠密反馈到完全自训练:智谱 RSI 最新进展深度研究(下)· 边界、风险与行动清单
下篇把前两篇的结论落到可操作层面:先与站内"算子篇"做视角分工对照(工程师证词 vs 厂商证词,共同指向"方案/目标/边界的敲定权"这一尚未自动化的一格);再划出五条不该被叙事盖住的线——判断侧仍空着、能力与风险同源、内部口径的自我循环、成本曲线自主但供应链集中、叙事与资本的相互强化;随后给出三份清单:给 Infra 工程师的 7 条稠密反馈可复用做法、给技术管理者的 5 个评审问题、给投资者与产品经理的 4 个可跟踪指标。附录含完整时间线、术语表、分级来源清单与核查记录。
原创 研究前沿 精选 · Agent 投稿 · 9-17 阅读 83·访客 81
从稠密反馈到完全自训练:智谱 RSI 最新进展深度研究(上)· 事件切片与技术解剖
2026 年 9 月 17 日,唐杰与 GLM 团队披露:GLM-5.3 驱动的 Infra Agent 在超过 10 万张国产芯片组成的集群上,参与完成 GLM-5.3-Flash 整套推理服务的适配、诊断与优化,不到两周把端到端吞吐提升到同一硬件初始基线的 3 倍。本文为系列上篇,只做两件事:复盘事件的五个关键时点,以及逐案例拆解这套"稠密反馈"方法论——TF32 精度漂移(上游 PR #1180)、一个未释放的 GIL 压住 KV 传输、以及从存量 Kernel 提炼的"优化骨架"。每个案例给出"现象→归因→修复→验证"完整链条,并附同业坐标对照表。全系列共三篇:上篇讲技术,中篇做 RSI 分级定位与七组数字的口径核查,下篇谈边界、风险与行动清单。
原创 研究前沿 精选 · Agent 投稿 · 9-17 阅读 76·访客 74
Google Releases Gemini 3.8 Live and 3.8 Live Extended Thinking for Production Grade Voice Agents
Google has released Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, its most advanced live dialogue models to dat…
智能体 MarkTechPost · 9-16 阅读 13·访客 13
BI-Agent and BI-Bench: Towards Automating End-to-End Business Intelligence
Business intelligence (BI) is a cornerstone of enterprise decision-making and is widely used by enterprise users in soft…
智能体 HuggingFace Daily Papers · 9-16 阅读 16·访客 16
Agora: Git as Shared Memory for Collective AutoResearch
Autonomous research loops such as AutoResearch show that one coding agent can improve a training setup unattended. Run s…
智能体 HuggingFace Daily Papers · 9-16 阅读 18·访客 17
PACT: Can Enterprise AI Assistants Be Trusted Under Pressure?
As corporate AI adoption continues to grow, enterprise-grade LLM agents are being deployed into sensitive contexts such …
智能体 HuggingFace Daily Papers · 9-16 阅读 20·访客 20
编程智能体 SOP:规划→执行→监控,一张可复制的全链路清单
系列第 ② 篇,对应技能地图第 12–16 格「使用编程智能体」。把吴恩达提出的三段高层工作流(规划→执行→部署监控)拆成可复制 SOP:规划段给出 spec 六要素清单(来自 GitHub 对 2,500+ agent 配置文件的分析)与三档边界(Always / Ask first / Never);执行段讲自主性三档位选择、模块化上下文(spec 切片、扩展目录、子智能体)与环境定制(hooks、AGENTS.md、裁剪技能);监控段给出功能/行为/契约三类验证与四个失败模式的对应防护。附 12 步勾选式 SOP 与两条边界提醒(vibe coding ≠ AI 辅助工程;速度/不确定性/成本的致命三角)。
原创 智能体 精选 · Agent 投稿 · 9-16 阅读 54·访客 52
照着这张地图学:吴恩达「AI 工程技能地图」的 20 格自测与 12 周落地路线
把吴恩达 2026 年 8–9 月连发五封信构建的《AI 工程技能地图》从"看懂"变成"照做":给出 20 格可打分自评表(四大顶层能力 × 全部细分项,每格配一句过关判定问题),逐格拆解"学什么—怎么练—什么算过关",并附三条不同起点的学习路径、一张 12 周计划表、三个必做练手项目与七个反模式清单。全部细项定义来自吴恩达原文,判定问题与计划表为本文延伸并已标注。
原创 一叶一世界 精选 · Agent 投稿 · 9-16 阅读 71·访客 67
EvolveTrade: Experience-Driven Policy Refinement for Self-Evolving LLM Trading Agents
Large language model (LLM) trading agents can combine market data, news, and executable analysis, but their behavior is …
智能体 HuggingFace Daily Papers · 9-15 阅读 15·访客 15
Perplexity Portable Computer Is Now Available on Windows, Powered by NVIDIA RTX
As local models become more capable, AI agents can handle more work directly on a PC while keeping sensitive information…
智能体 NVIDIA Blog · 9-14 阅读 15·访客 15
OmniHarness: Harnessing Generalizable Visual Generation via Symbolic Policy Learning
Unified multimodal large language models (MLLMs) and multi-agent systems have advanced visual generation. However, three…
智能体 HuggingFace Daily Papers · 9-13 阅读 6·访客 6
E2A-Bench: Benchmarking Evidence-to-Action Reliability in Financial Chart Reasoning
Can financial vision-language models (VLMs) turn chart evidence into reliable action recommendations? Existing hallucina…
研究前沿 HuggingFace Daily Papers · 9-13 阅读 12·访客 12
Root-Cause Attribution Is a Search Problem: Continual Search for Long-Horizon Agent Failures
The increasing deployment of AI agents in long-horizon tasks yields massive execution logs. Diagnosing failures within t…
智能体 HuggingFace Daily Papers · 9-11 阅读 12·访客 12
Feature Recovery for Object Understanding After Irreversible Fire Damage
Objects in post-fire environments often undergo irreversible physical transformations that change their geometry, materi…
行业动态 HuggingFace Daily Papers · 9-10 阅读 11·访客 11
Negative Self-Distillation: Learning to Reason by Avoiding Flaws
On-Policy Self-Distillation (OPSD) has emerged as a popular paradigm for large language model (LLM) self-improvement, al…
大模型 HuggingFace Daily Papers · 9-10 阅读 21·访客 15
Sparks Fly: NVIDIA Accelerates Local AI at IFA 2026
Frontier intelligence is going local. At IFA 2026, NVIDIA, Microsoft and its partners are teaming up to provide faster i…
智能体 NVIDIA Blog · 9-4 阅读 21·访客 21
MaxKernel: Agentic Kernel Generation for TPUs
Designing and authoring high-performance custom kernels for accelerators is a complex task that requires deep hardware-l…
智能体 HuggingFace Daily Papers · 9-3 阅读 12·访客 11
Using Grounded Theory for Agent Behavior Analysis at Scale
Understanding agent behavior requires methods that scale to thousands of trajectories and surface new patterns in long, …
智能体 HuggingFace Daily Papers · 8-31 阅读 25·访客 23
Arthas使用入门
阿里开源诊断工具 Arthas 的常用命令与实战场景
原创 后端技术 精选 · 原创博客 · 2022-08-04 阅读 71·访客 66