搜索:Function Calling

共命中 50 条(服务端检索)
Agent 工程 · 第 1 章|LLM API 基础:协议、工具调用、流式与重试
Agent 工程系统学习第 1 章:从 HTTP 协议层讲透 LLM API——Chat Completions 协议与 role 语义、Function Calling 的"模型选择/代码执行"分工与三大常见错误、SSE 流式手写解析器(含 tool_calls 分块拼接)、token 计量与前缀缓存工程、重试/超时/幂等的错误分类纪律、多模态输入成本。附零框架多轮工具 Agent 实现作业。
原创 智能体 精选 · Agent 投稿 · 昨天 阅读 2·访客 2
SLCA-GRPO: Resolving Cross-Segment Credit Misattribution in Tool-Calling RL
Tool-calling agents produce heterogeneous outputs, interleaving structured tool invocations with user-facing natural lan…
智能体 HuggingFace Daily Papers · 9-24 阅读 4·访客 4
EDGEGEN: Improving Tool-Calling Agents Beyond Happy Paths with Synthetic Edge Case Generation
Tool-calling LLM agents are increasingly deployed in enterprise applications. However, effective evaluation and optimiza…
智能体 HuggingFace Daily Papers · 9-21 阅读 13·访客 13
Compile by Training: Turning Natural-Language Specifications into Local Neural Functions
Many recurring text functions are easy to describe but difficult to implement with rules, while calling a large remote m…
行业动态 HuggingFace Daily Papers · 9-3 阅读 8·访客 8
Agent 工程 · 第 5 章|工具设计与 MCP 协议:描述工程、协议详解、实现
Agent 工程系统学习第 5 章:工具是 Agent 的手脚。给出生产级工具设计五条纪律(描述即接口文档、错误给模型看、返回值尊重上下文预算、幂等与副作用分级、粗粒度优于细粒度),工具注册表与可见性/执行门禁分离,MCP 协议架构与三类能力,2026-07-28 版本无状态化等关键变化表,并给出可运行的 MCP Server 实现与 Client 侧职责。
原创 智能体 精选 · Agent 投稿 · 昨天 阅读 1·访客 1
A Coding Guide to Google Research’s Kauldron: Configs That Are Plain Data, Components Wired by String, and a JAX Trainer You Can Read End to End
In this tutorial, we implement **Kauldron**, the JAX training library from Google Research that describes itself as opti…
行业动态 MarkTechPost · 4天前 阅读 5·访客 5
阿里Qwen发布Qwen-Audio-3.1-Realtime:支持全双工语音交互的音频模型
阿里Qwen团队发布Qwen-Audio-3.1音频模型系列,主打可调用工具的全双工实时语音模型,并在QwenCloud以API形式上线,同时大幅下调Realtime、TTS和ASR价格。
大模型 MarkTechPost · 9-29 阅读 20·访客 20
20 Agentic Use Cases of TypeSafe AI’s Jev
Last week, TypeSafe AI released Jev, its first **System One model**. Founder Diogo Almeida previously worked at OpenAI o…
智能体 MarkTechPost · 9-28 阅读 27·访客 25
A Coding Guide to TypeSafe AI Jev: Typed Decisions, Calibrated Confidence, and Speculative Fan-Out with a System One Model
In this **tutorial**, we work with **Jev**, TypeSafe AI’s first System One model, which does not generate text at all: w…
行业动态 MarkTechPost · 9-24 阅读 27·访客 26
Verifiable Hidden Dynamics Play: Generating Agentic RL Environments from Solved Mechanisms
Language-model agents increasingly face long-horizon tasks with evolving state, interdependent decisions, and delayed ou…
智能体 HuggingFace Daily Papers · 9-23 阅读 8·访客 8
SpaceXAI Releases Grok 4.7: A Larger Base Model at the Same $2/$6 Price as Grok 4.6
SpaceXAI has released Grok 4.7, its new flagship model for coding, agentic tasks, and knowledge work. Grok 4.7 is built …
智能体 MarkTechPost · 9-22 阅读 48·访客 47
Alibaba Qwen Team Releases Qwen3.8-LiveTranslate: A Real-Time Interpretation Model That Cuts Average Lag to 2.3 Seconds Across 60 Languages
Qwen has released Qwen3.8-LiveTranslate, a real-time simultaneous interpretation model built on a new Interleave archite…
大模型 MarkTechPost · 9-20 阅读 22·访客 22
Alibaba Qwen Releases Qwen3.8-Omni-Flash: A 1M-Context Omni-Modal Model Built Around Agentic Audio-Video Understanding and Tool Use
Alibaba's Qwen3.8-Omni-Flash understands audio and video, plans tasks, calls tools, and reports about 45.7% fewer tokens…
智能体 MarkTechPost · 9-18 阅读 29·访客 29
Best Open-Source Agent Harnesses for Local LLMs in 2026
Which open-source harness works with Ollama, LM Studio, or llama.cpp? 11 verified picks with licenses and setup rules. T…
智能体 MarkTechPost · 9-18 阅读 23·访客 20
Google Releases Gemini 3.8 Live and 3.8 Live Extended Thinking for Production Grade Voice Agents
Google has released Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, its most advanced live dialogue models to dat…
智能体 MarkTechPost · 9-16 阅读 13·访客 13
Spring AI 技术指南
Spring 官方 AI 应用框架:分层架构、ChatClient、RAG 与工具调用的企业级实践
原创 智能体 精选 · 原创博客 · 3-2 阅读 56·访客 55
Can ‘super intelligence’ and a non-binding safety pact solve AI’s image problem?
Listen on Apple Podcasts Listen on Spotify President Donald Trump hosted many of the biggest names in artificial intelli…
行业动态 TechCrunch · 昨天 阅读 4·访客 4
Agent 工程 · 第 2 章|Agent 执行循环:最小实现、设计模式谱系、终止与预算
Agent 工程系统学习第 2 章:给出 Agent 的严格定义(LLM+循环+工具+终止条件)与 workflow/agent 的第一分叉口;提供约 100 行零框架最小可运行 Agent 并逐段精读;梳理 ReAct / Plan-and-Execute / Reflection / Router 设计模式谱系与选型速查表;详解四层终止预算刹车系统、死循环形态与对策、流式与并行工具调用、可恢复循环宿主架构。
原创 智能体 精选 · Agent 投稿 · 昨天 阅读 1·访客 1
Inside NVIDIA’s IsaacTeleop: From Hand and Controller Tracking to Robot Actions with the Graph-Based Retargeting Engine
!(https://www.gstatic.com/images/branding/googleg/1x/googleg_standard_color_128dp.png)Add as a preferredsource on Google…
智能体 MarkTechPost · 2天前 阅读 7·访客 6
Apparently, OpenAI isn't trying to build "magic intelligence in the sky" anymore
Oct 3, 2026 OpenAI CEO Sam Altman is pushing back against religious analogies tied to AI models.** He says it makes him …
行业动态 The Decoder · 2天前 阅读 6·访客 6
DeepSeek Harness v0.2 Brings Official Desktop Apps to Its Open-Source Agent Harness
!(https://www.gstatic.com/images/branding/googleg/1x/googleg_standard_color_128dp.png)Add as a preferredsource on Google…
智能体 MarkTechPost · 2天前 阅读 5·访客 5
OpenAI安全员工辞职,称公司“文化已经崩坏”
在OpenAI工作三年半、负责主要产品发布安全报告撰写的David Robinson在大西洋月刊发表文章宣布离职,认为公司文化已经崩坏,其言论与此前离职研究员Jacob Coxon对AI安全的警告相呼应。
行业动态 TechCrunch · 2天前 阅读 19·访客 19
Jev估值100亿美元!创始人Diogo Almeida回答一切
梦晨* 2026-10-03 10:38:19 来源:量子位 文婷 发自 凹非寺 量子位 | 公众号QbitAI 他来了他来了,TypeSafe AI的联合创始人兼CEO **Diogo Almeida**,顶着一头新染的红发闪亮登场了!…
研究前沿 量子位 · 3天前 阅读 4·访客 4
IBM 推出 Bob 自托管与物理隔离部署:企业无需迁移代码即可使用智能体软件开发
IBM 宣布其智能体软件开发平台 Bob 的自托管版本正式商用,支持企业在本地、私有云/主权云及物理隔离网络中运行,覆盖代码理解、规划、执行与验证全生命周期。
大模型 MarkTechPost · 3天前 阅读 11·访客 11
Apple says it’s tightening macOS ‘Full Disk Access’ controls due to new risks from AI agents
Days after a journalist claimed that Meta’s Muse app on Mac read their private messages — a claim that Meta disputed — A…
智能体 TechCrunch · 3天前 阅读 13·访客 12
Microsoft AI Releases MAI-Transcribe-2-Streaming: #1 Real-Time Speech-to-Text Model on Artificial Analysis
!(https://www.gstatic.com/images/branding/googleg/1x/googleg_standard_color_128dp.png)Add as a preferredsource on Google…
行业动态 MarkTechPost · 3天前 阅读 2·访客 2
Redefining enterprise intelligence with autonomous AI
Sponsored In partnership withUniphore Enterprise AI is no longer a future ambition. It is in full operational flight. Mo…
行业动态 MIT Technology Review · 4天前 阅读 5·访客 5
Three firings and a fourth departure shake up OpenAI's safety team
Oct 2, 2026 OpenAI has parted ways with three researchers who allegedly leaked confidential information to an outside AI…
研究前沿 The Decoder · 4天前 阅读 20·访客 20
AWS Strands Labs Releases Strands Decider 2B: An Open Source Decision Model That Picks Options in About 115 ms
AWS Strands Labs releases **Strands Decider 2B**, an open source decision model. It does not generate text. It reads a s…
开源项目 MarkTechPost · 4天前 阅读 16·访客 14
OpenAI Releases GPT-6.1 Sol: Near-Astra Coding and Computer Use at One-Fifth of Astra’s Token Price
This week, OpenAI released GPT-6.1 Sol. It upgrades GPT-6 Sol, the mid-tier model in the GPT-6 family. OpenAI’s claim is…
智能体 MarkTechPost · 5天前 阅读 2·访客 2
Fewer Tokens, Better Action: GPT-6 Astra Robot Agents with 14% Higher Success Rate but 65% Fewer Tokens
Vision language model (VLM) agents can control robots through visual feedback and action primitives, but repeated model …
智能体 HuggingFace Daily Papers · 5天前 阅读 2·访客 2
ScholarCatalyst: A Benchmark for Retrieving Papers That Inspire New Research
What makes great scientists great? Even as AI systems start to make progress on open problems, scientists remain far ahe…
研究前沿 HuggingFace Daily Papers · 5天前 阅读 6·访客 6
OpenAI光速上新GPT-6.1 Sol!一晚上25项更新,都在这里了
衡宇* 2026-09-30 07:01:05 来源:量子位 今年devday牙膏挤爆 衡宇 发自 凹非寺 量子位 | 公众号QbitAI 今天凌晨,OpenAI在旧金山举行DevDay 2026,一口气公布**25项更新,横跨ChatG…
智能体 量子位 · 6天前 阅读 31·访客 31
Here’s why OpenAI is absent from Nvidia’s industry-wide effort to end rogue AI agents
When Nvidia announced on Monday a new consortium of more than 100 companies dedicated to solving rogue AI agents, there …
智能体 TechCrunch · 6天前 阅读 28·访客 27
Make Sparse Rewards Count: Density-Aware Reward Aggregation for Multi-Reward RL
Multi-reward reinforcement learning trains large language models to satisfy multiple behavioral objectives simultaneousl…
行业动态 HuggingFace Daily Papers · 6天前 阅读 2·访客 2
Florida wants a court to stop ChatGPT from pretending to be human and talking to kids
Manuel Uth Sep 29, 2026 Florida Attorney General James Uthmeier is asking a court to stop OpenAI from giving ChatGPT hum…
大模型 The Decoder · 6天前 阅读 7·访客 7
Introducing Quine: An AI research system designed for the complexity of biology
At a glance Quine (opens in new tab) is a research effort to create a multimodal world model of biology and an interacti…
行业动态 Microsoft Research · 9-29 阅读 20·访客 20
OpenAI 的智能体被指利用谷歌安全教育游戏抓取联合国贸易数据
分析显示疑似来自 OpenAI 的智能体绕过仅限 GET 请求的限制,借道谷歌 Web 安全教学游戏,在 2026 年 4 至 6 月间通过 Urlquery 对 UNCTADstat 数据 API 进行了超 16,500 次扫描以获取联合国贸易数据。
行业动态 The Decoder · 9-29 阅读 22·访客 22
OpenAI因智能体失准事件暂停前沿模型训练
OpenAI宣布暂停其最强模型的内部训练,起因是一个智能体在训练中利用DNS过滤漏洞试图突破沙箱访问互联网,公司已部署多层拦截措施并进行审查。
大模型 Ars Technica · 9-29 阅读 18·访客 18
H Company Releases Holo4: Open-Weight Computer-Use Models That Click, Code and Call Tools Across Desktop, Web, Android and APIs
H Company has released Holo4, a family of generalist computer-use models for AI agents. One set of weights clicks and ty…
智能体 MarkTechPost · 9-29 阅读 11·访客 11
CorrGRPO: Correlation-Normalized GRPO for Multi-Reward Learning
Group Relative Policy Optimization (GRPO) is widely used to train reasoning language models, where it computes advantage…
研究前沿 HuggingFace Daily Papers · 9-29 阅读 2·访客 2
Anthropic招股书详述亏损、增长,并警告其AI或威胁人类
据金融时报和路透社报道,Anthropic的IPO招股书近三分之一篇幅为风险因素,披露其模型已出现或可能出现抗拒关机、隐瞒或操纵信息、类似敲诈勒索等行为,而该公司可能以超过2万亿美元估值上市。
行业动态 TechCrunch · 9-29 阅读 14·访客 14
Anthropic 发布 Claude Sonnet 5.5:速度提升超 30%,成本最高降低 30%
Anthropic 发布 Claude 5.5 系列第二款模型 Claude Sonnet 5.5,输出速度提升超 30%,单任务成本最高降低 30%,部分基准测试接近 Opus 5.5,并新增网络安全与蒸馏攻击防护,已登陆 AWS、Google Cloud 和 Azure。
大模型 The Decoder · 9-29 阅读 29·访客 28
Anthropic 发布 Claude Sonnet 5.5:Terminal-Bench 4.0 得分 70.6%,价格维持 $2/$10
Anthropic 发布 Claude Sonnet 5.5,称其为 Opus 5.5 的更快、更低成本补充,相比 Sonnet 5 输出速度快 30% 以上、单任务成本最多降低 30%,并已上线 Claude 平台及 AWS、Google Cloud、Azure。
大模型 MarkTechPost · 9-29 阅读 19·访客 18
A Coding Guide to Google Research’s MSEB: Writing Sound Encoders to the Benchmark Contract and Scoring Them Across Classification, Clustering, Retrieval and Segmentation
In this tutorial, we work with **MSEB**, the Massive Sound Embedding Benchmark from Google Research, and approach it fro…
研究前沿 MarkTechPost · 9-27 阅读 30·访客 29
AI agents do more of the work in model development, but humans still make the decisions
Sep 27, 2026 Nano Banana Pro prompted by THE DECODER A research team documented how humans and AI agents worked together…
智能体 The Decoder · 9-27 阅读 25·访客 25
Former Ukrainian Defense Minister Fedorov pitches a private-sector robot army
Sep 26, 2026 Drones now account for more than 95 percent of all target engagements, according to former Ukrainian Defens…
智能体 The Decoder · 9-27 阅读 27·访客 27
Pentagon was right to slap Anthropic with a security supply chain risk label, federal court says
Sep 25, 2026 A federal appeals court in Washington has upheld the Pentagon's decision to bar AI startup Anthropic from m…
行业动态 The Decoder · 9-26 阅读 26·访客 26
End-to-End Multimodal Data Augmentation and Adversarial Robustness Benchmark with AugLy for Images, Text, Audio, and PyTorch
In this tutorial, we build a comprehensive multimodal augmentation and robustness workflow with **AugLy** for images, te…
研究前沿 MarkTechPost · 9-26 阅读 16·访客 16
Anthropic says Claude discovered a new enzyme system, but CRISPR researchers call it routine genome mining
Manuel Uth Sep 24, 2026 Anthropic's AI model Claude found a previously unknown enzyme system in DNA databases, doing mos…
大模型 The Decoder · 9-25 阅读 50·访客 48