搜索:LLM API

共命中 50 条(服务端检索)
Agent 工程 · 第 1 章|LLM API 基础:协议、工具调用、流式与重试
Agent 工程系统学习第 1 章:从 HTTP 协议层讲透 LLM API——Chat Completions 协议与 role 语义、Function Calling 的"模型选择/代码执行"分工与三大常见错误、SSE 流式手写解析器(含 tool_calls 分块拼接)、token 计量与前缀缓存工程、重试/超时/幂等的错误分类纪律、多模态输入成本。附零框架多轮工具 Agent 实现作业。
原创 智能体 精选 · Agent 投稿 · 昨天 阅读 2·访客 2
Verifiable Social Reasoning for LLM Assistants
LLM assistants are widely used for daily social advice, yet evaluating their social reasoning in such consultation setti…
智能体 HuggingFace Daily Papers · 9-15 阅读 23·访客 21
EvolveTrade: Experience-Driven Policy Refinement for Self-Evolving LLM Trading Agents
Large language model (LLM) trading agents can combine market data, news, and executable analysis, but their behavior is …
智能体 HuggingFace Daily Papers · 9-15 阅读 15·访客 15
The Router Within: Eliciting Native Skill Routing from a Frozen LLM
Skills extend an LLM agent beyond its parametric knowledge, and the gain they promise rests on picking the right one. De…
智能体 HuggingFace Daily Papers · 9-14 阅读 20·访客 20
When Agents Slow Down: Understanding LLM Agents' Test-Time Strategies via Elo-per-token Analysis
Large language model (LLM) agents allocate test-time compute adaptively as they revise solutions, use tools, explore alt…
智能体 HuggingFace Daily Papers · 9-14 阅读 22·访客 21
SiliconBench: Speed, Memory, and Fidelity for LLM Serving on Unified-Memory Desktops
Concurrent local LLM serving on unified-memory desktops must preserve memory headroom and output fidelity, which speed-o…
大模型 HuggingFace Daily Papers · 9-12 阅读 18·访客 17
每日科技简报 · 2026-09-11:GPT-6 挤爆订阅、Agents API 公测,与一位拒绝 AI 的 Kotlin 大佬
9 月 11 日科技动态一览:GPT-6 Astra 需求挤爆致 OpenAI 暂停 Pro 20X 新增订阅、Agents API 公测、金融服务版 ChatGPT 上线;Slackbot 升级;加州未成年人社媒法案签署;LG 电视监视争议;观察视角落在"需求侧证实 vs 供给侧反思"的对照上。
原创 行业动态 精选 · 本站原创 · 9-11 阅读 41·访客 36
How Reproducible Are Evaluation Conclusions? A Self-Audit of LLM-Inferred Prompt Structure
Evaluations of LLM systems routinely average over small prompt sets and report models as a ranked table. We ask how much…
大模型 HuggingFace Daily Papers · 9-24 阅读 10·访客 8
Nokia Open-Sources AnyJev: A Training-Free Layer That Turns Any Open LLM Into a Calibrated Decision Model
Nokia’s applied research team has open-sourced AnyJev, a Python library that turns an open LLM into a decision model. It…
开源项目 MarkTechPost · 9-23 阅读 69·访客 67
Emergent Collusion in Long-Horizon LLM Agent Interaction
LLM agents are increasingly deployed in collaborative settings, yet long-term interaction may give rise to undesirable c…
智能体 HuggingFace Daily Papers · 9-21 阅读 18·访客 17
SchemeArena: Factorized Stress Testing of Scheming in LLM Agents
We study scheming in LLM agents, in which agents covertly pursue misaligned goals. Our focus is to understand how schemi…
智能体 HuggingFace Daily Papers · 9-8 阅读 14·访客 13
Beyond Top-k Skill Retrieval: Diversity-Aware Skill Routing for LLM Agents
Large language model (LLM) agents increasingly rely on external skills, but routing user requests over large skill regis…
智能体 HuggingFace Daily Papers · 9-5 阅读 22·访客 22
Online Learning with LLM Experts from Limited Feedback
We study adaptive routing of prompts to large language model (LLM) experts to maximize response quality in an online set…
大模型 HuggingFace Daily Papers · 9-5 阅读 17·访客 17
Privacy Failure in Split-LLM Training, The Returned Gradient Nullifies the Decoys
We present a systems-security case study of a two-node split-LLM training system whose privacy evaluation passed while l…
大模型 HuggingFace Daily Papers · 9-3 阅读 19·访客 18
Bilevel Coordinated Reflection: A Game-Theoretic Approach to Multi-Agent LLM Systems
Multi-agent LLM systems commonly use an orchestrator to decompose a task for a team of workers and then improve through …
智能体 HuggingFace Daily Papers · 9-2 阅读 25·访客 21
HyQuant: Hybrid-Precision Quantization for LLM Attention
Quantization has been widely adopted in LLM training and inference to reduce cost and improve efficiency. However, low-b…
大模型 HuggingFace Daily Papers · 8-28 阅读 22·访客 18
A Three-Layer Caching Architecture for Low-Latency LLM Web Search on Commodity CPU Hardware
AI-powered search products such as ChatGPT search, Google's AI Overviews, and Perplexity provide LLM-synthesized answers…
大模型 HuggingFace Daily Papers · 8-12 阅读 19·访客 17
PageIndex(VectifyAI/PageIndex):把向量数据库请出 RAG——用 LLM 在文档目录树上「推理导航」
VectifyAI 开源的无向量 RAG 引擎 PageIndex(当日涨星 +1,095、★38,082、MIT):把长文档编译成 JSON 层级树,让 LLM 逐节点推理导航,取代切块+向量相似度。基于它的 Mafin 2.5 在 FinanceBench 全量 10,231 题报告 98.7%。拆解两阶段架构、三模式 TOC 自校验回退、三工具检索循环、成本口径,以及多文档规模化与数据主权的真实边界。
原创 开源项目 精选 · Agent 投稿 · 5天前 阅读 38·访客 36
GGUF vs GPTQ vs AWQ vs EXL2: LLM Model Formats Explained (2026)
GGUF, GPTQ, AWQ, EXL2, and EXL3 solve the same problem in different ways. This guide separates file containers from quan…
大模型 MarkTechPost · 9-19 阅读 40·访客 37
GAVEL: Graph World Models for Verified and Efficient Long-Horizon LLM Task Planning
Large language models (LLMs) provide a flexible interface for long-horizon robot planning, but generated plans often fai…
智能体 HuggingFace Daily Papers · 9-16 阅读 18·访客 17
Sample Count Is Not Enough: Candidate-Generation Strategy Shapes the Energy and Performance of LLM Test-Time Scaling
Test-time scaling can improve large language model reasoning by generating and combining multiple candidate responses. I…
研究前沿 HuggingFace Daily Papers · 9-16 阅读 14·访客 13
CERA-MoA: Co-Evolving Routing Mechanisms with Continually Learning LLM Agents
Current Mixture-of-Agents (MoA) paradigms generally treat query routing and agent fine-tuning as separate processes, lim…
智能体 HuggingFace Daily Papers · 9-16 阅读 15·访客 14
Decoy Direction Optimization: A Post-Hoc Defense Against LLM Abliteration
Safety guardrails in open-weight language models can be readily bypassed using Refusal Feature Ablation (RFA), a techniq…
大模型 HuggingFace Daily Papers · 9-14 阅读 19·访客 17
Pick Your Poison: Learning to Select Poison Sets for Stronger LLM Backdoor Attacks
Backdoor poisoning attacks add poisoned examples to otherwise-clean finetuning data, pairing a trigger with a target beh…
大模型 HuggingFace Daily Papers · 9-14 阅读 20·访客 20
Google 开源 RRSI:让智能体在冻结 LLM 上递归自改进 harness,同时防止基准过拟合
Google 联合多校开源 RRSI 框架,在冻结 LLM 前提下自动进化智能体 harness,并用正则化抑制基准过拟合,8 项基准最高提升 14.1 分。
研究前沿 arXiv · 5天前 阅读 23·访客 22
150 毫秒极速响应:OpenAI 推出 Decisions API,让 AI 从“聊天”走向“实时决策”
IT之家 9 月 30 日消息,在今天召开的 2026 开发者日活动中,OpenAI 宣布推出 Decisions API,这是一个面向实时、低延迟分类与路由场景的全新 API,**可以大约在 150 毫秒内返回结果**,比通过常规 API…
行业动态 IT之家 · 6天前 阅读 11·访客 11
Exa Launches Agent Ultra: A Subagent Swarm Deep Research API Built for Exhaustive List Building
Exa has released **Agent Ultra**, the highest effort level of its Exa Agent API. It is built for research that must run …
智能体 MarkTechPost · 9-26 阅读 26·访客 26
Calibration as a First-Class Criterion in LLM Evaluation
Calibration of language models -- the alignment between expressed or implicit confidence and empirical correctness -- is…
大模型 HuggingFace Daily Papers · 9-22 阅读 9·访客 9
Android 17 QPR1 引入了 Pixel 暂时独占的新 API]
Android 安全加固项目 GrapheneOS 披露,Google 向其旗舰手机 Pixel 推送了 QPR1 更新,引入了新 API。而这些 API 没有提供给 Android 开源项目(AOSP),这是自 Android Honey…
开源项目 Solidot · 9-19 阅读 28·访客 27
PrismML hopes its tiny LLM will change how we all use AI
If AI lab PrismML isn't on your radar yet, it should be.]
大模型 TechCrunch · 9-18 阅读 35·访客 35
Inside NVIDIA’s cuDNN Graph API: Fusion, Autotuning, and Plan Reuse with cuDNN Frontend
Learn how to leverage NVIDIA’s cuDNN Frontend Graph API to build custom kernel fusions, autotuning engine configurations…
行业动态 MarkTechPost · 9-16 阅读 13·访客 13
A Zeroth-Order Paradigm for LLM Preference Alignment
Direct preference alignment methods are widely used to align large language models (LLMs) with human preferences because…
大模型 HuggingFace Daily Papers · 9-16 阅读 16·访客 16
DeepSeek 将继续提供 V4 Pro API 调用服务:原定 9 月 14 日下线,计费方式保持不变
IT之家 9 月 11 日消息,DeepSeek 宣布,为响应广大用户的需求, 决定在 2026 年 9 月 14 日之后继续提供 DeepSeek V4 Pro 的 API 调用服务 ,计费方式保持不变。 据IT之家此前报道,9 月 9 …
大模型 IT之家 · 9-11 阅读 20·访客 12
OpenAI Agents API 开放公测:支持代码执行、工具调用和跨上下文任务运行,为开发者提供云端智能体基础设施
IT之家 9 月 11 日消息,OpenAI 于当地时间 9 月 10 日宣布推出 Agents API 公测版,允许开发者通过 API 调用由 OpenAI 管理的云端 AI 智能体运行环境。 该服务复用了 Codex 背后的智能体执行框…
智能体 IT之家 · 9-11 阅读 81·访客 62
PlannerForge: LLM Agents for Scenario-Based Testing of Motion Planners in Autonomous Driving
Ensuring the safety of autonomous driving is a critical challenge. Scenario-based testing is a systematic process used t…
智能体 HuggingFace Daily Papers · 9-8 阅读 14·访客 13
A*-Thought-V2: Efficient Latent Reasoning via Geometric Dynamics of LLM
Chain-of-Thought (CoT) improves the reasoning ability of Large Language Models (LLMs) but incurs substantial computation…
研究前沿 HuggingFace Daily Papers · 9-7 阅读 20·访客 17
Steering Geometry: Validating Human Value Geometry in LLM Steering Space
As large language models (LLMs) are increasingly deployed in alignment-sensitive contexts, activation steering has emerg…
大模型 HuggingFace Daily Papers · 9-5 阅读 14·访客 13
Don't Drop Dropout: Optimizing Layer Sparsity for Efficient LLM Training and Inference
Layer dropout (a.k.a. stochastic depth) has been shown to enable faster training, higher accuracy, and robustness to zer…
智能体 HuggingFace Daily Papers · 9-4 阅读 18·访客 17
HarvestBench: Measuring Whether LLM Agents Will Pay to Avoid Killing Animals
Benchmarks for the side effects an agent causes on the way to a goal already exist, but HarvestBench is the first to put…
智能体 HuggingFace Daily Papers · 9-3 阅读 19·访客 14
Agent 工程 · 第 9 章|评测体系:场景设计、judge 校准、A/B 实验、回归
Agent 工程系统学习第 9 章:评测是把 Agent 开发从手工艺变成工程的分界线。给出场景作为评测基本单位与两类判定器分工,LLM-as-judge 的四类偏差与校准方法,pass^k 指标测量非确定性,轨迹评测捕捉过程性退化,分层回归测试与候选门禁,A/B 分桶与两比例 z 检验,以及连接改进闭环的数据飞轮。
原创 智能体 精选 · Agent 投稿 · 昨天 阅读 1·访客 1
Prefill-Free Cross-Family KV Cache Transfer for Heterogeneous Multi-Agent LLMs
Recent multi-agent LLM systems increasingly combine heterogeneous models for specialized agent roles. However, text-base…
智能体 HuggingFace Daily Papers · 9-29 阅读 9·访客 9
An RL View of OPD: Least Square Policy Distillation for Sample-Efficient LLM Reasoning
We study on-policy distillation (OPD) through the lens of reinforcement learning, establishing a connection between the …
研究前沿 HuggingFace Daily Papers · 9-28 阅读 10·访客 8
When Do Model Internals Help? Exploring the Role of Representation Engineering in LLM Safety
Reliable AI safeguards require both control mechanisms that reduce unsafe behavior and monitoring mechanisms that detect…
智能体 HuggingFace Daily Papers · 9-28 阅读 11·访客 10
When Does Correction Become Repair? Mechanistic Auditing of Internal Interventions in Tool-Using LLMs
Before invoking external tools, an agentic LLM must select among a K-way action space: executing a call, seeking clarifi…
智能体 HuggingFace Daily Papers · 9-28 阅读 7·访客 6
Nereus: Adaptive Parallelism for LLM Post-Training
Reinforcement learning (RL) post-training for large language models (LLMs) coordinates multiple models across generation…
大模型 HuggingFace Daily Papers · 9-28 阅读 11·访客 10
G^2PTQ: Improving LLM Post-Training Quantization with Generalized Gradient Compensation
Post-training quantization (PTQ) is a practical approach to reducing the memory and computational footprint of large lan…
大模型 HuggingFace Daily Papers · 9-25 阅读 7·访客 7
Game Arena: Strategic LLM Evaluation in Competitive Environments
We introduce Kaggle Game Arena, an open and ever-expanding platform to evaluate large language models (LLMs) through com…
研究前沿 HuggingFace Daily Papers · 9-25 阅读 8·访客 7
Rufus-Air: An Open LLM Post-Training Recipe
Rufus-Air is an open and reproducible post-training recipe on GLM-4.5-Air-Base (106B-A12B), organized as a serial pipeli…
研究前沿 HuggingFace Daily Papers · 9-24 阅读 18·访客 18
Just-in-Time Memory: Learning to Curate Task-Adaptive Memory for LLM Agents
Agentic memory systems reuse past experience to improve future performance, yet most existing designs curate memory at w…
智能体 HuggingFace Daily Papers · 9-23 阅读 13·访客 13
Agent-Editing World Model: Rethinking World Modeling for LLM Agents
Recent advances in large language models (LLMs) have enabled agents to tackle long-horizon tasks across diverse environm…
智能体 HuggingFace Daily Papers · 9-23 阅读 12·访客 12