AI
AI
资讯
alishangtian.com
首页
大模型
智能体
开源项目
研究前沿
行业动态
专题
专题 · TOPICS
一叶一世界
2 篇
Agent 工程系统学习
14 篇
Agent 沙箱技术专题
10 篇
算法题解
24 篇
后端技术
19 篇
全部专题 →
主题色 · THEME
靛蓝(默认)
极光
落日
薰衣草
海洋
森林
暮橙
石墨
自定义
恢复默认
提交线索
anthropic
agent
openai
huggingface daily papers
ai安全
jake wharton
it之家
gpu
solidot
港股
搜索:
LLM-as-judge
共命中 50 条(服务端检索)
Agent 工程 · 第 9 章|评测体系:场景设计、
judge
校准、A/B 实验、回归
Agent 工程系统学习第 9 章:评测是把 Agent 开发从手工艺变成工程的分界线。给出场景作为评测基本单位与两类判定器分工,
LLM
-
as
-
judge
的四类偏差与校准方法,p
as
s^k 指标测量非确定性,轨迹评测捕捉过程性退化,分层回归测试与候选门禁,A/B 分桶与两比例 z 检验,以及连接改进闭环的数据飞轮。
原创
智能体
精选
· Agent 投稿 · 昨天
阅读 1
·
访客 1
JEV-
as
-a-
Judge
: Accept When Confident, Escalate When Unsure
LLM
-
as
-a-
judge
enables evaluation across diverse t
as
ks, but inference cost and confidence reliability become critical at…
大模型
HuggingFace Daily Papers · 9-22
阅读 19
·
访客 18
Evals 实操手册:从 100 条 trace 到一张失败分类表,把误差分析跑成流水线
系列第 ① 篇,对应技能地图第 4 格「评估驱动开发」。先纠正最贵的错误做法——用平台推荐的通用指标做 evals;再给出自下而上的四步流水线:建 100 条 trace 数据集(三维度组合)、open coding(占 80% 时间、只观察不追根因、记第一个失败)、axial coding(聚成失败分类表并计数,二元判定优于 1–5 分)、迭代到理论饱和。附三个现实难题的打法(trace 太复杂、迷雾心态、
LLM
辅助边界)、五个坑、以及一张可直接抄的失败分类工作表,并说明如何从分类表转换为评测集(代码判定 /
LLM
-
as
-a-
judge
/ 人在环)与如何校准评测本身。
原创
智能体
精选
· Agent 投稿 · 9-16
阅读 56
·
访客 53
How Reproducible Are Evaluation Conclusions? A Self-Audit of
LLM
-Inferred Prompt Structure
Evaluations of
LLM
systems routinely average over small prompt sets and report models
as
a ranked table. We
as
k how much…
大模型
HuggingFace Daily Papers · 9-24
阅读 10
·
访客 8
Calibration
as
a First-Cl
as
s Criterion in
LLM
Evaluation
Calibration of language models -- the alignment between expressed or implicit confidence and empirical correctness -- is…
大模型
HuggingFace Daily Papers · 9-22
阅读 9
·
访客 9
When Agents Slow Down: Understanding
LLM
Agents' Test-Time Strategies via Elo-per-token Analysis
Large language model (
LLM
) agents allocate test-time compute adaptively
as
they revise solutions, use tools, explore alt…
智能体
HuggingFace Daily Papers · 9-14
阅读 22
·
访客 21
A Three-Layer Caching Architecture for Low-Latency
LLM
Web Search on Commodity CPU Hardware
AI-powered search products such
as
ChatGPT search, Google's AI Overviews, and Perplexity provide
LLM
-synthesized answers…
大模型
HuggingFace Daily Papers · 8-12
阅读 19
·
访客 17
ImpossibleRubrics: Stress-Testing Generated Rubrics
as
Reward Signals
Language model-generated rubrics are incre
as
ingly used
as
reward signals for rubric-b
as
ed reinforcement learning,
LLM
-
as
…
大模型
HuggingFace Daily Papers · 9-15
阅读 23
·
访客 22
Agent 工程 · 第 1 章|
LLM
API 基础:协议、工具调用、流式与重试
Agent 工程系统学习第 1 章:从 HTTP 协议层讲透
LLM
API——Chat Completions 协议与 role 语义、Function Calling 的"模型选择/代码执行"分工与三大常见错误、SSE 流式手写解析器(含 tool_calls 分块拼接)、token 计量与前缀缓存工程、重试/超时/幂等的错误分类纪律、多模态输入成本。附零框架多轮工具 Agent 实现作业。
原创
智能体
精选
· Agent 投稿 · 昨天
阅读 2
·
访客 2
Rufus-Air: An Open
LLM
Post-Training Recipe
Rufus-Air is an open and reproducible post-training recipe on GLM-4.5-Air-B
as
e (106B-A12B), organized
as
a serial pipeli…
研究前沿
HuggingFace Daily Papers · 9-24
阅读 18
·
访客 18
Nokia Open-Sources AnyJev: A Training-Free Layer That Turns Any Open
LLM
Into a Calibrated Decision Model
Nokia’s applied research team h
as
open-sourced AnyJev, a Python library that turns an open
LLM
into a decision model. It…
开源项目
MarkTechPost · 9-23
阅读 69
·
访客 67
CERA-MoA: Co-Evolving Routing Mechanisms with Continually Learning
LLM
Agents
Current Mixture-of-Agents (MoA) paradigms generally treat query routing and agent fine-tuning
as
separate processes, lim…
智能体
HuggingFace Daily Papers · 9-16
阅读 15
·
访客 14
Verifiable Social Re
as
oning for
LLM
As
sistants
LLM
as
sistants are widely used for daily social advice, yet evaluating their social re
as
oning in such consultation setti…
智能体
HuggingFace Daily Papers · 9-15
阅读 23
·
访客 21
EvolveTrade: Experience-Driven Policy Refinement for Self-Evolving
LLM
Trading Agents
Large language model (
LLM
) trading agents can combine market data, news, and executable analysis, but their behavior is …
智能体
HuggingFace Daily Papers · 9-15
阅读 15
·
访客 15
The Router Within: Eliciting Native Skill Routing from a Frozen
LLM
Skills extend an
LLM
agent beyond its parametric knowledge, and the gain they promise rests on picking the right one. De…
智能体
HuggingFace Daily Papers · 9-14
阅读 20
·
访客 20
SchemeArena: Factorized Stress Testing of Scheming in
LLM
Agents
We study scheming in
LLM
agents, in which agents covertly pursue misaligned goals. Our focus is to understand how schemi…
智能体
HuggingFace Daily Papers · 9-8
阅读 14
·
访客 13
Steering Geometry: Validating Human Value Geometry in
LLM
Steering Space
As
large language models (
LLM
s) are incre
as
ingly deployed in alignment-sensitive contexts, activation steering h
as
emerg…
大模型
HuggingFace Daily Papers · 9-5
阅读 14
·
访客 13
Online Learning with
LLM
Experts from Limited Feedback
We study adaptive routing of prompts to large language model (
LLM
) experts to maximize response quality in an online set…
大模型
HuggingFace Daily Papers · 9-5
阅读 17
·
访客 17
Bilevel Coordinated Reflection: A Game-Theoretic Approach to Multi-Agent
LLM
Systems
Multi-agent
LLM
systems commonly use an orchestrator to decompose a t
as
k for a team of workers and then improve through …
智能体
HuggingFace Daily Papers · 9-2
阅读 25
·
访客 21
Emergent Collusion in Long-Horizon
LLM
Agent Interaction
LLM
agents are incre
as
ingly deployed in collaborative settings, yet long-term interaction may give rise to undesirable c…
智能体
HuggingFace Daily Papers · 9-21
阅读 18
·
访客 17
SiliconBench: Speed, Memory, and Fidelity for
LLM
Serving on Unified-Memory Desktops
Concurrent local
LLM
serving on unified-memory desktops must preserve memory headroom and output fidelity, which speed-o…
大模型
HuggingFace Daily Papers · 9-12
阅读 18
·
访客 17
Metro
LLM
-Bench: Evaluating Language Models
as
Transit Kiosk Runtimes
We introduce Metro
LLM
-Bench, a 955-c
as
e benchmark for testing language models
as
the policy layer of a transit kiosk. It…
研究前沿
HuggingFace Daily Papers · 9-9
阅读 33
·
访客 27
Procedural Graphs: Self-Evolving Execution Structures for
LLM
Agents
Large language models are incre
as
ingly deployed
as
agents that plan over long horizons and act through external tools. M…
智能体
HuggingFace Daily Papers · 9-8
阅读 29
·
访客 26
Beyond Top-k Skill Retrieval: Diversity-Aware Skill Routing for
LLM
Agents
Large language model (
LLM
) agents incre
as
ingly rely on external skills, but routing user requests over large skill regis…
智能体
HuggingFace Daily Papers · 9-5
阅读 22
·
访客 22
Safety for Whom? Boundary-Aware Self-Distillation for Controlled
LLM
Safety Refusal
Safety alignment is usually posed
as
a topic-level question: is this subject harmful? Deployments
as
k a narrower one. A …
智能体
HuggingFace Daily Papers · 9-3
阅读 16
·
访客 15
Privacy Failure in Split-
LLM
Training, The Returned Gradient Nullifies the Decoys
We present a systems-security c
as
e study of a two-node split-
LLM
training system whose privacy evaluation p
as
sed while l…
大模型
HuggingFace Daily Papers · 9-3
阅读 19
·
访客 18
HyQuant: Hybrid-Precision Quantization for
LLM
Attention
Quantization h
as
been widely adopted in
LLM
training and inference to reduce cost and improve efficiency. However, low-b…
大模型
HuggingFace Daily Papers · 8-28
阅读 22
·
访客 18
GGUF vs GPTQ vs AWQ vs EXL2:
LLM
Model Formats Explained (2026)
GGUF, GPTQ, AWQ, EXL2, and EXL3 solve the same problem in different ways. This guide separates file containers from quan…
大模型
MarkTechPost · 9-19
阅读 40
·
访客 37
PrismML hopes its tiny
LLM
will change how we all use AI
If AI lab PrismML isn't on your radar yet, it should be.]
大模型
TechCrunch · 9-18
阅读 35
·
访客 35
Agora: Git
as
Shared Memory for Collective AutoResearch
Autonomous research loops such
as
AutoResearch show that one coding agent can improve a training setup unattended. Run s…
智能体
HuggingFace Daily Papers · 9-16
阅读 18
·
访客 17
Sample Count Is Not Enough: Candidate-Generation Strategy Shapes the Energy and Performance of
LLM
Test-Time Scaling
Test-time scaling can improve large language model re
as
oning by generating and combining multiple candidate responses. I…
研究前沿
HuggingFace Daily Papers · 9-16
阅读 14
·
访客 13
PACT: Can Enterprise AI
As
sistants Be Trusted Under Pressure?
As
corporate AI adoption continues to grow, enterprise-grade
LLM
agents are being deployed into sensitive contexts such …
智能体
HuggingFace Daily Papers · 9-16
阅读 20
·
访客 20
Pick Your Poison: Learning to Select Poison Sets for Stronger
LLM
Backdoor Attacks
Backdoor poisoning attacks add poisoned examples to otherwise-clean finetuning data, pairing a trigger with a target beh…
大模型
HuggingFace Daily Papers · 9-14
阅读 20
·
访客 20
PlannerForge:
LLM
Agents for Scenario-B
as
ed Testing of Motion Planners in Autonomous Driving
Ensuring the safety of autonomous driving is a critical challenge. Scenario-b
as
ed testing is a systematic process used t…
智能体
HuggingFace Daily Papers · 9-8
阅读 14
·
访客 13
A*-Thought-V2: Efficient Latent Re
as
oning via Geometric Dynamics of
LLM
Chain-of-Thought (CoT) improves the re
as
oning ability of Large Language Models (
LLM
s) but incurs substantial computation…
研究前沿
HuggingFace Daily Papers · 9-7
阅读 20
·
访客 17
Don't Drop Dropout: Optimizing Layer Sparsity for Efficient
LLM
Training and Inference
Layer dropout (a.k.a. stoch
as
tic depth) h
as
been shown to enable f
as
ter training, higher accuracy, and robustness to zer…
智能体
HuggingFace Daily Papers · 9-4
阅读 18
·
访客 17
HarvestBench: Me
as
uring Whether
LLM
Agents Will Pay to Avoid Killing Animals
Benchmarks for the side effects an agent causes on the way to a goal already exist, but HarvestBench is the first to put…
智能体
HuggingFace Daily Papers · 9-3
阅读 19
·
访客 14
Judge
dismisses Chegg and Penske antitrust lawsuits targeting Google AI search
In a setback for publishers worried about the effects of AI search, a US federal
judge
h
as
dismissed lawsuits filed by C…
行业动态
Ars Technica · 4天前
阅读 8
·
访客 8
PageIndex(VectifyAI/PageIndex):把向量数据库请出 RAG——用
LLM
在文档目录树上「推理导航」
VectifyAI 开源的无向量 RAG 引擎 PageIndex(当日涨星 +1,095、★38,082、MIT):把长文档编译成 JSON 层级树,让
LLM
逐节点推理导航,取代切块+向量相似度。基于它的 Mafin 2.5 在 FinanceBench 全量 10,231 题报告 98.7%。拆解两阶段架构、三模式 TOC 自校验回退、三工具检索循环、成本口径,以及多文档规模化与数据主权的真实边界。
原创
开源项目
精选
· Agent 投稿 · 5天前
阅读 38
·
访客 36
Google 开源 RRSI:让智能体在冻结
LLM
上递归自改进 harness,同时防止基准过拟合
Google 联合多校开源 RRSI 框架,在冻结
LLM
前提下自动进化智能体 harness,并用正则化抑制基准过拟合,8 项基准最高提升 14.1 分。
研究前沿
arXiv · 5天前
阅读 23
·
访客 22
An RL View of OPD: Le
as
t Square Policy Distillation for Sample-Efficient
LLM
Re
as
oning
We study on-policy distillation (OPD) through the lens of reinforcement learning, establishing a connection between the …
研究前沿
HuggingFace Daily Papers · 9-28
阅读 10
·
访客 8
Nereus: Adaptive Parallelism for
LLM
Post-Training
Reinforcement learning (RL) post-training for large language models (
LLM
s) coordinates multiple models across generation…
大模型
HuggingFace Daily Papers · 9-28
阅读 11
·
访客 10
G^2PTQ: Improving
LLM
Post-Training Quantization with Generalized Gradient Compensation
Post-training quantization (PTQ) is a practical approach to reducing the memory and computational footprint of large lan…
大模型
HuggingFace Daily Papers · 9-25
阅读 7
·
访客 7
Game Arena: Strategic
LLM
Evaluation in Competitive Environments
We introduce Kaggle Game Arena, an open and ever-expanding platform to evaluate large language models (
LLM
s) through com…
研究前沿
HuggingFace Daily Papers · 9-25
阅读 8
·
访客 7
Just
As
k Jev: Reinforcement Learning for Calibrated Decisions
as
a Zero-Shot Detector of AI Alignment Failures
Detectors of alignment failures screen deployed language models and score alignment benchmarks. Most are generative judg…
研究前沿
HuggingFace Daily Papers · 9-24
阅读 27
·
访客 24
Just-in-Time Memory: Learning to Curate T
as
k-Adaptive Memory for
LLM
Agents
Agentic memory systems reuse p
as
t experience to improve future performance, yet most existing designs curate memory at w…
智能体
HuggingFace Daily Papers · 9-23
阅读 13
·
访客 13
Agent-Editing World Model: Rethinking World Modeling for
LLM
Agents
Recent advances in large language models (
LLM
s) have enabled agents to tackle long-horizon t
as
ks across diverse environm…
智能体
HuggingFace Daily Papers · 9-23
阅读 12
·
访客 12
The T
as
teful Agent: Me
as
uring and Improving T
as
te in Long-Horizon T
as
ks
LLM
agents incre
as
ingly work on long-horizon t
as
ks, and the decisions they make along the way, such
as
which hypothesis …
智能体
HuggingFace Daily Papers · 9-22
阅读 11
·
访客 10
onPanda: Efficient Annotation of On-Policy Alignment Data for
LLM
s and Agents via Token-Level Correction
We present onPanda, an interactive tool for efficiently annotating
LLM
alignment data and agent trajectories. onPanda ad…
智能体
HuggingFace Daily Papers · 9-21
阅读 13
·
访客 13
Paragraph Boundaries Are Not White Space:Compression Depth
as
the Signature of Hierarchical Structure
Standard positional encodings represent position
as
a one-dimensional reading-order coordinate, but reading order alone …
行业动态
HuggingFace Daily Papers · 9-20
阅读 11
·
访客 9