搜索:benchmark

共命中 5 条(服务端检索)
Last Translation Benchmark
For scientific progress, we need benchmarks that test the limits of state-of-the-art models, and evaluation methods that…
研究前沿 HuggingFace Daily Papers 6天前
MOLE: Detecting Insider Threats in AI Agents
Model misalignment, prompt injection, or operator misuse could lead AI agents operating frontier-lab accounts to exfiltr…
智能体 HuggingFace Daily Papers 2天前
Ask Before You Optimize: Dynamic Pre-Formulation Clarification for Interactive Optimization
Large language models (LLMs) are increasingly used to formulate optimization models from natural-language problem descri…
大模型 HuggingFace Daily Papers 5天前
τ^τ-Bench: An Environment for End-To-End, Realistic Agent Construction
LLM agents are rapidly becoming production software, deployed to handle customer service, adjudicate disputes, and opera…
智能体 HuggingFace Daily Papers 5天前
Knowing What Not to Answer: Selective Non-Compliance in Vision-Language Models
Vision-language models (VLMs) are expected to respond helpfully to appropriate requests while withholding compliance wit…
行业动态 HuggingFace Daily Papers 5天前