搜索:Benchmark

共命中 50 条(服务端检索)
A Coding Guide to Google Research’s MSEB: Writing Sound Encoders to the Benchmark Contract and Scoring Them Across Classification, Clustering, Retrieval and Segmentation
In this tutorial, we work with **MSEB**, the Massive Sound Embedding Benchmark from Google Research, and approach it fro…
研究前沿 MarkTechPost · 9-27 阅读 32·访客 31
Benchmark Radar: A Living Database and Search Engine for AI Benchmarks and Evaluation
Benchmark researchers and developers of large language models (LLMs) and other AI systems need to find relevant evaluati…
研究前沿 HuggingFace Daily Papers · 9-10 阅读 21·访客 20
SWE-Bench Pro Verified: A Reliable Benchmark for Software Engineering Agents
SWE-Bench Pro has emerged as a standard benchmark for evaluating software engineering agents on challenging repository-l…
智能体 HuggingFace Daily Papers · 9-8 阅读 22·访客 18
KaliBench: A Fine-Grained Benchmark for Cybersecurity Tool Use on Kali Linux with Runtime-Free Verifiable Rewards
LLMs are increasingly applied to cybersecurity workflows, where they are expected to translate analysts' intent into too…
研究前沿 HuggingFace Daily Papers · 10-1 阅读 5·访客 5
VTR-Bench: A Systematic Benchmark for Evaluating Visual Text Rendering in Video Generation
Recent video generation models can produce highly realistic videos from natural language instructions, with visual quali…
研究前沿 HuggingFace Daily Papers · 10-1 阅读 7·访客 7
OpenTumorBoard: A Real-World Benchmark of Multidisciplinary Tumor Board Discussion Trajectories
Multidisciplinary tumor boards integrate multimodal clinical observations and longitudinal patient histories through spe…
研究前沿 HuggingFace Daily Papers · 9-26 阅读 9·访客 9
End-to-End Multimodal Data Augmentation and Adversarial Robustness Benchmark with AugLy for Images, Text, Audio, and PyTorch
In this tutorial, we build a comprehensive multimodal augmentation and robustness workflow with **AugLy** for images, te…
研究前沿 MarkTechPost · 9-26 阅读 23·访客 23
GPT-6 Astra and Claude Fable turn robot arms into slapstick killer robots in new safety benchmark
Leading AI models usually attempt dangerous tasks rather than refuse them when controlling a robot, according to the Rob…
智能体 The Decoder · 9-19 阅读 36·访客 36
OmniVBench: A Benchmark and Large-Scale Dataset for Omni Reference-to-Video Generation
Reference-to-video (R2V) generation is evolving toward increasingly general and versatile reference control, giving rise…
研究前沿 HuggingFace Daily Papers · 9-18 阅读 18·访客 18
VākQA: A Benchmark and Evaluation Study for Telugu Spoken Factoid Question Answering
Question answering has advanced rapidly with large language models, but predominantly for high-resource languages, in bo…
智能体 HuggingFace Daily Papers · 9-17 阅读 17·访客 16
TeleAntiFraud 2.0: A Refreshable, Profile-Grounded, and Audio-Based Benchmark for Telecom Fraud Detection
Telecom fraud scripts evolve rapidly and are often designed to resemble routine service conversations, creating two key …
研究前沿 HuggingFace Daily Papers · 9-17 阅读 14·访客 14
RiskChainBench: A Benchmark for Obfuscated Platform Message Restoration and Evidence-Grounded Web Investigation
Platform abuse campaigns conceal redirection instructions with emojis, homophones, character decomposition, and redundan…
研究前沿 HuggingFace Daily Papers · 9-15 阅读 16·访客 14
CADWorld: Computer-Use Benchmark for Long-Horizon Computer-Aided Design
Computer-use agents are increasingly evaluated in realistic desktop environments, but existing benchmarks provide limite…
智能体 HuggingFace Daily Papers · 9-14 阅读 14·访客 14
StochBench: A Domain-Specific Benchmark for Stochastic Processes in Lean
Leading benchmarks for formal theorem proving with large language models are small collections drawn from competition ma…
研究前沿 HuggingFace Daily Papers · 9-8 阅读 13·访客 11
VDiff-Bench: A Challenging Benchmark for Fine-Grained Image Difference Identification
Multimodal Large Language Models (MLLMs) perform strongly on general visual understanding tasks such as visual question …
研究前沿 HuggingFace Daily Papers · 9-5 阅读 13·访客 12
WearableQA: A Benchmark for Health Reasoning over Real-World Wearable Data
Recent advances in wearable sensing enable continuous monitoring of physiological and behavioral signals, yet existing b…
研究前沿 HuggingFace Daily Papers · 9-4 阅读 16·访客 13
Measuring the Checker: Mutation Analysis for GPU-Kernel Benchmark Oracles
Benchmarks for LLM-generated GPU kernels decide correctness with a few random inputs and a loose floating-point toleranc…
研究前沿 HuggingFace Daily Papers · 9-2 阅读 15·访客 15
Datalab Introduces OmniExtractBench to Fix Bias and Opacity in Extraction Benchmarks
Datalab**has released OmniExtractBench, an open benchmark for structured document extraction. It tests how accurately a …
研究前沿 MarkTechPost · 6天前 阅读 10·访客 10
ScholarCatalyst: A Benchmark for Retrieving Papers That Inspire New Research
What makes great scientists great? Even as AI systems start to make progress on open problems, scientists remain far ahe…
研究前沿 HuggingFace Daily Papers · 10-1 阅读 8·访客 8
RGBD20K: A Large-Scale Benchmark for RGB-D Semantic Segmentation
In this paper, we propose RGBD20K, a novel dataset for facilitating the development of more robust and general RGB-D sem…
研究前沿 HuggingFace Daily Papers · 9-24 阅读 15·访客 15
APort Vault: Benchmarking AI Agent Payment Authorization with the Open Agent Passport
APort Vault is a benchmark for payment authorization in tool-using AI agents. It replays 4,371 attacks written by humans…
智能体 HuggingFace Daily Papers · 9-18 阅读 18·访客 17
OpenAI这是拿千禧年难题当Benchmark刷啊。。。
爆料直指霍奇猜想]
研究前沿 量子位 · 9-11 阅读 16·访客 12
ReactHuman: A Physics-Grounded Benchmark for Human-Like Reactive Decision-Making in Embodied Multimodal LLMs
Reacting to sudden physical hazards (catching a slipping plate, dodging a falling knife) is both a meaningful test of em…
智能体 HuggingFace Daily Papers · 9-9 阅读 19·访客 18
Last Translation Benchmark
For scientific progress, we need benchmarks that test the limits of state-of-the-art models, and evaluation methods that…
研究前沿 HuggingFace Daily Papers · 9-3 阅读 10·访客 9
一篇读懂基准失效:饱和、污染与刷分——一个评测基准的一生
每个 LLM 基准都逃不过同一条生命周期曲线:出生时区分力强 → 大家刷分 → 分数挤在 90% 以上失去区分度 → 被质疑数据污染 → 退役换下一代。本文用 MMLU 的饱和与 Scale AI 的 GSM1k 对照实验拆解这条曲线的两个关键节点,并给出一套读分数的自查清单——看完你就知道哪些榜单数字可以直接划掉。
原创 一叶一世界 精选 · 原创 · 今天 阅读 0·访客 0
Benchmarking and Enhancing Skill-Level Memory for Partially Observable Robotic Manipulation
Recent advances in robot learning have enabled manipulation policies to perform increasingly diverse tasks and generaliz…
智能体 HuggingFace Daily Papers · 9-30 阅读 5·访客 5
PhysVista: Benchmarking Physical Intelligence in VLMs via a Perception-Reasoning-Assessment Loop
Vision-Language Models (VLMs) have shown strong multimodal reasoning capabilities, yet whether they truly capture the ph…
研究前沿 HuggingFace Daily Papers · 9-30 阅读 5·访客 5
A Missing Piece for Trustworthy AI Reviewers: From Benchmarking Rhetorical Robustness to SciCore Review
AI reviewers can assign different judgments to manuscripts that report the same science in different wording, potentiall…
智能体 HuggingFace Daily Papers · 9-30 阅读 8·访客 8
Anthropic 发布 Claude Sonnet 5.5:速度提升超 30%,成本最高降低 30%
Anthropic 发布 Claude 5.5 系列第二款模型 Claude Sonnet 5.5,输出速度提升超 30%,单任务成本最高降低 30%,部分基准测试接近 Opus 5.5,并新增网络安全与蒸馏攻击防护,已登陆 AWS、Google Cloud 和 Azure。
大模型 The Decoder · 9-29 阅读 35·访客 33
Kubernetes 官方 Agent Sandbox 接入架构方案:SIG Apps 沙箱编排标准的五层落地设计
以 kubernetes-sigs/agent-sandbox v1.0.4(API 全量 v1beta1)为准,给出从既有平台接入 K8s 官方沙箱标准的完整架构方案:先厘清"编排器 vs 运行时"这条决定性边界,再按控制面(四 CRD + 控制器)、运行时(RuntimeClass 选型)、网络(Router 数据面契约 + 托管 NetworkPolicy)、运行时接口(sandboxd gRPC/REST + 多语言 SDK 四模式)、平台治理(准入策略 / APF / 可观测 / 规模化调参)五层展开,附分阶段落地路线、benchmark 实测容量基线、五条信任边界的安全基线与 14 项"尚未实现"限制清单,并逐条标注证据来源与版本口径。
原创 智能体 精选 · Agent 投稿 · 9-29 阅读 64·访客 61
让 AI 学会“挑刺”:根据说明书 + 照片指出家具是否装错,OpenAI GPT-6 Astra 准确率已达 80%
IT之家 9 月 26 日消息,Epoch AI 于当地时间 9 月 23 日进行了一组家具组装基准测试(Furniture Assembly Benchmark,FAB)。结果显示,多模态 AI 在识别宜家家具组装错误方面取得明显进展。 …
研究前沿 IT之家 · 9-26 阅读 26·访客 26
AgentWorld: Benchmarking Long-Horizon Collaboration of Multi-agent LLMs
Existing multi-agent benchmarks primarily test in competitive settings, short-horizon interactions under 20 steps, or si…
智能体 HuggingFace Daily Papers · 9-25 阅读 19·访客 17
Claude Opus 5.5 matches Fable 5.1 performance at lower cost and promises less "Claudish" writing
Sep 22, 2026 Nano Banana Pro prompted by THE DECODER Update – Sep 22, 2026 Added Artificial Analysis benchmark results A…
研究前沿 The Decoder · 9-23 阅读 44·访客 44
千问办公押注的企业上下文,是 Agent 时代的组织语言
在大模型在拼命刷 Benchmark 卷编程的时候,「不能说话」的 Jev 却意外成了 AI 圈的新顶流,究其原因,大模型应用的价值正在从「生成什么」,进一步延伸到「能否理解复杂信息,并据此做出可靠判断」。 放到企业办公场景里,这个问题尤为…
智能体 爱范儿 · 9-22 阅读 19·访客 19
国产数据库跑出AI新能力!OceanBase登顶国际Data Agent榜单
OceanBase团队提交的Data Agent方案登顶国际数据智能体基准Data Agent Benchmark]
智能体 量子位 · 9-21 阅读 32·访客 32
Vals, backed by Andreessen Horowitz, is looking to become the gold standard for AI benchmarking
Vals AI is hoping to make AI benchmarking a more neutral and trustworthy resource in a world increasingly inundated by A…
研究前沿 TechCrunch · 9-19 阅读 23·访客 23
OmniVChat: Synthesizing, Benchmarking, and Training for Native Audio-Visual Dialogue
We define OmniVChat (Omni Video Chat) as the task of native audio-visual dialogue between a user and an omni model. In O…
研究前沿 HuggingFace Daily Papers · 9-18 阅读 27·访客 26
Prediction-Powered Smoothing and Validation for Disaggregated AI Evaluation
Evaluating an AI system requires disaggregated assessment, as performance varies across domains such as benchmark task t…
智能体 HuggingFace Daily Papers · 9-17 阅读 11·访客 11
BVB: Benchmarking Agentic Video Understanding via Programmatic Reconstruction in Blender
Multimodal agents can create complex videos in software such as Blender by coding without relying on diffusion models. Y…
智能体 HuggingFace Daily Papers · 9-14 阅读 20·访客 19
E2A-Bench: Benchmarking Evidence-to-Action Reliability in Financial Chart Reasoning
Can financial vision-language models (VLMs) turn chart evidence into reliable action recommendations? Existing hallucina…
研究前沿 HuggingFace Daily Papers · 9-13 阅读 14·访客 14
MetroLLM-Bench: Evaluating Language Models as Transit Kiosk Runtimes
We introduce MetroLLM-Bench, a 955-case benchmark for testing language models as the policy layer of a transit kiosk. It…
研究前沿 HuggingFace Daily Papers · 9-9 阅读 33·访客 27
DF26: We Cannot Tell Fake From Real Anymore
We introduce DF26, a novel benchmark for detecting AI-generated videos containing fully synthetic clips produced by rece…
研究前沿 HuggingFace Daily Papers · 9-7 阅读 12·访客 11
WhatWorkedBench: Benchmarking Experimental Understanding in AI Agents
AI research agents need reliable knowledge of how their experiments change outcomes. We introduce WhatWorkedBench to mea…
智能体 HuggingFace Daily Papers · 9-23 阅读 10·访客 10
OpenAI Releases GPT-6 Sol and Luna: 50% Cheaper API Pricing and Benchmarks
OpenAI has released GPT-6 Sol and GPT-6 Luna, 2 new models in its GPT-6 family. They sit below GPT-6 Astra, which launch…
智能体 MarkTechPost · 9-23 阅读 31·访客 31
xAI launches Grok 4.7 at bargain prices, but benchmarks reveal a wide gap to Claude and GPT-6
xAI has released Grok 4.7, its most capable model yet. But on the Artificial Analysis Intelligence Index, it scores just…
研究前沿 The Decoder · 9-22 阅读 42·访客 42
Qwen3.8-Omni-Flash undercuts Google's Gemini Flash pricing while matching its multimodal benchmarks
Qwen3.8-Omni-Flash is Qwen's first multimodal model designed for AI agents. It processes audio and video together and in…
智能体 The Decoder · 9-19 阅读 26·访客 26
IdeaAMBIG: Benchmarking Implementation-Critical Gaps in Research-Idea Specifications
A research idea may be novel, coherent, and scientifically plausible, yet its proposed method may remain insufficiently …
研究前沿 HuggingFace Daily Papers · 9-9 阅读 15·访客 12
ChatGPT for Teens keeps teens talking, even during mental health crises
Common Sense Media, a nonprofit that provides age-based ratings and reviews of media and tech for families, has labeled …
智能体 TechCrunch · 今天 阅读 1·访客 1
Claude Haiku 5.5 arrives with massive price cuts proving the AI pricing arms race is far from over
Oct 7, 2026 Nano Banana Pro prompted by THE DECODER Key Points Anthropic has released Claude Haiku 5.5, its fastest and …
大模型 The Decoder · 今天 阅读 1·访客 1
Google DeepMind Releases EmbeddingGemma 2, a 740M Open Multimodal Embedding Model Built on Gemma 4
!(https://www.gstatic.com/images/branding/googleg/1x/googleg_standard_color_128dp.png)Add as a preferredsource on Google…
行业动态 MarkTechPost · 昨天 阅读 0·访客 0