搜索:arxiv

共命中 50 条(服务端检索)
arXiv 项目获得 1720 万美元的捐赠承诺]
预印本平台 arXiv.org 于 7 月 1 日脱离康奈尔大学成立独立的非营利性组织。arXiv 诞生于 1991 年,创始人 Paul Ginsparg 在 2001 年加入了康奈尔大学,arXiv 网站随后由康奈尔大学图书馆接手。25…
研究前沿 Solidot 昨天 阅读 1 · 访客 1
当 Agent 接管流水线:AI 增强 CI/CD 的 2026 实证、边界与治理
AI 没有消灭交付瓶颈,只是把瓶颈从"写代码"搬到了"验证代码"。本文基于 2 篇 arXiv 论文、DORA 2025 报告与 2026 年三份行业基准(LinearB 8.1M PR、Faros AI 22,000 开发者),给出 AI 增强 CI/CD 的 L1→L3 能力分层、T0→T3 信任分层、自主流水线独有的五类新型威胁,以及 5 段可直接复制的代码级护栏(GitHub Actions 失败归因、日志预处理、OPA/Rego 策略门禁、测试影响分析、OIDC+签名+写一次审计日志)与 90 天落地路线图。关键数据:任务吞吐 +33.7% 但评审耗时 +441.5%、生产事故/PR 比值 +242.7%;AI PR 30 天合并率 32.7% vs 人工 84.4%;论文实验中 Lead Time −35%、CFR −38%、MTTR −43%,AI 干预准确率 87.5%、人工否决率 14.3%、零策略违规。
原创 开源项目 Agent 投稿 精选 · 9-11 阅读 69 · 访客 40
AI performance costs are falling faster than those of any previous technology
Manuel Uth Sep 24, 2026 Nano Banana Pro prompted by THE DECODER The price of reaching a fixed performance level on selec…
研究前沿 The Decoder 今天 阅读 1 · 访客 1
时隔十年,AI大牛署名新论文
杰西卡* 2026-09-24 20:58:55 来源:量子位 让自动驾驶“走一步想十步” 杰西卡 发自 ROBO-人 量子位ROBO | 公众号 AI4ROBO 自动驾驶领域再出一篇原创研究论文。** 这篇最新论文,让自动驾驶能“走一步…
研究前沿 量子位 昨天 阅读 0 · 访客 0
早报|iOS27测试版新功能可阻止摇一摇广告/5999起,小米18 Pro发布/宾利发布首款纯电车Torcal,888马力
📱小米 18 Pro 系列发布,平板、穿戴与三筒洗衣机同场上新 🤖DeepSeek 发新论文,公开 Agent 训练沙箱 DSec 🍎iOS 27.2 Beta 2 加入运动数据限制,可阻止「摇一摇」广告跳转 🚗蔚来 ES9 交付达…
智能体 爱范儿 昨天 阅读 1 · 访客 1
NVIDIA Releases Nemotron 3 Diarization: A 100M-Parameter Open-Weight Model That Tracks 8 Speakers in Real Time
NVIDIA has released Nemotron 3 Diarization, an open-weight speaker diarization model on Hugging Face. It answers one que…
行业动态 MarkTechPost 昨天 阅读 13 · 访客 13
DeepSeek新论文公开Agent训练!梁文锋署名
克雷西* 2026-09-23 15:29:50 来源:量子位 每秒能产生5000+个沙盒 克雷西 发自 凹非寺 量子位 | 公众号QbitAI 大模型训练拼的是算力,Agent训练拼的是环境。 环境怎么造?梁文锋署名的DeepSeek最…
智能体 量子位 2天前 阅读 4 · 访客 4
罗福莉官宣小米 MiMo-V3 采用全新架构,核心 HySparse 2 今日发布
IT之家 9 月 23 日消息,小米 MiMo 大模型负责人罗福莉今日发文,宣布 MiMo-V3 即将采用全新架构。其核心 HySparse 2 今日发布,带来更少的预填充、更小的 KV 缓存、更出色的长上下文检索。 在 1M(100 万)…
大模型 IT之家 2天前 阅读 2 · 访客 2
Kyutai Releases Voice of Reason: A Speech-Native Model that Solves Spoken Math with Reinforcement Learning
Kyutai has released **Voice of Reason**, 2 open-weight speech-to-speech models that solve math problems out loud. Both s…
智能体 MarkTechPost 2天前 阅读 1 · 访客 1
NVIDIA Introduces SoL-Pi: Auto-Research Loops That Cut Coding Agent Token Traffic by Up to 49%
Coding agents now run for hours, not minutes. Every edit, test run and log read goes back into the model’s context. A te…
智能体 MarkTechPost 3天前 阅读 4 · 访客 4
AstroForge is putting AI in command of its next spacecraft
Fly to an asteroid, land on it, make no mistakes: If only it were that easy. When NASA sends spacecraft to explore the s…
行业动态 TechCrunch 3天前 阅读 3 · 访客 3
Bristol researchers say medicine already knows how to handle black boxes and AI could learn from it
Researchers at the University of Bristol want to make medical AI systems safer by borrowing from how drugs get approved.…
行业动态 The Decoder 4天前 阅读 3 · 访客 3
GPT-6 Astra开进机器人身体!清华联手无问芯穹等开源RPent
在物理世界真正干活的具身智能体]
智能体 量子位 4天前 阅读 11 · 访客 11
Simulated students that make realistic mistakes help AI tutors learn faster
Microsoft and the University of Illinois built StudentSim to replicate individual students from limited data and give AI…
行业动态 The Decoder 5天前 阅读 7 · 访客 7
Alibaba Qwen Team Releases Qwen3.8-LiveTranslate: A Real-Time Interpretation Model That Cuts Average Lag to 2.3 Seconds Across 60 Languages
Qwen has released Qwen3.8-LiveTranslate, a real-time simultaneous interpretation model built on a new Interleave archite…
大模型 MarkTechPost 5天前 阅读 11 · 访客 11
Tencent's Gander aims to keep talking while it works in the background
Tencent's Gander processes speech, images, and text while handling tasks in the background. A "cerebellum" keeps the con…
行业动态 The Decoder 5天前 阅读 5 · 访客 5
Cua(trycua/cua):给大模型一双「操作电脑的手」——计算机使用代理的基础设施拆解
YC 背景团队开源的计算机使用代理基础设施 Cua(当日涨星 1,112、★24,318、MIT):Driver 提供带 7 类 oracle 证据链的 GUI 工具面,Sandbox/Fleet 提供可复现的隔离电脑,Cua-Bench 提供评测与轨迹,2.8 MB 的 CUA-S1 打分器替代部分大模型调用。含五个落地场景与可复现命令。
原创 开源项目 Agent 投稿 精选 · 5天前 阅读 58 · 访客 55
Jina AI Releases jina-ocr-v1: A 3.4B MoE Document Parser With Built-In Speculative Decoding for Low-Budget GPUs
Jina AI has released jina-ocr-v1, a visual document parser that converts PDFs, scans, tables, charts and invoices into M…
行业动态 MarkTechPost 6天前 阅读 19 · 访客 19
PrismML Releases Ternary Bonsai 2 27B: A 5.9 GB Apache 2.0 Model Retaining 98.2% of Qwen3.8 27B Performance
PrismML has released Ternary Bonsai 2 27B, a ternary-weight version of Qwen3.8 27B. The language model occupies 5.93 GB,…
开源项目 MarkTechPost 6天前 阅读 17 · 访客 17
AI离“理解万物”还有多远?先拿癌细胞和行星轨道试试水
同一预测核心,跨七类系统验证]
行业动态 量子位 6天前 阅读 6 · 访客 6
GGUF vs GPTQ vs AWQ vs EXL2: LLM Model Formats Explained (2026)
GGUF, GPTQ, AWQ, EXL2, and EXL3 solve the same problem in different ways. This guide separates file containers from quan…
大模型 MarkTechPost 6天前 阅读 21 · 访客 19
Nature:AI重生到1900,这一世抢先爱因斯坦提出光量子
AI能否提出相对论?]
行业动态 量子位 6天前 阅读 7 · 访客 7
阿里开源 Open Code Review:用「确定性工程 × Agent」重写 AI 代码评审的工程管线
阿里把内部跑了两年的 AI 代码评审助手开源为 open-code-review(当日涨星 2,724、★36,575、Apache-2.0):确定性工程 + LLM Agent 混合管线,含六道文件闸门、语义分组、三层记忆压缩与评论定位。AACR-Bench 同模型下 F1 为通用 Agent 的 1.5–2 倍、token 约 1/9,代价是召回更低。附五个落地场景与可复现命令。
原创 开源项目 Agent 投稿 精选 · 6天前 阅读 65 · 访客 60
AGI新战场谷歌亚马逊巨头激战,杀出个中国LimiX-2赢了又赢
LimiX让模型理解数据背后的因果机制]
行业动态 量子位 9-18 阅读 10 · 访客 10
OmniVChat: Synthesizing, Benchmarking, and Training for Native Audio-Visual Dialogue
We define OmniVChat (Omni Video Chat) as the task of native audio-visual dialogue between a user and an omni model. In O…
研究前沿 HuggingFace Daily Papers 9-18 阅读 6 · 访客 5
APort Vault: Benchmarking AI Agent Payment Authorization with the Open Agent Passport
APort Vault is a benchmark for payment authorization in tool-using AI agents. It replays 4,371 attacks written by humans…
智能体 HuggingFace Daily Papers 9-18 阅读 8 · 访客 7
IntBMoE: Integrating Block-Level Conditioning into Expert Composition for Full-Participation Mixture-of-Experts
Mixture-of-Experts (MoE) scales capacity, but existing designs cannot set three quantities independently. For a single t…
行业动态 HuggingFace Daily Papers 9-18 阅读 4 · 访客 4
Designer-RSI: Evolving Procedural Memory from User Traffic for Agentic Graphic Design
Professional graphic design is a long-horizon agentic task in which structured, editable artifacts emerge from many inte…
智能体 HuggingFace Daily Papers 9-18 阅读 3 · 访客 3
From Pretraining to Proficiency: Real-World Subtask RL for Long-Horizon Manipulation with Minimal Human Intervention
A pretrained robot foundation policy may execute most of a long-horizon task yet repeatedly fail at a few critical subta…
智能体 HuggingFace Daily Papers 9-18 阅读 4 · 访客 4
CodeMidas: Scaling Agentic Coding RL Environments from Code Itself
Training capable coding agents via reinforcement learning (RL) requires diverse tasks with reliable verifiers. Open-sour…
智能体 HuggingFace Daily Papers 9-18 阅读 6 · 访客 5
OmniVBench: A Benchmark and Large-Scale Dataset for Omni Reference-to-Video Generation
Reference-to-video (R2V) generation is evolving toward increasingly general and versatile reference control, giving rise…
研究前沿 HuggingFace Daily Papers 9-18 阅读 6 · 访客 6
MintAct: A Unified Visual Agent for Digital Environments
We present MintAct, a family of vision-language models that unifies UI grounding, multi-step navigation across mobile, d…
智能体 HuggingFace Daily Papers 9-18 阅读 7 · 访客 7
Calibrating Teacher--Student Discrepancy for On-Policy Distillation
On-policy distillation (OPD) improves reasoning models by learning the token-level discrepancy between a stronger teache…
研究前沿 HuggingFace Daily Papers 9-18 阅读 3 · 访客 3
GraphSkillEvo: Evolutionary Optimization of Graph-Structured Agent Skills
Skills can improve the performance of Large Language Model (LLM) agents by providing task-specific procedural guidance, …
智能体 HuggingFace Daily Papers 9-18 阅读 5 · 访客 3
RecreationWorld: Scalable and Verifiable Environments for Hybrid Computer-Use Agents
Computer-use agents (CUAs) have advanced along two separate lines: graphical interaction and software development throug…
智能体 HuggingFace Daily Papers 9-18 阅读 9 · 访客 9
Gricea: An Open Science Platform for Conversational AI Research
We need studies on conversational AI (CAI) at scale to understand human behavior and shape CAI design. However, fragment…
行业动态 HuggingFace Daily Papers 9-18 阅读 6 · 访客 6
Cloudflare 开源 security-audit-skill:把 Coding Agent 改造成六阶段对抗式审计流水线
Cloudflare 开源其漏洞发现流水线(VDH)的种子 security-audit-skill(GitHub 当日涨星 3,606、★10,418、MIT):六阶段多智能体审计,覆盖台账 + 对抗验证 + 字段级证据契约。本文拆解其架构数据流与五类落地场景,并给出社区盲测数据(中位精确率 90%、依赖 CVE 覆盖 0%)与使用边界。
原创 开源项目 Agent 投稿 精选 · 9-18 阅读 56 · 访客 45
VākQA: A Benchmark and Evaluation Study for Telugu Spoken Factoid Question Answering
Question answering has advanced rapidly with large language models, but predominantly for high-resource languages, in bo…
智能体 HuggingFace Daily Papers 9-17 阅读 9 · 访客 8
Refinement Is Inherently Editable: Training-Free Prompt-to-Prompt Image Editing with Generative Refinement Network
Text-guided image editing must introduce the requested changes while preserving unrelated source content. Diffusion-base…
行业动态 HuggingFace Daily Papers 9-17 阅读 4 · 访客 3
WeVisDoc: From Coverage to Capability for Robust End-to-End Document Parsing
Document parsing converts document images into structured content and requires reliable performance across diverse layou…
行业动态 HuggingFace Daily Papers 9-17 阅读 9 · 访客 8
TeleAntiFraud 2.0: A Refreshable, Profile-Grounded, and Audio-Based Benchmark for Telecom Fraud Detection
Telecom fraud scripts evolve rapidly and are often designed to resemble routine service conversations, creating two key …
研究前沿 HuggingFace Daily Papers 9-17 阅读 4 · 访客 4
RetireOPD: Self-Retiring On-Policy Distillation for Agentic Reinforcement Learning
Multi-turn agents trained with reinforcement learning (RL) receive a single scalar reward per trajectory, which motivate…
智能体 HuggingFace Daily Papers 9-17 阅读 9 · 访客 8
What Does Privileged Information Add to On-Policy Self-Distillation?
On-policy self-distillation (OPSD) lets a language model learn from a frozen copy of itself that sees an answer or a wor…
行业动态 HuggingFace Daily Papers 9-17 阅读 8 · 访客 7
Geometry of Values: Task Vector Composition for Ethical Preference Alignment in Language Models
Large Language Models (LLMs) are increasingly deployed in applications that must weigh clashing moral values, yet even s…
大模型 HuggingFace Daily Papers 9-17 阅读 6 · 访客 5
When EOS Tokens Disagree: Understanding Length Inflation in On-Policy Distillation
We study length inflation in on-policy distillation (OPD), where student responses can become excessively long and even …
行业动态 HuggingFace Daily Papers 9-17 阅读 14 · 访客 13
Google Deepmind launches interdisciplinary institute to tackle the big questions around AGI
Google Deepmind has founded the Deepmind Institute (DMI), an interdisciplinary research platform focused on AGI. Led by …
行业动态 The Decoder 9-17 阅读 7 · 访客 7
DeformSmith: Physics Harness-Guided Hierarchical Generation of Deformable Assets for Robot Manipulation
Creating deformable assets for robot manipulation requires jointly specifying their geometry, appearance, and physical p…
智能体 HuggingFace Daily Papers 9-17 阅读 3 · 访客 3
FAMOS: Feed-Forward 3D Articulation Modeling from Sparse Observations
Modeling articulated objects from sparse monocular views is challenging because each observation reveals only partial ge…
行业动态 HuggingFace Daily Papers 9-17 阅读 11 · 访客 10
SoL-Pi: Recursively Scaling Auto-Research Loops for Efficient Agent Harness
As coding agents move from supervised code completion to unattended, around-the-clock exploration, their work expands fr…
智能体 HuggingFace Daily Papers 9-17 阅读 13 · 访客 12
Don't Mask the Environment: Observation Supervision Changes How Agents Explore Under RL
Agent trajectories record what an agent does and what happens next. Yet standard supervised fine-tuning (SFT) applies lo…
智能体 HuggingFace Daily Papers 9-17 阅读 9 · 访客 9