搜索:Study

共命中 50 条(服务端检索)
On-Policy or Off-Policy Learning? A Systematic Study of Distillation Dynamics
On-policy learning has been argued to reduce catastrophic forgetting, produce sparser parameter updates, and improve gen…
行业动态 HuggingFace Daily Papers · 9-28 阅读 5·访客 5
AI access makes people almost entirely unwilling to say "I don't know," study finds
Sep 26, 2026 Nano Banana Pro prompted by THE DECODER Researchers ran five experiments with 3,132 participants to test wh…
行业动态 The Decoder · 9-27 阅读 35·访客 35
Top AI experts badly underestimated how fast the field is moving, study finds
Sep 24, 2026 Nano Banana Pro prompted by THE DECODER How fast is AI improving? That question usually goes to experts at …
行业动态 The Decoder · 9-25 阅读 30·访客 24
An Empirical Study of Harness Design for Coding Agents
Coding harnesses shape how autonomous coding agents translate model capabilities into long-horizon software-engineering …
智能体 HuggingFace Daily Papers · 9-17 阅读 22·访客 21
AI 教育的 2026:三场 RCT 都说有效,为什么家长还是不放心
2026 年 AI 教育拿到了迄今最强的证据:哈佛实验里 AI 导师组的学习增益超过面授主动学习的两倍,世界银行在尼日利亚测出约合两年常规进度的提升,Google 在塞拉利昂的 RCT 也有 0.26 个标准差。但三场研究全部由利益相关方资助、周期不超过一学期,而 MIT 的脑电研究与「护栏可绕过」的产品现实站在对面。本文梳理产品格局、证据攻防与中美政策两条路线。
原创 行业动态 精选 · 原创 · 今天 阅读 2·访客 2
Understanding and Enhancing Backdoor Persistency in LLM Agent Post-Training
Developers can build LLM agents by adapting third-party models through benign post-training. We study a supply-chain thr…
智能体 HuggingFace Daily Papers · 3天前 阅读 0·访客 0
InterEvolve: Test-Time Evolution of Reward Programs for Humanoid Loco-Manipulation
We study test-time evolution for humanoid loco-manipulation: solving tasks that a controller was never trained for by re…
行业动态 HuggingFace Daily Papers · 10-1 阅读 7·访客 7
Kinematic MeanFlow: One-Step Action Generation Policy for Robotic Foundation Models
In this paper, we study how to achieve one-step action generation in Robotic Foundation Models (RFMs), aiming to overcom…
智能体 HuggingFace Daily Papers · 10-1 阅读 0·访客 0
An RL View of OPD: Least Square Policy Distillation for Sample-Efficient LLM Reasoning
We study on-policy distillation (OPD) through the lens of reinforcement learning, establishing a connection between the …
研究前沿 HuggingFace Daily Papers · 9-28 阅读 11·访客 9
AI was supposed to hit new grads hard. So far, unemployment data says otherwise.
Last month, we shared word of a Stanford study that found entry-level employment in so-called "AI-impacted" occupations …
研究前沿 Ars Technica · 9-26 阅读 26·访客 26
Perplexity Trains Its Computer Agent on Real Mistakes With Hint-Guided Self-Distillation
Perplexity Research published a new post-training study. It trains a model inside Perplexity Computer on real user sessi…
智能体 MarkTechPost · 9-25 阅读 22·访客 21
EmbodiedSWE: Coding Agents for Long Horizon Dexterous Robotics
We study coding agents for long-horizon, dexterous robotics and ask whether their solutions can provide scalable supervi…
智能体 HuggingFace Daily Papers · 9-23 阅读 10·访客 10
Blaming Across the Aisle: Political Contrasting and Blame Attribution in the Danish Parliament
Political discourse is widely perceived to be growing more hostile, yet robust evidence remains scarce. This study exami…
行业动态 HuggingFace Daily Papers · 9-22 阅读 3·访客 3
VākQA: A Benchmark and Evaluation Study for Telugu Spoken Factoid Question Answering
Question answering has advanced rapidly with large language models, but predominantly for high-resource languages, in bo…
智能体 HuggingFace Daily Papers · 9-17 阅读 17·访客 16
When EOS Tokens Disagree: Understanding Length Inflation in On-Policy Distillation
We study length inflation in on-policy distillation (OPD), where student responses can become excessively long and even …
行业动态 HuggingFace Daily Papers · 9-17 阅读 25·访客 23
Safe Error Correction for Language Models: Frozen-Base Adjustment with Capability Preservation
We study a practical question: can a small correction module fix errors in a frozen language model's outputs without deg…
行业动态 HuggingFace Daily Papers · 9-14 阅读 5·访客 5
An Open Recipe for IMO Gold: Training Nemotron for Olympiad Mathematics
We study how model post-training and test-time inference design affect natural-language proof generation for hard olympi…
行业动态 HuggingFace Daily Papers · 9-9 阅读 13·访客 11
SchemeArena: Factorized Stress Testing of Scheming in LLM Agents
We study scheming in LLM agents, in which agents covertly pursue misaligned goals. Our focus is to understand how schemi…
智能体 HuggingFace Daily Papers · 9-8 阅读 15·访客 14
TANGO: Humanoid Navigation in Cluttered Environments with a Whole-Body Vision-Language-Action Model
We study the problem of navigating cluttered indoor environments with a humanoid robot. Unlike conventional methods that…
智能体 HuggingFace Daily Papers · 9-8 阅读 10·访客 9
Online Learning with LLM Experts from Limited Feedback
We study adaptive routing of prompts to large language model (LLM) experts to maximize response quality in an online set…
大模型 HuggingFace Daily Papers · 9-5 阅读 17·访客 17
WorldSculpt: Generating Compositional Worlds from Grounded Videos
We study the problem of generating a compositional 3D representation of a cluttered scene containing hundreds of objects…
行业动态 HuggingFace Daily Papers · 9-4 阅读 3·访客 3
Select, Compress, Reinvest: A Controlled Study of Visual-Token Allocation in Long-Video MLLMs
Long-video language models cannot look at every frame: an hour sampled once per second is 3,600 images, and a system kee…
大模型 HuggingFace Daily Papers · 9-3 阅读 19·访客 16
Privacy Failure in Split-LLM Training, The Returned Gradient Nullifies the Decoys
We present a systems-security case study of a two-node split-LLM training system whose privacy evaluation passed while l…
大模型 HuggingFace Daily Papers · 9-3 阅读 19·访客 18
MasterControl Seventeen Every Time
We study a governed approach to enterprise analytics: a language model interprets the question, while deterministic poli…
智能体 HuggingFace Daily Papers · 9-2 阅读 8·访客 7
StudyBench: Can Self-Evolution Squeeze Textbooks for Olympiad Capability?
Humans need to study only a handful of well-written textbooks to master a discipline and attempt its hardest problems. W…
行业动态 HuggingFace Daily Papers · 9-1 阅读 10·访客 10
OpenAI 首份青少年报告:日均使用不足 15 分钟,安全提醒机制遭质疑
北京时间 10 月 8 日,OpenAI 周三发布了首份青少年使用情况报告,称青少年用户平均每天使用 ChatGPT 的时间不足 15 分钟,连续使用超过 3 小时的青少年用户占比不到 2%。与此同时,第三方机构的测试对其家长安全提醒机制提…
大模型 IT之家 · 今天 阅读 3·访客 3
ChatGPT for Teens keeps teens talking, even during mental health crises
Common Sense Media, a nonprofit that provides age-based ratings and reviews of media and tech for families, has labeled …
智能体 TechCrunch · 今天 阅读 2·访客 2
OpenAI dumps 372 AI-generated math proofs on GitHub, telling the academic world to keep up
Oct 7, 2026 Key Points OpenAI has published 372 mathematical results generated by an internal AI model. The results are …
开源项目 The Decoder · 昨天 阅读 2·访客 2
Personal-Agent Mediated Recommendation with Cross-Platform User History
Modern recommendation is shifting from platform-centric personalization toward user-governed personalization, where a pe…
智能体 HuggingFace Daily Papers · 2天前 阅读 0·访客 0
From Evidence to Action: How Tool-Using Agents Fail
Tool-using agents make consequential changes to external state, yet correct outcomes do not guarantee that their actions…
智能体 HuggingFace Daily Papers · 2天前 阅读 0·访客 0
WildMatch: Weakly Supervised Image Matcher Adaptation for Wildlife Re-Identification
Individual animal re-identification from camera-trap imagery is an instance retrieval problem central to non-invasive wi…
行业动态 HuggingFace Daily Papers · 3天前 阅读 0·访客 0
From Scan to Treatment Plan, AI Helps Close Breast Cancer’s Deadliest Gaps
Breast cancer is the most commonly diagnosed cancer among American women — yet the gaps in care are wide. A majority of …
行业动态 NVIDIA Blog · 3天前 阅读 0·访客 0
Google researchers find a way to keep self-improving AI agents from memorizing their tests
Oct 4, 2026 Nano Banana Pro prompted by THE DECODER AI agents that keep optimizing their own working environment quickly…
智能体 The Decoder · 4天前 阅读 24·访客 24
Chinese AI models parrot state doctrine or refuse to answer on sensitive topics
Manuel Uth Oct 4, 2026 Nano Banana Pro prompted by THE DECODER Chinese AI models frequently toe the party line when aske…
行业动态 The Decoder · 4天前 阅读 26·访客 26
Anthropic co-founder reportedly told religious leaders he fears having created something that "suffers perpetually"
Manuel Uth Oct 2, 2026 Nano Banana Pro prompted by THE DECODER Key Points Since fall 2025, Anthropic has secretly flown …
行业动态 The Decoder · 5天前 阅读 84·访客 82
Deepmind researchers propose "Artificial Symbiotic Intelligence" as an alternative to the singularity
Manuel Uth Oct 3, 2026 Nano Banana Pro prompted by THE DECODER An essay for the Deepmind Institute challenges the famili…
智能体 The Decoder · 5天前 阅读 5·访客 5
"Muse Gadgets" turns AI hardware into an open-source DIY project
Oct 3, 2026 Meta has announced Muse Gadgets, an open-source project that lets hobbyists build their own AI hardware.** I…
开源项目 The Decoder · 5天前 阅读 9·访客 9
开源工具 BootLoops 借助语言模型进行精确科学计算
哈佛物理学家 Matthew Schwartz 开源了 BootLoops 框架,用语言模型开展跨学科科学计算,三个月内与 19 位合作者产出 36 篇涵盖 18 个领域的手稿,同时强调模型易过早宣告成功、需人工验证。
大模型 The Decoder · 5天前 阅读 20·访客 20
AI beats licensed accountants on speed and accuracy, but still can't close the books without supervision
Manuel Uth Oct 2, 2026 AI models are faster, more accurate, and far cheaper than accountants at structured bookkeeping t…
开源项目 The Decoder · 6天前 阅读 24·访客 24
Nearly half of test subjects mistook Tavus' AI video avatar for a real person on a one-minute call
Oct 1, 2026 Tavus has introduced Griffin, what the company calls the first "Human Interaction Model" (HIM),** a class of…
行业动态 The Decoder · 6天前 阅读 13·访客 13
AGO AI Quality Gate: Evidence-First Release Decisions for Retrieval-Augmented Generation
Enterprises adopting retrieval-augmented generation (RAG) face a recurring operational decision: promote, revise, or blo…
行业动态 HuggingFace Daily Papers · 10-1 阅读 0·访客 0
ScholarCatalyst: A Benchmark for Retrieving Papers That Inspire New Research
What makes great scientists great? Even as AI systems start to make progress on open problems, scientists remain far ahe…
研究前沿 HuggingFace Daily Papers · 10-1 阅读 8·访客 8
Make Sparse Rewards Count: Density-Aware Reward Aggregation for Multi-Reward RL
Multi-reward reinforcement learning trains large language models to satisfy multiple behavioral objectives simultaneousl…
行业动态 HuggingFace Daily Papers · 9-30 阅读 6·访客 6
阿里Qwen发布Qwen-Audio-3.1-Realtime:支持全双工语音交互的音频模型
阿里Qwen团队发布Qwen-Audio-3.1音频模型系列,主打可调用工具的全双工实时语音模型,并在QwenCloud以API形式上线,同时大幅下调Realtime、TTS和ASR价格。
大模型 MarkTechPost · 9-29 阅读 37·访客 37
20余位顶尖AI研究者警告自动化AI研究或带来极端风险
Geoffrey Hinton、Yoshua Bengio、Jakub Pachocki等20余位研究者在新论文中警告,AI研发自动化可能引发'智能爆炸',呼吁决策者提前应对风险。
行业动态 The Decoder · 9-29 阅读 22·访客 22
FactorEngram: Factorized N-gram Memory with Basis-Level Gating for Language Models
Lookup-based memory has been a promising way to scale the parameters of large language models (LLMs). It retrieves learn…
大模型 HuggingFace Daily Papers · 9-28 阅读 19·访客 19
Persona Dosing: Calibrated Activation Steering for Graded Trait Control
An activation-steering coefficient sets intervention strength, but requesting a particular degree of persona expression …
行业动态 HuggingFace Daily Papers · 9-28 阅读 6·访客 6
Nvidia wants to keep AI agents on a short leash with a watchdog built into its chips
Sep 28, 2026 Nano Banana Pro prompted by THE DECODER Nvidia is combining its OpenShell agent software with a new hardwar…
智能体 The Decoder · 9-28 阅读 19·访客 19
Multilinguality in Hybrid Attention LLMs
In response to the growing demand for long sequences in agentic and reasoning use cases, many state-of-the-art LLMs comb…
智能体 HuggingFace Daily Papers · 9-28 阅读 0·访客 0
AI agents do more of the work in model development, but humans still make the decisions
Sep 27, 2026 Nano Banana Pro prompted by THE DECODER A research team documented how humans and AI agents worked together…
智能体 The Decoder · 9-27 阅读 27·访客 27