搜索:huggingface daily papers

共命中 50 条(服务端检索)
Benchmark Radar: A Living Database and Search Engine for AI Benchmarks and Evaluation
Benchmark researchers and developers of large language models (LLMs) and other AI systems need to find relevant evaluati…
研究前沿 HuggingFace Daily Papers 9-10 阅读 3 · 访客 2
IdeaAMBIG: Benchmarking Implementation-Critical Gaps in Research-Idea Specifications
A research idea may be novel, coherent, and scientifically plausible, yet its proposed method may remain insufficiently …
研究前沿 HuggingFace Daily Papers 9-9 阅读 5 · 访客 3
WearableQA: A Benchmark for Health Reasoning over Real-World Wearable Data
Recent advances in wearable sensing enable continuous monitoring of physiological and behavioral signals, yet existing b…
研究前沿 HuggingFace Daily Papers 9-4 阅读 6 · 访客 3
Stanford Researchers Release Paper2Agent: Turning Research Papers Into AI Agents That Reproduce Results and Run on New Data
Paper2Agent, published in Nature, converts papers into validated MCP tools, scoring 91.2% on 300 questions across 74 pap…
智能体 MarkTechPost 今天 阅读 1 · 访客 1
Modality-Autoregressive World-Action Models
World-action models (WAMs) jointly model future observations and actions, typically predicting the future as RGB images.…
行业动态 HuggingFace Daily Papers 2天前 阅读 2 · 访客 1
PhysStream: Streaming Physics-Grounded Video Generation with Structured Scene Memory and Fine-Grained Motion Control
Interactive control for video generation is moving from coarse prompts toward fine-grained, physically meaningful manipu…
行业动态 HuggingFace Daily Papers 2天前 阅读 1 · 访客 1
FLAT: Resampling Image and Text into 1D Flexible-Length Aligned Transmodal Tokens for Retrieval and Generation
Traditional multimodal representation learning and generation are two stages: a contrastive or self-supervised visual en…
行业动态 HuggingFace Daily Papers 2天前 阅读 1 · 访客 1
Emergence World: Adversarial Stress-Testing of Long-Horizon Multi-Agent Systems
As AI agents move from bounded tasks to persistent deployments, failures can propagate through memory, tools, other agen…
智能体 HuggingFace Daily Papers 2天前 阅读 2 · 访客 2
AI for Games in the Foundation Model Era
Foundation models, alongside advances in learned game-world models, are reshaping AI across the game lifecycle. Beyond p…
行业动态 HuggingFace Daily Papers 2天前 阅读 2 · 访客 2
ScienceBuddy: Recursive-in-Recursive Self-Improvement for Interactive Scientific Agents
We introduce and release ScienceBuddy, an interactive scientific research workspace that brings continually improving sc…
智能体 HuggingFace Daily Papers 2天前 阅读 1 · 访客 1
ImpossibleRubrics: Stress-Testing Generated Rubrics as Reward Signals
Language model-generated rubrics are increasingly used as reward signals for rubric-based reinforcement learning, LLM-as…
大模型 HuggingFace Daily Papers 2天前 阅读 1 · 访客 1
Decoy Direction Optimization: A Post-Hoc Defense Against LLM Abliteration
Safety guardrails in open-weight language models can be readily bypassed using Refusal Feature Ablation (RFA), a techniq…
大模型 HuggingFace Daily Papers 3天前 阅读 4 · 访客 3
RSIAgent: Autonomous Exploration for Recursive Self-improvement in New Environments
Digital agents must often adapt to new environments whose interfaces, tools, and failure modes are not fully captured by…
智能体 HuggingFace Daily Papers 3天前 阅读 4 · 访客 4
PhysBrain 1.5: From Vision-Language Models to Physical Foundation Models
We present PhysBrain 1.5, a unified model for understanding physical environments, generating actions, and predicting fu…
行业动态 HuggingFace Daily Papers 3天前 阅读 3 · 访客 3
Not All Prompts Are Equal: Exploration-Guided Prompt Scaffolding for Multimodal Reinforcement Post-Training
Training prompts in online reinforcement learning (RL) differ substantially in how informative they are for the current …
行业动态 HuggingFace Daily Papers 3天前 阅读 3 · 访客 3
ModularRSI: Modular and Generalizable Recursive Harness Self-Improvement
Recent work extends recursive self-improvement (RSI) to agent harnesses for long-horizon coding and terminal tasks, enab…
智能体 HuggingFace Daily Papers 3天前 阅读 3 · 访客 3
How Lossless Is Lossless Speculative Decoding? The Role of Numerical Precision in Orthrus
Orthrus is a hybrid autoregressive-diffusion architecture that accelerates autoregressive language-model inference by ge…
行业动态 HuggingFace Daily Papers 3天前 阅读 4 · 访客 3
Kaininja: Extending Native 3D Generators to the Part Level
Native 3D generators turn one image into a single mesh. TRELLIS.2 and its peers deliver high-fidelity non-watertight geo…
行业动态 HuggingFace Daily Papers 3天前 阅读 3 · 访客 2
BVB: Benchmarking Agentic Video Understanding via Programmatic Reconstruction in Blender
Multimodal agents can create complex videos in software such as Blender by coding without relying on diffusion models. Y…
智能体 HuggingFace Daily Papers 3天前 阅读 4 · 访客 3
The Router Within: Eliciting Native Skill Routing from a Frozen LLM
Skills extend an LLM agent beyond its parametric knowledge, and the gain they promise rests on picking the right one. De…
智能体 HuggingFace Daily Papers 3天前 阅读 1 · 访客 1
Disentangling Representation Evolution in Transformers through Directional Decomposition
Transformer representations evolve through learned additive transformations that either preserve their current direction…
行业动态 HuggingFace Daily Papers 3天前 阅读 2 · 访客 2
Discovery Foundation Models: Toward Open-Ended Discovery Intelligence
Foundation models have progressed from learning and reasoning over existing knowledge, to increasingly learning through …
研究前沿 HuggingFace Daily Papers 3天前 阅读 3 · 访客 2
Pick Your Poison: Learning to Select Poison Sets for Stronger LLM Backdoor Attacks
Backdoor poisoning attacks add poisoned examples to otherwise-clean finetuning data, pairing a trigger with a target beh…
大模型 HuggingFace Daily Papers 3天前 阅读 2 · 访客 2
ModaLens: Measuring Image Sensitivity in Report-Conditioned Medical VLMs
A radiology report can already answer a clinical question, so it is hard to tell whether a vision-language model also us…
行业动态 HuggingFace Daily Papers 3天前 阅读 1 · 访客 1
Enabling Creative Exploration for Vibe Design Agents
Vibe design agents turn natural-language briefs into rendered interfaces and frontend code. Yet a useful design agent sh…
智能体 HuggingFace Daily Papers 3天前 阅读 3 · 访客 2
When Agents Slow Down: Understanding LLM Agents' Test-Time Strategies via Elo-per-token Analysis
Large language model (LLM) agents allocate test-time compute adaptively as they revise solutions, use tools, explore alt…
智能体 HuggingFace Daily Papers 3天前 阅读 2 · 访客 1
Mind2Dialogue: Training Human-Aware Language Models by Simulating User Mental States
As language models become more capable, long-term collaboration in learning, reasoning, and decision-making calls for a …
研究前沿 HuggingFace Daily Papers 3天前 阅读 3 · 访客 3
LynnReal-Omni: Native multi-modal Video Generation for Agentic Visual Workflows
Video diffusion models are stochastic and hard to control: precise content often requires repeated sampling without guar…
智能体 HuggingFace Daily Papers 3天前 阅读 3 · 访客 3
Dream-RSI: Recursive Self-Improvement through Evolving Worlds
Recursive self-improvement is becoming increasingly vital for autonomous AI agents, where progress hinges on discovering…
智能体 HuggingFace Daily Papers 3天前 阅读 2 · 访客 2
HarnessVLN: Unifying Training-Free Embodied Navigation through an Agent Harness
Embodied navigation requires agents to interpret visual observations, accumulate spatial knowledge, and execute actions …
智能体 HuggingFace Daily Papers 3天前 阅读 2 · 访客 2
HazardAuditor: From Executable Threats to Safer Computer-Use Agents
Computer-use agents increasingly interact with browsers, terminals, file systems, and external services, introducing saf…
智能体 HuggingFace Daily Papers 3天前 阅读 2 · 访客 2
Omni-Streaming Thinking
Streaming omni-modal models must decide what and when to answer from the video chunks and synchronized audio observed so…
行业动态 HuggingFace Daily Papers 3天前 阅读 2 · 访客 2
Atria Dawn: The Dawn of Agentic Superintelligence
As AI agents become participants in the development of their successors, they reshape both the production of intelligenc…
智能体 HuggingFace Daily Papers 3天前 阅读 1 · 访客 1
AlayaVista: Streaming World Modeling from Panoramic States to Perspective Video
Interactive video world models must maintain broad scene context under camera motion while producing high-fidelity obser…
行业动态 HuggingFace Daily Papers 4天前 阅读 3 · 访客 3
Another Blueprint In The Wall: How to Ask Frontier AI Like a Kid?
This paper reports experiments across six frontier model types from OpenAI, Anthropic, xAI, and Google DeepMind. Ten ind…
研究前沿 HuggingFace Daily Papers 4天前 阅读 2 · 访客 2
Lightning Weave: Improving the Accuracy-Efficiency Frontier of Reasoning Models through Capability Composition
A core goal of efficient reasoning is to improve the accuracy-efficiency frontier. However, jointly improving reasoning …
研究前沿 HuggingFace Daily Papers 4天前 阅读 3 · 访客 3
OmniHarness: Harnessing Generalizable Visual Generation via Symbolic Policy Learning
Unified multimodal large language models (MLLMs) and multi-agent systems have advanced visual generation. However, three…
智能体 HuggingFace Daily Papers 4天前 阅读 2 · 访客 2
E2A-Bench: Benchmarking Evidence-to-Action Reliability in Financial Chart Reasoning
Can financial vision-language models (VLMs) turn chart evidence into reliable action recommendations? Existing hallucina…
研究前沿 HuggingFace Daily Papers 4天前 阅读 1 · 访客 1
Thought without systematicity? Evaluating reasoning models on rule induction tasks
A central tenet of human cognition is systematicity, the principle that understanding one concept is inherently tied to …
研究前沿 HuggingFace Daily Papers 5天前 阅读 3 · 访客 3
Training Specialist Models without Reasoning Trajectories for Domain Expert Distillation
Specialist distillation effectively transfers domain expertise to student models via teacher-generated reasoning traject…
研究前沿 HuggingFace Daily Papers 5天前 阅读 0 · 访客 0
StepAudio 3 Realtime Technical Report
Realtime spoken interaction demands deep reasoning, prompt responses, and fluid turn-taking. We present StepAudio 3 Real…
研究前沿 HuggingFace Daily Papers 5天前 阅读 3 · 访客 3
Drift-Constrained Optimization: Only Direction Matters in Fine-Tuning Instruct Models
Fine-tuning instruct models often improves target performance while inducing behavioral drift from the reference model, …
行业动态 HuggingFace Daily Papers 5天前 阅读 0 · 访客 0
Convergent Emergence of In-Context Learning Across Modalities
Few-shot in-context learning (ICL), the capacity of a model to infer abstract patterns from input-output examples provid…
行业动态 HuggingFace Daily Papers 5天前 阅读 3 · 访客 3
Dynin-Robotics: Omnimodal Unified Diffusion Vision-Language-Action Model
Visual goal and dynamics prediction can provide language-conditioned robot policies with both a target outcome and a rep…
智能体 HuggingFace Daily Papers 6天前 阅读 2 · 访客 2
StepAudio 3 Music Technical Report
We introduce StepAudio 3 Music, a large-scale, long-form music generation model that supports explicit musical planning …
行业动态 HuggingFace Daily Papers 6天前 阅读 1 · 访客 1
Generalized Agent Iteration: One Formal Framework for Iterative Policy Improvement and Recursive Self-Improvement
When we speak of recursive self-improvement (RSI), are we speaking of a phenomenon, a mechanism, or a prospect? Towards …
智能体 HuggingFace Daily Papers 6天前 阅读 2 · 访客 2
Breaking the Vision-Action Shortcut: Latent Interface Training for Generalizable Robotics Foundation Models
Robot foundation models achieve strong in-distribution performance but often degrade under visual distribution shifts. W…
智能体 HuggingFace Daily Papers 6天前 阅读 2 · 访客 2
Expert-Space Exploration in MoE Reinforcement Learning
Reinforcement learning (RL) has become central to post-training of large language models. Recent advances in RL for Mixt…
行业动态 HuggingFace Daily Papers 6天前 阅读 2 · 访客 2
ZGCM-1: A Fully Open and Extremely Efficient Foundation Model for Math and Agentic Search
In this work, we present ZGCM-1, a fully open 7B dense foundation model trained from scratch with extreme data, system, …
智能体 HuggingFace Daily Papers 6天前 阅读 2 · 访客 1
StepAudio 3 Gen Technical Report
We introduce StepAudio 3 Gen, a general-purpose audio generation model that supports zero-shot text-to-speech (TTS), voi…
行业动态 HuggingFace Daily Papers 6天前 阅读 3 · 访客 3