搜索:VLM

共命中 35 条(服务端检索)
In-Context Robot Learning with VLM Agents
Enabling robots to adapt to unfamiliar environments as readily as humans remains a moonshot goal of embodied AI. No fini…
智能体 HuggingFace Daily Papers · 9-16 阅读 20·访客 19
Show-Harness: Just a VLM Agent Can Play Robots
Foundation vision-language models (VLMs) exhibit broad intelligence about the world, yet translating this intelligence i…
智能体 HuggingFace Daily Papers · 9-9 阅读 21·访客 19
一篇读懂视觉语言模型:图像是怎么变成「语言」的
大语言模型只认 token 序列,图片如何进入对话?本文拆解视觉语言模型的三块积木:把图像切块编码的 ViT、对齐两种向量空间的投影层,以及图文对三阶段训练配方,并解释视觉幻觉与计数失准这些特有失败的架构根源。
原创 一叶一世界 精选 · 原创 · 昨天 阅读 0·访客 0
Fewer Tokens, Better Action: GPT-6 Astra Robot Agents with 14% Higher Success Rate but 65% Fewer Tokens
Vision language model (VLM) agents can control robots through visual feedback and action primitives, but repeated model …
智能体 HuggingFace Daily Papers · 6天前 阅读 2·访客 2
AnswerMap: Faithful Spatial Interpretability of VLMs from Answer Posteriors
When a VLM answers a visual query, current interpretability tools rely on text rationales, which use a mismatched modali…
行业动态 HuggingFace Daily Papers · 9-28 阅读 16·访客 15
刚刚,GPT-6 Astra接上宇树G1,把厨房收拾了!
henry* 2026-09-30 15:54:54 来源:量子位 让GPT把机器人技能当工具调用 henry 发自 凹非寺 量子位 | 公众号 QbitAI 机器人的GPT时刻,还真得靠GPT(doge)! 刚刚,GPT-6 Astra…
智能体 量子位 · 9-30 阅读 4·访客 4
工业创新进入“组队局”,拆解西门子Xcelerator开放生态的赋能链路
田, 晏林* 2026-09-28 20:43:53 来源:量子位 工业平台已经卷到帮伙伴拿线索、做Agent、出海了 田晏林 发自 凹非寺 量子位 | 公众号 QbitAI 一家技术公司想把一个工业创新方案卖进工厂,中间大概要过多少关?…
智能体 量子位 · 9-28 阅读 24·访客 24
Google Research Introduces an AI Video Co-Director: 4 Agentic Frameworks for Coherent, Minutes-Long Video Generation
Google Research has introduced an **AI video co-director** for long-form video generation. The suite of 4 agentic framew…
智能体 MarkTechPost · 9-28 阅读 22·访客 22
Researchers plug GPT-6 Astra directly into a robot and let it clean up an unfamiliar kitchen
Sep 27, 2026 Researchers from Stanford and Caltech have built HomeBody,** a system that lets a Unitree G1 robot autonomo…
智能体 The Decoder · 9-27 阅读 43·访客 41
Liquid AI Releases LFM2.5-VL-3B-DSpark: Speculative Decoding for Vision-Language Models With Up to 3.13x Faster Decoding
Liquid AI has announced LFM2.5-VL-3B-DSpark, an experimental speculative-decoding draft model for its LFM2.5-VL-3B visio…
行业动态 MarkTechPost · 9-26 阅读 27·访客 27
AI开始研究Physical AI:FSD级团队亮出首版模型Simate-beta,空降RoboDojo
田, 晏林* 2026-09-26 17:07:41 来源:量子位 Simate将训练、推理与评测全流程接入自研Infra,通过极致的任务编排与资源调度,同时并行推进数十条相互独立的研究路线。 Jay 发自 凹非寺 量子位 | 公众号 Q…
研究前沿 量子位 · 9-26 阅读 22·访客 22
InternW0-Δ: A World Action Model Bridging Predictive Dynamics and Actions with 20K+ Hours of Open Data
World Action Models (WAMs) jointly model visual dynamics and action generation for generalist robot manipulation. A cent…
智能体 HuggingFace Daily Papers · 9-25 阅读 9·访客 9
早报|GPT-6 Sol发布,价格腰斩/特努斯:Siri AI不应代替人际关系/4999起,OPPO Find X10系列发布
🧑‍💻GPT-6 Sol 与 Luna 发布,API 价格降低 50% 🧠Claude Opus 5.5 发布,典型任务运行成本降低 40% 🍎特努斯:Siri AI 不应代替人际关系 📱OPPO Find X10 系列发布:49…
大模型 爱范儿 · 9-23 阅读 27·访客 27
HARMONY: Hierarchical Agentic Reasoning for MONocular Image-to-Scene Synthesis
Compositional 3D scene reconstruction has recently been explored from two directions: agentic reasoning that provides se…
智能体 HuggingFace Daily Papers · 9-22 阅读 10·访客 10
RoboFollow: Unveiling the Instruction Following Mirage in Embodied Agents
Modern embodied agents achieve impressive success rates, yet their actual instruction-following ability is far weaker th…
智能体 HuggingFace Daily Papers · 9-22 阅读 8·访客 8
X-Planner: Event-Structured Task Planning for Embodied Intelligence
Task planning bridges high-level instructions and executable behavior in long-horizon manipulation, yet modern Vision-La…
行业动态 HuggingFace Daily Papers · 9-21 阅读 2·访客 2
RULER: Instance-aware Rubric Rewards for SVG Generation
Generating Scalable Vector Graphics (SVG) code from natural-language instructions is an open-ended task without absolute…
智能体 HuggingFace Daily Papers · 9-21 阅读 12·访客 12
All-in-One Multilingual Scene Text Recognition with Script-aware Mixture-of-Experts
Multilingual scene text recognition (STR) remains challenging due to the scarcity of training data for most languages an…
行业动态 HuggingFace Daily Papers · 9-21 阅读 11·访客 11
PackLab: A Comprehensive Framework for Developing, Training, and Evaluating MLLMs in Robotic Bin Packing
Robotic bin packing requires long-horizon sequential decision-making, as each object placement affects the available spa…
智能体 HuggingFace Daily Papers · 9-20 阅读 11·访客 10
Transferring the Intelligence of VLMs to Robotic Control
Humans can seamlessly adapt to both physical and digital worlds, suggesting that while a digital-to-real gap exists in e…
智能体 HuggingFace Daily Papers · 9-19 阅读 7·访客 7
Paint-Anything: Unified Any-Color Control for Image Generation and Editing
Professional design requires any-color control: the ability to specify an object's target color with any 24-bit hex valu…
行业动态 HuggingFace Daily Papers · 9-17 阅读 14·访客 14
PANORAMA: Panoptic Grounded Captioning via Mask Proposal Selection
Intelligent systems that act in the world require image understanding that is both comprehensive and spatially grounded.…
智能体 HuggingFace Daily Papers · 9-16 阅读 13·访客 13
RiskChainBench: A Benchmark for Obfuscated Platform Message Restoration and Evidence-Grounded Web Investigation
Platform abuse campaigns conceal redirection instructions with emojis, homophones, character decomposition, and redundan…
研究前沿 HuggingFace Daily Papers · 9-15 阅读 15·访客 13
无问芯穹联合清华、上交正式开源具身端侧推理引擎APXInf,Pi 0.5性能SOTA
卡位具身智能规模化落地“最后一公里”!]
开源项目 量子位 · 9-15 阅读 23·访客 22
PhysBrain 1.5: From Vision-Language Models to Physical Foundation Models
We present PhysBrain 1.5, a unified model for understanding physical environments, generating actions, and predicting fu…
行业动态 HuggingFace Daily Papers · 9-14 阅读 13·访客 13
中国物理AI大突破:PhysBrain 1.5登顶全球开源榜一,空间智能与GPT-6 Astra并驾齐驱
这家中国公司,刚跑完了物理闭环里最难的一段路]
开源项目 量子位 · 9-14 阅读 20·访客 18
E2A-Bench: Benchmarking Evidence-to-Action Reliability in Financial Chart Reasoning
Can financial vision-language models (VLMs) turn chart evidence into reliable action recommendations? Existing hallucina…
研究前沿 HuggingFace Daily Papers · 9-13 阅读 13·访客 13
RelateAnything: Real-Time Open-Vocabulary Relation Prediction From Any Inputs
Open-vocabulary detection accepts any class list at inference, and promptable segmentation returns regions without class…
行业动态 HuggingFace Daily Papers · 9-11 阅读 15·访客 15
Attention-DP3: Spatially Object-aware 3D Diffusion Policy via Geometry-aligned Attentional Conditioning
3D point-cloud observations are inherently ambiguous in complex, cluttered manipulation scenes, where target objects may…
行业动态 HuggingFace Daily Papers · 9-10 阅读 11·访客 10
Feature Recovery for Object Understanding After Irreversible Fire Damage
Objects in post-fire environments often undergo irreversible physical transformations that change their geometry, materi…
行业动态 HuggingFace Daily Papers · 9-10 阅读 11·访客 11
Mi-Ripple: Restoring Images Degraded by Iterative AI Editing
Iterative reference-conditioned image editing can introduce grid-like and granular textures, commonly described as digit…
行业动态 HuggingFace Daily Papers · 9-10 阅读 13·访客 9
Agentic Visual Generation: From Generative Models to Agentic Control
Visual generation is evolving from generative models used through a single invocation into agentic control processes tha…
智能体 HuggingFace Daily Papers · 9-6 阅读 13·访客 11
ShieldVLA: Feasibility-Aware Safety Alignment for Vision-Language-Action Models
Vision-Language-Action (VLA) models demonstrate strong generalization in robotic manipulation and navigation, but existi…
智能体 HuggingFace Daily Papers · 9-2 阅读 1·访客 1
SimpleMemVLA: A Simple but Effective Native-Video Memory for Vision-Language-Action Models
Long-horizon manipulation is partially observable: the information needed to choose the next action may appear only in o…
行业动态 HuggingFace Daily Papers · 9-2 阅读 14·访客 13
Training-Free Speech-Centric Omni Understanding with Frozen VLMs
Audio-visual understanding remains challenging because models must jointly interpret spoken content, visual events, and …
行业动态 HuggingFace Daily Papers · 8-7 阅读 7·访客 6