AI 资讯站上线:内容开源 + 本地索引检索
本站(alishangtian.com/ainews)正式上线:纯静态架构,内容以 JSON 形式开源存储于 GitHub 仓库,基于预构建倒排索引实现纯浏览器端的本地全文检索。
阅读全文 →
行业动态
共 80 条TransNormal-2: Geometry-Grounded Rectified Flow with Edge-Aware Decoding for Precise Normal Estimation
Diffusion-based models enable monocular geometry estimation, yet their pixel-space precision is limited by a shared, und…
Multi-Grid Post-Training for Long-Form Multi-Shot Video Generation
Generating long-form multi-shot videos requires coherent within-shot motion and visually consistent narratives across sh…
DriveZero: End-to-End Driving Beyond Human Demonstrations
Most end-to-end autonomous-driving systems learn by imitating human driving logs, leaving their learned behavior constra…
WorldSculpt: Generating Compositional Worlds from Grounded Videos
We study the problem of generating a compositional 3D representation of a cluttered scene containing hundreds of objects…
Knowing What Not to Answer: Selective Non-Compliance in Vision-Language Models
Vision-language models (VLMs) are expected to respond helpfully to appropriate requests while withholding compliance wit…
Refuse without Refusal: A Structural Analysis of Safety-Tuning Responses for Reducing False Refusals in Language Models
Striking a balance between helpfulness and safety remains a fundamental challenge in aligning large language models. To …
Scal3R: Learning Efficient Multi-Relative Pose Query for Scalable Online 3D Reconstruction
Online 3D reconstruction models perform poorly on long videos. This happens because regressing poses relative to a fixed…
When Quantization Breaks Memory: Recurrent-State Write-Back in Low-Precision Temporal Inference
Quantization is widely used to reduce the computational and memory demands of neural-network inference. In recurrent net…
The Attention Triangle in Audio-Video Models
Audio-video diffusion models rely on cross-modal attention to coordinate text, sound, and visual content, yet this same …
The 2026 PNPL Competition: Word Classification and Efficient Cross-Subject Generalisation in LibriBrain100
The ambition of the 2025 PNPL competition (Landau et al., 2025) was to launch a multi-year curriculum for non-invasive s…
Compile by Training: Turning Natural-Language Specifications into Local Neural Functions
Many recurring text functions are easy to describe but difficult to implement with rules, while calling a large remote m…
One Editor, Many Edits: A Unified Training-Free Framework for Diverse Video Editing
Video editing spans diverse editing paradigms, yet achieving high-quality instruction-guided and subject-guided editing …