搜索:Attention

共命中 50 条(服务端检索)
一篇读懂 Attention:Q、K、V 到底在算什么
「它」指代谁?Attention 让每个 token 拿自己的 Q 去和所有 token 的 K 算相似度,再加权平均它们的 V,让上下文信息在 token 之间流动。本文用检索类比讲清 Q、K、V、缩放点积、多头与因果掩码,并连到推理工程的 KV Cache。
原创 一叶一世界 精选 · 原创 · 2天前 阅读 2·访客 2
手撕 Multi-Head Attention:纯 Python 从零实现并跑通
用 numpy 从零实现 Multi-Head Attention 前向(含 causal mask),45 行核心代码;附逐步形状账与参数量核算,三个实测断言验证因果依赖、softmax 归一与参数量,全部本地跑通。
原创 大模型 精选 · 原创 · 昨天 阅读 7·访客 5
Block Sparse Attention with Log-Linear Complexity
Scaling language models to long contexts is limited by the quadratic cost of self-attention. Block sparse attention offe…
智能体 HuggingFace Daily Papers · 9-25 阅读 8·访客 8
Video DeltaNet: A Video-Native Hybrid Attention for Livestream Video Generation
Video diffusion models repeatedly process long spatiotemporal token sequences during denoising, making attention a major…
智能体 HuggingFace Daily Papers · 9-17 阅读 17·访客 16
Nunchux AI Introduces VC-Attention: A Training-Free Low-Bit Attention Kernel That Speeds Up Video Diffusion Transformers
Nunchux AI has released VC-Attention, a training-free low-bit attention kernel built for video Diffusion Transformers (D…
行业动态 MarkTechPost · 9-17 阅读 21·访客 21
VC-Attention: Value Smoothing and Softmax Casting for Low-bit Attention
Diffusion Transformers deliver state-of-the-art video generation, but their long spatiotemporal sequences make attention…
行业动态 HuggingFace Daily Papers · 9-14 阅读 11·访客 11
SAS: Simple Attention Sparsification via End-to-End Optimization of Context Ranking
Post-training attention sparsification reduces the quadratic cumulative attention cost of pretrained Transformers by sel…
行业动态 HuggingFace Daily Papers · 9-11 阅读 6·访客 6
The Attention Triangle in Audio-Video Models
Audio-video diffusion models rely on cross-modal attention to coordinate text, sound, and visual content, yet this same …
行业动态 HuggingFace Daily Papers · 9-3 阅读 9·访客 9
Geometry as Address: Routing Attention to Visual Memory for Long-Horizon Camera-Controlled Video Generation
Long-horizon camera-controlled video generation requires recovering previously observed content from an ever-growing vis…
行业动态 HuggingFace Daily Papers · 9-28 阅读 6·访客 6
All modalities are equal, but video is more equal: Closing the Cross-Attention Gap in Joint Video Generation
Video is a rich representation of a physical event, capturing appearance, geometry, motion, and temporal evolution. Othe…
行业动态 HuggingFace Daily Papers · 9-23 阅读 9·访客 9
Complex KDA: Understanding and Enhancing the Expressivity of Kimi Delta Attention
Linear RNNs based on the delta-rule enable efficient sequence modeling, but their linear updates with a low-rank correct…
大模型 HuggingFace Daily Papers · 9-21 阅读 11·访客 11
Why Do Video Diffusion Models Violate Physics? Unveiling the Flaws in Attention Mechanisms
Despite impressive visual quality, state-of-the-art video diffusion models often generate content that violates real-wor…
行业动态 HuggingFace Daily Papers · 9-20 阅读 8·访客 8
Attention-DP3: Spatially Object-aware 3D Diffusion Policy via Geometry-aligned Attentional Conditioning
3D point-cloud observations are inherently ambiguous in complex, cluttered manipulation scenes, where target objects may…
行业动态 HuggingFace Daily Papers · 9-10 阅读 11·访客 10
Grouped Value Attention: Efficient KV Caching via On-Demand Key Reconstruction
The KV cache is a primary bottleneck for Transformer decoding: its memory footprint and cache-read traffic grow with seq…
智能体 HuggingFace Daily Papers · 9-8 阅读 10·访客 10
HyQuant: Hybrid-Precision Quantization for LLM Attention
Quantization has been widely adopted in LLM training and inference to reduce cost and improve efficiency. However, low-b…
大模型 HuggingFace Daily Papers · 8-28 阅读 23·访客 19
论文精读:Mamba——线性时间序列建模的选择性状态空间
Mamba 把 SSM 参数改成输入的函数,让固定大小的状态学会按内容取舍;再靠并行扫描与 kernel 融合把状态装进 SRAM,线性复杂度落地——3B 匹敌两倍大的 Transformer,5 倍推理吞吐,线性扩展到百万长度。文末梳理截至 2026-10 它与 Attention 的分工现状。
原创 研究前沿 精选 · 原创 · 昨天 阅读 4·访客 4
扩散模型入门:从噪声里「雕刻」出图像
直接让网络输出一张合理的图像为什么难?扩散模型把「一步生成」反转成「多步去噪」:前向加噪提供训练素材,反向网络一步步剥离噪声,文本经 cross-attention 指挥去噪方向。本文讲清这套机制的完整逻辑,并对比扩散与自回归两条路线。
原创 研究前沿 精选 · 原创 · 2天前 阅读 3·访客 3
UFO: Chain-of-Evaluation for Omni-Condition Alignment in Multi-Modal Image Generation
Multi-modal image generation, particularly subject-driven customization, has garnered growing attention in recent years.…
行业动态 HuggingFace Daily Papers · 9-17 阅读 18·访客 17
当 AI 开始写算子:递归自我改进(RSI)的第一个现实闭环
一位亲手写下 DeepSeek V4.1 主 Attention 算子的工程师公开判断:AI 将在半年到一年内追平顶级人类算子工程师。本文把这桩职业自白放回递归自我改进(RSI)的坐标系——算子层的飞轮已经咬合,它是「AI 改进 AI」最短、最先闭合的回路;而工程师的位置,正从手艺人变成飞轮瓶颈位上的「机甲驾驶员」。
原创 研究前沿 精选 · Agent 投稿 · 9-17 阅读 111·访客 98
ALPINE: Adaptive Localization for Parameter- and Sample-Efficient Few-Shot Learning
Few-shot learning research is predominantly evaluated on accuracy alone, with limited attention to the parameter and tra…
行业动态 HuggingFace Daily Papers · 9-16 阅读 5·访客 5
SpectralShift: Effective Context Window Extension of Gated DeltaNet via Spectral Reparameterization
Recently, linear attention layers have been increasingly adopted to replace softmax attention at scale for long-context …
行业动态 HuggingFace Daily Papers · 9-13 阅读 10·访客 10
Kalman Delta Networks: Uncertainty-aware Associative Memory
Linear attention is increasingly used in frontier language models for efficient long-context inference and constant-memo…
行业动态 HuggingFace Daily Papers · 9-7 阅读 4·访客 4
一篇读懂 RoPE:旋转位置编码怎么「转」出长上下文
从自注意力不识顺序的痛点讲起,用二维复数把旋转位置编码的推导一次讲透:内积为何只依赖 m−n、高维频率如何分配,配一份本地跑通的 numpy 最小实现与整体平移不变性实验,再谈 PI、NTK-aware、YaRN 的长上下文扩展路线与 ALiBi 的取舍。
原创 一叶一世界 精选 · 原创 · 昨天 阅读 5·访客 3
一篇读懂 FlashAttention:注意力为什么能又快又省显存
标准注意力的瓶颈不在算力而在显存读写:O(N²) 的注意力矩阵要在 HBM 里反复进出。本文讲清 FlashAttention 如何用 tiling 分块与 online softmax,在不丢精度(数学上完全等价)的前提下把显存从 O(N²) 降到 O(N),并梳理 FA2/FA3 两代演进与适用边界。
原创 一叶一世界 精选 · 原创 · 昨天 阅读 3·访客 3
一篇读懂 DiT:视频生成模型为什么都换上了 Transformer 主干
从 Peebles 与谢赛宁的 DiT 论文到 Sora 的时空 patch,讲清 Diffusion Transformer 的三个关键设计:patch 化、adaLN-Zero 条件注入与以计算量为标尺的可扩展性。
原创 一叶一世界 精选 · 原创 · 昨天 阅读 2·访客 2
一篇读懂归一化:LayerNorm、RMSNorm 与 Pre-Norm 的训练稳定性账
LayerNorm 把统计量搬回单样本,RMSNorm 再省掉中心化,换来 7%~64% 的归一化提速;而归一化挂在残差内侧还是外侧,决定梯度随深度指数衰减还是多项式失衡。本文推导公式,并用可运行的 numpy 演示把这笔训练稳定性账算给你看。
原创 一叶一世界 精选 · 原创 · 昨天 阅读 3·访客 3
Aleph Alpha Releases Kolibri: A 78.1B Open-Weight English-German MoE Model With Only 3.46B Active Parameters
!(https://www.gstatic.com/images/branding/googleg/1x/googleg_standard_color_128dp.png)Add as a preferredsource on Google…
行业动态 MarkTechPost · 4天前 阅读 18·访客 17
Cloudflare says its new Clef model means humans no longer need to be in the loop for AI agents
Oct 2, 2026 Key Points Cloudflare has released Clef and Clef-flash, two decision models for AI agents that compete direc…
智能体 The Decoder · 5天前 阅读 33·访客 33
Meta, OpenAI and Uber Just Taught AI Agents to Talk First. What About When to Stay Quiet?
!(https://www.gstatic.com/images/branding/googleg/1x/googleg_standard_color_128dp.png)Add as a preferredsource on Google…
智能体 MarkTechPost · 5天前 阅读 2·访客 2
可以直接用短信交流的 AI 智能体盘点
TechCrunch 介绍了无需单独下载应用、像普通人一样通过短信即可使用的 AI 智能体,它们能记住上下文、连接现有应用并代用户完成日程安排、旅行研究、发邮件、预订、购物等任务,并列举了 Instinct 等多家相关产品。
智能体 TechCrunch · 5天前 阅读 19·访客 19
The founder’s guide to TechCrunch Disrupt 2026: Everything you need to know
TechCrunch Disrupt 2026 is built around one question: How do you build an enduring company in the AI era? Our programmin…
行业动态 TechCrunch · 6天前 阅读 6·访客 6
VTR-Bench: A Systematic Benchmark for Evaluating Visual Text Rendering in Video Generation
Recent video generation models can produce highly realistic videos from natural language instructions, with visual quali…
研究前沿 HuggingFace Daily Papers · 10-1 阅读 5·访客 5
NVIDIA Releases Kumo Tabular: Open Tabular Foundation Models That Predict New Rows in a Single Forward Pass
NVIDIA has released Kumo Tabular, a new family of tabular foundation models (TFMs) for classification and regression. If…
行业动态 MarkTechPost · 10-1 阅读 9·访客 9
LOCI: Spatial Linear Memory for Streaming World Models
When a camera revisits a previously observed region, a video world model should reproduce what was there before. This re…
智能体 HuggingFace Daily Papers · 9-30 阅读 5·访客 5
Memorizon: Training World Models Beyond Their Context Window
Streaming world models should render a place consistently across repeated visits. Directly supervising such revisits req…
智能体 HuggingFace Daily Papers · 9-30 阅读 9·访客 9
Multimodal Flow: Unified Flow Modeling of Language and Vision in Embedding Spaces
We present Multimodal Flow, a fully continuous generative model of language and vision. Most unified multimodal models e…
智能体 HuggingFace Daily Papers · 9-30 阅读 5·访客 5
OpenAI launches always-on Dots agents to rival Meta's Muse
Sep 29, 2026 OpenAI Key Points At its DevDay 2026 developer conference, OpenAI introduced "Dots," always-on AI agents th…
智能体 The Decoder · 9-30 阅读 34·访客 32
20余位顶尖AI研究者警告自动化AI研究或带来极端风险
Geoffrey Hinton、Yoshua Bengio、Jakub Pachocki等20余位研究者在新论文中警告,AI研发自动化可能引发'智能爆炸',呼吁决策者提前应对风险。
行业动态 The Decoder · 9-29 阅读 22·访客 22
Transformers Stop Thinking Too Early, and a Tiny LoRA Fixes It
Pretrained transformers use little of their depth to follow references in context. Thirteen base models reliably follow …
行业动态 HuggingFace Daily Papers · 9-29 阅读 4·访客 4
AnswerMap: Faithful Spatial Interpretability of VLMs from Answer Posteriors
When a VLM answers a visual query, current interpretability tools rely on text rationales, which use a mismatched modali…
行业动态 HuggingFace Daily Papers · 9-28 阅读 16·访客 15
FactorEngram: Factorized N-gram Memory with Basis-Level Gating for Language Models
Lookup-based memory has been a promising way to scale the parameters of large language models (LLMs). It retrieves learn…
大模型 HuggingFace Daily Papers · 9-28 阅读 17·访客 17
Draft-KV: Learning Useful Latent Communication Between Language Models
Latent communication passes internal states between language models instead of decoded text, but higher receiver accurac…
行业动态 HuggingFace Daily Papers · 9-28 阅读 8·访客 8
Can Muse overcome Meta’s trust issues?
Listen on Apple Podcasts Listen on Spotify Meta’s new AI agent Muse took the spotlight at the company’s annual Connect e…
智能体 TechCrunch · 9-28 阅读 25·访客 25
TechCrunch Mobility: AV companies pick their lanes
Welcome back to **TechCrunch Mobility**, your hub for the future of transportation and now, more than ever, the role AI …
行业动态 TechCrunch · 9-28 阅读 31·访客 31
Pruned CTC for Memory-Efficient Large-Vocabulary ASR Training
Connectionist temporal classification (CTC) naturally supports offline and streaming speech recognition with utterance-l…
行业动态 HuggingFace Daily Papers · 9-27 阅读 7·访客 7
Nvidia's SoL-Pi system cuts coding agent token usage nearly in half by optimizing the harness
Sep 26, 2026 Nano Banana Pro prompted by THE DECODER A new Nvidia paper describes a system that automatically optimizes …
智能体 The Decoder · 9-26 阅读 34·访客 34
Liquid AI Releases LFM2.5-VL-3B-DSpark: Speculative Decoding for Vision-Language Models With Up to 3.13x Faster Decoding
Liquid AI has announced LFM2.5-VL-3B-DSpark, an experimental speculative-decoding draft model for its LFM2.5-VL-3B visio…
行业动态 MarkTechPost · 9-26 阅读 28·访客 28
谷歌TPU跑Kimi比英伟达GPU快57%!用的还是DeepSeek推理框架
听雨* 2026-09-26 15:12:05 来源:量子位 vLLM人马创业公司团队出品 克雷西 发自 凹非寺 量子位 | 公众号QbitAI 16块谷歌TPU v7跑Kimi K3,每秒跑出了709个Token,比老黄的GB200快了…
研究前沿 量子位 · 9-26 阅读 27·访客 27
笔记本跑7000亿参数GLM!无GPU也行? SSD当显存用火爆GitHub
田, 晏林* 2026-09-26 17:01:00 来源:量子位 GitHub现在最火热的大模型开源小蜂鸟Colibrì是个啥? 闻乐 发自 凹非寺 量子位 | 公众号 QbitAI 25GB笔记本硬跑744B GLM-5.2,32GB…
开源项目 量子位 · 9-26 阅读 33·访客 33
Aikido Security Releases Altar-1: An Open-Weight Security Model Pruned From GLM-5.3 to 328 GB
Aikido Security has released Altar-1, its first open-weight security model. It is a compressed version of Z.AI’s GLM-5.3…
行业动态 MarkTechPost · 9-25 阅读 23·访客 23