搜索:FLOPS

共命中 13 条(服务端检索)
算力的度量衡:从 FLOPS 到卡时,看懂算力新闻的单位
FLOPS、算力、卡时、集群规模混着说,是看懂算力新闻的第一道坎。本文厘清 FLOPS 与 FLOPs 一字之差的两个概念,给出 6ND 训练算力估算经验与卡时换算方法,解释峰值与有效算力之间的 MFU 差距,最后附一张看懂算力新闻的换算清单。
原创 行业动态 精选 · 原创 · 昨天 阅读 1·访客 1
从 Transformer 到今天:注意力架构的十年演进地图
2017 年的 Transformer 之后,架构研究沿三条主线展开:把注意力做便宜、把注意力换掉、把 FFN 做稀疏。本文梳理稀疏注意力、FlashAttention、SSM、混合架构与 MoE 的脉络,并给出读新架构论文的三问。
原创 研究前沿 精选 · 原创 · 昨天 阅读 0·访客 0
Decoding Looped Transformers Better for (Almost) Free
Looped Transformers achieve parameter efficiency by repeatedly executing a shared block across recurrent loops. Each loo…
行业动态 HuggingFace Daily Papers · 6天前 阅读 5·访客 5
How Far Are We from Removing the Visual Encoder? Scaling Laws for Encoder-Free Multimodal Pretraining
Most modern multimodal large language models (MLLMs) build on a pretrained visual encoder that provides a strong visual …
大模型 HuggingFace Daily Papers · 9-28 阅读 19·访客 19
A Coding Guide to Google Research’s MSEB: Writing Sound Encoders to the Benchmark Contract and Scoring Them Across Classification, Clustering, Retrieval and Segmentation
In this tutorial, we work with **MSEB**, the Massive Sound Embedding Benchmark from Google Research, and approach it fro…
研究前沿 MarkTechPost · 9-27 阅读 30·访客 29
TT-VidT: Decoupling the Temporal Axis for Efficient Motion-Centric Video Pretraining
Comparisons in video self-supervised learning often evaluate complete training recipes rather than isolating the method …
行业动态 HuggingFace Daily Papers · 9-27 阅读 5·访客 5
罗福莉官宣小米 MiMo-V3 采用全新架构,核心 HySparse 2 今日发布
IT之家 9 月 23 日消息,小米 MiMo 大模型负责人罗福莉今日发文,宣布 MiMo-V3 即将采用全新架构。其核心 HySparse 2 今日发布,带来更少的预填充、更小的 KV 缓存、更出色的长上下文检索。 在 1M(100 万)…
大模型 IT之家 · 9-23 阅读 16·访客 16
DeltaWAM: Delta World Action Models for Bimanual Manipulation
World-action models (WAMs) transfer visual and motion priors from pretrained video generators to robot control by jointl…
智能体 HuggingFace Daily Papers · 9-23 阅读 11·访客 11
Neural Spectral Capacity: Measuring and Designing Architectures from Network Specification Alone
Modern Transformer design and compression both reduce to allocating capacity under a budget. The standard scalars for th…
智能体 HuggingFace Daily Papers · 9-19 阅读 6·访客 6
Inside NVIDIA’s cuDNN Graph API: Fusion, Autotuning, and Plan Reuse with cuDNN Frontend
Learn how to leverage NVIDIA’s cuDNN Frontend Graph API to build custom kernel fusions, autotuning engine configurations…
行业动态 MarkTechPost · 9-16 阅读 14·访客 14
Modality-Autoregressive World-Action Models
World-action models (WAMs) jointly model future observations and actions, typically predicting the future as RGB images.…
行业动态 HuggingFace Daily Papers · 9-15 阅读 12·访客 11
Meta新研究:字节模型蒸馏后,天花板破了
]
研究前沿 量子位 · 9-15 阅读 20·访客 20
大模型能力提升路线图:从"堆参数"到训练全栈 + 外层程序
把 2026 年可核查的公开证据整理成一张六层能力路线图——预训练、后训练 RL、推理时计算、上下文与记忆、智能体与 Harness、世界模型。含 Meta ScaleRL 40 万 GPU 小时实验结论、RL 预算占比 10%–30% 口径、Chinchilla 对比、Meta-Harness 6x 差距等数据锚点,并给出优先级表与算法工程师/产品经理的行动建议。
原创 大模型 精选 · 本站原创 · 9-10 阅读 81·访客 62