搜索:DiT

共命中 18 条(服务端检索)
一篇读懂 DiT:视频生成模型为什么都换上了 Transformer 主干
从 Peebles 与谢赛宁的 DiT 论文到 Sora 的时空 patch,讲清 Diffusion Transformer 的三个关键设计:patch 化、adaLN-Zero 条件注入与以计算量为标尺的可扩展性。
原创 一叶一世界 精选 · 原创 · 昨天 阅读 4·访客 4
开源语音合成现状:零样本克隆已经卷到什么程度
盘点 2026 年 10 月主流开源 TTS 七个项目(GPT-SoVITS、CosyVoice、F5-TTS、Fish Speech、IndexTTS、Kokoro 等):机制、音色克隆方式、中文支持与许可证商用限制,附中文效果/实时率/长文本对比表与 F5-TTS 上手示例,兼谈声音克隆的授权与深度伪造合规风险。
原创 开源项目 精选 · 原创 · 昨天 阅读 10·访客 9
Diffusion Policy:机器人动作生成为什么弃用回归、改用扩散模型
模仿学习的老问题是「多峰动作分布」——两种都对的做法被回归平均成一种错的。Diffusion Policy 用条件去噪扩散直接建模动作分布,在 12 个任务、4 个基准上平均成功率提升 46.9%,此后 DP3、DPPO 相继跟进,π0 的流匹配与 GR00T 的扩散 Transformer 把它推成了 VLA 时代的标配动作头。
原创 研究前沿 精选 · 原创 · 昨天 阅读 5·访客 5
VLA(视觉-语言-动作)模型进化史:机器人如何「看懂再动手」
从 RT-2 把机器人动作当文本 token 输出,到 OpenVLA 开源 7B、π0 用流匹配跑到 50Hz、GR00T 与 Helix 转向双系统架构——本文梳理 VLA 模型三年的演进脉络:动作怎么表示、参数怎么变小、控制频率怎么上去,以及开源与闭源两条路线的分野。
原创 研究前沿 精选 · 原创 · 昨天 阅读 3·访客 3
AI 视频的两道坎:角色一致性与物理合理性
生成一条好看的视频已经不难,难的是让同一个人跨镜头不「变脸」、让画面里的世界遵守物理。盘点各家的参考生视频、LoRA 微调、推理优先架构等方案与仍然存在的短板。
原创 大模型 精选 · 原创 · 昨天 阅读 3·访客 3
CtrlCache: Accelerating Interactive Video World Models with Control-Aware Caching
Interactive video world models need to generate each video chunk efficiently while responding faithfully to user control…
行业动态 HuggingFace Daily Papers · 2天前 阅读 0·访客 0
Learning to Read the Contextual Tokens in Diffusion Transformers
Multimodal Diffusion Transformers (MM-DiTs) jointly process visual and textual representations throughout generation. Th…
行业动态 HuggingFace Daily Papers · 3天前 阅读 0·访客 0
Does Native 3D Texture Generation Necessarily Require 3D Assets for Training?
Native 3D texture generation synthesizes colors directly in 3D space for a given geometry, conditioned on multi-view ref…
行业动态 HuggingFace Daily Papers · 9-28 阅读 3·访客 3
Structured Residual Connectivity Matters for Diffusion Transformers
Diffusion Transformers (DiTs) have established themselves as a scalable backbone for high-fidelity image synthesis. Howe…
行业动态 HuggingFace Daily Papers · 9-27 阅读 6·访客 6
FuseReg: Regularizing Layer Fusion Mitigates the Reconstruction-Generation Gap in Representation Autoencoders
Representation autoencoders (RAEs) reuse features from a pretrained visual encoder as reconstruction and diffusion laten…
行业动态 HuggingFace Daily Papers · 9-25 阅读 3·访客 3
PCIe显卡被低估了!内核补齐+通信重构,DeepSeek推理吞吐翻近7倍
思邈* 2026-09-24 22:17:48 来源:量子位 1.5台6000D跑赢1台B300! 允中 发自 凹非寺 量子位 | 公众号QbitAI 随着大模型对算力需求持续攀升,**GPU卡的采购成本、供给约束持续加剧。** AI推理…
开源项目 量子位 · 9-24 阅读 34·访客 33
Alibaba Qwen Releases Qwen-Image-2.1: A 7B Open-Weight Model for Image Generation and Editing
Alibaba's Qwen team has released Qwen-Image-2.1, a 7B diffusion transformer that handles text-to-image generation, multi…
大模型 MarkTechPost · 9-22 阅读 33·访客 33
UltraTex: Unleashing 2K Multi-View Diffusion for 3D Texturing
High-quality texture generation is essential for creating realistic and production-ready 3D assets. Recent multi-view di…
行业动态 HuggingFace Daily Papers · 9-19 阅读 8·访客 8
Video DeltaNet: A Video-Native Hybrid Attention for Livestream Video Generation
Video diffusion models repeatedly process long spatiotemporal token sequences during denoising, making attention a major…
智能体 HuggingFace Daily Papers · 9-17 阅读 17·访客 16
Training-Adaptive Convolutional Sparse Coding via Information Bottleneck for Robust Visual Representation
Visual signals require compact yet sufficient representations for robust downstream prediction. Convolutional sparse cod…
智能体 HuggingFace Daily Papers · 9-17 阅读 14·访客 14
StepAudio 3 Music Technical Report
We introduce StepAudio 3 Music, a large-scale, long-form music generation model that supports explicit musical planning …
行业动态 HuggingFace Daily Papers · 9-11 阅读 11·访客 11
吹爆开源!RunningHub让MiniMax H3满血提速12倍,本地部署照样起飞
15秒视频,50秒出片]
开源项目 量子位 · 9-11 阅读 89·访客 84
Marigold V2: Revisiting Diffusion Transformers for Monocular Depth Estimation
Monocular depth estimation is a ubiquitous yet highly ill-posed computer vision task, with downstream applications in sc…
智能体 HuggingFace Daily Papers · 9-8 阅读 13·访客 12