搜索:GQA

共命中 2 条(服务端检索)
Mistral 7B 开源:小模型的高效革命
法国初创公司 Mistral AI 以 Apache 2.0 协议开源 7B 模型,以更小参数量比肩 Llama 2 13B,GQA 分组查询注意力等设计影响深远。
开源项目 Mistral AI 精选 · 2023-11-21 阅读 2 · 访客 0
Grouped Value Attention: Efficient KV Caching via On-Demand Key Reconstruction
The KV cache is a primary bottleneck for Transformer decoding: its memory footprint and cache-read traffic grow with seq…
智能体 HuggingFace Daily Papers 9-8 阅读 0 · 访客 0