SiliconBench: Speed, Memory, and Fidelity for LLM Serving on Unified-Memory Desktops
Concurrent local LLM serving on unified-memory desktops must preserve memory headroom and output fidelity, which speed-o…
大模型
HuggingFace Daily Papers · 9-12
阅读 18·访客 17