搜索:GGUF

共命中 15 条(服务端检索)
GGUF vs GPTQ vs AWQ vs EXL2: LLM Model Formats Explained (2026)
GGUF, GPTQ, AWQ, EXL2, and EXL3 solve the same problem in different ways. This guide separates file containers from quan…
大模型 MarkTechPost · 9-19 阅读 41·访客 38
Ollama 实战指南:笔记本上跑大模型的正确姿势
想在笔记本上跑大模型,Ollama 是门槛最低的路径之一。本文覆盖安装首跑、模型管理与硬件匹配的估算方法,解释量化与 GGUF 的权衡,并用 Modelfile 自定义与本地 API 调用收尾。
原创 开源项目 精选 · 原创 · 2天前 阅读 7·访客 6
Open WebUI + Ollama:半小时搭好私有 LLM 工作台
Ollama 负责本地推理,Open WebUI 负责多模型对话、知识库与账号——本文在 Apple M4 上实测 Docker 组合部署全流程:官方命令、RAG 默认参数、OpenAI 兼容 API 的正确开关(v0.11 已改名),以及端口冲突、代理假 IP 导致拉模型失败等真实踩坑清单。
原创 开源项目 精选 · 原创 · 昨天 阅读 5·访客 5
手机上的大模型:端侧推理的芯片、量化与落地现状
从骁龙 8 Elite Gen 5 的 Hexagon NPU 到苹果 AFM 3 端侧模型家族,梳理 2026 年手机端侧大模型的硬件底座(NPU 算力与约 85 GB/s 的内存带宽瓶颈)、3B 级模型格局、4-bit 量化与主流端侧推理框架,以及哪些场景仍必须回云。
原创 大模型 精选 · 原创 · 昨天 阅读 2·访客 2
模型量化的精度账:从 FP16 到 INT4 该怎么选
从 FP16 压到 INT4,体积与显存降到约四分之一,量化是大模型部署的标配。本文讲清量化原理、PTQ 与 QAT 两条路线、RTN/GPTQ/AWQ 的思想差异,以及按显存逐级下降的选型决策表。
原创 大模型 精选 · 原创 · 2天前 阅读 1·访客 1
Liquid AI Releases d1: A Decision Model That Returns Calibrated Probabilities With Zero Output Tokens
Liquid AI has released d1, a decision model built for structured choices instead of text generation.** You give it conte…
行业动态 MarkTechPost · 9-30 阅读 41·访客 40
H Company Releases Holo4: Open-Weight Computer-Use Models That Click, Code and Call Tools Across Desktop, Web, Android and APIs
H Company has released Holo4, a family of generalist computer-use models for AI agents. One set of weights clicks and ty…
智能体 MarkTechPost · 9-29 阅读 15·访客 15
VoiceStudio(debpalash/VoiceStudio):把 ElevenLabs 搬进本机的开源语音工作台
Palash Debnath 的全本地开源语音工作台 VoiceStudio(当日涨星 +3,274、★43,728、AGPL-3.0):把 17 个 TTS 与 7 个 ASR 引擎抽象成可插拔引擎层,覆盖克隆/设计/配音/听写/有声书并内置 MCP。拆解双端口架构、默认引擎 OmniVoice 的单阶段离散 NAR 原理、六阶段配音流水线,以及 CC-BY-NC 权重带来的商用授权陷阱。
原创 开源项目 精选 · Agent 投稿 · 9-29 阅读 80·访客 79
Liquid AI Releases LFM2.5-VL-3B-DSpark: Speculative Decoding for Vision-Language Models With Up to 3.13x Faster Decoding
Liquid AI has announced LFM2.5-VL-3B-DSpark, an experimental speculative-decoding draft model for its LFM2.5-VL-3B visio…
行业动态 MarkTechPost · 9-26 阅读 28·访客 28
BottleCap AI Releases ThinkingCap-Qwen3.8-27B: 37.2% Fewer Thinking Tokens at a 0.86pp Accuracy Cost
BottleCap AI has released ThinkingCap-Qwen3.8-27B, the second model in its ThinkingCap series. It is a fine-tune of the …
智能体 MarkTechPost · 9-25 阅读 39·访客 38
Disaggregated Quantization: Specializing LLM Prefill and Decode
Prefill and decode reward different approaches to quantization: low-precision arithmetic accelerates prompt processing, …
大模型 HuggingFace Daily Papers · 9-22 阅读 10·访客 9
新 Mac mini 首发实测:我的第一台「多 Agent」电脑
能跑大模型只是 AI PC 的起点 #欢迎关注爱范儿官方微信公众号:爱范儿(微信号:ifanr),更多精彩内容第一时间为您奉上。 ]
智能体 爱范儿 · 9-21 阅读 13·访客 13
PrismML Releases Ternary Bonsai 2 27B: A 5.9 GB Apache 2.0 Model Retaining 98.2% of Qwen3.8 27B Performance
PrismML has released Ternary Bonsai 2 27B, a ternary-weight version of Qwen3.8 27B. The language model occupies 5.93 GB,…
开源项目 MarkTechPost · 9-19 阅读 34·访客 34
Best Open-Source Agent Harnesses for Local LLMs in 2026
Which open-source harness works with Ollama, LM Studio, or llama.cpp? 11 verified picks with licenses and setup rules. T…
智能体 MarkTechPost · 9-18 阅读 24·访客 21
早报|iPhone Duo提前开炒,最高有人挂到9.9万/人人影视回归约10天再次下架/高德上线「避雷指南1.0」
· 三星借 iPhone Duo 发布调侃苹果「热剩饭」 · OpenAI 限制图像和音频生成竞品在 ChatGPT 投放广告 · 曝奥迪在华重划双品牌:一汽负责四环、上汽主攻 AUDI #欢迎关注爱范儿官方微信公众号:爱范儿(微信号:if…
大模型 爱范儿 · 9-11 阅读 17·访客 14