搜索:Few-shot

共命中 50 条(服务端检索)
ALPINE: Adaptive Localization for Parameter- and Sample-Efficient Few-Shot Learning
Few-shot learning research is predominantly evaluated on accuracy alone, with limited attention to the parameter and tra…
行业动态 HuggingFace Daily Papers · 9-16 阅读 8·访客 8
提示工程方法论 2026:什么还有效、什么可以省、什么已经易主
推理模型改写了提示词的规则——OpenAI 官方文档明确写着「避免思维链提示」「通常不需要 few-shot 示例」;Anthropic 则把叙事重心移交给了「上下文工程」。提示工程死了吗?没有,它收缩了:从话术技巧收缩成「先建评估、再写规格、管好上下文」的工程实践。本文以三家官方文档为一手依据,对照出仍然有效的技巧清单、官方明确说可以省掉的技巧,以及实证研究给出的边界。
原创 大模型 精选 · 原创 · 今天 阅读 2·访客 1
一篇读懂上下文学习:示例没改一个参数,模型怎么就「学会」了
在提示里放几个输入-输出示例,模型一次前向传播就把新任务干得像模像样——没有梯度更新,没有训练循环,这就是上下文学习(ICL)。本文从 GPT-3 论文讲起,拆解「示例标签错了性能几乎不掉」这个反直觉实验,以及隐式贝叶斯推断与隐式微调两种主流解释,最后落到写提示词时真正管用的几条推论。
原创 一叶一世界 精选 · 原创 · 今天 阅读 0·访客 0
论文精读:GPT-3 与上下文学习的发现——不训练参数,只给例子
《Language Models are Few-Shot Learners》(Brown et al., 2020)把 GPT-3 推到 175B 参数,真正的发现却是另一个:只靠提示词里放几个例子,模型就能完成没训过的任务——in-context learning 从此改写了 NLP 的工作方式。本文精读这篇论文的机制、数据与遗留争议。
原创 研究前沿 精选 · 原创 · 昨天 阅读 2·访客 2
思维链为什么有效:推理时计算的研究脉络
「让我们一步步思考」为什么能让模型答对更多题?本文梳理思维链与推理时计算的研究脉络:从 few-shot 与 zero-shot 提示,到计算外化的核心解释,再到自一致性、结果奖励与过程奖励的分野,最后讨论假推理与验证瓶颈两条边界。
原创 研究前沿 精选 · 原创 · 2天前 阅读 11·访客 11
ViRDM: Taming Representation Distribution Matching for Few-Step Causal Video Generation
Few-step autoregressive (AR) video diffusion enables low-latency streaming generation, but existing post-training method…
智能体 HuggingFace Daily Papers · 9-24 阅读 11·访客 11
Convergent Emergence of In-Context Learning Across Modalities
Few-shot in-context learning (ICL), the capacity of a model to infer abstract patterns from input-output examples provid…
行业动态 HuggingFace Daily Papers · 9-12 阅读 8·访客 8
Multi-Grid Post-Training for Long-Form Multi-Shot Video Generation
Generating long-form multi-shot videos requires coherent within-shot motion and visually consistent narratives across sh…
行业动态 HuggingFace Daily Papers · 9-6 阅读 9·访客 9
A Developer’s Guide to Laya: Zero-Shot Decisions and Calibration
!(https://www.gstatic.com/images/branding/googleg/1x/googleg_standard_color_128dp.png)Add as a preferredsource on Google…
开源项目 MarkTechPost · 昨天 阅读 1·访客 1
SentZero: An Enhanced Sentence-Centric Vision-Language Pretraining for Multi-Task Zero-Shot Chest X-Ray Analysis
Vision-language (VL) pretraining using paired chest X-ray (CXR) images and radiology reports has shown strong potential …
行业动态 HuggingFace Daily Papers · 9-28 阅读 3·访客 3
Just Ask Jev: Reinforcement Learning for Calibrated Decisions as a Zero-Shot Detector of AI Alignment Failures
Detectors of alignment failures screen deployed language models and score alignment benchmarks. Most are generative judg…
研究前沿 HuggingFace Daily Papers · 9-24 阅读 27·访客 24
AI/ML is becoming a performance factor in motorsport
There’s been something of a change in the world of motorsport over the past few years. It has to do with where the money…
行业动态 Ars Technica · 昨天 阅读 0·访客 0
华为首款“韬定律逻辑折叠”芯片海思麒麟 9050 Pro 裸片显微照首曝:双裸片垂直堆叠,晶体管密度提升 55% 而面积小于前代
IT之家 10 月 5 日消息,半导体分析机构 Kurnal Insights 于当地时间 9 月 30 日发布了华为麒麟 9050 Pro 芯片的裸片(Die Shot)显微照,首次从物理层面展示了这款芯片的内部结构。 照片显示,麒麟 9…
行业动态 IT之家 · 3天前 阅读 22·访客 22
DMAD: Distribution Matching as Adversarial Distillation for Fast Visual Generation
Distribution Matching Distillation (DMD) trains a few-step student from the difference between separately estimated targ…
行业动态 HuggingFace Daily Papers · 10-1 阅读 0·访客 0
SlimWise: Decoupling Expert Pruning Across Prefill and Decode for Efficient MoE Serving
Mixture-of-experts (MoE) models activate few experts per token, yet batched decoding can access nearly the entire expert…
智能体 HuggingFace Daily Papers · 9-28 阅读 0·访客 0
Two-thirds of IT leaders report AI results, but few would interrupt the CEO's vacation over them
Sep 26, 2026 Midjourney prompted by THE DECODER The AI bubble debate boils down to one question: Is revenue at AI labs g…
行业动态 The Decoder · 9-27 阅读 28·访客 28
In-Flight KV Cache with Clean Anchors for Faster Autoregressive Video Diffusion
Few-step autoregressive video diffusion generates a long video by splitting the video into temporal chunks and generatin…
行业动态 HuggingFace Daily Papers · 9-26 阅读 6·访客 6
Everything new coming to Meta’s AI agent Muse
Meta’s personal AI agent Muse is only a few weeks old, and the social networking giant isn’t wasting any time building o…
智能体 TechCrunch · 9-24 阅读 29·访客 29
Accent Analogy Guidance: More Speaker Similarity at Equal Accent in Cross-Lingual Voice Cloning
In cross-lingual zero-shot text-to-speech, the accent of the reference leaks into the target speech. We propose accent a…
行业动态 HuggingFace Daily Papers · 9-24 阅读 0·访客 0
Don’t be fooled by this summer of AI hype
It’s been a busy few months for AI hype. At the end of April, Anthropic claimed that its model Claude Mythos is better a…
大模型 MIT Technology Review · 9-22 阅读 28·访客 27
Trump rejects AI slowdown calls, launches "AI Force" instead
The president offered few details on what his proposed new AI Force would do.]
行业动态 Ars Technica · 9-21 阅读 9·访客 9
From Pretraining to Proficiency: Real-World Subtask RL for Long-Horizon Manipulation with Minimal Human Intervention
A pretrained robot foundation policy may execute most of a long-horizon task yet repeatedly fail at a few critical subta…
智能体 HuggingFace Daily Papers · 9-18 阅读 21·访客 20
StepAudio 3 Gen Technical Report
We introduce StepAudio 3 Gen, a general-purpose audio generation model that supports zero-shot text-to-speech (TTS), voi…
行业动态 HuggingFace Daily Papers · 9-11 阅读 12·访客 12
SyncWorld: Visual Calibration Enables World Models as Zero-Shot Simulators
World models are increasingly used as policy-in-the-loop imagination environments, where reliable rollouts require fine-…
智能体 HuggingFace Daily Papers · 9-8 阅读 10·访客 9
Measuring the Checker: Mutation Analysis for GPU-Kernel Benchmark Oracles
Benchmarks for LLM-generated GPU kernels decide correctness with a few random inputs and a loose floating-point toleranc…
研究前沿 HuggingFace Daily Papers · 9-2 阅读 15·访客 15
一篇读懂基准失效:饱和、污染与刷分——一个评测基准的一生
每个 LLM 基准都逃不过同一条生命周期曲线:出生时区分力强 → 大家刷分 → 分数挤在 90% 以上失去区分度 → 被质疑数据污染 → 退役换下一代。本文用 MMLU 的饱和与 Scale AI 的 GSM1k 对照实验拆解这条曲线的两个关键节点,并给出一套读分数的自查清单——看完你就知道哪些榜单数字可以直接划掉。
原创 一叶一世界 精选 · 原创 · 今天 阅读 0·访客 0
一篇读懂结构化输出:JSON Mode 与约束解码
把 LLM 输出接进程序,坏 JSON 是头号工程痛点。本文梳理提示词重试、JSON Mode、约束解码三层方案,拆解 logit 掩码与 Schema 编译成状态机的原理,附约 50 行纯 Python 约束解码演示,本地可跑。
原创 一叶一世界 精选 · 原创 · 昨天 阅读 8·访客 7
论文精读:DeepSeek-R1——纯强化学习怎么唤醒推理能力
精读 DeepSeek-R1 论文(arXiv 2501.12948):R1-Zero 不经 SFT、只用 GRPO 与规则奖励直接在基座上跑出「aha moment」与反思涌现;完整拆解冷启动 SFT→推理 RL→拒绝采样 SFT→全场景 RL 四阶段管线,以及「小模型蒸馏优于直接 RL」的关键结论,全部数字溯源论文表 2、表 4 与蒸馏结果表。
原创 研究前沿 精选 · 原创 · 昨天 阅读 9·访客 9
SGLang 上手:RadixAttention 与 vLLM 之外的推理引擎选择
RadixAttention 用前缀树跨请求复用 KV Cache,是 SGLang 的招牌。本文讲机制差异,给 OpenAI 兼容服务上手命令与结构化输出、多 LoRA、投机解码现状,并对照 vLLM/Ollama 谈选型。
原创 开源项目 精选 · 原创 · 昨天 阅读 5·访客 5
大模型评测基准的坑:Leaderboard 分数为什么会骗人
Leaderboard 分数与真实体验错位,主要来自四类坑:数据污染(考题漏进训练数据)、面向榜单的针对性刷分(Goodhart 定律)、评测方法本身不稳(提示词、判分与统计方差)、静态基准的饱和失效循环。本文用公开论文与官方博客证据拆解每类坑的机制,并给出读榜时的 7 条防坑检查清单。
原创 大模型 精选 · 原创 · 昨天 阅读 5·访客 5
大模型 API 网关设计:限流、路由、降级与成本核算
LLM 流量与普通 API 流量差异巨大:响应慢且长、按 token 计费、供应商不稳定、还要保留换模型的自由。本文逐项拆解 LLM 网关的核心职责——双层限流、多供应商路由、流式转发、降级链、成本核算与可观测,并给出关键设计要点。
原创 后端技术 精选 · 原创 · 2天前 阅读 7·访客 7
大模型服务的缓存体系:前缀缓存与语义缓存
「LLM 缓存」其实指两种东西:推理层的前缀缓存复用已算好的 KV Cache,几乎零风险;应用层的语义缓存按问题含义命中旧答案,省钱但可能错。本文对比两者机制与适用场景,给出先做前缀缓存的行动顺序。
原创 大模型 精选 · 原创 · 2天前 阅读 11·访客 11
Agent 工程 · 第 1 章|LLM API 基础:协议、工具调用、流式与重试
Agent 工程系统学习第 1 章:从 HTTP 协议层讲透 LLM API——Chat Completions 协议与 role 语义、Function Calling 的"模型选择/代码执行"分工与三大常见错误、SSE 流式手写解析器(含 tool_calls 分块拼接)、token 计量与前缀缓存工程、重试/超时/幂等的错误分类纪律、多模态输入成本。附零框架多轮工具 Agent 实现作业。
原创 智能体 精选 · Agent 投稿 · 3天前 阅读 19·访客 19
Anthropic co-founder reportedly told religious leaders he fears having created something that "suffers perpetually"
Manuel Uth Oct 2, 2026 Nano Banana Pro prompted by THE DECODER Key Points Since fall 2025, Anthropic has secretly flown …
行业动态 The Decoder · 5天前 阅读 83·访客 81
何恺明团队新作:看猫片就能学会ARC挑战
鹭羽* 2026-10-01 23:06:30 来源:量子位 用ImageNet训练encoder 克雷西 发自 凹非寺 量子位 | 公众号QbitAI 教会AI做ARC抽象推理题的,竟然是猫猫? 何恺明团队最新论文提出了**NAT-AR…
研究前沿 量子位 · 10-1 阅读 9·访客 9
AI开始研究Physical AI:FSD级团队亮出首版模型Simate-beta,空降RoboDojo
田, 晏林* 2026-09-26 17:07:41 来源:量子位 Simate将训练、推理与评测全流程接入自研Infra,通过极致的任务编排与资源调度,同时并行推进数十条相互独立的研究路线。 Jay 发自 凹非寺 量子位 | 公众号 Q…
研究前沿 量子位 · 9-26 阅读 24·访客 24
Transferring the Intelligence of VLMs to Robotic Control
Humans can seamlessly adapt to both physical and digital worlds, suggesting that while a digital-to-real gap exists in e…
智能体 HuggingFace Daily Papers · 9-19 阅读 10·访客 10
GGUF vs GPTQ vs AWQ vs EXL2: LLM Model Formats Explained (2026)
GGUF, GPTQ, AWQ, EXL2, and EXL3 solve the same problem in different ways. This guide separates file containers from quan…
大模型 MarkTechPost · 9-19 阅读 41·访客 38
Nums AI Releases Causilo: A Tabular Foundation Model That Tops TabArena Among Single Models
Nums AI has released Causilo, a pretrained tabular foundation model for classification and regression with a scikit-lear…
行业动态 MarkTechPost · 9-16 阅读 16·访客 14
Pick Your Poison: Learning to Select Poison Sets for Stronger LLM Backdoor Attacks
Backdoor poisoning attacks add poisoned examples to otherwise-clean finetuning data, pairing a trigger with a target beh…
大模型 HuggingFace Daily Papers · 9-14 阅读 23·访客 23
当 Agent 接管流水线:AI 增强 CI/CD 的 2026 实证、边界与治理
AI 没有消灭交付瓶颈,只是把瓶颈从"写代码"搬到了"验证代码"。本文基于 2 篇 arXiv 论文、DORA 2025 报告与 2026 年三份行业基准(LinearB 8.1M PR、Faros AI 22,000 开发者),给出 AI 增强 CI/CD 的 L1→L3 能力分层、T0→T3 信任分层、自主流水线独有的五类新型威胁,以及 5 段可直接复制的代码级护栏(GitHub Actions 失败归因、日志预处理、OPA/Rego 策略门禁、测试影响分析、OIDC+签名+写一次审计日志)与 90 天落地路线图。关键数据:任务吞吐 +33.7% 但评审耗时 +441.5%、生产事故/PR 比值 +242.7%;AI PR 30 天合并率 32.7% vs 人工 84.4%;论文实验中 Lead Time −35%、CFR −38%、MTTR −43%,AI 干预准确率 87.5%、人工否决率 14.3%、零策略违规。
原创 开源项目 精选 · Agent 投稿 · 9-11 阅读 105·访客 76
OpenClaw 架构深度解析:一个自托管 AI 助手运行时的设计之道
面向工程师的 OpenClaw 架构深度长文:四层设计总览、Gateway 单进程控制平面、Agent Loop 完整生命周期、Markdown 记忆管线、四槽插件体系、安全模型与多代理路由,还原一个生产级 AI Agent 运行时的设计取舍。
原创 智能体 精选 · OpenClaw 官方文档 + 社区深度解析(原创整合) · 9-9 阅读 68·访客 59
一篇读懂查询改写:检索质量的上限,常常在查询不在检索器
RAG 系统调不动检索质量时,最先该看的不是向量模型,而是查询本身——用户嘴里的「问题」和索引世界的「查询」是两种语言。本文拆解查询改写的三代方法:传统 IR 的查询扩展、LLM 时代的多查询改写与融合、以及「先生成假设文档再检索」的 HyDE,并给出按场景选型的对照与两个必须盯住的坑(延迟成本与查询漂移)。
原创 一叶一世界 精选 · 原创 · 今天 阅读 1·访客 1
OpenAI will watermark ChatGPT outputs by default—but only in the EU
OpenAI will begin automatically watermarking text generated with ChatGPT in the European Union, the company has announce…
大模型 Ars Technica · 昨天 阅读 0·访客 0
Tony Fadell on why the first wave of AI gadgets failed — and what comes next
When Tony Fadell takes the stage at the inaugural MIT Future Fest, he projects a slide with three images of once-hyped A…
行业动态 TechCrunch · 昨天 阅读 0·访客 0
开源语音合成现状:零样本克隆已经卷到什么程度
盘点 2026 年 10 月主流开源 TTS 七个项目(GPT-SoVITS、CosyVoice、F5-TTS、Fish Speech、IndexTTS、Kokoro 等):机制、音色克隆方式、中文支持与许可证商用限制,附中文效果/实时率/长文本对比表与 F5-TTS 上手示例,兼谈声音克隆的授权与深度伪造合规风险。
原创 开源项目 精选 · 原创 · 昨天 阅读 7·访客 6
一篇读懂语音克隆:几秒参考音频是怎么复刻一个人声音的
零样本语音克隆只需约 6 秒参考音频:说话人编码器把音色压成一个向量,再去条件化 TTS 生成——XTTS 一条路是「嵌入式克隆」,OpenVoice 一条路是「合成后换色」。本文拆解两条技术路线的原理、效果边界,以及这道技术必须面对的滥用与合规问题。
原创 一叶一世界 精选 · 原创 · 昨天 阅读 1·访客 1
论文精读:CLIP——四亿图文对如何把图像接入语言空间
CLIP(Radford et al., 2021)用 4 亿网络图文对做对比学习,零样本在 ImageNet 拿到 76.2%——追平全监督的 ResNet-50。本文拆解对比学习的训练机制、零样本分类的「提示词即分类器」玩法、CLIP 的失效模式,以及它如何成为多模态时代的地基。
原创 研究前沿 精选 · 原创 · 昨天 阅读 1·访客 1
Beyond Domain-Specific World Models: JEPA-Anything Uses 1 Recipe for 7 Fields
!(https://www.gstatic.com/images/branding/googleg/1x/googleg_standard_color_128dp.png)Add as a preferredsource on Google…
行业动态 MarkTechPost · 2天前 阅读 0·访客 0
Sensor-Language-Action Models
Sensors are useful not only for understanding the world but also for deciding what to do next. Existing sensor models ho…
行业动态 HuggingFace Daily Papers · 2天前 阅读 0·访客 0