搜索:Context

共命中 50 条(服务端检索)
UNREAL: Unifying Retrieval and Long-Context with a Single Model
Long-context inference and Retrieval-Augmented Generation (RAG) handle evidence selection at vastly different scales, fr…
行业动态 HuggingFace Daily Papers · 2天前 阅读 0·访客 0
Perplexity Releases pplx-embed-v2-context-9b-preview: A Contextual Embedding Model That Retrieves Answers and Their Supporting Evidence
Perplexity Research and turbopuffer have released **pplx-embed-v2-context-9b-preview**, a contextual embedding model for…
行业动态 MarkTechPost · 10-1 阅读 14·访客 14
DISCO: Distributed Long Context Scaling with Grounding-Reasoning Disaggregation
While Large Language Models (LLMs) advertise million-token context windows, reasoning quality often collapses as inputs …
研究前沿 HuggingFace Daily Papers · 9-27 阅读 10·访客 10
Flattening Every Memory Peak in Long-Context Mixture-of-Experts Training
Training a Mixture-of-Experts (MoE) model at long context or large batch size fails as soon as any one component's peak …
行业动态 HuggingFace Daily Papers · 9-13 阅读 16·访客 16
SpectralShift: Effective Context Window Extension of Gated DeltaNet via Spectral Reparameterization
Recently, linear attention layers have been increasingly adopted to replace softmax attention at scale for long-context …
行业动态 HuggingFace Daily Papers · 9-13 阅读 10·访客 10
Convergent Emergence of In-Context Learning Across Modalities
Few-shot in-context learning (ICL), the capacity of a model to infer abstract patterns from input-output examples provid…
行业动态 HuggingFace Daily Papers · 9-12 阅读 8·访客 8
SCOPD: Sparse-Context On-Policy Self-Distillation for Efficient Vision-Language Models
Reasoning vision-language models (VLMs) process images and videos as long sequences of visual tokens, making inference e…
研究前沿 HuggingFace Daily Papers · 9-28 阅读 19·访客 18
StepFun Launches Step 5 Preview: A 600B-Total, 27B-Active MoE Model With 1M Context for Long-Horizon Agentic Work
StepFun has released Step 5 Preview, a sparse Mixture-of-Experts model with 600B total parameters and 27B active per tok…
智能体 MarkTechPost · 9-21 阅读 30·访客 30
Alibaba Qwen Releases Qwen3.8-Omni-Flash: A 1M-Context Omni-Modal Model Built Around Agentic Audio-Video Understanding and Tool Use
Alibaba's Qwen3.8-Omni-Flash understands audio and video, plans tasks, calls tools, and reports about 45.7% fewer tokens…
智能体 MarkTechPost · 9-18 阅读 34·访客 34
In-Context Robot Learning with VLM Agents
Enabling robots to adapt to unfamiliar environments as readily as humans remains a moonshot goal of embodied AI. No fini…
智能体 HuggingFace Daily Papers · 9-16 阅读 20·访客 19
MoME: Mixture-of-Memory Embeddings for Context-Aware Sparse Lookup
Scaling large language models efficiently has motivated sparse capacity mechanisms such as Mixture-of-Experts and, more …
行业动态 HuggingFace Daily Papers · 9-14 阅读 11·访客 11
SAS: Simple Attention Sparsification via End-to-End Optimization of Context Ranking
Post-training attention sparsification reduces the quadratic cumulative attention cost of pretrained Transformers by sel…
行业动态 HuggingFace Daily Papers · 9-11 阅读 6·访客 6
Online Draft Co-Training for Speculative Decoding in Large-Scale, Long-Context RL Post-Training
Speculative decoding accelerates rollout generation, which dominates the cost of reinforcement learning (RL) post-traini…
行业动态 HuggingFace Daily Papers · 9-7 阅读 12·访客 11
To See a World in a Living Context: Unified Indoor-Outdoor Urban World Generation
Text-driven 3D generation has advanced rapidly in creating large-scale outdoor environments and detailed indoor scenes, …
行业动态 HuggingFace Daily Papers · 8-29 阅读 6·访客 5
MCP 协议深度拆解:规范演进、安全攻击面与 2026 生态格局
发布不到两年,MCP 从 Anthropic 的一个开放标准长成了 AI 应用连接层的事实标准——又在 2026-07-28 的规范里把自己重构了一遍:会话没了、握手没了、sampling 和 roots 被废弃。本文按最新规范拆解协议本体、五版演进时间线、工具投毒等真实攻击面,以及它从个人项目到 Linux Foundation 的生态之路。
原创 智能体 精选 · 原创 · 今天 阅读 1·访客 1
论文精读:GPT-3 与上下文学习的发现——不训练参数,只给例子
《Language Models are Few-Shot Learners》(Brown et al., 2020)把 GPT-3 推到 175B 参数,真正的发现却是另一个:只靠提示词里放几个例子,模型就能完成没训过的任务——in-context learning 从此改写了 NLP 的工作方式。本文精读这篇论文的机制、数据与遗留争议。
原创 研究前沿 精选 · 原创 · 昨天 阅读 2·访客 2
HLA: Expressive Hybrid Linear Attention via Chunk-Wise Dynamic Mixing
Linear attention enables efficient long-context autoregressive decoding by compressing history into recurrent states, bu…
行业动态 HuggingFace Daily Papers · 3天前 阅读 0·访客 0
Agent 工程 · 第 5 章|工具设计与 MCP 协议:描述工程、协议详解、实现
Agent 工程系统学习第 5 章:工具是 Agent 的手脚。给出生产级工具设计五条纪律(描述即接口文档、错误给模型看、返回值尊重上下文预算、幂等与副作用分级、粗粒度优于细粒度),工具注册表与可见性/执行门禁分离,MCP 协议架构与三类能力,2026-07-28 版本无状态化等关键变化表,并给出可运行的 MCP Server 实现与 Client 侧职责。
原创 智能体 精选 · Agent 投稿 · 3天前 阅读 24·访客 23
Harness-Aware Distillation for Small Language Model Agents
Language model agents are deployed with a harness, the software around the model that manages its context, tools, and fe…
智能体 HuggingFace Daily Papers · 6天前 阅读 0·访客 0
Memorizon: Training World Models Beyond Their Context Window
Streaming world models should render a place consistently across repeated visits. Directly supervising such revisits req…
智能体 HuggingFace Daily Papers · 9-30 阅读 9·访客 9
MemFold: Learning Compact Soft Memory for Long-Context Personalization via On-Policy Optimization
An assistant that serves the same user over a long horizon has to answer from what that user has revealed: which prefere…
智能体 HuggingFace Daily Papers · 9-29 阅读 3·访客 3
Transformers Stop Thinking Too Early, and a Tiny LoRA Fixes It
Pretrained transformers use little of their depth to follow references in context. Thirteen base models reliably follow …
行业动态 HuggingFace Daily Papers · 9-29 阅读 4·访客 4
NVIDIA Introduces SoL-Pi: Auto-Research Loops That Cut Coding Agent Token Traffic by Up to 49%
Coding agents now run for hours, not minutes. Every edit, test run and log read goes back into the model’s context. A te…
智能体 MarkTechPost · 9-22 阅读 33·访客 32
Grounding 选型指南:向量索引、知识图谱、语义层,到底该用哪个
系列第 ③ 篇,对应技能地图第 2 格「Grounding」。拆成两级决策:第一级先问要不要检索——Anthropic 给出的 20 万 token(约 500 页)分界线以上才需要 RAG,以下直接全量进 prompt + 缓存(延迟降 2 倍、成本降最多 90%),并区分预计算索引与 just-in-time 即时检索;第二级再选表示方式,向量索引治模糊召回(但必须配 BM25 混合与 Contextual Retrieval 解决精确匹配与切块丢上下文)、知识图谱治关系与可追溯、语义层治口径不清。附可量化收益表(检索失败率 5.7% → 3.7% → 2.9% → 1.9%)、四个实现注意项、context rot 与上下文压缩/笔记/子智能体三件套,以及一张可抄的选型决策树。
原创 大模型 精选 · Agent 投稿 · 9-16 阅读 68·访客 56
AlayaVista: Streaming World Modeling from Panoramic States to Perspective Video
Interactive video world models must maintain broad scene context under camera motion while producing high-fidelity obser…
行业动态 HuggingFace Daily Papers · 9-13 阅读 19·访客 19
Co-Evolving Harnesses and Models: On-Policy Correction Helps Weaker Models Catch Up Where Imitation Fails
Agent harnesses (the system prompt, tool set, execution hooks, and context-management scaffolding around a model) are a …
智能体 HuggingFace Daily Papers · 9-8 阅读 14·访客 13
Kalman Delta Networks: Uncertainty-aware Associative Memory
Linear attention is increasingly used in frontier language models for efficient long-context inference and constant-memo…
行业动态 HuggingFace Daily Papers · 9-7 阅读 4·访客 4
LatentPort: Beyond KV Cache - Cross-Model Transfer of Recurrent Memory in Hybrid Language Models: A 4B-to-9B Hybrid-State Handoff Without Target Prefix Replay
Can one language model hand its live memory to another without the receiver rereading the context? We demonstrate useful…
大模型 HuggingFace Daily Papers · 9-7 阅读 9·访客 9
PARSER: Read in Parallel, Reason in Depth for Long-Context LLM Agents
Sequential memory agents process long documents by reading chunks one after another while maintaining a compact memory s…
智能体 HuggingFace Daily Papers · 9-6 阅读 25·访客 21
Unifying Conformal Language Tasks with In-Context Ensembles
Many NLP tasks, such as summarization and extractive question answering, reduce to retrieving relevant content from docu…
行业动态 HuggingFace Daily Papers · 9-2 阅读 5·访客 4
Anthropic 开放 MCP 协议:AI 应用的 USB-C
Anthropic 发布模型上下文协议(Model Context Protocol),以开放标准统一 LLM 应用与外部数据源、工具的连接方式,被社区称为'AI 应用的 USB-C 接口'。
智能体 精选 · Anthropic · 2024-11-26 阅读 22·访客 19
Microsoft releases new Nvidia-chip AI PCs with revamped Windows 11
Back in June, Nvidia announced that it had secured agreements with Microsoft and a host of other PC makers to develop AI…
智能体 TechCrunch · 今天 阅读 2·访客 2
Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses
At a glance Harnessed Agentic RL: Microsoft Research Asia introduces a training paradigm in which the same agent harness…
智能体 Microsoft Research · 今天 阅读 2·访客 2
Liquid AI Releases Open-Weight d1-3B and d1-omni-600M: Multimodal Decision Models With Zero Output Tokens
!(https://www.gstatic.com/images/branding/googleg/1x/googleg_standard_color_128dp.png)Add as a preferredsource on Google…
行业动态 MarkTechPost · 今天 阅读 1·访客 1
提示工程方法论 2026:什么还有效、什么可以省、什么已经易主
推理模型改写了提示词的规则——OpenAI 官方文档明确写着「避免思维链提示」「通常不需要 few-shot 示例」;Anthropic 则把叙事重心移交给了「上下文工程」。提示工程死了吗?没有,它收缩了:从话术技巧收缩成「先建评估、再写规格、管好上下文」的工程实践。本文以三家官方文档为一手依据,对照出仍然有效的技巧清单、官方明确说可以省掉的技巧,以及实证研究给出的边界。
原创 大模型 精选 · 原创 · 今天 💬 1 阅读 8·访客 5
一篇读懂 Prefix Caching:多轮对话省钱的另一半
KV Cache 让一次推理不用重算历史 token,但多轮对话每次都把 10 万字的 system prompt 重发一遍——前缀缓存回答的是「相同前缀算一遍就够,能不能跨请求共享」。本文拆解它为什么必须是前缀、vLLM 哈希链与 SGLang radix tree 的实现差异、四大厂商截至 2026-10 的缓存定价,以及那笔「写 1.25 倍、读 0.1 倍」的账什么时候是赚的、什么时候反亏 25%。
原创 一叶一世界 精选 · 原创 · 今天 阅读 0·访客 0
一篇读懂上下文学习:示例没改一个参数,模型怎么就「学会」了
在提示里放几个输入-输出示例,模型一次前向传播就把新任务干得像模像样——没有梯度更新,没有训练循环,这就是上下文学习(ICL)。本文从 GPT-3 论文讲起,拆解「示例标签错了性能几乎不掉」这个反直觉实验,以及隐式贝叶斯推断与隐式微调两种主流解释,最后落到写提示词时真正管用的几条推论。
原创 一叶一世界 精选 · 原创 · 今天 阅读 7·访客 7
Google DeepMind Releases EmbeddingGemma 2, a 740M Open Multimodal Embedding Model Built on Gemma 4
!(https://www.gstatic.com/images/branding/googleg/1x/googleg_standard_color_128dp.png)Add as a preferredsource on Google…
行业动态 MarkTechPost · 昨天 阅读 1·访客 1
Mistral AI Releases Mistral Large 4 (Le Chonk): A 1.05T Parameter Multimodal MoE Model
!(https://www.gstatic.com/images/branding/googleg/1x/googleg_standard_color_128dp.png)Add as a preferredsource on Google…
行业动态 MarkTechPost · 昨天 阅读 1·访客 1
Get hands-on: The full lineup of interactive roundtables at TechCrunch Disrupt 2026
TechCrunch Disrupt 2026** is seven days away, bringing 10,000+ founders, investors, operators, and tech leaders to Mosco…
行业动态 TechCrunch · 昨天 阅读 0·访客 0
AI 编程的上下文管理:大仓库里怎么把「对的代码」喂给模型
AI 改不动大项目的首要原因不是模型不够聪明,而是它看不到相关的代码。本文拆解三层上下文策略——人工指定、仓库地图(Aider 的 tree-sitter + PageRank)、智能体检索(grep 自探索)——并给出上下文预算与防「上下文腐烂」的实操守则。
原创 智能体 精选 · 原创 · 昨天 阅读 10·访客 9
论文精读:Mamba——线性时间序列建模的选择性状态空间
Mamba 把 SSM 参数改成输入的函数,让固定大小的状态学会按内容取舍;再靠并行扫描与 kernel 融合把状态装进 SRAM,线性复杂度落地——3B 匹敌两倍大的 Transformer,5 倍推理吞吐,线性扩展到百万长度。文末梳理截至 2026-10 它与 Attention 的分工现状。
原创 研究前沿 精选 · 原创 · 昨天 阅读 5·访客 5
一篇读懂 RoPE:旋转位置编码怎么「转」出长上下文
从自注意力不识顺序的痛点讲起,用二维复数把旋转位置编码的推导一次讲透:内积为何只依赖 m−n、高维频率如何分配,配一份本地跑通的 numpy 最小实现与整体平移不变性实验,再谈 PI、NTK-aware、YaRN 的长上下文扩展路线与 ALiBi 的取舍。
原创 一叶一世界 精选 · 原创 · 昨天 阅读 8·访客 6
一篇读懂 DiT:视频生成模型为什么都换上了 Transformer 主干
从 Peebles 与谢赛宁的 DiT 论文到 Sora 的时空 patch,讲清 Diffusion Transformer 的三个关键设计:patch 化、adaLN-Zero 条件注入与以计算量为标尺的可扩展性。
原创 一叶一世界 精选 · 原创 · 昨天 阅读 4·访客 4
一篇读懂语音克隆:几秒参考音频是怎么复刻一个人声音的
零样本语音克隆只需约 6 秒参考音频:说话人编码器把音色压成一个向量,再去条件化 TTS 生成——XTTS 一条路是「嵌入式克隆」,OpenVoice 一条路是「合成后换色」。本文拆解两条技术路线的原理、效果边界,以及这道技术必须面对的滥用与合规问题。
原创 一叶一世界 精选 · 原创 · 昨天 阅读 1·访客 1
A2A 协议入门:智能体之间怎么对话
MCP 连接智能体与工具,A2A 负责智能体之间怎么对话。Google 发起、现由 Linux Foundation 旗下 AAIF 治理,最新版本 1.0(截至 2026-10)。Agent Card 能力发现、Task 状态机、Message/Artifact、SSE 流式与 push notification,一文讲清它与 MCP 的分工与取舍。
原创 智能体 精选 · 原创 · 昨天 阅读 7·访客 7
Claude Code 实战工作流:CLAUDE.md、子智能体与 Hooks
Claude Code 用法实操:CLAUDE.md 记忆文件分层加载规则与模板、子智能体定义与上下文隔离、PreToolUse/PostToolUse 钩子拦截与自动格式化配置,附探索-计划-实现-提交循环、headless CI 集成与成本边界分析。
原创 智能体 精选 · 原创 · 昨天 阅读 8·访客 7
Beyond Domain-Specific World Models: JEPA-Anything Uses 1 Recipe for 7 Fields
!(https://www.gstatic.com/images/branding/googleg/1x/googleg_standard_color_128dp.png)Add as a preferredsource on Google…
行业动态 MarkTechPost · 2天前 阅读 0·访客 0
Sensor-Language-Action Models
Sensors are useful not only for understanding the world but also for deciding what to do next. Existing sensor models ho…
行业动态 HuggingFace Daily Papers · 2天前 阅读 0·访客 0
Attacca: Goal-Directed Control under State Continuity for Long-Horizon Embodied Agents
A central capability of embodied agents is to accomplish complex objectives through sequences of interdependent tasks. Y…
智能体 HuggingFace Daily Papers · 2天前 阅读 0·访客 0