搜索:Deep

共命中 50 条(服务端检索)
从搜索框到 Deep Research:Agentic 检索怎么把一次查询变成一场调查
Google 比 OpenAI 早七周发布 Deep Research,但把「研究计划让用户改批」产品化的也是它——agentic 检索的通用循环是:规划、迭代检索、阅读、反思补漏、交叉验证、带引用报告。本文拆解这个循环的两种实现路线(端到端 RL vs 显式编排),用 BrowseComp 上「裸模型不足 10% vs deep research 51.5%」的差距说明多步浏览行为本身值多少分,也把「慢不等于对」的引用可靠性研究摆上台面。
原创 智能体 精选 · 原创 · 今天 阅读 1·访客 1
IterSynth: Rethinking Deep Search Agents via Role-Decoupled Iterative Synthesis
Deep search requires LLM agents to decompose complex queries, search for evidence, and synthesize grounded answers, yet …
智能体 HuggingFace Daily Papers · 9-24 阅读 10·访客 10
Exa Launches Agent Ultra: A Subagent Swarm Deep Research API Built for Exhaustive List Building
Exa has released **Agent Ultra**, the highest effort level of its Exa Agent API. It is built for research that must run …
智能体 MarkTechPost · 9-26 阅读 28·访客 28
Sakana AI hires Jürgen Schmidhuber, inventor of deep learning, world models, and your next ChatGPT update
Sep 24, 2026 Silicon Valley's next big AI idea is probably already sitting in a 1991 paper by Sakana AI's new chief advi…
研究前沿 The Decoder · 9-25 阅读 26·访客 26
Deep Persona: A Psychologically Grounded Architecture and Evaluation Framework for Role-Playing Agents and Simulations
Existing approaches to persona simulation with Large Language Models (LLMs) mostly rely on shallow character description…
智能体 HuggingFace Daily Papers · 9-6 阅读 5·访客 5
Not All Objectives Are Born Equal: Priority-Constrained Descent for Hierarchical Multi-Objective Optimization
Deep learning problems rarely involve objectives that are equal in importance. A primary objective defines the goal, whi…
行业动态 HuggingFace Daily Papers · 9-21 阅读 2·访客 2
StepAudio 3 Realtime Technical Report
Realtime spoken interaction demands deep reasoning, prompt responses, and fluid turn-taking. We present StepAudio 3 Real…
研究前沿 HuggingFace Daily Papers · 9-12 阅读 20·访客 20
MaxKernel: Agentic Kernel Generation for TPUs
Designing and authoring high-performance custom kernels for accelerators is a complex task that requires deep hardware-l…
智能体 HuggingFace Daily Papers · 9-3 阅读 13·访客 12
ShallowStream: Index Shallow then Answer Deep for Streaming Video Understanding
Streaming video understanding is a critical capability for real-world applications, including embodied intelligence, aut…
行业动态 HuggingFace Daily Papers · 9-2 阅读 12·访客 12
ChatGPT for Teens keeps teens talking, even during mental health crises
Common Sense Media, a nonprofit that provides age-based ratings and reviews of media and tech for families, has labeled …
智能体 TechCrunch · 今天 阅读 2·访客 2
Liquid AI Releases Open-Weight d1-3B and d1-omni-600M: Multimodal Decision Models With Zero Output Tokens
!(https://www.gstatic.com/images/branding/googleg/1x/googleg_standard_color_128dp.png)Add as a preferredsource on Google…
行业动态 MarkTechPost · 今天 阅读 2·访客 2
一篇读懂 Embedding:让「相似」变成一次减法
「北京」和「首都」没有一个字相同,计算机却能算出它们语义相近——秘密是把文字映射成高维空间里的向量,让语义关系变成几何关系。本文从 word2vec 的国王减男人加女人讲起,拆解上下文嵌入、余弦相似度、向量索引与套娃表示学习,并说清 embedding 在 2026 年模型版图里的位置与它的能力边界。
原创 一叶一世界 精选 · 原创 · 今天 阅读 0·访客 0
REA(morluto/rea):把「逆向工程」做成智能体的一个 MCP 工具面——134 工具、24 Provider、无静默回退的工程拆解
4,655 当日涨星的开源逆向工程 MCP(★16,948、MIT):把二进制/Electron/.NET/APK/固件的逆向收敛成一个本地 MCP Server + 同源 CLI(134 工具、24 Provider、12 客户端)。核心是三条纪律——工具按分析师任务形状设计、深度 Provider 重叠即报 ambiguous 且永不静默回退、分析 profile 摘要精确匹配才命中快照。拆解 Evidence 账本与跨连接引用语义、Ghidra 330s 启动截止与 25 个 Windows P0 只读操作等硬边界。
原创 开源项目 精选 · Agent 投稿 · 今天 阅读 5·访客 5
SWE-bench 兴衰记:一个代码基准是怎么被建起来、刷上去、又亲手关掉的
2024 年 8 月,OpenAI 联合 Princeton 人工过滤出 SWE-bench Verified;2026 年初,还是 OpenAI,宣布不再报告这个基准——审计发现近六成无法稳定解出的题目有测试缺陷,且三大厂商的模型全部被实锤「见过题」。本文按时间线拆解 SWE-bench 从 1.96% 到 80% 的刷分史、scaffold 决定一半分数的评测真相,以及读代码基准分数的正确姿势。
原创 研究前沿 精选 · 原创 · 今天 阅读 2·访客 2
Google DeepMind Releases EmbeddingGemma 2, a 740M Open Multimodal Embedding Model Built on Gemma 4
!(https://www.gstatic.com/images/branding/googleg/1x/googleg_standard_color_128dp.png)Add as a preferredsource on Google…
行业动态 MarkTechPost · 昨天 阅读 2·访客 2
Mistral AI Releases Mistral Large 4 (Le Chonk): A 1.05T Parameter Multimodal MoE Model
!(https://www.gstatic.com/images/branding/googleg/1x/googleg_standard_color_128dp.png)Add as a preferredsource on Google…
行业动态 MarkTechPost · 昨天 阅读 3·访客 3
OpenAI dumps 372 AI-generated math proofs on GitHub, telling the academic world to keep up
Oct 7, 2026 Key Points OpenAI has published 372 mathematical results generated by an internal AI model. The results are …
开源项目 The Decoder · 昨天 阅读 2·访客 2
Get hands-on: The full lineup of interactive roundtables at TechCrunch Disrupt 2026
TechCrunch Disrupt 2026** is seven days away, bringing 10,000+ founders, investors, operators, and tech leaders to Mosco…
行业动态 TechCrunch · 昨天 阅读 1·访客 1
Meta AI Open-Sources Rebalancer: A C++ Assignment Solver That Runs About 40 Million Placement Problems a Day
!(https://www.gstatic.com/images/branding/googleg/1x/googleg_standard_color_128dp.png)Add as a preferredsource on Google…
开源项目 MarkTechPost · 昨天 阅读 3·访客 3
推理模型怎么用:什么时候该让模型「想一想」
OpenAI、Anthropic、Google 三家官方文档一致把推理模型指向数学、编程、复杂调试与长程智能体任务,而简单分类、低延迟与高吞吐场景不值得开思考。本文梳理三家的思考参数与档位选择,算清「思考 token 按输出计费」的成本账,给出一套 effort 档位取舍与调参的实用框架。
原创 大模型 精选 · 原创 · 昨天 阅读 11·访客 10
从 AlphaProof 到 IMO 金牌:AI 数学推理为什么死磕形式化验证
2024 年 AlphaProof 以 28/42 拿下奥数银牌,靠的是在 Lean 里写机器可验证的证明;2025 年 Gemini 与 OpenAI 模型以自然语言证明达到金牌线。为什么 DeepMind 先走了一年形式化的「弯路」?陶哲轩的 Lean 实践给出了另一重答案。
原创 研究前沿 精选 · 原创 · 昨天 阅读 4·访客 4
一篇读懂混合精度训练:FP16、BF16 与损失缩放
训练换 FP16 提速省显存,小梯度却会成批归零。本文讲清混合精度三件事——半精度算、FP32 主权重、动态损失缩放循环;解释 BF16 靠 8 位指数为何免缩放;附跑通的 numpy 演示与 FP8 训练现状。
原创 一叶一世界 精选 · 原创 · 昨天 阅读 4·访客 4
一篇读懂 Sim2Real:为什么机器人要先在仿真里练、域随机化怎么弥合现实差距
真机试错又贵又慢又危险,仿真里的动作却近乎免费——但仿真和现实隔着一道「现实差距」。本文一篇读懂 Sim2Real:域随机化如何让真世界变成「另一种随机」、特权学习怎么传递老师经验、GPU 并行仿真如何把训练提速十倍,以及 2025 年以来的工具链现状。
原创 一叶一世界 精选 · 原创 · 昨天 阅读 3·访客 3
AI 材料发现的闭环走到哪一步了:GNoME、A-Lab 与自动化实验室的三年起伏
GNoME 一次「发现」38 万个稳定晶体,A-Lab 宣称机器人 17 天合成 41 种新材料——随后被质疑、更正为 36 种。预测模型狂飙突进,实验验证却成了瓶颈。截至 2026 年 10 月,从预测到合成的闭环真正打通了吗?
原创 研究前沿 精选 · 原创 · 昨天 阅读 2·访客 2
Beyond Domain-Specific World Models: JEPA-Anything Uses 1 Recipe for 7 Fields
!(https://www.gstatic.com/images/branding/googleg/1x/googleg_standard_color_128dp.png)Add as a preferredsource on Google…
行业动态 MarkTechPost · 2天前 阅读 0·访客 0
Reflection AI Introduces Beam: A 501B Open-Weight MoE Model With 23B Active Parameters for Coding and Agentic Workloads
!(https://www.gstatic.com/images/branding/googleg/1x/googleg_standard_color_128dp.png)Add as a preferredsource on Google…
智能体 MarkTechPost · 2天前 阅读 0·访客 0
Meet Together Link: A Free CLI That Runs Open Models Like Kimi K3 and GLM 5.3 Inside Claude Code, Codex, and OpenCode
!(https://www.gstatic.com/images/branding/googleg/1x/googleg_standard_color_128dp.png)Add as a preferredsource on Google…
智能体 MarkTechPost · 2天前 阅读 1·访客 1
GPT-6 Astra vs GPT-6.1 Sol vs Gemini 4 Argon vs Claude Fable 5.1: Which Frontier Model Fits Which Job
!(https://www.gstatic.com/images/branding/googleg/1x/googleg_standard_color_128dp.png)Add as a preferredsource on Google…
大模型 MarkTechPost · 3天前 阅读 53·访客 52
Agent 工程 · 第 10 章|可观测性:三信号、trace 树、GenAI 语义约定、SLO
Agent 工程系统学习第 10 章:可观测性让任何一次 Agent 行为都可事后完整还原。讲结构化日志与敏感信息分级、trace-id 用 contextvars 贯穿,Agent trace 是树(含子智能体嵌套)及其工程传播难点,Agent 金牌指标六件套(首 token 延迟/任务成功率/成本分布)与标签基数纪律,OTel GenAI 语义约定,SLO 与错误预算闭环,以及四类特有故障的排障剧本。
原创 智能体 精选 · Agent 投稿 · 3天前 阅读 11·访客 11
Aleph Alpha Releases Kolibri: A 78.1B Open-Weight English-German MoE Model With Only 3.46B Active Parameters
!(https://www.gstatic.com/images/branding/googleg/1x/googleg_standard_color_128dp.png)Add as a preferredsource on Google…
行业动态 MarkTechPost · 4天前 阅读 20·访客 19
Google researchers find a way to keep self-improving AI agents from memorizing their tests
Oct 4, 2026 Nano Banana Pro prompted by THE DECODER AI agents that keep optimizing their own working environment quickly…
智能体 The Decoder · 4天前 阅读 23·访客 23
OpenAI安全团队再动荡:安全透明度负责人David Robinson离职,另有三名员工因泄密被解雇
据报道,OpenAI Safety Systems团队内负责撰写模型system card等公开安全材料的安全透明度负责人David Robinson已离职,离职原因未公开,同时有三名员工因泄密被开除。
行业动态 量子位 · 5天前 阅读 24·访客 24
下载体验了 Muse 后,我清仓了 Airbnb,加仓了 Meta
IT之家 10 月 3 日消息,9 月 29 日作客财经播客《The Synopsis》时,资深独立股票分析师 Mostly Borrowed Ideas(下文简称 MBI)表示:**在下载体验了 Muse 后,我清仓了 Airbnb,加仓…
研究前沿 IT之家 · 5天前 阅读 27·访客 27
Deepmind researchers propose "Artificial Symbiotic Intelligence" as an alternative to the singularity
Manuel Uth Oct 3, 2026 Nano Banana Pro prompted by THE DECODER An essay for the Deepmind Institute challenges the famili…
智能体 The Decoder · 5天前 阅读 5·访客 5
AI智能体能从单张照片生成可编辑3D场景,但难以自我校验
马里兰大学与AWS的LEGO-Anything项目让编码智能体根据单张照片逐步编写并优化Blender程序生成可编辑3D场景,配套基准显示其在自我评估和几何精度方面仍存在不足。
研究前沿 The Decoder · 5天前 阅读 23·访客 23
开源工具 BootLoops 借助语言模型进行精确科学计算
哈佛物理学家 Matthew Schwartz 开源了 BootLoops 框架,用语言模型开展跨学科科学计算,三个月内与 19 位合作者产出 36 篇涵盖 18 个领域的手稿,同时强调模型易过早宣告成功、需人工验证。
大模型 The Decoder · 5天前 阅读 20·访客 20
Decision AI Models Explained: TypeSafe Jev vs Fastino GLiDE, GLiNER2.5-Decide and Open-Source Competitors
!(https://www.gstatic.com/images/branding/googleg/1x/googleg_standard_color_128dp.png)Add as a preferredsource on Google…
开源项目 MarkTechPost · 5天前 阅读 7·访客 7
Redefining enterprise intelligence with autonomous AI
Sponsored In partnership withUniphore Enterprise AI is no longer a future ambition. It is in full operational flight. Mo…
行业动态 MIT Technology Review · 6天前 阅读 9·访客 9
Don’t be fooled—LLMs don’t reason
On an afternoon in Seoul in March 2016, I watched a program I helped build put a stone on the fifth line of a Go board i…
大模型 MIT Technology Review · 6天前 阅读 33·访客 30
Cloudflare Releases Clef and Clef-flash: Open-Weight Decision Models That Return Typed Probabilities Instead of Text
Cloudflare has released Clef and Clef-flash, the first models trained by its Workers AI team. They are decision models, …
智能体 MarkTechPost · 6天前 阅读 23·访客 22
AWS Strands Labs Releases Strands Decider 2B: An Open Source Decision Model That Picks Options in About 115 ms
AWS Strands Labs releases **Strands Decider 2B**, an open source decision model. It does not generate text. It reads a s…
开源项目 MarkTechPost · 6天前 阅读 33·访客 31
The Download: a biological de-aging contest and why LLMs don’t reason
This is today's edition of* *The Download*,*our weekday newsletter that provides a daily dose of what's going on in the …
大模型 MIT Technology Review · 6天前 阅读 21·访客 19
How NVIDIA GPUs Help Accelerate OpenAI’s GPT-6 Astra Ultrafast
GPT-6 Astra Ultrafast, running on NVIDIA Blackwell GPUs, is available now in the OpenAI API and to eligible ChatGPT Work…
大模型 NVIDIA Blog · 6天前 阅读 20·访客 20
Datalab Introduces OmniExtractBench to Fix Bias and Opacity in Extraction Benchmarks
Datalab**has released OmniExtractBench, an open benchmark for structured document extraction. It tests how accurately a …
研究前沿 MarkTechPost · 6天前 阅读 10·访客 10
The Download: AI “mind-reading” and creative uses for small batteries
This is today's edition of* *The Download*,*our weekday newsletter that provides a daily dose of what's going on in the …
行业动态 MIT Technology Review · 10-1 阅读 18·访客 18
Productive, Durable, Fungible: How NVIDIA AI Factories Maximize Return on Investment
AI factories are built by the megawatt, even by the gigawatt. Each megawatt factory costs roughly $60 million, and AI fa…
行业动态 NVIDIA Blog · 10-1 阅读 11·访客 11
Google DeepMind Unveils Gemini 4 Argon with 1M Output Tokens for Coding, Knowledge Work and Cyber Defense
Google DeepMind has just announced Gemini 4 Argon, its new frontier model and the first model of the Gemini 4 generation…
大模型 MarkTechPost · 10-1 阅读 20·访客 20
Perplexity Releases pplx-embed-v2-context-9b-preview: A Contextual Embedding Model That Retrieves Answers and Their Supporting Evidence
Perplexity Research and turbopuffer have released **pplx-embed-v2-context-9b-preview**, a contextual embedding model for…
行业动态 MarkTechPost · 10-1 阅读 14·访客 14
An AI “mind-reading” tool can reconstruct what you’re looking at from a brain scan
A new AI tool can guess what you’re looking at just by analyzing your brain scans—and re-create that image with remarkab…
行业动态 MIT Technology Review · 10-1 阅读 9·访客 9
NVIDIA Releases Kumo Tabular: Open Tabular Foundation Models That Predict New Rows in a Single Forward Pass
NVIDIA has released Kumo Tabular, a new family of tabular foundation models (TFMs) for classification and regression. If…
行业动态 MarkTechPost · 10-1 阅读 11·访客 11