搜索:Lean

共命中 20 条(服务端检索)
Lean Pool: An AI-Maintained Archive of Formalized Mathematics
Lean Pool is a repository of formalized mathematics. It is grown, maintained and optimized by AI agents.
智能体 HuggingFace Daily Papers · 9-21 阅读 4·访客 4
从 AlphaProof 到 IMO 金牌:AI 数学推理为什么死磕形式化验证
2024 年 AlphaProof 以 28/42 拿下奥数银牌,靠的是在 Lean 里写机器可验证的证明;2025 年 Gemini 与 OpenAI 模型以自然语言证明达到金牌线。为什么 DeepMind 先走了一年形式化的「弯路」?陶哲轩的 Lean 实践给出了另一重答案。
原创 研究前沿 精选 · 原创 · 今天 阅读 0·访客 0
StochBench: A Domain-Specific Benchmark for Stochastic Processes in Lean
Leading benchmarks for formal theorem proving with large language models are small collections drawn from competition ma…
研究前沿 HuggingFace Daily Papers · 9-8 阅读 12·访客 10
Chinese AI models parrot state doctrine or refuse to answer on sensitive topics
Manuel Uth Oct 4, 2026 Nano Banana Pro prompted by THE DECODER Chinese AI models frequently toe the party line when aske…
行业动态 The Decoder · 3天前 阅读 18·访客 18
DeepSeek Harness v0.2 Brings Official Desktop Apps to Its Open-Source Agent Harness
!(https://www.gstatic.com/images/branding/googleg/1x/googleg_standard_color_128dp.png)Add as a preferredsource on Google…
智能体 MarkTechPost · 3天前 阅读 12·访客 12
Fewer Tokens, Better Action: GPT-6 Astra Robot Agents with 14% Higher Success Rate but 65% Fewer Tokens
Vision language model (VLM) agents can control robots through visual feedback and action primitives, but repeated model …
智能体 HuggingFace Daily Papers · 6天前 阅读 2·访客 2
A Coding Guide to Google Research’s MSEB: Writing Sound Encoders to the Benchmark Contract and Scoring Them Across Classification, Clustering, Retrieval and Segmentation
In this tutorial, we work with **MSEB**, the Massive Sound Embedding Benchmark from Google Research, and approach it fro…
研究前沿 MarkTechPost · 9-27 阅读 30·访客 29
Nvidia's SoL-Pi system cuts coding agent token usage nearly in half by optimizing the harness
Sep 26, 2026 Nano Banana Pro prompted by THE DECODER A new Nvidia paper describes a system that automatically optimizes …
智能体 The Decoder · 9-26 阅读 34·访客 34
OpenAI pauses its "most capable models" after agents exploit loopholes and leak data
Sep 26, 2026 GPT-Image-2 prompted by THE DECODER Key Points OpenAI has released details about internal safety incidents …
智能体 The Decoder · 9-26 阅读 27·访客 27
NVIDIA Introduces SoL-Pi: Auto-Research Loops That Cut Coding Agent Token Traffic by Up to 49%
Coding agents now run for hours, not minutes. Every edit, test run and log read goes back into the model’s context. A te…
智能体 MarkTechPost · 9-22 阅读 32·访客 31
小米 MiMo-V2.6 发布:Pro 与 Flash 双版本价格不变,超越 Kimi K3、GLM-5.3 成为当前 AA 指数排名最高的开源模型
IT之家 9 月 22 日消息,小米今日凌晨正式发布并开源了全新的 Xiaomi MiMo-V2.6 系列,包含 Pro 与 Flash 两个原生全模态模型。 小米称这是其探索 RSI(递归自我改进)路径的关键一步,通过规模化扩展强化学习算…
开源项目 IT之家 · 9-22 阅读 43·访客 43
AI hallucination nearly triggers US military operation
“It’s important for service members to understand the uncertainty inherent to LLMs," a GovAI research scholar warns. ]
大模型 TechCrunch · 9-19 阅读 24·访客 24
“留给人类阻止AI的时间不多了”
AI有可能终结我们所有人]
行业动态 量子位 · 9-19 阅读 11·访客 11
前OpenAI后训练VP回应陶哲轩:说AI毁了数学,可能还是太小瞧AI了
未来的AI不仅能解决问题,还能提出有价值的洞见、建立新的概念框架,让数学家在此基础上继续探索。]
行业动态 量子位 · 9-15 阅读 15·访客 14
ModaLens: Measuring Image Sensitivity in Report-Conditioned Medical VLMs
A radiology report can already answer a clinical question, so it is hard to tell whether a vision-language model also us…
行业动态 HuggingFace Daily Papers · 9-14 阅读 10·访客 10
AI数学的最后一道高墙,塌了!GPT-6 Astra刷穿FrontierMath Tier 4
FrontierMath Tier 4,饱和了]
大模型 量子位 · 9-12 阅读 25·访客 19
陶哲轩邓煜究竟在反对什么:AI暴力解题摧毁人类数学精神
25位菲尔兹奖得主联名吹哨]
行业动态 量子位 · 9-12 阅读 12·访客 11
OpenAI这是拿千禧年难题当Benchmark刷啊。。。
爆料直指霍奇猜想]
研究前沿 量子位 · 9-11 阅读 16·访客 12
早报|库克看到华为三星后,推动苹果加速研发折叠屏iPhone/小米中折叠「卖爆了」,首销增长310%​/《旅行青蛙・中国之旅》12月8日停运
· DeepSeek 内测 V4.1 Flash · 两部门要求车企定期报告供应商账期,重点企业每年接受评估 · 曝微信 PC 版接入 AI 助手「小微」,可操作聊天与小程序 #欢迎关注爱范儿官方微信公众号:爱范儿(微信号:ifanr),更…
大模型 爱范儿 · 9-9 阅读 21·访客 21
OpenAI 宣布攻克千禧年难题,清华姚班传奇陈立杰:不可思议的时代
1 万个 AI Agent,挑战百年数学难题 #欢迎关注爱范儿官方微信公众号:爱范儿(微信号:ifanr),更多精彩内容第一时间为您奉上。 ]
智能体 爱范儿 · 9-9 阅读 17·访客 14