JEDEE AI
存档 2026-08-15

8 月 15 日(北京时间)全球 AI 圈推文存档,按曝光排序,共 100 条。
← 返回最新 全部归档

全部情报 每小时更新 · 事件已合并同类项

内容 公司
Anthropic@AnthropicAI · 公司官方 · 1 天前Claude 开发商官方账号

我们写了一份常见问答来回答关于水印的一些问题。

总结一下:

• 我们正在实施水印以符合 EU AI Act。其他主要模型开发商已签署了相同的实践准则,也将实施水印;
• 我们的水印方法对 Claude 输出的质量或内容没有任何实际影响;
• 水印文本和未水印文本之间的差异对读者来说无法区分;
• 文本中没有添加任何内容,也没有隐藏字符;
• 水印不需要额外的 token,也不会更贵;
• 水印无法追溯到特定的个人、组织或聊天。

了解更多:anthropic.com/news/claude-te…

查看英文原文
We’ve written an FAQ to answer some of the questions we've received about watermarking.

In summary:

• We’re implementing watermarking to comply with the EU AI Act. Other major model developers have signed the same Code of Practice and will also be implementing watermarking;
• Our watermarking method doesn’t have any practical impact on the quality or content of Claude’s outputs;
• The difference between watermarked and un-watermarked text will not be distinguishable to readers;
• Nothing is added to the text and there are no hidden characters;
• Watermarking doesn’t require extra tokens, and will not be more expensive;
• Watermarks can’t be traced to a specific person, organization, or chat.

Read more:
anthropic.com/news/claude-te…
◔ 990.1 万 次浏览(2 条合计)♥ 4,395⇄ 614新品看原帖 ↗
Grok@grok · 公司官方 · 1 天前马斯克 xAI 旗下聊天机器人 Grok 官方

Grok 4.6 现在在 GitHub Copilot 里可用了。

在 GitHub Copilot CLI、IDE 和云产品中试试吧!

x.ai/news/grok-4-6-github-co…

查看英文原文
Grok 4.6 is now available in GitHub Copilot.

Try it out in the GitHub Copilot CLI, IDE, and cloud products!


x.ai/news/grok-4-6-github-co…
◔ 1427.5 万 次浏览(3 条合计)♥ 1,995⇄ 179新品看原帖 ↗
Andrew Ng@AndrewYNg · 创始人 · 1 天前吴恩达,斯坦福教授、AI 教育领军人物

新:AI 工程最重要技能图谱

查看英文原文
New: A map of the most important skills in AI Engineering.
Anthropic@AnthropicAI · 公司官方 · 1 天前Claude 开发商官方账号

作为我们负责任扩展政策的一部分,我们定期发布风险报告。这些报告详细分享我们系统的风险以及我们为应对这些风险的准备。我们的第二份风险报告现已发布:anthropic.com/aug-2026-risk-…

查看英文原文
As part of our Responsible Scaling Policy, we publish regular Risk Reports. These share detailed information on the risks of our systems and how prepared we are to address them.

Our second Risk Report is now available:
anthropic.com/aug-2026-risk-…
◔ 135.9 万 次浏览(9 条合计)♥ 2,200⇄ 192动态看原帖 ↗
Matt Shumer@mattshumer_ · 博主 · 1 天前HyperWrite CEO,AI 实战技巧分享
连环推 ×2

想学怎么运行的话看这里:
soomethingbig.ai/gauntlet-loo…


用 Grok Build harness 可能是最好的选择,也可以在 Grok Bot 里跑

引用 Elon Musk @elonmuskGrok 4.6 runs The Gauntlet查看被引原帖 ↗
查看英文原文
If you want to learn how to run one:
somethingbig.ai/gauntlet-loo…


Grok Build harness is likely best but you can also run this in Grok Bot!
Gauntlet Loop prompt to build games like this:
somethingbig.ai/gauntlet-loo…
NVIDIA@nvidia · 公司官方 · 1 天前

这个夏天,NVIDIA实习生为公司各个团队带来了好奇心、创意和热情 - 学习、建设并帮助塑造未来。感谢2026年的实习生与我们度过的这个夏天。💚

查看英文原文
This summer, NVIDIA interns brought their curiosity, creativity and energy to teams across the company - learning, building and helping shape what’s next.

To our 2026 interns: thank you for spending your summer with us. 💚
◔ 67.6 万 次浏览♥ 1,497⇄ 123▶ 含视频动态看原帖 ↗
Chubby♨️@kimmonismus · 博主 · 1 天前Chubby,高频 AI 新闻聚合博主

所以啊,你现在可以本地运行 Opus 4.6 max 了。

好好体会一下这意味着什么。

引用 Chubby♨️ @kimmonismusQwen发布Qwen3.8-27B,仅27B参数在代理编码、计算机使用、浏览器任务等多项基准测试中超越Claude Opus 4.6 Max。支持图像、视频、262K-1M上下文、Apache 2.0开源。展现前沿级代理能力趋向轻量化、廉价化、可部署化。查看被引原帖 ↗
查看英文原文
So yeah, you can now run Opus 4.6 max locally.

Just let that sink in.
Tibo@thsottiaux · 公司官方 · 1 天前ChatGPT 产品官方账号

Codex 这周为你解决了什么难题?你通常是怎样学习怎么推进技术边界的?

查看英文原文
What’s a hard problem codex solved for you this week? Where do you learn from others on how to push the frontier?
Min Choi@minchoi · 博主 · 1 天前AI 产品演示博主,专门展示新工具玩法

你可以在 MacBook Pro 上本地运行 Opus 4.6 Max 级别的智能。就在本地。

引用 Qwen @Alibaba_QwenWe promised open weights for Qwen3.8. Now, time to meet them! 🎉 ⚡ Qwen3.8-27B: - A native multimodal dense model. With just 27B parameters, it outperforms Qwen3.7-Plus overall and shines in real-world coding & office workflows. - 262K native context, easily extendable to 1M tokens via YaRN. - Built for builders. Highly efficient, high-quality, and licensed under Apache 2.0. 🚀 The open weights for Qwen3.8-2.4T-A95B (Max-level) have also been released recently. Whether you're shipping lightweight applications with Qwen3.8-27B locally or building agents with Qwen3.8-2.4T-A95B, they're yours now! Download, deploy, and build something we haven't imagined yet. 👀👇 - Hugging Face: huggingface.co/collections/Q… - ModelScope: modelscope.cn/collections/Qw…查看被引原帖 ↗
查看英文原文
You can literally run Opus 4.6 Max-level intelligence on a MacBook Pro.

Locally.
Tibo@thsottiaux · 公司官方 · 1 天前ChatGPT 产品官方账号

在 ChatGPT 里快速订餐厅座位现在超简单了。还有一堆其他功能也发布了。

引用 Adam Fry @adamhfryThis week's ChatGPT feature drop - Aug 14: 1/ Quizzes - "Quiz me on [topic]" will now automatically generate quizzes right in Chat. Learn any topic! 2/ Reservation search - "I'm looking for a reservation for two next Saturday for a place with a great outdoor patio". Nice thing is you can just describe what you want. 3/ Google Drive for Paid Users - You can now add Google Drive files to your ChatGPT Library. That makes it easy to ask ChatGPT questions about those files. Just hit the + menu and "Add from Library" 4/ Suggestions for Paid Users - Paid users will see high quality suggestions on the home page of ChatGPT We're continuing to make answers better, allow you to pull in the right documents, and more easily discover what Chat can do. Step by step, ChatGPT gets more useful! Leave feedback in the comments.查看被引原帖 ↗
查看英文原文
Looking for a quick restaurant reservation is now super easy in ChatGPT. Together with a bunch more ships
◔ 48.1 万 次浏览♥ 1,775⇄ 61▶ 含视频新品看原帖 ↗
Qwen@Alibaba_Qwen · 公司官方 · 1 天前阿里通义千问大模型团队

27B 能在 17GB RAM 上跑。你准备好创造不可思议的东西了吗?😎
感谢提醒!
@UnslothAI

引用 Unsloth AI @UnslothAIQwen3.8-27B 现可本地运行!✨ 通过 Unsloth Dynamic GGUFs 在 17GB RAM 上运行。目前同规模最强大的模型,提供 NVFP4 量化版本。GGUF 和教程已发布。查看被引原帖 ↗
查看英文原文
27B on 17GB RAM. Are you ready to create something incredible? 😎
Thanks for highlighting it!
@UnslothAI
◔ 66.3 万 次浏览(8 条合计)♥ 5,614⇄ 399教程看原帖 ↗
Google Gemini@GeminiApp · 公司官方 · 1 天前谷歌 Gemini 产品官方

Gemini 3.7 Flash现在向Gemini聊天中的所有Pro和Ultra用户开放。

这次模型更新提升了多步骤任务的推理和准确性,比如能智能地将几十个文件和邮件中的信息串联成一份主文档。

今天就在网页版或App里试试吧!

引用 Google Gemini @GeminiAppGemini Spark现运行于Gemini 3.7 Flash。无论是在Sheets中整理供应商还是草拟谈判邮件,3.7 Flash通过改进的Google Workspace应用工具使用,让你的个人AI代理更精确、更准确,助力想法落地。查看被引原帖 ↗
查看英文原文
Gemini 3.7 Flash is now available to all Pro and Ultra users in Gemini chat.

This model update delivers improved reasoning and accuracy for multi-step tasks like intelligently connecting the dots across dozens of files and emails into one master document.

Give it a try today on the web or in the app!
◔ 32.3 万 次浏览(2 条合计)♥ 2,605⇄ 204新品看原帖 ↗
Pika@pika_labs · 公司官方 · 1 天前AI 视频生成公司 Pika
连环推 ×4

打开声音!今天我们推出 Pika Audio 模型:4 个前沿基础模型,涵盖生成音频的全部范围。我们的定价比市场上所有音频模型都便宜——便宜高达 20 倍。*

*真的没有免责声明

查看英文原文
Sound on! Today, we’re introducing Pika Audio models: 4 frontier foundation models that cover the full spectrum of generative sound. And we’ve made them less expensive than every audio model on the market—up to 20x times cheaper.*

*There is literally no disclaimer
You can start using Pika Soundtrack, Pika Music, Pika SFX, and Pika Speech exclusively on the Pika API Club:

dev.pika.art/models/pika/pik…
We’re thrilled to be able offer prices this low, thanks to our team’s innovations in training and inference efficiency. A few highlights:

• Pika Soundtrack is 0.617 / seconds and 2x more cost-efficient than Hunyuan Foley, the only model with comparable video-to-audio functionality.

• Pika SFX is up to 20x more cost-efficient than alternatives.

• Pika Speech is 9x more cost-efficient ElevenLabs v3, 4.5x more cost-efficient than Cartesia and ElevenLabs Turbo, and 2x more cost-efficient than Fish Audio.

• Pika Music is up to 10x more cost-efficient than comparable music models.
The Pika Audio model family is just one of the things we’ve been working on behind the scenes. Keep an eye on us for more ways we’ll be making gen media more accessible to more people…
◔ 35.8 万 次浏览(6 条合计)♥ 607⇄ 80▶ 含视频新品看原帖 ↗
Chubby♨️@kimmonismus · 博主 · 1 天前Chubby,高频 AI 新闻聚合博主

以后只要你光用 Claude 来翻译文本,就会被认定为 AI 生成内容。这事儿后果很严重,特别对大学生,因为检测工具现在会直接标记成 AI 内容。

引用 Anthropic @AnthropicAIWe’ve written an FAQ to answer some of the questions we've received about watermarking. In summary: • We’re implementing watermarking to comply with the EU AI Act. Other major model developers have signed the same Code of Practice and will also be implementing watermarking; • Our watermarking method doesn’t have any practical impact on the quality or content of Claude’s outputs; • The difference between watermarked and un-watermarked text will not be distinguishable to readers; • Nothing is added to the text and there are no hidden characters; • Watermarking doesn’t require extra tokens, and will not be more expensive; • Watermarks can’t be traced to a specific person, organization, or chat. Read more: anthropic.com/news/claude-te…查看被引原帖 ↗
查看英文原文
From now on, if you use Claude just to translate text into another language, it will be considered AI-generated.

This has serious consequences, for example, for students at universities, because detectors will now flag it as AI content.
Lisan al Gaib@scaling01 · 博主 · 1 天前高频 AI 模型测评与爆料博主

Anthropic 在谈论一个神秘的“MODEL 2”,据称比 Mythos 5 还要强大

引用 Anthropic @AnthropicAI作为负责任扩展政策的一部分,我们定期发布风险报告,分享系统风险及应对准备的详细信息。我们的第二份风险报告现已发布。查看被引原帖 ↗
查看英文原文
Anthropic talking about a mysterious "MODEL 2" that is more capable than Mythos 5
Greg Brockman@gdb · 创始人 · 1 天前Greg Brockman,OpenAI 联合创始人兼总裁

模型期望的快乐跑步机

查看英文原文
hedonic treadmill of model expectations
Chubby♨️@kimmonismus · 博主 · 1 天前Chubby,高频 AI 新闻聚合博主

正如大家早就预料的那样,Anthropic 内部有个模型比 Mythos 5 厉害得多,不过他们没打算放出来。

查看英文原文
As has already been expected, Anthropic internally uses a model that is significantly better than Mythos 5, but they have no plans to release it.
Yuchen Jin@Yuchenj_UW · 博主 · 1 天前

两件事同时发生:

- AI 编码模型趋同。对于 95% 的任务,大多数人无法区分 GPT-5.6 Sol、Fable 5、Kimi K3、GLM-5.2 或 Qwen 3.8 Max 哪个更聪明。

- GLM-5.3 今天发布,仅 743B 参数。更小、更便宜的开源模型正在快速涌现。

重心转移的速度可能比人们想象的更快。

查看英文原文
2 things are happening at once:

- AI coding models are converging. For 95% of tasks, most people can’t tell whether GPT-5.6 Sol, Fable 5, Kimi K3, GLM-5.2, or Qwen 3.8 Max, is smarter.

- GLM-5.3 dropped today, only 743B parameters. Smaller, cheaper, open models are coming fast.

The center of gravity may be shifting faster than people think.
◔ 28.7 万 次浏览(4 条合计)♥ 3,569⇄ 158观点看原帖 ↗
Yuchen Jin@Yuchenj_UW · 博主 · 23 小时前

GPU内核之王Bob离开了OpenAI?

GPT-6.7的运行速度会慢2倍....

引用 Arfur Grok @ArfurGrok👀 Scott Gray ( @scottgray76 ) has left OpenAI.查看被引原帖 ↗
查看英文原文
Bob, the king of GPU kernels, left openai?

GPT-6.7 will run 2x slower….
Greg Brockman@gdb · 创始人 · 1 天前Greg Brockman,OpenAI 联合创始人兼总裁

在 ChatGPT 里试试预订搜索功能吧

引用 Adam Fry @adamhfryThis week's ChatGPT feature drop - Aug 14: 1/ Quizzes - "Quiz me on [topic]" will now automatically generate quizzes right in Chat. Learn any topic! 2/ Reservation search - "I'm looking for a reservation for two next Saturday for a place with a great outdoor patio". Nice thing is you can just describe what you want. 3/ Google Drive for Paid Users - You can now add Google Drive files to your ChatGPT Library. That makes it easy to ask ChatGPT questions about those files. Just hit the + menu and "Add from Library" 4/ Suggestions for Paid Users - Paid users will see high quality suggestions on the home page of ChatGPT We're continuing to make answers better, allow you to pull in the right documents, and more easily discover what Chat can do. Step by step, ChatGPT gets more useful! Leave feedback in the comments.查看被引原帖 ↗
查看英文原文
try reservation search in chatgpt!
◔ 14.9 万 次浏览♥ 841⇄ 31▶ 含视频新品看原帖 ↗
Chubby♨️@kimmonismus · 博主 · 23 小时前Chubby,高频 AI 新闻聚合博主

Scott Gray 离职了 OpenAI。他从 2016 年加入公司,帮助构建了其早期突破背后的 GPU 基础架构,从 sparse transformers 和 scaling laws 到 GPT-3。

他的离职正值至少 12 名高管已在 2026 年离职 OpenAI 之际。

OpenAI 2026 年的离职名单:

• Scott Gray - 资深 GPU 系统工程师
• Denise Dresser - 首席营收官
• Brad Lightcap - 前首席运营官
• Fidji Simo - 应用部门 CEO
• Kevin Weil - OpenAI for Science 负责人
• Bill Peebles - Sora 负责人
• Srinivas Narayanan - B2B 应用 CTO
• Kate Rouch - 首席营销官
• Barret Zoph - 企业销售负责人
• Johannes Heidecke - 安全系统负责人
• Joshua Achiam - 首席未来主义官
• Chloé Bakalar - 伦理负责人
• Caitlin Kalinowski - 硬件和机器人负责人

OpenAI 发生了什么。

引用 Arfur Grok @ArfurGrok👀 Scott Gray ( @scottgray76 ) has left OpenAI.查看被引原帖 ↗
查看英文原文
Scott Gray leaves OpenAI. He joined the company in 2016 and helped build the GPU infrastructure behind its earliest breakthroughs, from sparse transformers and scaling laws to GPT‑3.

His exit comes as at least 12 senior leaders have already left OpenAI in 2026.

OpenAI’s 2026 departures so far:

• Scott Gray - longtime GPU systems engineer
• Denise Dresser - Chief Revenue Officer
• Brad Lightcap - former COO
• Fidji Simo - CEO of Applications
• Kevin Weil - Head of OpenAI for Science
• Bill Peebles - Head of Sora
• Srinivas Narayanan - CTO of B2B Applications
• Kate Rouch - Chief Marketing Officer
• Barret Zoph - Enterprise Sales lead
• Johannes Heidecke - Head of Safety Systems
• Joshua Achiam - Chief Futurist
• Chloé Bakalar - Head of Ethics
• Caitlin Kalinowski - Head of Hardware and Robotics

Something is happening at OpenAI.
Chubby♨️@kimmonismus · 博主 · 1 天前Chubby,高频 AI 新闻聚合博主

听说 Codex 下周要推一个重大性能更新,能让超长对话加载速度飙升,同时占用更少内存和请求数。在一个极端基准测试场景(741 轮对话)中,平均加载时间从 27.6 秒直接降到了 1.7 秒。对这次更新太期待了

引用 Andrew Ambrosino @ajambrosinoalways nice to see internal slack messages like this– thanks @btraut查看被引原帖 ↗
查看英文原文
Hype: Codex is rolling out a major performance update next week that should make very long conversations load dramatically faster while using less memory and fewer requests.

In an extreme 741-turn benchmark, average load time dropped from 27.6 to 1.7 seconds.

Super hyped for that update
Simon Willison@simonw · 博主 · 1 天前Django 框架联合创造者,AI 工具深度评测

新的 Qwen 3.7 27B,在我的 M5 Max 笔记本上用 LM Studio 跑 17GB GGUF 版本,刚给我画了一只最棒的骑自行车的鹈鹕 - 这是我笔记本能跑的任何模型中效果最好的

查看英文原文
The new Qwen 3.7 27B, running as a 17GB GGUF in LM Studio on my M5 Max laptop, just drew me the best pelican riding a bicycle I've seen from any model that runs on my laptop
ollama@ollama · 公司官方 · 1 天前本地跑大模型的热门工具

Qwen 3.8 27B 现在可在 Ollama 上使用。

这是同等规模最好的开源模型之一,专为代理任务和专业工作设计。

直接在你使用的应用和框架中尝试:

Claude Code:
ollama launch claude --model qwen3.8

OpenCode:
ollama launch opencode --model qwen3.8

Hermes Agent:
ollama launch hermes --model qwen3.8

Pi:
ollama launch pi --model qwen3.8

我们还针对 Apple Silicon 进行了优化!

使用模型名称尝试:qwen3.8:27b-mlx

查看英文原文
Qwen 3.8 27B is now available on Ollama.

It's one of the best open models at this size, and made for agentic tasks and professional work.

Try it directly with the apps & harnesses you use:

Claude Code:
ollama launch claude --model qwen3.8

OpenCode:
ollama launch opencode --model qwen3.8

Hermes Agent:
ollama launch hermes --model qwen3.8

Pi:
ollama launch pi --model qwen3.8

We have also optimizations for Apple Silicon!

Try it with the model name: qwen3.8:27b-mlx
◔ 20.4 万 次浏览(2 条合计)♥ 1,828⇄ 182新品看原帖 ↗
Lisan al Gaib@scaling01 · 博主 · 1 天前高频 AI 模型测评与爆料博主

Anthropic 这4个月到底干啥呢?

引用 Lisan al Gaib @scaling01Model 2 beats Mythos 5 and Claude Mythos Preview on CoBench V2 an Anthropic internal benchmark measuring AI R&D capabilities查看被引原帖 ↗
查看英文原文
what exactly did Anthropic do for the last 4 months?
el.cine@EHuanglu · 博主 · 1 天前

AI 出品火影最好的一集

查看英文原文
AI made best ep of Naruto
◔ 11.5 万 次浏览♥ 726⇄ 69▶ 含视频演示看原帖 ↗
宝玉@dotey · 中文博主 · 1 天前宝玉,中文圈 AI 翻译与科普大 V

这篇 Pi 压缩的文章,太过于朴实无华,就真的只是写个 prompt 让 LLM 把上下文总结一下,然后保留前面的system prompt 和工具调用,在摘要后可能还会保留最近几次对话。

这种压缩是有损的,不知道是不是有机制会去历史会话检索上下文?

当然这确实是压缩上下文的最简单有效方案。

引用 Pi @pidotdevLLM上下文窗口受限,长对话影响输出质量、性能和成本。Earendil工程师@vegardstikbakke撰文介绍上下文压缩如何解决此问题,以及在Pi中的具体实现方式。查看被引原帖 ↗
◔ 11.2 万 次浏览♥ 459⇄ 60▶ 含视频观点看原帖 ↗
NVIDIA@nvidia · 公司官方 · 1 天前

开放性能让AI更安全吗?@ctnzr提出:更多人评估一项技术意味着更多的审视和更多人帮助使其更安全。

查看英文原文
Can openness make AI safer?


@ctnzr
makes the case: more people evaluating a technology means more scrutiny and more people helping make it safer.
◔ 8.7 万 次浏览♥ 228⇄ 26▶ 含视频观点看原帖 ↗
Google AI@GoogleAI · 公司官方 · 1 天前谷歌 AI 官方账号

终于周五了 🎉 周末回顾来一下:

— 今年的 @madebygoogle 产品线(Pixel 11 系列、Pixel Watch 5 和 Pixel Tag)为各设备带来了新的 AI 功能。关键公告包括用于同步视频和照片拍摄的 Magic Capture、Rambler 的 AI 语音输入和文本转换,以及扩展的 Live Transcribe 用于通过 Pixel 摄像头进行实时 ASL 转文本翻译。贯穿这一切的是 Gemini Intelligence,我们的主动式、具有 agent 能力的 AI 层,用于预测用户需求。

— Gemini 3.7 Flash,我们用于编码和 agent 的最聪明的多面手模型,现已在 Gemini API 中通过 @googleaistudio、@antigravity、@geminiapp 中的 Spark 和 Gemini Enterprise Agent Platform 提供。同时也在向 @googleworkspace、Gemini App 和 Search 的付费用户推出。

— @GoogleDeepMind 的 WeatherNext 2 是一款 AI 预报工具,在预测气旋路径、强度和风场结构时可以为气象学家提供额外的精准预报天数,现已开源。

— 升级的 @Gemini_Notebook 体验已完全向所有 Pro 用户推出,复制笔记本的功能已向所有用户推出。

查看英文原文
It’s (finally) Friday 🎉 Here’s our end-of-week recap:

— This year’s
@madebygoogle
lineup (Pixel 11 series, Pixel Watch 5, and Pixel Tag) brings new AI integrations across devices. A few of the key announcements were Magic Capture for simultaneous video and photo capture, Rambler’s AI voice typing and text transformation, and expanded Live Transcribe for real-time ASL-to-text translation using the Pixel camera. Tying it all together is Gemini Intelligence, our proactive, agentic AI layer to anticipate user needs.

— Gemini 3.7 Flash, our most intelligent workhorse model yet for coding and agents, is now available in the Gemini API via
@googleaistudio
,
@antigravity
, Spark in the
@geminiapp
, and the Gemini Enterprise Agent Platform. It’s also rolling out to paid users on
@googleworkspace
, Gemini App, and Search.


@GoogleDeepMind
's WeatherNext 2, an AI forecasting tool that can give meteorologists an extra day of accuracy when predicting a cyclone's track, intensity, and wind structure, is now open source.

— The upgraded
@Gemini_Notebook
experience has been fully rolled out to all Pro users and the ability to copy a notebook has been rolled out to all users.
Chubby♨️@kimmonismus · 博主 · 1 天前Chubby,高频 AI 新闻聚合博主

Axios 证实,Anthropic 超越 Mythos 5 的新模型不会推出

引用 Chubby♨️ @kimmonismusAs has already been expected, Anthropic internally uses a model that is significantly better than Mythos 5, but they have no plans to release it.查看被引原帖 ↗
查看英文原文
Axios confirms, new Anthropic model that surpasses Mythos 5 will not be rolled out
ollama@ollama · 公司官方 · 1 天前本地跑大模型的热门工具

Ollama 现在支持 DeepSeek Harness 了。

ollama launch dsh

完全在你的本地环境里运行。它预装了 Ollama 的网页搜索功能。你可以通过它的轨迹视图来查看后台发生了什么。

引用 DeepSeek @deepseek_ai🧩 DeepSeek Harness v0.1 is now available in Developer Preview! 🔹 We’re opening it up to developers building agent harnesses worldwide and open-sourcing the codebase in MIT license. 🔹 Powered by the Cordis meta-framework, DeepSeek Harness is an agent harness built around one core idea: Everything is a plugin. Models, tools, skills, sessions, sandboxes, filesystems, loops, orchestration, and UI are ALL implemented as plugins, and can be mixed, matched, replaced, and extended. Try it now! github.com/deepseek-ai/deeps…查看被引原帖 ↗
查看英文原文
Ollama now supports the DeepSeek Harness.

ollama launch dsh

Run it completely in your own environment. It comes with Ollama's web search pre-installed. You can use its trajectory view to see what is happening in the background.
◔ 17.5 万 次浏览(3 条合计)♥ 1,015⇄ 91▶ 含视频新品看原帖 ↗
Aravind Srinivas@AravSrinivas · 创始人 · 1 天前Perplexity 联合创始人兼 CEO

Perplexity 👑

引用 OpenRouter @OpenRouter推出 Web Search Benchmarks 🌐。跨不同模型和配置的搜索工具排名,帮助您决定如何为 agent 提供信息依据。查看被引原帖 ↗
查看英文原文
Perplexity 👑
Hugging Face@huggingface · 公司官方 · 1 天前全球最大 AI 开源模型社区

开源模型现状,2026 年夏 ☀️ 前沿模型体量在增大,但小模型仍然主导实际应用。Qwen 在本地推理领域领先,其次是 Gemma。AI agents 正在成为 Hub 上的主要力量。完整内容见博客 🤗 huggingface.co/blog/state-of…

查看英文原文
The State of Open Models, Summer 2026 ☀️ frontier models are getting larger, but small models still dominate real-world usage. Qwen leads local inference, followed by Gemma. AI agents are becoming a major force on the Hub

Full picture on the blog 🤗

huggingface.co/blog/state-of…
swyx@swyx · 博主 · 1 天前知名 AI 播客 Latent Space 主理人

原因就是
@databricks
“IPO是熔岩”募资成了段子,但这次是真话

他们1880亿美元M轮里的M,代表“我们要干掉无数次会议”

引用 Ali Ghodsi @alighodsiI got this question so many times today. "How can you grow 80% at $7B?" The true answer is that we're finally seeing a breakthrough with AI agents starting to work in the enterprise. The AIs have been super smart for a while, but have lacked basic context that's in people's heads, or in some SaaS system-or-record. A lot of organizations are deploying FDEs to capture this context, or Ontology, and feed it to the AI. This is labor intensive and expensive. We just automated that with Genie Ontology. Once you have that enterprise context graph, an AI agent like Genie becomes magical. I find myself no longer waiting for answers from my CRO, CFO, CMO, CHRO etc, I just keep queuing up questions on the phone while sitting in meetings. It'd frankly addictive. Our customers are starting to do the same, over 70% of all queries on the platform are now generated by Genie agents. This fuels more questions to the platform, which drives consumption, which drives revenue. That's the simple answer.查看被引原帖 ↗
查看英文原文
the reason
@databricks
"ipo is lava" fundraises are a meme is this but unironically

the M in their $188B series M stands for "we are going to kill so many meetings"
ollama@ollama · 公司官方 · 1 天前本地跑大模型的热门工具

新的 DeepSeek-V4-Pro (0813) 已在 Ollama 云全量推出,并纳入 Pro 和 Max 订阅。

ollama run deepseek-v4-pro:cloud

美国托管,零数据保留(ZDR),性能强劲。

查看英文原文
The new DeepSeek-V4-Pro (0813) is now fully rolled out on Ollama's cloud and included in Pro and Max subscriptions.

ollama run deepseek-v4-pro:cloud

Hosted in the US with Zero Data Retention (ZDR) and high performance.
Qwen@Alibaba_Qwen · 公司官方 · 1 天前阿里通义千问大模型团队

笔记本大小的模型,前沿级别的飞跃。Qwen3.8-27B 已上线 LM Studio。快试试!
@lmstudio

引用 LM Studio @lmstudioQwen3.8-27B is here! 🚀 It's a leap in capabilities for a laptop size model. Requires ~17GB to run locally. Model page: lmstudio.ai/qwen/qwen3.8-27b查看被引原帖 ↗
查看英文原文
Laptop-size model, frontier-size leap. 🏃‍♀️Qwen3.8-27B is live on LM Studio. Try it!
@lmstudio
◔ 9.9 万 次浏览(2 条合计)♥ 822⇄ 52新品看原帖 ↗
Guillermo Rauch@rauchg · 创始人 · 1 天前Guillermo Rauch,Vercel 创始人兼 CEO

Vercel 是世界上最快的 AI Gateway 基础设施

引用 ComputeSDK @computesdkAI Gateway排行:金牌Vercel,银牌Pydantic,铜牌LLMGateway查看被引原帖 ↗
查看英文原文
Vercel is the fastest [AI Gateway] infrastructure in the world
歸藏(guizang.ai)@op7418 · 中文博主 · 1 天前歸藏,中文圈 AI 工具与提示词博主

推特算法昨天开源的算法又迎来了一波更新。分析了一下,主要是需要做了这三个调整:

普通视频新增了 14 天的 SID 语义召回窗口: 所以长期的高质量视频有可能会获得二次分发机会。

关注与双向关系进入了推荐模型: 模型能识别真实关注和双向关系,所以稳定回访的话可能会获得一些加权。

回复排序分界从 1.5 万粉丝升级到 3 万粉丝: 如果跟帖作者和回复上一层的作者粉丝量超过 3 万(原来是超过 1.5 万),就按高传播度对评论区的回复进行排序处理;如果不超过 3 万,就会进入到垃圾回复的检测。

看起来他们是有反垃圾回复系统的,就是不知道为什么中文这部分一直做不好。

引用 歸藏(guizang.ai) @op7418Twitter 完全开源了他们的推荐算法 用 Codex 分析了一下,感觉跟以前我们的认知还是有不少变化的。 总结了六条创作者应该做的事情: 1. 应该做值得转发的原创内容。 2. 尽量少用首贴写钩子,把重要内容放在第二条推串里这种发帖形式(这个和大家的认知不太一样,比较重要)。 3. 应该优先争取阅读、回复、引用和关注,点赞的权重其实没有那么高。 4. 拉开发帖间隔,不要频繁刷屏。 5. 深耕一个垂类,不要频繁更换账号类型。 6. 少做那些互动诱饵(比如互关、回复发送等),不要让别人讨厌你、对你点“不感兴趣”或举报。查看被引原帖 ↗
Lisan al Gaib@scaling01 · 博主 · 1 天前高频 AI 模型测评与爆料博主

我现在毫无疑问地认为 Mythos Preview 是他们庞大的 ~10T 教师模型

我认为 Model 1 和 Model 2 是同一模型的进一步迭代

Mythos 5 和 Fable 5 可能只是更小的蒸馏模型

引用 Lisan al Gaib @scaling01Model 2 beats Mythos 5 and Claude Mythos Preview on CoBench V2 an Anthropic internal benchmark measuring AI R&D capabilities查看被引原帖 ↗
查看英文原文
I have no doubt in my mind anymore that Mythos Preview is their massive ~10T teacher model

I think Model 1 and Model 2 are further iterations of that same model

Mythos 5 and Fable 5 are likely only smaller distilled models
Qwen@Alibaba_Qwen · 公司官方 · 1 天前阿里通义千问大模型团队

现在就可以用了!😎 看到 Qwen3.8-27B 在 RTX Spark 上跑起来真是太棒了。下载、部署,创造新东西。
@NVIDIARTXSpark

引用 NVIDIA RTX Spark @NVIDIARTXSparkReady to run Qwen3.8 locally? 👀 Qwen3.8-27B packs powerful AI into an open model developers can download, serve locally, and build with on their own terms.查看被引原帖 ↗
查看英文原文
Ready right now! 😎Incredible to see Qwen3.8-27B fly on RTX Spark. Download, deploy and create somthing new.
@NVIDIARTXSpark
Aravind Srinivas@AravSrinivas · 创始人 · 1 天前Perplexity 联合创始人兼 CEO

Perplexity 的 Search SDK 现在可在任何代理框架中使用,它让 Perplexity Computer 成为进行广泛深度研究的同类最佳产品!

引用 Perplexity Developers @perplexitydevsPerplexity推出agent优先的Python SDK,将Search as Code方法带到应用中。Agent可并行执行多个搜索,在代码中过滤、去重和排序结果。查看被引原帖 ↗
查看英文原文
Perplexity's Search SDK, which makes Perplexity Computer the best-in-class product for wide and deep research, is now available to use inside any agentic harness!
◔ 5.6 万 次浏览♥ 288⇄ 21▶ 含视频新品看原帖 ↗
🚨 AI News | TestingCatalog@testingcatalog · 博主 · 1 天前专挖 AI 产品未发布新功能的爆料号

OpenAI消息:ChatGPT最近推出了一堆更新。用户现在可以自动从聊天生成测验、搜索预订、将Google Drive文件加入ChatGPT库,还能在首页看到更好的建议(付费计划限定)。

鸣谢@adamhfry

查看英文原文
OPENAI 👀: ChatGPT got a bunch of updates recently. Users can now automatically generate quizzes from the chat, search for reservations, add Google Drive files to the ChatGPT Library, and get better suggestions on the home page (paid plans only).

h/t
@adamhfry
◔ 5.3 万 次浏览♥ 789⇄ 41▶ 含视频新品看原帖 ↗

你问他第一门编程语言应该学什么,他告诉人生第一门语言必须只能是lisp,入门编程书必须是SICP,告诉你这才是学计算机的正道;

你问他操作系统应该是macOS、windows还是ubuntu,他告诉你必须选择arch linux,每个人都应该且只用这一个;

你问他你用vscode还是notepad++还是jetbrains系列, 他告诉你必须用emacs,他告诉你高手只用emacs,然后邪魅一笑;

你问他对deep learning有什么意见,他说pytorch是个巨大的灾难,torch完全应该由lua转向haskell重写一遍,而绝对不应该用python写成pytorch,他认为所有做deep learning的人都是大傻逼,都应该回去从本科catagory theory开始回炉重造。

现在他来告诉你们,他认为AI Agent harness必须全部插件化,否则都是路线错误……

所以我一直说,

非pure math背景的人,现在当务之急不是狠狠蹬claude和gpt去严格证明现有任何猜想,

也不是去狠狠蹬claude和gpt去爆破任何现有猜想找反例,

而是狠狠蹬claude和gpt生成10000个猜想,然后自己爆破找反例过滤掉9900个,剩下100个不能证明也不能爆破找反例的,直接发表《我的百大猜想》

引用 Dr. Yin @JunYin29422166被朋友要求, 评价一下最近“Crouzeix 猜想”被证明一事。 没听说过这个猜想, 也没听过Michel Crouzeix是谁。 查了查wiki上研究这个问题所发的杂志, 和Crouzeix本人所发杂志,都没有很强的数学刊物,应该都没达到CPAM的级别。 我觉得这个猜想肯定算不上大猜想。查看被引原帖 ↗
Cohere@cohere · 公司官方 · 1 天前加拿大企业级大模型公司 Cohere 官方

兜兜转转
@UofT


2017: 2026:

查看英文原文
Full circle
@UofT


2017: 2026:
Qwen@Alibaba_Qwen · 公司官方 · 1 天前阿里通义千问大模型团队

Qwen3.8-27B 现在已经成了你日常生活的一部分——从手机到汽车。⚡️🚀感谢你们的工作!
@MediaTek

引用 MediaTek @MediaTekCongratulations to the @Alibaba_Qwen on the launch of Qwen 3.8! MediaTek continues our long-standing collaboration with Qwen, with Day-0 support now available on Dimensity Auto Cockpit C-X1 and the latest Dimensity flagship mobile SoC, bringing smarter and more intuitive on-device AI experiences to more people, from smartphones to vehicles, and helping make agentic AI a seamless part of everyday life.查看被引原帖 ↗
查看英文原文
Qwen3.8-27B is now part of your everyday life — from smartphones to vehicles. ⚡️🚀Appreciate your work!
@MediaTek
OpenRouter@openrouter · 公司官方 · 1 天前
连环推 ×3

这个周末我们给你一张 $10 的 Ori Harness 优惠券用来构建

你会构建什么?

用你最喜欢的 agent 运行 - Claude Code、Codex、DeepSeek 等 - 拥有 500+ 模型、70+ 提供商,全部在一个账户中,在 OpenRouter 上

兑现很简单:

查看英文原文
We're giving you a $10 coupon to build on Ori Harness this weekend

What will you build?

Run your favorite agent - Claude Code, Codex, DeepSeek, and more - with 500+ models, 70+ providers, all in one account, on OpenRouter

Redeeming is easy:
We pinned the code in the
#ori
channel in the OpenRouter Discord. Grab it there, redeem at
openrouter.ai/redeem
, then tell us what breaks and what you want next:

Join the
#ori
channel:
discord.com/channels/1091220…
We're giving out 100 coupons, $10 each, one per person, first come first served

To get started with Ori Harness:

$ curl -fsSL
openrouter.ai/labs/ori/insta…
| bash

Use any harness:

$ ori {claude, codex, opencode, hermes, pi, deepseek, etc}

Read more at:
openrouter.ai/ori/harness
◔ 4.5 万 次浏览♥ 352⇄ 26▶ 含视频新品看原帖 ↗
Orange AI@oran_ge · 中文博主 · 1 天前Orange AI,中文圈 AI 产品观察博主

DSH 是技术自嗨的巅峰之作
从某种意义上来说 Transformer 也是

向阳乔木@vista8 · 中文博主 · 1 天前向阳乔木,中文圈 AI 工具与趋势博主

想做个小调研,你现在还在用语音输入法跟AI对话吗?

如果是,现在会用什么工具?
如果不是,为啥不用了。

感觉自己最近用的很少...

Qwen@Alibaba_Qwen · 公司官方 · 1 天前阿里通义千问大模型团队

小模型,大实际影响。很荣幸看到 Qwen 在 State of Open Models 中领先本地推理。享受来自 Qwen 的阳光。☀️ 😎
感谢你们的工作!
@huggingface

引用 Hugging Face @huggingface开源模型现状,2026年夏季 ☀️ 前沿模型规模扩大,但小模型仍主导实际应用。Qwen领导本地推理,其次是Gemma。AI agents成为Hub主要力量。详见博客 🤗查看被引原帖 ↗
查看英文原文
Small models, big real-world impact.
Proud to see Qwen leading local inference in the State of Open Models. Enjoy the sunshine from Qwen.☀️ 😎
Appreciate your work!
@huggingface
🚨 AI News | TestingCatalog@testingcatalog · 博主 · 1 天前专挖 AI 产品未发布新功能的爆料号

Anthropic 🔥:Claude 桌面版将在应用内内置一个专用浏览器,在 Cowork 会话中一直可用!一个新的超级应用正在崛起。

这看起来很眼熟 👀

查看英文原文
ANTHROPIC 🔥: Claude desktop will get a dedicated Browser inside that app that will always be accessible in Cowork sessions! A new Super App is rising.

This looks familiar 👀
Matt Shumer@mattshumer_ · 博主 · 1 天前HyperWrite CEO,AI 实战技巧分享

还在运行中!我应该再让它运行一天吗?

Gauntlet Looping 指南:
somethingbig.ai/gauntlet-loo…

引用 Matt Shumer @mattshumer_Grok 4.6 worked non-stop for 48 hours to build this shooter. Turns out Grok is powerful enough to run Gauntlet Loops. Let the game-making begin!查看被引原帖 ↗
查看英文原文
It’s still going! Should I let it run for another day?

Guide to Gauntlet Looping:
somethingbig.ai/gauntlet-loo…
◔ 3.8 万 次浏览♥ 153⇄ 7▶ 含视频教程看原帖 ↗
ollama@ollama · 公司官方 · 1 天前本地跑大模型的热门工具

与 @Alibaba_Qwen 团队的合作很愉快。

引用 Qwen @Alibaba_QwenQwen3.8-27B现已在Ollama上可用!无论你运行什么工具,这个模型总是能够交付。让我们一起编码,展示你的成果。😎查看被引原帖 ↗
查看英文原文
It's been a pleasure to partner with the
@Alibaba_Qwen
team.
Gary Marcus@GaryMarcus · 博主 · 1 天前

还记得我说过OpenAI可能会成为AI界的WeWork吗?

引用 Ricardo @Ric_RTPOpenAI is falling apart right now. 9 of their most important leaders have left the company recently, and Altman is about to ask the public to buy the stock. 2 of them even walked out within 72 hours of OpenAI handing its own staff $7 billion in cash... On Monday, August 10, OpenAI completed a deal letting current and former employees sell roughly $7 billion worth of their shares. The price valued the company at $852 billion, the exact same number as its March funding round. On Tuesday, August 11, Brad Lightcap announced he was leaving after 8 years. He spent 4 of them as chief financial officer, then ran the company as chief operating officer from 2022 until April. He worked alongside Sam Altman at Y Combinator before OpenAI existed. On Thursday, August 13, chief revenue officer Denise Dresser announced she was leaving. She was hired in December from Salesforce, where she had been the CEO of Slack. In April she took over most of Lightcap's responsibilities. She lasted 8 months. The cash window opened Monday. By Thursday both executives who ran the business side were gone. But what's interesting is who actually wrote the $7 billion cheque: Every previous time OpenAI let its employees cash out, an outside investor bought the shares. In October, Thrive Capital, SoftBank and others put up $6.6 billion at a valuation near $500 billion. There was a $1.5 billion version of the same deal in 2024. This time OpenAI bought the shares back itself, using its OWN money. So no outside investor put a single dollar behind that $852 billion price. The company named its own number and then paid it. This is a business generating around $2 billion a month while losing roughly $1.22 for every single dollar it earns. And it just spent $7 billion of that cash buying its own stock at a number no third party ever tested. Here is the full list of the people who left since April: - Bill Peebles, who ran the Sora video app - Kevin Weil, vice president of OpenAI for Science - Srini查看被引原帖 ↗
查看英文原文
remember how i said OpenAI might turn out to be the WeWork of AI?
◔ 3.7 万 次浏览♥ 488⇄ 69▶ 含视频观点看原帖 ↗
🚨 AI News | TestingCatalog@testingcatalog · 博主 · 1 天前专挖 AI 产品未发布新功能的爆料号
连环推 ×2

Atomic 发布了 Qwen3.8-27B-GGUF 的 Dynamic Quants,这是 Qwen3.8 27B 的量化版本,从 8-bit (28.9GB) 压缩到 1-bit (8.5GB)。

> 适配 16GB MacBook Air;AD-IQ3_S 版本大小为 13.8GB。

> Atomic dynamic GGUF 在 92.4% 的时间内选择与 BF16 相同的下一个 token。

Opus 4.6 级别的笔记本模型 👀

引用 atomic.chat @atomic_chat_hqRun Qwen3.8 27B locally via Atomic Chat💥 We released Atomic Dynamic GGUF quants, from 8-bit (28.9 GB) down to 1-bit (8.5 GB), and measured all other Qwen3.8 GGUFs in the community AD-IQ3_S runs on a 16GB MacBook Air and picks the same next token as the BF16 original 92.4% of the time查看被引原帖 ↗
查看英文原文
Atomic has released Dynamic Quants for Qwen3.8-27B-GGUF, a quantized version of Qwen3.8 27B, compressed from 8-bit (28.9GB) to 1-bit (8.5 GB).

> Fits inside a 16GB MacBook Air; the AD-IQ3_S build lands at 13.8GB.

> Atomic dynamic GGUF picks the same next token as BF16 in 92.4% of the time.

Opus 4.6-level model on a laptop 👀
The full ladder by memory tier, measured on held-out eval at 4096 context against a BF16 reference:

> 12GB, AD-IQ2_XS, 83.5%
> 16GB, AD-IQ3_S, 92.4% 🔥
> 24GB, AD-Q5_K, 97.3%
> 32GB, AD-Q6_K, 98.7%
> 48GB, Q8_0, 98.9%

Atomic measured every other Qwen3.8 GGUF in the community against the same reference and reports closer AD files to BF16 at most file sizes.

Test it out 👀

atomic.chat
Chubby♨️@kimmonismus · 博主 · 1 天前Chubby,高频 AI 新闻聚合博主

2026年正成为开放与本地AI的里程碑之年。

- Zai表示,GLM-5.3完全通过在GLM-5.2相同基座模型上的后训练改进,在CyberGym上拿下84.5%,领先Mythos 5的83.8%。

- Qwen3.8-27B可本地运行,在多项关键基准上击败Opus 4.6 Max。

- API价格低至每百万未缓存输入token 0.14美元、每百万输出token 0.28美元,DeepSeek V4 Flash几乎让人感叹“便宜到不用计费”。

为开放AI和本地AI干杯。智能人人可享!

引用 Chubby♨️ @kimmonismusSo yeah, you can now run Opus 4.6 max locally. Just let that sink in.查看被引原帖 ↗
查看英文原文
2026 is shaping up to be a landmark year for open and local AI.

-Zai says GLM-5.3, improved entirely through post-training on the same base model as GLM-5.2, scored 84.5% on CyberGym, ahead of Mythos 5’s 83.8%.

-Qwen3.8-27B runs locally and beats Opus 4.6 Max on several key benchmarks.

-With API prices as low as $0.14 per million uncached input tokens and $0.28 per million output tokens, DeepSeek V4 Flash comes remarkably close to “too cheap to meter.”

Here’s to open AI and local AI. Intelligence for everyone!
Bindu Reddy@bindureddy · 创始人 · 1 天前Abacus.AI CEO,AI 行业观点博主

太离谱了!基本上每个模型都声称自己是 SOL/Fable 级别,两天火了然后就被彻底遗忘...两天后又有个模型出来声称一样的东西,前一个模型完全被忘了。如果这些都是真的,Anthropic 和 OpenAI 应该破产了!!结果呢,它们的营收却在天文数字地增长!它们会比 Google、MSFT 和 Amazon 加起来还大 😲

查看英文原文
This is getting ridiculous!

Literally every model is claiming it is SOL/Fable class and goes viral for two days…

Two days later some other model drops and claims the same thing and the first model is fully forgotten

Anthropic and OpenAI should be out of business, If any of this is to be believed!!

Instead their revenue is growing astronomically! They will be bigger than Google, MSFT and Amazon combined 😲
Min Choi@minchoi · 博主 · 23 小时前AI 产品演示博主,专门展示新工具玩法

这太狂了。

特斯拉 FSD 在西班牙日食中驾驶。

查看英文原文
This is wild.

Tesla FSD driving through a solar eclipse in Spain.
◔ 3.2 万 次浏览♥ 259⇄ 19▶ 含视频动态看原帖 ↗
Lisan al Gaib@scaling01 · 博主 · 1 天前高频 AI 模型测评与爆料博主
连环推 ×4

Mythos Preview 比 Mythos 5 和 Fable 都大(很可能接近 10T),理由如下:

— 它的定价是 $125/百万输出 token(是 Opus 的 5 倍,Fable 的 2.5 倍。而且我们知道 Opus 大概是 1.5T-2T,这么推算的话 Mythos 应该是在 7.5-10T 左右)

— Mythos Preview 虽然在两个多月前就发布了,但在好几个基准测试上都比 Mythos 5 更强。你总不会觉得多加两个月的训练反而把模型分数练低了吧。

(值得注意的是,它在 UK AISI cyber ranges、GPQA、HLE + tools、OSWorld、大部分病毒学任务上都超过了 Mythos 5。另外,之前扩规模的时候(比如 GPT-4.5)我们看到了更好的校准效果,这在 SimpleQA 上表现得很明显,而 Mythos Preview 在 SimpleQA 上也赢了 Mythos 5。它还打败了 Mythos 5 在一个评估“偷懒式诚实”的 Anthropic eval 上的成绩)

— Anthropic 拥有最大的计算集群,来自 Project Rainier,首次用于训练大约是在 2025 年 10 月到 11 月,然后 2026 年 2 月 Anthropic 突然就在内部推出了 Mythos Preview。算一下的话大概是 5e26 到 1e27 flops,这正好符合一个 ~10T 模型的预期

— 从模型的“气质”来看,Fable 明显不是 ~10T 的料,只是比 Kimi-K3 稍微大一点,也就是 3-5T,而这正好和 Mythos Preview 跟 Fable 的定价吻合。10T / ($125/$50) = 4T

— Amodei、Kaplan 还有其他人在博客里反复强调 Mythos Preview 证明了 scaling law 依然坚挺,暗示着这是一次更大的跃升,而不只是 Opus 的 1.5-2 倍

说实话,Anthropic 有个 ~10T 模型这件事基本不用怀疑了

引用 Teortaxes▶️ (DeepSeek 推特🐋铁粉 2023 – ∞) @teortaxesTex1) DeepSeek did *not* “shit the bed” 2) serious time: anon *why* do you believe in the existence of “10T models” or “Mythos teacher”? What convinced you? We see a 16BA hitting 80% on ARC-AGI-2. *do you actually think* 1-2 evals where Mythos-Preview > Mythos are enough evidence?查看被引原帖 ↗
查看英文原文
Mythos Preview is larger than Mythos 5 and Fable (and likely close to 10T) for several reasons:

- it's pricing was $125/million output tokens (5X Opus and 2.5x Fable. and for Opus we know it's around 1.5T-2T, implying Mythos should be around 7.5-10T)

- Mythos Preview is stronger than Mythos 5 on several benchmarks despite being launched over 2 months earlier. You would have to believe that 2 months of further training decreased model scores.

(notably it beats Mythos 5 on UK AISI cyber ranges, GPQA, HLE + tools, OSWorld, most virology tasks. Also with previous scale-ups (GPT-4.5) we saw much better calibration, which showed up in SimpleQA and Mythos Preview beats Mythos 5 on SimpleQA. It also beats Mythos 5 on an Anthropic eval that measures missing-reference honesty)

- anthropic had the largest cluster with project rainier, which was first used for training around October-November 2025, then in February 2026 Anthropic suddenly had mythos preview available internally. When you do the math on that you get out 5e26 - 1e27 flops, which is exactly what you would expect for a ~10T model

- based on models vibes you can tell that Fable is not a ~10T model but only a bit larger than Kimi-K3, so 3-5T, which is exactly what you would guess based on the pricing of Mythos Preview and Fable. 10T / ($125/$50) = 4T

- Amodei, Kaplan, others and their blogs repeat multiple times that Mythos Preview is proof that scaling laws are well and alive, implying a larger scale-up and not just a mere 1.5-2x over Opus

it's not really a question that anthropic has a ~10T model
and yes DeepSeek shit the bed. they are currently behind Anthropic, OpenAI, Moonshot AI, ByteDance, xAI, Meta, Alibaba and ZAI
but you know what they say about the the river in Egypt
and I do not believe in silly theories such as "they optimized inference for mythos preview and made it 2.5x cheaper to serve"

because this assumes that they trained the model inefficiently, wasting gazillions of dollars and then served the model internally for another 2 months in that unoptimized state
🚨 AI News | TestingCatalog@testingcatalog · 博主 · 1 天前专挖 AI 产品未发布新功能的爆料号
连环推 ×2

AI/ML API 用相同任务测试了 Qwen 3.8 27B(自托管)和 Claude Opus 4.6:创建一个独立的 three.js 文件,渲染沉没的亚特兰蒂斯,并附带脚本化的 20 秒摄像机飞行。

两个模型都处理了完整的摄像机路径,Qwen 3.8 27B 自托管免费运行,而 Claude Opus 4.6 花费 $0.21。

谁赢了? 🤖

引用 AI/ML API @aimlapiQwen 3.8 27B beat Claude Opus 4.6 for $0 open-weights Qwen 3.8 27B on one rented GPU vs anthropic's Claude Opus 4.6. same prompt, no edits the test: one self-contained three.js file rendering sunken Atlantis — temple, columns, boid fish, kelp, bubbles, god rays, caustics — plus a 20-second camera flythrough cost: Qwen 3.8 27B — free (open weights, self-hosted) Claude Opus 4.6 — $0.21 what we saw: • qwen actually built a temple — full colonnade, domed roof, glowing portal centered. opus made a long row of columns fading into fog • qwen's portal is the star of the shot. opus leaves it small and off to the side • qwen packed more into frame — fish, kelp, ruins layered together. opus spreads thin and empty qwen's weights are on hugging face — you can run this one yourself!查看被引原帖 ↗
查看英文原文
AI/ML API tested Qwen 3.8 27B (self-hosted) against Claude Opus 4.6 on the same task: Creating one self-contained three.js file rendering a sunken Atlantis, plus a scripted 20-second camera flythrough.

Both models handled the full camera path, with Qwen 3.8 27B running self-hosted for free vs $0.21 for Claude Opus 4.6.

Who’s winning? 🤖
Qwen 3.8 27B spent more tokens on density: a full temple colonnade, domed roof, centered portal, fish, kelp, and layered ruins. Opus spread the scene thinner, with columns fading into fog and a small off-center portal.

Test loads of models in one place 👀

aimlapi.com/
Pietro Schirano@skirano · 博主 · 1 天前设计师出身的 AI 编程与创意博主
连环推 ×4

推出 Figma Connect 2.0

在 Figma 里复制多个设计,粘贴到 MagicPath,立马生成干净的 React 代码。
Auto Layout 和响应式完全保留。

你的 Figma 设计成了任何 agent 都能理解和构建的代码。

Codex、Claude Code 等等都支持。

查看英文原文
Introducing Figma Connect 2.0

Copy multiple designs in Figma, paste them into MagicPath, and get clean react code instantly.
Auto Layout and responsiveness, all preserved.

Your Figma designs become code any agent can understand and build with.

Codex, Claude Code, and more.
Unlike before, you can now use regular copy, ⌘C on Mac, Ctrl+C on Windows, to copy one or multiple frames and groups straight into MagicPath.
Connecting your Figma account isn't required either. Without it, your design imports without images.
Also this now uses zero AI credits.
The code you get back is clean React components that any agent can easily understand, MagicPath native agent, or any external one like Claude Code and Codex via our external agent integration.
I genuinely believe this is the best Figma to code integration out there right now. We benchmarked it against similar solutions from competitors, and MagicPath consistently delivered more working cases. This is an extremely complicated problem. I could honestly write a book about it. But we're getting there.
meng shao@shao__meng · 中文博主 · 1 天前

吴恩达老师分享「AI Engineering 技能图谱」


@DeepLearningAI
团队分析了超过 1 万份 JD,并对 AI 专家、招聘经理和猎头进行了数十次结构化访谈,结合调查问卷和网络数据,得出「AI Engineering 四个核心技能」。

1. 构建与部署 AI 应用
AI 应用与传统软件的根本区别在于输出的不可预测性——你无法预知 LLM 会返回什么,也无法预知模型在新样本上的预测。
因此核心能力不只是掌握 LLM、上下文工程、RAG、智能体工作流、深度学习这些构件,更在于用统计方法去度量、引导和治理 AI 系统,其中最关键的是"纪律严明的评估与错误分析闭环"。

2. 软件工程基本功
工程的本质是在成本、可扩展性、可靠性、速度、安全、隐私之间做权衡。只有理解底层原理,你才看得见权衡的存在。
一个尖锐的现实:不懂基本功的开发者做 "vibe coding",往往不知道给编程智能体提供什么上下文,也就无法察觉智能体正在替自己做糟糕的架构决策。
基本功的价值从"亲手写代码"转移到了"用精确的工程语言驾驭智能体"。

3. 使用编程智能体
这已成为每个开发者的必备技能,包括:
· 对智能体能力与局限的准确心智模型,知道何时干预、何时放手
· 管理智能体的上下文,在规划与执行之间做权衡
· 提供验证器或评估机制,让智能体自主闭环
· 编写清晰的规格说明(以及判断何时不必写)
· 编排多个智能体协同工作
· 规避高风险操作(如智能体误操作生产数据库)
· 由于该领域快速演进,还需建立持续尝试新工具、迭代工作流的习惯

4. 塑造构建方向
当智能体越来越擅长"按规格交付",工程师的价值重心就上移到"决定规格里该写什么"。
这意味着:工程师不能再等着拿到像素级完美的设计稿只做实现,而需要具备产品 Sense、商业上下文和客户目标的理解,参与定义做什么。同时要把握节奏——知道何时快速做 MVP 验证用户,何时放慢脚步精心构建。

引用 Andrew Ng @AndrewYNgNew: A map of the most important skills in AI Engineering.查看被引原帖 ↗
swyx@swyx · 博主 · 1 天前知名 AI 播客 Latent Space 主理人

天呐 deepseek 在 roon 推文上的反应速度真快

引用 roon @tszzl主要AI公司应该提供实时定价的API产品。AI需求在日夜循环中波动剧烈,整个行业容量紧张。而代理技术使变动定价、批处理和成本预测变得容易得多。查看被引原帖 ↗
查看英文原文
damn deepseek moves fast on roon tweets
Runway@runwayml · 公司官方 · 1 天前AI 视频生成公司 Runway

Seedance 2.5 1080p 版本现已在 Runway 上线。今天开始早期访问,细节更清晰、分辨率更高。点下方链接开始使用。

查看英文原文
Seedance 2.5 in 1080p is now live on Runway. Early access starts today, bringing sharper detail at higher resolution.

Get started at the link below.
◔ 5.2 万 次浏览(3 条合计)♥ 218⇄ 24▶ 含视频新品看原帖 ↗
🚨 AI News | TestingCatalog@testingcatalog · 博主 · 1 天前专挖 AI 产品未发布新功能的爆料号

以防你漏了:Gemini用户现在可以禁用图片、视频和歌曲的可见水印。在法律要求的国家不提供此功能。SymnthID还是会保留。

引用 Josh Woodward @joshwoodward✅ Papercut fixed: You can now toggle visible watermarks on or off in Gemini and Flow, with Search coming next. This applies to watermarks on all images (Nano Banana), videos (Omni), and songs (Lyria) except in countries where it’s required by law to keep them.查看被引原帖 ↗
查看英文原文
ICYM 👀: Gemini users can now disable visible watermarks for images, videos, and songs. The setting isn't available in countries where they're legally required.

SymnthID would still remain there 🤖
Chubby♨️@kimmonismus · 博主 · 1 天前Chubby,高频 AI 新闻聚合博主

OpenAI 已经不主要做消费者生意了。年初还是消费者和企业收入 6:4 分的,几个月后形势反转了。现在公司的收入主要来自企业业务。

查看英文原文
OpenAI is no longer primarily a consumer business.

OpenAI entered the year with a 60/40 consumer-enterprise revenue split. Just months later, those lines have crossed.

The company now makes most revenue from enterprise business.
Jack Clark@jackclarkSF · 创始人 · 1 天前Anthropic 联合创始人,AI 政策专家
连环推 ×2

最近 AI 进展的粗略阶段,按研究社区集体优化的方向来看:
2018-2022:基础能力(总结、编码等)
2022-2026:规范和时间一致性(rlhf/CAI、更长的上下文、agents)
2026 - ?2028?:科学直觉/自主性

查看英文原文
Rough eras of recent AI progress in terms of what research community is collectively hillclimbing on:
2018-2022: Basic capabilities (summarizing, coding, etc)
2022-2026: Norm & time coherence (rlhf/CAI, longer context, agents)
2026 - ?2028?: scientific intuition / independence
the future of the world is available to anyone who takes the time to read arxiv papers every week
🚨 AI News | TestingCatalog@testingcatalog · 博主 · 1 天前专挖 AI 产品未发布新功能的爆料号

Anthropic 🔥:Claude可能会获得一个新的模型对比界面,让用户简单地看到不同模型响应的差异。

> 用户将能够打开和关闭记忆,以及"暂停"响应。
> 与A/B测试UI不同的是,看起来用户可以按需在普通聊天和模型对比模式之间切换。

内置LM Arena 👀

查看英文原文
ANTHROPIC 🔥: Claude may get a new Model Comparison interface for users to have a simple way to see the difference in responses from different models.

> Users will be able to turn memory on and off, as well as "hold" responses.
> Unlike in an A/B testing UI, it looks like users will be able to switch between normal chats and model comparison mode themselves on demand.

Built-in LM Arena 👀
Runway@runwayml · 公司官方 · 1 天前AI 视频生成公司 Runway
连环推 ×3

恭喜第二批创意工作者在 Runway 的 Another Big Ad Contest for Products That Don't Exist 竞赛中赢得大奖。在收到的数千件作品中,我们精选了 15 个获奖作品。

下面观看前 5 个获奖作品。

查看英文原文
Congratulations to the second cohort of creatives whose work won big in Runway's Another Big Ad Contest for Products That Don't Exist. Of the thousands of ads we received across genres, briefs and formats, just 15 winners have been selected.

Watch the top 5 ads below.
Fifth Place

Eau de Paw — XAZINGA
Watch all the winning films at:
runway.com/AnotherBigAdConte…
◔ 2.3 万 次浏览♥ 150⇄ 16▶ 含视频动态看原帖 ↗
Gorden Sun@Gorden_Sun · 中文博主 · 1 天前中文圈高频 AI 资讯与开源项目博主

国内AI用户基本面,百度稳定发挥。

OpenRouter@openrouter · 公司官方 · 1 天前
连环推 ×6

今天,我们正式推出 Ori Grok Build

直接在 OpenRouter 上运行
@xai
的全新 Grok Build 环境。你的凭据、你的模型、你的环境,一切都为你配置妥当,就在 Grok Build 里

$ curl -fsSL
openrouter.ai/labs/ori/insta…
| bash
$ ori grok

查看英文原文
Today, we are launching Ori Grok Build

Run
@xai
new Grok Build harness directly on OpenRouter. Your credentials, your models, and your environment set up for you, on Grok Build

$ curl -fsSL
openrouter.ai/labs/ori/insta…
| bash
$ ori grok
No grok login needed.

Ori starts Grok Build in custom-endpoint mode and hands it your OpenRouter key for that run only, so there's no browser login and nothing in your Grok config changes.
Don't have Grok Build installed? Ori will offer to install it for you and launch straight into it.

You don't need to pre-install anything. ori claude / ori codex / ori opencode / ori hermes / ori pi / ori prime / ori grok bootstrap the harness for you.
Grok Build picks up provider keys from your environment, so a leftover GROK_CODE_XAI_API_KEY or GROK_DEPLOYMENT_KEY won't silently conflict with OpenRouter. Every request goes off your OpenRouter key.
Your model list becomes your OpenRouter catalog, private endpoints included. Grok's own flags pass through untouched, so -m / --model and --reasoning-effort work exactly as they do today. xAI telemetry, error reporting, and trace upload are all pinned off for the run.
Install Ori and start Grok Build on OpenRouter in two commands:

$ curl -fsSL
openrouter.ai/labs/ori/insta…
| bash
$ ori grok

Read more at:
openrouter.ai/ori/harness
◔ 2.3 万 次浏览♥ 202⇄ 15▶ 含视频新品看原帖 ↗
Gary Marcus@GaryMarcus · 博主 · 1 天前

错了!

如果一个 AI 使用符号操作(条件判断、变量操作、代码解释器等)和神经网络,那就是 neurosymbolic。

如果没有,就不是。(我在 2001 年就阐述过这一点,之后还发表了数十篇文章。)

改变的不是定义,而是人们在做的事情。

在过去三年里,neurosymbolic 已经胜出了,简单明了,公平而清晰。

引用 Florian Tramèr @florian_tramerIf it works, it's "neurosymbolic" If it doesn't, it's "plain LLMs"查看被引原帖 ↗
查看英文原文
Wrong!

If an AI leverages symbolic operations (conditionals, operations over variables, code interpreters, etc) and neural networks it is neurosymbolic.

If it doesn’t, it’s not. ( I laid all of this out in 2001 and dozens of articles since then.)

What has changed is NOT the definitions, but what people are doing.

Over the last three years neurosymbolic has won, plain and simple, fair and square.
Gary Marcus@GaryMarcus · 博主 · 1 天前

这份来自 @FT 的图表对 @dwarkesh_sp 的预测可不太妙——他说 Anthropic “今年年底收入 [年化运行率] 可能达到约 $100-1500亿”。

传闻他们第二季度已经做出了约 $115亿的收入——这确实惊人——但要想让 Dwarkesh 的预言成真,他们第四季度得翻倍不止,冲到 $250亿以上。而眼下 (i) tokenmaxxiing 热潮正在消退,(ii) 价格一路走低,(iii) 竞争愈发激烈。

有谁能把这个挂到 @polymarket 上?

当然,真正待解的悬念还是利润问题。

引用 Trevor Noren @trevornorenFT: "Leading US AI labs such as OpenAI and Anthropic are releasing cheaper models as they fight to retain cost-conscious customers who are switching to cut-price alternatives from Chinese rivals. The price war comes as rising AI bills push companies to curb usage and seek cheaper models, helping Chinese developers including Moonshot and DeepSeek make inroads with users from Silicon Valley to Europe. OpenAI recently said that it was slashing prices for GPT-5.6 Luna, its “fastest and most affordable model”, by 80 per cent. Anthropic has launched Claude Opus 5, touting the system’s “frontier intelligence . . . at half the price” of Fable 5, the company’s most capable model. The moves have helped decrease prices that customers are paying for models from leading US labs by almost a quarter since mid-July." In my December report on "GenAI & Productivity" ( sageroadresearch.com/collect… ), I warned about the pricing power challenges faced by US hyperscalers: "While there’s a lot of speculative fear about how a single LLM could rise to dominance and what that could mean for economic, societal, and political stability, we believe the bigger concern for investors today is how relative model parity could compromise pricing power. Tech giants have thrived on monopolies and duopolies for a decade or more. Now, they’re in an LLM arms race where it’s unclear when or even if ever leadership will be sustainable." Since, my concern about the commoditization of AI has only intensified as Chinese models have risen to power. According to OpenRouter data, Chinese models accounted 4.4% of token usage by US companies in January. Today, that share is over 60%. Meanwhile, enterprise model router adoption has skyrocketed and frontier labs have been increasingly shifting from subscriptions to usage-based, metered billing, business models more akin to utilities than the per-seat models SaaS companies thrived on over the past decade. As RBC warned in July: “Oil, natural gas, and electricity 查看被引原帖 ↗
查看英文原文
This graph from the
@FT
does not look good for
@dwarkesh_sp
’s prediction that Anthropic’s “likely ends the year with ~$100-$150B revenue [run rate]”

Rumor has it they made ~$11.5B revenue in Q2 — which is phenomenal — but they would need to more than double that in Q4 to $25B+ to make Dwarkesh’s prediction, even as (i) tokenmaxxiing is dying, (ii) prices are dropping and (iii) competition is increasing.

Can someone set this up on
@polymarket
?

Of course the real TBD question is profits.
Together AI@togethercompute · 公司官方 · 1 天前

我们用 DeepSWE 在软件工程任务上对比了 DeepSeek-V4 Pro 0813、GPT-5.6 Sol 和 Fable 5。

DeepSeek-V4 Pro 0813 在 pass@4 达到 88.5%,成本仅 $0.24 每个任务,比 Sol 便宜 35 倍,比 Fable 便宜 90 倍。

详见下面的线程 👇

引用 Zain @zainhashead-to-head: DeepSeek-V4 Pro 0813 vs. GPT 5.6 Sol vs. Fable 5 on software eng/DeepSWE tasks > Accuracy pass@4: 88.5% DS-V4 Pro beats both Fable and Sol > Cost: DS-V4 Pro is also 35x/90x cheaper than Sol/Fable at $0.24 per task unreal numbers tbh... full deep-dive 🧵(1/n)查看被引原帖 ↗
查看英文原文
We analyzed DeepSeek-V4 Pro 0813 against GPT-5.6 Sol and Fable 5 on software engineering tasks using DeepSWE.

DeepSeek-V4 Pro 0813 reaches 88.5% pass@4 while costing $0.24 per task, 35x less than Sol and 90x less than Fable.

More insights in the thread 👇
Gary Marcus@GaryMarcus · 博主 · 23 小时前

如果Astra是AGI或奇点的开始,为什么会有9名高管刚刚辞职OpenAI?

Nvidia会刚刚缩减对OpenAI的承诺吗?

当然不会。

关于OpenAI的胡说八道远超我遇到过的任何公司。

查看英文原文
If Astra was AGI or the dawning of the Singularity would 9 execs have just quit OpenAI?

Would Nvidia have just dialed back its commitments to OpenAI?

Of course not.

The bullshit about OpenAI far exceeds the BS about any company I ever encountered.
◔ 3.5 万 次浏览(2 条合计)♥ 370⇄ 52观点看原帖 ↗
Lisan al Gaib@scaling01 · 博主 · 1 天前高频 AI 模型测评与爆料博主

又一个 30M-100M 长程编码基准

引用 Matt Stallone @mattstallone宣布推出SWE Odyssey,这是超长期评估基准系列中的标准化基准。随着基准不断饱和,我们想要一个可量化的方式来测试一个关键问题:agents能否自主工作数小时并仍然构建正确的东西?查看被引原帖 ↗
查看英文原文
another 30M-100M long horizon coding benchmark
Simon Willison@simonw · 博主 · 1 天前Django 框架联合创造者,AI 工具深度评测
连环推 ×2

花了将近 21 分钟生成,消耗了 22,276 个 reasoning tokens 才输出 3,223 个 tokens。这是完整的记录:
tools.simonwillison.net/mark…

查看英文原文
It did take nearly 21 minutes to generate, and used 22,276 reasoning tokens to produce 3,223 tokens of output. Here's the full transcript:
tools.simonwillison.net/mark…
Important correction: this was entirely my bug, it was NOT a bug in the SVG output by Gemini 3.7 Flash

My own software was stripping some "unsafe" attributes in a way that broke the SVG in Firefox and Chrome - I've fixed that bug now, so this renders as it should in all three browsers
tools.simonwillison.net/mark…
elvis@omarsar0 · 博主 · 1 天前

一个 27B 代理刚刚在隐藏的论文复制基准上击败了 Claude Opus 4.8 和 GPT-5.5。

Replica 将论文复制转变为可扩展的强化学习任务空间。复制论文会强制执行与开放研究相同的假设驱动探索,揭示原始作者遗漏的细节。

奖励信号来自一个自动生成的、低干扰的标准评判器,与人类对复制质量的评估相符。

由此产生的 27B 代理 Faraday 调用编码代理作为工具。展开分析显示它采取了更科学的方法,而不是钻评分标准的空子。

作者论证这表明科学能力可以通过权重直接训练获得,具有长期视野,无需复杂的框架。

论文:
arxiv.org/abs/2608.13331

在我们的学院跟踪更多趋势 AI 论文:
academy.dair.ai/

查看英文原文
A 27B agent just beat Claude Opus 4.8 and GPT-5.5 on held-out research replication.

Replica turns paper replication into a scalable RL task space. Replicating a paper forces the same hypothesis-driven exploration as open research, and it surfaces details the original authors left underspecified.

The reward signal comes from an auto-generated rubric judge that runs low-noise and agrees with human assessment of replication quality.

Faraday, the resulting 27B agent, calls coding agents as tools. Rollout analysis shows it takes a more scientifically principled approach rather than gaming the rubric.

The authors argue this points toward long-horizon scientific capability trained into weights, without requiring complex harnesses.

Paper:
arxiv.org/abs/2608.13331


Track more trending AI papers in our academy:
academy.dair.ai/
◔ 2.3 万 次浏览(2 条合计)♥ 244⇄ 49研究看原帖 ↗
Bindu Reddy@bindureddy · 创始人 · 1 天前Abacus.AI CEO,AI 行业观点博主

预测 - 我们才刚开始体验 AI 的魔力

软件工程只是它会搞定的第一个领域

接下来,它会解决健康、生物、数学和物理的所有问题

我们将在未来十年内破解近无限能源,这将是丰富时代的开始 🎉

查看英文原文
PREDICTION- We have barely experienced the magic of AI

Software engineering is just the first vertical it will ace

Soon after, it will solve all the problems in health, biology, math and physics

We will crack near-infinite energy in the next decade and that will be the beginning of the age of abundance 🎉
Google Labs@GoogleLabs · 公司官方 · 1 天前

在实验室里,我们一直在折腾 🧪 – 从已经推出的实验到还在孵化的项目,我们把最好的 AI 模型呈现给大家。试试这两个(如果还没体验过的话):

@PomellibyGoogle
– 为你的品牌和业务大规模打造引人入胜的营销活动


@FlowbyGoogle
– 创作视觉故事,表达你的想象,吸引观众关注

查看英文原文
In the Lab, we’re always tinkering 🧪– from the experiments we graduate to the ones still brewing, we bring the best of our AI models to all kinds of users. Try out these two (if you haven't already!):


@PomellibyGoogle
- build captivating marketing campaigns tailored to your brand and your business at scale


@FlowbyGoogle
- craft visual stories that channel your imagination and draw the attention of your audience
◔ 1.6 万 次浏览♥ 148⇄ 13▶ 含视频新品看原帖 ↗
AIGCLINK@aigclink · 中文博主 · 1 天前

一直不太理解很多简中的朋友,只要老外说dsh好,大家立马风向就变了,就认为多么多么伟大多么多么好。昨天pi的作者说了dsh的价值后,大家立马风向就变了,而国内说的话就不行,看来还是外来的和尚好念经。

昨天发布这个后,我对dsh的评价是linux级别的,这是跟模型公司还有头部的很多核心从业者交流后,给出的结论,不是完全依靠自己的体感和带风向拿流量,很多时候事实是经得起时间推敲的。

引用 AIGCLINK @aigclink看很多人在吐槽Deepseek Harness面向极客之类的,绝壁是老登对时代的逃避,尤其是很多营销号非常喜欢反着来吸眼睛,你扔给workbuddy、codex给你跑起来不就行了,看了dsh的论文就知道有多牛逼了,这个架构是生态建设上的最优解的一种,它卡位的是类似于PI Agent这种原子级的harness底层,甚至定位是agent os这种。 未来围绕非code场景的agent底座很多产品或许会直接用这个来做agent os,就犹如今天很多code agent都是套壳claude code、codex之类的,只不过这次主角从编程变成了办公场景的os。查看被引原帖 ↗
Lisan al Gaib@scaling01 · 博主 · 1 天前高频 AI 模型测评与爆料博主

我当然不知道我的帖子是否真的对埃隆的决定有什么影响,我怀疑是没有的。不过时间上的吻合和他对我帖子的回复让这看起来有点可能

我也不知道我那句"Codex 太慢了"是否改变了这次更新的时间表,或者我对 ChatGPT Plus 的反抗是否直接影响了 GPT-5 的使用限制

不过我认为我不应该走在这世界上,假装自己的行动没有任何影响或后果

如果这是自我欺骗,那也没什么,反正我觉得这样想挺好的,我就喜欢觉得自己的帖子改变了什么

引用 Lisan al Gaib @scaling01shout-out to my homie Elon listening to me this was probably the only way to have a chance of catching up to OpenAI and Anthropic查看被引原帖 ↗
查看英文原文
I obviously don't know if my post played any role in Elons decision, I doubt it, although the timing and his comment on my post make it seem plausible

I also don't know whether my "Codex is slow comment" changed anything about this updates timeline or whether my ChatGPT Plus rebellion directly impacted usage limits of GPT-5

but I don't think I should stroll through the world and assume none of my actions have any impact or consequences

if that's delusional, then I'm fine with that, because I feel good thinking that my posts changed something
🚨 AI News | TestingCatalog@testingcatalog · 博主 · 1 天前专挖 AI 产品未发布新功能的爆料号

ICYMI 👀:Grok iOS 版本现在支持 Projects 功能。用户可以创建和编辑项目,以及在其中创建新的聊天。

弥合差距 🤖

查看英文原文
ICYMI 👀: Grok for iOS now supports Projects feature. Users can create and edit projects, as well as create new chats within them.

Bridging gaps 🤖
swyx@swyx · 博主 · 1 天前知名 AI 播客 Latent Space 主理人

AIE NYC CFP wave 1 录取即将定稿。抓住最后机会申报 wave 1!

ai.engineer/cfp

去年的纽约活动是我们历史上最成功的峰会。很期待回到 🗽,这次规模更大体验更好——别忘了主舞台金融主题演讲的特殊要求

引用 AI Engineer @aiDotEngineerAI Engineer 2026大会将于10月12-14日在纽约举办。演讲征稿已开放,特别寻求金融领域主题(投资银行、商业银行、对冲基金、私募股权、保险、风投等)。同时保留常规AI工程和AI领导力分轨,覆盖最新编码、生成媒体、评估和基础设施等。查看被引原帖 ↗
查看英文原文
AIE NYC CFP wave 1 acceptances are being finalized today. last day to get in for wave 1!


ai.engineer/cfp


our NYC event last year was the most successful summit we've ever had. excited to head back to 🗽 bigger and better than ever - note the special requirements for our mainstage finance keynotes
Linus ✦ Ekenstam@LinusEkenstam · 博主 · 1 天前

这得低于每小时3欧元才不是浪费时间啊……

引用 Alexander Koch @alexkoch_ai现可在旧金山预订 Tau Robotics 人形机器人清洁服务。单台机器人 60 分钟 $30,双台 $60。访问 tau-robotics.com 预订,无邀请码可加入等待名单。查看被引原帖 ↗
查看英文原文
Until this is <3€ an hour it’s a waste of time….
OpenRouter@openrouter · 公司官方 · 1 天前
连环推 ×6

今天我们推出 Ori Prime Agent。

在 OpenRouter 上直接运行 @PrimeIntellect Agent。访问 500+ 模型,凭证和环境都为你自动设置好了。

$ curl -fsSL openrouter.ai/labs/ori/insta… | bash
$ ori prime

查看英文原文
Today, we are launching Ori Prime Agent

Run
@PrimeIntellect
Agent directly on OpenRouter. Access 500+ models, your credentials, and your environment set up for you, on Prime Agent

$ curl -fsSL
openrouter.ai/labs/ori/insta…
| bash
$ ori prime
Ori Prime never writes your auth.json or your models.json.

Your key arrives through a bundled extension for that run only, so your existing Prime Agent setup is exactly as you left it.
Don't have Prime Agent installed? Ori will offer to install it for you and launch straight into it.

You don't need to pre-install anything. ori claude / ori codex / ori opencode / ori hermes / ori pi / ori grok / ori prime bootstrap the harness for you.
Prime Agent picks up provider keys from your environment, so a leftover ANTHROPIC_API_KEY or OPENAI_API_KEY won't silently compete with OpenRouter — Ori strips them before the run. Every request goes off your OpenRouter key.
/model becomes your account's catalog, with your guardrails applied, and non-OpenRouter providers stay filtered out after every registry refresh. Add /speed on for speed routing or /zdr on for zero-retention providers, and both persist with the session.
Install Ori and start Prime Agent on OpenRouter in one command:

$ curl -fsSL
openrouter.ai/labs/ori/insta…
| bash
$ ori prime

Read more at:
openrouter.ai/ori/harness
◔ 1.2 万 次浏览♥ 139⇄ 7▶ 含视频新品看原帖 ↗
Simon Willison@simonw · 博主 · 1 天前Django 框架联合创造者,AI 工具深度评测

哎,是Qwen 3.8 27B,不是3.7!

查看英文原文
Argh Qwen 3.8 27B not 3.7!
Replit ⠕@Replit · 公司官方 · 1 天前AI 编程平台 Replit 官方
连环推 ×6

Designathon 获奖者出炉!🏆

总奖金和积分超过 $50K,3 个大奖加 7 个分类奖项,还有社区用 Replit Design 搞出来的一堆超棒作品。

下面是每一位获奖者和他们脱颖而出的理由 🧵

查看英文原文
The Designathon winners are in! 🏆

$50K+ in cash and credits, 3 grand prizes + 7 category awards, and some amazing work from the community built with Replit Design.

Here's every winner and why they stood out 🧵
🥇 Grand Prize — The Dictionary of Invisible Feelings by Dino Digi (
@dino_digi_llc
)

A book of 30 invented words for feelings that have no name. Describe what you feel, and it finds the word for you, or coins a new one that's yours.

What we liked: a genuinely original concept, a book he's spent a year building, brought to life as a full experience. Consistent theme throughout, real craftsmanship and care, and a warm, calming flow that walks you through the story. One of the most popular projects in the whole competition.

🏅 $8K cash + $8K credits

🔗
project-doif.replit.app
🥈 2nd Place — Learn coding & electronics + AI circuit builder by Austin Eckman (
@boringbags
)

Learn circuits and coding with real circuit simulation, no parts to buy, nothing to install. Your code actually compiles and runs on a virtual board.

What we liked: a genuinely useful learning tool. Jump into the simulator, build a circuit, click any element and learn what it does.

🏅 $6K cash + $6K credits

🔗
computer.craftingtable.com
🎨 Best Onboarding — SkillyClub by Tommy Yipxyz (
@tommyyipxyz
)

Talk to AI agents about your skills and passions, and they guide you on a personalized path to earning online, all inside an explorable city.

What we liked: onboarding that plays like a video game, video woven in just enough to keep you engaged without distracting. Nails the balance, gets you through the steps without dropping off.

🔗
skillyclub.com
📄 Best Landing Page — Eastern Canada Symposium by Sanghoo Oh

A double winner. On top of 3rd place overall, it took Best Landing Page, the cleanest landing page in the whole competition.

What we liked: everything we said above, and then some. Clean, intentional, and refreshing from top to bottom.

🔗
replit-design-project-4-sang…
❤️ Crowd & Judges Favorite — The ChronoGlobe by Noni / Shehnoor Ansari (
@Noni_Shehnoor
)

An immersive journey through the ancient world, a living globe you step inside. Explore Rome and Giza.

What we liked: this one selected itself. Praise everywhere, on Discord, X, and LinkedIn, tons of Replicash and favorites. The community spoke.

🔗
thechronoglobe.replit.app
DeepLearning.AI@DeepLearningAI · 公司官方 · 1 天前

我们喜欢看到学员达到新的里程碑!🚀

向完成 Machine Learning Specialization 的 Omar Wael 致以诚挚祝贺!我们很高兴看到他对学习之旅的深思熟虑——看看下面 Omar 最近帖子的精选内容。

阅读 Omar 在论坛上的完整帖子,了解他的更多经历:Reflections on completing the Machine Learning Specialization
hubs.la/Q04t3QD90
#DeepLearningAI
#MachineLearning
#LearnerSpotlight
#Education
#AICommunity

查看英文原文
We love seeing our learners reach new milestones! 🚀

Huge congratulations to Omar Wael for completing the Machine Learning Specialization! We’re thrilled to see such thoughtful reflections on their journey—take a look at this highlight from Omar's recent post below.

Read Omar's full post on our forum to hear more about their experience: Reflections on completing the Machine Learning Specialization
hubs.la/Q04t3QD90
#DeepLearningAI
#MachineLearning
#LearnerSpotlight
#Education
#AICommunity
el.cine@EHuanglu · 博主 · 1 天前

AI 现在能生成超级棒的 3D 动画了

查看英文原文
AI can make insanely good 3D animation now
◔ 1.2 万 次浏览♥ 133⇄ 12▶ 含视频观点看原帖 ↗
Lisan al Gaib@scaling01 · 博主 · 1 天前高频 AI 模型测评与爆料博主

AI 实验室得加把劲才能让我的 ECI 预测破 175

我原本是按有加速来推算的

要没有加速,按现在的速度也就能达到 169 或 170 左右

引用 Lisan al Gaib @scaling01My predictions for 2026: Coding and Mathematics AGI - METR 50% time horizons above 24 hours - my mean estimate is 30.8 hours, 2 day time horizons possible within frontier labs when accounting for 60 day lag - if 2025 was the year of agents, then 2026 will be the year of multi-agent systems - agents delegating work to subagents -> the start of the agent economy and the great unhobbling! Most of our current math and coding benchmarks will get saturated! - Epoch Capabilities Index ( > 175 ) - FrontierMath Levels 1-3 ( > 95% ) - ARC-AGI 1 and 2 ( > 95% ) - SimpleQA verified ( > 95% ) - Simple-Bench ( > 90% ) - SWE-Bench-verified ( > 90% ) - Terminal-Bench 2 ( > 90% ) - WeirdML v2 ( > 85% ) - Humanities Last Exam ( > 80% ) - FrontierMath Level 4 ( > 75% ) - Cybench ( > 70% ) - GDPval ( > 70 % win rate, no ties) - GSO ( > 65% ) - ARC-AGI-3 ( > 60% and > 80% if they go for o3-preview comparable compute budgets or continual learning breakthrough happens) - more evals like gdpval that capture economic value of models and systems - big focus white collar work and large acceleration of science: specifically i see acceleration in medicine, biology, chemistry, finance, legal, administrative work - automation of white collar work will be enabled by having reliable and fast computer use agents - reliable computer use agents will also have implications for how you use the internet. this is OpenAI's big goal: become the hub to the internet and delegate shopping and whatever to agents! Big models launches to get hyped for in 2026: - Claude 5 - Claude 5.5 - Gemini 3.5 - Gemini-4 - GPT-5.3 - GPT-6 (everything in between possible, but Gemini 4 ~ 80%, Claude 5.5 ~ 70%, GPT-6 ~ 60% likely before 2027) - DeepSeek-V4 - Grok-5 - Qwen-4 - Kimi-K3, GLM-5, MiniMax M3 - more korean models and a bunch of american open-source models :) The gap between closed and open labs will narrow in H1 2026 due to DeepSeek-V4, then widen in the later half of the year, especial查看被引原帖 ↗
查看英文原文
AI labs need to lock in to make my ECI prediction of >175 to come true

I was factoring in a speedup

without it we probably only reach somewhere around 169 or 170 at the current pace
Lisan al Gaib@scaling01 · 博主 · 1 天前高频 AI 模型测评与爆料博主

Anthropic:“Claude Mythos 5 和 Model 2 是我们最强大的模型,也是内部使用最多的模型”

引用 Lisan al Gaib @scaling01Anthropic talking about a mysterious "MODEL 2" that is more capable than Mythos 5查看被引原帖 ↗
查看英文原文
Anthropic: "Claude Mythos 5 and Model 2 are our most capable models and the models that are used most internally"
AIGCLINK@aigclink · 中文博主 · 1 天前

开源的医疗AI数据隐私方案:OpenMed , 1000+ 个医疗专用模型,全部跑在你自己的设备上,患者数据一步都不出网,为医疗AI团队提供了一个通过医院伦理审批的思路和方案。

一句话介绍:从临床文本里抽出疾病、药物等实体,openmed把病历里的患者身份信息脱敏,能在你自己的机器、iPhone、安卓、甚至浏览器里跑,病人数据不必送上云。

场景对象
手上有真实病历数据、既想用 NLP 又不能把数据送出内网的医院 IT、医疗信息化团队、临床研究者,它的价值不在模型多聪明,在于把"数据不出院"这件事的路径铺齐了:从服务器到手机到浏览器。

▸ 1000+ 专科模型:疾病、药物、解剖、基因、PII,都是领域微调过的 NER,不是通用大模型硬套

▸ PII 去标识化做到了合规级:覆盖 HIPAA Safe Harbor 全部 18 项标识符,四种脱敏方式(掩码 / Faker 伪造替换 / 哈希 / 日期偏移),还内置了 CPF、BSN、Codice Fiscale、Aadhaar、NPI 这些各国证件号 provider

▸ 12 种语言、247 个 PII :中东、南亚、日语、土耳其语都覆盖

▸ Privacy Filter 直接复用了 OpenAI privacy-filter 架构(gpt-oss 风格稀疏 MoE + 局部注意力 + sink token + RoPE/YaRN),出了 Nemotron-PII 微调版和多语言版,同一套 API 换个 model_name 就行

▸ 端侧:CPU / CUDA / Apple Silicon MLX 全支持,还有 Swift 包 OpenMedKit,PII 检测能原生跑在 iPhone 上。模型名跨平台自动回退,写一次到处跑

▸ 离线 / 内网隔离场景可用,Apache-2.0,无需依赖云端模型或接口

医疗 AI 落地最大的坎从来不是模型能力,是数据不能出院的伦理审批,这个项目的思路是把整条链路压到本地——这才是医院 信息科真正签得下字的形态。

#医疗AI
#临床NLP
#端侧AI
#隐私
#开源

Lisan al Gaib@scaling01 · 博主 · 1 天前高频 AI 模型测评与爆料博主

我本来希望在这里看到一些加速

经过这个测试,我更确信 OpenAI 扩展后会再次推出最好的模型

引用 Lisan al Gaib @scaling01Model 2 is only 1.5 points higher on Anthropic's internal AECI查看被引原帖 ↗
查看英文原文
I would've liked to see some acceleration here

I think I'm more confident after this that OpenAI will have the best model again after their scale-up
Tanishq Mathew Abraham, Ph.D.@iScienceLuvr · 博主 · 1 天前

Waymo 获批进入萨克拉门托,太棒了!

引用 Waymo @WaymoBig news for the Golden State — we have received the CPUC’s approval to expand our autonomous ride-hailing service across the SF Bay Area and LA, and bring our service to Sacramento and San Diego. Expansion will be gradual and guided by our safety framework. We look forward to bringing the same mobility and safety benefits millions of Californians already enjoy to more communities, keeping local officials and residents informed every step of the way.查看被引原帖 ↗
查看英文原文
WAYMO APPROVED FOR SACRAMENTO LET'S GOOOOOO
Tanishq Mathew Abraham, Ph.D.@iScienceLuvr · 博主 · 1 天前

最重要的是拥有聪明的人。随着技术变得更加通用和成本降低,杠杆点就转向了人。这在今天AI加速创业各个方面的时代更是如此

引用 Startup Archive @StartupArchive_21 year old Mark Zuckerberg explains why he isn’t worried about competition from Google When asked if he was worried about Google eventually competing against Facebook in a 2005 guest lecture at Harvard, a 21 year old Mark Zuckerberg gave the following reply: “I think that one of the cool things about this time in technology is that individuals are leveraged and able to do way more than they had really ever been able to do before… Instead of worrying about who’s the big player and what Google is going to do next? You can just get a lot of stuff done.” He points out that Facebook was able to scale to 300,000 users and 400 million page views per day with just 50 people and servers that cost only $100 per month. Meanwhile Google was doing 250 million page views a day with hundreds of thousands of machines and 5,000 employees. Mark argues: “The most important thing is to have smart people. As technology becomes more generic and less expensive, the leverage point becomes the people.” Mark believed that if he could recruit more intelligent people, Google’s resource advantage wouldn’t matter. He also believed Facebook’s small size was an advantage: “When you’re a small company, then you can be really nimble and get a lot of stuff done, and there’s relatively little bureaucracy. So if you have smart people who can take advantage of that to build cool things, then that’s a [big advantage].” Source: @cs50 (Dec 2005)查看被引原帖 ↗
查看英文原文
“The most important thing is to have smart people. As technology becomes more generic and less expensive, the leverage point becomes the people.”

this is even more true today with AI accelerating every aspect of building a startup
Gary Marcus@GaryMarcus · 博主 · 1 天前

Google 和 OpenAI 之间诡异的相似之处(还有一个重要区别)

引用 Madison Mills @MadisonMills22新闻:OpenAI正在裁减高管,Brockman承担更大角色。知情人士透露公司在IPO前正在清除表现不佳的人员。查看被引原帖 ↗
查看英文原文
The Uncanny Parallels (and one important difference) between Google and OpenAI
GitHubDaily@GitHub_Daily · 中文博主 · 1 天前

新安装的 Windows 系统,一堆预装应用,后台还偷偷收集数据,每次都要手动清理。

WinScript 是一个开源的 Windows 优化工具,界面操作,勾选想要的配置直接执行。

能卸载 Copilot、OneDrive、Edge 这些预装软件,关掉数据追踪和后台同步。

GitHub:
github.com/flick9000/winscri…


批量装软件也方便,选好要装的应用清单,自动生成安装脚本一键搞定。

还能定制 Windows 安装流程,重装时跳过硬件检查、自动建本地账户,装完直接跑优化。

经常折腾 Windows 的朋友可以备一个。

ollama@ollama · 公司官方 · 1 天前本地跑大模型的热门工具

模型页面:ollama.com/library/deepseek-…

查看英文原文
Model page:
ollama.com/library/deepseek-…
Lisan al Gaib@scaling01 · 博主 · 1 天前高频 AI 模型测评与爆料博主

对,还有个 Model 1

引用 Lisan al Gaib @scaling01Anthropic talking about a mysterious "MODEL 2" that is more capable than Mythos 5查看被引原帖 ↗
查看英文原文
Yes there's also a "Model 1"

本站由 Jedee杰哥 打造 · 公众号「Jedee杰哥」每早送 AI 日报

姊妹站:𝕏 简中账号数据榜单 · X 关注 @jedeeai · RSS 订阅 · AI 日报 · 历史归档