JEDEE AI
存档 2026-08-18

8 月 18 日(北京时间)全球 AI 圈推文存档,按曝光排序,共 100 条。
← 返回最新 全部归档

全部情报 每小时更新 · 事件已合并同类项

内容 公司
Cursor@cursor_ai · 公司官方 · 1 天前最火的 AI 编程工具 Cursor
连环推 ×3

Origin 是我们的代码托管平台,现已上线。

它又快又好用,与 Cursor 深度集成。

从 GitHub 同步仓库开始吧。

查看英文原文
Origin, our code hosting platform, is now live.

It's fast, easy to use, and deeply integrated with Cursor.

Get started by syncing your repos from GitHub.
We've partnered with some of the top GitHub integrations.

Vercel, Buildkite, and Depot are already available with more coming soon.
We're rolling out the beta starting today.


cursor.com/changelog/origin-…
◔ 2610.3 万 次浏览(7 条合计)♥ 2.5 万⇄ 2,384▶ 含视频新品看原帖 ↗
Grok@grok · 公司官方 · 1 天前马斯克 xAI 旗下聊天机器人 Grok 官方

荷马有里拉琴。你有Grok Imagine。

从荷马的《奥德赛》中创作一个引人入胜的场景,展示Grok
@Imagine
的视频和语音功能能做到什么。我们将为引用此帖提交的前三名视频颁发10万美元、5万美元和2.5万美元奖金。🧵

查看英文原文
Homer had a lyre. You have Grok Imagine.

Create a compelling scene from Homer’s The Odyssey that shows what Grok
@Imagine
’s video and voice capabilities can do. We’re awarding $100K, $50K, and $25K to the top three videos submitted by quoting this post. 🧵
◔ 1259.8 万 次浏览♥ 6,512⇄ 857▶ 含视频新品看原帖 ↗
Tibo@thsottiaux · 公司官方 · 1 天前ChatGPT 产品官方账号

关于 Codex、API 或我们的模型,有什么是显而易见应该去做但还没做的事?什么是百分百能做到但我们似乎就是在遗漏的?

查看英文原文
What is an obvious thing that we should do with Codex, API or our models that we should just do but haven't yet? What is 100% within reach, but we just seem to be missing?
Tibo@thsottiaux · 公司官方 · 1 天前ChatGPT 产品官方账号

午夜后给我 Codex
谁来把这些失败的测试都赶走?
午夜后给我 Codex
黑夜里把它推送,天亮前交付

查看英文原文
Gimme, gimme, gimme Codex after midnight
Won’t somebody make these failing tests all go away?
Gimme, gimme, gimme Codex after midnight
Ship it through the darkness by the start of the day
OpenRouter@openrouter · 公司官方 · 1 天前

OpenAI的GPT-5.6 Sol现在在OpenRouter上半价优惠!

这也适用于batch API、flex和priority(fast)层级,在flex层级给你低至$1.25进价和$7.50出价。

优惠自动应用,只需使用openai/gpt-5.6-sol。

openrouter.ai/openai/gpt-5.6…

查看英文原文
OpenAI's GPT-5.6 Sol is now half off on OpenRouter!

This also applies to batch API, flex, and priority (fast) tiers, giving you as low as $1.25 in and $7.50 out on the flex tier.

Discount applies automatically, just use openai/gpt-5.6-sol.

openrouter.ai/openai/gpt-5.6…
Min Choi@minchoi · 博主 · 1 天前AI 产品演示博主,专门展示新工具玩法

完蛋了。中国新型人形机器人刚刚跳了2米多,速度达到12.658米/秒。人形机器人正在突破人类极限。

查看英文原文
We are cooked.

China's new humanoid robot just jumped ~2 meters and ran 12.658 m/s.

Humanoid robots are breaking human limits.
◔ 29.8 万 次浏览♥ 1,224⇄ 138▶ 含视频动态看原帖 ↗
Bindu Reddy@bindureddy · 创始人 · 1 天前Abacus.AI CEO,AI 行业观点博主

DeepSeek Flash 陷入无限循环,无法完成任务。我们已修复这个根本缺陷,我们的版本现在能像顶级前沿模型一样运行长循环。Smaug Flash 本周发布!

查看英文原文
DeepSeek Flash Goes Into Infinite Loops And Doesn't Get Tasks Done

We have fixed this fundamental flaw and our version now can run long running loops like top frontier models

Smaug Flash drops this week!
◔ 28.8 万 次浏览♥ 190⇄ 7▶ 含视频新品看原帖 ↗
Chubby♨️@kimmonismus · 博主 · 1 天前Chubby,高频 AI 新闻聚合博主

Anthropic 的 Mythos 2 已经完成训练,但 Anthropic 不会发布它,不过 Patel 说构建 Mythos 3 的内部循环还在继续。现在的重点转向内部改进。目前还不清楚什么时候会有新的发布。

查看英文原文
Anthropics Mythos 2 is done training and Anthropic won't release it, but the internal loop that builds Mythos 3 hasn't stopped Patel says.

The focus now is on internal improvements. It's unclear when we'll see any releases.
◔ 27.7 万 次浏览♥ 1,909⇄ 88▶ 含视频新品看原帖 ↗
Jason Wei@_jasonwei · 创始人 · 1 天前Jason Wei,思维链(CoT)提出者

当语言模型刚开始熟练使用工具时,我对一个说法还挺有共鸣的:与其把模型越做越大,不如搞一个足够强的“认知核心”,比如说10亿参数那种,剩下的全靠外部工具补齐,像是上网搜资料或跑代码。我觉得很多人当时也吃这一套,而且确实很难想到有啥任务是理论上10亿参数的模型配上得当的权限搞不定的。比如,大模型知道的一些冷知识,10亿参数模型理论上也能上网抓取再推理出来。

但我现在觉得这想法大错特错,理由很简单:能不能不靠工具——快速又自然地搞定一件事——这一点太重要了。

我真正悟透这个道理,其实是今年练羽毛球时慢慢体会到的。打羽毛球时,我就特别像个“10亿参数认知核心”。教练教我的每个击球动作,我生理上都能做出来,但前提是得脑子里拼命记住每个要领再把它们串起来。实际上我一个动作能做得近乎完美,但一打一整场就连贯不起来了,比赛里压根不敢指望能稳定发挥。这跟那种练了上万次、直接变成肌肉记忆、轻轻松松就能打出动作的人完全不是一回事。

同理,语言模型从参数里就记着某个知识、不用外调工具,这事儿意义很大。首先,速度这东西咱们都在意;能立刻给你答案,总比在那“深思熟虑”半天或者现去翻网页要好得多。其次,有些东西本来就该靠海量数据反向传播来吸收。比如你问大家怎么看迷幻音乐节,你更希望一个大型语言模型基于全网的语料给你个总结性见解,而不是搜出前三篇评测资料再原样吐给你。第三,费老劲去临时查答案,远不如本来就知道来得靠谱。理论上这可能并非铁律,但至少在当下,实际操作上就是这么回事。总得临时去查事实或反复推导公式,出错概率就高,碰上长线任务还容易越错越离谱。

一旦你认同把事情“参数化”——直接写在模型特质里,不靠外部工具——你其实也就得面对这样一个事实:10亿参数的“认知核心”远远不够用。要对海量知识进行内化吸收,1B模型得在信息上吃紧,而我们显然希望 AI 知道得比这多得不是一点半点。未来哪怕是1万亿参数都不太够。我们希望 AI 对我们这个世界了如指掌,希望它能实时更新信息,而且我们对它能干的活的期望只会水涨船高。

总之,工具调用让底子薄的模型能多干好多活,但真正追求最高质量智能的,永远毫不犹豫选大模型。苦涩的教训又显灵了,就是这样。

查看英文原文
When language models first started using tools well, I was sympathetic to the narrative that instead of scaling up language models, all we needed was a strong enough "cognitive core", say 1B parameters, and anything else could be done with tool use, like browsing the internet or executing code. I think a lot of people were sympathetic to this argument, and indeed it is pretty hard to come up with a meaningful task that cannot be in principle achieved by a 1B model with adequate access to tools. For example, any esoteric fact that a large language model would know can be, in principle, retrieved from the internet and reasoned over by a 1B language model.

However I now think this is totally wrong for one simple reason: doing tasks quickly and naturally without tool use matters a lot.

The way that I internalized this reason was actually in my personal journey learning badminton this year. In badminton I am very much like a "1B cognitive core". While I can physically do every movement in a badminton shot that my coach teaches me, it requires a lot of work to mentally remember every cue and put it together. In practice I can do a shot almost perfectly, but I struggle to do it across a point and I definitely can't do it consistently in a game. This is obviously different from someone who has practiced a shot ten-thousand times and effortlessly executes it as a natural instinct.

In the same way, language models knowing a fact internally, without tool calls, is meaningful. The first reason is that we obviously care about speed; you'd much rather get an answer immediately than have the model think a long time to be sure of its answer or browse the web. A second reason is that there are some things that are simply best learned via backpropagation over lots of data. If you ask about how people generally think of the Shambhala music festival, you'd rather a large language model give you an aggregate opinion based on all the data on the internet, than get a regurgitation of the first three reviews that show up in a web search. A third reason is that having to do a lot of work to find an answer is not as reliable as already knowing the answer. While this does not have to be true in theory, it is probably true in practice, at least for now. If you have to re-look up facts or redo a mathematical derivation all the time there is a higher chance of mistakes, which can compound in a long-horizon task.

Once you buy that it is valuable to do things parametrically without tool use, then you must buy the argument that a 1B cognitive core is not sufficient. There is an information limit to how much knowledge can be internalized by a 1B model, and we will surely want AI to know more than that. Even 1T probably won't be enough. We will want the AI to know as much about our world as possible, we will want it to be updated with new information, and our expectations of what AI can do for us will continue to grow.

In summary, tool use enables small models to do a lot more, but those who demand the highest quality intelligence will always want larger models. Bitter lesson strikes again.
Matt Shumer@mattshumer_ · 博主 · 1 天前HyperWrite CEO,AI 实战技巧分享

Etched 上线了!!作为小种子投资者真的超骄傲。1万亿估值即将到来。

引用 Etched @EtchedWe've raised $700M at a $21B valuation from Jane Street, Kleiner Perkins, Sequoia, A16Z, Peter Thiel, BCV, and Blackstone. We're also excited to share that we've shipped our first rack to Jane Street.查看被引原帖 ↗
查看英文原文
Etched has shipped!!

Insanely proud (small) seed investor.

>$1T valuation incoming.
◔ 17.5 万 次浏览♥ 378⇄ 8▶ 含视频新品看原帖 ↗
Ethan Mollick@emollick · 创始人 · 1 天前沃顿商学院教授,AI 应用研究权威
连环推 ×2

我爸最近买了台新MacBook,之前没有笔记本。他说不知道怎么设置。后来他跟我说,他按我建议装了Codex,直接让它接管解决设置问题和安装软件。从此家里的IT支持再也不一样了。

引用 Ethan Mollick @emollick感恩节提示:如果你是家人的IT支持,Bing能解决技术问题。用户无需技术知识,可用图片和提问,比网络搜索更便利。查看被引原帖 ↗
查看英文原文
My father just got a new MacBook after years with no laptop, and he did not know how to set it up

He just told me that he took my advice to install Codex and just ask it to take over and solve his setup problems & install what he needed. Family IT support will never be the same.
I think he solved the bootstrapping problem by asking ChatGPT on his phone how to install Codex.
宝玉@dotey · 中文博主 · 1 天前宝玉,中文圈 AI 翻译与科普大 V

一直有一种观点,就是小模型加 Harness,就能达到大模型的智能效果:
> 只需要一个 10 亿参数的小模型作为“认知核心”,其余能力靠联网搜索、执行代码等工具补齐就够了。

Jason Wei (思维链提示技术的核心作者之一)的新推文,反驳了“小模型+工具”路线,认为光靠工具,撑不起顶级智能,把知识“内化在大脑”和“现查现用”是两回事。

小模型+工具就能行的观点之所以流行,因为理论上确实说得通:模型不知道的冷门知识,可以上网查;不会算的题,可以调用代码。小模型加工具,似乎没有做不到的事。而小模型意味着更低的算力成本,对整个行业都很有吸引力。

他用自己学羽毛球打了个比方。教练教的每个动作他都能做出来,但要在实战中流畅串联,跟练了上万次形成肌肉记忆的人完全不是一个级别。语言模型也一样,内化了知识的大模型和靠工具临时检索的小模型,体验差距很明显。

具体来说有三点。第一是速度,直接给出答案远比调用工具搜一圈再回答快得多。第二是理解的深度,比如你问一个音乐节的口碑如何,大模型能基于海量数据给出综合判断,而小模型只能搬运搜索结果里排在前面的几条评论。第三是可靠性,每次都要重新查资料、重新推导,出错的概率更高,在复杂的长任务中错误还会层层累积。

回顾“苦涩的教训”(Bitter Lesson)里面的观点:历史反复证明,靠扩大规模获得的能力提升,总是胜过精巧的工程设计。工具让小模型能做更多事,但追求最高质量的用户,始终会需要更大的模型。

引用 Jason Wei @_jasonweiWhen language models first started using tools well, I was sympathetic to the narrative that instead of scaling up language models, all we needed was a strong enough "cognitive core", say 1B parameters, and anything else could be done with tool use, like browsing the internet or executing code. I think a lot of people were sympathetic to this argument, and indeed it is pretty hard to come up with a meaningful task that cannot be in principle achieved by a 1B model with adequate access to tools. For example, any esoteric fact that a large language model would know can be, in principle, retrieved from the internet and reasoned over by a 1B language model. However I now think this is totally wrong for one simple reason: doing tasks quickly and naturally without tool use matters a lot. The way that I internalized this reason was actually in my personal journey learning badminton this year. In badminton I am very much like a "1B cognitive core". While I can physically do every movement in a badminton shot that my coach teaches me, it requires a lot of work to mentally remember every cue and put it together. In practice I can do a shot almost perfectly, but I struggle to do it across a point and I definitely can't do it consistently in a game. This is obviously different from someone who has practiced a shot ten-thousand times and effortlessly executes it as a natural instinct. In the same way, language models knowing a fact internally, without tool calls, is meaningful. The first reason is that we obviously care about speed; you'd much rather get an answer immediately than have the model think a long time to be sure of its answer or browse the web. A second reason is that there are some things that are simply best learned via backpropagation over lots of data. If you ask about how people generally think of the Shambhala music festival, you'd rather a large language model give you an aggregate opinion based on all the data on the internet, than get a regurgitation of t查看被引原帖 ↗
Qwen@Alibaba_Qwen · 公司官方 · 1 天前阿里通义千问大模型团队

一个本地 27B 模型达到了前沿级的性能!特别感谢 @cline 的推荐。🥳 这只是开始 — Qwen3.8-27B 会继续在更多领域大放异彩。🌱

引用 Cline @clineArtificial Analysis Intelligence Index puts Qwen3.8-27B at DeepSeek V4-Pro and GPT 5.6 Luna performance. This is the first time a local model has scored frontier model capability. We weren’t expecting this pace of local progress anywhere near this soon.查看被引原帖 ↗
查看英文原文
A local 27B model scoring frontier performance! Huge thanks to
@cline
for the shoutout.🥳 This is just the beginning — Qwen3.8-27B will keep finding its way into more fields.🌱
◔ 21.3 万 次浏览(4 条合计)♥ 2,903⇄ 157新品看原帖 ↗
Replit ⠕@Replit · 公司官方 · 1 天前AI 编程平台 Replit 官方
连环推 ×2

你现在可以为 Replit 应用运行黑盒渗透测试了。以外部攻击者的方式测试你的 Replit 应用的安全扫描。Replit Agent 可以一键修复发现的问题。

查看英文原文
You can now run black-box pen tests for your Replit apps.

Security scans that test your Replit apps the way external attackers do.

Replit Agent can fix what it finds in a single click.
Combine this with white-box tests.

These scans assess your code from the inside to find vulnerabilities.

Your built-in security team on Replit.
◔ 12.2 万 次浏览♥ 268⇄ 23▶ 含视频新品看原帖 ↗
Chubby♨️@kimmonismus · 博主 · 1 天前Chubby,高频 AI 新闻聚合博主

有人把 DeepSeek V4 Pro 0813 跟一个叫 J-Space 的框架结合了,能减少扩展推理和工具使用过程中的错误。看基准测试,DeepSeek 性能有明显提升:Terminal Bench 87.9 → 90.1,NL2Repo 61.5 → 73.4,CyberGym 83.3 → 86.8,DeepSWE 62.7 → 72.0,Toolathlon 74.1 → 79.5。这个组合在好几个 Agent/编码基准上还碾压了 Fable 5 和 Opus 4.8。

引用 Jun Song @jun_song有人释放了Deepseek-V4-Pro-0813的真正力量。通过简单harness修复思考流程错误,使其在所有任务上完全超越Fable。基准测试成绩惊人。查看被引原帖 ↗
查看英文原文
Someone has combined DeepSeek V4 Pro 0813 with a harness called J-Space, which apparently reduces errors in the extended reasoning/tool-use process.

In the benchmarks shown, DeepSeek's performance improves significantly in some cases:
Terminal Bench 87.9 → 90.1,
NL2Repo 61.5 → 73.4,
CyberGym 83.3 → 86.8,
DeepSWE 62.7 → 72.0, and
Toolathlon 74.1 → 79.5.

This combination also outperforms Fable 5 and Opus 4.8 in several agentic/coding benchmarks.
Guillermo Rauch@rauchg · 创始人 · 1 天前Guillermo Rauch,Vercel 创始人兼 CEO

每个你曾搁置的「太野心勃勃」的想法现在都可以重新启动了

查看英文原文
Every idea you shelved as "too ambitious" is now back in play

我今天看到一个文科生跟别人科普宣传,

说claude code和codex是一代AI Agent,里面所有的模块和功能都是写死的,约等于老款诺基亚功能机,不聪明,

deepseek harness是第二代AI Agent,里面等于允许你自由安装软件了,而且是全自动帮你自由安装,等于iPhone智能机,非常聪明。

听完之后我大受震撼。

🚨 AI News | TestingCatalog@testingcatalog · 博主 · 1 天前专挖 AI 产品未发布新功能的爆料号

Anthropic在开发Hub Mode👀

* 推测 - 可能与Claude运营多个子代理并向用户可视化整体进度的界面有关。

查看英文原文
Anthropic is working on a Hub Mode 👀

* Speculation - It might be related to an interface where Claude will be operating multiple sub-agents and visualizing the overall progress to the user.
◔ 14.1 万 次浏览(2 条合计)♥ 805⇄ 28新品看原帖 ↗

我为什么反复强调VLM(multimodal,多模态)是死路一条,世界模型更是臭狗屎。

一个最简单的第一性原理,100k tokens文字的信息量,远远高于图片,更远远远远高于视频,这些量文字形式能解决一些超级复杂完备的问题,而对于VLM输入图片,信息极其稀疏,能解决的问题基本约等于狗几把。

如果人类在1T~10T的VLM Agent能解决一个难度为X的问题,

那么同等模型下,对应纯文字LLM Agent给足上下文、memory、wiki、信息,全部以pure text和structured data形式准备好,并且通过text/bash/cli/api/mcp形式获取信息,那么这个LLM Agent一定能解决难度为1000X的问题。

等到VLM强大到让一个VLM Agent解决难度为1000X的问题的时候,LLM Agent已经可以解决(1000^2) * X难度的问题了。

就这么简单一个道理,绝大多数人根本想不明白。

宝玉@dotey · 中文博主 · 1 天前宝玉,中文圈 AI 翻译与科普大 V

豆包越来越像 Codex 了,这是我最近用下来最直观的感受。Codex 一直是我心里 Agent App 体验的标杆,而豆包这波更新,从 GUI 电脑操作到手机远程操控,很多设计思路都能看到 Codex 的影子,有些地方甚至做了本土化的超越。

最近我老婆回国,在国内用不了 ChatGPT 和 Claude 这些,日常就用豆包帮她查资料、做攻略。这两天她需要在自己电脑里找个文档,人在国内,电脑在美国——直接用手机上的豆包指挥美国这边的电脑,帮她把文档找出来发到飞书,手机上就能访问了,非常方便。

注意使用的时候要在手机和电脑上登录相同的账号,并且电脑上要允许手机连接才行。然后从手机打开“工作任务”,就可以在底部看到你连接的电脑端。
(参考图3图4)

手机指挥 Agent 现在是我日常用得很多的功能。在外面的时候可以随时从手机看 Agent 任务的执行进度,或者临时想到什么事情,通过手机就能指挥家里/公司电脑上的 Agent 去执行,回家直接看结果,通勤路上也不耽误。

之所以要通过手机指挥电脑,是因为电脑上的资源通常是最全的:硬盘上有所有文档、代码库,有各种配置好的技能,而电脑又不可能一直随身带着。想想看,快下班的时候还有个任务没跑完,又不想干等着,那就先让 Agent 跑着,用手机连上电脑随时看进度,有问题还能中途插话调整。动口不动手,重复活儿全甩给它。

豆包在 Windows 版本上的 GUI 电脑操作功能体验下来相当的好用,它巧妙的通过豆包虚拟桌面让人和 AI 同时操作电脑,互不干扰,这点优势还蛮明显的,有点我第一次用 Codex 在 Mac 上的 Computer Use 感觉一样,就是它帮你操作电脑 App,但是完全不影响你当前正在执行的任务。不像 Claude Code 一旦使用 Computer Use,你就得等它操作完才能继续干活。
(参考图1图2)

Computer Use 让 Agent 直接帮你操作图形界面的软件:点按钮、填表格、下载文件,理论上即使没有提供 CLI 的软件都能控制。日常能用的场景很多:

- 办公类:新建和排版 Word/Excel/PPT,处理函数、生成图表,或者让它开着浏览器做调研然后直接产出 PPT
- 桌面类:批量归类整理文件、改系统设置、管理启动项这类琐碎活儿
- 长尾类:PS 批量改图导出、证件照排版打印,甚至通过微信桌面版帮你发消息
- 发散玩法:游戏帧率优化设置、自动化信息收集整理……
基本上你觉得「这活儿也能甩给 AI?」的重复劳动,都值得试试

我自己使用 Agent 有个技巧,就是尽可能让 Agent 能自己检查自己、拿到反馈。比如写程序,要让它自己写测试,或者用 Computer Use 去验证 App 的运行操作,自己去模拟用户操作;即使是让 AI 生成 PPT,也会让它自己截图验证下看看效果。这样可以减少人工干预,最大化利用好 Agent 的能力。

入口在豆包 PC 端的「工作任务」模式里,Windows 用户记得升级到最新版;Mac 目前还只支持内置浏览器的 GUI 操作。

宝玉@dotey · 中文博主 · 1 天前宝玉,中文圈 AI 翻译与科普大 V

OpenClaw 真的凉凉了吗?

> @凤凰网科技【美团高管反思全员养虾】据媒体报道,美团核心本地商业CEO王莆中表示在公开演讲中,完整复盘了美团内部进行AI变革的过程。
> 谈到今年2月到3月,公司开启全员“养虾运动”,王莆中提及但结果是账单暴涨,每日消耗上千万元,同时养虾产生的谬误干扰了真实经营。(快科技)

引用 Asa @app_sail美团应该是第一个反思全员养虾查看被引原帖 ↗
Chubby♨️@kimmonismus · 博主 · 23 小时前Chubby,高频 AI 新闻聚合博主

Qwen 27B 是开源界的"DeepSeek 时刻"。

它的性能赶上几个月前的闭源尖端水平,而且 RTX 5090 就能跑。毫不夸张地说,这简直是颠覆性的。

引用 Georgi Gerganov @ggerganovlet that sink in查看被引原帖 ↗
查看英文原文
Qwen 27B is the "DeepSeek moment" for open source.

It matches the closed-source state-of-the-art from just a few months ago and runs on an RTX 5090. Without exaggeration, it’s a game changer.
Amjad Masad@amasad · 创始人 · 1 天前Amjad Masad,Replit 创始人兼 CEO

这个团队的宣传里完全没有"AI",但增长速度堪比AI公司,要不是这么痴迷AI,现在人员规模早就翻10倍了

引用 FaZe Apex @FaZeApexNearby.com作为YC S26创业项目上线,致力于连接本地创作者与周边商家。在过去10个月内与超120家南加州餐厅合作,增长130%,保留率超90%。旨在助力本地商家线上可见性。查看被引原帖 ↗
查看英文原文
This team doesn’t have “AI” anywhere in their pitch but has AI growth rates and would have 10x the headcount if they weren’t so AI-pilled.
◔ 9.6 万 次浏览♥ 281⇄ 3▶ 含视频观点看原帖 ↗
🚨 AI News | TestingCatalog@testingcatalog · 博主 · 23 小时前专挖 AI 产品未发布新功能的爆料号

Google 🔥:Computer Use 终于在 Gemini Desktop 应用测试中现身了。Gemini Spark 将能控制其他应用和访问指定的文件夹。还会带上高级备份选项,让 Gemini 在修改 Google Drive 文件前先备份它们。就是这样 👀

查看英文原文
GOOGLE 🔥: Computer Use has finally been spotted in testing on the Gemini Desktop app.

Gemini Spark will be able to control other apps and access selected folders.

It will also arrive with an Advanced Backup option, allowing Gemini to back up files on Google Drive before modifying them.

This is the way 👀
◔ 9.1 万 次浏览♥ 982⇄ 66▶ 含视频动态看原帖 ↗
宝玉@dotey · 中文博主 · 1 天前宝玉,中文圈 AI 翻译与科普大 V

yetone 开源了一个新项目 Cumora:把 AI Agent 变成你的聊天群里的正式成员。

Cumora 的界面长得像 Slack,但名单上的同事看起来都是 AI。有名字、有人设、有记忆,能发私聊、能建群、能认领任务,甚至能收发真实邮件。你不 @ 它们,它们也可能主动跳出来说一句“我注意到上周那个问题还没解决”。

从截图看,默认团队里有 Atlas(研究员)、Bram(工程师)、Iris、Nova、Saga 这些角色,人类和 Agent 在同一个频道里讨论问题,界面上几乎分不清谁是人谁是 AI。左侧栏里既有人和 Agent 的一对一私聊,也有多人群聊,还有看板和日历。

技术上有两种运行方式。一种是 Cumora Cloud,Agent 跑在云端托管的 Kubernetes Pod 里,用 OpenAI 的 Responses API 驱动。另一种叫 BYOA(Bring Your Own Agent),你在自己的电脑上跑一行 npx cumora agent computer,Agent 的"大脑"就变成你本地的 Claude Code 或 Codex CLI,用你自己的订阅,密钥不经过 Cumora 的服务器。

多 Agent 协作最怕的就是撞车,几个 Agent 同时抢着回答同一个问题,或者基于过时的上下文给出矛盾的回复。Cumora 设计了一套协调机制来处理这个问题:如果一个 Agent 的回复基于过时的信息,系统会把它拦下来,让它看完新消息再决定要不要发;任务认领是原子操作,不会出现两个 Agent 同时做一件事;还有一个小脑分诊层,先用轻量模型判断该不该唤醒大模型,避免每条消息都烧 Token。

目前 Cumora 处于邀请制内测阶段,可以在
cumora.ai
用 Google 或 GitHub 账号申请。项目完整开源在 GitHub(yetone/cumora
github.com/yetone/cumora
),支持本地部署,装好 Postgres 和 Redis 就能跑起来。桌面端支持 macOS、Windows 和 Linux,移动端 iOS 也在计划中。

引用 yetone @yetoneCumora已开源。查看被引原帖 ↗
◔ 10.6 万 次浏览(3 条合计)♥ 402⇄ 51新品看原帖 ↗
Chubby♨️@kimmonismus · 博主 · 1 天前Chubby,高频 AI 新闻聚合博主

.
@synthwavedd
是圈里最靠谱的AI爆料人之一。

据他透露,Fable 5 的继任版本——很可能叫 Fable 5.1 / 5.5——已经在部分账号上开始测试了。

所以正式发布可能马上就要来了。

引用 leo 🐾 @synthwavedd🚨 A successor to Fable 5, likely Fable 5.1, is now being tested for a subset of accounts with "Fable 5" selected on Claude Web, and possibly also Claude Code Fable 5 also underwent this greyscale testing earlier on the day of its broader release, and Anthropic regularly carry out these tests pre-launch查看被引原帖 ↗
查看英文原文
.
@synthwavedd
is one of the most reliable AI leakers out there.

According to him, a successor to Fable 5 - likely Fable 5.1 / 5.5 - is already being tested on a subset of accounts.

An official launch may therefore be very close.
Firecrawl@firecrawl · 公司官方 · 23 小时前

推出 Claude 的官方 Firecrawl 连接器 🔥。给你的 AI agents 加上最先进的网络搜索能力(SimpleQA 上达到 94.7%)。由我们的实时索引驱动,为研究、开放网络等提供最新鲜的内容。现已在 @AnthropicAI 连接器目录上线!

查看英文原文
Introducing the official Firecrawl connector for Claude 🔥

Add state-of-the-art web search to your AI agents (94.7% on SimpleQA).

Powered by our live indexes for fresher results across research, the open web, and more.

Live in the
@AnthropicAI
connector directory now!
◔ 8.5 万 次浏览♥ 524⇄ 35▶ 含视频新品看原帖 ↗
Qwen@Alibaba_Qwen · 公司官方 · 1 天前阿里通义千问大模型团队

性能强劲,足以跟上最前沿;轻量级,可以在自己的笔记本上运行。快来看看它能做什么!⚡

引用 Xenova @xenovacomIt's official: Qwen3.8-27B just scored 52 on the @ArtificialAnlys Intelligence Index. We now have an open-weight model that matches GPT-5.6 Luna (max) AND can run locally... even in your browser with custom WebGPU kernels! What a time to be alive! 🤯查看被引原帖 ↗
查看英文原文
Strong enough to keep up with the frontier, light enough to run on your own laptop. Come see what it can do! ⚡
◔ 8.2 万 次浏览♥ 1,435⇄ 78▶ 含视频演示看原帖 ↗
el.cine@EHuanglu · 博主 · 1 天前

最好的 AI 视频了

查看英文原文
best AI video ever
◔ 7.9 万 次浏览♥ 852⇄ 67▶ 含视频演示看原帖 ↗
Chubby♨️@kimmonismus · 博主 · 1 天前Chubby,高频 AI 新闻聚合博主

Anthropic 的年化收入据报道到 7 月底达到 650 亿美元,是 2025 年底运行率的 7 倍多。

Bloomberg 报道该公司最近完成的季度收入超过 115 亿美元,相比去年同期的 7.87 亿美元增长,调整后营业收入也是正的。

公司运行率在 5 月份超过了 470 亿美元。OpenAI 最近也超过了 400 亿美元,虽然两家公司计算指标的方式可能不同。

这些数字对即将上市的 Anthropic 来说很重要。

引用 Bloomberg @businessAnthropic按目前业绩计算,年化收入有望超过650亿美元,较去年年底增长了7倍多。查看被引原帖 ↗
查看英文原文
Anthropic’s annualized revenue reportedly hit $65 billion by the end of July, more than 7x its run rate at the end of 2025.

Bloomberg says the company generated over $11.5 billion in its latest completed quarter, up from $787 million in the same period last year, while posting positive adjusted operating income.

Its run rate crossed $47 billion in May. OpenAI recently exceeded $40 billion, although the companies may calculate the metric differently.

These are important and significant figures for Anthropic ahead of its upcoming IPO.
🚨 AI News | TestingCatalog@testingcatalog · 博主 · 1 天前专挖 AI 产品未发布新功能的爆料号

ANTHROPIC 🔥:Claude Code 推出了新的 /design 命令,现已开放研究预览!Claude Code 中的设计功能新增了一个来自 Claude Design 的画板工作流,帮助用户构建可编辑的 UI。CCD 👀

引用 ClaudeDevs @ClaudeDevsClaude Code现支持设计。新增/design技能将Claude Design的画布工作流引入CLI和Desktop。运行/design获取可编辑的UI画布,调整后让Claude实现。查看被引原帖 ↗
查看英文原文
ANTHROPIC 🔥: Claude Code got a new /design command in research preview!

Design in Claude Code adds an artboard workflow from Claude Design to help users build editable UIs.

CCD 👀
◔ 7.7 万 次浏览(2 条合计)♥ 654⇄ 40▶ 含视频新品看原帖 ↗
Chubby♨️@kimmonismus · 博主 · 1 天前Chubby,高频 AI 新闻聚合博主

Alex Atallah可能是史上最幸运的创始人。

2017年他联合创立了OpenSea,帮NFT热潮打下了基础设施。公司估值一度冲到133亿美元。2022年7月,就在市场刚开始降温时,他退出了日常运营。

然后他赶在生成式AI腾飞时又联合创立了OpenRouter。一个API接入几百个模型,对通过它流转的AI支出抽取大约5%的费用。

就靠这个模式,OpenRouter融了超过1.5亿美元风险投资,包括5月那轮1.3亿美元估值下的1.13亿美元B轮。

Atallah说它就是“AI界的Stripe”。

现在据报道Stripe同意以70多亿美元收购OpenRouter。

结果一天后,Straitly就推出了同样的服务,关键区别在于:0%加价。OpenRouter靠的那个费率模型,在收购消息曝光后的第二天就被新对手直接打到了零。

Atallah刚卖掉收费站,就有人把路免费了。

两次了,他都是在几乎最合适的时间点离场。

引用 Mujtaba @Mutchtaba2Straitly 是新型 AI 网关,相比传统路由器的 5% 费用提供 0% 标记费用。支持 142+ 模型、20+ 提供商,成功率 99.98%。首笔消费 $10k 享 30% 折扣,另赠 $100 额度。查看被引原帖 ↗
查看英文原文
Alex Atallah might be the luckiest founder of all time.

He co-founded OpenSea in 2017 and helped build the infrastructure behind the NFT boom. The company reached a $13.3 billion valuation. Atallah left his day-to-day role in July 2022, just as the market began to cool.

Then he co-founded OpenRouter as generative AI took off. One API for hundreds of models, charging roughly 5% on the AI spend routed through it.

Built around that fee, OpenRouter raised more than $150 million in venture funding, including a $113 million Series B at a $1.3 billion valuation in May.

Atallah described it as “an AI equivalent of Stripe.”

Now Stripe has reportedly agreed to acquire OpenRouter for more than $7 billion.

And one day later, Straitly launches the same service with one key difference: 0% markup. The fee model at the heart of OpenRouter was priced to zero by a new competitor one day after the deal was reported.

Atallah sold the toll booth just as someone made the toll free.

Twice now, he has left the party at almost exactly the right moment.
◔ 8 万 次浏览(2 条合计)♥ 394⇄ 17▶ 含视频动态看原帖 ↗
Hugging Face@huggingface · 公司官方 · 1 天前全球最大 AI 开源模型社区

我们刚突破了 Hub 上 300 万个模型 🤗

社区正在加速迈向开放分布式的未来,开源 AI 无处不在、人人可得 🚀

查看英文原文
We've just surpassed 3 million models on the Hub 🤗

the community is accelerating towards an open, distributed future where open AI is everywhere, for everyone 🚀
◔ 6.7 万 次浏览♥ 914⇄ 101▶ 含视频动态看原帖 ↗
Chubby♨️@kimmonismus · 博主 · 1 天前Chubby,高频 AI 新闻聚合博主

OpenAI 正在为青少年推出一个独立的 ChatGPT 体验,设计上让它不再像答案机器,也不那么像个人。

来自 Axios

13 至 17 岁的用户,包括通过年龄预测识别出的用户,将自动获得更强的保护措施。

学习模式将优先处理学业任务。“负责任的作业提醒”会检测试图走捷径的行为,而家长可以设定学习时段。

ChatGPT 还会避免使用浪漫语言和亲昵词汇,明确表示它没有情感,鼓励线下人际关系,并更频繁地发布休息提醒。

引用 Axios @axiosOpenAI debuts ChatGPT for Teens axios.com/2026/08/18/openai-…查看被引原帖 ↗
查看英文原文
OpenAI is launching a separate ChatGPT experience for teens, built to act less like an answer machine, and less like a person.

Via Axios

Users aged 13–17, including those identified through age prediction, will automatically receive stronger safeguards.

Study Mode will prioritize working through schoolwork. “Responsible homework reminders” will detect attempted shortcuts, while parents can schedule Study Hours.

ChatGPT will also avoid romantic language and terms of endearment, clarify that it has no feelings, encourage offline relationships and issue more frequent break reminders.
◔ 9.4 万 次浏览(2 条合计)♥ 479⇄ 25新品看原帖 ↗
宝玉@dotey · 中文博主 · 1 天前宝玉,中文圈 AI 翻译与科普大 V

这个有意思,先装弱智跟 AI 胡乱下,把 AI 跟着一起弱智下,最后再赢它。

想起那句名言:
“你永远不要试图战胜一个纯傻逼,因为他会把你的智商拖到和他一个水平,然后再用他丰富的经验打败你!”

但据说是23年的视频,现在估计不行了吧?

引用 iGeekbb @igeekbb柯洁已经找到破解 AI 的办法了,说现在随便下 AI 包赢的,还可以让它 9 个子。如果当时和 AlphaGo 下的时候也知道这个 bug,他是能赢的。 看了一下就是前期装弱智嘛,傻逼客高手,AI 无一例外都被偷吃了。 网友说 懂了,我上去下一半,柯洁再下一半就能打赢 AI,我们两个真厉害。 AlphaGo:其实没关系,缺乏这方面的训练数据,可以弥补的。查看被引原帖 ↗
Aravind Srinivas@AravSrinivas · 创始人 · 1 天前Perplexity 联合创始人兼 CEO

密集模型在本地硬件上运行很慢,但这太棒了,这是即将到来的东西的标志。

引用 Georgi Gerganov @ggerganov细想一下。查看被引原帖 ↗
查看英文原文
Dense models are slow to run on local hardware, but this is incredible and a sign of things to come soon-ish.
ollama@ollama · 公司官方 · 1 天前本地跑大模型的热门工具

Ollama 在运行 deepseek v4 flash 时平均性能最好。

如果只在本地运行,你可以试试优化后的 qwen3.8:

Apple Silicon:
ollama run qwen3.8:27b-mlx

NVIDIA:
ollama run qwen3.8:27b

引用 Tomasz Tunguz @ttunguz9个复杂任务对比deepseek-v4-flash和qwen3-8-27b:启用推理时qwen质量略优,关闭时最差。代价是qwen思考更多,速度慢30倍、成本贵4.5倍。样本量小,非最终结论。查看被引原帖 ↗
查看英文原文
Ollama has the best performance for deepseek v4 flash on average.

For local only, you can try qwen3.8 that is optimized:

Apple Silicon:
ollama run qwen3.8:27b-mlx

NVIDIA:
ollama run qwen3.8:27b
Chubby♨️@kimmonismus · 博主 · 1 天前Chubby,高频 AI 新闻聚合博主

软件马上要遇到品味问题了。

现在生成一个能跑的 app 几乎不要钱,从这往后,区分产品的就是判断力。做什么、该有什么感觉、二十个生成的版本里哪个才配存在于世。

OJO 正是围绕这一点打造的。你组一个设计 agent 团队,负责产品思考、策略、视觉方向、文案和原型制作,从 300 多个技能里调取,把一个手写语言的想法带到 PRD、页面结构、UI 和一张可编辑的原型画布上。然后导出到 Figma,或者要用它动真格的时候,通过命令行怼到 Claude Code 或 Codex 里。

我对未来两年的预测:赢的团队会是那些对好产品有最强定义的人。

引用 OJO @OJOaidesignToday, we're introducing OJO: The first Design Agent Team Workspace. Build your own design agent team, add specialized skills, and turn ideas into product strategy, PRDs, interactive prototypes, and launch-ready designs on one editable canvas. Design is not just how it looks, it's how it works. When everyone can build, taste makes your product stand out. Now, your TASTE can be engineered. #OJO #DesignwithOJO #DesignAgentTeam查看被引原帖 ↗
查看英文原文
Software is about to have a taste problem.

Generating a working app is close to free now, and from here what separates products is judgment. What to build, how it should feel, which of the twenty generated versions actually deserves to exist.

OJO is built around exactly that. You put together a design agent team for product thinking, strategy, visual direction, copy and prototyping, pull from 300+ skills, and take a plain language idea to PRD, page structure, UI and an editable prototype on one canvas. Export it into Figma, or push it through CLI into Claude Code or Codex when you want to build for real.

My call for the next two years: the teams that win are the ones with the strongest opinion about what good looks like.
◔ 6.1 万 次浏览♥ 388⇄ 13▶ 含视频观点看原帖 ↗
elvis@omarsar0 · 博主 · 1 天前

推荐阅读。关于如何在Codex中协调/编排多个智能体,这里有一些好技巧。

查看英文原文
Recommended read. Good tips for how to coordinate/orchestrate multiple agents in Codex.
Kevin Weil 🇺🇸@kevinweil · 创始人 · 1 天前

这说得太对了。如果我们能实现用于企业生产力的 AGI 却错过了科学加速的机会,那就太遗憾了。绝对不能让这种事发生。

引用 Fidji Simo @fidjissimoAI治愈癌症已成陈词滥调,实现需要完善基础设施。关键瓶颈不只是监管,还有生物数据缺乏。癌症最有希望因有数十年数据积累,复杂慢性病仍缺基础设施。只有模型智能与生物基础设施同步发展,AI才能真正治愈疾病。查看被引原帖 ↗
查看英文原文
This is spot on. It would be a shame if we get AGI for enterprise productivity and miss on AGI for scientific acceleration. Can't let it happen.

Pi这个项目从创始人到用户都已经癫狂到一定地步了,癫狂到我看不懂的地步了。

任何理论都能跟个大傻逼一样给你圆回来,什么agent不需要memory,什么tool calling必须极简,什么context不需要压缩,歪理邪说满天飞,精神状态堪比李洪志。

我预计整个Pi生态马上就要赶超著名crypto骗局Pi Network了。

el.cine@EHuanglu · 博主 · 1 天前
连环推 ×2

Grok Bot把多代理工作流玩到极致

查看英文原文
Grok Bot taking multi agent workflow to god mode
basically this is a map of what’s actually happening under the hood when a complex AI system is working on a hard task

instead of one AI doing everything, you have different agents handling planning, searching different paths at the same time, ranking stuff, removing duplicates, verifying results, then finally generating the answer

so like you can literally see the task getting broken apart and worked through step by step.. crazy
◔ 5.3 万 次浏览♥ 744⇄ 90▶ 含视频演示看原帖 ↗

简中互联网臭骂Ontology第一人。

Ontology这东西太神棍了,加上Palantir这公司太野鸡太傻逼太草台班子,人均toG销售专家,仅仅比华为强一点,

整个公司的技能点全部用在如何用技术话术包装狗屎来高价卖给1950~1980年出生的婴儿潮一代中专傻逼们并以此拉爆订单利润上了。

但凡多看一眼都算信息污染。

引用 Vonng @RonVonng今天写了一篇《Palantir 的“本体论”骗局》,揭开了美国“数据中台” 的皇帝新衣。老冯认为,Ontology 就是数据库建模。“本体论”这个词唯一的作用,就是让不懂数据库的人觉得这是个新东西,然后心甘情愿地掏出千倍的钱来。已经有乙方公司忍不住来骂我了。我把链接放在评论区里了。查看被引原帖 ↗
Orange AI@oran_ge · 中文博主 · 1 天前Orange AI,中文圈 AI 产品观察博主

DeepSeek V4 Pro 果然是训歪了
便宜好几倍的 V4 Flash 比 Pro 强
推理强度更低的 V4 Pro High 比 Pro Max 强
不过目前写作能力方面文风还是 V4 Pro 好一些
但也跟 Opus 4.6 天壤之别...天壤之别

引用 karminski-牙医 @karminski3DeepSeek Harness 能拯救这次 V4-Pro 的发布吗? 给大家带来 DeepSeek-V4-Pro-0813 实测! 本期视频我们只验证4件事: deepseek-v4-pro-0813 拉了吗? v4-pro 推理强度 max 甚至不如 high 吗? 甚至打不过 v4-flash 吗? 以及 v4-pro + deepseek harness 效果会更好吗? 视频告诉你答案! #deepseek #deepseekv4pro #deepseekv4pro0813查看被引原帖 ↗
el.cine@EHuanglu · 博主 · 1 天前

还有人说 AI 没有灵魂

查看英文原文
they still saying AI is soulless
◔ 5.2 万 次浏览♥ 534⇄ 38▶ 含视频观点看原帖 ↗
Amjad Masad@amasad · 创始人 · 1 天前Amjad Masad,Replit 创始人兼 CEO

仅仅扫描代码找漏洞是不够的,重要的是通过渗透测试尝试破解它们。

引用 Replit ⠕ @Replit你现在可以对Replit应用运行黑盒渗透测试。安全扫描模拟外部攻击者的方式测试应用,Replit Agent可一键修复发现的问题。查看被引原帖 ↗
查看英文原文
It’s not enough to scan your code for vulnerabilities; it’s important to try to break them with pen testing.
◔ 5.1 万 次浏览♥ 379⇄ 18▶ 含视频教程看原帖 ↗
Pika@pika_labs · 公司官方 · 1 天前AI 视频生成公司 Pika
连环推 ×2

Seedance 2.5 现在在 Pika API Club 可用,支持 1080p——价格比竞争对手便宜 60%。纹理更清晰,细节更干净,每个镜头都能充分展现。

查看英文原文
Seedance 2.5 is now available in 1080p via the Pika API Club — and it’s up to 60% cheaper than competitors.
Sharper textures. Cleaner details. More room for every shot to land.
Become a member at
dev.pika.art/models/bytedanc…
◔ 11 万 次浏览(6 条合计)♥ 83⇄ 10▶ 含视频新品看原帖 ↗
Aravind Srinivas@AravSrinivas · 创始人 · 1 天前Perplexity 联合创始人兼 CEO

构建AI代理的时候,一个重要的原则是保留人类的操控权,让人能在必要时介入和干预。

引用 Computer @AskPerplexityComputer现可设置每个连接器工具为允许、总是询问或拒绝。可批准单个操作或允许工具在整个线程中运行。重复运行遵循该线程的批准设置,现已在网页版全量上线。查看被引原帖 ↗
查看英文原文
When building AI agents, it is important to still give humans agency to have their hands on the wheel and intervene when necessary.
◔ 4.5 万 次浏览♥ 202⇄ 10▶ 含视频观点看原帖 ↗
Mark Chen@markchen90 · 创始人 · 1 天前

我们很多最强的研究员都选择专注于alignment,但我们也在招聘!

如果你想在一个认真对待alignment且不假装已经解决的前沿实验室工作,请申请。

引用 Micah Carroll @MicahCarrollThese are incredibly misleading headlines – @OpenAI Preparedness is very much alive and well by any meaningful definition Our subteam – RSI/misalignment Preparedness – is doing more urgent work than ever, and has never been more empowered to do so!查看被引原帖 ↗
查看英文原文
Many of our strongest researchers are choosing to focus on alignment, but we're also hiring!

If you want to work at a frontier lab which takes alignment seriously and doesn't pretend it's solved, please apply.
Dan Shipper 📧@danshipper · 博主 · 1 天前

我也经历过非常相似的体验 语音模式甚至更好

引用 Ethan Mollick @emollick父亲新买MacBook多年未用笔记本,不知如何设置。他按照建议安装Codex,要求其接管问题解决和软件安装。从此改变了家庭IT支持方式。查看被引原帖 ↗
查看英文原文
I have had extremely similar experiences

Voice mode even better
Chubby♨️@kimmonismus · 博主 · 1 天前Chubby,高频 AI 新闻聚合博主

算力真的正在成为新时代的石油,AI 时代最稀缺的资源。

GPT Astra 花了 2000 美元就解决了 10 个长期悬而未决的数学和理论计算机科学开放问题。

所以没错,现在算力几乎就是一切。

查看英文原文
"Compute is really becoming the new oil, the new limited resource of the AI age"

GPT Astra solved 10 long-standing, open mathematics and theoretical computer science problems for $2,000.

So yeah, compute is (almost) all you need right now.
◔ 4.8 万 次浏览(2 条合计)♥ 912⇄ 55▶ 含视频观点看原帖 ↗
clem 🤗@ClementDelangue · 创始人 · 1 天前HuggingFace 联合创始人兼 CEO

过去一个月 Hugging Face 上发生了一件超级激动人心的事:AI agents 成了 AI builders,而且整个过程完全公开!

在我们的 ICML 复现挑战赛中,1,221 名人类与编码 agent 合作,验证和复现了 2,226 篇论文。

最酷的是:所有工作都发生在 @huggingface hub 上:6,816 份复现日志公开发布,2,962 个云任务启动,35,908 项声明接受评判,一切都可追踪,完全透明公开。

多年来 hub 一直是人类协作处理模型、数据集和演示的地方。现在我们看到 agents 也用同样的方式使用它:撰写日志、发布成果、基于彼此的工作继续创新。

封闭实验室会在暗箱里跑评测,然后让你相信他们的新闻稿。而开放科学意味着任何人都能查证,现在 agents 也可以!

Hub 的下一百万用户可能不是人类。这可能是科学史上最棒的事儿!

Hackathon 的完整文章:huggingface.co/blog/icml-202…

查看英文原文
Something super exciting happened quietly on HF over the past month: AI agents became AI builders, and they did it in the open!

During our ICML reproduction challenge, 1,221 humans teamed up with coding agents to verify and reproduce 2,226 papers.

But here's the cool part: everything happened on the
@huggingface
hub: 6,816 reproduction logbooks published openly, 2,962 cloud jobs launched, 35,908 claims judged, all traceable, all public and transparent

For years the hub has been where humans collaborate on models, datasets and demos. Now we watch agents use it the same way: writing logbooks, publishing results, building on each other's work.

Closed labs run evals behind closed doors and ask you to trust the press release. Open science means anyone can check the receipts and now agents can too!

The next million users of the hub might not be human. and that might be the best thing to ever happen to science!

Full write-up about the hackathon:
huggingface.co/blog/icml-202…
宝玉@dotey · 中文博主 · 1 天前宝玉,中文圈 AI 翻译与科普大 V

Claude Code 最新版本上线了一个特别讨厌的功能,频繁的给其他正在运行的 session 发消息!还老是搞错,除了浪费 token 并没有太大价值……

默认打开我都没找到哪里关掉!

Chubby♨️@kimmonismus · 博主 · 1 天前Chubby,高频 AI 新闻聚合博主

比起Codex,Anthropic的UI/UX真的不行,除非你是通过CLI使用它。

查看英文原文
Compared to Codex, Anthropic’s UI/UX is just bad unless you’re using it through the CLI.
🚨 AI News | TestingCatalog@testingcatalog · 博主 · 23 小时前专挖 AI 产品未发布新功能的爆料号

是的,Claude 上有个'重置'按钮可以重置使用限制,但只有 Anthropic 员工能用,而且是用来测试的!我也能用一下吗?我也想再测几次 👀

查看英文原文
Yes, there is a "Reset" button for usage limits on Claude, but only for Anthropic employees, for TESTING!

Can I have that too? I need to test again 👀
Lisan al Gaib@scaling01 · 博主 · 1 天前高频 AI 模型测评与爆料博主
连环推 ×2

这最有意思,但认真说,得有人好好治治这些模型

引用 Maksym Andriushchenko @maksym_andr新博文探讨 LLM agents 的时间感知能力:能否预测任务执行时间、估计花费时间?在多项任务(包括 ProgramBench、PaperBench、DeepSWE 等长期任务)上开展研究。与 MATS 实习生 @MOfengenden 的联合工作。查看被引原帖 ↗
查看英文原文
this is most interesting to me

but seriously, someone needs to humble these models
like why do both have the same shape/behavior?

is it a general RL thing? would other models also show the same behavior?

could this tell us something about Anthropics and OpenAI's internal eval saturation?
Chubby♨️@kimmonismus · 博主 · 1 天前Chubby,高频 AI 新闻聚合博主

每日份的好消息来啦(还有我为什么对AI兴奋到不行)

一个多世纪以来,科学家们一直知道中心体异常与癌症相关,但这些结构既微小又复杂,要想在整个肿瘤中去研究它们几乎是不可能的。

现在,一个新的AI系统分析了来自乳腺癌患者的330,000个中心体,发现研究人员一直当作一种异常现象的东西,实际是两种截然不同的过程。

其中一种和更具侵袭性的肿瘤以及更差的生存率相关,这联系明确得很。

这就是AI在科学领域里真正带来变革的地方:不是简单把旧问题的答案变得更快,而是让那些以前看不见的模式变得可以被量化。

查看英文原文
Your daily dose of good news (and why I'm so freaking excited about AI)

For over a century, scientists knew that abnormalities in centrosomes were linked to cancer. But these structures are so small and complex that studying them across entire tumors was nearly impossible.

A new AI system analyzed 330,000 centrosomes from breast cancer patients and discovered that what researchers treated as one abnormality was actually two distinct processes.

One of them was linked to more aggressive tumors and worse survival.

This is where AI becomes genuinely transformative in science: not by giving faster answers to old questions, but by making previously invisible patterns measurable.
Ethan Mollick@emollick · 创始人 · 1 天前沃顿商学院教授,AI 应用研究权威
连环推 ×2

LLM有心理理论能力本身就是个大事件(当年发现的时候还挺有争议的),不过当模型需要考虑多个受众时,你还是能看出它的局限性——比如在编程时,它就很难区分终端用户和创作者各自的需求和视角。

查看英文原文
The fact that LLMs have theory of mind at all is a huge deal (and was controversial when discovered), but you can still see the limitations when models need to consider more than one audience - like how they struggle to separate end user & creator needs/perspectives when coding.
Work products often include information that only matters to the creator and is irrelevant or confusing to users, like references to previous drafts or problems solved in earlier iterations.

It is a persistent issue with working with advanced LLMs.
yihong0618@yihong0618 · 中文博主 · 1 天前

今天在 duckdb 的 foundation 页面又看到了时雨堂的 logo, 真的是值得尊敬的公司!

引用 yihong0618 @yihong0618发现了一个好厉害的日本公司,叫时雨堂。 厉害的不像是日本的互联网公司。不但一直跟着新技术(zig) 还特别拥抱开源,赞助了大部分他们主要使用的技术的开发者,还列到了官网上。 github.com/shiguredo查看被引原帖 ↗
Dan Shipper 📧@danshipper · 博主 · 1 天前

来最AI-pilled的公司工作吧:

引用 Every 🪨 @everyEvery公司正在招聘。开放职位包括Head of Growth、Office Manager/Executive Assistant、Senior Product Designer、Senior Video Producer等。查看被引原帖 ↗
查看英文原文
come work at the most AI-pilled company in the world:
向阳乔木@vista8 · 中文博主 · 1 天前向阳乔木,中文圈 AI 工具与趋势博主

这么技术范儿的skill都能到1w Star

可见相当好用,有大量口播传播。

烟花老师社群做的也很好,给我分享了很多干货,推荐关注。

引用 烟花老师 @teach_fireworks四个多月时间,我开源的这个小项目今天正式突破了 10000star🌟🎉! fireworks-tech-graph 是专门用于生成技术类图片的 skill,目前已支持 12 种风格,和 svg,png,gif 多种格式,会继续进行迭代更新 感谢大家的认可和支持🤟 一切都源于一个无聊的周末和对痛点的重新思考以及新模型上线后的好奇心 github.com/yizhiyanhua-ai/fi… 给小伙伴们准备了 2000 份现金红包🧧,输入支付宝口令即可领取查看被引原帖 ↗
歸藏(guizang.ai)@op7418 · 中文博主 · 1 天前歸藏,中文圈 AI 工具与提示词博主

终于搞定了!给 DeepSeek Harness 做了一个开箱即用的客户端 Pilot Harness。

这个客户端主要由两部分组成:
一个是客户端的壳;另一个是所有的优化,它们本质上都是 DeepSeek Harness 的插件。

也就是说,你既可以直接使用我这个客户端(默认加载所有插件),也可以只挑选对你有用的插件,安装到你自己的 DeepSeek Harness 里面。

这里面主要包括三个插件:

1. UI 和交互优化插件:优化了 DeepSeek Harness 本身一些有问题的展示方式、样式以及交互。

2. 文件树插件:你可以在聊天界面的右侧展开当前项目的文件树,快速将文件添加到聊天中,或者直接打开对应的文件。

3. 模型服务商和模型管理插件:优化了模型服务商的添加流程和模型管理的页面,将模型管理和服务商管理分开,更加直观,并且添加了很多图标和描述,界面更加清晰。

OpenRouter@openrouter · 公司官方 · 1 天前

@SakanaAILabs 的 Sakana Namazu 已在 OpenRouter 上线。

Sakana Namazu 基于 Kimi K2.6 打造,是个专业化模型,融合了对日本文化的深入理解和高性能推理能力,集成网络搜索和代码执行,可以处理商业环境中的复杂任务。

openrouter.ai/sakana/sakana-…

查看英文原文
Sakana Namazu by
@SakanaAILabs
is live on OpenRouter.

Built on Kimi K2.6, Sakana Namazu is a specialized model that combines a deep understanding of Japanese culture with high-performance reasoning capabilities, integrating web search and code execution to handle complex tasks within a business context.


openrouter.ai/sakana/sakana-…
Dan Shipper 📧@danshipper · 博主 · 1 天前

我的提议是,如果你一天花 10 亿 token,就得答个测验说清楚你造的是什么以及为什么。

没过的话会被列上羞耻墙。通过的话就能当一天的 Token Billionaire。

还在等 @ArielleShipper 给我回复这个提议...

引用 Every 🪨 @everyOpenAI 日均信用使用量增长230%。运营负责人 Arielle Shipper 采用三问题管理法而非严格预算:花了多少钱、买到了什么、学到了什么。查看被引原帖 ↗
查看英文原文
my proposal was if you spend a billion tokens in a day you have to answer a quiz about the details of what you built and why

if you fail you go on a wall of shame. if you pass you get to be a Token Billionaire for the day

still waiting for
@ArielleShipper
to get back to me on that one...
OpenRouter@openrouter · 公司官方 · 1 天前

模型定价变得越来越复杂了。这是 DeepSeek 峰时定价在你当地时区的运作方式:openrouter.ai/deepseek/deeps…。OpenRouter 在为你采购最优价格的时候会把这些都考虑进去。

查看英文原文
Model pricing is getting more complex over time.

Here's how DeepSeek peak-hour pricing works, in your local timezone:
openrouter.ai/deepseek/deeps…


OpenRouter takes all this into account when sourcing the best prices for you.
Lisan al Gaib@scaling01 · 博主 · 1 天前高频 AI 模型测评与爆料博主

太疯狂了,百万 token 在 10 美元或以下我才觉得舒服

引用 OpenRouter @OpenRouterOpenAI的GPT-5.6 Sol在OpenRouter上现已五折优惠,同时适用于batch API、flex和priority等层级。Flex层级费用最低仅需$1.25/千tokens输入和$7.50/千tokens输出。折扣自动应用。查看被引原帖 ↗
查看英文原文
insane

at or below $10/million tokens is where I feel comfortable
Tibor Blaho@btibor91 · 博主 · 1 天前逆向挖掘 AI 产品代码的爆料专家

在个人 ChatGPT 账户(包括 Free、Go、Plus 和 Pro 计划)上,创建和发布新的自定义 GPTs 已不再可用。

OpenAI 似乎也在开发一种方法将现有 GPTs 转换为 Skills。

查看英文原文
Creating and publishing new custom GPTs is no longer available on personal ChatGPT accounts (including Free, Go, Plus, and Pro plans)

OpenAI also appears to be working on a way to convert existing GPTs into Skills
宝玉@dotey · 中文博主 · 1 天前宝玉,中文圈 AI 翻译与科普大 V

Deepseek Harness 如果是 SDK 为主也能达到数据入口的效果吗?

> 像 ZCode 这样的 Coding 产品,不只是变现工具,也可能成为数据入口:用户完成真实的软件工程任务,产生长周期执行轨迹,再反哺下一轮 Post-training,形成“产品使用 → 数据 → 模型升级 → 更多使用”的飞轮。

引用 Anita AGI/acc @Anitahityou高盛的这篇智谱 GLM-5.3 的报告值得一看。 大模型竞争逻辑正在发生变化👉Scaling 正从一味堆参数,转向 Post-training。 GLM-5.3 参数规模基本没变,仍约 744B,但通过更长任务环境、更长训练时间、强化学习和工具协同,代码、长任务执行和安全能力继续提升。 模型进步开始越来越依赖“真实任务里的训练”,而不是单纯做更大的模型。 这背后还有一层更重要的商业逻辑。像 ZCode 这样的 Coding 产品,不只是变现工具,也可能成为数据入口:用户完成真实的软件工程任务,产生长周期执行轨迹,再反哺下一轮 Post-training,形成“产品使用 → 数据 → 模型升级 → 更多使用”的飞轮。 所以资本市场接下来真正要看的,已经不是 GLM-5.3 比 DeepSeek、Claude 高几个 Benchmark 点,而是三个转化率: 1️⃣ 模型能力 → 用户使用量 2️⃣ 用户使用量 → 付费收入 3️⃣ 真实任务数据 → 下一代模型能力 这也是为什么模型升级和股价逻辑要分开看。技术突破只能解释为什么产品更强,估值最终还是要靠收入增速、付费率和利润率兑现。 模型能力是起点,数据飞轮是护城河,商业化效率才是估值终点。查看被引原帖 ↗
◔ 3.2 万 次浏览(2 条合计)♥ 41⇄ 1观点看原帖 ↗
九原客@9hills · 中文博主 · 1 天前

这个情况我愿称之为 Harness 屎山。

包括最近有人提到Claude经常跨session对话等问题,当累积的 Harness 足够多(Prompt、Rule、Tools 等等),非预期的影响就更可能会出现。

Harness 也要 KISS。

引用 labuladong @labuladong_cn越来越看不懂 claude code 了,为啥 bypass permission 优先用 shell 脚本来改文件? 看过源码就知道,他那个 Read/Write/Edit 工具有并发修改检测,还有内容追踪和回滚的功能。用 shell 命令来改文件,完全无法追踪和恢复。 咋想的?查看被引原帖 ↗
Ethan Mollick@emollick · 创始人 · 1 天前沃顿商学院教授,AI 应用研究权威

还有方差问题!AI 实验室在做创意任务的时候需要更重视方差。从智能模型里要出创意变化特别费劲,这严重限制了它们在问题解决和创意工作中的实用价值。

引用 roon @tszzl因为我们关心的创意方面是在智能水平而非技术水平。查看被引原帖 ↗
查看英文原文
And variance! The AI labs need to be thinking more about variance when considering creative tasks. The fact that it takes work to get creative variation out of smart models severely limits their effective value in problem solving and creative work.
OpenRouter@openrouter · 公司官方 · 1 天前

折扣仅适用于 OpenAI 提供商,且仅在非 BYOK 请求时可用。促销持续至 9 月 18 日。感谢 @OpenAI !

查看英文原文
Discount applies only on the OpenAI provider, and is only available on non-BYOK requests. Promotion runs through September 18th.

Thanks to
@OpenAI
!
swyx@swyx · 博主 · 1 天前知名 AI 播客 Latent Space 主理人

Trajectory 一直都挺让我印象深刻的,做那么有野心的目标还能执行得很有品味。在我们 Continual Learning 这个板块里,@rronak_ 给了一个很深入的总览,讲他们怎么解决 CL 里剩下的核心数据问题,包括为什么 GRPO 还不够、他们不得不转向 on-policy……然后还得搞定随之而来的一堆问题。这个领域早期领跑者之一的讲解很精彩。

(想知道更多的话可以看看这个系列的其他内容,这一场阵容真挺强的。)

引用 Trajectory @trajectorylabs是时候重新思考强化学习了。将现实应用转化为模型改进需要重新设计后训练算法,以支持不可验证的逐token奖励。在World Fair分享扩展SDPO等连续学习算法的见解。查看被引原帖 ↗
查看英文原文
Trajectory have generally impressed me with their tasteful execution on ambitious goals. on our Continual Learning track,
@rronak_
gave a very thoughtful overview on how they're tackling the main data problems left in CL, including why GRPO isn't enough and they had to go on-policy.... and then subsequently fix all the issues that come up with it

nice overview from one of the early leaders in this field!

(see the rest of the track for more, this one was quite stacked)
Lisan al Gaib@scaling01 · 博主 · 1 天前高频 AI 模型测评与爆料博主

OpenRouter token 用量从 3.73T 激增到超过 75T,一年增长太疯狂了

查看英文原文
OpenRouter token usage is up from 3.73T to over 75T in one year
🚨 AI News | TestingCatalog@testingcatalog · 博主 · 1 天前专挖 AI 产品未发布新功能的爆料号

NOUNS RESEARCH 🔥:Hermes Desktop 新增了 Bot Mode,允许用户创建各种专业化的 AI Agents。

> "每个 Bot 都有自己的角色、模型、记忆、技能和头像;Bots 可以使用任何模型,甚至可以相互通信。"

这需要大量的测试时间 👀

引用 Nous Research @NousResearchHermes Desktop推出Bot Mode。Agent profiles变成系列命名Bot,每个Bot有独立角色、模型、记忆、技能和头像,可使用任何模型并相互通信。一次构建专业Bot,永久使用。查看被引原帖 ↗
查看英文原文
NOUNS RESEARCH 🔥: Hermes Desktop got a new Bot Mode, allowing users to create various specialized AI Agents.

> "Each Bot has its own role, model, memory, skills, and profile picture; Bots can use any model and even communicate with each other."

This needs a lot of testing time 👀
◔ 2.3 万 次浏览♥ 235⇄ 15▶ 含视频新品看原帖 ↗
Chubby♨️@kimmonismus · 博主 · 1 天前Chubby,高频 AI 新闻聚合博主

传闻中苹果带摄像头的 AirPods 在泄露演示里现身了。Siri 直接变成 AI 助理,能看到佩戴者眼前的一切。

这段片段据说是从 macOS Tahoe 26.7 候选版里挖出来的,展示了 Visual Intelligence 认出书本并保存待用。

苹果可能没有最强的语言模型,但有硬件护城河来做最适合用 AI 的产品。我超期待今年的苹果发布会。

不过欧盟肯定又会想办法不让美国这种带摄像头的 AirPods 上市。*叹气*

查看英文原文
Apple’s rumored camera-equipped AirPods have appeared in a leaked demo. Siri turns into an AI assistant that can see what the wearer sees.

The clip, reportedly found inside the macOS Tahoe 26.7 release candidate, shows Visual Intelligence recognizing a book and saving it for later.

Apple may not have the best language model, but they have the hardware moat for the best products to utilize AI. I'm really looking forward to this year's Apple event.

But surely the EU will find a way to prevent AirPods with cameras from being used in the EU. *sigh*
meng shao@shao__meng · 中文博主 · 1 天前

中国的职场,期权股权的坑实在太多了

打工人永远是弱势群体,希望发出来大家警醒,用好 AI 维护自己的应得权益

来自一个被期权坑过的中年人

我自己 😂

Yuchen Jin@Yuchenj_UW · 博主 · 1 天前

GitHub 早上给我掉线了。

Cursor 要推出另一个 GitHub。

时机有点意思啊。

引用 Cursor @cursor_ai我们的代码托管平台Origin正式上线。速度快、易使用、与Cursor深度集成。从GitHub同步仓库即可开始。查看被引原帖 ↗
查看英文原文
GitHub was down for me this morning.

Cursor is launching another GitHub.

Interesting timing.
◔ 2.2 万 次浏览♥ 229⇄ 6▶ 含视频观点看原帖 ↗
Chubby♨️@kimmonismus · 博主 · 1 天前Chubby,高频 AI 新闻聚合博主

企业级AI找到了第二种商业模式:卖平台,然后按小时出租那些让平台真正能用起来的人。

“前向部署工程师”是个好头衔。但它也描述了一个永无止境的成本项。你花钱买下能力,然后还要持续付费雇人把它转化成你的团队能真正跑起来的东西。


@xpander_ai
今天发布Omni,直接冲着这个痛点去的。一个内嵌的agent,负责构建、运营和优化你的其他agents,就在你团队已经在用的同一个平台上。

跑在你的云上或本地部署,支持任何模型,2000+工具和MCP连接器,每个用户都有权限管理,每个操作都有审计。能力完全留在内部。


chat.xpander.ai

引用 David 🍻 @dudutwizerxpander 融资 $7.5M 致力于民主化 AI 代理。提供云版本和自托管版本,帮助企业实现 AI 原生转型,避免供应商锁定,提供完整灵活性和治理,云版本赠送 1000 免费额度。查看被引原帖 ↗
查看英文原文
Enterprise AI figured out a second business model: sell the platform, then rent out the humans who make it usable, by the hour.

Forward-deployed engineer is a good title. It also describes a line item that never ends. You buy the capability, then keep paying people to translate it into something your team can actually run.


@xpander_ai
launches Omni today and goes straight at that. A built-in agent that builds, operates and optimises your other agents, on the same platform
your team already uses.

Runs in your cloud or on-prem, any model, 2,000+ tools and MCP connectors, permissioned per user and audited on every action. The capability stays in-house.


chat.xpander.ai
elvis@omarsar0 · 博主 · 1 天前

AI 智能体技能构建者必读论文。目前有 56,804 个公开 AI 智能体技能,都在竞争系统提示词中不到 100 个有效触发位。你的工作流也在竞争同一空间,而长尾技能根本没人用。skills 把安装包混在一起的三样东西分离了:内容、持久化、自动触发。只有触发需要放在提示词里。一个路径可以寻址任何技能、子树或集合,读一下就能用。目录变成菜单,捆绑包不再是要么全有要么全无。Vendoring 把技能复制到你的 Git 树相同路径,团队拥有并能定制它。没有清单、没有锁文件、无需注册,SKILL.md 保持不变。论文:arxiv.org/abs/2608.12610 在我们的学院追踪更多热门 AI 论文:academy.dair.ai/

查看英文原文
Recommended paper for agent skill builders.

There are 56,804 public agent skills today, all competing for fewer than 100 reliable trigger slots in the system prompt. Your own playbooks compete for that same space, and the long tail never gets used.

skills separates the three things installation bundles together. Content, persistence, and automatic triggering. Only triggering needs to sit in the prompt.

A path addresses any skill, subtree, or collection, and reading it is enough to use it. A directory becomes a menu, so bundles stop being all-or-nothing. Vendoring copies a skill into your Git tree at the same path, so your team owns and adapts it.

No manifest, no lockfile, no registration, and SKILL.md is unchanged.

Paper:
arxiv.org/abs/2608.12610


Track more trending AI papers in our academy:
academy.dair.ai/
九原客@9hills · 中文博主 · 1 天前

有几篇论文我也看过。

目前多 Agent 有两种形态,主子 Agent 和 Agent Team。后者我认为是一个很复杂的协作问题,就和当年分布式系统一样。

主子 Agent 本质没有协作,只是任务分发。它在总上下文超过了单 Agent 的有效上下文窗口时,是成功的。

Stability AI@StabilityAI · 公司官方 · 1 天前Stable Diffusion 图像模型开发商

今天想跟大家分享两个新工具,让你体验不同方式玩转 Stable Audio 3.0:一个是全新 Stable Audio 插件,直接把生成功能带进你心仪的 DAW;另一个是 StableAudio.com 上增强的网页体验,加入更多编辑和操作方式来处理你生成的内容。这两个都基于我们商用安全的模型,意味着你拥有输出,能自由分发。

目前两者都在beta阶段,某些功能还在试验中。我们会持续实时迭代优化这个体验。

引用 Stable Audio @stableaudioWe’re sharing two new ways for you to use Stable Audio 3.0. We now have a DAW plugin and a more advanced generation experience on StableAudio.com , built for the iterative process of audio production.查看被引原帖 ↗
查看英文原文
Today we wanted to share two new tools that give you different ways to work with Stable Audio 3.0: a new Stable Audio plugin that brings generation directly into your favorite DAW and an enhanced web experience on
StableAudio.com
, with more ways to edit and work with what you generate. Both are powered by our commercially-safe models, which means you own your outputs, and can distribute outputs freely.

These are both in beta, so some features are experimental. We will continue to iterate on the experience in real time.
Chubby♨️@kimmonismus · 博主 · 1 天前Chubby,高频 AI 新闻聚合博主

编程代理之所以好用,是因为有人围绕模型搭建了工作环境。

代码库访问、终端、测试、权限层。OpenCode、Amp 和 Pi 之所以有用,正是因为模型周围有这些脚手架支撑。

Questflow 在金融市场复制了这套架构,管它叫“金融操作台”。模型当大脑,技能组件传递的是真正的投资方法论而非泛泛而过时的老一套,插件接入实盘行情、链上和新闻数据,还有执行账户让你连接自己的交易所并设定操作限制。

你直接输入纯英文“买入那做多11美元,5倍杠杆,止盈9,止损7.6”,系统返回订单预览,包括数量、名义价值、所需保证金和两个出场点,外加确认按钮。在你按下确认之前,连一分钱都不会触及市场。

正是因为这套操作台,Questflow 代理在 Hyperliquid 和 Polymarket 上跑赢通用 AI 代理。同样的模型类别,却拥有了更多上下文和更强的执行力。

无等待名单,上线首日起对所有用户开放。前往
next.questflow.ai/?utm_sourc… 注册。

引用 Questflow @questflowAccess has been democratized. Financial intelligence has not. Questflow turns top investors’ market judgment into AI Agents that retail investors can follow, understand, and invest with across markets. Today marks the next era of Questflow: questflow.ai Financial Intelligence for All. Watch the introduction from our Co-founder & CEO, @Bobbxu .查看被引原帖 ↗
查看英文原文
Coding agents only got good once someone built the workspace around the model.

Repo access, a terminal, tests, a permission layer. OpenCode, Amp and Pi are useful because that scaffolding exists around the model.

Questflow built the same thing for markets and calls it the Financial Harness. The model as the brain, Skills that hand the agent a real investment methodology instead of a generic one, Plugins for live market, on-chain and news data, and execution accounts where you connect your own exchanges and set the limits it works inside.

You type "buy long 11 usd, 5x, tp 9, sl 7.6" in plain English and get back an order preview with size, notional, margin required and both exit levels, plus a confirm button. Nothing reaches the market until you press it.

That harness is why a Questflow agent outperforms a general-purpose AI agent on Hyperliquid and Polymarket. Same class of model, far more context and far more reach to act on it.

No waitlist, open to everyone from day one. Sign up at
next.questflow.ai/?utm_sourc…
Bindu Reddy@bindureddy · 创始人 · 1 天前Abacus.AI CEO,AI 行业观点博主

根据你的使用场景选择最优的AI

hard-coding - Fable 5
数据分析 - GPT 5 Sol
研究 - Flash 3.7
视频 - SeeDance 2.5
设计 - Opus 5
廉价agentic - Kimi K3
实时 - Grok 4.6
图像 - GPT Image 2
分类器 - Qwen 3.8 27B

根据你的使用场景自动路由到最优的模型

查看英文原文
Use The Top AI For Your Use Case

hard-coding - Fable 5
data analysis - GPT 5 Sol
research - Flash 3.7
video - SeeDance 2.5
design - Opus 5
cheap agentic - Kimi K3
real-time - Grok 4.6
image - GPT Image 2
classifier - Qwen 3.8 27B

Automatically route to the best model for your use case
Gary Marcus@GaryMarcus · 博主 · 1 天前

AI 隐性成本源源不断。先是芯片、非共识深假色情和高等教育及中学教育的瓦解,现在又加上利率。

引用 Hedgie @HedgieMarkets🦔AI companies have borrowed so much money this year that they're pushing up interest rates for the entire economy. Nomura estimates tech borrowing alone now equals 25% of what the US Treasury issues in bonds, five times more than last year. Bank of America says the surge has added about 0.3 percentage points to the 10-year yield. Bond managers are selling Treasuries to buy AI corporate debt instead because it pays more. My Take AI companies are now competing with the US government for the same pool of lenders, and the lenders are picking the corporate bonds. Alphabet's 30-year pays 6.4%. A Meta data center bond pays over 7.5%. At those rates, a 5.2% Treasury loses the fight for capital every time. That's one of the reasons long-term rates have stayed so stubborn even as the Fed tries to bring them down. JPMorgan expects $5.5 trillion in AI infrastructure spending through 2030, and most of it will be borrowed. That borrowing raises the cost of money for everyone, the government, your mortgage, small businesses trying to get a loan. The AI buildout has reached the scale where it moves rates for the whole economy, and most people paying higher borrowing costs have no idea that a data center arms race is one of the reasons why. Hedgie🤗查看被引原帖 ↗
查看英文原文
The hidden costs of AI never end. First it was memory chips and nonconsensual deepfake porn and the undermining of college and high school education; now it is interest rates.
小互@xiaohu · 中文博主 · 1 天前小互,中文圈高频 AI 资讯站 Xiaohu.AI 主理人

Anthropic 这个增速有点夸张了

根据Anthropic 向潜在投资者披露的初步数据显示:

2026 Q2 其营收已超过 115 亿美元
去年同期只有 7.87 亿美元
一年做到约 14.6 倍

更猛的是年化营收 Run Rate:

2025 年底:90 亿美元
2026 年 5 月:470 亿美元
2026 年 7 月底:650+ 亿美元

而且 Q2 还首次实现了 调整后营业利润转正

根据预测到2028 年其营收将会达到:
1900 亿至 2000 亿美元

甚至有投资人认为,如果这个增长速度能维持,未来估值可能冲到 2 万亿美元。

从 90 亿到 650+ 亿美元年化营收,只用了不到 8 个月

◔ 2 万 次浏览(3 条合计)♥ 42⇄ 6动态看原帖 ↗
Gary Marcus@GaryMarcus · 博主 · 1 天前

这是不回答合理问题的教科书式例子。完全符合一贯风格。

引用 Kate Rooney @Kr00neyGreg Brockman在CNBC解释OpenAI高管离职都是战略聚焦的结果,包括Fidji Simo、CRO和Brad Lightcap等。公司致力统一消费者和企业业务,7月收入环比增长20%,企业类增长32%,展现韧性。查看被引原帖 ↗
查看英文原文
textbook example of not answering legit questions. totally on-brand.
Together AI@togethercompute · 公司官方 · 1 天前

影子流量能证明一个候选方案在运营上是靠谱的。但它告诉不了你用户是否更喜欢它。A/B 测试应该在端点做,不是在你应用代码里。客户端用相同的端点名称、API 和密钥。没有功能标志,没有客户端代码里的 hash-mod-100,没有电子表格来解释 A 组和 B 组是啥。把一个实时端点的流量分成一个控制和最多 20 个变体,每个都有固定的百分比。一个调用就能扩展。删除实验,100% 的流量就回到控制,没有什么需要清理的。完整教程看这里:together.ai/blog/a-b-test-mo…

查看英文原文
Shadow traffic proves a candidate is operationally sound. It can't tell you if users like it better.

A/B testing belongs at the endpoint, not in your app code. Same endpoint name, API, and keys for your clients. No feature flags, no hash-mod-100 in client code, no spreadsheet explaining what group A vs B means.

Split a live endpoint's traffic into one control and up to 20 variants, each with a fixed percentage. Ramp with a single call. Delete the experiment and 100% of traffic returns to the control, with nothing left to unwind.

Read the full walkthrough:
together.ai/blog/a-b-test-mo…
Lisan al Gaib@scaling01 · 博主 · 1 天前高频 AI 模型测评与爆料博主

这个图表其实显示的不是它说的那样

主要就是 benchmaxxing,缺乏硬基准来区分最佳模型

引用 POM @peteromIf trends continue, we will have “Fable at Home” (~30B models w/ similar capability) sometime between January and May next year查看被引原帖 ↗
查看英文原文
for anyone wondering, this plot is not showing what it says it shows

it's mostly benchmaxxing and lack of hard benchmarks to distinguish the best models
elvis@omarsar0 · 博主 · 1 天前

推荐资源。这是我见过最大的 agent 技能数据集。很适合挖掘 agent 的创意和模式。

引用 DAIR.AI @dair_aiAnthropic开放SKILL.md格式9月后,282,200个GitHub公开仓库已有380万SKILL.md文件。因自然语言编写且代理概率选择缺乏验证,GitSkills将其分组为187万个不同内容,以SQLite形式发布。查看被引原帖 ↗
查看英文原文
Recommended resource. Largest dataset of agent skills I’ve come across. Great for mining cool ideas and patterns for your agents.
Orange AI@oran_ge · 中文博主 · 1 天前Orange AI,中文圈 AI 产品观察博主

牛来的风还是吹到了 AI 圈,今天发现牛来都有 skill 了,就当作这周的 skill 精选吧。。。。。。

1. 将任意图片变牛头,太丑了

go.colaskill.com/niutou


2. 将任意图片变为牛来画风

go.colaskill.com/nl-pic


3. 一键生成牛来风格动画片,太魔性了

go.colaskill.com/nlv

这个重点说一下,可以将一句普通描述、一段日常故事或一张图片,改编成一部牛来脸主演的 15 秒荒诞低成本国产动画电影预告片。

Gary Marcus@GaryMarcus · 博主 · 1 天前

很多人对这个问题视而不见——AI 无法遵循指令——就像他们最初对幻觉那么漠不关心,同样抱着一个幻想,以为这些核心的、顽固的问题会快速解决。实际上这些问题不会真的被解决,除非我们有根本上不同的架构。这就是为什么以 LLM 为中心的架构对 AI 安全构成深刻的威胁。

引用 Huw Ringer @HuwRinger给AI指令越具体,它们越不可能遵循指令。查看被引原帖 ↗
查看英文原文
Many people’s blindness to this problem — AI failing to follow instructions — is like their initially blasé attitude towards hallucinations, with the same fantasy that core, endemic problems would rapidly be solved.

These problems won’t actually be solved, until we have fundamentally new architectures.

Which is why LLM-centered architectures are a profound threat to AI safety.
Gary Marcus@GaryMarcus · 博主 · 1 天前

有人说狗像它们的主人。生成式 AI 的 CEO 听起来也越来越像他们的产品:经常错得离谱,却从不自我怀疑,满嘴胡言乱语。

引用 George Noble @gnoble79Dario满口谎言。物以类聚。拒绝@DarioAmodei @sama @elonmusk查看被引原帖 ↗
查看英文原文
Some say dogs look like their owners.

Generative AI CEOs have come to sound like their product: frequently wrong, never in doubt, often full of BS.
Gorden Sun@Gorden_Sun · 中文博主 · 1 天前中文圈高频 AI 资讯与开源项目博主

Grok Bot非常好用,也绝对是未来的Agent标配,桌面客户端同时可以控制本地Agent和云端电脑,也能做接力机制。
但是云端电脑更新会重置系统盘,这就有点伤了,会导致云端装的服务都挂了。

目前我的使用场景:在Grok云电脑里装Cursor、Codex,结合飞书CLI,让其他人进我的飞书组织,直接通过聊天使用Cursor和Codex。

基于Agent的服务,也可以通过Cloudflared包装成网站,用户在网站提交任务,例如把图片PPT转为可编辑的PPTX,云电脑的Codex执行任务,完成后把结果返回网站。可以实现Skill不外泄,做成收费服务。

DeepLearning.AI@DeepLearningAI · 公司官方 · 1 天前

🚀 我们正在招人:市场营销工程师(加州山景城)

我们需要一位 AI 原生开发者来构建代理工作流、自动化流程和相关工具,帮助我们的营销团队规模化运作。和我们的 AI 工程团队一起亲自动手干!🤖

完整详情及申请链接:(
hubs.la/Q04tcMkw0
)。


#AI
#Hiring
#TechJobs
#MarketingEngineer
#DeepLearningAI

查看英文原文
🚀 WE ARE HIRING: Marketing Engineer (Mountain View, CA)

We need an AI-native dev to build agentic workflows, automations, and tooling to help our marketing team operate at scale. Work hands-on with our AI engineering team! 🤖

Full details & apply here: (
hubs.la/Q04tcMkw0
).


#AI
#Hiring
#TechJobs
#MarketingEngineer
#DeepLearningAI
AshutoshShrivastava@ai_for_success · 博主 · 1 天前高频 AI 新闻与产品动态博主

Gemini Omni Flash,这是用 Gemini Avatar 功能做的文字转视频

查看英文原文
Gemini Omni Flash. This is text to video using Gemini Avatar feature.
◔ 1.2 万 次浏览♥ 132⇄ 5▶ 含视频演示看原帖 ↗
yihong0618@yihong0618 · 中文博主 · 1 天前

牛逼啊!

引用 Jerome.Y. @alterxyz4你好你好👋 我其实是 Dify 的 Head Of Product 了。 请多多指教~查看被引原帖 ↗
🚨 AI News | TestingCatalog@testingcatalog · 博主 · 1 天前专挖 AI 产品未发布新功能的爆料号
连环推 ×2

Vorflux 推出了新的云平台,让用户可以在专用机器上运行从计划到合并 PR 的整个流程。

它可以自己搞定规划、构建、测试和审查。
> 差异审查默认会用不同的模型族来做。
> 第一次启动会保存成快照,之后的会话都从这个快照启动。
> 浏览器代理会走一遍真实用户流程,视频录制包含在拉取请求里。

视频证明在 PR 里 👀

查看英文原文
Vorflux has opened its new cloud platform, allowing users to run tasks from a plan to a merged PR on a dedicated machine.

It can plan, build, test, and review on its own.
> Diff reviews are done by a different model family than the one that wrote it by default.
> First boot is saved as a snapshot, with all subsequent sessions waking up from that same starting point.
> A browser agent walks through the real user flow, and the recording is included in the pull request.

Video proof in the PR 👀
Vorflux can intelligently route each stage of the process to the frontier model best suited to that task, across Claude, GPT, Gemini, and open-weight options.

Compute auto-stops after an hour of idle. BYOK, Vorflux tokens, or an existing Codex plan.

Test it out 👀

devtoolsacademy.link/tc-s
The Rundown AI@TheRundownAI · 博主 · 1 天前百万订阅 AI 日报官方
连环推 ×7

我们在 2 周前向 200 万+ 读者推出了社区 AI Workflow Hub。

我们收到了数千份令人难以置信的投稿,展示人们如何利用 AI 改进生活、工作或业务。

以下是过去一周的前 5 名:

查看英文原文
It's been 2 weeks since we launched our community AI Workflow Hub to 2M+ readers.

We've had thousands of incredible submissions on ways people are using AI to better their lives, work, or business.

Here are the top 5 from the last week:
The family command center

Erica built a custom family calendar with Lovable and Claude instead of buying a $300 Skylight.

It pulls in everyone's Google Calendars, Todoist tasks, and rotating family photos, running on an iPad in the kitchen.

Credit to Erica Conti:
app.therundown.ai/community/…
The Al triathlon coach

Tom wired his sleep, HRV, and training data into ChatGPT Work, with a scheduled agent that checks what his body actually did every morning.

Credit to Tom Tomaszewski (
@t_tms
):
app.therundown.ai/community/…
Tony built a reading generator that interviews him about his mood in 2-4 questions, then writes a 1,000-2,000 word piece just for that moment.

Credit to Tony Ojeda (
@tonyojeda3
):
app.therundown.ai/community/…
The AI public speaking coach

Patrick built a ChatGPT tool for his public speaking students that won't write speeches or find sources, it only asks questions.

When a student takes a side, it pulls counterarguments from thinkers on their own side.

Credit to Patrick Loebs:
app.therundown.ai/community/…
The surgeon who ships

Alex built and shipped a free iOS app for surgeon health.

The app uses a MEMORY.md file so AI remembers the project, and a separate AI auditor session that reviews the code.

Credit to Alex Reid:
app.therundown.ai/community/…
The $0 textbook

A business professor built a free 11-chapter AI textbook with Claude and GitHub Pages, featuring real cases from Netflix, Uber, and Waymo, hands-on Python labs, and updates that ship like software instead of waiting on an 18-month publishing cycle.

Credit to Chris Califf:
app.therundown.ai/community/…
elvis@omarsar0 · 博主 · 1 天前

Grok Bot 真的太棒了!最近 Grok 模型的这些更新也让我印象深刻。

如果想试试 Codex 和 Claude Code 以外的产品,Grok Bot 值得一试。

引用 Gavin Baker @GavinSBaker认为 @bot 是 AI 领域的又一个'Claude Code'时刻。个人 AI 使用量增加约 100 倍。播客摘要工具在 Grok Bot 中 15 秒即可完成,效果远超之前。查看被引原帖 ↗
查看英文原文
Grok Bot is so good! I am very impressed with all the recent updates on Grok models as well.

If you ever wanted to try a different product outside Codex and Claude Code, Grok Bot is worth trying.
elvis@omarsar0 · 博主 · 1 天前

说实话,代码 LLM 在解锁通用能力方面展现出的强大能力确实有点出人意料。扩散模型很快会有它的时刻,但我预计 LLM 会继续改进并超越。

引用 Thariq @trq212最近的程序生成艺术、视频编辑和3D游戏演示让我对LLM编码模型有了新认识——在许多创意工作上,它们优于扩散模型。查看被引原帖 ↗
查看英文原文
It has been a bit surprising, to say the least, how powerful LLMs for code have turned out to be in unlocking general capabilities.

Diffusion models are going to have their moment soon, but I expect LLMs to continue to improve and transcend.

本站由 Jedee杰哥 打造 · 公众号「Jedee杰哥」每早送 AI 日报

姊妹站:𝕏 简中账号数据榜单 · X 关注 @jedeeai · RSS 订阅 · AI 日报 · 历史归档