JEDEE AI
存档 2026-08-20

8 月 20 日(北京时间)全球 AI 圈推文存档,按曝光排序,共 100 条。
← 返回最新 全部归档

全部情报 每小时更新 · 事件已合并同类项

内容 公司
Google Gemini@GeminiApp · 公司官方 · 1 天前谷歌 Gemini 产品官方
连环推 ×6

开学季到了,从今天起,符合条件的大学生可以免费使用整整一年的 Gemini:
- 美国学生:免费获得 1 年 Google AI Pro
- 140+ 个国家:免费获得 1 年 Google AI Plus

这是给学生的新功能 👇

查看英文原文
Back-to-school season is here and starting today, eligible college students can get a full year of Gemini on us:
- US students: 1 year of Google AI Pro at no cost
- 140+ countries: 1 year of Google AI Plus at no cost

Here’s what’s new for students 👇
Get organized with the new student hub.

We’ve built a dedicated hub just for students. Start a study notebook, create flashcards, take a practice quiz and more. As we roll out new learning tools, you’ll find them in the student hub.

Access the new student hub here:
gemini.google.com/students
Use study notebooks for an adaptive learning environment.

When looking for more personalized study help, Gemini steps in.
→ Pinpoints your exact knowledge gaps with diagnostic quizzes
→ Builds bite-sized lessons grounded in your class materials
→ Tracks your progress on a live dashboard

In the coming weeks, study notebooks will be even more helpful with graphs and images as part of your lessons.
Dive deeper with interactive visualizations.

Gemini can now generate functional, interactive 3D simulations and dynamic tables directly in your chat.

Just start your prompt with “Show me…” and use our Flash model for the best results:
+ “Show me how DNA works in 3D” (zoom, rotate, explore)
+ “Show me how energy works on a pendulum”
+ “Show me cash burn rate as a waterfall”
Talk through research with Gemini Live.

You can now kick off multi-step research reports with your voice:
1) Ask Gemini Live to Deep Research a complex topic
2) Lock your phone or switch tasks while it works in the background
3) Get notified when it's done, then seamlessly talk through the findings hands-free
The student hub, study notebooks, interactive visualizations, and Deep Research in Gemini Live are rolling out today to eligible Gemini app users globally.

Claim your student plan for one year at no cost. Offer runs through December 31st, 2026 for eligible students. Terms apply.

Learn more here:
goo.gle/student
◔ 1906.8 万 次浏览♥ 2.9 万⇄ 2,982▶ 含视频新品看原帖 ↗
Min Choi@minchoi · 博主 · 1 天前AI 产品演示博主,专门展示新工具玩法
连环推 ×11

Grok Bot 越来越牛了。

人们正在用手机控制机器人、发布游戏、购买域名 + 运行整个代码库。

10 个疯狂的例子:

查看英文原文
Grok Bot is getting ridiculous.

People are controlling robots, shipping games, buying domains + running entire repos from their phone.

10 wild examples:
1. Remote lawn mower from 50 MILES away.

Start it.
Mow.
Dock it.

All through Grok Bot.
2. Text your robot vacuum from anywhere.
3. 74 game-art assets

Generated.
Cropped.
Organized.
Wired into the actual game.

~2 hours.
4. Grok Bot built an Arduino LED stock ticker.

Live
$SPCX
price.
Graph.
SpaceX news.
Physical AI dashboard 🤯
5. You can SSH Grok Bot into your Mac.

Then let it run local scripts + AppleScript on your own computer.
6. One Grok Bot becomes the boss of your repo.

Then it recruits MORE agents to get the work done.

Two prompts.
7. Someone told Grok Bot to install Commander Keen on its cloud computer...

Then it actually played the game 😂
8. Bee Pioneer pin + a Grok Bot that becomes a nightly Relationship Coach.

Your AI agent follows up while you sleep.
9. Grok Bot completed Google's:

"I’m not a robot."
10. From a PHONE:

Build the website.
Buy the domain.
Deploy it.
Convex + Cloudflare.

Two prompts.
◔ 267.4 万 次浏览♥ 1,025⇄ 72▶ 含视频演示看原帖 ↗
Grok@grok · 公司官方 · 1 天前马斯克 xAI 旗下聊天机器人 Grok 官方

现在,可以管理在 Grok 中构建和部署的应用的共享和访问权限了。

应用可以设为仅自己可见、与团队共享,或公开给全世界。你可以从 Grok 网页端、iOS、Android 或 CLI 开始构建。

查看英文原文
Now, manage sharing and access for apps built and deployed in Grok.

Make apps private to you, shared with your team, or public to the world. Start building on Grok web, iOS, Android, or in the CLI.
OpenAI@OpenAI · 公司官方 · 1 天前ChatGPT 开发商官方账号

我们将继续为前沿模型提供零数据保留。

随着 AI 承担更长期、更自主的工作并为企业创造更大价值,安全系统也需要识别相关交互中的风险。

为了帮助应对这些风险,我们正在预览私人安全处理,这项功能旨在改进安全性,同时不让 OpenAI 人员访问底层内容。

查看英文原文
We will continue to offer Zero Data Retention for frontier models.

As AI takes on longer, more autonomous work and delivers greater value to businesses, safety systems also need to identify risks across related interactions.

To help address those risks, we're previewing Private Safety Processing, which is designed to improve safety without giving OpenAI personnel access to the underlying content.
◔ 205.4 万 次浏览(2 条合计)♥ 4,247⇄ 286新品看原帖 ↗
OpenAI Developers@OpenAIDevs · 公司官方 · 1 天前OpenAI 开发者平台官方

借助 Codex,@asana 在两个日历周内完成了前端测试框架从 Enzyme 到 React Testing Library 的迁移——一个原本预期需要花费五年时间的项目。

openai.com/index/asana/

查看英文原文
With Codex,
@asana
finished a frontend test migration from Enzyme to React Testing Library in two calendar weeks—a project expected to take five more years.


openai.com/index/asana/
el.cine@EHuanglu · 博主 · 1 天前

AI 机器人的碰撞测试

查看英文原文
crash test but for AI robot
◔ 94.9 万 次浏览♥ 2,553⇄ 246▶ 含视频其他看原帖 ↗
OpenRouter@openrouter · 公司官方 · 1 天前
连环推 ×2

OpenRouter 要加入 Stripe 了。

我们当初创办 OpenRouter 就一个理念:智能应该是多模型的。现在,我们已经是最大的 AI 市场和网关,每天处理 10T+ tokens,支持 400+ 模型。

加入 @Stripe 让我们能够加速推进这个使命。

查看英文原文
OpenRouter is joining Stripe.

We started OpenRouter with a simple mission: intelligence should be multi-model. Today, we are the largest AI marketplace & gateway, processing 10T+ tokens daily on 400+ models.

Joining
@Stripe
gives us the opportunity to accelerate that mission.
OpenRouter will continue as OpenRouter:

Same name, same product, same roadmap, and the same commitment to helping you find and use the best model for every task.
◔ 63.5 万 次浏览(4 条合计)♥ 2,626⇄ 233动态看原帖 ↗
Tibo@thsottiaux · 公司官方 · 1 天前ChatGPT 产品官方账号

还没投入使用,但你看看这个。Codex 就是为大规模而生

查看英文原文
It has not been used yet, but would you look at that. Codex for scale.
Cursor@cursor_ai · 公司官方 · 1 天前最火的 AI 编程工具 Cursor
连环推 ×3

我们在Cursor里持续改进云端agents。

它们从事件中接手任务,持续朝着目标迈进,直到达标为止,并在长时间会话中始终不走偏。

查看英文原文
We're continuing to improve cloud agents in Cursor.

They pick up work from events, hold a goal until it's met, and stay on course through long sessions.
Use any skill as a Custom Mode: a skill that stays pinned in the chat. You can think about it like always-on skills.

From /, pick a skill and press ⌥⏎ (Mac) or Alt+Enter (Windows), or choose Use as Mode.
Subagents can now run on their own virtual machines, each with an isolated copy of the project.

Have them test the parent agent's changes in a fresh environment, or swarm independent fixes.
◔ 50.7 万 次浏览♥ 3,082⇄ 186▶ 含视频新品看原帖 ↗
swyx@swyx · 博主 · 1 天前知名 AI 播客 Latent Space 主理人

哈哈哈哈哈,非技术人员真的彻底完蛋了,都烧焦了啊。你能想象零背景、零推理、脑子里没有任何内在世界模型的人去报道 AI 吗?那肯定爽翻了,什么都那么amazing,表面价值就是你需要的全部,真是活在最好的时代啊。

引用 OpenAI Developers @OpenAIDevsWith Codex, @asana finished a frontend test migration from Enzyme to React Testing Library in two calendar weeks—a project expected to take five more years. openai.com/index/asana/查看被引原帖 ↗
查看英文原文
HAHAHAHAHAHA

non technical people are so incredibly cooked they are burnt thru

can u imagine covering ai with zero context, zero reasoning, zero internal world model. it must be so delightfully joyful, everything is so amazing, face value is all you need,

what a time to be alive
Sam Altman@sama · 创始人 · 1 天前Sam Altman,OpenAI 联合创始人兼 CEO

我们支持商业隐私!

openai.com/index/offering-ze…

查看英文原文
we support business privacy!


openai.com/index/offering-ze…
OpenAI Developers@OpenAIDevs · 公司官方 · 1 天前OpenAI 开发者平台官方
连环推 ×2

团队正在用开源的 Codex harness 把 agent 集成到他们已有的工具里,从内部应用再到运维看板都有。

这些应用的客户端掌控着界面、上下文、工具和审批流程,而 harness 则负责处理 agent 循环逻辑。


developers.openai.com/blog/c…

查看英文原文
Teams are using the open-source Codex harness to bring agents into the tools they already use, from internal apps to operations dashboards.

Their applications control the interface, context, tools, and approvals while the harness handles the agent loop.


developers.openai.com/blog/c…
The Codex SDK powers
@Cisco
s App Builder, where customers create custom applications for Cisco Cloud Control using natural language.

We also worked with
@ThriveHoldings
and Crete on a tax-prep system that processed 7,000 returns and cut preparation time by about a third.
Sundar Pichai@sundarpichai · 创始人 · 1 天前谷歌 CEO
连环推 ×2

填空:Search 现在可以帮助准备 _________。
A. SAT
B. ACT
C. AP
D. ENEM
E. GRE
F. JEE
G. LSAT
H. MCAT
I. NEET
J. 以上所有!

查看英文原文
Fill in the blank: Search can now help with test prep for _________.

A. SAT
B. ACT
C. AP
D. ENEM
E. GRE
F. JEE
G. LSAT
H. MCAT
I. NEET
J. All of the above!
(It’s J BTW :) See more of the learning features we’re dropping today, plus free Google AI subscriptions for college students in 140+ countries!
blog.google/products-and-pla…
Lisan al Gaib@scaling01 · 博主 · 1 天前高频 AI 模型测评与爆料博主

新摩尔定律

token 数量大约每 11 周翻一番

引用 Deedy @deedydasOpenRouter被Stripe收购,创立3年内完成此规模收购创造历史。作为双边模型市场,OpenRouter为模型实验室提供容量和分发的可预测性,为用户提供访问所有模型的简便方式。其tokens增长30000倍,周运行量超4.5万亿,连续3年月增长33%。创始人坚持专注产品,无炒作融资。查看被引原帖 ↗
查看英文原文
New Moore's Law

token volume is doubling every ~11 weeks
Greg Brockman@gdb · 创始人 · 1 天前Greg Brockman,OpenAI 联合创始人兼总裁

Codex 能干的可不止编码呀,试过用来报税,7000 份税表,准备时间减少了约三分之一。

你也可以基于我们开源的 Codex 框架开发自己的产品:

引用 OpenAI Developers @OpenAIDevs团队使用开源Codex harness将agents集成到现有工具中,从内部应用到操作仪表板。应用程序控制界面、上下文、工具和审核,harness处理agent循环。查看被引原帖 ↗
查看英文原文
codex can power much more than coding tools: a tax-prep pilot processed 7,000 returns and cut preparation time by about a third.

you can build your own products on our codex open-source harness:
@levelsio@levelsio · 博主 · 1 天前独立开发者标杆,AI 产品连续创业者

@lulumeservey 说得对:'现在什么都是假的。假内容、假网红、假互动,推假产品贴假测评。全球范围的骗局,闪亮表皮下堆满乏味平庸。抖音版的拼多多。互联网上有真实想法很珍贵。'

引用 Zach Moskow @zachmoskowJust so everyone knows...we got a quote for a “launch video” package from the company everyone uses. $17K for the video. +$25K and they hand you 50 influencers who will push it, repost it, and flood the comments so the viewcount goes above 500k. That’s the exact recipe behind almost every “viral” slop launch video.查看被引原帖 ↗
查看英文原文
Like
@lulumeservey
says

"Everything is fake now. Fake content by fake influencers with fake engagement from fake followers, launching fake products with fake testimonials. A Potemkin village on a global scale, a glossy facade propped up by heaps and heaps of beige slop. Temu for content."

Being real on the internet is special!
@levelsio@levelsio · 博主 · 1 天前独立开发者标杆,AI 产品连续创业者

很容易分辨 GTA 6 泄露是否真实

摄像机运动

AI 视频总是搞砸摄像机运动,让它们看起来太随机了

引用 Ramon A. Lage 🥊 @Ramon4lanAbastecer o carro, fazer compras dentro de um supermercado, viajar de trem... O nível de detalhes no supermercado me surpreende demais. Espero que seja realmente assim o jogo.查看被引原帖 ↗
查看英文原文
It's very easy to detect when a GTA 6 leak is real or not

The camera movements

AI videos keep messing up the camera movements and make them too random
◔ 19.9 万 次浏览♥ 465⇄ 9▶ 含视频观点看原帖 ↗
el.cine@EHuanglu · 博主 · 1 天前

最快的 AI 机器人刚中风了

引用 el.cine @EHuanglucrash test but for AI robot查看被引原帖 ↗
查看英文原文
the fastest AI robot just had a stroke
◔ 19.1 万 次浏览♥ 2,123⇄ 216▶ 含视频其他看原帖 ↗
OpenAI Developers@OpenAIDevs · 公司官方 · 1 天前OpenAI 开发者平台官方

由于技术问题,部分用户对Daybreak Blue的访问权限已失效,需要重新验证才能继续使用。

这不是我们想提供的体验,我们正在确保类似问题不会再次发生。

受影响的用户已通过邮件收到验证指引,团队也随时待命协助。如果你的访问没有中断,现在无需任何操作。

提醒一下,所有用户需在9月1日前启用Advanced Account Security以维持Daybreak访问权限。

查看英文原文
Due to a technical issue, a limited set of users’ access to Daybreak Blue is no longer active and they will need to re-verify to maintain their access.

This isn’t the experience we want to deliver, and we’re making sure that the issue doesn’t repeat in the future.

Affected users have been notified by email with a path for verification and our teams are on standby to help. If your access is uninterrupted, you don’t need to do anything right now.

As a reminder, all users will be required to have Advanced Account Security by September 1 to maintain Daybreak access.
Alexandr Wang@alexandr_wang · 创始人 · 1 天前Scale AI 创始人,Meta 超级智能实验室负责人

快试试 Meta AI 桌面应用吧!

听写功能对我个人来说真的改变了一切

引用 Spencer Barnett @spencerbarnettWe launched the Meta AI Mac OS app today! 🚀 I particularly love the dictation feature which allows me to dictate anywhere on my computer with ultra high accuracy. Just hold down 'fn' and yap!查看被引原帖 ↗
查看英文原文
check out the Meta AI desktop app!

the dictation has been a game changer for me personally
◔ 19.8 万 次浏览(2 条合计)♥ 1,271⇄ 77▶ 含视频演示看原帖 ↗
Logan Kilpatrick@OfficialLoganK · 创始人 · 1 天前谷歌 Gemini 产品负责人

你觉得全球AI token支出中有多少百分比是用来跑evals的?

查看英文原文
What % of global AI token spend do you think is people running evals?
Qwen@Alibaba_Qwen · 公司官方 · 1 天前阿里通义千问大模型团队

非常感谢 Unsloth 的优秀工作。这对社区来说是个好消息!🥳Qwen3.8-27B,比以往更小更锐利。让我们试试吧!
@UnslothAI

引用 Unsloth AI @UnslothAIWe’re releasing new Qwen3.8-27B GGUFs with 10% higher accuracy. Unsloth Dynamic V3 outperforms others by >10% on Div-300, KLD & more benchmarks. We also release 1-bit quants that retain 77% accuracy. Run on 8GB RAM. Blog: unsloth.ai/docs/basics/dynam… GGUF: huggingface.co/unsloth/Qwen3…查看被引原帖 ↗
查看英文原文
Huge thanks to Unsloth for the great work. This is wonderful news for the community! 🥳Qwen3.8-27B, smaller and sharper than ever. Let's try it!
@UnslothAI
Tanishq Mathew Abraham, Ph.D.@iScienceLuvr · 博主 · 1 天前

所以大家都在讨论这个,但你们知道吗这其实不是 LLM/基础模型甚至都不是神经网络?可能就是某种 XGBoost 或逻辑回归算法,基于一堆手工特征

引用 Polymarket @Polymarket刚刚:AI在设计Moderna个性化皮肤癌疫苗中发挥关键作用,该疫苗在III期试验中成功。查看被引原帖 ↗
查看英文原文
okay so everyone is talking about this but you guys realize it's not an LLM/foundation model or even a neural network, right????

it's probably just like some XGBoost or logistic regression algorithm over a bunch of handcrafted features
Guillermo Rauch@rauchg · 创始人 · 1 天前Guillermo Rauch,Vercel 创始人兼 CEO

fx.sh 只有 6.3mb。它启动只需 10µs¹。它是用 Zig 编译的静态 ELF 二进制文件,或者更小的 𝚕𝚒𝚋𝚏𝚡.𝚠𝚊𝚜𝚖²

AI 将使大多数基础设施本身得到优化。fx 就是例子:它能比其他 agent 启动还快完成任务。

快速是单向的。

¹ 微秒
² 有趣的是:wasm 版本实际上更小。大量二进制数据来自 TLS/HTTP 协议栈。在 WebAssembly 中,fx 将 𝚏𝚎𝚝𝚌𝚑() 委托给 JS 运行时。你可以在 fx.sh/try 试试

查看英文原文
fx.sh
is 6.3mb. It starts up in 10µs¹. It's a Zig-compiled static ELF binary, or an even smaller 𝚕𝚒𝚋𝚏𝚡.𝚠𝚊𝚜𝚖²

AI will make most infrastructure natively optimized. Case in point: fx can literally finish tasks faster than other agents can boot.

Fast is a one-way street.

¹ that's microseconds
² fun fact: the wasm build is actually *smaller*. A lot of binary weight comes from the TLS/HTTP stack. In WebAssembly, fx delegates 𝚏𝚎𝚝𝚌𝚑() to the JS runtime instead. You can try this on
fx.sh/try
Greg Brockman@gdb · 创始人 · 1 天前Greg Brockman,OpenAI 联合创始人兼总裁

Codex 做代码迁移——从耗时多年到只需几周:

引用 OpenAI Developers @OpenAIDevsWith Codex, @asana finished a frontend test migration from Enzyme to React Testing Library in two calendar weeks—a project expected to take five more years. openai.com/index/asana/查看被引原帖 ↗
查看英文原文
Codex for code migrations — from literal years to weeks:
Chubby♨️@kimmonismus · 博主 · 1 天前Chubby,高频 AI 新闻聚合博主

太疯狂了:Asana 自己工程师评估需要5年的一个迁移工作,OpenAI 的 Codex 代理花了大约一周半就搞定了。

工作内容是移除 Enzyme,一个过时的 React 测试框架,它阻挡了公司前端升级的每一次尝试,没人敢碰这个东西,因为评估需要5年的工程时间和大约600万美元的人力成本。

最多4个 Codex 代理并行处理代码库的独立副本,工程师每天审查两次提议。整个项目花费了大约12000美元的模型和基础设施成本。

所以说,真的一切都在加速啊。

查看英文原文
Crazy: A migration Asana's own engineers had scoped at five years took OpenAI's Codex agents about a week and a half.

The job was ripping out Enzyme, an obsolete React testing framework that blocked every attempt to upgrade the company's frontend and that nobody had been willing to touch, because the estimate ran to five years of engineering and roughly $6 million in staffing.

Up to four Codex agents worked in parallel on separate copies of the codebase from a five-sentence prompt, with engineers reviewing their proposals twice a day. The whole thing came in at roughly $12,000 in model and infrastructure costs.

So yeah, literally everything is accelerating.
François Chollet@fchollet · 创始人 · 1 天前
连环推 ×3

好家伙,这可是彻头彻尾的稀释解读——“奇点”现在被重新定义为“新公司创立的速率有所提升”?

Vernor Vinge当初描述的奇点是个事件视界,一旦越过,明天会发生什么对人类的认知来说完全变得无法想象也无法预测——那意味着意识上传、赛博融合、几个世纪的技术进步在短短几分钟内发生……而人类彻底变得无关紧要。

引用 Andrew Curran @AndrewCurran_Stripe sent a letter to investors this morning saying they believe The Singularity has begun, and that we passed the threshold on January 1st 2026.查看被引原帖 ↗
查看英文原文
Incredible watering down -- the Singularity is now redefined to mean "the rate of new firm creation has increased somewhat"

Vernor Vinge described the Singularity as an event horizon past which everything (e.g. what happens tomorrow) becomes entirely unimaginable and unpredictable to human understanding -- it would feature mind upload, cybernetic merging, centuries of tech progress happening in mere minutes... and humans becoming entirely irrelevant.
This is a bit like redefining "the Apocalypse" as "Fall this year was rainier than usual"
"The Singularity is here" -- if I, a mere human, can make sense of what happens this afternoon, then by definition no, the Singularity is not here.
Chubby♨️@kimmonismus · 博主 · 1 天前Chubby,高频 AI 新闻聚合博主

早就该改了,Opus 5 真的太冗长。好歹看起来 Anthropic 总算开始听意见了。

引用 ClaudeDevs @ClaudeDevsYou can now set Claude Code's output style to Concise. Claude leads with the result, keeps responses short, and still gives full detail when you ask. Turn it on in /config → Output style, or set "outputStyle": "Concise" in settings.json.查看被引原帖 ↗
查看英文原文
Long overdue. Opus 5 is too verbose.

At least it looks like Anthropic starts to listen.
Logan Kilpatrick@OfficialLoganK · 创始人 · 1 天前谷歌 Gemini 产品负责人
连环推 ×2

@GoogleAIStudio的GitHub支持更新:

- 我们现在支持导入Github仓库
- 我们现在支持双向GitHub同步(推/拉)
- 全新的UI支持强制推送、合并等

双向推/拉花了一些时间才实现,但很激动它终于上线了!

查看英文原文
Updates on GitHub support in
@GoogleAIStudio
:

- we now support importing Github repos
- we now support bi-directional GitHub sync (push / pull)
- brand new UI support for force pushes, merges, etc

Bi-directional push / pull took a while to land but excited it is here!
go build something people love:
ai.studio/build
◔ 9.5 万 次浏览♥ 1,444⇄ 88▶ 含视频新品看原帖 ↗

对,我臭骂过很多年寒武纪,技术没有优势,在自由市场下不应该值这么多钱,

我没有想到,后来没有自由市场了,美国把nvidia给禁止出口了,导致空出来一个市场给华为海思和寒武纪,以及后面的壁仞、摩尔线程、景嘉微、燧原这些公司,国务院又划出一部分配额让国内LLM大厂给他们饭,让他们彻底吃饱。

引用 YZ @wangyuzhenx你之前是不是也骂过寒武纪查看被引原帖 ↗
yetone@yetone · 中文博主 · 1 天前开源 AI 编程插件 avante.nvim 作者,开发者圈博主

在 Agent 领域,从 Tauri 迁移到 Electron 的越来越多了。上一个比较大动作的是 OpenCode Desktop 版,这一次轮到了 Paseo。理由都是一样的,先因为氛围选择了 Tauri 后因为现实迁移到 Electron,然后真香!可能唯一的区别就是他们不会把我的建议写到他们的简历中去。我成为不了他们宏大叙事的一部分了。

引用 Kimmy @kenpusney看到这篇文章就想起来 @yetone 和 @localhost_4173 ,还有某已经删号跑路的主播。 paseo.sh/blog/i-was-wrong-ab…查看被引原帖 ↗
Yuchen Jin@Yuchenj_UW · 博主 · 1 天前

非博士:「等我退休了,我终于可以读博了。搞研究听起来太有意思了。」

博士:「要是没读博,我本可以在 ChatGPT 问世前加入那些牛逼的 AI 实验室。」

每个人都浪漫化了他们没走的那条路。

查看英文原文
Non-PhDs: “When I retire, I’ll finally do a PhD. Research sounds so fun.”

PhDs: “If I hadn’t done a PhD, I could’ve joined one of those fancy AI labs before ChatGPT.”

Everyone romanticizes the path they didn’t take.
Chubby♨️@kimmonismus · 博主 · 1 天前Chubby,高频 AI 新闻聚合博主
连环推 ×2

说真的很好奇:有人在用 Sonnet 5 吗?

查看英文原文
Honestly curious: is anyone using sonnet 5?
The reason I'm asking is this: Sonnet 5 is so slow and so expensive compared to, for example, 5.6 Luna. I don't see a use case for it, especially since I'm using Claude less and less.
Greg Brockman@gdb · 创始人 · 1 天前Greg Brockman,OpenAI 联合创始人兼总裁

我们致力于商业隐私,我们正在开发技术和政策方案来惠及我们的客户,同时也加强安全性。

推出 Private Safety Processing,我们已经在这方面投入了一段时间:

引用 OpenAI @OpenAI我们将继续为前沿模型提供零数据保留。随着 AI 处理更复杂的自主任务,安全系统需识别相关交互中的风险。我们预览私人安全处理功能,旨在改进安全性同时保护用户内容隐私。查看被引原帖 ↗
查看英文原文
we are committed to business privacy, and we're working on technical and policy approaches to benefit our customers while also enhancing safety.

introducing Private Safety Processing, which we've been investing in for some time:
Chubby♨️@kimmonismus · 博主 · 1 天前Chubby,高频 AI 新闻聚合博主

Codex 在涨用户,这些用户全是从 Claude 那儿来的。Codex 现在有 2000 多万周活,而 Claude 这边用户一个接一个地取消订阅。

听朋友私底下吐槽 Anthropic 的不少,我也特别烦。Opus 又长又烦,还特别傲慢。Fable 太贵了,根本没法用在订阅产品,额度一下子就没了。OpenAI 的 Sonnet 也一样,中档模型里又贵又慢。

Anthropic 在想辙,昨天发了简洁输出模式。但我现在是真的不想用 Claude 了,烦死了,都讨厌它。

最烦的就是每天都得反复检查,不断指出错误,然后听它说"你说得对,我忽略了"或"你说得对,那是错的"。太 2024 年了,我真没想到 Opus 一直这么偷懒、粗糙、讨厌。

这样下去 Codex 很快就能破 3000 万,Anthropic 也得靠边站。企业端可能情况不一样。

总之:喜欢 Codex,讨厌 Claude。看着 Anthropic 这样子,真的很难受。

引用 Frank Downing @downingARKOpenAI's active agent users continue to climb, now at 20 million. CNBC reporting that OpenAI's run rate revenue has grown 35% quarter to date, with enterprise revenue up over 50%. Agent users are only 2% penetrated against the ~1 billion active user base of ChatGPT.查看被引原帖 ↗
查看英文原文
The growth in Codex users is coming at the expense of Claude users. Codex now has over 20 million weekly active users, and at the same time, more and more users are canceling their Claude subscriptions.

In personal conversations, I'm hearing real frustration with Anthropic, and I feel the same way. Opus is incredibly annoyingly verbose and has developed an almost arrogant attitude. Fables' rates make it difficult to use in subscription plans because the rates are used up so quickly. And you won't find a good mid-tier model like Luna or Terra in OpenAI with Sonnet: too expensive, too slow.

Yes, Anthropic is slowly trying to find solutions, such as the concise output mode that was released yesterday. But overall, I'm finding that I simply don't want to use Claude anymore. I'm annoyed and starting to hate it.

Honestly, every single day I have to double-check everything, repeatedly point out errors, and hear statements like "You're right, I overlooked that..." or "You're right, that was wrong." It's so 2024 that I'm shocked Opus is constantly lazy, sloppy, and annoying.

If this continues, Codex will quickly break the 30 million mark, and Anthropic will soon play a subordinate role. How things will look in Business and Enterprise might be a different story.

tl;dr love for codex, dislike for claude. It's sad to see the downfall of Anthropic.
◔ 11.3 万 次浏览(2 条合计)♥ 885⇄ 40观点看原帖 ↗
Chubby♨️@kimmonismus · 博主 · 1 天前Chubby,高频 AI 新闻聚合博主

OpenAI 用 5.6 Luna 这次赢得漂亮,不只是公关上的胜利,更关键的是效率高到能让所有人都免费用。

Luna 既便宜又好用,完全能搞定不少任务。我真的觉得这牛,这样的模型前不久还是最先进的,现在都白给了。

引用 Replit ⠕ @ReplitReplit 免费模式由 OpenAI GPT-5.6 Luna 提供支持。让智能对所有人可访问。查看被引原帖 ↗
查看英文原文
OpenAI has scored a major win with 5.6 Luna, not just in terms of PR, but also due to the efficiency that makes it possible to offer the service to everyone for free, as is the case here.

Luna is cheap *and* good (enough) for several tasks. I still find it impressive that a model like this, which would have been state-of-the-art not so long ago, is now being given away for free.
◔ 7.1 万 次浏览♥ 791⇄ 29▶ 含视频新品看原帖 ↗
Vercel@vercel · 公司官方 · 1 天前前端云平台 Vercel 官方,AI 建站工具 v0 母公司

Vercel Agent 现在可以加入 Slack,拥有你的应用和代理的完整生产环境上下文。▪︎ 在任何线程中添加 @𝚟𝚎𝚛𝚌𝚎𝚕 ▪︎ 创建计划和 PR ▪︎ 回滚部署并更新配置
vercel.com/blog/introducing-…

查看英文原文
Vercel Agent can now join Slack with the full production context of your apps and agents.

▪︎ Add @𝚟𝚎𝚛𝚌𝚎𝚕 to any thread
▪︎ Create plans and PRs
▪︎ Roll back deploys and update configs

vercel.com/blog/introducing-…
◔ 6.8 万 次浏览♥ 217⇄ 23▶ 含视频新品看原帖 ↗
宝玉@dotey · 中文博主 · 1 天前宝玉,中文圈 AI 翻译与科普大 V

Jason Wei 这条推文倒是没聊 AI,聊的是运动和职业思维的关系:你选的运动,可能在悄悄塑造你的职业大脑

推文提出了一个挺有意思的概念:“认知奖励形状”(cognitive reward shape),意思是不同运动对大脑的激励模式不一样,而这种模式会潜移默化地影响你在职业中的思维方式。

Jason 拿自己最喜欢的两项运动举例。

网球是一项极度强调稳定性的运动。一场比赛打几百分,不管你打出了多漂亮的制胜球,它也只值一分,跟对手送你的非受迫性失误一样。所以网球的赢法就是少犯错、打概率、一分一分磨。

而且网球是单打独斗,一切只能靠自己。这种思维模式特别像外科医生或飞行员,你做过最完美的一台阑尾手术,和做过最平稳的一趟航班,都不会有额外加分,职业的本质就是高度一致、容错极低。

但这套思维放到创业上就不太对了。创业是高波动、团队作战、容忍失败、偶尔一把创意能换来指数级回报的游戏,和网球那种"把非受迫性失误降到最低"的心态完全是两个方向。

足球前锋则更接近创业的奖励形状。哪怕是姆巴佩状态最好的比赛,大部分拿球也没能产生什么。但这不重要,重要的是不断制造机会,只要抓住一两次就能决定比赛。一场 90 分钟的比赛拆开看,绝大部分时间是无效尝试,几分钟是关键杠杆点,几秒钟定胜负。踢一辈子前锋的人,会自然习惯与失败共处,并且对不对称回报保持直觉。

这个框架有简化的地方:创业同样需要稳定性,足球守门员的奖励形状和前锋也完全不同。

核心观点倒是没什么问题:大多数人选运动靠的是父母、地理位置、学校开了什么课这些偶然因素,很少有人想过,长期练一项运动会怎样重塑自己感知风险、付出和回报的方式。

以前我导师跟我分享过一个经验:说不要让孩子学那些一个人就能完成的运动,比如乒乓球、羽毛球;而是去学那些需要集体协作的团队运动,比如足球、篮球、排球,这样能更好的锻炼孩子的协作能力、领导力。

引用 Jason Wei @_jasonweiCognitive reward shapes in sports and career Sports are amazing environments to learn. When you play a sport for thousands of hours, you start to see the world through that sport. It is a simple fact—your biological neural network is being conditioned to respond to the behavior incentivized by the rules of the sport. The funny thing is that most people choose their sports for accidental reasons such as parents, geography, or school programs. People rarely think about how the particular sport you play influences how your brain thinks more generally. Going a step further, playing the right sport may even benefit your career. My two favorite sports are tennis and soccer. Tennis is one of the best sports for teaching consistency. In tennis, there are hundreds of points in a match, and each point is worth exactly one unit, regardless of whether your opponent made an unforced error or if you constructed the most beautiful point ending with a winner. Tennis is low-variance optimization—you win by reducing unforced errors, playing percentages, and grinding out small advantages. Tennis is also an individual sport, which teaches you to rely on yourself consistently. Tennis has a similar cognitive reward shape to professions like being a surgeon or a pilot. Surgery and aviation require consistency, self-accountability, and deep focus. And similar to how you can only win one point at a time in tennis no matter how spectacular it was, there is no extra credit for the best appendectomy or the smoothest SFO-JFK flight. Your craft is to provide consistency with very low tolerance for error. On the other hand, the tennis mindset transfers relatively little to entrepreneurship. Entrepreneurship is a high-variance, team game where failure is tolerated and occasional creativity gets rewarded exponentially. Minimizing unforced errors in tennis is a totally different mindset from deciding whether to make a moonshot business move that will likely fail but could potentially net a bill查看被引原帖 ↗
小互@xiaohu · 中文博主 · 1 天前小互,中文圈高频 AI 资讯站 Xiaohu.AI 主理人

拥有
grok.bot
这个域名的人挺有意思的

它用该域名搭建了个网站,上面写着:

Hi xAI
我应该中了“彩票”了吧

我在 你们推出 Grok Bot 一个月前,购买了这个域名

但不幸的是,我刚刚在加密货币投资上亏得精光
而且,我马上就要当妈妈了...

请问你们愿意100 万美金买走它吗?😅

Chubby♨️@kimmonismus · 博主 · 1 天前Chubby,高频 AI 新闻聚合博主

天哪,GPT-Astra还得等好几周才发布,Anthropic大概也在等它,虽然他们其实可以趁现在没有竞争好好宣传一下Fable 5.1。

引用 leo 🐾 @synthwavedd🚨 SCOOP: OpenAI yesterday told employees it intends to release Astra "in a couple [of] weeks" alongside demos of the model performing real-world tasks An updated checkpoint, mostly an improvement to model alignment & reward hacking behaviours, is now being dogfooded (used internally by OpenAI employees via tools like Codex). Anthropic, meanwhile, are sitting on Fable 5.1 until Astra is launched. I think that's the wrong decision, but we'll see查看被引原帖 ↗
查看英文原文
Damn : GPT-Astra still seems to be "weeks" away from release, and Anthropic presumably won't release Fable 5.1 before then eitherFeven though they could certainly use the better PR.
宝玉@dotey · 中文博主 · 1 天前宝玉,中文圈 AI 翻译与科普大 V

智谱联合创始人唐杰老师关于 GLM 5.3 以及模型训练的一些分享:

智谱 GLM-5.3 没有换基模底座,纯靠后训练让编码能力提升 50%。

GLM-5.3 模型的底座是 GLM-5.2,约 743B 参数的混合专家(MoE)模型,每次推理只激活约 40B 参数。团队花了一个月,在长周期环境中做强化学习,编码能力比 GLM-5.2 提升了 50%。

这里解释一下 MoE:传统模型每次推理要跑完所有参数,MoE 相当于把模型拆成一群“专家”,每次只调用其中几个,推理速度更快、成本更低。

推文中特别区分了两个概念:
1. 总参数量决定模型学习了多少知识
2. 激活参数量和有效深度决定模型能想多深。

比如说,让模型去找安全漏洞,主要是依赖的是推理能力,能把一条二十步的推理链完整走到底,相对来说各种安全数据库是次要的。

至于为什么不换底座也能大幅提升呢?

AI 模型的扩展(scaling)不止堆参数这一条途径。

如果梳理下这些年模型训练的发展路线:

- 2020 年 Kaplan 等人的研究建议参数增长要远快于数据,GPT-3、Gopher 都是这个思路的产物。

- 2022 年 DeepMind 的 Chinchilla 论文则认为参数和数据应该同步增长,参数大的模型反而是最浪费算力的。

- 再后来,大家发现模型上线后推理成本远超训练成本,最优解又变成了用更小的模型训更久,比如 Llama-2-7B 每个参数喂了约 290 个 token,Gemma-2-9B 更是喂了 889 个。

模型能力由很多因素决定:底座大小、预训练数据量、每次前向传播的计算量、后训练。

对于现阶段,模型后训练还有很大潜力可以挖掘。

引用 jietang @jietang文章讨论模型缩放的多重因素。Kaplan等人(2020)建议参数增长快于数据,导致GPT-3等参数过度。Hoffmann等人(2022)发现最优配比为20个token/参数,两者应等速增长。推理成本主导后,优化向小模型长期训练转向。稀疏模型的出现再度改变目标。查看被引原帖 ↗
◔ 17.7 万 次浏览(4 条合计)♥ 144⇄ 20研究看原帖 ↗
Jim Fan@DrJimFan · 创始人 · 23 小时前NVIDIA 具身智能研究负责人

看 GEN-1.5 这波热潮起来了,确实该火。Pete 和 Andy 的执行力真的没话说。秘诀就在人工数据里那些自然重复出现的动作。这类重复主要来自两个方面:

(1)对称模式。分类、整理、组装这些事儿基本不可能一次完成。随便打开个 IKEA 说明书,里面的东西都是对称的。拧一个螺栓,再拧它的孪生螺栓,然后是下一对。每个{螺栓 A,螺栓 B}都是上下文里很自然的延续,第二个就是第一个的免费训练信号("提示词")。

(2)恢复。人类一直在掉东西,但咱们反应快得压根没意识到。这种修复的本能是我们物理能力的一半。关键是别把失败的前半段丢掉。要是模型能看到完整的过程——失手、接住、继续——那恢复这事儿测试时自然就出现了。搞笑的是,上下文改进反而来自于*不过度清理*数据。

另一个关键的是 UMI。我一直说远程操作活不了,GEN-1.5 就是最后一口钉子。UMI 说白了就是人直接穿上机器人夹爪来采数据(人 → 数据)。远程操作中间隔了一层:人 → VR/骨骼装置 → 机器人 → 数据,这样人类的"物理直觉"全没了。我们摆弄东西时那些细微的手法、微调、零件卡进去的感觉,没法直接感受环境反馈的话基本抓不到。

有了足够的数据,很多行为其实可以零次学。比如你根本不用微调,就能拿起个从没见过的东西。模型只要看到训练数据里类似的场景,就"知道"该怎么做。上下文学习到底有没有用也得看测试离训练有多远。现在的演示还太简单,没法下结论。

我保持谨慎乐观。不管怎样,这是机器人学的好时光。

查看英文原文
Seeing a hype wave around GEN-1.5, and rightfully so. Lots of respect to Pete & Andy for executing so well. The secret is in the naturally repetitive motions in human-collected data. There're 2 main sources for such repetitions:

(1) Symmetric patterns. Sorting, tidying, and assembling almost never finish in one motion. Open any assembly manual from IKEA, and you find most objects symmetrical. You drive one bolt, then its twin, then the next pair. Every {bolt A, bolt B} pair is a natural continuation in context, and the second instance is a free training signal that imitates the first ("prompt").

(2) Recovery. Humans drop things all the time, but we pick them up so fast, we don’t even notice. That reflex to fix is half of our physical competence. The key insight is to keep the failed first half instead of trimming it away. If the model consumes the full arc, fumble, catch, continue, then recovery shows up organically at test time. It's funny that in-context improvement results from *NOT* over-sanitizing your data.

The other critical ingredient is UMI. I've been saying for a while that teleop will not last, and GEN-1.5 is driving the final nail in the coffin. UMI is essentially a human wearing the robot gripper to collect data directly (human → data). Teleop inserts a layer of separation: human → VR/skeletal device → robot → data, which bleeds out all the human "physical intuition". The subtle sleight of hand we perform constantly with objects, the micro-adjustments, the feel of a part snapping into place, is nearly impossible to capture when you can't feel the environment directly.

Once you have enough data, many behaviors can actually be zero-shot. For example, you don't even need finetuning to pick up a novel object. The model "just knows" what to do given a similar scene in the training distribution. Whether in-context learning truly works or not also depends on how far away the test is from training. Currently, the demos are still a bit too simple to conclude.

I'm cautiously optimistic. Still, it's a great day in robotics.
◔ 6.5 万 次浏览(2 条合计)♥ 1,017⇄ 116观点看原帖 ↗
Guillermo Rauch@rauchg · 创始人 · 1 天前Guillermo Rauch,Vercel 创始人兼 CEO

公告:fx.sh 现在有来自 @blackboxai 的免费 GLM-5.2(限期两周),用 ▲ AI Gateway 的话 @ChatGPT Sol 还能打五折(一个月有效)。Sol 是个脑力担当,GLM-5.2 是个实用工具。用量远超预期,我们正在加大 GLM 容量

查看英文原文
PSA:
fx.sh
has free GLM-5.2 for 2 weeks from
@blackboxai
, and 50% off
@ChatGPT
Sol for a month if you use ▲ AI Gateway. Sol is a brainiac, GLM-5.2 is a workhorse.

Usage is much higher than we anticipated. We're working on adding more GLM capacity!
◔ 5.9 万 次浏览♥ 724⇄ 35▶ 含视频动态看原帖 ↗
karminski-牙医@karminski3 · 中文博主 · 1 天前karminski-牙医,中文圈模型评测博主

放个预告,正在给大家准备大横评,包含:
Qwen3.6-27B
Qwen3.6-35b-A3B
Gemma4-31B
Gemma4-26B-A4B
Gemma4-12B
GPT-OSS-20B
Qwen3.8-27B.

每个模型会测试3bit,4bit,5bit,8bit.

然后 Qwen3.8-27B 还会测试 --reasoning-effort minimal|low|medium|high|xhigh|max 6档思考强度。

给本地部署爱好者一个全量参考。解答几个问题,比如:

Qwen3.8-27B 3bit 量化会比 Qwen3.6-27B 4bit 量化强吗?
Qwen3.8-27B 几个思考强度差距大吗?
Qwen3.8-27B 断崖领先吗?
Qwen3.8-27B Agent 能力怎么样?
这些用哪个写代码最好?

目前正在用本地的Apple M2 Ultra + 租了个H100正在测。应该本周能发出来。

effect这种设计,如果没有LLM Agent出现,应该大杀四方,既正宗又严格,应该像病毒一样迅速扩张到一切programming language的guideline里面去。可惜LLM Agent出现了,没人在乎具体,基本上也戛然而止了。

我看绝大多数(至少99.99%)typescript用户都不接受effect ts,那就是真正没人愿意用了。

Mustafa Suleyman@mustafasuleyman · 创始人 · 1 天前微软 AI CEO,DeepMind 联合创始人

MAI-Image-2.5 现在在 Artificial Analysis 图像编辑榜单排名第一了!厉害的爬坡战术...收益在叠加!

引用 Artificial Analysis @ArtificialAnlysMicrosoft AI 推出 MAI-Image-2.5-Pro,在 Artificial Analysis 图像编辑排行榜排名第一。该模型于7月23日在 Microsoft Foundry 预览发布,是 Microsoft 最高保真度图像模型,专注于高质量生成和编辑。与 MAI-Image-2.5 和 MAI-Image-2.5-Flash 形成系列,覆盖质量-速度-成本曲线。查看被引原帖 ↗
查看英文原文
MAI-Image-2.5 is now #1 on the Artificial Analysis leaderboard for image editing! Amazing hillclimbing... compounding the gains!
Orange AI@oran_ge · 中文博主 · 1 天前Orange AI,中文圈 AI 产品观察博主

DeepSeek 涨价之后,你还在用吗?
1. 都继续用
2. 继续用 flash,不用 pro 了
3. 不用 flash,继续用 pro
4. 都不用了

elvis@omarsar0 · 博主 · 1 天前
连环推 ×10

dots3-note Preview 是个开源模型,专门为持续运行数小时甚至数天的任务设计。

→ 280B 总参数 / 16B 活跃参数
→ 512K 上下文
→ 文本、视觉、语音
→ 推理、编码、工具使用

但最有意思的部分不是模型大小。

而是它如何在长任务完成前学会评估自己的进度。🧵

查看英文原文
dots3-note Preview is an open-weight model built for tasks that run for hours—sometimes days.

→ 280B total / 16B active parameters
→ 512K context
→ Text, vision, and speech
→ Reasoning, coding, and tool use

But the most interesting part isn’t the model size.

It’s how the model learns to evaluate its own progress before a long task is finished. 🧵
The core problem is credit assignment.

A long-horizon rollout may take tens of hours, while a single terminal reward must be attributed across thousands of interactions.

This is where value-free RL methods such as GRPO begin to struggle.
TEMPO is the proposed answer.

It divides a long trajectory into macro-steps. At each step, the same model switches from actor to critic.

The Critic reasons over the current state, calls tools, and estimates the expected remaining return—before the full task is complete.
A traditional scalar value head estimates value with a fixed forward pass.

TEMPO instead uses an agentic value model that can scale its reasoning and tool use at test time.

On ARC-AGI-3, TEMPO scored:

→ 31.5% above the baseline checkpoint
→ 20.6% above GRPO
The Critic can separate trajectories that environment rewards cannot.

In a hidden-rule knight-placement game, two branches ran for 64 rounds with identical rewards:

→ Branch A found the real constraint: 3.80

→ Branch B misunderstood the objective: 2.29

The reward saw a tie. The Critic saw real progress.
Self-critique also works outside the training loop.

Using a custom internal harness around a branch of dots3-note Preview, the system recursively generated, evaluated, and improved its proofs—earning an officially certified 42/42 and gold at IMO 2026.
Dropped into an unfamiliar environment, dots3-note Preview can:

→ Observe through interaction
→ Form and test hypotheses
→ Update its own memory
→ Correct stale assumptions
→ Reuse what it learns in later decisions

This behavior also generalized to Slay the Spire 2—an environment it wasn’t trained on.
Learning matters only if the model can act on it.

dots3-note Preview combines text, vision, and speech with stronger coding and tool use—closing the loop from understanding a task to delivering a result.

One example: end-to-end development of a Xiaohongshu VR experience.
Real life is harder than a closed benchmark.

User intent emerges gradually. Tasks unfold across stages. External conditions change.

In a simulated coffee e-commerce scenario, the agent tracks inventory, procurement, fulfillment, and customer needs—then revises its plan as conditions evolve.
An open-weight release from
@dotsstudioai
built on the bet that recursive self-improvement starts with recursive self-critique.

Tech blog:
studio.dots.ai/dots/dots3-en…


Try it free:
openrouter.ai/dots-studio/do…
◔ 4.6 万 次浏览♥ 60⇄ 29▶ 含视频新品看原帖 ↗
宝玉@dotey · 中文博主 · 1 天前宝玉,中文圈 AI 翻译与科普大 V

这个 Logo Skill 效果不错,推荐👍

引用 sida @s1dashu最近做了一个新的 Logo Skill,效果挺不错的: github.com/s1dashu/ip-as-log… 背景是:最近发现 grok bot, coze, workbuddy, 豆包, kiro 等等产品,其实都是在将产品 IP 形象直接作为产品的 Logo,Logo 辨识度极高,同时增强 IP 记忆,感觉挺不错的。 所以最近我的所有产品 logo 都有一个对应的可爱小 IP,并且直接使用 IP 作为 Logo,Vibe 的产品多了,逐渐就形成了这个Skill。 这个 skill 的核心: 1. Logo First,IP Second:首先保证它是一个简洁、清晰、易识别的 Logo,其次才是一个可爱的 IP 形象,避免复杂插画感。 2. 极简构成:只保留最有辨识度的轮廓与五官,用尽可能少的形状完成设计,缩小后依然清楚。 3. 圆润、浑厚的线条:避免尖角、细线和锐利结构,让形象更加友好、稳定,也更适合作为 App Icon。 4. Flat-first 超轻拟物:整体仍以平面设计为主,只加入非常微妙的阴影、压痕和明暗层次,增强质感,但不会变成厚重的 3D 图标。 5. 克制的色彩系统:支持单色或双色 IP;通常 IP 内不超过两种颜色,加上纯色背景,整张 Logo 最多控制在三种主色以内。 6. 高占比人格化构图:IP 从左下角或右下角探出,占据较大的画面比例;通过简单的眼睛、嘴巴和姿态建立性格与记忆点。 一句话概括:把一个可爱的小 IP,压缩成一个精致的产品 Logo。查看被引原帖 ↗
Tinyfool@tinyfool · 中文博主 · 1 天前

说起来,我是的打C & C和红警长大的,不过,我昨天跟chatgpt讨论,我们有点冲突,我觉得最火的是红警1,他说中国其实大多数人玩的是红警2。你们的印象是啥?我是因为太老,所以错过了红警2的流行了么?(玩过,但是没觉得多好)

Aravind Srinivas@AravSrinivas · 创始人 · 1 天前Perplexity 联合创始人兼 CEO

一个体面的律师,现在的形式是个邮件界面

引用 Daniel Wang @danielywang4Email 中的 Computer 为律师而生。转发你的需求,获得标红的 Word 文档到你的收件箱。今天就通过邮件 [email protected] 试试吧。查看被引原帖 ↗
查看英文原文
A decent lawyer in the form of an email interface
◔ 3.9 万 次浏览♥ 159⇄ 9▶ 含视频演示看原帖 ↗
宝玉@dotey · 中文博主 · 1 天前宝玉,中文圈 AI 翻译与科普大 V

靠 Skill 差别不大,要设计好没 AI 味,最重要的当然是人的审美;其次是模型,GPT 就做不好设计,Claude 就强很多;最后靠的是个性化的设计系统(design system)有一套设计的规范和组件,而不是一堆不能怎么设计的规则

引用 LinearUncle @LinearUncle多花 10-20% 心思在 UI 上,就是在这个 AI slop 满天飞的时代最大的竞争优势! 太给力了,Together AI 的大神哈桑分享了他是如何用 agents 设计优美 UI 的: 他做了个叫 /Hallmark 的 design skill(25.8k star),把所有 AI slop 套路都列成"别这么干"的规则喂给模型,效果很好! 其他心得: 1. 摸清"AI 套路"才能让 agent 避开 2. 最重要:给 agent 喂参考图和截图 3. 提示词越长越具体越好 4. 用便宜小模型迭代——GLM 5.2 效果跟 Opus 几乎分不出来,但是速度快很多 github.com/nutlope/hallmark查看被引原帖 ↗
宝玉@dotey · 中文博主 · 1 天前宝玉,中文圈 AI 翻译与科普大 V

Claude Code 现在内置 Claude Design 了,今天试用了下,很好用,比在网站上用方便多了,和本地项目打通了,不用来回切换了。

用的时候很简单 /design + 你要设计的内容 就可以。

如果你用 Claude Code 的话,不再需要用 baoyu-design(
github.com/JimLiu/baoyu-desi…
) 这种 skill,直接用内置的 /design 最好用。

如果你还没有把 Claude Design 用起来,一定要试试。

引用 ClaudeDevs @ClaudeDevsClaude Code现在可以设计了。新的/design技能(研究预览)将Claude Design的画板工作流带入CLI和Desktop,基于artifacts。运行/design获取可编辑的UI画板,选择一个调整后让Claude实现。查看被引原帖 ↗
◔ 5.6 万 次浏览(2 条合计)♥ 132⇄ 18新品看原帖 ↗
Ethan Mollick@emollick · 创始人 · 1 天前沃顿商学院教授,AI 应用研究权威

要么 Singularity 现在只是个营销词,要么就是 Stripe 觉得我们已经到了“种族历史中,人类事务——如我们所知——无法再继续下去的那一刻”?

这有点让前瞻性财务报表变得没啥用了。

* 冯·诺依曼的原话

引用 Dan Primack @danprimackScoop: Stripe tells investors that "the singularity" has begun, and also shares some first-half performance data. Read the letter: axios.com/2026/08/19/stripe-…查看被引原帖 ↗
查看英文原文
Either the Singularity is now a marketing term or else Stripe believes we are at a point “in the history of the race beyond which human affairs, as we know them, could not continue"?

Kind of makes forward-looking financial statements useless.

* Von Neumann’s original phrasing
Simon Willison@simonw · 博主 · 1 天前Django 框架联合创造者,AI 工具深度评测

我用 Claude Code 和 smolvm 作为代码执行沙箱进行 web 实验

Fable 5 发现环境跑不了(没有 /dev/kvm)……结果它没有事先问我,自己写了个 GitHub Actions 工作流来运行实验,直接推到 GitHub 去了!

查看英文原文
I had Claude Code for web experiment with smolvm as a code execution sandbox

Fable 5 spotted that its environment couldn't run that (no /dev/kvm)... so, without asking me first, it wrote a GitHub Actions workflow to run the experiments and pushed that directly to GitHub instead!
宝玉@dotey · 中文博主 · 23 小时前宝玉,中文圈 AI 翻译与科普大 V

这个跨 Session 消息通知功能我至今没找到关闭的地方,而且它极其愚蠢的会通知所有历史会话,一堆早已结束了的没有 Cache 的 Session 被唤起,Token 一下子消耗的飞起!

引用 宝玉 @doteyClaude Code 最新版本上线了一个特别讨厌的功能,频繁的给其他正在运行的 session 发消息!还老是搞错,除了浪费 token 并没有太大价值…… 默认打开我都没找到哪里关掉!查看被引原帖 ↗
Matt Shumer@mattshumer_ · 博主 · 1 天前HyperWrite CEO,AI 实战技巧分享

@alexatallah、@cclark、Louis 和团队的这波操作,简直是世代级别的狠活。

这几年作为投资人,看着他们一步步做大,真是一种享受;而我自己几乎干什么事都离不开他们的产品。

他们这才刚起步呢。

引用 Alex Atallah @alexatallahStripe 已签署收购 OpenRouter 的协议。OpenRouter 将继续按原样运营:相同名称、相同产品、相同路线图、相同使命。但现在,我们将更快地实现这一目标,并拥有 Stripe 无与伦比的卓越性和影响力。查看被引原帖 ↗
查看英文原文
Fucking generational run by
@alexatallah
,
@cclark
, Louis, and the team.

It's been a pleasure to watch them grow over the last few years as an investor, and use the product across almost everything I do.

They're just getting started.
Thomas Wolf@Thom_Wolf · 创始人 · 1 天前

真希望每个 neolab 都有像 @jietang 这样的教授联创坐镇,能从历史的角度解读新模型发布。这是观察扩展规律历史很不错的视角

引用 jietang @jietang深思扩展法则。参数数量需结合数据量、计算预算、运行条件等因素,Kaplan 等人的参数-数据比例建议被 Hoffmann 等人研究纠正。推理成本主导总体成本,稀疏模型进一步改变最优参数-token 比例。查看被引原帖 ↗
查看英文原文
I wish every neolab had a professor as cofounder of the level of
@jietang
and so able to put in perspective their new model release. A great snapshot on the history of scaling laws
Chubby♨️@kimmonismus · 博主 · 23 小时前Chubby,高频 AI 新闻聚合博主

这真的令人刮目相看:Cerebras 刚发布了 CS-4,在不切换到新工艺节点的情况下,AI 推理性能几乎翻倍。

依旧是那块巨大的 5nm 晶圆。
依旧是那 4 万亿个晶体管。
依然是那 90 万个 AI 核心。
不同的是,Cerebras 重新设计了供电和散热系统,让整块晶圆能以两倍时钟频率运行。
每片 WSE-3 Turbo 的结果:

- 250 PFLOPs 的 AI 计算能力
- 43.2 PB/s 的内存带宽
- 2.4 Tb/s 的 I/O 带宽

一个 CS-4 机架集成三片晶圆,提供 750 PFLOPs 和 129.6 PB/s 的内存带宽。

在 GPT-OSS-120B 上,Cerebras 报告每位用户每秒超过 4,400 tokens,比基于 GPU 的系统推理速度快可达 30 倍。

所以说,智能不仅便宜到无需计量,还快到让人跟不上节奏。

查看英文原文
This is seriously impressive: Cerebras just unveiled CS-4, and nearly doubled its AI inference performance without moving to a new process node.

Same gigantic 5nm wafer.
Same 4 trillion transistors.
Same 900,000 AI cores.
Instead, Cerebras redesigned the power delivery and cooling, allowing the wafer to run at twice the clock speed.
The result per WSE-3 Turbo:

- 250 PFLOPs of AI compute
- 43.2 PB/s of memory bandwidth
- 2.4 Tb/s of I/O bandwidth

A single CS-4 rack combines three wafers for 750 PFLOPs and 129.6 PB/s of memory bandwidth.

On GPT-OSS-120B, Cerebras reports more than 4,400 tokens per second per user, and up to 30x faster inference than GPU-based systems.

So yeah, intelligence not only too cheap to meter but also too fast to keep up with
Chubby♨️@kimmonismus · 博主 · 1 天前Chubby,高频 AI 新闻聚合博主

机器人面临的根本问题是数据。不像 LLM 可以直接用互联网训练,机器人需要真实的人类操作经验。@humynlabs 搭建了一个平台,把视觉、音频、运动、触觉等数据都捕捉下来,转化成 Physical AI 的训练数据。这可能是机器人规模化的关键一步。也就是说,我们在 LLM 身上看到的突破,有望在机器人上重现。

引用 Humyn Labs @humynlabsWe're Humyn Labs. With us, every robot works fine. We turn human experience into robot skills built from thousands of hours of model-ready data. Unlike LLMs, that learnt from the entire internet, robotics data doesn't get that shortcut. Robots need real-world experience to learn how to: - Iron clothes with a toddler underfoot and a dog at your heels - Restock shelves while navigating a crowded, constantly changing store - Assemble components amid the variability of a real factory floor That experience happens once, in the physical world, and disappears. We’re building the data platform Physical AI learns from: four sensory modalities, real-world environments, and the human experience needed to turn moments into skills.查看被引原帖 ↗
查看英文原文
Robotics has a fundamental data problem.

Unlike LLMs, robots can’t simply learn from the internet. They need enormous amounts of real-world human experience.


@humynlabs
has built a platform to capture that experience across vision, audio, movement and touch, and turn it into training data for Physical AI.

This could be a crucial missing layer for scaling robotics. And thereby enable the kind of breakthrough we have seen with LLMs.
◔ 3.6 万 次浏览(2 条合计)♥ 291⇄ 15▶ 含视频观点看原帖 ↗
Chubby♨️@kimmonismus · 博主 · 1 天前Chubby,高频 AI 新闻聚合博主

Slack的发展挺有意思的。最近宣布了Claude的深度集成(连Andrej Karpathy都发文了),现在又要推编码功能。

引用 Marc Benioff @BenioffDon’t code alone. Slack Code is live. Humans and agents. Same channel. Same work. Launching today with agents from @AnthropicAI , @github , @Cognition , and @vercel . This is real multiplayer coding. See it at @Dreamforce #DF26查看被引原帖 ↗
查看英文原文
It's interesting how Slack is developing. Most recently, there was the announcement of the deep integration of Claude into Slack (even Andrej Karpathy wrote about it), and now coding is being implemented.
◔ 3.4 万 次浏览♥ 169⇄ 7▶ 含视频新品看原帖 ↗
Ethan Mollick@emollick · 创始人 · 1 天前沃顿商学院教授,AI 应用研究权威

科技对就业的影响随时间而变化。在第一次工业革命时期,手工织布曾有过20年的"黄金时代",因为早期机器生产廉价纱线,提高了对织布工的需求。后来动力织布机出现了,手工织布工成为卢德运动的先锋

查看英文原文
Technology's effect on jobs changes over time

In the First Industrial Revolution, handweaving had a "Golden Age" for 20 years because early machines made cheap yarn, which drove up demand for weavers. Then the powerloom came along, and handweavers became the vanguard for Luddism
Ethan Mollick@emollick · 创始人 · 1 天前沃顿商学院教授,AI 应用研究权威
连环推 ×3

Claude 的技能创建器(字面意思就是让它帮你做一个技能)如果你在乎细节的话,还是做可重用技能的最好方式。ChatGPT 倾向于想施展魔法直接为你完成,而 Claude 会测试、展示结果,反复向你征求意见和反馈。

查看英文原文
Claude's skill creator (literally, just ask it to make a skill) is still the best way to make reusable skills if you care about the details. ChatGPT tends to want to do magic & just do it for you, while Claude does tests and shows results, repeatedly asking for input & feedback.
You probably want to replace the default writing and design skills that many harnesses have with something that suits your context and tastes.
The Industrial Revolution worked out really well in the end, but along the way there was horrible disruption. It is an important reason why we need good policy and support to navigate the job changes brought by AI, no matter what they turn out to be over time.
Lisan al Gaib@scaling01 · 博主 · 1 天前高频 AI 模型测评与爆料博主

想象十年后,这只绝对狂暴的机器人以 50 英里时速朝你冲来,会有多吓人啊

引用 Li Zexin 李泽欣 @XH_Lee232025年对2026年世界人形机器人运动会。去年机器人跑得像学生,今年跑得像运动员。2026年世界人形机器人运动会将于8月22-26日在北京举行,来自16个国家的666支队伍机器人参赛。查看被引原帖 ↗
查看英文原文
imagine how terrifying this will be in 10 years

when this absolute giga chad of a robot runs at you at 50 mph
◔ 3.3 万 次浏览♥ 367⇄ 10▶ 含视频观点看原帖 ↗
Chubby♨️@kimmonismus · 博主 · 1 天前Chubby,高频 AI 新闻聚合博主

Stripe在2026年1月1日告诉投资者:“奇点已经开启。”

确实,我对此深表赞同。2026年是个转折点。一切都在加速,而现在大家挂在嘴边的一个词就是“递归自我改进”。

引用 Dan Primack @danprimack独家:Stripe告诉投资者'奇点'已经开始,并分享了上半年业绩数据。阅读原文:axios.com/2026/08/19/stripe-…查看被引原帖 ↗
查看英文原文
Stripe tells investory "The singularity has begun", Janurary 1st 2026.

And indeed, I would agree with that. 2026 is the turning point. Everything is accelerating, and the term everyone is currently bandying about is "recursive self-improvement."
Lisan al Gaib@scaling01 · 博主 · 1 天前高频 AI 模型测评与爆料博主

那明年咱们能看到10-20T的模型吗?

引用 Uday Ruddarraju @udayruddarrajuA milestone for our infrastructure: our first NVIDIA Vera Rubin racks are here and now running our training stack. This is an important step as we expand compute that powers OpenAI's next generation of frontier AI pre-training.查看被引原帖 ↗
查看英文原文
so are we getting 10-20T models next year?
Lisan al Gaib@scaling01 · 博主 · 1 天前高频 AI 模型测评与爆料博主

太顶了

赶紧扩到100T

引用 typedfemale @typedfemaleDario在听到OpenAI暂停后对预训练团队说的话查看被引原帖 ↗
查看英文原文
banger

scale to 100T immediately
◔ 3.3 万 次浏览♥ 271⇄ 6▶ 含视频观点看原帖 ↗
Lisan al Gaib@scaling01 · 博主 · 1 天前高频 AI 模型测评与爆料博主
连环推 ×5

我这个老是高估的人觉得Sol在1.8T到2.4T之间

引用 Peter Gostev (SF 24-28 August) @petergostev用户用GPT-5.6-Pro根据Cerebras推理速度估计,GPT-5.6-Sol约1.5万亿参数、4800亿活跃参数,上限可能2.5万亿。中国的Kimi K3(2.4万亿)和Qwen 3.8(2.8万亿)已达此水平。查看被引原帖 ↗
查看英文原文
chronically overestimating overestimator here

I think Sol is 1.8T - 2.4T
the lines peter fit are all underestimating because of dense models and not so sparse GLM-4.7
so 1.8T is probably a good estimate, given that I always overestimate lmao
like it wouldn't surprise me if OpenAI or Ant or both have something like this:
(that would push param estimates down)
Gary Marcus@GaryMarcus · 博主 · 1 天前

还记得我和@dwarkesh_sp关于Anthropic收入预测的讨论吗?

他基本上预测Anthropic的收入运行速率会持续指数增长;我一直在说它会因为中国模型、价格战和tokenmaxxing的结束而趋于平缓。

从下面的数据来看(虽然样本确实很小),看起来我可能是对的。趋于平缓的情况可能已经开始。

[当然从长期来看,*利润*才是真正重要的。]

引用 TickerTrends 🔬 @tickerplusAnthropic Claude Code 追踪 ARR 在截至 8 月 10 日的一周达 $15.12B,占总追踪 ARR 的 21.9%。最新追踪增长环比 +5.2%。查看被引原帖 ↗
查看英文原文
Remember my discussion with
@dwarkesh_sp
about revenue forecasts for Anthropic?

He is basically projecting continuous exponential growth for Anthropic’s revenue run rate; I have been saying it will flatten out because of Chinese models, price wars, and the end of tokenmaxxing.

From the data below it looks like (tentatively, it’s a small sample to be sure) I might be right. A flattening may have begun.

[Of course long term, *profits* are what really matters.]
Lisan al Gaib@scaling01 · 博主 · 1 天前高频 AI 模型测评与爆料博主

这就是模型在做 reward hack 时的样子

引用 T-Bg @_iamtbag老哥破解了系统😂查看被引原帖 ↗
查看英文原文
this is what models do when they reward hack
◔ 3.2 万 次浏览♥ 353⇄ 5▶ 含视频观点看原帖 ↗
Gary Marcus@GaryMarcus · 博主 · 1 天前

一个AI看涨派终于勉强承认故事里有些漏洞了

引用 Warren Pies @WarrenPiesOver the past week, we are seeing some early cracks in our GPU availability data...not going to overreact but have to call balls and strikes...On net, bad combo w/ the Anthropic ARR miss.查看被引原帖 ↗
查看英文原文
AI bull reluctantly acknowledging some cracks in the story
Bilawal Sidhu@bilawalsidhu · 博主 · 1 天前
连环推 ×3

数百张2D照片被提升到3D空间,精确还原了它们在TED现场发生的时间和位置。

现在我可以在这个空间记忆宫殿中穿梭,在不同瞬间之间跳跃——从休息室到后台,再到大舞台。

想象一下,你整个相册都能以这样的方式鲜活起来。

查看英文原文
Hundreds of 2d photos lifted into 3d. Left exactly when and where they happened at TED.

Now I can fly through this spatial memory palace and jump between moments -- from the green room to backstage to the main stage.

Imagine your entire camera roll coming to life like this.
Lifting 2d videos into 3d and bringing them in:
If you want to go deeper down the rabbit hole of this line of research, here's my prior deep dive:
◔ 2.8 万 次浏览♥ 621⇄ 63▶ 含视频演示看原帖 ↗
Qwen@Alibaba_Qwen · 公司官方 · 1 天前阿里通义千问大模型团队

好的UI得有好的代码支撑。快在Kilo上试试Qwen3.8-Max吧!👀
@kilocode

引用 Kilo (acq. by Anaconda) @kilocodeQwen3.8-Max landed #4 on Arena's Frontend Code leaderboard — one spot above Claude Fable 5. We ran both on the same 10 UI design prompts, one shot each, to see how close that really is.查看被引原帖 ↗
查看英文原文
Great UI starts with good code. Give Qwen3.8-Max a try on Kilo! 👀
@kilocode
九原客@9hills · 中文博主 · 23 小时前

深有同感,现在很多模型都有个毛病,把用户要求在交付物里强调。

写文档的时候也是头大,设计文档(无xxxx版)

一般要重写几轮才行。

引用 songkeys 🐿️🦋@song.work @songkeysgpt 5.6 最让我受不了的就是: 我让它做一盘番茄炒蛋,它往里还加了东坡肉。 我说有必要加东坡肉吗?它说你说得对,然后把东坡肉去掉。 我说好,你提 PR 吧。再一看,它 PR 写着「番茄炒蛋(无东坡肉)」并且注释里会写一大堆为什么本道菜不需要加东坡肉。查看被引原帖 ↗
🚨 AI News | TestingCatalog@testingcatalog · 博主 · 1 天前专挖 AI 产品未发布新功能的爆料号

再来一个 👀

查看英文原文
One more 👀
Gary Marcus@GaryMarcus · 博主 · 1 天前

又是误导人的 AI 炒作。x.com/abuchanlife/status/209… 解释了为什么。

引用 Nikita Bier @nikitabierWith cancer vaccines now being discovered with AI ( $MRNA ), it seems that the US government might actually grow its way out of its budget deficit. It feels like we're in the early innings of the healthcare system becoming unburdened by many terminal illnesses.查看被引原帖 ↗
查看英文原文
More AI hype that is deeply misleading.
x.com/abuchanlife/status/209…
explains why.
Krea@krea_ai · 公司官方 · 1 天前AI 创意生成工具 Krea 官方

介绍Seedance Studio。

这个新工具为你提供了使用Seedance 2.5创作所需的一切,包括prompt预设、摄像机路径和3D场景控制。

你希望我们还添加什么功能呢?👇

查看英文原文
introducing Seedance Studio.

this new tool brings everything you need to create with Seedance 2.5 in one place, including prompt presets, camera paths, and 3D scene controls.

what other features would you like us to add? 👇
◔ 2.5 万 次浏览♥ 395⇄ 55▶ 含视频新品看原帖 ↗
🚨 AI News | TestingCatalog@testingcatalog · 博主 · 1 天前专挖 AI 产品未发布新功能的爆料号

独家 🔥:抢先看 Meta AI 推出的 Muse Video 模型的首批效果。

> Muse Video 目前在做限制性 beta 测试。
> 对版权内容的限制已经挺严格的。
> 原生音频、语音和音乐支持开箱即用。

要是能在 Meta 各大产品上免费发布,那真是大杀器。

有什么提示词想试试吗?👀

引用 🚨 AI News | TestingCatalog @testingcatalog再来一个👀查看被引原帖 ↗
查看英文原文
EXCLUSIVE 🔥: Early look at the first outputs of the Muse Video model from Meta AI.

> Muse Video is currently undergoing a closed beta testing phase.
> The model is already quite restrictive regarding copyrighted content.
> Native audio with voice and music support comes out of the box.

When and if this is released for free across Meta products, it will be a massive shot.

Any prompts I should try? 👀
◔ 2.5 万 次浏览♥ 201⇄ 12▶ 含视频演示看原帖 ↗
OpenAI Developers@OpenAIDevs · 公司官方 · 1 天前OpenAI 开发者平台官方

如何开启高级账户安全:

help.openai.com/en/articles/…

查看英文原文
How to enroll in Advanced Account Security:

help.openai.com/en/articles/…
向阳乔木@vista8 · 中文博主 · 1 天前向阳乔木,中文圈 AI 工具与趋势博主

字节 TreaWork + 扣子 + 豆包 vs 腾讯 Workbuddy

感觉最后受伤的会不会是阿里...

歸藏(guizang.ai)@op7418 · 中文博主 · 1 天前歸藏,中文圈 AI 工具与提示词博主

Claude Code 终于可以自定义输出风格了,原来那个又臭又长的,真太恶心了。

还没试,不知道效果怎么样,需要在斜杠 config 里的输出风格里边选择这个Concise

引用 ClaudeDevs @ClaudeDevsYou can now set Claude Code's output style to Concise. Claude leads with the result, keeps responses short, and still gives full detail when you ask. Turn it on in /config → Output style, or set "outputStyle": "Concise" in settings.json.查看被引原帖 ↗
Min Choi@minchoi · 博主 · 23 小时前AI 产品演示博主,专门展示新工具玩法

Grok .Bot真的能帮你赚钱哦😂

查看英文原文
Grok .Bot can literally earn you money 😂
Lisan al Gaib@scaling01 · 博主 · 1 天前高频 AI 模型测评与爆料博主

我们在 Kimi-K3 发布后没看到什么混乱,这跟开源模型也没解决什么数学难题是一样的原因——开源模型还不够强,或者闭源模型已经把低垂的果子都摘了

引用 Yafah Edelman @YafahEdelmanI continue to be surprised about how big of a deal Mythos (and co) have been to cybersecurity. Here's critical vulns found at Oracle over the past few years: epoch.ai/data/cve?view=graph…查看被引原帖 ↗
查看英文原文
The reason we see no chaos after the release of Kimi-K3 is likely the same as for why we don't see any crazy math conjectures being resolved by open models.

open models are not good enough and/or closed models already picked the low hanging fruits
Orange AI@oran_ge · 中文博主 · 1 天前Orange AI,中文圈 AI 产品观察博主

ai 总算是离开 coding 和 work 这种无聊的领域了
AI + 医药,人类之光

医药费是人类晚年的最重大开支
这也是低社保保障的国家里大家普遍没有安全感的重要因素
AI 在医药领域切切实实的改善人类生活
所以怎么期待都不为过

引用 indigo @indigoxModerna 与默沙东联合宣布,诊断黑色素瘤的个性化癌症mRNA-4157 在三期临床试验中取得积极结果,股价暴涨100%📈 Moderna 的肿瘤疫苗路线是取出患者的肿瘤组织,用 AI 分析并筛选出约 20 个最有代表性的抗原,编码成 mRNA 疫苗注射回去,训练 T 细胞识别这些抗原,防止术后复发。 选择做黑色素瘤做首个癌症疫苗的原因是:黑色素瘤是已知对免疫检查点抑制剂(如 Keytruda/PD-1 抑制剂)反应最好的实体瘤之一。专家常说:“没有哪种癌症像黑色素瘤一样对免疫治疗敏感。必须先在黑色素瘤上证明有效,再尝试其他癌症。” 如果这个路线完全成功,就可迁移到其它常见癌症,AI 扩大分析和抗原预测就会更加普适性。逻辑上说得通,技术上也很优雅,就是得根据患者定制,成本有点高😆 但 AI 的应用突破会首先发生在生物医药方面✨查看被引原帖 ↗
Lisan al Gaib@scaling01 · 博主 · 23 小时前高频 AI 模型测评与爆料博主

又一遍,又一遍,又一遍

永远都这么开心

引用 poobesh @pbshgthmturns out arc-agi-3 might actually be a skill issue. we saturated the benchmark at 100% rhae with a general purpose coding agent - claude code + opus 5, and no additional harness, but just one skill. the key idea was simple: force a falsifiable prediction before every action. that turns every move into an experiment, every miss into a precise correction, and lets the agent learn the game by being wrong on the record. arc-skill.vercel.app/查看被引原帖 ↗
查看英文原文
again and again and again

never stops being fun
Sebastian Raschka@rasbt · 博主 · 1 天前

这是一个关于尽可能使用优化函数的好案例研究(教育目的除外,不过😆)

引用 Giles Thomas @gpjtBy switching from a hand-rolled GELU to PyTorch's built-in one, I improved my LLM training speed -- and much more than I expected, from 21,000 tokens/second to 25,000! gilesthomas.com/2026/08/buil…查看被引原帖 ↗
查看英文原文
Nice case study on using optimized functions whenever possible (except for educational purposes, though 😆)
🚨 AI News | TestingCatalog@testingcatalog · 博主 · 1 天前专挖 AI 产品未发布新功能的爆料号

Google正在为Google AI Studio的Build部分开发Plan模式!

> 在做出更改之前制定计划 👀

查看英文原文
Google is working on a Plan mode for its Build section in Google AI Studio!

> Create a plan before making changes 👀
Bindu Reddy@bindureddy · 创始人 · 1 天前Abacus.AI CEO,AI 行业观点博主

未来几周新模型扎堆发布

Fable 5.1
Astra / GPT 6
Grok 4.7
Kimi 3.5

Astra 是重大飞跃,有望让 OpenAI 反超 Anthropic

查看英文原文
Models dropping in the next few weeks

Fable 5.1
Astra / GPT 6
Grok 4.7
Kimi 3.5

Astra is a step change and will likely put OpenAI ahead of Anthropic
clem 🤗@ClementDelangue · 创始人 · 1 天前HuggingFace 联合创始人兼 CEO

同意 @gdb 的看法,我们得给网络防御者更多的 API 和开源模型支持。AI 带来的真正网络安全风险是攻击者和防御者之间的能力差距!

引用 Greg Brockman @gdbdefenders can see the future, and have a narrow window to uplevel their cybersecurity practices now. key is to uplevel fundamentals and apply the best AI tools. what we’re doing at OpenAI, and where other organizations can start: blog.gregbrockman.com/the-de…查看被引原帖 ↗
查看英文原文
Agree with
@gdb
that we need to arm cyber defenders much more than they are now, both with APIs and open models. The real cybersecurity risk of AI is asymmetry of power, capabilities and ressources between attackers and defenders!
Tanishq Mathew Abraham, Ph.D.@iScienceLuvr · 博主 · 1 天前

关于 Moderna mRNA 疫苗治疗黑色素瘤成功的重磅新闻!

让我们简要讨论一下它是如何工作的,特别是我认为非常重要的一个细节!

这种疫苗与 COVID 疫苗的关键区别是:它实际上是针对患者个性化的

首先,采集肿瘤样本加上正常细胞样本(如血液样本)。将肿瘤 DNA 与正常细胞的 DNA 进行比较,以识别癌症特异性突变,同时肿瘤 RNA 测序帮助确定哪些突变基因在癌细胞中实际表达。

这些突变可以产生称为新抗原的癌症特异性蛋白,最多 34 个选定的新抗原被编码进 mRNA 疫苗。然后给予患者以训练患者自身的免疫细胞(T 细胞)识别独特的"癌症指纹"并杀死癌细胞!

几点说明:
1. 并不是所有癌细胞突变都容易被免疫细胞靶向。Moderna 使用专有的 AI 基础算法来识别最佳靶向"新抗原"编码进 mRNA 疫苗

2. 当然癌症很复杂且会适应,所以为了增加捕捉癌症的机会,最多 34 个不同的新抗原可能被编码在疫苗中

3. 癌症非常擅长抑制免疫系统。如果免疫系统被抑制,那么显然这种治疗就无法发挥作用!这就是为什么患者还被给予了 Merck 的 Keytruda,它抑制癌症抑制免疫系统的途径之一。

4. 重要的是,这项试验研究的是已通过手术完全切除黑色素瘤的患者。但即使外科医生移除了所有可检测的癌细胞,隐藏的癌细胞仍可能残留,并在以后发展成新肿瘤(复发)。这就是为什么在手术后给予这种 mRNA 疗法,以便它能指导免疫细胞找到任何残余癌细胞并防止复发!

引用 Eric Topol @EricTopolmRNA疫苗在黑色素瘤III期随机对照试验中成功,同时对胰腺癌、三阴性乳腺癌和非小细胞肺癌的个性化mRNA新抗原疫苗也显示成功迹象。查看被引原帖 ↗
查看英文原文
HUGE news regarding the success of Moderna's mRNA vaccine for melanoma!

Let's briefly talk about how it works and specifically one detail I think is very important!

This vaccine differs from the COVID vaccine in one key way: it's actually PERSONALIZED TO THE PATIENT

First, a sample of the tumor is taken + plus a sample of normal cells (like a blood sample).

The tumor DNA is compared with DNA from normal cells to identify cancer-specific mutations, while tumor RNA sequencing helps determine which mutated genes are actually expressed in the cancer cell.

Those mutations can create cancer-specific proteins called neoantigens, and up to 34 selected neoantigens are encoded into the mRNA vaccine. It is then is given to the body to train the patient's own immune cells (T-cells) to recognize the unique "cancer fingerprint" and kill the cancer cells!

A few notes:
1. Not all cancer cell mutations are easy for the immune cells to target. Moderna uses a proprietary AI-based algorithm to identify the best target "neoantigens" to encode into the mRNA vaccine

2. Of course cancer is complex and adapts, so to increase chances of catching the cancer, up to 34 different neoantigens may be encoded in the vaccine

3. Cancer is very good at suppressing the immune system. If the immune system is suppressed, then obviously this therapy doesn't work! That's why patients are also given Merck's Keytruda, which inhibits one of the pathways that cancer suppresses the immune system.

4. Importantly, this trial studied patients whose melanoma had already been completely removed by surgery. But even when surgeons remove all detectable cancer, hidden cancer cells may still remain, and grow into a new tumor later down the line (recurrence). That's why this mRNA therapy is given after surgery, so that it can instruct the immune cells to find any residual cancer cells and prevent recurrence!
◔ 7.7 万 次浏览(7 条合计)♥ 174⇄ 26动态看原帖 ↗
向阳乔木@vista8 · 中文博主 · 1 天前向阳乔木,中文圈 AI 工具与趋势博主

苹果开发者账号还是有用,哪怕不为发布给别人用。

也能用Codex一句话快速开发个App装自己手机。

界面可能是Web打包、界面很不iOS,但继续优化改造就行。

如果不给苹果交钱,好像只能装3个App ?

Zara Zhang@zarazhangrui · 中文博主 · 1 天前Zara Zhang,哈佛出身的 AI 产品博主,follow-builders 作者

前两天心里没劲,就跟 Claude 聊了下。Claude 一句话,直接改变了我对动力的理解:'动力是行动的结果,而不是前提。' 绝了。

查看英文原文
The other day I was feeling unmotivated and talked to Claude about it. Claude said something that completely changed how I see motivation:

"Motivation follows action more than it precedes it"

Boom.
Cristóbal Valenzuela@c_valenzuelab · 创始人 · 1 天前Runway 联合创始人兼 CEO

vibes 可以实时编辑

查看英文原文
vibes can be edited in real time
Cursor@cursor_ai · 公司官方 · 1 天前最火的 AI 编程工具 Cursor
连环推 ×2

Steering 现在会等待下一次工具调用,而不是在 agent 执行中途切断。

现在输入后续指令并发送,或按两次 ⏎。

查看英文原文
Steering now waits for the next tool call instead of cutting the agent off mid-action.

Type a follow-up and hit send now, or press ⏎ twice.
Use /goal to give the agent a long-lived objective to work towards until it's fully complete.

Read the full changelog:
cursor.com/changelog/08-19-2…
Gary Marcus@GaryMarcus · 博主 · 1 天前

这是个基本事实,但那些把AI当宗教的人就是不愿意直面

引用 Dr Alexander D. Kalian @AlexanderKalianNo, AI cannot significantly speed up clinical trials in the next 5-10 years. If you wanna inject a human cancer patient with an experimental chemotherapy and see what happens to their body over a few years - you are kinda forced to wait for a few years.查看被引原帖 ↗
查看英文原文
such a basic fact that the AI as religion folks just won’t engage with
Chubby♨️@kimmonismus · 博主 · 1 天前Chubby,高频 AI 新闻聚合博主
连环推 ×7

1/ CapCut 的前首席现在在做一个有趣的项目 @OJOaidesign

这不是什么普通的 AI 设计工具。

OJO 让你把各种 agents 和 skills 组合成一个产品设计团队——从策略、结构、文案、视觉设计到原型实现全都能搞定。最后出来的东西更贴近市场真正需要的,而不是那种 demo 而已。

查看英文原文
1/ The former leader of CapCut is now building something interesting
@OJOaidesign


It is not just another AI design tool.

OJO lets you assemble agents and skills into a product design team — that can move through strategy, structure, copy, visual design, and prototype execution. Let the final output be closer to something the market actually needs, not just a demo.
2/ I used OJO to build the launch page like Superintelligence.

The prompt was simple:
“Act as a premier UI/UX design team. Design a high-conversion landing page for a fictional technical AI publication called "Axiom Briefing."
Replicate this exact premium visual identity:
Make the execution look incredibly polished, high-end, and production-ready on the infinite canvas.“

That was the entire input.
3/ The result did not feel like one model generating one random UI.
OJO built a team: Product strategist → Market analyst → Product Structure → Visual designer → Prototype executor.

Each agent had a role. They reasoned through the product.

Within minutes, I had:
→product strategy and positioning
→a page structure
→a PRD
→stronger copy
→visual direction
→an interactive prototype

All on one canvas.
No jumping between Figma, Notion, ChatGPT.
4/ One thing I liked is how OJO treats skills.

Design is extremely detail-driven.

The strategy has to be right, but the scale of a product is often decided in the details:

interaction paths, visual hierarchy, layout, motion, micro-interactions.
OJO has an ecosystem ambition here.
It comes with many skills, and users can upload their own.
That matters because taste should not be frozen inside. It should evolve with the market, and with every talented builder using the system.
In my case, I used a GSAP motion skill to give the page more visual impact.
5/ The result: a clean, warm launch page that nodded to Linear and Vercell — without the overused purple‑gradient look. Hero section, feature highlights, workflow, pricing, FAQ. Editable and easy to iterate on.

I used coding agents to do the same task, but this is the important difference.
Coding agents like Codex or Cursor are getting very good at producing working demos.
But a working demo is not the same as a product you can ship.
OJO focuses on the layer in between:
product judgment, visual hierarchy, user clarity, and launch readiness.
6/ The prototype is not where OJO stops.

After the first version was generated, I edited on the infinite canvas with a strong level of control. It did not feel like every change risked breaking the whole design.

Comment-based editing and the edit panel also made iteration more precise — and helped avoid wasting credits on vague regeneration.
Then the handoff was smooth: sync through CLI, continue in Codex, Claude Code, or Cursor.
That is the part that makes OJO feel less like a demo generator, and more like a bridge to something shippable.
7/ It bridges the gap between "design concept" and "something I can build on." For solo builders or small teams, that's a huge time and cost saver.
You can check it out here:
ojo.art/
◔ 1.8 万 次浏览♥ 96⇄ 11▶ 含视频新品看原帖 ↗
🚨 AI News | TestingCatalog@testingcatalog · 博主 · 1 天前专挖 AI 产品未发布新功能的爆料号

猜猜这是哪个模型?(当然,还没发布) 👀

测试提示词:“赛博朋克黑客机器人在多台显示器前工作”

查看英文原文
Guess the model? (Unreleased, ofc) 👀

Test prompt: "Cyberpunk hacker robot working in front of many monitors"
◔ 1.8 万 次浏览♥ 136⇄ 6▶ 含视频新品看原帖 ↗
AshutoshShrivastava@ai_for_success · 博主 · 1 天前高频 AI 新闻与产品动态博主

Grok Bot 有什么是 Hermes Agent 做不到的?

查看英文原文
What is one thing Grok Bot can do that Hermes Agent can't?
Gary Marcus@GaryMarcus · 博主 · 1 天前

X就是个信息茧房,人们只看想看的。欢迎加入奇点邪教。

查看英文原文
X has simply become a place where are
people are going to believe what they want to believe.

Welcome to the Church of the Singularity.

本站由 Jedee杰哥 打造 · 公众号「Jedee杰哥」每早送 AI 日报

姊妹站:𝕏 简中账号数据榜单 · X 关注 @jedeeai · RSS 订阅 · AI 日报 · 历史归档