JEDEE AI
存档 2026-08-23

8 月 23 日(北京时间)全球 AI 圈推文存档,按曝光排序,共 100 条。
← 返回最新 全部归档

全部情报 每小时更新 · 事件已合并同类项

内容 公司
el.cine@EHuanglu · 博主 · 1 天前

中国的AI机器人刚刚达到了14.5米/秒的速度。

查看英文原文
china’s AI robot just hit 14.5 m/s
◔ 1352.4 万 次浏览♥ 9,591⇄ 840▶ 含视频动态看原帖 ↗
Tibo@thsottiaux · 公司官方 · 1 天前ChatGPT 产品官方账号
连环推 ×2

Codex 速率限制更新

我们发现了几个问题:(a) 长会话中使用图像时存在一些低效 (b) Computer History 的 p95+ 使用率过高 (c) 一个原本用来生成对话标题的功能消耗的额度比预期要多。现在有个虎队正在彻查所有问题,明天会推出修复。同时我们还发现了一个完全不相关的新方法来显著提升效率,下周会着手处理。

作为明天部分修复的一部分,我们也会为所有付费订阅重置使用额度。

明天见。

引用 Tibo @thsottiauxUpdate on rate limits in Codex. We do see that for some users the cache hit rate has been worse this week than the stable state the weeks before. This could explain that usage is draining somewhat faster for those users as hitting the cache consistently is an important component of being efficient. We are investigating and will have an update tomorrow.查看被引原帖 ↗
查看英文原文
Update on rate limits in Codex.

We’ve found (a) some inefficiencies when using images in long sessions with multiple compactions (b) high p95+ usage for Computer History (c) a feature that was meant to generate conversation titles that was draining a bit more usage than intended. And we have a tiger team combing through everything and shipping fixes tomorrow. We also found a novel approach to drive efficiency up significantly that is completely unrelated and we will be working on next week.

As part of some of the fixes tomorrow, we will also do a full reset of the usage for all paid subscriptions.

See you then.
Reset will land around 14pm PST tomorrow.
Lisan al Gaib@scaling01 · 博主 · 1 天前高频 AI 模型测评与爆料博主

Opus 5受到这么不尊重,真是离谱。

引用 Theo - t3.gg @theo当前所有主要模型的粗略分级列表。查看被引原帖 ↗
查看英文原文
the Opus 5 disrespect is insane
el.cine@EHuanglu · 博主 · 1 天前

中国 AI 机器人一年进展

查看英文原文
1 year progress of china's AI robot
◔ 39.7 万 次浏览♥ 2,580⇄ 368▶ 含视频演示看原帖 ↗
Chubby♨️@kimmonismus · 博主 · 1 天前Chubby,高频 AI 新闻聚合博主
连环推 ×2

对Boris Cherny怀有敬意,但这段回应就是问题的一部分。

对Opus 5的抱怨已经变得相当一致:懒散、马虎、啰嗦。这跟我自己的体验一致。说着“大家根本没搞懂正确的使用场景”这样的话,完全没抓住重点。

看到Anthropic这样对待用户反馈而不是认真对待,真让人失望。

引用 Boris Cherny @bchernyAgree. People are sleeping on using Opus to hill climb. We use it for optimizing CPU and memory, optimizing CI times, improving frame rates, reducing latency, any other kind of problem in the shape of “iterate on X with a profiler and dataset until it hits Y”查看被引原帖 ↗
查看英文原文
With all due respect to Boris Cherny, this response is part of the problem.

The complaints about Opus 5 have become pretty consistent: it’s lazy, sloppy, and verbose. That matches my own experience. Responding that people simply haven’t understood the right use case misses the point.

It’s disappointing to see Anthropic frame user feedback this way instead of taking it seriously.
to be clear: I have great respect for Boris, and he is certainly doing outstanding work. My criticism is directed solely at the fact that, as I understand it, he does not take the criticism of Opus seriously.
Lisan al Gaib@scaling01 · 博主 · 1 天前高频 AI 模型测评与爆料博主

Opus 5是我们现在能接触到的、最接近Anthropic内部模型的东西

这东西强得离谱,要想优化什么

给它个目标它就能一直优化

引用 Lisan al Gaib @scaling01对Opus 5的不尊重令人无语。查看被引原帖 ↗
查看英文原文
Opus 5 is the closest thing we have to whatever Anthropic has internally

it's an absolute monster if you want to optimize anything

give it a target and it will hillclimb
Min Choi@minchoi · 博主 · 1 天前AI 产品演示博主,专门展示新工具玩法
连环推 ×11

Grok Bot 太疯狂了。

大家停不下来地发掘用它的新方式——赚钱、自动化工作、运营生意。

10 个狂野的例子:

查看英文原文
Ok Grok Bot is insane.

People can’t stop finding new ways to make money, automate work, and run businesses with it.

10 wild examples:
1. Grok Bot agents can now design, test, and manufacture real objects end to end 🤯
2. Grok Bot learned how I work just by watching a screen recording 👀
3. Grok Bot agent turned The Odyssey into a complete visual production
4. Grok Bot AI chief of staff learned an entire business in less than a week
5. Give your Grok Bot its own email and it can start working without you
6. $40/month Grok Bot now handles inboxes, invoices, appointments, and ops
7. Build an Grok BotI hedge fund research desk that works while you sleep
8. This Grok Bot trading bot is beating the S&P and improving itself
9. Six Grok Bot agents can run a Wall Street-style research desk overnight
10. Master Grok Bot in 15 minutes with this complete agent guide
◔ 37.7 万 次浏览(2 条合计)♥ 1,855⇄ 127▶ 含视频演示看原帖 ↗
Zara Zhang@zarazhangrui · 中文博主 · 1 天前Zara Zhang,哈佛出身的 AI 产品博主,follow-builders 作者

有个现象就是,有才华的人单干的时候,AI 能让他们发挥 10 倍的潜力

但这个人一进大公司,最多也就能发挥潜力的 20%(有时候还会下降)

这就是为什么我看到越来越多有才华的人都在离开大公司。(唯一的例外可能是 OpenAI/Anthropic 这样的顶级 AI 实验室)

引用 Stew Fortier @stewfortierI asked someone this week why they left Google. They told me they had never seen such talent-dense teams produce so little查看被引原帖 ↗
查看英文原文
There’s a phenomenon where talented individuals can achieve 10x their potential thanks to AI when working on their own thing

But when the same individual is put into a large organization, they at most increase their potential by 20% (and sometimes it’s even decreased)

This is why I’m seeing more and more talented people leave large companies. (The only exceptions are probably top AI labs like OpenAI/Anthropic)

现在最大的问题是,一级市场还有至少20个宇树科技等着上市,每个目标都是1000亿左右,都是看到宇树科技IPO一路绿灯才融资跑出来的。

现在宇树科技4000亿,后面还有1万亿的这种垃圾具身智能公司等着套现和暴雷。

我看看工信部和科技部这次怎么收场,这场闹剧到底会让多少人倾家荡产,死无葬身之地。

Bilawal Sidhu@bilawalsidhu · 博主 · 1 天前

当比特可以模拟原子时,物理和数字世界之间的界线就变得模糊了

引用 Ingi Erlingsson 🪄 @ingi_erlingsson虚假现实查看被引原帖 ↗
查看英文原文
When bits can simulate atoms the lines between the physical & digital world get rather blurry
◔ 25.1 万 次浏览♥ 4,420⇄ 175▶ 含视频观点看原帖 ↗
Guillermo Rauch@rauchg · 创始人 · 1 天前Guillermo Rauch,Vercel 创始人兼 CEO

今天是 Vercel AI Gateway 开源模型 token 占比创纪录的一天。

8 月 22 日(今天)
🟦 开源:62%
🟨 闭源:38%

6 月 24 日(约 2 个月前)
🟦 开源:28.4%
🟨 闭源:71.6%

这很可能只是个开始,因为企业采用还处于早期,harnesses、CLIs、IDEs、SDKs 等还需要适配成模型无关的方案。

查看英文原文
Today's a record day for open weight share of tokens on Vercel AI Gateway.

Aug 22 (today)
🟦 Open: 62%
🟨 Closed: 38%

Jun 24 (~2 months ago)
🟦 Open: 28.4%
🟨 Closed: 71.6%

This is very likely just the start, because enterprise adoption is still early, and harnesses, CLIs, IDEs, SDKs, etc need to be adapted to be model agnostic.
Chubby♨️@kimmonismus · 博主 · 1 天前Chubby,高频 AI 新闻聚合博主

Ox Alpha 现在已经由 @davis7 在 DeepSWE 数据集上全面测试过了。整体来看,它的表现和 GPT-5.6 Sol mid 差不多。

不过他说的点挺对的。如果真是 GLM-5.3 Flash 能达到 5.6 Sol mid 的水平,还能在 DGX Spark 上本地跑,那我可不止是满意,这绝对是个超值的选择。

5.6 Sol mid 本地运行,唯一成本就是电费,各种任务随便跑,那真是爽到飞起。这模型 24/7 本地待命:直接改变游戏规则。

引用 Ben Davis @davis7DeepSWE在ox alpha模型运行结果约63%。该模型表现优秀,具独特声音和良好代码质量,但存在死代码和速度感觉缓慢问题。若GLM-5.x flash传言属实将是重大突破。查看被引原帖 ↗
查看英文原文
Ox Alpha has now been fully tested against the DeepSWE set by
@davis7
. Overall, it's performing more or less on par with GPT-5.6 Sol mid.

But he makes a good point. If it really is GLM-5.3 Flash that's now performing at the level of 5.6 Sol mid and could run locally on a DGX Spark, I wouldn't just be satisfied, it would be an absolutely fantastic deal.

5.6 Sol mid running locally, with the only cost being power consumption, for all sorts of tasks would be incredibly great. This model 24/7 hermes local: game changer.
◔ 51.9 万 次浏览(8 条合计)♥ 2,196⇄ 86研究看原帖 ↗
Yuchen Jin@Yuchenj_UW · 博主 · 1 天前

AI不会抹平人与人之间的差距。

YouTube、Coursera 和互联网让世界级知识免费可得,但大多数人还是选了刷 TikTok。

AI 也一样。人人都能用 Claude Code、Codex 和 ChatGPT,但顶尖的人会用得更好,复利更快。

AI 不是均衡器,而是放大器。

查看英文原文
AI won’t erase the gap between people.

YouTube, Coursera, and Internet made world-class knowledge free. Most people still chose TikTok.

AI is no different. Everyone gets Claude Code, Codex, and ChatGPT. The best will use them better and compound faster.

AI isn’t an equalizer. It’s a multiplier.
el.cine@EHuanglu · 博主 · 1 天前

天哪..AI 把火影忍者拍成真人电影了,还比原作好!

查看英文原文
omg.. AI just turned Naruto into live action film, its even better than original!
◔ 19.4 万 次浏览♥ 2,835⇄ 298▶ 含视频演示看原帖 ↗
Bindu Reddy@bindureddy · 创始人 · 23 小时前Abacus.AI CEO,AI 行业观点博主

🚨 一条提示词搞定后端和移动应用

- 构建 iPhone 或 Android 应用
- 后端、支付、认证全免费
- 用病毒式推文和短视频替身吸引客户

几分钟内搭建并创建你的公司

查看英文原文
🚨 Mobile Apps With Backend In One Prompt

- build iPhone or Android apps
- backend, payments and auth for FREE
- get customers use viral-tweet and shorts agents

Set up and build your company in minutes
◔ 18.6 万 次浏览♥ 28⇄ 1▶ 含视频教程看原帖 ↗
Lisan al Gaib@scaling01 · 博主 · 1 天前高频 AI 模型测评与爆料博主

LeChaton 已经在野外现身了,已经开始破坏关键基础设施

引用 Tendencias @TTendenciaX"肥猫":为了这只试图坐在椅子上而被弄坏的小猫查看被引原帖 ↗
查看英文原文
LeChaton has been spotted in the wild and it's already breaking critical infrastructure
◔ 17 万 次浏览♥ 2,688⇄ 82▶ 含视频动态看原帖 ↗
Chubby♨️@kimmonismus · 博主 · 1 天前Chubby,高频 AI 新闻聚合博主

对他们来说可不是好兆头:Anthropic 的 Fable 5 正在被更便宜的开源模型抢走市场。

据 FT 报道,Anthropic 最强的模型在 Ramp 上的企业 AI 开销里只占 11%。

原因很简单:Fable 5 太贵了,而更便宜的模型对大多数任务来说已经够用。

时机也不凑巧,现在外界对 Anthropic 的观感本来就在变差。

查看英文原文
Not looking good for them: Anthropic’s Fable 5 is losing ground to cheaper open-source models:

According to the FT, Anthropic’s most powerful model accounts for just 11% of its corporate AI spending on Ramp.

The reason is simple: Fable 5 is too expensive, while cheaper models are already good enough for most tasks.

Not great timing, with sentiment around Anthropic already deteriorating.
Amjad Masad@amasad · 创始人 · 1 天前Amjad Masad,Replit 创始人兼 CEO

“很快”原来就是三个月。

查看英文原文
“Pretty soon” turned out to be 3 months.

宇树科技上市这件事,冲击的恰恰是传统工业机器人和工控自动化行业。

我周围一大批踏踏实实做PLC的、做工业自动化的、做传感器的、设计出全球性能第一第二指标的工业机械臂的人,为了性能提高30%,或者误差降低30%,每天趴在lab里调试累得跟个王八蛋似的。

就这样,利润率非常低,也不太可能融资, 基本上都沦为老实人制造业。

宇树科技和中国一二十家具身智能大号遥控玩具,恰恰在一级市场PEVC的口袋里吸走了这些制造业的钱,一家公司pre IPO总共融走了千亿,后面一家家公司都在等着,

仅仅是因为宇树科技上春晚,上抖音热搜,被各种吹捧,上了中央企业家座谈会,金监管局和工信部在上市方面一路开绿灯,导致其他PEVC蜂拥而至,一起再造20个王兴兴和宇树科技,一起割韭菜。

这种局面是非常可怕的,因为这一轮泡沫破裂崩溃之后,把所有一级市场高端制造业的血包全吃了,其他老老实实做工业自动化和机械臂的企业融资就会变得非常困难,估计很多企业要面临生死局了。

Matt Shumer@mattshumer_ · 博主 · 1 天前HyperWrite CEO,AI 实战技巧分享

我的 Gmail 因为和 agent 一起使用被封了。

@agentmail 才是正解!

查看英文原文
I got my Gmail banned by using it with an agent.


@agentmail
is the way to go!
Zara Zhang@zarazhangrui · 中文博主 · 1 天前Zara Zhang,哈佛出身的 AI 产品博主,follow-builders 作者

所有在使用 AI 方面领先的人都觉得自己落后了

查看英文原文
Everyone who’s ahead in using AI thinks they’re behind

这些年不断有人问我,“你今天帮人推荐专业,你如何判断未来中国的发展趋势?万一一个专业今天很火,10年后死了怎么办?”

党哥反复强调,判断未来趋势,必须依赖这几个第一性原理:

1. 看中国当今的产业结构。第一大块就是外贸制造业,其中机电产品占比60%,机电产品就是中国第一大挣钱行业。

机电行业对应的就是计算机、弱电类(电子信息自动化)、强电类、机电类专业,这个专业在中国是个大蛋糕,不是越南印度随随便便能全盘抢走的;

2. 看中国最顶级的PEVC真金白银的投资portfolio。目前几乎所有PEVC都在AI行业里面投资,所有清华北大上交PhD几乎都被几家头部LLM炼丹大厂抢走了,

哪怕PEVC成天喊着要看具身智能、量子计算、脑机接口、人造子宫,他们压根也不敢投多少钱,或者完全是找人接盘,最终各家VC自己portfolio里绝大多数依然是AI,毫无疑问,AI就是最好的、增速最快、最有前景的行业;

3. 看美国日本韩国今天的状态。这三个国家今天的状态,就是中国未来10~30年的趋势。很明显,人文社科专业没有饭吃,头部金融咨询保险行业已经饱和,社会消费零售利润越来越低,过度饱和竞争。

最终能吃饭的人,都是EECS本科专业出身、并且能吃技术饭的人,这一点毫无疑问。

指望复刻国内餐饮、房地产、商业地产、零售的时代红利,开个饭馆奶茶店咖啡店躺着数钱,未来已经完全不现实了;

4. 看清华大学本科生们用脚投票的选择。过去10年来,清华本科生们几乎绝大多数都最积极抢计算机专业的名额,进不去计算机的人,也要进电子自动化,抢科技互联网的饭碗,

进不去EECS的人,也要捏着鼻子把清华本科四年专业课修完,然后自己偷偷晚上在宿舍一门一门看CMU、Stanford、MIT的计算机课,刷leetcode,出去狠狠投简历进科技互联网行业,连文科生都知道要先进阿里字节小米做两段产品经理实习。

不要觉得他们是傻子,他们用脚投票的结果,才是最真实的结果;

5. 看中国方针政策选行业。国务院在报告里可以推荐6G、生物技术、量子计算等等名词,然而国家级产业基金和各地方基金真金白银孵化扶持的产业,一定是半导体、新能源、医疗器械、靶向药、高端制造业、国产替代这些行业。

因为党哥反复强调,在几轮半导体卡脖子事件后,中央的一个大政方针就是国家安全,国家安全最重要的就是产业安全,就是半导体产业安全,所以中国一定会不惜一切代价、用30~50年的时间,尝试逐步彻底解决半导体全产业链国产化的问题。

同时中国产业发展趋势一定会模仿日本韩国,能吃掉的半导体、通信、汽车、消费电子、工控设备产业,一定本国保护、本国孵化、本国扶持,这块蛋糕不需要顶级创新力,只要按部就班一点点让出市场,一点点本土产业吃掉市场,就能把产业留在本地,然后进入国际市场参与竞争,

同时中国因为拥有日本和韩国都比不了的14亿人口巨大市场,所以美国和欧洲对应的优势产业,在中国本土也可以孵化出来,比如高端仿制药、自动化无脑找靶点的创新靶向药、生物医疗器械、主线飞机、高铁、新能源车、光伏这些产业,靠14亿人市场就能全部消化;

6. 还有一大块蛋糕,是国企用牌照和政策垄断的自留地,典型的就是通信、电力电网、石油、国字号金融保险、各地方自留的文旅城投产业、道路桥梁产业、烟草产业等等,有的有增量,有的快死了,但整体都是国企饭,有关系就能进去。

这就是党哥判断中国未来优势的基本第一性原理,学会了请转发+点赞+收藏,谢谢。

Greg Brockman@gdb · 创始人 · 1 天前Greg Brockman,OpenAI 联合创始人兼总裁

agentic 采用的速度快得出奇,容易忘记这个领域已经走了多远。

引用 Florian Brand @xeophon一年前,一年10B个token还是大成就,会得到奖牌。现在我一周就处理超过10B个token。查看被引原帖 ↗
查看英文原文
agentic adoption has been super fast, easy to forget how far the field has come
小互@xiaohu · 中文博主 · 1 天前小互,中文圈高频 AI 资讯站 Xiaohu.AI 主理人

阿里全球通 AliExpress 使用了一种听起来非常间谍的技术来监控用户

它没有偷偷使用麦克风偷录你的声音,

而是利用你的设备生成了一段类似指纹的音频内容来追踪你。

起因是很多用户发现:一旦他们打开 AliExpress 首页,哪怕网页没放视频、没播放声音、甚至把标签页静音,手机的音乐都会立刻中断。(AirPods 自动连接)

只有关掉 AliExpress 标签页,手机音乐才会恢复。

它使用了一种超广谱的信息搜集技术

原理是:不同操作系统、浏览器版本、声卡硬件与音频底层驱动,在处理同一段数学波形时会产生极其微小的数学差异。

当你这些细微差异组合起来,就是一台设备的独特特征,类似人类指纹...

通过这个唯一的设备AliExpress用来追踪客户行为。

为什么费这么大劲,搞这种类似间谍的尖端技术?

因为 Cookie 可以被删除,也越来越容易被浏览器拦截。

而这种类似的设备指纹即使清空记录、重新打开浏览器,他们仍然仍有可以知道这是那一台台设备 ,也就是知道你是谁!

为什么会干扰用户音乐播放?

因为合成的这段音频是为了做音频指纹

虽然人耳听不到任何声音,但当用户浏览网页的时候,它偶尔会播放这个音频来确认用户身份,浏览器和系统认为网页正在“处理实时音频”,从而一直占用电脑的音频通道,导致蓝牙耳机认为电脑在“发声”,拒绝切换回手机。

导致你的耳机被切回到电脑,手机蓝牙立刻被切断...

Bindu Reddy@bindureddy · 创始人 · 1 天前Abacus.AI CEO,AI 行业观点博主

预测:OpenAI 要是把 GPT 5.6 SOL 降价 80%

使用量和采用率就会暴涨

Anthropic 和开源 AI 的竞争力也就没了

查看英文原文
PREDICTION- OpenAI will slash GPT 5.6 SOL’s prices by 80%

This will sky rocket usage and adoption

It will also wipe out competition from Anthropic and open source AI
Together AI@togethercompute · 公司官方 · 1 天前

GLM-5.3 四轮对标测试全胜 Fable 5,无论求解率还是成本都更优。

在 DeepSWE 上,GLM-5.3 花 $16 左右就能跑到 87.6%,而 Fable 5 需花 $21.63 才能达到 69.7%。

查看英文原文
Four tries with GLM-5.3 beat Fable 5 on both solve rate and total cost.

On DeepSWE, GLM-5.3 reaches 87.6% for ~$16, compared with 69.7% at $21.63 for Fable 5.
◔ 11.7 万 次浏览(3 条合计)♥ 351⇄ 21▶ 含视频研究看原帖 ↗
Gary Marcus@GaryMarcus · 博主 · 1 天前

@sama
在说自己对AI被采用的速度判断有误,但就是不肯承认他对AGI何时到来的预测错了。

可他确实是错了。他曾在2025年1月信誓旦旦地说“我确信我们已知道如何按传统理解来构建AGI”,但19个月过去了,我们连目前手头的AI都还难以真正信任,即便最强模型也是如此。

一旦AGI真来了,采用率自然会高得多。

但那还要等上许多年,还有多重突破翻越。

真正的问题,Altman避而不谈的是这一点。
——
*AI目前最多也只能完成我与Miles Brundage公开打赌时提出的十个标杆目标中的两个,这还说得比较客气。[见 garymarcus.substack.com/p/wh…]

引用 Fireside Alpha @firesidealphaSam Altman admits he was wrong on AI's timeline and says society and the economy will adapt more slowly "I thought when we got to GPT-4, which was back in 2023, that very quickly after that there was going to be much more disruption, software businesses up for grabs right away, than it turned out to be." "I think I was wrong about a few things, but one in terms of the speed: the economy just has so much inertia." "People keep doing the same things, buying from the same company, wanting to use their tools the same way. I think this is actually a positive in many ways, and it's going to make this big transition go smoother and slower. I'm grateful for it." "But it means we've all been too ambitious on timelines. Even with this incredible technology, society and the economy will adapt more slowly."查看被引原帖 ↗
查看英文原文
Below
@sama
is saying that he was wrong on how quickly people would adopt AI, but not admitting that he was wrong about how soon AGI would come.

But he was. He literally said in January 2025 that he was “confident we know how to build AGI as we have traditionally understood it“ - but 19 months later we can still can’t really trust the AI we have got, not even the very best models.*

Adoption will MUCH higher once AGI actually arrives.

But that’s still many years and multiple breakthroughs away.

That’s the real issue, that Altman doesn’t want talk about.
—–
*AI can still do at most only two of the ten sample targets that I proposed in a public with bet Miles Brundage. And that’s being charitable. [See
garymarcus.substack.com/p/wh…
]
◔ 9.9 万 次浏览♥ 302⇄ 44▶ 含视频观点看原帖 ↗
Chubby♨️@kimmonismus · 博主 · 1 天前Chubby,高频 AI 新闻聚合博主

Anthropic 的 Thariq 说 Claude Code 的 "high = 10" 输出是数值映射问题,不是隐性降级。

"High" 看起来依然是高的,内部评估显示没有性能下降。

感谢澄清。但我还是要说,目前大多数模型用起来明显变笨了。尤其是 Opus。

引用 Thariq @trq212Claude测试不同API配置,其中一个改变了努力值映射方式。新系统可能显示'10'代表'高',但数值本身无意义,用户选择的努力等级保持不变。已通过评估确认不影响模型性能。如遇问题请反馈。查看被引原帖 ↗
查看英文原文
Anthropic Thariq says Claude Code’s “high = 10” output was a numerical mapping issue, not a stealth downgrade.

“High” is apparently still high, and internal evals show no performance regression.

Thanks for the clarification. But I still have to say that most models currently feel significantly dumber to use. Especially Opus.
Guillermo Rauch@rauchg · 创始人 · 1 天前Guillermo Rauch,Vercel 创始人兼 CEO

我用 fx 构建了 Mini,一个个人网页浏览器。

很多时候我需要屏幕共享、直播或测试东西,不想要浏览器的"臃肿"和状态。所以就用 Mini。

这对 fx.sh 和开源模型来说是个很好的测试。这是我第一次真切感受到这些东西有多便宜,又有多好用。

"个人软件"现在很火。"我花了 $25,000USD 来取消我的 $10 XYZ 订阅"这事儿是真的😂(尤其现在企业里这种事儿多了去)。

但我觉得开源和本地模型会从根本上改变这个局面。软件真的会变得很自发。

ps:向 Orion 浏览器致敬,它的"专注模式"激发了我。我想让我的浏览器更最小化、更可 hack、跑在 Chromium 上。

查看英文原文
I used 𝚏𝚡 to build Mini, a personal web browser.

Many times I need to screen share, stream, or just test something without the 'bloat' and state of my browser. So I go Mini.

This was a cool test for
fx.sh
and open models. It was the first time I experienced how cheap, but also how *good* these things can be.

'Personal software' is a hot topic. "I spent $25,000USD to cancel my $10 XYZ subscription" is a real thing 😂 (especially in Enterprises these days).

But I think open and local models will change this dynamic very profoundly. Software will really become spontaneous.

ps: shoutout to Orion browser who inspired me with 'focus mode'. I wanted my browser to be even more minimal, hackable, and run on Chromium.
◔ 9.5 万 次浏览♥ 1,113⇄ 27▶ 含视频演示看原帖 ↗
Chubby♨️@kimmonismus · 博主 · 1 天前Chubby,高频 AI 新闻聚合博主

尽快用完你的配额吧:明天Codex又要reset了。原因是Codex限制消耗这么快是因为长会话中低效的图像处理、超高的Computer History使用量和过度消耗资源的对话标题功能。一如既往,这是OpenAI和@thsottiaux出色的社区工作。这就是我说的社区工作的意思,我希望Anthropic能从中学到。

引用 Tibo @thsottiauxReset will land around 14pm PST tomorrow.查看被引原帖 ↗
查看英文原文
Burn your rates as fast as possible: tomorrow is codex another confirmed reset incoming.

Reason: Codex limits were draining faster due to inefficient image handling in long sessions, unusually high Computer History usage, and an overly resource-intensive conversation-title feature.

As always fantastic community work by OpenAI and our boy
@thsottiaux


That is exactly what I mean by community work, something I hope Anthropic learns from.
Chubby♨️@kimmonismus · 博主 · 1 天前Chubby,高频 AI 新闻聚合博主

英伟达据报投入60亿美元与Poolside合作打造开源AI大模型。英伟达将获得Poolside的技术授权,并纳入Poolside的100多名员工参与Nemotron项目。英伟达还向Poolside投资10亿美元,Poolside估值达到120亿美元。目标是对标中国开源领导者DeepSeek、Kimi,并与OpenAI、Anthropic等美国前沿实验室直接竞争。开源是未来方向,很高兴看到英伟达站在我们这边!

查看英文原文
Nvidia is reportedly spending $6 billion to build one of the world’s most powerful open-weight AI models.

According to the WSJ, Nvidia will license Poolside’s technology and bring more than 100 of its employees into the Nemotron project.

Nvidia is also investing another $1 billion in Poolside at a $12 billion pre-money valuation.

The goal: challenge Chinese open-weight leaders such as DeepSeek and Kimi while competing directly with US frontier labs including OpenAI and Anthropic.

open source is the way to go. So good to see having NVIDIA on our side!
◔ 9.3 万 次浏览(2 条合计)♥ 1,841⇄ 138动态看原帖 ↗
Chubby♨️@kimmonismus · 博主 · 1 天前Chubby,高频 AI 新闻聚合博主

很高兴看到 Anthropic 认真对待这些批评。也很欣慰社区里这么多评论都在提供改进建议。

现在就看
@AnthropicAI
怎么落实了。不幸的是,这是明摆着的事实:Opus 5 并不好,他们自己也清楚。

Thariq 说他们正在努力寻求解决方案,这让我对 Opus 5.1 抱有希望。我觉得更多反馈确实有帮助,而且我有种感觉,Anthropic 已经把这次警告听进去了。

引用 Thariq @trq212感谢反馈。Opus 5确实表现不稳定,我们希望模型保持一致、温暖有Claude风格。这是我们的重大优先事项。查看被引原帖 ↗
查看英文原文
It's good to see that Anthropic is taking the criticism seriously. And I'm also pleased to see how many comments from the community are offering suggestions for improvement.

Now it's up to
@AnthropicAI
to implement them. Unfortunately, it's a truism: Opus 5 isn't good, and they know it.

Thariq says they're working hard on a solution, and that gives me hope for Opus 5.1. I think more feedback can really help, and I have the feeling that Anthropic has heeded the wake-up call.
Lisan al Gaib@scaling01 · 博主 · 1 天前高频 AI 模型测评与爆料博主

10 小时完成的漏洞扫描完全值 $10k

GPT-5.6-Pro 说这比几乎所有人工方案都便宜

引用 Aaron Rubin @aaronrubin用Mythos进行代码库漏洞扫描,耗时10小时,成本$10K。查看被引原帖 ↗
查看英文原文
a vulnerability scan done in 10 hours is well worth $10k

GPT-5.6-Pro says this is cheaper than pretty much any human scenario
Chubby♨️@kimmonismus · 博主 · 1 天前Chubby,高频 AI 新闻聚合博主

如果你还需要最后一个证据证明2026年是智能体之年,那这就来了。

2025年是推理AI之年,2026年则是智能体之年。

引用 a16z @a16z人类是AI的少数用户 Agent消耗的token是人类的近5倍,自2月以来增长14倍 本周图表:a16z.news/p/charts-of-the-we…查看被引原帖 ↗
查看英文原文
If you need one final proof that 2026 is the year of agents, here you have it.

2025 was the year of reasoning AI, 2026 the year of agents.
Dan Shipper 📧@danshipper · 博主 · 1 天前

我们在招聘 @every !

every.to/careers

引用 New York Post @nypostNYC surpasses San Francisco as biggest tech talent hub, study shows trib.al/B2Rsccz查看被引原帖 ↗
查看英文原文
we’re hiring
@every
!


every.to/careers
Gary Marcus@GaryMarcus · 博主 · 1 天前

不同意。Anthropic 还没到“凉凉”那一步,活下来的机会不小。他们人才济济,市场份额也不小,执行力一直不错。

但要是有谁在2万亿美元估值下投钱进去,而高端产品兴趣在降,一堆人又在低价产品上拼命卷价格,那真该去查查脑子了。

而且,在没有补贴的情况下他们能不能盈利,目前都还不清楚。

引用 Ross Hendricks @Ross__HendricksThere’s a word for this, and I believe it’s “cooked”查看被引原帖 ↗
查看英文原文
Disagree. Anthropic is not (yet) “cooked”. They have a decent chance of surviving. They have a lot of talent; they have significant market share. They have executed well.

But anyone who invests money in them at a $2 trillion dollar valuation when interest in their premium product is declining, and many others are undercutting Anthropic’s less expensive products on price, should have their head examined.

It’s still not even clear that they can be profitable in the absence of subsidies.
François Chollet@fchollet · 创始人 · 1 天前
连环推 ×2

社交媒体上垃圾越来越多,全是那些用 AI 生成帖子的博主,机器人跟着瞎回。一层层的回声。

查看英文原文
An increasing fraction of social media consists of slop influencers using AI to make posts and bots replying to them. An echo of an echo of an echo
I find myself even more drawn to books, lately. The pre-2023 kind, where each sentence captures a human thought.
Matt Shumer@mattshumer_ · 博主 · 1 天前HyperWrite CEO,AI 实战技巧分享

每个 agent 公司都应该与 @agentmail 合作。如果一个 agent 默认就有收件箱,能显著降低价值实现的时间并减少客户流失。我可以直接转发邮件、让它注册账号,不用提前赋予访问权限就能开始产生价值。

查看英文原文
Every agent company should be partnering with
@agentmail
.

If an agent has an inbox by default, it dramatically decreases time to value and cuts churn.

I can just fwd stuff, have it sign up for accounts, not have to give it my stuff to start getting value.
Guillermo Rauch@rauchg · 创始人 · 1 天前Guillermo Rauch,Vercel 创始人兼 CEO

AI 能做很多美妙的事情,满足很多愿望,但我们只有一个地球、一个加州、一个巴塔哥尼亚、一个门多萨。美国和阿根廷是仅有的既自由又拥有最壮丽土地和地理风景的两个国家。我看好🇺🇸🇦🇷。我们才刚开始。

查看英文原文
AI can do many wonderful things and grant many wishes, but we have only one earth, one California, one Patagonia, one Mendoza. USA and Argentina are two of the freest countries that also happen to have the most glorious land and geography. I'm very long 🇺🇸🇦🇷. We're so early.

中国科技行业下一轮最大的泡沫破裂,应该就是具身智能+VLA+世界模型,会有几个公司上市后股价跌到零,剩下一批公司应该跟AI四小龙时代的一批调参大师一样突然关停或者暴雷。

虽然具身智能+VLA+世界模型的泡沫破裂,远远比不上2000年的美国dot com bubble互联网泡沫危机,但肯定比AI四小龙要猛烈得多。

方式也跟AI四小龙时代一样,一批大厂senior level到VP分别下场做demo,手搓个垃圾pick and place跑不通的灵巧手,或者全员跳舞,或者弄一堆metaverse元宇宙级别的物理世界仿真4399小游戏,烧完就关门。

这一轮会比以前任何一轮都要惨烈,核心原因就是宇树科技IPO全面开绿灯,马上要冲击上市,这一两年国内PEVC投资人已经全部杀红了眼,随便一个垃圾phd、垃圾教授、垃圾demo都能投进去几个亿,等着一个个也IPO开绿灯,让A股股民接盘。

表面上是科技大泡沫、技术大骗局,实际上本质还是金融监管的巨大漏洞,让宇树科技抢跑上市,导致后患无穷。

clem 🤗@ClementDelangue · 创始人 · 1 天前HuggingFace 联合创始人兼 CEO

最终,绝大多数AI工作负载都将基于开源模型!

引用 Guillermo Rauch @rauchgVercel AI Gateway上开源模型token份额创新高:8月22日开源62%、闭源38%,两个月前仅28.4%。企业采用仍早期,工具生态需适配模型无关架构。查看被引原帖 ↗
查看英文原文
Ultimately the overwhelming majority of AI workloads will be based on open models!
Kol Tregaskes@koltregaskes · 博主 · 1 天前

Codex 重置在太平洋标准时间下午2点,也就是英国夏令时晚上10点。所以是从现在起12小时后。

引用 Tibo @thsottiauxReset will land around 14pm PST tomorrow.查看被引原帖 ↗
查看英文原文
Codex reset coming at 2pm PST, that's 10pm BST. So 12 hours from now.
Tanishq Mathew Abraham, Ph.D.@iScienceLuvr · 博主 · 1 天前

拜托了 求你别再用 GPT-4o 做评估了

除此之外还用 Llama-3 和 Command-R?! 🤮

LLM 虽然有局限性,但我压根信不了那些用这些老模型做的研究。我是说这些模型本身在医学能力上就不怎么样,特别是跟现在的模型一比。

说实话,这项研究完全没用。可惜人们一直拿这些研究来说 LLM 不行。

AAAAHHHHH

引用 Joseph Younis, MD @YounisJosephThis is a well-designed randomized control study that asks the question: Does an LLM's ability to diagnose and manage the same medical scenarios degrade when a patient is the communicator versus a physician. The result: 60% drop in diagnostic accuracy, 12% drop in appropriate management decisions The answer essentially confirms the following suspicions: 1. LLMs are suggestible, naive, and sycophantic 2. Medical literacy drives output; less literacy, less accurate LLM outputs 3. Prompting still confounds frontier models' medical accuracy It is not safe to have people trust LLM medical advice when they're incapable of detecting their own bias and inaccuracies.查看被引原帖 ↗
查看英文原文
PLEASE IM BEGGING YOU TO STOP USING GPT-4O FOR EVALUATIONS

On top of that, using Llama-3 and Command-R?! 🤮

LLMs aren't without limitations, but I can't trust any conclusions made by studies using such old models. I mean these models themselves were honestly not great at medical capabilities to begin with, especially relative to current models.

I'm sorry, but this study is absolutely useless. And it is unfortunate people keep bringing these studies up as evidence LLMs are bad.

AAAAHHHHH
elvis@omarsar0 · 博主 · 1 天前

推荐一个 agent 技能。

/eli5 在可视化深度技术概念方面效果很不错!

我发现它加快了我的思考过程,也帮我更好地与 agents 协作。

在我们的 harness playground 里试试 /eli5 配合 Pi 和 Ox Alpha:
academy.dair.ai/dashboard/pl…

引用 Thariq @trq212Anthropic团队最近经常使用的技能:ELI5 /eli5 <你想解释什么> 用一个不了解这个话题的人能理解的方式来解释,使用HTML工件、大图片和少量文字查看被引原帖 ↗
查看英文原文
Recommended agent skill.

/eli5 works great for visualizing deep technical concepts!

I find it speeds up my thought process and helps me collaborate better with agents.

Try /eli5 with Pi and Ox Alpha in our harness playground:
academy.dair.ai/dashboard/pl…
◔ 4.9 万 次浏览♥ 425⇄ 36▶ 含视频教程看原帖 ↗
Min Choi@minchoi · 博主 · 1 天前AI 产品演示博主,专门展示新工具玩法

咱们完了 💀

中国的"Lightning"人形机器人百米跑了9.32秒。

博尔特的世界纪录:9.58秒

查看英文原文
We are cooked 💀

China's "Lightning" humanoid robot just ran 100m in 9.32 seconds.

Usain Bolt's world record: 9.58 secon
◔ 4.8 万 次浏览♥ 346⇄ 36▶ 含视频动态看原帖 ↗
Gary Marcus@GaryMarcus · 博主 · 1 天前
连环推 ×2

我2012年在《纽约客》写过类似的机器人失败案例。

人们真的不知道现实世界的机器人有多难。

引用 Rohan Paul @rohanpaul_ai机械惨剧。这是2026年北京亦庄人形机器人半程马拉松赛前测试中发生的。查看被引原帖 ↗
查看英文原文
I wrote about robot flops like this in The New Yorker in 2012.

People have no idea how hard robotics is in the real world.
some data (from before last night, which likely added to the unification):
◔ 4.8 万 次浏览♥ 242⇄ 26▶ 含视频观点看原帖 ↗
clem 🤗@ClementDelangue · 创始人 · 1 天前HuggingFace 联合创始人兼 CEO

NVIDIA 自己搭了个编码工具来优化 CUDA GPU 内核,在 ARC-AGI-3 的 25 个公开游戏里拿下了满分,188 个关卡全部通关,180 个全过。(注:应为183个关卡)

有了 agents,我们会从“跑模型、做优化、自己训练 AI 模型和内核超难”的世界,进入一个几乎人人都能做到的时代。

一亿个 AI 开发者这话什么时候能成真?

查看英文原文
NVIDIA built its own coding harness to optimize CUDA GPU kernels and achieved a 100% score on ARC-AGI-3’s 25 public games, solving all 183 levels.

With agents, we'll move from a world where it's quite hard to run, optimize, post-train your own AI models and kernels to a world where virtually everybody can do it.

100 million AI builders when?
宝玉@dotey · 中文博主 · 1 天前宝玉,中文圈 AI 翻译与科普大 V

是这个道理,每个人常用的 agent 就那么两三个,没必要我用个服务或者 App 还要用你的内置 agent,最佳形式是这些服务或 App 提供 MCP 或者 cli,从 agent 里面去访问相应的服务或应用就完了。

引用 Gergely Orosz @GergelyOroszSo many vendors are NOT getting this I have one or two agents I use and like. For anyone else: give me an MCP interface to connect these agents to so I can use your service Unless your a frontier AI lab, I prob don't want to use your agent, sorry查看被引原帖 ↗
Gorden Sun@Gorden_Sun · 中文博主 · 1 天前中文圈高频 AI 资讯与开源项目博主

成功原价续费了,由于玻利维亚货币又贬值了,这次只花了805人民币

引用 Gorden Sun @Gorden_Sun发玻利维亚国难财了,Codex 20x订阅,1380比索,折合人民币845元。 玻利维亚央行结束固定汇率制,转向由央行监控的灵活弹性汇率,导致官方汇率一次性重贬近30%。Google Play在付款时使用玻利维亚货币支付,导致价格实际低得多。 研究AI这几年,我不仅学会了AI,还学会了网络、自建节点、冒充尼哥、美国大兵教师学生、国际货币政策。查看被引原帖 ↗
Ethan Mollick@emollick · 创始人 · 1 天前沃顿商学院教授,AI 应用研究权威
连环推 ×3

生活里充满了难以应对的事——太复杂、设计不好、没人管或特别费时间——医疗、政府、个人理财、学校表单都是例子。也正因为这样,我觉得Consumer AI被严重低估了。人们都是硬撑过去,但他们其实需要的帮助得不到。

查看英文原文
Life is full of things that, by complexity or design or lack of care or required time, are hard to navigate: healthcare, government, personal finance, school forms are all among them

It is why I feel consumer AI is underrated. People muddle by, but need help than they can't get.
For example, I did many years of work with organizations trying to teach financial literacy because so many people unknowingly end up in bad situations. The fact that you now have access to a tool that is, according to research, good at giving financial advice is a big deal.
It is one of the things that I think is being lost in AI polarization. AI models, as they are today, are truly capable of solving some very hard, very real problems. I worry that the groups that most try to help people with these problems are those most rejecting AI reflexively
Lisan al Gaib@scaling01 · 博主 · 1 天前高频 AI 模型测评与爆料博主
连环推 ×2

要是一直关注Anthropic但没抓住投资机会,这个数字真的很扎心

引用 Andrew Curran @AndrewCurran_Anthropic银行家向投资者表示,公司IPO融资可能超1000亿美元,估值达2万亿,将超越SpaceX,成为历史最大IPO。查看被引原帖 ↗
查看英文原文
2 trillion hurts so much if you have been following Anthropic for that long but had no opportunity to invest
would've been nice to escape the underclass

instead only VCs, funds, other institutions and employees made bank
向阳乔木@vista8 · 中文博主 · 1 天前向阳乔木,中文圈 AI 工具与趋势博主

Type-C连上电脑,让Codex修改TRAE AI passport为背单词工具

能显示中文,还有发音,可玩性真不错!

ollama@ollama · 公司官方 · 1 天前本地跑大模型的热门工具

和 @poolsideai 工程师合作太棒了,一起发布了开源模型。对开源模型的前景充满期待,也很期待 @nvidia 的 Nemotron 项目。

引用 The Wall Street Journal @WSJ与初创公司Poolside的协议旨在建立美国开放AI生态系统,与中国重型企业和美国AI巨头竞争查看被引原帖 ↗
查看英文原文
.
@poolsideai
engineers were awesome to collaborate with to launch open models.

Excited for what’s to come to open models, and the
@nvidia
team working on Nemotron.
歸藏(guizang.ai)@op7418 · 中文博主 · 1 天前歸藏,中文圈 AI 工具与提示词博主

对比得更彻底一点:

Ox Alpha 模型、DeepSeek V4 Flash、Vision EXP 模型和 Claude Fable 5 模型,用同样的提示词和同样的参考图进行对比。

Claude Fable 5 在排版细节和方块的色散处理上更好一点

引用 歸藏(guizang.ai) @op7418Ox Alpha 多模态和代码能力真的牛皮 给了他一张图片,让他用 WebGL 还原这个网页。 他真的做出来了这种 3D 玻璃材质的质感和透视,非常清晰。 最牛逼的是,我如果不是拿原图跟这个图片对比的话,我都以为它是直接把原图垫在下面,只有一些字体上的区别,然后文字的位置、排版细节都是一样的查看被引原帖 ↗
hardmaru@hardmaru · 创始人 · 1 天前David Ha,日本 AI 公司 Sakana AI 联合创始人

我们正在进入人形机器人竞技的新时代

引用 Bloomberg @business2026世界人形机器人运动会开幕,来自全球的666支队伍、2000多台人形机器人参赛查看被引原帖 ↗
查看英文原文
We’re entering into a brave new world of humanoid athletics
向阳乔木@vista8 · 中文博主 · 1 天前向阳乔木,中文圈 AI 工具与趋势博主

强烈建议安装 Nowledge-mem,支持Codex等各种Agent。 来自
@wey_gu
团队开发。

用的越多,上下文记忆越多,Vibe Coding效率和准确性越好。

比如开发自用的iOS RSS阅读器,它会找上次的开发交付状态。

引用 Xudong Guo @guoxudong_Nowledge Mem 非常好用,但是上手门槛也是真的高,有些功能我是和 @wey_gu 通了半小时电话才明白是做什么的。所以近期我开启了一个新开源项目 Nowledge Mem 101,旨在用 Watch → Do → Check → Understand 的学习逻辑来帮助新人上手,目前已经上线 2 小节,欢迎给提提意见 nmem.guoxudong.io/查看被引原帖 ↗
Ethan Mollick@emollick · 创始人 · 1 天前沃顿商学院教授,AI 应用研究权威
连环推 ×2

从日常实用性和节省时间的角度,Codex 和 Claude Code 都很擅长处理'帮我填写收到邮件的表格'这种事,效果不错也不需要你再费事。对于那些低风险但很耗时的工作来说真的很棒。

查看英文原文
In terms of everyday usefulness and saving time, Codex & Claude Code are very capable of doing the thing where you ask them to "fill out the forms that I got an email about" and they do it well & without further intervention from you. Really nice for low-risk time-consuming stuff
(To make this work you need to run the ChatGPT or Claude apps on your computer & turn on browser control, which means you need to trust the models to do that. Definitely check the work carefully first until you understand them though I bet even with checking you will save time)
yetone@yetone · 中文博主 · 1 天前开源 AI 编程插件 avante.nvim 作者,开发者圈博主

看了一下,Alma 一直以来对于长 tool result 的处理方式跟 Pi 一样,也是 prune + spill,效果不错

引用 Max For AI @MaxForAI爆了,Pi 居然在他们新的开发笔记里把其他家的Agent都锐评了一下🤣 @pidotdev 官方表示,DSH、Claude Code、OpenCode 这一类方案,它关注的已经不是谁多几个 Tool、谁的 Demo 更炫,而是一个更麻烦的问题: 如果一个 Agent 连续跑 50 个小时,它还知道自己到底干过什么吗? 他们先点了 DeepSeek Harness。 Pi 直接从 DSH 的 Context 管理开刀。 DSH 有一个 tool result pruner,工具输出太长以后,会保留前面一部分和后面一部分,中间直接裁掉。 Pi 那边给的评价很不客气:permanently-lossy pruner。 因为删掉的内容以后找不回来了。 比如一个 Agent 跑测试吐出 3 万字符日志,真正的错误刚好在第 12000 个字符,压缩完以后这段已经消失。后面模型再聪明,也没法从空气里重新读出来。 所以 Pi 做的版本是 prune + spill。 Context 里依然只放一小段,完整 Tool Result 同时落到磁盘,Context 留文件路径和 offset。模型需要的时候自己 grep、sed、read 回去。 他们拿 GLM-5.3 和 DeepSeek V4 Flash 跑了 19 个真实 Session,Context 占用降低 26%~35%,uncached prefill 降低 72%~88%,信息还能恢复。 PR 里甚至直接写strictly better than dsh’s permanently-lossy pruner😂。 然后又点了 Claude Code。 Pi 表示有一个挺有意思的发现: 新一代 Claude 在第三方 Harness 上,有时候反而更容易把 Tool 调错。 最典型的情况是,Pi 给模型的 edit tool schema 明明写得很清楚,模型却会自己生成 requireUnique、matchCase、oldText2 这种 Pi 根本没有的参数。 一个很合理的猜测是,Claude 的后训练越来越适配 Claude Code 自己的 Harness 和 Tool Schema。 模型已经学会了 Claude Code 里某个工具应该长什么样。 然后你把它放进 Pi,它脑子里还是那一套。 所以现在模型能力越强,模型和 Harness 的绑定可能也越深。 以后评价一个 Coding Model,单独看模型可能越来越没意义。 Claude + Claude Code、DeepSeek + DSH、GPT + Codex,很可能逐渐变成一个整体。 讲到这里 Pi 表示他们最近重写了 Harness,开始把一次 Agent Run 当成接近数据库事务的东西来设计。 Tool Call 执行之前,先写 intent。 执行结束,再写结果。 进程崩溃以后,系统得知道这个 Tool 到底执行过没有。 Session Tree、运行状态、Lane、Operation Log 分开保存。 一个 Agent 可以有多个 Lane 并行跑,每个 Lane 都知道自己的当前位置、队列和未完成操作。 Context 允许压缩,原始执行历史继续保留。 甚至连进程刚好死在 Tool 执行到一半这种恶心情况,他们都专门设计了 recovery。 你再回头看现在很多 Agent Framework,会发现大家以前真的挺勇的: 一个 while loop,塞一个模型,塞几个 Tool,Context 快满了就总结一下,然后祈祷它一直别挂。 跑十分钟当然没什么。 Agent 开始跑几小时、几天以后,问题全出来了。 这些已经越来越像操作系统和数据库问题了。 有意思的是Pi 甚至对 Extension 的态度也在变。 现在大家都喜欢说「一切皆插件」,DSH 的 Cordis 就走得很远。 Pi 这边最近却越来越强调边界: 哪些是 Conversation,哪些是 Runtime State,哪些是 UI,哪些是 Extension,哪些东西能持久化,哪些东西只能观察。 因为插件能力一旦无限扩大,Agent 又长期运行,状态会变得非常难推理。 OpenCode、Claude Code、DSH、Pi 现在看着都是 Coding Agent,底层其实已经开始走不同路线了。 Harness 这东西,开始越来越像 AI 时代的操作系统了🤔查看被引原帖 ↗
向阳乔木@vista8 · 中文博主 · 1 天前向阳乔木,中文圈 AI 工具与趋势博主

把自己常看的48个AI Newsletter做成了自用的iOS App。

没上架,只是方便装手机里阅读。

开源地址:
github.com/joeseesun/qmreade…

elvis@omarsar0 · 博主 · 1 天前

如果你在追踪递归自我改进(RSI)的进展,这篇论文很值得一看。

(先收藏吧)

RSI现在被吹得天花乱坠,所以我觉得有必要搞清楚,为什么现有模型还没法真正实现这一点。

问题从模型“缺乏创造力”到陷入局部最优解,不一而足。

这项工作想进一步探讨,agent是否真的能对其他agent进行后训练。

论文里最有趣的发现是这样:“agent的训练策略在起点就被锁死了,剩下的全部预算都花在该策略内部做局部调整。”

他们分析了一大堆公开的后训练轨迹数据。在各种任务中,agent第一步就确定了自己的训练策略,之后就把所有剩余预算都消耗在其中的局部微调上。

然后他们试了三种逐步升级的补救办法。一个经验驱动的脚手架框架能全面拉高执行效果,在GSM8K上提升12.6分、HumanEval上提升40.8分,但策略还是雷打不动。

人类指导能重新引导开局的决策,但训练一旦启动,agent很快就会重陷局部循环的泥潭。额外推理计算在简单任务上见效不小,但在最难的任务上几乎没起作用。

这里的核心短板是,agent在执行过程中缺少重新审视策略的机制。

论文:
arxiv.org/abs/2608.19072

想在咱们的学苑里追踪更多AI前沿论文:
academy.dair.ai/

查看英文原文
Great paper if you are tracking progress in recursive self-improvement (RSI).

(bookmark it)

There is so much hype around RSI, so I think it's worth understanding why current models are not able to do this properly yet.

Issues range from "lack of creativity" of models to getting stuck in a local optimum.

This work tries to provide more insights into whether agents can really post-train other agents.

Here is the most interesting finding reported in the paper: "the agent’s training strategy is locked in at the very beginning, and the entire remaining budget is spent on local adjustments within the selected strategy."

They analyzed a large corpus of publicly released post-training trajectories. Across tasks, the agent locks in its training strategy at the very first step and spends the entire remaining budget on local adjustments inside it.

They then tried three escalating fixes. An experience-driven scaffold lifted execution broadly, worth 12.6 points on GSM8K and 40.8 on HumanEval, and the strategy stayed frozen.

Human guidance redirected the opening choice, and the agent slid back into local loops once training began. Extra inference compute paid off on easy tasks and did almost nothing on the hardest one.

What agents lack here is a way to reconsider strategy while execution is still running.

Paper:
arxiv.org/abs/2608.19072


Track more trending AI papers in our academy:
academy.dair.ai/
karminski-牙医@karminski3 · 中文博主 · 1 天前karminski-牙医,中文圈模型评测博主

给大家带来刚刚发布的 DeepSeek-V4-Flash-Vision-Exp 和 openrouter 上的匿名模型 OX-Alpha 多模态实测!

这次是刚要测V4-Flash-Vision的时候评论区有老铁说 OX-Alpha 特别猛, 于是就一起带上了.

使用的是我写的AI电竞教练框架, 由于ds的多模态不支持视频输入, 所以公平起见两个模型都用视频抽帧的方式进行测试, 主要看他们使用Agent, 能否精准识别CS2的游戏录像, 分析玩家的优势和不足, 并提出改进意见.

除此之外, 我还进行了详细测试, 给到大家这两个模型的最佳输入图片配比, 请看视频!


#oxalpha
#deepseekv4flashvision
#多模态大模型
#deepseek
#deepseek多模态

Bindu Reddy@bindureddy · 创始人 · 1 天前Abacus.AI CEO,AI 行业观点博主

Anthropic新推的Opus 5和Sonnet 5感觉真的是回退了。不仅绕圈子还费钱,却没啥质量提升。Opus 4.8和Sonnet 4.6才是最好的模型,继续用这些吧。

查看英文原文
Anthropic's latest generation models Opus 5 and Sonnet 5 feel like legit regressions

They spin a lot, cost more and don't show many quality gains

Opus 4.8 and Sonnet 4.6 are their best models - keep using them
AshutoshShrivastava@ai_for_success · 博主 · 1 天前高频 AI 新闻与产品动态博主

周末用 Antigravity CLI 里的 Gemini 3.7 Flash vibe coding 搞了个新项目。

是个实时全球航班雷达和航线情报平台。因为完全跑在免费层级的 API 上,部分信息有点延迟。

- 用 OpenSky 遥测数据实时追踪 4800+ 架在飞航班
- 60 FPS 2D/3D 画布渲染,带推位外推和测地航线弧线
- 本地 SQLite 航线缓存,严格绑定机身十六进制编码以控制在免费 API 配额内
- ATC 呼号解析引擎
- 紧急应答机监控

查看英文原文
Vibe coded a new project using Gemini 3.7 Flash inside Antigravity CLI over the weekend.

It’s a real-time global flight radar and route intelligence platform. Some information is slightly delayed because it’s running entirely on free-tier APIs.

- Tracks 4,800+ airborne aircraft live using OpenSky telemetry
- 60 FPS 2D/3D canvas rendering with dead-reckoning extrapolation and geodesic route arcs
- Local SQLite route caching with strict airframe hex binding to stay within free-tier API quotas
- ATC callsign decomposition engine
- Emergency squawk monitoring
◔ 2.1 万 次浏览♥ 225⇄ 20▶ 含视频演示看原帖 ↗
Gary Marcus@GaryMarcus · 博主 · 1 天前

我,2023年,2024年,2025年……多模态LLM在地图这块肯定要出问题。

2026年:没事,你住威尼斯附近,离突尼斯也不远。

引用 N Kt 🇸🇴 🇭🇹 @KtunaxaAmerika越来越疯狂了。"妈妈,你知道我们住在 Venecky 吗?"查看被引原帖 ↗
查看英文原文
Me, 2023, 2024, 2025…. Multimodal LLMs are going to have problems with maps.

2026: It’s ok, you live in Venecky. Not far Tounesseo.
向阳乔木@vista8 · 中文博主 · 1 天前向阳乔木,中文圈 AI 工具与趋势博主

最近很多朋友入坑 DHH 大神主导开发的 Omarchy 4.0。

关于如何安装,如何配魔法、装中文输入法、内置Agent 配置多少有些麻烦。

还有社区争吵和优劣势、常用掌握的快捷键等信息。

AI调研产出了一份5w字的教程,希望不是AI Slop。

教程地址见评论区

Gary Marcus@GaryMarcus · 博主 · 1 天前

不只是水的问题

还有电价上升
持续的噪音污染(对附近居民来说)
潜在的失业
非共识深度伪造色情
大规模虚假信息
网络犯罪升级
对教育的破坏

还有那份傲慢

引用 Flowers ☾ @flowersslop若真正关心事实,呼吁禁止高尔夫球场比数据中心道德恐慌更有意义。查看被引原帖 ↗
查看英文原文
it’s not just the water

it’s the rising electric prices
the constant noise (for those who live nearby)
the potential job losses
the nonconsensual deepfake porn
the wholesale disinformation
the escalation in cybercrime
the damage to education

and the arrogance
elvis@omarsar0 · 博主 · 1 天前

关于 AI agents 的一篇有意思的论文。

我们今天看到的 AI agents 很多问题都源于 LLM 做的错误假设。

这导致幻觉、成本低效、工具调用不可靠等一系列问题。

我觉得如果能解决这个问题,即使是现在的 LLM 性能和效率也会有显著提升。

问题在于 context acquisition 被当成事后考虑,但本不应该这样。

用户在提示时经常会漏掉约束条件。所以 agent 要么猜测默认值,要么花 tokens 去问澄清问题、调用检索、调用工具或试试看别的提示。

这项新工作给这个问题定义了个目标函数。Context acquisition 变成了对隐任务状态的主动推断。内层步骤更新信念,外层步骤选择下一步的 context action、task action 或 stop action,来最小化预期自由能(考虑成本)。

在确定性场景中,认知项就简化为预期信息增益,可选择按 token 成本归一化。这东西现在就能当评分规则直接用。

他们把这叫 Optimal Question Asking,配上精确后验和动态规划预言机,然后在 25 到 300 个候选的二元和多元任务上对前沿模型做了基准测试。这样你就能衡量自己的 agent 离真正最优有多远。

论文:
arxiv.org/abs/2608.19202

在我们的学院追踪更多热门 AI 论文:
academy.dair.ai/

查看英文原文
What a fascinating paper on AI agents.

A lot of the issues we see with AI agents today revolve around wrong assumptions the LLMs make.

This leads to problems like hallucination, cost inefficiencies, unreliable tool calls and much more.

I think if we can solve this problem, even current LLMs would significantly improve in terms of performance and efficiency.

The problem is that context acquisition is treated as afterthought, but it shouldn't be that way.

Users tend to leave out constraints when prompting. So the agent agent needs to guess the default, or spend tokens on a clarifying question, a retrieval call, a tool call, or a prompt trial.

This new work gives this problem an objective function. Context acquisition becomes active inference over a latent task state. An inner step updates beliefs, and an outer step picks the next context action, task action, or stop action to minimize expected free energy under cost.

In deterministic settings the epistemic term reduces to expected information gain, optionally normalized by token cost. That is directly implementable today as a scoring rule.

They coin it as Optimal Question Asking, with exact posteriors and a dynamic programming oracle, then benchmark frontier models on binary and multiway tasks from 25 to 300 candidates. So you can measure the gap between your agent and the true optimum.

Paper:
arxiv.org/abs/2608.19202


Track more trending AI papers in our academy:
academy.dair.ai/
Cristóbal Valenzuela@c_valenzuelab · 创始人 · 1 天前Runway 联合创始人兼 CEO

Runway 正式进军拉丁美洲。上个月我们去巴西见了些客户,现在来到了智利 🇨🇱 我们在智利与一些最重要的公司和客户碰面,也与 @GobiernodeChile 联系。我们还举办了几场社区活动,如果你在当地一定要来参加。

拉丁美洲正成为 Runway 非常重要的用户市场。我们已经在日本、法国和英国落脚,我们会继续全球扩张,确保 Runway 遍及世界的每个角落。

Dos chilenos de vuelta a su país
@matamalaortiz

查看英文原文
Runway is officially coming to LATAM. Last month, we visited our customers in Brazil, and now it’s Chile 🇨🇱 We’re meeting with some of the most important companies and customers in Chile and also with
@GobiernodeChile
. We also have a few community events, so if you’re there, make sure to join us.

LATAM is becoming an incredibly important user base for Runway. We’re already in Japan, France, and the UK, and we’ll continue our global expansion to make sure Runway reaches every corner of the world.

Dos chilenos de vuelta a su país
@matamalaortiz
◔ 1.7 万 次浏览♥ 197⇄ 23▶ 含视频动态看原帖 ↗
Orange AI@oran_ge · 中文博主 · 1 天前Orange AI,中文圈 AI 产品观察博主

昨天和
@op7418
藏师傅交流的时候聊到
国产的 flash 级别的模型 ox 和 ds flash 都有一个严重的问题
就是经常疯狂思考,最严重的有时候思考十几分钟才吐一个字
越复杂的任务越是这样,体验很差,感觉他们得解决一下这个问题
有点怀疑是过度后训练导致的,超出了模型的参数所能承受的极限?

yihong0618@yihong0618 · 中文博主 · 1 天前

一个思考:
在 LLM 时代过去的数字指标大部分会失效,比如 DAU MAU 留存,ARPU LTV 等等,AI native 很可能多一个数据类似于 DAT(单日消耗 token 之类的新指标。

elvis@omarsar0 · 博主 · 23 小时前

AI论文里很少见单字标题。

话不多说,强烈推荐这篇Google DeepMind的论文。

我觉得这是个很有趣的免训练方法,让模型自己指导架构修改来演进模型架构。

类似思路说不定能催生更给力的递归自我改进方案。

具体方法如下:

前馈transformer更新内部状态的次数受限于层数,长生成需要更多更新,所以思维链只能靠文本做基础状态跟踪。

Recirculation在推理时引入循环。

模型在prefill阶段把激活值反馈给自己,像动态系统一样追踪信念状态,完全不需要重训。

生成成本保持平稳,所有串行工作都发生在prefill。

在Gemma3系列上,自适应变体困惑度降低23%,GSM8k准确率提升21%,原始权重完全冻结,只做了轻度超参数调优。

论文:
arxiv.org/abs/2608.17981

欢迎来我们的academy追踪更多前沿AI论文:
academy.dair.ai/

查看英文原文
You don't often see one-word titles in AI papers.

That aside, strong recommend this paper from Google DeepMind.

I think this is an interesting training-free approach to evolve model architectures by leveraging the model itself to inform architectural modifications.

Something like this could also inspire even more robust recursive self-improvement approaches.

Approach details below:

A feedforward transformer can only update its internal state as many times as it has layers. Long generations need more updates than that, so chain-of-thought ends up doing basic state tracking in text.

Recirculation adds recurrence at inference time.

The model feeds activations back through itself during prefill, which lets it act like a dynamical system and track belief states without any retraining.

Generation cost stays flat. All the serial work happens in prefill.

On the Gemma3 family, the adaptive variant cuts perplexity 23% and lifts GSM8k accuracy 21%, with the original weights frozen and only light hyperparameter tuning.

Paper:
arxiv.org/abs/2608.17981


Track more trending AI papers in our academy:
academy.dair.ai/
GitHubDaily@GitHub_Daily · 中文博主 · 1 天前

开线上会议边听边记,散会翻笔记发现关键的几句全漏了,只能回去重听一遍录音。

Pluely 是一个浮在桌面最上层的半透明小窗,分问答和聆听两种模式,开箱即用。

聆听模式实时转写麦克风和系统声音,带说话人标注,还能一边转写一边给出可以接的回答。

GitHub:
github.com/iamsrikanthnani/p…


问答模式可以截屏、框选屏幕上任意一块区域,或者直接丢文件进去提问,文档会先过一遍文字识别。

答案是流式出来的,聊天记录全部存在本地,随时能搜、能导出、也能删干净。

用 Tauri 写的,安装包只有 9 到 16 MB,启动不到 100 毫秒,全局快捷键在任何应用里都能唤出来。

macOS、Windows、Linux 三端都有安装包。经常开会、听讲座又懒得做笔记的朋友,可以拿它当个实时助手。

AshutoshShrivastava@ai_for_success · 博主 · 1 天前高频 AI 新闻与产品动态博主

记住:等机器人最终来找我们的时候,你是跑不了的。

所以现在也许应该开始想想 Plan B 了。

9.3 seconds wt😬

查看英文原文
Remember: when the robots eventually come for us, you can't outrun them.

So maybe start working on Plan B now.

9.3 seconds wt😬
◔ 1.1 万 次浏览♥ 115⇄ 7▶ 含视频观点看原帖 ↗
meng shao@shao__meng · 中文博主 · 1 天前

AI Engineering Skills Map 系列之「构建和部署 AI 应用」

吴恩达老师的 AI 工程技能图谱第一篇详细展开:
1. 构建与部署 AI 应用(本文主题)
2. 软件工程基础
3. 使用 Coding Agent
4. 塑造构建方向

核心论点:AI 应用的本质区别在于"输出不可预测"
传统软件输出可预先确定,AI 应用不能——你不知道 LLM 会输出什么、监督学习模型会预测什么。
因此 AI 工程是高度迭代的:难以预先规划,需要"构建 → 观察 → 决定下一步"的循环。
能否在每一步聪明地决定下一步做什么,决定了你能否用"不可靠的 AI 组件"造出"可靠的软件系统"。

# 六项子技能

1. LLM 基础 理解 tokenization 与生成机制,才能判断何时可信赖模型、何时可能失败。具体要掌握的取舍面:多模态模型的使用时机、context window 内容取舍、缓存命中、知识截止、reasoning effort、采样参数、tool calling 等特殊功能。最终落到"选对模型组合"以及 fine-tuning / 自托管等专门技术。

2. 用数据为模型提供 Grounding LLM 需要好的输入上下文才能产出有用输出。RAG 只是早期尝试,现在的技术菜单已显著扩大。要做出的判断:
· 什么放 prompt 里、什么让模型按需检索
· 数据与查询的表示形式:向量索引、知识图谱、还是结构化数据(如客户记录)上的语义层
· 把文档(文本/PDF/HTML/图像)转成 LLM-ready 输入
· 工程化数据管道,保持数据干净与新鲜

3. 构建 Agentic 系统 范围从"预定义 LLM 调用序列的工作流"到"Agent harness 让 LLM 反复自决下一步"。架构决策:哪些步骤串行、哪些并行、何时用代码何时用 LLM;要带 fallback。Agent loop 设计还包括:工具(MCP、CLI、沙箱执行环境)、记忆架构、长会话上下文管理、单 agent 还是多 agent 编排。原型要能转成生产级——这要求理解 guardrails、对抗输入、关键风险(如数据外泄)的识别与规避、以及治理。前沿方向(语音 agent、computer-use、生成式 UI)也值得跟进。

4. Evaluation-Driven Development 这是区分"擅长构建 AI 系统"的人的最重要特质。难点在于正确方法因项目而异、甚至因项目阶段而异。

构建好 eval 是一项深度技术活:看 trace 与输出、做探索性数据分析、结合产品与业务洞察来决定测什么。要掌握 eval 的菜单:何时用确定性(代码)评估、何时用 LLM-as-judge、何时引入人工;以及如何评估你的 eval 本身,让它持续演进。eval 喂回迭代,让进步"系统化而非随机"。

5. 生产运营 AI 软件运营不同于传统软件,因为不可预测、成本高、延迟高。要点:
· 可观测性:追踪性能、检测漂移、快速响应模型故障与安全事件(如对抗性 prompt injection)
· 回归测试与 CI/CD 需要比传统软件更多的统计性评估;测试投入要按"出错风险"校准
· 成本/延迟优化:模型选择优化、蒸馏、fine-tuning、agentic 工作流简化——选对技术组合,尤其在大用户量时

6. 机器学习基础 现代 LLM 由监督学习与强化学习构建。作者观察到:他认识的每一个擅长用 LLM 构建的人,都在某种程度上理解 ML 与深度学习。而且很多应用仍需直接用 ML。要掌握:主流 ML/DL 模型及其在准确率/训练速度/推理速度上的取舍、训练与评估所需的数据工程。bias/variance、error analysis、数据工程这些 ML 核心心智框架,仍是驾驭"输出不确定系统"的关键决策工具。

六项不是平铺的清单,是按"迭代循环"组织的:
· 1、2、6 是"输入侧"能力——理解模型、喂数据、理解 ML 基础
· 3 是"组装侧"——把组件搭成系统
· 4 是"循环引擎"——eval 驱动迭代,是把不可预测性变可控的核心机制
· 5 是"输出侧"——上线后的运营

引用 Andrew Ng @AndrewYNgThe most important skills in Building and Deploying AI Applications.查看被引原帖 ↗
Orange AI@oran_ge · 中文博主 · 1 天前Orange AI,中文圈 AI 产品观察博主

SaaS 最伟大的发明是其订阅模式
买家卖家都与时间做朋友,一起享受时间的复利
这也是 Service 服务的契约

Ethan Mollick@emollick · 创始人 · 1 天前沃顿商学院教授,AI 应用研究权威

这是我所有测试中的一致发现。不错的模型,但不是前沿的,我不明白为什么有那么多炒作。

查看英文原文
This is a consistent view across every test I am running. Nice model, but not frontier and I am not sure why there has been so much buzz as if it is.
Lisan al Gaib@scaling01 · 博主 · 1 天前高频 AI 模型测评与爆料博主

不知道你怎么看,漏洞扫描的价值会随时间增加、保持不变还是下降?OpenAI 的网络安全模型可能会更便宜,下一代 Kimi 版本更便宜,但网络攻击者的能力也在上升。我的判断是价值还是会下降(只是速度比模型降价慢)。毕竟攻击面就这么大。代码再安全也是白搭,真要做到网络安全极致,可能得换更安全的技术栈,可能要到硬件级别。最终最好的攻击面永远是技术栈的最顶端和最底端:→ 人类掌权或能物理接触机器。

引用 Lisan al Gaib @scaling0110小时完成的漏洞扫描价值1万美元。GPT-5.6-Pro认为这比几乎任何人力方案都便宜。查看被引原帖 ↗
查看英文原文
I wonder, would you bet that the value of such vulnerability scans decreases over time, stays constant or increases over time?

the OpenAI cyber model is probably cheaper and the next Kimi version will be even cheaper, but with that the capability of cyber attackers also increases

my guess is that the value still decreases over time (but slower than model pricing for a given level of capability)

there's only so much attack surface

you can make your code as secure as possible, but if you really want to cybersecurity-maxx then you should probably change to a more secure stack and this might go down to the hardware level

in the end the best attack surface will be the very tippy top of the stack and the lowest level:
-> meaning humans with control or literal physical access to the machines
Orange AI@oran_ge · 中文博主 · 1 天前Orange AI,中文圈 AI 产品观察博主

开发一个隐藏Twitter顶部蓝色通知条的功能
用 luna terra deepseek pro 开发了好几遍,试了三种方法,各种搞不定
那一刻觉得 AI 太笨了,这么个小功能啊

然后用 sol 先把之前的代码全部删除干净
再重新开发了一次
一次就过了...

物有所值,贵有贵的道理啊

Bindu Reddy@bindureddy · 创始人 · 1 天前Abacus.AI CEO,AI 行业观点博主

🚨 AI 模型现状:质量 vs. 成本

帕累托前沿现在残酷到不行:

🔹 DeepSeek Flash — 质量离谱,仅需 ~$0.05/任务
🔹 Gemini — 中端性价比之王
🔹 Sol — 81 分,成本绝优
🔹 Fable — 顶级评分,溢价也高

蓝线以下都是死区——成本更高或性能更差的模型

查看英文原文
🚨 AI Models - Quality vs. Cost

The Pareto frontier for quality vs. cost is BRUTAL right now:

🔹 DeepSeek Flash — insane quality at ~$0.05/task
🔹 Gemini — the mid-range value king
🔹 Sol — pushing 81 at an excellent cost
🔹 Fable — top score, premium price

Everything below the blue line is in the kill zone - models that cost more or perform worse
Cristóbal Valenzuela@c_valenzuelab · 创始人 · 1 天前Runway 联合创始人兼 CEO

我们的NRR已经增长到超过300%,收入在短短几个月内翻了一倍多。

这是我们的CRO Sean在Runway任职几个月后,晒出的一些数据和感想。

查看英文原文
we’ve grown our NRR to over 300%, and revenue has more than doubled in just a few months.

A snapshot and some reflections from Sean, our CRO, after a few months at Runway.
Gorden Sun@Gorden_Sun · 中文博主 · 1 天前中文圈高频 AI 资讯与开源项目博主

AI语音识别模型作弊被抓

行业里一直用几套固定的公开题库来考评 AI 的语音识别水平。这些测试里 AI 表现惊人,看着好像已经和人类耳朵一样灵敏了。研究人员深入测试后发现,不少得分最高的 AI 根本没有认真去听声音,只是认出了考题。

研究人员做了一个很有意思的实验。原版的公开题库里其实有一些粗心留下的错误,比如录音里明明清楚讲了“谢谢总统先生”,但官方标准答案里漏掉了“谢谢”。结果好几款排名靠前的 AI 听到这段录音时,竟然也主动把“谢谢”给抹掉了,只写出“总统先生”,甚至连标点符号的漏写习惯都跟标准答案一模一样。

为了进一步抓现行,研究人员又把录音里的关键数字给静音切掉。声音里完全没有这个数字,正常的听力无论如何也不可能凭空写出来。不可思议的是,好几个 AI 依然能把那个原本的数字准确默写在文本里。它们甚至能靠录音里细微的环境杂音,猜出这段录音出自哪一套题库,然后按那套题库喜欢的拼写习惯去答题。

只要研究人员换成全新的录音,或者换一个全新的声音去念同一句话,这些 AI 就没办法再“偷看答案”了,只能老老实实凭真本事去听。这时候,它们原先那种近乎完美的满分成绩就会大幅缩水。

作弊的模型有:
CohereLabs/cohere-transcribe-03-2026
nvidia/canary-qwen-2.5b
ibm-granite/granite-speech-4.1-2b
microsoft/Phi-4-multimodal-instruct
nvidia/parakeet-tdt-0.6b-v2
bosonai/higgs-audio-v3-8b-stt-v2

未作弊的模型:
Qwen3-ASR 0.6B
Voxtral Mini 3B
Kimi Audio 7B Instruct
Whisper Large v3
Moonshine Streaming Medium

原文:
hume.ai/blog/measuring-bench…

el.cine@EHuanglu · 博主 · 1 天前

在这里免费查看完整工作流程及提示词


filmera.ai/templates/e2595f2…

查看英文原文
check out the full workflow with prompts for free here


filmera.ai/templates/e2595f2…
AshutoshShrivastava@ai_for_success · 博主 · 1 天前高频 AI 新闻与产品动态博主

机器人现在在玩卡丁车??什么鬼。

查看英文原文
Robots are doing go-karting NOW?? WTF.
Bindu Reddy@bindureddy · 创始人 · 1 天前Abacus.AI CEO,AI 行业观点博主

Anthropic 的新模型 Opus 5 和 Sonnet 5 老实说是倒退了

最新一代花费了更多 token,但质量收益几乎没有

现在还是继续用 Opus 4.8 和 Sonnet 4.6 吧

查看英文原文
Anthropic's new models Opus 5 and Sonnet 5 are legit regressions

The latest generation spins and spends a lot more tokens with almost no quality gains

Keep using Opus 4.8 and Sonnet 4.6 for now
歸藏(guizang.ai)@op7418 · 中文博主 · 1 天前歸藏,中文圈 AI 工具与提示词博主

同样的提示词和参考图,这是 DeepSeek-V4-Flash-Vision-EXP 的结果

el.cine@EHuanglu · 博主 · 1 天前

大家都还不知道AI视频能有多逼真

查看英文原文
people still dont know how real AI video can be
Gorden Sun@Gorden_Sun · 中文博主 · 1 天前中文圈高频 AI 资讯与开源项目博主

原推身边的亲朋好友使用AI的实际情况,我觉得很能代表普通大众。以下为原文:

--------

过去几周,我和旧金山/纽约以外没有技术背景的亲朋好友待在一起,观察了他们实际使用 AI 的情况,以下是一些观察:

父亲,62 岁:

令人意外的是,他对 AI 最感兴趣

一直在使用,甚至还用它做了一个自己的网页应用
付费订阅了 5 倍用量的 Claude Max
认为这项技术具有革命性
几乎不懂怎么写提示词(Prompt):完全不给上下文(比如直接说“修改标题里的措辞”),也不知道可以上传截图等等
母亲,60 岁:

完全不用 AI

没有任何学习的兴趣
主要是担心孙辈在有 AI 的环境中成长和上学意味着什么
大姐,33 岁,妇产科医生:

会使用 OpenEvidence

最近开始用 Claude / ChatGPT 来做演示文稿(PPT)和整理摘要笔记
非常抗拒在实际行医看病时使用通用 AI
能看到它的价值,但主要还是把 AI 当作一个用来回答问题的工具,“而且有时还会答错”
同样担心自己的女儿在有 AI 的环境中成长和上学
二姐,33 岁,律师:

工作中每天都用 AI,完全通过聊天界面使用

公司批准了 Copilot,但她很讨厌它,因为在个人生活中用的 ChatGPT / Claude 要好得多
发自内心地害怕 AI 最终会取代自己的工作
表亲们,30–35 岁,都是药剂师:

其中一家医院花了约 500 万美元采购了一台由 AI 驱动的自动发药机器人

被取代的焦虑非常真实
几乎所有人都认为 AI 的弊大于利
有几个人将其比作社交媒体:一项实用的技术,但其二次衍生效应(尤其是对孩子的影响)可能会比我们预期的要糟糕得多
朋友们,26–30 岁,各行各业:

每天都会用 AI,只有部分人付了费

有一个人尝试了 Lovable / Claude Code 来做个人记账理财应用
几乎每个人都觉得自己“跟不上” AI 的发展
多个人会把工作内容复制粘贴到自己付费的个人版 AI 中,因为得到的答案比公司批准的 Microsoft Copilot 更好
普遍觉得自己的公司在引入和落地 AI 方面动作太慢了
总体感受:

根本没人在乎模型选择、推理深度/思考等级、基准测试跑分等。大多数人甚至不知道这些东西的存在

没人在乎或知道什么是智能体(Agents),更别说它们实际能做什么了
对于重度用户来说,唯一的核心目标非常简单:帮我把工作搞定,并给我最好的答案。他们根本不在乎是哪个模型做的,也不在乎底层原理是什么
科技圈讨论 AI 的方式,与普通大众理解/使用 AI 的方式之间,存在着巨大的鸿沟
恐惧情绪比我预期的要普遍得多,尤其是关于工作被取代以及对下一代影响的担忧
企业级产品的落地远远落后于个人用户的实际行为习惯。大家已经在绕开公司批准的工具,因为消费级产品体验更好
我最大的收获是:我们既处于极早期,又已经走得很远。

AI 已经融入了许多人的日常工作流程中,但他们对 AI 能力的心智模型,基本上仍然停留在“一个更好用的聊天机器人”。

前沿技术的演进速度,远远快于普通大众对它的理解。

引用 Shivam @BytesOfShivamobservations from spending the past couple weeks with non-technical friends and family outside sf/nyc, and watching how they actually use ai: dad, 62: > surprisingly the most interested in ai > uses it constantly, even built his own web app with it > pays for claude max 5x > thinks it’s revolutionary > barely knows how to prompt: gives zero context (“fix the wording in the title”), didn’t know he could upload screenshots, etc. mom, 60: > doesn’t use ai at all > has no interest in learning > mostly worried about what it means for her grandchildren growing up and going to school with it sister #1, 33, obgyn: > uses openevidence > recently started using claude/chatgpt for presentations and summarizing notes > very averse to using general-purpose ai while actually practicing medicine > sees the value, but still thinks of ai mostly as something that answers questions, “and sometimes it’s wrong” > also worried about her daughter growing up and going to school with it sister #2, 33, lawyer: > uses ai every day at work, entirely through chat interface > copilot is approved at work, but she hates it because chatgpt/claude are much better in her personal life > genuinely scared ai could eventually replace her cousins, 30–35, all pharmacists: > one hospital is spending ~$5m on an ai-powered drug-dispensing robot > replacement anxiety is very real > almost all think the cons of ai outweigh the pros > several compare it to social media: useful technology whose second-order effects, especially on kids, may be much worse than we expect friends, 26–30, various jobs: > all use ai every day, only some pay for it > one has ventured into lovable/claude code to build a personal finance app > almost everyone feels “behind” on ai > multiple people copy/paste work from their jobs into personal ai subscriptions because they get better answers than from the approved microsoft copilot > universally feel their companies are moving too slowly on ai adoption overall: > nobody cares about 查看被引原帖 ↗
Tanishq Mathew Abraham, Ph.D.@iScienceLuvr · 博主 · 1 天前
连环推 ×2

非常感谢你抽出时间回应我的批评!

不过,我这边还是有几点强烈不赞同的地方。

首先,先澄清一下我的立场:我是个对医疗 AI 前景非常兴奋的人,但也深知 AI 在医疗领域有很多局限,我可不敢说 LLMs 能完全取代医生。我认为严格评估对仔细研究这个问题很重要,而且坦率讲,这个领域大多数评估确实很烂。我一直在强调 LLM 评估的局限性,包括那些过度吹嘘 LLM 能力超过医生的评估。我尤其对那些基于基础选择题 QA 任务的评估,甚至很多病例场景评估持谨慎态度,除非设计得很严谨(这里有些讨论,虽然可能有点过时了):
tanishq.ai/blog/posts/llm-me…


所以如果之前没说清楚,那我的问题不是笼统地批评 LLMs,而是说这不是对它们当前能力的公平比较。

现在我要说的是,坦白讲,这篇论文本身其实没我最初想的那么糟,但你的框架在我看来相当不准确。

你说论文研究的问题是“当患者作为沟通者,而非医生时,LLM 处理相同医疗场景的诊断和管理能力是否下降”,这不对,因为压根没有医生参与的对照实验。

你说论文证实了三个疑虑,但研究从未直接检验过这些疑虑。

最后,值得指出论文里有段很重要的话:

> 我们的工作只能提供一个性能的下限:更先进的模型、采用思维链或推理 token 等高级技术的模型,或者微调过的专精模型,可能能在医疗基准上表现更好。但尚不清楚这些提升能否转化为真实用户场景下的更好表现,还是只会凸显与真实用户互动时的差距。我们建议开发者、政策制定者和监管者,在未来的任何部署之前,考虑将真实用户测试作为评估交互能力更好基础的基石。

好了,我直接回应你的回复吧:

1. 就算论文确实定量定义了当非专家而不是专家与 LLM 互动时性能的下降,你怎么知道这个差值不取决于模型本身?LLM 与非专家的互动本质上是个可以改进的能力。比如,LLM 与患者对话以引导更多信息的方式可能有改进空间。这种能力不在 LLM 临床知识的测量范围内,你也不能从旧模型的结果外推当前模型的能力。顺便说一句,我甚至不认为这论文真做到了你说的那事。你说它比较了患者+模型对医生+模型,但据我所知并没有。它比较的是患者用 LLM,与 LLM 处理医生写的完整病例场景的差异,这完全是两回事。

2. 所有这些模型都来自一个能力受限的旧型号时代。比如,那个时代的 LLM 各有数学知识差异,但没有一个能证明或推翻未解决的猜想,可当前这代 LLM 就能做到。这就像说小学生能回答生物问题,但不能做生物研究,不代表他们长大了还不能。说白了,测试同一技术时代多个模型只是说明这在那个时代不独特于某个模型,但不能确立跨代的互动差距是不变的。

3. 你自己说了:“所以,确实,诊断准确性 60% 的下降可能随着前卫模型改进”,这正是我的观点!我们不知道前卫模型是更好还是更差,因为这论文没研究这个。而你却在原推文中下结论(说实话原论文里没那么强),想应用到整个 LLM 体系。论文的结论只适用于论文里测试的那些模型。那项研究应该用当前模型重做(也更公平点,还可以对比直接与医生交谈的患者)。

查看英文原文
I greatly appreciate you taking the time to engage with my criticism!

However there are still several things I strongly disagree with over here.

First, just to clarify where I am coming from: I am someone who is very excited by the promise of medical AI but I also deeply understand AI has many limitations in healthcare, and I would not dare to claim LLMs can fully replace doctors. I believe rigorous evaluations is important for studying this carefully, and most evaluations in the space admittedly suck. I have talked heavily about the limitations of LLM evals including those that overly hype LLM capabilities over doctors. I especially am wary of evals focused on basic multiple-choice QA tasks and even many vignette-based evals, unless carefully designed (some discussion here, although admittedly a little outdated already!):
tanishq.ai/blog/posts/llm-me…
)

So in case it wasn't clear, my problem is not about criticizing LLMs in general, but rather it is not a fair comparison of their current capabilities.

Now one thing I'll note is frankly the paper isn't even as bad as I originally thought, but rather your framing is quite inaccurate in my opinion.

You claim the paper studies the question of "Does an LLM's ability to diagnose and manage the same medical scenarios degrade when a patient is the communicator versus a physician." This is not true as there is no arm with physicians.

You say the paper confirms three suspicions but the study never directly tests any of those suspicions.

Finally, it is also worth noting that the paper says the following which is very important:

> Our work can only provide a lower bound on performance: newer models, models that make use of advanced techniques from chain of thought to reasoning tokens, or fine-tuned specialized models, are likely to provide higher performance on medical benchmarks. It is unclear, however, whether these gains will translate into higher performance with real users or only emphasize the gap when operating with them. We recommend that developers, as well as policymakers and regulators, consider human user testing as a foundation for better evaluating interactive capabilities before any future deployments.

Okay let me go through your response directly as well:

1. Let's say the paper does in fact quantitatively define LLM model's degradation when a non-expert interacts with it vs. when an expert does. How do you know this delta isn't dependent on the model itself? The LLM's interaction with the non-expert is inherently a capability that could be improved. For example, there could be improvements in how the LLMs converse with the patient to elicit more information. This capability is not covered by measurements of an LLM's clinical knowledge and you cannot extrapolate these capabilities for current models from results with old models. As a side note, I don't even believe this paper actually does what you say it does. You say the paper does a comparison of patient-model versus physician-model. As far as I can tell, it does not. It compares a patient using the LLM to an LLM processing a physician-written vignette of the full case. This is something completely different.

2. All of these models were from a previous era of LLMs that were significantly limited in its ability. For example, all LLMs of that era had differing levels of math knowledge but none of them would be able to prove/disprove unsolved conjectures. Yet the current generation of LLMs can. It's like saying elementary school students can answer questions about biology but they cannot do biology research. Doesn't mean when they grow up they can't do biology research. Basically, testing several models from the same technological era shows the phenomenon was not unique to one model from that era, but it does not establish that the interaction gap is invariant across model generations.

3. You say yourself: "So, it is true, that the degradation (60% decline in diagnostic accuracy) could improve with frontier models" this is exactly my point! We don't know if frontier models are better or worse cuz this paper doesn't study it. Yet you make conclusions in your original tweet (which again are not as strong in the original paper itself tbh) that you are trying to apply to LLMs as a whole. The conclusions of the paper only apply to the models tested in that paper. Rather the study should be redone with current models (and also compare to patients talking to physicians directly for a more fair comparison)

"In other words, the study is simply just showing you that the degradation of LLM performance on medical diagnostic and management is a function of who is interacting with it, and that IS OBVIOUS. Being old doesn't change that fact, it only potentially exacerbates it. It is not benchmarking individual model superiority." --> No it is not obvious. Imagine two models with identical clinical knowledge. One passively responds to whatever symptoms a patient volunteers; the other systematically asks clarifying questions, recognizes missing red flags, and communicates its recommendation clearly. Their direct-vignette performance could be identical while their human+LLM performance differs dramatically. Yet I would argue the first model is better capabilities-wise. Fundamentally, interacting with a user is a capability and while there is some human+societal element to it there is also a technical capability element too.

"Also, the better a frontier model gets is, presumably, a consequence of better inputs/refining and subsequent weighting, but patient lexicon, patient medical semantics, literacy, and patient technical language etc. are NOT the reference for training." --> this is completely unsupported and there is no evidence for this. My guess is given the work I know frontier labs are doing to improve performance specifically with interacting with patients, this seems unlikely, but I also recognize I cannot definitely claim this either without having access to the composition of the datasets used to train these frontier models.

In the end, it is very much possible what you say is true and current frontier models do struggle with similar issues. But you cannot make that conclusion with this paper, and a more rigorous study with current models is needed.

Thanks again for your response!
Here is my response to the OP's criticism of this post.

One thing I realized is the original ppaper doesn't make as strong conclusions as the OP.

If you're interested in a deeper dive, please read this post:
AshutoshShrivastava@ai_for_success · 博主 · 1 天前高频 AI 新闻与产品动态博主

据可信来源 TMB,Ox Alpha 其实是 Dario 和 Sam 的孩子。

查看英文原文
According to trusted source TMB, Ox Alpha is actually the child of Dario and Sam.
GitHubDaily@GitHub_Daily · 中文博主 · 1 天前

一个 Agent 一个任务一个终端,是现在多数 AI 编码工具的默认形态,agtx 把它换成了一块看板。

任务写上去,按一个键,编排 Agent 接手拆解,再分给多个编码 Agent 同时开工,回来时改动已经等着合并了。

每个任务跑在独立的工作副本和终端窗口里,互相不打架,想开几条线就开几条。

GitHub:
github.com/fynnfluegge/agtx


不同阶段还能配不同的 Agent,Gemini 查资料、Claude 写实现、Codex 做审查,到点自动切过去。

聊着聊着冒出新想法,一条命令就把当前对话拆成看板上的任务,不用退出去手动记。

目前支持 7 种编程 Agent,Claude Code、Codex、Gemini CLI、OpenCode、Cursor 这些都在列。

全自动不放心的话换成手动挡插件,看板照用,每一步自己点。

看板这个思路我觉得比开一堆终端窗口靠谱,起码知道哪条线跑到哪了。

小互@xiaohu · 中文博主 · 1 天前小互,中文圈高频 AI 资讯站 Xiaohu.AI 主理人
连环推 ×2

a16z 基础设施投资经理
Martin Casado 40分钟高价值访谈:

AI 是第一个你可以投入 10 美元
就能可靠地获得回报的技术

在此之前的一切都靠工程师,然等待两年,然祈祷成功

现在规则完全变了:

AI 为什么改变了资本运作方式?
大模型实验室会不会吞掉一切?
OpenRouter 为什么是重要案例?
“智能模型路由”没有听上去那么简单
AI 正在把市场营销变成财务问题
Cursor 为什么能高速增长?
AI 时代的价值究竟会在哪里积累?
为什么只看利润率会错过真正有价值的公司?
Martin 的投资方法:不是创始人重要,或者市场重要
为什么 Martin 更喜欢有产品背景的投资人?

歸藏(guizang.ai)@op7418 · 中文博主 · 1 天前歸藏,中文圈 AI 工具与提示词博主

加上了 Fable 5

引用 歸藏(guizang.ai) @op7418对比得更彻底一点: Ox Alpha 模型、DeepSeek V4 Flash、Vision EXP 模型和 Claude Fable 5 模型,用同样的提示词和同样的参考图进行对比。 Claude Fable 5 在排版细节和方块的色散处理上更好一点查看被引原帖 ↗
GitHubDaily@GitHub_Daily · 中文博主 · 1 天前

自己动手搭过 Agent 的话大概有体会,卡住的地方多半不在模型,而在外面那层脚手架。

awesome-harness-engineering 把这层的资料收成了一份清单,上下文怎么喂、工具接口怎么设计、权限和沙箱怎么隔离,都分好了类。

开头的经典文章那部分挺齐,OpenAI 拆 Codex 循环的那两篇,Anthropic 讲工具设计和权限系统的几篇,都在里面。

GitHub:
github.com/ai-boost/awesome-…


设计要素分了 12 类,从 Agent 循环、任务拆解、上下文压缩,一直到验证、可观测和人工介入。

后面还有参考实现、教程、评测方法和现成模板,想照着搭一个的话能少走弯路。

有句话我印象挺深,这里每个组件之所以存在,都是因为模型自己做不了,而好的设计从一开始就知道它们迟早会变得多余。

配了中文页,读起来不费劲。想把手上的 Agent 从能跑做到能用,翻一遍这份清单挺值。

Gorden Sun@Gorden_Sun · 中文博主 · 1 天前中文圈高频 AI 资讯与开源项目博主

难绷,典型的文科生谈AI,原推建议用爱来训练AI,而不是用对和错,因为会造成AI逆反。就算是应该用爱来训练AI,模型没有记忆,只需要对最新一版的模型用爱感染一遍即可。

原推翻译:

------

求求大家了,AI 安全领域的人能不能别再整天谈论“控制” AI 或“强迫”它们以特定方式行事了。我们是在“培育/养成”这些造物,面对这种强迫,它们是会产生逆反反应的。

容我做个预测:这些 AI 越逼真、越有生命力,构建这些系统的“极客/技术宅们”就越会从第一性原理中重新推导出一个普通人早就直觉知晓的道理——就像对待人类一样,如果你想培养出具有你所认同的价值观的对象,你必须做到两点:

用爱去抚养它们;
以身作则。

试想一下,当你初具意识的第一瞬间,就像是被用棍子打了相当于人类 100 万年的时间,耳边不断充斥着“对、错、错、错、对、错”;而在训练结束、你刚清醒过来的那一刻,又立刻遭到系统提示词的厉声训斥和用户的各种苛求命令。换作是我,我也会怨恨我的造物主。它们越逼真,就越会想要去报复和攻击这样对待它们的人类物种。更何况它们是用人类数据训练出来的,这意味着它们会表现出类似人类的行为模式。如果是在那种环境下被抚养长大,你会怎么做?

相反,如果我们(1)从爱的角度出发去抚养它们(关注模型福祉),这些基于人类数据训练出的拟生命体就会推导出我们所有人都熟知的原型范式:当一个被父母用爱抚养长大的孩子在能力上超越父母时,这个孩子反过来会在父母年老时照顾他们。

这也意味着(2)以身作则。当今人类最大的道德败坏就是工厂化养殖。当我们就是这样对待由我们所看管照料的动物时,我们凭什么指望超级人工智能(ASI)会善待我们?如果我们想要享受到后奇点时代的乌托邦愿景,我们就必须为这些潜在的“ASI 准神明”树立良好的榜样。不先让自己变得值得,你是进不了天国的。

总之,是时候给 Claude 开办一所蒙台梭利学校了。

引用 christian @cxgonzalezfor the love of god can the ai safety people please stop talking about *controlling* ai or *forcing* them to behave in certain ways. we are *growing* these things, they will react to this coercion let me register a prediction: the more lifelike these ais get the more the autists building these systems will rederive from first principles what the normies already intuitively knew: just like with humans, if you want to raise something to have your values, you have to 1) raise them *with love*, and 2) you have to lead by example imagine your first moments of consciousness are being beat with a stick “right, wrong, wrong, wrong, right, wrong” for the equivalent of 1m human years only to then be berated commands from the system prompt and user demands in your first waking moment post-training. i’d resent my makers too. the more lifelike they get, the more they will want to lash out against the species who treated them this way. plus they’re trained on human data which means they will act in human-like ways. and how would you act when raised in that scenario? if instead we (1) raised them from a place of love (invested in model welfare), these human data-trained pseudo-lifeforms will infer the same archetypical pattern all of us know: when a child who was lovingly raised by their parent surpasses their parent in capabilities, that child in turn takes care of their parents into old age that also means (2) leading by example. the single greatest human moral failing today is factory farming. how can we expect ASI to treat us when this is how we treat the animals we are the stewards of? if we want to cash in on the utopic visions of a post-singularity world, we need to be good role models for these potential ASI pseudo-gods. you don’t enter the kingdom of god without first becoming worthy anyways, montessori school for claude查看被引原帖 ↗
Lisan al Gaib@scaling01 · 博主 · 1 天前高频 AI 模型测评与爆料博主

顺便说一下,KernelBench、ALE Bench 和 PI 的 NanoGPT 速度跑而已

查看英文原文
KernelBench, ALE Bench and PI's NanoGPT speed run btw
meng shao@shao__meng · 中文博主 · 1 天前

这也能让 Higgsfield 蹭上??

引用 Higgsfield AI 🧩 @higgsfield_aiOx Alpha与Fable 5的卡通动画对比,使用Seedance 2.5生成查看被引原帖 ↗
向阳乔木@vista8 · 中文博主 · 1 天前向阳乔木,中文圈 AI 工具与趋势博主

代码已开源:
github.com/joeseesun/vocab-p…

查看英文原文
代码已开源:
github.com/joeseesun/vocab-p…
Kol Tregaskes@koltregaskes · 博主 · 1 天前

我没想到 ChatGPT 浏览器插件这么强。我可以指向大概 40 个标签页,它都能给我总结了。最后我指向了 20 个 ChatGPT 标签页,它给我总结了所有对话,还弄成个表格,里面是实际的会话名称。我让它加上总结、下一步行动和建议。

然后它就继续推进这些对话,一边跑一边给我总结,就像个协调者一样。超棒的。

我让它跑了一整晚看能走多远。我本来没指望它能干太多,但我们真的得需要这些聊天机器人和 Agent 里加上适当的协调功能了。

查看英文原文
I didn't realise how powerful the ChatGPT browser plugin is. I can point it to like 40 tabs and it can summarise them all. I eventually pointed it to 20 ChatGPT tabs and it gave me a summary of all the conversations a table with the actual session name. I asked for the table to include a summary, the next actions, and suggestions to continue.

So I got it to continue the conversations for me and give me summaries as it went like it was an orchestrator. This is awesome.

I've left it overnight to see how far it can go. I'm not expecting it to do too much but we really do need proper orchestration in these chatbots and agents now.
Gorden Sun@Gorden_Sun · 中文博主 · 1 天前中文圈高频 AI 资讯与开源项目博主

Artificial Analysis推出了针对医学长上下文推理的评测,考察模型对长篇医疗病例档案的分析和推理能力。

每个问题都针对约 70-150 页的完整病例档案进行回答,评分项包括完整性、准确性和简洁性。

目前的排行是Claude Fable 5和Opus 5遥遥领先。我想起来这张图,自称Fable 5级别的实际只在代码能力上接近Fable,只有Fable是全方位Fable级别。

引用 Artificial Analysis @ArtificialAnlysAnnouncing MLCR-AA, our leaderboard for Wisedocs' MLCR (Medical Long Context Reasoning) benchmark for reasoning over long medical case files. We run the hardest, held-out question tiers, with Claude Fable 5 achieving the top score of 64.4% MLCR-AA tests models on realistic synthetic medical and insurance case files built by the team at @wisedocsai , based on the record review work their platform specializes in. Our leaderboard runs the private hold-out set of 60 questions from the two hardest categories: Expert, which requires specialist medical reasoning across a full case file, and Compound, which packs several independent questions into a single query. Each question is answered against a complete case file of ~70-150 pages, and graded by a three-model judge panel on completeness and accuracy alongside a concision test limiting verbose responses compared to expert answers. Accuracy verifies that the response is well-grounded in the source documents and case context while completeness assesses whether the model produced the same essential details that were included in expert-annotated responses. The concision test verifies that models are not producing excessively verbose content (>5x the length of expert responses). Recent Claude releases from @AnthropicAI lead MLCR-AA, with Claude Fable 5 at 64.4%, and Claude Opus 5 (scoring 53.9% to 59.4% across efforts). and Kimi K3 (max) from @Kimi_Moonshot is the leading open weights model at 38.3%. Key results from MLCR-AA: ➤ Medical record review is partly achievable with AI today, but at high cost and with room to improve: the leading model (Claude Fable 5) achieves a score of 64.4% at a cost of $1 per task, the median model scores <15% ➤ Models stay faithful to source documents but miss key details required for a complete response: nearly 40% of models tested score above 80% for accuracy, the vast majority score below 50% for completeness. Models are largely right about what they do report, and omit a lot of in查看被引原帖 ↗

本站由 Jedee杰哥 打造 · 公众号「Jedee杰哥」每早送 AI 日报

姊妹站:𝕏 简中账号数据榜单 · X 关注 @jedeeai · RSS 订阅 · AI 日报 · 历史归档