JEDEE AI
存档 2026-08-24

8 月 24 日(北京时间)全球 AI 圈推文存档,按曝光排序,共 100 条。
← 返回最新 全部归档

全部情报 每小时更新 · 事件已合并同类项

内容 公司
@levelsio@levelsio · 博主 · 1 天前独立开发者标杆,AI 产品连续创业者

Airbnb 删掉差评,所以所有住的地方都常年 4.7 分

Google Maps 在很多国家把差评当诽谤直接移除,所以也永远都是 4.7 分

订房平台想让你下单,所以也会乐呵呵地帮合作的酒店删差评,结果依然全是 4.7 分

我做了个网站叫
Hotelist.com
(或者 HotelList.com,看你喜欢,两个域名我都有!)

它会聚合所有订房平台的评分,然后归一化成真实分数(比如某个订房网站上大部分酒店普遍都在 4 到 4.7 之间,那就能把它换算成 0 到 5 的真实区间)

还用 AI 视觉模型根据酒店照片来评分

同时会去网上搜真实入住体验(比如博客和 Reddit 帖子)获取那些点评之外的真实评价

现在已经开始有效果了——我刚在雅典住了一家叫 Woo Suites 的精品公寓酒店,每晚 $400。干净、极简风装修很舒服、空调制冷给力、地段居中、员工超好、健康早餐也棒,而且我们就是通过 Hotelist 找到它的!

完全免费,我不靠它赚钱(其实每个月光用 API 和爬虫收集数据就亏几千刀),也没有其他订房网站那种联盟返佣的破玩意儿(那些平台只要酒店付更高佣金就会给它们拉排名!)

Hotelist.com

快推荐给你的朋友们吧!!!!!!🙏

引用 Pascal Lindenau @PascalLindenau这样的人怎么能有4.7星?他试图胁迫我给好评💀。我想知道有没有客人在评价后真的拿回钱🥀。我通常很好说话,但这太过分了,我得举报。查看被引原帖 ↗
查看英文原文
Airbnb deletes negative reviews so everything is always 4.7

Google Maps removes negative reviews in many countries as defamation so everything is always a 4.7

Booking sites want you to book so will also happily remove negative reviews for their hotels so everything is always a 4.7

I made a site called
Hotelist.com
(or
HotelList.com
whatever you prefer, I have both!)

It gets ratings from all booking platforms, normalizes them to show the actual score (eg a booking site where most hotels are always between 4 and 4.7 means you can normalize that to a 0 and 5)

It also uses AI vision models to rate hotels based on how they look

And it browses the internet for actual experiences (think blogs and Reddit posts) to get the real opinion about them outside reviews

And it's starting to work, I just stayed in this beautiful boutique apartment hotel called Woo Suites in Athens for $400/night. Clean, nice minimalist interior, working cold AC, central, excellent staff, great healthy breakfast, and we found it through Hotelist!

It's free, I don't make any money on it (actually I lose thousands per month on it to collect data via APIs and web scrapers) and there's no affiliate money bs like other booking sites (they will increase ranking of hotels if they pay sites a higher commission!)


Hotelist.com


Tell your friends please!!!!!! 🙏
Linus ✦ Ekenstam@LinusEkenstam · 博主 · 1 天前

没想到看机器人奥运会这么有趣😂

查看英文原文
I did not expect it to be THIS fun watching robot olympics 😂
◔ 144.6 万 次浏览♥ 1.8 万⇄ 1,556▶ 含视频演示看原帖 ↗
Tibo@thsottiaux · 公司官方 · 1 天前ChatGPT 产品官方账号

周日好。Reset已推送至各账户,我们修复了昨天提到的一些使用问题。你应该能感受到明显改进。明天还会有更多更新,我们会持续沟通。

引用 Tibo @thsottiauxCodex 速率限制更新。发现三个问题:长会话图像处理低效、Computer History 使用过高、对话标题生成功能耗用超预期。明天发布修复并重置所有付费订阅使用数据,下周推进另一优化方案。查看被引原帖 ↗
查看英文原文
Good Sunday. Reset has been propagated to accounts and we landed some fixes to usage for things mentioned yesterday as issues we found. You should feel a positive difference. More to come tomorrow and will keep communicating.
Pietro Schirano@skirano · 博主 · 1 天前设计师出身的 AI 编程与创意博主

没人能准备好迎接即将到来的东西。下一代模型将带来本体论震撼。

查看英文原文
No one is ready for what’s coming. The next generation of models will be an ontological shock.
Tibo@thsottiaux · 公司官方 · 1 天前ChatGPT 产品官方账号

2026年将是企业开始认真关注模型效率与可靠性的年头,因为它们正成为关键基础设施。

查看英文原文
2026 is the year companies start seriously caring about model efficiency and reliability as it becomes critical infrastructure
@levelsio@levelsio · 博主 · 1 天前独立开发者标杆,AI 产品连续创业者

我能证实@xAI是我在中国旅游时唯一能用的西方AI

因为我的eSIM是通过香港路由的

Anthropic和OpenAI完全屏蔽了香港

再强调一遍,不是香港屏蔽他们,而是他们屏蔽了香港

对了,你知不知道Elon的妈妈住在上海?

引用 Damon Chen @damonchen忘记在手机上启用VPN,仍能收到Grok Bot通知,甚至可以发送消息。Elon对中国真友好!🫶查看被引原帖 ↗
查看英文原文
I can confirm
@xAI
was the only Western AI I was able to use when I was traveling in China

Because my eSIM was routed through Hong Kong

Both Anthropic and OpenAI fully block Hong Kong

Again, it's not Hong Kong blocking them, it's them blocking Hong Kong

Also did you know Elon's mom lives in Shanghai?
@levelsio@levelsio · 博主 · 1 天前独立开发者标杆,AI 产品连续创业者

我就喜欢他这样乘机发挥

妥妥的狠人 😂

引用 Jonathan Wilke @jonathan_wilke推出outbid.lol/today。不必击败历史最高$15k出价也能被看到。今天花点钱就能占据排行榜顶端直到午夜,明天所有人重新开始。查看被引原帖 ↗
查看英文原文
I love how he's milking this

Absolute chad 😂
Cristóbal Valenzuela@c_valenzuelab · 创始人 · 1 天前Runway 联合创始人兼 CEO

前几天听到个数据挺惊人的。在国内,AI token 消耗里视频生成占了大约七成,这还没算短内容和机器人相关的。相比之下,国内视频模型的增长速度简直甩美国那边 Claude Code 好几条街。

美国是被 LLM 主导,中国则是 World Model 主导。

引用 Shaun Maguire @shaunmmaguire美国在机器人领域落后严重,这是生死攸关的问题。查看被引原帖 ↗
查看英文原文
Heard a crazy stat the other day. Video generation represents around 70% of the share of AI token consumption in China. A combination of short form content and robotics. In China, video models are growing much faster than Claude Code grew in the America.

America is LLM-pilled. China is World-Model-Pilled.
歸藏(guizang.ai)@op7418 · 中文博主 · 1 天前歸藏,中文圈 AI 工具与提示词博主

小米发了一个 AI 的本地的主机,搭载他们新发布的三个芯片,O3、O100 和 D100,支持 120B 和 3B 双模型,支持快慢系统的切换。

Greg Brockman@gdb · 创始人 · 1 天前Greg Brockman,OpenAI 联合创始人兼总裁

OpenAI 已经养成了进行长期研究下注的能力:

引用 Kundan Kumar @kundan2510在OpenAI工作14个月间,领导层全力支持全双工模型(gpt-live系列)最令人惊喜。虽然研究充满挑战,但每次深入理解问题后都能找到解决方案。这样的研究环境在世界上极为罕见。查看被引原帖 ↗
查看英文原文
OpenAI has built the muscle of making long-term research bets:
Andrew Ng@AndrewYNg · 创始人 · 1 天前吴恩达,斯坦福教授、AI 教育领军人物

在维护AI开放性的努力中,Marin项目是展示模型训练开放性的珍贵范例,包含开放的代码、数据、方案,甚至实验结果。公开发布AI研究曾经是常态;我很感谢@percyliang的开放实验室做法。

引用 Percy Liang @percyliangMarin 535B-A23B本周启动训练,全程开放。计划:18.75T tokens在11个GB200 NVL72上进行预训练(占80%)和中期训练(占20%),耗时约3个月,FLOPs为2.7e24。前期完成4层scaling ladder调试验证(1.6B-27.7B)用于预测。这是迄今规模最大的运行。查看被引原帖 ↗
查看英文原文
In the fight to defend openness in AI, the Marin project is a precious demonstration of openness in model training, with open code, data, recipes, even experimental results. Releasing AI research openly used to be the norm; I'm grateful for
@percyliang
's open lab approach.

宇树科技如果股价一直崩下去,

就要在二级市场圈子不断发酵成金融舆情了,到时候就会出现100个峰哥这种顶级大愣子,流窜全国到处踹机器人了,

如果这事儿闹大了,估计证监会就要紧急叫停其他具身智能IPO了,接下来十几家机器人公司就要IPO全面受阻、政策上全面收紧IPO了,

一批基金要彻底套死了。

Bindu Reddy@bindureddy · 创始人 · 1 天前Abacus.AI CEO,AI 行业观点博主

接下来几天Anthropic会连发好几款新模型

- Opus 5.1
- Sonnet 5.1
- Fable 5.1

Opus和Sonnet将修复此前的回退,性能绝对优于Opus 4.8和Sonnet 4.6。

Fable 5.1会直接登顶排行榜。

查看英文原文
Multiple new Anthropic models dropping in the next few days

- Opus 5.1
- Sonnet 5.1
- Fable 5.1

Opus and Sonnet will reverse the regressions and will be strictly be better than Opus 4.8 and Sonnet 4.6.

Fable 5.1 will top the leaderboards.
OpenAI Developers@OpenAIDevs · 公司官方 · 1 天前OpenAI 开发者平台官方

如果你能在办公桌前少花点时间,却完成更多工作,会怎样?

提示:你已经可以做到了


@AlexFinn
明天来我们工作室,展示他如何在 Codex 里靠语音实现“无手操作”——桌面端和移动端都用。


x.com/i/broadcasts/1OGwbnWnB…

查看英文原文
What if you could spend less time at your desk and still get more done?

Hint: you already can


@AlexFinn
joins us in the studio tomorrow to show how he builds hands-free using voice in Codex—on desktop and mobile.


x.com/i/broadcasts/1OGwbnWnB…
@levelsio@levelsio · 博主 · 1 天前独立开发者标杆,AI 产品连续创业者

新项目的事儿一定要一直说,说到所有人都知道才行!

引用 @levelsio @levelsio多个酒店评分平台删除负面评价导致评分虚高(通常4.7分)。Hotelist.com通过汇合多个订票网站的评分并标准化处理、使用AI视觉模型评估、从博客和Reddit获取真实体验来解决此问题。平台免费使用,无会员费或佣金,帮助找到真正优质酒店。查看被引原帖 ↗
查看英文原文
Never stop talking about your new projects until everyone finds out!
@levelsio@levelsio · 博主 · 1 天前独立开发者标杆,AI 产品连续创业者

这正是让我讨厌 Claude Code 的地方

你被当小孩子,诱导你开宝箱才能工作,还得花一大笔钱买 token

查看英文原文
This is exactly what annoys me about Claude Code

You're treated like a child getting lootboxes just to be able to do your work while you pay massive amounts of money for extra usage tokens
Chubby♨️@kimmonismus · 博主 · 1 天前Chubby,高频 AI 新闻聚合博主

新的Claude模型已被发现:Marshmallow和Melon。这说明发布可能迫在眉睫。

很可能是 Opus 更新,甚至可能有个新的 Haiku。我想我们很快就会知道了。

引用 J A Z I I @notjazii爆料:Anthropic两个新模型现身——claude-mashmallow-eap和claude-melon-eap,不是Fable级别但还不错,很可能下周看到Opus或Sonnet级别的新模型。查看被引原帖 ↗
查看英文原文
New Claude models have been spotted: Marshmallow and Melon. This suggests a release is likely imminent.

Probably Opus update and maybe even a new Haiku. I guess we’ll find out soon
◔ 12.2 万 次浏览(2 条合计)♥ 1,241⇄ 48新品看原帖 ↗
NVIDIA@nvidia · 公司官方 · 23 小时前

📢 首次在硅上测出NVIDIA Vera Rubin的真实性能。

⚡ 相比GB300 NVL72,单位功耗吞吐量提升30倍,token成本降低35倍。

代理会话和聊天、摘要任务完全不同。上下文能跨越数百步,达到数十万token。

NVIDIA用@SemiAnalysis_的AgentX工作负载测试了Vera Rubin,这个负载包含DeepSeek V4 Pro模型的真实代理编码轨迹。

查看英文原文
📢 First ever on-silicon NVIDIA Vera Rubin performance measured on how agents actually run.

⚡ Up to 30x more throughput per megawatt and up to 35x lower token cost than GB300 NVL72.

Agentic sessions are nothing like chat or summarization workloads. Context grows across hundreds of steps and can reach hundreds of thousands of tokens.

NVIDIA measured Vera Rubin performance on
@SemiAnalysis_
AgentX workload consisting of real-world agentic coding trajectories using DeepSeek V4 Pro model.
◔ 10 万 次浏览(2 条合计)♥ 763⇄ 67新品看原帖 ↗
Chubby♨️@kimmonismus · 博主 · 1 天前Chubby,高频 AI 新闻聚合博主

博主继续发谜语。不过,a16z合伙人Martin Casado倒是用上了一个未发布的新模型,给他留下了极深的印象。

我猜这八成是Sutskever的“SSI”模型,尤其因为a16z投了Ilya的初创公司。

要么,我也能轻易想象他们提前看了Astra一眼。不管怎样,这竞争看起来是越演越烈了。

引用 martin_casado @martin_casado我刚获得的新模型让我大开眼界。我认为这将是今年最具意义的发布之一。很兴奋,抱歉说得这么模糊。查看被引原帖 ↗
查看英文原文
The vague posting continues. However, a16z partner Martin Casado had access to an unreleased model that left a deep impression on him.

I suspect it is Sutskever’s "SSI" model, not least because a16z has invested in Ilya’s startup.

Alternatively, I could easily imagine they got an early look at Astra. In any case, the competition seems to be intensifying further.
Yuchen Jin@Yuchenj_UW · 博主 · 1 天前

2024年:一台配RTX 4090的游戏电脑,用4-bit精度跑Llama 3 70B,速度约2 tok/s。

2026年:一台配RTX 5090的游戏电脑,原生支持mxfp4,跑DeepSeek-V4-Flash 284B,速度约24 tok/s。

预测:到2027年底,一台游戏电脑就能以不错的速度本地运行Fable 5级别的智能。

查看英文原文
2024: A gaming PC with a RTX 4090 ran Llama 3 70B at ~2 tok/s in 4-bit.

2026: A gaming PC with a RTX 5090 runs DeepSeek-V4-Flash 284B at ~24 tok/s with native mxfp4.

Prediction: by the end of 2027, a gaming PC will run Fable 5-class intelligence locally at decent speed.
Chubby♨️@kimmonismus · 博主 · 1 天前Chubby,高频 AI 新闻聚合博主

我们知道Opus并不完美,修复它是团队的首要任务。”

昨天,Anthropic的Thariq也说了同样的话。这应该清楚地表明,我们的批评意见确实已经传达到了Anthropic。我真心希望很快能看到一个像当年4.6版本那样令人惊艳的Opus。期待值拉满,就在前方!

我觉得如果我们继续向 @bcherny 和 @trq212 提供建设性的批评意见,会有所帮助 :)

引用 Boris Cherny @bcherny我经常用Opus处理工作,举例说明它擅长的任务类型。每个模型都有不同优势——Opus在长期任务和编码上表现优异,但也有特点,冗长问题尤为突出。我们认真对待,团队优先级是改进。可运行:claude /config outputStyle=concise查看被引原帖 ↗
查看英文原文
"We know Opus is not perfect, and it is a big priority for the team to fix it."

Yesterday, Thariq from Anthropic said the same thing. That should make it clear that our criticism has definitely reached Anthropic. I truly hope we'll soon have an Opus that feels like version 4.6 did back then. Hope is high and there!

I think it helps if we continue to offer constructive criticism to
@bcherny
and
@trq212
:)
elvis@omarsar0 · 博主 · 1 天前

这可是好东西!想了解AI芯片架构格局的话,强烈推荐读一读。依我看,一旦RSI(递归自我改进)成为现实,变化会非常大,因为模型架构可能会飞速演变,形态跟今天完全不一样,对算力的需求也会大相径庭。

引用 Jacob @jacobpeake分享新博客文章讨论 AI 芯片架构。内容涵盖 Nvidia、AMD、TPU、Trainium、Cerebras、Groq 等领先芯片的架构设计、纵向和横向扩展方案以及软件栈,帮助读者理解不同架构的权衡。查看被引原帖 ↗
查看英文原文
This is gold! Recommend reading if you want to understand the landscape of AI chip architectures. IMO, a lot can change once RSI comes around, as model architectures could evolve rapidly, looking completely different from today, hence requiring different compute needs.
David@DavidSHolz · 创始人 · 1 天前Midjourney 创始人

这还没发生?粗略算一下,大概需要8个GB300机架过个周末(服务器成本约15万美元),就能把所有约3000个主要开放数学猜想一网打尽。拜托!咱们干起来吧!

引用 David @DavidSHolz想知道哪个大实验室何时会把所有开放猜想投入 10,000 个随机 agent 会话,让它们跑一个周末。查看被引原帖 ↗
查看英文原文
this still hasn't happened? quick math suggests 8x GB300 racks for a weekend (roughly ~150k$ server cost) to do a end-run at all ~3000 major open math conjectures. cmon! lets do it!
Guillermo Rauch@rauchg · 创始人 · 1 天前Guillermo Rauch,Vercel 创始人兼 CEO

我们对扩展𝚏𝚡的理念:开放协议。
∙ MCP (modelcontextprotocol⁠.io)
∙ Skills (agentskills⁠.io)
∙ Plugins (agent-plugins⁠.org)

最棒的还有Unix:

① 小而精的程序,各司其职,通过调用其他程序来组合。

② 𝚕𝚒𝚋𝚏𝚡让它能嵌入更复杂的程序。你应该能打造自己的CLI、后台agent、软件工厂,无论是本地还是云端都行。

查看英文原文
Our philosophy on extending 𝚏𝚡: open protocols.
∙ MCP (modelcontextprotocol⁠.io)
∙ Skills (agentskills⁠.io)
∙ Plugins (agent-plugins⁠.org)

And the best one, Unix:

① Small programs that do one thing well and compose by calling other programs.

② 𝚕𝚒𝚋𝚏𝚡 that enables embeddability into more complex programs. You should be able to build your own CLI, background agent, software factory. Local or cloud.
宝玉@dotey · 中文博主 · 1 天前宝玉,中文圈 AI 翻译与科普大 V

软件行业有一条康威定律:“你有什么样的团队沟通结构,就会造出什么样的系统架构”,反过来,系统架构也会决定组织架构。

传统软件开发的组织架构是围绕传统软件工程打造的,所以有需求分析 → 产品设计 → 架构 → 编码 → 测试 → 发布这样的开发流程,也产生了产品经理、架构师、软件工程师、QA 工程师、运维这样的角色。

这样分工不仅是因为软件开发的流程,还因为随着软件系统越来越复杂,个人很难完整的掌握所有能力,也几乎没有精力去做所有的事情,所以必须依赖团队分工协作。

这个问题在软件工程神作《人月神话》里面有专门讨论:n 个人有 n(n-1)/2 条沟通路径,沟通成本随人数平方增长。

作者 Brooks 也试图给出解决方案,他从外科手术室找到了灵感:把整个系统的设计决策集中在一个人(外科医生)的大脑中,其他人全是支持角色。沟通路径从网状变星型,层级减少平方值也大幅降低。

这个架构的好处在于:一个 10 人左右小团队,真正做设计决策的只有 1 人,但整个团队的产出却远超 1 人所能达到的上限。因为外科医生的所有认知带宽都被释放出来专注于最核心的事情——思考和决策。

但当年为什么这个模式没流行?我估计很多人都第一次听说这个模式。

1. 外科医生级别的人才极其稀缺,而且这种模式高度依赖个人,如果这个人离开或判断失误,整个团队就会瘫痪。

2. 随着软件规模和领域复杂度的爆炸,要求一个人掌握整个系统的所有设计细节变得越来越不现实。

3. 工具的进步(版本控制、IDE、CI/CD)让一部分辅助角色自然消亡了,团队结构也随之进化成了我们今天更熟悉的敏捷模式。

但 AI Agent 时代来了,Coding 能力越来越强,这个模式倒是可以拿出来讨论,也许变得可行。

Agent 可以是那个支持团队。 一个能定义问题、有判断力的人,带一群 Agent,就是一个完整的交付单元。

“1 个人 + AI” 正在逼近 Brooks 想象中的外科手术团队的产出。

一个技术判断力强的工程师或者产品经理或者任何其他角色,配合 Claude Code、Codex 这类 Agent,可以去思考架构和核心逻辑(外科医生的角色),让 AI 生成实现代码、编写测试、处理样板文件、重构、写文档。

这样决策权高度集中在一个人脑中,执行力被极大放大,沟通成本极低,因为 AI 不需要对齐上下文的会议,它直接读代码。

但这不意味着执行力可以无限扩大。

外科医生需要在脑中维持整个系统的一致模型。AI 加速了执行,但并没有扩大人类工作记忆的容量。一个人能用 AI 更快地写出代码,但他能同时驾驭的系统复杂度并没有同比例增长。

这意味着,AI 时代的外科医生模式可能在中小规模系统上极其高效,甚至于一个人顶一个传统团队,但在超大规模系统上,仍然需要某种形式的分工,只不过分工的粒度和方式会发生变化。

比如让几个“外科医生”各自带着自己的 AI 团队,分头推进不同的模块,模块之间靠设计好的接口契约来保持松耦合。

无论当前的 AI Agent 多强大,最终的瓶颈始终是那个做判断、做取舍、在脑中维持系统一致性的人类大脑。 也许只有等到 AGI 真正到来、AI 能完全替代人类做系统级决策的那一天,这个瓶颈才会被真正突破。

引用 Xiaowen @ixiaowenz十年前需要写一个多月的项目,如今在AI辅助下一天就能完成;过去必须先搭建项目骨架、写基础类才能开始业务逻辑,现在打开笔记本先写提示词,测试通过后AI已经承担了大部分实现工作。 这种体验是实实在在的,开发者第一次感受到“一个人就是一支队伍”的可能。 也正因如此,组织会本能地选择批量购买 Token,把AI工具分发到每个基层工程师手里,认为100个人提效 20%,乘以100至少就相当于团队多了 20 个人啊。 说穿了,个人效率解决的是“更快地完成眼前这件事”,组织效率解决的是“省掉哪些不值得做的事”。 AI的终极价值不是让所有基层员工在细分场景里变得更快,而是让高一级角色拥有“直接做出完整结果”的能力——这种能力在AI出现之前只属于一个完整的团队。 当组织把AI的杠杆交给那些本来就能定义问题、整合资源的人,他们就能用一天时间完成过去需要协调数周才能交付的东西。 反之,把杠杆交给那些只会执行局部指令的人,他们只是把原本一小时的事缩短到半小时,而组织层面看不到任何变化。查看被引原帖 ↗
el.cine@EHuanglu · 博主 · 1 天前

中国AI机器人跳远7.97米后摔倒

查看英文原文
china’s AI robot long jumped 7.97m and fell
◔ 6.1 万 次浏览♥ 453⇄ 32▶ 含视频动态看原帖 ↗

乌兰察布这种承接的是传统阿里云、AWS这类云计算,还有一堆软件外包和传统公司的企业云,

因为LLM时代大家根本就不在乎inference后网络延时,一个LLM本身first token latency钻出来都要几秒钟时间,距离的延迟已经完全没关系了。

要知道无数人顶着claude codex和codex的延时天天用,比新疆远多了。

引用 Robinson · 鲁棒逊 @python_xxt东数西算,最近有个很有意思的现象: 乌兰察布突然挤满了算力中心 2024年底,当地还只有36个数据中心项目,到2026年6月已经变成89个,签约总投资超过5000亿元,运行算力达到16.5万P,其中约6.5万P在承接北京外溢算力。 最直接的优势,是离北京近。 乌兰察布距离北京只有300多公里,直达北京的光缆单向最低时延约2.1毫秒,双向约4.2毫秒。 庆阳同样是“东数西算”核心节点,但距离北京接近900公里。过去做存储、备份和大模型训练,多几毫秒影响不大;现在AI算力逐渐从训练向推理扩散,实时调用、在线模型这些业务对时延越来越敏感,乌兰察布的位置就突然变得很值钱:拿着西部的成本,却接近北京的时延。 第二个优势是电。 庆阳绿电到户电价已经压到0.4元/度以内,乌兰察布部分算力中心还能做到约0.358元/度。只差4分钱,但一个100MW的数据中心一年耗电约8.76亿度,对应就是3500万元成本差;如果做到500MW级别,就是接近1.75亿元。 几分钱电价差,最后是巨额现金流。 而且内蒙古本身就是能源大区,过去最擅长的是把电送到东部,现在开始变成:把GPU搬到电厂旁边,再把Token送到东部。 第三个优势是气候。乌兰察布年平均气温只有4℃左右,大部分时间可以利用自然冷源。数据中心除了GPU耗电,冷却系统同样是能耗大户,低温可以进一步降低用电成本。 乌兰察布离北京只有300多公里, 有低价绿电,有低温气候,还有直连北京的低时延网络。 如果把北京看成一个巨大的AI算力消费中心, 乌兰察布可能会成为北京的的“算力卫星城”。 草原负责发电, GPU负责烧电, 最后顺着光纤送出去的,已经不是电了。 是Token~!查看被引原帖 ↗
OpenRouter@openrouter · 公司官方 · 23 小时前
连环推 ×2

Ox Alpha 预计今天会达到近 6 万亿 tokens。

现在可以通过 ori 在你的编码代理中试用:
$ ori [your favorite harness] --model stealth/ox-alpha

查看英文原文
Ox Alpha is on track to hit nearly 6 trillion tokens today.

Try it now via ori in your coding agents:
$ ori [your favorite harness] --model stealth/ox-alpha
Track Ox Alpha's data live:
openrouter.ai/stealth/ox-alp…
◔ 7.9 万 次浏览(4 条合计)♥ 1,010⇄ 50新品看原帖 ↗
Chubby♨️@kimmonismus · 博主 · 1 天前Chubby,高频 AI 新闻聚合博主

新款Claude模型表现相当不错,即便在中级推理上也很有水准(Opus 5.1?)

昨天Thariq和Boris重申已经把批评听进去了,这让人看到希望——说不定不仅是对Opus 5的升级,甚至可能还有新的Haiku版本,毕竟Luna那边收获了不少好评。

引用 Lentils @Lentils80我从claude-melon-eap和claude-marshmallow-eap得到了一些输出😉。他们确实在大力推进3D RL?建筑物放置得很好。我注意到这两个模型都使用了大量思考token,我多次达到了上限。查看被引原帖 ↗
查看英文原文
The new Claude models show quite good results, even on Medium reasoning. (Opus 5.1?)

After Thariq and Boris reiterated yesterday that they had taken the criticism to heart, it gives hope that these are good updates to Opus 5 and perhaps even a new Haiku, since Luna is receiving a lot of love.
Aravind Srinivas@AravSrinivas · 创始人 · 23 小时前Perplexity 联合创始人兼 CEO

很高兴欢迎 Andrew Gordon Wilson 加入我们的研究团队。他将向 Denis 汇报,领导在持续学习、合成数据、长期 RL 环境和架构方向上的新研究工作。我们正在招聘!有兴趣的话请联系 Andrew!

引用 Andrew Gordon Wilson @andrewgwils很高兴宣布加入 Perplexity AI 担任研究负责人!将从事雄心勃勃的范式转变工作,在开放中推进前沿。欢迎加入重新想象持续学习、智能体协作等领域的工作!查看被引原帖 ↗
查看英文原文
Excited to welcome Andrew Gordon Wilson to our research team. He will be reporting to Denis and lead new research efforts on continual learning, synthetic data, long horizon RL environments and architectures. We’re hiring! Please reach out to Andrew!
Chubby♨️@kimmonismus · 博主 · 1 天前Chubby,高频 AI 新闻聚合博主

有趣的事还没完呢:看来还有个新的 Qwen 模型——Qwen 4——已经在测试中。现在咱们有了:
- 新的 Claude 模型(可能是 Opus 5.1 和 Haiku)
- Ox Alpha(可能是 GLM-5.3 Flash)
- Qwen 4(?)
已经在测试中/被发现了。+ GPT Astra 已确认。没错,有趣的一周(几周)在前方。

引用 White @White1637402Qwen的最新模型"paloma"在Arena中出现。其前端编码能力似乎可能与5Opus不相上下。查看被引原帖 ↗
查看英文原文
The fun doesn't stop: apparently there is also a new Qwen model - Qwen 4 - already being tested.

So now we got:

- New Claude models (probably Opus 5.1 and Haiku)
- Ox Alpha (possibly GLM-5.3 Flash)
- Qwen 4 (?)

Already being tested and / or spotted.

+ GPT Astra confirmed.

Yep, funny week(s) ahead.
@levelsio@levelsio · 博主 · 1 天前独立开发者标杆,AI 产品连续创业者

假酒店评论这个坑你想象不了有多深。整个付费假评论产业链,连我女友爸都被找过,不过他拒绝了。

引用 Anna Lux @theannalux作者父亲收到WhatsApp邀请,通过发表Google Maps酒店评价赚钱,反映酒店评分越来越不可信。Hotelist.com通过汇合多平台评分、关注真实体验和AI辅助来解决此问题。项目寻求全球经常旅行、标准高、给真实评价的评论者。查看被引原帖 ↗
查看英文原文
You don't know how deep the rabbit hole of fake hotel reviews goes

Entire paid schemes of fake reviewers, even my gf's dad was hired for it but he declined!
歸藏(guizang.ai)@op7418 · 中文博主 · 1 天前歸藏,中文圈 AI 工具与提示词博主

我实在搞不懂那些用 AI 回复的人是什么心理和心态。

这样只会加大你被别人拉黑的概率,而且你的内容一眼就能看出来是 AI 生成的,AI 圈所有人都能看懂。

你转发,我不管,再强调一遍:所有用 AI 回复我的内容的人,我都会拉黑

这是一个非常吃力不讨好、闲得没事干、浪费 token 的行为,怎么还有人锲而不舍地这么干?

hardmaru@hardmaru · 创始人 · 1 天前David Ha,日本 AI 公司 Sakana AI 联合创始人

让我想起了在 MuJoCo 环境中发现的新颖运动策略,除了这些在真实世界中真的能用!

引用 Eren Chen @ErenChenAI世界人形机器人运动会400米冠军跑步姿态的慢镜头展示。查看被引原帖 ↗
查看英文原文
Reminds me of novel locomotion policies discovered in MuJoCo environments, except that they work in the real world!
◔ 4.4 万 次浏览♥ 449⇄ 15▶ 含视频观点看原帖 ↗
Guillermo Rauch@rauchg · 创始人 · 1 天前Guillermo Rauch,Vercel 创始人兼 CEO

智能的成本正在不断下降。


@OpenAI
Sol的价格下调以及Vercel AI Gateway的折扣,让Sol成为了我们增长最快的前沿模型。

这说明① 智能的需求极具弹性:推理成本一降,使用量就飞速上涨。

② 如果你还没用gateway,你就错过这波猛烈的价格波动了,它能降你的运营成本、抬高你的利润空间。

难怪routing这块越来越热……gateway是躲不开的趋势了。

查看英文原文
Intelligence is getting cheaper.


@OpenAI
Sol's price reductions & discounts on Vercel AI Gateway have made Sol our fastest-growing frontier model.

This shows ① that the demand for intelligence is highly elastic: as inference costs fall, usage grows rapidly.

② If you're not using a gateway, you're missing out on this incredible price volatility, which lowers your operating costs and increases your margins.

It's no wonder the router space has heated up… gateways are inevitable.
Runway@runwayml · 公司官方 · 1 天前AI 视频生成公司 Runway

用 Runway 上的任何模型生成,然后用 Ruby 转换到交付规格。Seedance 2.5、Gen-4.5、MiniMax H3,无论你用什么,都能直接转换成 16-bit EXR 或 10 和 12-bit ProRes 和 HEVC。

立即在下方链接试试。

查看英文原文
Generate with any model on Runway, then take it to delivery spec with Ruby. Seedance 2.5, Gen-4.5, MiniMax H3, whatever you use, converts straight to 16-bit EXR or 10 and 12-bit ProRes and HEVC.

Try it now at the link below.
◔ 3.8 万 次浏览♥ 183⇄ 20▶ 含视频教程看原帖 ↗
🚨 AI News | TestingCatalog@testingcatalog · 博主 · 1 天前专挖 AI 产品未发布新功能的爆料号

GOOGLE 🔥:Avatar 支持即将登陆 Gemini Desktop,同时还有即将推出的 Customize 标签页,让用户能发现 Apps、Skills 和 Plugins。

除此之外,Gemini 4 的准备工作已经启动;去年 Gemini 3 也出现过类似情况——我们在 7 月就检测到了首批痕迹。

引用 Bedros Pamboukian @bedros_pGoogle 终于准备发布 Gemini 4。约 2-3 天前某平台出现 120+ 条关于 Gemini 4 的新提及(之前为 0)。不到 24 小时前某产品又获得 6 条类似提及。传播已开始。查看被引原帖 ↗
查看英文原文
GOOGLE 🔥: Avatar support is coming to Gemini Desktop, along with an upcoming Customize tab that lets users discover Apps, Skills, and Plugins.

Besides that, Gemini 4 preparation has begun; a similar pattern was seen for Gemini 3 last year, when we started detecting the first traces in July.
Together AI@togethercompute · 公司官方 · 1 天前

给GLM-5.3和Fable 5相同的100美元预算,GLM-5.3能完成的工作量是前者的5倍多。在DeepSWE评测中,GLM-5.3解决了大约17个任务,而Fable 5只有3个,尽管两者首次尝试的性能几乎一样。

查看英文原文
Give GLM-5.3 and Fable 5 the same $100 budget, and GLM-5.3 gets over 5x as much work done.

On DeepSWE, that works out to about 17 solved tasks with GLM-5.3 vs. 3 with Fable 5, even though they perform almost the same on the first try.
◔ 3.4 万 次浏览♥ 234⇄ 15▶ 含视频研究看原帖 ↗
@levelsio@levelsio · 博主 · 1 天前独立开发者标杆,AI 产品连续创业者

Photo AI 的一个绝妙新功能,特别是考虑到你在巴西的家装设计和写给葡萄牙邻居关于狗叫的信

引用 Klaas @forgebitzAI聊天中的记忆功能真是最烦人的东西。我在做各种不同的事情,它试图提取一件事的上下文,混淆后变得几乎无法使用。查看被引原帖 ↗
查看英文原文
Great new feature for Photo AI especially considering your home interior design in Brazil and the letter to the Portuguese neighbor about the barking dog

峰哥踹机器人,开启了一个人类当代史上一个非常恶劣的先河。

接下来中国会出现1000个像峰哥一样的顶流网红、压抑男大、反女拳斗士、B站红小将、反科技进步人士、炒股被套大师、灵活就业民办三本男大等等各种人群,因为各种原因,在网络上零零散散汇集到一起,流窜各个场所,专门到处踹机器人。

以前社会流窜分子喜欢欺负小孩、抢小孩钱、踹老头、踹老太太、踢猫、踢狗、嘲笑外地人、欺负垃圾桶、掰断交通路牌。

但这些行为无一例外,要么伤害人,要么伤害动物,要么伤害公共设施,给公众带来麻烦,很多人还是有所忌惮。

从峰哥开始,踹机器人成为了流行风尚,而且机器人公司也不敢把你怎么样,于是踹机器人成为了最爽、最有效果、最有流量、最不用承担责任、最肆意妄为的发泄方式和施暴方式。

到最后的结果,就是这些小嘉豪们一个个模仿峰哥,到处都是踹机器人的视频。

AshutoshShrivastava@ai_for_success · 博主 · 1 天前高频 AI 新闻与产品动态博主

看起来中国已经赢得了人形机器人竞赛。🔥

查看英文原文
Looks like China already won the humonoid race. 🔥
Ethan Mollick@emollick · 创始人 · 1 天前沃顿商学院教授,AI 应用研究权威

当人们问:“我需要一个agent来处理哪些个人任务?”我想这是因为他们把自动化理解为AI做常规的简单任务。

但现在前沿AI在处理非规律、耗神的事情上真的很在行:处理医疗账单、解决信用纠纷、填写繁琐的表格……

引用 Ethan Mollick @emollick生活中充满了难以处理的事情:医疗保健、政府、个人财务、学校表格等。这就是为什么我认为消费级AI被低估了。人们勉强过日子,但需要他们无法获得的帮助。查看被引原帖 ↗
查看英文原文
When people ask: “what personal tasks do I need an agent for?” I think that it is because they think of automation as a simple regular task for the AI.

But frontier AI is now really good at helping with irregular, draining tasks: a medical bill, a credit dispute, tedious forms…
Chubby♨️@kimmonismus · 博主 · 1 天前Chubby,高频 AI 新闻聚合博主

开源AI在Vercel上的token份额仅仅两个月就从28%猛增到了62%。

就在今天,《金融时报》报道说Fable 5使用量在下降,因为它太贵了,而开源模型正变得日益出色。与此同时,我觉得越来越多的用户感到有必要在本地方能对AI拥有主权控制。

开源的成本效益比不断改善,这让我很高兴。

引用 Gavin Baker @GavinSBakerVercel数据显示,开源AI在过去两个月token份额从28%升至62%,超过OpenAI和Anthropic。表明AI基础设施需求增长更快。预计最终闭源frontier模型占经济价值60-90%,但仅占15-25%的token数。查看被引原帖 ↗
查看英文原文
Open-source AI’s token share on Vercel surged from 28% to 62% in just two months.

Just today, the Financial Times wrote that Fable 5 usage is declining because it's too expensive and open source is becoming increasingly better. At the same time, I think that more and more users feel the need to have sovereign control over their AI locally.

The cost-benefit ratio of open source is constantly improving, and that makes me happy.
Lisan al Gaib@scaling01 · 博主 · 1 天前高频 AI 模型测评与爆料博主

随机鹦鹉能有多幸运?

引用 levent @__alpoge__欢迎新几何对象面世,这是我一直热爱的问题。claude真是包罗万象😄 S^6能承认复杂结构吗?能。查看被引原帖 ↗
查看英文原文
how lucky can the stochastic parrot get?
Ethan Mollick@emollick · 创始人 · 1 天前沃顿商学院教授,AI 应用研究权威

发表研究仅涉及旧版本模型对AI影响本身没有错,但需要详尽谨慎的讨论,并且对非技术读者来说必须极其明确。

如果能证明AI能做某事,那没问题。总体上,一旦AI获得了某种能力,它不会在后续版本中退化。因此,那些表明AI跨过某个门槛(比如“跟人类一样好”)或旨在确立某种最低影响或效果的研究,即使用的是GPT-4,依然站得住脚。

如果研究发现AI在某个方面表现糟糕或存在偏见,那就得格外小心了。不能因为GPT-4做不好就断言AI在此任务上不行,因为当前或未来的模型可能做到。那到底能怎么说呢?这只是些例子:
1)直接了当:“GPT-4无法做X。”这种方式少见,因为很多情况下这不算是篇有意义的可发表论文。但如果提供了复现所需的细分信息,就可能成为一个衡量进展的实用基准。
2)追踪趋势:对比GPT-4、GPT-5、GPT-5.6 Sol之类的(至少包一个推理模型真的很关键。这样你就能就相关能力做出更有力的论断。
3)提出a strong且有理有据的论点,指AI存在天然缺陷或限制,使其无法做X,同时展示佐证证据。
4)聚焦于调节变量或中介因素:比如某种提示方式、方法、社交情境或连接方式会影响GPT-4完成X的能力,这点对未来必须予以关注。
5)着眼于人类:人们如何应对AI?成功和失败在哪里?它给我们带来哪些危险、优势或变化?

再次强调,这些而非全部可能性,但多数情况下,负面能力的论断持久性远不及正面论断。

查看英文原文
It is not inherently bad to publish research on the impact of AI that only refers to older models, but it requires a very careful discussion and has to be very clear to non-technical readers.

If you show AI can do something, its fine. Generally, once AI has gained an ability, it does not regress in future generations. So papers that show that AI has crossed a threshold (such as "good as a human") or that seek to establish some sort of minimum impact or effect still hold up even if they used GPT-4.

If the finding is that AI is bad or biased at something, you need to be much more careful. You cannot claim that because GPT-4 fails at something, that AI is bad at that task, because current or future models may do it. So what can you claim? These are just some examples:
1) You can just be explicit: "GPT-4 could not do X." This is a rare way to frame things because it is not really an interesting publishable paper in many cases. However, if you provide the information needed to reproduce your work, it can become a useful benchmark to measure progress.
2) You can measure trends: compare GPT-4 to GPT-5 to GPT-5.6 Sol or whatever (it is really important to include at least one reasoning model). You can then make better claims about abilities relating to this task.
3) You can make a strong, grounded argument that AI has a natural flaw or limitation that means it cannot do X, and then demonstrate evidence that this might be the case.
4) You can focus on a moderator or mediator: this sort of prompting or approach or social context or connection impacts the ability of GPT-4 to do X, and needs to be a concern for the future.
5) You can focus on humans: how do people react to AI? what are the failures and successes? what dangers or advantages or changes does it bring to us?

Again, these aren't exhaustive, but generally negative capability claims have been much less durable than positive ones.
Chubby♨️@kimmonismus · 博主 · 1 天前Chubby,高频 AI 新闻聚合博主

快速提醒:从8月23日起,DeepSeek将不再区分周末API的高峰期和低谷期定价。

查看英文原文
Quick reminder: DeepSeek will stop distinguishing between peak and off-peak API pricing on weekends from Aug. 23
Gary Marcus@GaryMarcus · 博主 · 1 天前

“我们并不是在经历所谓的‘AI就业冲击’。充其量,只是一些不懂技术的高管在减少招聘,他们吸收了即将发生‘AI就业冲击’的氛围。这些氛围从哪来的?来自两拨人:一拨是骗钱的骗子,另一拨是脑子里听到‘数字上帝’即将降临幻听狂人”。

这是来自
@delong
newsletter的一则精彩引语。

查看英文原文
“We are not experiencing an “AI Jobs Shock”. At most, we are experiencing a reduction in hiring by executives who do not understand the technology but who have absorbed vibes that there will be a real-soon-now “AI Jobs Shock”. Where do these vibes come from? From a combination of grifters seeking money, and madmen hearing the voices in their heads of a forthcoming Digital God”.

Brilliant quote from
@delong
’s newsletter.
Linus ✦ Ekenstam@LinusEkenstam · 博主 · 1 天前

新的斯普特尼克时刻来了,中国在仿人机器人领域领先全球,不止一星半点。

查看英文原文
A new Sputnik moment, China is world leader in humanoid robotics by more than a mile
◔ 2.5 万 次浏览♥ 531⇄ 54▶ 含视频观点看原帖 ↗
elvis@omarsar0 · 博主 · 1 天前

这可能是我看过最有用的关于在生产环境里维持 LLM judge 效果的文章了。

(收藏一下)

Netflix 每周要对数十万条节目级的推荐解释跑 judge,直接服务给百万级的移动用户。

他们没把 judge 当成一个一验证就完的东西,而是当成一个四阶段的生命周期。

> Birth(诞生):定义多个评估标准,用人工标注和解释构建精选的测试基准。

> Training(训练):通过推理对齐的标准调优(Reasoning-Aligned Rubric Tuning)来优化 judge 的评分标准,用元-judge 来指导推理输出作为学习信号。

> Deployment(部署):让一个 judge 同时承担质量把关和反思式生成两个角色。

> Monitoring(监测):持续进行人机在环的对齐,检测漂移并在审查通过后触发重新调优。

他们在数千万用户上做了为期五周的 A/B 测试,结果是用户看了更多之前没看过的内容,浏览到播放的成功转化率相比无解释的对照组有所提升,而且没有质量相关的下架问题。

论文:arxiv.org/abs/2608.18300

学院里可以跟踪更多趋势 AI 论文:academy.dair.ai/

查看英文原文
This is one of the most useful writeups I have seen on keeping an LLM judge effective in production.

(bookmark it)

Netflix runs judges over hundreds of thousands of show-level recommendation explanations per week, served to millions of members on mobile.

They describe the judge as a lifecycle with four phases rather than an artifact that you only validate once.

> Birth defines multiple evaluation criteria and builds curated benchmarks with human labels and rationales.

> Training refines the judge's rubric through Reasoning-Aligned Rubric Tuning, using a meta-judge over reasoning output as the learning signal.

> Deployment puts one judge in two roles, quality gating and reflective generation.

> Monitoring runs continuous human-in-the-loop alignment that detects drift and triggers re-tuning behind a review gate.

A five-week A/B test over tens of millions of members shifted viewing toward previously unwatched content and increased successful browse-to-play sessions against a no-explanation control, with no quality-related takedowns.

Paper:
arxiv.org/abs/2608.18300


Track more trending AI papers in our academy:
academy.dair.ai/
Gary Marcus@GaryMarcus · 博主 · 1 天前

你没法把这些报道跟2万亿美元的估值对得上号。真的对不上。

引用 Guillermo Rauch @rauchgVercel AI Gateway 开源模型的 token 份额创新高,从两月前的 28.4% 升至 62%。作者认为这只是开始,随着企业级应用推进和各类工具实现模型无关性,开源模型份额还将增长。查看被引原帖 ↗
查看英文原文
You cannot square a $2 trillion valuation with all the reports like these. You just can’t.
🚨 AI News | TestingCatalog@testingcatalog · 博主 · 1 天前专挖 AI 产品未发布新功能的爆料号

NVIDIA 爱上 Perplexity

根据 The Information 报道,NVIDIA 以 30 亿美元估值向 Perplexity 投资。

> Nvidia 还考虑过支付许可费来授权 Perplexity 的部分技术,或者直接招聘他们的员工。

引用 Jessica Lessin @Jessicalessin独家新闻:Nvidia正在洽谈投资Perplexity。来自@theinformation团队今晚的意外独家爆料。查看被引原帖 ↗
查看英文原文
NVIDIA ❤️ Perplexity

NVIDIA is investing in Perplexity at $30B valuation, according to The Information.

> Nvidia considered paying Perplexity to license some of the startup’s technology and hire some of its staff.
◔ 4.3 万 次浏览(2 条合计)♥ 267⇄ 13动态看原帖 ↗
小互@xiaohu · 中文博主 · 1 天前小互,中文圈高频 AI 资讯站 Xiaohu.AI 主理人

TRAE、扣子并入豆包
字节将推统一办公品牌 "豆包工作"

字节跳动宣布对旗下的办公AI产品完成了一轮团队整合,TRAE、扣子(Coze)团队将整体并入豆包体系

其中TRAE Work、扣子将与豆包在工作场景的产品能力进行整合

TRAE IDE及CLI 将作为豆包品牌下的编程产品线持续发展

此次调整后,字节也进一步明确了豆包的 AI 主干业务定位,将 AI 办公类产品的重心集中到豆包。

此外,豆包最快将于本周内推出独立的AI办公产品 "豆包工作",作为字节面向 AI 办公场景的统一产品及品牌。

针对该消息,字节方面回应称,此次调整旨在更好地协同产品和技术资源,为用户提供更优质的AI工作体验,现有用户权益不会受到影响。

◔ 2.5 万 次浏览(2 条合计)♥ 50⇄ 2新品看原帖 ↗
Bindu Reddy@bindureddy · 创始人 · 1 天前Abacus.AI CEO,AI 行业观点博主

除非接下来几周AI领域出现重大突破

AI正以疯狂的速度加速发展,多家隐身实验室都在推进极其重要的进展

一个持续学习的小型模型可能即将问世

查看英文原文
Excepting a very huge AI breakthrough in the next few weeks

AI is accelerating at an insane pass and multiple stealth labs are making very meaningful progress

A small continually learning model may be imminent
Ethan Mollick@emollick · 创始人 · 1 天前沃顿商学院教授,AI 应用研究权威

这网站上AI机器人多了以后有个烦人的副作用——它们全都“博览群书”,所以能条理清晰地回复我那类小众梗推。以前这种推文只会吸引一小撮感兴趣的人,现在却招来LLM的回复。

感觉就像再也分不清以法莲人和基列人了一样。

查看英文原文
Annoying side effect of all the AI bots on the site is that they are all “well-read” and therefore reply cogently my niche reference tweets that would normally attract a small but interested group now get LLM replies.

Its like you can’t tell Ephraimites from Gileadites anymore.
🚨 AI News | TestingCatalog@testingcatalog · 博主 · 1 天前专挖 AI 产品未发布新功能的爆料号

Hugging Face 在探索一项 13 亿美元的潜在出售,据 Business Insider 报道。

> Google、Amazon 和 Nvidia 都是现有投资者,但潜在买家还未确定。

说实话,对 NVIDIA 来说这笔买卖特别划算(观点)。

"All your weights belong to us"

引用 Business Insider @BusinessInsiderAI开发者平台Hugging Face一直在探索潜在的130亿美元出售,突显其在AI生态系统中的关键地位。查看被引原帖 ↗
查看英文原文
Hugging Face is exploring a potential $13 billion sale, according to Business Insider.

> Google, Amazon, and Nvidia are existing investors, however its potential buyer is not yet known.

Tbh, for NVIDIA it would make a lot of sense (opinion).

“All your weights belong to us”
◔ 2.5 万 次浏览(2 条合计)♥ 272⇄ 9动态看原帖 ↗
@levelsio@levelsio · 博主 · 1 天前独立开发者标杆,AI 产品连续创业者

欧洲如何正在扼杀创客和小微企业家

查看英文原文
How Europe is killing makers and micro-entrepreneurs

lectronz.com/u/lectronz/arti…


by
@alainpannetrat
歸藏(guizang.ai)@op7418 · 中文博主 · 1 天前歸藏,中文圈 AI 工具与提示词博主

Codex 重置了,提到的 Token 过度消耗问题也得到了修复

引用 Tibo @thsottiaux周日好。重置已推送到账户,昨天发现的问题修复已落地。应该能感受到明显改善。更多修复明天推出,持续更新。查看被引原帖 ↗
🚨 AI News | TestingCatalog@testingcatalog · 博主 · 1 天前专挖 AI 产品未发布新功能的爆料号

Grok 的语音模式最近新增了 2 个声音——Aurora 和 Liora。

哪个更好用?👀

引用 ege @aegeantic在最新更新中为Grok添加了2个新声音,试试Aurora和Liora吧!查看被引原帖 ↗
查看英文原文
Voice mode on Grok recently got 2 new voices - Aurora and Liora.

Which one is better? 👀
◔ 1.9 万 次浏览♥ 187⇄ 7▶ 含视频新品看原帖 ↗
Linus ✦ Ekenstam@LinusEkenstam · 博主 · 1 天前

订阅我的newsletter

insidemyhead.ai(进我脑子的门户)

查看英文原文
subscribe to my newsletter

insidemyhead.ai
elvis@omarsar0 · 博主 · 1 天前

NVIDIA 这篇新论文很有意思。

(先收藏为敬)

它深入探讨了如何评估 agent 技能。

企业团队开始使用共享技能库,而审核关卡通常是个扫描器,负责检查结构、风格和安全性。

NVIDIA 测试了这个关卡能否预测实际效果。

他们用了内部和公开目录里的 145 个真实技能,结果发现结构扫描得分与 LLM 裁判的质量相关性仅 Spearman rho 0.14。

ACES 改用“技能提升”(Skill Lift)作为替代方案。

简单来说,就是在同一个模型、沙箱、工作区和评分器下,同一任务跑两遍——一次加载技能,一次不加载。然后对比 agent 完成任务结果的差异。

他们用四个测试框架,从 58 个生产技能中评估了 947 组配对用例,将轨迹标准化为共享的 Agent Trajectory Interchange Format,这样结果可以在不同框架间纵向对比。

他们发现,最大的流程指标提升出现在技能执行、行为检查和技能效率上。

论文:
arxiv.org/abs/2608.20614


在我们的 academy 查看更多热门 AI 论文:
academy.dair.ai/

查看英文原文
Very interesting new paper from NVIDIA.

(bookmark it)

It takes a closer look at evaluating agent skills.

Enterprise teams are starting to leverage shared skill libraries, and the review gate is typically a scanner that checks structure, style, and security.

NVIDIA measured whether that gate predicts anything.

Across 145 real skills from internal and public catalogs, structural scan scores correlate with LLM-judge quality at a Spearman rho of 0.14.

ACES proposes Skill Lift instead.

In other words, run the same task twice under the same model, sandbox, workspace, and scorer, once with the skill loaded and once without. Then you measure the difference in what the agent completed.

They scored 947 paired cases from 58 production skills across four harnesses, normalizing trajectories into a shared Agent Trajectory Interchange Format, so results compare across harnesses.

They fins that the largest process-metric gains appear in skill execution, behavior check, and skill efficiency.

Paper:
arxiv.org/abs/2608.20614


Track more trending AI papers in our academy:
academy.dair.ai/
Gary Marcus@GaryMarcus · 博主 · 1 天前

如果 Sam 在董事会解雇他时离开了 OpenAI,这事会更有可信度。

引用 NIK @ns123abcSam Altman 批评某些 AI 安全观点:若 AI 过于强大而无法控制、权力过度集中,就应将控制权交给 AI 模型。他认为这种基于对人类不信任的做法是反人性且反人道的。查看被引原帖 ↗
查看英文原文
this would be more credible if Sam had left OpenAI when the board fired him.
◔ 1.9 万 次浏览♥ 131⇄ 15▶ 含视频观点看原帖 ↗
🚨 AI News | TestingCatalog@testingcatalog · 博主 · 22 小时前专挖 AI 产品未发布新功能的爆料号

SPACEXAI 🔥:Grok Bot 即将推出模板、多账户支持、出口隧道和 Chrome 浏览器配置导入功能。

目前还不清楚模板是在任务级别还是机器人级别,但它们将在设置中有一个专门的发现选项卡。

查看英文原文
SPACEXAI 🔥: Grok Bot will soon get Templates, multi-account support, an egress tunnel, and Chrome profile import features.

It is still unclear whether Templates will be at the task level or the bot level, but they will have a dedicated discovery tab in settings.
◔ 2.1 万 次浏览(2 条合计)♥ 279⇄ 12新品看原帖 ↗
OpenAI Developers@OpenAIDevs · 公司官方 · 1 天前OpenAI 开发者平台官方

对 Codex 或 Alex 的语音设置有问题?在下面留言吧。

查看英文原文
Got questions about voice in Codex or Alex’s setup? Drop them below.
🚨 AI News | TestingCatalog@testingcatalog · 博主 · 1 天前专挖 AI 产品未发布新功能的爆料号

Grok Imagine 的 discovery 页面更新了。现在用户可以探索并试用各种功能,比如照片编辑、图片调整大小等等!

引用 @blankspeakerGrok Imagine推出新的Discovery视图,可访问grok.com/imagine查看。查看被引原帖 ↗
查看英文原文
Grok Imagine got an updated discovery page. Users can explore and test different features like photo editing, image resizing, and more!
◔ 1.6 万 次浏览♥ 159⇄ 5▶ 含视频动态看原帖 ↗
Gary Marcus@GaryMarcus · 博主 · 1 天前

奥运级的双标。Altman真是啥都敢说。

今天他在这儿暗戳戳讽刺Dario,搞得好像Sam自己几个月前没干过一模一样的事似的。

Sam今天说:“AI领域有些人在本质上讲,我们会给世界带来治愈一切疾病的方法,我们会把东西搞得很便宜,但交换条件是让人们放弃自主权、对未来和权力的影响力,还打着安全的旗号。”

而仅仅几个月前,Sam自己就发了个声称癌症能被治愈的截图。

引用 Sahil @sahilypatelSam Altman批评某些AI领导者(暗指Anthropic):以'安全名义'要求民众放弃自主权换取疾病治愈和低成本,由精英决策未来。他讽刺这如'亲爱平民,我们赐予癌症治愈和财富,停止抱怨,信任我们好于独裁者',认为这是极差的营销策略。查看被引原帖 ↗
查看英文原文
Olympic-level hypocrisy. Altman will say absolutely anything.

Here is he today taking a shot at Dario as if Sam hadn’t done EXACTLY the same thing just a few months earlier.

Sam today: “there are some people in the AI field who effectively say we're going to give the world a cure to all disease and we're going to make stuff really cheap in exchange for people giving up their autonomy and impact over the future and power, in the name of safety”

and a screenshot of Sam himself alleging an cancer cure just a few months earlier:
◔ 1.5 万 次浏览♥ 141⇄ 25▶ 含视频观点看原帖 ↗
Orange AI@oran_ge · 中文博主 · 1 天前Orange AI,中文圈 AI 产品观察博主

最近 Flash 级别的模型能力都上来了,又一个全新模型 August 来了
小型,快速,高性能。
Ox 牛来在海外已经火到暴量,现在全网的资源紧张,在官方公布谜底之前,应该都是这样的状态了。
Cola 首发,限时免费 2 天,欢迎来玩~

Linus ✦ Ekenstam@LinusEkenstam · 博主 · 1 天前

得有人告诉他们前面有堵墙,不过好在现在我们有办法了。

查看英文原文
somebody needs to tell them there is a wall, at least now we have a survival strategy.
Bindu Reddy@bindureddy · 创始人 · 1 天前Abacus.AI CEO,AI 行业观点博主

6 个月内,95% 的任务都会由小型 250B 开源模型完成。大模型会处理非常复杂的任务,包括训练这些小模型。大部分 token 消耗来自小模型。

查看英文原文
In 6 months, 95% of tasks will be done by small 250B open source models

The large models will perform very complex tasks including train the smaller models

The majority of token usage will be small models
🚨 AI News | TestingCatalog@testingcatalog · 博主 · 1 天前专挖 AI 产品未发布新功能的爆料号

Perplexity 可能会在 Perplexity Computer 上支持 GPT-5.6 Sol Fast。

快了吗?👀

查看英文原文
Perplexity may receive support for GPT-5.6 Sol Fast on Perplexity Computer.

Soon? 👀
Cristóbal Valenzuela@c_valenzuelab · 创始人 · 1 天前Runway 联合创始人兼 CEO

16-bit EXR 或 10 和 12-bit ProRes 和 HEVC

引用 Runway @runwayml在 Runway 上用任何模型生成,再用 Ruby 调整至交付规格。Seedance 2.5、Gen-4.5、MiniMax H3 等均可直接转换为 16 位 EXR 或 10/12 位 ProRes 和 HEVC。现在就试。查看被引原帖 ↗
查看英文原文
16-bit EXR or 10 and 12-bit ProRes and HEVC
Chubby♨️@kimmonismus · 博主 · 23 小时前Chubby,高频 AI 新闻聚合博主

现阶段患者应该在看医生前用 AI 来更好地理解自己的病情和治疗方案,医生也应该强制用 AI 来复核治疗方案。

查看英文原文
At this point, patients should be encouraged to use AI before seeing a doctor to better understand their condition and treatment options, while for doctors it should be mandatory to use it to double-check their treatment plans.
elvis@omarsar0 · 博主 · 1 天前

关于多智能体代码评审的优秀论文。

要知道该用多少个编码智能体来解决问题,确实挺难的。

针对弱智能体代码评审,默认方案就是加更多智能体。但结果表明,在仓库级任务上,扩展智能体数量到很大规模时,收益会递减。

这项新工作尝试了结构化对抗。Adversarial Review 运行三个智能体:一个主编码智能体负责编写,一个评审者进行评估,还有一个批评者在任何编辑前审查评审意见。

在 LiveCodeBench 上,它用三个智能体就击败了五个智能体的基线。

在 SWE-PRBench 上,朴素版本暴露了一个失败模式:智能体在没有足够证据支撑的情况下就趋于达成一致。而把“分歧”作为显式指令后,在所有测试方法中恢复了最高的 F1 分数。

他们还发现,当分歧最小、结构化且基于证据时,协作式评审才有效。

论文:
arxiv.org/abs/2608.18167

在学院里追踪更多热门 AI 论文:
academy.dair.ai/

查看英文原文
Great paper on multi-agent systems for code review.

It's challenging to know how many coding agents to use to address a problem.

The default fix for weak agentic code review is more agents. In turns out that scaling agents to a large number gives diminishing returns on repository-level tasks.

This new work tries structured conflict instead. Adversarial Review runs three agents. A main coding agent writes, a reviewer evaluates, and a critic audits the review before any edit are done.

On LiveCodeBench it beats a five-agent baseline while using three agents.

On SWE-PRBench the naive version exposed a failure mode. The agents converged on agreement without enough evidence behind it. Making disagreement an explicit instruction recovered the highest F1 among tested methods.

They also find that cooperative review works when the disagreement is minimal, structured, and grounded in evidence.

Paper:
arxiv.org/abs/2608.18167


Track more trending AI papers in our academy:
academy.dair.ai/
Mustafa Suleyman@mustafasuleyman · 创始人 · 23 小时前微软 AI CEO,DeepMind 联合创始人

历史上最重要和最鼓舞人心的引言之一:"对人和人类命运的关切要永远是所有技术努力的首要关注…这样我们思想的结晶才能成为人类的祝福而不是诅咒。"(阿尔伯特·爱因斯坦,1931)

查看英文原文
One of the most important and inspiring quotes of all time: “The concern for man and his destiny must always be the chief interest of all technical effort… in order that the creations of our mind shall be a blessing and not a curse to mankind.” (Albert Einstein, 1931)
🚨 AI News | TestingCatalog@testingcatalog · 博主 · 1 天前专挖 AI 产品未发布新功能的爆料号

Perplexity 正在为 Perplexity Computer 开发一个新的细化力度选择器。

能够在低力度下运行 Computer 应该会对积分消耗产生积极影响。

这看起来有点眼熟👀

查看英文原文
Perplexity is working on a new granular effort selector for Perplexity Computer.

Being able to run Computer on low effort should have a positive impact on credit consumption.

This looks familiar 👀
NVIDIA@nvidia · 公司官方 · 23 小时前

了解更多:
nvda.ws/46iThZz

查看英文原文
Learn more:
nvda.ws/46iThZz
AshutoshShrivastava@ai_for_success · 博主 · 1 天前高频 AI 新闻与产品动态博主

很多人还不明白一个 20 块的 AI 订阅能给你带来什么样的超能力。

查看英文原文
Many people still don’t understand the kind of superpower you get for $20 AI subscription.
OpenRouter@openrouter · 公司官方 · 23 小时前
连环推 ×3

当智能便宜 100 倍会怎样?

@shashankgoyal95 与 @MenloVentures 的 @deedydas 聊了聊 AI 经济学的几个真相:

- Token 价格跌了 40%,使用量反而增长 4 倍。
- 成本优化太急了可能反而是错的。
- 最好的 AI 产品会越来越多地跨多个模型构建。

查看英文原文
What happens when intelligence gets 100X cheaper?


@shashankgoyal95
sits down with
@deedydas
from
@MenloVentures
to unpack a few truths about AI economics:

- Token prices are down 40%, while usage has grown 4X.
- Optimizing costs too early can be the wrong move.
- The best AI products will increasingly be built across multiple models.
The conversation goes beyond cost: how fast model cycles are changing, why benchmarks only tell part of the story, and what happens when AI usage is still nowhere near its ceiling.
This is New Defaults, our event series exploring the shifts reshaping how AI gets built and used.

Watch the full episode on Youtube:
youtube.com/I7da3Fe9xXA?si=cRYv…
向阳乔木@vista8 · 中文博主 · 1 天前向阳乔木,中文圈 AI 工具与趋势博主

信用卡账单公司Ramp统计的Anthropic公司模型使用占比。

看来性价比高,大家最爱用的模型是opus4.8 和 sonnet4.6

fable5只能排第三。

elvis@omarsar0 · 博主 · 1 天前

这只是一个数据点,但我觉得开放模型的 token 份额会继续上涨。

如果要更深入洞察,可以看看各领域正在采用的多智能体系统/模式。一个管理者主导一堆执行子代理(前沿开放模型在这里表现卓越)。这是当前最有效且广泛使用的模式之一,子代理占据主要 token 消耗(从前沿开放模型获取时价格和效率都更优)。这个趋势会持续下去。真的没必要事事都用 Fable 5,更实惠的模型完全够用,还能更快更省。

引用 Gavin Baker @GavinSBaker开源AI在过去两个月内从28%的token份额增长到62%,正在抢占OpenAI和Anthropic的市场。即使两者在7月仍在加速增长,整体AI基础设施需求增长更快。预计frontier tokens最终将占经济价值的60-90%,但仅占token数的15-25%。查看被引原帖 ↗
查看英文原文
It's one data point, but I think token share for open models will continue to rise.

You can get more insights by looking more closely at the multi-agent systems/patterns being adopted across domains. A supervisor leads a bunch of execution subagents (frontier open models work beautifully here). This is one of the most effective and widely used patterns today, where subagents dominate token use (which you can get from frontier open models at a much better price and efficiency). This trend will continue. There really is no point in using a Fable 5 for everything, where cheaper models will do just fine and at a faster and cheaper rate.
Gary Marcus@GaryMarcus · 博主 · 1 天前

如今几乎没人再信任美国的大学了。

但对比之下,信任 Sam Altman 的人更是少得离谱,差了十倍都不止。

引用 Gary Marcus @GaryMarcus你更信任谁查看被引原帖 ↗
查看英文原文
Scarcely anyone trusts American universities anymore.

but by a factor of 10 even fewer people trust Sam Altman:
elvis@omarsar0 · 博主 · 1 天前

重要的讨论。用测试框架来衡量模型已经完全失效了。我更倾向于用Pi和Hermes Agent这样极简的框架来测试模型质量。但这也并非完美,因为框架中存在着偏向某些模型而不是其他模型的偏差。

这种标准化的测试方法缺失,但至关重要,因为前沿AI公司都在聚焦框架工程。可惜努力还不够,事情发展太快了,而且框架优化大多是各自为战。

另一方面,我觉得模型最终将能根据具体任务动态生成框架。Claude模型已经在某种程度上做到了这点,虽然挺不稳定,还像个谜。但这可能意味着框架就像系统提示词那样,只是一个可调节的产物。那在这个层面上,我们在评估什么,又该如何评估?基准测试只会从此变得更模糊。

引用 Onur Solmaz @onusoz作者主张建立标准化的 AI 模型评估体系,如教育考试一样统一规范。虽然不同模型在专有基准上表现差异大,但应使用统一、相对稳定的测试框架(如 mini-swe-agent 等),让所有模型在相同条件下接受评估,解决当前基准测试的复杂度问题。查看被引原帖 ↗
查看英文原文
Important discussion. Measuring models against harnesses is completely broken. I prefer to test model quality against minimal harnesses like Pi and Hermes Agent. This is not perfect, as there are biases in the harnesses that favor some models and not others.

A standard way to do this is missing but important, as harness engineering is where leading AI companies are focusing efforts. Not enough effort here, as things are moving fast and harness optimization is individualistic.

On the flip side, I feel like models will eventually have the ability to dynamically generate harnesses on the fly as per task. Claude models do this already to some extent, though pretty inconsistently and remain a mystery. But this could mean that a harness is just a tunable artifact like a system prompt. In that realm, how are we assessing it, and exactly what? Benchmarking will only get murkier from here onwards.
小互@xiaohu · 中文博主 · 1 天前小互,中文圈高频 AI 资讯站 Xiaohu.AI 主理人

ChatGPT Search 底层搜索机制突变:

ChatGPT 开始大规模直接对特定网站进行“定点搜刮”
影响网站GEO策略

8 月 8 日后,Promptwatch 监控到的ChatGPT site: 查询占比从日均 0.368% 升至 16.78%,约 45.5 倍。

ChatGPT Search的新流程可能变更为先全网发现候选来源,再锁定特定域名寻找文档、规则和更新;

这对GEO 策略因此可能变成两层竞争:域名先入选,页面再争取引用。

网站应优先建设可抓取、可核验的一手内容体系


best.xiaohu.ai/article/chatg…

Gary Marcus@GaryMarcus · 博主 · 1 天前
连环推 ×2

帮我画个 Sam Altman 前言不搭后语的漫画

查看英文原文
“Draw a cartoon of Sam Altman talking out of both sides of his mouth”
Drop your favorite examples here, for my next essay.
Gary Marcus@GaryMarcus · 博主 · 1 天前

脑海里蹦出来的词是“手忙脚乱”,

要么就是“白日做梦”。

引用 Rohan Paul @rohanpaul_aiSam Altman阐述OpenAI战略:应从产品公司转向平台公司。目标是为个人和企业提供统一AGI接口及开放API,让开发者可在其上构建任意应用。公司将在成本-性能曲线各点提供最优AI,涵盖高端科研用途到廉价大规模处理。查看被引原帖 ↗
查看英文原文
“flailing” is the word that comes to mind.

either that or “dreaming”
elvis@omarsar0 · 博主 · 23 小时前

用 harness 解决递归自我改进。

我使用的 agent harness 的大问题是:它如何支持 RSI 并在每次迭代上复合?

我需要拥有 harness 的什么功能?

现在最好的解决方案是什么?

在使用 harness 时要理解的一些重要事情。一旦 agent 开始重写自己的 prompts、tools 和 memory,你需要在它下面有它无法修改的持久状态,加上一种方式可以把一次运行回滚到更早的检查点。

exo 是一个新的开源 agent harness,专门为解决这个问题而建立。它把一个 agent 分成三层。

1) exoharness 存储所有持久的东西。对话历史是一个 agent 无法改变的仅追加事件日志,它与 artifacts、secrets 和 sandbox 生命周期一起保存。

2) executor 决定 agent 的行为方式。它组装 prompt、调用模型、分发 tools 并管理 memory。agent 可以重写其中任何一个,你可以为另一个 harness 交换 executor。

3) sandbox 运行重要的工作。packages、files 和 commands 在一个隔离的机器中执行,你可以快照和回滚它。

通过这样的设置你可以做很多事情。

- 从任何事件分叉一次对话并运行同一任务的两个版本

- 回滚到 agent 破坏自己之前的那个事件

- 在几周后用完整的历史和自己的挂载点恢复一次对话

- 读取哪些命令实际从 tool_requested 和 tool_result 运行

随着 agents 承担越来越多自己的配置,这提高了下面 harness 的标准。它到达了一个事件日志、分叉和回滚是必须品的地步。

查看英文原文
Solving recursive self-improvement with a harness.

The big question with the agent harnesses I use is: how does it support RSI and compound on every iteration?

What capabilities do I need to own the harness?

What are the best solutions right now?

Something important to understand when using harnesses. Once an agent starts rewriting its own prompts, tools, and memory, you need durable state underneath it that the agent cannot modify, plus a way to roll a run back to an earlier checkpoint.

exo is a new open-source agent harness built to solve exactly that. It splits an agent into three layers.

1) The exoharness stores everything durable. Conversation history is an append-only event log the agent cannot alter, and it holds artifacts, secrets, and sandbox lifecycle alongside it.

2) The executor decides how the agent behaves. It assembles the prompt, calls the model, dispatches tools, and manages memory. The agent can rewrite any of that, and you can swap the executor for another harness.

3) The sandbox runs the important work. Packages, files, and commands execute in an isolated machine you can snapshot and rewind.

You can do many things with that setup.

- Fork a conversation from any event and run two versions of the same task

- Roll back to the event right before an agent broke itself

- Resume a conversation weeks later with its full history and its own mount

- Read which commands actually ran from tool_requested and tool_result

Agents will keep taking on more of their own configuration, and that raises the bar for the harness underneath them. It reaches a point where an event log, forking, and rollback are must-haves.
OpenRouter@openrouter · 公司官方 · 23 小时前
连环推 ×3

新功能:图像模型的视觉基准评测🌅

查看每个模型在各种难题提示词上的表现:
openrouter.ai/benchmarks/med…

查看英文原文
New: visual benchmarks for understanding the capabilities of image models 🌅

See how every model performs across a variety of challenging prompts:
openrouter.ai/benchmarks/med…
Can the model arrange multiple input images into a scene? Hold up the correct number of fingers? Remove the 2nd cup from a row?

Sort a grid of images for each model on OpenRouter by price and generation time to understand the tradeoffs with their capabilities.
Start using the image generation api:
openrouter.ai/docs/guides/ov…


See the full announcement:
openrouter.ai/blog/announcem…
Tibor Blaho@btibor91 · 博主 · 1 天前逆向挖掘 AI 产品代码的爆料专家

OpenAI 和 Anthropic 本周动态:青少年版 ChatGPT、训练暂停、蛋白质结合剂(2026年第34周)

OpenAI 开始在全球推出面向青少年的 ChatGPT,更新了 Model Spec,并与 CodeAI 达成了教育合作

他们还披露,在最新模型上暂停了两周的 RL 训练,以加固研究环境——此前有初步证据表明,即将推出的模型 Astra 可能触及“关键网络威胁”阈值;Greg Brockman 发布了《守护者的窗口》一文

OpenAI 加入了俄亥俄州的 8 千兆瓦 PORTS-Pike 数据中心项目,将 GPT-5.6 Sol 价格下调超 20%(为期三个月),把 ChatGPT Ads 扩展到 31 个欧洲市场,并预告了 Private Safety Processing,以继续提供“零数据保留”服务

他们还资助了 14 个政策项目,上线了 AI Futures 博客,宣布 DevDay Exchange 将在八座城市举办,并让首批 NVIDIA Vera Rubin 机架跑了起来

Codex 活跃用户突破 2000 万,GitLab 支持进入 beta;ChatGPT 新增共享对话线程、Apple Messages 插件,以及欧洲版 Computer History

另外,我发现了一个 Share prompt 操作、个人账户上自定义 GPT 创建的终止、Sketch 编辑器、ChatGPT with Friends,以及 Agent Email 收件箱

Anthropic 透露,Claude 设计出了针对 15 个靶点中 14 个的有效蛋白质结合剂,命中率远超常规水平,并在分析化学上与一家合同实验室打平

他们还通过 Claude Security 扫描为防护者带来了 Claude Mythos 5,并设立了 3500 万美元的 Defender Advantage Fund 支持开源安全

Claude Cowork 登陆移动端和网页端,Claude 现在可以发送 Gmail 邮件、管理 Google Drive 文件;computer use、浏览器工具,以及 Skills 和 Files API 全面正式可用

Claude Code 新增了 /design 技能、Concise 输出风格并修复了 Remote Control 问题;Claude Academy 上线,对齐研究涵盖了 CHIVE 和测谎器相关课题

查看英文原文
OpenAI and Anthropic weekly: teen ChatGPT, training pause, protein binders (Week 34, 2026)

OpenAI started rolling out ChatGPT for Teens globally, with an updated Model Spec and an education partnership with CodeAI

They also disclosed a two-week pause of RL training on their latest models to harden research environments, after early evidence that upcoming model Astra may meet the Critical cyber threshold, and Greg Brockman published The Defender's Window essay

OpenAI joined the 8-gigawatt PORTS-Pike data center project in Ohio, cut GPT-5.6 Sol pricing by over 20% for three months, expanded ChatGPT Ads to 31 European markets, and previewed Private Safety Processing to keep offering Zero Data Retention

They also funded 14 policy projects, launched the AI Futures blog, announced DevDay Exchange in eight cities, and got their first NVIDIA Vera Rubin racks running

Codex passed 20 million active users and got GitLab support in beta, and ChatGPT gained shared threads, an Apple Messages plugin, and Computer History in Europe

Plus, I spotted a Share prompt action, the end of custom GPT creation on personal accounts, a Sketch editor, ChatGPT with Friends, and Agent Email inboxes

Anthropic shared that Claude designed working protein binders against 14 of 15 targets, with hit rates well above what's typical, and matched a contract lab on analytical chemistry

They also brought Claude Mythos 5 to defenders through Claude Security scans, with a $35 million Defender Advantage Fund for open-source security

Claude Cowork arrived on mobile and web, Claude can now send Gmail emails and manage Google Drive files, and computer use, the browser tool, and the Skills and Files APIs became generally available

Claude Code got a /design skill, a Concise output style, and Remote Control fixes, Claude Academy launched, and alignment research covered CHIVE and lie detectors
meng shao@shao__meng · 中文博主 · 1 天前

稍等,eli5 Skill ?

这不就是咱们刚开始学提示词工程时,很经典的「给我 5 岁的孩子解释***」、「让我奶奶也能看懂」?

引用 Thariq @trq212Anthropic最近常用的技能:ELI5。/eli5 <你想解释的内容>。像向一个对该话题一无所知的人讲解,使用包含大图片和少文字的HTML呈现。查看被引原帖 ↗
Tanishq Mathew Abraham, Ph.D.@iScienceLuvr · 博主 · 1 天前

扩展Muon用于Diffusion Transformers

Meta和USC的新论文。

“我们首先在1.3B到15B参数的DiT上确立了Muon的扩展行为,表明其相对于AdamW的优化和生成质量优势在不同模型规模下保持一致。然而,在大规模下,每一步优化都执行的5步Newton-Schulz迭代(NS5),加上完整的动量具体化,引入了大量计算和通信开销,这可能会抵消Muon的步效率优势。我们引入了周期行级Muon,它每K步执行一次完整的NS5谱更新,并在其余步骤基于当前动量应用一个低计算和通信成本的行级约束更新。”

查看英文原文
Scaling Muon for Diffusion Transformers

New paper by Meta and USC.

"We first establish Muon's scaling behavior on DiTs from 1.3B to 15B parameters, showing that its optimization and generative quality advantages over AdamW persist across model scales. However, at scale, the 5-step Newton--Schulz iteration (NS5) performed at every optimization step, together with full-momentum materialization, introduces substantial computation and communication overhead that can offset Muon's step-efficiency advantage. We introduc Periodic Row-wise Muon, which performs a full NS5 spectral update once every K steps and applies a low compute and communication cost row-wise constrained update based on the current momentum at the remaining steps."


arxiv.org/abs/2608.20818
Gorden Sun@Gorden_Sun · 中文博主 · 1 天前中文圈高频 AI 资讯与开源项目博主

韩国一家做AI视频聚合服务的网站被中国做短剧的薅羊毛了,薅了大概几十万人民币。不过显然是程序来自动执行的,导致提示词、参考图、生成结果都自动分享到了社区。

相当有价值的一线AI自动化短剧的实战经验。

1、不是制作单部高品质艺术品,而是像工厂一样自动化作业

不会为了追求 100% 的一致性而消耗额外的时间与金钱成本。
单次生成即达到 90% 的一致性:提示词结构高度聚焦于“最小化 AI 重抽(Re-roll)成本”和“最小化后期剪辑时间”。
无需后期剪辑工序,实现镜头与镜头之间过渡连接的自动化。
即使是 10 秒以内、画面表现质量要求不高的短镜头,其提示词在结构上也经过了极其严密的推敲。
在制作超过 80 集的整部短剧中,该团队对提示词几乎不做大改动:
· 各分镜角色动作:短叙述句,不把资源浪费在过度的镜头调度演出上。
· 各分镜角色台词:直接复制粘贴原作台词。
· 起始与结束镜头:各角色状态仅用简短的一行字明确定义。

2. 故意破坏角色参考图,用于过人脸和提升特写表现

为防止视频生成模型直接机械地照搬参考图中的整张脸,通常会搭配使用“无脸全身图”和“面部特写图”。
但该团队更进一步:故意破坏只保留身份识别特征的角色面部区域,强制模型在特写镜头下也不直接从参考图中死板硬套。
(备注:这是韩国团队对SD不了解,实际主要是为了过人脸)

3. 将“光影与材质”定义为锁定画面质感(Look)一致性的最高优先级元素

仅光影设计与各材质高光部分就占据了提示词预算的 1/5。
不使用“刀刃闪闪发光”这种模糊的形容词,而是精确指定 3 种材质的相对比例:
长刀金属高光×1.5、皮肤高光×0.5、布料高光×0.3
压暗皮肤和布料、仅突出金属材质,通过指定相对比例而非绝对数值,确保无论重新生成多少次,画面材质的关系与层级始终保持一致。

4. 锁定每个生成的起始与结束画面状态,最大限度压缩剪辑耗时

不使用参考图作为起始帧与结束帧,而是纯靠文本固定视频开始与结束时各人物的状态。
有团队甚至会在豆包(DuoBao)提示词中直接写明“承接下一集的内容与交接状态”。
不需要逐个确认视频去手动对齐下一镜头的连贯性,彻底消除了制作瓶颈。只需一键点击首次生成,下一段视频的首个镜头、首句台词以及角色的状态变化就会呈链式自动传递。

5. 依靠配音驱动整体叙事与节奏,而非依赖视觉调度

专为快节奏、重台词交互的微短剧定制。
团队极少把提示词资源耗费在镜头布局与运镜上。
将配音色调的一致性与情感演绎直接固化为可复用的文本结构:
音色与性格:标记为绑定在角色身上的“常量”。
状态与语调:标记为随分镜变化的“变量”。

角色的动作表演严格与台词对齐,依据台词中具体词汇的出现时间节点依次指示动作。

过人脸图片:

AK@_akhaliq · 博主 · 23 小时前HuggingFace 研究员,每日 AI 论文速递

InfinityEdit

使用轻量级编辑点火适配器的无限视频编辑

论文:
huggingface.co/papers/2608.2…

查看英文原文
InfinityEdit

Infinite Video Editing with a Lightweight Edit-Ignition Adapter

paper:
huggingface.co/papers/2608.2…
elvis@omarsar0 · 博主 · 1 天前

掌握你自己的智能技术栈。

这是我对今天所有正在构建东西的人的建议,无论你是在起步阶段还是正在进行中。

这其中的一部分,就是要懂得如何打好基础技术,并在其上构建应用和体验。

引用 Paul Graham @paulg有人问我如果17岁会做什么。我会学习从零开始构建LLM,然后用能获得的任何硬件训练尽可能强大的模型。查看被引原帖 ↗
查看英文原文
Own your intelligence stack.

That’s what I would advise anyone building today. Whether you are starting or in the process.

Part of that is understanding how to build the foundational technology and the application/experiences on top of it.
DeepLearning.AI@DeepLearningAI · 公司官方 · 23 小时前

Grok 4.6 配合 Cursor 数据,智能体效率起飞。🧠 新模型完成长期知识工作任务,轮数只要其他顶级模型的一半。📉 轮数少了,复杂智能体应用成本就低了。💸

深入了解架构细节:
hubs.la/Q04v2nWD0

查看英文原文
Grok 4.6 with Cursor data = a massive leap in agentic efficiency. 🧠 The new model completes long-running knowledge work tasks in half the turns of other leading models. 📉 Fewer turns mean lower costs for complex agentic apps. 💸 

Dive into the architecture details:
hubs.la/Q04v2nWD0
◔ 1.2 万 次浏览(3 条合计)♥ 77⇄ 1新品看原帖 ↗
小互@xiaohu · 中文博主 · 1 天前小互,中文圈高频 AI 资讯站 Xiaohu.AI 主理人

为什么大部分 CEO 的 AI 尝试都在“假装进步”?

企业 AI 转型 CEO 指南

Bain & Company 最近发布了一套由总论、七个决策和结论构成的 CEO 指南

这份指南盯住一个尴尬现象:企业做了越来越多 AI 试点,但能把它变成竞争优势的仍然很少

核心痛点:为什么大部分 CEO 的 AI 尝试都在“假装进步”?

最典型的症状有五种。

1. 把试点数量当作转型深度
客服团队做一个回复草稿助手,财务团队做一个报表问答机器人,法务团队再做一个合同摘要工具。年底一汇总,公司有几十个 AI 用例,汇报材料非常漂亮。

但把这些项目拿走,企业的核心工作流可能一点没变:客服仍要在五个系统之间复制信息,财务仍在月底手工对口径,法务仍靠资深员工判断例外。

2. 把旧流程加一个 AI 按钮,当成重构业务
许多 AI 转型只是在低效流程里加快一个步骤。过去员工花 20 分钟写邮件,现在模型 2 分钟写完;可邮件为什么存在、前后为什么需要四次交接、谁有权做决定,这些问题一概不碰。

结果是局部快了,端到端未必更快。上游数据仍然混乱,下游审批仍然拥堵,员工还要花时间核对模型输出。

3. 把短期 ROI 当成唯一筛选器
约 85% 的 CEO 主要用 AI 追求近期降本和效率改善,希望借此为后续转型筹资。

致命之处在于,一旦所有项目都被要求当年证明 ROI,组织就会系统性选择最容易量化、最接近旧做法、也最不可能改变竞争位置的任务。需要跨部门重做的客户旅程、供应链或产品开发流程,因为周期更长、责任更复杂,反而最先被预算机制淘汰。

4. 把供应商能力误当成自己的能力
给全员开通一套带 AI 的 SaaS,不等于企业拥有了 AI 能力。竞争对手明天也能买同一套功能。更麻烦的是,如果数据语义、工作流逻辑、工具连接和使用反馈全部留在供应商平台里,企业每使用一次,可能都在增强对方的产品,却没有给自己的下一次部署留下多少资产。

5. 战略口号很激进,运行机制却极其保守
最常见的荒诞场景是:CEO 宣布 AI 是公司未来三年的最高优先级,项目却仍被要求逐个证明当年回报;团队说要建设专有能力,关键工程仍完全外包;管理层要求快速迭代,每次上线却要经过数月串行审批。

这些症状指向同一个问题:企业把“做了更多项目”当成“拥有了更强能力”。

项目结束后,失败没有变成新的测试题,数据仍然各说各话,老员工的判断仍留在个人脑中,其他团队也不知道这次踩过什么坑。下一次部署还是接近从零开始。项目铺得越广,重复劳动也跟着铺得越广。

“专有智能”护城河?

领先企业之所以拉开差距,在于打造了对手用钱买不来的三样核心资产:

专属数据(Proprietary Data):只属于你自己的客户、业务运营和交易记录,随着业务运转越滚越多、越用越值钱。

编码化工作流(Encoded Workflows):把资深员工的专业判断和打法沉淀下来,写成智能体(Agents)直接规模化执行。

闭环学习机制(Learning Architecture):人机协同形成反馈闭环。人指导 AI,AI 提升人的产出,沉淀的数据再反哺模型,形成持续拉开差距的“飞轮效应”。

Gorden Sun@Gorden_Sun · 中文博主 · 1 天前中文圈高频 AI 资讯与开源项目博主

巨头造火箭,老板骑电驴

Anthropic推出地表最强模型Fable 5已经两个月了,但是营收只占Anthropic所有模型的11%,而且增长几乎停滞,完全卖不动。

过去大家总觉得,既然要用 AI,肯定要买最强、最聪明的版本。现在的企业老板们一算账,Fable是强一些,但不值这么贵的价格。所以Opus 5的销量增长迅速,已经超过了Fable 5。

还有隔壁OpenAI的性价比极高的 GPT 5.6 Sol和中国的便宜开源模型,Fable 5的故事不好讲。

AI智力达到一定水平之后,企业追求的就不是最高智力了,而是更具性价比的智力。

数据来源:
ft.com/content/5ee49718-c258…

Ethan Mollick@emollick · 创始人 · 1 天前沃顿商学院教授,AI 应用研究权威

多亏这条帖子,算法现在净给我推荐 AI 发现的各种晦涩数学公式,我还得装作能看懂。

查看英文原文
Thanks to this post, the algorithm is showing me ever more obscure mathematical discoveries due to AI which I also have to pretend to understand.
GitHubDaily@GitHub_Daily · 中文博主 · 1 天前

最近发现一款可清除 AI 生成的图片或视频里水印的开源工具:Remove-AI-Watermarks。

支持移除 Gemini、Nano Banana 的星标,还有豆包、即梦、Qwen、可灵、元宝、百度这些平台的角标。

而移除视频水印,则支持 Sora、Veo、Seedance、Hailuo、Kling 这些平台。

GitHub:
github.com/wiltodelta/remove…


无论是可见的水印,还是隐藏的水印都能处理,包括谷歌的 SynthID。

隐藏水印靠模型重绘打散,效果好但要有 CUDA 显卡,只清元数据的话 CPU 就够。

希望大家只用于处理自己拥有的内容,不针对图库预览图这类保护他人付费内容的水印。

另外作者还提供了一个在线体验 Demo ,那些可见水印和元数据移除可以直接使用。

Chubby♨️@kimmonismus · 博主 · 1 天前Chubby,高频 AI 新闻聚合博主

theinformation.com/articles/…

查看英文原文
theinformation.com/articles/…
Tanishq Mathew Abraham, Ph.D.@iScienceLuvr · 博主 · 1 天前

管理自动研究的 agents 就像给一群博士后学生当导师。

我在想如果把怎样使用这些 agents 的方式说得更明确,会不会有帮助。

查看英文原文
Managing agents for autoresearch feels a lot like being a PI advisor for a bunch of junior PhD students.

I wonder if making this more explicit in how I use these agents might help.

本站由 Jedee杰哥 打造 · 公众号「Jedee杰哥」每早送 AI 日报

姊妹站:𝕏 简中账号数据榜单 · X 关注 @jedeeai · RSS 订阅 · AI 日报 · 历史归档