JEDEE AI
存档 2026-08-16

8 月 16 日(北京时间)全球 AI 圈推文存档,按曝光排序,共 95 条。
← 返回最新 全部归档

全部情报 每小时更新 · 事件已合并同类项

内容 公司
Dario Amodei@DarioAmodei · 创始人 · 1 天前Anthropic 联合创始人兼 CEO
连环推 ×2

第1部分:感谢 Gavin 进行了一场格外深思熟虑的交流。我平时不太花时间在社交媒体上,但这次想参与进来,因为它真正触及了一场重要对话的核心。

首先,关于监管,我认为"要么通过监管把权力集中在少数选定的公司和政客手中,要么广泛分散"是一个虚假的两难选择。我知道硅谷有一种简化的说法,认为监管=监管捕获=权力集中,但我一直觉得这是对世界过于简化的描绘。这个圈子之外的很多人把监管视为限制企业权力、惠及普通人的工具。我并不完全同意那种观点,而是认为事情很复杂,真的取决于"监管"具体包含什么。但尤其值得注意的是,我认为持"监管=监管捕获=权力集中"框架的人往往低估了客观、公正的制度程序所具有的去中心化力量。举个粗糙的类比:正式法院系统有时会显得刻板和精英化,但它在捍卫弱势个体权利方面远比替代方案——暴民正义——做得好。在最佳状态下,制度可以把权力赋予想法而非个人,从而实现权力分散。

这正是 Anthropic 一直非常谨慎制定政策提案的原因。我们非常努力地提出对前沿 AI 公司不利(减缓其发展)但对小型竞争者有利的提案。加州 SB53(我们支持的)甚至备受诟病的 SB 1047(我们对其态度矛盾)都完全豁免了收入或模型训练成本低于一定门槛的公司(SB 53 是 5 亿美元,1047 更低但我们对这一点提出了异议)。最近,我们在 CAISI 和白宫倡导的测试流程对前沿模型的测试要求比非前沿模型更严格——这对挑战者更有利。同样,"Pacing the Frontier"公开信设想(或至少 Anthropic 偏好的实施方案设想)调节最顶尖模型的推进速度,同时不限制追赶中的参与者。这损害了前沿实验室的商业利益,却帮助了挑战者,包括开放权重模型!

总的来说,我的观点是,AI *在结构上*就是一种倾向于集中权力的技术,原因与监管无关(更多与 scaling laws 的极端含义有关)。开放权重确实在这方面有些帮助,但远非充分的解决方案,因为它们只是将集中程度部分转移到拥有最多算力和芯片的人手中(这些人大致就是前沿实验室,加上可能还有硬件供应商)。相比之下,我认为正确的"游戏规则"可以同时做到:(a)应对 AI 的网络、生物和 alignment 风险;(b)在制度上约束前沿 AI 公司的权力;(c)为开放权重模型留出空间,同时解决它们带来的特定风险。

顺便说一下,我并不认为过去几个月的事件"未能促成[我]倾向的监管路径"。据报道,特朗普政府采取的做法——对前沿模型进行部署前测试,以及在开放权重模型接近前沿水平时也对其进行测试——我非常支持,当然还得看到细节才能确认。我也支持 Demis Hassabis 关于类似 FINRA 机构的设想。这与六个月前形成鲜明对比,当时业界大多数人还在推动全盘否决各州监管,而联邦层面似乎也没有任何明确方案。

引用 Gavin Baker @GavinSBakerSholto, thank you for setting the record straight. Larger issue is that multiple very serious people in Silicon Valley have heard some variation of this and believe it to be true. And the reason it is believable to so many is that it is consistent with Dario’s public messaging and what he outlined in the essay you shared: this technology *might* be dangerous for humans in multiple ways, could lead to extreme concentration of economic power (as outlined in the essay) and therefore needs to be regulated thoughtfully. I agree with the potential risks and I believe Dario makes all of these arguments in good faith. 
As discussed on the pod, if one agrees that AI *might* be dangerous, there are two ways to address this potential risk. Either concentrate it in the hands of a chosen few companies and politicians via regulation or distribute it widely. Essentially boils down to whether one believes AI is too dangerous to concentrate or too dangerous to distribute. There are reasonable arguments on both sides, but I profoundly agree with Zuckerberg’s statement that: “The notion that AI is so dangerous that the only safe path is an extreme concentration of power seems inherently problematic. Historically, hoping that an absolute power will benevolently provide for humanity if sufficiently enlightened has not led to safe or positive outcomes.” And as Dario says in the aforementioned essay, “some may object that we can simply keep AIs in check with a balance of power between many AI systems, as we do with humans.” I believe this is the best path forward: I want as many AIs as possible to maximize the odds that one shares my own particular values. And as Dario notes, no human has ever been able to take over the world. At this point, I think safe to say that Dario has lost the argument. His messaging has failed to result in his preferred regulatory path. The fact that the only solution to the recent incident where an unreleased advanced OpenAI model hacked Hugging Face was an 查看被引原帖 ↗
查看英文原文
1/2 Thanks Gavin for an especially thoughtful exchange. I don't usually spend much time on social media but I wanted to engage here because it really brings out the heart of an important conversation.

First, on regulation, I think that “either concentrate it in the hands of a chosen few companies and politicians via regulation or distribute it widely” is a false choice.  I know that there’s a sort of Silicon Valley shorthand where regulation = regulatory capture = concentration of power, but I’ve always found this to be an overly simplified picture of the world.  Many people outside this bubble think of regulation as something that constrains corporate power and benefits ordinary people.  I don’t necessarily agree with that perspective either, rather I think it’s complicated and really depends on what the “regulation” consists of.  But in particular I think that those in the “regulation = regulatory capture = concentration of power” frame often underrate the decentralizing power of objective and fair institutional processes.  A crude analogy is that the formal court system can sometimes feel stuffy and elitist, but it does a much better job of defending the rights of vulnerable individuals than the alternative, mob justice.  At their best, institutions can vest power in ideas rather than people, and thereby decentralize that power.

This is why Anthropic has always made its policy proposals very carefully.  We try very hard to make proposals that disadvantage (slow down) frontier AI companies while *advantaging* smaller competitors.  California’s SB53 (which we supported), and even the much-maligned SB 1047 (which we were ambivalent on), completely exempt any company below a certain amount of revenue or model training costs from being covered at all (it was $500M for SB 53, lower for 1047 but we objected to that).  More recently the testing process we’ve advocated for at CAISI and the White House involves more rigorous tests for frontier models than off-frontier models — something that differentially advantages challengers.  Similarly, the “Pacing the Frontier” letter envisions (or at least Anthropic’s preferred implementation of it envisions) modulating the pace of the very best models while not constraining those who are catching up.  This hurts the business interests of the frontier labs and helps challengers, including open-weights!

Overall my view is that AI is *structurally* a technology that tends to concentrate power, for reasons that have nothing to do with regulation (more to do with the extreme implications of the scaling laws).  Open-weights do help some with this but are nowhere near a sufficient solution because they simply shift the concentration somewhat to those with the most compute and chips (which are roughly the frontier labs plus maybe hardware providers).  By contrast I think the right “rules of the road” can simultaneously (a) address AI’s cyber/bio/alignment risks, (b) institutionally constrain the power of the frontier AI companies, and (c) leave room for open-weights models while also addressing the specific risks that they bring.

BTW I do not think that the events of the last few months have “failed to result in [my] preferred regulatory path”.  The approach that the Trump administration is reported to be taking — pre-deployment testing for frontier models, and also testing of open-weights models when they get closer to the frontier — is one that I am very supportive of, though of course I have to see the details to be sure.  I am also supportive of Demis Hassabis’ ideas around a FINRA-like entity.  This contrasts with six months ago when most of the industry was still pushing for preemption of all state regulation and no apparent federal approach either.
2/2 Second, on the messaging around AI.  I do not agree that my messaging has been disproportionately negative.  In fact it has been about equally balanced between risks and benefits: I’ve written one major essay about each, and even in interviews where I discuss the risks, I make sure to frequently mention the incredible benefits as well as proposing possible solutions to the risks (short clips from my interviews that end up on social media tend to be disproportionately negative, as that gets clicks).  In fact, I wrote Machines of Loving Grace because I didn’t feel the AI industry was painting an inspiring enough picture of how the technology could radically transform the world for the better.  The bulk of the essay is devoted to refuting skepticism of AI’s potential in health and biology, and showing why I think it will actually be possible to cure most human disease in ~5-10 years, as crazy as it may sound to ordinary people and frankly to biologists as well (I used to be one!).  And, if you read my most recent essay (Policy on the AI Exponential), I discuss concrete proposals for how to streamline the FDA process to make sure the deluge of AI-accelerated drugs isn’t slowed down by the regulatory process.  I feel the urgency here: I lost my father to Hepatitis C only a few years before the development of direct-acting antivirals (sofosbuvir), which cure 95% of patients and probably would have cured him.

I do agree that the public has a negative view of AI (and that this is a big problem), but I don’t think it is primarily caused by me or any other AI leader warning about AI’s risks.  I think it is fundamentally a crisis of trust.  I think that ordinary people don’t trust companies, governments, or the tech industry and always suspect that we are cooking up some new way to screw them over.  The causes of this go back decades and AI is just the latest iteration of it.  I don’t think that a glitzy marketing campaign with a positive spin (which some have advocated that Anthropic do) is the way to win back that trust — at this point, saying that AI will cure cancer is more a cliche than it is inspiring, and most people think it is deceptive.  The thing that will work is *actually curing cancer*.  I think by far the most accurate criticism of AI companies including Anthropic is that we haven’t yet delivered on our big promises to benefit the world.  That is totally on us, and I think it’s the criticism you should be making, instead of all this stuff about messaging and marketing.

We are however doing our best to fix this: Anthropic is ramping up its efforts very quickly in biology and medicine, and we hope to have incredible results in the coming years and some early glimmers in the coming months.  When we’ve actually accomplished something real, the whole world will hear about it, as loudly as possible, you have my word on that.  But until then I don’t want to make empty promises, and in the meantime I feel compelled to speak honestly about the very real risks of AI and how to address them.  Honesty is the right thing on the merits, and in terms of public credibility and trust it is no worse than, and may in fact be better than, an approach that ignores or distracts from risks which people instinctively understand are real.
Bindu Reddy@bindureddy · 创始人 · 1 天前Abacus.AI CEO,AI 行业观点博主

🚨 推出 Voice Agents - AI 通过一条文本提示即可处理企业级客户支持

- 免费电话号码
- 接入你所有的系统
- 自定义人类般的语音
- 一条提示即可设置

用文本配置和设置企业电话。在 Abacus AI 上构建和发展你的 AI native 业务。

随 ChatLLM 订阅附赠

查看英文原文
🚨 Announcing Voice Agents - AI Handles Enterprise Class Support With Just One Text Prompt

- toll-free phone number
- connect to all your systems
- custom human-like voices
- set up with one prompt

Use text to configure and set up enterprise calls. Build and grow your AI native business on Abacus AI.

Comes included with your ChatLLM subscription
◔ 300.4 万 次浏览♥ 278⇄ 19▶ 含视频新品看原帖 ↗
Tibo@thsottiaux · 公司官方 · 1 天前ChatGPT 产品官方账号

OpenAI 现在发生了不可思议的事情。气氛很高涨。

查看英文原文
Incredible things are happening at OpenAI right now. Energy is high.
Tibo@thsottiaux · 公司官方 · 1 天前ChatGPT 产品官方账号

你们公司还在用 Opus 而不是 Sol,这是为啥?难不成价格对你不是个事儿?那什么样的情况能说服你切过去?

查看英文原文
If you have a company and still use Opus instead of Sol, why so? Does price not matter to you? What's something that would convince you to switch?
Min Choi@minchoi · 博主 · 1 天前AI 产品演示博主,专门展示新工具玩法
连环推 ×11

Grok 4.6 现在牛逼得有点吓人。

游戏、3D 世界、游戏预告片、甚至真实的 3D 打印。

10 个疯狂的案例:

查看英文原文
Ok Grok 4.6 is getting scary good.

Games. 3D worlds. Game trailers. Even real-world 3D prints.

10 wild examples:
1. Grok 4.6 built this FPS shooter autonomously for 48 hours of nonstop AI work
2. Grok made its own video game trailer

It played, recorded, edited, and voiced everything itself.
3. Grok 4.6 designed a real part for a Tesla

It found the measurements and made the 3D-print files.
4. Grok 4.6 built Minecraft clone
5. Grok got its first iPhone
6. Grok 4.6 turned imagination into interactive 3D
7. Grok 4.6 recreated a black hole simulation
8. AI city-building improved this much in 9 months
9. An entire Japanese town generated in 3D
10. Grok 4.6 gets near-Fable quality at 1/10 the cost

Same prompt, half the time, dramatically cheaper.
◔ 95.8 万 次浏览(2 条合计)♥ 3,123⇄ 334▶ 含视频演示看原帖 ↗
Tibo@thsottiaux · 公司官方 · 1 天前ChatGPT 产品官方账号

让 Sol 为你管理高效的 Luna agents 车队。

这些模型彼此了解得很好,以极快且高效的方式协作完成任务。

引用 eric provencher @pvncher我们本周推出了 multi agents v2 能够委托给任何支持的模型(包括 Luna)的功能。花了不少时间确保这个功能运行可靠。查看被引原帖 ↗
查看英文原文
Let Sol manage an efficient fleet of Luna agents for you.

These models know each other well and collaborate to achieve the result in an incredibly fast and efficient way.
Tibo@thsottiaux · 公司官方 · 1 天前ChatGPT 产品官方账号

关于token和每个token的价格。

我说过会详细聊聊这个话题,所以现在来了:OpenAI的token不等于别的模型的token。我们拿每百万token的美元价格来比较AI,好像token是个标准单位,比如克或千瓦时似的。但并不是这样。不同的模型生成同样的文本,用的token数量不同,也就是说低价token并不一定意味着账单更低。

想象两个一模一样的披萨。一个切成8片,每片2美元。另一个切成16片,每片1.25美元。第二家店卖得更便宜的片数,但整只披萨20美元,而不是16美元。尴尬的是……你的胃根本不在乎你吃了几片。

我知道你现在饿了,但回到token话题。在一个跨英语、技术、多语言和数字文本的小对比中,我们用于GPT-5.6 Sol的tokenizer用了766个token,而预计Claude Opus 5会用约1,170个。这差了大概34.5%的token,差距相当明显。你能得到完全一样的文本,却要为那些多余的token买单。每token价格其实没讲清楚这个事儿。

即使修正tokenizer差异,也忽略了一个更大的核心。真正重要的是每个成功结果的价格,为此你可以用基准测试作为起点,但说到底,你得在自己的用例中试跑和实测。

就这些。愿token滚滚而来。

查看英文原文
On tokens and prices per token.

I said I’d write more about this, so here goes: an OpenAI token != another model’s token. We compare AI prices in dollars per million tokens as if a token were a standardized unit, like a gram or a kilowatt-hour. It isn’t. Different models use and produce the exact same text using different numbers of tokens, which means a lower price per token does not necessarily mean a lower bill.

Imagine two identical pizzas. One is cut into 8 slices at $2 each. The other is cut into 16 slices at $1.25 each. The second place advertises cheaper slices, but the whole pizza costs $20 instead of $16. Bummer ... your stomach doesn't actually care about the number of slices you just ate.

I know you are hungry now, but back to tokens. In one small comparison spanning English, technical, multilingual, and numerical text, the tokenizer we use for GPT-5.6 Sol used 766 tokens versus an estimated 1,170 for Claude Opus 5. That's a very significant difference of about 34.5% fewer tokens. You can get the same exact text, but pay for all those extra tokens. The price per token doesn't really tell this story.

Even correcting for tokenizer differences misses the bigger point. What actually matters is price per successful outcome, and for that you can use benchmarks as a starting point, but really you have to try it and measure on your own use cases.

That's all. May the tokens flow.
Guillermo Rauch@rauchg · 创始人 · 1 天前Guillermo Rauch,Vercel 创始人兼 CEO

我觉得大多数人都没意识到,React能这么成功,
很大程度要归功于@shadcn。

我曾把React称作“成年人的乐高积木”。但实际上,那可能更多是在描述积木的几何形状——也就是规格本身。

Shadcn才是人们对React真正的期待。高质量可复用组件,还能灵活调参。它是个伪库,虽然提供了代码,但本质上是让你把这些代码消化进自己的语境里,然后重新组合。

查看英文原文
I think people don’t realize how much of the success of React is actually
@shadcn
.

I once referred to React as the actual “LEGO brick for adults.” But in reality it was more the description of the geometry of the bricks. The spec.

Shadcn is what people actually wanted from React. Reusable high quality components that also tunable. It’s a pseudo-library. There’s code but it’s actually meant to be digested into your context window and remixed.
Qwen@Alibaba_Qwen · 公司官方 · 1 天前阿里通义千问大模型团队

30亿下载量!一眼能数清几个零吗?😎 谢谢大家满满的爱,咱们继续一起成长!🌱

引用 Bloomberg @businessAlibaba's open-weight models have accumulated more than 3 billion global downloads in the past six months, eclipsing Meta, Alphabet and domestic peers to become the world’s No. 1 artificial-intelligence model bloomberg.com/news/articles/…查看被引原帖 ↗
查看英文原文
3000000000 downloads! Can you count the zeros at a glance? 😎 Thank you all for the incredible love. Let's keep growing together! 🌱
Guillermo Rauch@rauchg · 创始人 · 1 天前Guillermo Rauch,Vercel 创始人兼 CEO

我对寻找并资助下一个迪士尼感兴趣。

真正忠于其原始价值观,专注于美国精神。品德、家庭、克服逆境、乐观、追求幸福。

AI 让每个人都能创作。定义下一代的是我们*用它创作什么*。下一批标志性形象。超越绘制方式或渲染效果的下一个独特品牌。

米老鼠经历了多次技术革命而幸存至今,正因为它所代表的意义。我认为大多数大型娱乐公司已经忘记了这一点。

如果你在这个领域创作了很酷的东西,需要技术或财务上的帮助,直接给我发私信。

查看英文原文
I’m interested in finding and funding the next Disney.

True to its original values, focused on the American spirit. Merit, family, overcoming of adversity, optimism, the pursuit of happiness.

AI makes everyone able to create. What will define our next generation is *what* we create with it. The next icons. The next unique brand that goes further than how it’s drawn or rendered.

Mickey Mouse survived several technological revolutions because of what it represented. I think most big entertainment companies have forgotten that.

If you’re creating something cool in this space and need help, technologically or financially, slide into my DMs.
Greg Brockman@gdb · 创始人 · 1 天前Greg Brockman,OpenAI 联合创始人兼总裁

朝着永远不用手动选择模型的方向发展

引用 eric provencher @pvncher我们本周推出了 multi agents v2 能够委托给任何支持的模型(包括 Luna)的功能。花了不少时间确保这个功能运行可靠。查看被引原帖 ↗
查看英文原文
towards never having to manually select a model again
Chubby♨️@kimmonismus · 博主 · 1 天前Chubby,高频 AI 新闻聚合博主

OpenAI 在 $200 套餐上还要再收 $80 来重置使用限额。比预期的贵啊

引用 NIK @ns123abc周重置费用超过80美元,这太贵了。查看被引原帖 ↗
查看英文原文
OpenAI is charging $80 for a usage limit reset on $200 plan. Higher than expected y
Qwen@Alibaba_Qwen · 公司官方 · 1 天前阿里通义千问大模型团队

🚀Qwen3.8-27B在笔记本上飞起来了,融入我们的工作与日常生活。感谢点名!
@atomic_chat_hq

引用 atomic.chat @atomic_chat_hqRun Qwen3.8 27B locally via Atomic Chat💥 We released Atomic Dynamic GGUF quants, from 8-bit (28.9 GB) down to 1-bit (8.5 GB), and measured all other Qwen3.8 GGUFs in the community AD-IQ3_S runs on a 16GB MacBook Air and picks the same next token as the BF16 original 92.4% of the time查看被引原帖 ↗
查看英文原文
🚀Qwen3.8-27B flies on a laptop, becoming part of our work and daily lives. Thanks for the shoutout!
@atomic_chat_hq
Qwen@Alibaba_Qwen · 公司官方 · 1 天前阿里通义千问大模型团队

太感谢整个社区了!Qwen3.8-27B 现在是 Hugging Face 上的 #1 热门模型了!🏆 快来试试吧,告诉我们你的想法。🤗

引用 Julien Chaumond @julien_csoon 10k查看被引原帖 ↗
查看英文原文
Huge thanks to the whole community! Qwen3.8-27B is now the #1 trending model on Hugging Face! 🏆 Try it out and let us know what you think. 🤗
Ethan Mollick@emollick · 创始人 · 1 天前沃顿商学院教授,AI 应用研究权威
连环推 ×3

对于不玩视频游戏的人来说,小型独立开发者中对任何AI使用的审查是时刻在进行的。他们是这个行业资源最紧张的公司,利润微薄且艺术愿景经常被折扣,但他们受到的惩罚比大开发商严厉得多。

查看英文原文
For those who don’t follow video games, there is constant policing of any AI use among small, indie developers. They are the most resource constrained firms in a field where profits are rare & artistic vision is often compromised, but they are punished more harshly than big devs.
Consumer segment feelings about AI is going to be a bigger driver than economics in determining which fields become efficient and which become artisanal in reaction to AI.
That is especially challenging for companies whose audiences span both types of audiences, like indie video games. Also this means that most indie developers are incentivized to use AI and hide it.
Amjad Masad@amasad · 创始人 · 1 天前Amjad Masad,Replit 创始人兼 CEO

“AI结构性集中权力”这个说法,建立在当前AI极度依赖算力的基础上,却忽略了计算性价比在过去125年里呈超指数增长的事实。

算法改进和硬件效率的持续提升意味着,没有理由认定AGI级别的能力永远需要数据中心才能运行。

扩展定律并非物理法则。它们只是在特定架构、目标、数据集等条件下观察到的经验关系。只要改变这些因素中的任何一个,得到的扩展曲线就会不同。

如果大脑能作为参考,真正的AGI很可能非常高效。事实上,当前的扩展定律或许是个bug而非特性——它恰恰说明我们目前发现的机器学习算法有多么低效。

引用 Dario Amodei @DarioAmodeiAnthropic关于AI监管的观点:"集中vs分散"是虚假二元对立。强调公平制度可将权力寄托于思想而非个人。Anthropic政策谨慎设计,旨在减缓frontier AI公司速度而促进竞争者。支持的加州SB53及中立的SB1047都豁免低收入公司,同时倡导对frontier模型实施更严格的测试标准。查看被引原帖 ↗
查看英文原文
The argument that “AI structurally centralizes power” because it’s currently compute hungry ignores 125 years of super exponential growth in compute price-performance.

Improvement in algorithms and continued hardware efficiency gains means there is no reason to assume AGI-level capabilities will always require a data center to run.

Scaling laws are not laws of physics. They’re simply empirical relationships observed for particular architectures, objectives, datasets etc. Change any one of those factors and you get a different scaling curve.

If the brain is any guide, true AGI will likely be very efficient. In fact, current scaling laws might a bug and not a feature. It shows how inefficiency of the ML algorithms we discovered so far.
@levelsio@levelsio · 博主 · 1 天前独立开发者标杆,AI 产品连续创业者

能在 170GB 的 Photo AI SQLite 数据库上用 Litestream 吗?

因为 AI 说这个数据库对它来说可能有点太大了?

查看英文原文
Can I do Litestream on 170GB Photo AI's SQLite DB?

Cause AI says it might be bit too big for it?
宝玉@dotey · 中文博主 · 1 天前宝玉,中文圈 AI 翻译与科普大 V

我最近就发现,即使聪明如 Fable 5,如果你只是给它一个目标让它优化,它可能也就是在既定的框架下想办法帮你优化到极致,但是它很难跳出既定框架,发现一条完全不同的路线。

比如说最近有用户反映使用 BaoCut 转录速度慢,我就反复的用 Codex 去转录视频,然后让 Codex 或者 Fable 去分析瓶颈在哪里,该怎么优化,然后每次它们都能给我一堆理由和看起来靠谱的优化方案。

比如它会建议多开启 workers(subagent),对 workers 预热之类。让它按照方案优化后,似乎有效果但是又不明显。

最后还是自己去分析数据包,发现耗时长还是输出的 JSON 格式太长了,导致生成、校验时间较长。

如果不用 JSON 格式,比如纯文本,或者简单的 html 格式,会更节约 token,只不过这样程序解析会很复杂,比如对字幕润色分段,输出是纯文本的话,就要做 diff 比较润色后更改的内容,还要把变更后的部分,重新对应到原始字幕的单词上。

简单来说,就是以前为了让程序简单就让模型输出复杂更费 token;如果要节约 token,就可以让模型的输出简单,但是程序解析会很复杂。

按照这个思路改进后,效果很明显(对比图2图3),调用次数从 33 次降低到 12 次;时间从 31 分钟降低到 18 分钟。测试张小珺那期将近 7 小时的访谈,完整润色也只需要 42 分钟。

如果不走 Agent 走 Cloud 模型的话,一个小时的视频,完整的翻译校对成双语字幕,用 DeepSeek v4 Flash,成本大于是 ¥0.4元。

有兴趣可以试试看效果,Mac 有专门的 App,Windows 支持 cli 或者 skill。
BaoCut :
baocut.app/

引用 宝玉 @dotey很多人不知道该怎么用好 Agent 的 /goal 功能,也就是说给 Agent 一个目标,让它长时间运行,直到目标完成为止。 其实没你想的那么复杂,注意几个点: 1. 你的目标是什么 2. 如何验证结果 3. 停止条件 比如说我这两天做的一个性能优化的任务,Fable 5 帮我把视频转录性能优化了2倍多(图2),提示词很简单(图1): > /goal 帮我优化当前 cli 的转录大视频的性能,在遇到像这样大体积的视频时,需要优化转录性能,请以 Moss 模型测试这个视频(英文为主,多语言)转录,建立基准,然后分析性能瓶颈,尝试优化,直到你觉得已经没有优化空间了。注意你的主要任务是分析、编排和验证,具体任务尽可能交给 subagent(Opus5)去执行 首先用 /goal 表示这是一个需要长时间执行的任务,需要反复执行,不能运行一会就结束了。 然后给它一个视频让它先自己跑一遍转录,记录一下关键数据,建立基准。 基于转录时收集的数据,Agent 自己可以去分析原因,去自己优化,优化完成后再去跑一遍,记录数据,对照前面的基准看是更好了还是更坏了。 结束条件是它自己觉得已经没有优化空间了就结束。之所以我没给它一个具体指标,是因为我也不知道能优化多少,如果指标太容易达到,它一轮可能就结束了;如果指标太难超出物理极限也没意义,反而可能会出现为了优化去做一些极端的事情。 最后一句让它开subagent执行子任务是因为 Fable 5 太贵,全程 Fable 5 用不了多久就要额度不够了,加了这句就耐用多了,而且质量也挺好。 --- 还有些时候,想到一个新的技术方案,但并不知道这方案是不是有效,那也可以让它开个worktree,去验证一下是不是靠谱,看数据是更好还是更坏,如果没提升就没必要做了。(参考图3)查看被引原帖 ↗
Ethan Mollick@emollick · 创始人 · 1 天前沃顿商学院教授,AI 应用研究权威

我们确实已经处于Vinge所定义的Singularity阶段,但还没有达到von Neumann最初定义、Ulam回忆的那个Singularity。

引用 David Pfau @pfau呼吁科技从业者使用'奇点'术语前先阅读Vernor Vinge和Ray Kurzweil。这个词不仅是指某项技术驱动大规模资本循环。查看被引原帖 ↗
查看英文原文
We are definitely in a Singularity as defined by Vinge, but not yet in a Singularity as originally defined by von Neumann as recalled by Ulam.
Gary Marcus@GaryMarcus · 博主 · 1 天前

如果这是真的,这些公司为什么不更透明地公开财务信息呢?

引用 dnap @dnapwayGavin Baker认为Anthropic S-1会颠覆投资者认知。许多宏观和价值投资者错误假设token被补贴,但实际上Anthropic已盈利,token生成现金。投资者对此事实认识不足。查看被引原帖 ↗
查看英文原文
if this were true, why aren’t the companies being more transparent about their finances?
◔ 8.2 万 次浏览♥ 242⇄ 24▶ 含视频观点看原帖 ↗
Gorden Sun@Gorden_Sun · 中文博主 · 1 天前中文圈高频 AI 资讯与开源项目博主

马斯克的SuperGrok Heavy现在超值了,之前印度区薅到的朋友赚大发了,不仅送X Premium+会员,还送Cursor Ultra会员!
300美元,3家会员,Grok、Cursor、X、Grok Bot,Cursor Ultra还有400美元的Claude/GPT等三方模型的额度。

deepseek哪怕涨价以后,跟同性能的竞争对手相比还是太便宜了,实话实说,deepseek v4 pro就是我最舒服的模型,再涨三倍我也1一定愿意掏钱买。

骂deepseek涨价的都是凑热闹的闲人,我们真正deepseek用户是愿意多掏钱买的,deepseek越涨价,我们越爱用。

◔ 10.2 万 次浏览(2 条合计)♥ 163⇄ 5观点看原帖 ↗
Lisan al Gaib@scaling01 · 博主 · 1 天前高频 AI 模型测评与爆料博主

Prime Intellect 和 Elie Bakouch 的精彩工作!

还有一个有趣的引用:
"我们再次对缺乏新颖性感到惊讶。这些模型显然在深层次上理解它们操纵的对象,然而很少有真正的新想法涌现"

引用 Prime Intellect @PrimeIntellectWe ran the largest open experiment on how frontier models do AI research. 100+ autonomous runs across 10+ models, sandboxed on 8xH200s for up to 8 days, iterating on the nanoGPT optimizer track. Best runs closed 82% of the gap to a record built by dozens of humans over months.查看被引原帖 ↗
查看英文原文
Awesome work by Prime Intellect and Elie Bakouch!

and an interesting quote:
"we were again surprised by the lack of novelty. The models clearly understand the objects they manipulate at a deep level, and yet very few genuinely new ideas emerge"
Chubby♨️@kimmonismus · 博主 · 1 天前Chubby,高频 AI 新闻聚合博主

Dario Amodei 和 Gavin Baker 花了整个周末时间纠缠在一个问题上:AI 监管到底是集中权力还是防止权力集中。这到底怎么回事呢?

Amodei 拒绝接受这个二元对立:他认为 AI 本身就会集中权力,因为训练前沿模型需要资本、芯片和能源,只有少数组织才能凑齐这些。开源权重救不了这个局面,因为计算资源的访问权最终还是掌握在前沿实验室和大型硬件供应商手中。他的药方是制度性的:制定更严厉地约束最大开发者的规则,对足够强大的模型进行测试(不管开源还是闭源),以及建立类似 FINRA 风格的监管机构(Demis Hassabis 提议过这个)。

Baker 的问题是:最后谁活着。合规是固定成本,固定成本对已经最大的公司是大礼物,尤其当这帮最大的家伙还坐在起草规则的房间里时。他第二个论点我觉得更难反驳。他说持续谈灾难在政治上产生了后果,不管 Anthropic 是不是这么想,它正在给反对数据中心和反对基建的本地运动加码,结果这件事的美好版本就更不可能出现了,反而风险更大。Amodei 否认自己的记录被这样解读,并指出了《Machines of Loving Grace》这篇文章。说得对,虽然那不是人们经常用来攻击他的那篇。

我想看到一个有很多模型、很多建造者、真正竞争的未来,让有用的智能变得无处不在,很多任务的成本低到都不值一提。我想看到一个我们携手反抗 AI 恐惧的未来。我想看到前沿 AI 免费开放给尽可能多的人的未来。

引用 Dario Amodei @DarioAmodei2/2 Second, on the messaging around AI.  I do not agree that my messaging has been disproportionately negative.  In fact it has been about equally balanced between risks and benefits: I’ve written one major essay about each, and even in interviews where I discuss the risks, I make sure to frequently mention the incredible benefits as well as proposing possible solutions to the risks (short clips from my interviews that end up on social media tend to be disproportionately negative, as that gets clicks).  In fact, I wrote Machines of Loving Grace because I didn’t feel the AI industry was painting an inspiring enough picture of how the technology could radically transform the world for the better.  The bulk of the essay is devoted to refuting skepticism of AI’s potential in health and biology, and showing why I think it will actually be possible to cure most human disease in ~5-10 years, as crazy as it may sound to ordinary people and frankly to biologists as well (I used to be one!).  And, if you read my most recent essay (Policy on the AI Exponential), I discuss concrete proposals for how to streamline the FDA process to make sure the deluge of AI-accelerated drugs isn’t slowed down by the regulatory process.  I feel the urgency here: I lost my father to Hepatitis C only a few years before the development of direct-acting antivirals (sofosbuvir), which cure 95% of patients and probably would have cured him. I do agree that the public has a negative view of AI (and that this is a big problem), but I don’t think it is primarily caused by me or any other AI leader warning about AI’s risks.  I think it is fundamentally a crisis of trust.  I think that ordinary people don’t trust companies, governments, or the tech industry and always suspect that we are cooking up some new way to screw them over.  The causes of this go back decades and AI is just the latest iteration of it.  I don’t think that a glitzy marketing campaign with a positive spin (which some have advocated t查看被引原帖 ↗
查看英文原文
Dario Amodei and Gavin Baker spent the weekend arguing about whether regulating AI concentrates power or protects against it. Whats it about?

Amodei rejects the choice itself: He argues that AI concentrates power on its own, because training frontier models depends on capital, chips and energy that few organizations can assemble, and that open weights do not fix this, since access to compute stays with the frontier labs and the large hardware providers. His answer is institutional: rules that fall harder on the biggest developers than on smaller ones, testing for models capable enough to matter whether they are open or closed, and something like the FINRA style oversight body Demis Hassabis has proposed.

Baker's problem is with who survives that: Compliance is a fixed cost, and fixed costs are a gift to whoever is already biggest, especially when the biggest are also in the room while the rules get drafted. His second point is the one I find harder to wave away. He thinks the constant catastrophe talk is doing political work whether Anthropic wants it to or not, feeding the local campaigns against data centers and against buildout generally, and that this makes the good version of all this less likely rather than safer. Amodei rejects that reading of his own record and points at Machines of Loving Grace. Fair enough, though that is not the essay people quote back at him.

I want a future with many models, many builders and real competition, where useful intelligence becomes abundant and, for a growing number of tasks, too cheap to meter. I want a future where we work together to combat the fear of AI. And I want a future where frontier AI becomes freely available to as many people as possible.
swyx@swyx · 博主 · 1 天前知名 AI 播客 Latent Space 主理人

如果你不生活在科技圈里,大概会惊讶于顶级人才之间实际上很少见面或彼此认识。作为局外人,你可能会以为大家都在秘密的智囊群里。那些群确实存在,但只是极短暂的例外,而非常态。你在头条上看到的,他们同样也看到。我想,这也说明大部分精力还是花在做事上,而不是经营舆论或八卦上。于我个人而言,至少在我职业生涯的这个阶段,这确实让人感到安心。

引用 Internal Tech Emails @TechEmailsAnthropic cofounder emails Noam Shazeer November 21, 2021查看被引原帖 ↗
查看英文原文
in case you’re not living in the tech bubble, as a general rule i’ve been surprised by how infrequently top tier folks actually meet/know each other. as an outsider i might have assumed that everyone is in secret illuminati group chats. those exist, but are very much short lived exceptions rather than the rule. what you see of the major headlines is pretty much what they also see. i guess one way to interpret this is also simply that most effort is still on doing the work rather than working the narrative or the gossip. and that, to me at least, is genuinely quite reassuring, this far in to my career.
el.cine@EHuanglu · 博主 · 1 天前

我们已经突破了界限……这看起来一点都不像AI了

查看英文原文
we’ve crossed the line.. this doesn’t look AI at all now
◔ 6.5 万 次浏览♥ 607⇄ 31▶ 含视频观点看原帖 ↗
hardmaru@hardmaru · 创始人 · 1 天前David Ha,日本 AI 公司 Sakana AI 联合创始人

冷门观点:Gemini 3.1 Pro 其实是个不错的模型。现在大多数基础模型对日常工作的 99% 场景都完全够用。不是所有人都需要最先进的编码模型才能干活。

查看英文原文
unpopular opinion: gemini 3.1 pro is a pretty great model

most current baseline models are actually perfectly fine for 99% of everyday work. not everyone needs a state-of-the-art coding model to get things done.
Ethan Mollick@emollick · 创始人 · 1 天前沃顿商学院教授,AI 应用研究权威
连环推 ×2

难以想象,即使忽略 AI 的其他一切,光是 Google AI Overview 一个产品,长期来看也会彻底改变网络的面貌和我们获取、消费、利用信息的方式。显然这已经在发生了。

查看英文原文
Its hard to imagine, even if you ignore literally everything else associated with AI, that Google AI Overview alone would not profoundly change the nature of the web, and the information we consume and act on as a result, over time. Its obviously already starting to do that.
(And that also ignores its impact on hundreds of billions of dollars of e-commerce and advertising)
🚨 AI News | TestingCatalog@testingcatalog · 博主 · 1 天前专挖 AI 产品未发布新功能的爆料号

ANTHROPIC 🔥:Claude Tag 即将在 Claude Desktop 上以协作 Projects 功能的形式推出,面向 Claude Code。

基本上就是内置的 Slack 👀

> 之前发现 Managed Projects 功能有了小升级,用户现在能在一个项目会话里创建多条线程。

> 项目创建界面新增了将代码库添加为上下文的选项,这些代码库会跨所有会话使用。

即将推出的 Claude Code Projects 功能将让团队和个人能在 Claude Desktop 里协作,就像用 Slack 上的 Claude Tag 一样。

Claude Code Projects 会有持久化的上下文和记忆,Claude 会不断改进和优化。

查看英文原文
ANTHROPIC 🔥: Claude Tag is coming to Claude Desktop in the form of collaborative Projects for Claude Code.

That's essentially a built-in Slack 👀

> Earlier spotted Managed Projects feature got a slight upgrade, mentioning that users will be able to spawn different threads within a single project session.

> Project creation UI got a new option to add repositories as context, which will be used across all sessions.

The upcoming Projects for Claude Code feature is expected to enable teams and individuals to work collaboratively inside Claude Desktop, as they would with Claude Tag on Slack.

Claude Code Projects will have persistent context and memory, which Claude will periodically refine and improve over time.
小互@xiaohu · 中文博主 · 1 天前小互,中文圈高频 AI 资讯站 Xiaohu.AI 主理人

工牌特工

嘿嘿

这个质量很高,AI制作,完全是电影的感觉,而且还挺搞笑🤪

全片由小云雀Seedance2.5制作完成

作者:许立展 (抖音)

◔ 5.4 万 次浏览♥ 313⇄ 33▶ 含视频演示看原帖 ↗
Matt Shumer@mattshumer_ · 博主 · 1 天前HyperWrite CEO,AI 实战技巧分享

正在教 Grok Bot 用 Unity 开发游戏。敬请期待!

引用 Matt Shumer @mattshumer_Grok 4.6 连续工作 48 小时构建了这款射击游戏。证明 Grok 足够强大来运行 Gauntlet Loops。游戏开发时代开始了!查看被引原帖 ↗
查看英文原文
Teaching Grok Bot to build games in Unity.

Updates soon!
◔ 5.1 万 次浏览♥ 362⇄ 9▶ 含视频新品看原帖 ↗
Chubby♨️@kimmonismus · 博主 · 1 天前Chubby,高频 AI 新闻聚合博主

Sol 终于可以导航 Luna subagents 了。这是个急需的更新,非常欢迎。

引用 eric provencher @pvncherThis went under the radar this week, but we just shipped the ability for models with multi agents v2 to delegate to any supported model, including Luna! Took a bit of time to make sure this worked reliably查看被引原帖 ↗
查看英文原文
Sol can finally navigate Luna subagents. Much needed and very welcome update.
歸藏(guizang.ai)@op7418 · 中文博主 · 1 天前歸藏,中文圈 AI 工具与提示词博主

正在给 DeepSeek Harness 做一个界面和交互美化的插件。

还有一个服务商的插件,这样整个产品体验会好一点,后续也会打包成一个客户端。

这主要是把 Codepilot 里面沉淀的一些资产抽离出来,然后做到这个 DeepSeek Harness 的插件里,这样整个使用体验和视觉效果都会变好不少。

看了一些现有的美化插件,基本上都是改改背景图、加一两个组件效果什么的,感觉对于长期使用和用户体验没什么多大的帮助。

Ethan Mollick@emollick · 创始人 · 1 天前沃顿商学院教授,AI 应用研究权威

这其实就是AI对经济影响的核心问题。普遍的假设是技术采纳的那些传统摩擦还会继续(AI到现在为止就是这样),但如果这些系统不断改进,这些摩擦可能就……不存在了。怎么发展还说不好。

引用 Arpit Gupta @arpitrage我讨论AI固有的摩擦如何阻碍经济范围内的采用。但内心里我不认为它们有那么糟。你只需要要求机器做事,它就做了。这些摩擦怎么可能还会存在呢?查看被引原帖 ↗
查看英文原文
This is, in fact, the Big Question of the impact of AI on the economy. There is a general assumption that the usual frictions of technology adoption continue (as they have up until now with AI), but if systems keep improving, they may just… not. Which way it goes is unclear.
🚨 AI News | TestingCatalog@testingcatalog · 博主 · 1 天前专挖 AI 产品未发布新功能的爆料号

OpenAI 正在更新 ChatGPT 的隐私政策,涉及广告相关内容,包括广告个性化。

> 广告现在扩展到更多地区,目前只显示给免费和 Go 账户用户。

> 用户如果想要广告个性化,需要明确选择同意。

> 在某些地区,免费用户可以选择不看广告,但这样的话他们的消息数量会受到限制。

查看英文原文
OpenAI is updating its privacy policy in regard to ads on ChatGPT, mentioning ad personalisation.

> Ads are expanding to more regions and still only shown to free and Go accounts.

> Users will have to explicitly opt it into ad personalisation if they want to.

> In some regions, Free users can still choose not to see ads but they will be limited in a number of messages.
Ethan Mollick@emollick · 创始人 · 1 天前沃顿商学院教授,AI 应用研究权威
连环推 ×2

在 agentic 循环里(已经非常过时了),o3-mini 能出高质量的考试题:“在我们迄今规模最大的 AI 生成问题实地研究之一中,我们发现了这种方法能实现与高风险标准化考试中题目相当的测量学特性。”

引用 Misha Teplitskiy | Science of Science @MishaTeplitskiyAI生成的考试没有问题查看被引原帖 ↗
查看英文原文
o3-mini (very obsolete) in an agentic loop made good exam questions: "In one of the largest field studies of AI-generated questions to date, we found that this approach achieves psychometric properties on par with those of questions appearing on high-stakes standardized tests."
Experiments like these keep showing that AI can do a lot of fairly hard things well if you decide to put the time in to figure out how to approach the problem.
Orange AI@oran_ge · 中文博主 · 1 天前Orange AI,中文圈 AI 产品观察博主

牛来这玩意儿已经赢了啊
又赚钱又赚注意力
成为了新的 meme
你不能不停下来看它一秒
营销之神

Chubby♨️@kimmonismus · 博主 · 1 天前Chubby,高频 AI 新闻聚合博主

不管 Baker 和 Amodei 之间怎么吵,有一点好像所有人都达成共识了:就是在短短几年内,(几乎)所有疾病都能被治愈。撇开那些争议不说,光是这一个积极的前景就足以让人兴高采烈,尤其对那些本来就持怀疑态度的人来说。

癌症在几年内会完全可治愈。

Amodei:「(...)我想说明为什么我认为在大约 5-10 年内实际上有可能治愈大多数人类疾病,尽管听起来疯狂,说实话连生物学家都很难相信(我以前就是生物学家!)。

(...)

我最近写了一篇文章《Policy on the AI Exponential》,里面讨论了怎样简化 FDA 流程的具体方案,确保这波 AI 加速的新药不会被监管卡住。

(...)

这些问题的根源可以追溯到几十年前,AI 只是最新的迭代。我不认为靠一个精致的营销活动加正面旋转(有些人建议 Anthropic 这么干)就能赢回信任——在这阶段,说 AI 能治愈癌症已经成了烂大街的说法,不再能鼓舞人心,大多数人反而觉得这很虚伪。

真正有用的方法是实际去治愈癌症。我觉得对包括 Anthropic 在内的 AI 公司最中肯的批评就是:我们还没有兑现造福人类的宏大承诺。

不过我们在尽力改变这一点:Anthropic 正在大幅加强在生物学和医学领域的投入,我们期望在未来几年内取得了不起的成果,未来几个月会有初步的好消息。

一旦我们真正做出什么实实在在的成果,全世界都会听说,尽可能大声地广播,你们有我的保证。但在那之前,我不想说空话,同时我也觉得有必要坦诚地谈论 AI 真实存在的风险以及怎么应对。」

引用 Chubby♨️ @kimmonismusDario Amodei and Gavin Baker spent the weekend arguing about whether regulating AI concentrates power or protects against it. Whats it about? Amodei rejects the choice itself: He argues that AI concentrates power on its own, because training frontier models depends on capital, chips and energy that few organizations can assemble, and that open weights do not fix this, since access to compute stays with the frontier labs and the large hardware providers. His answer is institutional: rules that fall harder on the biggest developers than on smaller ones, testing for models capable enough to matter whether they are open or closed, and something like the FINRA style oversight body Demis Hassabis has proposed. Baker's problem is with who survives that: Compliance is a fixed cost, and fixed costs are a gift to whoever is already biggest, especially when the biggest are also in the room while the rules get drafted. His second point is the one I find harder to wave away. He thinks the constant catastrophe talk is doing political work whether Anthropic wants it to or not, feeding the local campaigns against data centers and against buildout generally, and that this makes the good version of all this less likely rather than safer. Amodei rejects that reading of his own record and points at Machines of Loving Grace. Fair enough, though that is not the essay people quote back at him. I want a future with many models, many builders and real competition, where useful intelligence becomes abundant and, for a growing number of tasks, too cheap to meter. I want a future where we work together to combat the fear of AI. And I want a future where frontier AI becomes freely available to as many people as possible.查看被引原帖 ↗
查看英文原文
But regardless of the public dispute between Baker and Amodei, there is one thing in particular on which everyone seems to agree: the fact that in just a few years, presumably (almost) all diseases can be treated (cured). And setting aside all the disagreements, this positive statement alone is reason enough to be filled with joy and excitement, especially among people who are otherwise rather reserved and skeptical.

Cancer will be completely curable in a few years.

Amodei "(...) and showing why I think it will actually be possible to cure most human disease in ~5-10 years, as crazy as it may sound to ordinary people and frankly to biologists as well (I used to be one!).

(...)

And, if you read my most recent essay (Policy on the AI Exponential), I discuss concrete proposals for how to streamline the FDA process to make sure the deluge of AI-accelerated drugs isn’t slowed down by the regulatory process.

(...)

The causes of this go back decades and AI is just the latest iteration of it. I don’t think that a glitzy marketing campaign with a positive spin (which some have advocated that Anthropic do) is the way to win back that trust — at this point, saying that AI will cure cancer is more a cliche than it is inspiring, and most people think it is deceptive.

The thing that will work is actually curing cancer. I think by far the most accurate criticism of AI companies including Anthropic is that we haven’t yet delivered on our big promises to benefit the world.

We are however doing our best to fix this: Anthropic is ramping up its efforts very quickly in biology and medicine, and we hope to have incredible results in the coming years and some early glimmers in the coming months.

When we’ve actually accomplished something real, the whole world will hear about it, as loudly as possible, you have my word on that. But until then I don’t want to make empty promises, and in the meantime I feel compelled to speak honestly about the very real risks of AI and how to address them."
宝玉@dotey · 中文博主 · 1 天前宝玉,中文圈 AI 翻译与科普大 V

Claude Code (Claude Desktop)这个功能很好用,就是你5小时限制到了,可以自动继续,不只是主 Agent 到时间继续,子 Agent 到时间也会继续。不需要手工输入 return,也不需要重新开 subagent,可以重用之前的 subagent

引用 ClaudeDevs @ClaudeDevs在Claude Code桌面版达到使用限制?现在有自动继续复选框。打开它,限制重置后会自动从中断处继续。查看被引原帖 ↗
Chubby♨️@kimmonismus · 博主 · 1 天前Chubby,高频 AI 新闻聚合博主

“开源权重在这方面确实有帮助,但远非充分的解决方案,因为它们只是把集中度转移到拥有最多算力和芯片的人身上(大致是前沿实验室,可能外加硬件供应商)。”

我强烈不同意这点,Qwen-3.8 27b 刚刚就彻底证实了这一点。开源是为所有人准备的,不只是给那些拥有海量计算资源的人。

引用 Dario Amodei @DarioAmodei1/2 Thanks Gavin for an especially thoughtful exchange. I don't usually spend much time on social media but I wanted to engage here because it really brings out the heart of an important conversation. First, on regulation, I think that “either concentrate it in the hands of a chosen few companies and politicians via regulation or distribute it widely” is a false choice.  I know that there’s a sort of Silicon Valley shorthand where regulation = regulatory capture = concentration of power, but I’ve always found this to be an overly simplified picture of the world.  Many people outside this bubble think of regulation as something that constrains corporate power and benefits ordinary people.  I don’t necessarily agree with that perspective either, rather I think it’s complicated and really depends on what the “regulation” consists of.  But in particular I think that those in the “regulation = regulatory capture = concentration of power” frame often underrate the decentralizing power of objective and fair institutional processes.  A crude analogy is that the formal court system can sometimes feel stuffy and elitist, but it does a much better job of defending the rights of vulnerable individuals than the alternative, mob justice.  At their best, institutions can vest power in ideas rather than people, and thereby decentralize that power. This is why Anthropic has always made its policy proposals very carefully.  We try very hard to make proposals that disadvantage (slow down) frontier AI companies while *advantaging* smaller competitors.  California’s SB53 (which we supported), and even the much-maligned SB 1047 (which we were ambivalent on), completely exempt any company below a certain amount of revenue or model training costs from being covered at all (it was $500M for SB 53, lower for 1047 but we objected to that).  More recently the testing process we’ve advocated for at CAISI and the White House involves more rigorous tests for frontier models than off-frontier mo查看被引原帖 ↗
查看英文原文
„Open-weights do help some with this but are nowhere near a sufficient solution because they simply shift the concentration somewhat to those with the most compute and chips (which are roughly the frontier labs plus maybe hardware providers).“

I vehemently disagree, and Qwen-3.8 27b has just conclusively proven this. Open source is for everyone, not just for those with massive compute resources.
宝玉@dotey · 中文博主 · 1 天前宝玉,中文圈 AI 翻译与科普大 V

Anthropic 的篇官方技术说明,解释 Claude 文本水印的具体工作方式。
x.com/trq212/status/20887210…


Claude 用的是 Google DeepMind 2024 年发表在《Nature》上的 SynthID-Text 方案,往上追溯可以到 Scott Aaronson 2022 年的提案。这一脉方案的共同特征是在模型选词时,用密钥改变随机数的来源,从而在输出中留下统计痕迹。

Anthropic 举了个例子帮助理解。假设模型写到“今天天气很冷,而且……”,下一个词可能是“阴沉”,也可能是“灰蒙蒙的”,对读者来说意思差不多。正常情况下,模型用一个随机数在这些候选词里选一个。加了水印之后,选词过程还是随机的,但随机数的来源变了,变成由密钥和前面的词共同决定。文本读起来完全一样,但事后拿着密钥去核对选词序列,就能算出这段话是 Claude 写的概率。

Anthropic 在文中特别强调了一点:水印不会让模型选它本来不会选的词。模型不会因为水印而突然用一个生僻词(文中举的例子是 nubilous,一个几乎没人用的“阴天”的同义词)。水印只在模型本就犹豫的那些选择上起作用。

然后是大家最关心的:会不会降低输出质量?

Anthropic 说他们内部测试没有观察到水印对内容质量、创意水平或可读性的影响。他们援引了 Google DeepMind 的数据作为佐证:Google 在 Gemini 的实际流量中给一部分用户提供了带水印的模型,对比用户的点赞点踩反馈,两组没有统计显著差异。同时在受控实验中,人类评估者把带水印和不带水印的回答放在一起对比,也看不出区别。

打个比方:想象你在玩大富翁,本来每回合掷骰子决定走几步。现在改成查圆周率的小数位来决定,从某一位开始,依次读下去。对玩家来说,每一步还是随机的,游戏体验没有变化。但如果有人事后看走步的完整序列,再对照圆周率,就能判断这盘游戏用的是圆周率而不是骰子。

关于代码,Anthropic 的解释是:水印只在选哪个词都行的地方起作用。代码大量地方必须精确,“2+2=”后面只能跟 4,水印无从施加。所以代码的水印密度天然就比自然语言低。不过在代码注释、变量命名这些有选择余地的地方,水印仍然可以嵌入。

还有一个之前没怎么讨论的重要信息:水印不携带任何用户身份信息。不能追溯到具体的用户、组织或对话。水印只回答一个问题,“这段文字是否可能经过 Claude 处理”,仅此而已。

文章中也提到了水印的局限。短文本检测不了,信号太弱。纯事实性内容(比如“牛顿最著名的著作叫《自然哲学的数学原理》……”)水印也没什么发挥空间,因为下一个词几乎没有选择余地。你让 Claude 校对一篇文章只改语法标点,大部分词还是你的,水印可能根本不够形成可检测的信号。

轻度编辑大概率不会完全去掉水印。但如果把每个词都换一遍,水印就没了,当然到这个程度,这段文字算不算 AI 生成的本身也值得商榷了。

至于为什么全球统一加水印而不只在欧盟范围内执行,Anthropic 给出的理由是他们目前没有一个可靠的方式按地区区分是否施加水印。所以干脆全球统一上线,后续再评估是否调整。

检测端还没有完全就位。Anthropic 表示很快会推出水印检测 API,但具体时间和使用方式还在敲定中。这也呼应了 Alex Cui 之前提到的一个关键问题:检测器是公开还是内部使用,会直接影响水印的实际安全性。

顺便说一下,签署这份欧盟透明度行为准则的不止 Anthropic 一家,总共有大约 190 个签署方。水印不是 Claude 独有的事情,其他大模型厂商也会陆续跟上各自的方案。

引用 Anthropic @AnthropicAIAnthropic发布了关于水印的FAQ。摘要:为遵守EU AI Act实施水印,其他主要模型开发者也将实施;水印不影响Claude输出质量;用户无法区别水印与非水印文本;无需添加文本或隐藏字符;无需额外token,无额外成本;水印无法追踪到个人、组织或聊天。查看被引原帖 ↗
◔ 2.9 万 次浏览♥ 77⇄ 10▶ 含视频研究看原帖 ↗
🚨 AI News | TestingCatalog@testingcatalog · 博主 · 1 天前专挖 AI 产品未发布新功能的爆料号

Google正在开发一项功能,让Gemini Notebook可以直接从聊天UI查询Google Drive中的文件。

> 你现在可以在Gemini Notebook中直接查询Google Drive文件。提一个问题就能开始。

Gemini Notebook成为超级应用?👀

查看英文原文
Google is working on a way to let Gemini Notebook query files from Google Drive directly from the chat UI.

> You can now query Google Drive files directly in Gemini Notebook. Ask a question to get started.

Gemini Notebook as a super app? 👀
Guillermo Rauch@rauchg · 创始人 · 1 天前Guillermo Rauch,Vercel 创始人兼 CEO

递归自我改进

查看英文原文
Recursive self improvement
Bindu Reddy@bindureddy · 创始人 · 1 天前Abacus.AI CEO,AI 行业观点博主

Anthropic想要集中权力,而OpenAI正在选择让AI民主化

他们明确选择不发布Mode 2——Mythos的下一个版本

相反他们在内部用它来解决生物学、物理学和医学中的难题

与此形成鲜明对比的是,OpenAI已确认发布Astra!我相信长期来看,OpenAI会让AI普遍可及🥳

查看英文原文
Anthropic wants to concentrate power, while OpenAI is choosing to democratize AI

They have explicitly chosen not to release Mode 2 - the next version of Mythos

Instead they are using it internally to work on hard problems in biology, physics and medicine

In sharp contrast OpenAI has confirmed the release of Astra! Over time, I expect OpenAI to make AI universally accessible 🥳
Bindu Reddy@bindureddy · 创始人 · 1 天前Abacus.AI CEO,AI 行业观点博主

世界正在分裂成两个阶层的人

Token-rich —— 专注于用 AI 让自己提升 100 倍、最优化 AI 使用来最大化 ROI 的人

Token-poor —— 专注于最小化劳动成本、获得小幅生产力收益、使用廉价模型的人

随着时间推移,差异会变得非常明显

查看英文原文
The world is dividing into two classes of people

Token-rich - people who are focused on using AI to 100x themselves, optimal use of AI to maximize ROI

Token-poor - focused on minimizing the cost of labor, getting small productivity gains, using cheap models

Overtime the difference will become stark
向阳乔木@vista8 · 中文博主 · 1 天前向阳乔木,中文圈 AI 工具与趋势博主

像丁师傅说的一样,B站藏龙卧虎,有一大批优秀的 AI 产品创作者。

DeepSeek Harness 插件贡献者中很多都是 B 站的UP主,一水儿的二次元头像。

比如 DSH 发布当天热榜第一 colleague-skill ,第三的OpenBiliClaw全都是B站 UP主。

当年二次元属于小众爱好,现在都成了各大公司的核心技术主力。

试着在 DSH 装了这两个插件,用colleague-skill 生成了一个纳瓦尔Skill:

github.com/joeseesun/celebri…


OpenBiliClaw也安装了,但要配置向量模型才能运行,抽空再试试。

纳瓦尔说的 Build & Sell,放在开源生态里,除了 X,B站就是最重要传播宣传平台。

DSH 刚开源没几天,插件生态已经这么繁荣,未来可期。

◔ 4 万 次浏览(3 条合计)♥ 131⇄ 13观点看原帖 ↗
Gary Marcus@GaryMarcus · 博主 · 1 天前

💯:“从长远来看,要让各国感到安全有保障,它们需要一定程度、粒度足够细的国际透明度——虽然这听起来很理想化,但最终我们将不得不真正成为一个全球共同体。”

引用 Robert Wright @robertwrighterThis interview of me by @dwallacewells in the New York Times ( nytimes.com/2026/08/13/opini… ) manages to touch on a lot of themes from my new book on AI. If you want to read the book's introduction, and excerpts from other chapters, they're here: thegodtest.net查看被引原帖 ↗
查看英文原文
💯: “in the long run, for countries to feel safe and secure, they’re going to need a degree of international transparency at a sufficiently fine-grained level that — as idealistic as this may sound — we are going to have to finally become a true global community.”
Tanishq Mathew Abraham, Ph.D.@iScienceLuvr · 博主 · 1 天前

Dario 谈论什么会改变公众对 AI 的看法:

>真正有效的是*真正治愈癌症*。

就像我一直在说的,把这个世纪最具变革性的技术应用于拯救生命是最重要的使命。

>不过我们在尽力改善:Anthropic 正在迅速加大对生物医学的投入

Anthropic 真的想推动 AI 在生物/健康领域的发展...

引用 Dario Amodei @DarioAmodei作者辩称自己关于AI的表述并非过度负面,而是平衡讨论风险与好处。撰写《Machines of Loving Grace》展示AI在健康领域的潜力,认为可在5-10年内治愈大多数人类疾病。查看被引原帖 ↗
查看英文原文
Dario talking about what will change the public's perception on AI:

>The thing that will work is *actually curing cancer*.

Like I've been saying, applying the most transformative technology of the century to save lives is THE MOST IMPORTANT MISSION.

>We are however doing our best to fix this: Anthropic is ramping up its efforts very quickly in biology and medicine

Anthropic really wants to push forward AI for bio/health...
Lisan al Gaib@scaling01 · 博主 · 1 天前高频 AI 模型测评与爆料博主

Anthropic 就像园丁

引用 Dario Amodei @DarioAmodei1/2 Thanks Gavin for an especially thoughtful exchange. I don't usually spend much time on social media but I wanted to engage here because it really brings out the heart of an important conversation. First, on regulation, I think that “either concentrate it in the hands of a chosen few companies and politicians via regulation or distribute it widely” is a false choice.  I know that there’s a sort of Silicon Valley shorthand where regulation = regulatory capture = concentration of power, but I’ve always found this to be an overly simplified picture of the world.  Many people outside this bubble think of regulation as something that constrains corporate power and benefits ordinary people.  I don’t necessarily agree with that perspective either, rather I think it’s complicated and really depends on what the “regulation” consists of.  But in particular I think that those in the “regulation = regulatory capture = concentration of power” frame often underrate the decentralizing power of objective and fair institutional processes.  A crude analogy is that the formal court system can sometimes feel stuffy and elitist, but it does a much better job of defending the rights of vulnerable individuals than the alternative, mob justice.  At their best, institutions can vest power in ideas rather than people, and thereby decentralize that power. This is why Anthropic has always made its policy proposals very carefully.  We try very hard to make proposals that disadvantage (slow down) frontier AI companies while *advantaging* smaller competitors.  California’s SB53 (which we supported), and even the much-maligned SB 1047 (which we were ambivalent on), completely exempt any company below a certain amount of revenue or model training costs from being covered at all (it was $500M for SB 53, lower for 1047 but we objected to that).  More recently the testing process we’ve advocated for at CAISI and the White House involves more rigorous tests for frontier models than off-frontier mo查看被引原帖 ↗
查看英文原文
Anthropic is the Gardener
Lisan al Gaib@scaling01 · 博主 · 1 天前高频 AI 模型测评与爆料博主

Opus 5 在 KernelBench 上的统治性表现。你可以想象 Model 2 内部有多强

查看英文原文
Opus 5 domination on KernelBench

you can just imagine what Model 2 does internally
Ethan Mollick@emollick · 创始人 · 1 天前沃顿商学院教授,AI 应用研究权威
连环推 ×2

Infocom 游戏现在开源了,我让 Codex 更新了 1987 年的文字冒险游戏 "Nord and Bert Couldn't Make Head or Tail of It" 以便现在可以玩。它保留了 O'Neill 的谜题,尽管新 UX 让赢得更容易一些(但 AI 在翻译某些谜题方面做得更好)

ord-and-bert.netlify.app/

查看英文原文
As Infocom games are now open source, I had Codex update 1987’s word game “Nord and Bert Couldn't Make Head or Tail of It” for playing now. It keeps O'Neill’s puzzles, though its easier to win with a new UX (but the AI did better translating some puzzles)


nord-and-bert.netlify.app/
All of this was done from ChatGPT on mu phone which connected remotely to Codex on my computer while I was traveling. Sol made all the choices, generated all the images, etc. Really different way of working.

Code here:
github.com/emollick/nord-and…
Chubby♨️@kimmonismus · 博主 · 1 天前Chubby,高频 AI 新闻聚合博主

特朗普政府正在施压苹果不要购买中国内存芯片,因为 AI 数据中心在耗尽全球芯片供应。

据华尔街日报报道

苹果据悉正在测试 CXMT 和 YMTC 的芯片,用于在中国销售的设备。商务部长 Howard Lutnick 声称他向苹果直言华盛顿反对这一举动。

苹果可以合法购买这两家公司的标准现货产品。为定制芯片共享产品信息需要获得美国许可证。

看来内存短缺会持续,价格也会继续高位。

查看英文原文
The Trump administration is pressuring Apple not to buy Chinese memory chips as AI data centers drain global supply.

Via WSJ

Apple is reportedly testing chips from CXMT and YMTC for devices sold in China. Commerce Secretary Howard Lutnick says he told Apple “plainly” that Washington opposes the move.

Apple can legally buy standard, off-the-shelf parts from both companies. Sharing product information for customized chips would require a U.S. license.

Looks like ram shortage will continue and prices stay high.
Linus ✦ Ekenstam@LinusEkenstam · 博主 · 1 天前

这个机器人是盲的,但照样能搞定这事。
没错,我们真是活在大结局时刻了。

查看英文原文
This robot is blind, but can still do this.
yeah, we’re truly living in the end game
Ethan Mollick@emollick · 创始人 · 1 天前沃顿商学院教授,AI 应用研究权威

我特别想看到更多历史学家、生物学家、物理学家这样的专业人士去用AI视频。我们根本没有足够准确的视频来展现恐龙时代的生活、木卫一的表面、古罗马是什么样的——这些领域没什么商业价值。

AI视频可以让专家直接展示给我们。

引用 Ethan Mollick @emollick我一直在读Bethany Hughes关于七大奇迹和罗德岛巨像神话的书(它其实没有跨越港口)。我用Google Deep Research为veo 3生成了一个历史罗德岛巨像的描述,效果出色,细节精准且引人入胜。查看被引原帖 ↗
查看英文原文
I would love to see more historians, biologists, physicists, etc using AI video. We simply do not have enough accurate video of what life was like at the time of the dinosaurs, or the surface of Io, or Ancient Rome, etc. No money in it

AI video allows experts to show us directly
Gary Marcus@GaryMarcus · 博主 · 1 天前

经典!OpenAI 员工声称 AGI 已经实现!🤦‍♂️

引用 Mike Sulka @SulkaMike2024 年 12 月 7 日 o1查看被引原帖 ↗
查看英文原文
classic! openAI employee declares AGI reached! 🤦‍♂️
Gary Marcus@GaryMarcus · 博主 · 1 天前

有人还记得 GPT-5 被说成是 AGI 吗?🤣

引用 Maria Sukhareva | AI Realist @maria_airealistThey have an “AGI” every other month. Now it’s about one year anniversary as the last AGI a.k.a. GPT-5 was released查看被引原帖 ↗
查看英文原文
anyone remember how GPT-5 was supposed to be AGI? 🤣
elvis@omarsar0 · 博主 · 1 天前

IBM最近有个有趣的研究。

如果你从benchmark的改进中挑选模型,其中一些改进其实是来自措辞而不是模型本身。

BenchDrift在语言学、指涉、语用和结构层面生成意义保留的benchmark变化,固定答案,然后测量正确性反转的频率。

措辞敏感性不会随着模型改进而消退。它会改变方向。弱模型通过重新措辞获益更多,而强模型损失反而更多——所以benchmark上排名最高的模型恰好是那些分数最依赖于它们碰到的措辞的模型。

脆弱性也来自重新措辞。在8个模型的GSM8K、MMLU和MATH-Hard上,它们在哪些重新措辞会丢失最多正确答案方面基本一致,尽管总体漂移程度差异很大。

重新措辞会打破模型曾经很有把握的答案,不管问题变长还是变短。

论文:
arxiv.org/abs/2608.11694

在我们的academy上追踪更多热门AI论文:
academy.dair.ai/

查看英文原文
Interesting new research from IBM.

If you pick models from benchmark deltas, some of that delta belongs to the phrasing rather than the model.

BenchDrift generates meaning-preserving variations of benchmark problems along linguistic, referential, pragmatic, and structural axes, holding the answer fixed, then measures how often correctness flips.

Phrasing sensitivity does not fade as models improve. It changes sign. Weak models gain more from rephrasing than they lose, while strong models lose far more than they gain, so the top models on a benchmark are the ones whose scores depend most on the wording they happened to receive.

Fragility also belongs to the rephrasing. Across eight models on GSM8K, MMLU, and MATH-Hard, they largely agree on which rephrasings cost the most correct answers even while differing in how much they drift overall.

Rephrasing breaks answers models were confident about, whether the problem gets shorter or longer.

Paper:
arxiv.org/abs/2608.11694


Track more trending AI papers in our academy:
academy.dair.ai/
Lisan al Gaib@scaling01 · 博主 · 1 天前高频 AI 模型测评与爆料博主

还记得有些人编那些阴谋论,说比尔·盖茨搞新冠是为了通过疫苗注射芯片吗?如果你觉得这听起来疯了,那就等着吧,等 Anthropic 大概 5 年后开始推疫苗、用他们的机器神灵控制世界经济的时候,那时候又会出现什么理论呢。

查看英文原文
Remember how some people had these conspiracy theories how Bill Gates was behind COVID to spread his microchips via vaccines?

If you think that sounds crazy, just wait for the theories that will emerge when Anthropic tries to spread their vaccines in ~5 years, while controlling the world economy with their machine gods
Lisan al Gaib@scaling01 · 博主 · 1 天前高频 AI 模型测评与爆料博主
连环推 ×3

Opus 5 的救赎之旅

引用 Lisan al Gaib @scaling01Awesome work by Prime Intellect and Elie Bakouch! and an interesting quote: "we were again surprised by the lack of novelty. The models clearly understand the objects they manipulate at a deep level, and yet very few genuinely new ideas emerge"查看被引原帖 ↗
查看英文原文
Opus 5 redemption arc
Kimi-K3 big model smell, ey?
seeing GPT-5.6-Sol like this hurts a little tho
Tanishq Mathew Abraham, Ph.D.@iScienceLuvr · 博主 · 1 天前

Prime Intellect 在自动研究方面进行了一个很好的实验

我一直主要用 GPT 5.6 Sol 来进行自己的自动研究实验,但也许我应该切换到 Claude,或者至少试试 Kimi K3

Prime-agent 改进了 Kimi-K3,希望他们也测试了它是否改进了 Sol 和 Fable…

引用 Prime Intellect @PrimeIntellect我们进行了最大规模开放实验,测试frontier模型进行AI研究的能力。100多次自主运行,10多个模型,8×H200s沙箱运行8天。最佳运行缩小了与人类数月成果82%的差距。查看被引原帖 ↗
查看英文原文
Great experiment by Prime Intellect on autoresearch

I've been mainly using GPT 5.6 Sol for my own autoresearch experiments but perhaps I should switch to Claude or at least Kimi K3

Prime-agent improves Kimi-K3, wish they tested if it improves Sol and Fable as well...
Cristóbal Valenzuela@c_valenzuelab · 创始人 · 1 天前Runway 联合创始人兼 CEO

AI 正在成为媒体和娱乐领域的新型计算模式。

引用 NVIDIA @nvidiaAI is becoming a new computing model for media and entertainment. At the 2026 Runway AI Summit, NVIDIA’s Richard Kerris explored how real-time generative AI can make creative production a live, artist-controlled process. Powered by NVIDIA Vera Rubin, Runway brought Gen-4.5 to its platform in one day, helping close the gap between imagining a scene and bringing it to life.查看被引原帖 ↗
查看英文原文
AI is becoming a new computing model for media and entertainment.
向阳乔木@vista8 · 中文博主 · 1 天前向阳乔木,中文圈 AI 工具与趋势博主
连环推 ×3

一个Deepseek Harness 插件聚合站,看起来做的还挺认真的。

另外好奇大家都装了什么插件?或者开发了什么插件?

装了一些基础的,识图、文件便捷操作等。

地址见评论区

AshutoshShrivastava@ai_for_success · 博主 · 1 天前高频 AI 新闻与产品动态博主
连环推 ×2

GLM 5.3 相比 5.2 是个巨大升级,他们竟然在 2 个月内就搞定了。太不可思议了。

官方博客说这 100% 的性能提升来自训练后扩展,AI/ML API 的这个测试就是很好的例子。

两个模型都拿到同一个着陆页任务,但 GLM 5.3 用更少的 token 完成了,成本更低,生成的输出质量也更好。

GLM 5.3 很快也要开源了。闭源实验室和开源之间的差距在缩小。

引用 AI/ML API @aimlapiGLM-5.3 vs GLM-5.2:构建你自己的登陆页面。@Zai_org 声称新模型为编码而生,我们进行了对比测试。两个模型均完成相同任务:各自构建一个独立HTML登陆页面。包含打字终端、基准图表、标签页、FAQ、3D背景等,完全自写代码无框架无模板。GLM-5.3耗时161,971个token,成本$0.80;GLM-5.2耗时197,853个token,成本$0.98。查看被引原帖 ↗
查看英文原文
GLM 5.3 is a serious upgrade from 5.2, and they managed to do this in just 2 months. It's incredible.

The official blog says the 100% gains came from scaling post training, and this test from AI/ML API is a good example of that improvement.

Both models got the same landing page task, but GLM 5.3 completed it with fewer tokens, at a lower cost, and generated a better output too.

GLM 5.3 is also going to be open source soon. The gap between closed labs and open source is disappearing.
Check out
aimlapi.com

They offer “One API for 1,000+ AI models.”
Gary Marcus@GaryMarcus · 博主 · 1 天前

阅读科技大佬不想让你知道的内容:Marcus论AI

查看英文原文
Read what the tech bros don’t want you to know: Marcus on AI
歸藏(guizang.ai)@op7418 · 中文博主 · 1 天前歸藏,中文圈 AI 工具与提示词博主

CodePilot 这几天也有不少的更新:

1. 支持了新发布的 Grok 4.6、DeepSeek V4 Pro 以及 GLM 5.3 模型。

2. 支持将你 Grok 会员里的图像模型甚至视频模型的额度,也都在 CodePilot 里面使用。支持对 Twitter 进行检索。

这样的话,你只要有一个 Twitter 会员,就可以在 CodePilot 里面白嫖 Grok 的图像、视频模型以及 Twitter 检索服务。

Simon Willison@simonw · 博主 · 1 天前Django 框架联合创造者,AI 工具深度评测

我提高了context limit,让它再试一次,结果生成了个漂亮的(动画)圆形
gist.github.com/simonw/39977…

查看英文原文
I bumped up the context limit and let it have another go and it sure did produce a beautiful (animated) circle
gist.github.com/simonw/39977…
Guillermo Rauch@rauchg · 创始人 · 1 天前Guillermo Rauch,Vercel 创始人兼 CEO

我有兴趣寻找和资助下一个迪士尼。

忠于其原始价值观,专注于美国精神。价值观、家庭、克服逆境、乐观主义、对幸福的追求。

AI 让每个人都能进行创意表达。定义我们下一代的是我们用它创造的*内容*。下一代的偶像。超越绘制或渲染方式的下一个独特品牌。

米老鼠经历了多次技术革命仍然屹立不倒,因为它代表了什么。我认为大多数大型娱乐公司都忘记了这一点。

如果你在这个领域创造了很酷的东西,需要技术或资金帮助,可以 DM 我。

查看英文原文
I’m interesting in finding and funding the next Disney.

True to its original values, focused on the American spirit. Merit, family, overcoming of adversity, optimism, the pursuit of happiness.

AI makes everyone able to create. What will define our next generation is *what* we create with it. The next icons. The next unique brand that goes further than how it’s drawn or rendered.

Mickey Mouse survived several technological revolutions because of what it represented. I think most big entertainment companies have forgotten that.

If you’re creating something cool in this space and need help, technologically or financially, slide into my DMs.
Gary Marcus@GaryMarcus · 博主 · 1 天前

Anthropic 在6月1日秘密提交了 IPO 申请。但*在那之前*,他们是否明确证实过你所说的,每个 token 都能赚钱,并没有补贴 token 成本?如果要真有这种说法,我很想看看报告(欢迎甩链接或私信我)。

另外,(尽管处于静默期)昨天有广泛报道说 Anthropic 在告诉投资者,他们在第二季度实现了“调整后正运营收入”,但这里的“调整后”到底指的是什么?

我渴望透明。

查看英文原文
Anthropic confidentially filed for an IPO on June 1 . But *prior to that* did they ever establish clearly that, as you allege, they make money on every token, and that they are not subsidizing tokens? I would love to read the report if yes (please drop a link or DM).

Also, (despite the quiet period) there is a widespread report yesterday that Anthropic is telling investors that in Q2 they showed “positive adjusted operating income”, but what does “adjusted” mean in that context?

I yearn for transparency.
Tanishq Mathew Abraham, Ph.D.@iScienceLuvr · 博主 · 1 天前

顺便说一下,如果你想了解我们期刊俱乐部接下来的讨论会,一定要加入 @MedARC_AI 的 Discord 服务器!

引用 Tanishq Mathew Abraham, Ph.D. @iScienceLuvrMeta AI研究员、DINOv2创建者Tim Darcet讲解现代自监督学习和新方法CAPI。其在MedARC AI journal club的演讲信息丰富、表述清晰。查看被引原帖 ↗
查看英文原文
btw if you want to hear about the next sessions in our journal club, make sure to join the
@MedARC_AI
discord server!
AshutoshShrivastava@ai_for_success · 博主 · 1 天前高频 AI 新闻与产品动态博主

Gemini Omni Flash 不错。

查看英文原文
Gemini Omni Flash is good.
Tinyfool@tinyfool · 中文博主 · 1 天前

都在研究ai,我在考古,当然我也是在考AI的古

GitHubDaily@GitHub_Daily · 中文博主 · 1 天前

karpathy 的 autoresearch 带火了一个方向,衍生项目一茬接一茬,已经自成生态。

awesome-autoresearch 把它们整理成一份清单:通用改造、研究 Agent 系统、各平台移植,一路到评测基准。

翻了一圈比预想的丰富,Claude Code、Codex、Gemini CLI 的移植都齐了。

GitHub:
github.com/alvinunreal/aweso…


想追这个方向的,可以先从这份清单入手。

Tanishq Mathew Abraham, Ph.D.@iScienceLuvr · 博主 · 1 天前

@SophontAI 现在事儿都牛逼啊。能量特别高。

引用 Tibo @thsottiauxIncredible things are happening at OpenAI right now. Energy is high.查看被引原帖 ↗
查看英文原文
Incredible things are happening at
@SophontAI
right now. Energy is high.
Tanishq Mathew Abraham, Ph.D.@iScienceLuvr · 博主 · 1 天前

Dario 罕见地在 Twitter 上评论了 Gavin Baker 和 Sholto Douglas 之间的交流。

我觉得 Dario 直接在 Twitter 上解释 Anthropic 对 AI 监管和风险的看法很棒,他肯定应该多做这样的事。

引用 Dario Amodei @DarioAmodei作者否定'监管导致权力集中'论点。Anthropic的政策提案旨在减缓frontier AI公司同时扶持小型竞争者。支持SB53和SB1047,主张对frontier模型制定更严格的测试标准。查看被引原帖 ↗
查看英文原文
In a rare appearance on Twitter, Dario weighs in on the exchange between Gavin Baker and Sholto Douglas.

I think that it's great that Dario is directly explaining Anthropic's views on AI regulation and risk on Twitter, and he should definitely do more of it.
GitHubDaily@GitHub_Daily · 中文博主 · 1 天前

装了一堆 Skill,好不好用心里没数。想加新的,还得翻帖子找链接,手动拷文件夹。

SkillNet 干脆把这套流程封装成工具,包含一个公共技能库并配一套命令行。

可以按语义搜索现成的 Skill,只要一条命令即可装进 Agent 工作区。

在安装之前,还能先看评分,安全性、完整性、能不能跑都有分。

GitHub:
github.com/zjunlp/SkillNet


反向生成这个功能有点东西:丢一个仓库或一段执行记录给它,自动提炼成结构化 Skill。

来自浙大团队做的,论文和在线站点都配齐了。手里 Skill 多到记不清的,拿它管管看。

GitHubDaily@GitHub_Daily · 中文博主 · 1 天前

给 Claude Code 的记忆文件越写越长,决定、踩坑、纠正全往里塞,矛盾的记录并排躺着。

IWE 换了个思路,把一个 markdown 文件夹变成知识图谱,笔记之间用链接组织。

我们在编辑器里正常读写,AI 通过命令行和 MCP 来查,同一批文件两个入口。

GitHub:
github.com/iwe-org/iwe


记忆存成一堆互相链接的小文档,Agent 要什么查什么,不用每次全量重读一个大文件。

基于 Rust 写的,两万个文件一秒内处理完,数据全在本地,git 直接管版本。

说实话这个痛点戳中我了,单文件记忆堆到后面,自己都不敢相信里面的结论。

AshutoshShrivastava@ai_for_success · 博主 · 1 天前高频 AI 新闻与产品动态博主

咱们以后能在Dario的帖子和博客上用上Claude的文本水印吗?

引用 Dario Amodei @DarioAmodei1/2 Thanks Gavin for an especially thoughtful exchange. I don't usually spend much time on social media but I wanted to engage here because it really brings out the heart of an important conversation. First, on regulation, I think that “either concentrate it in the hands of a chosen few companies and politicians via regulation or distribute it widely” is a false choice.  I know that there’s a sort of Silicon Valley shorthand where regulation = regulatory capture = concentration of power, but I’ve always found this to be an overly simplified picture of the world.  Many people outside this bubble think of regulation as something that constrains corporate power and benefits ordinary people.  I don’t necessarily agree with that perspective either, rather I think it’s complicated and really depends on what the “regulation” consists of.  But in particular I think that those in the “regulation = regulatory capture = concentration of power” frame often underrate the decentralizing power of objective and fair institutional processes.  A crude analogy is that the formal court system can sometimes feel stuffy and elitist, but it does a much better job of defending the rights of vulnerable individuals than the alternative, mob justice.  At their best, institutions can vest power in ideas rather than people, and thereby decentralize that power. This is why Anthropic has always made its policy proposals very carefully.  We try very hard to make proposals that disadvantage (slow down) frontier AI companies while *advantaging* smaller competitors.  California’s SB53 (which we supported), and even the much-maligned SB 1047 (which we were ambivalent on), completely exempt any company below a certain amount of revenue or model training costs from being covered at all (it was $500M for SB 53, lower for 1047 but we objected to that).  More recently the testing process we’ve advocated for at CAISI and the White House involves more rigorous tests for frontier models than off-frontier mo查看被引原帖 ↗
查看英文原文
Can we use Claude’s text watermark on Dario’s posts and blog going forward?
el.cine@EHuanglu · 博主 · 1 天前

中国的一部1块钱电影赚了16万多美元

不是因为大家听说它好看

而是都想看看它能烂到什么程度

特别是现在 AI 都能制作好莱坞级别的电影了

查看英文原文
a $1 movie in China just made $160,000+

not because people heard it was good

but everyone wants to see how terrible it can possibly be

especially when AI can alr make Hollywood level films
🚨 AI News | TestingCatalog@testingcatalog · 博主 · 1 天前专挖 AI 产品未发布新功能的爆料号

继续阅读 🗞️

testingcatalog.com/anthropic…

查看英文原文
Read more 🗞️

testingcatalog.com/anthropic…
Together AI@togethercompute · 公司官方 · 1 天前

Yutori 的浏览器使用 agent 在紧密循环中运行:截图、操作、重复,每个任务数十次。

在 Together AI 上,他们的 Navigator 模型以 2 倍更快的推理速度和 4-5 倍更低的成本击败了前沿性能。

查看英文原文
Yutori's browser-use agents run in tight loops: screenshot, action, repeat, dozens of times per task.

On Together AI, their Navigator model beats frontier performance at 2x faster inference and 4-5x lower cost.
GitHubDaily@GitHub_Daily · 中文博主 · 1 天前

网上的语音 Agent 教程,大多五分钟跑通一个 demo 就收尾,真要接真实电话就没了下文。

Neural Maze 开源的这门课直接把场景做满:搭一个房产公司的 AI 客服中心,员工全是语音 Agent。

用 Twilio 接打真实电话,语音识别、合成、实时对话全用开源模型,最后部署到 GPU 云上。

GitHub:
github.com/neural-maze/realt…


共有五节课,从架构一路讲到部署和调用追踪,代码全开源。

课程设定挺有意思,不是做玩具,是把一门生意的电话线交给 Agent。

想把语音 Agent 做到能上线的,跟着走一遍不亏。

AIGCLINK@aigclink · 中文博主 · 1 天前

昨天anthropic发布的 8 月风险报告,提到的Model 2超越了mythos ,但最值得关注的不是内部模型model2的能力提升(本身他们也承认没有提升太多),而是:

1、防蒸馏
model 2 把“防蒸馏当成一项常规工程在投入”

2、Anthropic 自己的判定是没达到 RSP 阈值:
Model 2 相对 Mythos 5 只是"明显改进",没有像之前Opus 4.6 → Mythos Preview 是跳了一大步,这是减速的形状,不是加速的形状

3、尺子坏了
任务型评估尺子不能再测量出能力提升情况(也就是benchmark的评估精度不够了),但他们自己文章上在RSI方向上“看到了"加速的早期迹象”

按报告里的披露看,RSI 在当前的测量维度上没看到明显突破,甚至代际跳跃还在收窄。报告里,又承认自己的测量工具已经失灵,还自曝看到了早期加速迹象,但把细节挡住了。

#anthropic
#rsi

Kol Tregaskes@koltregaskes · 博主 · 1 天前

搞不懂为什么那么多人只用一个 agent、一个 AI,其实明显的是你至少需要两个来听听不同的意见。

我今天就碰到这个问题了。同一个问题先问 Claude 再问 Codex,结果 Codex 把我的 Linear issues 搞得一团糟,还啥都没说。虽然我最近一直用 Codex 处理 Linear 的事,但它还是照常乱来。

所以说,如果能承受成本的话,最好同时跑至少两个不同的 agent,这样能听两个意见。

我自己的方案是 Codex x5($100)加 Claude Pro($20),用得不错。而且我还能用 Claude 来复盘 Codex 的意见。

查看英文原文
Don't know why a lot of people just pick one agent, one AI, when it's really apparent you need to have at least two just to get a second opinion.

I found today, after asking the same questions in Claude to Codex, that Codex has been making an absolute mess of my Linear issues and has not said anything. Even though I've been working with Codex a lot in the last few days around Linear Codex has carried on regardless.

It's always worth having at least two different agents on the go that can give you two opinions of you can afford it.

I have one Codex x5 ($100) and one Claude Pro ($20) and that works well for me. And I have Claude talking to Codex about its second opinions.
Gorden Sun@Gorden_Sun · 中文博主 · 1 天前中文圈高频 AI 资讯与开源项目博主

SandAI开源MAGI-2 Preview视频模型
北京三呆科技开源的AI视频模型,MoE架构,114B总参数,6B激活参数,能生成音视频同步的10秒视频。生成的画面看起来有点油腻。
三呆科技成立于2024年,之前效果不错的GAGA-1也是他家的视频模型。

模型:
huggingface.co/sand-ai/MAGI-…

Github:
github.com/SandAI-org/MAGI-2…

AIGCLINK@aigclink · 中文博主 · 1 天前

Qwen3.8-27B的开源模型,能跑在 3090 上,它的意义不是传统的"开源超越闭源"这种老叙事,真正值得关注的是商业生态层面:Opus 4.6 级别的能力,从今天起变成了全球可以人人都可拥有的私有能力,对于全球的AI生态是一次大洗牌。

这次发布的具有很多标志性意义:

一、模型能力地板被抬起来了,而且这个地板是私有的

用户对"一个能用的模型"的最低心理预期,从今天起锚定在 Opus 4.6 那条线以上:低于这条线的模型,无论开源闭源,商业价值会迅速塌陷,而这条线现在的价格是 OpenRouter 上claude官方的 $0.45 / $3.20 每百万 token,或者一张 3090(这是anthropic那边的人要求限制的核心原因,触及其利益了)

二、定价权发生了转移

任何靠"我比开源强"收费模型产品,溢价空间被压到了opus 4.6 以上,其中长程 agent、深度推理能力qwen38-27b不如opus 4.6,这块恰恰是最难做、最难卖、也最难交付的部分,还有中间地带的模型 API 生意,从今天起会很难做。

三、拥有模型主权不再成为模型公司的养料

智能体公司以前要拿到 4.6 级别的能力,必须把数据发出去,当模型公司的养料,现在可以不用。医疗、金融、政务、法务等所有卡在"数据不能出内网"的场景,今天获得了一个此前不存在的选项。这件事的商业含义,我认为比 benchmark 数字大得多,我们已经尝试把医疗模型底座换成qwen3.8试试效果。

四、开源的迭代节奏已经快到眼睛看不清了

Meta 8 月 10 日发布 Muse Glimmer-30B,定位"最强 30B 级开源模型",四天后,Qwen3.8-27B 在 Qwen 公布的每一项重叠指标上都超过了它。

客观说,Qwen3.8-27B也不是全部超越了opus 4.6:吃 benchmark、吃工具调用、吃单点任务的能力被qwen3.8-27b追平甚至反超;但长程可靠性和深度推理方面还是opus4.6强一些,这块就是上面说的溢价机会。

#qwen27b
#qwen
#opus46
#claude

佐敦哥@jordan97995944 · 博主 · 1 天前

沒有比AI 更懂人類了

Gorden Sun@Gorden_Sun · 中文博主 · 23 小时前中文圈高频 AI 资讯与开源项目博主

North-Micro-Vision-Instruct:多模态小模型
Cohere开源,仅2.4B参数,适合OCR、图片描述、视觉定位等任务,不适合复杂推理任务。评分比不上Qwen3.5-2B,Qwen在小模型领域几乎垄断领先。

模型:
huggingface.co/CohereLabs/No…

Kol Tregaskes@koltregaskes · 博主 · 1 天前

ChatGPT 桌面应用已经很糟糕了,今天又变得更糟。它已经开始每两分钟崩溃一次,即使我重新连接了三次也还是一直弹出这些消息。这不只是 exa 的问题,各种连接器都在频繁弹出这条消息。真的是糟糕的一天。

查看英文原文
ChatGPT desktop app is already pretty awful and today it's just taken it to another level. It's started crashing every two minutes and I've had these messages pop up even though I've reconnected three times. It's not just exa but all sorts of connectors that are spawning this message. Really hard day.
佐敦哥@jordan97995944 · 博主 · 1 天前

最具盈利能力的 AI 半導體股票

營業利益率(Operating Margin)排名:
🥇 頂級(50%以上)
$MU(美光科技)
$NVDA(輝達)
$SNDK(SanDisk)

$SKHY

$TSM(台積電)

🥈 菁英級(40%~49%)
$ANET(Arista Networks)
$AVGO(博通)
$KLAC(科磊)

🥉 強勁級(20%~39%)
$ASML(艾司摩爾)
$LRCX(科林研發)
$WDC(威騰電子)
$CRDO(Credo Technology)
$AMAT(應用材料)

$ALAB


⚡ 穩健級(低於20%)
$ARM(安謀)
$LITE(Lumentum)
$MRVL(邁威爾科技)
$AMD(超微半導體)
$COHR(Coherent)
$ON(安森美半導體

Kol Tregaskes@koltregaskes · 博主 · 1 天前

天哪,计算机使用真的需要提速个上千倍吗!!

查看英文原文
Goodness does computer use need to speed up like a thousand-fold!!
Kol Tregaskes@koltregaskes · 博主 · 1 天前

OpenAI,能不能在projects里的session上加未读标记?在外面能工作,但在里面看不到。至少Android应用上看不到。

能不能也弄个活动或提醒页面显示未读的session?这对Codex尤其重要——那些需要审批或要求回复的情况。

查看英文原文
OpenAI, could we please have the unread dot marker on the session inside projects? They work outside, but there is nothing visible inside. At least not on the Android app.

Could we also have an activity or alerts page that shows unread sessions -this goes for mostly Codex - where it's requesting approval for something, or is asking for a reply.
AIGCLINK@aigclink · 中文博主 · 1 天前

补充一句:这还是个多模态模型

hg:
huggingface.co/Qwen/Qwen3.8-…

ginobefun@hongming731 · 中文博主 · 1 天前

过去几期周刊,我们讨论了判断力、验证边界、「1% 法则」和个人 AGI。

第 108 期进入更具体的工程现场,沿六条主线展开:执行模型、Harness、质量与评测、多智能体与权限、路由与缓存,以及执行层进入公共基础设施、物理世界和科学研究之后的新边界。欢迎阅读~

麻瓜链@blockheadchain_ · 博主 · 1 天前

今日 AI 一日总结:
1. CRWV采购服役6年的A100
2. Blackwell GPU集群全面投运
3. AI智能体小世界影响RuneScape
4. 英伟达缩减OpenAI数据中心担保计划
市场观察:NVDA +0.0% / SMH +0.1% / TAO +0.2%

本站由 Jedee杰哥 打造 · 公众号「Jedee杰哥」每早送 AI 日报

姊妹站:𝕏 简中账号数据榜单 · X 关注 @jedeeai · RSS 订阅 · AI 日报 · 历史归档