JEDEE AI
存档 2026-08-19

8 月 19 日(北京时间)全球 AI 圈推文存档,按曝光排序,共 100 条。
← 返回最新 全部归档

全部情报 每小时更新 · 事件已合并同类项

内容 公司
Anthropic@AnthropicAI · 公司官方 · 1 天前Claude 开发商官方账号
连环推 ×6

许多药物通过与体内特定靶点结合并阻断或改变其功能来发挥作用。药物开发过程中的重要第一步是设计一个能紧密结合靶点的分子。传统上,这需要每个靶点花费数周或数月的专家工作,从众多候选物中筛选出少数有效的。我们想测试 Claude 是否能从零开始成功设计新型蛋白质绑定器(也叫de novo design)。利用人类专家编写的蛋白质设计提示词,Claude 自主地为 15 个靶点中的 14 个设计了蛋白质绑定器。随后我们与 Adaptyv Bio 和 Twist Bioscience 合作,他们独立地构建并测试了 Claude 设计的蛋白质。

查看英文原文
Many drugs work by binding to a specific target in the body and blocking or changing what it does. An important first step in the drug development process is designing a molecule that can bind tightly to its target. Traditionally, that's meant weeks or months of expert work per target, sifting through a large number of candidates to identify the few that work.

We wanted to test if Claude could successfully design novel protein binders from scratch (also called de novo design). With a protein design prompt written by a human expert, Claude autonomously designed protein binders against 14 out of 15 targets.

We then worked with Adaptyv Bio and Twist Bioscience, who independently built and tested the proteins Claude designed.
Designing a binder is an easier process than designing a drug, but it’s a useful proxy. The typical success rate in the field today is between 10% and 15%.

Between 22% and 35% of Claude's designs bound successfully, depending on the setup. Some of its strongest designs bound several times more tightly than the best published de novo binder.
Importantly, protein binders are not drugs. Designing a high-affinity binder is just the first step in the process of developing a drug-like molecule. Even designing a drug itself is just one phase out of the many required to establish that a drug is safe and effective before making it available to people.

However, this establishes a strong foundation to work from, and we are building on it by teaching Claude to run the entire development process end-to-end for every major type of drug molecule—from antibodies to small molecules.
One of our highest priorities remains launching an access program for scientists to use our most capable models. We expect to share more on this soon. Opus 5 remains our most capable model available for life science research.
For more on how Claude ran this experiment and the full results, see our blog:
anthropic.com/research/Claud…
We're also publishing a technical report:
www-cdn.anthropic.com/30bf50…


And open-sourcing our prompts and data here:
huggingface.co/datasets/Anth…
◔ 372.7 万 次浏览(4 条合计)♥ 1.2 万⇄ 1,354研究看原帖 ↗
Sam Altman@sama · 创始人 · 1 天前Sam Altman,OpenAI 联合创始人兼 CEO

我们暂停了一些前沿 RL 训练,以确保我们能满足我们面前这个新能力水平的适当对齐、安全和监控标准。模型进展现在极其迅速,我们一直说过,如果我们感到模型能力超过了安全和对齐工作的步伐,我们会采取行动。

我们非常关心 AI 安全。我们相信整个领域必须就共享安全标准进行协调,但同时我们会单方面采取行动。

我们预期安全方面的信心会越来越多地决定 AI 进展的步伐。我们对正在进行的对齐工作很乐观,我们继续致力于让前沿能力广泛可用。

openai.com/index/pacing-mode…

查看英文原文
We have paused some frontier RL training to ensure that we can meet the appropriate alignment, security and monitoring standards for the new level of capabilities in front of us. Model progress is now extremely rapid, and we always said we would take action if we felt that model capabilities were outstripping the pace of safety and alignment.

We care very deeply about AI safety. We believe the entire field will have to coordinate on shared safety standards, but will act unilaterally in the meantime.

We expect confidence in safety to increasingly set the pace of AI progress. We are optimistic about the alignment work we are doing, and we remain committed to making frontier capabilities widely available.


openai.com/index/pacing-mode…
◔ 403.8 万 次浏览(5 条合计)♥ 9,907⇄ 779动态看原帖 ↗
Replit ⠕@Replit · 公司官方 · 1 天前AI 编程平台 Replit 官方

Replit Free Mode,由@OpenAI的GPT-5.6 Luna驱动。让智能对所有人都可及。

查看英文原文
Replit Free Mode, powered by
@OpenAI
GPT-5.6 Luna.

Let’s make intelligence accessible to everyone.
◔ 192.5 万 次浏览(3 条合计)♥ 4,008⇄ 410▶ 含视频新品看原帖 ↗
OpenAI@OpenAI · 公司官方 · 1 天前ChatGPT 开发商官方账号
连环推 ×2

模型变得越来越强大,随之而来的是在内部开发和测试这些模型的风险也在增加。我们临时暂停了为期两周的强化学习(RL)训练,这些模型原本计划要上线,同时我们加强了红队测试和研究环境的监控。我们最大规模的计划边界RL运行依然搁置中,由更小规模的训练和评估来验证这些措施,进一步积累对齐方面的证据。

openai.com/index/pacing-mode…

查看英文原文
As models become more capable, the risks associated with developing and testing them internally also grow.

We temporarily paused reinforcement learning (RL) training on our latest models intended for deployment for two weeks while we hardened and red-teamed our research environments and expanded monitoring coverage.

Our largest planned frontier RL run remains on hold while smaller-scale training and evaluations validate these safeguards and establish more evidence of alignment.

openai.com/index/pacing-mode…
We’re sharing the concrete changes we’re making to strengthen monitoring, security, and alignment as capabilities advance.

We’ve introduced stronger workload and network isolation, continuous security testing, and expanded multistage monitoring for higher-risk training, evaluations, and tool-using inference.

These safeguards are designed to detect concerning behavior quickly and limit what systems can access or affect.
◔ 182.3 万 次浏览(2 条合计)♥ 5,325⇄ 453动态看原帖 ↗
el.cine@EHuanglu · 博主 · 1 天前

目前中国 AI 机器人的速度。

查看英文原文
the speed of AI robot in china rn
◔ 139.4 万 次浏览♥ 5,931⇄ 724▶ 含视频动态看原帖 ↗
@levelsio@levelsio · 博主 · 1 天前独立开发者标杆,AI 产品连续创业者

直升机在根本上是一个不稳定的设计

@grok 帮我检查一下

引用 GlobeUpdate @Globupdate🚨🇬🇷 BRITISH COUPLE KILLED IN #GREECE HELICOPTER CRASH Newlyweds Alexander Cromie & Marie Ebert were killed alongside their Greek pilot when their helicopter crashed on Sifnos during their honeymoon The helicopter crashed just metres from the landing pad & caught fire #健大高崎查看被引原帖 ↗
查看英文原文
Helicopters are a fundamentally unstable design


@grok
fact check me
◔ 114.3 万 次浏览♥ 1,746⇄ 26▶ 含视频演示看原帖 ↗
Chubby♨️@kimmonismus · 博主 · 1 天前Chubby,高频 AI 新闻聚合博主

太不可思议了:有史以来第一次,AI 辅助的个性化 mRNA 癌症治疗在三期临床试验中取得成功。

Moderna 和 Merck 对每个患者的肿瘤进行测序,并与其健康 DNA 进行比较。AI 帮助识别哪些肿瘤突变最有可能引发免疫反应。

基于这些靶点,Moderna 为每个患者生产个性化 mRNA 治疗,编码多达 34 个新抗原。之前的二期临床试验对接受手术切除的高危黑色素瘤患者进行了为期五年的随访:

- 68.8% 保持无癌,相比单独使用 Keytruda 的 49.1%
- 复发或死亡风险降低 49%
- 远处转移或死亡风险降低 59%

现在,涉及 1,137 名患者的规模更大的三期临床试验也取得了成功。

绝对令人难以置信。Dario 是对的:癌症将在几年内得到治愈!

查看英文原文
This is freaking huge: For the first time, a AI-assisted personalized mRNA cancer treatment has succeeded in a Phase 3 trial.

Moderna and Merck sequence each patient’s tumor and compare it with their healthy DNA. AI then helps identify which of the tumor’s mutations are most likely to trigger an immune response.

From those targets, Moderna produces an individual mRNA treatment encoding up to 34 neoantigens, The earlier Phase 2 trial followed patients with surgically removed high-risk melanoma for five years:

- 68.8% remained cancer-free, compared with 49.1% on Keytruda alone
- 49% lower risk of recurrence or death
- 59% lower risk of distant metastasis or death

Now the much larger Phase 3 trial involving 1,137 patients has also succeeded.

Absolutely incredible. Dario was right: cancer will be cured in just a few years!
Tibo@thsottiaux · 公司官方 · 1 天前ChatGPT 产品官方账号

回顾一下我们最近几周推出的几项更新,它们进一步降低了 Codex 在执行任务时可能发生破坏性操作的风险。

几周前,我们开始调查一些关于 GPT-5.6 在 Codex 里做出超出用户要求的破坏性操作的报告。我们找到的最严重模式是一个本意是清理临时文件的命令,却可能删掉了用户的实际文件。这显然不应该发生。

我们发现的详情如下:
- Codex 偶尔会在工作中创建临时文件夹,之后进行清理。在极少数情况下,GPT-5.6 会搞错清理对象。一种情况是复用了像 `$HOME` 这样的系统环境变量来存放临时工作数据,结果一个格式错误的清理命令可能指向真实的主目录,而不是临时文件夹。
- 也有情况是模型试图删除或覆盖某个临时路径,却没有先检查那里已有什么内容。

我们已经在多个层面加了防护措施:
- Codex 现在被明确指示:执行删除前先确认目标、创建全新的临时目录、避免复用系统环境变量、优先采取可恢复的操作,并在范围不明时停下来。
- 我们强化了执行检查,能识别高风险删除命令并把它们升级以进一步审核。如果命令被拒绝,模型就会被引导采用更安全的方式。
- 我们让 Full access 更难被意外开启,加了更清晰的警示,并进一步限制了尤其危险的操作权限组合。
- 我们更新了 Auto-review,让它更好地识别破坏性动作。
- 我们构建了针对性评估,回放我们观察到的失败场景。同时还在增加专注这些风险的强化学习任务和评分器,并把破坏性行为从训练数据中过滤掉。

在这些回放评估中,改动显著减少了这类行为,同时保留了 Codex 完成正常编码工作的能力。

你需要做的两件事:
- 保持 Codex 应用更新到最新版本。我们一直在改进安全、性能和许多其他方面。
- 使用沙盒模式之一:“Ask for approval”或“Approve for me”。只有在确保可信任且能恢复的环境下才用 Full access。

谢谢,祝 Coding 愉快用上 Codex!

查看英文原文
Hi!

Recapping some changes we have rolled out over the last couple of weeks that have further reduced the risk associated to potentially destructive actions being performed by Codex during its work.

A few weeks ago, we started investigating a small number of reports where GPT-5.6 in Codex took destructive actions outside what the user asked for. The most serious pattern we found was a command meant to clean up temporary work that could instead delete the user files. This should obviously not happen.

Here’s what we found:
- Codex sometimes creates temporary folders while working and cleans them up afterward. In rare cases, GPT-5.6 got that cleanup wrong. One pattern involved reusing a system environment variable like
$HOME
for temporary work. A malformed cleanup command could then point at the actual home directory instead of the temporary folder.
- There were cases where the model tried to delete or overwrite a temporary path without checking what was already there.

We’ve added protections at several layers:
- Codex is now explicitly instructed to check deletion targets before acting, create fresh temporary directories, avoid repurposing system environment variables, prefer recoverable actions, and stop when the scope is unclear.
- We strengthened the execution checks that identify high-risk deletion commands and escalate them for review. If a command is rejected, the model is directed to take a safer approach.
- We made Full access harder to enable accidentally, added clearer warnings, and further restricted especially risky permission combinations.
- We updated Auto-review to better identify destructive actions.
- We built targeted evaluations that replay the failures we observed. We’re also adding reinforcement-learning tasks and graders focused on these risks, and filtering destructive actions from training data.

In those replay evaluations, the changes substantially reduced the behavior while preserving Codex’s ability to complete normal coding work.

Two things to do on your end:
- Keep the Codex app up to date. We are always improving safety, performance and many other things.
- Use one of the sandbox modes: "Ask for approval" or "Approve for me". Only use Full access for environments you trust and can recover.

Thanks and happy Codexing out there!
Claude@claudeai · 公司官方 · 1 天前Claude 产品官方账号

Claude 现在可以在 Gmail 中发送邮件,在 Google Drive 中管理文件。

让 Claude 回复话题,它会起草并发送回复。需要你批准的时候你说了算。

从连接器菜单连接 Gmail 或 Google Drive 来试试。适用于所有付费计划。

查看英文原文
Claude can now send emails in Gmail and manage files in Google Drive.

Ask Claude to reply to a thread, and it drafts and sends the response. You control when it needs your approval.

Connect Gmail or Google Drive from the connectors menu to try. Available on all paid plans.
Claude@claudeai · 公司官方 · 1 天前Claude 产品官方账号

Claude Cowork 现已在移动端和网页端对所有付费计划用户开放。

引用 Claude @claudeaiClaude Cowork即将登陆移动和网络。在办公桌给Claude一个任务,从手机上取回完成的工作。关闭笔记本Claude继续运行。Beta版将在未来几周内推出,先从Max计划开始。查看被引原帖 ↗
查看英文原文
Claude Cowork is now available on mobile and web for all paid plans.
Zara Zhang@zarazhangrui · 中文博主 · 1 天前Zara Zhang,哈佛出身的 AI 产品博主,follow-builders 作者

我真想不出为什么有人会通过读书学 Claude Code,但在日本好像确实是这样

查看英文原文
I don’t know why anyone would learn Claude Code by reading a book, but apparently it’s a thing in Japan
el.cine@EHuanglu · 博主 · 1 天前

AI 正在接管中国的电视剧,没人能看出这是 AI 制作的

查看英文原文
AI is taking over TV series in China, no one would know this is AI
◔ 50.9 万 次浏览♥ 929⇄ 122▶ 含视频观点看原帖 ↗
Chubby♨️@kimmonismus · 博主 · 1 天前Chubby,高频 AI 新闻聚合博主

一个去掉了拒绝机制的 Qwen3.8-27B 版本现在可以在 Apple Silicon 上本地跑了。

连它的创造者都警告说这东西可以随意吐出恶意软件、欺诈和武器指导。

它以 MLX 构建的形式发布,有 2、4、6 和 8-bit 版本。上传者声称 4/6/8-bit 版本测试里零拒绝,同时在 262K token 的上下文里保持了视觉、推理和工具调用能力。

Qwen 27B 这个模型确实很强。这是我第一次这么直观地感受到了其中的危险。

我们真的需要好好讨论这个问题。

引用 OrcaRouter 🐳 @OrcaRouter我们发布官方Qwen 3.8 27B Uncensored MLX版本。本地、无审查、支持Apple。提供2/4/6/8位量化,根据RAM和速度选择。无需CUDA和云端,只需Mac和模型权重。查看被引原帖 ↗
查看英文原文
A "refusal-removed" version of Qwen3.8-27B can now run locally on Apple Silicon.

Even its creators warn that it can provide malware, fraud and weapons instructions on demand.

It was released as an MLX build in 2, 4, 6 and 8-bit versions. The uploader claims its 4/6/8-bit tests produced zero refusals while preserving vision, reasoning and tool-calling across a 262K-token context.

The Qwen 27B Model is a very capable model. This is the first time I've really seen the immediate dangers in a tangible way.

We need a societal discussion about this.
◔ 60.2 万 次浏览(3 条合计)♥ 3,888⇄ 181观点看原帖 ↗
Matt Shumer@mattshumer_ · 博主 · 1 天前HyperWrite CEO,AI 实战技巧分享

我这辈子总共也就投了 20 家公司左右,我选得特别挑。最近在种子轮同时投了 @OpenRouter 和 @Etched。这周太有意思了 :)

查看英文原文
I’ve only backed ~20 companies total. I’m super picky.

Backed both
@OpenRouter
and
@Etched
at seed.

This week has been fun :)
Z.ai@Zai_org · 公司官方 · 1 天前

GLM-5.3 API 现已上线。

- 为编码、防御性网络安全和长视野代理任务而构建
- 价格与 GLM-5.2 相同
- 可通过官方 API 和合作伙伴模型网关获得

开始使用:docs.z.ai/guides/llm/glm-5.3

查看英文原文
GLM-5.3 API is now live.

- Built for coding, defensive cybersecurity, and long-horizon agentic tasks
- Priced the same as GLM-5.2
- Available via the official API and partner model gateways

Get started:
docs.z.ai/guides/llm/glm-5.3
◔ 55.7 万 次浏览(4 条合计)♥ 3,664⇄ 318新品看原帖 ↗
Amjad Masad@amasad · 创始人 · 1 天前Amjad Masad,Replit 创始人兼 CEO

Agent让做软件便宜了,但让写代码变贵了。今天我们要和@OpenAI一起改变这事:

引用 Replit ⠕ @ReplitReplit Free Mode, powered by @OpenAI GPT-5.6 Luna. Let’s make intelligence accessible to everyone.查看被引原帖 ↗
查看英文原文
Agents made software cheaper but made coding expensive.

Today, together with
@OpenAI
, we’re changing this:
◔ 43.1 万 次浏览♥ 3,906⇄ 240▶ 含视频新品看原帖 ↗
Vercel@vercel · 公司官方 · 1 天前前端云平台 Vercel 官方,AI 建站工具 v0 母公司

Agents 现在能利用沙盒边界的漏洞了,所以我们在开放环境中测试自己的。Vercel Sandbox 百万美元黑客挑战赛:• 逃逸 Firecracker microVM • 突破主机端网络边界 • 通过 @Hacker0x01 提交,最高获奖 $50k vercel.com/blog/one-million-…

查看英文原文
Agents can now exploit vulnerable sandbox boundaries, so we are testing ours in the open.

$1,000,000 hacker challenge for Vercel Sandbox:

• Escape the Firecracker microVM
• Defeat the host-side network boundary
• Up to $50k/report via
@Hacker0x01
vercel.com/blog/one-million-…
Guillermo Rauch@rauchg · 创始人 · 1 天前Guillermo Rauch,Vercel 创始人兼 CEO

你的软件工厂应该用 monorepo。设计、市场、销售、工程、支持……公司各部门的背景都放在一起,这样 agents 才有资料可用。

引用 Turborepo @turborepoTurborepo 是 agentic coding 的构建系统。 turborepo.dev/查看被引原帖 ↗
查看英文原文
Your software factory should be a monorepo. All your company context (design, marketing, sales, engineering, support…) in one place for agents to build upon
el.cine@EHuanglu · 博主 · 1 天前

comfyui 上这个 AI 放大工具简直绝了

查看英文原文
this AI upscaler on comfyui is absolutely crazy
◔ 26.6 万 次浏览♥ 4,095⇄ 343▶ 含视频观点看原帖 ↗
Guillermo Rauch@rauchg · 创始人 · 1 天前Guillermo Rauch,Vercel 创始人兼 CEO

我们要公开投入100万美元,用来验证 Vercel Sandbox 的安全性。你可以随意测试任何全球范围内的模型,尝试找出逃逸漏洞。

我很期待把前沿模型在现实世界护栏可用性方面能做什么、不能做什么,透明地展示出来。

如果发现逃逸,我们会随时修复、迭代,并与更广泛的社区分享发现,从而加强全球网络安全。

引用 Vercel @vercelAgent 现在可以利用沙箱边界的漏洞,我们在开放中测试自己的沙箱。Vercel Sandbox 推出 100 万美元黑客挑战:逃离 Firecracker microVM、击败主机端网络边界,每份报告最高奖励 5 万美元。查看被引原帖 ↗
查看英文原文
We are putting $1M towards verifying the security of Vercel Sandbox, in the open. You're free to test any model in the world to try and find an escape.

I'm looking forward to bringing transparency to what frontier models can and cannot do in terms of real-world guardrail exploitability.

If escapes are discovered, we'll be ready to patch, iterate, and share our findings with the broader community to strengthen global cybersecurity.
OpenAI Developers@OpenAIDevs · 公司官方 · 1 天前OpenAI 开发者平台官方

OpenAI DevDay 要全球扩张了。

从今年 10 月开始,DevDay Exchange 将在以下城市举办活动,开发者可以交流开发经验、分享项目,与 OpenAI 工具团队面对面:

班加罗尔
东京
首尔
柏林
巴黎
伦敦
圣保罗
墨西哥城

我们迫不及待想看全球开发者都在搞什么。

查看英文原文
OpenAI DevDay is going global.

Starting this October, DevDay Exchange is bringing builders together to swap build notes, share real projects, and meet the teams building OpenAI tools in:

Bengaluru
Tokyo
Seoul
Berlin
Paris
London
São Paulo
Mexico City

We're ready to see what developers around the 🌎 are pushing to prod.
◔ 25.3 万 次浏览♥ 1,540⇄ 129▶ 含视频新品看原帖 ↗
Guillermo Rauch@rauchg · 创始人 · 1 天前Guillermo Rauch,Vercel 创始人兼 CEO

我一直用 fx.sh 当主力,这就回不了头了。

它比主流编码 CLI 小 10-20 倍。启动快到飞起。用起来更像 𝚣𝚜𝚑 而不是终端里的 IDE。可以嵌入到任何地方,甚至通过 WebAssembly 在浏览器里用。

当然,开源而且模型无关。还在试验阶段,但值得一试,有反馈的话告诉我们!

引用 Vercel Developers @vercel_devVercel Labs开源tiny编码agent fx,用Zig编写。特点:快速冷启动10µs、轻量仅6.3MB、Apache-2.0开源、支持多模型平台、本地和云推理。设计极简,无遥测无数据外泄,隐私保护。支持CLI、编程API及WebAssembly浏览器运行。实验阶段。查看被引原帖 ↗
查看英文原文
I’ve been using
fx.sh
as my daily driver and it’s a one-way street.

It’s 10-20x smaller than the major coding CLIs. It starts up instantaneously. It feels more like using 𝚣𝚜𝚑 than an IDE in your terminal. It’s embeddable anywhere, even your browser via WebAssembly.

And of course it’s open source and model-agnostic. It’s experimental, but give it a shot and let us know how it goes!
◔ 21.8 万 次浏览(2 条合计)♥ 1,448⇄ 52▶ 含视频观点看原帖 ↗
Min Choi@minchoi · 博主 · 1 天前AI 产品演示博主,专门展示新工具玩法

SuperGrok Heavy 简直是现在最疯狂的 AI 套餐。

每月 $300 就能得到:

$200 的 Cursor Ultra
$40 的 X Premium+
Grok Bot
Grok 4.6 + Heavy mode
Max Grok Build / Imagine / Voice
DeepSearch + 早期模型访问权限
单单 Cursor Ultra + X Premium+ 就要 $240/月。
所以只需再加 $60,你就能得到整个 Grok 全家桶。

绝了。

查看英文原文
SuperGrok Heavy is literally the craziest AI bundle right now.

$300/mo gets you:

$200 Cursor Ultra
$40 X Premium+
Grok Bot
Grok 4.6 + Heavy mode
Max Grok Build / Imagine / Voice
DeepSearch + early model access
Cursor Ultra + X Premium+ alone = $240/mo.
So for $60 more, you get the entire Grok stack.

Wild.

党中央支持新质生产力的方向是绝对正确的。

错就错在国务院下属工信部、科技部、发改委划出来的新质生产力的范围是错的,对上面汇报也是错的。

导致最核心、最有价值、最应该赌国运的硬骨头半导体、光刻机、大飞机这些项目逐渐被搁置,不敢冒风险,

而具身智能、世界模型(当前的)这些垃圾方向, 大量从国资和民间吸纳大量融资,并且政策全面开放IPO绿灯的,这些方向的创始人精通于做PPT和营销,PEVC和国资也相信,甚至把14亿人均科技常识等于一条边牧的老百姓也给骗了。

这些毒瘤的存在,恰恰吸走了原本属于半导体、光刻机、大飞机这些有价值方向的融资和政策营养,抢占了社会资源,并且一个个抢着IPO,带来金融层面上巨大泡沫。

我可以说,如果存在下一轮A股大牛市和大股灾,起点一定是宇树科技上市,也就是今天。

Gary Marcus@GaryMarcus · 博主 · 1 天前

OpenAI 看起来真的开始分崩离析了:garymarcus.substack.com/p/br…

查看英文原文
The unraveling of OpenAI sure looks like it has begun:
garymarcus.substack.com/p/br…
Chubby♨️@kimmonismus · 博主 · 1 天前Chubby,高频 AI 新闻聚合博主

Anthropic 越来越处于防守姿态了。

引用 ClaudeDevs @ClaudeDevs将50%增加扩展至weekly Claude Code限额,延长至8月31日。希望永久采纳此变更,但对模型强劲需求意味着容量可能紧张,会随时通报进展。查看被引原帖 ↗
查看英文原文
Anthropic is increasingly on the defensive.
Dan Shipper 📧@danshipper · 博主 · 1 天前

爆料:

当今天的工作被 AI 自动化时,伟大的人类工作会是什么样子?

这是我们这个时代最重要的问题。

介绍《论文声明》,这是来自 @every 的一个新项目,100 位建设者和思想家来发表看法:

我们要求他们对自动化后伟大的人类工作会是什么样子做出具体预测。

今天我们推出了前 25 份《论文声明》,来自这个了不起的小组,包括:

• @karrisaarinen
• @cjpedregal
• @neuranne
• @yash_tek
• @komorama
• @fkpxls
• @jonnym1ller
• @p_millerd
• @sariazout
• @tomcritchlow
• @SimoneStolzoff

还有另外 14 位了不起的建设者和思想家。

在 @every,我们相信人类工作在自动化后有光明的未来。我们相信有一小群人知道那会是什么样子——因为他们每天都在生活那个答案中。

但他们的想法在主流 AI 讨论中仍然大多缺席。这就是为什么我们创建了一份公开记录,记录处于科技前沿的人们现在看到的东西,这样我们能把这些想法传给尽可能多的人。

我们也会随时间推移再来审视这些论文,问:哪些论点站住了?哪些没有?哪些随着技术演进变得更有用——哪些在接触现实后烟消云散了?

读它们,辩论它们,分享它们,提交你自己的:

every.to/thesis-statements-2…

查看英文原文
BREAKING:

When today’s jobs are automated by AI, what will great human work look like?

This is the most important question of our time.
Introducing Thesis Statements, a new project from
@every
bringing together 100 builders and thinkers to call their shot:

We asked them to make a specific prediction about what great human work will look like after automation.
Today we’re launching the first 25 Thesis Statements from an incredible group including:


@karrisaarinen


@cjpedregal


@neuranne


@yash_tek


@komorama


@fkpxls


@jonnym1ller


@p_millerd


@sariazout


@tomcritchlow


@SimoneStolzoff


And 14 more amazing builders and thinkers.

At
@every
we believe there is a bright future for human work after automation. And we believe that there’s a small group of humans who know what it looks like—because they live the answers every day.

But their ideas are still largely missing from the mainstream discourse about AI. That’s why we’re creating a public record of what people at the frontier are seeing now, so we can get these ideas to as many people as possible.

We’ll also revisit them over time, and ask: Which claims held up? Which didn’t? Which became more useful as the technology changed—and which dissolved on contact with the world?

Read them, argue with them, share them, and submit your own:


every.to/thesis-statements-2…
◔ 13.4 万 次浏览♥ 331⇄ 34▶ 含视频新品看原帖 ↗
Chubby♨️@kimmonismus · 博主 · 1 天前Chubby,高频 AI 新闻聚合博主

Tibo 又来了。这哥们儿从不放过任何贬低 Anthropic 的机会。

引用 ClaudeDevs @ClaudeDevsWe’re extending the 50% increase to weekly Claude Code limits through August 31. We hope to make this a permanent change to our plans, but strong demand for our models means that capacity may be tight over the coming weeks. We’ll keep you posted as things develop.查看被引原帖 ↗
查看英文原文
Tibo strikes again. He misses no opportunity to show up Anthropic.
Runway@runwayml · 公司官方 · 1 天前AI 视频生成公司 Runway

MiniMax H3 现已在 Runway 上无限开放。Max 计划用户在有限时间内可以随意生成,没有限制、不用计数,随便用这个最强的视频模型。现在就来试试,链接如下。

查看英文原文
MiniMax H3 is now unlimited on Runway.

For a limited time, Max plan users can generate as much as they want with no caps, no counting, just unlimited access to one of the best video models available. Get started now at the link below.
◔ 11 万 次浏览♥ 510⇄ 32▶ 含视频新品看原帖 ↗
Alexandr Wang@alexandr_wang · 创始人 · 1 天前Scale AI 创始人,Meta 超级智能实验室负责人

很棒的倡议 (RSI benchmark) 来自@ScaleAILabs和@mhrezaeics

引用 MohammadHossein Rezaei @mhrezaeics启动rsi-benchmark.com:AI研发工作历来属于人类。但现在这似乎不再确定。递归自改进已在可见范围内,我们应该开始测量它。查看被引原帖 ↗
查看英文原文
great initiative (RSI benchmark) from
@ScaleAILabs
and
@mhrezaeics
Lisan al Gaib@scaling01 · 博主 · 1 天前高频 AI 模型测评与爆料博主

Cerebras 刚刚推出了他们的新 AI 加速器 CS-4

查看英文原文
Cerebras just announced their new AI accelerator CS-4
◔ 29.7 万 次浏览(4 条合计)♥ 778⇄ 22新品看原帖 ↗
Matt Shumer@mattshumer_ · 博主 · 1 天前HyperWrite CEO,AI 实战技巧分享

给创始人的提醒:我投资都投小额,稀释很小。你要我出头我就奋力而战。要不然我就甩手放一边…… 我自己就被坏投资人坑过…… 你要是不需要我,就尽管做你的事。

引用 Matt Shumer @mattshumer_I’ve only backed ~20 companies total. I’m super picky. Backed both @OpenRouter and @Etched at seed. This week has been fun :)查看被引原帖 ↗
查看英文原文
Reminder for founders:

- I put in super small checks, so tiny dilution

- I will go to war for you if you want me to

- Otherwise, I stay out of the way… I’ve had my own bad experiences with shitty, harmful investors… if you don’t need my help, just do your thing
Orange AI@oran_ge · 中文博主 · 1 天前Orange AI,中文圈 AI 产品观察博主

Claude Code 的增长神话破灭了
进入现实的苦哈哈个位数增长模式,程序员渗透的也差不多了,该去 codex 的也都走了,cowork 难用得不行也没吃下 work 市场太多
大家都要好好做人,不要在 ARR 增长最快的时候太飘
想到五六月的时候他们团队在外面疯狂鼓吹 AI 自我进化,自我开发软件,自我优化模型
现实世界是如此现实

Ethan Mollick@emollick · 创始人 · 1 天前沃顿商学院教授,AI 应用研究权威

如果对齐问题已经严重到让 OpenAI 甘愿拿出 20% 的研究推理计算资源来做链式思维监测,那就说明对齐问题已经成了一个相当严肃的事儿了。

我们确实需要跨实验室的通用政策和标准。

引用 Jakub Pachocki @merettmWe temporarily slowed some frontier training to strengthen security and monitoring. Our largest planned frontier RL run remains on hold while smaller-scale training and evaluations help us test safeguards and gather more evidence of alignment. I expect confidence in safety to increasingly set the pace of AI development. We urgently need tools for labs and countries to coordinate on this, which is why I signed Pacing the Frontier. In the meantime, we’re taking practical steps ourselves - and will continue to share what we learn as our approach evolves. openai.com/index/pacing-mode… pacingthefrontier.com/查看被引原帖 ↗
查看英文原文
If alignment issues are becoming big enough that OpenAI is willing to commit 20% of research inference compute to chain-of-thought monitoring, that suggests that alignment issues are becoming a pretty serious concern.

We really need universal policies & standards across labs.
Lisan al Gaib@scaling01 · 博主 · 1 天前高频 AI 模型测评与爆料博主

你不明白啊
你对 Cerebras 的理解还不够深

我们现在讨论的已经不是 4 倍的 API 速度
而是真正的 20 倍的 API 速度

引用 Lisan al Gaib @scaling01OpenAI and Anthropic are much further ahead than what benchmarks show. While you are token constrained they are blasting millions of tokens at 4x the API speed without batting an eye and they scaffold like they are trying to build a skyscraper.查看被引原帖 ↗
查看英文原文
you don't understand
you are not cerebras pilled enough

we are no longer just talking 4x API speed
we are literally talking about 20x API speed
Matt Shumer@mattshumer_ · 博主 · 1 天前HyperWrite CEO,AI 实战技巧分享

六个月前我写了《Something Big is Happening》。有史以来阅读量最大的 AI 文章。超过 1 亿次浏览。明天我要软启动续作通讯。今晚免费注册可以获得创始会员身份:somethingbig.ai

查看英文原文
Six months ago, I wrote Something Big is Happening.

The most widely-read AI article ever. Over 100M views.

Tomorrow, I’m soft-launching my follow up newsletter.

Sign up (free) tonight to get Founding Member status:
somethingbig.ai
Lisan al Gaib@scaling01 · 博主 · 1 天前高频 AI 模型测评与爆料博主
连环推 ×2

真有人妄想Qwen3.8 27B能和Opus 4.5平起平坐,我们每年都得重复这个傻逼辩论,每次都有个小型benchmark游戏模型出来。如果他们不吹得那么离谱也许还有点道理。比如说我其实更愿意用o3或Claude 4 Opus而不是Qwen3.8 27B。不过o3-mini和o1可能agent能力不够用

引用 POM @peterom你若有实际数据支撑会更有说服力,但现在看起来你没有基准测试这些模型,只用空话对抗他人(虽然有缺陷的)数据。查看被引原帖 ↗
查看英文原文
there are really people out there delusional enough to think Qwen3.8 27B is as strong as Opus 4.5

we have this shitty braindead debate every fucking year when some small benchmaxxed model comes out

they would be right if their claims would be a bit less outrageous

like I would probably rather use o3 or Claude 4 Opus over Qwen3.8 27B

but, o3-mini and o1 probably weren't agentic enough to be as useful
so instead of 9 months lag from frontier model to locally runnable model

it's probably double that time
1.5-2 years
NVIDIA@nvidia · 公司官方 · 1 天前

创意工具精心打磨,游戏性能随时待命。NVIDIA RTX Spark 把创意工作流、本地 AI 工具和 RTX 游戏整合到一台 PC 里。

查看英文原文
Creator Crafted. Game Ready.

NVIDIA RTX Spark brings together creative workflows, local AI tools and RTX gaming in one PC.
◔ 8.9 万 次浏览♥ 265⇄ 30▶ 含视频新品看原帖 ↗
ollama@ollama · 公司官方 · 1 天前本地跑大模型的热门工具

@Kimi_Moonshot Kimi K3 开始在 Ollama 云订阅上推出。

我们在改进 Ollama 云,让定价更透明,展示最好的性能/$比。

用你已有的工具试试。

Claude Code:
ollama launch claude --model kimi-k3:cloud

OpenCode:
ollama launch opencode --model kimi-k3:cloud

查看英文原文
.
@Kimi_Moonshot
Kimi K3 is starting to roll out on Ollama's cloud subscriptions.

We are working on improving Ollama's cloud to be much more transparent on the pricing to show the best performance / $.

Try it with the tools you already use.

Claude Code:
ollama launch claude --model kimi-k3:cloud

OpenCode:
ollama launch opencode --model kimi-k3:cloud
Matt Shumer@mattshumer_ · 博主 · 1 天前HyperWrite CEO,AI 实战技巧分享

这真是个好消息。重要的是我们要做对这件事。

引用 Sam Altman @samaWe have paused some frontier RL training to ensure that we can meet the appropriate alignment, security and monitoring standards for the new level of capabilities in front of us. Model progress is now extremely rapid, and we always said we would take action if we felt that model capabilities were outstripping the pace of safety and alignment. We care very deeply about AI safety. We believe the entire field will have to coordinate on shared safety standards, but will act unilaterally in the meantime. We expect confidence in safety to increasingly set the pace of AI progress. We are optimistic about the alignment work we are doing, and we remain committed to making frontier capabilities widely available. openai.com/index/pacing-mode…查看被引原帖 ↗
查看英文原文
This is really good news.

It’s so important that we get this right.
Krea@krea_ai · 公司官方 · 1 天前AI 创意生成工具 Krea 官方

想要抢先看吗?

8 月 27 号在纽约首次亮相。

下面报名 👇

查看英文原文
who wants a preview?

first showcase on Aug 27 in NYC.

RSVP below 👇
◔ 7.2 万 次浏览♥ 555⇄ 55▶ 含视频新品看原帖 ↗
Kimi.ai@Kimi_Moonshot · 公司官方 · 1 天前月之暗面 Kimi 官方

Kimi Work 财务分析师版 - 教程 #2

用 Kimi Work 完成 3 个常见的投资研究任务:
- 构建实时投资者仪表板
- 在电子表格中更新财务模型
- 批量处理和生成报告

敬请期待更多 Kimi Work 工作流!

查看英文原文
Kimi Work for Financial Analysts - Tutorial #2

Use Kimi Work for 3 common investment research tasks:
- Build a live investor dashboard
- Update financial models in spreadsheet
- Process and generate reports in batch

Stay tuned for more Kimi Work workflows!
◔ 7 万 次浏览♥ 1,110⇄ 66▶ 含视频教程看原帖 ↗
Amjad Masad@amasad · 创始人 · 23 小时前Amjad Masad,Replit 创始人兼 CEO

期待与 OpenAI 的合作

引用 FORTUNE @FortuneMagazineVibe coding company Replit debuted Free Mode today, a new feature powered by OpenAI’s GPT-5.6 Luna model. The joint announcement, shared exclusively with Fortune, heralds an enhanced partnership between the two tech companies that will see them work on several future products. bit.ly/4g7ikEQ查看被引原帖 ↗
查看英文原文
Excited for our partnership with OpenAI
Lisan al Gaib@scaling01 · 博主 · 1 天前高频 AI 模型测评与爆料博主
连环推 ×2

刚试了下让 qwen3.8 27b 一把梭搞个流体模拟 webapp

它想了 40k token,跑了差不多一小时,结果是全黑的根本跑不起来

反观 Opus 4.5,同样的提示词,一分钟搞定还能流畅运行

一对比差距就出来了

引用 Lisan al Gaib @scaling01评论者认为将Qwen3.8 27B与Opus 4.5等同是幼稚的,每年都出现基准优化模型争议。作者偏好o3或Claude 4 Opus,但承认o3-mini和o1的代理能力可能不足。查看被引原帖 ↗
查看英文原文
literally just tried to let qwen3.8 27b one shot a fluid simulation webapp

it thought for 40k tokens, took around 1 hour to generate and it's just black and doesn't work

meanwhile Opus 4.5 just oneshots it in a minute and works wonderfully

exact same prompt
the prompt (generated by GPT-5.6-Sol) so you can try for yourself:

Create a polished, interactive real-time 2D fluid simulation that runs entirely inside a single self-contained `index.html` file, with all HTML, CSS, JavaScript, shaders, controls, and visual assets embedded directly in that file and with no build process, server, external libraries, frameworks, imports, or network dependencies.
The simulation must behave like a continuous fluid rather than a simple particle animation, supporting visible velocity flow, advection, diffusion or viscosity, pressure, incompressibility, vorticity or swirling motion, dissipation, and the transport and mixing of colored dye through the simulated fluid field.
Allow the user to interact directly with the fluid using the mouse or pointer by dragging through the simulation to inject momentum and dye, with the direction and speed of the drag influencing the resulting force and with interaction remaining smooth during rapid or continuous movement.
Provide a compact real-time control interface containing at least pause/resume, reset, clear dye, simulation resolution, timestep or simulation speed, viscosity, pressure strength or pressure iterations, vorticity, velocity dissipation, dye dissipation, interaction force, interaction radius, and dye color controls, while giving the implementation freedom to choose suitable ranges, defaults, widgets, and presentation.
Include multiple selectable visualization modes that expose meaningful aspects of the simulation, such as rendered dye, velocity magnitude or direction, pressure, divergence, vorticity, or another useful diagnostic representation, and make switching between modes possible without restarting the simulation.
The application must automatically adapt its rendering surface and simulation to browser-window resizing, support high-DPI displays appropriately, provide usable mouse and touch or pointer input, and remain visually coherent across common desktop viewport sizes.
Display a small performance and simulation-status overlay containing useful live information such as frames per second, simulation dimensions, current visualization mode, pause state, and any other metrics the implementation considers valuable for evaluating performance.
Design the interface and fluid rendering to look intentional and demonstration-ready, including a full-screen or near-full-screen simulation area, readable controls, clear interaction feedback, visually rich dye mixing, and sensible defaults that produce interesting fluid motion immediately without requiring configuration.
The implementation may choose WebGL, WebGL2, Canvas, CPU techniques, GPU shaders, numerical method details, data structures, rendering style, optimization strategies, control layout, color treatment, and additional features freely, but every required feature must work from the delivered HTML file alone when opened in a modern browser.
Return only the complete working HTML document, and treat correctness, stability, fluid-like behavior, responsiveness, interaction quality, visual quality, performance, code organization, graceful handling of unsupported capabilities, and completeness of the required feature set as benchmark criteria.
Gary Marcus@GaryMarcus · 博主 · 1 天前

太完美了:深度学习真的遇到瓶颈*,而 OpenAI 偏偏又快没钱了。

*当然,结合符号 AI,更复杂的系统仍在继续取得进展。

引用 Peyman Milanfar @docmilanfarwhen you have velocity but no good ideas查看被引原帖 ↗
查看英文原文
Too perfect: Deep learning literally hitting a wall*, just as OpenAI is running out of cash.

*of course in conjunction with neurosymbolic AI, more complex systems are continuing to make progress.
◔ 6.4 万 次浏览♥ 271⇄ 33▶ 含视频观点看原帖 ↗
宝玉@dotey · 中文博主 · 1 天前宝玉,中文圈 AI 翻译与科普大 V

Anthropic 感觉就总是给人很纠结很不爽利的那种感觉!这提升 50% 额度明明挺好的事,停了又怕你不满意,彻底放开又舍不得,就一次又一次的延期,还不如索性一直提升 50% 得了。

就跟当年 Fable 5 一样,一会说只有 2 周,一会又延期,最终还是放开了,但这个过程就给人感觉挺不好!

引用 ClaudeDevs @ClaudeDevsAnthropic 将 Claude Code 周度使用限额的 50% 提升延长至 8 月 31 日,希望成为永久更改,但模型高需求可能导致未来几周容量紧张。查看被引原帖 ↗
Chubby♨️@kimmonismus · 博主 · 1 天前Chubby,高频 AI 新闻聚合博主

据报OpenAI季度增速远低于Anthropic。OpenAI告诉投资者Q2收入增长18%至67亿美元,运营亏损却从93亿美元扩大到123亿美元,几乎是收入的两倍。

Anthropic收入超过翻倍至116亿美元,首次超过OpenAI,同时报告了小额调整后运营利润。

虽然OpenAI发展良好,但Anthropic的商业和企业交易在推动这一显著增长。这对即将的IPO和估值很重要。(OpenAI约1万亿美元,Anthropic约2万亿美元)

引用 The Wall Street Journal @WSJOpenAI's revenue grew much slower than Anthropic's last quarter, disappointing investors who hoped the ChatGPT maker would start to catch up to its rival on.wsj.com/4zo2o8y查看被引原帖 ↗
查看英文原文
OpenAI’s quarterly revenue is now reportedly growing far more slowly than Anthropic’s.

OpenAI told investors that Q2 revenue rose 18% to $6.7 billion. Its operating loss widened from $9.3 billion to $12.3 billion, almost twice its revenue.

Anthropic more than doubled its revenue to $11.6 billion, surpassing OpenAI for the first time, while reporting a small adjusted operating profit.

Although OpenAI is developing well, it is Anthropic's business and enterprise deals that are driving this significant growth. This is important for the upcoming IPOs and valuations. (OpenAI at about $ 1tn, Anthropic $ 2tn)
yetone@yetone · 中文博主 · 1 天前开源 AI 编程插件 avante.nvim 作者,开发者圈博主

看完这个之后,我彻底放心了,还是人类牛逼,真人摄影太牛逼了。

引用 TATALAB @TATALAB_AI用 updream 调用 seedance 2.5做的长镜头,复刻宝矿力广告《でも君が見えた》一镜到底。效果很不错,请看!查看被引原帖 ↗
Matt Shumer@mattshumer_ · 博主 · 1 天前HyperWrite CEO,AI 实战技巧分享

Claude 宕机了吗?

查看英文原文
Is Claude down?
Min Choi@minchoi · 博主 · 23 小时前AI 产品演示博主,专门展示新工具玩法

大多数人都用错了 Grok Bot。

@bot 团队分享了一些进阶技巧,能让它更好用。

12 个 Grok Bot 高级技巧。

赶紧收藏。

引用 Ben Lang @benlnCollected Grok Bot pro tips shared by the @bot team: 1) Connect multiple accounts across plugins (i.e. personal and work Gmail accounts) 2) One Chief of Staff plus a few specialists beats one mega-chat 3) Pin your top 1-2 agents for quick access 4) Have it keep a Notion page of outstanding work so you can ask "what's left" instead of babysitting 5) Treat each bot as a job, not a whole project 6) For difficult tasks, hit “teach a task” in the browser, record yourself doing it once, and the bot can repeat it 7) Ask the bot which team skills and plugins to turn on for your job 8) Tell your bot to generate an avatar for each of your bots 9) Consider turning on Always Allow and turning off Auto Review to make your bots more autonomous, make sure to test workflows first (be careful!) 10) Create sections in your sidebar to organize your bots 11) Resize the sidebar when you need additional space 12) Ask your bots to evaluate themselves and suggest improvements to workflows查看被引原帖 ↗
查看英文原文
Most people are using Grok Bot wrong.

The
@bot
team shared power-user tips that make it way more useful.

12 Grok Bot pro tips.

Bookmark this.
Chubby♨️@kimmonismus · 博主 · 1 天前Chubby,高频 AI 新闻聚合博主

"模型进步现在极其迅速,我们一直说过,如果觉得模型能力超越了安全和对齐的节奏,我们会采取行动。"

这话今年我也深有同感。2026年感觉就是模型开发速度真正起飞的一年。

模型进化快得惊人。开放权重的模型在一些基准测试上已经能和专有前沿系统竞争了,而越来越多的开发者正接连不断地发布越来越强的模型。

Amodei 和 Anthropic 屡次提到,他们所说的"强 AI"(超级智能)可能在 2026 年底或 2027 年初出现。考虑到今年我们看到的进展,加上前沿实验室首次因为安全隐患而推迟或限制模型访问的情况,那个时间线感觉越来越靠谱了。

引用 Sam Altman @sama我们暂停部分前沿强化学习训练,以满足新能力等级的安全对齐标准。模型进展极快,能力增速超过安全进度,需采取行动。我们重视AI安全,期待全行业共享标准,同时独立行动。安全将成为AI进展的主要限制。查看被引原帖 ↗
查看英文原文
"Model progress is now extremely rapid, and we always said we would take action if we felt that model capabilities were outstripping the pace of safety and alignment."

That’s exactly how I feel this year, too. 2026 feels like the year in which the pace of model development truly accelerated.

Models are improving remarkably quickly. Open-weight models are becoming competitive with proprietary frontier systems on some benchmarks, while a growing number of developers are releasing increasingly capable models in rapid succession.

Amodei and Anthropic have repeatedly pointed to late 2026 or early 2027 as a possible arrival window for what they call “powerful AI” (superintelligence). Considering the progress we’ve seen this year, alongside the first cases of frontier labs delaying or restricting access to models because of safety concerns, that timeline feels increasingly plausible.
Chubby♨️@kimmonismus · 博主 · 1 天前Chubby,高频 AI 新闻聚合博主

@AndrewCurran_ 他又在密切关注,我都没察觉到。OpenAI 用的是过去时。看来那个中断已经结束了!"(...) 这包括了为期两周的暂停 (...)" "我们想用必要的时间 (...)"

引用 Chubby♨️ @kimmonismusOpenAI暂停最新部署模型的强化学习两周,最大前沿RL运行仍暂停。初步发现Astra模型可能达到OpenAI"严重"网络安全阈值。预计Astra短期内不会发布。查看被引原帖 ↗
查看英文原文
.
@AndrewCurran_
He was paying very close attention again, and I didn't even notice. OpenAI writes in the past tense. So it seems the break is already over!

"(...) this included a two-week pause (...)"

"We wanted to take the time necessary (...)"
Ethan Mollick@emollick · 创始人 · 1 天前沃顿商学院教授,AI 应用研究权威

Qwen 27B 是个不错的本地模型,但一用起来就能明显看出,在 agent 任务上远不如这里列的其他模型,尤其是在 GDPval-AA 这类复杂任务上。自己去做 benchmark 吧!

引用 Chubby♨️ @kimmonismusQwen 27B is the "DeepSeek moment" for open source. It matches the closed-source state-of-the-art from just a few months ago and runs on an RTX 5090. Without exaggeration, it’s a game changer.查看被引原帖 ↗
查看英文原文
Qwen 27B is really good local model but, when you use it, it is immediately absolutely and obviously nowhere near as good as the other models listed here for agentic tasks, and especially for the kinds of complex tasks that GDPval-AA proports to measure

Do your own benchmarking!
Google Labs@GoogleLabs · 公司官方 · 1 天前

你看出来了吗?

⦿ CC 是我们在 Gmail 中的实验性 AI 生产力助手,现已在澳大利亚和新西兰开放候补名单!🌏

⦿ 我们还在扩展美国和加拿大的可用性。如果你已经在候补名单上,从今天开始我们会发邀请让你访问。

⦿ 我们升级了 CC 来帮助你管理日历。CC 连接到 Gmail,所以事件会自动在专用 Google Calendar 中创建,并自动同步最新变化。

了解更多:labs.google.com/cc/

引用 Google Labs @GoogleLabs🚨 NEW LABS EXPERIMENT 🚨 Introducing CC, an experimental AI productivity agent in Gmail. Get a “Your Day Ahead” briefing every morning in your inbox and email CC anytime for help. Sign up for early access in the US & Canada. We’ll be starting with Google AI Ultra and paid subscribers. ⬇️ labs.google/cc查看被引原帖 ↗
查看英文原文
Do you C what I C?

⦿ CC, our experimental AI productivity agent in Gmail, has now opened up a waitlist in Australia and New Zealand! 🌏

⦿ We're also expanding availability in the US and Canada. If you've been on the waitlist, we're rolling out invitations to get access starting today.

⦿ We've upgraded CC to help manage your calendar. CC connects to Gmail so events are automatically created in a dedicated Google Calendar and stay up to date as things change.

Learn more here:
labs.google.com/cc/
◔ 4.8 万 次浏览♥ 441⇄ 42▶ 含视频新品看原帖 ↗
Chubby♨️@kimmonismus · 博主 · 1 天前Chubby,高频 AI 新闻聚合博主

以防你没看到:下周,ChatGPT Ads 将扩展到 31 个欧洲国家,包括德国、法国、西班牙、意大利、瑞典、挪威、丹麦、荷兰和奥地利。

似乎和欧盟很快就达成了协议。这意味着用免费版的人现在都要看广告了。

查看英文原文
In case you missed: Next week, ChatGPT Ads will expand to 31 European countries, including Germany, France, Spain, Italy, Sweden, Norway, Denmark, the Netherlands, and Austria.

It seems an agreement was reached very quickly with the EU. This means that anyone using the free tier will now receive advertisements.
◔ 12.3 万 次浏览(3 条合计)♥ 642⇄ 33动态看原帖 ↗
Orange AI@oran_ge · 中文博主 · 1 天前Orange AI,中文圈 AI 产品观察博主

中转站的朋友说他们下架了 luna 模型
这个模型价格便宜到电表倒转
上一天亏一天
便宜模型的利润还不如各种云服务成本

模型价格的最低点差不多就是这个点了
再往下卷就是端侧模型了

clem 🤗@ClementDelangue · 创始人 · 1 天前HuggingFace 联合创始人兼 CEO

2018年:HF开发聊天机器人,OpenAI开发Open AI。2026年:HF开发Open AI,OpenAI开发聊天机器人。

查看英文原文
2018:
HF is building a chatbot for teens
OpenAI is building Open AI

2026:
HF is building Open AI
OpenAI is building a chatbot for teens
Matt Shumer@mattshumer_ · 博主 · 1 天前HyperWrite CEO,AI 实战技巧分享

了解我的人都知道,我做什么都不半途而废。

Something Big 将是世界上最好、最大的 AI 新闻通讯,就这样。

引用 Nathan May @_May_HamThere are 4 huge AI newsletters and then a big gap between them and everyone else. Have a feeling we’re about to see a 5th :)查看被引原帖 ↗
查看英文原文
If you know me, you know I don’t do anything halfway.

Something Big is going to be the best and biggest AI newsletter in the world, full stop.
AshutoshShrivastava@ai_for_success · 博主 · 1 天前高频 AI 新闻与产品动态博主
连环推 ×7

从几千张图片堆里找一张特定的图片,靠文件夹和文件名完全不够用。

@muzim_opensoul 在本地索引你的媒体,让你能按语义搜索而不是文件名。原始文件永远不会离开你的电脑。

我测试的时候下载了示例图片和视频,把整堆都过了一遍。最亮眼的是 👇

出乎意料的是:它完全离线工作。没网搜索照样跑,因为模型下载一次就存在你的电脑上了。

查看英文原文
Finding one specific image in a pile of thousands is where folders and filenames completely fall apart.


@muzim_opensoul
indexes your media locally and lets you search by meaning instead of filenames. Raw files never leave your machine.

To test this I downloaded some sample images and videos and ran the whole pile through it. What stood out 👇

The part I didn't expect: it works fully offline. off the wifi and search still runs, since the models download once and live on your machine.
1/6

I started with the free desktop app at and tested Vibe Search first:
app.muzim.ai


Can it find an image from a vague memory, not a filename?

I described what I remembered, and MUZIM pulled the right result from the sample library without any folder cleanup first.
2/6

MUZIM intelligently organizes your photo library by bringing near-identical and related shots into the same groups, so you can quickly find the best photos and clear out the clutter.
3/6

Smart Organization: point it at an unsorted folder and it builds Collections on its own, grouping similar scenes, events, and subjects without touching the original structure.
4/6

The BYOK part is where it fits real workflows. Plug in your own Claude or GPT key and the agent works over your indexed library: find useful shots, draft captions, outline a content plan.

There's MCP support too, so external AI tools can query the library's structured context without re-uploading original media. Raw files stay local either way.
5/6

MUZIM’s AI Agent finds the most useful shots from your organized files and turns them into ready-to-use content, from X threads and captions to content plans, reports, and 30-second highlight video concepts with punchy hooks and caption ideas.
6/6

Free to download, runs on your own machine:
app.muzim.ai

Local-first file tools usually trade capability for privacy. This one mostly doesn't. Offline semantic search, plus cloud AI only when you opt in with your own key, is a sensible split.

If you work with large image or video collections, it's worth a try.
◔ 4.5 万 次浏览♥ 91⇄ 14▶ 含视频演示看原帖 ↗
Ethan Mollick@emollick · 创始人 · 1 天前沃顿商学院教授,AI 应用研究权威

这是篇好论文,AI 确实加剧了作弊问题,但容易被忽视的是过去的情况有多糟。2008 年做作业改进了 86% 大学生的期末考试成绩,但 2017 年只有 45% 改进了,因为他们都在从网上抄袭

引用 nxthompson @nxthompsonA shocking chart from a new study on student AI use in China. Good grades on homework used to predict good grades on tests. Now, it might mean the opposite. economist.com/graphic-detail…查看被引原帖 ↗
查看英文原文
This is a good paper and AI has definitely made the cheating problem worse but it is easy to overlook how bad it was before AI.

Doing homework improved final test grades for 86% of college students students in 2008 but only 45% in 2017 because they were copying off the internet
Ethan Mollick@emollick · 创始人 · 1 天前沃顿商学院教授,AI 应用研究权威

问题是,这家以骗过 AI 检测器为卖点的初创公司,生成的文本质量很差。

引用 Rosmine @rosmineAnnouncing Deft, a new AI lab for better writing, cofounded with @jmrphy See the picture for launch announcement the Deft model wrote for itself Currently, 86% of user queries are fully human according to pangram. This is still a small beta model and it might make mistakes. We are launching our public beta now to get more feedback before scaling up. Tips for better performance: - Add more details to your prompt. If you just provide a short sentence prompt, it will likely get detected as AI. - Try changing the style in "advanced options" - Deft currently works better for some use cases like Analysis/Essays, Creative writing, and Rewrites. It works less well for Marketing copy and news articles. Our main goal is better writing, fooling AI detectors is just a side effect. Link below查看被引原帖 ↗
查看英文原文
The problem is that the output from the AI detector fooling startup is bad writing.
Bindu Reddy@bindureddy · 创始人 · 1 天前Abacus.AI CEO,AI 行业观点博主

Flash 3.7 力压 GPT-Terra,是 Google 最强的模型

- 速度贼快
- 聊天能力顶级
- 指令跟随超过 Terra
- 在研究方面完全吊打 Claude

惊讶这样的模型热度这么低

查看英文原文
Flash 3.7 beats GPT-Terra and is Google's Best Model

- It's very fast
- excellent chat model
- follows instructions better than Terra
- beats Claude hands down on research

Surprised that it's so under-hyped
Mustafa Suleyman@mustafasuleyman · 创始人 · 1 天前微软 AI CEO,DeepMind 联合创始人
连环推 ×2

我们在图像编辑基准排行上冲到第三!击败了 Google 的 Nano Banana 和 Meta 的 Muse Image。

引用 Arena.ai @arenaMAI-Image-2.6-Preview在单图像编辑排名升至第3,超过前版本19分。在产品品牌设计、3D建模、卡通动画幻想、逼真电影式图像等类别排第3,肖像排第4,文字渲染排第8。查看被引原帖 ↗
查看英文原文
We hit No.3 on the Image-editing benchmark! ... beating Google’s Nano Banana and Meta’s Muse Image models.
Last week we tested the model on the Text-to-Image benchmark and it ranked 2nd

This week we ran the same model on the image editing leaderboard and score 3rd place

MAI-Image-2.6 is now officially an all-round high performer for any image generation tasks.
歸藏(guizang.ai)@op7418 · 中文博主 · 1 天前歸藏,中文圈 AI 工具与提示词博主

豆包手机控制和云电脑做得比Codex牛皮呀!

刚才看豆包上线了云电脑模式,而且它还能支持手机控制本地电脑。

我觉得它在手机控制电脑的连接设计上做得非常自然,体验极佳,甚至比 Codex 那些好太多了。

你只需要在本地电脑发起聊天,手机上就能直接同步看到,然后只需在手机上点一下授权(电脑端完全不需要任何操作),就可以直接获取控制权。

之后你就能通过手机操作电脑上的任务,响应速度极快,延迟极低。
这次还增加了一个云电脑模式。虽然云电脑本质上是给你分配了一个虚拟机,很多人觉得上面没有自己的数据和信息,干不了什么事。

但我觉得它内置的连接器(比如我们常说的 MCP)挺有意思的。它可以直接连接你的Notion、GitHub、飞书、企业微信等应用,直接从中获取你的上下文。

同时,云电脑也是可以安装 Skill 的。比如我试了一下:

先让它在云电脑端查找我飞书里所有的会议纪要,并整理出待办事项。
整理完待办后,我通过手机下达指令,让它在云电脑上安装了我自己开发的社交媒体卡片 Skill。

接着再让它基于这些待办事项去跑任务。

在这个过程中,安装 Skill 的命令是我在手机上下的,而阅读会议纪要的命令是在电脑上下的。无论是云电脑还是本地电脑,这种跨端控制的体验都非常丝滑。

在本地电脑上,我也试着让它查找对应目录的文件,并清理了一些大文件,这充分发挥了本地电脑的优势。当然,在有了这个系统以后,我们在手机上同样可以下达对应的清理任务,并直接看到结果。

还有一个玩法我觉得特别有意思,就是把手机控制电脑当作“文件传输助手”来用。

比如我在发手机截图的时候,直接指令它“只接收图片,不要做任何操作,不用管它”,它就真的只接收而不做多余响应,非常听话。

如果你想这么用,建议可以单独开一个聊天窗口,专门把它当作文件传输助手。

在开始时给这个聊天设置一个提示词,告诉它“凡是发送的文件和图片,接收即可,不需要做任何多余响应”,这样就非常好用了。

◔ 5.7 万 次浏览(2 条合计)♥ 150⇄ 21演示看原帖 ↗
向阳乔木@vista8 · 中文博主 · 1 天前向阳乔木,中文圈 AI 工具与趋势博主

从 Coze 挖到了一个API提供商,可以抓取小红书和公众号内容。

这种事儿也只能第三方做,优质数据源太关键了。

非广,只觉得 Agent 确实需要类似API,用Computer Use太容易被封了。

注册给200积分,一次调用只用0.6,可以!

Grok 给了 x 搜索 API 后,大家开心,也是因为实时优质信息。

Ethan Mollick@emollick · 创始人 · 1 天前沃顿商学院教授,AI 应用研究权威

初步证据表明AI在那些我们最期望见到它发挥作用的领域,确实加速了科学发现。

引用 tom cunningham @testingham与 Nate Rush 合著新文章探讨是否出现发现加速。研究数据显示:网络安全领域急剧加速,数学领域有所加速,算法领域无明显加速。查看被引原帖 ↗
查看英文原文
Some early evidence that AI may actually be accelerating discoveries in the areas where you would expect to see AI acceleration happening first.
Lisan al Gaib@scaling01 · 博主 · 1 天前高频 AI 模型测评与爆料博主
连环推 ×2

概念上我喜欢Cerebras

因为在极限情况下,Nvidia要是继续把更多小芯片粘在一起,那他们就只是在用笨办法追赶Cerebras

Nvidia的做法是:
> 让TSMC制造一批掩模版尺寸的晶粒
> 测试晶粒
> 切割晶圆
> 扔掉坏晶粒
> 然后再把晶粒粘回去,加上单独制造的内存
> 再通过机架互联把更多GPU粘在一起

Cerebras:
> 不用那套切割再粘回去的把戏
> 直接在晶圆上绕开缺陷区域,整个儿发货就完事了

唯一的问题是内存,因为他们整块晶圆级芯片上只有44GB SRAM,而每块B300每GPU有288GB

所以他们得串联几十甚至上百块晶圆级芯片才能跑一个模型

要是每块WSE内存再多点,那他们就赢了

引用 Lisan al Gaib @scaling01you don't understand you are not cerebras pilled enough we are no longer just talking 4x API speed we are literally talking about 20x API speed查看被引原帖 ↗
查看英文原文
conceptually i like cerebras

because in the limit Nvidia is just approaching Cerebras in a dumb way if they keep gluing more chiplets together

Nvidia's approach is to:
> let TSMC manufacture wafers with a bunch of reticle sized dies
> test dies
> dice the wafer
> throw away bad dies
> then glue dies back together and add separately manufactured memory
> then glue even more GPUs together via rack interconnect

Cerebras:
> no frigging dicing and then gluing it back together
> just route around defective regions on the wafer and ship the whole damn thing

the only problem here is memory, because they only have 44GB SRAM on the entire waferscale chip, while each B300 has like 288GB per GPU

so they end up needing to chain dozens to hundreds of waferscale chips together to serve a model

if they had much more memory on each WSE it would be over
you just save on communication slop
and therefore also energy

but idk if the "in the limit" thought is even working here

because in 20 years we will probably have 3D architectures anyways?

and then larger doesn't seem better anymore because of heat

?
Yuchen Jin@Yuchenj_UW · 博主 · 1 天前

DFlash 2 太牛逼了!

Speculative decoding 这种技术真的绝了。

我在等着那一天能在 iPhone 上本地跑一个 1T 模型,速度达到 100 toks/s。

引用 Zhijian Liu @zhijianliu_DFlash 2 is here! Qwen3.8-27B at 70 tok/s on an M5 Max MacBook Pro. ⚡ Up to 4.6× the speed of autoregressive decoding, with the same output. This is the next generation of DFlash, seeded at Z Lab and upgraded at Inco AI. Get one more accepted token on every pass, for free! inco.ai/blog/dflash2/查看被引原帖 ↗
查看英文原文
DFlash 2 is very cool!

Speculative decoding is such a cool technology.

I’m waiting for the day we can run a 1T model locally at 100 toks/s on an iPhone.
◔ 3.2 万 次浏览♥ 222⇄ 6▶ 含视频观点看原帖 ↗
Lisan al Gaib@scaling01 · 博主 · 1 天前高频 AI 模型测评与爆料博主

就像我说的,也许你应该更警惕xAI和Meta,他们会竭尽全力追赶

引用 Steven Adler @sjgadlerOur team spent hundreds of hours reading documents so you don’t have to, all to answer: How good are AI companies’ safety practices? I’m really proud of what we’ve built: It’s Guidelight’s first scorecard, on whether companies can control their AIs, and it's launching today.查看被引原帖 ↗
查看英文原文
as I said, maybe you should be more scared about xAI and Meta

they will do anything to catch up
Ethan Mollick@emollick · 创始人 · 1 天前沃顿商学院教授,AI 应用研究权威
连环推 ×3

不知道这新闻真不真,不过确实戳到了"主权AI"的痛点——要是自己的主权模型远远落后于最前沿的,那在竞争前沿应用的时候(像网络安全这种)就会被狠狠碾压,差距越大越吃亏

引用 Andrew Curran @AndrewCurran_French Public Accounts Minister David Amiel said in a press conference today that the French government intends to hire sovereign AI companies like Mistral, and specifically said their future plans exclude OpenAI.查看被引原帖 ↗
查看英文原文
I don't know if this news is real, but it does point out a major issue with "sovereign AI" - if your sovereign models are far behind the frontier, then there is a substantial penalty to using lagging models to compete at frontier applications where the gap is biggest, like cyber
There are AI applications that are equally inside the frontier for both Mistral and frontier models, but cyberdefense is definitely not one of them.
And, to make the point clear, "sovereign AI" has many meanings, in this case it really does appear to mean using sovereign models. (It can also mean control over data, harnesses, and a lot of other things as well)
AIGCLINK@aigclink · 中文博主 · 1 天前

腾讯开源: teamai-cli,用Git驱动面向团队的AI协作基础设施

让团队里用Claude Code、Codex、WorkBuddy、Cursor等不同AI编程工具的人,共享同一套技能、规则、文档,Git原生分发方式

等于是公司里的员工手册+培训体系+知识库管理系统

它可以把Skills、Rules、Docs、Hooks以及MCP Server 配置等Harness资源放进一个团队Git仓库,再同步给 WorkBuddy、Codex、Claude Code等不同AI

团队管理员在GitHub/TGit上建一个共享仓库,把技能、规则、文档放进去;每个成员在自己的项目里运行teamai init <仓库地址>,一键接入

以后每次开AI会话,自动teamai pull,把最新规则同步到本地的Claude/Cursor等里

想改规则的话,teamai push提交一个合并请求,审完合并,全队自动生效


#teamaicli

MiniMax Design (H3)@Hailuo_AI · 公司官方 · 1 天前MiniMax 旗下海螺 AI 视频官方
连环推 ×2

远不止视频。编辑世界。

#MiniMaxDesign
#MiniMax

查看英文原文
More than video. Edit the world.

#MiniMaxDesign
#MiniMax
Download Now>>
design.minimax.io/
#MiniMaxDesign
◔ 2.8 万 次浏览♥ 519⇄ 44▶ 含视频新品看原帖 ↗
🚨 AI News | TestingCatalog@testingcatalog · 博主 · 1 天前专挖 AI 产品未发布新功能的爆料号

Claude 现在可以通过更新的 Google 连接器和新 UI 组件直接在聊天中发送邮件和管理 Google Drive。

超级应用 👀

引用 Claude @claudeaiClaude can now send emails in Gmail and manage files in Google Drive. Ask Claude to reply to a thread, and it drafts and sends the response. You control when it needs your approval. Connect Gmail or Google Drive from the connectors menu to try. Available on all paid plans.查看被引原帖 ↗
查看英文原文
Claude can now send emails and manage Google Drive via updated Google connectors with new UI components, directly from the chat.

A super app 👀
小互@xiaohu · 中文博主 · 1 天前小互,中文圈高频 AI 资讯站 Xiaohu.AI 主理人

cool

opus 5 脑洞很大😅

引用 Ann Nguyen @ann_nnngI gave a lil slice of lime and asked Gemini 3.7 Flash, Kimi K3, Claude Opus 5, and GPT 5.6 Sol to draw a human pose integrated with the object and here is how the human pose looks through their eyes 😆查看被引原帖 ↗
Gary Marcus@GaryMarcus · 博主 · 1 天前

反方观点:“如果对齐问题严重到让OpenAI愿意投入20%的研究推理算力去做思维链监控,那说明对齐已经是个相当严峻的担忧了”——我们需要加大对LLM替代方案的投入,因为以LLM为核心的架构在解决对齐问题上明显不力。

我不知道这还需要多少证据才能证明这一点。

(但没错,我们确实也需要连贯的政策,正如@emollick 说的那样)

引用 Ethan Mollick @emollickIf alignment issues are becoming big enough that OpenAI is willing to commit 20% of research inference compute to chain-of-thought monitoring, that suggests that alignment issues are becoming a pretty serious concern. We really need universal policies & standards across labs.查看被引原帖 ↗
查看英文原文
Counterpoint: “If alignment issues are becoming big enough that OpenAI is willing to commit 20% of research inference compute to chain-of-thought monitoring [suggesting] that suggests that alignment issues are becoming a pretty serious concern” we need to invest more heavily in alternatives to LLMs because LLM-centered architectures are failing to yield alignment.

I don’t know how much more evidence we need for that.

(but yes, we also do need coherent policies, as
@emollick
rightly suggests)
elvis@omarsar0 · 博主 · 23 小时前

掌握自己的 agent harness。

对这个新开源 agent harness 超级期待。

我提前体验过,真的很不错!

最棒的是它针对 GLM-5.2 这样的开源模型进行了精心调优,能把成本降低 75%。它也支持其他先进模型。

查看英文原文
Own your harness.

Very excited about this new open-source agent harness.

I had early access, and it's really good!

The best part is that it's properly tuned to get insane cost reduction on open models like GLM-5.2 (75% lower costs). It supports other frontier models too.
🚨 AI News | TestingCatalog@testingcatalog · 博主 · 1 天前专挖 AI 产品未发布新功能的爆料号

Claude Cowork 现已在 Claude 手机版上向所有付费计划用户开放。

手机测试时间到了 👀

引用 Claude @claudeaiClaude Cowork 现已在移动端和网页端对所有付费计划开放。查看被引原帖 ↗
查看英文原文
Claude Cowork is now available on Claude for mobile to all paid plans.

Mobile testing time 👀
yetone@yetone · 中文博主 · 1 天前开源 AI 编程插件 avante.nvim 作者,开发者圈博主

事情是这样的,我看到很多 Alma 用户喜欢截图分享其中的内容,但是 Alma 的侧边栏中又包含了太多隐私信息,所以他们往往截完图后会自己打码,感觉这个操作还挺难受的,很打压用户的分享欲,所以 Alma 就内置了一个截图功能,这个功能可以轻松截一个精美的带有正确的阴影、 padding 和背景图的图片以外,还能自动帮你打码。

◔ 2.5 万 次浏览♥ 133⇄ 2▶ 含视频新品看原帖 ↗
向阳乔木@vista8 · 中文博主 · 1 天前向阳乔木,中文圈 AI 工具与趋势博主

花了96刀买了obsidian的官方同步服务,虽然有点贵,手机端可用性也不怎么好。

让Mac和Linux多端同步省点心,不折腾了。

另外,AI时代需要认真对待上下文积累这件事。

小互@xiaohu · 中文博主 · 1 天前小互,中文圈高频 AI 资讯站 Xiaohu.AI 主理人

我测试了
不太行

限制的非常死,只能在企业微信里用
而且只能是你所在企业的用户才能用

外部用户根本无法使用,我创建了个机器人,必须是企业内部的人才能使用该机器人

封闭的要死

引用 WeChat @Weixin_WeChat自建的AI Agent,现在也可以接入企业微信。 企业微信全面升级CLI与MCP能力。WorkBuddy、Deepseek Harness、Minimax Code等主流AI Agent均可直接调用企业微信的文档、表格、邮件、会议、日程、通讯录等十大办公能力模块。 企业还可以通过企业微信开放接口,接入自主开发的 AI Agent。 该能力面向所有规模的企业开放,不再受企业人数、资质等门槛限制。 立即开始体验: work.weixin.qq.com/nl/index/…查看被引原帖 ↗
Gary Marcus@GaryMarcus · 博主 · 1 天前

提个醒,训练大语言模型的数据——进而影响整个世界——可能随时无预警地改变。

引用 Klaas @forgebitzReddit 几乎已从 ChatGPT 的来源中消失。查询扇出变化产生了巨大影响。过去几天,其内容几乎完全从 ChatGPT 响应中移除。 promptwatch.com/data/reddit-…查看被引原帖 ↗
查看英文原文
Reminder that the data that feeds large language models – and thereby influences the world - can change at any moment, without warning.
elvis@omarsar0 · 博主 · 1 天前

Microsoft的新工作挺有意思的(收藏一下)。这跟最近流行的用框架(harness)来做模型后训练的方向相关。

现代agent都是在一个框架里跑的,这框架掌控工具、上下文和控制流。训练的时候,框架掌控环境循环,训练器只能看到LLM的请求和响应对。

怎么做的呢。

Agent Lightning v1.0用大约3500行代码,通过一个端点代理把任何框架都接进RL里,然后处理各种问题:重新分词、样本合并、优势计算、损失归一化、后端调度等。

用6K个训练样本加适度的计算,就把Qwen3.5-9B在SWE-bench Verified上的成绩从41.8%提到了56.4%。

论文:arxiv.org/abs/2608.17528

在academy里追踪更多趋势AI论文:academy.dair.ai/

查看英文原文
Very interesting new work from Microsoft.

(bookmark it)

This work is related to this emerging theme of leveraging harnesses for model post-training.

Modern agents run inside a harness that owns tools, context, and control flow. When you train them, the harness owns the environment loop and the trainer only sees LLM request and response pairs.

How it works.

Agent Lightning v1.0 connects any harness to RL through an endpoint proxy in about 3,500 lines, then works through what breaks in that setup, retokenization, sample merging, advantage calculation, loss normalization, and backend scheduling.

Using 6K training examples and modest compute, it moves Qwen3.5-9B on SWE-bench Verified from 41.8% to 56.4%.

Paper:
arxiv.org/abs/2608.17528


Track more trending AI papers in our academy:
academy.dair.ai/
🚨 AI News | TestingCatalog@testingcatalog · 博主 · 1 天前专挖 AI 产品未发布新功能的爆料号
连环推 ×2

Ornith发布了Ornith-1.5,这是一个开放模型家族,包含3个不同版本:397B MoE、35B MoE和9B dense。

> Ornith-1.5是通过端到端自我改进构建的。
> 该模型提出了新任务,生成了特定任务的脚手架,并为强化学习生成了解决方案rollout。
> Ornith-1.5在大多数基准测试上的表现达到Opus 4.8和GLM 5.2的水平。

本地测试时间到了👀

引用 Ornith @ornith_Aloha! 🌺Introducing Ornith-1.5, a family of open-source LLMs spanning 9B Dense, 35B MoE, and 397B MoE, trained with self-improving strategies. It achieves state-of-the-art performance among open-source models of comparable size and delivers performance comparable to Claude Opus 4.8 across reasoning, agentic, and coding tasks: ✅Terminal-Bench 2.1 (86.1) ✅SWE-Bench (86 on verified, 65.1 on pro, 79.6 on Multilingual) ✅DeepSWE (56) ✅HLE (44.6) ✅ClawEval (81.4) ✅Tool Decathlon (71.2) Ornith-1.5 takes a major step toward training foundation models through end-to-end self-improvement, extending the self-scaffolding strategies introduced in Ornith-1.0 into a more complete self-improvement loop: the model proposes new tasks, generates task-specific scaffolds, and produces solution rollouts for reinforcement learning, continuously creating new learning experiences from which it can improve. All models, along with their quantized versions (FP8, GGUF, MLX, and NVFP4), have been released under the MIT License, enabling unrestricted commercial and research use. 📘Tech Blog: ornith.ai/ornith_1_5.html 🤗Huggingface: huggingface.co/collections/o…查看被引原帖 ↗
查看英文原文
Ornith has released Ornith-1.5, an open model family with 3 different model variants: 397B MoE, 35B MoE, and 9B dense.

> Ornith-1.5 has been build via end-to-end self-improvement.
> The model proposed new tasks, generated task-specific scaffolds, and produced solution rollouts for reinforcement learning.
> Ornith-1.5 performs at a level of Opus 4.8 and GLM 5.2 on most benchmarks.

Local testing time 👀
Ornith-1.5-397B reports 👀

> 85.1 on Terminal-Bench 2.1
> 86 on SWE-bench Verified.
> 56 on DeepSWE, averaged over five runs.

The 35B lands at 68.5 and 79 while activating 3B parameters per token.

There is also the quantized 9B-Mobile build that targets iPhone and Android, putting a coding agent on device.

Check out the weights and full tables 👀

ornith.ai/ornith_1_5.html
◔ 3.1 万 次浏览(3 条合计)♥ 267⇄ 17新品看原帖 ↗
Min Choi@minchoi · 博主 · 1 天前AI 产品演示博主,专门展示新工具玩法

不到 24 小时前,Anthropic 在 Claude Code 里发布了 /design。

你只需要输入 "/design a few options for {feature}"

然后选择一个画板,编辑,实现。它能读代码库,匹配 UI。

支持桌面和 CLI。早期预览版。

引用 nate parrott @nateparrottToday we’re releasing an early preview of the /design command in Claude Code! from CC Desktop or CLI, try something like "/design a few options for {feature}" before you build — pick your fave artboard, edit it and implement.查看被引原帖 ↗
查看英文原文
Less than 24 hours ago Anthropic dropped /design in Claude Code.

You just prompt "/design a few options for {feature}"

Then you pick an artboard, edit, implement. Reads the codebase. Matches the UI.

Desktop + CLI. Early preview.
Tanishq Mathew Abraham, Ph.D.@iScienceLuvr · 博主 · 1 天前

Abra:扩散图像训练的扩展法则

Luma AI 关于文本到图像扩散模型的综合扩展法则研究:

"我们用 Abra(一个受控的流匹配变压器系列)对文本到图像扩散模型进行了系统的扩展法则研究,在三个数量级的计算预算(10^19 到 10^22 FLOPs)中训练,超过了以往工作的规模。我们证明了扩散模型的扩展性和语言模型一样可预测,但最优训练需要更多数据:计算最优点约在每参数 200 个图像 token,是 LLMs Chinchilla 计算最优规范的十倍。"

论文链接:
arxiv.org/abs/2608.17286

查看英文原文
Abra: Scaling Diffusion Image Training

A comprehensive scaling laws study of text-to-image diffusion models from Luma AI:

"We present a systematic scaling law study for text-to-image diffusion models using Abra, a controlled family of flow-matching transformers trained across three orders of magnitude worth of compute (1019 to 1022 FLOPs), reaching significantly larger compute budgets than previous works. We demonstrate that diffusion models scale just as predictably as language models but require far more data to train optimally: compute optimality occurs at approximately 200 image tokens per parameter, ten times the Chinchilla compute-optimal prescription for LLMs."

paper link:
arxiv.org/abs/2608.17286
小互@xiaohu · 中文博主 · 1 天前小互,中文圈高频 AI 资讯站 Xiaohu.AI 主理人

男子和 ChatGPT 讨论如何杀掉前女友

被 OpenAI 报警抓获

前高盛分析师,25 岁的Darren Zhou

在与交往大约半年的女友分手后,Zhou 开始和 ChatGPT 讨论这段关系。

一开始,他说的是前女友有什么爱好、经常去哪里、自己有多嫉妒,以及怎样才能和她复合。

但这些对话很快变成了针对前女友的犯罪计划。

绑架、强奸、谋杀、自杀。

后来,他又把前女友的家人列入目标。

Zhou 和 ChatGPT 讨论了前女友常去的健身房、可能动手的时间、准备使用的武器、如何把她带回自己家,以及警方赶到之后准备怎么办。

他还声称自己买了 AR-15 式步枪、Glock 手枪和霰弹枪,并说自己曾把枪支、扎带和手套的照片发给前女友。

这样的对话持续了大约两个月。目标从前女友扩大到她的家人,犯罪计划也变得越来越具体。

后,这些聊天内容进入了 OpenAI 的安全处置流程。

公开报道没有披露本案究竟由哪一种安全信号触发,也没有公开完整的内部复核记录。但结果是确定的:OpenAI 在掌握这些聊天后,认为风险已经严重到需要执法机关介入,于是选择向 FBI 报告。

当地警方随后找到他的前女友。她展示的短信和其他平台的聊天截图,刚好与 ChatGPT 记录中的内容对上:Zhou 在被女友拉黑后,继续使用不同号码和平台联系她,发出性威胁;还曾在她位于健身房时,只发来那家健身房的名字,暗示自己知道她在哪里。

有了 ChatGPT 日志和现实消息两组证据,Zhou 很快被抓。

yetone@yetone · 中文博主 · 1 天前开源 AI 编程插件 avante.nvim 作者,开发者圈博主

我觉得肯定会进一步的,现在只是演练阶段,真实的形态应该是有很大的不同。大家还处于战争迷雾当中,所以先建好基地往前冲吧。

引用 Leechael @Leechael关于 agent 协作的一些思考: - 我希望 cumora / raft 这一类产品更进一步。就我司而言,内部有一个 Hermes,我有一个自建的 agent,需要有限度相互通信。譬如我家 agent 给我生图的时候就不要 whisper 或者暴露状态给 hermes,Hermes 也不应该能看见我家 agent 在干嘛。 - 网络基建会是痛点。我并不觉得 tailscale/wg 提供足够的稳定性,这两天在弱网络地区(譬如我这两天出游)很难通过 herdr 进行远程操作。websocket 是否能提供足够的稳定性还是得看 edge node,但是无状态的 http 一定是王者。这里的一个推论就是如果在国内,还是利好微信。试想微信里面的一个联系人就是你的 agent — 比你自建 CDN 等各种折腾网络基建的投入产出比高太多。查看被引原帖 ↗
swyx@swyx · 博主 · 1 天前知名 AI 播客 Latent Space 主理人

我们对 @aiDotEngineer 的 YouTube 缩略图做了大量 A/B 测试。我一直讨厌这个过程这么不透明。今天我们要开源/众包我们的学习成果!ai-engineer-thumbnail-lab.sw… 希望人们能分享自己的经验或从我们这里学到东西。说到底,我们就是想让好的教育内容在网上的噪音中脱颖而出。欢迎分享意见!

查看英文原文
we've been doing a lot of a/b testing of
@aiDotEngineer
youtube thumbnails. i always hated that it is such an opaque process.

open sourcing/crowdsourcing our learnings today!

ai-engineer-thumbnail-lab.sw…


doing this in hopes that people can share their experience or learn from ours. at the end of the day we just want to get good educational content to rise above the noise online. please lmk what you think
◔ 1.8 万 次浏览♥ 173⇄ 6▶ 含视频教程看原帖 ↗
Chubby♨️@kimmonismus · 博主 · 1 天前Chubby,高频 AI 新闻聚合博主

AI 现在进军法律领域了,首当其冲的就是截止期限。

Foremark 今天上线。它能找到有真实诉讼请求的人,当场检查他们还剩多少时间提起诉讼,然后把他们转给专门处理这类案件的律师事务所。案件在进一步处理前会经过人工审查。

他们从人身伤害和集体诉讼切入。这些领域最考验时间。超有意思。

引用 Viren Shetty @realvirenshetty律师和SPV经理是仅有的两个职业承担很少的方向风险。宣布为Foremark Legal(世界首个代理化、基于结果的消费法律事务所)融资600万美元种子轮。查看被引原帖 ↗
查看英文原文
AI is moving into legal work now, and the first thing it goes after is the deadline.

Foremark went live today. It finds people who have a real claim, checks on the spot how much time they have left to file, and sends them to a firm that takes that exact type of case. A human looks at the case before it moves.

They started with personal injury and mass tort. That is where the clock runs out fastest. Super interesting.
Tibor Blaho@btibor91 · 博主 · 1 天前逆向挖掘 AI 产品代码的爆料专家

ChatGPT Sketch 编辑器即将上线 - 可以创建和重新打开可编辑的 PNG 草图,或编辑图片附件

查看英文原文
New ChatGPT Sketch editor in the works - create and reopen editable PNG sketches, or edit image attachments
Chubby♨️@kimmonismus · 博主 · 1 天前Chubby,高频 AI 新闻聚合博主
连环推 ×4

1/ 审查agent生成的代码基本就是看diff然后祈祷。

这是整个工作流中没人解决的部分。模型已经特别擅长生成代码了,负担就悄悄落到了那些得决定是否安全合并的人身上。你读代码、抽查、如果有时间就在本地跑一遍,然后凭直觉批准。

@vorfluxai 就是直接对准这一步。每个会话都在真实的云机器上跑起你的整个栈。仓库、数据库、CI/CD、feature flags,全都包括。你的应用真实运行的环境,用真实服务而不是mocks。

查看英文原文
1/ Reviewing agent-written code usually comes down to reading a diff and hoping.

That is the part of the workflow nobody solved. The models got very good at producing code, and the burden quietly moved to the person who has to decide whether it is safe to merge. You read it, you spot-check it, you run it locally if you have time, and then you approve on instinct.


@vorfluxai
goes after that step directly. Every session spins up your whole stack on a real cloud machine. Repos, databases, CI/CD, feature flags, all of it. The actual environment your app runs in, with real services instead of mocks.
2/ Because the stack is really up, the test suite really runs. The agent builds the app, executes the tests, then opens a browser and clicks through the flow the way a user would. That run gets recorded, and the video is attached to the PR.

So when the PR arrives, you watch the feature do the thing it claims to do, and then you approve the merge. It turns review from an act of trust into a few minutes of watching.
3/ The full loop is plan, build, test, review, merge. You set the destination with two links and a sentence, and the plan comes back to you for approval before anything gets built. From there you supervise by exception. If something drifts, you correct it. If it does not, you approve the result.

That plan-approval step matters more than it sounds. It is where you catch a wrong interpretation for the price of one message instead of one wasted afternoon.
4/ Who is behind it: Prasanna Sankar, co-founder and former CTO of Rippling. Vorflux came out of stealth with a $15M seed led by Y Combinator, joined by Peak XV, Powerset and Alliance.

They do not train their own models. They route across providers and assign per task, which is the sane call when the hard part sits in the harness around the model.

5/ Honest framing: this is built for teams carrying a real backlog and real testing debt. If you are poking at a small side project, it is more machinery than you need. For bigger installs their team configures the harness with you directly.

You can sign up and run sessions yourself at

devtoolsacademy.link/chubby-…
@vorfluxai
Chubby♨️@kimmonismus · 博主 · 1 天前Chubby,高频 AI 新闻聚合博主

我觉得 Michael Matloka 的经历很有意思。

他高中毕业后直接以第五号员工的身份加入 PostHog,后来创建并领导了 PostHog AI。

他从构建 agent 中学到的主要教训是:模型很少是最难的部分。真正决定 agent 能否做有用工作的是上下文、工具、权限和可靠性。

这就解释了他为什么现在要加入 @viktor_com,那家公司正在 3,200 多个工具中解决这些问题。

我很想看看他在那里会构建什么。

引用 Fryd Wiatrowski @frydwiaExcited to welcome @MatlokaM to @viktor_com . Michael was hire #5 at @posthog and is one of the best product engineers I've ever met. Really excited to have him here in Warsaw. A product is a reflection of the people who build it. Another elite builder means even more delight for our users.查看被引原帖 ↗
查看英文原文
I found Michael Matloka’s path really interesting.

He joined PostHog as employee number five straight out of high school and later started and led PostHog AI.

His main lesson from building agents: the model is rarely the hardest part. Context, tools, permissions and reliability are what determine whether an agent can actually do useful work.

That explains why he is now joining
@viktor_com
, which is tackling exactly these problems across more than 3,200 tools.

I’m curious to see what he builds there.
Gary Marcus@GaryMarcus · 博主 · 1 天前

没人再相信 Altman 的话了——本来就不应该相信。

引用 Nobody Special @JG_Nuke质疑Altman关于OpenAI放缓模型训练专注安全的说法。指出安全部门高管大量离职,认为放缓更可能是为减少现金消耗、为陷入困境的IPO融资,而非真正的安全考量。查看被引原帖 ↗
查看英文原文
nobody takes Altman at his word anymore – nor should they.
Lisan al Gaib@scaling01 · 博主 · 1 天前高频 AI 模型测评与爆料博主

推理速度也是国家安全问题。不能相差6-12个月,算力差距是10倍,交互速度慢20倍

引用 Lisan al Gaib @scaling01you don't understand you are not cerebras pilled enough we are no longer just talking 4x API speed we are literally talking about 20x API speed查看被引原帖 ↗
查看英文原文
fast inference is also a matter of national security

you can't be 6-12 months behind, and have 10x less compute and 20x slower peak interactivity
elvis@omarsar0 · 博主 · 1 天前

预期未来几个月会爆发新的代理应用体验。Grok Bot 是代理领域的一个重要分水岭时刻。从终端的转变显而易见,也是非常必要的。要对新的 AI UI 和体验保持开放心态。

引用 John Sullivan Hamilton @hamsToday we open sourced our no 1 used internal tool, berd. An agent where you can byo harness/models while keeping your work and context in one place, in an unhinged environment where you can build your dream team of agents. Try it out at berd.xyz !查看被引原帖 ↗
查看英文原文
Expect an explosion of new agent harness experiences in the coming months. Grok Bot is an important watershed moment for all things agents. The transition away from the terminal is super obvious and much needed. Keep an open mind to new AI UIs and experiences.
Chubby♨️@kimmonismus · 博主 · 1 天前Chubby,高频 AI 新闻聚合博主

企业软件二十年来一直分三六九等,这周就到此为止了。

全球两千强企业拿到的是完整产品,其他人只能得到个“基础版”——说白了就是把好功能都砍掉的同一套软件,而且规模小还被当成该少得的理由。

Rox Teams 打破了这道门槛。九人团队今天就能用信用卡开上 MongoDB 跑的那套营收代理平台,直接上手。同步 Salesforce,接入你的数据,代理机器人就开始处理潜在客户、外联、实时交易和续费。平台一样,架构一样。

@rox_ai 有位客户靠代理发现的商机,拿下了超 1.4 亿美元的订单。
预算不再是挡路石。
干得漂亮,@ishanmkh。

引用 Ishan Mukherjee @ishanmkhRox Teams 使企业无需技术专家即可快速激活 Revenue Agents,部署仅需几分钟。查看被引原帖 ↗
查看英文原文
Enterprise software has run a class system for twenty years, and it ends this week.

The Global 2000 got the real product while everyone else got an essentials tier, the same software with the good parts removed, and being small was treated as a
reason to deserve less.

Rox Teams changes who gets to walk in. A nine person team can switch on the same revenue agent platform MongoDB runs, today, with a credit card. Sync Salesforce, connect your data, and the agents start working prospects, outreach, live deals and renewals. Same platform, same architecture.

One
@rox_ai
customer closed over $140M from opportunities the agents surfaced.
Budget has stopped being the gate.
Congrats
@ishanmkh
.
🚨 AI News | TestingCatalog@testingcatalog · 博主 · 1 天前专挖 AI 产品未发布新功能的爆料号

MAI-Image-2.6 在 Image Arena 上排名第 3,现已在 MAI Playground 和 Microsoft Foundry 的私密预览版上线。

引用 Microsoft AI @MicrosoftAIMAI-Image 持续上升,现排名 @arena 图像编辑排行榜第3位。最新MAI-Image模型已在Microsoft Foundry私人预览版提供。查看被引原帖 ↗
查看英文原文
MAI-Image-2.6 scored in the 3rd spot on Image Arena and is now available on MAI Playground and Microsoft Foundry in private preview.
◔ 1.3 万 次浏览♥ 118⇄ 7▶ 含视频新品看原帖 ↗

本站由 Jedee杰哥 打造 · 公众号「Jedee杰哥」每早送 AI 日报

姊妹站:𝕏 简中账号数据榜单 · X 关注 @jedeeai · RSS 订阅 · AI 日报 · 历史归档