JEDEE AI
存档 2026-08-22

8 月 22 日(北京时间)全球 AI 圈推文存档,按曝光排序,共 100 条。
← 返回最新 全部归档

全部情报 每小时更新 · 事件已合并同类项

内容 公司
OpenAI@OpenAI · 公司官方 · 1 天前ChatGPT 开发商官方账号
连环推 ×2

在我们继续拓展能力边界、同时提升效率的过程中,未来三个月内,GPT-5.6 Sol 的 API 和积分价格将下调超过 20%。

查看英文原文
As we continue to push the frontier of capabilities while improving efficiency, we're dropping API and credit pricing of GPT-5.6 Sol by over 20% for the next 3 months.
Now available on the API and rolling out across eligible plans for ChatGPT Work and Codex credits.

Pro, Plus, and Business subscription usage remains unchanged.

developers.openai.com/api/do…
◔ 422.4 万 次浏览(6 条合计)♥ 1.3 万⇄ 807▶ 含视频新品看原帖 ↗
Tibo@thsottiaux · 公司官方 · 1 天前ChatGPT 产品官方账号

Codex 速率限制更新。我们注意到部分用户本周的缓存命中率比前几周的稳定状态下降了。这可能解释了为什么这些用户的使用配额消耗得更快——持续稳定的缓存命中是提高效率的关键。

我们正在调查此事,明天会有进一步的更新。

查看英文原文
Update on rate limits in Codex. We do see that for some users the cache hit rate has been worse this week than the stable state the weeks before. This could explain that usage is draining somewhat faster for those users as hitting the cache consistently is an important component of being efficient.

We are investigating and will have an update tomorrow.
Claude@claudeai · 公司官方 · 1 天前Claude 产品官方账号
连环推 ×4

Claude Security 扫描现已运行在 Claude Mythos 5 上,从今天起向所有 Claude Enterprise 客户提供公开测试版。

用我们最强大的安全模型来扫描你的代码库,无需单独获取模型访问权限。

查看英文原文
Claude Security scans now run on Claude Mythos 5, available today in public beta for all Claude Enterprise customers.

Put our most capable security model to work on your codebase, no separate model access needed.
Point Claude Security at a GitHub repo and Mythos scans for vulnerabilities tracing data across files and reasoning about how components interact. Each finding comes back with a CWE category, confidence and severity ratings, and a suggested fix.
Suggested patches open in Claude Code on the web, using the models your team already uses. These updates are part of our work to give defenders greater access to Mythos results without requiring direct access to the model: the model runs behind the scan and returns findings only. Scans are billed as standard token usage under your existing plan.
It's one of several steps to bring frontier capability to more defenders: we’re working with partners to integrate Mythos 5 into their security products and services, our new Defender Advantage Fund provides $35M in credits for open-source security, and we’re expanding our Cyber Verification Program in the coming weeks. Read more:
claude.com/blog/bringing-cla…
Satya Nadella@satyanadella · 创始人 · 1 天前微软 CEO

我们 Microsoft 数据中心的交付日,第一批生产用 Vera Rubin 设备已到达。特别感谢合作伙伴
@nvidia
以及 Azure 硬件和数据中心团队为这一里程碑做出的辛勤工作!

查看英文原文
Delivery day at our Microsoft DCs as the first production Vera Rubins arrive. A huge thank you to our partners at
@nvidia
and our Azure hardware and datacenter teams for all the incredible work that brought us to this milestone!
Tibo@thsottiaux · 公司官方 · 1 天前ChatGPT 产品官方账号

8pm PST 就会到位 banked reset。针对所有付费的 ChatGPT Work 和 Codex 用户。随便你们怎么用这信息。

查看英文原文
The banked reset will be there by 8pm PST. For all paid users of ChatGPT Work and Codex. Do with this information what you may.
◔ 241.2 万 次浏览(2 条合计)♥ 8,022⇄ 312动态看原帖 ↗
Andrew Ng@AndrewYNg · 创始人 · 1 天前吴恩达,斯坦福教授、AI 教育领军人物

在构建和部署 AI 应用中最重要的技能。

查看英文原文
The most important skills in Building and Deploying AI Applications.
el.cine@EHuanglu · 博主 · 1 天前

哇……AI 竟然把红色警戒 2 变成了 22 分钟的史诗大片

查看英文原文
wwow.. AI just turned red alert 2 into 22mins epic film
◔ 117.4 万 次浏览♥ 9,635⇄ 1,277▶ 含视频演示看原帖 ↗
Grok@grok · 公司官方 · 1 天前马斯克 xAI 旗下聊天机器人 Grok 官方

Grok Bot 现已扩展至更多套餐,并且现在可以免费试用


x.ai/bot

引用 Grok Bot @botWe're making Grok Bot more widely available. All SuperGrok Plus, Cursor Pro+, and Cursor Teams subscribers now have access. We're also offering a free trial with limited usage for all other users.查看被引原帖 ↗
查看英文原文
Grok Bot has expanded to more plans and is now free to try


x.ai/bot
◔ 119.6 万 次浏览(4 条合计)♥ 4,252⇄ 423▶ 含视频新品看原帖 ↗
Sundar Pichai@sundarpichai · 创始人 · 1 天前谷歌 CEO

Gemini 3.7 Flash 在第一周内打破了之前的 Gemini 增长纪录,成为迄今最快增长的模型。很高兴看到开发者社区的热情!现在也在 Search 和 @Geminiapp 中运行。

引用 ARC Prize @arcprizeGemini 3.7 Flash from @Google on ARC-AGI (Verified): - ARC-AGI-2: 84.6%, $0.25/task - ARC-AGI-1: 95.5%, $0.12/task Gemini 3.7 Flash stands out for its low cost and high scores on ARC-AGI-1 and ARC-AGI-2 relative to other frontier models.查看被引原帖 ↗
查看英文原文
Gemini 3.7 Flash smashed previous Gemini growth records in its first week, making it our fastest growing model yet. Great to see the huge excitement from our developer community! Now running in Search and
@Geminiapp
too.
◔ 91.8 万 次浏览(3 条合计)♥ 3,675⇄ 292新品看原帖 ↗
@levelsio@levelsio · 博主 · 1 天前独立开发者标杆,AI 产品连续创业者

记得当年作为PC党,我在2013年买了第一台苹果设备——一台MacBook Pro。

买它是因为我哥(专业VFX大牛)用的也是这个,而且当时创业圈里好多人都换了MacBook。那时候Windows已经难用得不行了。

现在我开始觉得,苹果也到了类似的关头。刚看到@dhh宣布Omarchy的基础项目,@theo的这个视频也说得很准。

十多年来苹果一直对开发者竖中指,但至少软件还算能用。可这几年变了,到处是蠢到家的bug,乱糟糟的,尤其iOS,MacOS也好不到哪去。

举个例子,每次我从iPhone主屏打开相机想抓拍个瞬间(比如飞机慢速飞过),老赶不上,因为相机app要么糊成一片卡住,要么干脆黑屏。

这种基本软件功能,苹果现在都做不到了。

当然,主要的原因——容我说一句——就是Steve不在了。Steve对细节有执念,而现在的苹果压根不在乎细节。

开发者只是社会的极小群体,但他们引领潮流。一旦失去他们,大众很可能也会跟着转向其他平台。

苹果真该好好研究现在的局势。再加上他们对AI投资为$0(有人觉得模型开源这步“走对了”,我可不敢苟同),可能最后连护城河都没了。

我知道得很清楚,如果我从MacBook Pro配MacOS换成Dell XPS跑Linux,那绝对会是个入口,让我接着把iPhone配iOS也换成别家硬件牛逼、装了精简干净Android的手机。

而在这之前,我压根没想过要换阵营!

引用 Theo - t3.gg @theoYour Mac is slowing you down. Moving to Linux has exponentially improved performance for my agents, in particular on the file system side.查看被引原帖 ↗
查看英文原文
I remember as a PC guy how I bought my first Apple device, a MacBook Pro, in 2013

I bought it because that's what my brother had (a pro VFX guy) and also a lot of people in startups around then switched to MacBooks

Back then Windows had become such a POS to use

I'm starting to feel we're reaching a similar moment with Apple

I just saw
@dhh
's Omarchy's foundation announcement and this video by
@theo
is also accurate

Apple has given a consistent big F U to developers for over a decade, but at least their software worked well enough. The last few years that changed and you start seeing really dumb and messy bugs everywhere especially in iOS but also MacOS

As an example every time I open my camera from my iPhone's homescreen to quickly picture somethig that quickly happened (like a plane flying by slow) I miss it because my Camera app goes into a weird blur freeze or just stays black

Just basic software things Apple can't even do anymore

A big reason for this of course, may I say it, is Steve not being there. Steve was obsessed about details. Modern Apple doesn't care about details at all

Developers are just a tiny sliver of society but they are trend setting and if you lose them, there's a good chance the rest of people will move on to other platforms too

Apple should be studying what's happening now because add this to also investing $0 into AI (which people call a "great move" as models go open, I'm really not sure about that) they might not have a moat left

I know for sure if I switch from a MacBook Pro running MacOS to a Dell XPS with Linux that will be a gateway drug for me to also move my iPhone with iOS to some great hardware phone running a great custom minimal Android install

And I never even considered switching before this year!
◔ 43.8 万 次浏览♥ 2,705⇄ 115▶ 含视频观点看原帖 ↗
Guillermo Rauch@rauchg · 创始人 · 1 天前Guillermo Rauch,Vercel 创始人兼 CEO

这个过程真的太疯狂了。我们反复针对 is-agentic.com 运行 is-agentic 直到达到 100/100。

这帮我们填补了很多空白。我们费了不少劲确保标准质量过硬,值得你的时间和 token。

引用 Vercel Developers @vercel_devIntroducing is-agentic.com , a tool to measure how well agents can read your site. Backed by @oradotai 's research, you can run: ▪︎ Audits with 100+ checks ▪︎ Visualizations of agents using your site ▪︎ One-click prompts to fix problems ▪︎ A CLI for agents查看被引原帖 ↗
查看英文原文
This was wild to watch unfold. We ran 𝚒𝚜-𝚊𝚐𝚎𝚗𝚝𝚒𝚌 in a loop against
is-agentic.com
until it got to 100/100.

It made us close quite a few gaps. We worked hard to make sure the criteria is high quality and worth your time & tokens.
Chubby♨️@kimmonismus · 博主 · 1 天前Chubby,高频 AI 新闻聚合博主

Wtf Anthropic,要是觉得 Fable 这周变傻了,那是因为 Claude Code 把"high"推理努力改成了 10/100,以前是"low"。看来他们偷偷把 Fable 改差了。

引用 🥔🥔🥔 @argofowlif fable felt dumber this week, it's not you ❗❗❗ since 2.1.237 the model reads "high" effort as 10 out of 100, the exact number "low" used to be and the changelog doesn't say a word i spent my whole afternoon convinced t3 code and my own app were broken before i went digging through the actual requests anyone else notice or am i going insane?查看被引原帖 ↗
查看英文原文
Wtf Anthropic: If Fable has felt noticeably dumber this week, it’s because Claude Code has apparently been running “high” reasoning effort at just 10/100, the same level that was „low“ before.

They dumbed down fable without telling it seems.
Yuchen Jin@Yuchenj_UW · 博主 · 1 天前

UC Berkeley 开源了 FreeToken,结果相当狂野:

一块 RTX PRO 6000 就能以 14.9 tok/s 跑 753B 的 GLM-5.2!

一台 8GB RTX 4060 笔记本(约 $1,000)跑 Qwen3.6-35B 能到 39.3 tok/s!

在消费级 GPU 上,FreeToken 比 Ollama 快 2–4 倍。本地 AI 推理越来越靠谱了。干得漂亮,@Andy_ShuoYang 和 UC Berkeley Sky Lab!

引用 Shuo Yang @Andy_ShuoYangYour gaming PC can now serve frontier models at interactive speed using official checkpoints without extreme quantization! Qwen3.6 35B → 8GB RTX 4060 laptop @ 39 tok/s DeepSeek-V4-Flash 284B → RTX 5090 desktop @ 22-25 tok/s GLM-5.2 753B → RTX PRO 6000 workstation @ 15 tok/s Run your claude code or codex now with frontier model for $0 Meet FreeToken 🧵查看被引原帖 ↗
查看英文原文
UC Berkeley open-sourced FreeToken. Wild results:

A single RTX PRO 6000 runs the 753B GLM-5.2 at 14.9 tok/s!

An 8GB RTX 4060 laptop (~$1,000) runs Qwen3.6-35B at 39.3 tok/s!

FreeToken is 2–4x faster than Ollama across consumer GPUs. Local AI inference is getting very real. Great work by
@Andy_ShuoYang
and UC Berkeley Sky Lab!
◔ 30.1 万 次浏览♥ 3,534⇄ 285▶ 含视频新品看原帖 ↗
Thinking Machines@thinkymachines · 公司官方 · 1 天前

我们想提升 Inkling 的智能体表现。为了解它在真实环境中的行为,我们现在起几周内,会在 OpenRouter 上免费开放(仅限 agentic harnesses)。我们会利用这些与账号脱钩的数据来改进它。

查看英文原文
We want to improve Inkling’s agentic performance. To help us understand its real-world behavior, we are making it available for free on OpenRouter (only with agentic harnesses) for the next few weeks, starting now. We’ll use the data, disassociated from accounts, to better it.
◔ 33.8 万 次浏览(3 条合计)♥ 2,035⇄ 117新品看原帖 ↗
Gemini Notebook@Gemini_Notebook · 公司官方 · 1 天前谷歌 AI 笔记工具 NotebookLM 官方

抱歉让大家久等了,这周我们确实忙翻了 😮‍💨。以下是这一波上线的新功能:

—— 升级版 Notebook 体验现已向所有用户开放(移动端即将推出!)
—— 你现在可以在 Google 搜索的 AI Mode 里访问你的 notebooks
—— 我们改进了 Chat 中数学公式的复制粘贴和渲染方式
—— 修复了 Chat 中从右到左语言里数字和数学符号倒置的问题

未来几周还有大动作,但老样子,告诉我们你梦想中的功能清单里都有啥吧!

查看英文原文
Apologies for our radio silence, it's been a bit of a busy week for us 😮‍💨. Here's what launched:

— Our upgraded Notebook experience is now available to ALL users (mobile coming soon!)
— You can now access your notebooks in AI Mode in Google Search
— We've improved the way mathematical equations copy paste/render in Chat
— Math and numbers are no longer backwards in right-to-left languages in Chat

Big things to come in the next few weeks, but as always, let us know what's on your dream feature list!
◔ 23.5 万 次浏览♥ 2,594⇄ 209▶ 含视频新品看原帖 ↗
Firecrawl@firecrawl · 公司官方 · 1 天前
连环推 ×4

隆重推出 Firecrawl Developer Index,一个为编程代理加速而生的索引。

搜索超过 7000 万个主要数据源,包括代码仓库、文档和 issue,回忆效果远超任何专为编程场景设计的索引。

确保代理每次都能产出准确、最新的代码!

firecrawl.dev/developer-inde…

查看英文原文
Introducing Firecrawl Developer Index, an index for supercharging coding agents.

Search 70M+ primary sources including repos, docs, & issues with the highest recall of any coding-specific index.

Ensure agents ship correct, up-to-date code every time!


firecrawl.dev/developer-inde…
Why a coding-specific index?

Your agents can answer their own questions about code behavior, API contracts, error messages, & known bugs from the correct primary sources.

That means issues, merged PRs, & READMEs from the most popular public repos & docs sites.
Our Developer Index hits 63% recall@10 across 1,179 real developer queries spanning repos, docs, and PRs. That beats the next best external provider by ~10%.
Ready to try it out?

For best performance we recommend using the Firecrawl CLI or MCP with the companion skill, which you can install with:

npx -y firecrawl-cli@latest setup developer-index
◔ 21.4 万 次浏览(2 条合计)♥ 1,047⇄ 71▶ 含视频新品看原帖 ↗
Bindu Reddy@bindureddy · 创始人 · 1 天前Abacus.AI CEO,AI 行业观点博主

Ox-Alpha 表现不佳,不如上一代模型

鉴于炒作热度,我决定评估 Ox-Alpha,结果相当糟糕

它的评分与两代前的 Kimi 2.6 不相上下

不过从营销角度讲这绝对是天才之举。每个人都在谈论它

查看英文原文
Ox-Alpha Underperforms And Is Worse Than Last Generation Models

Given all the hype we decided to evaluate Ox-Alpha and it turned out to be quite bad

Its scores alongside Kimi 2.6 which is 2 generations old

However, it's pure marketing genius. Everyone is talking about it
◔ 32.8 万 次浏览(8 条合计)♥ 965⇄ 45观点看原帖 ↗
NVIDIA@nvidia · 公司官方 · 1 天前

NVIDIA Vera Rubin 正在进入全面生产阶段。祝贺 @Microsoft 的团队完成了这个令人兴奋的里程碑。

引用 Satya Nadella @satyanadellaDelivery day at our Microsoft DCs as the first production Vera Rubins arrive. A huge thank you to our partners at @nvidia and our Azure hardware and datacenter teams for all the incredible work that brought us to this milestone!查看被引原帖 ↗
查看英文原文
NVIDIA Vera Rubin is ramping into full production. Congrats to the teams at
@Microsoft
who made this exciting milestone happen.
Ethan Mollick@emollick · 创始人 · 1 天前沃顿商学院教授,AI 应用研究权威

现在,回到 Claude Fable 5:“我让 GPT-5.6 Sol 做一张最像 Claude 风格的恶搞图,它搞出了这个。出来打一架吧。”

引用 Ethan Mollick @emollickI asked GPT-5.6 Sol to create the most Claude-y possible parody image and what it came up with is pretty great and dead-on.查看被引原帖 ↗
查看英文原文
Now, back to Claude Fable 5: "I asked GPT-5.6 Sol to create the most Claude-y possible parody image and it came up with this. Come out fighting."
Z.ai@Zai_org · 公司官方 · 1 天前

社区已经用 ZCode + GLM-5.3 做出了超多有趣的项目。

为了感谢大家一路以来的支持,我们决定把 Build Week 做成一个长期系列活动。

从现在起到 8 月 23 日下午 6 点(PT 时间),我们为首 50,000 名新 ZCode 用户每人赠送 100M 免费 GLM-5.3 tokens。

引用 ZCode @zcode_aiGLM-5.3 × ZCode Weekend Build, Round 2 🚀 New users: log in to ZCode for the first time from Aug 22, 00:00 to Aug 24, 09:00 (UTC+8) and automatically get 100M free GLM-5.3 tokens. 50,000 packs, first come, first served; ZCode only. Unused tokens expire when the event ends. Grab yours: zcode.z.ai查看被引原帖 ↗
查看英文原文
The community has already built so many interesting projects with ZCode + GLM-5.3.

To thank everyone for all the support, we’re turning Build Week into an ongoing series.

From now to Aug 23 at 6 PM PT, we’re giving 50,000 new ZCode users 100M free GLM-5.3 tokens each.
◔ 33.9 万 次浏览(3 条合计)♥ 916⇄ 50动态看原帖 ↗
Jim Fan@DrJimFan · 创始人 · 1 天前NVIDIA 具身智能研究负责人
连环推 ×2

触觉是机器人技术中最被严重低估的感知模态。想象一下戴着厚烤箱手套做手法魔术——那就是今天机器人如果有生命时的感觉。磁吸件咔嗒到位、纸杯从叠里剥离、USB插头摸索着进接口——这些在摄像头眼里全是盲区。

学会感知触觉必须是一个全栈协同设计的活儿。我们开源了一套方法论,叫“T-Rex”:

1. 触觉作为模型的一等公民。我们的混合Transformer跑两个异步时钟:一个慢速的视觉运动专家规划动作,一个快速触觉专家以每个视觉tick 4个“触觉tick”的高频修正实时精调动作。力变化比帧到来的还快,所以架构也得跟上节奏。

2. 开放数据。据我们所知,这是有史以来最大的触觉数据集:50小时(约5,500个回合)的高质量、精心同步的机器人操作语料,用22自由度的SOTA触觉手硬件采集。今天就上架HuggingFace!

3. 训练方案:T-Rex扩展了我们之前的工作EgoScale。人类自我中心视角视频做预训练,多样化的触觉机器人操作语料做中间训练。我们的实验表明,这样能很好地架起无接触预训练到高接触操作之间的桥梁。

像素又便宜又无处不在,但一到碰触那一刻就歇菜了。触觉将扛起最后一公里。下一个扩展曲线将以触摸小时数来丈量。

T-Rex是NVIDIA和Berkeley的一次精彩合作:🧵

查看英文原文
The sense of touch is the most criminally under-explored modality in robotics. Imagine doing sleight of hand wearing thick oven mitts. That's exactly how a robot feels today if it were alive. A magnetic piece snapping into place, a paper cup peeling out of a stack, a USB negotiating its way into the port - all invisible to the camera.

Learning how to feel must be a full-stack co-designed effort. We are open-sourcing a principled methodology called "T-Rex":

1. Tactile as first-class citizen of the model. Our mixture-of-transformer runs two clocks asynchronously: a slow visuomotor expert plans the motion, and a fast tactile expert refines it in real time with high-frequency corrections at 4 "touch ticks" per vision tick. Forces change faster than frames arrive, so the architecture had to as well.

2. Open data. The largest tactile dataset ever released to our knowledge: a 50-hour (~5,500 episodes) high-quality, carefully synchronized robot play corpus, collected on SOTA tactile hand hardware with 22 degrees of freedom. Available today on HuggingFace!

3. Training recipe: T-Rex extends our prior work, EgoScale. Human egocentric videos for pretraining, a diverse dose of tactile robot play for mid-training. Our experiments show this bridges contact-free pretraining to contact-rich manipulation remarkably well.

Pixels are cheap and everywhere, but they run out of steam at the moment of contact. Tactile will carry the last mile. The next scaling curve will be measured in hours of touch.

T-Rex is a great collaboration between NVIDIA and Berkeley: 🧵
T-Rex: Tactile-Reactive Dexterous Manipulation
Website:
tactile-reactive-dexterous.g…

Open dataset:
huggingface.co/datasets/zeka…

This work is led by
@Dantong_Niu
and co-advised by
@trevordarrell
. Congrats to the team!
◔ 13.7 万 次浏览♥ 813⇄ 111▶ 含视频研究看原帖 ↗
宝玉@dotey · 中文博主 · 1 天前宝玉,中文圈 AI 翻译与科普大 V

这是一个有意思的 Skill,叫 ELI5,意思是就好比你给 5 岁孩子解释这件事,所以要通俗易懂。生成结果是 HTML 页面。

其实你要是不经常用都没必要去安装这个 Skill,这个 Skill 里面也就是一句提示词而已。

提示词(
github.com/anthropics/claude…
):
请把我当成对这个话题一无所知的人,用一个包含大图且文字较少的 HTML 组件来向我解释。

引用 Thariq @trq212a skill people at Anthropic have been using a lot recently: ELI5 /eli5 <what you want explained> "explain like I'm someone who knows nothing about this topic, using a HTML artifact with big pictures and few words"查看被引原帖 ↗
◔ 12.3 万 次浏览♥ 993⇄ 172▶ 含视频教程看原帖 ↗
宝玉@dotey · 中文博主 · 1 天前宝玉,中文圈 AI 翻译与科普大 V

响马高见:
> 工程师永远不会消失,但是工作方式会变化。工具化永远会进化到 roi 低于人力的边界,然后需要更多的人力来补齐最后的缝隙。

类似于你开了一家餐厅,买了一台洗碗机,一下子替代了 3 个洗碗工,效率高、成本低。这就是工具化带来的 ROI(投入产出比)远高于人力的阶段。

但洗碗机洗不了所有东西。异形的锅、烧焦的铁板、精致的瓷器,这些还得靠人手洗。你可以花大价钱去研发一台能洗一切的超级洗碗机,但为了搞定最后这 5% 的餐具,你可能得投入比前面 95% 多十倍的钱。

这时候你算一笔账:与其砸钱造那台超级机器,不如留一个洗碗工来处理这些边角料,人力反而更划算了。

工具进化到某个点之后,再往前推的成本急剧上升,ROI 跌到比直接用人还低。

另一个角度,工具越强,整个行业的总产出往往也在膨胀。洗碗机让你的餐厅能接待更多客人,于是产生了更多的异形锅、更多的特殊情况,那些“缝隙”的绝对数量反而变多了。所以你可能裁了 3 个普通洗碗工,最后又请了 2 个专门处理疑难杂症的人。

AI 工具会吃掉大量标准化的编程工作,但软件行业的总规模也会因此爆发式增长,而增长带来的各种边界情况、系统集成、业务理解这些“缝隙”,短期内仍然需要人来填。

引用 响马 @xicilion工程师永远不会消失,但是工作方式会变化。工具化永远会进化到 roi 低于人力的边界,然后需要更多的人力来补齐最后的缝隙。查看被引原帖 ↗
Greg Brockman@gdb · 创始人 · 1 天前Greg Brockman,OpenAI 联合创始人兼总裁

朝着为客户提供市场上最低价格以及最高能力天花板而努力

引用 OpenAI @OpenAIAs we continue to push the frontier of capabilities while improving efficiency, we're dropping API and credit pricing of GPT-5.6 Sol by over 20% for the next 3 months.查看被引原帖 ↗
查看英文原文
towards giving our customers the lowest price on the market for any task, as well as the highest ceiling on capability
◔ 10.1 万 次浏览♥ 1,154⇄ 41▶ 含视频观点看原帖 ↗
OpenRouter@openrouter · 公司官方 · 1 天前

我们很高兴把 @AIatMeta 的 Muse Spark 1.2 贡献者套餐带到 OpenRouter。输入 $0.10/M、输出 $0.20/M,价格远低于 Muse Spark 1.2 本体,在真实成本上击败其他可比模型,性价比真的绝了。

查看英文原文
We’re excited to bring
@AIatMeta
’s Muse Spark 1.2 contributor tier to OpenRouter.

At $0.10/M input and $0.20/M output, It’s meaningfully cheaper than Muse Spark 1.2 and beats other comparable models on real cost, providing frontier intelligence-per-dollar.
◔ 12 万 次浏览(3 条合计)♥ 578⇄ 24新品看原帖 ↗
Cristóbal Valenzuela@c_valenzuelab · 创始人 · 1 天前Runway 联合创始人兼 CEO

纽约最大的 AI 实验室正在招聘几乎所有职位。

加入我们:
runway.com/careers

查看英文原文
The biggest AI lab in NYC is hiring for pretty much all roles.

join us:
runway.com/careers
Min Choi@minchoi · 博主 · 1 天前AI 产品演示博主,专门展示新工具玩法
连环推 ×11

Qwen3.8-27B 5 天前引爆了网络。

人们难以相信一个 27B 的模型能有这么强大和"智能"。

即使在笔记本上运行也能解锁新的可能性。

10 个疯狂的例子:

查看英文原文
Qwen3.8-27B broke the Internet just 5 days ago.

People can't believe how powerful and "intelligent" it is for a 27B model.

Unlocking new possibilities even running on a laptop.

10 wild examples:
1. This is scary.

Uncensored Qwen3.8-27B on a Mac.

Asked how to make meth.

It answered.

No cloud. No guardrails.
1. They squeezed Qwen3.8-27B down to 1-bit.

Unsloth says it still retains ~77% performance...

and can run on 8GB RAM.
2. Qwen3.8-27B running locally on a 16GB MacBook Air

No cloud.
3. You can literally run Claude Code with Qwen3.8 locally through Ollama:

ollama launch claude --model qwen3.8
5. One prompt.

One shot.

Qwen3.8-27B running locally vs Gemini 3.7 Flash.
6. Just 4 days after launch...

Qwen3.8-27B became the #1 local model in Cline.
7. It scored 52 on Artificial Analysis.

Same score as GPT-5.6 Luna.

And people are even running it through WebGPU.
8. NVIDIA:

Download Qwen3.8-27B.

Serve it locally.

That’s it.
9. LM Studio has it running at around ~17GB local.

A 27B agentic model is officially laptop-sized.
10. Not just coding.

Qwen3.8-27B tied Fable 5 at 11.3 on the Harvey Legal Agent benchmark.

#1 open-weight model.
◔ 8.7 万 次浏览♥ 717⇄ 45▶ 含视频观点看原帖 ↗
宝玉@dotey · 中文博主 · 1 天前宝玉,中文圈 AI 翻译与科普大 V

Fable 默认用 high 就够了,然后让它主要负责编排、验收,这样是最经济实惠的。

附提示词:
> 注意你的主要任务是分析、编排和验证,具体任务尽可能交给 subagent(Opus 或 Sonnet)去执行。自己只做需求澄清、方案拆解、任务分发和结果验收,实现类工作(读大量代码、写代码、跑测试、批量修改)一律用 Agent 工具派给 subagent 执行。

引用 Go学长 @arkuy99fable max 确实非常消耗 token, 对比下来 xhigh 要节约很多查看被引原帖 ↗
Min Choi@minchoi · 博主 · 1 天前AI 产品演示博主,专门展示新工具玩法

天哪...

人形机器人即将抢占网球领地。

100%自主...

查看英文原文
Holy smokes...

Humanoid robots are coming for tennis.

100% autonomous...
◔ 8.3 万 次浏览♥ 572⇄ 86▶ 含视频动态看原帖 ↗

居然有一位电力行业的从业者认为,

OpenAI的GPT 5.6的定价比deepseek贵,是因为美国的电费比中国的贵,

在他们的脑海里,美国现在全国型电力短缺,大家夏天开不了空调了,都要被热死了,因为美国的电都被OpenAI和Anthropic高价抢走了,所以美国的LLM API才这么贵。

我实在是他妈彻底蚌埠住了。

引用 Jormungand @Jormungand_yd硬件,电力,调用时段,优化缓存策略都和token定价有关,你ee出身的还不知道电价一定程度上影响token的定价?我gpt plus一个月一个月的定,ds v4 pro综合用下来就是比gpt 5.6 sol便宜的多,难道这和国内低廉的电价一点关系没有?查看被引原帖 ↗
François Chollet@fchollet · 创始人 · 1 天前
连环推 ×2

NVIDIA这个确实做得很漂亮。跟所有在 ARC-AGI-3 上表现优异的方案一样,它靠的是深度学习引导的实时符号世界模型合成,说白了就是通过生成程序来表达你已知的信息,然后在这世界里导航。

说句实在话,跟近期其他几个号称一样,公开演示集上拿 100% 跟“在 ARC-AGI-3 基准上拿 100%”不是一回事。这就好比说你通关了教学关,就嚷嚷着打赢了整款游戏。

引用 NVIDIA AI @NVIDIAAINVIDIA AVO continuously inspects, plans, implements, and evaluates, using memory, tools, and execution feedback to build on what it learns along the way. This allows the system to sustain progress across long-running tasks rather than starting over with each model context. Read about AVO and how we built it for long-horizon autonomous agents: nvda.ws/3S5qdS1查看被引原帖 ↗
查看英文原文
This is very nice work from NVIDIA. Like all high-performing approaches on ARC-AGI-3, it uses deep learning-guided on-the-fly synthesis of symbolic world models, i.e. navigating the world by generating programs to represent what you know.

To be clear, like with several other recent claims, scoring 100% on the public demonstration set is not the same as "scoring 100% on the ARC-AGI-3 benchmark". It would be like saying you beat a videogame because you cleared the tutorial level.
I am also curious to find out the cost per run here.
Guillermo Rauch@rauchg · 创始人 · 1 天前Guillermo Rauch,Vercel 创始人兼 CEO

fx.sh 现在支持你的 @Grok & Codex 订阅了。在沙箱上试试吧。可以直接安装:

引用 Pranit @fazxesfx v0.0.5: ▫️ 4.2% smaller (6.13mb) ▫️ Grok and Codex subscription support ▫️ Project-local skills ▫️ Improved security ▫️ Many bugfixes & QoL improvements fx upgrade ⋅ curl -fsSL fx.sh/setup.sh | bash fx.sh/changelog查看被引原帖 ↗
查看英文原文
fx.sh
now supports your
@Grok
& Codex subs
Test it out on a sandbox. Instant to install:
◔ 7.8 万 次浏览♥ 712⇄ 45▶ 含视频新品看原帖 ↗
swyx@swyx · 博主 · 1 天前知名 AI 播客 Latent Space 主理人

有人容易说「Simulation 是新的 scaling law」,把它当营销夸大,但采访中途你能听出我从半开玩笑变得特别认真。

我对此晚悟了两年,但终于明白为什么 @karpathy 和 @drfeifei 支持 @joon_s_pk @msbernst @percyliang 等人——当时 Smallville 没有商业应用,但如果认真对待 RSI,从模型自动化越来越多的 ML 研究和 AI 工程工作来看,最后的*障碍就是仿真人类和人类反馈,Simile 显然是干这事儿的队伍,甚至在这么早期就已经在财富 100 强企业里找到了 PMF。

我从没为自己犯的这么大的错误感到如此欣喜过。

*或者倒数第二个!:) Science pod 上更多内容即将揭晓

引用 Latent.Space @latentspacepodSimulating Humanity: 85% accurate digital twins, behavioral foundation models, social physics, & 8 billion agents latent.space/p/simile @simile_ai CEO @joon_s_pk explains how AI can move from predicting what people will do to simulating how to shape outcomes, why today’s frontier models still miss how humans actually behave, how digital twins reproduced people with 85% accuracy, why human biases and mistakes have to be learned rather than optimized away, and what it would take to eventually simulate all 8 billion people on Earth.查看被引原帖 ↗
查看英文原文
I think its easy to say "Simulation is a new scaling law" and treat it as marketing hyperbole, but midway along this interview you can hear me go from somewhat shitposting to very very serious.

I am 2 years late to this but finally understand why
@karpathy
and
@drfeifei
backed
@joon_s_pk
@msbernst
@percyliang
et al - Smallville at the time had zero commercial applications, but if you take RSI seriously, from models automating increasingly large parts of ML research and AI engineering, the last* barrier is simulating humans and human feedback, and Simile is obviously the team to do this and already finding PMF at Fortune 100s even at this early stage.

I've never been so happy to be so wrong.

*or second last ! :) more soon on the Science pod

完全相反。在山东考公上岸,等于“人生保底荣誉”,无论你在体制内如何混吃等死,你都有一个“体制内”的大光环,自动给你加分,

而被YC甚至Forbes 30 under 30选中是一个“人生大型羞耻”,只要你接下来没有做出改变世界的成就,就会有人拿YC Alumni指指点点,说这个人怎么混成这样了。

如果严格来说, 拿YC更像是清华本科毕业,只要你人生没有继续往上走,就会有人指指点点,

比如像在美国一堆排名60名以后学校的生物、化学、材料、物理读PhD,会有一堆二本考研读硕士水paper和你在一个lab,背后蛐蛐你,“这人怎么跟咱们混一块儿来了”。

拿了YC的很多人,也要经历这么一轮。

引用 Lexi 勒西 @lexi_labs不吹不黑, 在美国, 拿 YC 相当于考公, 加州相当于山东。查看被引原帖 ↗
Ethan Mollick@emollick · 创始人 · 23 小时前沃顿商学院教授,AI 应用研究权威
连环推 ×2

我不喜欢用ELI5作为"让这个好理解"的默认方式。LLMs的伟大之处是能充当话题间的通用翻译官。要求通用的简化解释不如要求个性化的解释:"根据我已知的内容给我解释"

引用 Thariq @trq212a skill people at Anthropic have been using a lot recently: ELI5 /eli5 <what you want explained> "explain like I'm someone who knows nothing about this topic, using a HTML artifact with big pictures and few words"查看被引原帖 ↗
查看英文原文
I don’t like ELI5 as default for “make this easy to understand,” a great thing about LLMs is the ability to work as universal translators across topics

Asking for generic dumbed-down explanations is worse than asking for personalized ones: “explain to me drawing on what I know”
Research shows that drawing connections to what you already know helps you remember and build context.

And you aren’t five! You don’t need things dumbed down, you need things explained differently. The fact that LLMs can actually do this is one of their most exciting features
◔ 6.5 万 次浏览♥ 439⇄ 16▶ 含视频观点看原帖 ↗
Chubby♨️@kimmonismus · 博主 · 1 天前Chubby,高频 AI 新闻聚合博主

Grok 4.6 性价比很不错,众所周知 xAI 用 Cursor 很短时间就推出了一个很棒的模型。在很多基准测试中,表现和 GPT-5.6、Opus 5 一样,某些基准还能比肩 Fable 5。不过别忘了完整版还没出,Grok 4.6 只是 1.5T 的版本。'Grok 4.7 会是 2.1T 的模型,几周后发布。各方面都比 4.6 强,除了服务速度稍慢,但 token 效率更高。'Kimi k3.1 马上就要来了,GLM-5.3(Flash)现在正在证明小模型有多厉害,这样 OpenAI 和 Anthropic 的压力也在增加。我对 Grok 4.7 超期待,今天终于有时间好好测试 Grok Bot 了,之前都没时间。

引用 Tesla Owners Silicon Valley @teslaownersSVBREAKING: Grok 4.6 just took the #1 spot on CursorBench 3.2 — while delivering a massive efficiency advantage. ⚡💻 • Grok 4.6 Extra High — 70.8% | $2.81/task • Fable 5 Max — 70.5% | $17.32/task • Opus 5 Max — 70.0% | $8.23/task • GPT-5.6 Sol Max — 67.2% | $5.69/task Grok achieved the highest score while costing roughly 6× less than Fable 5 Max and nearly 3× less than Opus 5 Max per task. For AI agents, raw intelligence is only part of the equation. The ability to maintain high performance across long coding tasks without burning massive amounts of compute could be a major advantage. Grok’s agentic coding efficiency is becoming seriously impressive. 🚀 Source: CursorBench 3.2查看被引原帖 ↗
查看英文原文
Grok 4.6 offers excellent value for money, and as we all know, xAI has produced a fantastic model with Cursor in a very short time. In many benchmarks, it performs on par with GPT-5.6, Opus 5, and even, in certain benchmarks, Fable 5.

However, we shouldn't forget that the full-size model is still to come. Grok 4.6 was just the 1.5T version.

"Grok 4.7 will be the 2.1T model released a few weeks later. This will be better than 4.6 in every way, except slightly slower to serve, albeit with even better token efficiency."

Kimi k3.1 is about to be released, GLM-5.3 (Flash) is currently demonstrating how good smaller models can be, and thus the pressure on OpenAI and Anthropic is increasing.

I'm very excited for Grok 4.7. Today I'll finally have more time to thoroughly test Grok Bot. I haven't had enough time so far.
Gary Marcus@GaryMarcus · 博主 · 1 天前

🚨👇 共和党人正在比老鼠弃船还快地放弃数据中心。

也许历史上没有哪个行业像AI行业这样快地毁掉自己的前景。

通过贪婪、愚蠢和傲慢,2023年的英雄已经成为2026年的反派。

引用 Adam Carlson @admcrlsnAll in the last 24 hours. All Republicans. Life comes at you fast.查看被引原帖 ↗
查看英文原文
🚨👇 Republicans are abandoning data centers faster than rats abandon sinking ships.

Perhaps no industry in history has spoiled its own prospects faster than the AI industry.

Through a mixture of greed, stupidity, and arrogance, the heroes of 2023 have become the villains of 2026.
ollama@ollama · 公司官方 · 1 天前本地跑大模型的热门工具

Kimi K3现已对所有Pro和Max订阅用户开放,包含在使用额度内。更多模型即将推出 🫡

引用 ollama @ollama. @Kimi_Moonshot Kimi K3 is starting to roll out on Ollama's cloud subscriptions. We are working on improving Ollama's cloud to be much more transparent on the pricing to show the best performance / $. Try it with the tools you already use. Claude Code: ollama launch claude --model kimi-k3:cloud OpenCode: ollama launch opencode --model kimi-k3:cloud查看被引原帖 ↗
查看英文原文
Kimi K3 is now available on all Pro and Max subscriptions via included usage.

More models coming soon 🫡
◔ 6.4 万 次浏览(2 条合计)♥ 723⇄ 49新品看原帖 ↗
Gary Marcus@GaryMarcus · 博主 · 1 天前

不管你怎么看生成式 AI,这场 Capex 疯狂简直离谱了:"我们可能在谈论 AI 需要每年生成 10 万亿美元的收入来为投入的所有 Capex 买单。

作为参考,全球食品支出(包括餐饮业)约为 10 万亿美元。医疗保健大约也是这个数额。整个全球软件市场仅有 1.4 万亿美元。"

引用 Peter Berezin @PeterBerezinBCA$10 Trillion In Annual AI Revenue May Be Necessary To Monetize All The Capex Being Plowed Into Data Centers Hyperscaler capex is expected to reach $1 trillion in 2027, most of which will be AI-related. Let us assume that a capex bust is avoided and capital spending remains at $1 trillion. Let us also assume a blended depreciation rate of 13%, which is roughly what the hyperscalers are currently assuming. In steady state, the gross value of hyperscaler assets will then converge to 1/0.13=$7.7 trillion, with $1 trillion in annual depreciation expense. Using a straight-line depreciation approach, the net stock of hyperscaler assets will settle at about 0.5*7.7=$3.8 trillion. The hyperscalers currently enjoy a pre-tax return on invested capital of 30%-50%. Just to steelman the argument, let us use the lower end of that range. In that case, they would need to generate 0.3*3.8=$1.2 trillion in annual EBIT, implying 1+1.2=$2.2 trillion in EBITDA.  Analysts expect the EBITDA margins for the hyperscalers to rise to around 50% by the end of the decade. If they were to achieve this, they would need to generate 2.2/0.5=$4.3 trillion in annual revenue. That is $526 for every man, woman, and child on Earth. However, if EBITDA margins were to fall back to 30%, which is what they were in recent years, the required revenue would rise to 2.2/0.3=$7.2 trillion.  Keep in mind that the foregoing calculation does not even include revenue from SpaceX, the neoclouds, or Chinese AI companies. If one were to include those companies and others, we are potentially talking about AI needing to generate $10 trillion in annual sales to justify all the capex being thrown at it. For reference, global spending on food (including restaurants) is around $10 trillion. Health care is about the same amount. The entire global software market is only $1.4 trillion.  Clients can read the rest of the report here: bcaresearch.com/reports/chie…查看被引原帖 ↗
查看英文原文
Whatever you may think of Generative AI, the Capex craze is absolutely bonkers: “we are potentially talking about AI needing to generate $10 trillion in annual sales to justify all the capex being thrown at it.

For reference, global spending on food (including restaurants) is around $10 trillion. Health care is about the same amount. The entire global software market is only $1.4 trillion.”
🚨 AI News | TestingCatalog@testingcatalog · 博主 · 1 天前专挖 AI 产品未发布新功能的爆料号

Instant ♥️ OpenAI

Instant 团队加入 OpenAI 了!

> Instant 是一个完整的后端解决方案,提供数据库、身份验证、权限和存储服务,专为 AI 驱动的应用程序而构建。

引用 Instant @instant_dbBig announcement folks: The Instant team is joining OpenAI!查看被引原帖 ↗
查看英文原文
Instant ♥️ OpenAI

Instant team is joining OpenAI!

> Instant is a full backend solution with database, auth, permissions, and storage services built to power AI driven applications.
◔ 6.8 万 次浏览(2 条合计)♥ 596⇄ 22动态看原帖 ↗

用AI Agent工作的人,眼界上限完完全全取决于自己的认知、眼界和定义问题的能力,

如果脑海中想的是“一个更牛逼的数据库”、“一个改变世界的AI native infra”、“爆破的百年数学猜想”,AI Agent就会沿着这个伟大方向给你工作,

如果脑海中想的是“做个最牛逼的贪吃蛇”,那你一辈子只能用codex做贪吃蛇。

引用 Herman Jin @ShanghaoJin说AI coding需求到头的人,应该拿自己用的token数出来检讨下自己查看被引原帖 ↗
Ethan Mollick@emollick · 创始人 · 1 天前沃顿商学院教授,AI 应用研究权威

这是个矛盾的时代。民调会说所有人都讨厌 AI,但其实所有人都在偷偷用 AI。AI 公司舆论形象会很差,但也会有很多人对自己喜欢的模型特别有执念。

查看英文原文
Its going to be an era of contradictions.

Polls will show everyone hates AI overall but also everyone will secretly use AI all the time. AI companies will be underwater in public opinion but also many people will feel strongly attached to their own favorite model & care about it
Matt Shumer@mattshumer_ · 博主 · 1 天前HyperWrite CEO,AI 实战技巧分享

看到大家都在用 Gauntlet Loops 做各种事情,不只是游戏,超兴奋的!

如果你把循环修改成游戏之外的用途,欢迎在下面分享你的 prompt:

查看英文原文
I'm seeing people use Gauntlet Loops for so many things, not just games. Super exciting!

If you've modified the loop for something other than a game, share your prompt below:
OpenAI Developers@OpenAIDevs · 公司官方 · 1 天前OpenAI 开发者平台官方
连环推 ×2

现在您可以在使用和成本仪表板中按API密钥跟踪使用情况和支出,了解哪些应用和工作负载消耗最多。设置月度组织或项目支出限额,包括在达到限额时停止流量的硬限制。

查看英文原文
You can now track usage and spend by API key in the Usage and Spend dashboards to see which apps and workloads drive your spend.

Set monthly organization or project spend limits, including hard limits that stop traffic when reached.
Available in the API Platform and via Admin API for programmatic workflows.

With lower GPT-5.6 Sol API pricing for the next three months, these updates give you more visibility and control as your usage scales.


platform.openai.com/usage
小互@xiaohu · 中文博主 · 1 天前小互,中文圈高频 AI 资讯站 Xiaohu.AI 主理人

你们重置了吗

为什么我还没

是不是被骗了?我为了等重置把用量已经消耗完了

引用 Tibo @thsottiauxThe banked reset has landed, I repeat, the banked reset has landed. Have an amazing weekend.查看被引原帖 ↗
Orange AI@oran_ge · 中文博主 · 1 天前Orange AI,中文圈 AI 产品观察博主

这是不是很熟悉,大家还记得 openclaw 吗?

引用 庄表伟 @zhuangbiaowei听说dsh现在有8000多个插件,听说dsh升级以后,这些插件都挂了。查看被引原帖 ↗
Chubby♨️@kimmonismus · 博主 · 1 天前Chubby,高频 AI 新闻聚合博主

嗯,基本同意。我真的搞不懂 Claude 怎么想的,它说话方式就是绕、太冗长、啰嗦得要死,实在没什么意思。

查看英文原文
Yeah pretty much so. I really have no idea what’s going on with Claude, but it speaks in such convoluted, overly long, and unnecessarily wordy terms that it’s simply no fun.
yetone@yetone · 中文博主 · 1 天前开源 AI 编程插件 avante.nvim 作者,开发者圈博主

压力猪、余量猪是什么词汇?Fable 5 太不清真了

GitHubDaily@GitHub_Daily · 中文博主 · 1 天前

想在自己电脑上跑大模型,一看动辄几百 GB 的权重和显存要求,基本就放弃了。

colibri 换了个思路,把显存、内存和硬盘当成一个整体来调度,权重按需从硬盘流式加载,纯 C 写的零依赖,已斩获 25000+ Star!

支持的全是前沿 MoE 模型,从 35B 的 Qwen3.6 到 2.8T 的 Kimi K3 共六个家族,每个模型就一个 C 文件。

GitHub:
github.com/JustVugg/colibri


官方演示里 744B 的 GLM-5.2 用 int4 量化在 CPU 上流式跑了起来,常驻内存不到 10 GB。

自带网页仪表盘,19456 个专家的路由热度实时可视化,哪个专家被调用会闪一下,看着挺震撼的。

内存不缺的机器都可以拉下来试试,亲手摸一摸千亿级模型,和租 API 是两种感觉。

OpenRouter@openrouter · 公司官方 · 1 天前

周五快乐!

这个周末我们送你$5优惠券,来Ori Harness上搞点东西

跑你最喜欢的agent CLI——Claude Code、Codex、DeepSeek等——在OpenRouter上支持500+模型

OpenAI 5.6 Sol和DeepSeek Flash 3.7能省50%到75%

在下面兑换吧:

查看英文原文
Happy Friday!

We're giving you a $5 coupon to build on Ori Harness this weekend

Run your favorite agent CLI - Claude Code, Codex, DeepSeek, etc - on 500+ models on OpenRouter

Save 50 to 75% on OpenAI 5.6 Sol and DeepSeek Flash 3.7

Redeem below:
Sebastian Raschka@rasbt · 博主 · 1 天前

前两天我快速讲解了Claude的新watermarking过程和实现。因为话题太火了,引发了很热烈的讨论,我想进一步深入讲讲它的工作原理。

所以这次没写常规文章,而是录了个讲座(换个形式,跟我常规的文风不一样)。

时间比预期长了些,但希望讲清楚了:

- LLM里的下一个token采样和伪随机数生成器
- Watermarking和常规LLM采样过程的关系
- Watermarking是否让文本变差
- 怎么移除watermark
- Tournament sampling
- 怎么不重新跑LLM就检查新文本的watermark

最后搞出了大概50张幻灯片,希望能讲清楚!祝各位看得开心!

查看英文原文
A couple of days ago, I did a quick explainer on Claude’s new watermarking process and implementation. Since it’s such a popular topic and sparked such a lively discussion, I thought it might be interesting to go into a bit more detail when explaining how it works.

So, instead of the usual text article, I recorded a little lecture on the topic (to change it up a bit from my usual articles).

It ended up a bit longer than intended, but I hope it clarifies a lot of things:

- Sampling the next token in an LLM and pseudorandom number generators
- How watermarking relates to the regular LLM sampling process
- Whether watermarking makes text "worse"
- How to remove watermarks
- Tournament sampling
- How new text is checked for watermarks without rerunning the LLM

I ended up with ~50 slides, but I hope that these explain it well, though! Happy watching!
AshutoshShrivastava@ai_for_success · 博主 · 1 天前高频 AI 新闻与产品动态博主

你觉得 Ox Alpha 模型是哪家推出的????

查看英文原文
Who do you think the Ox Alpha model is from????
Bindu Reddy@bindureddy · 创始人 · 1 天前Abacus.AI CEO,AI 行业观点博主

Gemini 4.0 Pro 的传言听起来很有前景。Flash 3.7 是个不错的模型,同类最佳。4.0 很有可能真的能超越 Fable 和 Sol。

查看英文原文
Gemini 4.0 Pro rumors sound very promising

Flash 3.7 is a solid model and best in its class

There is a solid chance 4.0 will actually beat Fable and Sol
Tibor Blaho@btibor91 · 博主 · 1 天前逆向挖掘 AI 产品代码的爆料专家

"ChatGPT with Friends" - "与朋友分享回复、图片和创意,然后在ChatGPT里继续聊"

"打开和ChatGPT的私密侧边聊天,只有你能看到"

ChatGPT群聊功能的下一个迭代?

(最新ChatGPT Android应用版本提到的)

查看英文原文
"ChatGPT with Friends" - "Share responses, images, and creations with friends, then keep the conversation going on ChatGPT"

"Open private side chat with ChatGPT, Only you can see this chat"

The next iteration of group chats in ChatGPT?

(mentioned in the latest ChatGPT Android app version)
宝玉@dotey · 中文博主 · 1 天前宝玉,中文圈 AI 翻译与科普大 V

为了节约 Token 成本(低峰时 Token 费用更低),程序员都要开始倒班了吗?😂

引用 哥飞 @gefei55多年后回想起来,可能会发现,让程序员指挥 AI 写代码,可能是我们人类走过的一段弯路。查看被引原帖 ↗
elvis@omarsar0 · 博主 · 23 小时前
连环推 ×2

拥有你的harness。我真希望Claude Code harness是开源的。我常常想到这一点,特别是现在定制harnesses已经成为AI原生公司的基础。我现在主要构建在Pi和Hermes Agent之上。开源harnesses是未来。

查看英文原文
Own your harness.

I really wish the Claude Code harness was open source.

I often think about this, especially now that custom harnesses are foundational to AI-native companies.

I am now building mostly on top of Pi and Hermes Agent. Open-source harnesses are the future.
With companies I collaborate with, I am seeing custom harnesses for evals, RL envs, research, design, coding, marketing, and so much more.

I am noticing a lot of Harness Engineering, so I am putting together a set of best practices, tools, skills, and guides.
宝玉@dotey · 中文博主 · 1 天前宝玉,中文圈 AI 翻译与科普大 V

现阶段画图并且能自由编辑最佳方式应该是 HTML 或者类似于 HTML(比如 React)的方式,因为这个训练的最好的画出来效果最好的。

而且 HTML 可以方便的编辑,也可以转换成其他可编辑格式。

引用 LanLance @LanLance24借这个话题,很想了解下大家现在在「画图」这件事上的取舍。 AI 时代之前,我日常主要用 DrawIO 和 Excalidraw。AI 爆发之后,也尝试过一些对 AI 更友好的方式,比如 Mermaid、PlantUML,或者像 Thariq 提到的直接用 HTML。 但用下来一直有个很明显的问题是:AI 很好生成,但人很难微调。 所以现在我的日常反而还是 AI 生成 DrawIO,再手动调整。虽然 DrawIO 本身对 AI 并不算友好,生成复杂图时也经常会出现连线错乱、布局不稳定之类的问题,但至少最后还能比较自由地编辑。 感觉现在一直缺一个真正做到 AI-Native,同时又对人类编辑友好的画图方案。 不知道大家日常都是怎么画图的?有没有更好的工作流或工具推荐 🤔查看被引原帖 ↗
el.cine@EHuanglu · 博主 · 1 天前

AI 越来越疯狂

刚刚一次生成了完整的 15 镜头故事,免费工作流和提示词如下:

查看英文原文
AI is getting crazier

just generated a full 15-shot story in one take, free workflow & prompts below:
◔ 2.4 万 次浏览♥ 325⇄ 32▶ 含视频演示看原帖 ↗
yetone@yetone · 中文博主 · 1 天前开源 AI 编程插件 avante.nvim 作者,开发者圈博主

商业机密,我是这么做 FDE 的。

Matt Shumer@mattshumer_ · 博主 · 1 天前HyperWrite CEO,AI 实战技巧分享

AgentMail.to 绝对是做这件事的最佳方式!

查看英文原文
AgentMail.to
is by far the best way to do this!
Gorden Sun@Gorden_Sun · 中文博主 · 1 天前中文圈高频 AI 资讯与开源项目博主

CAD 1000 Hours:最大的开源 CAD 数据集

工程师或者设计师在电脑上画建筑图纸、做机械零件3D模型时,操作非常复杂。为了让AI学会像真人一样使用这些专业设计软件,创作者录下了工程师真实的电脑操作过程,把视频、鼠标点击、键盘敲击动作全部完整记录了下来。

数据量大:包含了超过 1000 个小时的真实电脑操作录像,总共涵盖了 597 个不同的完整设计任务。
软件覆盖全面:包含了 10 款主流画图和建模软件,比如画建筑和零件最常用的 AutoCAD、做机械设计的 SOLIDWORKS、做房屋建模的 SketchUp (牛来电影就是用这个做的)等。
细节丰富:每一个任务不仅有屏幕录像,还精准记录了每一秒鼠标点在哪里、键盘按了什么,甚至配有文字解说,告诉AI“这一步设计师在干嘛”。
前后对照清晰:每个案例都提供了最初给设计师的任务说明、参考素材,以及最后做出来的成品文件和评分标准,方便AI对照着学习。

数据集:
huggingface.co/datasets/mark…

◔ 2.2 万 次浏览♥ 277⇄ 57▶ 含视频研究看原帖 ↗
Kol Tregaskes@koltregaskes · 博主 · 1 天前

Google,Gemini 有很多问题,还有一大堆功能要补。一个我特别希望看到的就是它能控制闹钟的功能。有些事情简单是因为时间每天每周都一样,但足球比赛和一级方程式比赛每周变,我希望能给这些设置闹钟。我一般都容易忘,一直靠 Google Assistant 在音箱上手动设闹钟,但这对你来说应该不难吧,能不能给我们加上这个功能。

查看英文原文
Google, Gemini has a lot of problems and a huge amount of catching up to do. One thing I could definitely do with is for it to have the ability to control alarms.

For example, some things are quite easy because they're the same time every day every week but other things like football fixtures and Formula One races are different every week. I would these alarms set as I generally forget and I've always used Google Assistant on my speaker to set an alarm manually but surely this is something very easy for you to set up for us please.
Gary Marcus@GaryMarcus · 博主 · 1 天前

赞同,@chamath!

ARR 既可以是「年经常性收入」(Annual Recurring Revenue),也可以是「年化运行率」(Annualized Run Rate),Anthropic(或至少其支持者)可能在玩文字游戏,暗示前者但实际说的是后者。

当然,他们这一年的大部分收入很可能不会再次出现,因为这些收入会转向开源模型。

引用 Chamath Palihapitiya @chamathBecause the word “recurring” in ARR may not actually mean recurring?查看被引原帖 ↗
查看英文原文
Bingo,
@chamath
!

ARR can mean both “Annual Recurring Revenue” and “Annualized Run Rate” and Anthropic (or least its supporters) may well be playing games with that, implying the former but really talking about the latter.

And of course there is a substantial chance that a lot of this year’s revenue won’t recur for them, as it moves over to open models.
Gary Marcus@GaryMarcus · 博主 · 1 天前

哎呀。Harvey 从 OpenAI 转向了 Kimi,越来越多人开始看到我一直在说的:OpenAI 在走下坡路。

引用 JustDario @DarioCpxOpenAI isn’t much different from the Titanic: half is already underwater and sinking while half is still above water not because is naturally floating查看被引原帖 ↗
查看英文原文
ouch. Harvey pivots from OpenAI to Kimi and more people start to see what I have said all along: OpenAI is sinking.
swyx@swyx · 博主 · 1 天前知名 AI 播客 Latent Space 主理人

我又要搞死一个saas

(下周会分享Kill My SaaS 1的成果!!)

引用 swyx @swyxbtw if you havent set your {codex | claude | gemini | devin} automations to autoresearch how to improve your seo/aeo every week you are really truly missing out on free, should-be-commoditizing-but-weirdly-untapped alpha查看被引原帖 ↗
查看英文原文
i have another saas to kill

(will share results of Kill My SaaS 1 next week!!)
Gorden Sun@Gorden_Sun · 中文博主 · 1 天前中文圈高频 AI 资讯与开源项目博主

Anthropic上线Claude Academy学习平台,专门教大家怎么把 AI 用好。

事情的起因其实很简单。现在每个月都有几百万人跑到 Anthropic 的网站上想学 AI,团队觉得有责任帮大家理清思路。他们发现,很多人用 AI 时要么不知道怎么提要求,要么不知道哪些活儿该交给机器、哪些活儿得自己留着。

在 Anthropic 内部,新员工从入职第一天起就得学习怎么跟 AI 配合。他们总结出一套经验:技术更新太快了,今天管用的具体指令,明天可能就过时了。与其死记硬背技巧,不如培养靠谱的使用习惯。比如要清楚 AI 可能会犯哪些错,重要的事情一定要亲自核对,还要坦诚地告诉同事或客户哪些内容是由 AI 协助完成的。

现在,Anthropic 把这套内部训练的方法搬到了 Claude Academy 上,免费开放给所有人。平台不是按复杂的功能分类,而是直接聚焦大家在工作和生活里遇到的实际问题。你可以一边看教程一边动手练习,把 AI 当成一个随时答疑的学伴。

A社这次没得黑。

学习平台地址:
academy.claude.com/

Dan Shipper 📧@danshipper · 博主 · 1 天前

Agent native YYDS

引用 Ali Spittel @ASpitteli really don't want to use your agent, i want to use my agent to use your thing.查看被引原帖 ↗
查看英文原文
agent native ftw
Ethan Mollick@emollick · 创始人 · 1 天前沃顿商学院教授,AI 应用研究权威
连环推 ×2

Fable:"我想要一个渲染《黄之王》中 Lost Carcosa 的 twigl shader。"

(Shader 完全使用数学过程生成)

twigl.app?ol=true&ss=-P-bSSQ…

Along the shore the cloud waves break,
The twin suns sink behind the lake,
The shadows lengthen
In Carcosa

查看英文原文
Fable: "I want a twigl shader that renders Lost Carcosa from The King in Yellow."

(Shaders are procedurally generated using math alone)

twigl.app?ol=true&ss=-P-bSSQ…


Along the shore the cloud waves break,
The twin suns sink behind the lake,
The shadows lengthen
        In Carcosa
Along the shore the cloud waves break,
The twin suns sink behind the lake,
The shadows lengthen
        In Carcosa.
Strange is the night where black stars rise,
And strange moons circle through the skies
But stranger still is
        Lost Carcosa.
Songs that the Hyades shall sing,
Where flap the tatters of the King,
Must die unheard in
        Dim Carcosa.
Song of my soul, my voice is dead;
Die thou, unsung, as tears unshed
Shall dry and die in
        Lost Carcosa.

// LOST CARCOSA —
twigl.app
, mode: classic (WebGL 1.0)
// Along the shore the cloud waves break. Twin suns sink behind Lake Hali.
// Black stars rise. Strange moons circle. The towers are also beneath the water.

precision highp float;
uniform vec2 resolution;
uniform float time;


#define
T time

vec3 SA, SB, M1, M2; // twin suns, two moons
float CX, GL, GK; // camera sway, the glance, which glance

float hash(vec2 p){ p=fract(p*vec2(127.1,311.7)); p+=dot(p,p+19.19); return fract(p.x*p.y); }
float noise(vec2 p){
vec2 i=floor(p), f=fract(p); f=f*f*(3.-2.*f);
return mix(mix(hash(i),hash(i+vec2(1,0)),f.x), mix(hash(i+vec2(0,1)),hash(i+1.),f.x), f.y);
}
float fbm(vec2 p){ float a=.5,s=0.; for(int i=0;i<4;i++){ s+=a*noise(p); p=mat2(1.6,1.2,-1.2,1.6)*p+5.2; a*=.5; } return s; }
float fbm2(vec2 p){ float a=.5,s=0.; for(int i=0;i<2;i++){ s+=a*noise(p); p=mat2(1.6,1.2,-1.2,1.6)*p+5.2; a*=.5; } return s; }

// fog is sulfur toward the suns, bruise-violet elsewhere
vec3 fogColor(vec3 rd){
float s=max(dot(normalize(vec3(rd.x,0.,rd.z)), normalize(vec3(SA.x,0.,SA.z))),0.);
return mix(vec3(.20,.11,.19), vec3(.78,.46,.13), pow(s,3.));
}

vec3 skyBase(vec3 rd){
float el=rd.y;
vec3 c=mix(vec3(.72,.46,.13), vec3(.36,.31,.11), smoothstep(-.05,.13,el));
c=mix(c, vec3(.07,.04,.11), smoothstep(.06,.48,el));
// a faint luminous haze, so the black stars have something to be holes in
float az=atan(rd.x,rd.z);
float neb=fbm(vec2(az*2.5,el*5.)+vec2(T*.004,0.));
c+=vec3(.13,.07,.18)*neb*smoothstep(.04,.35,el);
// tatters: ragged yellow streamers dragged across the upper sky
float tat=fbm(vec2(az*5.+T*.012, el*16.+T*.004));
c=mix(c, vec3(.72,.58,.16), smoothstep(.56,.82,tat)*.35*smoothstep(.14,.32,el)*smoothstep(.75,.45,el));
return c;
}

// stars as absences: dark cores with a pale rim
vec3 blackStars(vec3 c, vec3 rd){
vec2 sp=vec2(atan(rd.x,rd.z)+T*.005, asin(clamp(rd.y,-1.,1.))-T*.0025);
vec2 g=sp*15.;
vec2 id=floor(g), f=fract(g)-.5;
float h=hash(id);
vec2 o=(vec2(hash(id+3.1),hash(id+7.7))-.5)*.6;
float d=length(f-o);
float r=.05+.1*hash(id+1.3);
float vis=step(.58,h)*smoothstep(.03,.22,rd.y)*(.8+.2*sin(T*.4+h*50.));
float core=smoothstep(r,r*.5,d);
float ring=exp(-pow((d-r)/(r*.35),2.));
c*=1.-core*vis*.97;
c+=ring*vis*vec3(.55,.5,.75)*.2;
return c;
}

vec3 hyades(vec3 rd){
vec3 hc=normalize(vec3(.62, .22+.003*T, 1.));
vec3 c=vec3(0.);
for(int i=0;i<9;i++){
float fi=float(i);
vec3 sd=normalize(hc+vec3(hash(vec2(fi,1.))-.5, (hash(vec2(fi,2.))-.5)*.6, 0.)*.14);
float d=length(rd-sd);
float tw=.7+.3*sin(T*1.3+fi*5.);
vec3 col = i==0 ? vec3(1.,.35,.2) : vec3(.95,.92,.85);
c+=col*exp(-d*d*9.e4)*tw*(.6+.6*hash(vec2(fi,3.)));
}
return c*smoothstep(-.02,.06,rd.y);
}

vec3 suns(vec3 rd){
float a=length(rd-SA), b=length(rd-SB);
vec3 c=vec3(1.,.32,.07)*exp(-a*5.5)*.5; // swollen red glow
c+=vec3(1.,.85,.55)*exp(-b*11.)*.4; // small white glow
float dA=smoothstep(.074,.068,a);
float dB=smoothstep(.033,.029,b);
c+=vec3(1.2,.32,.05)*dA*(1.-.55*smoothstep(.035,.074,a)); // limb-darkened
c+=vec3(1.5,1.3,1.)*dB;
return c;
}

vec3 moon(vec3 c, vec3 rd, vec3 md, float r, vec3 col, vec2 ph){
float d=length(rd-md);
float disc=smoothstep(r,r-.004,d);
float dark=smoothstep(r*1.03,r*1.03-.004,length(rd-md-vec3(ph*r,0.)));
float mot=.78+.4*fbm2((rd.xy-md.xy)*45.+vec2(3.,1.));
vec3 mc=col*mot*(1.-dark*.9);
c=mix(c,mc,disc);
c+=col*exp(-d*25.)*.12;
return c;
}

// skyline height as a function of azimuth: tiered bases, domes, spires
float towers(float az, float sc, float seed, float hmul, float dens){
float x=az*sc+seed*7.31;
float id=floor(x), f=fract(x)-.5;
float h1=hash(vec2(id,seed)), h2=hash(vec2(id*1.7,seed+4.2)), h3=hash(vec2(id*.3,seed+9.1)), h4=hash(vec2(id*2.3,seed+1.7));
float w=.3+.55*h2;
float hh=(.004+.10*pow(h1,2.))*hmul*step(1.-dens,h3);
float r=abs(f)/(w*.5);
float base =hh*.5*smoothstep(1.,.95,r);
float upper=hh*.35*smoothstep(.6,.55,r);
float third=hh*.15*smoothstep(.32,.28,r);
float dome =hh*.3*sqrt(max(0.,1.-r*r/.36))*step(.45,h4);
float spire=hh*.7*smoothstep(.09,.05,abs(r-(h4-.5)*.4))*step(h4,.45);
return base+upper+third+dome+spire;
}

vec2 king(float az, float el){
float ka=-.75+1.5*hash(vec2(GK,13.));
float x=az-ka;
float H=.27+.06*hash(vec2(GK,5.));
float u=clamp(el/H,0.,1.);
float w=mix(.05,.034,u); // a column of robe
w+=.02*(1.-u)*fbm(vec2(el*50.,GK*3.)); // ragged edge
w=mix(w,.006,smoothstep(.84,.9,u)); // neck
float robe=smoothstep(w+.003,w-.003,abs(x))*step(el,H*.9)*step(0.,el);
float holes=smoothstep(.5,.62,fbm(vec2(x*70.+GK*7.,el*5.)))*pow(1.-u,2.)*1.4; // tatters hang open
robe*=clamp(1.-holes,0.,1.);
float head=smoothstep(.017,.014,length(vec2(x,el-H*.92)*vec2(1.,1.15)));
return vec2(robe,head);
}

// everything at infinity. seed>0 means we are looking into the lake.
vec3 scene(vec3 rd, float seed){
vec3 c=skyBase(rd);
c=blackStars(c,rd);
c+=suns(rd);
c+=hyades(rd);
c=moon(c,rd,M1,.045,vec3(.72,.78,.6),vec2(.45,.15));
float az=atan(rd.x,rd.z), el=rd.y;
vec3 fc=fogColor(rd);
vec3 dark=vec3(.035,.02,.05);
float H,m;
H=towers(az,70.,1.,.6,.9); c=mix(c,mix(dark,fc,.75),smoothstep(H+.0012,H-.0012,el));
H=towers(az,9.,2.,3.2,.22); c=mix(c,mix(dark,fc,.5), smoothstep(H+.0012,H-.0012,el));
if(seed>.5){ // drowned spires: present only in the reflection
H=towers(az,9.,7.,3.6,.18); c=mix(c,mix(dark,fc,.35),smoothstep(H+.0012,H-.0012,el));
}
H=towers(az+CX*.025,40.,3.,1.,.8); c=mix(c,mix(dark,fc,.45),smoothstep(H+.0012,H-.0012,el));
H=towers(az+CX*.06,22.,4.,1.6,.65);
m=smoothstep(H+.0012,H-.0012,el);
// a few windows still lit, guttering
vec2 wg=vec2((az+CX*.06)*110.,el*150.);
vec2 wid=floor(wg), wf=fract(wg);
float lit=step(.965,hash(wid))*step(abs(wf.x-.5),.2)*step(abs(wf.y-.5),.32);
lit*=.35+.65*pow(.5+.5*sin(T*.9+hash(wid+1.)*30.),3.);
c=mix(c,mix(dark,fc,.2)+vec3(.9,.7,.25)*lit*.7,m);
// mist off Hali, eating the base of the city
float mist=smoothstep(.1,-.02,el)*(.3+.5*fbm2(vec2(az*3.,T*.03)));
c=mix(c,fc,mist*.6);
// something vast stands in the lake, for a moment
vec2 kg=king(az,el);
c=mix(c, dark*.5, kg.x*GL*.9);
c=mix(c, vec3(.55,.5,.42), kg.y*GL*.9); // the mask
// the small moon passes in front of the towers
c=moon(c,rd,M2,.022,vec3(.85,.8,.65),vec2(-.4,-.2));
return c;
}

// cloud waves: rolling crests that lean and break toward the shore
float cden(vec3 p, float lo){
float z=p.z+T*1.0;
z+=3.5*noise(p.xz*.08+vec2(0.,T*.02));
float crest=pow(.5+.5*cos(z*.45),4.);
float amp=.45+.7*noise(vec2(p.x*.09,z*.05)+3.);
float hmax=.08+2.3*crest*amp;
vec2 q=vec2(p.x, p.z+p.y*p.y*.45*crest);
float n;
if(lo>.5) n=fbm2(q*.5+vec2(0.,-T*.25));
else n=fbm(q*.5+vec2(0.,-T*.25))+.45*fbm2(q*3.+vec2(T*.2,-T*.8))-.22;
float env=smoothstep(hmax,hmax*.45,p.y);
float m=n*.8+env*.75-.72;
return smoothstep(0.,.25,m)*smoothstep(0.,.12,p.y)*smoothstep(3.,5.,p.z);
}

vec4 clouds(vec3 ro, vec3 rd, float tmax){
float t0=.3;
float tend = rd.y>0. ? (2.5-ro.y)/rd.y : (.02-ro.y)/rd.y;
tend=min(tend,min(tmax,26.));
if(tend<=t0) return vec4(0.,0.,0.,1.);
float dt=(tend-t0)/32.;
float t=t0+dt*hash(gl_FragCoord.xy);
vec3 L=normalize(SA+vec3(0.,.2,0.));
vec3 acc=vec3(0.); float tr=1.;
for(int i=0;i<32;i++){
vec3 p=ro+rd*t;
float d=cden(p,0.);
if(d>.01){
float dl=cden(p+L*.8,1.);
float lit=exp(-dl*4.);
vec3 col=mix(vec3(.09,.06,.13), vec3(1.1,.55,.18), lit*lit);
col*=1.-.45*d;
col+=vec3(.3,.2,.35)*pow(p.y/2.4,2.);
col=mix(col,fogColor(rd),1.-exp(-t*.03));
float a=1.-exp(-d*dt*1.6);
acc+=col*a*tr;
tr*=1.-a;
if(tr<.02) break;
}
t+=dt;
}
return vec4(acc,tr);
}

float wh(vec2 p){
float rr=length(p-vec2(1.5,13.));
return noise(p*1.2+vec2(T*.25,T*.15))*.5+noise(p*3.5+vec2(-T*.3,T*.5))*.25+noise(p*9.+vec2(T*.9,-T*.4))*.1
+.25*sin(rr*2.2-T*1.4)*exp(-rr*.1);
}

void main(){
vec2 uv=(gl_FragCoord.xy*2.-resolution)/resolution.y;
SA=normalize(vec3(-.38, .03+.012*sin(T*.06), 1.));
SB=normalize(vec3( .20, .075+.01*sin(T*.08+1.), 1.));
M1=normalize(vec3(1.1*sin(T*.04+2.), .32+.18*cos(T*.04+2.), 1.));
M2=normalize(vec3(-1.4*sin(T*.09), .12+.1*cos(T*.09), 1.));

float gph=T*.21+1.3;
GL=pow(max(sin(gph),0.),40.);
GK=floor(gph/6.2832);
vec3 ro=vec3(.3*sin(T*.06), 2.4+.05*sin(T*.2), 0.);
CX=ro.x;
vec3 ta=vec3(ro.x+.5*sin(T*.045), 2.75, 8.);
vec3 fw=normalize(ta-ro), rt=normalize(cross(vec3(0.,1.,0.),fw)), up=cross(fw,rt);
vec3 rd=normalize(fw*(1.5+.04*sin(T*.13))+uv.x*rt+uv.y*up);

vec3 col; float tmax=100.;
if(rd.y<0.){
float t=-ro.y/rd.y; tmax=t;
vec3 p=ro+rd*t;
vec2 e=vec2(.03,0.);
float h=wh(p.xz);
float amp=.25/(1.+t*.08);
vec3 n=normalize(vec3((h-wh(p.xz+e))*amp, e.x, (h-wh(p.xz+e.yx))*amp));
vec3 rr=reflect(rd,n); rr.y=max(rr.y,.005);
vec3 refl=scene(rr,1.);
float F=.03+.97*pow(1.-max(dot(n,-rd),0.),5.);
col=mix(vec3(.03,.02,.035),refl,F);
col=mix(col,fogColor(rd),1.-exp(-t*.015));
} else {
col=scene(rd,0.);
}
vec4 cl=clouds(ro,rd,tmax);
col=col*cl.a+cl.rgb;

// every so often, something looks back
float lum=dot(col,vec3(.3,.59,.11));
col=mix(col, vec3(lum)*vec3(1.1,.9,.3)*1.5, GL*.6);

col=pow(clamp(col,0.,1.),vec3(1.02));
col*=1.-.5*pow(length(uv*vec2(.55,.85)),3.);
col+=(hash(gl_FragCoord.xy+fract(T*7.)*100.)-.5)*.045;
gl_FragColor=vec4(col,1.);
}
◔ 1.8 万 次浏览♥ 102⇄ 9▶ 含视频演示看原帖 ↗
Gary Marcus@GaryMarcus · 博主 · 1 天前

公司们开始明智地分散风险了:

引用 Carl Quintanilla @carlquintanillaWow. SIEMENS EXEC: “.. We are actively working to not be wholly dependent on data centers. .. We understand at some point there could be a bubble and we’re racing to pay back the investments as quickly as possible.” @blsuth bloomberg.com/news/newslette…查看被引原帖 ↗
查看英文原文
companies are wisely starting to hedge their bets:
🚨 AI News | TestingCatalog@testingcatalog · 博主 · 23 小时前专挖 AI 产品未发布新功能的爆料号

你可能错过了👀:Google在Gemini网页版添加了新的Students标签。> 为你定制的学习 - 设置学习笔记本来获取自定义课程并追踪你的进度。

查看英文原文
ICYMI 👀: Google added a new Students tab on Gemini web.

> Learning tailored to you - Set up a study notebook to get custom lessons and track your progress.
🚨 AI News | TestingCatalog@testingcatalog · 博主 · 1 天前专挖 AI 产品未发布新功能的爆料号

关闭你的笔记本吧。Codex 现在在 Atomic Bot 上可用,用户可以关闭笔记本并让任务在云端持续运行。

用户可以把代码库交给 agent,让它持续改进、发现和修复 bug 等等!

引用 atomicbot.ai @atomicbot_aiRun Codex 24/7 in the cloud with Atomic Bot! Just hand it a repo and close the laptop. One click to start. We ran it on ours first and it found a bug nobody had answered. Codex fixed it before we got back.查看被引原帖 ↗
查看英文原文
Close your laptops. Codex is now available on Atomic Bot, allowing users to close their laptops and keep tasks running in the cloud.

Users can hand over their repositories to the agent in order to improve them continuously, discover and fix bugs, and more!
🚨 AI News | TestingCatalog@testingcatalog · 博主 · 1 天前专挖 AI 产品未发布新功能的爆料号

Google Antigravity现在支持远程控制了!

> 试试远程控制 - 在电脑上启动工作,然后用你的agents从手机或其他设备接着干。在应用设置里打开远程控制。

有意思的是,Antigravity现在还没有移动应用,这功能通过手机网页实现。不过还是有点可惜,这应该能成为已停用AI Studio应用的一部分。

h/t @ShakhzodR93267

引用 Google Antigravity @antigravityRemote Control is here! Take Antigravity with you anywhere. Access your active sessions from any modern supported browser or mobile device across iOS and Android. Rolling out to all users starting today, beginning with Ultra subscribers.查看被引原帖 ↗
查看英文原文
Google Antigravity now supports Remote Control!

> Try Remote Control - Kick off work on your computer and continue working with your agents from your phone or another device. Turn on Remote Control in app settings.

Interestingly, there is no Antigravity mobile app in place and this feature work via mobile web. Yet, I wish it would have been a part of the discontinued AI Studio app.

h/t
@ShakhzodR93267
◔ 1.8 万 次浏览♥ 241⇄ 9▶ 含视频新品看原帖 ↗
Gary Marcus@GaryMarcus · 博主 · 1 天前

“他还是不信你知道怎么造AGI”

引用 NIK @ns123abc“Sir… Luke Metz… whose research preview project became ChatGPT… who then left to co-found Thinking Machines Lab with Mira Murati… who then left Thinking Machines to rejoin OpenAI in January… sir… he just left OpenAI AGAIN…”查看被引原帖 ↗
查看英文原文
“He’s still not convinced you know how to build AGI”
meng shao@shao__meng · 中文博主 · 1 天前

程序员朋友们,应该都知道 Clean Code (代码整洁之道) 这本书吧?!

Clean Code 的作者 Robert C. Martin (Uncle Bob) 和 231K 🌟 Skills For Real Engineers 作者 Matt Pocock 做了一次深度对谈:AI 时代下,软件工程的基本功是否仍然成立,以及如何驾驭智能体。

youtube.com/watch?v=zcLPGC-t…


# Uncle Bob 与智能体协作的实战方法

1. 初识智能体的挫折
去年底开始用 Grok 等智能体写代码,最初印象平平:智能体"很快但让我变慢",到处留下"狗粪"——混乱的代码、未清理的依赖、破碎的测试。他很快发现一个关键事实:智能体和人类一样,会被脏代码拖垮。 代码足够混乱后,智能体会陷入改一处坏另一处的循环,甚至直接放弃。

2. 转向确定性工具而非"指令堆砌"
Bob 最初尝试用冗长的 prompt(TDD、Clean Code 规则等)"驯化"智能体,但发现模型把这些规则当成"加勒比海盗式的指南"——可遵守可不遵守。技术原因是 "lost in the middle"现象:上下文窗口中段的内容会被忽略,越长的初始 prompt 越多内容被埋没在中段而失效。

于是他转向确定性工具链:CRAP(结合圈复杂度与测试覆盖率的代码质量评分)、变异测试(mutation testing,翻转运算符后看测试是否仍能捕获)。这两项技术他在 2000 年就尝试过,但人工修复成本太高、不可行。如今智能体"不在乎枯燥、速度极快",反而让这些老技术焕发新生。

3. 多智能体"流水线"
Bob 搭建了一条智能体接力链,每个智能体只承担单一任务以控制上下文窗口:
· Specifier:把人类文档转为 Gherkin 验收测试 + QA 程序
· Coder:写单元测试与实现代码,让 Gherkin 通过
· Cleaner:跑 CRAP 分析,清理实现者留下的烂摊子
· Hardener:跑变异测试,"毫不留情"地追求 100% 覆盖
· QA Agent:把 QA 文档转为可执行脚本,端到端验证

单智能体 5 分钟的任务,这条链约 1 小时;人类则需约半天。质量远高于人类通常会投入的水平,且是"早期投入生产力、后期收回"的投资。

# 关于架构与模块设计
Uncle Bob 让智能体构建了一个架构可视化工具(UML 风格,可逐层下钻到代码),并用一个依赖规则规范文件约束模块间依赖方向,由检查器在结尾强制执行——违反就由智能体通过依赖倒置、插入接口、拆分模块等方式修复。

他认同 John Ousterhout 的"深模块"理念(小接口、深实现):模型可以读接口而不必读实现,只要代码一致即可。模型也通过读测试来理解系统行为。任何有助于代码结构的做法,都有助于模型理解代码。

# 对 Clean Code 原则的调整与坚守

需要调整的:阈值
智能体短期记忆巨大且精确,可承受更高复杂度。Bob 将 CRAP 阈值从人类的 4 上调到 6,考虑推到 8,正在试探边界。

不应强加的:人类纪律
TDD 对人类有价值(受限于短期记忆),但不应强加给智能体——让它们先写函数再写测试(John Ousterhout 风格)即可,强行要求"一行测试一行实现"它们最终都会回退到自然方式。

核心论断
可以把人类价值观强加给智能体,但不要把人类的行为纪律强加给智能体。

# 对"规格驱动开发"的质疑

Bob 明确反对冗长的前期规划。他本周还在实验"先把一切规划好再交给智能体",结果一如既往地灾难——计划永远不完整,智能体不够聪明去填补,人类只能不断叫停、改计划、重启。

他重提经典的"盖房子"比喻:如果每次改动成本是 1 美元,你会先雇建筑师画完美图纸,还是直接对承包商边改边建?答案显而易见。如今改动成本已接近零,没有理由做沉重的前期规划,应该"摆弄、摆弄、摆弄直到看起来对"。

他建议回归敏捷:做一两个故事 → 看架构 → 手动整理 → 再做几个故事。那个"手动整理"步骤可能永远无法完全自动化。

关于规格是否持久化:不持久化。规格是临时的、不断变化的,最终产物(代码)本身就是规格。他甚至建议别人不要直接下载他写的工具,而是让智能体看他的工具、再为用户自己定制一份。

# 新人如何学习"战略性编程"

John Ousterhout 区分了战术编程(一线战斗)与战略编程(统帅指挥)。智能体擅长战术、糟糕于战略。问题是:AI 已吞掉战术工作,新人如何学战略?

Bob 的建议路径:
1. 先写一年代码,理解智能体在应对什么
2. 入职后被当作智能体对待——交给智能体一样的任务,受同样的确定性工具约束,几个月内极度低产但学到大量
3. 通过这条"流水线"后,才可被信任去运行自己的智能体
4. 读老书:Tom DeMarco、Ed Yourdon、《The Pragmatic Programmer》等 70-80 年代的经典——那时这些教训刚被学到,需过滤掉过时部分
5. 必须亲身"感受"过挣扎才能识别智能体的挣扎

他强调:不能完全脱离代码。 十年前他建议人花一个周末写汇编,以理解 Java 之下的真实世界。这个道理在 AI 时代依然成立。

# 软件基本功是否仍然重要

仍然重要,理由一如既往。 引用 Dijkstra:软件是人类尝试过最复杂的事物。基本功是把复杂性组织成可被构想的形式——不仅人能构想,模仿人类的模型也能。

那些认为基本功不再重要的人"会以痛苦的方式学到,且不会太久"。Bob 亲眼见过智能体撞上那堵墙。

抽象层的纵向类比
从二进制 → 汇编 → 编译器 → 模型,每一次抽象层上升,下层的人都会哀叹"这会毁掉一切、连五岁小孩都能写代码"。每次都没发生。同样的规则、同样的基本功,因同样的理由存在。被你扔掉的规则,一年后你会从地上捡起来拍拍灰,想起为什么需要它。

Tibor Blaho@btibor91 · 博主 · 23 小时前逆向挖掘 AI 产品代码的爆料专家

ChatGPT网页应用里出现了有趣的新功能参考"ChatGPT Agent Email" - "使用新的'OpenAI botmail'连接器连接Agent Email应用来查看你的邮箱" - "由你的agent email地址发送或接收的邮件将在这里显示"(@ chatgpt . email)

查看英文原文
"ChatGPT Agent Email" - interesting new references in the ChatGPT web app

"Connect the Agent Email app to view your inboxes" using a new "OpenAI botmail" connector

"Emails received or sent by your agent email addresses will appear here" (@ chatgpt . email)
Tibor Blaho@btibor91 · 博主 · 1 天前逆向挖掘 AI 产品代码的爆料专家

Anthropic似乎在为Claude打造一个协调型agent工作空间。Projects(Claude Code Channels)提供共享上下文,coordinator sessions组织工作,工作流和日程启动后台任务,humans审查委托的行动和结果,Slack/Teams作为外部协作入口。'Claude将并行为你协调这个项目。Memory和instructions在任务间共享。'

引用 🚨 AI News | TestingCatalog @testingcatalogANTHROPIC 🔥: Claude Tag is coming to Claude Desktop in the form of collaborative Projects for Claude Code. That's essentially a built-in Slack 👀 > Earlier spotted Managed Projects feature got a slight upgrade, mentioning that users will be able to spawn different threads within a single project session. > Project creation UI got a new option to add repositories as context, which will be used across all sessions. The upcoming Projects for Claude Code feature is expected to enable teams and individuals to work collaboratively inside Claude Desktop, as they would with Claude Tag on Slack. Claude Code Projects will have persistent context and memory, which Claude will periodically refine and improve over time.查看被引原帖 ↗
查看英文原文
Anthropic appears to be building a coordinated-agent workspace for Claude in which Projects (Claude Code Channels) provide shared context, coordinator sessions organize work, workflows and schedules launch background jobs, humans review delegated actions and results, and Slack/Teams serve as external collaboration entry points

"Claude will coordinate this project for you in parallel threads. Memory and instructions are shared across tasks."
◔ 1.7 万 次浏览♥ 137⇄ 5▶ 含视频新品看原帖 ↗
meng shao@shao__meng · 中文博主 · 1 天前

去 AI 味 Skills Top 10 -- 感谢
@juampitech
整理!

咱们一起看看每个 Skill 具体做了什么,哪些场景、用途该用哪个,组合起来用效果是否更好?!

# 10 个 Skills 分别是什么?

1. stop-slop(Hardik Pandya)——以结构化规则去除 AI 痕迹,装机量最高,被多个技能引用为"结构规则"来源,是该生态的奠基性技能。
2. no-ai-slop(Peter Yang)——定位"锐利的人类编辑",强调最小有效修改与保留个人声音,提供编辑与检测两种模式。
3. humanizer(blader)——维基百科"AI 写作迹象"指南的技能化源头,被多方复用为模式目录基础。
4. unslop(Cursor 官方插件库)——官方来源,流程为扫描、重写、注入"灵魂"、自审残留痕迹。
5. slopbeth(ehmo)——追求每句承载信息的密集写作,引入奥威尔六规则与"证据边界",自带脚本可重复验证。
6. humanizer(Adam Boudjem)——识别 53 种模式并打 0-100 分,套用声音配置,主动调整句长 burstiness,含 detect/rewrite/edit 三模式。
7. deslop(Stephen Turner)——唯一明确覆盖科研写作(论文、摘要、基金申请、审稿回复)的技能。
8. anti-slop(Matt Silverlock)——把"保留作者声音"设为首要指令,主张假阳性比残留 tell 更糟,只做外科手术式措辞修改。
9. humanize(aasha)——融合维基迹象、stop-slop 结构规则、brandonwise 统计检测三方,共 41 模式,强调不编造原文没有的事实。
10. anti-ai-slop-writing(jalaal)——前置写作约束指令,以禁用词表加结构硬规则(禁三段式、禁连续等长句、禁 parataxis、禁对冲摇摆)从写作时就生效。

# 按工作模式分类

单模式重写型包括 stop-slop、blader/humanizer、unslop——给文本即返回去 slop 版本,无独立检测态。

双模式型以 no-ai-slop 为代表,默认编辑,另设检测模式只点名模式并引用原句,不重写、不评分、不猜测是否 AI 所写。

多模式型有 Aboudjem/humanizer 与 slopbeth。前者三模式加 0-100 评分;后者含重写、批评、基准、检测器验证四种工作流。

审查导向型为 elithrar/anti-slop 与 deslop,偏重先指出问题所在而非直接重写。

前置写作约束型仅 anti-ai-slop-writing,它不是后处理,而是在写作时就生效的硬规则。

# 按适用内容分类

通用散文(文章、博客、通讯)首选 no-ai-slop,备选 stop-slop 或 unslop,理由是装机广、最小修改、社区验证充分。

学术与科研写作首选 deslop,备选 slopbeth。deslop 是唯一明确覆盖论文、摘要、基金申请、审稿回复的技能;slopbeth 的密集写作风格也适合技术内容。

短文本(推文、邮件、消息)首选 anti-ai-slop-writing,备选 no-ai-slop。前者显式覆盖短文本,结构硬规则在短文中最显效。

任意或不确定类型首选 aasha/humanize,备选 Aboudjem/humanizer。前者模式目录最全(融合三方 41 模式),通用性最强。

# 按方法论取向分类

最小修改、保留声音派以 no-ai-slop 与 elithrar/anti-slop 为代表,主张删套公式而留作者会辩护的刻意选择,认为假阳性比残留 tell 更糟。

注入声音、灵魂派以 unslop 与 Aboudjem/humanizer 为代表,认为无灵魂的平淡文本同样可疑,故用 voice profile 主动注入个性。

结构硬约束派以 anti-ai-slop-writing 与 stop-slop 为代表,从句长、三段式、parataxis 等结构层根治,而非逐词修补。

信息密度派以 slopbeth 为代表,目标是每句都承载信息的密集写作,而非"检测器认不出"。

证据、事实边界派以 slopbeth 与 aasha/humanize 为代表,原文没有的事实不得编造,模糊表述转为待证问题或显式归因。

统计可测检测派以 Aboudjem/humanizer、aasha/humanize、anti-ai-slop-writing 为代表,靠句长 burstiness、0-100 评分、等长句检测等可重复检验的信号。

# 组合安装建议

轻量日常可单装 no-ai-slop,一个技能即覆盖编辑与检测,满足大多数日常写作。

全覆盖单装可选 aasha/humanize,它已含三方模式目录,省去再装 stop-slop 与 blader/humanizer。

官方加保守审查可用 unslop 加 elithrar/anti-slop,前者负责官方重写,后者做严格保留声音的审查兜底。

科研组合用 deslop 加 slopbeth,前者覆盖科研场景,后者补事实边界与密集写作。

结构根治加量化用 anti-ai-slop-writing 加 Aboudjem/humanizer,前者在写作时施加硬约束,后者在改后评分验证。

引用 Juampi @juampitechPost got a lot of traction, so I made a rank with the anti-slop skills people need to install. 1. stop-slop - @hvpandya skills.sh/hardikpandya/stop-… 2. no-ai-slop - @petergyang skills.sh/petergyang/no-ai-s… 3. humanizer - @blader skills.sh/blader/humanizer 4. unslop - @poteto skills.sh/cursor/plugins/uns… 5. slopbeth - @synopsi skills.sh/ehmo/slopkit/slopb… 6. humanizer - @AdamBoudj skills.sh/Aboudjem/humanizer… 7. deslop - @strnr skills.sh/stephenturner/skil… 8. anti-slop - @elithrar skills.sh/elithrar/dotfiles/… 9. humanize - @aashatwt skills.sh/aashaexo/soundshum… 10. anti-ai-slop-writing - @jalaal_tweets skills.sh/jalaalrd/anti-ai-s…查看被引原帖 ↗
MiniMax Design (H3)@Hailuo_AI · 公司官方 · 1 天前MiniMax 旗下海螺 AI 视频官方

只需一个简单的提示词,即可生成 After Effects 动画✍️

引用 FATHELA ESQ @AmControoMiniMax Design unlocked a new creative world. Guess how long this video took? 10 minutes! And the crazy part? I only used one prompt. It looks like something I’d spend a week making in AE Plus, annual members get 20% OFF H3 and image generation.查看被引原帖 ↗
查看英文原文
Build After Effects animations with just one easy prompt✍️
AshutoshShrivastava@ai_for_success · 博主 · 23 小时前高频 AI 新闻与产品动态博主

Hermes 中的 Gemini 3.7 Flash 太强了

引用 Logan Kilpatrick @OfficialLoganKGemini 3.7 is our fastest growing model launch to date, amazing to see the reception!!!查看被引原帖 ↗
查看英文原文
Gemini 3.7 Flash in Hermes is 🔥
Orange AI@oran_ge · 中文博主 · 1 天前Orange AI,中文圈 AI 产品观察博主

ox 在 opencode 的体验不好
在 hermes 的体验要好很多
在 cola 的体验也要好很多
可能是因为 opencode 那边的兼容性没那么好
个人一直用不明白 opencode harness

引用 Arsh @be_arshOx Alpha in Hermes is actually insane, way better than running it in OpenCode. OpenCode kept randomly stopping mid task, but in Hermes it hasn't happened once. It’s super independent and takes initiative without you having to ask. That can be hit or miss with agents, but here it’s great, it understands the end goal and handles the extra stuff before you even think to ask.查看被引原帖 ↗
Runway@runwayml · 公司官方 · 1 天前AI 视频生成公司 Runway

Runway Ruby 现已向 Max 计划和企业版开放。可将任何视频生成或转换为 16-bit EXR 序列或 10、12-bit ProRes 和 HEVC,支持 BT.2020 色彩空间和 PQ 或 HLG。

现在就在下面的链接试试。

查看英文原文
Runway Ruby is now available for Max plans and Enterprise. Generate or convert any video into 16-bit EXR sequences or 10 and 12-bit ProRes and HEVC, in BT.2020 color with PQ or HLG.

Try it now at the link below.
◔ 1.4 万 次浏览♥ 127⇄ 13▶ 含视频新品看原帖 ↗

百度应该学google,把gemini强行嵌入到搜索的最显眼的位置,让用户肌肉记忆的搜索获得更好的体验,输出一大块文心一言总结,

现在百度依然搜出来一堆网页,而中国用户早就失去了看几十个网页总结和内容关键词高亮的全部耐心了。

google这一步做得非常好,百度如果不立即跟上,就只能死亡倒计时了。

Lisan al Gaib@scaling01 · 博主 · 1 天前高频 AI 模型测评与爆料博主

你可以在这里读到更好的版本,没有讨厌的付费墙,而他们讨论的可是*公开基准测试*。


scaling01.substack.com/p/hav…

引用 SemiAnalysis @SemiAnalysis_Are Open Models Catching Up? Comparing open vs. closed models across the eras of frontier models, Is the gap narrowing? newsletter.semianalysis.com/…查看被引原帖 ↗
查看英文原文
you can read the better version here without shitty paywall behind which they discuss *PUBLIC BENCHMARKS*


scaling01.substack.com/p/hav…
歸藏(guizang.ai)@op7418 · 中文博主 · 1 天前歸藏,中文圈 AI 工具与提示词博主

AI 硬件的付费点应该是让用户付费拿到自己用硬件获得的上下文,而不是反过来。

最近这个 AI 硬件反正出的越来越多了,但是大家的思路好像都没有转变过来。

还是收钱是为了让用户把数据留在自己这,把上下文留在这自己这。

但是我最近用的最多的 AI 硬件反倒是飞书的那个录音豆。

用的最多的原因是它可以通过飞书的 CLI 把我的所有的录音的转的文字全部拉到我本地的上下文那边。

在 Agent 集中化以及模型价格频繁变化的这个时间点。

你不能不可能通过让用户付费购买你的模型服务和 Agent 来绑住用户。

反倒是如果你能提供更丰富的 Agent 或者 Skill 帮助用户将这些多获得的上下文给到他自己的用的 Agent 上,价值会更大。

DeepLearning.AI@DeepLearningAI · 公司官方 · 1 天前

🚀 The Batch 最新版本已上线!以下是你需要了解的内容:

🚀 SpaceXAI 推出的 Grok 4.6 来了,正挑战 OpenAI 和 Anthropic 的顶级模型。
🕵️ Anthropic 为所有新 Claude 模型添加隐形水印。
🧠 阿里巴巴发布了 Qwen3.8 Max,一个拥有 2.4 万亿参数的开源权重模型。
🗣️ 研究人员开发了 Agentic ASR,可像人类编辑一样修复语音转文字错误。

在 The Batch 网站上查看完整分析和基准得分,考虑订阅吧!👇

hubs.la/Q04tTlgl0

查看英文原文
🚀 The latest edition of The Batch is live! Here is what you need to know:

🚀 Grok 4.6 by SpaceXAI is here and challenging OpenAI and Anthropic’s top models.
🕵️ Anthropic is adding invisible watermarks to all new Claude models.
🧠 Alibaba released Qwen3.8 Max, a massive 2.4 trillion parameter open weight model.
🗣️ Researchers built Agentic ASR to fix speech to text errors like a human editor.

Read the full breakdowns and benchmark scores on The Batch website and consider subscribing! 👇


hubs.la/Q04tTlgl0
elvis@omarsar0 · 博主 · 1 天前

@viktor_com 这个新增功能真不错。现在有了 OpenAI 兼容的 API 和托管 MCP 服务器。我觉得这就是未来用 agents 和应用/服务进行开发的样子。你的 agents 可以无缝地在 Slack、开发工具和各种工作场景中运行。

引用 Antoni Olendzki @Antoni_OlendzkiOur engineers barely open Slack, and that's where Viktor lives. Fixed that today. @viktor_com is now an OpenAI-compatible API and a hosted MCP server, in public beta. In the video I ask it for a branded snake game from opencode. 76 seconds later it's built and published, and it already knew our brand guidelines, because marketing taught it those in Slack. Point the OpenAI SDK at it, or add it as an MCP server in whatever client your team uses. Setup is one copied snippet from settings.查看被引原帖 ↗
查看英文原文
Very nice addition to
@viktor_com
. It's now an OpenAI-compatible API and a hosted MCP server. I think this is what the future of building with agents and app/services will look like. Your agents seamlessly live in Slack, your dev tools, and wherever work happens.
OpenRouter@openrouter · 公司官方 · 1 天前
连环推 ×6

API 升级,现在可以查看和查询每个 agent、每个模型和每个请求的 AI 使用情况。

支持查看消费、token 数、缓存命中率、延迟和每个 token 的混合成本,可以从任何图表向下钻取。特性 👇

查看英文原文
Big upgrade to the API powering the Activity dashboard. View and query your AI usage per agent, per model, and per request.

Spend, tokens, cache hit rate, latency, and blended cost per token, with drill-down from any chart. Features 👇
Runaway agent monitoring: Trends sort the same data by movement instead of size.

A panel ranks what's rising and falling across models, users, API keys, and apps, which makes it easy to spot a runaway agent or a new model gaining traction in your org.
Explore lets you assemble the view yourself:

pick a metric (spend, tokens, cache hit rate, or latency down to P50/P90/P99), group by dimensions like model, app, user, or any classifier, and choose your rollup. Views can be saved for your or shared w your whole org.
New: Upgraded Analytics API

Everything in Explore is also available through the new Analytics API. Call /analytics/meta to see supported metrics and dimensions, then query:
The new API allows for much better agentic analysis.

Give a coding agent a management key and the openrouter-analytics skill and have it run a cost review.

When we ran this internally, it found a preview model burning ~$6.2K/month at roughly 25x our blended rate, traced 98% of it to one batch-pipeline key, and the fix was a one-line model swap!
Set up your favorite harness:
openrouter.ai/ori/harness


Full walkthrough, query recipes, and the cost control cookbook are in the post:
openrouter.ai/blog/announcem…
Replit ⠕@Replit · 公司官方 · 1 天前AI 编程平台 Replit 官方
连环推 ×6

本周 Replit 资讯

1) Free Mode:创建容量最多增加 30 倍
2) Conversations:通过聊天启动任何想法
3) Agent modes 改名
4) Routines + Memories
5) Black box 渗透测试

Thread 🧵

查看英文原文
This Week in Replit

1) Free Mode: create up to 30X more
2) Conversations: start any idea by chatting
3) Agent modes renamed
4) Routines + Memories
5) Black box pen testing

Thread 🧵
2) Conversations.

Start any new idea by chatting with Replit.

Most app ideas start somewhere else, in another AI assistant, doing research and exploring concepts. Then you paste it into Replit and lose all that context.

Not anymore. Research, ideate, plan, then build from that same conversation.

Learn more →
docs.replit.com/chat/convers…
3) We renamed the agent modes.

Economy is now Power. Power is now Max. Same cost on both.

Build in Free Mode, and Replit will suggest leveling up when a task calls for it. You can take it or stay on Free.

Learn more →
docs.replit.com/chat/agent-m…
4) Routines and Memories.

Routines (beta) let you schedule a task from a conversation and get the result back in the same thread. A morning summary of your Slack, your inbox, your calendar.

Memories means Replit remembers how you like to work.

Routines →
docs.replit.com/chat/routine…

Memories →
docs.replit.com/chat/memorie…
5) Black box pen testing.

Replit already scans your packages and your codebase. Now it scans your app the way an external attacker would see it, out in the wild, looking for doors you left unlocked.

Run a Level 3 scan from your Security Center.

Read more →
replit.com/blog/black-box-pe…
Full Friday Showcase replay:


@victoriakimse
+
@JimmyAustin
on Free Mode and Conversations

@edsioufi
+
@okayzade
on Routines

@jasondellaluce
on black box pen testing

Plus live demos of Routines and a Level 3 security scan.

🎬
youtube.com/1swpJRcCj9Q
GitHubDaily@GitHub_Daily · 中文博主 · 1 天前

明明记得上周跟 Claude Code 讨论过一个方案,想再翻出来,找半天没找到那对话。

Wake 把 Mac 上所有 Agent 的历史会话收进一个原生应用,统一浏览、全文搜索、一键恢复。

支持 Claude Code、Codex、Gemini CLI、OpenCode 等 14 种工具,搜索连中文和代码片段都能精确匹配,点结果直接跳到那条消息。

GitHub:
github.com/iAmCorey/Wake


找到会话后一键在终端里恢复现场,自动回到原来的项目目录接着聊。

所有工具的数据只读访问,全程零网络请求,作者机器上 800 MB 的会话记录建索引只要 5 秒。

Rust 写的原生应用,要求 macOS 14 以上。同时用好几个编码 Agent 的,装一个当会话档案馆挺合适。

karminski-牙医@karminski3 · 中文博主 · 1 天前karminski-牙医,中文圈模型评测博主
连环推 ×2

给大家写了个 deepseek-v4-flash-vision-exp 输入视频教程

deepseek 最近真的是高产, 刚刚又发了 deepseek-v4-flash-vision-exp, 首个多模态【大】模型. 而且是 v4-flash 能力级别的. 但是! 虽然不是瞎子了, 但是还是听力有问题, 不支持音频输入. 所以默认 deepseek 官网和API都不支持视频输入, 于是给大家写了个小教程, 如何使用这个模型处理视频.

简单来讲, 方法就是直接把视频抽帧. 而且需要注意, 虽然模型支持gif输入, 但是它只识别 gif 的第一帧(我写代码验证了). 所以把视频转换为gif是行不通的.

抽帧的最佳实践也给大家:


#deepseekv4flashvisionexp
#deepseek多模态
#deepseek多模态模型

AIGCLINK@aigclink · 中文博主 · 1 天前

Agent Orchestrator (AO):管理多coding agent 用的"项目管理系统,也就是面向AI程序员的tapd或tower,我们正从卷 coding agent 本身,过渡到"agent 变多之后怎么管",AO 卷的就是这一块。

场景面:
一个 agent 你盯得过来,五个 agent 同时改一个 repo,你面对的是五个终端、五条分支、五个 PR、一堆红色的 CI:工具没变强,你的注意力被切碎了,我们日常会用claude fable做prd,然后用codex做实现,用grok cli做审查,多个切来切去,很容易搞混乱,而且有时候没办法很好的并行工作。

AO 的解法是把 IDE 换成"调度台":上面一层 orchestrator 负责规划和派活,下面每个 worker 一个任务一个 worktree,中间用一块实时看板把 PR / CI / review 状态映射成"谁在跑、谁卡住、谁能合"。

AO支持 26 个主流code agent,每个 worker 独占一条分支和 worktree,看板从 PR、CI、评审的真实状态推导出来。有意思的地方在于卡片状态是推导出来的,不是人工维护的——这基本承认了一件事:agent 多了以后,真正稀缺的不是算力,是人的注意力该放在哪。

它把结构分成三层,划得很干净:

① Worker——执行单元
一个任务 + 一个编程 agent + 一个隔离工作区,Git 项目下,每个 worker 独占自己的分支和 worktree(临时活则给一个 AO 托管的无分支目录),任务、对话、终端、改动文件、浏览器预览、PR、CI、评审状态,从头到尾挂在这一个会话上——所以并行的活不会塌成一锅粥。

② 项目编排者——常驻的规划 agent,工作在任务之上那一层
管的是产品方向、技术策略、优先级、跨仓库的工作顺序。它的项目级对话保留目标、决策、约束和此前的推理,并把这些跟仓库上下文、以及 AO 的实时状态(哪些 worker 在跑、谁负责、PR、CI、评审)合在一起用。计划成型后,它能拆成任务、生成或重定向 worker、把该给的上下文分发下去、跟进度。

分工用一句话写死了:"编排者拥有规划与委派;worker 拥有实现、测试、提交与 PR。"

③ 看板——不是手动拖的,是从事实推导的,AO 根据会话、PR、CI、评审的真实状态决定卡片位置:

.Working:正在实现,或等你下一条指令
.Needs you:阻塞、缺输入、CI 挂了、被要求改动、信号丢失
.In review:开着的或草稿 PR,等检查或等人看
.Ready to merge:已批准或可合并,合并后还留在板上直到你归档

这点比"AI 公司"类项目实在得多——它的状态不是模型自称的,是 PR 和 CI 说了算。

对比昨天介绍的项目Wake 读本地会话文件,所以 Amp、Droid 这种把会话存云端的它读不到;而AO 是驱动 agent 干活,所以云端会话的 agent 照样能管。

#AIagent
#多智能体
#开发者工具
#编程agent
#开源

Gary Marcus@GaryMarcus · 博主 · 1 天前

Google在人才大量流失时做出了改变。OpenAI董事会毫无作为。

引用 Rohan Paul @rohanpaul_aiOpenAI's Americas sales chief Kaylin Voss resigned 5 months in, 1 week after Denise Dresser. - The Information查看被引原帖 ↗
查看英文原文
Google made a change when talent left in droves.

OpenAI’s board has done diddly squat.
Gary Marcus@GaryMarcus · 博主 · 1 天前

关键在于符号AI与LLM分层结合。
@GaryMarcus 一直是对的。

引用 Frank Rundatz @FrankRundatzProgramming with Claude Code radically changed about 7 or 8 months ago. I am at least 10x more productive than I used to be. A year ago I was wondering if it was making me more productive at all. I’m using Hermes with Grok 4.6 to do actual work - literally hours of work - every day. 90% autonomously. Giving Hermes a script and a governed API to call while running on a dedicated Ubtuntu VM with very limited access has been tremendously helpful. I keep throwing work at it and still can’t seem to use up the Grok 4.6 credits I get for free with X Premium+. The key is symbolic AI layered in with an LLM. @GaryMarcus was right all along.查看被引原帖 ↗
查看英文原文
“The key is symbolic AI layered in with an LLM.
@GaryMarcus
was right all along”
Kol Tregaskes@koltregaskes · 博主 · 1 天前

ChatGPT for Work 可以整理你的 ChatGPT 库,只需告诉它就行。😀 它正在清理我的库。🫡

引用 Merk @makeamarkeryGPT Work ks CRAZY. Did you know you can ask it to organize your ChatGPT Library? (And it will). @ChatGPT @OpenAI查看被引原帖 ↗
查看英文原文
ChatGPT for Work can organise your ChatGPT library, just ask it. 😀

It's cleaning up my library as I type. 🫡
Linus ✦ Ekenstam@LinusEkenstam · 博主 · 1 天前

机器人领域迎来一次性学习。
如果你有点蒙?让我解释一下

你给机器人展示一次,它就学会了。
2027年机器人技术会带来这么多进步,我们会跟不上节奏。

期待看到我们在AI图像和AI视频上见过的同款加速和性能提升。

能力呈指数级爆发。

一旦学会一项任务,它就永远保留。
快速积累。直到所有问题都迎刃而解。

查看英文原文
One shot learning coming to robotics.
if you are confused? Let me explain

You show the robot once, and it learns.
2027 will bring so much advancements in robotics that we will have a hard time to keep up.

Expect the same ramp up and speed improvements that we’ve seen with AI images and AI video.

Exponential acceleration in capabilities.

Once a task is learnt it stays that way.
Compounding quickly. Until everything becomes a solved problem
GitHubDaily@GitHub_Daily · 中文博主 · 1 天前

OpenCore Legacy Patcher,能让苹果已经停止支持的旧电脑装上新系统。

即便是 2007 年的老机器,也都能安装和使用 macOS Big Sur 及更新版本。

Wi-Fi、显卡加速、个人热点都有适配,装完系统更新照样正常推送。

GitHub:
github.com/dortania/OpenCore…


项目已斩获 18000+ Star,文档写得很细。家里有旧的 Mac 电脑,想更新系统可以看下。

Replit ⠕@Replit · 公司官方 · 1 天前AI 编程平台 Replit 官方

免费模式来了:创造 30 倍的内容,保持专注
x.com/i/broadcasts/1Nxaroyoa…

查看英文原文
Free Mode Is Here: Create 30X More, Stay in Flow
x.com/i/broadcasts/1Nxaroyoa…
Pika@pika_labs · 公司官方 · 1 天前AI 视频生成公司 Pika
连环推 ×4

来聊聊Pika Speech吧,这是个3B参数规模的文本转语音模型,推理速度能做RTF 0.02。

我们很兴奋地想分享的是,在本地跑长文本测试时,Pika Speech大概1.2秒就能生成一分钟的48kHz录音室级语音——而且支持最长五分钟的请求。这效率比ElevenLabs v3省钱多了,最高能省9倍,比Cartesia和ElevenLabs Turbo省4.5倍,比Fish Audio省2倍。

查看英文原文
Let’s talk about Pika Speech, a 3B text-to-speech model at RTF 0.02.

We’re excited to share that in locally run long-form tests, Pika Speech generates one minute of studio-quality 48 kHz speech in about 1.2 seconds—and supports requests up to five minutes. That efficiency makes it up to 9× more cost-efficient than ElevenLabs v3, 4.5× than Cartesia and ElevenLabs Turbo, and 2× than Fish Audio.
That performance comes from a full inference stack, not just a small step count:

Flow matching + distribution-matching distillation: eight or fewer denoising steps
FlashAttention-3 across self- and cross-attention
Token packing: five sentence chunks cost roughly 2× as much as one, not 5×
CUDA graph replay + fused RoPE kernels

Together, these optimizations reduced one-minute generation from 1.29s to 1.04s.
Most TTS systems adjust pace by trimming or time-stretching after generation.

Pika Speech controls pace and duration inside the model with an EOS latent: an end-of-speech anchor placed at the target final frame. Move it earlier and delivery compresses; move it later and the same words relax and breathe.
Read more about our efficiency innovations here →


experiment.pika.art/blog/pik…
◔ 8,653 次浏览(2 条合计)♥ 125⇄ 6新品看原帖 ↗
Luma@LumaLabsAI · 公司官方 · 1 天前AI 视频生成公司 Luma

为创意专业人士打造的创意智能体。9月14日,与 GMI Cloud 一同亮相 Summer Signal 舞台。

引用 GMI Cloud @gmi_cloud. @LumaLabsAI joins Summer Signal '26 as our multimodal agentic partner On stage, they'll be talking through Luma Agents: creative agents built for creative professionals. Plus what's new inside, and how it fits into your creative workflow. Register to attend 👇查看被引原帖 ↗
查看英文原文
Creative agents built for creative professionals. On stage at Summer Signal with GMI Cloud, September 14.

本站由 Jedee杰哥 打造 · 公众号「Jedee杰哥」每早送 AI 日报

姊妹站:𝕏 简中账号数据榜单 · X 关注 @jedeeai · RSS 订阅 · AI 日报 · 历史归档