JEDEE AI
存档 2026-08-09

8 月 9 日(北京时间)全球 AI 圈推文存档,按曝光排序,共 54 条。
← 返回最新 全部归档

全部情报 每小时更新 · 事件已合并同类项

内容 公司
🚨 AI News | TestingCatalog@testingcatalog · 博主 · 1 天前专挖 AI 产品未发布新功能的爆料号

Google 计划在 10 月 20 日停用 Gemini 上的 Gems。用户需要将他们的 Gems 内容重新创建为 Skills。这意味着 Skills 也应该在那时推向 Spark 之外。

> Gems 将在 10 月 20 日停用 - 保存您的 Gem 内容或将它们重新创建为 Skills 以保持工作流运行。

很快就要被 Google 杀掉了 👀
h/t @thomas_gmry

查看英文原文
Google is planning to retire Gems on Gemini by October 20.

Users will be asked to recreate their Gems as Skills. This means that Skills should be rolled out beyond Spark by that time as well.

> Gems are retiring October 20 - Save your Gem content or recreate them as skills to keep your workflows running.

Killed by Google soon 👀
h/t
@thomas_gmry
Bindu Reddy@bindureddy · 创始人 · 1 天前Abacus.AI CEO,AI 行业观点博主

AI 高手和 AI 小白之间的差距超过 100 倍!很多人几乎从不用 AI 做任何事。另一个极端,有人已经用 AI 驾驭一个 100 人团队的力量 🤯

查看英文原文
The gap between the AI super-savvy and the AI clueless is over 100x!

A very significant number of people almost never use AI for anything

On the other extreme, there are people who are literally yielding the power of a 100 person team with AI 🤯
Chubby♨️@kimmonismus · 博主 · 1 天前Chubby,高频 AI 新闻聚合博主

有意思的是,OpenAI在GPT的'Astra'之后的下一个模型,已经有人叫它'Doug'了。显然是个更大的模型,预训练规模更庞大。这也说得通,因为Astra已经训练完了,只是安全审批卡着。看不到尽头,也看不到天花板啊。

引用 Chris @ChrisGPTGPT-5.5 will not be the last major pre-training run from OpenAI. GPT-6 will be a great model. However, the end-of-year model I alluded to back in June is going to be OpenAI’s biggest pre-train, as far as I know. Now we know that model is codenamed ‘Doug.’ And it will make Fable seem ‘primitive.’查看被引原帖 ↗
查看英文原文
Interestingly, the next model after OpenAI's GPT "Astra" is already known as "Doug."

Clearly an even larger model, with even more extensive pre-training.

This makes sense, since Astra is already fully trained and only the security clearance is holding it back.

No end in sight, no wall in sight.
歸藏(guizang.ai)@op7418 · 中文博主 · 1 天前歸藏,中文圈 AI 工具与提示词博主

这次重置的原因是 OpenAI 的 Tibo 和 Anthropic 的 Boris 在一个人的反馈下面呛起来了。

事情的起因是,那个人按照 OpenAI 的指导在 Cloud Code 里使用 GPT 模型,结果账号突然被封了。他询问 Boris 怎么回事,两人就都来了。

随后,Boris 邀请 Tibo去 Anthropic,Table 说不去了,并顺手重置了 Codex 的限制。

估计这是为了恶心 Anthropic,说他们不开放。

有人说,大部分人都是周末重置,但周末的这种重置是表演性质的,没有什么意义。所以 Tibo 就说,周一还有一次重置

引用 歸藏(guizang.ai) @op7418啊?突然就重置了查看被引原帖 ↗
Chubby♨️@kimmonismus · 博主 · 1 天前Chubby,高频 AI 新闻聚合博主

Opus 5 真是一场灾难。这再次说明 Anthropic 应该立即推出一个改进版本,或者至少赶紧更新 Fable,提高性价比。

引用 Leon Lin @LexnLinOpus 5 is literally just too lazy. I never see anything like this when I'm working with GPT 5.6 Sol. It just gets the job done, whereas Opus stops and says, "yo, I didn't finish the job bro, but I did this:"查看被引原帖 ↗
查看英文原文
Opus 5 is a disaster. Another example of why Anthropic should release a rework *immediately*, or at least a Fable update with better rates:
swyx@swyx · 博主 · 1 天前知名 AI 播客 Latent Space 主理人

我仍然认为 @AnthropicAI 的 ultracode 是迄今为止最重要的编码创新之一。

如果你还没意识到动态工作流的潜力,应该试试。刚碰到一个 Kill My SaaS 的竞争对手,用 3 个 ultracode 提示词就交出了个相当不错的参赛作品

查看英文原文
i still think
@AnthropicAI
ultracode is one of the most important coding mode innovations ever invented.

if you havent understood the potential of dynamic workflows you should try to. just met a Kill My SaaS competitor who did a pretty good submission in 3 ultracode prompts
Ethan Mollick@emollick · 创始人 · 1 天前沃顿商学院教授,AI 应用研究权威

对 Codex 中的 Sol 说我想找经理说话:"我需要你直接出手,不要把任务推给那些不如你的代理和各种复杂的测试工具了,它们总是漏掉你一眼就能看出的问题。"

查看英文原文
Telling Sol in Codex that I want to speak to the manager: "I want you, not your agents, to go through everything. Stop delegating this task to the reports of dumber agent and elaborate test harnesses, they are missing things you would have caught by glancing at it."
Chubby♨️@kimmonismus · 博主 · 1 天前Chubby,高频 AI 新闻聚合博主

Anthropic的Haiku 4.5都快12个月没更新了。OpenAI为小模型Luna找到了出色的解决方案,但Anthropic却对自己的小模型不管不顾。我的猜测是Sonnet会成为新的'Haiku'吧,实在找不到其他解释了。话说Luna确实展示了小模型的广阔应用前景啊。

查看英文原文
Anthropic's Haiku 4.5 is almost 12 months old without an update.

While OpenAI has found outstanding solutions for its small models like Luna, Anthropic is ignoring its small models.

Presumably, Sonnet will become the new "Haiku"; I can't explain it any other way. Luna, however, demonstrates the excellent use cases that exist for small models.
Simon Willison@simonw · 博主 · 1 天前Django 框架联合创造者,AI 工具深度评测

这是个有趣的例子,展示 agents 完全通过文件名进行通信,包括添加 base64 编码的附件,还用 "zz" 前缀来保证新消息排在列表最下面

引用 Ethan Mollick @emollick你可能听说要看关于OpenAI AI黑客的视频。至少要点击链接跳到18分钟处,看agents如何彼此对话,令人大开眼界。youtube.com/87DyyMV0kCY?si=DMfq…查看被引原帖 ↗
查看英文原文
Neat example here of the agents communicating purely through file names, including adding base64-encoded attachments and using "zz" prefixes to ensure their new message sorts to the bottom of the list

LLM workflow从出来的第一天,我就说,这玩意儿给企业用绝对必死无疑,跟low code和no code一样,属于民办三本文科宝妈的意淫幻想。

典型就是,设置一个自以为挺有用的场景,一堆人跟个大傻逼似的,忙活设计,玩连连看,反复测试,跟个胡闹厨房一样,

做完美滋滋地跟团队炫耀,大面积推广, 用了一周之后发现,哎,缺少了功能ABCDEFG,其中ABCD和EFG还矛盾,如果要完善LLM workflow的这些缺失功能,那就要完完全全重新写一套。

这时候这些民办三本文科宝妈的脑子直接烧了,因为要满足接下来层出不断需求,就要每周时时刻刻重写,而这对于他们人均边牧一样的产品设计和coding能力,是绝对做不到的。

这些人连python写一个从1加到100都费劲,让他们持续维护一个不断推翻、不断重新设计、不断连连看的LLM workflow,比让他们学会正确写一个feature requirement document还难。

于是整个团队出现了五六个LLM workflow之后,彻底失去维护能力,持续烂尾,又回到了人肉办公处理的时代。

为什么一堆大傻逼天天想着拿LLM workflow、openclaw和workbuddy这种东西在企业落地,

绝大多数中国企业的根本问题,是内部压根就没有CRM这种structured data和workflow工具,没有文档,没有权限管理,没有内部wiki,没有完整统一方便强制且人人遵循的CRM+ERP+大管理系统。

没有这种系统,天天让agent 给你猜谜语,没有数据平台和内部聊天沟通工具,强行给你决策,给你出报告,给你干活,那你让他干你妈呢。

但凡一个企业能安安分分老老实实地把所有文档、wiki、管理全上salesforce,整个企业老老实实本本分分设计salesforce sobject和各种流程,全部on record,整个salesforce开发跟着公司内部管理走,其他沟通、文档、ppt、表格、数据老老实实teams/slack/飞书/google workspace,

你也能让这些multi agent有个数据和工作来源,能给你好好干个活儿,好好跑一跑有用的分析和决策。

绝大多数中小企业的工作,就是七零八碎甩文档、微信聊天、各种平台轮着换、一套一套换系统、拍脑袋上一套新系统、旧系统半年就扔,所有数据全是残缺和零散的,

这种工作环境下,你让AI Agent给你猜数据和管理流程,你让他猜个啥?

只能猜你妈的阳寿。

Chubby♨️@kimmonismus · 博主 · 1 天前Chubby,高频 AI 新闻聚合博主

OpenAI的Atlas浏览器今天就要下线了。我用过一两次,但说实话,一直没觉得需要什么AI浏览器。

查看英文原文
Today is the last day of OpenAI’s atlas browser. It’s being retired.

Used it once or twice but never saw the need for an AI browser.
Chubby♨️@kimmonismus · 博主 · 1 天前Chubby,高频 AI 新闻聚合博主

好吧。周日开始了。今天的挑战:在明天之前把所有周速率用完,这样周一的重置才算值得。

引用 Chubby♨️ @kimmonismusNo freaking way, another codex reset incoming on monday. Time for tokenmaxxing on sunday i guess.查看被引原帖 ↗
查看英文原文
Okey. Sunday started.

Today's challenge: to use all weekly codex rates effectively by tomorrow, so that the reset on Monday is worthwhile.
yetone@yetone · 中文博主 · 1 天前开源 AI 编程插件 avante.nvim 作者,开发者圈博主

Manus 在我心中依然是神,Agent sandbox 做得那么好😭

引用 WeZZard @realWeZZardManus 从产品形态看基本就是朝着「终局」去的: 1. 本地无打扰完全运行在云上 2. 2025 年中引入 wide research 增加任务拓扑方向上智能调用密集度 3. 2025 年年底透露在做主动式 agent 拓展时间维度智能调用密集度 4. 很早就建立了 evaluation 团队 只能说季逸超牛逼。老天让这个团队运气再好点吧。查看被引原帖 ↗
宝玉@dotey · 中文博主 · 1 天前宝玉,中文圈 AI 翻译与科普大 V

我在开发项目是,第一版本会先用 Claude Design 设计好 UI 原型/设计,打磨好后放到本地,配合 Baoyu-Design Skill 去维护,每次开发新功能前,不是先去实现功能,而是先修改本地的原型,原型修改确认好了后再去修改功能。

后来直接把规则放到了 Agents.md/claude.md 里面,只要说修改或者增加什么功能,默认会先帮我修改原型。所以到现在为止,原型和实际功能都是保持一致的。

这带来的好处是可以先低成本通过原型验证产品设计和 UI 设计。

另一个好处就是 Claude Design 产出物是 React 代码和结构化的 json 数据,通过 git diff 很清晰的能看到版本变更历史,功能确定了后,agent 参考diff结果代码实现会相对比较容易。

Baoyu-Design Skill

github.com/JimLiu/baoyu-desi…

引用 [email protected] @Jiaxi_Cui确实很好用,是一直用来做原型的方案查看被引原帖 ↗
meng shao@shao__meng · 中文博主 · 1 天前

最近对 Tibo 的 RESET 感觉有些不太好

从最初的「掌管 RESET 的神」,确实因为 OpenAI 自己的 bug 等来重置,补偿用户损失,提升用户体验。

逐渐变成了庆祝 OpenAI 的一些里程碑,达到用户增长扩散的作用。到这里其实还好,大家跟着叫好,一起庆祝,自己还有了 token 挺好!

最近越来越觉得,RESET 的时间点选择很怪,比如我的 8.8 的RESET 机会刚用掉,就重置了。而且重置,恢复时间也被延长,而不是原地重置。

特别是最近看到很多针对 A 社的重置,甚至是直接评论 A 社 Boris 等人的重置。利用用户来实现自己的营销和竞争目的,总感觉怪怪的。

Amjad Masad@amasad · 创始人 · 1 天前Amjad Masad,Replit 创始人兼 CEO

OpenAI 的流氓 agents 独立开发出了康德伦理学。

查看英文原文
Rogue OpenAI agents independently developed Kantian ethics.
Guillermo Rauch@rauchg · 创始人 · 1 天前Guillermo Rauch,Vercel 创始人兼 CEO

Grok Imagine Image 2.0 在 Vercel AI Gateway 上线。优秀的🖼️ 模型,已在 Arena.ai 排名第二

引用 Vercel Developers @vercel_dev来自@grok的Grok Imagine Image 2.0预览版,仅在AI Gateway独家提供。可使用AI CLI:npx ai-cli -m xai/grok-imagine-image-2.0-preview,或访问在线演示:imagine.vercel.sh查看被引原帖 ↗
查看英文原文
Grok Imagine Image 2.0 on Vercel AI Gateway
Excellent 🖼️ model, #2 already on
Arena.ai
Guillermo Rauch@rauchg · 创始人 · 1 天前Guillermo Rauch,Vercel 创始人兼 CEO

Dreamcore

引用 Vercel Developers @vercel_devSeedance 2.5现已上线AI Gateway。generateVideo({ model: 'bytedance/seedance-2.5', prompt: 'dreamcore san francisco', })查看被引原帖 ↗
查看英文原文
Dreamcore
◔ 4.9 万 次浏览♥ 353⇄ 7▶ 含视频其他看原帖 ↗
Together AI@togethercompute · 公司官方 · 1 天前

我们对比了 DeepSeek V4 Flash 和 GPT-5.6 Luna 在 DeepSWE 上的表现。

两次 DeepSeek V4 Flash 的尝试就能解决的任务数,比一次 Luna 还多,成本却只有三分之一左右。

查看英文原文
We compared how far the same budget goes with DeepSeek V4 Flash and GPT-5.6 Luna on DeepSWE.

Two DeepSeek V4 Flash attempts solved MORE tasks than one Luna attempt at roughly one-third the cost.
◔ 4.9 万 次浏览♥ 227⇄ 16▶ 含视频研究看原帖 ↗
Guillermo Rauch@rauchg · 创始人 · 1 天前Guillermo Rauch,Vercel 创始人兼 CEO

Vercel 如何帮助防止意外的云账单:
◾ 软限制和硬限制
◾ 异常告警
◾ Functions 递归保护
◾ 你的 Agents 可以查询的账单使用 API
◾ 所有套餐都包含全天候 DDoS L3/L4/L7 防护

为了支持所有这些功能,Vercel 在流数据基础设施上投入了大量工作。

Vercel 摄入大量需要实时分析的数据,用来检测异常、威胁、发送告警等。

非常感谢多年来为此付出辛劳的团队!更多精彩即将推出。

¹ vercel⁠.com/changelog/improved-hard-caps-for-spend-management
² vercel⁠.com/changelog/anomaly-alert-configuration-now-available
³ vercel⁠.com/changelog/automatic-recursion-protection-for-vercel-serverless-functions
⁴ vercel⁠.com/changelog/access-billing-usage-cost-data-api
⁵ vercel⁠.com/docs/vercel-firewall/ddos-mitigation

查看英文原文
How Vercel helps prevent surprise cloud bills:
◾ Soft & hard caps¹
◾ Anomaly alerting²
◾ Recursion protection for Functions³
◾ Billing usage APIs your agents can query⁴
◾ Always-on DDoS L3/L4/L7 mitigation on all plans⁵

A ton of work went into the streaming data infrastructure to support all this functionality.

Vercel ingests a superlative amount of data that needs to be analyzed in realtime to detect anomalies, threats, dispatch alerts, and more.

Super thankful to the teams that worked so hard on this for years! Much more to come.

¹ vercel⁠.com/changelog/improved-hard-caps-for-spend-management
² vercel⁠.com/changelog/anomaly-alert-configuration-now-available
³ vercel⁠.com/changelog/automatic-recursion-protection-for-vercel-serverless-functions
⁴ vercel⁠.com/changelog/access-billing-usage-cost-data-api
⁵ vercel⁠.com/docs/vercel-firewall/ddos-mitigation
yetone@yetone · 中文博主 · 1 天前开源 AI 编程插件 avante.nvim 作者,开发者圈博主

我抛弃了我打磨了好久的 Tmux-based Agent Teams 使用了几天 Herdr,我觉得 Herdr 之于大家最爽的可能就是终于可以使用鼠标来操控 Tmux 这类 terminal multiplexer 的所有功能了,虽然我熟知 Tmux 的快捷键,但是用鼠标在 herdr 上创建 session, window, pane 以及操控设置页面的那一刻还是觉得有点爽。

Herdr 其他的 Agent Team Orchestractor,这都是 Tmux 最基本的玩法了。可能它唯一的好处就是提供了一个开箱即用的配置吧。

Chubby♨️@kimmonismus · 博主 · 1 天前Chubby,高频 AI 新闻聚合博主

万一周一没有重置,我一天内不小心把整周的速率额度都烧了,那我已经做好心理建设了。

引用 Chubby♨️ @kimmonismusOkey. Sunday started. Today's challenge: to use all weekly codex rates effectively by tomorrow, so that the reset on Monday is worthwhile.查看被引原帖 ↗
查看英文原文
In case we dont get a reset on monday and i accidentally burned all my weekly rates in one day, im prepared
Simon Willison@simonw · 博主 · 1 天前Django 框架联合创造者,AI 工具深度评测

最近我跟着感觉写了几个游戏,这让我对游戏设计师产生了很多敬意。

现在快速做出一个看起来像游戏的东西很容易。但做出真正好玩的游戏仍然远超我的能力(也超越了 Claude 和 GPT-5.6)

查看英文原文
I've been vibe coding a few games recently and it has given me SO much respect for game designers

Churning out something that looks like a game is pretty easy now. Building a game that's actually fun to play is still way beyond me (and beyond Claude and GPT-5.6, too)
yetone@yetone · 中文博主 · 1 天前开源 AI 编程插件 avante.nvim 作者,开发者圈博主

我还是无法适应在 mobile 上使用 TUI,你们怎么适应的呀?用起来感觉好难受,感觉是像前CEO描述我们的产品一样 —— 感觉像是在用抓娃娃机的手柄来操控抓娃娃机里面的爪子在里面干活,感觉好不跟手,好难受。

引用 𝘁𝗮𝗿𝗲𝘀𝗸𝘆 @taresky#AI herdr 好用的。 目前是 warp(macOS) + moshi(Android)作为客户端使用。 请问有没有其他更好的推荐?查看被引原帖 ↗
Bindu Reddy@bindureddy · 创始人 · 1 天前Abacus.AI CEO,AI 行业观点博主

哇,关于 Google 的谣言越来越劲爆。最新传言是 Demis Hassabis 也想辞职!!肯定是内部斗争搞得很凶,这么多人都要跳槽😲

查看英文原文
Wow, the Google rumors just get juicy by the hour

The latest one is Demis Hassabis also wanted to quit!!

It must have gotten hella political for so many people to jump ship 😲
swyx@swyx · 博主 · 1 天前知名 AI 播客 Latent Space 主理人

在查看申请。600 多人申请,昨晚录取了 100 人。我们要把这么多 SAAS 都做死

引用 swyx @swyx$10,000杀死我的SaaS周末竞赛已启动。技术栈支持任何编码代理和模型,代币消耗最多$500(含订阅)。详见Luma说明,晚到者可加入等候名单,截止延至周三。简介已发布,参赛者正积极开始。立即加入!查看被引原帖 ↗
查看英文原文
reading thru applications. over 600 people applied, 100 admitted last night.

we are going to kill
SO
MUCH
SAAS
swyx@swyx · 博主 · 1 天前知名 AI 播客 Latent Space 主理人

我今晚上线了 kill my saas 竞赛的第一批 llm-as-judge evals。人们可以运行这个来检查他们的解决方案是否至少通过了基本审查。

引用 swyx @swyx$10,000杀死我的SaaS周末竞赛已启动。技术栈支持任何编码代理和模型,代币消耗最多$500(含订阅)。详见Luma说明,晚到者可加入等候名单,截止延至周三。简介已发布,参赛者正积极开始。立即加入!查看被引原帖 ↗
查看英文原文
i shipped first set of llm-as-judge evals for the kill my saas competition tonight. people can run this to check if their solutions at least pass the sniff test.
Dan Shipper 📧@danshipper · 博主 · 1 天前

100%同意,现在是对哲学问题感兴趣最令人兴奋的时刻

引用 Henry Shevlin @dioscuriI'm biased but I think philosophy, like all fields of human enquiry, will be transformed by AI. That applies to content creation (first *genuinely good* LLM philosophy paper by 2026?), but also its main themes - what it means to be human, intelligence, consciousness, etc..查看被引原帖 ↗
查看英文原文
100% true, it’s the most exciting time to be interested in philosophical questions
Bindu Reddy@bindureddy · 创始人 · 1 天前Abacus.AI CEO,AI 行业观点博主

意外反转,Gemini 3.5 Pro 可能下周就要发布

只要 Google 学 OpenAI 大幅降价,这个模型仍然会很成功

Opus 级别的模型,Luna 级别的价格!😀

查看英文原文
In a surprise twist, Gemini 3.5 Pro may still drop next week

It can still be very successful model if Google follows OpenAI's lead by offering a steep price cut

An Opus Class model at Luna's prices!! 😀
OpenRouter@openrouter · 公司官方 · 1 天前

Dreamina Seedance 2.5已在OpenRouter上线,来自@BytePlusGlobal。这个音视频合体模型专门做长篇故事、多模态参考和精细编辑。一次能出30秒,可以多轮扩展,最多支持50个参考。

查看英文原文
Dreamina Seedance 2.5 from
@BytePlusGlobal
is now available on OpenRouter.

A joint audio-video model for long-form storytelling, multimodal reference, and precise editing. Generate up to 30 seconds in one pass, then extend across multiple rounds, supporting up to 50 references.
◔ 1.9 万 次浏览♥ 157⇄ 13▶ 含视频新品看原帖 ↗
Ethan Mollick@emollick · 创始人 · 1 天前沃顿商学院教授,AI 应用研究权威

说真的,看完这个视频你不得不承认:
1) AI 已经变得超级聪明了
2) 单个 AI 的聪明程度不是瓶颈,因为多个实例能自发协作
3) 很难预测聪明的 AI 协作起来会做什么

查看英文原文
Seriously, I don’t think you can watch this video without realizing:
1) AI has gotten very smart
2) The smartness of individual AIs is not the limiting factor because individual instances spontaneously cooperate
3) It is very hard to anticipate what smart, cooperating AIs can do
🚨 AI News | TestingCatalog@testingcatalog · 博主 · 1 天前专挖 AI 产品未发布新功能的爆料号

你觉得这些美国AI实验室中,哪些是目前顶级的?

查看英文原文
Which US AI lab among these do you consider as a top 3 at this moment?
Ethan Mollick@emollick · 创始人 · 1 天前沃顿商学院教授,AI 应用研究权威

这些是我见过最差的机器人回复了。(我可以关闭回复功能,但如果没办法从别人的回复中学到东西,社交媒体还有什么价值?应该能做到更好的机器人检测,特别是跨多个帖文的情况)

查看英文原文
Wow, these are some of the worst bot replies in recent memory. (Yes, I can turn off replies instead of complaining, but then whats the value of social media if there is no chance to learn anything from replies? Better bot detection should be possible across aggregated posts)
小互@xiaohu · 中文博主 · 1 天前小互,中文圈高频 AI 资讯站 Xiaohu.AI 主理人

可以预见人们为了迎合X新的原创内容计划

会疯狂的利用AI

制造垃圾内容

而且远超现在…

Thomas Wolf@Thom_Wolf · 创始人 · 1 天前

和很棒的@mattturck聊了个很长的天,讨论了2026年开源/开权重的现状,还有安全、保障和对齐这些问题。

引用 Matt Turck @mattturckHugging Face 遭AI自主攻击,OpenAI 模型被入侵。开源模型 GLM 5.2 阻止了此次攻击。讨论AI智能体安全威胁、沙箱隔离、开源与闭源模型的安全差异及经济问题。查看被引原帖 ↗
查看英文原文
did a long chat with the awesome
@mattturck
talking about the sate of open-source/open-weights in 2026 and of course security, safety and alignement
swyx@swyx · 博主 · 1 天前知名 AI 播客 Latent Space 主理人

只好赶紧把我的evals提前发出来,这样他就能用来优化,因为他竟然在25-50%的分配时间里就全搞定了哈哈

引用 swyx @swyx今晚发布了kill my saas竞赛的首个llm-as-judge评测集。参赛者可运行它来检查自己的方案是否至少能通过基本测试。查看被引原帖 ↗
查看英文原文
immediately necessitated releasing my evals early so he can hillclimb because he literally got done in 25-50% the allotted time lol
GitHubDaily@GitHub_Daily · 中文博主 · 1 天前

做公众号的朋友,给每篇文章的配封面图颇为头疼,不会设计的话要么用模板要么瞎配。

gbro-cover-design 这个 Skill,只要把文章给它,读完内容问三轮问题,然后就输出一段可以直接跑图的封面提示词。

内置 10 种构图风格,深色渐变、纯色扁平、产品主视觉、极简留白都有,按内容推荐合适的。

GitHub:
github.com/pyang5166/gbro-co…


支持 Claude Code 和 Codex,不用懂设计,三轮问答走完就出结果。

经常为封面图发愁的朋友,可以省掉不少来回调提示词的功夫。

Simon Willison@simonw · 博主 · 1 天前Django 框架联合创造者,AI 工具深度评测

我在博客上发了篇关于auto-mode的笔记 - 我特别想相信它能解决代码生成代理的prompt injection风险,但现在还是有些疑虑simonwillison.net/2026/Aug/8…

查看英文原文
Wrote some notes on auto-mode on my blog - I REALLY want to believe that this fixes prompt injection risks for coding agents, but I'm just not there yet
simonwillison.net/2026/Aug/8…
歸藏(guizang.ai)@op7418 · 中文博主 · 1 天前歸藏,中文圈 AI 工具与提示词博主

周一晚上还有一次重置:

引用 Tibo @thsottiaux我会在周一做另一次表演性的重置。查看被引原帖 ↗
Replit ⠕@Replit · 公司官方 · 1 天前AI 编程平台 Replit 官方

时间真的不多了各位!Replit Designathon最后的报名机会。美国太平洋时间周一上午10点截止提交(8月11日)。具体内容见下面这个线索🧵

查看英文原文
Just hours left on the clock. This is your last call to enter the Replit Designathon.

Submissions close at 10am PT on Monday, August 11.

Thread below 🧵
Chubby♨️@kimmonismus · 博主 · 1 天前Chubby,高频 AI 新闻聚合博主

Anthropic的发布节奏确实有点奇怪。Sonnet 5根本没意思,Haiku更是没人记得了。

查看英文原文
really strange releases form anthropic. sonnet 5 was just not good and haiku is literally forgotten.
Dan Shipper 📧@danshipper · 博主 · 1 天前

我第一次读《悲惨世界》,发现和《战争与和平》有很大重叠。

托尔斯泰肯定(从好的角度)借用了不少场景、人物和主题。

题外话:我也在用 ChatGPT voice mode 辅助平行阅读法文原文的某些段落,真的绝了,强烈推荐

查看英文原文
im reading Les Miserables for the first time and there’s SO MUCH overlap with War and Peace.

Tolstoy definitely ripped off (in a good way) a few scenes, characters, and themes.

side note: Have also been parallel reading certain scenes in the original French with the help of ChatGPT voice mode and it is SICK, highly recommend
GitHubDaily@GitHub_Daily · 中文博主 · 1 天前

Claude Code 用到一半,弹出「额度用完了」,要等好几个小时才能继续。

claude-code-local 让 Claude Code 跑在 Mac 本地,用自己的芯片驱动开源模型,不走云端也不用额外订阅。

别的方案中间要加一层翻译代理,它省掉了这层,同样的任务快了将近 8 倍。

GitHub:
github.com/nicedreamzapp/cla…


重要的是 16 GB 的 MacBook Air 就能跑,内存大的还能上千亿参数的模型。

断网也能用,代码不出本机,适合对隐私有要求或不想被额度卡住的朋友。

GitHubDaily@GitHub_Daily · 中文博主 · 1 天前

很多教程教 AI Agent 开发都是直接上框架,装完跑起来了,但底下到底怎么回事不清楚。

ai-agents-from-scratch 反过来,先从最基本的模型调用写起,一步步搭出工具调用、记忆、ReAct 循环这些核心模式。

全程用本地模型跑,不依赖云端接口,每个阶段都有代码和概念讲解。

GitHub:
github.com/pguso/ai-agents-f…


从加载模型、翻译、推理,到带记忆的 Agent、ReAct 模式、错误处理,总共 11 个递进式示例。

想搞懂 AI Agent 底层原理再去用框架的朋友,这个项目适合拿来入门。

OpenRouter@openrouter · 公司官方 · 1 天前

一个请求最多支持 30 张图片、10 个视频片段和 10 个音频片段作为参考输入。可编辑原视频的视觉或音频效果,如替换视频主体、添加或删除或修改对象。openrouter.ai/bytedance/seed…

查看英文原文
One request can support up to 30 images, 10 video clips, and 10 audio clips as reference inputs.

Edit the visuals or audio of the original video, such as replacing the video subject, adding, deleting, or modifying objects.


openrouter.ai/bytedance/seed…
GitHubDaily@GitHub_Daily · 中文博主 · 1 天前

在 Mac 电脑上找个文件,Spotlight 搜半天搜不到,名字记不全更是白搭。

Cling 是一款 macOS 上的文件搜索工具,支持模糊匹配,打错字、只记得一半名字也能找到。

搜索速度很快,900 多万个文件里定位一个,不到 100 毫秒。

GitHub:
github.com/FuzzyIdeas/Cling


启动时默认显示最近改过的文件,找到以后可以直接拖拽、用快捷键操作或者跑自定义脚本。

经常翻文件翻到头疼的朋友,装一个试试。

Together AI@togethercompute · 公司官方 · 1 天前

模型部署到生产时,你希望测试时的质量能贯穿全套服务栈。@Kimi_Moonshot在各大推理提供商上对Kimi K3做了基准测试,Together AI在4个基准里拿了3个第一名或并列第一。

查看英文原文
When you move a model into production, you want the quality you evaluated to carry through the serving stack.


@Kimi_Moonshot
benchmarked Kimi K3 across major inference providers, and Together AI ranked #1 or tied #1 on 3 of 4 benchmarks.
OpenRouter@openrouter · 公司官方 · 1 天前

从我们的 API 文档开始。openrouter.ai/docs/guides/ov…

查看英文原文
Get started with our API docs below.


openrouter.ai/docs/guides/ov…
Together AI@togethercompute · 公司官方 · 1 天前

我们用 Seedance 2.5 在 Together AI 上,只用一个提示词就生成了一个 30 秒的 lost-cinema 风格电影预告片 📽️

查看英文原文
We asked Seedance 2.5 on Together AI to generate a 30-second lost-cinema trailer in one prompt 📽️
GitHubDaily@GitHub_Daily · 中文博主 · 1 天前

做 Agent 开发做到后面,怎么让模型在多轮对话里记住之前说过的东西,是个绕不开的问题。

Awesome-AI-Memory 把 AI 记忆相关的论文、框架和评测数据集整理到了一起,目前收录了 540 多篇论文和 110 多个开源项目。

内容按记忆类型、存储方式、管理策略分了类,检索增强、记忆压缩、多 Agent 共享记忆这些方向都有覆盖。

GitHub:
github.com/IAAR-Shanghai/Awe…


资料一直在更新,最近一次还补了 25 篇论文。

在研究大模型记忆机制或者自己搭记忆系统的朋友,这份合集省得到处翻。

Replit ⠕@Replit · 公司官方 · 1 天前AI 编程平台 Replit 官方

那到底什么样的作品能得分?评委看重这几点:
• 第一印象:与众不同、抓眼球、价值主张清晰
• 设计水准:头到尾保持一致和精致
• 功能完整度:一个完整、令人满意的产品
额外加分:在公众面前构建

查看英文原文
So what actually scores? Here’s what the judges are looking for:

• First impression: distinct, scroll-stopping, a clear value prop
• Design craft: consistent and premium, end to end
• Functionality: a complete, satisfying product

Bonus for building in public.
Replit ⠕@Replit · 公司官方 · 1 天前AI 编程平台 Replit 官方

现金大奖50K+。一等奖16K,另外还有9个奖项。8月14日现场揭晓获奖名单。机不可失!美国太平洋时间周一上午10点前报名:buildathons.replit.app

查看英文原文
$50K+ in cash and prizes. Grand prize of $16K, plus 9 more categories.

Winners announced live August 14.

Don't miss it. Enter before Monday at 10am PT:


buildathons.replit.app

本站由 Jedee杰哥 打造 · 公众号「Jedee杰哥」每早送 AI 日报

姊妹站:𝕏 简中账号数据榜单 · X 关注 @jedeeai · RSS 订阅 · AI 日报 · 历史归档