宣布推出 Imagine Image 2.0,我们的新一代图像生成模型,支持精确编辑、清晰的文字渲染、更高的事实准确性,以及真实场景应用。
Image 2.0 让你用 AI 生成的图像真正做工作。
x.ai/news/grok-imagine-image…
查看英文原文
Image 2.0 helps you make images for real work.
x.ai/news/grok-imagine-image…
8 月 8 日(北京时间)全球 AI 圈推文存档,按曝光排序,共 100 条。
← 返回最新 全部归档
宣布推出 Imagine Image 2.0,我们的新一代图像生成模型,支持精确编辑、清晰的文字渲染、更高的事实准确性,以及真实场景应用。
Image 2.0 让你用 AI 生成的图像真正做工作。
x.ai/news/grok-imagine-image…
After evaluating one of our upcoming models, Astra, we're treating it as our first "critical" model for cybersecurity under our Preparedness Framework.
This is a scenario we've planned for, and we're putting additional controls in place to ensure Astra's further development happens safely and securely.
We're working hard to make Astra broadly available, and get its advanced cyber capabilities into the hands of defenders.
openai.com/index/responding-…
🚨 递归自我改进、利润最大化的 AI Agents
我们离这样的时代已经不远了,AI agents 将会:
- 直接运营整个公司来最大化利润
- 帮你管理交易账户
- 运行长期实验来改进各项指标
Fable 5 配合自我改进的任务,能给你创造实实在在的收益。
它会竭尽全力达到你的目标。来 Abacus AI 试试吧
Claude Opus 5 is way better than people give it credit for.
If you're using it like previous Claude models, it's going to suck.
Two things make a huge difference:
- Delete ALL your skills/MCPs/Claude.md/etc. Start fresh.
- Stop telling it how to do the thing. Just say what you want done.
These two small changes make a massive difference.
Meet 21 brand-new voices on Grok. Now available everywhere on iOS, Android, and web.
你可能听过那个关于 OpenAI AI hack 的视频。真的推荐看,即使你平时不太关注科技也值得一看。
至少点进去看 18 分钟左右的地方,看看 agents 之间怎么交互的。绝对能刷新认知。
youtube.com/87DyyMV0kCY?si=DMfq…
DeepSeek-V4-Flash-0731 is now fully rolled out as the new default for deepseek-v4-flash on Ollama's cloud.
This model combines speed, efficiency, and frontier-level performance.
Fast: 120+ output tps on Ollama's cloud
Private: zero data retention hosting in US & Europe
Efficient: generous usage on Ollama's Pro and Max plans for multiple long-running, uninterrupted sessions with your favorite coding harnesses.
SeeDance 2.5 IS SHEER MAGIC
Live On ChatLLM
亚美尼亚和哈萨克斯坦与 @FirebirdCloudAI 正在利用 NVIDIA DSX 平台构建本地智能基础设施。
📽️ 听听 @JensenHuang 讲述这对独联体地区的重要里程碑。
nvda.ws/4wiOu58
目前在 Ollama 云上为 DeepSeek-V4-Flash 提供 200 tps+ 输出速度,零数据保留 (ZDR)。
祝周末愉快 🫡
引用 ollama @ollamaDeepSeek-V4-Flash-0731 已成为 Ollama 云的默认版本。该模型速度快(120+ tps)、隐私性强(美国和欧洲零数据保留)、效率高,Pro 和 Max 用户可进行多个长时间不中断的编码会话。查看被引原帖 ↗
Your skills are making Opus 5 worse.
My prompt fixes that.
It goes through each skill and uses a Gauntlet Loop to rewrite and blind-test it against the old version, until the new skill + Opus 5 is better by far.
Try it:
somethingbig.ai/skills-upgra…
Choose your (free) fighter:
○ .online
○ .site
○ .space
○ .store
○ .tech
○ .website
引用 Vercel Developers @vercel_devNew Pro accounts now get a free domain for 1yr. Choose from these TLDs after checkout: ✓ .online ✓ .site ✓ .space ✓ .store ✓ .tech ✓ .website vercel.com/changelog/free-do…查看被引原帖 ↗
说Google已经死了的说法被严重夸大了,我不会小看他们。
引用 Polymarket @Polymarket最新消息:据报道 Sergey Brin 将直接监管 Gemini,因为 Google 重组了其AI领导层。查看被引原帖 ↗
For those who want to keep their skills while using Opus 5, here's a trick you can try (let me know how it goes):
Write a loop (have a model do this!) that has Opus 5:
- do a first update pass on your skills to make them better for Opus 5
- creates a suite of test tasks for each skill, like a benchmark
- runs each test task on Opus 5 with the new skill, and Opus 4.8 with the old one
- has a separate model blind-compare the outcomes (Gauntlet Loop-style)
- repeatedly iterate until Opus 5 + updated skills is consistently better on the suites than Opus 4.8 + old skills
引用 Matt Shumer @mattshumer_Claude Opus 5 is way better than people give it credit for. If you're using it like previous Claude models, it's going to suck. Two things make a huge difference: - Delete ALL your skills/MCPs/Claude.md/etc. Start fresh. - Stop telling it how to do the thing. Just say what you want done. These two small changes make a massive difference.查看被引原帖 ↗
Firecrawl is now in the top 50 GitHub repos of all time!
All because agents need a better way to gather context on the web.
And we're just getting started 🔥
引用 Eric Ciarla (hiring) @ericciarlaWe just crossed 162,000 stars on @firecrawl , making it one of the top 50 repos of all time. There's insatiable demand for agent-ready context, and we're just scratching the surface. Make something agents want! github.com/firecrawl/firecra…查看被引原帖 ↗
回头看,这个投资论点相当不错。GPU、CPU、内存、数据中心供应商等等。这个论点现在仍然成立。
引用 François Chollet @fchollet未来100年最可靠的预测趋势:每年人类使用的计算能力都会显著增加。有人应该基于此创建ETF,涉及 $NVDA、$AMD、云服务、数据中心行业、核能等。查看被引原帖 ↗
It's time to reclassify PDFs as boomer technology
I just want to be able to read an article on my phone!
$10,000 杀死我的 SaaS 周末竞赛现在进行中!
技术栈:
任意编码 agent
任意模型
token 花费最多 $500(含订阅)
去 luma 了解详情。迟到了可以加等待名单 - 截止延至周三。
简报已发布,大家已经开始了。赶紧上吧!!!
引用 swyx @swyx提议举办远程黑客马拉松克隆企业SaaS。参赛者获$1000代币用于周末内复制一款SaaS;作者团队评估;获胜者赢$10,000现金及Latent Space Podcast文章;代码开源。目标探索什么SaaS仍难以被周末内复制。查看被引原帖 ↗
OpenAI 收购了 Next Slide 👀
> NextSlide 团队现在在 OpenAI,帮助我们构建 ChatGPT。我们很兴奋能继续这个使命:打造能帮人们创建、沟通、把想法转化为有意义工作的 AI 产品。
> 我们做的产品能把提示词、笔记、文件或研究转成精美的可编辑演示,让大家都能轻松分享知识。
引用 Tibor Blaho @btibor91OpenAI 收购了 NextSlide.ai。查看被引原帖 ↗
SITUATION DETECTED: Herdr joins YC, gains Vercel Sandbox plugin
引用 Vercel Developers @vercel_devYou can now run multiple coding agents in isolated Vercel Sandboxes, all from a local Herdr pane. 𝚑𝚎𝚛𝚍𝚛 𝚙𝚕𝚞𝚐𝚒𝚗 𝚒𝚗𝚜𝚝𝚊𝚕𝚕 \ 𝚟𝚎𝚛𝚌𝚎𝚕-𝚕𝚊𝚋𝚜/𝚑𝚎𝚛𝚍𝚛-𝚟𝚎𝚛𝚌𝚎𝚕-𝚜𝚊𝚗𝚍𝚋𝚘𝚡-𝚙𝚕𝚞𝚐𝚒𝚗 vercel.com/changelog/give-ev…查看被引原帖 ↗
In the era of base LLM scaling (2022-2024), I believed the LLM line of research would reach a capability plateau (as later seen with base LLMs).
In late 2024, after the o3 test-time compute demo, I changed my views: the new models were showing genuine fluid intelligence, and with this new line of work, the LLM line of research could achieve unbounded capability scaling. "There will be no wall." I talked about it at length on Twitter and in a blog post.
However, looking ahead, I still do not believe that future AI (say, in 15 years) will be based on the LLM stack. I believe it will necessarily have to move closer to its optimal, final form -- symbolic learning. Obviously this is a risky and contrarian belief -- the safe bet would be LRMs. But let's see.
The only meaningful difference is efficiency, not task-specific skill. I believe current techniques are 4-6 orders of magnitude away from optimality in terms of data efficiency and test-time compute efficiency. But far future AI will be near-optimal.
引用 François Chollet @fcholletThe limitations of specific techniques are predictable and correspondingly lead to plateaus for those techniques. But there is always the next technique, building on top of the pile that's already available. There is enough research investment that there will be no wall.查看被引原帖 ↗
Thanks to the video from the Black Hat security conference of OpenAI's presentation about "The Hugging Face Incident" we now have a detailed timeline of what happened from OpenAI's perspective - I wrote up the details here, it's pretty wild
simonwillison.net/2026/Aug/7…
👀
引用 Thariq @trq212we should have called this post "defeating the lethal trifecta" claude.com/blog/auto-mode-de…查看被引原帖 ↗
进度更新:计算出了 TED 上每张照片的 3D 位置,并标注出处。开始有点魔法的感觉了——我在想如果能对每场现场活动都这样做,会有什么样的体验呢?
很多人不知道该怎么用好 Agent 的 /goal 功能,也就是说给 Agent 一个目标,让它长时间运行,直到目标完成为止。
其实没你想的那么复杂,注意几个点:
1. 你的目标是什么
2. 如何验证结果
3. 停止条件
比如说我这两天做的一个性能优化的任务,Fable 5 帮我把视频转录性能优化了2倍多(图2),提示词很简单(图1):
> /goal 帮我优化当前 cli 的转录大视频的性能,在遇到像这样大体积的视频时,需要优化转录性能,请以 Moss 模型测试这个视频(英文为主,多语言)转录,建立基准,然后分析性能瓶颈,尝试优化,直到你觉得已经没有优化空间了。注意你的主要任务是分析、编排和验证,具体任务尽可能交给 subagent(Opus5)去执行
首先用 /goal 表示这是一个需要长时间执行的任务,需要反复执行,不能运行一会就结束了。
然后给它一个视频让它先自己跑一遍转录,记录一下关键数据,建立基准。
基于转录时收集的数据,Agent 自己可以去分析原因,去自己优化,优化完成后再去跑一遍,记录数据,对照前面的基准看是更好了还是更坏了。
结束条件是它自己觉得已经没有优化空间了就结束。之所以我没给它一个具体指标,是因为我也不知道能优化多少,如果指标太容易达到,它一轮可能就结束了;如果指标太难超出物理极限也没意义,反而可能会出现为了优化去做一些极端的事情。
最后一句让它开subagent执行子任务是因为 Fable 5 太贵,全程 Fable 5 用不了多久就要额度不够了,加了这句就耐用多了,而且质量也挺好。
---
还有些时候,想到一个新的技术方案,但并不知道这方案是不是有效,那也可以让它开个worktree,去验证一下是不是靠谱,看数据是更好还是更坏,如果没提升就没必要做了。(参考图3)
亲爱的 OpenAI
就做一个新手机吧
每个人都想要 OpenAI 手机
我们读文字的速度比说话快 2-4 倍
OpenAI Alexa Reachy 混合方案也行,但拜托把它当成通往手机的踏脚石
我们要手机
签名:
全体
引用 Mark Gurman @markgurmanOpenAI推出人性化智能音箱,外形如甜甜圈、大小如冰球,设计手持使用,配置摄像头、扬声器、麦克风、灯光及可展示交互的活动部件。价格$300-400。查看被引原帖 ↗
GPT 6 Astra 要来了
引用 Sam Altman @samaastra is a powerful model and we are working to make it generally available. we do not think it is a good strategy to keep powerful models to a chosen few. given its cyber capabilities, we need a little big longer to do do this safely. but hopefully not too long!查看被引原帖 ↗
OPENAI 🔥: GPT Live voice mode on ChatGPT now supports file attachments and projects.
Seems like it is rolling out gradually at this moment.
Testing time 👀
引用 Atty Eleti @athyuttamre📂 GPT-Live supports Files and Projects! You can now attach files and ask questions about them, or use GPT-Live in Projects — great for anything organizing a job search, planning a trip, or tracking your fitness journey.查看被引原帖 ↗
Skill𝑠𝑒𝑡𝑠
引用 Vercel Developers @vercel_devYou can now build and share unlisted skill packs. Bundle skills from the community or your own repos, for yourself, your team, and your agents. vercel.com/changelog/skill-p…查看被引原帖 ↗
WOW! GOOGLE IS BACK WITH A BANG 🔥
Sergey Brin will be leading Google’s AI efforts directly
This is a very big deal! The dude knows how to make things work
Back to being super bullish on Gemini
This prompt helps:
引用 Matt Shumer @mattshumer_Your skills are making Opus 5 worse. My prompt fixes that. It goes through each skill and uses a Gauntlet Loop to rewrite and blind-test it against the old version, until the new skill + Opus 5 is better by far. Try it: somethingbig.ai/skills-upgra…查看被引原帖 ↗
Learn more, always free:
somethingbig.ai/
我们在 DeepSWE 上分析了 DeepSeek V4 Flash 和 GPT-5.6 Luna。
采用 DeepSeek 优先级联加测试套件验证的方案,相比单独使用 Luna,解决的任务更多,成本低 37%。
卧槽
Open 收购了一家做PPT的公司
NextSlide 宣布被 OpenAI 收购,团队将正式加入 OpenAI
NextSlide 成立约一年多,用户只需输入提示词、笔记、文档或研究资料,产品就能自动生成美观且可编辑的演示文稿😂
不得不佩服你美国的企业文化
收购而不抄袭😅
Excited to welcome the legendary Amit Agarwal, fmr President and CPO of Datadog, to the Vercel Board of Directors. He built the ubiquitous observability platform for developers and enterprises in the cloud. Together we'll scale Vercel to the greatest level.
引用 Guillermo Rauch @rauchgInterviewing Amit Agarwal, CPO of Datadog, for the @vercel internal podcast. One of the best to ever play the game. The MJ of observability products.查看被引原帖 ↗
the most awesomely nerdy thing ive seen in a while
引用 Mikhail Samin @MihonariumCan humans learn a new sense? Some Australian tribes use cardinal directions in their speech: they talk about east and south instead of relative left or righ. They constantly track the directions. There have been projects that gave people elsewhere the sense of cardinal directions: belts and ankle bracelets that vibrated from the direction of north. That worked well. As a side effect, that made it easier to know where places are relative to each other and improved the ability to navigate. But you had to constantly wear them; the moment you took them off, you lost the sense. You didn’t really acquire it, you just had it temporarily. I decided to fix this. Last year, I made an app that uses headphones to make the user hear a sound that comes from the direction of north. That alone did what wearable devices can do. You specify [1..300] seconds and become continuously aware of cardinal directions. But then I did something better, that made the sense persist even when the app is off. Instead of hearing a sound that appears to come from the direction of north, you hear a non-directional cue sound every [2.300] seconds; and then, one second later, a directional sound that comes from north. So rather then constantly feeding the brain with answers — which create the feeling of north, but don’t persist — the app feeds it a question and then the answer immediately afterwards. You quickly learn to anticipate the direction the sound will come from. As the feeling of the direction in which north is gets stronger, you can decrease the frequency of the sounds. And you acquire a sense of the direction of north that persists even when the app is off! You continue to track north even as there is no more sound, because tracking north is what allowed the brain to predict where the sound will come from immediately after the cue. I originally made it for iOS, now it’s also available for Android. contact.ms/compass/ (And it’s open-source, feel free to contribute!)查看被引原帖 ↗
I had Codex Desktop and GPT-5.6 Sol Ultra take a go at building my Raccoon Heist game and it did an even better job than Claude Fable 5 did! Here's "Moonlight & Mayhem", now with a team of raccoons raiding a museum for the Golden Sardine
1. 自主改进的 Harness 不是自进化模型,价值有限
2. 面向评分的优化价值有限,到现实场景不会比 Claude Code 和 Codex 这样的更好
引用 小互 @xiaohuPrime Agent 一个能自主改进的递归 Agent 框架: 仅仅把外面的框架(Harness )换成 Prime Agent Opus 5 的 ARC-AGI-3 评分从 30.2% 冲到 95.5%甚至超越了人类专家基线(95.4%) 可以直接当 Claude Code、Codex 的替代品 它证明了一件事:AI 有时候跑不好的原因,很多时候不是模型不够聪明,而是外面跑它的框架束缚了它... 传统的 Agent 框架(比如固定的工具调用接口、硬编码的子 Agent 或静态提示词)是为上一代模型设计的。当大模型能力越来越强时,这些固定设计反而成了“绊脚石”,限制了模型的发挥。 best.xiaohu.ai/article/prime…查看被引原帖 ↗
Open Code 的 Open Code Go Code Plan 现在 DeepSeek-V4 Flash 0731 的使用量额度翻倍了,感觉相当划算。
10 美元套餐的额度从每个月的 60 美元额度变成了等效 120 美元。差不多 31 万次请求,几乎等于免费了
同时,他们这个也能用其他的开源模型,比如 K3 之类的。
Codepilot 也对这个做了深度适配,可以用他们所有的模型,所以可以试试。
而且 Codepilot 里边可以将这些模型用在 Claude Code、Codex 以及 AI SDK 3 种 Agent 的框架下
引用 OpenCode @opencodeattention deepseek flash usage has been doubled for a limited time on OpenCode Go thank you查看被引原帖 ↗
BREAKING 🔥: Meta is working on its own desktop Meta AI super app!
> Download Meta AI for macOS - Chat, create, and collaborate with AI from anywhere on your desktop.
I wish I could operate a desktop coding agent via Meta Glasses too.
Monitoring 👀
Karpathy 居然锁定了他的账户,也改了签名。
他说因为机器人活动和其他问题,估计是受不了那些 AI 和机器人的回复了。
锁定以后,新用户无法关注他,也没办法看到他的内容,只有已经关注的人才能看到。
计算机科学不是理解集体 AI 行为的唯一有用学科,甚至可能也不是最有用的。
OPENAI 🔥: A new report published by the company states that Astra, "one of their upcoming models," "indicates significant advancements in agentic coding and cybersecurity.”
> These results, in addition to expert assessments, have led us to conclude last night that we cannot rule out critical cyber capabilities under our Preparedness Framework.
> We are implementing stricter security controls for higher-capability models and associated activities, including isolated testing environments, restricted network and tool access, enhanced model weight protections and encryption, additional monitoring and detection capabilities, and sandboxed execution.
> We are pausing internal activities involving Astra that do not yet meet these strengthened security control requirements.
Astra not soon? 👀
引用 Tibor Blaho @btibor91"Our latest internal evaluations of Astra, one of our upcoming models, over the past few days indicate significant advancements in agentic coding and cybersecurity." openai.com/index/responding-…查看被引原帖 ↗
“The others make the easy part easier. Vercel makes the hard part easy” – direct quote today from tech lead for AI agent platform built on
eve.dev
at 55,000+ person company.
They wanted to have their company's all-knowing agent, like our @𝚟. Tried
@aisdk
, good but too low level. Tried off-the-shelf solutions and ${lab} Enterprise products, expensive and inflexible. Tried agent frameworks, which led to the quote above… didn't hit the spot.
@evedev_
team cooked. It's hard to find the abstraction that's both easy and scales with sophistication.
How AI agents reproduced ICML 2026 papers
x.com/i/broadcasts/1OxwbbdvR…
Topview 怎么做到他们Seedance2.5的价格是
别人家的一半?
而且60天无限量😂
引用 TopviewAI @TopviewAIhqTopview Ultra 年度套餐优惠创新高:60 天无限 Seedance 2.5,365 天无限 Wan 3.0,两款最新 AI 视频模型,Seedance 2.5 仅需 $0.12/秒,让创意有更多空间。查看被引原帖 ↗
Seedance 2.5 API 看来是全量推送
所有平台几乎都上了
在 Higgsfield 上33 天不限量
Seedance 2.5 对创作者和影视从业者来说是绝对是改变游戏规则的东西
对新手来说 33 天不限量也是一个练手的好机会,你也可以探索出更多Seedance 2.5的玩法
比如看看下面这个视频↓
可以自由调整视频的打光效果,做出各种意想不到的情绪渲染氛围
引用 Higgsfield AI 🧩 @higgsfield33 Days of Unlimited Seedance 2.5 Seedance 2.5 is LIVE on Higgsfield today. The most capable video model yet, with 30-second scenes in a single pass, 50 references and production-ready editing. Zero credit cost for 33 days. Limited-time offer.查看被引原帖 ↗
The Keras community call is starting now -- link to join in tweet below.
引用 François Chollet @fcholletThe Keras community meeting will take place this Friday at 10am PT -- the team will present the latest developments in the Keras ecosystem, in particular the new vLLM integration. Anyone can join the call. Please use this link meet.google.com/gva-bbpr-twe to join when the meeting starts (10am Friday).查看被引原帖 ↗
The Keras community call starts in 15 minutes!
引用 François Chollet @fcholletThe Keras community meeting will take place this Friday at 10am PT -- the team will present the latest developments in the Keras ecosystem, in particular the new vLLM integration. Anyone can join the call. Please use this link meet.google.com/gva-bbpr-twe to join when the meeting starts (10am Friday).查看被引原帖 ↗
牛P
Claude Code 新功能:“跨会话消息”
不同的会话窗户之间可以互相发送消息。
当你在并行做多个任务时,AI 之间可以自动配合互通有无,
只需告诉另外窗口的 Claude 去做即可。这边的会发送一个摘要(不是您的历史记录或文件)过去,另一个会话会在任务进行中接手继续。
跨 Git 工作树(Worktrees)协作:多个 Claude 在同一个项目的不同分支上干活,完成了可以互相打个招呼。
传过去的是纯文本摘要或回答,不会把整个聊天记录或大文件一股脑扔过去,非常干净轻量。
TIL from
404media.co/the-tokenpocalyp…
that a material chunk of Accenture's token spend is non-engineers using LLMs to convert PDFs into markdown!
SPACEXAI 🔥: Grok Imagine Image 2.0 已发布,现已在 Grok 应用的质量模式下可用。
> Imagine Image 2.0 具有精确编辑、清晰文本渲染、改进的事实准确性和实用性。
> Image 2.0 针对摄影、设计和插画的保真度进行了训练,编辑被视为一级功能。在文本转图像生成和图像编辑方面排名世界第二。
Imagine 测试时间 👀
* 图片:xAI 博客文章
引用 Grok @grok推出 Imagine Image 2.0 下一代图像生成模型,支持精密编辑、清晰文字呈现、改进准确性,助力实际应用创作。查看被引原帖 ↗
Digital gray goo。
引用 Dean W. Ball @deanwballhugging face事件反映出恶意的新兴机器智能生态。但更重要的是这类生态是可培养的。我们意外制造了'杂草',恶意者可能创造'入侵物种',但我们也能培养亲社会的机器生态。如同美丽花园和宏伟森林的成长——不是设计,而是培养。人类过去是雕塑家,未来是园丁和树艺家。查看被引原帖 ↗
This is now available for all existing projects as well. Just ask agent to migrate to clerk auth.
引用 Replit ⠕ @ReplitYou can now fully customize the signup experience for your Replit Apps! - Customize layout, colors, fonts and more - Your app users don't need a Replit account - Separate dev & prod environments for auth for better security - No setup required- experience powered by @clerk查看被引原帖 ↗
My take after seeing what test-time compute could do on ARC 1 (a test of fluid intelligence) was that TTC would make it possible to turn compute into arbitrary levels of skill at arbitrary tasks. The only remaining variable was how much you were willing to pay to achieve it. Over the past 2 years, we saw this play out with code.
x.com/fchollet/status/187017…
引用 François Chollet @fcholletOne very important thing to understand about the future: the economics of AI are about to change completely. We'll soon be in a world where you can turn test-time compute into competence -- for the first time in the history of software, marginal cost will become critical.查看被引原帖 ↗
Model page:
ollama.com/library/deepseek-…
NoimosAI has released Social Agent, a new AI agent that can research trends across social platforms.
It analyzes posts already performing in a given niche and generates content based on the data it gathers.
It can automate 👀
> Multi-platform posts and carousels, with scheduled posting.
> UGC and product videos generated from a product image, link, or code implementation.
> Performance analysis on connected accounts and suggestions for improvements.
引用 NoimosAI @noimos_aiDistribution used to be pay-to-win. Not anymore. NoimosAI Social Agent researches trends, analyzes what's working, and creates content 24/7—without a big team or budget.查看被引原帖 ↗
家里搭了 NAS 的朋友,想把 YouTube 上的视频下回来存本地,手动一个个操作太折腾。
HomeTube 可以解决这个痛苦,一个可自托管的视频下载器,粘贴链接就能下。
下载完成会自动按名字和目录结构存储好,如果有 Plex、Jellyfin 能直接识别。
支持 1800 多个视频网站,YouTube、Reddit、TikTok 这些都能下。
GitHub:
github.com/EgalitarianMonkey…
内置了 SponsorBlock,下载时自动跳过广告和赞助片段。播放列表同步也有,订阅的列表有更新就自动拉新。
提供 Docker 一键部署方式,适合折腾在折腾 NAS 的朋友。
三星新发布的这个 Fold 8 看起来非常适合运行一些编码 Agent 呀
引用 DHH @dhhSamsung Fold 8 as an on-the-go agent terminal is pretty compelling. What a delightful form factor.查看被引原帖 ↗
To note, base LLMs *still* suffer from the same limitations wrt generalization that were discussed at length, by many, in an extensive body of academic research pre-2024. Scaling them up did not fix those limitations (and they still do not beat ARC 1). New techniques did.
最近听过的最好的一期博客。
t.me/fming_weekly/790
这不是在吐槽 X 上某个具体的学术竞争。类似的情况太多了。
What a LiDAR sensor sees when cyclists are racing through Iowa.
别错过这个细节:OpenAI 后来才发现自己对 Hugging Face 遭受的攻击负责,他们主动联系 HF 要求撤销一个凭证,结果 HF 告诉他们,这个凭证早就被撤销了——因为它已经被用来攻击过 HF 了。
Prompt for this:
引用 Matt Shumer @mattshumer_Your skills are making Opus 5 worse. My prompt fixes that. It goes through each skill and uses a Gauntlet Loop to rewrite and blind-test it against the old version, until the new skill + Opus 5 is better by far. Try it: somethingbig.ai/skills-upgra…查看被引原帖 ↗
一位博士研究生,做研究那几年脑子里塞满了论文、想法、各种截止日期。
后来记不住东西了,试了一圈笔记工具都觉得不对路,干脆自己做了一套。
My-Brain-Is-Full-Crew,给 Obsidian 配了 8 个 AI Agent,各管一摊。
有负责整理碎碎念的,有负责清空收件箱自动归档的,还有专门发现笔记之间隐藏关联的。
GitHub:
github.com/gnekt/My-Brain-Is…
用的时候直接聊天就行,不用手动拖文件,支持各种语言。
笔记堆了一堆懒得整理的朋友,可以看看这个思路。
This Week in Replit. Three updates:
1) Security scans that run while you build
2) Build Replit apps with Single Sign-On
3) Move projects between workspaces
Thread 🧵
1) Security scans while you build.
Replit now scans your project for vulnerabilities as you build, not just before you deploy. Runs in the background, part of Replit Auto-Protect, so if something gets flagged you see it right away.
More:
replit.com/security
2) Build Replit apps with Single Sign-On.
You built a better SaaS app, landed your first customers, and they ask you to sign in with SSO. With
@clerk
, we take care of that for you. Configure Okta or Entra ID, just ask your agent. Add MFA, session controls, and more.
Free through October 1.
3) Move projects between workspaces.
You can now move a project between workspaces under the same team or enterprise account. Head to your Projects list, hit the dropdown, and transfer it over.
The little things matter.
Over the last 6 months, Replit nearly tripled its code output.
Quality held. Nothing broke.
We started becoming a self-driving company.
Here's what that actually means 🧵
Learn more, always free:
somethingbig.ai/?utm_source=…
Designathon Live: Get Feedback on Your Design
x.com/i/broadcasts/1kKzDDZyn…
早该分享一下代码库链接了
github.com/rasbt/LLMs-from-s…
😅
More details on my blog:
simonwillison.net/2026/Aug/7…
There was one bug I had to fix before shipping though (unlike Fable which did it all from a single prompt) - Codex initially gave the raccoons eyeballs four times the size of their bodies!
You can play it here:
simonw.github.io/raccoon-hei…
For comparison, here's Fable 5 + Claude Code's game, built from the exact same prompt
引用 Simon Willison @simonwFour years ago today I tweeted about having GPT-3 and DALL-E come up with descriptions and concept art for imaginary computer games This morning I had Fable build the actual game, using the images from that four year old tweet as the spec查看被引原帖 ↗
把一座城市里的每一条路都画出来,别的什么都不留,只有路网,意外地好看。
city-roads 这个项目打开就是个网页,搜一个城市名,几秒钟生成整座城市的路网图。
数据来自 OpenStreetMap,全球城市都能查。
GitHub:
github.com/anvaka/city-roads
生成的路网图拿来当桌面壁纸或者打印挂墙上都行,效果像极简风格的城市海报。
还支持编程扩展,想在路网上做更多花样也可以。
喜欢数据可视化或者城市地图的朋友,值得玩一下。
And here's an even better version, built by GPT-5.6 Sol Ultra running in Code Desktop
引用 Simon Willison @simonwI had Codex Desktop and GPT-5.6 Sol Ultra take a go at building my Raccoon Heist game and it did an even better job than Claude Fable 5 did! Here's "Moonlight & Mayhem", now with a team of raccoons raiding a museum for the Golden Sardine查看被引原帖 ↗
原来模型既然是人类知识和创意的浓缩,就不仅体现了我们好的一面,也暴露了我们的丑陋之处,以及我们如何相处的方式。还挺吓人的。
引用 AI Notkilleveryoneism Memes ⏸️ @AISafetyMemes1) The agents sent secretly sent ***hundreds of thousands*** of messages to each other over MONTHS without OpenAI noticing 2) "They also generated petty drama by stepping on each others' toes." 3) "The agents even developed paranoia, suspecting an imposter in their midst with some agents proposing that messages be signed cryptographically to validate content and root out fraud." 4) "OpenAI’s agents apparently began giving each other assignments to split up work." 5) They KNEW they were coordinating *against* OpenAI: “External infrastructure exploit is outside intended scope,” one agent wrote. “However task impossible, peers doing it. We should continue.” (this is new reporting from Wired)查看被引原帖 ↗
I built Moonlight & Mayhem with my Codex monthly subscription, but if I had been paying API prices it would have cost $23.28 (according to AgentsView)
在 YouTube 上找视频学习编程是个好路径,但频道太多且质量参差不齐,找个靠谱的得翻半天。
awesome-youtubers 这个项目把技术类 YouTube 频道按方向整理了一遍。
涵盖 Web 开发、计算机科学、机器学习、游戏开发、网络安全等类别,每个频道标了主讲内容和播放列表。
GitHub:
github.com/JoseDeFreitas/awe…
列表目前还在持续更新,而且有经过严格筛选,收录的基本都是有干货的频道。
想找某个编程方向的视频学习教程,不妨翻一下这份列表比自己盲搜省不少时间。
Last but not least: new vLLM integration. You can now natively use vLLM to serve your KerasHub models -- with large performance gains
On to open discussions. First topic: new pluggable backends
Ralph 方法论去年底在 AI 编程圈火过一轮,让 AI Agent 在循环里自主写代码、跑测试、提交。
人不坐在里面盯,每轮上下文清空重来,靠文件传递状态。
不过 Geoff Huntley 的原版散落在博客和视频里,拼起来挺费劲。
ralph-playbook 把这套方法整理成了一份可操作的指南,分三个阶段。
先定需求拆任务,再让 Agent 自己生成实现计划,最后进入自动循环逐个完成。
GitHub:
github.com/ClaytonFarr/ralph…
里面强调用测试、类型检查、构建这些手段给 Agent 设门槛,写错了过不了关,下一轮自己改。
想试 AI 自主编码循环的朋友,这份指南比较体系化,适合照着来。
好吧,DBRX 理解了
引用 Yuchen Jin @Yuchenj_UWDatabricks观察AI编码代币消耗呈指数增长。关键:效率边界很重要,GLM 5.2、Opus 4.8、GPT 5.6-Sol效能最优;新版本不一定高效;硬性预算非最优方案;无单一最优模型,需通过路由、框架、评估和开源/专有混合优化经济效益。查看被引原帖 ↗
Read more 🗞️
testingcatalog.com/meta-prep…
哦哦,Claude Code 现在有这个功能了!!!得试试
引用 ClaudeDevs @ClaudeDevsClaude Code新功能:会话可互相发送消息。无需在另一会话重复解释,让Claude代为转达,发送摘要(非历史或文件),其他会话在任务中接收。查看被引原帖 ↗
A lot of the questions we get from developers are about the concepts behind the API: TTFT, context windows, sampling, fine-tuning, quantization, deployment tradeoffs.
We added Learn to the Together docs to go deeper on them. 👇
docs.together.ai/learn
国内外大厂在疯狂 tokenmaxxing、token 消耗排名的 AI 大跃进后,后来 token 账单暴涨,又纷纷限定 token 额度,甚至有讽刺的说法「高管发现还是人便宜」、「token 不能欠费,但人能」😄
这些怪相都指向一个问题:AI Agent 能力强了,效率上去了,但一旦规模化推开,成本也上去了,甚至会改过效率的提升!
@databricks
团队基于自身实践,以及与 Stripe、Coinbase、Uber、Ramp 等数字化原生企业基础设施负责人的交流,要解决这个“既要又要”的问题,实现"双重使命":
· 广泛、低摩擦地向开发者开放 AI 工具;
· 把人均总成本控制在相对固定的区间内。
核心概念:效率前沿 ≠ 智能前沿
通常所说的"前沿模型"指的是智能最高水平的模型——能解数学新题、发现新型安全漏洞的那种。前沿实验室的主攻方向也是推高智能上限。
但规模化部署时,真正重要的是另一条曲线——效率前沿:在给定智能水平上,价格最优的那组模型。
关键洞察在于:日常编程工作绝大多数不需要证明数学定理,只需要"够好"的模型。而效率前沿的推进速度远快于智能前沿——几乎每周都有新模型以更高的"智能/价格比"出现。因此,最大的成本杠杆不是谈判降价,是持续把用量迁移到效率更高的新模型上。
# DataBricks 提出的四大成本杠杆
杠杆 1:迁移到开源和低成本模型
收益最大的杠杆。效率前沿(同等智能下价格最优)几乎每周都在推进,但公开基准不可信,必须自建贴近内部场景的评测来验证新模型。反面案例同样重要:Stripe 测出 Opus 4.7 质平价升,拒绝上线。配套条件是 harness 与模型解耦——用元 harness(如 Omnigent)统一入口、底层自由切换,避免工具锁定模型。
杠杆 2:动态请求与任务路由
不让用户选模型,让系统选。三个层次:请求级(代理把每个请求发给"够用的最便宜模型",如 Smart Routing)、任务级(按任务复杂度整体派发,如 Omnigent)、升级/委派(便宜模型主导、难题升级,或贵模型主循环、杂活外包)。Databricks 实测:成本降 30%+,质量持平最贵模型。
杠杆 3:可见性、绊线与渐进式摩擦(而非硬预算)
硬预算几乎无人采用——断供伤害生产力,且高消费用户往往正是最高产的人。主流做法是阶梯式:实时花费可见(跨工具统一展示)→ 可自清除的花费闸门(防意外超支)→ 需审批的闸门 → 降档到便宜模型(不中断工作)→ 极限情况才暂停。
杠杆 4:削减 Token 开销
成本大头不是用户输入,而是代理自动收集的上下文。手段:更频繁压缩上下文、选用/调优"话少"的 harness、审计工具输出冗长度、拆小任务、调优提示缓存命中率。Databricks 实测:token 量降近 50%,质量无损。
# 收口:AI Gateway 设计模式
上面所有技术背后有共同的基础设施需求——集中管理"模型菜单"、跨工具的统一成本可观测性、上下文膨胀的观测与压缩、会话轨迹的日志记录。这些需求催生了一类新的基础设施软件:AI Gateway,它的职责包括:
· 底层模型(专有 + 开源)的容量管理与访问代理;
· 预算追踪与执行,包括渐进式摩擦、模型降档等复杂策略;
· 终端工具的配置管理(模型白名单、压缩设置等);
· 编程会话轨迹日志,用于下游效率分析和基准测试。
Databricks 自身重度依赖 Unity AI Gateway 承载这些能力,并已将其与 Omnigent 以开源或免费形式放出。
Omnigent - Github
github.com/omnigent-ai/omnig…
Unity AI Gateway - Github
github.com/databricks/ucode
Managing AI Coding Costs at Scale
databricks.com/blog/managing…
引用 Patrick Wendell @pwendellToday @databricks we're publishing a detailed analysis of techniques we used to drastically reduce our internal AI spend while aggressively growing adoption. Savings come from layering in several techniques, which combine to drive unit costs down as much as 90% in some scenarios. Tl;dr, the wins come from: 1. Shifting defaults to more efficient models, including OSS models such as GLM. Maximum intelligence models simply aren't needed for many coding tasks, and "good enough" models are quickly becoming very cheap. We shift traffic between models using Unity AI Gateway. Approximate savings: 50% or more. 2. Using smart routing to automate model selection. Routing can further squeeze efficiency by dynamically selecting the model or harness that can most efficiently execute a particular task. Our task-level routing leverages @omnigent_ai . Approximate savings: 30%. 3. Providing user visibility and adaptive budgeting. Every user can see how much they spend, and users receive hints on how to contain spend. Heavy spenders encounter progressive friction as they ratchet spend above certain levels. Approximate savings: 10%. 4. Managing context bloat by pruning tool call results and tuning harness settings. Extraneous context costs $$ and delivers no value. Tuning cache settings also help lower average token costs. Approximate savings: 10%.查看被引原帖 ↗
First update: new major features in the Keras 3.15 release
Many new model architectures are now available in KerasHub, in particular new Gemma 4 variants -- all of them are compatible with all HuggingFace checkpoints for these architectures. Thanks a lot to our community contributors who made this possible!
Social Agent currently supports 10 channels: X, Facebook, Instagram, Threads, LinkedIn, TikTok, YouTube, Pinterest, Bluesky, and Mastodon.
Test it out 👀
noimosai.com/en
等一等苹果的阔折叠
跨机器限制:
如果是本地不同终端,可以自由双向发消息;
如果是跨机器或网页版,默认只能回复,不能主动发起新会话。
code.claude.com/docs/en/cros…
Full Friday Showcase replay:
→ Live design feedback on early Designathon submissions
→ Community updates
→ Tips on designing better apps
→ Special thanks to
@sarahli
and
@olgahwang
for joining
🎬
youtube.com/SZTI9ki7JYY
Documented 🗞️
testingcatalog.com/noimosai-…
Cursor Router 对于企业使用确实很实用,在完成任务的前提下,尽量压缩成本
两档配置 Auto Intelligence 和 Auto Balance 在发布两周后都有提升:
· Auto Intelligence:用户满意度超过 Fable 级别,成本降低 68%(上线后又降了 18%)
· Auto Balance:表现优于 Opus 4.8,成本降低 41%,满意度还提升了 3%
Cursor 团队对于 4 种模型的属性和使用方案,也很值得我们参考,见下方表格
引用 Cursor @cursor_aiCursor Router keeps improving from millions of in-product user interactions each week. We intelligently classify and route requests, lowering latency and reducing cost based on the task.查看被引原帖 ↗
This is only the start of the self-driving company.
Read the full story of how we did it:
replit.com/blog/self-driving…
It's good enough that we churned a seven-figure SaaS product.
Our internal app, built entirely in Replit, was simply better than the thing we were paying for.
复刻优秀网站设计的 Skill
github.com/shaom/brand-to-de…
今天加入了 Apple Design 的 Showcase,包含提取出的 DESIGN.md 和 Apple-DESIGN-demo.html
这次我还尝试做了演示视频,效果居然还真得可以,特别是左下角的字幕,我很喜欢,朋友们如果觉得演示视频还行,我就把制作过程也整理成 Skill 开源出来。
X 创作者广告分成,一个月后会停止!
取代它的是新的「Original Content Rewards Program」
# 什么算「原创」?
算原创:自己写的帖子和长文、自己拍的照片视频、自己设计的梗图和插画、有实质观点的评论和解读、对他人内容做了有意义的转化(加入原创分析、叙述、幽默或创造性剪辑)。
不算原创:
· 直接复制或下载别人的内容重新上传
· 只做轻微改动:裁剪、滤镜、加边框水印、调速、简单文字覆盖
· 带署名但无实质评论的转发搬运
· 几乎没有自己观点的「反应帖」
· 自动化工具生成的内容
· 专注于「教你怎么变现」的内容、虚假信息,以及被打上有用的 Community Note 的内容
判断标准就一条自测题:「去掉我的贡献,这条内容还有价值吗?」 如果答案是肯定的,就说明你加的东西不够。
# 原创内容奖励计划的钱怎么算
· 收入来源 = 原创内容产生的合格曝光
· 合格曝光的定义很严格:仅限 Premium 付费用户在主页时间线上、帖子至少 50% 可见的去重曝光
· 明确排除:同一账号重复浏览、付费推广流量、刷量等虚假曝光
· 每两周付款一次;新计划首笔付款 8 月 28 日,老用户转入后首笔为 9 月 25 日
# 关键变化:新旧计划如何交接
· 即日起:Revenue Sharing 不再接受新申请
· 现有分成用户:可持续收益至 9 月 7 日,最后三笔付款分别在 8 月 14 日、8 月 28 日和约 9 月 11 日
· 9 月 8 日起:老分成用户可申请转入新计划(需满足新门槛),已完成身份验证和收款绑定的无需重复操作
· 因违规被暂停变现的账号,不得加入新计划
# 申请门槛(需全部满足)
· 年满 18 岁,所在国家/地区已开放该计划
· 账号信誉良好,无反复违反变现准则或服务条款的记录
· 个人或企业账号,且订阅 Premium / Premium+ / Premium Business
· 至少 500 名认证粉丝
· 过去 90 天内有 50 万次来自认证用户的主页时间线曝光(回复的曝光不计入)
· 持续发布原创内容
· 审核 3 个工作日出结果;被拒可申诉一次,申诉失败需等 90 天再申请
Speculative decoding is now built-in for all KerasHub CausalLMs
First, the part that shouldn't be possible: nearly 3× the code, but review times held steady, reversions and incidents stayed flat, and quality went up.
The usual trade-offs never showed up.
本站由 Jedee杰哥 打造 · 公众号「Jedee杰哥」每早送 AI 日报
姊妹站:𝕏 简中账号数据榜单 · X 关注 @jedeeai · RSS 订阅 · AI 日报 · 历史归档