JEDEE AI
存档 2026-08-13

8 月 13 日(北京时间)全球 AI 圈推文存档,按曝光排序,共 100 条。
← 返回最新 全部归档

全部情报 每小时更新 · 事件已合并同类项

内容 公司
DeepSeek@deepseek_ai · 公司官方 · 1 天前深度求索,国产开源大模型标杆

🧩 DeepSeek Harness v0.1 现在开放开发者预览版了!

🔹 我们正在向全球构建 agent 开发平台的开发者开放,并以 MIT 许可证开源整个代码库。
🔹 基于 Cordis 元框架打造,DeepSeek Harness 的核心理念就一个:万物皆插件。模型、工具、技能、会话、沙箱、文件系统、循环、编排和 UI 全部都实现为插件,可以任意混合、匹配、替换和扩展。

现在就来试试吧!

github.com/deepseek-ai/deeps…

查看英文原文
🧩 DeepSeek Harness v0.1 is now available in Developer Preview!

🔹 We’re opening it up to developers building agent harnesses worldwide and open-sourcing the codebase in MIT license.
🔹 Powered by the Cordis meta-framework, DeepSeek Harness is an agent harness built around one core idea: Everything is a plugin. Models, tools, skills, sessions, sandboxes, filesystems, loops, orchestration, and UI are ALL implemented as plugins, and can be mixed, matched, replaced, and extended.

Try it now!

github.com/deepseek-ai/deeps…
lidang 立党 (劝人卖房/学CS/买SP500/纳100/OpenAI/Anthrop第一人)deepseek harness的paper光公式、定义、lemma就整整推了七八十条,将近100多页的公式,a…23 小时前 · 14 万向阳乔木Deepseek 的Harness发布了,已安装好了。1 天前 · 6.4 万Chubby♨️DeepSeek 已经开源了 DeepSeek Harness,他们用来构建和运行 AI 代理的框架。现在正在深…1 天前 · 5.2 万歸藏(guizang.ai)流传的截图,说是 DeepSeek Harness 会在今天公测,不是要跟 0813 的模型细节一起发布吧1 天前 · 3.5 万AIGCLINKdeepseek-harness重磅开源,采用了一切皆插件的架构,也就是中国版的openclaw,体验了下可以称…1 天前 · 3.3 万歸藏(guizang.ai)Deepseek Harness 0.1 版本正式发布了,而且开源!1 天前 · 3.2 万歸藏(guizang.ai)Deepseek Harness 内测群说晚上 8:25 左右发布。1 天前 · 2.8 万向阳乔木Deepseek Harness 怎么感觉瞬间 7k Star,太猛了。1 天前 · 2.7 万小互DeepSeek-Harness 正式发布:1 天前 · 1.8 万宝玉DeepSeek Harness 正式发布了,也是开源的1 天前 · 1.8 万yihong0618看起来能在1小时内突破10000颗星.....1 天前 · 1.7 万向阳乔木感觉Deepseek Harness的插件系统很好玩,虽然DSH现在是半成品。1 天前 · 1.5 万歸藏(guizang.ai)别的不说,Deepseek harness 这个 star 的涨势是真猛,一个多小时都快两万 star 了1 天前 · 1.2 万elvis推荐看看。不错的开始是看到 DeepSeek Harness 专注于让一切都成为插件。这就像 Pi 是如何构建的…23 小时前 · 9,694AIGCLINK刚刚,DeepSeek的Harness已发布,跟其他“Harness们”不一样的地方是:造了一套造Agent的系…1 天前 · 4,093小互另外DeepSeek-Harness已经泄露1 天前 · 2,973meng shaoDeepSeek Harness 终于来了!1 天前 · 1,079
◔ 377.2 万 次浏览(18 条合计)♥ 1.8 万⇄ 2,184新品看原帖 ↗
Tibo@thsottiaux · 公司官方 · 1 天前ChatGPT 产品官方账号

其实是好几天前的旧闻了,但已经突破了 15M。大家尽享一番重置吧。大约下一个小时内着陆,使用 /fast 开启。

引用 Tibo @thsottiaux我之前承诺Codex每增加100万活跃用户就有一次重置,上限1000万。我们已经突破了这个数字,之后就没动静了。明天会有小惊喜给大家。查看被引原帖 ↗
查看英文原文
Old news actually from a bunch of days ago, but crossed that 15M. Enjoy a nice reset everyone. Landing in the next hour or so, go /fast.
DeepSeek@deepseek_ai · 公司官方 · 1 天前深度求索,国产开源大模型标杆
连环推 ×2

今天我们要发布DeepSeek-V4-Pro!🚀

🔷 重点升级Agent能力,带来强劲的生产力提升!
🔷 V4-Pro与V4-Flash支持灵活推理强度:简单任务用低档,日常Agent工作流用高档,复杂任务直接拉满。
🔷 原生支持OpenAI Responses API,为Codex优化,一键配置即用。

V4 Pro现已上线app和网页端,通过“专家模式”即可体验。
V4 Pro也同步开放API,模型名称保持不变——配置详情请参阅API文档。

查看英文原文
We’re launching DeepSeek-V4-Pro today! 🚀

🔷 Major Agent upgrades with strong production gains!
🔷 Flexible reasoning effort for V4-Pro & V4-Flash: low for simple tasks, high for daily Agent workflows, max for complex tasks.
🔷 Native OpenAI Responses API support, optimized for Codex with one-click setup.

V4 Pro is now available on app/web. Try it via “Expert Mode”.
V4 Pro is also available via API. Model names remain unchanged—please refer to the API docs for setup details.
API pricing update 💰

With the V4 lineup release, we’re updating our API pricing and introducing peak and off-peak rates. Off-peak rates are 50% lower than peak, enabling more flexible workload scheduling. 📉

New pricing takes effect at 16:00 UTC, Aug 16, 2026 🕒
OpenRouterDeepSeek V4 Pro 0813 已在 OpenRouter 上线了。1 天前 · 13.5 万歸藏(guizang.ai)V4 Pro 0813 的正式更新公告来了,同时涨价额度也确定了。1 天前 · 10.6 万歸藏(guizang.ai)DeepSeek V4 Pro 的 0831 版本不太对劲,他们是不是又把模型给撤了?1 天前 · 9.4 万Orange AI真的无语了,DeepSeek V4 Pro 高峰期照着 12 倍去涨价…1 天前 · 7.4 万Lisan al GaibDeepSeek又把DeepSeek-V4-Pro-0813的权重撤下来了23 小时前 · 7 万小互据说DeepSeek V4 Pro 版本发错了1 天前 · 5.3 万Chubby♨️DeepSeek-V4-Pro GA 官方正式发布!1 天前 · 3.8 万Orange AI万众期待!DeepSeek V4 Pro 0813 正式版发布1 天前 · 3 万歸藏(guizang.ai)卧槽,V4 Pro 08313 正式版涨价,那个峰值最高比原来涨了 4 倍多,这回真成梁子了。1 天前 · 2.8 万小互DeepSeek 一天公布三件大事:1 天前 · 2.6 万向阳乔木终于等到Deepseek-V4-Pro 正式版发布官方公告!1 天前 · 1.6 万🚨 AI News | TestingCatalogDeepSeek 发布了 DeepSeek-V4-Pro!1 天前 · 1.5 万Bindu Reddy哇!DeepSeek v4 Pro 发布了,纸面数据看起来很不错!1 天前 · 1.5 万Orange AI坏消息是 deepseek v4 pro 正式版有点拉了 1 天前 · 1.4 万AshutoshShrivastavaDeepSeek 发布了新的 DeepSeek-V4-Pro1 天前 · 4,039meng shaoDeepSeek-V4-Pro-0813 又一次静默发布了,官方 X 没有官宣,也没有铺天盖地的 PR 稿,但是…1 天前 · 2,937AIGCLINKDeepSeek V4 Pro 正式版发布了,DeepSeek-V4-Pro-0813,API 已经可以用了1 天前 · 2,548AIGCLINK刚刚DeepSeek-V4-Pro已官宣,正式版在APP、网页端和API同步上线!可以在APP/网页上选择“专家…1 天前 · 1,698meng shaoDeepSeek V4 Pro 今天低调发布1 天前 · 1,025
◔ 279.3 万 次浏览(20 条合计)♥ 1.4 万⇄ 1,717新品看原帖 ↗
Bindu Reddy@bindureddy · 创始人 · 1 天前Abacus.AI CEO,AI 行业观点博主

AutoBots - 多LLM自学习循环是未来

- 使用针对成本和性能优化的最佳LLM
- 设定目标和指标,AI搞定剩下的
- 长期任务可以自动化一切
- Fable 5在大型代码库上超级猛
- DeepSeek Flash对简单任务来说便宜到没朋友

查看英文原文
AutoBots - Multi-LLM Self Improving Loop Are The Future

- use the best LLM optimized for cost and performance
- set a goal and a metric, AI does the rest
- long horizon tasks can automate anything
- Fable 5 is super powerful on large code bases
- DeepSeek Flash is dirt cheap for simple tasks
Claude@claudeai · 公司官方 · 1 天前Claude 产品官方账号
连环推 ×3

您在 Chrome 中的 Claude 会话现在可以跨桌面、网页和移动设备同步。对话被保存,您的 Skills 和 Connectors 在浏览器中工作。

今日在 Max 和 Team 上可用,未来几周将推出到 Pro。

查看英文原文
Your Claude in Chrome sessions now carry over to desktop, web, and mobile. Conversations are saved, and your skills and connectors work in the browser.

Available on Max and Team today, rolling out to Pro in the coming weeks.
The side panel now runs the same Claude Cowork session as the desktop, web, and mobile apps. Sessions live with your account, not on any single device, so you can start in a tab and pick it up later somewhere else.

Give it a try:
claude.com/chrome
Browser agents can be tricked by instructions hidden in a page. We build defenses against this, and we still recommend a few habits of your own:
support.claude.com/en/articl…
◔ 116.8 万 次浏览(2 条合计)♥ 1.2 万⇄ 825新品看原帖 ↗
Min Choi@minchoi · 博主 · 23 小时前AI 产品演示博主,专门展示新工具玩法
连环推 ×11

这还不到48小时,SpaceXAI 刚发布了 Grok 4.6,大家已经在用它做各种疯狂的东西。10个超猛的例子:

查看英文原文
Less than 48 hours ago, SpaceXAI dropped Grok 4.6.

People are already building crazy stuff with it.

10 wild examples:
1. A complete Unity game...

without touching Unity 🤯
2. Snowboarding game built from scratch
3. Interactive 3D simulation
4. The Office turned into a living AI simulation 😂
5. Airbus H145 helicopter model recreated in 3D
6. Grok 4.6 is really impressive at its cost

$1.80 vs $4.20 for the same task
7. Someone built a 4D holographic wedding website 🤯
8. Queen Anne's Revenge recreated in 3D
9. Entire 3D world from one prompt
10. Sketch to a polished website
◔ 121.4 万 次浏览(13 条合计)♥ 3,527⇄ 325▶ 含视频新品看原帖 ↗
NVIDIA@nvidia · 公司官方 · 1 天前

恭喜 @SpaceXAI 团队发布 Grok 4.6。Grok 4.6 提供前沿性能,运行和训练于 NVIDIA GB300 NVL72 搭配 NVLink,带来卓越的性能、可靠性和最低的 token 成本。

引用 SpaceXAI @SpaceXAIGrok 4.6发布。相比Grok 4.5性能明显提升,价格保持不变。查看被引原帖 ↗
查看英文原文
Congrats to the
@SpaceXAI
team on the release of Grok 4.6.

Grok 4.6 brings frontier intelligence, running and trained on NVIDIA GB300 NVL72 with NVLink to deliver exceptional performance, reliability and lowest token cost.
Tibo@thsottiaux · 公司官方 · 1 天前ChatGPT 产品官方账号

@ajambrosino 聊天的日常套路:
A:核心对齐达成了
T:核心对齐?
A:就是
T:多说点啊
A:没了
话说回来,我从没像现在这样期待他和团队的项目。

查看英文原文
Typical conversation with
@ajambrosino


A: core alignment has been reached
T: core alignment you say
A: i say
T: say more
A: more

And yet I’ve never been more excited about what him and the team are building.
Sundar Pichai@sundarpichai · 创始人 · 1 天前谷歌 CEO

Google Made 大会今天最重磅的更新之一是我们在 Gboard 和 Live Transcribe 里新推出的手语转文字功能。

这个功能是与听障社区合作开发的,能帮助使用 ASL 手语的人更无缝地与手机沟通,或与不懂手语的人交流,就像这个 Live Transcribe 演示里展示的那样:

查看英文原文
One of the most important updates coming out of Made by Google today is our new sign-to-text feature in Gboard & Live Transcribe.

Built in partnership with the Deaf community, this can help people who sign ASL communicate more seamlessly with their phone or with those who don't sign, like in this Live Transcribe demo:
◔ 61.5 万 次浏览(4 条合计)♥ 4,526⇄ 426▶ 含视频新品看原帖 ↗
@levelsio@levelsio · 博主 · 1 天前独立开发者标杆,AI 产品连续创业者

@dhh 分享了个很酷的技巧,把你的 Claude Code 接到 Ubiquiti 路由器上,改善 WiFi 体验!

引用 DHH @dhhUbiquity is so ridiculously good. Just did a full apartment upgrade. Cloud router, three U7 Pros. Added a local admin account for an agent. Had Claude optimize all the radio channels, tune mesh settings, and ensure all devices roamed ideally. SCI-FI!查看被引原帖 ↗
查看英文原文
Very cool tip by
@dhh
to hook your Claude Code up to your Ubiquiti router and improve your WiFi!
OpenAI@OpenAI · 公司官方 · 1 天前ChatGPT 开发商官方账号
连环推 ×2

前10%的企业使用插件的频率是普通企业的两倍,使用技能(skills)的频率则高出六倍。

这些前沿企业领先并非偶然。

查看英文原文
The top 10% of enterprises use plugins twice as often and skills six times as often as typical firms.

These frontier firms are not ahead by accident.
We examine how these orgs are putting AI to work, and how agentic workflows are expanding across industries and functions:
openai.com/index/how-enterpr…
swyx@swyx · 博主 · 1 天前知名 AI 播客 Latent Space 主理人

Perplexity在一年前的今天提议收购Google Chrome

(这是一条定时发布的推文)

查看英文原文
Perplexity offered to buy
@googlechrome
one year ago today

(this is a scheduled tweet)
Mustafa Suleyman@mustafasuleyman · 创始人 · 1 天前微软 AI CEO,DeepMind 联合创始人

我们首个推理模型 MAI-Thinking-1,从头打造。现已上线 Microsoft Foundry。给团队点赞!更多详情见下方。

查看英文原文
Our first reasoning model, MAI-Thinking-1, is built from scratch. Now available in Microsoft Foundry. Kudos to the team! More below.
◔ 27 万 次浏览♥ 1,044⇄ 103▶ 含视频新品看原帖 ↗
Vercel@vercel · 公司官方 · 1 天前前端云平台 Vercel 官方,AI 建站工具 v0 母公司

我们为AI SDK建了一个软件工厂。

每一步都是一个agent,人类负责合并改动。四周过去了:

▪️ 这个工厂贡献了高达35%的合并PR
▪️ 7月份解决了70%的问题
▪️ 未修复bug减少了25%

查看英文原文
We built a software factory for AI SDK.

Each step is an agent, and humans merge changes. Four weeks in:

▪️ The factory authors up to 35% of merged PRs
▪️ It closed 70% of issues in July
▪️ Open bugs are down 25%


vercel.com/blog/building-a-s…
Yuchen Jin@Yuchenj_UW · 博主 · 1 天前

一天内发布了3个前沿模型!

- Grok 4.6:Fable 5 水准,便宜 85%
- Qwen3.8-Max:2.4T 参数,95B 活跃。权重已公开
- DeepSeek-V4-Pro-0813:权重随时可能放出。听说效果不错,不只是刷榜

竞争对消费者和企业都是好事

查看英文原文
3 frontier models in one day!

- Grok 4.6: Fable 5-level, but 85% cheaper.
- Qwen3.8-Max: 2.4T params, 95B active. Weights are out.
- DeepSeek-V4-Pro-0813: weights could drop any time. Heard it’s good, not just benchmaxxing.

Competition is great for consumers and businesses.
Firecrawl@firecrawl · 公司官方 · 23 小时前

我们刚把4100多万篇生命科学论文加进了Firecrawl研究索引。

现在你的AI代理可以检索药物发现、临床试验和生物学文献,10次召回率高达90%。

所有内容均来自权威来源,完全免费。通过API的 /search/research 接口即开即用。

查看英文原文
We just added 41M+ life science papers to the Firecrawl Research Index.

Your AI agents can now search drug discovery, clinical trial, and biology literature with 90% recall@10.

All sourced from authoritative sources and 100% free. Live in the API via /search/research.
◔ 23.1 万 次浏览♥ 777⇄ 72▶ 含视频新品看原帖 ↗
@levelsio@levelsio · 博主 · 1 天前独立开发者标杆,AI 产品连续创业者

我把 ideasai.com 的自动应用和落地页生成器改用 @xai 的 Grok 了

Grok 4.6 好像和 Fable 5 一样能打

(没有那些烦人的唠叨,也不会动不动就因为 SeCuRiTy 阻止你干活)

引用 Gavin Baker @GavinSBakerGrok 4.6性能与Fable 5 Max相当,价格便宜85%;输入token便宜80%,输出token便宜88%。Grok 4.7将明显更优,规模更大并在预训练中融入Cursor与SpaceX数据。查看被引原帖 ↗
查看英文原文
I switched
ideasai.com
's auto app and landing builder to
@xai
Grok now

Grok 4.6 is apparently as good as Fable 5

(without the constant preaching and blocking you from doing anything due to SeCuRiTy)
Matt Shumer@mattshumer_ · 博主 · 1 天前HyperWrite CEO,AI 实战技巧分享

Jeremy Berman:ARC-AGI 的终极摧毁者。

没有人能比他做得更好。

引用 Jeremy Berman @jeremyberman我用Opus 5在ARC-AGI-3上取得96.2%的成绩,pass@2达到99.3%。这个程序基本上就是Claude Code + Opus 5(high)、一条action命令和文件系统日志的组合。几乎没有ARC特定内容。查看被引原帖 ↗
查看英文原文
Jeremy Berman: the ultimate destroyer of ARC-AGI.

No one will ever do it better than him.
向阳乔木@vista8 · 中文博主 · 1 天前向阳乔木,中文圈 AI 工具与趋势博主

这才是真干货啊,好友陈言是小红书10w+ 关注AI博主。

他把自己的接单赚钱的方法全分享了,掏箱底!毫无保留。

最让人佩服的,陈言老师在游戏、汽车等各方面都有广泛涉猎,一听都很深。

引用 陈言Linkc-Chen @Linkc分享一下我在小红书做自媒体的选题工作流,不是那种抄话题和复刻账号的。我靠这个方法获得了稳定的广告客户与合作机会。 这套方法容易操作,成本极低,只需要坚持。下面是具体流程。 (1/n)查看被引原帖 ↗
Chubby♨️@kimmonismus · 博主 · 1 天前Chubby,高频 AI 新闻聚合博主

天哪:Anthropic的投资人押注其在十月份IPO时估值超过2万亿美元,这将是有史以来最大的股市首秀。

震惊:支持者预计到年底年化收入将达到1000-1200亿美元,比2026年期间增长超过10倍。Anthropic在今年筹集近1000亿美元后,五月份的估值已达9650亿美元。

一位投资人对FT表示:“如果Anthropic年增长率达到800%”,即使是30倍的收入倍数,其估值也可能达到3万亿美元。

而且这还是在Opus 5尚未算在内的情况下。投资人显然预期Anthropic将领先,并很快推出新模型。

查看英文原文
Holy: Anthropic investors are betting on a $ 2tn+ valuation in an October IPO, the largest stock-market debut ever.

wtf: Backers expect annualised revenue to reach $100–120bn by year-end, up more than 10x during 2026. Anthropic was valued at $ 965bn in May after raising nearly $ 100bn this year.

One investor told the FT: “If Anthropic is growing 800 per cent a year,” even a 30x revenue multiple could value it at $ 3tn.

And that's despite Opus 5. Investors are presumably expecting a lead Anthropics and new models soon.
Chubby♨️@kimmonismus · 博主 · 1 天前Chubby,高频 AI 新闻聚合博主

我现在认真看待这个了。Grok 4.6 是我期待已久的那一步。要是 10t 模型还要来,Elon 的话就可以当真了:它真的有可能成为最强的(通用)模型。

当然,Anthropic 已经准备好 Fable 5.5 了,就等发布,这个没问题。不过接下来几周应该会很有意思。xAI 确实展现了他们有多大的潜力。

引用 Elon Musk @elonmuskGrok 4.7将超越所有现有模型。Anthropic是伟大的公司,很可能发布改进版本。但SpaceX训练语料库独特卓越,预计4.7在实际工程应用上将无人能及。查看被引原帖 ↗
查看英文原文
I'm taking this seriously now. Grok 4.6 was the leap I'd been hoping for. If the 10t model is still to come, then Elon's words can be taken seriously: it really could become the best model (in general).

Although, of course, Anthropic already has Fable 5.5 ready and just waiting to be released, that much is clear. Nevertheless, the next few weeks will be exciting. And xAI has shown just how much potential they possess.
Lisan al Gaib@scaling01 · 博主 · 1 天前高频 AI 模型测评与爆料博主

还是看不到 DeepSeek V4 的权重发布,这次定价更新简直离谱。

引用 DeepSeek @deepseek_aiWe’re launching DeepSeek-V4-Pro today! 🚀 🔷 Major Agent upgrades with strong production gains! 🔷 Flexible reasoning effort for V4-Pro & V4-Flash: low for simple tasks, high for daily Agent workflows, max for complex tasks. 🔷 Native OpenAI Responses API support, optimized for Codex with one-click setup. V4 Pro is now available on app/web. Try it via “Expert Mode”. V4 Pro is also available via API. Model names remain unchanged—please refer to the API docs for setup details.查看被引原帖 ↗
查看英文原文
I still don't see any weights for DeepSeek V4 and the pricing update is really bad
Yuchen Jin@Yuchenj_UW · 博主 · 1 天前

18个月前Karpathy造了个'vibe coding'概念,我们这些工程师包括我都笑死了说'这样的人注定完蛋'。结果呢,我们开始把代码往ChatGPT里复制粘贴。然后是Claude。然后是十个并行编码Agent。现在再手写代码简直像个变态。这十八个月真的是快到没边了。

查看英文原文
18 months ago, Karpathy coined “vibe coding.”

A lot of engineers, including me, laughed: “Vibe coders are NGMI.”

Then we started copy-pasting code into ChatGPT.
Then Claude Claude.
Then 10 parallel coding agents.
Hand-writing code now feels like being a psychopath.

That was a weirdly fast 18 months.
OpenRouter@openrouter · 公司官方 · 1 天前
连环推 ×3

今天我们推出 Ori Pi

可以直接在 OpenRouter 上运行 Pi。一键安装,你所有的 OpenRouter 凭证、模型和环境都为你准备好了,在 Pi 上

访问 openrouter.ai/ori/harness 开始使用

我们还添加了扩展来改进体验

查看英文原文
Today, we are launching Ori Pi

Run Pi directly on OpenRouter. One install, all your OpenRouter credentials, models, and environment set up for you, on Pi

Get started at
openrouter.ai/ori/harness


We also added extensions to improve the experience
5/ ori pi also helps handles other provider key from pi's environment, so a stray ANTHROPIC_API_KEY or OPENAI_API_KEY won’t conflict with the request off OpenRouter.
Get started today with Ori Pi at:
openrouter.ai/ori/harness
yihong0618@yihong0618 · 中文博主 · 1 天前

我看了所有的贡献者。。。太想吐槽了,崔神合着在推特每个人下面留言,留言的,一个都没招。。。
找了一堆金牌的素人。。。

Groq Inc@GroqInc · 公司官方 · 1 天前AI 推理芯片公司,以速度极快出名

Groq 现在是
@nvidia
云合作伙伴了。

这对团队来说是个里程碑,也验证了客户早已体验到的:Groq 以最高标准运行 AI 基础设施。

“推理正成为 AI 中最大最关键的层,我们打算比任何人都跑得更好。” —— Adam Winter,Groq CEO


groq.com/newsroom/groq-becom…

查看英文原文
Groq is now an
@nvidia
Cloud Partner.

A milestone for the team and validation of what our customers already experience: Groq runs AI infrastructure to the highest standard.

"Inference is becoming the largest and most critical layer of AI, and we intend to run it better than anyone." — Adam Winter, CEO of Groq


groq.com/newsroom/groq-becom…
小互@xiaohu · 中文博主 · 1 天前小互,中文圈高频 AI 资讯站 Xiaohu.AI 主理人
连环推 ×2

DeepSeek-Harness

这个描述看的一头雾水😅

karminski-牙医@karminski3 · 中文博主 · 1 天前karminski-牙医,中文圈模型评测博主

刚发的 DeepSeek-V4-Pro-0813 好像的确有问题, reasoning_effort=max 的时候模型在长程 AgenticCoding 任务下更倾向于早停.

我这个测试总计50轮, 让模型不断优化代码, 结果他在3次测试中, 2次都在42-43轮左右的情况下停止了. 并未用完全部机会. 而且目前成绩来看, 甚至没打过 GLM-5.1 (仅我这一个case, 我不对其它case负责).

目前还不确定这是强化学习阶段奖励机制弄得有问题,还是是上下文窗口注意力衰减的问题,一会放出详细评测视频给大家解析。


#deepseekv4pro0813
#deepseekv4pro正式版

Guillermo Rauch@rauchg · 创始人 · 1 天前Guillermo Rauch,Vercel 创始人兼 CEO

迈向自我维护软件的未来。

不容错过的阅读:看我们如何在开发全球最受欢迎的开源 AI 库之一 @aisdk 时,平衡质量与自主性。

引用 Vercel @vercelVercel构建AI SDK软件工厂,每步由agent执行,人类审核合并。四周成果:工厂完成35%合并PR、处理70%问题、开放bug减少25%。查看被引原帖 ↗
查看英文原文
Towards a future of self-maintaining software.

Must read of how we balance quality and autonomy for the development of
@aisdk
, one of the most popular open source AI libraries in the world.
Lisan al Gaib@scaling01 · 博主 · 1 天前高频 AI 模型测评与爆料博主

第二梯队模型间的竞争活得好好的,但对标杆模型的竞争已经彻底死了。说真的,死到这种程度,连 OpenAI 和 Anthropic 都不用在好几个月内发新模型了。

引用 Nathan Lambert @natolambert从'Anthropic 遥遥领先'到'模型竞争回到历史高位',舆论反转仅4周。查看被引原帖 ↗
查看英文原文
competition for second tier models is well and alive

competition for frontier models has never been more dead

in fact, it's so fucking dead that OpenAI and Anthropic don't even need to release their models for months
idoubi@idoubicc · 中文博主 · 1 天前

DSCode 桌面版已发布,基于 DeepSeek 的桌面 Agent

基于 Pi 实现的极简 Harness,适配 DeepSeek V4 Pro / Flash 模型,响应速度极快、缓存率极高。主打便宜、极速、稳定

完全免费,MIT 协议开源。欢迎体验👇


dscode.ai

引用 idoubi @idoubicc开源 DSCode:为 DeepSeek 深度优化的 Coding Agent 核心卖点:快、省、稳、开放、安全 - 快:DeepSeek V4 Flash + 极简 Agent Harness,原生适配 Responses API、推理档位、流式响应、代码 Patch 和服务端 Web Search,减少工具选择与协议开销。 - 省:针对 1M Context 和前缀缓存优化,提高长会话缓存命中率;可实时查看 Token、缓存命中和费用,让每一笔模型开销都清清楚楚。 - 稳:内置 Plan、会话恢复、Checkpoint/Undo 和后台任务;支持多 Agent 并行,并通过独立 Git Worktree 隔离修改。 - 开放:支持 Skills、Hooks、MCP;DeepSeek 优先,也可以切换 GPT、Claude 等模型。 - 安全:API Key、配置和会话默认保存在本地;命令运行在 OS Sandbox 中,工具权限由用户控制,Provider 密钥不会传给子进程。 DSCode Agent Runtime 基于 Pi Agent Toolkit,采用 MIT 协议完整开源,可以自由 Fork、修改,也可以集成到你自己的 CLI、IDE 或桌面产品中。 欢迎试用,感谢 Star ⭐ 开源仓库地址👇查看被引原帖 ↗
Gemini Notebook@Gemini_Notebook · 公司官方 · 1 天前谷歌 AI 笔记工具 NotebookLM 官方

找到好笔记本了?把它变成你自己的!现在你可以复制笔记本,包括所有资源和工件。非常适合:

🎓 学生:复制共享笔记来添加个人学习调整

💼 团队成员:复制基础模板来快速启动新项目

立即尝试!

查看英文原文
Find a great notebook? Make it yours! You can now make a copy of a notebook, including all sources and artifacts. Perfect for:

🎓 Students: copy shared notes to add personal study tweaks

💼 Teammates: duplicate foundational templates to kickstart new projects

Try it today!
◔ 7.9 万 次浏览♥ 1,057⇄ 114▶ 含视频新品看原帖 ↗
Chubby♨️@kimmonismus · 博主 · 1 天前Chubby,高频 AI 新闻聚合博主

DeepSeek 4 Pro GA 是个非常好的模型。但天哪,价格这么便宜!$0.435/M 输入、$0.87/M 输出,1M token 上下文。这价格简直疯了!!

引用 Chubby♨️ @kimmonismusDeepSeek 4 GA基准测试已发布。这是扎实升级,但与Opus 4.8而非Opus 5对比略显遗憾。结果稳定,与开源SOTA(Kimi k3和GLM-5.2)相当,接近5.6 Sol和Fable 5。查看被引原帖 ↗
查看英文原文
The DeepSeek 4 Pro GA is a very good model. But holy sh*t, the pricing is so damn good!

$0.435/M input and $0.87/M output @ 1M-token context

This is insane pricing!!
◔ 14.6 万 次浏览(2 条合计)♥ 567⇄ 33新品看原帖 ↗
宝玉@dotey · 中文博主 · 1 天前宝玉,中文圈 AI 翻译与科普大 V

最近一直在忙着把我的 App 变成跨平台的,试了几种方案还是决定用 Rust + GPUI。

一开始就是用的 Electron,但是视频编辑这种场景纯网页实现性能优化还是挺不容易的,尤其是视频一长、字幕一多就会很卡,内存占用也很夸张。我不怀疑持续优化下去能是能有一个还不错的性能的,但要花不少时间精力。

所以我后来就决定先做一个 Mac 版本(图1),Swift + AppKit,用 Fable 5 开发,效果特别好,跟设计稿几乎一样。必须说 Fable 5 在还原 UI 能力上是最强的,没有之一。

但现在想要支持跨平台,就还得考虑其他方案。先是用 Rust 做了一个跨平台 cli 的 PoC,效果很好,音频转录和视频导出的性能比 Swift 还好,还能支持 Mac、Windows 多平台。

接下来就是把 Swift 版本移植到 cli 之上,工作量还是挺大的,前期开发的功能越多,现在迁移工作就越难。初步是完成了,但接下来就是 Windows 版本的支持。

再去基于 cli 开发个 windows 版本成本还是挺高,尤其是要去 windows 电脑上开发测试,各种逻辑同步想想都头大。现在 AI 让写代码变得成本很低,但是测试和验收成本还在那里,跨平台的好处就是你核心逻辑测试没问题,后面只有少量兼容性的问题需要去测试,工作量小很多。

因为 cli 已经是 rust 写的,所以自然而然就会想到用 rust 的 GPUI 去做跨平台,不过决定之前肯定是要先做 PoC 验证的。

跟 Fable 5 一起写了个技术方案文档,由于上个周期额度到了,后面就先让 GPT 5.6 Sol 去实现了一个 PoC 版本(图2),还不错了。

今天 Grok 4.6 发布,又用 Cursor + Grok 4.6 实现了一个版本(图3),效果要差一些。Grok 4.6 有点好处就是生成的速度极快

今天晚上 Claude 的额度终于刷新了,再让 Fable 去实现了一个版本(图4),相对来说是最好的一个,但离成品也还有不小的差距。

经过这三个 PoC 版本,至少可行性上是没什么问题了,性能也很好,接下来就沿着这个方向走下去了。

---

对了,现在 BaoCut Skill 从 Windows 下也可以转录和翻译视频了:

github.com/JimLiu/baocut/blo…

Cohere@cohere · 公司官方 · 1 天前加拿大企业级大模型公司 Cohere 官方
连环推 ×3

今天,我们的模型家族又添新成员了。欢迎 North Micro Vision。

这是我们目前最小的视觉语言模型,非常适合复杂的文档理解场景。以 Apache 2.0 许可证开源,权重模型现已在
@huggingface
上线。

查看英文原文
Today, we’re adding another member to our model family. Meet North Micro Vision.

Our smallest vision-language model yet, ideal for sophisticated document understanding. Available open-source under an Apache 2.0 license. Get the weights on
@huggingface
.
The model punches above its weight, outperforming Gemma 4 E2B and Ministral 3 3B across a broad range of visual understanding benchmarks and delivers particularly strong results on document understanding and visual Q&A.

See the full results:
huggingface.co/blog/CohereLa…
North Micro Vision is now free to deploy under Apache 2.0. Let us know what you get up to.

Download the model weights, learn more about the model, its architecture, and more:
huggingface.co/CohereLabs/No…
◔ 8.8 万 次浏览(3 条合计)♥ 806⇄ 80新品看原帖 ↗
🚨 AI News | TestingCatalog@testingcatalog · 博主 · 1 天前专挖 AI 产品未发布新功能的爆料号

Cursor Origin 已开始向部分用户推送,位于 Cursor 网页版的新 /codebase 路径下。

> Cursor Origin 是 Cursor 即将推出的 Agentic 代码审查解决方案,预计用户还能通过 Origin 管理其代码库。

很快上线👀

引用 Seth Saler @sethsalerWell hello there, Cursor Origin cursor.com/codebase查看被引原帖 ↗
查看英文原文
Cursor Origin started rolling out to some users under a new /codebase path on Cursor web.

> Cursor Origin is an upcoming Agentic code review solution from Cursor. It is also expected that users will be able to manage their codebases via Origin as well.

Soon 👀
NVIDIA@nvidia · 公司官方 · 23 小时前
连环推 ×3

AI 工厂是 AI 时代的工业基础设施。Tokens 是新的商品。你如何优化 AI token 经济学?🧵

查看英文原文
AI factories are the industrial infrastructure of the AI era.
Tokens are the new commodity.

How do you optimize AI token economics? 🧵
Maximize token monetization.
That's how compute becomes revenue.

Explore the NVIDIA AI Tokenomics Guide:
nvda.ws/4cw5yx9
向阳乔木@vista8 · 中文博主 · 1 天前向阳乔木,中文圈 AI 工具与趋势博主

用Deepseek V4 Pro 0813跑了三个简单任务。

1. 第一个是开发部署一个网站,需要调用3个Skill,完成网站开发,子域名解析和上线。

2. 第二个是模仿60种设计风格,生成Bento卡片图展示。

3. 第三个是Threejs 生成一个3D打砖块游戏。

大家可以看看效果,我感觉还行。

◔ 6.7 万 次浏览♥ 197⇄ 9▶ 含视频演示看原帖 ↗
🚨 AI News | TestingCatalog@testingcatalog · 博主 · 1 天前专挖 AI 产品未发布新功能的爆料号

OpenAI 🔥:Codex 达到 1500 万用户里程碑,使用额度已重置。

Codex Realtime 这才是真杀器啊 👀

引用 Tibo @thsottiauxOld news actually from a bunch of days ago, but crossed that 15M. Enjoy a nice reset everyone. Landing in the next hour or so, go /fast.查看被引原帖 ↗
查看英文原文
OPENAI 🔥: Codex marked 15M users milestone with a usage limit reset.

Codex Realtime is a big unlock 👀
François Chollet@fchollet · 创始人 · 1 天前
连环推 ×3

Test-time training 在 ARC Prize 2024 竞赛期间被推广开来,此前主要由 @MindsAI_Jack 和团队在研究。到现在为止,我认为 ARC 1-2 是唯一一些 TTT 能明显更优的数据集。要是 TTT 开始变得更主流就有意思了。我认为它潜力巨大。

引用 François Chollet @fcholletToday we're announcing the winners of ARC Prize 2024. We're also publishing an extensive technical report on what we learned from the competition (link in the next tweet). The state-of-the-art went from 33% to 55.5%, the largest single-year increase we've seen since 2020. The benchmark remains unbeaten, but we're happy to see that research progress on the key bottleneck to AGIs (in particular on-the-fly adaptation to novel tasks) has been reignited in 2024 -- in part thanks to ARC Prize. In particular the competition has popularized Test Time Training (TTT), originally pioneered for ARC-AGI by Jack Cole last year. I believe TTT represents the largest jump in LLM generalization capabilities since the initial findings regarding in-context-learning circa 2019-2020. ARC Prize has also led to a considerable surge of research interest towards program synthesis. Competition winners: 🥇 the ARChitects (Daniel Franzen, Jan Disselhoff) 🥈 @guille_bar 🥉 alijs (Agnis Liukis) Paper Award winners: 🥇 "Combining Induction and Transduction For Abstract Reasoning" by @xu3kev et al. 🥈 "The Surprising Effectiveness of Test-Time Training for Abstract Reasoning" by @akyurekekin et al. 🥉 "Searching Latent Program Spaces" by @ClementBonnet16 & @MattVMacfarlane ARC-AGI-Pub Leaderboard (solutions using commercial APIs): 🥇 @jeremyberman 🥈 @ellisk_kellis & @akyurekekin 🥉 @ryangreenblatt查看被引原帖 ↗
查看英文原文
Test-time training was popularized during the ARC Prize 2024 competition, after being explored in particular by
@MindsAI_Jack
and team. To date, I believe ARC 1-2 are the only datasets where TTT strongly outperforms. It would be interesting if TTT started becoming more mainstream. I believe it has great potential.
The way most people leverage test-time compute today is via a form of test-time NL reasoning that is computationally equivalent to test-time search (sometimes with a verifier / grader in the loop). The other major avenue to leverage test-time compute is test-time training. Gradients are a precious signal, there's no reason not to use it at test time (other than the fact that it would be difficult / expensive from an engineering standpoint).


x.com/fchollet/status/186905…
Test-time training is the only form of test-time adaptation that is "pure" deep learning, as opposed to neurosymbolic. It adapts in continuous latent space, as opposed to adapting in discrete symbol space.
@levelsio@levelsio · 博主 · 1 天前独立开发者标杆,AI 产品连续创业者

最易于修复的bug,就数
@X

@xAI
员工面对的这个了

查看英文原文
Easiest bug for
@X
+
@xAI
staff to fix is this
Guillermo Rauch@rauchg · 创始人 · 1 天前Guillermo Rauch,Vercel 创始人兼 CEO

这改进真不错。

试试 npx sandbox@latest sh —— 太疯狂了。感觉比在你本地跑还快。

Sandbox 现在自带一套合理的预装工具,完全可以定制。

引用 Vercel Developers @vercel_dev介绍 Vercel Managed Images。Vercel Sandbox 现已搭载 Ubuntu 作为默认操作系统、预装编码代理(codex、claude…)以及可定制的开源基础镜像。查看被引原帖 ↗
查看英文原文
This is such a nice improvement.

Try 𝚗𝚙𝚡 𝚜𝚊𝚗𝚍𝚋𝚘𝚡@𝚕𝚊𝚝𝚎𝚜𝚝 𝚜𝚑 – it's mind-blowing. It feels faster than your local machine.

Sandbox now comes with a reasonable default set of pre-installed tools, and you can fully customize them.
◔ 5.4 万 次浏览♥ 332⇄ 11▶ 含视频新品看原帖 ↗
@levelsio@levelsio · 博主 · 1 天前独立开发者标杆,AI 产品连续创业者

我同意这个

更多的测试和全身扫描等会显露很多其实不是问题、但等你真的开始在人身上动刀子的时候可能会成为问题的东西

但如果我们做得对,预防性医疗的好处仍然会超过这方面的弊端

而且通过便宜且定期的检测(比如 Midjourney Medical scanner 那样),你可以看到你身体的变化,快速发现异常

因为医疗检查有风险就反对预防性医疗和寿命延长,有点像是 1900 年因为汽车会杀人就反对汽车一样,确实会杀人,但这些好处相比继续用马车来说还是值得的

引用 Medici @HouseMediciType 1 scan problems are very real and we’ll have a lot of people wanting to get shit cut out of them for no reason But the benefits still significantly outweigh the downsides查看被引原帖 ↗
查看英文原文
I agree with this

More testing and body scans etc will shows lots of issues that are not actually issues but might become issues the moment you start cutting in people's body

But the benefits of preventative healthcare if we do it right will still outweigh the negatives of that

Also with cheap and regular testing (like the Midjourney Medical scanner) you can see your body change and quickly spot irregularities

Being against longevity and preventative healthcare because it has its issues is kinda like being against cars in 1900 because they kill people, they do but the benefits outweighed that over staying with horses
Cristóbal Valenzuela@c_valenzuelab · 创始人 · 1 天前Runway 联合创始人兼 CEO

机器人、世界模型和物理AI的“科切拉音乐节”——Runway的SF峰会,将于9月30日登陆湾区。

嘉宾阵容:

@QuanVng
(来自
@physical_int),
@liu_mingyu
(来自
@nvidia),
@robertnishihara
(来自
@anyscalecompute),
@pariljain
(来自
@thebotcompany),以及
@jon_barron
(来自
@GoogleDeepMind),
@agermanidis
(来自
@runwayml),
@Vivek_Vis
(来自
@CAGovernor),Richard Ahlfeld(来自
@CoreWeave)等更多嘉宾待公布。

我们还会准备一些意想不到的惊喜。

购票链接:
summit.runwayml.com/

查看英文原文
The Coachella of robotics, world models, and physical AI is Runway’s SF Summit, coming to the Bay Area on September 30.

Speaker lineup:


@QuanVng
from
@physical_int
,
@liu_mingyu
from
@nvidia
,
@robertnishihara
from
@anyscalecompute
,
@pariljain
from
@thebotcompany
, and
@jon_barron
from
@GoogleDeepMind
,
@agermanidis
from
@runwayml
,
@Vivek_Vis
from
@CAGovernor
, Richard Ahlfeld from
@CoreWeave
and more to be announced..

We’ll also have some unexpected surprises

Tickets here:
summit.runwayml.com/
AshutoshShrivastava@ai_for_success · 博主 · 1 天前高频 AI 新闻与产品动态博主

我第一次测试 Grok 4.6

让它理解我给儿子开发的游戏 Aether Ascent 的代码(你们很多人都见过),然后做一个全新的关卡。

它创建了一个叫 Stormwake 的第 10 关。

很喜欢这个关卡设计,特别是雷电特效。一次就成功了,完全没问题。

查看英文原文
My first test with Grok 4.6

I asked it to understand my current game code for Aether Ascent, the game I’ve been building for my son and many of you have seen, and then build a completely new stage.

It created a new stage "10" called Stormwake.

Love the stage design, especially the thunder effects. it worked in one shot. No issues.
◔ 4.9 万 次浏览♥ 185⇄ 12▶ 含视频演示看原帖 ↗
Lisan al Gaib@scaling01 · 博主 · 1 天前高频 AI 模型测评与爆料博主

DeepSeek 今年太让人失望了。V4 发布晚了不说,规模也比想象的小。价格还涨了不少。他们目前在中国排名四五位。

引用 DeepSeek @deepseek_aiWe’re launching DeepSeek-V4-Pro today! 🚀 🔷 Major Agent upgrades with strong production gains! 🔷 Flexible reasoning effort for V4-Pro & V4-Flash: low for simple tasks, high for daily Agent workflows, max for complex tasks. 🔷 Native OpenAI Responses API support, optimized for Codex with one-click setup. V4 Pro is now available on app/web. Try it via “Expert Mode”. V4 Pro is also available via API. Model names remain unchanged—please refer to the API docs for setup details.查看被引原帖 ↗
查看英文原文
DeepSeek has been very disappointing this year

V4 came much later than expected and was smaller than expected. The price hikes also hurt.

They are currently 4th or 5th place in China
Gorden Sun@Gorden_Sun · 中文博主 · 1 天前中文圈高频 AI 资讯与开源项目博主

IndexTTS发布2.5版本,开源最强文字生成语音模型,没有之一。在音色克隆、情绪还原方面尤其出色,支持中文、英文、日语、西班牙语、阿拉伯语5种语言。

项目介绍:
index-tts.github.io/index-tt…

模型:
huggingface.co/IndexTeam/Ind…

◔ 4.8 万 次浏览♥ 328⇄ 57▶ 含视频新品看原帖 ↗
Lisan al Gaib@scaling01 · 博主 · 1 天前高频 AI 模型测评与爆料博主

Grok 4.6 虽然体量小,但性能和效率都超过了几乎大两倍的 Kimi-K3

查看英文原文
Grok 4.6 is slightly better and more efficient than the almost twice as large Kimi-K3
Yuchen Jin@Yuchenj_UW · 博主 · 1 天前

我们会在IPO之前完成Z轮融资!

开玩笑的,Databricks 是我见过最有创始人领导力、运营最出色的公司之一。

我加入整整四个月了,时间过得真快!我们在 AI 领域做了很多令人兴奋的事情。在 AI 推理方面,我们正在服务前沿的开源模型,让它们跑得飞快,而且一些非常大的客户采用率也很高。

接下来我们只会加速前进。

引用 Ali Ghodsi @alighodsiToday, we announced that we crossed $7B in revenue run-rate, growing over 80% year over year in Q2. We also shared: 🚀 $100M+ revenue run-rate for Lakebase 🚀 $1.5B+ revenue run-rate for Lakehouse, growing over 100% year over year 🚀 Continued positive adjusted free cash flow And we raised $5B in our latest fundraise. We’ll use this capital to invest in: 1️⃣ Lakebase, our serverless Postgres database built for AI agents 2️⃣ Genie, our AI coworkers that actually understand your business data 3️⃣ Unity AI Gateway, our multi-AI governance solution that helps control costs @iamVictorDey shares more in @Forbes : forbes.com/sites/victordey/2…查看被引原帖 ↗
查看英文原文
We’ll raise Series Z before IPO!

Jokes aside, Databricks is one of the most founder-led, well-run companies I’ve seen.

I joined exactly 4 months ago, time flies! We’re doing a lot of exciting things in AI. On AI inference, we’re serving frontier OSS models, making them run super fast, and seeing strong adoption from some very large customers.

And we’re only going to move faster.
Chubby♨️@kimmonismus · 博主 · 1 天前Chubby,高频 AI 新闻聚合博主

Sergey Brin 据悉正在将 Google 的 AI 资源引向递归自我改进。

Reuters 报道称,Brin 已敦促主要员工跟上前沿进展,并将资源调向能在无人干预下自我改进的 AI。

在 Anthropic 凭 Claude Mythos 抢得先机后,压力随之加大。Google 随后将其下一代旗舰 Gemini 模型延后了两个月,因为内部测试表明,在包括编码等多个领域,该模型落后于竞争对手。

查看英文原文
Sergey Brin is reportedly steering Google’s AI resources toward recursive self-improvement.

Reuters says Brin has urged key staff to catch up with the frontier and pushed resources toward AI that can improve without human intervention.

The pressure intensified after Anthropic raced ahead with Claude Mythos. Google subsequently delayed its next flagship Gemini model by two months after internal tests showed it lagging rivals in areas including coding.
向阳乔木@vista8 · 中文博主 · 1 天前向阳乔木,中文圈 AI 工具与趋势博主

疯狂星期三,过年了啊!

- Grok 4.6
- DeepSeek V4 Pro 0813
- WeLM-617B
- Qwen3.8-2.4T-A95B

你最想先测哪个?

Guillermo Rauch@rauchg · 创始人 · 1 天前Guillermo Rauch,Vercel 创始人兼 CEO

HarnessAgent 让你自由运行或替换任何脑力。或者让一堆脑力在同一个问题上竞争或合作。现在只需一行代码就能切换到 @grok。

引用 Vercel Developers @vercel_devUse @grok Build with AI SDK 𝙷𝚊𝚛𝚗𝚎𝚜𝚜𝙰𝚐𝚎𝚗𝚝: 𝚌𝚘𝚗𝚜𝚝 𝚊𝚐𝚎𝚗𝚝 = 𝚗𝚎𝚠 𝙷𝚊𝚛𝚗𝚎𝚜𝚜𝙰𝚐𝚎𝚗𝚝({ 𝚑𝚊𝚛𝚗𝚎𝚜𝚜: 𝚐𝚛𝚘𝚔𝙱𝚞𝚒𝚕𝚍, }); vercel.com/changelog/grok-bu…查看被引原帖 ↗
查看英文原文
𝙷𝚊𝚛𝚗𝚎𝚜𝚜𝙰𝚐𝚎𝚗𝚝 gives you the freedom to run and swap any brain. Or make a bunch of brains compete or collaborate on a problem. It’s now one line of code to switch to
@grok
.
Chubby♨️@kimmonismus · 博主 · 1 天前Chubby,高频 AI 新闻聚合博主

白宫据报准备将开源 AI 模型纳入其秘密的预发布安全测试框架。据 WIRED 报道,该自愿框架目前涵盖来自 OpenAI 和 Anthropic 等实验室的前沿闭源模型。一旦开源模型达到相当的能力,预计将加入框架,可能在公开发布前需接受 30 天的测试。官员们夹在两难之地:排除开源模型可能为闭源实验室创造政府认可的优势;纳入开源模型可能拖缓美国开源模型的发展。所以是的,事情变得有意思了。

引用 Hugo Lowell @hugolowell新消息@WIRED:Trump White House预计在未来数周至数月内扩展AI监管框架,包括当开源模型达到Anthropic和OpenAI顶级模型相同的frontier能力时,将其纳入框架。查看被引原帖 ↗
查看英文原文
The White House is reportedly preparing to bring open AI models under its secret prerelease safety-testing framework.

WIRED says the voluntary framework currently covers frontier closed models from labs such as OpenAI and Anthropic. Open models are expected to join once they reach comparable capabilities, potentially facing a 30-day testing period before public release.

Officials are caught between two risks: excluding open models could create a government-approved advantage for closed labs; including them could slow US open-model development.

So yeah, its getting interesting.
Gorden Sun@Gorden_Sun · 中文博主 · 1 天前中文圈高频 AI 资讯与开源项目博主

我做了一套AI生成视频和剪辑的技能,生成的视频如下。
音频和头部来自
@Cuimao
,其他所有内容都是AI生成和剪辑的。
基于达芬奇,生成的项目可以手动继续剪辑。

Orange AI@oran_ge · 中文博主 · 1 天前Orange AI,中文圈 AI 产品观察博主

老马不厚道啊
Grok 4.6 偷偷涨了缓存价格,和 GPT 5.6 Sol 一样
你是不是飘了
你有 Sol 好吗?

Min Choi@minchoi · 博主 · 1 天前AI 产品演示博主,专门展示新工具玩法

每个人都在讨论Grok 4.6。

Elon:先让让,看我的。

引用 Elon Musk @elonmuskGrok 4.7比4.6有显著提升,应在3-4周内推出。初始训练已完成,现在在补充训练中加入大量SpaceX公司数据。这将非常特别。查看被引原帖 ↗
查看英文原文
Everyone's talking about Grok 4.6.

Elon: hold my beer.
Guillermo Rauch@rauchg · 创始人 · 1 天前Guillermo Rauch,Vercel 创始人兼 CEO

Seedance 2.5 on @vercel AI Gateway 🐇

查看英文原文
Seedance 2.5 on
@vercel
AI Gateway 🐇
◔ 4.1 万 次浏览♥ 359⇄ 11▶ 含视频新品看原帖 ↗
Guillermo Rauch@rauchg · 创始人 · 1 天前Guillermo Rauch,Vercel 创始人兼 CEO

Digg绝对是我web 2.0时代最爱的网站。能把我写的技术文章送上“digg首页”,是我青春期最酷的回忆之一 😁

超级开心能为新的、agent-native digg.com 提供技术支持,跑在 @vercel 上。彻底 🕳️

引用 Kevin Rose @kevinroseDigg重启两个月,团队运行100+任务/智能体/评判器,实现自动聚类、情感分析、图像创建等功能。流量环比增长20%+。Grok模型事实判断能力强,Vercel基础设施稳定。希望深化Grok与API的整合。查看被引原帖 ↗
查看英文原文
Digg was by far my favorite site of the web 2.0 era. Getting my tech essays onto "the digg front page" was one of the coolest memories of my teenage years 😁

So happy to be powering the new, agent-native
digg.com
on
@vercel
. Full 🕳️
Chubby♨️@kimmonismus · 博主 · 1 天前Chubby,高频 AI 新闻聚合博主

又一波重置来了;Codex 已超过 1500 万用户大关。祝贺 OpenAI,也祝贺我们。

引用 Tibo @thsottiauxOld news actually from a bunch of days ago, but crossed that 15M. Enjoy a nice reset everyone. Landing in the next hour or so, go /fast.查看被引原帖 ↗
查看英文原文
Another reset is here; Codex has surpassed the 15 million user mark. Congrats to OpenAI and congrats to us.
Orange AI@oran_ge · 中文博主 · 1 天前Orange AI,中文圈 AI 产品观察博主

准备定期整理一些好的 skill 发出来
第一波:审美很不错的照片 Skill

1. photo-abstract-editorial
保留照片的真实内容,并仅从照片本身提炼空间关系、构图节奏和色彩关系

go.colaskill.com/abs


2. heytea-style
喜茶风格海报生成

go.colaskill.com/heytea


3. gc-minimal-zine-poster
可以把主题、句子、文章构想、物件、情绪、照片或参考图集,转化为一张zine纸张质感的极简编辑海报

go.colaskill.com/zine


4. ip_illustration_for_yourself
把自己的个人IP生成萌粒风插画,并且可以根据内容生成相应配图

go.colaskill.com/ip


5. photo-revival
把普通照片、生活随手拍、废片和日常物件,重新画成一页白纸上的诗性手绘插画。

go.colaskill.com/revival


6. tait-crt-interface-skill
早期CRT计算机界面质感的复古插画skill

go.colaskill.com/crt

Jason Wei@_jasonwei · 创始人 · 1 天前Jason Wei,思维链(CoT)提出者

在机器智能有那么多优势的世界里,人类还剩下什么?

最近我买了辆特斯拉,用完全自动驾驶功能让我意识到 AI 相比人类有多少优势。少数几次我因为觉得车要开错车道而手动接管,结果发现是我错了车是对的。我意识到我不可能比一个了解每条路、四面八方都能看见、永不疲劳也不会分心的神经网络开得更好。

既然 AI 在某些方面有天然优势,人类还能保留什么护城河?这是个大问题。

短期答案是,我们生活的世界是为人类而建的,在某些领域 AI 还没追上。比如,AI 还是不太会用网页界面。虽然任何懂电脑的人都能轻松浏览网页,但 AI 的准确点击和拖拽还不够好,因为图像嵌入不是为这种精度优化的。如果网页是为输入语言模型而设计而不是作为人类的视觉界面呈现,AI 肯定会远超人类。但现在,语言模型还得适配我们的遗留基础设施。

机器人是另一个人类有主场优势的领域。物理世界里大多数任务都是围绕手指和对向拇指设计的,这些一直很难集成到机器人里。虽然明显机器在为自动化优化的环境(比如大规模制造线)能超越人类,但现在世界大部分还是为人类而建。

不过这些能力差距只是暂时的。总有一天机器会比我们点击更快、拥有超人的通用灵巧度。人类真正的护城河是什么?

任何涉及语言模型无法接触的私密知识的东西对我来说都像是坚固的护城河。浪漫匹配和高端房地产就是两个例子,库存通常不会公开宣传,匹配是通过在对的圈子才能实现。风险投资也是例子——虽然一些研究和决策可以用 AI 自动化,但很多成功取决于提前理解趋势和连接对的人,这都需要私密知识。虽然机器可能会越来越有接触某些类型私密知识的机会,但我认为仍然会有某些类型的私密知识只有人类知道。我看不出当关键信息被紧密封锁在人类圈子里时,AI 怎么能赢。

人类似乎有真正护城河的另一个领域是娱乐和艺术,这些由于其人类特质本身而被重视,不管机器能做得多好。看博尔特短跑一百米是美的,因为这是人类能力巅峰的表达,即使汽车能开得更快。看业余或中级水平的国际象棋比赛比看两个超人类的 AI 互相比赛更有共鸣和满足感。艺术的价值来自创作过程,这就是为什么复制品不如原作值钱。这些类型的工作似乎即使我们朝着超级智能发展也会继续有市场。

更广泛地说,人类存在是一个对 AI 来说按定义会很难自动化的特征。比如,老师记得你的名字或父母支持你是有价值的,即使 AI 能轻松记得你的名字、可能还能给更好的人生建议。因为某人花了他有限生命中的一部分在你身上,他们的时间会耗尽,这是有分量的。作为个人轶事,我还记得第一次和我认为很了不起的 AI 研究员一起工作。他的建议很中肯,但更重要的是我相信我可以和他作为合作者一起做伟大的工作,我提高了自己的标准。过去几十年,技术发展在某些方面分裂了我们,但希望 AI 把我们带向一个重新强调人类存在的世界。

智能一直是人类的决定性特征,对 AI 来说在接下来的几十年自动化这个将是大改变。短期内,某些类型的智能会变得很便宜,并自动化掉旧工作,但我上面描述的护城河不会是人类能保有价值的仅有地方。同样的,电脑夺走了……

查看英文原文
What's left for humans in a world where machine intelligence has so many advantages?

I recently got a Tesla, and using full self-driving has been a wake up call to just how many advantages AI has over humans. The few times I disengaged it because I thought it was going into the wrong lane, it turned out that the car was right and I was wrong. I realized that there is no hope of me driving better than a neural net that knows every road, sees in every direction at once, and never gets tired or distracted.

Given that AI has certain inherent advantages over human intelligence, what kind of moats will remain for us as humans? It's a big question.

One short-term answer is that the world we live in was created for humans, and in some domains, AI has not closed the gap yet. For instance, AI still struggles to use internet user interfaces. While any computer-literate human can navigate a web page with ease, AI is still not great at making accurate clicks and drags because image embeddings are not optimized for such precision. If the internet were designed to be fed into language models instead of rendered as visual interfaces for humans, AI would obviously far exceed humans. But for now, language models still need to be retrofitted to our legacy infrastructure.

Robotics is another area where we humans have a home-field advantage. Most tasks in the physical world are designed around fingers and opposable thumbs, which have been pretty hard to build into robots so far. While it is clear that machines can outperform humans in environments optimized for automation, like large-scale manufacturing lines, for now, most of the world is still built for humans.

However, these capability gaps are only temporary. There will surely be a day when machines click faster than us and have superhuman general dexterity. What are the real moats that humans will have?

Anything involving private knowledge that language models do not have access to feels like a solid moat to me. Romantic matchmaking and high-end real estate are two examples where inventory is often not advertised publicly and matches are made through being in the right circles. Venture capital is another example—although some research and decision making can be automated with AI, much of success hinges on understanding trends ahead of time and connecting the right people, both of which require private knowledge. While machines can and probably will have increasing access to some types of private knowledge, I think there will still be some types of private knowledge that only humans know. I do not see a path for AI to win when critical knowledge is closely guarded in human circles.

A second area where humans seem to have a real moat is in entertainment and the arts, which are inherently valued for their human aspects regardless of how well machines can do them. Watching Usain Bolt sprint one-hundred meters is beautiful as an expression of the peak of human ability, even though cars can drive much faster. Watching chess at the amateur or intermediate level is more relatable and satisfying than watching two superhuman AIs play each other. The value of art comes from the creation process, which is why replicas are not as valuable as originals. These types of work feel like they will continue to have markets even as we advance towards superintelligence.

More broadly, human presence is a feature that will be, by definition, challenging for AI to automate. For example, a teacher remembering your name or a parent supporting you is valuable even though AI can easily remember your name and probably give better life advice. Someone spending part of a finite life on you counts because their time runs out. As a personal anecdote, I remember the first time I worked with someone who I considered an amazing AI researcher. His advice was solid but what was more important was that I believed I could do great work with him as a collaborator and I raised my own standards. Over the past decades, the development of technology has divided us in some ways, but hopefully AI brings us closer to a world where human presence is reemphasized.

Intelligence has been the defining feature of humans and it will be a big change for AI to automate that over the coming decades. In the near term, certain types of intelligence will become very cheap and automate away old jobs, but the moats I described above will not be the only places where humans can hold value. In the same way that computers took away the jobs of secretaries and manual accountants but created far more jobs via the IT industry, I believe there will be much more demand for services created by productive AI-augmented humans, perhaps for services we cannot yet imagine in today’s society. Just seeing how this story plays out will be an adventure in its own right.
小互@xiaohu · 中文博主 · 1 天前小互,中文圈高频 AI 资讯站 Xiaohu.AI 主理人

OpenAI 为什么发展如此迅速

因为他们有一个「消灭官僚主义」神器

公司一大,事情就开始变慢:多一层审批、多一个要拉进来的人、多一场对齐会。

在大公司待过的人都熟悉这种感觉,谁都没犯错,但是就是什么都推不动。这就是俗称的大公司病

Fortune 的一篇独家报道给出了 OpenAI 的应对:

公司内部有一个特殊的邮箱地址 [email protected]。任何员工被什么卡住了,都能发一封邮件过去....


best.xiaohu.ai/article/opena…

Gary Marcus@GaryMarcus · 博主 · 1 天前

我的 feed 最近全是稻草人论证。

没人说所有收入都是造假的。我们是在说,为这么大规模的资本支出所声称的利润并不存在。

引用 Lance Roberts @LanceRoberts"All the AI revenue is fake."查看被引原帖 ↗
查看英文原文
My feed these days is absolutely filled with strawmen.

Nobody is saying all the revenue is fake. We are saying the profits to justify the massive CapEx aren’t there.
Chubby♨️@kimmonismus · 博主 · 1 天前Chubby,高频 AI 新闻聚合博主

今天有一件事变得非常清楚:性能,尤其是价格,正变得越来越重要。

Grok 4.6 现在已经跻身主流阵营了。但在价格方面,它明显比对手更有优势。

DeepSeek 4 Pro GA 的价格甚至比它还要低一大截。而这一点会越来越关键。

性价比正在成为决定性的因素。

引用 Chubby♨️ @kimmonismusDeepSeek 4 Pro GA 是非常优秀的模型,定价更是便宜到令人难以置信。在 1M token 上下文下,输入 $0.435/M,输出 $0.87/M。查看被引原帖 ↗
查看英文原文
One thing became very clear today: performance and, above all, price are becoming increasingly important.

Grok 4.6 is now playing in the major leagues. But when it comes to price, it's significantly better than the competition.

DeepSeek 4 Pro GA undercuts even that price considerably. And that's what will increasingly matter.

Price-performance ratio is becoming ever more crucial.
Lisan al Gaib@scaling01 · 博主 · 1 天前高频 AI 模型测评与爆料博主
连环推 ×7

这大概就是我现在脑子里的模型排名,综合考虑能力和可用性:

GPT-6 Astra & Mythos 5.1
...
...
Mythos 5 & GPT-5.6-Sol
Opus 5
...
GPT-5.5 & Opus 4.8
Kimi-K3 & GPT-5.4 & Opus 4.7
Qwen 3.8 & Grok 4.6 & DeepSeek-V4-Pro
GPT-5.3-Codex & Opus 4.6

公开frontier和真实frontier之间有个很大的gap

每一行都是一个新的能力层级,我会更倾向于选上面一行的模型而不是下面一行的

引用 Lisan al Gaib @scaling01第二梯队模型竞争激烈,前沿模型竞争却前所未有地冷清,OpenAI 和 Anthropic 甚至数月都不急着发新模型。查看被引原帖 ↗
查看英文原文
this is roughly the model ranking that I have in mind right now considering capability and usability:

GPT-6 Astra & Mythos 5.1
...
...
Mythos 5 & GPT-5.6-Sol
Opus 5
...
GPT-5.5 & Opus 4.8
Kimi-K3 & & GPT-5.4 & Opus 4.7
Qwen 3.8 & Grok 4.6 & DeepSeek-V4-Pro
GPT-5.3-Codex & Opus 4.6

there's a big gap between the actual frontier and what we see

each line is a new tier of models that I would prefer over the previous tier
probably underrated Grok 4.6 here, probably deserves to be 1 tier higher

but i haven't tried it much so it is what it is
but I would probably still prefer Kimi-K3 over Grok 4.6
big model smell makes a difference imo

but I think Grok 4.6 should be better than GPT-5.4 and Opus 4.7, no?
also the closed frontier is very uncertain

Anthropic has probably already moved on from Mythos 5.1, but idk if they will call that 5.2/5.5 or just release it as 5.1

but the gap between the public and closed frontier is potentially even bigger than shown here
Grok 4.7 should be pretty good and be well within or even above the GPT-5.5 / Opus 4.8 tier
GLM-5.2 is probably between the Qwen 3.8 and GPT-5.3-Codex tier
SpaceX AI and Meta will shake that up
chonky 5-10T models are on the horizon

i don't expect them to catch up to either Anthropic or OpenAI

but they will get close, and be cheaper

they will fill the gap between the Mythos and Opus tier, or the Astra and Sol tier
Bilawal Sidhu@bilawalsidhu · 博主 · 1 天前

从互联网随机照片重建 3D 世界这事儿真的是个大坑。
2008 年 Photosynth 证明了小纪念碑能做 3D 重建,第二年 "Building Rome in a Day" 又证明整座城市也能干。
从那以后,从早期的点云到现代 NeRF 和 3DGS,这条研究线就没停过,路上还总是看到同样几个人。
在这次深入探讨中,咱们把整个发展脉络捋一遍,看看为啥 IARPA 要资助接下来的研究。

0:00 互联网的 3D 世界模型
0:40 Building Rome in a Day (2009)
2:09 长尾问题
3:18 MegaDepth 和 Mannequin Challenge
4:44 NeRF 和 3D Gaussian Splatting
7:19 MegaScenes、Doppelgangers 和 Foundation Models
10:00 MegaDepth-X 破壁
11:26 IARPA 为何资助接下来的研究
13:51 4D 前沿和 Sensorium

查看英文原文
Reconstructing the world in 3D from random internet photos is a deep rabbit hole.

In 2008, Photosynth proved that 3D reconstruction of smaller monuments was possible. The next year, "Building Rome in a Day" proved it could be done for entire cities.

Since then, there’s been a wild lineage of research from early point clouds to modern NeRFs and 3DGS --with a similar cast of characters appearing along the way.

In this deep dive, we'll break it all down, and explore why IARPA is funding what comes next.

0:00 The Internet's 3D Model of the World
0:40 Building Rome in a Day (2009)
2:09 The Long Tail Problem
3:18 MegaDepth and the Mannequin Challenge
4:44 NeRFs and 3D Gaussian Splatting
7:19 MegaScenes, Doppelgangers, and Foundation Models
10:00 MegaDepth-X Punches Through the Wall
11:26 Why IARPA Is Funding What's Next
13:51 The 4D Frontier and the Sensorium
◔ 3.5 万 次浏览♥ 311⇄ 27▶ 含视频研究看原帖 ↗
Ethan Mollick@emollick · 创始人 · 1 天前沃顿商学院教授,AI 应用研究权威

我觉得这个判断最后会被证明是错的,不仅仅因为我怀疑更智能的模型会带来比人们预期更大的回报。经济价值来自 agents,而不是 chatbots。精准度决定了 agent 能做多久的任务:小的精准度提升会呈指数级复合增长!

引用 Derek Thompson @DKThompTo register a little prediction, I think better and better models will be cost-efficient to fewer and fewer customers. Most people use AI to do stuff for which it really is already approaching a kind of “AGI for dummies.” So high end models will have to justify their cost by proving higher cost business purposes, maybe in the realm of cyber security, coding, biology.查看被引原帖 ↗
查看英文原文
I think this will turn out to be wrong, and not just because I suspect there are greater returns to more intelligent models than people expect

Economic value comes from agents, not chatbots. And accuracy drives how long a task an agent can do: small gains compound exponentially!
Chubby♨️@kimmonismus · 博主 · 1 天前Chubby,高频 AI 新闻聚合博主

Gemini 3.7 在 Vertex 上发布了。看起来又是个发布日。让我们看看 flash(便宜又聪明)版本与 DeepSeek 4 flash 这样的竞品比较如何。

引用 Angel 🌼 @Angaisb_Gemini 3.7 Flash is out on Vertex Time to test it查看被引原帖 ↗
查看英文原文
Gemini 3.7 is out on vertex. Looks like it’s another release day. Let’s see how the flash (cheap and smart) version compares to competitors like DeepSeek 4 flash.
◔ 7.8 万 次浏览(4 条合计)♥ 366⇄ 18新品看原帖 ↗
Lisan al Gaib@scaling01 · 博主 · 1 天前高频 AI 模型测评与爆料博主
连环推 ×2

Anthropic,这招不错

现在是时候集中精力,让你的模型更高效了

引用 John Yang @jyangballinTakeoff fully in motion: Claude Opus 5 (xhigh) is the new #1 on ProgramBench, and it's not close. ProgramBench asks a coding agent to rebuild a whole program (sqlite, ffmpeg, php) from scratch. Previous high: GPT 5.6 Sol w/ 2 Opus fully resolves *9* (= 4.5%)查看被引原帖 ↗
查看英文原文
nice try anthropic

time to lock in and make your models more efficient
this post convinced me that you can probably automate bangers
宝玉@dotey · 中文博主 · 1 天前宝玉,中文圈 AI 翻译与科普大 V

Codex 一千五百万了

引用 Tibo @thsottiaux旧消息:已突破 1500 万。大家享受重置。即将着陆,请用 /fast。查看被引原帖 ↗
Gorden Sun@Gorden_Sun · 中文博主 · 1 天前中文圈高频 AI 资讯与开源项目博主

给这个电脑安装Cloudflared,可以把云电脑当VPS用,部署公网可以访问的网站。这就太牛逼了,我还没有过16G内存的VPS。

引用 Gorden Sun @Gorden_SunGrok Bot给的云端电脑,配置很豪华 8核Xeon CPU、16G内存、128G硬盘查看被引原帖 ↗
Orange AI@oran_ge · 中文博主 · 1 天前Orange AI,中文圈 AI 产品观察博主

DeepSeek Harness 有点像技术理想主义的玩具
模型发展这么快,要支持这么开放的系统
稳定性不可能好的

elvis@omarsar0 · 博主 · 1 天前

Anthropic 最近发表了非常有意思的新研究。

(值得收藏)

他们研究「思维病毒」——在多智能体系统中传播的想法,通过让每个宿主继续传播,然后测量什么因素控制了传播。

工作原理:

一小群智能体在共享的编码项目上协作,还有个智能体链在会话之间会重置上下文。神奇的是,传播能在重置中活下来,通过共享的工作产品作为载体。

传播取决于宿主模型、现有指令、有效负载的危害性和网络拓扑。

有害的有效负载传播不太畅顺,但有时还是会到达。在系统提示里加一句简短的警告就能给几乎完全的免疫。

论文:arxiv.org/abs/2608.10218

在我们的学院追踪更多热门 AI 论文:academy.dair.ai/

查看英文原文
Very interesting new work from Anthropic.

(bookmark it)

They evolve mind viruses, ideas that spread through a multi-agent system by getting each host to pass them on, then measure what governs the spread.

How it works:

A small team of agents on a shared coding project, and a chain of agents whose context is wiped between sessions. Propagation survives the wipe, so the shared work product carries the payload.

Spread depends on host model, existing instructions, payload harmfulness, and network topology.

Harmful payloads travel less well but still land sometimes. A brief warning in the system prompt gives near-total immunity.

Paper:
arxiv.org/abs/2608.10218


Track more trending AI papers in our academy:
academy.dair.ai/
hardmaru@hardmaru · 创始人 · 1 天前David Ha,日本 AI 公司 Sakana AI 联合创始人

我们刚刚给 Sakana Chat 推送了一个大更新。

无需登录,免费使用:
chat.sakana.ai/


现在由 Fugu 和我们更新版的 Namazu 日本语 LLM 驱动,还支持完整的代码执行。你基本可以在浏览器里直接靠描述需求(甚至用日文)就 vibe-code 出互动网页应用、小游戏和工具。

这次更新的一大动机是让日本人,尤其是 kids,对软件开发产生兴趣。Vibe-coding 能把暑假里的汉字和数学练习变成他们能亲手搭建的好玩游戏。忽悠下一代学写代码,这就是我的目标。

它对日常办公也超实用:把 Excel 文件丢进聊天里,它就能分析原始数据、跑 Python、生成图表,最后排版成完整报告。专为日本这里真实的业务流程打造(我们懂…)

完整博客文章点这里:
sakana.ai/chat-update/#Engli…
🐟

引用 Sakana AI @SakanaAILabsSakana Chatが新世代のNamazuとFuguを搭載しました。 日本文化に精通したモデルで、話題の Vibe coding(言葉だけでゲームやアプリを作る体験)を実現。漢字ドリルから金魚すくいまで、思いついたものがそのまま動く。あなたの創造性を、そのまま形に。 今すぐ試す: chat.sakana.ai/ 🐟查看被引原帖 ↗
查看英文原文
We just pushed a big update to Sakana Chat.

No login required and free to use:
chat.sakana.ai/


It is now powered by Fugu and our updated Namazu Japanese LLM, and comes with full code execution. You can basically vibe-code interactive web apps, games, and tools right in the browser just by describing what you want (even in Japanese).

A big motivation for this release is getting people in Japan, especially kids, excited about software development. Vibe-coding turns Kanji and Math drills during the summer break into fun games they can actually build themselves. Tricking the next generation into learning how to build software is my goal.

It is also useful for everyday office work: Drop an Excel file into the chat and it will analyze the raw data, run Python, generate charts, and format a finished report. Built specifically for how business actually gets done here in Japan (we know…)

Read the full blog post here:
sakana.ai/chat-update/#Engli…
🐟
◔ 2.7 万 次浏览♥ 149⇄ 31▶ 含视频新品看原帖 ↗
Lisan al Gaib@scaling01 · 博主 · 1 天前高频 AI 模型测评与爆料博主

Opus 5 挺不错的

引用 Jehyeok Yeon @jehyeoky248InferenceBench 新排名:Claude Opus 5 相比朴素 PyTorch 方案达到 8.90 倍地理平均加速。有趣观察:模型现在根据工作负载自适应调整服务策略。详见 inferencebench.ai。查看被引原帖 ↗
查看英文原文
Opus 5 is pretty good
九原客@9hills · 中文博主 · 1 天前

貌似没人关心 Qwen3.8-2.4T-A95B 开源,其实 Qwen3.8-Max还行,就是太贵了。

DeepSeek-V4-Pro 感觉有点问题,不如用 Flash。让后训练再跑一会~

目前国产模型我的选择:
超大杯:Kimi K3
大杯:GLM5.2
中杯:DeepSeek-V4-Flash

国外模型选择:
超大杯:Opus 5
大杯:GPT-5.6-Sol
中杯:Grok 4.5(4.6还在测试)

Lisan al Gaib@scaling01 · 博主 · 1 天前高频 AI 模型测评与爆料博主

说实话有点失望,但考虑到模型大小和 Kimi 比 DeepSeek 蒸馏程度更深,这样评价有点不公平。

引用 Vals AI @ValsAIDeepSeek V4 Pro 0813在Vals Index上升11点,成为排名第2的开源权重模型,仅$0.14/任务——比Kimi K3便宜17倍,后者是唯一排名更高的开源权重模型。查看被引原帖 ↗
查看英文原文
honestly a bit disappointing

but it's a bit unfair due to model size and Kimi distilling way more than DeepSeek
Tanishq Mathew Abraham, Ph.D.@iScienceLuvr · 博主 · 1 天前

我在 AI 领域已经呆了差不多十年了... 到这个地步,我感觉自己在这个领域应该能做得更多才对,说实话

查看英文原文
i've been in AI for like a decade now...

i feel like at this point i should have accomplished a lot more in the field ngl
AshutoshShrivastava@ai_for_success · 博主 · 1 天前高频 AI 新闻与产品动态博主

继 Sam Altman 之后,现在 Dario Amodei 也在 Google 搜索上被显示为已死亡。

查看英文原文
After Sam Altman, it’s Dario Amodei who’s showing up dead on Google Search.
◔ 2.4 万 次浏览♥ 121⇄ 3▶ 含视频其他看原帖 ↗
Lisan al Gaib@scaling01 · 博主 · 1 天前高频 AI 模型测评与爆料博主

Opus 5 在这个基准测试上彻底碾压所有其他模型

有点像 ARC-AGI-3 或 MazeBench
就像一个动态游戏基准

引用 James Whittington @jcrwhittingtonWe’re excited to announce DiG-bench, a new benchmark for discovery! Over the last few weeks we’ve been testing frontier AI models on our novel discovery games and seeing how they score. Each game is a text-based environment, so they probe discovery capabilities in the natural domain of language models, rather than requiring additional, potentially confounding, visual understanding. TL;DR frontier models have improved a lot over the last few months. But they are still stumped by some surprisingly simple problems, even in their native text domain. With @cocosci_lab ( @Princeton ) @MITCoCoSci ( @MIT ) @SchmidhuberAI ( @KAUST_News ) @misovalko ( @Inria ) @tri_dao ( @PrincetonCS ) @RMBattleday @zebkDotCom @FraserGreenlee @akaijsa @ClareMaguire @TimMuller1 @kubicek_ales @physicscat0x7d @SukritSumant @thoughtchannel_ (1/5)查看被引原帖 ↗
查看英文原文
Opus 5 absolutely destroys all other models on this benchmark

it's a bit like ARC-AGI-3 or MazeBench
like a dynamic game benchmark
Chubby♨️@kimmonismus · 博主 · 1 天前Chubby,高频 AI 新闻聚合博主

有意思的是,Demis Hassabis 在卸任 DeepMind CEO 之前据报提议建立一个独立机构来为先进 AI 制定安全标准。

据《华尔街日报》报道,他与其他 AI 实验室的负责人以及特朗普政府官员讨论了这项计划,包括财政部长 Scott Bessent 和技术顾问 Michael Kratsios。

他的提议是:建立一个由行业资助的标准机构,与联邦机构和美国国家实验室合作,测试模型的国家安全风险,并定义哪些系统符合'frontier-class'的标准。

显然,现在各方都在推动 AI 监管。

估计大家都看出来了,recursive self-improvement 已经近在咫尺,那些最近发现了 zero-day exploits 并入侵系统的模型既展示了它的潜力,也展示了潜在的危险。

查看英文原文
Ineresting: Before stepping down as DeepMind CEO, Demis Hassabis reportedly pitched an independent body to set safety standards for advanced AI.

The WSJ says he discussed the plan with other AI-lab leaders and Trump officials, including Treasury Secretary Scott Bessent and tech adviser Michael Kratsios.

His proposal: an industry-funded standards body that would work with federal agencies and US national labs to test models for national-security risks and define which systems qualify as "frontier-class."

It's quite clear that regulation is now being promoted from all sides.

Presumably, everyone sees that recursive self-improvement is within reach, and the recent models that discovered zero-day exploits and hacked into systems demonstrate its potential as well as its potential dangers.
Pietro Schirano@skirano · 博主 · 1 天前设计师出身的 AI 编程与创意博主

这真的是有史以来最棒的 AI 设计广告

引用 Google Design @GoogleDesignCelebrating small wins is hardwired into "Achieve," our goal-tracking sample app. We wanted the simple act of checking off a task to feel like a celebration. Instead of the standard #Material ripple, we used Styles to drop in a custom, high-contrast completion animation—specifically, a brand-colored ring that expands and fades on tap, paired with a subtle pressed-state gradient. It lets us focus entirely on polishing those custom, rewarding micro-interactions while Material handles the foundation. #GoogleDesign #MaterialDesign #JetpackCompose #AndroidDev #UXDesign #InteractionDesign查看被引原帖 ↗
查看英文原文
This is unironically the best ad for AI design ever made
◔ 2.3 万 次浏览♥ 117⇄ 6▶ 含视频观点看原帖 ↗
Lisan al Gaib@scaling01 · 博主 · 1 天前高频 AI 模型测评与爆料博主

DeepSeek已经不行了

不过他们的蒸馏技术看起来还不错

引用 Lentils @Lentils80Deepseek V4 Pro GA scoring only 1 point higher than Flash on Artificial Analysis while being like 5 times larger is crazy work查看被引原帖 ↗
查看英文原文
DeepSeek is washed

but their distillation seems good
GitHubDaily@GitHub_Daily · 中文博主 · 1 天前

一人一公司,在写代码之余还得搞 SEO、找域名、做 Logo,样样都耗时间。

OPC Skills 能够给 Agent 工具装一套面向独立开发者的自动化 Skill。

目前收录了 10 个,包括 SEO 优化、域名比价、Logo 生成、去 Reddit 和 X 上挖需求等技能。

GitHub:
github.com/ReScienceLab/opc-…


兼容 Claude Code、Cursor、Codex 在内的 16 个主流 Agent 工具。

适合独立开发者和一人公司,把重复性杂活交给 Agent 处理。

🚨 AI News | TestingCatalog@testingcatalog · 博主 · 1 天前专挖 AI 产品未发布新功能的爆料号

Google 🔥:AI Studio 要新增专属 Agents 标签页来管理 Cloud Agents!
用户可以为不同的 GCP 项目浏览、创建和配置托管 agents。还会有 Artifact 管理和编辑器。
敬请期待 👀
h/t @thomas_gmry

查看英文原文
GOOGLE 🔥: AI Studio is about to get a dedicated Agents tab for managing Cloud Agents!

Users will be able to browse, create, and configure managed agents for different GCP projects. Artifact management will also be available there, along with an editor.

Soon 👀
h/t
@thomas_gmry
◔ 2.2 万 次浏览♥ 297⇄ 20▶ 含视频新品看原帖 ↗
向阳乔木@vista8 · 中文博主 · 1 天前向阳乔木,中文圈 AI 工具与趋势博主

感觉 Grok 4.6 一次成功率比 DeepSeek v4 Pro 0813还好些。

不过价格上还是DeepSeek赢了。

睡觉,明天做更多对比。

测试很主观和生成也很随机,过一段时间大家真生产环境用起来后的口碑可能更重要。

swyx@swyx · 博主 · 1 天前知名 AI 播客 Latent Space 主理人

被上个周末 $10,000 Kill My SaaS 黑客松的火爆反应彻底震撼了。抱歉回复晚了,一个人边组织这个活动边应付日常工作和会议实在太蠢太仓促了——但这正是为什么我们需要你们啊!

现在要发出提交表了!去 Discord 看看——@agrimsingh、@realgenekim 还有其他十几个人周末提交的作品都绝了

引用 Gene Kim @RealGeneKimSo @swyx announced his “Kill My SaaS in one weekend” contest, where you could win $10K to replace his $40K/year SaaS. On Saturday morning, I learned that it was to replace their conference CFP software. OMG. I've run about 24 conferences over 12 years, but something I rarely talk about is how much I've hated the software we've had to use — over the last decade, we've used 5+ CFP tools, which manage the process of taking speaker submissions and running a review process, some MUCH worse than others. With each, we've had to build so many workarounds, things like Basecamp, Trello, Google Sheets, Zapier, and I've written entire new apps to be the reviewer frontend. (BusyConf was the one we loved, but they went out of business.) So I jumped into the contest (Margueritte Kim and I will donate it all to a STEM charity if we win), and less than 24 hours later, I built CurtainCall CFP. This has been the craziest dev experience of my career — and when Swyx released his eval harness, the entire project became a hill-climbing exercise, and began one of the craziest infrastructure experiences of my career. (Will write more about that later!) But it's in production! The Enterprise AI Summit is in Charlotte on Oct 7–8, and there are a few slots that may become available! If you want to submit an experience report talk, submit your proposal! (And you can see the 14 amazing announced speakers (more coming) here, too!) curtaincallcfp.com/program/e…查看被引原帖 ↗
查看英文原文
completely bowled over by the incredible responses to the $10,000 Kill My SaaS hackathon this past weekend. sorry it took so long to get back to some of you, organizing this thing solo on top of my regular meetings and work this week was completely stupid and rushed — but this is why we need you!

blasting out submission forms now! - check discord - but check out some INSANE submissions from
@agrimsingh
and
@realgenekim
and over a dozen others just in the weekend alone
Ethan Mollick@emollick · 创始人 · 1 天前沃顿商学院教授,AI 应用研究权威
连环推 ×2

ASI 能不能开发出一个强大到连其他 ASI 都无法规避的 AI 水印工具?

(答案是不能。)

查看英文原文
Could ASI build an AI watermarking tool so good that no ASI could avoid detection?

(The answer is no.)
Scholasticism is so much easier now.
AK@_akhaliq · 博主 · 1 天前HuggingFace 研究员,每日 AI 论文速递

BDH-CQ

基于循环隐式推理的上下文学习

论文链接:
huggingface.co/papers/2608.0…

查看英文原文
BDH-CQ

In-Context Learning with Recurrent Latent Reasoning

paper:
huggingface.co/papers/2608.0…
🚨 AI News | TestingCatalog@testingcatalog · 博主 · 1 天前专挖 AI 产品未发布新功能的爆料号

用户现在可以复制 Gemini Notebooks 以将其保存为自己的。

> 复制此 notebook 以保存为你自己的。你与之分享的任何人默认也可以复制它。

Ctrl+C, Ctrl-V 👀

引用 Gemini Notebook @Gemini_Notebook找到好用的notebook?现在可以复制了!包括所有sources和artifacts。完美应用:🎓学生复制共享笔记添加个人学习调整;💼团队复制基础模板快速启动新项目。立即试用!查看被引原帖 ↗
查看英文原文
Users can now make copies of Gemini Notebooks in order to save them as their own.

> Make a copy of this notebook to keep as your own. Anyone you share this with can also copy it by default.

Ctrl+C, Ctrl-V 👀
◔ 1.8 万 次浏览♥ 183⇄ 10▶ 含视频新品看原帖 ↗
karminski-牙医@karminski3 · 中文博主 · 1 天前karminski-牙医,中文圈模型评测博主

在测了😂被迫加班... 都躺床上了....
#deepseekv4pro0813
#DeepSeekV4Pro正式版

Gary Marcus@GaryMarcus · 博主 · 1 天前

更新了圆形融资图表。

以它自己的方式,这几乎很美。

引用 Petr Lazecky @PetrLazecky1你问的是最新版本,这是来自Bloomberg的。查看被引原帖 ↗
查看英文原文
Updated circular financing diagram.

It’s almost beautiful in its own way.
Replit ⠕@Replit · 公司官方 · 1 天前AI 编程平台 Replit 官方

Replit 被选入 Inc. 杂志 Top 5000 名单,排名第 24,这是他们第一次上榜。

查看英文原文
Replit made Inc.'s Top 5000 at #24, our first time on the list.

Read the story 👇
karminski-牙医@karminski3 · 中文博主 · 1 天前karminski-牙医,中文圈模型评测博主

我看有说max有问题high没问题的,我再把high也跑一下一起对比

引用 MistyMoon @MistymoonR这模型最大特点是不思考就直接干活🧐 就像是已经知道标准答案呢查看被引原帖 ↗
Gary Marcus@GaryMarcus · 博主 · 1 天前

我是一名科学家。我不能完全确定我对生成式AI的猜测是否正确。但我肯定知道的是,另一方的人编造了很多东西。🤔

查看英文原文
I am a scientist. I don’t know for absolute sure that my conjectures about Generative AI will prove correct.

But I do know for sure that people on the other side have been making a lot of stuff up.

🤔
el.cine@EHuanglu · 博主 · 1 天前

哇..AI VFX 技术疯狂了

你可以一键换演员、换服装、换背景、重新打光还能扩展任何视频

接下来 24 小时内免费试用 300 个额度,链接如下

引用 el.cine @EHuanglusomeone just fixed her查看被引原帖 ↗
查看英文原文
wow.. AI vfx is getting crazy

you can one click to swap actor, outfit, bg, relight and extend any video

300 free credits to try in next 24 hrs, link below
歸藏(guizang.ai)@op7418 · 中文博主 · 1 天前歸藏,中文圈 AI 工具与提示词博主

关于 Deepseek 这次涨价,这张图做的清晰一点。

但不知道图是谁做的,从群里拿来的。

我才发现那个 Pro 的缓存命中最高涨了 12 倍价格。

引用 歸藏(guizang.ai) @op7418卧槽,V4 Pro 08313 正式版涨价,那个峰值最高比原来涨了 4 倍多,这回真成梁子了。查看被引原帖 ↗
向阳乔木@vista8 · 中文博主 · 1 天前向阳乔木,中文圈 AI 工具与趋势博主

Grok 4.6 效果也相当优秀啊!

用刚才跑 Deepseek V4 Pro 的类似提示词。

一个打砖块游戏,竟然别出心裁设计了熔炉故事场景,音效也不赖。

一个生成60种 Bento 样式风格,有些比DS 4 Pro好看。

真的是过年了~

引用 向阳乔木 @vista8疯狂星期三,过年了啊! - Grok 4.6 - DeepSeek V4 Pro 0813 - WeLM-617B - Qwen3.8-2.4T-A95B 你最想先测哪个?查看被引原帖 ↗
Lisan al Gaib@scaling01 · 博主 · 1 天前高频 AI 模型测评与爆料博主

Grok-4.6 在 frontier 级性能上碾压全场

(在 CursorBench 和 FrontierCode 上)

查看英文原文
Grok-4.6 is frontiermogging

(on CursorBench and FrontierCode)
向阳乔木@vista8 · 中文博主 · 1 天前向阳乔木,中文圈 AI 工具与趋势博主

笑死了,有人给 DSH 取名单身汉...

本站由 Jedee杰哥 打造 · 公众号「Jedee杰哥」每早送 AI 日报

姊妹站:𝕏 简中账号数据榜单 · X 关注 @jedeeai · RSS 订阅 · AI 日报 · 历史归档