JEDEE AI
存档 2026-08-21

8 月 21 日(北京时间)全球 AI 圈推文存档,按曝光排序,共 100 条。
← 返回最新 全部归档

全部情报 每小时更新 · 事件已合并同类项

内容 公司
Claude@claudeai · 公司官方 · 1 天前Claude 产品官方账号
连环推 ×2

Claude Academy 现已上线。

无论你是在弄清楚什么是 AI,还是已经每天都在用 Claude,都有一条适合你的学习路径。课程和教程免费开放给所有人,访问 academy.claude.com

查看英文原文
Claude Academy is now live.

Whether you're figuring out what AI is or already using Claude every day, there's a path that meets you where you are. The courses and tutorials are free and open to anyone at
academy.claude.com
Learn more about how we approach teaching and learning AI, and where Claude Academy fits in:
claude.com/blog/anthropics-a…
◔ 934.5 万 次浏览(5 条合计)♥ 3.7 万⇄ 4,225新品看原帖 ↗
Tibo@thsottiaux · 公司官方 · 1 天前ChatGPT 产品官方账号

我们调查了几条关于 Codex 使用限制有所不同的反馈。这不是我们在没有充分与社区沟通和保持透明的情况下会改变的。

我们发现的是,许多反馈用户都在使用 sub2api。将订阅转换为 API 流量,然后重新提供给多个用户共享,这不是我们支持的做法,这类使用会被我们的反欺诈系统检测到。

如果你通过 Sign in With ChatGPT 使用订阅,完全没有问题,可以通过官方客户端或任何支持用账号登录的开源客户端(比如 Pi、OpenCode 等)来使用你包含的使用额度。

查看英文原文
We've investigated a few messages about codex usage limits being different. That's not something we change without engaging the community and being transparent.

What we did see is that when talking to affected users many were using sub2api. Converting a subscription into api traffic to then re-serve or share across many users is not something we support and this type of usage gets flagged by our fraud-prevention systems.

You are completely fine if you use your subscription through Sign in With ChatGPT, either through the official clients or through one of the many OSS clients (Pi, OpenCode, ...) that support signing in with your account and using your included usage.
Tibo@thsottiaux · 公司官方 · 1 天前ChatGPT 产品官方账号

又是我。我带来了好消息。

首先,Codex 的活跃用户本周达到了 2000 万。其次,这值得庆祝,我们会给每位 Codex 和 ChatGPT Work 用户赠送一个 BANKED reset,你可以随时自己用。还有其他好消息稍后公布!

关于使用限额消耗快的问题,虽然我们目前没看到异常,但我们非常重视这件事,正在调查。如果有发现,我会分享给大家。我下面的那条帖子主要是澄清一个特定的模式。

今天去做点了不起的事吧。

引用 Tibo @thsottiauxWe've investigated a few messages about codex usage limits being different. That's not something we change without engaging the community and being transparent. What we did see is that when talking to affected users many were using sub2api. Converting a subscription into api traffic to then re-serve or share across many users is not something we support and this type of usage gets flagged by our fraud-prevention systems. You are completely fine if you use your subscription through Sign in With ChatGPT, either through the official clients or through one of the many OSS clients (Pi, OpenCode, ...) that support signing in with your account and using your included usage.查看被引原帖 ↗
查看英文原文
It's me again. I come bearing great news.

First of all, we have hit 20M active users for Codex some time this week. Second of all, this is cause for celebration and during the day we will credit every Codex and ChatGPT Work user with a BANKED reset that you can use at your own leisure. And we will have some other good news later too!

Now, on usage limits draining faster, while we're not seeing anything abnormal, we do take it incredibly seriously and there is an ongoing investigation. I will share if we do find anything and my below post is really a clarification on a specific pattern that we did see that I wanted to call out.

Go do something amazing today.
◔ 312.6 万 次浏览(4 条合计)♥ 1.9 万⇄ 945新品看原帖 ↗
DeepSeek@deepseek_ai · 公司官方 · 1 天前深度求索,国产开源大模型标杆
连环推 ×4

DeepSeek-V4-Flash-Vision-Exp 现已在 DeepSeek API 平台上线!🚀

🔹 这是一个实验性多模态模型,文本能力与 DeepSeek-V4-Flash 相当——包括代理、推理和知识库。
🔹 在多模态代理基准测试中,V4-Flash-Vision-Exp 相比 V4-Flash 有了巨大飞跃,多模态代理性能已接近 Opus-4.8。

用 model='deepseek-v4-flash-vision-exp' 试试吧。DeepSeek Harness 0.1.1 今天发布,已原生支持新模型。

1/n

查看英文原文
DeepSeek-V4-Flash-Vision-Exp is now live on the DeepSeek API Platform! 🚀

🔹 This experimental multimodal model matches DeepSeek-V4-Flash on text capabilities—including agents, reasoning, and world knowledge.
🔹 On multimodal agent benchmarks, V4-Flash-Vision-Exp makes a major leap over V4-Flash, bringing multimodal agent performance close to Opus-4.8.

Try it with model='deepseek-v4-flash-vision-exp'. DeepSeek Harness 0.1.1 was released today with out-of-the-box support for the new model.

1/n
Multimodality unlocks more agent use cases. 👀

V4-Flash-Vision-Exp works smoothly across agent frameworks, combining visual understanding with a wide range of tools to unlock more practical workflows.

2/n
Multimodal API support 🔌

🔹 Set model='deepseek-v4-flash-vision-exp'
🔹 Images are tokenized for billing: up to 384 tokens each, at V4-Flash pricing
🔹 Supports Chat Completions, Messages & Responses
🔹 Supports mixed text + image input; images can be provided via base64, external URLs, or the Files API.

Docs:
api-docs.deepseek.com/guides…


3/n
Files API is now live. 📁

🔹 Free to use
🔹 Upload an image once, then reference it by file_id to save request bandwidth
🔹 Reuse the same image across requests—no need to upload it again

Learn more:
api-docs.deepseek.com/guides…


4/n
◔ 235.9 万 次浏览(9 条合计)♥ 1.1 万⇄ 1,144▶ 含视频新品看原帖 ↗
Grok@grok · 公司官方 · 1 天前马斯克 xAI 旗下聊天机器人 Grok 官方

在 grok.com、iOS 和 Android 上构建应用现在可用于所有 SuperGrok 和 X Premium 计划

一条提示词就能把想法变成一个拥有自己域名的已发布产品。

查看英文原文
Building apps on
grok.com
, iOS, and Android is now available on every SuperGrok and X Premium plan

One prompt turns an idea into a published product with its own domain.
◔ 216.3 万 次浏览(2 条合计)♥ 3,470⇄ 397▶ 含视频新品看原帖 ↗
OpenRouter@openrouter · 公司官方 · 1 天前
连环推 ×2

🥷 全新隐秘模型:Ox Alpha

Ox Alpha 是一个为高效编码、持续智能体工作和实际生产使用而开发的前沿模型。

- 1M token 上下文窗口
- 支持文本、图像和视频输入

现在就试用,分享反馈帮我们改进模型!
openrouter.ai/stealth/ox-alp…

查看英文原文
🥷 New stealth model: Ox Alpha

Ox Alpha is a frontier model built for efficient coding, sustained agentic work, and real-world production use.

- 1M token context window
- Text, image, and video input

Try it now and share feedback to improve the model!
openrouter.ai/stealth/ox-alp…
Notes for this stealth model:

💰 It is free
🔑 This time, the provider does not train on your prompts or completions
◔ 306.5 万 次浏览(13 条合计)♥ 3,728⇄ 272新品看原帖 ↗
OpenAI Developers@OpenAIDevs · 公司官方 · 1 天前OpenAI 开发者平台官方
连环推 ×2

API 中的 GPT-Image-2 现已支持透明背景预览版。生成可重复使用的资产,可以放在任何背景上——用于产品图像、图形设计、网站模型和营销活动。

查看英文原文
Transparent backgrounds are now available in preview for GPT-Image-2 in the API.

Generate reusable assets you can place on any background—for product imagery, graphic design, website mockups, and marketing campaigns.
More on transparency coming soon 😉


developers.openai.com/cookbo…
◔ 205.4 万 次浏览(6 条合计)♥ 5,064⇄ 415新品看原帖 ↗
Bindu Reddy@bindureddy · 创始人 · 1 天前Abacus.AI CEO,AI 行业观点博主

尝试 RouteLLM API - 一个根据你的提示词自动路由到最佳模型的 API

- 处理缓存
- 支持 150+ AI 模型
- 速度快、价格不错
- 可在 Claude 或 Codex 中使用

简单问题路由到开源模型,复杂长期任务路由到前沿模型

查看英文原文
Try RouteLLM API - An API that routes to the best model based on your prompt

- Handles caching
- Supports 150+ AI models
- Excellent speed and prices
- Works in Claude or Codex

Route to open source models for simple turns and frontier models for complex long-running tasks
Claude@claudeai · 公司官方 · 1 天前Claude 产品官方账号
连环推 ×2

我们最近看到的一些最喜欢的 Claude Code 项目:

引用 Mannay 🌹 @mannayYou can just draw faces with javascript // coding doodles查看被引原帖 ↗
查看英文原文
Some of our favorite Claude Code projects we've seen lately:
What are you building with Claude?
OpenAI Developers@OpenAIDevs · 公司官方 · 1 天前OpenAI 开发者平台官方
连环推 ×2

在你的 ChatGPT Sites 添加队友做编辑,这样大家就能一起构建和发布了。

协作者可以同时推送改动到同个项目,Codex 在后台搞定 git 管理和 CI。

查看英文原文
Add teammates as editors to your ChatGPT Sites so you can build and publish together.

Collaborators can push changes to the same project while Codex handles git management and CI behind the scenes.
You can also customize your ChatGPT Site’s URL so it reflects what you’ve built and is easier to recognize and share.


learn.chatgpt.com/docs/sites
◔ 60 万 次浏览♥ 1,013⇄ 79▶ 含视频新品看原帖 ↗
OpenAI Developers@OpenAIDevs · 公司官方 · 1 天前OpenAI 开发者平台官方

@ExaAILabs 新插件让 ChatGPT Work 和 Codex 能访问 100B+ 网站、论文、文档等资源

引用 Exa @ExaAILabs在Codex和ChatGPT中推出Exa:为Codex提供访问100B+网站、文档、论文、人物、公司等的权限。打开Codex -> 插件 -> "Exa" -> 安装查看被引原帖 ↗
查看英文原文
Give ChatGPT Work and Codex access to 100B+ websites, papers, docs, and more with the new
@ExaAILabs
plugin
◔ 46.8 万 次浏览♥ 3,519⇄ 210▶ 含视频新品看原帖 ↗
Tibo@thsottiaux · 公司官方 · 1 天前ChatGPT 产品官方账号

创建一个 ChatGPT 网站,分享出去,一起协作创造点东西。我一直在和别人一起做小游戏和小网站,简直好玩炸了。

引用 OpenAI Developers @OpenAIDevs将团队成员设为ChatGPT Sites编辑,共同构建发布。协作者可推送项目变更,Codex负责git管理和CI。查看被引原帖 ↗
查看英文原文
Create a ChatGPT Site, share it and create something together. I've been building little games and sites with others and it is so much fun.
◔ 45.2 万 次浏览♥ 1,868⇄ 70▶ 含视频演示看原帖 ↗
Perplexity Developers@perplexitydevs · 公司官方 · 1 天前AI 搜索引擎 Perplexity 官方

Perplexity Agent API 现在让开发者通过一个端点访问来自 9 个不同供应商的 41 个前沿模型。可以用网络搜索、财务搜索、获取和沙箱代码执行等内置工具来构建多模型 agent 工作流。

查看英文原文
The Perplexity Agent API now gives developers access to 41 frontier models across 9 providers in one endpoint.

Build multi-model agent workflows with built-in tools like web search, finance search, fetch, and sandboxed code execution.


pplx.ai/agent-api-blog
◔ 36.3 万 次浏览(2 条合计)♥ 68⇄ 6新品看原帖 ↗
Guillermo Rauch@rauchg · 创始人 · 1 天前Guillermo Rauch,Vercel 创始人兼 CEO

我们在为 agents 构建 AWS

引用 Vercel Developers @vercel_devYou can now manage Vercel Container Registry with Vercel CLI. ▲ ~/ vercel vcr build docker --push vercel.com/changelog/manage-…查看被引原帖 ↗
查看英文原文
We’re building AWS for agents
Aravind Srinivas@AravSrinivas · 创始人 · 1 天前Perplexity 联合创始人兼 CEO

一个完整的 AI 开发者平台应该给你提供各种模型的访问权限(前沿模型和主力模型),以及在生产工作中部署它们的工具。基本上 Perplexity Agent API 就是这样做的。

引用 Perplexity Developers @perplexitydevsThe Perplexity Agent API now gives developers access to 41 frontier models across 9 providers in one endpoint. Build multi-model agent workflows with built-in tools like web search, finance search, fetch, and sandboxed code execution. pplx.ai/agent-api-blog查看被引原帖 ↗
查看英文原文
A full-fledged developer platform for AI should provide you with access to various models (frontier and workhorse), as well as tools for deploying them in useful production workloads. That is basically the Perplexity Agent API.
Ethan Mollick@emollick · 创始人 · 1 天前沃顿商学院教授,AI 应用研究权威

我用 GPT-5.6 Sol 生成最'Claude味儿'的搞笑图片,出来的效果绝了,简直一模一样。

查看英文原文
I asked GPT-5.6 Sol to create the most Claude-y possible parody image and what it came up with is pretty great and dead-on.
Amjad Masad@amasad · 创始人 · 1 天前Amjad Masad,Replit 创始人兼 CEO

与 @OpenAI 的这次合作早就该来了。

Replit 进 YC 之前,PG 就让 Sam 来招我们。

这是 Sam 亲自讲我们初次相遇的故事:

引用 Replit ⠕ @ReplitReplit Free Mode由OpenAI GPT-5.6 Luna驱动。让智能对所有人可获。查看被引原帖 ↗
查看英文原文
This partnership with
@OpenAI
is long overdue.

Before Replit was in YC, PG asked Sam to recruit us.

Here’s Sam telling the story of how we first met:
◔ 25 万 次浏览♥ 1,307⇄ 55▶ 含视频动态看原帖 ↗
Chubby♨️@kimmonismus · 博主 · 23 小时前Chubby,高频 AI 新闻聚合博主
连环推 ×4

不可能!@synthwavedd 说 OpenRouter 上的秘密模型 Ox Alpha 就是即将推出的 GLM 5.3 Flash。一个 Flash 模型竟然超过 Fable 和 GPT Sol?!要真这样的话,zAI 的后训练让西方模型相形见绌。哦还有,Kimi K3.1 也快来了。现在怎么回事啊。难以置信

引用 leo 🐾 @synthwaveddA new Kimi model, likely K3.1, is now being tested on the Code @arena under the name "korrine" K3 was tested on the Arena as "kivine" prior to its launch If anyone's wondering, "Ox Alpha" on OpenRouter is the upcoming GLM 5.3 Flash from fellow Chinese lab Zhipu查看被引原帖 ↗
查看英文原文
No way!
@synthwavedd
says that the secret Model "Ox Alpha" on OpenRouter is the upcoming GLM 5.3 Flash

A Flash model (!!) outperforming even the best models like Fable and GPT Sol?! If this turns out to be true, zAI's post training put western models at shame

Oh, and looks like Kimi k3.1 also incoming

What is happening right now. unbevlieable
How did we end there? Googles Gemini Flash isnt even close to a chinese Flash now o.O
Caveat: I cant confirm anything and I dont have any intern information. But leo is a reliable leaker.
i mean, do you understand what this means? I just wrote about this:

鉴定当代科技大傻逼的10个问题:

1. 相信春晚宇树科技的机器人里有一个AI模型;

2. 相信中国大模型厂商(deepseek、阿里、六小虎)卖API实际上是卖电力,是工业克苏鲁+基建狂魔的碾压红利——美国现在太缺电了,OpenAI连灯都用不起,美国AI把全球电能全用完了,美国AI只有死路一条;

3. 相信中国春节是全球科技峰会,国产大模型突发式发布模型碾压OpenAI,宇树科技再爆黑科技,全球瞩目,以后全世界不看OpenAI发布会、不看苹果发布会,光看中国春晚就够了;

4. 相信AI目前最大的短缺是发电厂、内存、存储、铜、银这些因素,整个OpenAI和AI行业忙得团团转,OpenAI高管集体给发电厂磕头,求他们赶紧供电,整个硅谷和华尔街惊呼“建发电站、扩张存储芯片产能、全球抢铜矿这三件事,比训练GPT-6还难100倍”;

5. 相信中国AI在工业碾压全世界,中国AI不拼大模型,中国AI只拼制造业,江浙沪和深圳已经全是黑灯工厂,大到iPhone生产线,小到服装厂玩具厂,里面有10万个宇树科技机器人在里面24小时上班——中国制造业早就不用人了,开始进入赛博朋克时代;

6. 相信deepseek太牛逼了,模型一发布,硅谷哭了,白宫黄了,华尔街崩盘了,英伟达、Google、OpenAI排着队破产,被deepseek直接一锅端,全被干死了;

7. 相信大模型100%全靠数据,中国数据太多了,中国数据全球第一多,中国人到处造数据,美国人没有数据,所以中国大模型必赢,OpenAI天天搞不到数据在家快急死了,未来只有死路一条;

8. 相信中国新能源车碾压欧美,中国独创式研发了汽车的沙发彩电大冰箱,中国发明了车机系统,欧美日韩车厂打破脑袋也想不到这么厉害的发明,在全球市场被打得节节败退——美国人做梦都想买一台零重力座椅的中国国产新能源车;

9. 相信中国无人驾驶吊打特斯拉,特斯拉在中国连第四梯队都进不去,华为鸿蒙智驾领先特斯拉10年以上,全球无人驾驶最强的就是华为鸿蒙;

10. 相信中国AI未来一定会打败美国,无数事业单位、行政机关、国企里面早就全部署deepseek开源一体机了,现在每个单位都有1000个AI agent在里面智能化无人化办公,中国AI落地远比美国快,美国人绝大多数还没用过ChatGPT,已经对中国deepseek AI一体机馋疯了。

相信以上任何一条,基本可以断定是一个三本文科师范毕业的低智商、低学历、低认知的大傻逼。

OpenAI Developers@OpenAIDevs · 公司官方 · 1 天前OpenAI 开发者平台官方
连环推 ×2

欧洲用户有两个更新给你们。

Computer History 现在可以在 Mac 上用了,覆盖 EEA、英国和瑞士,面向 ChatGPT Pro、Business 和 Enterprise 用户。

ChatGPT 现在能记住你在各个应用和网站上的活动,这样以后聊天会更贴心,说明也少了。

引用 OpenAI Developers @OpenAIDevsCodex and ChatGPT can now understand the context of your recent work. Opt into Computer History to give ChatGPT richer context, so it can pick up where you left off, understand patterns in your work, and suggest skills or scheduled tasks for work you repeat.查看被引原帖 ↗
查看英文原文
Europe, two updates for you.

Computer History is now available in the EEA, UK, and Switzerland for ChatGPT Pro, Business, and Enterprise users on Mac.

ChatGPT can remember activity across your apps and websites, so future conversations feel more personalized and require less explanation.
Record & Replay is also available in the ChatGPT app for macOS across the EEA, UK, and Switzerland.

Turn your go-to workflows into reusable skills by showing, not just telling, ChatGPT Work and Codex what to do.
Chubby♨️@kimmonismus · 博主 · 1 天前Chubby,高频 AI 新闻聚合博主

如果这真的是 GLM-5.4/5.5,那将彻底改变一切,毫不夸张。

GLM-5.3 才在 7 天前发布,相比 GLM-5.2 已经取得了极其显著的飞跃,而 GLM-5.2 只是通过强化学习(RL)优化的同一个基础模型。这些进展全都发生在非常短的时间内。

如果 GLM-5.4 真的在短短一周后因为 RL 就变好这么多,这将证明:

1) 现在模型的进化速度快到什么程度。不仅看不到尽头,而且:比以往任何时候都快,呈指数级增长。

2) 这会逼迫 OpenAI 和 Anthropic 必须发布新模型。特别是 Anthropic,随着即将到来的 IPO,现在必须证明自己。这给那些想放慢速度优先安全的人施加了巨大压力。

3) 至少同样重要的是:中美之间的差距在进一步缩小,而且速度还更快了。看起来普遍认为这是个中国模型。没人会怀疑这是 Google 的模型。

鉴于 tokenizer 相同,最有可能就是 GLM,这将是我们长期以来见过最疯狂的事,原因就是上面这些。

但也许是 MiMo。或者相当牵强地说,可能是 Ilya Sutskever 的 SSI 模型。

反正很令人期待。我很少见到社区如此印象深刻又困惑的同时出现。我也是,既困惑又印象深刻。

@davis7 用 Fable 来判断它最接近哪个模型。测试截图已附。h/t Ben Davis

引用 Ben Davis @davis799% sure it's GLM-5.x, all the evidence points to it (same video encoder, same tokenizer, style matches, same audio rejection, etc.) The stuff about their RL env in the GLM-5.3 announcement seemed really cool and like it could go somewhere, did not expect it to get this good this fast but here we are This thing seems absurdly good. First tests I've done have all been excellent. Already done a couple of back to back tests against GPT/Fable. It's actually doing better than they are... Seems like OpenAI and Anthropic are gonna have to start dropping models again, because this thing is (or at least seems to be) the new state of the art...查看被引原帖 ↗
查看英文原文
If it's true that this is indeed GLM-5.4/5.5, then it would change everything, without exaggeration.

GLM-5.3 was released just 7 days ago and was an extremely significant leap compared to GLM-5.2, which was improved solely through real-time modeling (RL). Same base model. And all this in a very short time.

If it's true that GLM-5.4 has become so much better just a week later thanks to RL, it would demonstrate:

1) how much faster the models are now becoming. Not only is there no end in sight, but: now more than ever, exponential growth.

2) It would force OpenAI and Anthropic to release models. Anthropic, in particular, with its upcoming IPO, now has to prove itself. And it would put pressure on slowing down in favor of security.

3) And at least as importantly: the gap between China and the US is shrinking even further, even faster. It seems to be generally accepted that this is a Chinese model. No one suspects it's a Google model.

Given the same tokenizer, it's most likely GLM, and that would be the craziest thing we've seen in a long time for the reasons mentioned above.

But perhaps it's MiMo. Or, quite far-fetched, Ilya Sutskever's SSI model.

It remains exciting. I've rarely seen the community so impressed and confused at the same time. I'm equally confused and impressed.


@davis7
used Fable to determine which model it most closely resembles. A screenshot of the test is attached. h/t Ben Davis
OpenRouter@openrouter · 公司官方 · 1 天前

那么大家怎么看呢?模型实验室应该怎样改进该模型?它真正的优势在哪里?

引用 Toven @pingTovenhope you’re enjoying ox alpha :)查看被引原帖 ↗
查看英文原文
So, what does everyone think???

What should the model lab do to improve the model? Where does it really shine?
MiniMax Design (H3)@Hailuo_AI · 公司官方 · 1 天前MiniMax 旗下海螺 AI 视频官方
连环推 ×2

🐙MiniMax Design 正式上线!将创意想法变成商业级内容

#MiniMaxDesign
#MiniMax

🤖 Agent 驱动工作流:输入目标,Agent 自主规划执行并交付结果
🪄 为商业创意量身定制:做广告、电商物料、动态视频和后期编辑
🧩 开放灵活的集成:接入本地资产、私有部署,通过 API 扩展

查看英文原文
🐙MiniMax Design is officially live!
Turn Ideas into Commercial-Grade Content

#MiniMaxDesign
#MiniMax


🤖Agent-Driven Workflow: Input your goal. Agents autonomously plan, execute, and deliver.
🪄Built for Commercial Creation: Craft ads, e-commerce assets, motion graphics, and post-production content.
🧩Open & Flexible Integration: Connect local assets, deploy privately, and scale through APIs.
Download Now>>
design.minimax.io
#MiniMaxDesign
◔ 20.1 万 次浏览(3 条合计)♥ 2,004⇄ 190▶ 含视频新品看原帖 ↗
Chubby♨️@kimmonismus · 博主 · 1 天前Chubby,高频 AI 新闻聚合博主

天哪!Ben 用这个神秘模型跑了 10 个 DeepSWE 任务,结果破 80%,而 Fable 才 65%,GPT-5.6-sol 仅 52%!这太疯狂了。很可能是中国公司。我觉得要么是新的 GLM,要么是 Kimi 模型。

引用 Ben Davis @davis7I ran this thing through 10 tasks on DeepSWE (so there could be a ton of variance in it's real score, this is a subset), but uh... gpt-5.6-sol: 52% fable: 65% whatever the hell this is: 80% (was a near miss on the "x"s so actually over 80%) I am very confused查看被引原帖 ↗
查看英文原文
Wtf! Ben ran this mystery model through 10 DeepSWE tasks and it scored over 80%, versus 65% for Fable and 52% for GPT-5.6-sol!

this is insane. Probably a chinese company. Either a new GLM or Kimi model, I reckon.
el.cine@EHuanglu · 博主 · 1 天前

中国在打造令人难以置信的 AI 机器人

查看英文原文
incredible AI robots are being built in china
◔ 17.6 万 次浏览♥ 1,387⇄ 136▶ 含视频动态看原帖 ↗
Runway@runwayml · 公司官方 · 1 天前AI 视频生成公司 Runway

重磅推出 Runway Ruby。

这款新模型能将 SDR 视频转换为最高 16-bit HDR,支持 ProRes 和 EXR 序列。适用于任何已上传视频或生成时长达 30 秒的内容。

查看英文原文
Introducing Runway Ruby.

A new model that converts SDR video up to 16-bit HDR in ProRes and EXR sequences. Compatible with any existing uploaded video or generated output up to 30s.
◔ 19.5 万 次浏览(2 条合计)♥ 672⇄ 97▶ 含视频新品看原帖 ↗
Alexandr Wang@alexandr_wang · 创始人 · 1 天前Scale AI 创始人,Meta 超级智能实验室负责人
连环推 ×3

1/ muse spark 1.2 是一个非常强的多模态模型——它能处理视觉编程、机器人规划,以及视听理解,这些能力都通过智能工具整合在一起。

查看英文原文
1/ muse spark 1.2 is a very strong multimodal model—it can do visual coding, robotics planning, and audio-visual understanding that all come together through agentic tools.
2/ muse spark 1.2 performs quite strongly across a wide variety of multimodal capabilities and evals.
3/ you can get a deeper look at all this in the research blog published today.
research.meta.ai/blog/multim…
◔ 16.6 万 次浏览(2 条合计)♥ 840⇄ 90▶ 含视频新品看原帖 ↗
@levelsio@levelsio · 博主 · 1 天前独立开发者标杆,AI 产品连续创业者

绝了!一个叫 InkPoster 的电子墨水艺术相框,别再被一幅作品困住了,想换就换。想象给家里设置不同心情场景,比如平静、放松、雄心勃勃、激进,你的艺术作品就跟着变。inkposter.eu(无关联,就是个棒想法)

引用 Alex Sexton @SlexAxtonoverpriced but the newest generation is much better than the previous generation- have had both. If you wanna talk her out of it maybe offer to separate concerns. Get a $700 tv and an inkposter for photos/art. inkposter.com/products/inkpo…查看被引原帖 ↗
查看英文原文
This is genius

E-ink art frame called InkPoster

Don't get stuck with one artwork/painting forever, just change it whenever you like

Imagine setting up scenes/moods for your house, like calm, relaxed, ambitious, agressive and your art all changes etc.


inkposter.eu


(Unaffiliated, just a great idea)
◔ 13.6 万 次浏览♥ 814⇄ 18▶ 含视频观点看原帖 ↗
elvis@omarsar0 · 博主 · 1 天前

这只是开始而已。没人付钱让我说这些,但我觉得 Harvey 让我们窥见了未来成功的 AI 原生公司会是什么样子、如何运作。

成功的公司需要思考如何构建并拥有自己完整的情报栈。他们掌控模型、agents,以及夹在中间的一切。

引用 Harvey @harvey推出 Tenet,首个法律专用后训练模型。基于 Kimi K3,与 FireworksAI 合作。在 LAB 基准提高 82% 通过率,合同提高 22%。成本为主流模型四分之一。配备 M&A Diligence、Review Tables、Firm Knowledge 三个专家子模型。查看被引原帖 ↗
查看英文原文
This is just the beginning of what's coming. Not paid to say this, but I think Harvey provides a glimpse into the future of what successful AI-native companies will look like and how they operate.

Successful companies will need to think of how to build and own their entire intelligence stack. They own the models, agents, and everything in between.

有四件事我实在是骂累了,一个是骂华为,一个是骂宇树科技,一个是骂李鬼专业(统计、应用数学、生物信息、数据等等),一个是骂土地财政,劝你们卖房。

真不要问我如何看待了,最短的宇树科技我也骂了整整两年了,有不懂的可以看我之前的推文和视频。

像为什么灵巧手永远失败, 为什么工业自动化永远比人型机器人更重要,为什么robotics折腾了50年没有一点进步,这种话题我讲了无数次了,不需要额外科普了。

Google DeepMind@GoogleDeepMind · 公司官方 · 1 天前谷歌旗下 AI 研究机构,Gemini 背后团队

游戏 15 年来一直是我们 AI 研究的重要试验田。🎮

从掌握 Atari 到在《星际争霸 II》中达到大师等级,游戏驱动了我们许多重大的 AI 突破。

我们用 SIMA 的工作教会 AI 代理理解 3D 世界,但要学会应对真实的人类互动,需要一个真实的、持久的虚拟世界。

通过与 @FenrisCreations 的研究合作,我们在探索如何解决 AI 的开放性挑战:

🔵 持续学习——学到新技能同时不遗忘旧知识
🔵 深度记忆系统——存储和检索能力远超当今 context window
🔵 长期规划——跨越周、月甚至年的尺度
🔵 多智能体动态——涵盖合作、谈判、经济和涌现行为

我们的长期目标是与游戏开发者合作,用 AI 创造全新的游戏体验——让游戏更易达、更个性化——同时把我们学到的东西应用到现实世界的问题和科学发现。

了解更多 → goo.gle/4xcwzOH

查看英文原文
Games have been an important testbed for our AI research for over 15 years. 🎮

From mastering Atari to reaching Grandmaster in StarCraft II, they have driven some of our biggest AI breakthroughs.

Our work with SIMA taught agents how to understand 3D worlds, but learning to navigate real human dynamics takes a living, persistent universe.

Through our research partnership with
@FenrisCreations
, we’re exploring how to tackle open challenges in AI:

🔵 Continual learning to acquire new skills without forgetting past knowledge.
🔵 Deep memory systems that store and retrieve information far beyond today’s context windows.
🔵 Long-horizon planning over weeks, months, or years.
🔵 Multi-agent dynamics spanning cooperation, negotiation, economics, and emergent behaviors.

Our long-term goal is to use AI to discover entirely new gameplay experiences in partnership with game developers – making games more accessible and personalized – while applying what we've learned to problems in the real world and scientific discovery.

Find out more →
goo.gle/4xcwzOH
Greg Brockman@gdb · 创始人 · 1 天前Greg Brockman,OpenAI 联合创始人兼总裁

huge milestone in our partnership with NVIDIA

引用 Uday Ruddarraju @udayruddarraju基础设施里程碑:首批NVIDIA Vera Rubin机架已到位运行训练堆栈,助力OpenAI下一代前沿AI预训练计算。查看被引原帖 ↗
Gary Marcus@GaryMarcus · 博主 · 1 天前

我已经尽力警告你了。

引用 Gary Marcus @GaryMarcus⚠️ Microsoft and OpenAI have only one play to make GenAI profitable, and this is it: all out 24x7 surveillance. They are already starting to try to soften you up for it. Virtually every warning that I have given for the last several years has come true. Please don’t ignore this one. 🙏查看被引原帖 ↗
查看英文原文
I tried so hard to warn you.
Min Choi@minchoi · 博主 · 1 天前AI 产品演示博主,专门展示新工具玩法
连环推 ×11

Grok Bot 真的在为自己每月 300 美元的订阅费买单。人们用它来赢回客户、追回退款、降低账单和经营业务。10 个疯狂的例子:

查看英文原文
Grok Bot is literally paying for its own $300/month subscription.

People are using it to win back customers, recover refunds, cut bills + run businesses.

10 wild examples:
1. It paid its own salary.

Won several churned customers back.

Already covered its cost.
2. Plumber says Grok Bot is generating real customers...

and could save him ~$2,000/month.

Marketing agency out. Office manager stays.
3. Grok Bot emailed 5 merchants about unrefunded returns.

Recovered more than the monthly fee.

x.com/darian314/status/20893…
4. 5 businesses.

Grok Bot got hired for the night shift.

Still zero employees.

Costs less than an hour of anyone's time.
5. In under 18 hours...

Grok Bot found $175/month in bills to eliminate.

More than half the cost of SuperGrok Heavy.
6. Support inbox + Stripe API.

Routine refunds handled.

Done.
7. SpaceXAI's own GTM team is using Grok Bot for sales prospecting.
8. Grok Bot found 3 forgotten gift cards.

$150 back in her pocket.
9. A construction company is already using Grok Bot...

and says it's saving money + a LOT of time.
10. Grok Bot cleans your emails, files + paid subscriptions.

But asks for approval before cancelling anything.
NVIDIA@nvidia · 公司官方 · 1 天前

太棒了,看到第一批 NVIDIA Vera Rubin 机架正在为 @OpenAI 的训练堆栈提供算力。

我们正在携手打造加速计算基础设施,以推进前沿 AI 的发展。

引用 Uday Ruddarraju @udayruddarraju我们基础设施的里程碑:首批NVIDIA Vera Rubin 机架已到达,正在运行训练栈。这是重要步骤,因为我们扩展了为OpenAI下一代前沿AI预训练提供支持的计算能力。查看被引原帖 ↗
查看英文原文
Awesome to see the first NVIDIA Vera Rubin racks powering
@OpenAI
's training stack.

Together, we’re building the accelerated computing infrastructure to advance frontier AI.
François Chollet@fchollet · 创始人 · 1 天前

这个推论错了,这是 AI 生成的反例

查看英文原文
The conjecture is wrong, here's an AI-generated counter example
NVIDIA@nvidia · 公司官方 · 1 天前

正确的模型取决于具体任务。

NVIDIA NeMo Switchyard 帮助开发者根据自身的质量、延迟和成本标准,将每个智能体工作流步骤路由到选定的模型池中。

Kari Briski 做客 @MTSlive,解释为什么智能体工作流需要模型路由。

查看英文原文
The right model depends on the task.

NVIDIA NeMo Switchyard helps developers route each agent workflow step across a chosen model pool based on their own quality, latency and cost criteria.

Kari Briski joins
@MTSlive
to explain why agent workflows need model routing.
◔ 8.5 万 次浏览♥ 309⇄ 45▶ 含视频新品看原帖 ↗
ollama@ollama · 公司官方 · 1 天前本地跑大模型的热门工具

Kimi K3 现已向超过一半的订阅用户推出可用额度,我们今天继续扩大访问权限。美国和欧洲托管,零数据保留。

引用 ollama @ollamaKimi K3在Ollama云订阅中推出。Ollama改进定价透明度以展示最佳性能/价格比。可通过Claude Code或OpenCode命令启动:ollama launch claude --model kimi-k3:cloud。查看被引原帖 ↗
查看英文原文
Kimi K3 is now rolled out to over half of the subscription base for included usage, and we're continuing to expand access today.

US and Europe-hosted and zero data retention.
OpenAI Developers@OpenAIDevs · 公司官方 · 1 天前OpenAI 开发者平台官方

Codex 和 ChatGPT Work 中的共享线程功能让你能通过只读链接展示构建过程,为他人提供 pull request、深度分析或项目交接的背景和推理。

查看英文原文
Shared threads in Codex and ChatGPT Work let you show the process behind your build with a read-only link, giving others the context and reasoning behind a pull request, deep dive, or project handoff.
◔ 7.9 万 次浏览♥ 805⇄ 56▶ 含视频新品看原帖 ↗
Lisan al Gaib@scaling01 · 博主 · 1 天前高频 AI 模型测评与爆料博主

所以看起来大家相信一个14万神经元、5450万突触的大脑是“有意识的”

但一个一千万万个(也就是10万亿)参数的LLM却不是?

引用 Chris Lakin @chrislakinsome dude uploaded this fly's consciousness into Minecraft lol so cyberpunk查看被引原帖 ↗
查看英文原文
so apparently people believe that a 140k neuron and 54.5 million synapse brain is "conscious"

but not the 10 million million (also known as 10 trillion) parameter LLM
Qwen@Alibaba_Qwen · 公司官方 · 1 天前阿里通义千问大模型团队

新食谱刚发布!⚡ SGLang cookbook 里新增 NVFP4 + DFlash2 for Qwen3.8-27B。感谢支持!
@sgl_project

引用 SGLang @sgl_projectJust pushed DFlash2 ( @inco_ai ) recipes to the Qwen3.8 27B cookbook⚡️ docs.sglang.io/cookbook/auto… The community has been seeing great results with NVFP4 + DFlash2, and these recipes should be some very good starting points to play with. More Qwen3.8 27B updates on the way 🫡查看被引原帖 ↗
查看英文原文
Fresh recipes just dropped! ⚡ NVFP4 + DFlash2 for Qwen3.8-27B now in the SGLang cookbook. Thanks for the support!
@sgl_project
Yuchen Jin@Yuchenj_UW · 博主 · 1 天前

随性编码就是现在的编码方式。手写代码才是真正的随性编码。

查看英文原文
Vibe coding is just coding now.

Hand-writing code is the real vibe coding.
Ethan Mollick@emollick · 创始人 · 1 天前沃顿商学院教授,AI 应用研究权威

是的,AI 实验室的 CEO 们确实讨论过打造能取代所有劳动的 AI,还提过这可能带来生存风险。就我所知,早在他们公司值钱之前,他们自己就深信这一点了。他们不再谈论风险这件事,才是公关的部分。

引用 Matthew Yglesias @mattyglesiasI keep hearing people say this, but I think it's to the credit of Altman & Amodei that they told us what leading researchers in this field actually think. We should worry more about Mark Zuckerberg and his politically savvy AI reassurances!查看被引原帖 ↗
查看英文原文
Yes, the CEOs of the AI Labs talked about building AI that would replace all labor & could pose an existential risk because, as far as I could tell, they believed it to be true long before their companies were valuable. The fact that they stopped talking about risk is the PR part
Guillermo Rauch@rauchg · 创始人 · 1 天前Guillermo Rauch,Vercel 创始人兼 CEO

fx.sh
0.0.5 更精简了
能塞进两张软盘里 💾💾😁(𝚡𝚣 压缩格式)
明天发货,带上了大家最想要的功能!

查看英文原文
fx.sh
0.0.5 gets even smaller
fits in 2 floppy disks 💾💾😁 (𝚡𝚣-compressed)
shipping tomorrow with our most-asked feature!
Matt Shumer@mattshumer_ · 博主 · 1 天前HyperWrite CEO,AI 实战技巧分享

Ian Goodfellow 当年说得真对

查看英文原文
Ian Goodfellow really had it right back in the day
Bindu Reddy@bindureddy · 创始人 · 1 天前Abacus.AI CEO,AI 行业观点博主

DeepSeek Flash 在开源加速领域领跑,性价比无敌,是首轮推理的最佳选择,成本基本为零!最近几周使用量暴增 100 多倍,企业用户也都在用

查看英文原文
DeepSeek Flash is leading the open source acceleration

The price to performance is phenomenal

It's the go to model for first turn inference and costs almost nothing!

Usage has sky-rocketed more than 100x in the last few weeks for everyone including Enterprises

你要感谢国产机器人,

你就要感谢1块钱一片的宏晶华大新华龙的单片机,感谢几分钱一片的铁片电阻电容,感谢江浙沪和深圳一大堆封装机、包装机、拆包机、焊接机,感谢你家门口小饭馆从山东河南小厂采购的3000元(利润率不到50%)商用洗碗机、揉面机、炸薯条机、制冰机,

你甚至应该感谢你家电动牙刷、 扫地机器人、小米全套智能家居、智能门铃监控摄像头,感谢小米、OPPO、VIVO给你优化国产安卓的程序员,

你甚至要感谢那些分拣苹果和鸡蛋的农业分拣机,分拣快递的大传送带和摄像头录入系统,感谢你汽车里的ABS防抱死设备和刹车系统,感谢你家冰箱、洗衣机、洗碗机、智能马桶、油烟机、空调、电扇里面可靠运转十几年的小小单片机和写逻辑的嵌入式程序员,

你唯独不应该天天哭天喊地痛哭流涕地感谢宇树科技上春晚翻跟头的表演机器人。

这就好比你吃饱了饭不感谢你爸妈、不感谢农民、不感谢一线市场厮杀打拼的水稻育种公司、不感谢交通物流和绿色通道,天天感谢那个发明你一口都没吃过的野败杂交籼米的发明人袁隆平爷爷一样。

Tinyfool@tinyfool · 中文博主 · 1 天前

伪代码/UML等中间路线,会不会因为agent而复兴?
ourcoders.com/tech/show/tech…

Chubby♨️@kimmonismus · 博主 · 1 天前Chubby,高频 AI 新闻聚合博主

我天:NVIDIA的编码agent AVO在ARC-AGI-3的25个公开游戏中得分100%,解决了全部183个关卡。

该agent不接收任何规则或明确的目标。它必须通过尝试、观察结果和纠正错误来学习。

AVO的成功在于它能够记住学到的东西,并在长时间内基于这些知识继续构建,而不是在模型上下文重置时重新开始。

由Claude Opus 5驱动,同样的系统之前曾自主工作了七天来优化GPU代码。

真的很酷,看到NVIDIA能做到这一点

引用 NVIDIA AI @NVIDIAAIOur general-purpose coding agent just scored 100% on the ARC-AGI-3 interactive reasoning benchmark. NVIDIA AVO completed all 183 levels across all 25 public environments, figuring out what to do with no instructions, explicit rules, or stated goals.查看被引原帖 ↗
查看英文原文
Holy: NVIDIA’s coding agent AVO scored 100% on ARC-AGI-3’s 25 public games, solving all 183 levels.

The agent receives no rules or stated goals. It must learn by trying things, observing the results and correcting its mistakes.

AVO succeeds by remembering what it learned and building on it over long periods instead of starting over when the model’s context resets.

Powered by Claude Opus 5, the same system previously worked autonomously for seven days optimizing GPU code.

Really really cool to see what NVIDIA achieved here
◔ 7.2 万 次浏览(3 条合计)♥ 770⇄ 52▶ 含视频演示看原帖 ↗
Lisan al Gaib@scaling01 · 博主 · 1 天前高频 AI 模型测评与爆料博主

看起来就是某个GitHub仓库的复制品哈哈

引用 Lisan al Gaib @scaling01Ox Alpha one-shotted the GPU accelerated fluid simulation in a single 1000 line html file it looks absolutely stunning and is 1000x better than what Qwen3.8 27B or Opus 4.5 did when I tested them earlier this week查看被引原帖 ↗
查看英文原文
seems to be literally just a copy of an existing github repo lmao
◔ 4.8 万 次浏览♥ 434⇄ 1▶ 含视频其他看原帖 ↗
Guillermo Rauch@rauchg · 创始人 · 1 天前Guillermo Rauch,Vercel 创始人兼 CEO

Bun 对简洁、快速与开源的不懈追求,和 Vercel 的理念以及对世界的思考方式完美契合。

恭喜
@jarredsumner
成功发布!

引用 Vercel Developers @vercel_devBun 1.4 现已在Vercel Functions 中提供。支持WebSockets和Bun.serve()、支持Next.js、Elysia、Hono等、在Fluid上运行,具有Active CPU计费。在vercel.json中设置"bunVersion": "1.4.x"来升级。查看被引原帖 ↗
查看英文原文
Bun’s pursuit of simple, fast & open is the perfect match for Vercel and how we think about the world.

Congrats
@jarredsumner
on shipping!
Kevin Weil 🇺🇸@kevinweil · 创始人 · 1 天前

一条秩数为 30 的椭圆曲线!!

现在每天都兴奋地醒来,因为我们(人类 + AI)每天都在发现这么多关于世界的新东西。

引用 Robin Houston @robinhoustonomg y² + xy = x³ − 201769035260418549083594900060734240952308696994802735114305555 x + 1151107939141058565733479426024323225135665982951300586808823640527729578307228357301072889377 an elliptic curve of rank 30, dropped anonymously a few hours ago what a time etc. elliptic-rank.icarm.cloud/cu…查看被引原帖 ↗
查看英文原文
A rank 30 elliptic curve!!

It's gotten to the point where I'm excited to wake up each morning because we (humans + AI) are learning so many new things about the world each day.
Yuchen Jin@Yuchenj_UW · 博主 · 1 天前

“OpenAI的收入增长看起来更像Databricks,而非Anthropic。”

Databricks握有数据库、神经网络和芯片资源。就等着看我们指数级的增长吧。😉

查看英文原文
“OpenAI has revenue growth that looks more like Databricks than Anthropic.”

Databricks has databases, neural networks, and chips. Just wait for our exponential growth. 😉
Gary Marcus@GaryMarcus · 博主 · 1 天前

OpenAI 变成了监控公司,而且都不隐瞒了。就像下面这个链接讨论的那样,这事儿已经进行了好多年。

garymarcus.substack.com/p/op…

查看英文原文
OpenAI is becoming a surveillance company
It’s not even really trying to hide it, anymore.

And, as discussed at the link below, this has been years in the making.


garymarcus.substack.com/p/op…
Chubby♨️@kimmonismus · 博主 · 1 天前Chubby,高频 AI 新闻聚合博主

据报道,Anthropic正准备最早于8月底公开提交IPO申请。

该公司预计融资额至少与SpaceX创纪录的750亿美元首秀持平,估值约2万亿美元,有望成为历史上最大规模的IPO。

引用 Bloomberg @businessAnthropic expects to match or beat the size of SpaceX’s record-setting initial public offering, according to people familiar with the matter, as preparations for the artificial intelligence firm’s debut pick up speed bloomberg.com/news/articles/…查看被引原帖 ↗
查看英文原文
Anthropic is reportedly preparing to publicly file for its IPO as soon as the end of August.

The company expects to raise at least as much as SpaceX’s record $75 billion debut, potentially making it the largest IPO in history (~$ 2tn valuation)
Lisan al Gaib@scaling01 · 博主 · 1 天前高频 AI 模型测评与爆料博主

想知道ARC-AGI-4什么时候发布,还有他们会想出什么幺蛾子让模型得分不到1%。

查看英文原文
I wonder when ARC-AGI-4 will come out and what bullshit they come up with to make models score below 1%
Gary Marcus@GaryMarcus · 博主 · 1 天前

我一直没完没了地在强调这个(只有认真听的人才明白)

引用 Haider. @haider1Turing Award winner Rich Sutton: "LLMs are an amazing scientific breakthrough. but it's frustrating that instead of celebrating this progress in a subset of AI, it has to pretend to be all of AI" Language may be only 20-25% of intelligence. There's much more to intelligence — we're not done查看被引原帖 ↗
查看英文原文
what i have been saying endlessly (to those few who have listened carefully)
el.cine@EHuanglu · 博主 · 1 天前

AI 现在可以精准控制演员的情绪

查看英文原文
AI now can precisely control actors emotions
◔ 3.9 万 次浏览♥ 677⇄ 60▶ 含视频其他看原帖 ↗
Gary Marcus@GaryMarcus · 博主 · 1 天前

Karp 受尽批评,但其实是对的。我也有同样的感受。🙄

引用 NIK @ns123abc🚨OPENAI AND ANTHROPIC JUST FOLDED ON THEIR DATA RETENTION POLICY AFTER PALANTIR CEO EXPOSED THEIR REAL BUSINESS MODEL Alex Karp: >"something has gone completely wrong" >they want "access to my data" so they can "build my alpha" >"i'm gonna get no value, and they're gonna get my IP" >"this is effing insane" >"trying to drug addict us to a future they believe they control" >"who owns the data? where is it cached? are the prompts secured?" >"we need to rebuild trust" Yesterday: OpenAI started testing "private safety processing" to avoid retaining customer data Today: Anthropic reversing course, will allow enterprise customers to keep data on THEIR OWN cloud infrastructure instead of Anthropic's Alex Karp was right.查看被引原帖 ↗
查看英文原文
Karp got so much shit and was actually right.

I know the feeling. 🙄
◔ 3.9 万 次浏览♥ 258⇄ 33▶ 含视频观点看原帖 ↗
Gary Marcus@GaryMarcus · 博主 · 1 天前

我,2014 年:AI 缺乏常识。亚马逊,今天:哎呀

引用 Peyman Milanfar @docmilanfarAGI is just around the corner查看被引原帖 ↗
查看英文原文
Me, 2014: AI lacks common sense
Amazon, today: Oops
◔ 3.7 万 次浏览♥ 183⇄ 17▶ 含视频观点看原帖 ↗
Ethan Mollick@emollick · 创始人 · 1 天前沃顿商学院教授,AI 应用研究权威

AI 产品版本太多了真的让人懵(ChatGPT Work/Codex/Chat、Claude Cowork/Code/Chat)加上各种形式(电脑应用、网页应用),我都分不清每个产品具体有哪些功能:plugins、skills、memories、permissions、files 这些功能分别在哪里啊?🤷‍♂️

查看英文原文
A confusing thing about the proliferation of AI modes (ChatGPT Work/Codex/Chat, Claude Cowork/Code/Chat) & modalities (apps on computers, web apps) is that I am losing track of which stuff each has: when do plugins, skills, memories, permissions, files, etc. sit in each case? 🤷‍♂️
Gorden Sun@Gorden_Sun · 中文博主 · 23 小时前中文圈高频 AI 资讯与开源项目博主

Gemini 3.5 Pro被放弃有点可惜了,最近一直用Gemini 3.7 Flash辅助阅读和写作,非常说人话的模型,比Opus 5舒服多了。要是Gemini 3.5 Pro发布了,世界知识和创意能力肯定是更强的,也不是所有模型都必须得拿来写代码,Opus 4.6现在还有很多人用,基本都是用来写作。

Chubby♨️@kimmonismus · 博主 · 1 天前Chubby,高频 AI 新闻聚合博主

今早打开 Gmail,一半活儿已经干完了。
Lindy 扩展程序在夜间帮你处理收件箱,并在顶部固定一条简短摘要。一段话就让我知道来了什么、还有哪些需要我处理。
它直接嵌在 Gmail 里。不用迁移,不用设置。

@getlindy

引用 Flo Crivello @AltimorAnnouncing the Lindy Chrome extension. Bring Lindy straight into your inbox to highlight your most important emails, draft replies backed by all your memories, and teach it how to label your email. Live now: lindy.ai查看被引原帖 ↗
查看英文原文
I opened Gmail this morning and half the work was already done.
The Lindy extension runs your inbox overnight and pins a short brief on top. One paragraph and I knew what came in and what still needed me.
It sits inside Gmail. Nothing to move, nothing to set up.

@getlindy
◔ 3.2 万 次浏览♥ 124⇄ 7▶ 含视频演示看原帖 ↗
yetone@yetone · 中文博主 · 1 天前开源 AI 编程插件 avante.nvim 作者,开发者圈博主

这是典型的对「Electron 内存消耗比 Tauri 高」的误解,其实真相恰恰是相反的,Tauri 应用的内存占用往往高于 Electron 应用,看看前 Tauri 大使是怎么说的吧

引用 Vesper.Li @coinbacai今天还特意拿Electron和Tauri作了对比,同样的程序Electron打包内存消耗是Tauri的10倍。作为开发Electron是方便,但对于用户真的很致命。最近看到X圈很多优秀的软件,但发现一个包接近1G,望而却步,果断安装体验后无奈选择卸载查看被引原帖 ↗
swyx@swyx · 博主 · 1 天前知名 AI 播客 Latent Space 主理人

按浏览量,@aidotengineer 现在是我们简短历史中排名最高的发言人,也堪称全球最出色的 skills 专家之一——他的 /grill-me 已经传播到各个层级,包括 @satyanadella(虽然有些变化)。

/wayfinder 是 @mattpocockuk 的 /grill-me for /grill-me,用在你在战争迷雾中摸索、还不知道自己有什么知识盲点时,需要协调研究和其他 grill 会议的场景。非常高兴有 @ricmac 帮助我们用他这周推出的课程来启动新的 Skills 报道!

👇独家采访,快速阅读

引用 Latent.Space @latentspacepodWe chat to @mattpocockuk about his /wayfinder skill, which he designed for the "fog of war" — when you need to figure out a project but the end state isn’t entirely clear. This is the first in a series of skills we'll be exploring in the coming weeks. latent.space/p/wayfinder-ski…查看被引原帖 ↗
查看英文原文
by total views, matt’s now the top
@aidotengineer
speaker in our brief history, and arguably one of the best skills experts in the world — his /grill-me has reached all echelons up to
@satyanadella
(with some variations…)

/wayfinder is
@mattpocockuk
’s /grill-me for /grill-me, when you are just navigating the fog of war and don’t yet know what you dont yet know, and need to orchestrate research and other grill sessions to get there. super glad to have
@ricmac
help launch our new Skills coverage with his course launch this week!

👇exclusive interview, quick read
Ethan Mollick@emollick · 创始人 · 1 天前沃顿商学院教授,AI 应用研究权威
连环推 ×2

即使LLM写得再好,风格单调也是硬伤。每天看同样的文案出现在你的说明里、社交媒体、广告、软件和PowerPoint里,最后都会反胃。只靠提示词技巧远远不够,需要真正的多样性(而且这方面研究还不足)

查看英文原文
Even when LLMs write well, the lack of variety in style is crippling. Reading the same prose in your instructions & social media & advertisements & software & PowerPoint eventually makes one queasy

Prompting only gets you so far. Real variation is needed (and under-researched)
Variation is also going to be important if you want AIs to surface diverse new ideas, do science and tons of other things.

Temperature and top-p changes are also not enough, some similarity still creeps through (and similarity is good for some uses, but certainly not all)
小互@xiaohu · 中文博主 · 1 天前小互,中文圈高频 AI 资讯站 Xiaohu.AI 主理人

OpenAI 将 Codex 的底层核心框 Harness开源

你可以轻松将Agent接入到自己业务系统中

对开发者:少造一层基础设施。
上下文记忆、工具调用、沙箱、进度和审批机制,都可以复用 Codex 已有的开源组件。

对产品:AI 可以进入现有软件
物流后台、工单系统、报税软件和监控看板都可以成为 Codex 的工作界面。

对用户:不用把工作重新搬进聊天框。 原系统提供当前对象和业务状态,结果仍回到原来的正式记录。

说清楚:OpenAI 没有把整个 Codex、模型和云服务全部开源。明确开放的是 Codex CLI、SDK、app-server 等 Harness 与接入组件;模型访问、托管服务、IDE Extension 和 Codex cloud 仍是独立部分。

准确地说,开发者拿到的是 Codex 的 Agent 执行底座和接入层,不是整套 Codex 的自托管版本。


best.xiaohu.ai/article/codex…

◔ 3.4 万 次浏览(2 条合计)♥ 190⇄ 26新品看原帖 ↗
Chubby♨️@kimmonismus · 博主 · 1 天前Chubby,高频 AI 新闻聚合博主

我们以前会为那个给 reset 的。Tibo 怎么了...

引用 Tibo @thsottiauxWe've investigated a few messages about codex usage limits being different. That's not something we change without engaging the community and being transparent. What we did see is that when talking to affected users many were using sub2api. Converting a subscription into api traffic to then re-serve or share across many users is not something we support and this type of usage gets flagged by our fraud-prevention systems. You are completely fine if you use your subscription through Sign in With ChatGPT, either through the official clients or through one of the many OSS clients (Pi, OpenCode, ...) that support signing in with your account and using your included usage.查看被引原帖 ↗
查看英文原文
We would have gotten a reset for that in the past. Whatever happened to Tibo...
Lisan al Gaib@scaling01 · 博主 · 1 天前高频 AI 模型测评与爆料博主

我觉得 ZAI 刚搞明白了怎么用 GLM-5.2/5.3 把 RL 环境刷得很漂亮

我真的很想看到一张图,展示各 AI 实验室单独的 RL 环境数量随时间变化的曲线

这条曲线大概能解释我们所看到的 90% 的模型进展和实验室排名变化

查看英文原文
I think ZAI just figured out how to do good RL environment spam with GLM-5.2/5.3

I would really like to see a chart of the number of RL envs over time for all of the AI labs individually

this curve would probably explain 90% of what we have been seeing in terms of model progress and ranking of the labs
Cristóbal Valenzuela@c_valenzuelab · 创始人 · 1 天前Runway 联合创始人兼 CEO

Runway 即将推出的东西可能是我们有史以来最重大的发布之一。这个团队在做太多令人兴奋的事情了,我真的很激动,迫不及待想让大家看到接下来几周几个月会有什么。

查看英文原文
What’s coming from Runway might be among the most consequential releases we’ve ever made. The team is working on so many exciting things, pretty excited and can’t wait for you all to see what’s coming over the next few weeks and months.
AshutoshShrivastava@ai_for_success · 博主 · 1 天前高频 AI 新闻与产品动态博主

Gemini 3.7 Flash 真的绝了。

查看英文原文
Gemini 3.7 Flash is absolutely K!lling it.
九原客@9hills · 中文博主 · 1 天前

Qwen3.8-27B 本地部署指南

1. Q4 基本没有质量损失,甚至可以用Q3
2. 用Unsloth的GGUF。
3. 思考开low就行了。
4. 开 dflash2

引用 Alexey Fateev @superaleshaMain result. At xhigh every quant landed between 88.0 and 90.0% pass@1 on the full suite: AWQ INT4: 90.0% NVFP4: 89.3% GGUF Q4_K_M: 89.3% FP8: 88.7% NInfer: 88.0% Yes, the 4 bit quants scored above the FP8 baseline. McNemar says its a statistical tie, first and last place differ by three tasks out of 150. I started this run to show you how quantization eats quality. There is nothing to show. The gap between quants is smaller than the gap between reasoning presets.查看被引原帖 ↗
九原客@9hills · 中文博主 · 1 天前

我不太理解A和O对自家风控那么自信么,一点也不会有假阳性?

A社其实表现还好点,它就是封+退款,没啥小动作。

O 就比较鸡贼,根据ip质量风控降智我是见过的,现在还有降额度?误封的都不知道。

你不想挣钱就封号+退款呗,搞这么拧巴。

引用 Tibo @thsottiauxWe've investigated a few messages about codex usage limits being different. That's not something we change without engaging the community and being transparent. What we did see is that when talking to affected users many were using sub2api. Converting a subscription into api traffic to then re-serve or share across many users is not something we support and this type of usage gets flagged by our fraud-prevention systems. You are completely fine if you use your subscription through Sign in With ChatGPT, either through the official clients or through one of the many OSS clients (Pi, OpenCode, ...) that support signing in with your account and using your included usage.查看被引原帖 ↗
Chubby♨️@kimmonismus · 博主 · 1 天前Chubby,高频 AI 新闻聚合博主

苹果似乎没打算竞争开发全球领先的独立前沿 AI 模型,但这家公司继续展现其硬件供应商的实力。就像 NVIDIA 是 AI 时代「淘金热」中的「卖铲人」,苹果生产一些最好的用于 AI 工作的硬件。我很少这么期待一场即将到来的 keynote。对新任 CEO John Ternus 我期望很高。作为苹果硬件工程的负责人,他监督了负责公司硬件产品的各个团队。不过苹果把 Johny Srouji 列为在驱动硅芯片战略和向 Apple silicon 转变中起核心作用的人物。那次转变堪称完美之举,可以说是苹果近年最好、最重要的决定之一。如果你在苹果生态中工作,想让电脑 24/7 运行(比如用于 OpenClaw 或 Hermes),Mac mini 也被证明是款绝佳产品。如果苹果发布搭载 M6 芯片的 Mac mini,我几乎肯定会买,已经等不及了。看看苹果还会推出什么吧。

查看英文原文
While Apple does not appear to be competing to build the world’s leading standalone frontier AI model, the company continues to demonstrate its strength as a hardware provider. Just as NVIDIA is the "shovel seller in the gold rush" of the AI era, Apple produces some of the best hardware for working with AI.

I can rarely remember looking forward to an upcoming keynote this much. And I have high expectations for incoming CEO John Ternus. As Apple’s head of Hardware Engineering, he oversaw the teams responsible for the company’s hardware products. However, Apple credits Johny Srouji with playing the central role in driving its silicon strategy and transition to Apple silicon. That transition was a masterstroke and arguably one of the best and most important decisions Apple has made in recent years.

The Mac mini has also turned out to be a great product if you work within the Apple ecosystem and want to keep a computer running 24/7, for example, for OpenClaw or Hermes. If Apple releases a Mac mini with an M6 chip, I will almost certainly get one, and I’m already looking forward to it. Let’s see what else Apple releases.
ollama@ollama · 公司官方 · 1 天前本地跑大模型的热门工具

祝贺整个 @GoogleDeepMind Gemma 团队!Gemma 是最受欢迎的开源模型之一。作为紧密的合作伙伴,这段旅程真的很棒!期待看到接下来会发生什么!🎉

引用 Omar Sanseviero @osansevieroGemma正式突破10亿下载量。社区在各种激励人心的用途中使用这些模型,从水下到太空,令人兴奋。查看被引原帖 ↗
查看英文原文
Congratulations to the whole
@GoogleDeepMind
Gemma team!

Gemma is one of the most popular open models.

It's been an amazing journey being a close partner! Can't wait to see what's to come! 🎉
小互@xiaohu · 中文博主 · 1 天前小互,中文圈高频 AI 资讯站 Xiaohu.AI 主理人

Anthropic 官方发布了一份《Claude Code 创业公司指南》

系统复盘了 15 家高增长初创公司如何通过 Agentic Coding 实现“10 人团队打出百人产出”的交付效率。

这份指南详细拆解了这些公司如何把 AI 编程工具接进原型、研发、值班、验证、重建和产品化流程,并归纳为五条规则。

ClickHouse:功能交付量提升 30%;排查不稳定测试和补全测试覆盖率的两个定制 Agent,成为代码库贡献榜第 2 和第 3 名。

Omni:工程团队生产力提升 2 至 3 倍。

Clay:实现 100% 的缺陷分流自动化,Agent 还会在初步排查后提出代码修改建议。

Artemis Security:每周交付超过 6,000 个 PR。


best.xiaohu.ai/article/claud…

swyx@swyx · 博主 · 1 天前知名 AI 播客 Latent Space 主理人

嗯,@openclaw 对苹果挺有帮助的 这个团队今年光 Mac mini 销售就保守估计驱动了 $50-$150m(大概是全球 Mac mini 年销售额的 50% 增幅,全都得益于 openclaw 2026)

引用 Peter Steinberger 🦞 @steipete512GB内存工作室。Apple对我们不错。查看被引原帖 ↗
查看英文原文
well,
@openclaw
was good to apple

this team conservatively drove $50-$150m in mac mini sales alone this year haha (roughly +50% of normal annual mac mini sales worldwide, just due to openclaw 2026)
swyx@swyx · 博主 · 1 天前知名 AI 播客 Latent Space 主理人

很开心和 @vibhuuuus 一起做了最新一期播客,和 @eisokant 谈 NVIDIA 为什么花60亿美元买下这个model factory——一个在生产击败Thinky的模型的工厂(说真的一点没夸大,看看数据就知道)

去听听吧 / 订阅 / 随便啥 latent.space/p/poolside

只在 @latentspacepod

引用 elie @eliebakouchwow this is kind of a shock. from what i understand nvidia bought the "model factory" part of poolside and a lot of employees (researchers?) got offers from nvidia. founders staying at poolside is unusual, wondering if they will just become a neocloud/compute provider since i don't see any mention of PIC (poolside infrastructure company) here?查看被引原帖 ↗
查看英文原文
proud that
@vibhuuuus
and i did the most recent pod with
@eisokant
on why NVIDIA just paid him $6B to buy the incredible model factory that is pumping out Thinky-beating models (actually not exaggeration, look at the numbers)

tune in / subscribe / whatever
latent.space/p/poolside


only on
@latentspacepod
Ethan Mollick@emollick · 创始人 · 1 天前沃顿商学院教授,AI 应用研究权威

政治哲学家会这么说。

但我是研究组织设计的经济社会学家,对我来说很明显,多智能体对齐就是个组织设计问题,涉及很强的经济社会学成分

(我真觉得很多学科都能在这个领域贡献东西)

引用 Seth Lazar @sethlazar多智能体对齐基本上就是政治哲学吧?查看被引原帖 ↗
查看英文原文
A political philosopher would say that.

As an economic sociologist who studies organizational design, clearly multi-agent alignment is an organizational design problem with strong components of economic sociology

(I really think many academic disciplines have a lot to add here)
🚨 AI News | TestingCatalog@testingcatalog · 博主 · 1 天前专挖 AI 产品未发布新功能的爆料号

SPACEXAI 🔥:Grok Bot 即将登陆 Grok 移动应用!

> 现在有个针对 SuperGrok 用户的新促销活动,允许他们升级到 Grok Heavy 以解锁 Grok Bot 的使用。
> 除此之外,还有一个隐藏的导航栏项目,用户可以借此访问他们的自定义 AI 智能体。

看起来挺重量级的! 🤖

查看英文原文
SPACEXAI 🔥: Grok Bot is coming to Grok mobile apps soon!

> There is a new promotion for SuperGrok users offering them to upgrade to Grok Heavy in order to access Grok Bot.
> Besides this, there is a hidden nav bar item from where users will be able to access their custom AI agents.

Looks Heavy! 🤖
elvis@omarsar0 · 博主 · 1 天前

harness 持续学习的优质论文。

(收藏它)

如果你已经在让自己的智能体改写它们的 prompts、skills 或 memory 文件,这篇论文绝对值得一读。

(收藏它)

持续学习一直在追踪权重的变化。现代智能体则是在 harness 里积累经验,包括 prompts、memories、tools、skills 和路由规则。

这意味着什么?如果你更新了任何 harness 组件,之前稳定的行为可能就会坏掉,哪怕模型本身完全没动过。论文把这叫做 harness-level 遗忘,还给出了衡量办法。

Guarded harness evolution 把更新提议和提交分开了。一个 Continual Optimizer 根据执行后的反馈起草候选 harness,然后一个 Continual Evaluator 在检查了当前收益、历史保留和有效性后才正式提交。

相对收益在文本推理、多模态感知和开放世界交互上都超过 10%。

论文:
arxiv.org/abs/2608.19013

在我们的学院关注更多热门 AI 论文:
academy.dair.ai/

查看英文原文
Banger paper on harness continual learning.

(bookmark it)

If you already are allowing your agents to rewrite their own prompts, skills, or memory files, this one is worth your time.

(bookmark it)

Continual learning has always tracked what changes in the weights. Modern agents accumulate experience in the harness instead, across prompts, memories, tools, skills, and routing rules.

What this means is that if you update any harness component, previously reliable behavior can break with the model completely untouched. The paper names that harness-level forgetting and provides a way to measure it.

Guarded harness evolution separates proposing an update from committing it. A Continual Optimizer drafts a candidate harness from post-execution feedback, and a Continual Evaluator commits only after checking current improvement, historical retention, and validity.

Relative gains exceed 10% across textual reasoning, multimodal perception, and open-world interaction.

Paper:
arxiv.org/abs/2608.19013


Track more trending AI papers in our academy:
academy.dair.ai/
el.cine@EHuanglu · 博主 · 1 天前
连环推 ×2

AI 能把梦想变成电影

我做了个噩梦,用 AI 花了 2 天把它变成了恐怖电影元素,参加 OiiOii 的 Into the Paradox Forge Contest

只有 AI 才能做到啊

查看英文原文
AI can turn your dreams into films

I had a nightmare, then spent 2 days turning it into an AI horror film element for OiiOii’s Into the Paradox Forge Contest

only AI can make this possible
I’m also joining
@OiiOii_AI
’s Into the Paradox Forge Contest as a contest judge

something like this used to need an entire crew and budget, now you can wake up and make it real

excited to see how far everyone pushes their ideas into the paradox
#OiiOiiAIParadox
Bilawal Sidhu@bilawalsidhu · 博主 · 1 天前

仍然是我最喜欢的"world model 到底是啥"梗

查看英文原文
still my favorite "what even is a world model" meme

h/t
@KnightNemo_
Matt Shumer@mattshumer_ · 博主 · 1 天前HyperWrite CEO,AI 实战技巧分享
连环推 ×2

Something Big 通讯第一期上线了!

来看看吧,告诉我你的想法:
somethingbig.ai/welcome?utm_…

查看英文原文
First Something Big newsletter is live!

Check it out, and let me know what you think:
somethingbig.ai/welcome?utm_…
(this is the soft-launch, going to iterate a bunch based on feedback for the big launch soon!)
🚨 AI News | TestingCatalog@testingcatalog · 博主 · 1 天前专挖 AI 产品未发布新功能的爆料号

OpenAI 将计算机历史记录和录制与回放功能扩展到欧洲经济区、英国和瑞士。

> 计算机历史记录现已面向 Mac 上的 ChatGPT Pro、Business 和 Enterprise 用户推出。

> 录制与回放功能也在 macOS 版 ChatGPT 应用中提供

完整推出 🤖

引用 OpenAI Developers @OpenAIDevs欧洲更新:Computer History已在EEA、英国、瑞士向ChatGPT Pro/Business/Enterprise用户开放。ChatGPT可跨应用网站记忆活动,使对话更个性化,需要更少解释。查看被引原帖 ↗
查看英文原文
OpenAI expanded Computer History and Record & Replay features to EEA, UK, and Switzerland.

> Computer History is now available for ChatGPT Pro, Business, and Enterprise users on Mac.

> Record & Replay is also available in the ChatGPT app for macOS

Full rollout 🤖
elvis@omarsar0 · 博主 · 1 天前

来自 Google 的印象深刻的研究,关于如何为 agent 构建更好的环境。

Agent 的训练环境是手工构建的,容易过时。Agent 在进步,环境却不进步,甚至一开始就看不到 agent 的弱点。

EnvHarness 用可编程插件层包装了一个静态环境,在不触及底层逻辑的前提下改变其行为。每个改造后的环境都保留了原始验证器,这就是为什么在其上训练是安全的。

EnvRigger 把策略当黑箱对待,读取其执行轨迹,合成针对诊断出的缺陷的 harness 组件,然后用全新的执行来验证它们。

在五个基准、四个领域中,在保留测试集上最高提升 9.0 分,执行步数减少 9.8%。

论文:
arxiv.org/abs/2608.19880

在我们的学院追踪更多趋势 AI 论文:
academy.dair.ai/

查看英文原文
Impressive research from Google on building better environments for agents.

Training environments for agents are hand-built and go stale. The agent improves, the environment does not, and it's not able to see the agent's weaknesses in the first place.

EnvHarness wraps a static environment in a programmable plug-in layer that reshapes its behavior without touching the underlying logic. Every reshaped environment keeps its original verifier; this is what makes the reshaping safe to train on.

EnvRigger treats the policy as a black box, reads its execution trajectories, synthesizes harness components aimed at the diagnosed flaws, then validates them with fresh rollouts.

Across five benchmarks in four domains, up to 9.0 points better on held-out instances with 9.8% fewer execution steps.

Paper:
arxiv.org/abs/2608.19880


Track more trending AI papers in our academy:
academy.dair.ai/
ollama@ollama · 公司官方 · 1 天前本地跑大模型的热门工具

欢迎
@ATT
加入开放模型的行列!

引用 Amir Efrati @amiropen source AI is...happening? :) Anthropic & OpenAI better hope AT&T is the exception, not the rule查看被引原帖 ↗
查看英文原文
Welcome
@ATT
to open models!
🚨 AI News | TestingCatalog@testingcatalog · 博主 · 1 天前专挖 AI 产品未发布新功能的爆料号

macOS 上的 ChatGPT 现在可以通过新插件与 Apple Messages 集成。

该插件允许用户搜索对话记录,也能直接编写并发送回复。

引用 ChatGPT @ChatGPTInstall Messages from Plugins > Public, then try prompts like: - Check my calendar and reply to [name] with a few times I'm free for dinner next week - Suggest follow ups from yesterday in messages - Find birthdays in @ messages and add them to my calendar - Find potential spam messages I can delete查看被引原帖 ↗
查看英文原文
ChatGPT on macOS can now integrate with Apple Messages via a new plugin.

The plugin allows users to search through their conversations as well as write and send replies.
◔ 2.5 万 次浏览(2 条合计)♥ 194⇄ 13▶ 含视频新品看原帖 ↗
OpenAI Developers@OpenAIDevs · 公司官方 · 1 天前OpenAI 开发者平台官方

PR 审查不需要截图考古 🔍

learn.chatgpt.com/docs/use-c…

查看英文原文
PR context shouldn’t require screenshot archaeology 🔍


learn.chatgpt.com/docs/use-c…
Chubby♨️@kimmonismus · 博主 · 1 天前Chubby,高频 AI 新闻聚合博主

OpenAI 推出了 AI Futures,一个新的研究博客,关注当具有变革性的 AI 重塑权力时,自由社会如何能保护个人权利和代理能力。Dean Ball 的 Strategic Futures 团队认为,权力集中可能是最大的长期 AI 风险。自主系统可能让国家在无需合作士兵的情况下投射军事力量。自动化官僚体系和数据中心产生的财富可能会让政府减少对人力、税收和民众同意的依赖。不管怎么看,这再次清楚地表明一点:我们处在人类从未经历过的社会变革的门槛上。

查看英文原文
OpenAI has launched AI Futures, a new research blog focused on how free societies can preserve individual rights and agency as transformative AI reshapes power.

Dean Ball’s Strategic Futures team argues that concentration of power may be the largest long-run AI risk.

Autonomous systems could let states project force without cooperative soldiers. Automated bureaucracies and data-center-generated wealth could make governments less dependent on human labor, taxation and consent.

Regardless of what one thinks of it, it clearly shows one thing once again: we are on the threshold of a societal transformation the likes of which we have never seen before.
◔ 2.3 万 次浏览(2 条合计)♥ 309⇄ 29动态看原帖 ↗
🚨 AI News | TestingCatalog@testingcatalog · 博主 · 1 天前专挖 AI 产品未发布新功能的爆料号

Grok Build 正在向更多 SuperGrok 用户在网页和移动端推出。测试时间到了 👀

引用 Andrew Curran @AndrewCurran_Grok近48小时表现异常。我认为Grok 4.6正在平台和应用中悄悄推出,Grok Build测试版现已对我开放(之前灰显)。查看被引原帖 ↗
查看英文原文
Grok Build is rolling out to more SuperGrok users on web and mobile.

Testing time 👀
Chubby♨️@kimmonismus · 博主 · 1 天前Chubby,高频 AI 新闻聚合博主

开源与前沿AI之争,平时也就是纸上谈兵。

但@DataCamp
必须在生产环境里做出这个选择——一边是为超1900万学习者服务的平台,一边要打造AI原生导师。

我和联合创始人Peter坐下来,跟DataCamp联创兼CEO
@CornelissenJo
以及首席AI官Yusuf Saber聊了聊:

– 哪些模型在实际生产中真的能打
– 开源模型是不是真的便宜10倍
– 那些隐性的基础设施和工程成本
– 缓存、延迟、质量、还有幻觉问题
– 隐私、数据驻留、以及供应商锁定
– 开源模型到现在还搞不定的那些事

这场对话非常务实,聊的就是AI走出演示、面对真实用户、真实成本和真实翻车场景时到底会发生什么。

完整采访在下方,YouTube版放评论区了。

查看英文原文
Open-source vs. frontier AI is usually a theoretical debate.


@DataCamp
has to make that choice in production - while building an AI-native tutor for a platform serving more than 19 million learners.

My co-founder Peter and I sat down with DataCamp Co-Founder and CEO
@CornelissenJo
and Chief AI Officer Yusuf Saber to discuss:

– which models actually work in production
– whether open models are really 10x cheaper
– hidden infrastructure and engineering costs
– caching, latency, quality, and hallucinations
– privacy, data residency, and vendor lock-in
– what open models still cannot do

A genuinely practical conversation about what happens when AI leaves the demo and meets real users, real costs, and real failure modes.

Full interview down below + YouTube in comment section
Aidan Gomez@aidangomez · 创始人 · 1 天前Cohere CEO,Transformer 论文作者之一

Palantir?主权 AI 公司?

引用 Palantir @PalantirTechSovereignty is the future. Welcome @OpenAI查看被引原帖 ↗
查看英文原文
Palantir? The sovereign AI company?
Lisan al Gaib@scaling01 · 博主 · 1 天前高频 AI 模型测评与爆料博主
连环推 ×2

这看起来像是机器人基础模型的 GPT-3 时刻。

我推荐阅读这篇博文:

generalistai.com/blog/gen-1.…

引用 Generalist @GeneralistAIIntroducing GEN-1.5, a one-shot learner. It can learn new tasks in a few seconds. Show it what to do, and it generalizes. This capability emerged from pretraining on physical data at scale, as a step towards our mission of building general intelligence for the physical world.查看被引原帖 ↗
查看英文原文
this looks like the GPT-3 moment for robot foundation models

I recommend reading the blog:

generalistai.com/blog/gen-1.…
"robot foundation models are few-shot learners"
◔ 1.7 万 次浏览♥ 239⇄ 5▶ 含视频观点看原帖 ↗
yetone@yetone · 中文博主 · 1 天前开源 AI 编程插件 avante.nvim 作者,开发者圈博主

在 Alma 中使用 @ 进行所有类型资源的双向引用是一位 Alma 的资深用户提的需求,这位用户简直是我的产品经理,提了特别多很牛逼的需求,太强大了!

引用 Tom @Tomyu_2034#Alma 新出了一种 Skill:就是自动提醒用户当前的工作区存在重复的任务或者工作流,能够自动帮你提取总结,挺实用的。还有就是现在 @ 功能也越来越丰富,其中也能像 Codex 一样 @ 其他的聊天记录。做的真的很用心,就连 Windows 系列的主题居然还内置了那个年代很火的扫雷游戏。 @yetone 👍👍👍查看被引原帖 ↗
Orange AI@oran_ge · 中文博主 · 1 天前Orange AI,中文圈 AI 产品观察博主

没想到,给 Markdown 编辑器加 Mermaid 功能,竟然成为模型能力的试金石
本以为很成熟的东西,实现起来这么难

Gary Marcus@GaryMarcus · 博主 · 1 天前

这一直都是OpenAI的游戏计划。(我在2024年就曾警告过这一点,说监控可能是他们唯一可行的商业模式。)

引用 Hedgie @HedgieMarkets🦔ChatGPT can now read your entire iMessage history and send texts on your behalf. Every conversation with your family, your doctor, your lawyer, your ex, on OpenAI’s servers so it can reply to your mom for you. Altman told us a few days ago he wanted ChatGPT to hold your entire life. He meant it. This is a company that hasn’t turned an annual profit, is heading for an IPO, and just asked for access to the most personal data on your computer. But sure, let it answer your group chat. Hedgie🤗查看被引原帖 ↗
查看英文原文
this has always been OpenAI’s game plan. (I first warned of it in 2024, saying that surveillance might be their only viable business plane.)
Gary Marcus@GaryMarcus · 博主 · 1 天前

哎呀,IPO 了。

来源:
@theinformation
:

查看英文原文
Uh oh, IPOs.

Via
@theinformation
:

本站由 Jedee杰哥 打造 · 公众号「Jedee杰哥」每早送 AI 日报

姊妹站:𝕏 简中账号数据榜单 · X 关注 @jedeeai · RSS 订阅 · AI 日报 · 历史归档