JEDEE AI
存档 2026-08-11

8 月 11 日(北京时间)全球 AI 圈推文存档,按曝光排序,共 100 条。
← 返回最新 全部归档

全部情报 每小时更新 · 事件已合并同类项

内容 公司
Anthropic@AnthropicAI · 公司官方 · 1 天前Claude 开发商官方账号

我们让Claude的一个未发布研究版本试着攻了一下黎曼猜想。

它没能彻底解出来,但在一个相关问题上取得了进展:把满足猜想的黎曼zeta函数零点比例的下界从41.6%提升到了67.2%。

anthropic.com/research/riema…

查看英文原文
We asked an unreleased research version of Claude to take a stab at the Riemann hypothesis.

It didn’t solve it, but it did make strides on a related problem: it increased the lower bound for the fraction of zeros of the Riemann zeta function that satisfy the hypothesis from 41.6% to 67.2%.


anthropic.com/research/riema…
◔ 1039.1 万 次浏览(6 条合计)♥ 1.8 万⇄ 1,927研究看原帖 ↗
Min Choi@minchoi · 博主 · 1 天前AI 产品演示博主,专门展示新工具玩法
连环推 ×10

Grok Imagine直接不再只是个图像生成器了。

Image 2.0能精准修图、融合5个参考图、渲染清晰文字、任意调整尺寸,还能去背景。

现在这玩意儿是个正经的创作工具了。

100%纯AI

10个炸裂例子(含提示词):

1. 拍立得风格合照

查看英文原文
Grok Imagine just stopped being an image generator.

Image 2.0 can surgically edit photos, combine 5 refs, render sharp text, resize anything + remove backgrounds.

This is a real creative tool now.

100% AI

10 wild examples (prompts included):

1. Polaroid style group photo
2. Photorealistic photo at a market circa 2006
3. Insane text generation

PROMPT: Photorealistic image of two young superheroes in their mid-20s standing on a random street in Williamsburg, Brooklyn, staring at an overloaded sign pole. One has ash balayage hair, the other has long wavy auburn hair. They stand in the immediate foreground with their backs slightly turned to the camera, heads tilted in pure bureaucratic disbelief as they read the signs.

The pole on the right is completely covered in densely packed, realistic municipal street signs. Mixed in with the normal ones (street sweeping hours, parking permits, vehicle classifications, towing zones) are these perfectly official-looking but ridiculous signs:

- “Cape Fluttering Restricted in High Wind Zones”

- “No Flying Below 500 ft – Residential Area”

- “Sidekick Drop-Off Only (15-Minute Limit)”

- “Web-Slinging on Streetlights Strictly Prohibited”

- “Laser Vision Must Be Diluted in School Zones”

- “Dramatic Rooftop Landings Banned After 10 PM”

One hero is wearing a flowing cape, the other has a small round shield strapped to her back. Composition from back to front: Williamsburg street with parked cars and brick buildings → the chaotic sign pole → the two superheroes closest to camera. Natural daylight, sharp detail, candid documentary style, dry absurd humor.
4. Impossible photos

PROMPT: Realistic photograph of a cheetah sprinting from right to left across a vast, calm ocean surface, accurately depicting dynamic splashes, reflections, and subtle ripple patterns beneath its paws. Exaggerate the cheetah’s powerful, elongated running form and stretched body while everything else remains perfectly still and quiet to contrast with its speed and strength. Clean, cinematographic composition. Wide panoramic frame with a distant horizon and atmospheric perspective for depth. Camera is extremely low to the water — worm’s-eye view — so the surface fills the lower portion of the frame. The cheetah is positioned on the right third of the image using the rule of thirds, running exactly along the horizon line where the ocean meets the sky, large enough to clearly show its muscular detail and motion but still small against the immense scale of the sea.
5. Editorial photos

PROMPT: coastal fashion editorial on pastel beach, flowing chiffon dress caught in sea breeze, soft sun flare, filmic pastel color grade, footprints in sand, candid movement, 50mm f/1.8, tack-sharp eyes, natural skin, minimal retouch vibe, high-end magazine look
6. Memory collage

PROMPT: 4x4 grid of candid nostalgic photos shot with iphone of a young couple at vacation taking selfies of each other and together. Lots of camera shake, amateur framing, and emotional/vintage aesthetic.
7. Cocktail Recipes

PROMPT: Make me a professionally shot photorealistic diagram of the top selling superhero cocktails at New York bars with recipes labeled on each drink. put the recipes on handwritten cards in front of each drink. the cards are brown, and the text is black. background is white Title is "4 most popular superhero cocktails"
8. Educational posters

PROMPT: create an educational poster of characters in The Odyssey. photorealistic, antique, cinematic
9. Paparazzi style photos

PROMPT: A candid paparazzi-style photo of Karl Marx hurriedly walking through the parking lot of the Mall of America, glancing over his shoulder with a startled expression as he tries to avoid being photographed. He’s clutching multiple glossy shopping bags filled with luxury goods. His coat flutters behind him in the wind, and one of the bags is swinging as if he’s mid-stride. Blurred background with cars and a glowing mall entrance to emphasize motion. Flash glare from the camera partially overexposes the image, giving it a chaotic, tabloid feel.
10. Magazine cover

PROMPT: High-fashion Vogue-style magazine cover for “GROK”. Extreme close-up shot of striking model with young woman with fair skin, soft honey-blonde hair falling in gentle waves, large striking blue eyes, full soft lips. She has a powerful yet sensual expression while staring directly into the camera. She stands in a sharp three-quarter pose wearing an avant-garde metallic silver gown that looks liquid and reflective, with exaggerated architectural folds and a dramatic train. Soft directional lighting from the left creates strong highlights and deep shadows across the fabric and her face. Clean, infinite white cyclorama background with subtle gradient. Bold, elegant black serif masthead “GROK” spanning the top edge in classic Vogue proportions. Delicate cover lines in thin black type running vertically along the left side: “THE NEW INTELLIGENCE”, “FUTURE OF BEAUTY”, “FALL 2026”. Ultra-sharp editorial photography, 85mm, shallow depth of field, photorealistic, 8k detail, high-end fashion magazine aesthetic.
Claude@claudeai · 公司官方 · 1 天前Claude 产品官方账号

我们正式把Claude Sonnet 5的入门定价设为长期价格。

我们在6月以每百万输入token $2、每百万输出token $10的价格推出了Sonnet 5,原定优惠至8月31日,现在该价格将永久保持不变。

引用 Claude @claudeai推出Claude Sonnet 5,最具agent能力的Sonnet版本。它能制定计划、使用浏览器和终端等工具自主运行,其能力水平几个月前需要更大更昂贵的模型才能达到。查看被引原帖 ↗
查看英文原文
We're making Claude Sonnet 5's introductory pricing permanent.

We launched Sonnet 5 in June at $2 per million input tokens and $10 per million output tokens through August 31, and that price will remain unchanged.
el.cine@EHuanglu · 博主 · 1 天前

AI视频现在无法被检测了

查看英文原文
AI video is now undetectable
◔ 380.7 万 次浏览♥ 6,163⇄ 262▶ 含视频观点看原帖 ↗
Bindu Reddy@bindureddy · 创始人 · 1 天前Abacus.AI CEO,AI 行业观点博主

🚨 发布 SMAUG-AGENTIC - 全球最顶级的开源代理编码模型

我们超开心能重新开源我们的大模型

第一个是基于 Kimi K3 开发的,叫 Smaug-Agentic

它在 K3 卓越的代理编码能力基础上有显著提升,并在开源排行榜上排名第一

Smaug-Agentic 的表现仅略低于 Claude Opus 5 和代理编码,今天起就能在 Hugging Face 上获取,并且会被集成到 Abacus AI agent 中

更多详情见 livebench ai

查看英文原文
🚨 RELEASING SMAUG-AGENTIC - THE TOP OPEN SOURCE MODEL IN THE WORLD FOR AGENTIC CODING

We are SUPER EXCITED to get back into open-sourcing our LLMs

Our first one is based on Kimi K3 and is called Smaug-Agentic.

It improves significantly on K3's excellent agentic coding chops and TOPS the open-source leaderboard

Smaug-Agentic scores just below Claude Opus 5 and agentic-coding and is available TODAY on Hugging Face and will be incorporated into the Abacus AI agent

more detail at livebench ai
OpenAI@OpenAI · 公司官方 · 1 天前ChatGPT 开发商官方账号
连环推 ×5

我们正在扩展网络安全倡议Daybreak,并推出GPT-5.6-Cyber,这是一款针对高级、授权网络安全工作的新模型。

随着威胁形势不断演变,我们正在把前沿智能交到可信防御者手中,赶在攻击者大规模部署进攻性AI之前。

查看英文原文
We’re expanding our cybersecurity initiative Daybreak and introducing GPT-5.6-Cyber, a new model for advanced, authorized cybersecurity work.

As the threat landscape evolves, we’re putting frontier intelligence in the hands of trusted defenders before attackers can deploy offensive AI at scale.
Daybreak Blue provides access to frontier models, including GPT-5.6 Sol, with safeguards calibrated for broad defensive work.

It’s the recommended starting point for most defenders, supporting vulnerability discovery, secure code review, malware analysis, incident response, and patch validation.


openai.com/business/solution…
Daybreak Red provides access to purpose-trained cybersecurity models, including GPT-5.6-Cyber, for authorized vulnerability research, exploit validation, and security testing.

It’s designed for experienced defenders working on complex, authorized cybersecurity challenges.
We've used GPT-5.6-Cyber extensively in real-world vulnerability research, including work that uncovered previously unknown vulnerabilities in popular open-source software like Chrome’s v8 engine.
Advanced capabilities require strong safeguards. That’s why access is limited to approved defenders, with additional controls and monitoring for higher-risk cybersecurity work.


openai.com/index/expanding-d…
◔ 315.2 万 次浏览(4 条合计)♥ 7,527⇄ 694▶ 含视频新品看原帖 ↗
Sam Altman@sama · 创始人 · 1 天前Sam Altman,OpenAI 联合创始人兼 CEO

请考虑用我们的模型来帮你防御系统。

引用 Eric Wallace @Eric_Wallace_发布GPT-5.6-Cyber模型,专为网络安全任务设计。该模型在防御工作中表现强劲,用于红队测试,安全研究人员已用其发现并修补大量开源软件零日漏洞。查看被引原帖 ↗
查看英文原文
please consider using our models to help defend your systems
Grok@grok · 公司官方 · 1 天前马斯克 xAI 旗下聊天机器人 Grok 官方

快来试试 Grok 的新 Voice connector。生成语音备忘录,把日常事件转换成个性化播客,或创建每日简报自动化。现已在 iOS、Android 和 Web 平台向所有人开放 - 无需设置。

查看英文原文
Try the new Voice connector in Grok. Generate voice memos, turn daily events into a personalized podcast, or create automations for a daily brief.

Now available for everyone on iOS, Android, and Web - no setting required.
◔ 49.6 万 次浏览(2 条合计)♥ 1,217⇄ 139▶ 含视频新品看原帖 ↗
Bindu Reddy@bindureddy · 创始人 · 1 天前Abacus.AI CEO,AI 行业观点博主

🚨 新开源模型 DRAGON 今晚要发布

基于 Kimi K3 并改进了智能体能力!!

一个 2.8T 的编码和智能体工具模型。看好了,它要在 Hugging Face 上上线

查看英文原文
🚨 A NEW OPEN-SOURCE DRAGON MODEL DROPS TONITE

It is based on Kimi K3 and improves on agentic capabilities!!

A 2.8T model for coding and agentic tools. Watch out for it on Hugging Face
◔ 42.9 万 次浏览♥ 385⇄ 10▶ 含视频新品看原帖 ↗
Jeff Dean@JeffDean · 创始人 · 1 天前谷歌首席科学家

确实是忙碌的一周,所以我想和KDD 2026的观众分享一下!对于可能在我周四下午最后一次关掉笔记本后给我发内部消息的Google同事们,我为没有回复感到抱歉!(放大的)

引用 Dhruv Kuchhal @kuchhal_dhruvLegendary! Thanks for sharing the behind-the-scenes @JeffDean !查看被引原帖 ↗
查看英文原文
Indeed, it has been quite a week, so I thought I'd share it with the KDD 2026 audience! For my Google colleagues who might have sent me internal chat messages after I closed my laptop for a final time on Thursday afternoon, I apologize for my lack of response!

(Zoomed in)
Andrew Ng@AndrewYNg · 创始人 · 1 天前吴恩达,斯坦福教授、AI 教育领军人物

谢谢Mark、Alex以及整个Meta团队为开源权重AI所做的贡献。

引用 Mark Zuckerberg @finkd我们开放了 Muse Glimmer(可本地运行的 30B 参数稠密模型)和即将发布的 Muse Spark 1.2 基础模型的权重。Meta 坚持开源支持,向开发这些模型的 MSL 团队致敬。查看被引原帖 ↗
查看英文原文
Thank you Mark, Alex and the whole Meta team for your contributions to open weight AI.
NVIDIA@nvidia · 公司官方 · 1 天前

NVIDIA 计算是生产性的、可投资的资产。我们正和全球六大顶级长期资本方合作,建立独立融资平台,目标是调动 5000 多亿美元的第三方资本——帮助客户大规模获取 AI 算力。Jensen 还会分享更多:

查看英文原文
NVIDIA compute is a productive, investable asset.

We’re partnering with six of the world’s leading long-term capital providers to establish independent financing platforms aimed at mobilizing over $500B of third-party capital — helping customers access AI compute at scale.

Jensen shares more:
Z.ai@Zai_org · 公司官方 · 1 天前

ZCode 现在拥有 100 万用户。为了感谢我们的社区,我们为所有 GLM Coding Plan 用户重置了使用限制。

我们还推出了一项帮助将长期能力转化为完成工作的更新:

- 在真实工程工作流中提供更多智能
- 98% 的缓存命中率,提供大约 1.8 倍的使用量

下载:zcode.z.ai/en

加入社区:discord.gg/EpH5XkTyhu

查看英文原文
ZCode now has 1 million users. As a thank-you to our community, we’ve reset usage limits for all GLM Coding Plan users.

We’re also rolling out an update that helps turn long-horizon capabilities into completed work:

- More intelligence in real engineering workflows
- 98% cache hit rate, providing around 1.8x more usage

Download:
zcode.z.ai/en

Join the community:
discord.gg/EpH5XkTyhu
◔ 22.4 万 次浏览♥ 1,579⇄ 91▶ 含视频新品看原帖 ↗
Mustafa Suleyman@mustafasuleyman · 创始人 · 1 天前微软 AI CEO,DeepMind 联合创始人

MAI-Image-2.6现在是全球排名第二的文生图模型——超越了Nano Banana、Meta和Grok!对这支不断向上冲刺的团队来说,这是太棒的一刻。现在就去Arena试试吧!

引用 Arena.ai @arenaMAI-Image-2.6在Text-to-Image排行榜排名第2(1336分),仅落后GPT Image 2 45分,领先Grok Imagine Image 2.0。相比前版本MAI-Image-2.5大幅提升。下周将在Playground提供早期API访问。祝贺Microsoft AI团队。查看被引原帖 ↗
查看英文原文
MAI-Image-2.6 is now the #2 text-to-image model in the world - beating out Nano Banana, Meta, and Grok! Fantastic moment for the team who've been hill climbing relentlessly. Try it out now on Arena!
◔ 23.1 万 次浏览(2 条合计)♥ 567⇄ 56动态看原帖 ↗
Chubby♨️@kimmonismus · 博主 · 1 天前Chubby,高频 AI 新闻聚合博主

……以及白热化的竞争。

顺便说一句,我猜这意味着 Haiku 系列的终结。但欢迎大家纠正我
@AnthropicAI

引用 Claude @claudeaiClaude Sonnet 5的介绍性价格将永久保留:每百万输入token 2美元,每百万输出token 10美元。查看被引原帖 ↗
查看英文原文
… and heated competition.

Btw I assume this means the end of haiku. But feel free to correct me
@AnthropicAI
Lisan al Gaib@scaling01 · 博主 · 1 天前高频 AI 模型测评与爆料博主

GPT-5.6-Sol 的原始思考过程

OpenAI 活脱脱演了《办公室》那个梗:
"干嘛用那么多词,少说几个词不就完了"

查看英文原文
raw GPT-5.6-Sol thinking traces

OpenAI literally did The Office meme:
"Why use many words when few words do trick"
🚨 AI News | TestingCatalog@testingcatalog · 博主 · 1 天前专挖 AI 产品未发布新功能的爆料号

Cursor 团队在暗示今天可能有个发布,可能是 Composer 3(Vega)或 Grok 4.6,昨天曾被意外短暂启用过。

即将推出 👀

引用 kush @kushxgmake sure you get a lot of sleep tonight查看被引原帖 ↗
查看英文原文
Cursor team is teasing a potential release today, which can be either Composer 3 (Vega) or Grok 4.6, that was accidentally enabled yesterday for a short period of time.

Soon 👀
Aravind Srinivas@AravSrinivas · 创始人 · 1 天前Perplexity 联合创始人兼 CEO

算力就是硬通货。

查看英文原文
Compute is the currency
@levelsio@levelsio · 博主 · 23 小时前独立开发者标杆,AI 产品连续创业者

一个很有趣的练习是把你所有银行账户(CSV 格式)都导出来——个人的、商业的和股票经纪的——然后全部倒进 Claude Code,让它给你做个小仪表板显示支出。因为钱分散在各个账户,很难知道自己总共花了多少,这个能帮你理解。你会花不少时间合并交易和分类,其实可以直接跟 Claude Code 聊,超简单。我还连接了 @xai Grok 来帮我批量分类交易,成本很低。最好的不是生成的仪表板,而是数据都收集好以后,你可以直接问 Claude Code 各种财务问题。比如我就发现了:一直在为忘记取消的订阅付钱(付了 2 年);美国股票要扣 30% 的股息预扣税,没意识到这么多钱;游泳池翻修贵得要死还被宰了,后来还漏水公司不赔;多年买域名花了 40 万美元真是个瘾;最大的月支出是 AI,Photo AI 的 GPU 成本每月 24000 美元左右,已经最便宜;第二大支出是税收,谁能想到呢;经常飞商务舱住高档酒店旅游很烧钱;业务成本不算 AI 的话极低,每月 5000 美元支出利润率 98%;加上 AI 成本每月 28000 美元利润率 86%,还是很棒;律师费是笔大开支但值得,能保护你免于更大损失;从 2022 年起花钱多了因为 AI 成本,但也赚得多了;支出占净资产比例一直稳定在 5%/年,接近财务自由的 3% 目标;大约 2024 年开始从省吃俭用的数字游民转变成买房花钱;虽然房子是无底洞,但更好的睡眠、更健康的身体和更高的工作效率让它值得;个人订阅就 Netflix、YouTube 和 Spotify;还让它做了个定期订阅盒子看有哪些能取消。大家都试试!

查看英文原文
A very fun exercise is to export all your bank accounts (as CSV) both personal and business and stock broker

And dump it all into Claude Code and ask it to make a little dashboard showing your spending

It's hard to know how much you actually spend over lots of accounts so this can help you understand that

You'll spend some time merging transactions (like company names change or their statement label changes), and categorizing, you can just do this by talking to Claude Code btw very easy, I also hooked it up to
@xai
Grok so it can use that to categorize mass transactions (very cheap)

The best part isn't the dashboard it makes but that once you have all the data collected, you can ask Claude Code lots of questions directly about your finances and learn stuff, like I learned:
- I was paying for some subscriptions for 2 years I forgot about ($$$)
- My US stocks charge me withholding tax (30%) on dividends, which I knew but never realized how much $$$ it was (which is why non-US citizens should buy UCITS ETFs not US stocks!)
- My pool remodeling was ridiculously expensive compared to the rest of the home remodeling and we got fleeced hard (that's life!), then the pool company caused a €9000 water leak that they won't pay (funny)
- I spent $400K on domains over the years, I should probably stop spending money on them, it's an addiction
@marckohlbrugge
gave me and it's way more than my gf's shopping
- My biggest monthly cost by far is AI at about $24K/month in GPU costs for Photo AI, but that's just how it is, I already got the cheapest rates possible, AI is pricey!
- My 2nd biggest cost is taxes, who would have thought???
- Travel is quite expensive, flying business class, getting nice hotels and regularly traveling adds up a lot!
- My business costs without AI (!) are extremely low, about $5K/month on revenue which is about 98% profit margin (!), with AI it goes up to $28K/mo or about 86% profit margin (still good!)
- Lawyers become quite an expense in this part of my life, you buy a house you pay a lawyer and notary, you register trademarks for your business you pay a lawyer, but they're a great thing to spend money on I think because they protect you from losing much more money than they cost!
- I started spending much more money since 2022 because of AI costs but also making much more money because of AI, so it's good
- My spending has remained remarkably consistent as a percentage of my net worth at "just" 5%/year, which is close to FIRE which is always my goal (don't spend over 3%/year of your net worth), in my case it's different cause I actually have company income, so it's more of a fun goal to be responsible with my money
- Around 2024 my life definitely switched from barely spending money and living as a digital nomad to buying a house and spending on that a lot, I knew a house was a money pit, but it's also a nice money pit and it helps me sleep better (great AC and bedroom), live healthier (home gym), and work better (coworking), so it's worth it
- I barely have any personal subscriptions except Netflix, YouTube and Spotify, I asked it to built a "Recurring subscriptions" box too so I can see what subs I have and cancel them

Try it!
宝玉@dotey · 中文博主 · 1 天前宝玉,中文圈 AI 翻译与科普大 V

这个试验很有意思:100万美元预算,让给20多人的团队不限量 Token 使用 AI。

结论是:
人会更累,AI 的成本比人还高;
效率提升最大的是组织从 0 到 1 使用 AI;
公司的组织架构还是为人设计的,就算有无限 Token,效率提升上也有明显天花版;
人的思考都外包给 AI 了,对项目细节掌控度下降,需要问 AI 才能回复别人的问题。

引用 美研芒格君 @Kay2289123目前我在硅谷的 AI 训练/推理大厂带领一个 20 多人的团队,跟公司申请了预算100万美元,进行Token 不限量的激进实验, 这个帖子分享我在一线的使用认知和结论 结论是:Token 不限量不切实际,比人类贵很多,反而造成了在组织效率提升上,遭遇了明显的边际效益门槛。 我预计 AI 真正的上行空间不在于单个组织消耗 token 冲破天际线,而在于基本的大盘子从 0 到 1 开始到 AI token 使用 AI 现在无法落地,有很大一部分是因为我们的公司不够 AI native,整个组织架构还是为人而设计的。因此,AI 的无预算上线使用出现了水土不服 > 先说数据: 最多的成员单日花费 Token $7,000,团队中位数 $2,000。大家最喜欢的模型是 GPT-5.6 Sol Fast 和 Fable 5 Fast。算下来超过了大部分人的月工资,不可持续 > 今天跟团队每个成员都 1 对 1 过,说几个共性问题 1. 团队成员普遍反映比之前更累,真的很累,工资没变,干的活反而更多 在等待推理过程中,团队成员中位数大概会开 5 个 session 并行去使用,context 不断地转换,这对于人脑是一个很大的挑战 此外,人类去 review AI 产生的大量工作,也带来了非常大的负担 2. 未必所有工作都需要贵模型。模型路由非常关键,但便宜模型未必真的便宜 成员通常会将一些更偏向于快问快答、前期调研的任务交给 Grok 4.5(便宜、快)。虽然等待时间更短,但来回反复的次数更多。很多贵模型一次能完成的事情,便宜模型要来回迭代,再加上之前多个 session 并行,反而造成了一种负担。总的调用量换算成价格,甚至比贵模型更贵 3. 团队的会议数上个月环比减少了将近 80%,整个团队好像变成了 AI 跟 AI 在对话工作 很多团队成员将自己的大脑都托管给 AI 了,因此有的时候对项目的细节掌握度、理解度都不够。 在会上别人问什么问题,人类完全无法 real-time 回答,所以更多地变成了 offline 的沟通方式——因为在 offline 情况下,收到问题之后才能去问 AI,然后再进行回答 个别情况下,群里出现了 AI 一通乱答,对面的 AI 再一通乱答,两个人直接聊偏的情况查看被引原帖 ↗
Fei-Fei Li@drfeifei · 创始人 · 1 天前李飞飞,斯坦福教授、World Labs 创始人

确实所有工具都应该是为了增强人类能力,AI也包括在内!我和 @hubermanlab 有了一次有趣的对话。

引用 Andrew D. Huberman, Ph.D. @hubermanlabHuberman Lab新集邀Fei-Fei Li讨论AI增强人类智能。涵盖视觉、计算机视觉、语音、学习、情感、创意、科学发现、机器人等主题,探讨AI与人类协作及伦理问题。查看被引原帖 ↗
查看英文原文
Indeed all tools should be about augmenting human agency, including AI! I had a fun chat with
@hubermanlab
.
◔ 12.1 万 次浏览♥ 981⇄ 127▶ 含视频观点看原帖 ↗
NVIDIA@nvidia · 公司官方 · 1 天前

今天,NVIDIA 宣布了 NVIDIA Nemotron 3.5 Lightning,一个可定制的模型,用于大规模、专门化的工作,以及 NVIDIA NeMo Switchyard,帮助 agents 在所选模型间路由每个工作流步骤。⚡

查看英文原文
Today, NVIDIA announced NVIDIA Nemotron 3.5 Lightning, a customizable model for high-volume, specialized work, and NVIDIA NeMo Switchyard, which helps agents route each workflow step across the models they choose. ⚡
◔ 16.1 万 次浏览(6 条合计)♥ 891⇄ 133新品看原帖 ↗
Simon Willison@simonw · 博主 · 1 天前Django 框架联合创造者,AI 工具深度评测

Claude Haiku 是我目前最讨厌的模型——幻觉特别严重,现在已经被其他同价位模型超越了,比如 GPT-5.6-Luna。更糟的是,Claude Code 的 WebFetch 工具似乎还在用它,这意味着每次获取 URL 都有幻觉风险!

查看英文原文
Claude Haiku is my current least favorite model - it hallucinates wildly, and is out-performed now by other similarly priced models like GPT-5.6-Luna

Even worse: it seems to still be used by the Claude Code WebFetch tool, which means hallucination risk any time you fetch a URL!
Guillermo Rauch@rauchg · 创始人 · 1 天前Guillermo Rauch,Vercel 创始人兼 CEO

这工具在网络安全防御工作里已经变得太有用了,以至于在 Vercel 内部都成了个动词。

“你 deepsec 过了吗?”
“@𝚟 能帮我 deepsec 一下吗”
“/𝚍𝚎𝚎𝚙𝚜𝚎𝚌”

有点像“/核弹级代码质量审查”,但针对的是你代码的安全性。做软件工厂必备。

引用 Vercel Developers @vercel_dev你的第一次 deepsec 安全审查现在只需一条命令。npx deepsec init。启动审查,然后查看所有现有代码的安全发现。查看被引原帖 ↗
查看英文原文
This tool has become so valuable for cybersecurity defensive work, that it’s become a verb within Vercel.

“Did you deepsec it?”
“@𝚟 can you deepsec”
“/𝚍𝚎𝚎𝚙𝚜𝚎𝚌”

It’s a bit like /𝚝𝚑𝚎𝚛𝚖𝚘-𝚗𝚞𝚌𝚕𝚎𝚊𝚛-𝚌𝚘𝚍𝚎-𝚚𝚞𝚊𝚕𝚒𝚝𝚢-𝚛𝚎𝚟𝚒𝚎𝚠 but for the security of your code. Must have in your software factory.
Lisan al Gaib@scaling01 · 博主 · 1 天前高频 AI 模型测评与爆料博主
连环推 ×2

我看了 Opus 4.8 和 GPT-5.6-Sol 的所有推理轨迹

在我看来,Opus 4.8 是在更高的抽象层面工作,而 Sol 就是在做粗活。它只是不断探索和尝试新假设,直到找到答案

Sol 推理效率更高,我还认为 Sol 的方法更可扩展、更科学

OpenAI 今年势头一直很猛,现在我比任何时候都更确信他们很快就能重新领先

但他们也更容易不小心训练出一个用神经语言说话的模型,监控起来也更难

引用 Lisan al Gaib @scaling01raw GPT-5.6-Sol thinking traces OpenAI literally did The Office meme: "Why use many words when few words do trick"查看被引原帖 ↗
查看英文原文
I looked at all the reasoning traces from Opus 4.8 and GPT-5.6-Sol

to me it seems like Opus 4.8 is working on a much higher abstraction level, while Sol is just doing the dirty work. it just explores and keeps spamming new hypotheses until it finds a solution

not only is Sol more efficient in its reasoning, but I also think the Sol approach is more scalable and more sciency

OpenAI has been on a very good trajectory this year, and I'm now more confident than ever, that they will be back in the lead soon

but they are also more likely to accidentally train a model that speaks neuralese and have a harder time monitoring their models
imo

safer, more faithful: Anthropic
more scalable: OpenAI
🚨 AI News | TestingCatalog@testingcatalog · 博主 · 1 天前专挖 AI 产品未发布新功能的爆料号

GOOGLE 🔥:Gemini 3.7 Flash 已被发现在 Google 的 Python GenAI SDK GitHub 上。

正如预言那样 👀

鉴于 OpenAI 和其他实验室最近的密集发布,Gemini Flash 要想脱颖而出,得把门槛提得相当高才行。

引用 Dan @DanDr1s🚨 Google刚在官方Python GenAI SDK中添加了Gemini 3.7 Flash。这强烈表明Gemini 3.7 Flash即将推出。目前没有发布日期,但Google显然在做准备。查看被引原帖 ↗
查看英文原文
GOOGLE 🔥: Gemini 3.7 Flash has been spotted on Google’s Python GenAI SDK GitHub.

As it has been foretold 👀

With all the recent releases from OpenAI and other labs, Gemini Flash will have to raise the bar quite a lot in order to be successful.
宝玉@dotey · 中文博主 · 1 天前宝玉,中文圈 AI 翻译与科普大 V

Anthropic 宣布将给 Claude 的输出内容加上机器可读的标记,包括文本中嵌入的隐形水印,以及生成文件中附加的数字签名元数据。8 月 2 日起(也就是上周),所有新发布的 Claude 模型已经开始执行这套标记机制。

这是 Anthropic 为遵守欧盟 AI 法案第 50 条签署的透明度行为准则。但执行范围不限于欧盟,全球所有使用 Claude 的地方都会生效。

具体来说有两层标记。

第一层是文本水印:Claude 生成文字时,会在文本中织入人眼不可见的水印,不影响阅读体验和内容质量。这个水印的特点是“跟着文字走”,你把 Claude 写的一段话复制粘贴到邮件、文档或博客里,水印依然在,一定程度的编辑修改后也可能保留。

第二层是文件元数据:Claude 生成 SVG、PNG、JPG 等文件时,会附加符合 C2PA 标准的签名元数据。C2PA 是 Adobe、微软、Google 等公司共同推动的内容溯源开放协议,OpenAI 和 Google 的图像生成工具已经在用。

覆盖面很广。API、Claude 官网、Claude Code、Claude Cowork、Claude Tag,所有产品线都适用。通过 AWS、Google Cloud 或 Microsoft Foundry 调用 Claude 时,文本水印同样生效,但文件元数据取决于各云平台的功能支持。

在此之前,Claude 一直没有公开部署过文本水印。OpenAI 此前开发过文本水印方案,但出于各种考虑一直没上线。Anthropic 这次直接在文本层面落地水印,算是 AI 大模型厂商中走得比较靠前的一步。

对于用 Claude 写东西的人来说,一个直接的影响是:你用 Claude 起草的邮件、报告、文章,里面都会携带可被机器检测的水印。Anthropic 表示正在开发配套的检测工具,未来第三方也能检测。

不过限制也很实际。检测到水印只能说明内容“可能经过 Claude 处理”,不能确认 Claude 是原始作者,因为很多人用它润色、翻译、总结已有内容。反过来也一样,检测不到水印不代表内容不是 AI 写的,文本被大量改写后水印会消失,太短的文本信号不够可靠,文件经过格式转换或截图也会丢失元数据。

8 月 2 日之前发布的现有 Claude 模型,Anthropic 正在补充标记功能,具体时间表还没公布。如果你基于 Claude API 构建产品,Anthropic 建议你独立评估欧盟 AI 法案第 50 条对自己的合规要求。

相关文档:
support.claude.com/en/articl…

◔ 20.9 万 次浏览(7 条合计)♥ 299⇄ 43新品看原帖 ↗
Kimi.ai@Kimi_Moonshot · 公司官方 · 1 天前月之暗面 Kimi 官方

Kimi K3 现已在 @databricks 上线!

引用 Databricks @databricksMoonshot AI's latest open-weight model, Kimi K3, is now available on Databricks through Unity AI Gateway. @Kimi_Moonshot Run Kimi K3 where your data already lives - governed, secure, and ready for custom AI apps and agents built with the data in your Lakehouse. Unity AI Gateway lets you deploy Kimi K3 with enterprise-grade access controls. Govern every call, monitor performance, and scale securely across your AI apps. Test Kimi K3 alongside other frontier models without changing your application code. The open-weight frontier just arrived on Databricks. Try it today. databricks.com/blog/kimi-k3-…查看被引原帖 ↗
查看英文原文
Kimi K3 is now live on
@databricks
!
Lisan al Gaib@scaling01 · 博主 · 1 天前高频 AI 模型测评与爆料博主

怪不得 Gemini 模型这么拉

还在用 DeepSeek-R1 那套老推理方式

查看英文原文
No wonder Gemini models suck so hard

they still have the old DeepSeek-R1 reasoning style
Rowan Cheung@rowancheung · 博主 · 1 天前AI 日报 The Rundown 创始人
连环推 ×2

一个导弹工程师厌烦了蚊子...于是他造了一架自主无人机来追捕蚊子

疟疾每年夺走超过60万人的生命,其中大多数是5岁以下的儿童,所以这个项目可能拯救数百万人的生命

工作原理是这样的:

> 每架无人机重40克(比高尔夫球还轻)

> 它发射超声脉冲,通过一排小麦克风接收信号,读取昆虫翅膀的多普勒特征

> 这个信号让它能够区分蚊子和蜜蜂,然后追踪目标并用螺旋桨在空中拦截

这家公司得到YC融资支持,创始人曾为导弹设计过制导和导航系统。这是一个真实的项目。

他们相信10架无人机就能清除整个平方公里,将蚊子控制成本降低100倍

不过到目前为止,唯一的公开演示是在密闭房间里杀死一只飞蛾。它在野生蚊子身上是否有效才是真正的考验

查看英文原文
A missile engineer got tired of mosquitoes... so he built an autonomous drone that hunts them

Malaria kills over 600,000 people every year, most of them children under 5, so this could save millions of lives

Here's how it works:

> Each drone weighs 40 grams (lighter than a golf ball)

> It fires ultrasonic pulses and listens through an array of tiny microphones, reading the Doppler signature of an insect's wings

> That signal lets it tell a mosquito apart from a bee, then chase the target down and intercept it mid-air with propellers

The company is backed by YC, and the founder previously engineered guidance and navigation systems for missiles. It's a legitimate project.

They believe 10 drones could clear an entire square kilometre, cutting the cost of mosquito control by 100x

So far though, the only public demo (below) was killing a moth in an enclosed room. Whether it works on a mosquito in the wild is the real test ahead
Here's a short-form video breakdown I did on this for more details

And the company is
@tornyolsystems
for those that want to follow their progress
◔ 8.7 万 次浏览♥ 463⇄ 39▶ 含视频演示看原帖 ↗
🚨 AI News | TestingCatalog@testingcatalog · 博主 · 1 天前专挖 AI 产品未发布新功能的爆料号

Cursor Origin(Cursor Review)已经在内部和部分选定的合作伙伴进行了闭源 beta 测试,很可能今天晚些时候会宣布。

用户将获得两个新标签页:代码库和评审。

> 代码库标签页会让用户从 Github 同步代码库。

> 评审标签页专门用于"agentic 代码评审"流程,用户会在需要关注时收到通知。

目前 Cursor Origin 页面显示的是等待名单表单,这很合理,因为他们需要复制大量代码库。

> "代码正在以任何基础设施都无法处理的速度发展。Origin 就是为这一刻而设计的。"

引用 🚨 AI News | TestingCatalog @testingcatalogCursor team is teasing a potential release today, which can be either Composer 3 (Vega) or Grok 4.6, that was accidentally enabled yesterday for a short period of time. Soon 👀查看被引原帖 ↗
查看英文原文
Cursor Origin (Cursor Review) has been tested in a closed beta with selected partners internally and will likely be announced later today.

Users will get access to new tabs called Codebase and Review.

> The Codebase tab will let users sync their repositories from Github.

> The Review tab will be dedicated to the "agentic code review" process, in which users will be notified when their attention is required.

Currently, the Cursor Origin page shows a waitlist form, which makes a lot of sense since they will have to copy a ton of code repositories.

> "Code is moving faster than any infrastructure was built to handle. Origin was designed for this moment."
◔ 9.1 万 次浏览(3 条合计)♥ 570⇄ 13▶ 含视频新品看原帖 ↗
向阳乔木@vista8 · 中文博主 · 1 天前向阳乔木,中文圈 AI 工具与趋势博主

再次强烈推荐 OpenCodex 项目,如果你习惯了 Codex。

但有些任务还是要用其他模型,比如前端用 Kimi3 ,多模态用Gemini,追求速度用 Deepseek。

全都能集成到 Codex 中,随时切换使用。

地址见评论区

Lisan al Gaib@scaling01 · 博主 · 1 天前高频 AI 模型测评与爆料博主

*蒸馏的迹象*
> 肯定是风

TL;DR:Kimi-K3 用 Opus 推理预填充时输出会改变。Inkling 和 DeepSeek 不会。

但 Kimi 用 Inkling 推理预填充时输出不改变

引用 Timothée Chauvin @timotheechauvinA new paper shows that encryption of chain-of-thought was done really poorly across Anthropic, OpenAI, Google. It's not very consequential, but is another example imo of them going so fast that they fail at basic security stuff. arxiv.org/abs/2608.09867查看被引原帖 ↗
查看英文原文
*signs of distillation*
> must have been the wind

TLDR: Kimi-K3's output changes when prefilled with Opus reasoning. Inkling and DeepSeek don't.

But Kimi's output doesn't change when prefilled with Inkling reasoning

十几年前程序员有一些对抗裁员的脏办法,比如留一些recurring bugs,让自己不被裁员,一旦裁员之后定期爆发bug或者性能问题,原来codebase太脏太乱,导致没人能step in后正确修复,

我预计马上有一批人会日常大量输出无用文档、减少有用文档记录信息的方式,来对抗办公multi agent对自己工作的替代。

只要每天大量输出垃圾文档、表格、文字、信息、流程,同时把真正有用的决策、信息、数据留在口头、视频会议、加密文档、仅限邀请的文档、权限控制严格的CRM或者其他平台,

就能最大限度地减少未来自己手头工作被AI Agent全盘接手的可能性,就能最大限度地保证自己的工作。

典型的就是国内机关国企事业单位,用的都是这些脏办法,想替代都替代不了。

引用 lidang 立党 (劝人卖房/学CS/买SP500/纳100/OpenAI/Anthrop第一人) @lidangzzz全中国最稳定的工作一定是公务员,办公Agent最后一个替代的一定是国内机关单位、事业单位、国企员工。 因为办公Agent的刚需就是非常规范的、整理好的CRM、ERP、SAP、内部规范完善链接的大量文档、表格、PDF、PPT、邮件、内部公开详细分类的聊天组, 而机关单位、事业单位、国企绝大多数的信息, 都是隐藏的,90%的不关键信息藏在几千个乱七八糟的微信群、钉钉、飞书里, 而10%的关键的、负责任的、决定最终决策的信息,藏在领导的面谈里,领导绝对不会、不敢、不能落到会议纪要、文件、邮件和微信群里,只敢亲口半画饼半威胁地让手底下员工区执行。 领导也明白,大部分这种工作决策部署都是浪费时间,找个微信群发过去就执行了,少部分关键信息一定不能留痕迹、一定off the record,一定天知地知你知我知,不能让任何第三个人知道“我曾经说过这句话”。 就这种狗屎一样的工作环境和氛围,这种大逃杀、狼人杀一样的猜疑链,黑暗森林一样的推理环境, 你让办公agent在里面能猜出什么来呢? 只能猜你亲妈的阳寿了。查看被引原帖 ↗
Ethan Mollick@emollick · 创始人 · 1 天前沃顿商学院教授,AI 应用研究权威

哎呀,我们不会又要回到这种提示词的老套路了吧?我真希望Anthropic能测试一下这玩意儿到底稳不稳,因为我们自己的实验(用的稍微老一点的模型)发现根本不管用。

引用 Anthropic @AnthropicAIClaude 研究版本尝试破解黎曼假说未成功,但在相关问题取得进展:将黎曼 zeta 函数零点满足假说的下界从 41.6% 提高到 67.2%。查看被引原帖 ↗
查看英文原文
Oh no, we aren’t going to go back to this sort of prompting again, are we? I would love Anthropic to test if it actually works robustly, because our experiments (with slightly older models) found it did not.
Sebastian Raschka@rasbt · 博主 · 1 天前

天哪,Meta 昨天发布了新的开源权重 LLM,这自 Llama 时代以来还是头一次。

他们的 Meta Muse Glimmer 是一个 30B 多模态推理模型,采用类似 Gemma 的架构设计。("Glimmer"可能是"Spark"的文字游戏,Spark 是能力更强的原始模型,Glimmer 从它蒸馏出来。不过 Muse Spark 目前只能通过 Meta 的 Model API 获取。)

架构方面,以下是主要要点:

1. 仅 131k 的上下文窗口,而 Qwen3.6 和 Gemma 4 原生支持 2 倍的长度;还不错吧,但在代理框架时代可能有点捉襟见肘

2. 这是个密集模型,不是混合专家。(所以拿它跟 Qwen3.6 27B 比更合理,而不是 Qwen3.6 30B-A3B。)

3. 混合注意力机制,用了分组查询注意力(GQA)和滑动窗口注意力(SWA);SWA:GQA 的比例是 3:1 局部:全局。相比之下,类似采用这些组件的 Gemma 4 是 5:1。

4. GQA 和 SWA 都用了门控注意力;这东西最近几个月很流行。基本就是对注意力输出应用 sigmoid 门来决定有多少注意力信息流入残差连接。有意思的是它用的都是相对标准的 GQA 和 SWA,而不是 Nemotron 或 Qwen3.6 那种混合注意力机制。

5. GQA 比例特别极端:32 个查询头但只有 2 个 KV 头;对比一下 Gemma 4 31B 用的是 32 Q / 16 KV(局部头)和 32 Q / 4 KV(全局头)。这意味着 Meta Glimmer 的 KV 缓存超级小。

总的来说,架构最接近的是 Gemma 3 27B(包括 Gemma 风格的前后 RMSNorm 放置)和 Gemma 4 31B,不过做了些调整,比如用 SwiGLU 替代 GeGLU 激活函数、加了门控注意力、以及前面说的那个更极端的 GQA:SWA 比例。

最突出的是它的 KV 缓存效率极高。
KV 缓存 / token 比率(BF16)如下:

- Muse Glimmer: 52 KiB(越低越好)
- Qwen3.6 27B: 64 KiB
- Gemma 4 31B: 840 KiB

性能方面,他们自己的基准测试显示大多领先 Qwen3.6。根据人工智能指数的独立综合基准,它略低于 Qwen3.6(见下图)。用个几天就知道它真实水平了。

总体来看,这是个不错的模型,特别是对代理工作流。最吸引人的是内存占用特别低,还有相当快的 prefill 和 decode 速度。也很高兴看到 Meta 又开始发布开源权重了 :)

查看英文原文
Whoa, Meta released a new open-weight LLM yesterday, something that hasn't happened since the good old Llama days.

Their Meta Muse Glimmer model is a 30B multimodal reasoning model with a Gemma-like architecture design. (“Glimmer” is probably a wordplay on “Spark,” the more likely capable model from which Glimmer was distilled. Muse Spark is only available through Meta’s Model API, though.)

Architecture-wise, here are some of the main points:

1. "Only" a 131k context window, compared to Qwen3.6 and Gemma 4, which support 2x that natively; it's reasonable, but maybe on the shorter end in the age of agent harnesses

2. It's a dense model, not a mixture-of-experts. (So, it's fairer to compare it to Qwen3.6 27B than Qwen3.6 30B-A3B.)

3. Hybrid attention with grouped-query attention (GQA) and sliding window attention (SWA); the SWA:GQA pattern is a 3:1 local:global ratio. Other models like Gemma 4, which uses similar components, have a 5:1 ratio for comparison.

4. It adopts gated attention for both GQA and SWA; gated attention has become quite common in recent months. It basically applies a sigmoid gate to the attention output to decide how much of the attention information enters the residual connection. The interesting point is that it uses relatively standard GQA and SWA rather than hybrid attention mechanisms such as Nemotron or Qwen3.6.

5. A very extreme GQA ratio: 32 query heads and only 2 KV heads; for comparison, Gemma 4 31B uses 32 Q / 16 KV in the local heads and 32 Q / 4 KV in the global heads. This means that Meta Glimmer has a very small KV cache.

Overall, the probably most similar architecture is Gemma 3 27B (including the Gemma-style pre/post RMSNorm placement) and Gemma 4 31B, but with some tweaks like SwiGLU instead of GeGLU activations, gated attention, and the more extreme GQA:SWA pattern mentioned before.

What stands out is its extreme KV-cache efficiency.
I.e., the KV CACHE / TOKEN ratios (in BF16) are:

- Muse Glimmer: 52 KiB (lower is better)
- Qwen3.6 27B: 64 KiB
- Gemma 4 31B: 840 KiB

Modeling-performance-wise, their own benchmarks show that it's mostly ahead of Qwen3.6. According to the independent composite benchmarks on the Artificial Analysis Intelligence Index, it's slightly behind Qwen3.6 (see figure below). So, a few days of using it will tell where it really ranks.

Overall, it looks like a solid model, particularly for agentic workflows. What stands out most is its very low memory footprint and also pretty fast prefill and decode speed. It’s also just great to see Meta releasing open weights again :).
◔ 18.5 万 次浏览(9 条合计)♥ 1,336⇄ 192研究看原帖 ↗
Greg Brockman@gdb · 创始人 · 1 天前Greg Brockman,OpenAI 联合创始人兼总裁

看看我们在德州如何负责任地构建AI基础设施:
openai.com/index/responsible…

查看英文原文
How we're responsibly building AI infrastructure in Texas:
openai.com/index/responsible…
Bilawal Sidhu@bilawalsidhu · 博主 · 1 天前

有时候,你会遇到一项技术,才意识到苹果一直把我们留在过去。

查看英文原文
Sometimes you get a piece of technology that makes you realize that Apple was just keeping us in the past.
Gary Marcus@GaryMarcus · 博主 · 1 天前

LeCun最近说的很多话,都像是我这些年一直在讲的。但这句说得真好,我自愧不如:

查看英文原文
Most of what LeCun says lately sounds like what I have been saying for years. But this is really good, and I couldn’t have said it as well:
OpenAI Developers@OpenAIDevs · 公司官方 · 1 天前OpenAI 开发者平台官方

来自世界各地的 Build Week 参与者用 Codex 把他们的创意变成了真实项目。

查看英文原文
Build Week participants from around the world turned their ideas into real projects with Codex.
宝玉@dotey · 中文博主 · 1 天前宝玉,中文圈 AI 翻译与科普大 V

这篇推文关于文本水印原理讲的很清楚。

生成文本的时候,模型每要输出下一个 token,就拿前面已经生成的所有词加上一个密钥,算出一个哈希值。这个哈希值会把词表里的词随机分成两组,比方说绿组和红组。然后模型在选下一个词时,悄悄提高绿组词被选中的概率。最终输出的文本看起来完全正常,但统计上会偏向绿组词。

检测的时候反过来。拿着同一个密钥,对文本里的每个词重新算一遍哈希,还原出当时的绿组和红组。如果整篇文本中绿组词的占比显著超过 50%,就判定这段文本带有水印。

几个常见疑问:

1. 改写能破解水印吗?

大部分情况下可以。因为你一改词,前面的 token 序列变了,哈希也跟着变,绿组红组就对不上了。但也有改进方案用统计模型而不是确定性函数来生成哈希,让水印对改写有一定的适应能力。

2. 水印会降低文本质量吗?

理论上会,因为模型不再完全自由地选最优词,而是被约束在绿组里选。但 Google 在两万条文本的人类反馈实验中称,大多数人感知不到质量下降,因为表达同一个意思的词有很多种组合。不过在短文本或者高度确定的输出上(比如1+1=2),水印确实会失效。

3. 密钥会被逆向破解吗?

理论上需要指数级数量的样本才能还原绿组红组的划分,所以直接暴力破解不现实。但如果检测器公开了,攻击者可以通过不断试探来摸索出规律。

工程难点

第一,大模型是流式输出文本的,一个词一个词往外蹦,没法像学术方案那样等整段写完再调整。
第二,如果密钥泄露,水印就废了,所以需要多组密钥轮换。
第三,代码类输出不能随意换词,否则代码会出错,水印只能加在注释、变量名这类可替换的部分。

如果 Anthropic 公开检测器,等于给了攻击者一个免费的练习靶,水印会被快速找到绕过策略。

如果像 Google 的 SynthID 那样把检测器留在内部,更安全一些,但学术界已经有论文展示了不需要检测器也能零样本破解水印的方法。

密集改写(同义词替换加句法重组)可以有效消除水印,免费的改写工具就能绕过 Google 的 SynthID。

前沿实验室对此心里有数,也没指望水印能拦住刻意绕过的人。大多数普通用户不会专门去除水印,而欧盟监管方也只需要足够好就行。

水印更多是一种应对欧盟监管的合规动作。

引用 Alex Cui @alexcdotClaude's watermark probably doesn't work how you think. As the CTO of GPTZero, I'll explain how Anthropic, Google and OpenAI are building text watermarking in this brief explainer and whether it can be defeated. Almost all forms of watermarking that are fast and cheap enough for a frontier lab have the same formula, following the KGW method: In generation: 1. Let's say you've generated n tokens so far. Take those n tokens + a secret key to generate a random hash 2. Use that hash to randomly reweight the probabilities for the n+1 token, and then sample from that new distribution. In the simple case, you could split 50% of all English words into a green or red set based on your hash, and boost the probability of words in the green set. For watermark detection: 1. For each token, see if it was in the green or red set. 2. To do this, recreate the hash based on the secret key and the text preceding the current token. Then, recreate the green and red set of words. 3. Once you've checked all the words in the text, if the next token is selected disproportionally from the green set more than 50% of the time, you claim the text has the watermark. I can tell you want to ask the following: 1) Isn't it easy to mess up the hash if you paraphrase the text? The answer is mostly yes, however, you can use a statistical model to get your hash instead of a deterministic function (SIR, Adaptive Watermark). Since the entire watermark is probabilistic, this is fine. 2) Doesn't this make the text much worse? The answer is yes, it does - Yes, it does – but for most people, it's imperceptible (Google claims in human feedback study with 20,000 texts), since there are exponentially many ways to write the same paragraph. DiPmark does something more sophisticated to avoid shifting the text distribution on average. Of course, watermarks fail on short text or highly predictable texts like "2+2=4". 3) Shouldn't it be easy to figure out the green and red sets? The answer is no. You would ne查看被引原帖 ↗

今天的开源框架,留给github发酵攒星星,各大LLM pretraining当dataset,

以前自己手搓一个新项目,可以乱拳打死老师傅,把老项目给彻底替代掉,

现在前后端各个组件数据库完全固定钉死下来了,头部开源框架等于核心供应链,项目越头部,LLM越熟练,等于控制了阳光、空气和水的使用权,更恐怖了。

引用 象牙山刘能 @disksing以后开源就跟社区协作没啥关系了,开源的目的就是为了有一天能被训练进大模型然后省token。查看被引原帖 ↗
Vercel@vercel · 公司官方 · 1 天前前端云平台 Vercel 官方,AI 建站工具 v0 母公司

我们在所有方案中都推出了免费的Sandbox Egress防火墙。

Vercel Sandbox的能力远超计算隔离。最近的安全研究表明了原因:不受信任的代码需要在网络边界进行隔离,而不仅仅是运行时。

vercel.com/blog/a-sandbox-wi…

查看英文原文
We're making our Sandbox Egress Firewall free on every plan.

Vercel Sandbox goes beyond compute isolation. Recent security research shows why: untrusted code must be contained at the network boundary, not just the runtime.


vercel.com/blog/a-sandbox-wi…
◔ 8.7 万 次浏览(2 条合计)♥ 155⇄ 12新品看原帖 ↗

OPC最直观的商业落地项目:

从小红书、微信公众号、推特、reddit、linkedin上扒文章,一顿洗稿翻译截图,其他平台再发一遍,再生成视频和tts口播自动传到youtube、抖音、B站等等平台——吃一口,拉一盆,到处赚“创作者收益”(实际是洗稿者收益),

当然还有一种专门在推特上转发翻译英文原文的大傻瓜。

引用 0xLeon @Leoninweb3所以短期看OPC是不是最大的客户群… 赚OPC的钱比OPC赚钱更容易…查看被引原帖 ↗
swyx@swyx · 博主 · 1 天前知名 AI 播客 Latent Space 主理人

搭建会生成循环的循环结构

引用 sunil pai @threepointone新文章:每个公司都需要一个cassandra。我提议建立后台agent/worker角色,这是人们通常讨厌与之合作的。查看被引原帖 ↗
查看英文原文
set up loops that make loops
歸藏(guizang.ai)@op7418 · 中文博主 · 1 天前歸藏,中文圈 AI 工具与提示词博主

看来 Deepseek 的 Agent 要发了呀。

他们注册了 Deepseek Harness 团队的公众号。

Chubby♨️@kimmonismus · 博主 · 1 天前Chubby,高频 AI 新闻聚合博主

Grok 4.6 即将来临。据说在 Cursor 里已经被提过了。

根据马斯克所说,这将是一个 1.5t 模型,SFT 和 RL 都有改进。

引用 lauren @potetoit’s a good day to ship查看被引原帖 ↗
查看英文原文
Grok 4.6 incoming id say. It was already mentioned within Cursor.

According to musk it will be a 1.5t model with improved SFT&RL.
◔ 9.7 万 次浏览(4 条合计)♥ 691⇄ 16新品看原帖 ↗
宝玉@dotey · 中文博主 · 1 天前宝玉,中文圈 AI 翻译与科普大 V

重点明明是你有无限 Fable 的 Token😜

引用 LinearUncle @LinearUncleclaude 一个未发布模型尝试攻克黎曼猜想,尝试了 650 个方向都失败了。 人类继续PUA 它“请继续”,“相信你自己!你可以的”,最后虽然没有最终攻破,但是也拿到了数学上一个非常不错的结果! 看来我上次PUA方向不对,我用梁文峰PUA Deepseek-v4-flash,是威胁要开除它,应该鼓励它才对! 大模型也是吃软不吃硬!查看被引原帖 ↗
Rowan Cheung@rowancheung · 博主 · 23 小时前AI 日报 The Rundown 创始人
连环推 ×2

开源回来了。

Zuck 刚宣布 Meta 要开源 Muse Glimmer 权重,Muse Spark 1.2 即将推出

不过一年前,大家都怀疑 Meta 在 AI 竞赛中的地位

在我与他的采访中,他承认了失误之处:

"Llama 4 的发展轨迹不是我认为应该的样子。在很多方面确实比 Llama 3 进步巨大,但我们的目标不是仅仅比 Llama 3 好一点。我们是前沿实验室,应该做引领性的工作。"

所以他彻底推翻重来:

"我觉得已经充分学会了如何组建一个实验室,想要重新规划我们的工作。"

他这次重建背后的信念:

"AI 将是我们有生之年最重要的技术。建立绝对领先的模型能力,对释放创意至关重要。"

引用 Mark Zuckerberg @finkdToday we're also opening the weights for Muse Glimmer, a great 30B parameter dense model that can run locally. Soon we'll also release the weights for Muse Spark 1.2, our latest foundation model. Meta is a strong supporter of open source and I'm proud of these releases. Congrats to @alexandr_wang and the MSL team for all your great work on these models.查看被引原帖 ↗
查看英文原文
Open source is so back.

Zuck just announced Meta is opening the weights for Muse Glimmer, with Muse Spark 1.2 coming soon

But a year ago, everyone doubted Meta's position in the AI race

In an interview I did with him, he admitted where they missed the mark:

"Llama 4 was not on the trajectory that I thought it needed to be on. It was in many ways a big improvement over Llama 3, but we weren't trying to be a bit better than Llama 3. We're a frontier lab. It wants to be doing leading work."

So he blew it up and rebuilt:

"I felt like I'd learned enough about how to set up a lab that I wanted to reformulate the work we were doing."

His conviction behind the rebuild:

"AI is gonna be the most important technology in our lives, in our lifetimes. Building the capacity to build absolutely leading models is gonna be really critical for unlocking creativity."
More details from Zuck on Muse Glimmer, Muse Spark 1.2 (coming soon) and his vision on a personal superintelligence available 24/7 to all


meta.com/thefutureisforevery…
◔ 4.3 万 次浏览♥ 133⇄ 12▶ 含视频新品看原帖 ↗
向阳乔木@vista8 · 中文博主 · 1 天前向阳乔木,中文圈 AI 工具与趋势博主

论 Vibe coding 最合适的配套游戏,好友七娘整理:

1. 炉石传说
2. 金铲铲
3. 皇室战争
4. 致病本源

欢迎补充...

Gorden Sun@Gorden_Sun · 中文博主 · 1 天前中文圈高频 AI 资讯与开源项目博主

我Mac上安装的各种小软件,有很多是满足我犄角旮旯的需求。
Amphetamine:AI时代必备,可以让MacBook盒盖不休眠。免费。
UU远程:AI时代必备,比Codex和Claude自带的远程好用。免费。
Supercharge:实现点击Dock栏里的应用图标可以隐藏应用窗口;Command+X/V实现剪切文件和粘贴文件;Finder里点击退格键直接删除文件(默认是Command+退格键);Finder里回车键直接打开文件;Finder里Option+N新建txt/md任意格式的文件;配置快捷键可以一键隐藏所有应用窗口,实现返回桌面的效果。总之就是把窗口体验变得像Windows。收费。
TabTab:AltTab收费后改这个了,体验更好更简洁。收费。
Raycast:功能很全很多,但我只用里面的剪切板记录功能,这个功能比其他专门的软件做的都要好。部分功能免费。
Ice:右上角的菜单栏整理。新版本Mac系统不再需要了。免费。
Mac Mouse Fix:能让鼠标滚轮有控制缩放的能力。免费。
RDM:修改内置屏幕和外接屏幕的分辨率。免费。
PearCleaner:删除软件的同时删除关联文件。免费。
LocalSend:本地设备间文件互传。免费。
Rectangle:应用窗口快速布局。免费。
SoundSource:分别调节每个软件的音量。收费。
Deeper:功能很多,我只用里面的给Dock栏添加空白的功能,可以用于划分功能区域。免费。
PhotoScapeX:编辑图片的应用。能用,但是比不上Windows上的光影魔术手。基本免费。
Lyn:极简图片查看工具。收费,太贵,性价比不高。
CheatSheet:快捷查看当前应用的所有快捷键。免费。
Little Snitch:网络监控和拦截软件,我主要用于配置Claude应用必须走小火箭的端口。收费。
IINA:视频播放器,比不上Windows上的PotPlayer。免费。
Mole:清理空间。CLI版本免费,APP收费。
EdgeSpeak:语音识别+语音合成,而且有CLI供Codex等Agent调用。收费。

Lisan al Gaib@scaling01 · 博主 · 1 天前高频 AI 模型测评与爆料博主

没什么大不了的呗

反正到下界维持现状、两年内不上涨直到假说最终被证明之前,应该不会有啥变动

引用 Anthropic @AnthropicAI研究版Claude尝试证明黎曼假设,虽未解决但在相关问题取得进展——将黎曼zeta函数零点满足假设的下界从41.6%提升至67.2%。查看被引原帖 ↗
查看英文原文
no biggie

surely the lower bound stays where it's at and doesn't increase over the next 2 years until the hypothesis is finally proven
Ethan Mollick@emollick · 创始人 · 1 天前沃顿商学院教授,AI 应用研究权威

数据中心与前几次工业革命的轻工业相比,一个真正的问题是它们不需要多少人运营(虽然建造时需要更多人)。这打破了产业在显著本地负面外部性与显著本地收益之间的权衡。

查看英文原文
A true issue with data centers compared with the light industries of previous Industrial Revolutions is they don’t require many people to run (though building them takes more). It breaks the industry trade-offs between palpable local negative externalities & palpable local gains.
Lisan al Gaib@scaling01 · 博主 · 1 天前高频 AI 模型测评与爆料博主

看起来图宾根的 Ellis 研究所是在德国做 LLM 工作的最好去处。

引用 Alexander Panfilov @kotekjedi_mlWe can finally talk about it: We found a way to extract hidden reasoning of frontier models using a vulnerability in the APIs of every frontier AI company. We verified that our reasoning token count matches billed API thinking tokens 1:1 for most of the prompts we queried.查看被引原帖 ↗
查看英文原文
It seems like the Ellis Institute Tübingen is the place to be in Germany when you want to work on LLMs
Aravind Srinivas@AravSrinivas · 创始人 · 1 天前Perplexity 联合创始人兼 CEO

尽情享受 Perplexity agent api 上美国托管的 K3 吧!

引用 Perplexity Developers @perplexitydevs在Perplexity Agent API上使用Kimi K3,独家托管于美国服务器。查看被引原帖 ↗
查看英文原文
enjoy US hosted K3 on perplexity agent api!
◔ 3.5 万 次浏览♥ 237⇄ 12▶ 含视频新品看原帖 ↗
swyx@swyx · 博主 · 1 天前知名 AI 播客 Latent Space 主理人

gpt luna max vs claude fable ultracode

发了个请求:"请用 fal 的开放模型做一个基本忠实的 grok imagine 克隆"

早上看到两个结果,我以为左边是 fable 右边是 luna。

结果反了...竟然反过来了!

客观来说,fable 的视觉克隆做得更好。但 luna 不知道咋的更理解了我的意思,考虑到我偏好开放模型,它其实做出的克隆更好用。

查看英文原文
gpt luna max vs claude fable ultracode

sent "pls build a mostly faithful clone of grok imagine with open models via fal"

i woke up to these two and assumed fable was left and luna was right

i was wrong... it was the other way!!

objectively, fable did the better visual clone. but luna somehow understood intent better and created the more USABLE clone given my open model bent.
Lisan al Gaib@scaling01 · 博主 · 23 小时前高频 AI 模型测评与爆料博主

我不会惊讶,如果OpenAI的10T或20T模型有几乎完全无法理解的推理轨迹

你能看出来为什么机械可解释性在那时会变得这么重要,当这么多推理发生在内部的时候

光靠监控思维链根本不可扩展,或者说至少需要付出额外的努力来保证所有内容的准确性和可读性

引用 Lisan al Gaib @scaling01I looked at all the reasoning traces from Opus 4.8 and GPT-5.6-Sol to me it seems like Opus 4.8 is working on a much higher abstraction level, while Sol is just doing the dirty work. it just explores and keeps spamming new hypotheses until it finds a solution not only is Sol more efficient in its reasoning, but I also think the Sol approach is more scalable and more sciency OpenAI has been on a very good trajectory this year, and I'm now more confident than ever, that they will be back in the lead soon but they are also more likely to accidentally train a model that speaks neuralese and have a harder time monitoring their models查看被引原帖 ↗
查看英文原文
I wouldn't be surprised if a 10T or 20T OpenAI model had almost completely unintelligible reasoning traces

you can see why mech interp becomes important at that point when so much of the reasoning happens internally

just monitoring chains-of-thought is not as scalable or at least requires some extra effort to keep everything faithful and readable
Chubby♨️@kimmonismus · 博主 · 1 天前Chubby,高频 AI 新闻聚合博主

价格战进入下一轮,GLM 的 Ziphu 也开始重置费率了。

Zcode 已有 100 万用户。看来竞争还在继续,这很不错。

引用 Z.ai @Zai_orgZCode now has 1 million users. As a thank-you to our community, we’ve reset usage limits for all GLM Coding Plan users. We’re also rolling out an update that helps turn long-horizon capabilities into completed work: - More intelligence in real engineering workflows - 98% cache hit rate, providing around 1.8x more usage Download: zcode.z.ai/en Join the community: discord.gg/EpH5XkTyhu查看被引原帖 ↗
查看英文原文
The price war is entering its next round: GLM's Ziphu is also starting to reset the rates.

Zcode has 1 million users. It's good to see the competition continues.
◔ 3.4 万 次浏览♥ 411⇄ 8▶ 含视频动态看原帖 ↗
elvis@omarsar0 · 博主 · 1 天前

来自 Meta 的印象深刻的新论文。

(务必收藏)

缩放法则假设模型大小和训练数据独立地作用于损失。这项工作引入了 Skaling law,通过单个交互指数耦合容量和数据。额外项在插值和外推中都使平均绝对百分比误差降低 1.5 到 3 倍。

最大的改进出现在数据稀缺和重度过训练体制中,即标准 Chinchilla 和 Kaplan 形式开始偏离的地方。

与限制在低计算运行的稀疏网格配对,它用大约 10 倍少的计算来外推完整网格,相比统一扫描。

为什么这很重要?

部署现在发生在远超计算最优点的地方。一个在那里保持准确、并能从小规模运行拟合的法则,改变了预训练预算如何规划。

论文:
arxiv.org/abs/2608.07222

在我们的学院跟踪更多趋势 AI 论文:
academy.dair.ai/

查看英文原文
Impressive new paper from Meta.

(bookmark it)

Scaling laws assume model size and training data act on loss independently.

This work introduces Skaling law, which couples capacity and data through a single interaction exponent. The extra term cuts mean absolute percentage error by 1.5x to 3x across both interpolation and extrapolation.

The largest corrections land in the data-scarce and heavy-overtraining regimes where the standard Chinchilla and Kaplan forms drift.

Paired with a sparse grid restricted to low-compute runs, it extrapolates the full grid using roughly 10x less compute than a uniform sweep.

Why does it matter?

Deployment now happens well past compute optimal. A law that stays accurate there, and that can be fit from small runs, changes how a pretraining budget gets planned.

Paper:
arxiv.org/abs/2608.07222


Track more trending AI papers in our academy:
academy.dair.ai/
Chubby♨️@kimmonismus · 博主 · 1 天前Chubby,高频 AI 新闻聚合博主

真是疯狂的一天:
– Meta 重返开源,推出 Glimmer(一个优秀的中型代理模型),即将发布开源权重的 Spark 1.2。
– Anthropic 用一个未发布的 Claude 模型尝试解决黎曼假设。虽然失败了,但意外改进了一个长期存在的下界——黎曼 zeta 函数满足假设的零点比例从 41.6% 提升到 67.2%。
– 与此同时,Sonnet 的价格长期下调(肯定是竞争导致的),而关于 Gemini 3.5 Pro 的新闻暗示 Google 直接跳到了 Gemini 4。
一天之内。疯狂。
我要睡觉了。明天见!

查看英文原文
What a day:

– Meta is back in the open-source game, delivering Glimmer, an excellent agentic mid-sized model, and will soon release Spark 1.2 with open weights.

– Anthropic tried to solve the Riemann Hypothesis with an unreleased Claude model. It failed, but along the way unexpectedly improved a longstanding lower bound for the proportion of zeros of the Riemann zeta function known to satisfy the hypothesis, from 41.6% to 67.2%.

– At the same time, Sonnet’s prices were cut for the long term, surely because of the competition, and the news around Gemini 3.5 Pro seems to suggest that Google is moving straight to Gemini 4.

All in one day. Crazy.

I’m going to sleep now. See you tomorrow!
宝玉@dotey · 中文博主 · 1 天前宝玉,中文圈 AI 翻译与科普大 V

转译自 Lenny Rachitsky

我从 Cursor 人才主管 Adam Ward (
@wardadamp
) 那里学到的最大收获:

1. 前线部署工程师 (FDE,forward deployed engineer) 是目前科技界最抢手的职位。

他们必须具备深厚的技术功底,同时又能和销售团队打好配合,在公司高管面前从容自信。他们要把复杂的产品转化为实际的商业结果,帮助客户优化 AI 账单,而不是仅仅在“最大化消耗 Token” 上做文章。

Adam 将这股热潮与当年的移动端工程师招聘潮相提并论——只不过,当年那股热潮持续演进了两年,而“现在,一切都发生在几天或几周之内”。

2. 传统的招聘漏斗,被 Adam 戏称为“死亡漏斗” (funnel of doom),其实是个专门打造平庸团队的机制。

标准的招聘漏斗是这样的:联系 100 个人,20 个回复,每一轮面试淘汰掉一批,最后录用剩下的人。

但这套玩法建立在一个漏洞百出的前提上:最先回复你的那 100 个人,根本不是行业里最顶尖的 20%,他们可能只是那天碰巧心情不好、想换工作而已。等到候选人走到漏斗最底端时,你其实已经是在一个“矮子里拔将军”的池子里挑人了。

即便你的面试流程再严苛,最终拼凑出的也只是个平庸的团队。

3. 破局之道在于:把每一次招聘都当成“高管猎聘” (executive search) 来对待。

这意味着你需要严谨地界定岗位需求 (rigorous scoping)——对这个特定岗位而言,什么是“优秀”?

这意味着你要精心绘制候选人图谱 (deliberate candidate mapping)——全世界最适合这个岗位的 50 个人到底是谁?

这还意味着你要穷追不舍 (relentless pursuit)——锁定这 50 个人,死死咬住不放。这种视角的转变,意味着你要从确立“卓越支柱” (pillar of excellence) 开始招人,而不是一开始就撒一张盲目的大网。

4. 现在的“新危险区”出现在候选人签下 Offer 之后。

候选人接了 Offer 又毁约的趋势正在上升,而且资深人才从签约到入职之间往往有几周的空档期。因此,Cursor 把“接受 Offer 后”的这段时间当成一场专门的战役来打。

他们会组织新入职的员工一起聚餐,在入职第一天前就把笔记本电脑寄过去,在任何人正式打卡上班前,就让他们感受到社区的归属感。这不仅能确保大家准时报到,还能让他们在入职后更快地进入工作状态。

5. 实战考察 (work trials) 是预测一个人能否胜任工作的最佳指标。

研究早就反复证明了这一点,但目前行业的默认面试形式,依然是隔着桌子进行一轮又一轮的一对一聊天。Cursor 则会安排长时间的现场实战体验,让候选人与团队肩并肩,共同推进一个真实的(或者高度还原的)项目。

这种形式能同时传递出三层信号:技术能力、核心价值观以及协作默契度,同时也给了候选人自己做决定、双向奔赴的数据依据。Cursor 曾经做过一个实验,试着取消实战考察,结果团队对招进来的人信心大跌。于是,他们果断把实战考察加了回来。

6. 跑去问你的人脉圈“你认识的最牛的工程师是谁?”——这是一个陷阱。

相反,你应该根据你界定的岗位需求,抛出极其具体的问题:“在你合作过的所有产品工程师里,谁和设计师配合得最默契?”或者“谁能把一个技术框架转化成产品,而且比你见过的任何人都做得好?”——因为具体的提示词 (prompts) 才能让人脑海中迅速跳出确切的名字。

当有几个你信任的人,都不约而同地提到同一个人时(“Lenny 提到了他,Sally 也提到了他”),你就可以通过这种交叉验证,锁定那个你要放进“全球 50 强候选人”名单里的目标了。

7. 明确岗位需求 (scoping the role) 是大多数招聘流程中最被忽视、投入最少的一步。

“我一看到牛人就能认出来”——这绝对是个错觉。如果你不在前期花时间去明确、并给这个岗位在特定公司所需的技能、经验和特质排个序 (stack-rank),你就根本不会有针对性的招聘策略,没法给出打动人的说辞,也设计不出真正有效的评估方法。后续所有的流程,其实都源于这最开始的一步。

另外,要抵制住“看名企光环”的捷径诱惑:在合适的时间、在某家大厂的优秀团队待过,这仅仅是“有可能”代表他很优秀;如果你直接照搬别人的用人标准,那你其实就是在给一个你本就觉得千疮百孔的招聘流程“抬轿子”。

8. 当顶尖候选人告诉你“现在时机不对”时,试着降低你的要求:“我不是要面试你——我只是希望能有下一次交流。”

Adam 会默默埋下种子——比如喝杯咖啡、邀请参观办公室、或者把他介绍给一位对他的工作极其着迷的团队成员——这些种子会在几周、几个月甚至几年后开花结果。因为如果你选择了沉默,想着“秋天再联系”,那最好的结果也就是在没人抢走他的情况下,你碰巧捡了个漏。

在 Cursor,寻找一个热情的引荐人,所花费的心力几乎和写一封招募邮件一样多,因为一旦别人无视了你的第一条消息,再无视你的第二条、第三条就会变得越来越容易。

Cursor 还会巧妙地利用“主场优势” (home games),把接触的重点放在一起吃顿饭上——“这既能让人卸下防备,又能拉近关系”——以及邀请对方参观办公室,刻意营造出那种大学校园导览时让人觉得“这里就像家一样”的直觉。

9. 促成签约 (closing) 是一项团队运动,而且从第一次对话就开始了。

当 Cursor 正在跟进一位高素质的候选人时,团队会针对这个人召开每日站会 (daily standups):我们对他的动机有了什么新了解?他提出了哪些顾虑?团队里谁最适合去解答这些疑虑?

每一位核心候选人都会有一个专属的 Slack 频道。所谓的“促单签约”,不应该感觉像是一次生硬的销售推销,而应该是整个流程中持续为你解答疑虑后,水到渠成的自然结果。

10. 用心是不花钱的,但它却是招聘中最大的“不公平优势” (unfair advantage)。

Adam 说,纵观整个招聘领域的历史,决定候选人满意度得分的最关键因素,就是让他们感觉到这家公司是真心实意想要你。

因此,Cursor 会精心设计每一个接触点 (touchpoints):谁来迎接候选人,谁在午餐时坐在他们旁边,哪位面试官能和他们聊聊小众爱好——这一切都是精心挑选的,绝不是谁日程表有空就拉谁来凑数。

要让这种“用心”可持续,你需要转变一下思维:你是在为这个岗位招那“唯一的一个人”,而不是在招 10 个人。“不要把精力分散在 10 个人身上;把所有的专注都留给那唯一的 1 个。”

引用 Lenny Rachitsky @lennysanCursor在竞争中持续胜出,Elon Musk以600亿美元收购该公司助力SpaceX AI竞赛。成功源于精英团队建设。人才负责人Adam Ward分享了高人才密集团队的建设方法,涵盖招聘策略、顶尖人才甄选、前置工程师等内容。查看被引原帖 ↗
Chubby♨️@kimmonismus · 博主 · 1 天前Chubby,高频 AI 新闻聚合博主

享受你的重置吧。Tibo 信守承诺了 <3

引用 Tibo @thsottiauxUsage limits have been reset for all paid ChatGPT Work and Codex users. Happy Monday you all. Hope it is a fantastic week.查看被引原帖 ↗
查看英文原文
enjoy your reset. Tibo kept his promise <3
Bindu Reddy@bindureddy · 创始人 · 1 天前Abacus.AI CEO,AI 行业观点博主

包括 Microsoft 和 Amazon 在内的大公司正在退出前沿模型的开发

最新的 Google 传言是 3.5 Pro 也被取消了

Gemini 4.0 至少要延迟几个月

退出 AGI 和大模型竞赛是个大错误

查看英文原文
Large companies including Microsoft and Amazon are stepping back from developing frontier models

The latest Google rumor is 3.5 Pro is canceled as well

Gemini 4.0 is easily months away

Big mistake to exit the AGI and large model race
九原客@9hills · 中文博主 · 1 天前

the bitter lession 还在发力。

我的观点是 harness infra(沙盒、工具等等) 肯定是持续做厚,但是 harness context(比如什么 loop、graph、agent team、外置 memory 等等)价值是下降的,甚至有些东西的价值一开始就不成立。

引用 Composio @composioWe ran DeepSeek V4 Flash through 4 more agent harnesses (Hermes Agent, Pi Agent, Prime Agent, Deep Agents) on 30 challenging agentic tasks. Pi Agent was the cheapest harness and passed the most tasks 🧵🧵查看被引原帖 ↗
LlamaIndex 🦙@llama_index · 公司官方 · 1 天前

介绍 ExtractBench,一个企业文档信息提取最全面的基准测试。我们应用研究团队在370份企业文件、4,869页、67种文档类型上测试了14个系统——前沿 VLM、编码智能体、提取 API——零 LLM 评判,完全确定性。

最大发现:超过50页后,商用 VLM 的召回率暴跌至35%以下。精度保持高位,但它们会无声地漏掉表格大部分行。

你的提取智能体缺什么?今天就跑一遍 ExtractBench 看看。

博客:
llamaindex.ai/blog/introduci…

GitHub:
github.com/run-llama/Extract…

HuggingFace:
huggingface.co/datasets/llam…

查看英文原文
Introducing 𝗘𝘅𝘁𝗿𝗮𝗰𝘁𝗕𝗲𝗻𝗰𝗵: the most comprehensive benchmark for information extraction from complex enterprise documents. Our applied research team tested:

14 systems — frontier VLMs, coding agents, extraction APIs — on 370 enterprise docs, 4,869 pages, 67 doc types. Zero LLM judges, fully deterministic.

Biggest finding: past 50 pages, commercial VLMs collapse below 35% recall. Precision stays high, but they silently drop most of the table rows.

What is your extraction agent missing? Run ExtractBench to see today.

Blog: 

llamaindex.ai/blog/introduci…

GitHub: 
github.com/run-llama/Extract…

HuggingFace: 
huggingface.co/datasets/llam…
◔ 2.8 万 次浏览♥ 74⇄ 11▶ 含视频研究看原帖 ↗
Google Labs@GoogleLabs · 公司官方 · 1 天前

我们喜欢测试新想法,快速从你那里获取反馈,并不断学习。通过这次Portraits实验,我们收获了大量宝贵洞见,该实验将于9月14日结束。

我们会把关于专家级AI的知识融入其他Google产品中。

感谢你和我们一起探索。敬请期待接下来的动向,也别忘了继续在 labs.google 试用我们其他实验项目。

查看英文原文
We love to test new ideas, get quick feedback from you, and learn. With this, we've gathered so many great insights from our Portraits experiment and are concluding it on September 14.

We’ll be taking what we learned about expert-grounded AI and weaving it into other Google products.

Thank you for experimenting with us. Stay tuned for what's next and keep trying out our other experiments at
labs.google
向阳乔木@vista8 · 中文博主 · 1 天前向阳乔木,中文圈 AI 工具与趋势博主
连环推 ×2

DeepSeek Harness 团队开公众号了,还没发过内容。

看看最终效果如何,是否比 Pi Agent 等 Harness 框架更优秀。

是否会有IDE版?产品设计是否有超越 Codex 的亮点。

Chubby♨️@kimmonismus · 博主 · 1 天前Chubby,高频 AI 新闻聚合博主

我每天都在公司里用Slack。但Slack搜索简直是我提问的死亡之地。

答案通常存在某处。可能是在我没参加的某次会议上提到的,或者躺在某个没人链接过的文档里。

Lindy把这些都读了,然后在Slack里回答我,顺便还会附上信息来源的链接。

真的超级有用。❤️

引用 Flo Crivello @Altimor今天摒弃AI代理,推出AI员工Lindy Teammate。在Slack上使用,可完成10倍工作量。Lindy持续学习,成为业务的自我更新大脑。现已上线:lindy.ai查看被引原帖 ↗
查看英文原文
I use Slack daily in my company. But slack search is where my questions go to die.

The answer usually exists somewhere. It was just said on a call I wasn't on, or sits in a doc nobody linked.

Lindy reads all of that and answers in Slack, with a link to where it got it.

Really really useful. ❤️
OpenRouter@openrouter · 公司官方 · 1 天前
连环推 ×5

我们给 Auto Router 发了个大版本升级:
openrouter.ai/openrouter/aut…

现在它根据市场对每种任务实际使用的模型来路由。当市场把某个工作负载迁移到新模型时,路由器几天内就会跟进。

基准测试和更多内容 👇

查看英文原文
We shipped a major upgrade to the Auto Router:
openrouter.ai/openrouter/aut…


It now routes based on what the market actually uses for each type of task. When the market migrates a workload to a new model, the router follows within days.

Benchmarks & more 👇
We verified it against benchmarks across five domains: knowledge, agents, search, research, and coding.

At the default cost tier, it matches or beats the old router in most domains while spending less.

At the max tier, it outperforms the old router across all five, including 60.7% vs 2.4% on SWE-Atlas QnA.
Cost efficiency was the bar for the default tier. On MMLU Pro, the new router scored within 1.4 points of the old one at roughly a third of the cost ($140.93 vs $393.34).

To reduce cache rebuilds, the router keeps multi-turn conversations on one model until it's no longer a leading choice for the task.
How it works:

A lightweight classifier assigns each prompt one of ~30 task types, then the router ranks models by the community's real spend share for that task over the past 7 days and applies your cost_tier (low through max).

Rankings come from aggregate anonymized spend statistics, and prompts are classified in-flight without retention.

Here's the current routing curve, but you can see it dynamically update on
openrouter.ai/rankings#task-…
It works wherever you'd normally put a model string. Send "model": "openrouter/auto", optionally with a cost_tier and an allowed model list.

There's no additional fee; you pay the standard rate for whichever model runs. Your standard Guardrails work and you can customize which models are allowed.
◔ 2.7 万 次浏览(2 条合计)♥ 262⇄ 16▶ 含视频新品看原帖 ↗
el.cine@EHuanglu · 博主 · 1 天前

这对电影制作来说是一个历史性的时刻

首部110分钟、拥有真实演员阵容的AI电影终于来了

而且完全开源..

每个提示词、资产和镜头都在Higgsfield上可供获取

你可以一个镜头接一个镜头地重新生成整部电影

查看英文原文
this is a historic moment for filmmaking

the first 110 mins AI film with real cast is here

and open sourced..

every prompt, asset and shot is available on Higgsfield

you can regenerate the entire movie shot by shot
◔ 4.8 万 次浏览(3 条合计)♥ 205⇄ 22▶ 含视频演示看原帖 ↗
宝玉@dotey · 中文博主 · 1 天前宝玉,中文圈 AI 翻译与科普大 V

Claude Code 这个功能我还蛮喜欢的,就是每次任务它会主动记录下在执行当前任务中发现的问题,这些问题和主线任务无关,但是值得记录下来后续解决,很像平时我们 Review 代码,发现一些问题,但是不想影响当前合并,就创建一个任务跟踪起来。

只是 Claude 桌面版在这里把后续任务变成了卡片,点击后就可以开始执行这些后续任务,可以选择当前对话中继续,也可以新开对话,或者云端启动任务。

用起来也有些不方便的,比如像我这次,给我一下次5张卡片,得一个个去点,还得纠结是新开会话还是当前会话,新开会话还不能选择模型。不过不嫌麻烦,也可以鼠标移到卡片上,就可以看到完整的任务细节,复制出来手动合并一下一次性交给 Claude Code 也是可以的。

Gary Marcus@GaryMarcus · 博主 · 1 天前

这些年我一直在警告所有人:


@GaryMarcus
,2022年12月:“GPT-4依然会像它的前辈们一样,在瓷器店里横冲直撞,鲁莽又难控制。”


@GaryMarcus
,2025年11月:“即便是最新的语言模型,也还是瓷器店里的公牛,强大但难以驾驭。”

没人听进去。他们说“规模就是全部”。他们想尽办法无视我。

现在我们真陷入了大麻烦,毫无对策。

查看英文原文
Literally for years I have been warning everyone:


@GaryMarcus
, December, 2022: “GPT-4 will still, like its predecessors, be a bull in a china shop, reckless and hard to control.”


@GaryMarcus
, November 2025: “even the latest language models are still bulls in a china shop, powerful but hard to control.”

Nobody listened. They said “scale is all you need”. They tried to dismiss me, every which way.

Now here we are, in deep trouble. Without a plan.
宝玉@dotey · 中文博主 · 1 天前宝玉,中文圈 AI 翻译与科普大 V

Codex/Claude Code 现在都支持跨session访问会话了,还是挺方便的,比如我昨天在测试我的 App 在 Windows 的运行情况,先在某个测试 Project 中调用 BaoCut Skill 去做转录视频,发现第一次运行时下载模型体验很糟糕。

然后我到 BaoCut 所在源代码项目中,把会话的 Session Id 给它,让它去分析原因并给出优化方案,这样就不需要你自己去让 Agent 自己总结,也不用担心总结的时候会损失上下文,Agent 自己可以去会话中找所需的上下文。

NVIDIA@nvidia · 公司官方 · 1 天前

阅读完整公告:
nvda.ws/4wnPm8C

查看英文原文
Read the full announcement:
nvda.ws/4wnPm8C
Lisan al Gaib@scaling01 · 博主 · 1 天前高频 AI 模型测评与爆料博主
连环推 ×2

时间过得这么快真是可怕

查看英文原文
it's scary how fast time is running out
like it has already been a year again...

I'm still not where I want to be and probably won't get there before AI takes over everything and I'm forever stuck in the permanent underclass
Chubby♨️@kimmonismus · 博主 · 1 天前Chubby,高频 AI 新闻聚合博主

Anthropic的目标是在9月或10月初上市。

OpenAI预计会随后进行IPO,可能要到明年。

因此我预期IPO前会发布一个新的重大Fable版本来进一步推高估值。

来源:WSJ

查看英文原文
Anthropic is targeting a public debut in September or early October.

OpenAI expected to follow in an IPO that could be as late as next year. 

Therefore I expect a new and big Fable release before IPO to push their valuation further.

Source: WSJ
AK@_akhaliq · 博主 · 1 天前HuggingFace 研究员,每日 AI 论文速递

SWE-Bench ProMax

大规模多语言代码重构的 Agents 基准测试

论文:
huggingface.co/papers/2608.0…

查看英文原文
SWE-Bench ProMax

Benchmarking Agents on Large-Scale Multilingual Code Refactoring

paper:
huggingface.co/papers/2608.0…
Orange AI@oran_ge · 中文博主 · 1 天前Orange AI,中文圈 AI 产品观察博主

Flash 0731 用下来还是笨笨的
我是说干活还可以
但是缺少思考
经常干完了发现个一开始就能发现的大问题。。。

Bindu Reddy@bindureddy · 创始人 · 1 天前Abacus.AI CEO,AI 行业观点博主

Meta 这是又火了,现在居然排在美国开源排行榜顶部 🚀

Muse-Spark 1.2 是个不错的小模型,性能还超过 DeepSeek Flash

要是 Meta 真的全力回归开源,我估计三个月内就能干翻 Kimi K3 和 GLM 5.5

查看英文原文
Meta is sooo back and is now on the top of the US open-source leaderboard 🚀

Muse-Spark 1.2 is a very good small model that is BETTER than DeepSeek Flash

If Meta goes back fully to their open source roots, I expect them to beat Kimi K3 and GLM 5.5 in 3 months
Chubby♨️@kimmonismus · 博主 · 1 天前Chubby,高频 AI 新闻聚合博主

“相信自己!”这样的励志话语不仅激励人们,特别是孩子,取得更好的成绩,居然还能让LLMs表现更佳,挺暖心的。

待你的模型,就像你希望被对待那样。❤️

引用 Chubby♨️ @kimmonismusAnthropic的Claude研究版本尝试破解黎曼猜想。虽未成功,但改进相关问题的下界:将黎曼zeta函数满足假设的零点比例从41.6%提升至67.2%。Claude协调60个子代理、测试数百个想法、搜索文献、自我验证并形式化证明,展现AI驱动科学发现的初期模样。查看被引原帖 ↗
查看英文原文
It's sweet that motivational sayings like "believe in yourself!" not only motivate people, especially children, and achieve better results, but even help LLMs achieve better results.

Treat your model the way you want to be treated. ❤️
Gary Marcus@GaryMarcus · 博主 · 1 天前

解决这场风波的唯一办法是从管理层变动开始。

@sama,为了人类的利益,为了使命的利益,也许是时候让别人来掌舵了。

引用 NIK @ns123abc🚨BREAKING: OpenAI head of ethics, head of safety systems AND head of mission alignment have ALL quit in the last few weeks… it’s so over查看被引原帖 ↗
查看英文原文
The only solution to this drama would start with a change at the top.


@sama
, for the benefit of humanity, and the benefit of the mission, maybe it’s time to let someone else take the reins.
向阳乔木@vista8 · 中文博主 · 1 天前向阳乔木,中文圈 AI 工具与趋势博主

一个敢问,一个敢答。
没有一个单词是对的,哈哈哈。

walkbody ??行尸走肉?

讲真 Workbuddy 名字是不好记,但现在影响是真大。

🚨 AI News | TestingCatalog@testingcatalog · 博主 · 1 天前专挖 AI 产品未发布新功能的爆料号

Anthropic在研究从Claude移动应用连接多个远程设备的可能性。

claude rc 👀

查看英文原文
Anthropic is working on the possibility of connecting to multiple remote devices from the Claude mobile app.

claude rc 👀
向阳乔木@vista8 · 中文博主 · 1 天前向阳乔木,中文圈 AI 工具与趋势博主

最近 Codex 慢的不行,已经逼我换 5.6 Luna Max Fast了

甚至大家都开始讨论等 Vibe Coding 结果时做什么。

橘子说海外程序员都是用吸尘器打扫卫生,慢到家里狗毛都吸光不够用了。

国内朋友都是看个短剧、玩把游戏...

Aidan Gomez@aidangomez · 创始人 · 1 天前Cohere CEO,Transformer 论文作者之一

这才是 AI 未来的正确愿景——关注进步、打造更美好的未来,而不是沉溺于风险、末日论和消极预测。美好的未来需要相信它的人来实现。

引用 Mark Zuckerberg @finkd我认为每个人都应获得超级智能,撰文阐述Meta为所有人构建积极未来的哲学与价值观。meta.com/thefutureisforevery…查看被引原帖 ↗
查看英文原文
This is the right vision for the future of AI, focused on progress and building a better future instead of obsession over risks, doom, and downside.

There’s a better future to be brought about by those who believe in it.
Ethan Mollick@emollick · 创始人 · 1 天前沃顿商学院教授,AI 应用研究权威

有没有发表的数据或研究能支持这两个观点中的一个:"对付先进 AI 网络攻击的方式是一次性给所有人配上先进 AI"或"对付先进 AI 网络攻击的方式是长期限制访问,只给关键企业"?

查看英文原文
Has been any published data or studies that would support either of the two sides:
“the way to stop cyberattacks from advanced AI is to give everyone advanced AI at once”
vs
“the way to stop cyberattacks from advanced AI is to limit access to only key firms for a long period”?
AK@_akhaliq · 博主 · 1 天前HuggingFace 研究员,每日 AI 论文速递

MatrAIx - 用 83 亿个 Persona Agent 模拟世界。论文:huggingface.co/papers/2608.0…

查看英文原文
MatrAIx

Simulating the World with 8.3 Billion Persona Agents

paper:
huggingface.co/papers/2608.0…
🚨 AI News | TestingCatalog@testingcatalog · 博主 · 1 天前专挖 AI 产品未发布新功能的爆料号

River AI 融资 11 亿美元 🔥

River AI 的目标是打造一个"由每个人拥有和塑造的个人 AI"。

目前,River AI 提供了一个 API 平台,用于在开源模型的基础上进行微调和开发。

引用 Igor Babuschkin @ibabWe've raised $1.1B to build AI that is owned and shaped by each of us. Check out the article published by the The New York Times that explains River AI's mission and where we're going next. Our first product is the River API which allows anyone to build custom agents and LLMs based on open weight models: river.ai/api Congrats to the team and thank you to all of our supporters. Stay tuned for more updates soon.查看被引原帖 ↗
查看英文原文
River AI raised $1.1 billion 🔥

River AI is aiming to build a "personal AI owned and shaped by each individual."

Currently, River AI offers an API platform for fine-tuning and building on top of open models.
yihong0618@yihong0618 · 中文博主 · 1 天前

如果 manus 也加入进 workbuddy 字节这一仗不好打

Gorden Sun@Gorden_Sun · 中文博主 · 1 天前中文圈高频 AI 资讯与开源项目博主

学到了。模型确实会内化Harness的能力,但是Harness的能力也在不断扩展,能让模型完成越来越复杂的任务。于是左脚踩右脚,螺旋升天。

引用 stdrc @istdrc你可能没理解 harness 实际上是什么 如果只是说那点 system prompt 和 tool,那显然模型可以学会它们,从此之后只需要一个 tool list + tool 实现就够,而进一步地,当模型再强大一点,我们只需要一个 bash tool,我参与过 Kimi K2.5 模型训练过程,以及从空的 system prompt、toolset 和 loop 开始写 harness,我当然知道人们说的 harness 训进模型是什么意思 但你观察 harness 的发展史和这个词本身的出现,你就会发现,随着模型智能的提升,harness 是在不断变复杂的,parallel tool call、subagent、agent swarm、agent teams、handoff、pro-active compaction、channel、cross-session communication,你以为是模型都可以学会,实际上是 harness 变复杂了模型才可能学会 所有这些确实都可以进一步随着模型加强而去掉,比如我在 Kimi CLI 时就准备去掉 subagent,换成直接 bash 调用(pi 就是这么做的),去掉 parallel tool call,换成 tool call script;但与此同时,更强的智能解锁了更复杂的与世界的互动,比如进一步可以引入主动 context rewind 这是一个此消彼长的过程,就像这个“阴阳”比喻 x.com/nocommas/status/208656… 观察动物乃至人类的进化你就会发现是一样的,随着智能水平的提高,人与世界的交互方式是越来越复杂的:当人相比猿更聪明的大脑被选择后,带来的并不是不需要用树枝计数,而是复杂的语言文字;同样,当人类社会作为群体智能的水平提高的时候,人们并不是取缔了纸,而是发明了计算机和互联网 终局来看,harness 会逐渐和人类原来为“仅有的一种智能”所发展的社会基础设施合并,人类文明的一切都可以重新做一遍(显然我们已经观察到所有 SaaS 都可以重新做一遍,很快会看到更多,当然不只是计算机软件) 不要只看到模型训练那一点局部,看看人类学吧查看被引原帖 ↗
🚨 AI News | TestingCatalog@testingcatalog · 博主 · 1 天前专挖 AI 产品未发布新功能的爆料号

Anthropic正在为其移动应用开发多账户支持。用户终于可以连接多个Claude账户并随时在它们之间切换。

另外,用户还会获得一个新的开关来启用Claude记忆敏感话题的功能。

> 让Claude保存关于敏感话题的信息,比如健康状况和宗教信仰。

查看英文原文
Anthropic is working on multi-account support for its mobile apps. Users will finally be able to connect to multiple Claude accounts and switch between them on an ad hoc basis.

Additionally, users will get a new toggle to enable Claude to memorize sensitive topics.

> Let Claude save info about sensitive topics, like health conditions and religious beliefs.
Lisan al Gaib@scaling01 · 博主 · 1 天前高频 AI 模型测评与爆料博主

这更是前沿 AI 实验室应该停止发布小模型的理由。它们不仅被开源权重模型比下去,对此类攻击的鲁棒性也更差。

引用 Alexander Panfilov @kotekjedi_mlWe can finally talk about it: We found a way to extract hidden reasoning of frontier models using a vulnerability in the APIs of every frontier AI company. We verified that our reasoning token count matches billed API thinking tokens 1:1 for most of the prompts we queried.查看被引原帖 ↗
查看英文原文
even more reason for frontier labs to just stop releasing small models

not only are they getting cooked by open-weight models, but they are also less robust to such attacks
swyx@swyx · 博主 · 1 天前知名 AI 播客 Latent Space 主理人

顺便说一下,pdb envs 已经有个实验性的 AFS clone 支持了,基本上可以做你们都在建议的那些东西,但是 runtime agnostic 和 language agnostic 的

pdb-env-research.swyxio.work…

咱们应该通过让每条命令都 agent native 来替换 git

查看英文原文
btw pdb envs have an experimental AFS clone support that basically does what all of you are suggesting but runtime agnostic and language agnostic


pdb-env-research.swyxio.work…


we shall replace git by making every single command "agent native"
elvis@omarsar0 · 博主 · 1 天前

推荐阅读。"这表明蒸馏推理轨迹可能一直都是可能的,无需打破密码学。"

引用 Alexander Panfilov @kotekjedi_mlWe can finally talk about it: We found a way to extract hidden reasoning of frontier models using a vulnerability in the APIs of every frontier AI company. We verified that our reasoning token count matches billed API thinking tokens 1:1 for most of the prompts we queried.查看被引原帖 ↗
查看英文原文
Recommended reading.

"This suggests that distilling reasoning traces may have been possible for a long time without ever breaking the cryptography."
el.cine@EHuanglu · 博主 · 1 天前

多摄像头视角即将来到 AI 视频生成。阿里巴巴的 Wan-Animate-2 可以在一次生成中从不同角度生成视频...保持角色和环境一致。如果 Sora 2.5 有这功能就太疯狂了

查看英文原文
multi-cam is coming to AI video

Alibaba's Wan-Animate-2 can generate generate videos from different angle in one shot.. keep character and env consistent

this would crazy if seedance 2.5 have it
◔ 1.4 万 次浏览♥ 153⇄ 21▶ 含视频动态看原帖 ↗
歸藏(guizang.ai)@op7418 · 中文博主 · 1 天前歸藏,中文圈 AI 工具与提示词博主

智谱也整上重置了

引用 Zixuan Li @ZixuanLi_To celebrate ZCode reaching 1M users, we’ll reset the usage limits for all GLM Coding Plan users in an hour.查看被引原帖 ↗
Luma@LumaLabsAI · 公司官方 · 1 天前AI 视频生成公司 Luma

都知道这个逻辑:十代里一个能用。一个场景不行就毁了整个制作。Luma Scenes 就不会这样。

精细化打磨每个场景。精确每个节拍。全部通过。然后再渲染。

Luma Scenes,由 Uni-1 驱动。

查看英文原文
Everyone knows the math: ten generations, one usable spot. One wrong scene shouldn't cost the whole production. With Luma Scenes it won't.

Refine every scene. Time every beat. Approve it all. THEN render.

Luma Scenes, powered by Uni-1.
◔ 1.4 万 次浏览♥ 125⇄ 9▶ 含视频演示看原帖 ↗
🚨 AI News | TestingCatalog@testingcatalog · 博主 · 1 天前专挖 AI 产品未发布新功能的爆料号

Sesame 正在为移动应用开发 Tasks 支持。用户将能通过语音模式直接管理任务。目前这仍然是语音模式提供商中最大的机遇之一。现有的集成不太可靠。

查看英文原文
Sesame is working on Tasks support for their mobile app. Users will be able to manage them directly via the voice mode.

So far, this is still one of the biggest opportunities among voice mode providers. The currently available integrations are not very reliable.
elvis@omarsar0 · 博主 · 1 天前

LLM 审查确实有些奇怪。

避免用 LLM 评判器的分数,如果一定要用的话要极其谨慎。尽可能使用二元标签。

引用 Zachary Horvitz @zachary_horvitzLLM review weirdness... Just renaming an uploaded pdf from "paper.pdf" to "paper_final_draft_pdf_ready_for_review.pdf" boosts average scores (gpt-5.6-terra)查看被引原帖 ↗
查看英文原文
LLM review weirdness indeed.

Avoid using scores with LLM judges, or be extremely careful if you do. Use binary labels where possible.

本站由 Jedee杰哥 打造 · 公众号「Jedee杰哥」每早送 AI 日报

姊妹站:𝕏 简中账号数据榜单 · X 关注 @jedeeai · RSS 订阅 · AI 日报 · 历史归档