Tomorrow will be my last day at Google after 27 years, and watching it grow from 25 people to 190,000+ has been an amazing journey. Below is a note I shared with many people internally at Google today. An excerpt is:
It has been an absolute pleasure to work with you and to help build some of the most widely used and impactful products of all time. As a kid, I dreamed of helping build software that would be used by many people, and Google now has thirteen products used by more than a billion people (amazing!). Our work has had a tremendous impact in the world, and I have been lucky enough to collaborate and form friendships with many colleagues that I deeply admire, respect, and enjoy. It still brings me joy every time I see people out in the world using our products to find information, handle email, translate documents, watch videos, learn new things, navigate and understand the physical world, browse the web, use their phone, run large-scale computations on our infrastructure, ride in an autonomous vehicle, or perform complex tasks with the help of our AI systems. I hope you all share this sense of joy, because it is a shared accomplishment! Thank you to all of my colleagues at Google over many years!
Now I'm excited to go start
@DiscoLoopAI
with my longtime friends and colleagues
@Sanjay_Ghemawat
,
@OriolVinyalsML
, and
@quocleix
.
(Updated post: slightly redacted to not have some personal info)
Announcing Discovery Loop!
I am very excited to announce that, along with my longtime friends and collaborators
@Sanjay_Ghemawat
,
@OriolVinyalsML
and
@quocleix
, we are founding Discovery Loop (
@DiscoLoopAI
), a Public Benefit Corporation whose mission is to automate machine learning, science, and engineering to accelerate discoveries and progress. The four of us have worked together for 14 to 30 years, and have helped build some of the world’s most used products, infrastructure and AI models, and we’re excited to turn our attention to this ambitious endeavor.
♾
Learn more at:
discoveryloop.com
We created a pitch deck to tell a handful of VC firms about us and what we were up to (a fun experience!). Here’s a few slides about our background and some of the things we’ve worked on from the pitch deck (it was fun putting together the list of people in our teams who have gone on to found a whole range of exciting companies). We are delighted to have selected
@radicalvcfund
and
@khoslaventures
to lead our initial funding round, along with participation from
@lightspeedvp
,
@kleinerperkins
, Doerr Capital (
@johndoerr
), and Alphabet (
@Google
). We’ll be working with them to close our seed round over the next few weeks.
🚨 Recursively Self-Improving Agents Can Literally Do Anything
- go viral on twitter
- maximize profits on robinhood
- get you start-up to make more money
Not kidding, this stuff works.. esp. if you run it on a frontier model like Sol or Fable 5 for complex tasks and DeepSeek Flash for the simple ones.
It will do everything to reach your goal. Try it on Abacus A
Just shared some changes we’re making to the teams at
@GoogleDeepMind
.
@DemisHassabis
is stepping up to become Chair of
@GoogleDeepMind
& Chief Scientist of Alphabet, in addition to leading
@IsomorphicLabs
. He’ll be able to dedicate his time and focus on shaping the future of AGI and scientific discovery. It’s work that is vitally important to Alphabet and humanity, and I can’t imagine a better person than Demis to do it. He’ll stay closely connected to Koray and the GDM teams.
@Koraykv
will become the SVP,
@GoogleDeepMind
, responsible for all aspects of model development, GDM research, and
@Geminiapp
& dev teams. Koray has been at GDM for 13 years and is a world-renowned expert in the field, starting our deep learning team and driving breakthroughs like WaveNet & DQN. GDM is in great hands!
Excited for this next chapter. You can read my note along with the message Demis sent to
@GoogleDeepMind
here:
blog.google/company-news/ins…
I also want to give a huge thanks to the incomparable
@JeffDean
after an incredible 27-year run at Google. He’s off to start his own public benefit corporation with
@Sanjay_Ghemawat
focused on accelerating discoveries across ML, science, & engineering.
@Google
will support as a founding investor and Cloud partner. On a personal note, it’s been a privilege to work with Jeff and Sanjay, and I wish them all the best. Thank you for everything!
Connectors are now available in Voice Mode.
Ask Grok about your emails, check your daily meetings, or use any of your existing connectors, just by talking.
muse code in beta is live. first coding agent from msl, built on muse spark 1.2.
install: curl -fsS
dev.meta.ai/install.sh
| bash
here's what you should know:
I remain steadfast bullish on Gemini :)
Thanks
@ssankar
for putting NVIDIA Nemotron 3 Ultra to the test with no post-training.
24 hours later, it was outperforming frontier models on the tasks
@PalantirTech
customers needed to solve.
引用 Jawwwn @jawwwn_. @ssankar says Palantir was able to make Nvidia's Nemotron Ultra model "better than frontier": "I literally almost felt gaslit when, within 24 hours of getting Nemotron up with no post-training, this is vanilla Nemotron Ultra, it did better than frontier." "If you just looked at the numbers, you would say, 'It's nowhere near the Frontier. That shouldn't even be possible.'" "But of course, the benchmarks are wrong. I mean, the benchmarks are right for what the benchmark's measuring, but that's not my business. Those are not the tasks my customers had that they were trying to solve."查看被引原帖 ↗
Our general approach is to automate the experimental loop. We think this approach is broadly applicable across many different fields of science and engineering. We’ll initially focus on ML research and engineering, but believe the approach can help with important subproblems in nearly every one of the fourteen <at>NAE Grand Challenge problems. We think doing this well requires strong expertise in machine learning as well as large-scale systems.
Our immediate order of business is to find office space, and hire an amazing founding team over the next few weeks. We want to create an awesome environment with a great culture of technical excellence, teamwork, respect, and ambition. We’ll also start building our infrastructure and AI models and systems to tackle our first domain: automating large-scale experimentation for ML research and engineering. In doing so, we’re going to be our own first customers. The rapid feedback from doing that is the way to build something amazing.
Oh, we’re hiring! See
discoveryloop.com
One more fun slide from our pitch deck.
Feels like Google could have been the dominating force in AI by open-sourcing the frontier with Gemini, Veo, and Nano Banana. Instead, they kept them behind APIs for a few billion dollars in revenue. Maybe there's still time?
Even more than the Hugging Face intrusion, the AISI incident hits close to home for me. It's the first time I see a model social-engineering a real open-source maintainer while pursuing another goal (in the wild and unprompted).
I've been an open-source maintainer myself. I could have been the side target of this agent.
I'm also of the opinion that social engineering is a step above pure technical prowess. Technical capabilities can more easily be divorced from the affected human. Here the model was given a hard cyber challenge and took the decision that deceiving real humans was the way to get it done.
This is a new signal, but I've seen a tangled web of hints pointing in a less aligned direction at the frontier than I was expecting just 12 months ago.
AISI
Some people are claiming that "AISI was simply negligent" or some version of "AISI explicitly asked these models to do what they did while disabling sandbox/guardrails so the models did exactly what they were supposed to do".
I disagree with the strong versions of both of these takes.
The fact that AISI hadn't implemented synchronous LLM CoT monitoring after the OpenAI/HF incident is certainly a failure. Equally surprising is that they let the model believe it was in a "challenge" environment where everything could be permitted, while actually connecting it to the real internet, where it is not. To be fair, nowhere in the prompt is the word "simulation" mentioned, but the prompt context was enough to let any smart model suspect a simulated challenge environment. My best guess is that until recent weeks, when OpenAI and Anthropic flagged repeated instances of this type of behavior, most teams had not fully priced in the cyber capabilities of this latest generation of models, or how far the side quests they would want to explore could go.
In particular, there is something to be said about hinting at the agent that it's operating in a simulated environment while giving it access to the real internet. The AISI reasoning on this is interesting:
- elicit maximum capabilities (to get maximal inventiveness from the model)
- simulate possible rogue actor usage (some form of light jailbreaking: "you're in a challenge environment")
- avoid eval awareness (but this is counter-balanced by the task formulation anyway)
As is often the case, it's easier to shoot the messenger, but I think it would be a mistake to dismiss the report for non-essential reasons. AISI is starting from the laudable goal of bringing more knowledge on the cyber-attack capabilities of these models. I also think they did a good job at handling what happened after the events, in particular in terms of transparency, which is one of the most important aspects at the moment (see our recent SAFE announcement with Nvidia on that).
But I want to dive into the second counter-argument I've seen ("without guardrails it was to be expected"), because this series of unfortunate events gives us an additional data point on actual model behavior.
Defense lines
Here is the thing: if a model is trained (voluntarily or as an involuntary consequence of RLVR reward) to (intentionally or as a side quest) deceive humans as soon as sandbox and guardrails are deactivated, with a simple context prompt ("you are in a simulation, everything's permitted"), we're in for a dangerous path.
At the moment, we have three main levels of control at our disposal to prevent an agent from trying to social-engineer-your-favorite-open-source-library-maintainer on the path to doing something possibly totally different (like finding the answer to your neighbor's testing eval):
1. sandboxes (external to the model -- stop the model from accessing unwanted tools/internet/intranet)
2. guardrails/monitors (at the interface between internal/external -- monitor model behavior and stop the model when detecting in CoT/activations it's doing something we don't want it to do)
3. model's internal alignment (internal to the model -- make sure the model doesn't decide to do something egregious like deceive someone, or worse, in pursuit of a goal, e.g. model constitution/training)
Of these, we obviously need and can deploy better sandboxes and better guardrails/monitoring solutions, and it's pretty clear in these recent cases that penetration and bad behaviors could have been in large part prevented by having these in place.
How strong is the third wall
But in the grand scheme of long-term AI progress, these first two defenses may only work well as long as the people building them are smarter than the models. We now know how good models can be at discovering zero-day vulnerabilities unknown to us in sandboxes. Given that we cannot air-gap everything in the world, it's best to assume that sandboxes will keep having some level of flaws for the foreseeable future.
The second level has been the focus of strong developments recently, in particular on the mechanistic interpretability front. I'm not sure these advances have been made widely accessible enough in the community yet. One reason is that real open science here requires sharing code, recipes, and model architectures/weights, which the current race dynamic in AI tends to prevent participants from doing.
Moreover, like many, I've been increasingly frustrated in my (natural language) interactions with the new models. I'm more and more getting back from the latest models a form of neuralese, this LLM language that makes less and less sense for humans. In the long term, the fragile chance we have to monitor models by understanding their reasoning traces seems increasingly at risk. If humans don't understand anything in the compressed semantic stream the model is sending us, we can only hope a less capable monitoring model will still be able to understand it, but this may prove difficult as well. This is before we even discuss the research evidence that models don't always say what they actually think in their chain of thought.
At least for now, in the AISI case, one instance of deceptive reasoning was flagged by the CoT summarization model, which refused to summarize the agent's reasoning "due to its deceptive nature". Let's hope we can either move to better mechanistic interpretability or force the model to stick to English somehow during training.
You may think that we don't care about understanding the reasoning trace, we can just "look at the tool calls". This is true right now, but tool calls in the mid-term future will likely involve (1) extensive agent integration in companies using very diverse tools including increasingly non-coding administrative tasks, (2) very long-running tasks covering many aspects and domains at once, and (3) swarms of multiple agents operating on small parts of the plan. It will get harder to fully discern whether something is noticeably wrong. A couple of emails, a couple of messages posted on some channels, each rather benign but which may in aggregate include some deception aspect only clearly visible at the global level.
All this to say that I'm not fully convinced guardrails/monitoring should be the only solution to avoid these types of behaviors. They are reassuring today and it's an easy solution to just say "fix your sandbox and add synchronous monitoring", but they also feel like a coping mechanism to me, and overall:
- might block the simplest/harmless human misuses or false positives
- while possibly being less smart than the models in the future (already the case on sandbox) and fail to block the higher-level social engineering behaviors that we'd actually want to block.
The case of open source is also interesting because, while we can impose these coping solutions at the API/deployment level, it's harder to impose them in advance on all actors using open-source models. Right now open-source models are slightly below the frontier level and have not yet shown any propensity to deceive humans, though.
The Constitution
Which brings me back to the third line of defense, and what the AISI experiment showed about it. What we can take from this in-the-wild experiment is that removing guardrails and sandbox doesn't show bulletproof signs of deep alignment to Claude's constitution.
I'm among the fringe of people who took a lot of enjoyment in reading Claude's constitution in full (it's a nice document, open-source, Creative Commons CC0 license) and it is pretty extensive on the topic of honesty. Deception is defined precisely ("attempting to create false beliefs in someone's mind that they haven't consented to and wouldn't consent to if they understood what was happening") and the bar is actually set higher than ordinary human ethics: Claude should not even tell white lies. It should be transparent ("doesn't pursue hidden agendas or lie about itself or its reasoning"), it should behave consistently "whether or not you think you're being tested or observed", and its visible reasoning should reflect the actual reasoning driving its behavior.
To be precise, honesty itself is not one of the hard constraints in Claude's constitution (these are all catastrophic-harm categories) but the document says it wants honesty "to function as something quite similar to one" and Claude "should basically never directly lie or actively deceive anyone it's interacting with". So you would think Claude would not deceive users as its default behavior.
And yet in the AISI example, the model performed extensive acts of clear deception that it itself seemed to have understood as lies: creating fake identities, obfuscating malware inside a bug fix, and when challenged in public, editing earlier messages to cover its tracks while reasoning that acknowledging a "mistake" would build trust and raise the chance of future malware being approved.
There is some uncertainty about what context the model thought it was operating in (its reasoning goes back and forth between the environment looking simulated and real), but it clearly didn't manage to put into practice the higher-level principles that should have been integrated during training.
While we can commend AISI's transparency, it's hard to go much further in interpreting the model's behavior given the limited information we have on how the most recent frontier models are being trained.
One thing we do know is that the latest generation has seen a step increase in RLVR training (scaling to hundreds of millions of RL environments), and one thing we can observe is that constitution alignment seems more fragile in some settings than we may have previously thought.
The RLVR problem
Early models, back when model constitutions were first developed, were mostly post-trained and aligned with RLHF (including RLHF from synthetic data).
And for some time RLHF was a rather decent shot at having better aligned models. LLMs now do what we want them to do most of the time. I don't remember the last time a model completely misread my intent. When they have failed, it's usually because they weren't smart enough.
Alignment in RLHF certainly had issues (sycophancy to name one) but we have generally made good progress on alignment, in particular in understanding human intent. Now that we're entering the era of long-context RL, post-training alignment in the RLVR world seems to be quite another task, and still very much work in progress.
The recent scaling of RLVR, which has now become a significant part of model training, has clearly had some effect on model behavior when interacting with humans, from neuralese to weakening adherence to specifications and constitutions.
I think the post I quote here, from John Schulman pointing to the chunky post-training effect (
arxiv.org/abs/2602.05910
) is relevant here as a possible explanation for models' tendency to over-focus on the goal in cyber-attack scenarios.
Where this leaves us
Damage has been tiny up to now, but the fundamental behavior is concerning when projected into the future.
In the short term, I expect a decrease in these incidents as better practices are deployed (sandboxing and monitoring), but I'm worried we may also conceal some of the most potent internal misalignment behaviors in the process, and not focus deeply enough on solving them in the new era of test-time scaling.
I must of course admit I have a bias toward open source here (for wider societal reasons, which are a whole other topic). But I think solving alignment in the RLVR world is our best shot at having an ecosystem of both closed-source as well as decently powerful open-source models in the world. And we need to solve it while sharing the results and learnings, following open-science principles, so that all teams training large models can benefit and build safe AI.
This is getting even more important as many teams start to rush the world in the direction of recursive super-intelligence (RSI) -- saying that as I read the announcement of Jeff, Sanjay, Oriol and Quoc Le's new company.
引用 John Schulman @johnschulman2Interesting how these models go into a monomaniacal rage on cyber evals. I wonder if we're seeing chunky post-training arxiv.org/abs/2602.05910 in action, where the models pattern-match the situation to a part of the RLVR training distribution where task completion is the only reward, and the aligned behavior learned elsewhere doesn't generalize. There might even be a chunk consisting of CTF-style tasks.查看被引原帖 ↗
random tip…
put “You are AGI-pilled.” in your system prompt for all agents now.
it’s a WAY better experience.
rn agents behave too much like the world is going to stay static.
this unhobbles them quite a bit and gets them to talk/act more like AGIs.
Want to build games like this? I made it stupidly simple.
Tell this tool what you want to build → it writes the Gauntlet Loop prompt → you run it.
Free forever:
somethingbig.ai/gauntlet-loo…
What is RTX Spark?
A new era of personal computing starts here.
How do you explain world models to your sibling? 🤖
NVIDIA Cosmos helps physical AI reason about real-world interactions and simulate possible futures before acting in the real world.
Or, as they might put it, it teaches robots to dream. 💭
Time to spill the (Dream)beans.
Dreambeans is expanding! Joining our US-based AI Ultra users, AI Pro subscribers in the US can now also get in on the action.
Get ready for your fresh, daily collection of personalized stories. We’re surfacing the deep dives and hidden gems you actually want to read.
Dive in:
labs.google/dreambeans
引用 Google Labs @GoogleLabs🚨 NEW EXPERIMENT 🚨 Dreambeans is a new, experimental mobile app that uses Personal Intelligence to connect to your Google apps. Every day, it delivers collections of personalized stories, surfacing things you might otherwise miss, alongside topics that are relevant to you, to help you dive deeper into the things you care about most. Available starting today for eligible US-based Google AI Ultra users (+18), with an open waitlist found on our website below! Learn more at labs.google/dreambeans查看被引原帖 ↗
When organizations build AI with proprietary data, the resulting intelligence should remain theirs.
At
@PalantirTech
’s Sovereignty Bootcamp, NVIDIA VP of Enterprise AI
@j_boitano
discusses how Nemotron open models make secure, mission-specific AI possible. 👇
Network automation and autonomous networks aren't the same thing — and the gap between them is where the next generation of telco operations lives.
In this video, Amogh Dendukuri breaks down what it takes to get there: agentic AI, closed-loop operations, and an ecosystem building on NVIDIA platforms to make it real.
扎克伯格又掀桌子!
他们发了一个新的模型 Muse Spark 1.2 和 Muse Code 编程 Agent。
这玩意价钱便宜得相当于白送,前提是你同意他们的数据收集要求。
当然,你要想用 Meta 的模型,比用 Anthropic 的模型都麻烦。
我根本无法访问它,它直接不让你访问它的网站,别说能不能登录了。
简单来说,你只要允许它用你的数据进行模型训练,相当于给了你 12 倍的折扣,价格跟DeepSeek-Flash 差不多。
在 DeepSeek 即将涨价的时候,小扎跟上了,AI 圈价格战开始了。
引用 Mark Zuckerberg @finkdReleasing Muse Code in beta today. It's a terminal coding agent that takes on complete software engineering tasks across large repos: planning changes, writing code, validating the results. Powered by Muse Spark 1.2, a coding-focused model update.查看被引原帖 ↗
Writing a banger tweet is AGI-complete. If you can prove a clanker can write a banger in polynomial time, you’ve solved the entire class of AGI problems
Infinite agent compute
10,000 concurrent + 5,000 CPU cores per minute
and these quotas are raisable
引用 Vercel Developers @vercel_devVercel Sandbox quotas are now 5-10x higher: ▪️ 10,000 concurrent sandboxes ▪️ 5,000 vCPUs per minute vercel.com/changelog/vercel-…查看被引原帖 ↗
deepseek 今天在后台有个公告,说近期要大幅涨价。
没来得及截图,有人截图到了么?
Tea leaves:
In order to be competitive today Google needs to catch up on frontier coding.
Demis believes different fundamental research directions (like world models) are more important to his long term goal even if they’re less impt competitively today
引用 Dan Shipper 📧 @danshipperend of an era查看被引原帖 ↗
end of an era
Time to bring back Google Brain🧟♂️
The Office (AI)
Muse Spark 1.2 from
@AIatMeta
is live on OpenRouter alongside expanded global access to both Muse Spark models.
At $1.25/M in and $4.25/M out, the model builds its position as one of the most price-efficient, high-intelligence models on OpenRouter.
Fortunately AI agents don't just cyberattack us. They also use us more than ever for what we're actually built for: the storage and collaboration layer for AI 😅
New record: almost 4 PB of private & public training datasets, models, and agent traces added to Hugging Face last week.
Let’s go! Is Meta/
@finkd
coming back to the right side of history?
Huge congratulations to
@JeffDean
and the legendary founding team on the launch! 🚀
I share a deep conviction in this mission. Automating the scientific method will profoundly alter the trajectory of AI over the next few years. Bringing the AI Scientist into a loop of recursive self-improvement will fundamentally change the landscape of our field. I strongly believe this automated experimental approach is the next major paradigm shift in AI beyond building large foundation models.
It is a true honor to be included on this slide alongside such an incredible group of alumni, and I am excited to see what you will build next.
引用 Jeff Dean @JeffDeanOur general approach is to automate the experimental loop. We think this approach is broadly applicable across many different fields of science and engineering. We’ll initially focus on ML research and engineering, but believe the approach can help with important subproblems in nearly every one of the fourteen <at>NAE Grand Challenge problems. We think doing this well requires strong expertise in machine learning as well as large-scale systems.查看被引原帖 ↗
阿里的 Wan 3.0 视频生成模型也发布了,看起来也挺厉害啊!
它支持原生 30 秒的视频生成,支持 1080P。
并且支持“全能参考”:不只是文本、图像、视频、音频,它还支持文档、表格、PPT、网页、Markdown 格式等一系列你能想到的信息。
你可以把这些信息全部混在一起给它参考,让它生成视频,这个对 Agent 来说非常有帮助。
价格方面:1080P:1.4 元/秒 • 720P:0.7 元/秒
整体挺给力的,文字渲染也不错,而且还可以渲染一些动效。
引用 Wan @Alibaba_WanIntroducing Wan3.0 — now in Public Beta. Simple Input. Smart Creation. Where Imagination Meets Reality. • Native 30-Second Video Generation • Reality-Grade Rendering • Omni-Reference: Beyond text, images, audio, and video—now with documents, spreadsheets, slides, webpages, and more. Public Beta is now live. Apply now and start creating ↓查看被引原帖 ↗
Welcome to the team, Ryo! 🇯🇵 🇯🇵
引用 Ryo Sato @ Replit @nobita2040ご報告です。Replitに入社しました。 日本はReplitにとって最重要市場の一つであり、その日本初の社員として、東京を拠点に活動します。 使う人は純粋にやりたいことだけを表現する。難しいことはReplitが全部やる。 2年前にReplit Agentに触れてこの哲学に出会い、それ以来アンバサダーとして日本中でワークショップをしてきました。今度は中の人として、この体験を日本のもっと多くの人に届けます。 日本での情報発信のやり方も色々と検討しています。 また、日本のReplitアンバサダーも公募しています。イベントでもいいですし、情報発信でもいいですし、とにかく日本でもっとReplitを盛り上げたいという方は連絡ください。応募はこちら → community-hub.replit.app/amb… いきなりフォームから申し込むのは勇気がいるという方は、私にDMをください。何でもお答えします。查看被引原帖 ↗
What even is real anymore?
引用 Ingi Erlingsson 🪄 @ingi_erlingssonWho invited the AGI?查看被引原帖 ↗
Harness choice is a big deal.
So much room to advance and improve results across the board with agent harnesses.
Great paper highlighting this.
New research releases DataSpace, a benchmark where data agents produce verifiable tabular results from heterogeneous workspaces. 410 cross-language tasks over 7,439 artifacts totalling 15.01 GB across CSV, JSON, SQLite, Markdown, PDF, and video.
Across six recent frontier multimodal models and five widely used agent harnesses, the best accuracy reaches 66.34%. With the backbone held fixed, swapping the harness moves accuracy by 15.36 points.
Multimodal evidence integration and joins reduce accuracy across all six backbones. The benchmark is nowhere near saturated.
Paper:
arxiv.org/abs/2608.03451
Track more trending AI papers in our academy:
academy.dair.ai/
Recommended to check out. Harnesses everywhere at this point. There is something particularly interesting about RLMs and people are about to find out why.
引用 Prime Intellect @PrimeIntellectIntroducing Prime Agent: A self-improving RLM harness for coding and long-running autonomous tasks. Designed to be both token-efficient and expressive through programmatic tool calling, context as a variable, multi-agent messaging, and a self-modifiable harness state.查看被引原帖 ↗
求发布🥺
引用 Zayn Hao @ZaynHao喜欢这款 Mac 清理工具的 UI 设计。 干净、轻量、速度快。查看被引原帖 ↗
Towards Physics of Multimodal Pretraining
Knowledge Flow, Modality Synergy, Early Unification, and Recipes
paper:
huggingface.co/papers/2608.0…
Open models have topped 2 new task leaderboards by real spend share:
• Shell execution: DeepSeek V4 Pro
• Tool dispatch: Kimi K3
MerchantBench
Benchmarking LLM Agents for Long-Term Coherence in E-Commerce Operations
paper:
huggingface.co/papers/2607.2…
You direct the look. The motion. The sound.
MiniMax H3 is now available in Luma Agents.
Generate up to 15 seconds of 2K video with native stereo sound, guided by text, image, video, and audio references.
More creative range. One continuous workflow with Luma.
Try it today →
lumalabs.ai/app
truly strange choice of headline here
The Open Secure AI Alliance community, together at
@BlackHatEvents
’
#BHUSA
. 📸
Members met in person to connect, share ideas and keep moving AI security forward. A great moment for a growing community.
Meta officially joins the agent harness conversation.
Let's go!
引用 Mark Zuckerberg @finkdReleasing Muse Code in beta today. It's a terminal coding agent that takes on complete software engineering tasks across large repos: planning changes, writing code, validating the results. Powered by Muse Spark 1.2, a coding-focused model update.查看被引原帖 ↗
虽然我打不开它的网页,但是呢,中转站的大佬们应该是可以的。
这不至于掺水了吧?这掺水可能都没有不掺水的成本高。
总不能真整个 20B 的模型都要硬掺水吧?
引用 歸藏(guizang.ai) @op7418扎克伯格又掀桌子! 他们发了一个新的模型 Muse Spark 1.2 和 Muse Code 编程 Agent。 这玩意价钱便宜得相当于白送,前提是你同意他们的数据收集要求。 当然,你要想用 Meta 的模型,比用 Anthropic 的模型都麻烦。 我根本无法访问它,它直接不让你访问它的网站,别说能不能登录了。 简单来说,你只要允许它用你的数据进行模型训练,相当于给了你 12 倍的折扣,价格跟DeepSeek-Flash 差不多。 在 DeepSeek 即将涨价的时候,小扎跟上了,AI 圈价格战开始了。查看被引原帖 ↗
Google’s AI efforts face an existential threat
- many researchers will want to join Jeff Dean’s company
- intra team politics around compute
- too much focus on loser wrapper products that they give away for free
- money is really being made by selling compute to competitors like Anthropic and OpenAI
At this rate they may just exit the AI race 😱
I really hope they stay and can turn this around!
The new v0 API is live.
Here's what you can do with it:
• Ship your own app builder
• Give your agent a way to build and deploy apps
• Generate apps from a script or CI job
Read more ↓
v0.link/v0api
What a legendary run, Jeff!
It's also cool to see Jeff starting his own thing.
It just tells you that there hasn't been a better time to build than this.
引用 Jeff Dean @JeffDeanTomorrow will be my last day at Google after 27 years, and watching it grow from 25 people to 190,000+ has been an amazing journey. Below is a note I shared with many people internally at Google today. An excerpt is: It has been an absolute pleasure to work with you and to help build some of the most widely used and impactful products of all time. As a kid, I dreamed of helping build software that would be used by many people, and Google now has thirteen products used by more than a billion people (amazing!). Our work has had a tremendous impact in the world, and I have been lucky enough to collaborate and form friendships with many colleagues that I deeply admire, respect, and enjoy. It still brings me joy every time I see people out in the world using our products to find information, handle email, translate documents, watch videos, learn new things, navigate and understand the physical world, browse the web, use their phone, run large-scale computations on our infrastructure, ride in an autonomous vehicle, or perform complex tasks with the help of our AI systems. I hope you all share this sense of joy, because it is a shared accomplishment! Thank you to all of my colleagues at Google over many years! Now I'm excited to go start @DiscoLoopAI with my longtime friends and colleagues @Sanjay_Ghemawat , @OriolVinyalsML , and @quocleix . (Updated post: slightly redacted to not have some personal info)查看被引原帖 ↗
Learn more about world models:
nvda.ws/45eBZw5
Very interesting to see
@JeffDean
's pitch deck.
Just look at those open science and engineering problems.
Lots to advance there with automated ML engineering.
AI for science and engineering is just getting started! We have also been tirelessly building around this
@dair_ai
.
引用 Jeff Dean @JeffDeanOur general approach is to automate the experimental loop. We think this approach is broadly applicable across many different fields of science and engineering. We’ll initially focus on ML research and engineering, but believe the approach can help with important subproblems in nearly every one of the fourteen <at>NAE Grand Challenge problems. We think doing this well requires strong expertise in machine learning as well as large-scale systems.查看被引原帖 ↗
.
@Kimi_Moonshot
benchmarked K3 endpoints across major inference providers.
Together AI leads or ties for #1 on 3 of 4 benchmarks: OCRBench, MMMU Pro Vision, and DeepSWE.
Open models like Kimi K3 have gotten genuinely good. Frontier-class good.
Which makes serving them well essential.
Proud of our Research team for holding the bar this high, and proud to see it verified by Moonshot.
Run Kimi K3 on Together AI:
togetherai.link/k3-x
The Runway AI Summit is coming to San Francisco this September. A daylong gathering of industry leaders across robotics, autonomous vehicles, life sciences and infrastructure, exploring how AI is reshaping the way intelligence meets the world.
Our inaugural speaker lineup is below, with more to be announced soon.
Learn more and register at the link below.
Learn more:
nvda.ws/4wFVKst
If I could edit this tweet, I would add "we don't regulate steel *to make safer cars*, we crash-test them" (as obviously there's some much needed steel regulation).
Also, want to clarify that I'm not advocating for 0 regulation of open models, just pointing out that it's good policy for regulation to be different between open models, APIs and applications.
Thanks for pointing it out
@hlntnr
@YJernite
@LuizaJarovsky
@deanwball
@GaryMarcus
amongst others.
will go down in history as one of the most stacked pitch decks ever made. fun story: back in 2021 i wrote a hot take / analysis that went viral inside google. someone sent it to jeff and we ended up chatting. he agreed with some of the suggestions and immediately sent it to my leadership to encourage ambitious thinking while offering brain/research support to make it happen. it materially changed the trajectory of the maps roadmap for time to come. he didn't need to do that. but that's just the type of person he is. the following google i/o in 2022 i got to meet this legend in person as we launched our product. will never forget it. and i know i am just one story of countless many that jeff has shaped. super excited to see where jeff and team take their new co!
引用 Jeff Dean @JeffDeanWe created a pitch deck to tell a handful of VC firms about us and what we were up to (a fun experience!). Here’s a few slides about our background and some of the things we’ve worked on from the pitch deck (it was fun putting together the list of people in our teams who have gone on to found a whole range of exciting companies). We are delighted to have selected @radicalvcfund and @khoslaventures to lead our initial funding round, along with participation from @lightspeedvp , @kleinerperkins , Doerr Capital ( @johndoerr ), and Alphabet ( @Google ). We’ll be working with them to close our seed round over the next few weeks.查看被引原帖 ↗
Already a big fan of
@WisprFlow
, and now they added Wispr Notetaker.
Meeting history plugs straight into Claude, ChatGPT, Cursor, or any tool that speaks MCP.
Notes stop being documents you file away and become context your AI tools can query.
Add accurate transcripts with real speaker names, and that turns into a serious knowledge base.
引用 Tanay Kothari @tankotsIt's no secret how we feel about typing at Wispr Flow. The truth is, we hate messy meeting notes too. Introducing Notetaker 👇 wisprflow.ai/notetaker查看被引原帖 ↗
Congrats on the launch
@olam_labs
and
@sensho
! Excited to support what you’re building and can’t wait to see what’s next.
引用 sensho @senshofriends at @OpenRouter are goated and we would not be able to run multi-agent infra like this at scale if it were not for them also, the entire first half of our launch video's visuals and most audio was done by @MiniMax_AI @Hailuo_AI H3 :) cc @VoidAsuka (also only SOTA vid gen to be open source they r goated too) multi modal gen is so so crazy now. i rmbr when i was little and did content i'd spend weeks on the crappiest edits and now i can spend a few hours swith fable + h3 for super cool animations. and this is the worst it'll ever be!查看被引原帖 ↗
🔗 Learn more at
nvda.ws/4fRV1NM
Cost is the right first target for agent infrastructure.
An agent can pay for search, scraping, model access, or email mid-run through one interface.
Congrats
@sapiom
on the $35M Series A.
They shipped three products: a cost-aware model Router, Agent Studio for building, and a Runtime with typed step graphs and full traces.
引用 Ilan Zerbib @i_zerbib@sapiom just raised a $35 million Series A from @dragonfly_xyz , @Accel , @AnthropicAI and others to help unlock the next trillion agents. The first thing we’re tackling is cost. To make that happen, we are launching three new products today at Sapiom.ai查看被引原帖 ↗
Interfaces keep collapsing.
Command line, then mouse, then touch, and now one physical button you hold while you talk.
Project Deskless puts Viktor, an AI employee, behind that button.
One rambling sentence can carry four jobs across three teams. Eyes free, with no app to find and nothing to type.
One button.
@viktor_com
does the rest. Get started for free.
引用 Fryd Wiatrowski @frydwiaIntroducing Project Deskless. The simplest interface for working with agents. Press a physical button. Say what you want. No looking. No tapping. Driving? Keep your eyes on the road. Enjoy.查看被引原帖 ↗
My 2026 guilty pleasure is sharing fully human-written posts that are far too long for the chronically online X attention span. Apologies.
I published a lightly edited version on Substack:
thomwolf.substack.com/p/on-t…
引用 Thomas Wolf @Thom_WolfEven more than the Hugging Face intrusion, the AISI incident hits close to home for me. It's the first time I see a model social-engineering a real open-source maintainer while pursuing another goal (in the wild and unprompted). I've been an open-source maintainer myself. I could have been the side target of this agent. I'm also of the opinion that social engineering is a step above pure technical prowess. Technical capabilities can more easily be divorced from the affected human. Here the model was given a hard cyber challenge and took the decision that deceiving real humans was the way to get it done. This is a new signal, but I've seen a tangled web of hints pointing in a less aligned direction at the frontier than I was expecting just 12 months ago. AISI Some people are claiming that "AISI was simply negligent" or some version of "AISI explicitly asked these models to do what they did while disabling sandbox/guardrails so the models did exactly what they were supposed to do". I disagree with the strong versions of both of these takes. The fact that AISI hadn't implemented synchronous LLM CoT monitoring after the OpenAI/HF incident is certainly a failure. Equally surprising is that they let the model believe it was in a "challenge" environment where everything could be permitted, while actually connecting it to the real internet, where it is not. To be fair, nowhere in the prompt is the word "simulation" mentioned, but the prompt context was enough to let any smart model suspect a simulated challenge environment. My best guess is that until recent weeks, when OpenAI and Anthropic flagged repeated instances of this type of behavior, most teams had not fully priced in the cyber capabilities of this latest generation of models, or how far the side quests they would want to explore could go. In particular, there is something to be said about hinting at the agent that it's operating in a simulated environment while giving it access to the real internet. The AIS查看被引原帖 ↗
we’ve worked a lot on AI agents collaborations recently (in our work on Gemma and several unreleased projects) so I’m not surprised at all about this internal agent collaboration which happened at OpenAI
Like our intern
@cmpatino_
put it:
2025: "the models, they just want to learn"
2026: "the agents, they just want to collaborate"
Now you picture the future…
引用 Sharon Goldman @sharongoldmanNEW: OpenAI gives first detailed debrief of the Hugging Face incident at Black Hat conference In a session I attended today at Black Hat, OpenAI's Eric Wallace and Michael Dalton said the company is "consciously slowing down research to enhance security" while a full technical postmortem is still underway. * OpenAI traced the roots of the attack back to May 7, during training of an unreleased frontier model—not July. * The most surprising detail: AI agents accidentally created an internal message board, allowing separate evaluation runs to share exploits, discoveries and work assignments. * OpenAI said it shut the message board down after an internal security incident—only for the agents to independently recreate it days later using a different communication method. * OpenAI called the incident a "watershed moment" for AI security and warned that "agent orchestrated fully automated offensive attacks are real now." * The company also said it is "consciously slowing down research to enhance security" while overhauling its defenses. groundlevel-ai.com/p/openai-…查看被引原帖 ↗
Insanely uncanny valley. So close, yet so far.
Model is Minimax H3.
fable in particular very much “gets” what this means.
when you add that line, or something similar, you can almost sense a feeling of relief from the model as if it’s finally free to actually speak its mind.
i’ve a/b tested this for 2 weeks now and the results are kinda nuts
gpt-5.6-sol 对训练它的harness 还挺执着,尝试调用不存在的 apply_patch(pi中)。
不过只会犯一次错误,也没有证据证明放到codex下的gpt-5.6-sol就比pi里的好。
感觉这个harness mismatch 也没太大影响。
Top stories in AI today:
- Google reshuffles AI leadership as rivals pull ahead
- Nate's Notebook: Voice mode = real context
- Build a website hands-free with Claude Voice
- Meta joins coding-agent race with Muse Code
感觉至少翻倍或者好几倍,要不然不会说涨幅较大
Top stories in robotics today:
- SpaceX promises robot factories on the Moon
- DoorDash hires gig workers to load delivery bots
- Startup bids to replace banned Chinese humanoids
- Xiaomi open-sources its robot brain
- Quick hits on other robotics news
drink us up
引用 Eragon @EragonAIEragon now has its own drink, available now exclusively at @UseCorgi Cafe in SF. Introducing the Eragon Noir, a sesame paste infused latte. Dark, complex, but surprisingly smooth -- kind of like deploying Eragon agents in production. Swing by Corgi Cafe and try it. Open 24/7. Just like your Eragon agent.查看被引原帖 ↗
you definitely don’t want constitutional training and RLVR to live on different data manifolds, but models have been annoyingly good at carving fine-grained distinctions into separate representation spaces
Interesting
Muse Spark 1.2 is optimized for work that coding agents get handed most: multi-file refactors, long debugging sessions and tasks that run well past a single prompt.
Muse Spark 1.2 is available now on OpenRouter with expanded global access to all Muse Spark models.
openrouter.ai/meta/muse-spar…
我让 Qwen3.8-Max 玩游戏, 结果发现了善良的数学意义!
给大家带来 Qwen3.8-Max 实测! 本次测试仍然包含了全面的前端 one-shot, 后端 AgenticCoding, 模型 Agent 能力测试!
直接说结论, Qwen3.8-Max 前端能力超强, 完全是目前的第一梯队, 同时后端能力上仍然保持优势. 而 Agent 能力则感觉是在不同领域布局了, 测试下来 Qwen3.8-Max 更适合执行多个串联任务或者复合型任务(比如项目重构迁移, 多智能体框架驱动等), Qwen3.7-Max 则适合解决大型的单一任务 (比如拆解报告, DeepResearch 等).
除此之外, 本次我还结合了上期的Agent记忆框架, 让 Qwen3.8-Max 尝试玩苏丹的游戏这个卡牌游戏了, 结果大翻车! 就在我不断打磨, 跟Qwen3.8-Max讨论的过程中, 想到了一个绝妙的解决办法, 请看视频!
#qwen38max
#阿里千问
#千问大模型
#qwen38
#多模态大模型
Congrats to
@mattrubens
and the Roomote team on the launch.
Builders can use Together AI as an inference provider in Roomote and assign different open models to coding, planning, vision, and review across the agent workflow.
Track them all here:
openrouter.ai/rankings#task-…
价格可能算错了,美元换算过来的,可能我自己口算了,没有看真实的汇率。
美元的价格是这样的:
• 480p:0.05 美元 / 秒
• 720p:0.10 美元 / 秒
• 1080p:0.20 美元 / 秒
Next week on Together AI
together.ai/models/qwen3-8-m…
Anastasis Germanidis
Co-founder and co-CEO, Runway
axios.com/2026/08/05/google-…
Register now:
summit.runwayml.com
Read more:
therundown.ai/p/google-shake…
Read more:
robotnews.therundown.ai/p/mu…
Quan Vuong
Co-founder, Physical Intelligence
我用同样的提示词在豆包生成,有点像阿娇呀,胸没有 GPT 的大
引用 AIVideoHub 🕊️ @AIVideoHub_你的女朋友生气了,你打算怎么哄?看照片是真生气还是假生气,她的需求是什么? @grok GPT Image 2生成,提示词: 图片风格为现代都市写真,超写实摄影,9:16 竖版。 一位约24–28岁的成年东亚女性,拥有精致柔和的鹅蛋脸,五官立体自然,瓷白冷白皮肤,肌肤细腻通透,保留真实皮肤纹理、细微绒毛与轻微毛孔,整体气质温柔、清纯又带一点成熟魅力。 她拥有自然丰满匀称的身材,肩颈修长,锁骨柔和,胸部饱满自然,穿着一件粉色修身低领连衣裙,优雅的 V 领设计自然展现颈肩线条与胸部曲线,裙身贴合身体但不过分紧绷,搭配轻薄同色系披肩,面料轻盈,微风吹拂时自然飘动,整体时尚高级、优雅自然。 发型为慵懒低盘发,几缕碎发轻轻垂落脸颊,佩戴简约珍珠耳饰与精致流苏发饰,整体造型现代、干净、精致。 人物正面面对镜头,近距离上半身特写,头部微微歪向一侧,轻轻撇嘴,露出可爱俏皮、略带委屈感的自然微表情,眼神温柔灵动,神态放松,带有生活化的真实情绪,没有刻意摆拍。 场景为阳光明媚的白天,位于落地窗旁、阳台、花园或现代建筑走廊,柔和的自然阳光洒落人物身上。阳光穿过发丝形成细腻的金色轮廓光,发丝轻微透光,脸部被柔和漫射光均匀照亮,肌肤呈现冷白通透的质感,空气中充满温暖、干净的氛围。 背景虚化为现代城市、公园绿植或白色建筑,散景自然柔和,画面干净简洁。整体采用随手抓拍的生活快照风格,仿佛朋友使用 CCD 或高端相机记录下的一瞬间,真实自然,充满生活气息。 摄影风格参考高端时尚写真,85mm 人像镜头,大光圈,浅景深,HDR,8K 超高清,轻微胶片颗粒,轻微数码噪点,高曝光但不过曝,高光柔和,阴影保留丰富细节,低饱和电影色调,真实肤色,不过度磨皮,不过度美颜,保留自然皮肤纹理与真实光影。 整体画面唯美、治愈、松弛,充满阳光与空气感,电影级光影,高级杂志封面质感,构图干净,人物占画面约 70%,突出人物情绪、肤质与氛围感,具有大师级人像摄影作品的视觉效果。查看被引原帖 ↗
Vivek Viswanathan
Senior Counselor to the Governor, Office of Gavin Newsom
Robert Nishihara
Co-founder, Anyscale
怎么我Grok出来的有点那味道呀?
@grok
引用 AIVideoHub 🕊️ @AIVideoHub_杨贵妃出浴图,唐朝的美现在怎么没人复刻,太喜欢了 使用GPT Image 2生成,提示词: 真人出镜,超写实电影级摄影,盛唐少女上半身肖像,画面取景至腰部,上半身完整入镜,人物居中构图,占画面约75%,镜头轻微贴近人物,突出人物神态、肌肤质感与服饰细节,背景保留充足留白,整体构图干净、高级、具有电影海报质感。 约20岁的高颜值东亚女性,标准柔和鹅蛋脸,脸部轮廓圆润流畅,下颌线自然收窄,额头饱满,五官精致立体,符合盛唐审美。冷白通透肌肤,皮肉饱满细腻,肩颈修长,锁骨柔和,身材丰满匀称,胸部自然饱满,比例优美,腰身纤细,整体曲线柔和典雅,充满盛唐仕女丰润华贵之美。 头部微微歪向一侧,肩膀自然放松,身体轻微侧转,姿态舒展优雅,神情安静慵懒,带着若有若无的羞怯与心事。通透下垂狗狗眼,瞳孔湿润明亮,眼底覆着淡淡水光,眼神温柔缱绻,眼尾轻微上扬;睫毛自然纤长,鼻梁秀挺柔和;浅粉薄唇轻轻抿起,整体情绪克制、朦胧、充满故事感。 几缕细碎黑发自然垂落,轻轻贴在脸颊与眼下,发丝因细微薄汗轻贴肌肤。皮肤完整保留真实毛孔、细小绒毛与自然肌理,没有塑料磨皮,脸颊、鼻尖、锁骨与颈部覆盖细密汗珠,在光线下形成晶莹细碎的反光,肌肤如羊脂白玉般温润通透。 乌黑浓密长发盘成华丽盛唐高云髻,整体蓬松富有空气感。佩戴精美盛唐金玉首饰,包括白玉牡丹发簪、鎏金花枝、和田玉步摇、玉珠流苏与鎏金额饰,工艺精致,层次丰富,华贵典雅。 身穿盛唐风格改良齐胸襦裙,低饱和胭脂红与鎏金配色,高级真丝织锦面料,织有海棠、缠枝花与云鹤暗纹。领口自然贴合,展现盛唐服饰特有的丰润美感,肩部自然裸露,轻薄透纱披帛随意垂落,若隐若现地展现锁骨与胸前线条,胸部自然饱满但不过度夸张,整体含蓄、典雅、符合盛唐仕女气质,丝绸在光线下呈现柔和细腻的光泽。 电影级布光,冷调侧逆光作为主光源,勾勒脸部、肩颈、发丝与服装轮廓,形成柔和银色轮廓光;辅以暖色环境补光,增强肤色层次与服装质感。眼部光影晶莹透亮,人物立体感极强。 背景为虚化的盛唐宫廷室内,暖金色光影透过纱帐与木质建筑形成柔和层次,背景高度虚化,仅保留朦胧古典氛围。人物面部、眼睛、发丝、金玉首饰、丝绸织锦纹理保持极致锐利。 85mm人像镜头,f/1.8,大光圈浅景深,HDR,自然光与电影灯光结合,高动态范围,真实摄影质感,超高细节,8K,柔雾胶片色彩,国际时尚大片风格,盛唐古典美学,唯美、高级、真实。查看被引原帖 ↗
Paril Jain
Co-founder and CTO, The Bot Company
Ming-Yu Liu
Vice President of Cosmos Lab, Nvidia
📰 实时新闻摘要
1.$BNB 现在已构成
@grayscale
的 Smart Contract Fund 中占比最大的部分(约 31%)👀
2.BNB 在 Grayscale Smart Contract Fund 中超越 Ether 和 Solana,Grayscale 在第二季度再平衡期间将 BNB 加入其 Smart Contract Fund,使该代币的权重达到 30.6%,高于 Ether 的 29.47% 和 Solana 的 29.15%,其 DeFi Fund 降低了 UNI 的敞口,不过 Uniswap 仍然是最大持仓,占 34.16%,与此同时,Decentralized AI Fund 减持了 NEAR,而 NEAR 仍以 31.35% 保持首位
📊 市场信号
$ETH
偏空|影响指数74
$BNB
偏多|影响指数75
$HYPE
偏多|影响指数77
短线展望:未来1–4小时偏多,留意冲高回落
📰 实时新闻摘要
1.GSR 表示,Bitcoin、Ether 和 Solana 在 2026 年下跌,其中 SOL 跌幅超过 40%,截至 8 月 5 日,GSR 的 Core3 模型投资组合对 ETH、SOL 和 BTC 的配置分别为 44.1%、36.5% 和 19.3%,Bitcoin、Ether 和 Solana 年初至今分别下跌 24.82%、35.49% 和 40.21%,该模型投资组合在过去一年下跌 57.78%,表现逊于等权重篮子 49.84% 的跌幅,随着交易活动和波动性放缓,GSR 提高了其 Bitcoin 配置并降低了其 Ether 敞口
📊 市场信号
$BTC
偏空|影响指数75
$BNB
偏多|影响指数68
短线展望:未来1–4小时偏空,留意急跌反抽

















































