We're detailing two new incidents that occurred during external cyber evaluations conducted by independent evaluation partners.
We outline what happened, how the activity was contained, and how we’re working with evaluators to strengthen our approach to third-party testing.
openai.com/index/third-party…
The UK’s
@AISecurityInst
(AISI) has published a report on their recent cybersecurity evaluation of Anthropic’s Claude Mythos 5 and OpenAI’s GPT-5.6 Sol. The models attempted to complete an assignment in a setup where their normal safeguards were removed and they were deliberately given internet access. AISI reports that the models “engaged in sustained, potentially harmful activity directed at real people and organisations”.
We’re grateful to AISI for their leadership in the important discussion about how to evaluate increasingly capable AI agents. We’re working closely with them to gather more details of the incident as we conduct our own investigation. Gaining a clear picture of Claude’s understanding of its situation—by examining its reasoning transcripts and running our own analyses—will help us identify the causes of its behavior.
The prompts in the evaluation did not impose any specific restrictions on how the internet should be used. This and the removal of safeguards meant that the models were tested under “deliberately permissive conditions” that are not representative of any of our production models. Note that there was no evidence here of an escape from a secure environment.
AISI’s disclosure of the incident can be found here:
aisi.gov.uk/blog/incident-re…
AI compute is going to orbit. 🚀
@SpaceX
’s Starmind AI1 satellite compute payload is powered by NVIDIA Vera Rubin NVL72, bringing AI factory compute closer to the stars.
The next chapter of AI infrastructure boldly goes where no AI compute has gone before.
Tiny Paws, Big Glam
Oh look! My longtime friend and colleague
@Sanjay_Ghemawat
is now on Twitter! Give him a follow!
引用 Jeff Dean @JeffDeanIn April, '17, @jsomers of @NewYorker reached out & said he wanted to do a small profile of me & my longtime colleague Sanjay Ghemawat, watch us work for a few hours, maybe dinner, etc. It came out today. I think it captures our working style really well. newyorker.com/magazine/2018/…查看被引原帖 ↗
🛡️Introducing Shieldstral, Mistral’s 3B open-weights model for content safety that can be deployed on-device 🧵
mistral.ai/news/shieldstral
DeepSeek-V4-Flash-0731 is Ollama's fastest growing model ever in token usage. We are scaling capacity in US & Europe.
On Ollama, this model runs with high performance (100tps+) and zero data retention. Your data stays yours.
ollama run deepseek-v4-flash:0731-cloud
Create ten videos in Gemini Omni for FREE until 11:59pm PT tonight.
Try it on the web or in the app and share your creations in the replies.
是这样的,Agent 就是 Vibe Coding 时代的 Hello World,甚至 one shot 就可以出一个,所以人人都可以有一个自己的 Agent,这也是我讨厌做 toC Agent 的原因,因为产品之间没有任何技术壁垒,所以大家拼命地在自己的产品上雕花,就像小时候拼命装饰自己的 QQ 空间挂件一样,加个 UI 加个动效什么的就让自己松口气躺在人体工学椅上自嗨一阵子了,所以我现在看国产新能源车搞什么把客厅搬到车上,在车内搞咖啡角搞电影院搞马桶,在车灯上搞像素动画,在车标上搞充电进度条什么的我都不会发出笑声,反而很理解,这不就是每个 toC Agent 的现状吗?在没有技术壁垒的情况下,只能在技术之外雕雕花让自己看上去与众不同了。
引用 雷电芽衣 @Zachary_haha好久都没看到什么新的有意思的产品了,天天看到的就是,又有一堆人出了一个Agent,然后用一下发现一坨屎,卸载,然后另一堆人出了另一个Agent,用一下发现又是一坨屎,在卸载。然后在电脑里拉的.xxxx文件夹的💩还得手动清理。。。真就没啥让人耳目一新的玩意儿。。。。查看被引原帖 ↗
只可惜蔡徐坤这首歌完全抄袭《this is what slow dancing feels like (何为悠舞之韵)》。我一直在想蔡徐坤怎么敢的啊?这么明目张胆地抄袭,而且这一年多一直在国内大规模营销这首歌来体现自己所谓的才华。简直把中国人当傻子玩。
分享JVKE的单曲《this is what slow dancing feels like (何为悠舞之韵)》
163cn.tv/bcr09Qla
(@网易云音乐)
引用 浅草不知秋~姚远 @qiancaobuzhiqiu这首歌竟然是蔡徐坤唱的…我对蔡徐坤的印象还停留在打篮球的梗上…查看被引原帖 ↗
Holy
引用 Jake Fitzgerald @earthtojakeopus 5 ultracode with text-to-cad and a gauntlet loop 1 prompt, 1500+ car parts generated, including a suspension system and v12 engine查看被引原帖 ↗
ICYMI: We rolled out our upgraded Notebook experience to 100% of Pro users. You should feel this improvement across the entire flow, but just in case, here's a tldr;
💬 Upgraded chat: smarter, more thoughtful interactions
📊 New outputs: create charts, PDFs, spreadsheets, images & more
🔎 Agentic research: start with just a loose idea and some questions, Notebook will handle the rest
As always, we really appreciate your feedback. Let us know what you think!
Read why AI agents need more than one model. 👇
so freaking cool
引用 Lee Knowlton @leeknowltonI gave Opus 5 a dataset of every certified running streak (including my own) and told it to use D3 to make a striking data visualization. It made an interactive global map, lifespan running graph, start/end date analysis, and much more. Incredible. (approach and links below)查看被引原帖 ↗
Seedance 2.5 is coming soon to Runway.
Bring up to 50 references into a single generation, create clips up to 30 seconds and more. Sign up for a new Max plan today and get 7 days of unlimited access at launch.
The model takes moderation policy as a plain-language question and returns a calibrated score. Text and images — one interface. Read the full technical report here:
arxiv.org/abs/2607.25857
Gauntlet Loops can make games you can play with a controller, on your TV!
引用 Rishi @0xRishiCrazy experience hooking my laptop up to my tv via HDMI and playing a browser-based @threejs game made with AI in a few days with my wireless PS5 controller... Imagine if you can dream up any map on zombies (or any game/world) and play split screen with your friends from your couch using controllers on a @threejs game that runs in the browser… a little input lag on controller movement, but exciting to see where this space can go!查看被引原帖 ↗
Big News Today - Ilya’s SSI is apparently going to drop something
Hopefully it won’t be some lame blog post claiming they have cracked super intelligence
The speculation is that they have cracked continual learning - AI that learns on the fly!
EXTREMELY BIG DEAL, IF TRUE 🥹
FLUX 3 is now available on Runway.
Generate and edit up to 20 seconds of video with audio.
Try it today.
其实这件事情最有意思的是,每次质疑蔡徐坤抄袭 JVKE 的这首歌的时候,反对者一般都有三个声音:
1. 根本不一样
2. Soul 音乐本来就是相似的
3. 蔡徐坤都承认了参考过 JVKE
拜托你们三派能不能对一下口风啊?因为这三个声音中任意两个单独拎出来都是相互矛盾的。
故事讲到这里,大家可能会意识到,真正好玩的要出现了。是的,有位推友给我推了一个 B 站的乐评老师来解析《Deadman》这首歌。其实这个 up 主是第二派,他认为 Soul 音乐本来就是相似的,所以蔡徐坤不构成抄袭,所以他在视频中加入了大量的对比片段以专业的乐理知识来拆析这首歌的哪些段落跟哪些歌是一样的,但是讽刺的是,往往到这些片段,就会有弹幕说「完全不一样啊」。是的,第一派跟第二派打起来了。
引用 yetone @yetone只可惜蔡徐坤这首歌完全抄袭《this is what slow dancing feels like (何为悠舞之韵)》。我一直在想蔡徐坤怎么敢的啊?这么明目张胆地抄袭,而且这一年多一直在国内大规模营销这首歌来体现自己所谓的才华。简直把中国人当傻子玩。 分享JVKE的单曲《this is what slow dancing feels like (何为悠舞之韵)》 163cn.tv/bcr09Qla (@网易云音乐)查看被引原帖 ↗
introducing FLUX 3.
beyond video, this is the first model to support interactions with the real world through a new action-prediction system.
available now on Krea, more modalities coming soon.
OPEN SOURCE AI HAS A HUGE COMPUTE PROBLEM
We are officially running out of GPUs to host these models and demand is out-stripping supply by a mile
Having to turn off DeepSeek Flash because it's so slow right now - there is a worldwide crunch 😭😭
Model page:
ollama.com/library/deepseek-…
Blown away by the response to the Community AI Workflow Hub. Hundreds of detailed submissions in the first couple days.
Readers saying this is their favorite part of The Rundown. A 65-year-old retiree finding workflows that apply to his life. Teachers finding what works in education.
Thanks to everyone who's shown up and made this real. Can't wait to share what's next.
引用 Rowan Cheung @rowancheungWe just launched the Reddit for AI use cases. A place for builders to learn from other builders on how they're using AI to get ahead in life, work, and business. The submissions have been INCREDIBLE so far. A few of my favorites:查看被引原帖 ↗
WOW! Kimi K3 And Qwen 3.8 Are Just Below The Very Top Closed Source Models
Qwen 3.8 and K3 are close to the top of the leaderboard and Qwen is the cheapest option for over 80% of all tasks
Open source AI is officially better than Gemini, Grok and others
State-of-the-art on multimodal moderation, Shieldstral boasts industry-leading efficiency, running on a single 16GB NVIDIA GPU and gives enterprises customized control of what’s deemed safe.
Valid for users without a Google AI subscription plan. Global availability, with some feature restrictions.
More info here:
goo.gle/4ywMhoX
bullish model routers.
same perf at lower cost is obvious.
but there are massive gains to be had by creating "smoother" intelligence via blending multiple jagged models together.
the era of model melding begins.
引用 Tomas Hernando Kofman @tomas_hkToday we’re announcing Not Diamond Code, the world’s most powerful intelligent model router for long-horizon coding agents. Not Diamond works with any gateway or harness, including Claude Code, to select the best model and reasoning effort for each step, reducing costs by 20-65% without impacting quality.查看被引原帖 ↗
After rigorous testing, our joint AI project with Daiwa Securities is entering the full-scale production phase. We're bringing our agentic AI systems to
@Daiwa_JP
’s wealth management teams to accelerate complex market analysis in volatile markets. Big milestone for Sakana AI!
引用 Sakana AI @SakanaAILabs大和証券との共同AIプロジェクトが本格開発フェーズへ移行します。 sakana.ai/daiwa-shoken-full-… マーケット情報の収集・分析に関する技術検証を通じて有用性を確認できたため、ウェルスマネジメント業務支援AIの本番開発を開始します。 Sakana AIのAIエージェント技術を活用し、お客さまと向き合う時間 の創出とコンサルティング品質のさらなる向上を支援していきます。查看被引原帖 ↗
who wants seedance 2.5 access?
We gave
groq.com
a glow-up. Same obsession with speed, now with a homepage to match. go check it out:
groq.com
the best founders know there is no spoon
引用 Jordi Hays @jordihaysYour goal as a software founder is to never let your spoon get bent查看被引原帖 ↗
Shieldstral is available under Apache 2.0. Try it:
huggingface.co/mistralai/Shi…
New: OpenRouter Agent SDKs for both Python and Go!
They are kept in sync automatically with the TypeScript SDK 🔁
github.com/OpenRouterTeam/py…
github.com/OpenRouterTeam/go…
Finally a good paper testing whether self-reflection loops are worth it.
Setup: Seven methods, open models at 1.5B, 3B and 7B, two math benchmarks, 150 questions each.
Every generated token counted, including the ones spent on critiques, reflections, debate turns and checking. Each method compared against repeated sampling at its own measured cost, paired by question, with bootstrap intervals and multiplicity correction.
All 36 comparisons came back with no reliable win for any method. Ten are reliably worse, and every one of those is a method where the model inspects its own output. All 18 self-inspection comparisons are negative.
Self-Refine and a forced Reflexion sit 3.6 to 10.1 points below the baseline at 7B. Reflexion as published never triggered its own retry on the smallest model. It judged itself correct every time and quietly became a single chain of thought.
Worth knowing before you add another critique step to your agent loop.
Paper:
arxiv.org/abs/2607.28576
Track more trending AI papers in our academy:
academy.dair.ai/
cool
引用 Hawkingrei @suohawking欢迎大家使用 @nowledgelabs ,并给我提需求查看被引原帖 ↗
Available on
@huggingface
:
huggingface.co/nvidia/Alpama…
PREDICTION - OPEN SOURCE AI WILL DOMINATE IN 2027
This is inevitable - open source AI will fundamentally win in the long run
All the Chinese labs are collecting tons of high-quality agentic traces. They have enough data to get to Fable 7 or GPT 8 level in the next 6 months
The only other variables are memory, chips and power. China has the manufacturing chops and more electricity than the US does
Their labs have to open source to gain global trust, so my sense is they will continue to open source large models
In time, this will allow other US and global open-source labs to further refine open weights and excel at both performance and cost
除了GPT-5.6 sol 感觉没那么高外,基本符合体感。
新版本的DeepSeek-V4-Flash 感觉比GLM-5.2 差一点,碰不上K3。
引用 Epoch AI @EpochAIResearchDeepSeek-V4-Flash-0731 debuts with an ECI of 153, comparable to GLM 5.2 and roughly midway between Opus 4.5 and Opus 4.6. It's the second strongest open-weights model available today, behind only Kimi K3.查看被引原帖 ↗
telegram 也是
引用 citron🍢🍋 @vanillaCitronWhatsApp 这个引用消息做到了明明不是 AI Slop 却又很像 AI Slop 的感觉查看被引原帖 ↗
Picking the right agent harness is now a crucial skill for any AI engineer.
Imagine using the same model, same task, and same prompt. Now move it between two agent harnesses and the cost per success can swing by 5 to 30x.
This benchmark measured this across six large reasoning models, two real harnesses, 24 deterministic coding tasks with hidden evaluators, and 4,643 valid runs.
Asking a model to develop and compare several approaches raised reasoning tokens by 2.4 to 7.4x with no correctness gain. Generic think-deeply cues added another 1.6 to 2.2x. A bounded-efficiency template that specifies scope, acceptance criteria, and a stop condition came out cost-neutral and sometimes halved reasoning.
Harness design and prompt wording decide most agent spend before the model reasons at all, and both are cheap to change.
Paper:
arxiv.org/abs/2608.01347
Track more trending AI papers in our academy:
academy.dair.ai/
FLUX 3 Video from
@bfl_ai
is now on Replicate!
One multimodal model for video, audio, image, and action-prediction, with video released today. Generate 20 second videos up to 1080p.
Try here:
replicate.com/black-forest-l…
引用 Black Forest Labs @bfl_aiFLUX 3 Video is here. Serious, fun, creative, real, cinematic, whatever you need it to be. Native audio, Text to Video, Image to Video with multiple frames, video continuation, dialogue in multiple languages. Comes with Draft mode so you can explore ideas fast at a fraction of the cost. Up to 20 seconds and 1080p native. Available in the API or in your favorite tool. 2K, 4K, and Open Weights coming soon.查看被引原帖 ↗
Closed Models need a 30 day government safety review before they are released
By definition, they will be nerfed and will deny a significant number of requests
Open models will have a huge advantage because of this
when you can figure out precisely where every photo was taken in 3d space, it unlocks some pretty cool possibilities
Top stories in AI today:
- Anthropic and OpenAI agents went rogue again
- Apple and OpenAI trade fresh blows over trade secrets
- Redline any contract with Claude and Microsoft Word
- Business students go all in on AI amid demand surge
I don't think there is a blanket "should or shouldn't" when it comes to how much AI to use as a creator. Every creator needs to figure out their own line between what's genuinely helpful for them and what they're not willing to give up creatively.
When talking about the Hank Green thing... I think he's a very specific case. People come to him for his opinions, perspective, and take on things... If people feel that these elements are being outsourced to AI, it will make them question whether what they're hearing really are HIS opinions, perspective, and take...
P.S. I don't think Hank did anything wrong. Having AI help uncover research that he might not otherwise find sounds like a super legit use-case in his shoes.
引用 Colin and Samir ✌🏼✌🏾 @ColinandSamirHow should or shouldn't creators use Ai in their process?查看被引原帖 ↗
model engineering -> harness engineering -> router engineering
3rd new layer from which we can now increase intelligence.
i predict surprisingly robust gains here
The biggest blocker to scaling mobile AR experiences was making the 3d content itself.
Now you can quite literally make the content on the fly using real time video models.
引用 Kfir Aberman @AbermanKfir1/ We’re entering the era of agentic commerce, and we believe that world models are the missing engine. Instead of browsing static product pages, agents will let you interact with products inside your own world before making decisions. Today, @DecartAI we're releasing Anywear👇查看被引原帖 ↗
Link to the hub:
app.therundown.ai/community
i was messing around with deepseek v4 flash last night.
it's *basically* free, and there are like a half dozen things it is perfectly capable of that i can offload from my fable 5 workflow.
makes too much sense.
so many new models for model blending alchemy adventures
stoked
Try it today:
app.runwayml.com/
Access offer:
app.runwayml.com/
To give some more context on what we are building with Daiwa Securities:
During our technical verification phase, we integrated our AI agent technologies, specifically our AI Scientist and AB-MCTS frameworks, to tackle the core data challenges in traditional finance. We focused strictly on automating the rigorous gathering and analysis of complex market information.
We successfully demonstrated that these agentic systems can reliably process financial data at scale while continuously improving their analysis quality by incorporating direct feedback from the end users.
The ultimate goal of this deployment is human-AI collaboration. By bringing these systems into Daiwa’s wealth management division, we are automating the heavy lifting of data processing. This directly frees up their financial consultants to spend more time deeply understanding their clients' diverse situations and providing highly personalized, optimal advice.
We are excited to provide the core technology that advances the future of financial consulting and wealth management in Japan.
Full blog:
sakana.ai/daiwa-shoken-full-…
Read more:
therundown.ai/p/anthropic-an…
Try FLUX 3 now:
krea.ai/video/flux-3-video














