1B next!
引用 Olivier Lacombe @o_lacombeGemma model family crossed the 900M downloads. What a milestone!查看被引原帖 ↗
7 月 26 日(北京时间)全球 AI 圈推文存档,按曝光排序,共 13 条。
← 返回最新 全部归档
1B next!
引用 Olivier Lacombe @o_lacombeGemma model family crossed the 900M downloads. What a milestone!查看被引原帖 ↗
Automating AI research is going to look a lot more like data cleaning than it is going to look like inventing the transformer
Ollama is proud to sign
@satyanadella
's letter.
Our mission from day one has been to make open models accessible to every developer to unlock the next frontier in America and across the globe.
引用 Satya Nadella @satyanadellaOpen-weight models are essential to a healthy AI ecosystem. Together with others across our industry, we are outlining a path for open-weight models to strengthen American competitiveness and expand economic opportunity, while protecting national security. microsoft.com/en-us/corporat…查看被引原帖 ↗
The answer is probably simple, you’re just not doing the obviously correct things.
Yes, open-source / open-weight models are important for a healthy AI ecosystem. That's how we can verify things, check claims, and keep up outside the closed labs. Plus, it gives us the freedom to run AI on our own hardware if we are not ready to share personal data and IPs with closed labs through using their models. (Not that proprietary models are bad, actually I use them a lot as well, but it wouldn't healthy not to have any alternatives.)
Anyway, while pretty much everyone is waiting for the Kimi K3 and Ling 3.0 weights to land on the model hub any day now, there were quite a few other interesting new open-weight model releases the past week. Yes, one of those weeks!
So, here are the architecture pics along with some notes on what I found most interesting:
1) Nanbeige 4.2 3B uses looped depth sharing. This basically means it runs the same 22-layer (=transformer block) stack twice. So, it extends the 22-layer architecture to 44-layers, but without duplicating the weights. (2x the transformer block compute but same memory footprint.)
Why? The info is a bit sparse, but section 2.1 of the Nanbeige 4.2 technical report says two passes gave the best trade-off and retained about 75% of the token efficiency of a standard architecture. More passes gave barely any gains but made the training much slower and much more expensive.
2) Laguna S 2.1 is poolside's Laguna model in a really nice size: 118B sparse MoE with 8B active parameters and a 1M-token context window. Otherwise, the architecture is pretty standard. It uses 36 sliding-window and 12 global (gated-)GQA layers. However, given this size, and the fact that it (just barely) runs on my DGX Spark (uses about <80 GB of RAM), this is right now the most interesting model for me personally. It's 3x bigger and thus a tad slower but maybe a good candidate as daily-driver-Qwen3.6-35B-replacement. (Still waiting on some more independent performance benchmarks though.)
3) Motif-3-Beta is a new 314B-A13B sparse MoE that is somewhat based on DeepSeek V4 in terms of mHC and latent attention. But it uses a new component, Grouped Differential Latent Attention, which is inspired by Multi-head Latent Attention. I probably should write an article about this some time, but for now, the tl;dr is as follows. Regular MLA compresses the keys and values into a smaller latent representation to mainly reduce the KV cache size. GDLA does a similar low-rank compression but puts the attention heads into groups and also learns a noise head for each group where the noise gets subtracted for filtering purposes... Anyway, a topic for another day!
4) Solar Open 2 is a new 250B-A15B hybrid MoE by Upstage that interleaves three Kimi Delta Attention layers with one GQA layer.
5) Antares 1B is a small model (and there is also an even smaller 0.3B variant) from Cisco starts that with the IBM Granite 4.0 1B backbone and uses SFT plus GRPO for terminal-based cybersecurity stuff. It is a nice example of task-specific post-training on a genuinely small model.
6) BTL-3 is a rank-32 LoRA adapter for Qwen3.6-27B aimed at coding agents and structured tool use. The really strong benchmark performance suggests that LoRA adapters are still a useful tool/technique in 2026.
I added all six to the LLM Architecture Gallery for some additional details:
sebastianraschka.com/llm-arc…
Minis 竟然开源了,做的很细致。完全是付费 app 的水平。
引用 Ethan Wang @wsvn53After 3-4 months of iteration, Open Minis — "possibly the best Agent app on your phone" — has stabilized into a mature architecture, now powering tens of thousands of users' daily workflows. Today we're open sourcing all of it: the full iOS and Android code. A real on-device AI agent: a native Linux shell, browser automation, extensible skills, persistent memory, and deep system integration. Now, fork it, feel free to build your own on-device agent now. 🤗 github.com/OpenMinis/OpenMin…查看被引原帖 ↗
Open 24/7!
Fugu-Ultra now works with Claude Code 🐡
引用 Sakana AI @SakanaAILabsAnnouncing the Claude Code-compatible interface for our new Fugu-Ultra v1.1! 🐡 Put a dynamically coordinated team of frontier models to work inside the coding workflow you already know. Instead of relying on a single model to write, debug, and execute your code, you can now orchestrate a diverse pool of state-of-the-art models directly from your terminal. Put the whole school to work on your next task: console.sakana.ai/get-starte… 🐟查看被引原帖 ↗
ngl i miss a lot Karpathy’s voice on X
引用 am.will @LLMJunkyDid Karpathy remove Anthropic from bio? Or was it not there.查看被引原帖 ↗
I’m probably in a minority in the AI space but I don’t buy the narrative that these new models suck for game devs…
I agree with the Take-Two CEO that development itself was never the bottleneck for good games… Actually having good game ideas, making small judgement decisions, telling good stories, and all of the rest of the judgement decisions behind a game are the things that actually make a good game.
We’ve been telling the “anyone can make a good game now” story for a couple years at this point… yet most of what we’ve actually seen is one level demos that no one ever talks about again after the viral “look how cool this is” post disappears into obscurity…
引用 Matt Shumer @mattshumer_IMO it really sucks for those who have put their lives into game dev... but ppl burying their heads in the sand are doing themselves (and others) a disservice. The sooner people are aware of what's coming, the sooner they can position themselves well.查看被引原帖 ↗
There really is nothing like X.
Dennys knows how this social media thing is done!
Taking a break from cyber to chat to Terence and mathematician colleagues on the future of Math and AI in Philadelphia at
#ICM2026
本站由 Jedee杰哥 打造 · 公众号「Jedee杰哥」每早送 AI 日报
姊妹站:𝕏 简中账号数据榜单 · X 关注 @jedeeai · RSS 订阅 · AI 日报 · 历史归档