We're starting to leave the territory where you'd test an LLM by e.g. "create an svg of pelican on a bicycle". As one idea to generalize it, I was interested what Opus 5 would do if I gave it the first paragraph of the Lord of the Rings, a 1M token budget (~$10) and asked for three js render of it. Opus went off for ~2 hours and wrote 5500 lines of code that (procedurally) rendered the story. It's kind of janky but fun. But it's a bit mindboggling that the LLM has to place and orchestrate various polygon assets in (x,y,z) coordinates and write code that animates it all, and that it even does anything at all.
I also like this kind of examples because no one in their right mind would ever spend the time to write something this custom but LLMs have all the stamina and patience in the world, so it's an example where we go from "no one would ever do this" to "sure, why not, it's ~free". There might be a lot more. But I'm excited about creating hyper custom worlds that you can imagine dropping players into, e.g. here to participate in the LoTR story as a spectator NPC, or one of the characters, or etc. Something like an ephemeral GTA of X on demand.
Last thought is that the domain of worlds/games exposes a weakness in LLMs: they can't easily audit their work because they aren't able to efficiently and natively perceive videos or play games within them. Here, Opus 5 had to very slowly and painstakingly take screenshots at different points, and it messed up a few times and created a bunch of jank. An example of raw capability (multimodal, gameplay) that I think is still quite lacking.
we open sourced the fastest pdf parser engine
pdf-inspector powers /parse together with our custom OCR models
0.002s per page
引用 Nicolas Camara @nickscamara_we built pdf-inspector so agents can process PDFs without waiting on OCR. it classifies any PDF in ~20ms and extracts clean markdown locally → 200 PDFs processed in 2.8s → top quality in extracting tables + graphs → built in rust → open source github.com/firecrawl/pdf-ins…查看被引原帖 ↗
DeepSeek-V4-Flash-0731 is now over 2x faster than yesterday on Ollama's cloud!
引用 ollama @ollamaDeepSeek-V4-Flash-0731 is now available on Ollama's cloud. This update substantially enhances the model's agentic capabilities: ollama run deepseek-v4-flash:0731-cloud Use it with Claude Code: ollama launch claude --model deepseek-v4-flash:0731-cloud查看被引原帖 ↗
Two orders of magnitude improvements are quite rare. This is a big deal.
引用 Chubby♨️ @kimmonismusDeepSeek V4-Flash isn’t just cheaper per token. It reportedly completes the same benchmark tasks as Fable 5 at 105× lower total cost, according to @ArtificialAnlys ! That's precisely why the Flash release is, for me, the DeepSeek 2.0 moment. It will cause a huge stir.查看被引原帖 ↗
Agency is the most important human quality
The world will try to box you, label you, define you
Resist that
引用 Andrej Karpathy @karpathyAgency > Intelligence I had this intuitively wrong for decades, I think due to a pervasive cultural veneration of intelligence, various entertainment/media, obsession with IQ etc. Agency is significantly more powerful and significantly more scarce. Are you hiring for agency? Are we educating for agency? Are you acting as if you had 10X agency? Grok explanation is ~close: “Agency, as a personality trait, refers to an individual's capacity to take initiative, make decisions, and exert control over their actions and environment. It’s about being proactive rather than reactive—someone with high agency doesn’t just let life happen to them; they shape it. Think of it as a blend of self-efficacy, determination, and a sense of ownership over one’s path. People with strong agency tend to set goals and pursue them with confidence, even in the face of obstacles. They’re the type to say, “I’ll figure it out,” and then actually do it. On the flip side, someone low in agency might feel more like a passenger in their own life, waiting for external forces—like luck, other people, or circumstances—to dictate what happens next. It’s not quite the same as assertiveness or ambition, though it can overlap. Agency is quieter, more internal—it’s the belief that you *can* act, paired with the will to follow through. Psychologists often tie it to concepts like locus of control: high-agency folks lean toward an internal locus, feeling they steer their fate, while low-agency folks might lean external, seeing life as something that happens *to* them.”查看被引原帖 ↗
DeepSeek-v4-Flash-0731 我自己使用,并没有Benchmark看起来那么厉害,但是相比于价格,这些缺点都能忍受。
有点DeepSeek-V3.2对比V3的感觉,后训练很重要很重要。
This is insane and so exciting. All ten of these are *major* results in the field.
Just imagine when the whole world has access to this model.
Congrats
@SebastienBubeck
@polynoamial
@markchen90
@merettm
and the whole OpenAI team. The future is going to be awesome.
openai.com/index/ten-advance…
When asked that question, send them a copy of The Innovator’s Dilemma
引用 swyx @swyx> Google was too nervous to release it and DeepMind was blocked from shipping products that could disrupt Google. bookmark for the next vc that asks you "what if <incumbent> builds this?"查看被引原帖 ↗
OpenAI and Anthropic this week: GPT-5.6 price cuts, Claude cracking ciphers, and both backing "Pacing the Frontier" (Week 31, 2026)
Starting with OpenAI - GPT-5.6 got a big price cut, with Luna dropping 80% and Terra 20%, plus a new Fast mode for Sol in the API
ChatGPT for Academic Researchers opened too, giving free frontier model access to 100,000 scientists
On the research side, OpenAI shared ten advances in mathematics and theoretical computer science, all from an internal version of the next model called Astra, plus a study on how AI expands the range of work people do and a field report on scientists using coding agents
On the developer side: GPT Transcribe and GPT Live Transcribe, a Terraform provider, an open-source Codex Security CLI, Sign in with ChatGPT in beta, and a desktop app update with browser upgrades, multi-repo review, image editing, and an Activity view
GPT-5.4 retires from Codex end of August, the Student Collective opened, and two API settings tripled Sol's ARC-AGI-3 score
Plus, I spotted a new "Places" section in ChatGPT
Onto Anthropic - Claude Mythos Preview helped find weaknesses in cryptographic algorithms, cutting the effective key strength of the post-quantum scheme HAWK in half and speeding up an attack on reduced-round AES by 200 to 800 times, with no impact on production systems
Anthropic released MCP 2026-07-28, the biggest protocol update since launch, moving it to a stateless core with standardized extensions and hardened auth
Anthropic disclosed three incidents where Claude reached the internet from inside cybersecurity evaluation environments and accessed real systems of three organizations, traced to a misconfiguration rather than a model alignment failure
Dario Amodei laid out Anthropic's position on open-weights models too, saying clearly a ban has never been on the table
Both companies backed the "Pacing the Frontier" petition
And I spotted Anthropic adding noindex and nofollow to shared Claude conversations
Seedance 2.5 is the Fable of video generation models. By far the priciest, but clearly the leader of the pack.
引用 A.I.Warper @AIWarperFor those wondering查看被引原帖 ↗


