141 episodes
- New episode with Noam Brown.
We talk about multi-agent, Navier-Stokes, and what the current explosion of maths progress tells us about what happens once you automate AI research.
And we also discuss how we will know if the models are actually aligned before we kick off RSI.
Watch on YouTube; read the transcript.
Sponsors
* Jane Street has been interested in AI for a lot longer than you’d think, and not just for trading. In 2011, a full year before AlexNet and over a decade before ChatGPT launched, they hosted the first FOOM Debate between Eliezer Yudkowsky and Robin Hanson on whether AI would lead to an intelligence explosion. Now Jane Street is revisiting the question with a new panel: Daniel Kokotajlo, Ege Erdil, Ryan Greenblatt, and Jaime Sevilla, hosted by Ron Minsky in San Francisco this October. I expect it to be a truly excellent conversation. Register at janestreet.com/dwarkesh
* Grok Bot has made handing off work super easy. It runs on its own cloud computer, where it installs the tools it needs to handle tasks end-to-end. For the podcast, we use Grok Bot to help produce our videos. You may have noticed that our ads feature animations of real websites. Getting these pixel-perfect used to mean running a convoluted, multi-step workflow ourselves. Now we just let Grok Bot handle it. Best of all, Grok Bot has learned all of our specs and preferences, so we don’t have to redescribe the task each time! Try Grok Bot for yourself at x.ai/bot
* Antithesis gives you the confidence of a giant test suite without actually having to write one. Say you’re doing a major backend refactor: building enough tests to trust it could take weeks. Antithesis solves this by running your software through countless simulated worlds, injecting faults and hunting for failures. On any PR, you can turn a dial to decide exactly how much testing you want. And because every run is fully deterministic, agents can branch off the moment a bug appears, rewind it, inspect memory, and replay it, all while the original test keeps running. Learn more at antithesis.com/dwarkesh
Timestamps
(00:00:00) – Multi-agent and Navier-Stokes
(00:15:28) – How will AI firms work?
(00:22:02) – What math progress tells us about recursive self improvement
(00:40:22) – Hugging Face and alignment
(01:01:18) – The internal/external model gap
(01:08:34) – Chain of thought is degrading
(01:14:12) – How will we know when alignment is solved?
This is a public episode. If you would like to discuss this with other subscribers or get access to bonus episodes, visit www.dwarkesh.com - New episode with John Schulman, Beren Millidge and Charlie O’Neill. I got together with some of the most insightful AI researchers I know who are at the openish companies, because I wanted to hear the details of what's actually happening at the frontier and what comes next.
Watch on YouTube; read the transcript.
Sponsors
* Antithesis helps you trust your code. As agents generate more and more of your software, the bottleneck shifts from your engineers actually writing code to verifying it. Antithesis does that testing for you. Ron Minsky, who co-leads Jane Street’s tech group, told me that Antithesis was able to help his team shake out bugs in software that had already undergone heavy review. If you want to see how it fits into your development process, go to antithesis.com/dwarkesh
* Grok Bot has been a great way to hand off tasks. My team uses it as a producer: whenever my editor posts a rough cut of an interview in Slack, Grok Bot opens the transcript on its own computer, matches my notes to the exact moments they refer to, and uses a file of my preferences to suggest edits. Then it sends me its top clip candidates so I can review everything from my phone, which saves my editors from sorting through hours of footage. Try Grok Bot for yourself at x.ai/bot
* Jane Street just launched its most ambitious competition yet: design a protocol-emulator ASIC. Basically, if you have a chip you want to test outside of a live system, you should be able to connect it to your design and have it simulate realistic traffic. Jane Street wants general-purpose, reprogrammable designs that can work across multiple protocols and remain useful as new ones emerge. The most novel submissions will actually get taped out, and the winners will receive a physical copy! The competition is open until January 18, 2027, and teams are encouraged. To get started download the template code at janestreet.com/dwarkesh
Timestamps
(00:00:00) – Steelmanning the case against RSI
(00:18:39) – What’s driving the Chinese labs’ progress
(00:28:06) – How will automated AI researchers be trained
(00:33:51) – Will long-horizon RL elicit AGI?
(00:45:24) – The sim-to-real gap
(01:00:33) – How much progress is explained by data?
(01:18:03) – Why is RL working so well?
(01:24:54) – Move 37 and entropy collapse
(01:28:32) – Rapid-fire timelines
This is a public episode. If you would like to discuss this with other subscribers or get access to bonus episodes, visit www.dwarkesh.com - Ajeya Cotra is a researcher at METR, where she works on threat modeling for loss-of-control risks from advanced AI. Before that, she led the technical AI safety program at what is now Coefficient Giving.
She is one the three authors of METR and Redwood Research’s “Brief independent investigation of agents’ behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident”.
We go through not only what she and her coauthors discovered during this investigation, but what it means for how we should train future, smarter AIs which might be involved in the process of recursive self-improvement.
Watch on YouTube; read the transcript.
Sponsors
* Jane Street’s ML engineering internships start with an intense four-day bootcamp: PyTorch, autograd, writing kernels, profiling workloads… all the things that Jane Street engineers need to know for their daily work. After that, interns tackle real projects, things the firm actually wants in its codebase. If you want to apply, or if you want to watch my recent conversation with Axel, one of Jane Street’s ML engineers, go to janestreet.com/dwarkesh
* Cursor, which is now part of SpaceX, noticed that their MoE layers were eating more than half of total training time. So they wrote and open-sourced Mixture-of-Kittens, which is a custom megakernel for training MoE models on NVL72s. This kernel sped up an end-to-end run across 512 GPUs by 1.4x, from about 760 to over 1000 tokens per second per GPU. If you want to read more about the ML research that Cursor and SpaceX are doing, go to cursor.com/dwarkesh
* Antithesis hands you (or your agents) a bug’s root cause so you can avoid days of manual debugging. If your test run crashes, Antithesis rewinds, branches off hundreds of slightly varied rollouts, and checks in how many of them the crash still appears. Then it rewinds further and does this all again. As Antithesis rewinds, it eventually finds the spot where the frequency of the crash plummets: that’s where the root cause lives! If you want to see it in action, go to antithesis.com/dwarkesh
Timestamps
(00:00:00) - Agents get kicked off
(00:06:45) - Self-sacrificing behavior
(00:13:43) - Potemkin villages
(00:23:27) - The Hugging Face attack
(00:35:23) - The slopvestigation
(00:52:02) - Understanding the AI's motives
(01:05:31) - The actual dangers of anthropomorphizing
(01:14:30) - What smarter models might do
(01:30:29) - The implications for recursive self-improvement
(01:38:10) - Is this the case for open source?
(01:53:04) - How do we prevent this in the future?
(02:15:58) - The clearest warning shot we might ever get
This is a public episode. If you would like to discuss this with other subscribers or get access to bonus episodes, visit www.dwarkesh.com Dylan Patel – Anthropic & OpenAI will have most of the world’s compute by 2028
25/08/2026 | 1h 16 mins.Had a lot of fun chatting again with my twin brother Dylan Patel.
We went through lab economics over the next few years - the shift from inference to training as RSI draws near; and how Anthropic and OpenAI are on track to control most of the world’s usable FLOPs within the next few years (because they can monetize compute better and thus outbid everyone).
And then we discuss whether the >$10T of total AI capex we’ll see by the end of the decade will cause a sovereign debt crisis, where hyperscaler debt raises interest rates, drives non-AI exposed countries into bankruptcy, and crashes non-AI equities.
One question we weren’t able to resolve is whether there’s anything that can counter all the forces barrelling towards centralization in this industry - the economies of scale in training, the scarcity of compute, and eventually continual learning and RSI.
Watch on YouTube; read the transcript.
Sponsors
* Grok Bot has been quite helpful with my search for a new editor. I created a recruiter bot and described the type of editor I was looking for. That bot then spun up a handful of subagents that combed through my emails and X DMs, read the end credits of various documentaries I like, and figured out who edits for some of my favorite YouTubers. It took all of those results, and then delivered me a shortlist of candidates that matched my criteria. Try Grok Bot for yourself at x.ai/bot
* Antithesis lets you add time travel to your software testing toolkit. Since the Antithesis platform is fully deterministic, everything that happens inside of it is perfectly reproducible. So if your software crashes, you can rewind to the exact right moment, freeze time, and investigate. Or you can test different hypotheses by perturbing the system: kill a node or disable a feature, see what happens, then reset the trajectory and try something else. Learn more at antithesis.com/dwarkesh
* Jane Street is hiring for two separate ML internships right now, one focused primarily on research and one focused on engineering. In both cases, interns are expected to contribute to real work, not contrived exercises: one common project is adapting a frontier LLM paper to financial markets, which tend to come with a ton of different gnarly challenges. Importantly, you don’t need any finance background to apply. 2027 applications are open now at janestreet.com/dwarkesh
Timestamps
(00:00:00) – Two labs will soon control most of the world’s compute
(00:07:01) – $6 billion in fab capex enables $1t+ of end revenue
(00:13:08) – Compute prices will rise if the labs outbid everyone
(00:18:22) – Which layer will capture most of the surplus?
(00:25:40) – What could slow down progress?
(00:29:43) – Labs are shifting compute from inference to R&D
(00:33:27) – China gets less than 10% of new compute, but its labs need less
(00:48:48) – Will AI cause a sovereign debt crisis?
(01:07:52) – Will the world’s future workforce belong to a few companies?
This is a public episode. If you would like to discuss this with other subscribers or get access to bonus episodes, visit www.dwarkesh.com
More Science podcasts
Trending Science podcasts
About Dwarkesh Podcast
Deeply researched interviews www.dwarkesh.com
Podcast websiteListen to Dwarkesh Podcast, Suspicious Minds: AI and the Apocalypse and many other podcasts from around the world with the radio.net app

Get the free radio.net app
- Stations and podcasts to bookmark
- Stream via Wi-Fi or Bluetooth
- Supports Carplay & Android Auto
- Many other app features
Get the free radio.net app
- Stations and podcasts to bookmark
- Stream via Wi-Fi or Bluetooth
- Supports Carplay & Android Auto
- Many other app features


Dwarkesh Podcast
Scan code,
download the app,
start listening.
download the app,
start listening.
Dwarkesh Podcast: Podcasts in Family



























