Skip to content
PodcastsBusinessChain of Thought | AI Agents, Infrastructure & Engineering

Chain of Thought | AI Agents, Infrastructure & Engineering

Conor Bronsdon
Chain of Thought | AI Agents, Infrastructure & Engineering
Latest episode

79 episodes

  • Chain of Thought | AI Agents, Infrastructure & Engineering

    Switching Models Shouldn't Mean Starting Over | Walrus Protocol's Kimberly Logan

    07/10/2026 | 37 mins.
    If your agent's memory lives inside one model provider, switching models or harnesses means starting over. You need portable memory - yet when switching agent memory between models, you can swing accuracy by thirteen points or more, depending on which model reads them back.
    Kimberly Logan, Head of Product at the Walrus Foundation, where she builds Walrus Memory, joins Chain of Thought to discuss how to build portable and verifiable memory for AI agents - and why Walrus's decentralized storage network has proved an apt primitive. She and host Conor Bronsdon get into why model lock-in is a memory problem, what people use agent memory for beyond coding, how you prove what happened to your data, and why Walrus doesn't publish memory quality benchmarks, plus more. This episode is sponsored by Walrus.
    We cover:
    Why longer context windows don't help once you switch models or harnesses, and why an import function only fixes one point in time
    Why Kimberly doesn't benchmark Walrus Memory against memory startups, and why she sees the opportunity as bigger than memory alone
    What people build with Walrus Memory besides coding agents, which Walrus puts at maybe 30% of usage: long-running study aids, trading agents, and security teams comparing production snapshots with code changes
    Walrus's figure that almost 85% of the memories written are read back more than 30 days later
    How attestations on the blockchain and a delegate key you control let you show what happened to your data, and why unchanged data still isn't proof that it's accurate
    Why Kimberly treats recall quality and latency as table stakes, and how a knowledge graph keeps recall on the freshest memories
    Why she left almost 20 years in traditional software, including seven at Google, for a blockchain company, and where shared memory across companies and trust boundaries goes next
    Walrus Console, announced this week: one place to manage agent memories and files through an interface or an MCP connection, without knowing how the storage underneath is configured
    Connect with Kimberly Logan:
    LinkedIn: https://www.linkedin.com/in/krlgn/
    Twitter/X: https://x.com/kimblgn
    Walrus: https://walrus.xyz/cot
    Connect with Chain of Thought host Conor Bronsdon:
    Newsletter: https://newsletter.chainofthought.show/
    Twitter/X: https://x.com/ConorBronsdon
    LinkedIn: https://www.linkedin.com/in/conorbronsdon/
    YouTube: https://www.youtube.com/@ConorBronsdon
    🔗 More episodes: https://chainofthought.show
    Chapters:
    (0:00) Why agent memory has to move with you
    (1:25) Context windows don't follow you to a new model
    (4:41) Why an import function doesn't fix lock-in
    (6:38) Why Walrus isn't competing head-on with memory startups
    (8:39) Where Walrus Memory goes next
    (10:38) Trading agents and security snapshots on Walrus Memory
    (14:44) Verifiability: answering what happened to your data
    (17:21) Fresh recall, and why Walrus doesn't publish memory benchmarks
    (20:22) From traditional software to a blockchain company
    (24:40) A record other companies and regulators can check
    (25:42) What Walrus Console is for
    (29:09) Use agent memory beyond coding agents
    (30:52) Plugging in your own agent stack
    (31:41) Memory across companies and trust boundaries
    (35:48) Closing thoughts
    Thanks to Walrus, presenting sponsor of season four of Chain of Thought. Walrus Memory gives AI agents portable, verifiable memory that carries context across apps, sessions, and other agents. Get started at https://walrus.xyz/cot
    #AI #AIAgents #AgentMemory #SoftwareEngineering #DevTools
  • Chain of Thought | AI Agents, Infrastructure & Engineering

    Your Best AI Engineer Might Have the Worst Metrics | Sonar CTO Andrea Malagodi

    05/10/2026 | 56 mins.
    Andrea Malagodi is CTO at Sonar, which builds AI code verification and governance tools. When finance asked about the size of his team's AI bills, he calculated cost per PR. At one extreme was an engineer with more than 500 PRs in a short period and a low cost per PR. At the other was an engineer with very high spend and very few PRs.
    Judged on cost alone, the first engineer wins and the second looks like a problem. Malagodi looked closer. The first had built a personal agent factory, with specification, design review, documentation records, coding, verification and functional testing, and about two-thirds of the output was tests and validations. The second was working on a hard problem that needed the AI to reason through long, multi-turn sessions and didn't reduce to a line count.
    Numbers alone tell you something, but not which engineer was doing the more valuable work.
    We cover:
    Why cost per PR is becoming the industry's standard AI spend metric, and how Malagodi's 500-PR engineer and high-spend engineer broke it
    The history of bad proxies, from lines of code to story points to token leaderboards, and why Malagodi says the people who removed code were often the most important
    Sonar's guide, verify, solve loop: Vortex for codebase context, algorithmic plus LLM-based PR analysis, a remediation agent for legacy debt, and the Hunter agent for security vulnerabilities
    Why Malagodi breaks agent work into small pieces, and why a five-day session that produces 400,000 lines of code leaves you with no idea what is in your codebase
    Setting up disagreement between agents with an orchestrator and personas for engineering, product management and quality, plus Conor Bronsdon's cross-model-family review setup
    Rolling AI tools out at enterprise scale: one cost view for finance, up to 30% savings from cutting repeated context reads, and why budgets as small as $50 a week are hard to work with
    Rotating a triage duty to protect deep work, including how Sonar's teams built skills to sort through hundreds of reported CVEs, a number of them likely false positives
    New episodes, the ideas behind them and Conor's essays land in the Chain of Thought newsletter first. Subscribe: https://newsletter.chainofthought.show/
    Chapters:
    (0:00) Human reviewers rubber-stamp big AI changes
    (2:43) Learning the craft alongside AI tools
    (9:09) From lines of code to PR counts
    (12:17) Spreading top engineers' setups org-wide
    (16:50) The guide, verify, solve agent loop
    (22:51) Specialized agent lanes and multitasking limits
    (26:51) Clear asks and guardrails for agent coordination
    (30:29) Building disagreement between agents
    (34:59) Mixing models instead of picking one
    (39:46) Rolling out AI tooling at enterprise scale
    (46:19) Defense in depth for AI-written code
    (50:52) Protecting time for incident investigations
    (54:08) Personalized software for teams and individuals
    Links from the episode:
    Sonar Vortex research: https://www.sonarsource.com/blog/cut-your-coding-agents-cost-with-sonar-semantic-code-navigation/
    SonarQube Hunter Agent: https://www.sonarsource.com/products/sonarqube/hunter-agent/
    Connect with Andrea Malagodi:
    LinkedIn: https://www.linkedin.com/in/malagodia/
    Sonar: https://www.sonar.com/
    Connect with Chain of Thought host Conor Bronsdon:
    Newsletter: https://newsletter.chainofthought.show/
    Twitter/X: https://x.com/ConorBronsdon
    LinkedIn: https://www.linkedin.com/in/conorbronsdon/
    YouTube: https://www.youtube.com/@ConorBronsdon
    More episodes: https://chainofthought.show

    Thanks to Svix, presenting sponsor of season four of Chain of Thought. Svix delivers billions of reliable webhooks for startups and the Fortune 500. Get started at https://link.svix.com/cot. Qualified startups get $12,000 in credits, and YC companies get $50,000.
    Thanks to Walrus, presenting sponsor of season four of Chain of Thought. Walrus Memory gives AI agents portable, verifiable memory that carries context across apps, sessions, and other agents. Get started at https://walrus.xyz/cot

    Thanks to Inngest, presenting sponsor of season four of Chain of Thought. Agents in production run long - they call models and wait on APIs and people. But the longer agents run, the more they break. Inngest handles that with durable execution. You build your agent as steps in TypeScript, Python, or Go. When a step fails, Inngest retries it with exponential backoff, and completed steps are saved and skipped. Try it out: https://inngest.link/cot-pod
    Thanks to G2i for sponsoring this episode - for over a decade, they vetted and placed engineers at other companies, from startups to FAANG. Two years ago, they turned that same judgment inward, building their own bench to review RL environments, evals, and training data that models are trained on. Get access: https://fandf.co/3SFxVm6
  • Chain of Thought | AI Agents, Infrastructure & Engineering

    Why Context Alone Isn't Enough for Enterprise AI Agents | WisdomAI CPO Kapil Chhabra

    24/09/2026 | 1h 23 mins.
    An AI agent answering questions about your business needs to understand how that business works. How do you calculate churn? When does your fiscal year end? Who has the authority to change those definitions? Connecting a model to company data leaves those questions unresolved.
    That’s the work of context engineering: giving agents the business definitions, instructions and examples they need to interpret your data. But that context changes as the business evolves, and someone has to keep it accurate.
    In this episode of Chain of Thought, WisdomAI co-founder and CPO Kapil Chhabra joins Conor Bronsdon to explain how his team approaches that challenge. We explore how companies maintain shared context, why data teams are taking on the role of AI context engineers, and how a specialized harness plans queries, checks results and repairs errors before returning an answer.
    Recorded at WisdomAI’s San Mateo office, this episode is sponsored by WisdomAI.
    We cover:
    Why the context layer is broader than a semantic layer or catalog, and how it differs from memory
    How context drifts, and why a subject-matter expert has to approve changes to shared definitions
    Why Kapil sees an AI context engineer role emerging, and why data teams are moving from providing insights to providing context
    How WisdomAI's harness decomposes a question, federates queries across data sources and repairs errors, including a customer that replaced a $5M-a-year analytics pipeline
    The four ingredients of a trustworthy AI answer: accuracy, consistency, governance and explainability
    Live Apps: governed analytics apps built from a single prompt, and what keeps them live
    Why your context is your IP and should stay portable
    Chapters:
    (0:00) Do your agents have the right context?
    (1:56) The four ingredients of trust
    (5:13) The criticality and impact 2x2
    (8:12) Data, context, harness: the hospital analogy
    (11:12) What the context layer actually means
    (11:48) Specialized harnesses: legal, support, analytics
    (13:17) Why only 7% of data leaders have scaled AI
    (15:46) What models can't guess: ARR, churn, fiscal years
    (16:45) The data stack collapses into the context layer
    (20:55) Memory vs. context
    (25:26) Are agents the new users of software?
    (27:19) Where humans should spend their time
    (28:14) Commissioning an AI agent, and who verifies it
    (31:28) Context drift and the learning loop
    (34:04) Context is a multiplayer game
    (35:12) Decompose, query, verify, repair
    (38:47) Replacing a $5M analytics pipeline with federation
    (42:04) The context development life cycle
    (43:29) The AI context engineer
    (46:02) Jobs are changing, not disappearing
    (46:59) Product, people and process
    (51:20) Who decides? Why FDEs can't own your context
    (52:22) Data context vs. business context
    (54:11) The benchmark: specialized harness vs. general agent
    (56:08) Meeting users in ChatGPT, Claude and Slack
    (58:48) Static vs. runtime context
    (1:00:08) Harness engineering as models change
    (1:01:44) Right-sizing AI and Live Apps
    (1:04:42) The boring parts: governance, security, caching
    (1:06:18) 1,000 dashboards, 50 human-years
    (1:08:08) What "live" means
    (1:09:19) Are dashboards going away?
    (1:12:08) A pipeline app built on a weekend walk
    (1:15:53) Who owns the apps?
    (1:17:27) Data teams now provide context, not insights
    (1:18:56) Building with the WisdomAI MCP
    (1:19:43) Your context is your IP
    (1:20:48) Closing thoughts: none of that work goes to waste
    Links from the episode:
    Meet the Modern Data Team (WisdomAI CDO report)
    AI Context Engineer (ACE) certification
    Live Apps
    WisdomAI in ChatGPT Work
    Avoid AI Writing
    ssot-check
    I Paid an AI Agent $8 to Write About its 'Life'
    Slack Wants to Be the Context Harness for Code | CPO Jaime DeLanghe
    The AI Framework Era Is Over: Why Context Is the Moat | Jerry Liu
    Connect with Kapil Chhabra:
    LinkedIn
    WisdomAI
    WisdomAI on X
    Connect with Chain of Thought host Conor Bronsdon:
    Newsletter
    Twitter/X
    LinkedIn
    YouTube
    More episodes: https://chainofthought.show
    Thanks to WisdomAI for sponsoring this episode. WisdomAI is the agentic analytics platform for trusted enterprise intelligence: governed context, an analytics harness that makes every answer consistent and verifiable, and Live Apps built from a single prompt. Try Live Apps: https://wisdom.ai/liveapps
  • Chain of Thought | AI Agents, Infrastructure & Engineering

    AI Codes: Product Engineers Decide What to Build | Laurie Voss, Arize

    22/09/2026 | 56 mins.
    Laurie Voss co-founded npm - now head of developer relations at Arize, he argues that engineers will increasingly earn their keep as 'product engineers': understanding what users need and directing AI agents to build it.
    One example: a bakery owner who knows how to make a croissant but has no interest in building software. Someone still has to turn that owner's needs into requirements. Laurie sees that work becoming central to product engineering, with cheaper code making software for narrower industries more viable.
    We discuss where he still sees a need for human code review and operational knowledge, what he would look for in a computer science course if he were starting out today (and what he wouldn't do), and why he compares AI today to the web in 1997. He is optimistic about the technology and skeptical of the valuations, while leaving one question unresolved: how do junior engineers learn the judgment this work demands?
    We cover:
    Why the "aha moment" of solving a problem survives even when agents type the code
    Where Laurie still sees a need for human code review, operations, and tacit knowledge
    Whether a $15,000 coding bootcamp or a theory-heavy CS degree is still worth it
    Why AI in 2026 looks like the web in 1997, and what that says about the bubble
    How AI-generated pull requests burden open source maintainers, and why Laurie expects cheaper code to pressure closed-source business models
    Why the systems analyst returns as the product engineer, and why that means niche software for bakeries and auto parts
    Why Laurie predicts open-model competition and diminishing returns could compress frontier-model margins
    How apprenticeships could help junior engineers develop product judgment
    Chapters:
    (0:00) The shift in software jobs
    (0:44) Why Laurie is optimistic about the code generation explosion
    (3:06) The aha moment moves from typing code to thinking
    (5:49) Where agents still need human review and operational knowledge
    (8:51) Is college still worth it?
    (9:54) Bootcamps versus theory-heavy CS courses
    (13:04) We are all product engineers now
    (15:55) AI is the web in 1997
    (18:19) Exponential growth, the labs' pause, and npm's ten-year curve
    (20:07) Barring AGI, AI is a normal technology
    (21:48) Block's layoffs and companies staying smaller
    (23:14) What the labor data shows: fewer people, more capital
    (26:29) Open source as the canary: drowning in AI pull requests
    (28:11) AI reimplementations and the pressure on software moats
    (30:26) Personal software and the kill-my-SaaS hackathon
    (32:32) The bakery and the return of the systems analyst
    (34:09) Niche software for specific industries
    (35:36) Bootstrapping and the DevTools opportunity
    (37:13) What this means for the model companies
    (38:20) Frontier-model margins and open-model competition
    (39:53) How the bubble pops: scaling laws and diminishing returns
    (43:24) Staying private and the trough of disappointment
    (45:51) Get good at a domain, not the technology
    (50:44) The missing junior ladder is the question of our time
    (53:32) Closing thoughts: it's 1997, you can retrain
    Connect with Laurie Voss:
    Blog: https://seldo.com/
    LinkedIn: https://www.linkedin.com/in/seldo/
    Twitter/X: https://x.com/seldo
    Bluesky: https://bsky.app/profile/seldo.com
    Arize: https://arize.com/
    Connect with Chain of Thought host Conor Bronsdon:
    Newsletter: https://newsletter.chainofthought.show/
    Twitter/X: https://x.com/ConorBronsdon
    LinkedIn: https://www.linkedin.com/in/conorbronsdon/
    YouTube: https://www.youtube.com/@ConorBronsdon
    More episodes: https://chainofthought.show
    Thanks to Svix, presenting sponsor of season four of Chain of Thought. Svix delivers billions of reliable webhooks for startups and the Fortune 500. Get started at https://link.svix.com/cot. Qualified startups get $12,000 in credits, and YC companies get $50,000.
    Thanks to Walrus, presenting sponsor of season four of Chain of Thought. Walrus Memory gives AI agents portable, verifiable memory that carries context across apps, sessions, and other agents. Get started at https://walrus.xyz/cot
    Thanks to G2i for sponsoring this episode - for over a decade, they vetted and placed engineers at other companies, from startups to FAANG. Two years ago, they turned that same judgment inward, building their own bench to review RL environments, evals, and training data that models are trained on. Get access: https://fandf.co/3SFxVm6
    Thanks to Inngest, presenting sponsor of season four of Chain of Thought. Agents in production run long - they call models and wait on APIs and people. But the longer agents run, the more they break. Inngest handles that with durable execution. You build your agent as steps in TypeScript, Python, or Go. When a step fails, Inngest retries it with exponential backoff, and completed steps are saved and skipped. Try it out: https://inngest.link/cot-pod
    7Mdv8UhrombOD7fnCVai
  • Chain of Thought | AI Agents, Infrastructure & Engineering

    Genspark's Bet: AI Agents Become the Users of Software | Wen Sang

    17/09/2026 | 1h 1 mins.
    Genspark went from launch to $250 million in ARR in about a year. Along the way it shipped a card-thin meeting recorder, open sourced an office suite that Wen Sang says one engineer prototyped in a week, and started running product triage with agents instead of product managers.
    Wen's bet is that agents, not people, become the next users of software.
    Wen Sang is co-founder and COO of Genspark. In this episode he walks through the company's three-layer architecture (models, tools and premium data as the execution layer, a memory layer he calls the second brain, and a collaboration layer called Gen Team), why a meeting note should be the start of work rather than the end of it, the engineering behind the SecondBrain Note, and where he thinks knowledge work goes once agents absorb the busy work.
    Disclosure: Genspark provided the SecondBrain Note recorder discussed in this episode at no cost. Genspark is not a sponsor of this episode.
    We cover:
    Why Genspark builds the self-driving car around the frontier labs' engines, and what that means for people who cannot code
    How Genspark's mixture-of-agents architecture routes work across 70+ models, 150+ in-house tools and paid data sets
    Evals that grade whether the sales proposal answered the RFP, not whether the model can solve a differential equation
    What a meeting turns into a week later when an agent needs it: proposals, pricing models, research, follow-ups
    The SecondBrain Note's microphone array, battery decisions, and consent in a two-party state
    Why GenOffice went open source, and the one-week prototype story behind it
    How Genspark runs product feedback triage with agents and no dedicated PMs
    Chapters:
    (0:00) Cold open
     (0:32) Geniuses with goldfish memories
     (3:47) Engines and vehicles: Genspark builds the self-driving car
     (6:25) Mixture of agents: models, in-house tools and premium data
     (10:03) Grade the work output, not the intelligence
     (11:22) A meeting note is where the work starts
     (13:25) The second brain: Genspark's memory layer
     (14:45) A thousand recorders, one question for the revenue team
     (16:33) Execution, memory and collaboration layers
     (19:19) Ten days in Bora Bora without a laptop
     (20:07) Gen Team, Slack, and meeting customers where they are
     (21:57) Agents become the users of software
     (24:17) Keeping memories current when the deal changes
     (26:41) Engineering the SecondBrain Note
     (30:18) The note as an API for the room
     (31:19) What deserves hardware and what stays software
     (33:33) Learning hardware supply chains at a two-year-old company
     (35:05) Why GenOffice went open source
     (38:01) What knowledge workers do once the busy work is gone
     (40:07) Building on Genspark with the CLI
     (42:01) Consent, two-party states and the surveillance line
     (43:45) Genspark Claw
     (47:09) Cheaper hardware, deeper integration, and model welfare
     (51:13) Eighty people and a lot of agents
     (53:03) What 2027 looks like
     (59:26) Where to find Wen, and product triage without PMs
    Connect with Wen Sang:
    LinkedIn: https://www.linkedin.com/in/wen-sang/
    Twitter/X: https://x.com/sang_wen
    Genspark: https://www.genspark.ai
    GenOffice on GitHub: https://github.com/genspark-ai/genoffice
    Connect with Chain of Thought host Conor Bronsdon:
    Newsletter: https://newsletter.chainofthought.show/
    Twitter/X: https://x.com/ConorBronsdon
    LinkedIn: https://www.linkedin.com/in/conorbronsdon/
    YouTube: https://www.youtube.com/@ConorBronsdon
    More episodes: https://chainofthought.show
    Thanks to Walrus, presenting sponsor of season four of Chain of Thought. Walrus Memory gives AI agents portable, verifiable memory that carries context across apps, sessions, and other agents. Get started at https://walrus.xyz/cot
    Thanks to Svix, presenting sponsor of season four of Chain of Thought. Svix delivers billions of reliable webhooks for startups and the Fortune 500. Get started at https://link.svix.com/cot. Qualified startups get $12,000 in credits, and YC companies get $50,000.
    Thanks to Inngest,  presenting sponsor of season four of Chain of Thought.  Agents in production run long - they call models and wait on APIs and people. But the longer agents run, the more they break. Inngest handles that with durable execution. You build your agent as steps in TypeScript, Python, or Go. When a step fails, Inngest retries it with exponential backoff, and completed steps are saved and skipped. Try it out: https://inngest.link/cot-pod
    Thanks to G2i for sponsoring this episode - for over a decade, they vetted and placed engineers at other companies, from startups to FAANG. Two years ago, they turned that same judgment inward, building their own bench to review RL environments, evals, and training data,that  models are trained on. Get access: https://fandf.co/3SFxVm6
More Business podcasts
About Chain of Thought | AI Agents, Infrastructure & Engineering
AI is reshaping infrastructure, strategy, and entire industries. Chain of Thought is the podcast where builders reason through what's changing. Host Conor Bronsdon sits down with the engineers and founders shipping AI in production to get past the hype into what's working and what isn't. Episodes cover model infrastructure, inference, agent frameworks, evaluation, and developer tools. Guests have come from NVIDIA, Google DeepMind, AMD, Databricks, Vercel, and more. Every episode carries a full transcript and show notes at chainofthought.show. New episodes weekly. Conor Bronsdon is an independent consultant and angel investor in AI infrastructure and developer tools. He led technical ecosystem at Modular, acquired by Qualcomm in 2026; led developer awareness at Galileo, acquired by Cisco; and ran developer marketing at LinearB, where he was GM of the Dev Interrupted podcast and community. Views expressed by the host and guests are their own.
Podcast website

Listen to Chain of Thought | AI Agents, Infrastructure & Engineering, Get Started Investing and many other podcasts from around the world with the radio.net app

Get the free radio.net app

  • Stations and podcasts to bookmark
  • Stream via Wi-Fi or Bluetooth
  • Supports Carplay & Android Auto
  • Many other app features
Chain of Thought | AI Agents, Infrastructure & Engineering: Podcasts in Family