Your agentic app just ran a search. The tool returned 500 results as JSON. Your agent appended all of it and fired off an API call — 45,000 tokens to answer a question that needed maybe 4,500. Tejas Manohar, a senior engineer at Netflix, hit this problem every day. He was running out of tokens […]

Read More →

Most of the text flowing through an agent’s context window isn’t code, reasoning, or instructions. It’s logs. Table of contents The problem nobody talks about Here’s something I’ve been noticing while watching AI coding agents work. You ask Cursor or Claude Code to fix a failing test. The agent runs the test suite. The test […]

Read More →

Large Language Models (LLMs) all predict text, but they differ a lot in how they follow instructions, use context, handle tools, and optimize for safety, speed, or cost. If you treat them as interchangeable, you’ll ship brittle prompts. If you treat them as different runtimes with different affordances, you’ll get reliable results. This post explains the major differences across […]

Read More →