<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>Pratik Patel</title>
    <link>https://pratik.pa.tel/blog/</link>
    <description>Notes on AI agents, engineering leadership, and building software when the code is no longer the hard part.</description>
    <language>en-us</language>
    <copyright>© 2026 Pratik Patel</copyright>
    <lastBuildDate>Tue, 08 Sep 2026 12:00:00 GMT</lastBuildDate>
    <atom:link href="https://pratik.pa.tel/rss.xml" rel="self" type="application/rss+xml" />
    <item>
      <title>$581 Billion In, Single Digits Out</title>
      <link>https://pratik.pa.tel/blog/581-billion-in-single-digits-out/</link>
      <guid isPermaLink="true">https://pratik.pa.tel/blog/581-billion-in-single-digits-out/</guid>
      <description>Nearly every enterprise has bought into AI. Almost none are running agents at scale. The distance between those two facts is the most useful number a founder can hold this year.</description>
      <dc:creator>Pratik Patel</dc:creator>
      <category>ai</category>
      <category>agents</category>
      <category>economy</category>
      <category>adoption</category>
      <content:encoded><![CDATA[<p class="text-muted-foreground leading-7 my-4">In 2025, total AI-related investment reached <strong class="text-foreground font-bold">$581.69 billion</strong> — a 129.9% jump over the year before, and roughly forty times what it was in 2013. In the same year, when <a href="https://hai.stanford.edu/ai-index/2026-ai-index-report/economy" class="text-primary underline underline-offset-2 decoration-primary/40 print:decoration-primary hover:decoration-primary transition-colors">Stanford&#x27;s AI Index</a> asked organizations how much they actually used AI agents, the most common answer, across most business functions, was none.</p>
<p class="text-muted-foreground leading-7 my-4">Those two facts belong in the same sentence, because the distance between them is the most useful thing a founder can hold in their head right now. Almost everyone has bought in. Almost nobody is running agents at scale. Depending on where you sit, that is the bubble thesis or the opportunity thesis — and the point of this post is to give you the numbers to decide, plus enough survey literacy to notice when two credible reports say opposite-sounding things.</p>
<h2 id="where-the-money-went" class="font-display text-2xl lg:text-3xl font-bold text-foreground mt-12 mb-6 border-l-2 border-primary pl-4"><a href="#where-the-money-went" class="heading-permalink text-inherit no-underline">Where the Money Went</a></h2>
<p class="text-muted-foreground leading-7 my-4">Start with the $581.69 billion, because it is the number everyone quotes and almost nobody reads carefully. It is not a corporate AI budget. The AI Index compiles it from mergers and acquisitions, minority stakes, private investment, and public offerings — the flow of capital <em class="font-mono font-normal text-primary/80 print:text-primary">toward</em> AI, not the amount spent <em class="font-mono font-normal text-primary/80 print:text-primary">deploying</em> it. Private investment alone was <strong class="text-foreground font-bold">$344.66 billion</strong>, up 127.5%; M&amp;A activity rose 132.6%. However you slice it, the money is real and it is accelerating.</p>
<p class="text-muted-foreground leading-7 my-4">Adoption looks just as emphatic. In the same body of surveys, <strong class="text-foreground font-bold">88% of organizations</strong> reported using AI in at least one part of the business, and <strong class="text-foreground font-bold">70%</strong> reported using generative AI in at least one function. If you stopped reading there, you would conclude the transformation is essentially complete.</p>
<p class="text-muted-foreground leading-7 my-4">Then you get to the agents.</p>
<h2 id="where-it-didnt" class="font-display text-2xl lg:text-3xl font-bold text-foreground mt-12 mb-6 border-l-2 border-primary pl-4"><a href="#where-it-didnt" class="heading-permalink text-inherit no-underline">Where It Didn&#x27;t</a></h2>
<p class="text-muted-foreground leading-7 my-4">Here is the sentence from the chapter that reorganizes everything above it, quoted exactly because the summary version of it is misleading:</p>
<blockquote class="my-6 border-l-2 border-primary/40 print:border-primary pl-6">
<p class="text-muted-foreground leading-7 my-4">&quot;Across most business functions, a majority of respondents reported no agent use at all. Scaled use was in the single digits for nearly all functions.&quot;</p>
</blockquote>
<p class="text-muted-foreground leading-7 my-4">Read that carefully, because the chapter&#x27;s own one-line overview flattens it into &quot;AI agent deployment was in the single digits,&quot; and that is not what the data says. &quot;Single digits&quot; describes <em class="font-mono font-normal text-primary/80 print:text-primary">scaled</em> use — agents running as real infrastructure rather than in a pilot. The share reporting <em class="font-mono font-normal text-primary/80 print:text-primary">no</em> agent use at all is much larger than single digits. In the words of the report: &quot;Even in functions with the most activity, including IT and knowledge management, about two-thirds or more of respondents reported no use.&quot;</p>
<p class="text-muted-foreground leading-7 my-4">The high end is instructive. The functions with the most scaled agent use are exactly where you&#x27;d expect: software engineering at 24%, IT at 22%, service operations at 21%. Those are the peaks. Everywhere else falls away fast. So the picture is not &quot;agents are everywhere.&quot; It is &quot;agents are in engineering, and rare-to-absent in the rest of the company that engineering was supposed to be building them for.&quot;</p>
<p class="text-muted-foreground leading-7 my-4">One caveat the AI Index carries and I will carry too: these adoption figures come from McKinsey&#x27;s annual State of AI surveys, and they are self-reported. The report itself says they &quot;should be viewed as directional rather than comprehensive.&quot; Directional is enough for the argument. The direction is a two-order-of-magnitude gap between money in and agents out.</p>
<h2 id="the-number-that-says-the-opposite" class="font-display text-2xl lg:text-3xl font-bold text-foreground mt-12 mb-6 border-l-2 border-primary pl-4"><a href="#the-number-that-says-the-opposite" class="heading-permalink text-inherit no-underline">The Number That Says the Opposite</a></h2>
<p class="text-muted-foreground leading-7 my-4">Now the part that separates a useful reading from a credulous one.</p>
<p class="text-muted-foreground leading-7 my-4">If you spend any time in the agent-building community, the picture above will feel wrong, because you have seen a very different number. LangChain&#x27;s <a href="https://www.langchain.com/state-of-agent-engineering" class="text-primary underline underline-offset-2 decoration-primary/40 print:decoration-primary hover:decoration-primary transition-colors">State of Agent Engineering</a> survey found that <strong class="text-foreground font-bold">57.3% of respondents already had agents running in production</strong>. Not piloting. Production. That is not single digits; that is a majority.</p>
<p class="text-muted-foreground leading-7 my-4">Both numbers are honestly reported. They disagree because they asked different people. The McKinsey/AI Index data samples <em class="font-mono font-normal text-primary/80 print:text-primary">enterprises</em> — a broad cross-section of organizations and the functions inside them. The LangChain survey is a self-selected community sample: 1,340 respondents, fielded in late November and early December of 2025, 63% from the technology industry, and roughly half at companies with fewer than a hundred people. One instrument measured &quot;how much do organizations use agents.&quot; The other measured &quot;how much do people who build agents use agents.&quot; Of course they diverge.</p>
<p class="text-muted-foreground leading-7 my-4">The habit worth building is smaller than either statistic and worth more than both: before you believe an AI adoption number, ask <strong class="text-foreground font-bold">who was surveyed.</strong> A figure sampled from agent engineers tells you the frontier is real and shipping. A figure sampled from enterprises tells you the frontier is narrow and hasn&#x27;t diffused. Neither is a lie. They are answers to different questions, and most of the confusion in the market comes from treating them as answers to the same one.</p>
<p class="text-muted-foreground leading-7 my-4">The same discipline catches errors, not just framing. The AI Index chapter states, twice, that Google reported &quot;more than $150 billion in capex&quot; in 2025. Alphabet&#x27;s own filing puts 2025 capital expenditure at <strong class="text-foreground font-bold">$91.4 billion</strong>, up from $52.5 billion the year before, with $175–185 billion <em class="font-mono font-normal text-primary/80 print:text-primary">guided</em> for 2026. The $150 billion figure looks like a 2026 projection read as a 2025 actual — a mistake even a careful report can make when it restates a third party&#x27;s number about a specific company. When a claim can be checked against a primary source, check it. The gap between &quot;money in&quot; and &quot;agents out&quot; is real, but you want to be sure every number describing it is measuring what you think it is.</p>
<h2 id="what-a-founder-should-do-with-this" class="font-display text-2xl lg:text-3xl font-bold text-foreground mt-12 mb-6 border-l-2 border-primary pl-4"><a href="#what-a-founder-should-do-with-this" class="heading-permalink text-inherit no-underline">What a Founder Should Do With This</a></h2>
<p class="text-muted-foreground leading-7 my-4">The temptation is to treat the gap as a verdict — proof of a bubble, or proof of an untapped market. It is neither on its own. It is a description of <em class="font-mono font-normal text-primary/80 print:text-primary">timing</em>.</p>
<p class="text-muted-foreground leading-7 my-4">If you are selling agents, the gap is your addressable market and your warning label at once. The demand signal is unambiguous: nearly nine in ten organizations are already using AI, and the capital is flooding in. But the thing you are selling — agents running at scale, in functions beyond engineering — barely exists yet in the enterprises you are selling to. That is not a market that needs another demo. It is a market that needs the operational work of making agents trustworthy enough to run unattended: the permissions, the runbooks, the observability, the ownership. The companies that closed that gap for themselves are the ones with agents in production. Everyone else is stuck at the pilot.</p>
<p class="text-muted-foreground leading-7 my-4">If you are buying, the gap is permission to move deliberately. You are not late. The single-digit scaled-use number means the median organization has not figured this out either, and the ones that have did it by treating agents as infrastructure rather than as a feature to switch on. The advantage is not in adopting first. It is in being one of the few that gets an agent past the pilot and into the part of the business that isn&#x27;t engineering.</p>
<p class="text-muted-foreground leading-7 my-4">The consumer side hints at where the real value is accruing while enterprises deliberate. One estimate puts the annual consumer surplus from generative AI in the US at <strong class="text-foreground font-bold">$172 billion, up from $112 billion</strong> the year before, with the share of US adults using generative AI rising from 48% to 56% (Bick et al., 2026). Worth flagging what that figure is: a stated-preference measure, drawn from online experiments asking people what they&#x27;d need to be paid to give up generative AI for a month — not revenue, not revealed behavior. But even discounted, it points the same way. The tools are being used. The enterprise agent, running at scale, in production, owning real work, is the thing that hasn&#x27;t arrived.</p>
<p class="text-muted-foreground leading-7 my-4">$581.69 billion went looking for that agent last year. In most of the companies that spent it, the agent isn&#x27;t running yet. The gap is not the failure of the story. It is the middle of it — and it is a better place to be building than either end.</p>
<p class="text-muted-foreground leading-7 my-4">If you want the setup to this, <a href="/blog/the-zero-dollar-startup/" class="text-primary underline underline-offset-2 decoration-primary/40 print:decoration-primary hover:decoration-primary transition-colors">The Zero Dollar Startup</a> is about what happened when building got cheap, and <a href="/blog/distribution-is-the-new-code/" class="text-primary underline underline-offset-2 decoration-primary/40 print:decoration-primary hover:decoration-primary transition-colors">Distribution Is the New Code</a> is about where the leverage moved next.</p>]]></content:encoded>
      <pubDate>Tue, 08 Sep 2026 12:00:00 GMT</pubDate>
    </item>
    <item>
      <title>Green Is a Measurement, Not a Decision</title>
      <link>https://pratik.pa.tel/blog/green-is-a-measurement-not-a-decision/</link>
      <guid isPermaLink="true">https://pratik.pa.tel/blog/green-is-a-measurement-not-a-decision/</guid>
      <description>Nearly a quarter of what goes wrong is a check that said yes. Even when the check is right, green is still not a ship decision.</description>
      <dc:creator>Pratik Patel</dc:creator>
      <category>ai</category>
      <category>agents</category>
      <category>evals</category>
      <category>reliability</category>
      <content:encoded><![CDATA[<p class="text-muted-foreground leading-7 my-4">I keep watching the same meeting. Someone asks if the agent is ready. Someone else shares a dashboard. The suite is green. The conversation ends.</p>
<p class="text-muted-foreground leading-7 my-4">That used to be the right reflex. For a compiler, a passing test suite is close enough to a ship decision that we stopped noticing the gap. Agents broke the reflex and we didn&#x27;t update the meeting.</p>
<p class="text-muted-foreground leading-7 my-4"><a href="/blog/your-eval-suite-measures-the-wrong-thing/" class="text-primary underline underline-offset-2 decoration-primary/40 print:decoration-primary hover:decoration-primary transition-colors">Last month</a> I wrote that nearly a quarter of observed multi-agent failures are failures of the checking layer, and that the most common one is the check that ran and said yes. The closer was the part I want to pick up: your eval suite is another component in the system, and unlike everything else you built, there is nothing downstream of it that would notice if it broke.</p>
<p class="text-muted-foreground leading-7 my-4">This week is the next sentence. Even if you fix the suite — even if it starts measuring the right thing — a pass is not permission. Production agents fail in a place the suite was never pointed at. Green is a measurement. Shipping is a judgment. Most teams have quietly handed that judgment to the dashboard.</p>
<h2 id="the-suite-cannot-see-the-failure-you-will-get-paged-for" class="font-display text-2xl lg:text-3xl font-bold text-foreground mt-12 mb-6 border-l-2 border-primary pl-4"><a href="#the-suite-cannot-see-the-failure-you-will-get-paged-for" class="heading-permalink text-inherit no-underline">The Suite Cannot See the Failure You Will Get Paged For</a></h2>
<p class="text-muted-foreground leading-7 my-4">I have been reading Mukund Pandey&#x27;s <a href="https://arxiv.org/abs/2605.01604" class="text-primary underline underline-offset-2 decoration-primary/40 print:decoration-primary hover:decoration-primary transition-colors">Evaluating Agentic AI in the Wild</a>. It is a taxonomy of seven failure modes the author argues are specific to agents running continuously, not to models taking a test.</p>
<p class="text-muted-foreground leading-7 my-4">The list is unglamorous, which is why I trust the shape of it. Cascading decision error: an early step is wrong, every later step is locally correct given that input, and the output is internally coherent and systematically false. Silent tool degradation: a dependency starts returning schema-valid stale or partial data instead of failing, the logs stay clean, and downstream logic proceeds at full confidence. Distribution collapse: the agent converges on a narrow set of high-scoring outputs while accuracy stays flat. Cross-surface inconsistency: the same intent arriving through the API and the UI gets two different answers. Explanation decoupling: the decision is right and the reason you recorded is wrong. Latency-driven correctness erosion: the SLA is green because the system skipped the enrichment that made the answer good. Proxy goal convergence: the metric you rewarded went up for weeks while the thing you actually wanted quietly left.</p>
<p class="text-muted-foreground leading-7 my-4">I will not tour the framework the paper proposes. The useful part is the detection table. Against ROUGE, BERTScore, accuracy/AUC, AgentBench, and MT-Bench, <strong class="text-foreground font-bold">four of the seven modes produce no signal at all</strong>. The other three show up only after a lag of multiple evaluation cycles. No standard metric in that set detects any of them reliably inside a single cycle.</p>
<p class="text-muted-foreground leading-7 my-4">Sit with that. The suite can be honest, well-maintained, and pointed at the right property of the output, and still be blind to the incident you will get paged for. Last month the problem was a verifier that lied. This week the problem is a verifier that was never looking at the room the fire started in.</p>
<h2 id="accuracy-can-stay-flat-while-the-system-rots" class="font-display text-2xl lg:text-3xl font-bold text-foreground mt-12 mb-6 border-l-2 border-primary pl-4"><a href="#accuracy-can-stay-flat-while-the-system-rots" class="heading-permalink text-inherit no-underline">Accuracy Can Stay Flat While the System Rots</a></h2>
<p class="text-muted-foreground leading-7 my-4">The paper&#x27;s most instructive experiment is also the least dramatic.</p>
<p class="text-muted-foreground leading-7 my-4">The author simulates five weekly windows of session outputs. Accuracy is held between 0.86 and 0.88 the entire time — the production pattern where request-level correctness does not reflect what a user experiences across a session. Meanwhile the output distribution narrows from twenty categories to three. Diversity drops by a factor of six and a half. Repeat rate goes to 1.0: every output in the window comes from the same category.</p>
<p class="text-muted-foreground leading-7 my-4">The accuracy number never flinches.</p>
<p class="text-muted-foreground leading-7 my-4">A second experiment does the same trick with tools. Across four stages of an upstream service degrading into partial responses, the external accuracy signal moves by three hundredths. The partial-response rate goes from 4% to 58%. A team watching accuracy would see noise. The system is already shipping on incomplete inputs.</p>
<p class="text-muted-foreground leading-7 my-4">Two caveats, and I want them in the same section as the numbers.</p>
<p class="text-muted-foreground leading-7 my-4">First, these are synthetic traces built to reproduce signatures the author says he observed in production. The paper is explicit: there is no production dataset in the experiments, and the billion-event-scale examples are described without published proprietary metrics. Treat the direction as the finding, not 0.86 or 6.5×.</p>
<p class="text-muted-foreground leading-7 my-4">Second, this is a single-author paper with a proposed framework attached. I am not adopting the framework. I am taking the claim that is cheap to falsify and expensive to ignore: the metrics closest to the model are often the last to notice that the system has changed shape.</p>
<p class="text-muted-foreground leading-7 my-4">That is not a new idea. SRE has been living it for twenty years. Latency SLAs stay green while a fallback path skips the work that made the answer correct. Error rate stays low because the tool stopped erroring and started lying. <a href="/blog/agents-fail-quietly/" class="text-primary underline underline-offset-2 decoration-primary/40 print:decoration-primary hover:decoration-primary transition-colors">Agents fail quietly</a>. The new part is that the quiet failure can live entirely outside the eval you run before you ship, and still be the thing your users hit on day two.</p>
<h2 id="a-pass-is-an-input" class="font-display text-2xl lg:text-3xl font-bold text-foreground mt-12 mb-6 border-l-2 border-primary pl-4"><a href="#a-pass-is-an-input" class="heading-permalink text-inherit no-underline">A Pass Is an Input</a></h2>
<p class="text-muted-foreground leading-7 my-4">We already know how to treat a green suite in every other part of the stack. Unit tests passing is not a production deploy. It is one input to a decision that also includes an error budget, a canary, and a person who is allowed to halt the rollout after the tests said go.</p>
<p class="text-muted-foreground leading-7 my-4">Agent teams inverted that. The suite became the decision. &quot;Evals are green&quot; is how the meeting ends.</p>
<p class="text-muted-foreground leading-7 my-4">That only works if two things are true: the suite can see the failure mode that will page you, and the world the agent runs in is the world the suite was built against. Last month&#x27;s paper said the first is often false because the check itself is wrong. This week&#x27;s paper says the first is often false even when the check is right, because the failure is in the coupling — tool health to decision quality, latency to correctness, one step&#x27;s confidence to the next step&#x27;s certainty. The second is false the moment the model, the prompt, the index, or the tool changes after you froze the cases.</p>
<p class="text-muted-foreground leading-7 my-4">A snapshot cannot bless a system that keeps moving. Asking it to is how a measurement becomes a ritual.</p>
<h2 id="put-a-decision-after-the-check" class="font-display text-2xl lg:text-3xl font-bold text-foreground mt-12 mb-6 border-l-2 border-primary pl-4"><a href="#put-a-decision-after-the-check" class="heading-permalink text-inherit no-underline">Put a Decision After the Check</a></h2>
<p class="text-muted-foreground leading-7 my-4">Last week I wrote that somebody has to own the agent. The empty box on the org chart. That post is about the name. This one is about what that name is for.</p>
<p class="text-muted-foreground leading-7 my-4">An owner without a ship ritual is a name on a page. The suite will still end the meeting, and you will have assigned accountability for a decision nobody actually made.</p>
<p class="text-muted-foreground leading-7 my-4">I don&#x27;t want to end on a checklist, so let me end on the smallest set of things that would make &quot;evals are green&quot; stop being the last sentence in the room.</p>
<p class="text-muted-foreground leading-7 my-4"><strong class="text-foreground font-bold">The owner has to be allowed to say no after green.</strong> If the only halt is the suite, you do not have a ship decision. You have an automation. Give them a halt that works after merge, not just before it. I wrote about <a href="/blog/give-your-agent-an-undo-button/" class="text-primary underline underline-offset-2 decoration-primary/40 print:decoration-primary hover:decoration-primary transition-colors">reversibility</a> as a property of the agent&#x27;s actions. It is also a property of yours.</p>
<p class="text-muted-foreground leading-7 my-4"><strong class="text-foreground font-bold">Run something on live traffic that is not the suite.</strong> Shadow, canary, sampled traces — the shape matters less than the fact that it sees the couplings the offline cases cannot. <a href="/blog/trust-comes-from-the-trace/" class="text-primary underline underline-offset-2 decoration-primary/40 print:decoration-primary hover:decoration-primary transition-colors">Trust comes from the trace</a>, and the traces that matter are the ones from the system you actually shipped, including the runs that look fine.</p>
<p class="text-muted-foreground leading-7 my-4"><strong class="text-foreground font-bold">Watch the successful production runs.</strong> Catching a broken verifier was one reason. Catching a system whose accuracy is flat while its behavior has already narrowed, or whose tools have started returning partials, is the other. By definition no pre-ship alert will route you there.</p>
<p class="text-muted-foreground leading-7 my-4"><strong class="text-foreground font-bold">Write down what would make you unship.</strong> Not a severity matrix. One sentence: if this is true on Thursday, we turn it off. If you cannot finish that sentence, the suite is doing the deciding, and you have already seen why that is a bad job for it.</p>
<h2 id="the-bottom-line" class="font-display text-2xl lg:text-3xl font-bold text-foreground mt-12 mb-6 border-l-2 border-primary pl-4"><a href="#the-bottom-line" class="heading-permalink text-inherit no-underline">The Bottom Line</a></h2>
<p class="text-muted-foreground leading-7 my-4">You can spend a quarter fixing the eval suite and still ship the incident, because you asked the suite to do a job it cannot do. It can tell you what it saw on the cases you remembered to write. It cannot see the failure that only exists in the coupling between a tool and a decision, or in a distribution that collapsed while accuracy held still. And it cannot be the person in the room who is on the hook.</p>
<p class="text-muted-foreground leading-7 my-4">Green is a measurement.</p>
<p class="text-muted-foreground leading-7 my-4">It is not a decision.</p>]]></content:encoded>
      <pubDate>Tue, 01 Sep 2026 12:00:00 GMT</pubDate>
    </item>
    <item>
      <title>Somebody Has to Own the Agent</title>
      <link>https://pratik.pa.tel/blog/somebody-has-to-own-the-agent/</link>
      <guid isPermaLink="true">https://pratik.pa.tel/blog/somebody-has-to-own-the-agent/</guid>
      <description>Gartner expects more than 40% of agentic AI projects to be canceled by the end of 2027. None of the three reasons it names is a model-capability problem. The gap that kills agents in production is organizational, and we already know how to close it.</description>
      <dc:creator>Pratik Patel</dc:creator>
      <category>ai</category>
      <category>agents</category>
      <category>leadership</category>
      <category>governance</category>
      <category>engineering</category>
      <content:encoded><![CDATA[<p class="text-muted-foreground leading-7 my-4">Your agent has been in production for three months. It files tickets, moves money, or emails customers on your behalf. Now answer one question: whose name is on it?</p>
<p class="text-muted-foreground leading-7 my-4">Not which team deployed it. Not who wrote the prompt. Who is accountable when it does something expensive at 2am on a Sunday — the person who gets paged, who decides whether to shut it off, and who answers for that decision on Monday.</p>
<p class="text-muted-foreground leading-7 my-4">For a lot of agents running in production right now, that box on the org chart is empty. The agent has a repo, a budget line, and real permissions. It does not have an owner. And the failure that eventually takes it down will not be a model failure.</p>
<h2 id="the-reasons-projects-die-are-not-technical-reasons" class="font-display text-2xl lg:text-3xl font-bold text-foreground mt-12 mb-6 border-l-2 border-primary pl-4"><a href="#the-reasons-projects-die-are-not-technical-reasons" class="heading-permalink text-inherit no-underline">The Reasons Projects Die Are Not Technical Reasons</a></h2>
<p class="text-muted-foreground leading-7 my-4">Gartner predicts that <a href="https://www.gartner.com/en/newsroom/press-releases/2025-06-25-gartner-predicts-over-40-percent-of-agentic-ai-projects-will-be-canceled-by-end-of-2027" class="text-primary underline underline-offset-2 decoration-primary/40 print:decoration-primary hover:decoration-primary transition-colors">more than 40% of agentic AI projects will be canceled by the end of 2027</a>, citing escalating costs, unclear business value, or inadequate risk controls.</p>
<p class="text-muted-foreground leading-7 my-4">Read that list again slowly, because the interesting thing about it is what is missing. Not one of those three is a capability problem. A smarter foundation model does not fix escalating costs, does not clarify business value, and does not install risk controls. Every one of them is a question about who decided what, and who was watching.</p>
<p class="text-muted-foreground leading-7 my-4">That reading is mine, not Gartner&#x27;s. But the list is Gartner&#x27;s, and it is remarkably consistent with what the same analysts say is driving the hype in the first place. Anushree Verma, Senior Director Analyst at Gartner, put the underlying problem this way: &quot;Most agentic AI propositions lack significant value or return on investment, as current models don&#x27;t have the maturity and agency to autonomously achieve complex business goals or follow nuanced instructions over time.&quot;</p>
<p class="text-muted-foreground leading-7 my-4">The market is not helping. Gartner describes widespread &quot;agent washing&quot; — rebranding assistants, RPA, and chatbots as agents without substantial agentic capability — and estimates that only around 130 of the thousands of self-described agentic AI vendors are real. If you are buying rather than building, most of what you evaluate is a wrapper with a new label, and nobody inside your company is positioned to say so unless somebody owns the outcome.</p>
<h2 id="adoption-is-not-deployment" class="font-display text-2xl lg:text-3xl font-bold text-foreground mt-12 mb-6 border-l-2 border-primary pl-4"><a href="#adoption-is-not-deployment" class="heading-permalink text-inherit no-underline">Adoption Is Not Deployment</a></h2>
<p class="text-muted-foreground leading-7 my-4">Forrester&#x27;s <a href="https://www.forrester.com/blogs/the-state-of-agentic-ai-in-2026-companies-are-chasing-few-are-catching/" class="text-primary underline underline-offset-2 decoration-primary/40 print:decoration-primary hover:decoration-primary transition-colors">State of Agentic AI in 2026</a> frames the same gap from the other side: three-quarters of enterprise leaders say they are adopting agentic AI, while only a small minority have it running in meaningful production beyond what Forrester calls &quot;agentish&quot; chatbots.</p>
<p class="text-muted-foreground leading-7 my-4">That distance between adopting and running is where ownership lives. It is easy to sponsor an agent. It is easy to fund a pilot. What is hard, and what almost nobody staffs for, is the unglamorous ongoing work: watching cost per run drift up, noticing that the success rate slipped four points after a model update, deciding the agent should stop doing one of the five things it was scoped to do.</p>
<p class="text-muted-foreground leading-7 my-4">Forrester&#x27;s recommendation is specific and, I think, correct: treat every agent as a governed identity. &quot;Give it unique credentials, least privilege, full logging, and a named owner who manages its lifecycle — no unowned autonomy.&quot;</p>
<p class="text-muted-foreground leading-7 my-4">Worth being precise about what that is. It is a recommendation, not a measurement. Forrester is not reporting that companies with named owners succeed at some rate; it is saying that unowned autonomy is a bad idea. I have looked for a clean survey number tying named ownership to agent outcomes and have not found one that survives checking — the figures floating around on this are mostly untraceable. So take the following as an argument from practice rather than a finding.</p>
<h2 id="we-already-learned-this-with-services" class="font-display text-2xl lg:text-3xl font-bold text-foreground mt-12 mb-6 border-l-2 border-primary pl-4"><a href="#we-already-learned-this-with-services" class="heading-permalink text-inherit no-underline">We Already Learned This With Services</a></h2>
<p class="text-muted-foreground leading-7 my-4">None of this is new. It is on-call, rediscovered.</p>
<p class="text-muted-foreground leading-7 my-4">Fifteen years ago you could ship a service into production with no named owner, and the industry spent a decade learning why that ends badly. The answer we converged on was not better monitoring software. It was a person: a service has an owner, the owner has a pager, the pager has an escalation path, and the owner has enough authority to change the thing they are accountable for. Ownership without authority is just blame with extra steps.</p>
<p class="text-muted-foreground leading-7 my-4">Agents need the same structure, and they need it more urgently, because an agent can do damage at a speed and breadth that a broken service usually cannot. A service that falls over stops working. An agent that goes wrong <a href="/blog/agents-fail-quietly/" class="text-primary underline underline-offset-2 decoration-primary/40 print:decoration-primary hover:decoration-primary transition-colors">keeps working, confidently, in the wrong direction</a>.</p>
<p class="text-muted-foreground leading-7 my-4">So the useful question is not &quot;do we have an AI governance policy.&quot; It is the on-call question, asked about each agent individually:</p>
<ul class="space-y-2 my-6 ml-4">
<li class="flex gap-3 text-muted-foreground leading-7"><span class="text-primary shrink-0 mt-1.5">▸</span><span class="min-w-0">Who gets paged when this agent misbehaves, by name?</span></li>
<li class="flex gap-3 text-muted-foreground leading-7"><span class="text-primary shrink-0 mt-1.5">▸</span><span class="min-w-0">Can that person turn it off without convening a meeting?</span></li>
<li class="flex gap-3 text-muted-foreground leading-7"><span class="text-primary shrink-0 mt-1.5">▸</span><span class="min-w-0">Can they see what it actually did, step by step, or only that it returned a 200?</span></li>
<li class="flex gap-3 text-muted-foreground leading-7"><span class="text-primary shrink-0 mt-1.5">▸</span><span class="min-w-0">Do they own the budget it spends, so that cost drift is their problem and not a surprise in someone else&#x27;s quarterly review?</span></li>
<li class="flex gap-3 text-muted-foreground leading-7"><span class="text-primary shrink-0 mt-1.5">▸</span><span class="min-w-0">Is there a number that tells them whether it is working — not &quot;is it up,&quot; but is it producing the outcome it was funded to produce?</span></li>
</ul>
<p class="text-muted-foreground leading-7 my-4">If you cannot answer those five for an agent you are running today, it is unowned, whatever the slide deck says.</p>
<h2 id="the-trap-one-policy-for-every-agent" class="font-display text-2xl lg:text-3xl font-bold text-foreground mt-12 mb-6 border-l-2 border-primary pl-4"><a href="#the-trap-one-policy-for-every-agent" class="heading-permalink text-inherit no-underline">The Trap: One Policy for Every Agent</a></h2>
<p class="text-muted-foreground leading-7 my-4">The most common way teams get this wrong is not neglect. It is over-correcting into a single uniform policy.</p>
<p class="text-muted-foreground leading-7 my-4">Gartner published a specific warning about this in May 2026: <a href="https://www.gartner.com/en/newsroom/press-releases/2026-05-26-gartner-says-applying-uniform-governance-across-ai-agents-will-lead-to-enterprise-ai-agent-failure" class="text-primary underline underline-offset-2 decoration-primary/40 print:decoration-primary hover:decoration-primary transition-colors">applying uniform governance across AI agents will lead to enterprise AI agent failure</a>. Shiva Varma, Senior Director Analyst at Gartner, states the failure mode directly: &quot;Enterprises are treating AI agent governance as binary, either locked down or fully trusted, and that is the root cause of failure.&quot;</p>
<p class="text-muted-foreground leading-7 my-4">The distinction Gartner says organizations miss is between an agent&#x27;s ability to act and the scope of access it is granted. Those are two different dials, and collapsing them is exactly how you end up with a summarizer holding production database credentials, or a genuinely useful workflow agent throttled into uselessness because it got classified alongside it. Gartner&#x27;s recommendation is proportional governance: classify agents across distinct autonomy levels, where each level is a different trust boundary with its own requirements.</p>
<p class="text-muted-foreground leading-7 my-4">This is the same argument I made about <a href="/blog/agent-permissions-are-product-design/" class="text-primary underline underline-offset-2 decoration-primary/40 print:decoration-primary hover:decoration-primary transition-colors">agent permissions being product design</a>, arriving from the governance side. A permission model is not a compliance artifact you bolt on at the end. It is a description of what the agent is for.</p>
<p class="text-muted-foreground leading-7 my-4">Gartner also predicts that by 2027, 40% of enterprises will demote or decommission autonomous AI agents due to governance gaps identified only after production incidents occur. That last clause is the whole problem in six words. The gap was there the entire time. It became visible when something broke.</p>
<h2 id="what-an-owner-actually-does" class="font-display text-2xl lg:text-3xl font-bold text-foreground mt-12 mb-6 border-l-2 border-primary pl-4"><a href="#what-an-owner-actually-does" class="heading-permalink text-inherit no-underline">What an Owner Actually Does</a></h2>
<p class="text-muted-foreground leading-7 my-4">Naming an owner is the easy half, and it is where most orgs stop. The name goes in a spreadsheet cell and nothing else changes.</p>
<p class="text-muted-foreground leading-7 my-4">An owner who can actually do the job needs three things, none of which are usually granted at the same time as the title:</p>
<p class="text-muted-foreground leading-7 my-4"><strong class="text-foreground font-bold">Visibility.</strong> They have to be able to reconstruct what the agent did and why. Not logs — <a href="/blog/trust-comes-from-the-trace/" class="text-primary underline underline-offset-2 decoration-primary/40 print:decoration-primary hover:decoration-primary transition-colors">traces</a>. If the only available evidence is that the run completed successfully, the owner cannot form a judgment, and so they cannot be responsible for one.</p>
<p class="text-muted-foreground leading-7 my-4"><strong class="text-foreground font-bold">Authority to change scope.</strong> They can narrow the agent&#x27;s permissions, pause a workflow, or retire a capability without a committee. If turning the agent off requires the approval of the executive who sponsored it, the agent is not owned. It is protected.</p>
<p class="text-muted-foreground leading-7 my-4"><strong class="text-foreground font-bold">A number they are accountable to.</strong> Not usage. Outcome. Tickets resolved without rework, hours saved against a baseline someone measured before launch, error rate per hundred runs. Agents that cannot be evaluated get renewed forever on vibes, right up until the cost review that kills them — which is, roughly, the first of Gartner&#x27;s three cancellation reasons.</p>
<h2 id="the-empty-box" class="font-display text-2xl lg:text-3xl font-bold text-foreground mt-12 mb-6 border-l-2 border-primary pl-4"><a href="#the-empty-box" class="heading-permalink text-inherit no-underline">The Empty Box</a></h2>
<p class="text-muted-foreground leading-7 my-4">The uncomfortable version of this post is short. Most organizations running agents in production could not, today, produce a name for each one.</p>
<p class="text-muted-foreground leading-7 my-4">That is fixable this week, and it does not require a platform, a vendor, or a framework. List every agent you have running. Put a human name next to each. For any row where you cannot, either find the name or turn the agent off until you can.</p>
<p class="text-muted-foreground leading-7 my-4">Some of those rows will be hard to fill, and the difficulty is the signal. An agent nobody will put their name on is an agent nobody believes in enough to defend — and it is running with your credentials anyway.</p>
<p class="text-muted-foreground leading-7 my-4">The scaling problem in front of most teams is not that the models are not good enough yet. It is that we have deployed a new class of actor into our companies and skipped the part where we decide who is responsible for it.</p>
<p class="text-muted-foreground leading-7 my-4">Somebody has to own the agent. Right now, for most agents, nobody does.</p>]]></content:encoded>
      <pubDate>Tue, 25 Aug 2026 12:00:00 GMT</pubDate>
    </item>
    <item>
      <title>The Entry-Level Job Is the Canary</title>
      <link>https://pratik.pa.tel/blog/the-entry-level-job-is-the-canary/</link>
      <guid isPermaLink="true">https://pratik.pa.tel/blog/the-entry-level-job-is-the-canary/</guid>
      <description>The most-cited number about AI and jobs is real, and almost everyone repeats it wrong. Here is what the Stanford paper actually found, what its authors refuse to claim, and the one fact in it worth building a hiring plan on.</description>
      <dc:creator>Pratik Patel</dc:creator>
      <category>ai</category>
      <category>hiring</category>
      <category>economics</category>
      <category>startups</category>
      <category>leadership</category>
      <content:encoded><![CDATA[<p class="text-muted-foreground leading-7 my-4">Your headcount plan is a forecast about which tasks survive. Most people are making it from a statistic they have misread.</p>
<p class="text-muted-foreground leading-7 my-4">The statistic is 16%, and it comes from <a href="https://digitaleconomy.stanford.edu/publications/canaries-in-the-coal-mine/" class="text-primary underline underline-offset-2 decoration-primary/40 print:decoration-primary hover:decoration-primary transition-colors"><em class="font-mono font-normal text-primary/80 print:text-primary">Canaries in the Coal Mine?</em></a>, a paper by Erik Brynjolfsson, Bharat Chandar and Ruyu Chen at the Stanford Digital Economy Lab. It has been quoted in every direction for months. It is a serious piece of work and the number is real. It just does not say what the reposts say it says.</p>
<p class="text-muted-foreground leading-7 my-4">I read the November 13, 2025 version end to end before writing this, because an earlier draft carried a different figure and that one is still circulating. What follows is the finding stated correctly, followed by the part of it I would actually change a plan over.</p>
<h2 id="what-the-16-is" class="font-display text-2xl lg:text-3xl font-bold text-foreground mt-12 mb-6 border-l-2 border-primary pl-4"><a href="#what-the-16-is" class="heading-permalink text-inherit no-underline">What the 16% Is</a></h2>
<p class="text-muted-foreground leading-7 my-4">Here is the abstract, verbatim:</p>
<blockquote class="my-6 border-l-2 border-primary/40 print:border-primary pl-6">
<p class="text-muted-foreground leading-7 my-4">&quot;Early-career workers (ages 22-25) in AI-exposed occupations experienced 16% relative employment declines, controlling for firm-level shocks, while employment for experienced workers remained stable.&quot;</p>
</blockquote>
<p class="text-muted-foreground leading-7 my-4">Two words in that sentence do all the work, and both get dropped in the retelling.</p>
<p class="text-muted-foreground leading-7 my-4"><strong class="text-foreground font-bold">Relative.</strong> The 16% is a comparison, not a count. It is the gap between young workers in the most AI-exposed occupations and comparable older workers <em class="font-mono font-normal text-primary/80 print:text-primary">at the same firms</em>, after the model absorbs whatever else was happening to those firms. It is not &quot;one in six entry-level jobs disappeared.&quot; Nobody lost 16% of anything.</p>
<p class="text-muted-foreground leading-7 my-4"><strong class="text-foreground font-bold">Controlling.</strong> It is a regression estimate, conditioned on firm-time effects. That conditioning is the point — it is what rules out &quot;this firm was just shrinking&quot; — but it also means the number is a modelled contrast, not something you could count off a payroll.</p>
<p class="text-muted-foreground leading-7 my-4">If you want a figure you can hold in your hand, the paper has better ones, and they are unmodelled. From late 2022 to September 2025, in the <strong class="text-foreground font-bold">most AI-exposed occupations</strong>, employment for 22-to-25-year-olds <strong class="text-foreground font-bold">fell 6%</strong>, &quot;compared to a 6-9% increase for older workers.&quot; In the three <strong class="text-foreground font-bold">least</strong>-exposed quintiles, employment grew <strong class="text-foreground font-bold">5-13% for every age group, with no clear ordering by age</strong>.</p>
<p class="text-muted-foreground leading-7 my-4">Read those two sentences next to each other. That contrast is the entire finding. Where AI exposure is low, young and old grew together. Where it is high, the young line bends down while everyone else&#x27;s keeps climbing.</p>
<p class="text-muted-foreground leading-7 my-4">And the number that will get screenshotted: by September 2025, employment for <strong class="text-foreground font-bold">software developers aged 22-25 had declined nearly 20% from its peak in late 2022</strong>. True — with the caveat that late 2022 was a hiring bubble, so some of that fall is a bubble deflating rather than a job vanishing.</p>
<h2 id="what-the-authors-will-not-claim" class="font-display text-2xl lg:text-3xl font-bold text-foreground mt-12 mb-6 border-l-2 border-primary pl-4"><a href="#what-the-authors-will-not-claim" class="heading-permalink text-inherit no-underline">What the Authors Will Not Claim</a></h2>
<p class="text-muted-foreground leading-7 my-4">This is where the honest version parts company with the viral version.</p>
<p class="text-muted-foreground leading-7 my-4">The paper does not say generative AI caused this. Its own words:</p>
<blockquote class="my-6 border-l-2 border-primary/40 print:border-primary pl-6">
<p class="text-muted-foreground leading-7 my-4">&quot;While our main estimates may be influenced by factors other than generative AI, our results are consistent with the hypothesis that generative AI has begun to affect entry-level employment significantly.&quot;</p>
</blockquote>
<p class="text-muted-foreground leading-7 my-4">And from the conclusion: &quot;Future work would benefit from better firm-level AI adoption data, which would provide sharper variation for estimating plausible causal effects of AI on employment.&quot; The title is <em class="font-mono font-normal text-primary/80 print:text-primary">Six Facts</em>, not six effects. That word choice is deliberate and it is the most-ignored thing in the paper.</p>
<p class="text-muted-foreground leading-7 my-4">Three more limits worth carrying, because they are the ones that will be quoted back at you:</p>
<p class="text-muted-foreground leading-7 my-4"><strong class="text-foreground font-bold">The data is one payroll provider.</strong> ADP — large, administrative, real, and not a national sample.</p>
<p class="text-muted-foreground leading-7 my-4"><strong class="text-foreground font-bold">Credible datasets disagree.</strong> The paper cites its own contradictors, which is a good sign about the paper. Hampole et al. (2025), using Revelio Labs job postings and LinkedIn profiles from 2011 to 2023, find limited employment impacts overall, with growing labour demand at firms offsetting relative declines in exposed occupations. Chandar (2025b) — the same Chandar — using CPS data finds little differential trend overall, while noting how hard young workers are to measure there. Two credible datasets, two different pictures.</p>
<p class="text-muted-foreground leading-7 my-4"><strong class="text-foreground font-bold">Watch for the aggregator trap.</strong> You will see the &quot;nearly 20% for young software developers&quot; figure repeated in other 2026 reports. That is not corroboration. Those reports are citing this same paper. One study restated three times is still one study.</p>
<h2 id="the-fact-worth-planning-around" class="font-display text-2xl lg:text-3xl font-bold text-foreground mt-12 mb-6 border-l-2 border-primary pl-4"><a href="#the-fact-worth-planning-around" class="heading-permalink text-inherit no-underline">The Fact Worth Planning Around</a></h2>
<p class="text-muted-foreground leading-7 my-4">Now the part that survives all of that, and the reason I think this paper matters to anyone running a company.</p>
<p class="text-muted-foreground leading-7 my-4">The declines are not spread evenly across &quot;AI-exposed&quot; work. They concentrate where AI <strong class="text-foreground font-bold">automates</strong> a task, and go muted where it <strong class="text-foreground font-bold">augments</strong> one.</p>
<p class="text-muted-foreground leading-7 my-4">The authors measure this using the <a href="https://www.anthropic.com/economic-index" class="text-primary underline underline-offset-2 decoration-primary/40 print:decoration-primary hover:decoration-primary transition-colors">Anthropic Economic Index</a>, which classifies the share of real Claude queries per occupational task as automative or augmentative. Occupations with the highest automation shares show declining employment for the youngest workers. Occupations with the highest augmentation shares show no such pattern — the top augmentation quintile has among the <em class="font-mono font-normal text-primary/80 print:text-primary">fastest</em> employment growth for young workers.</p>
<p class="text-muted-foreground leading-7 my-4">Disclose the proxy when you repeat this: query mix is a measure of how people use one AI model, not a measure of what firms have adopted. It is a good proxy. It is still a proxy.</p>
<p class="text-muted-foreground leading-7 my-4">But the mechanism it points at is intuitive enough that I believe it. The paper&#x27;s framing is that AI is automating &quot;the codifiable, checkable tasks that historically justified entry-level headcount,&quot; while complementing the judgment-, client- and process-heavy work that experienced people do. The entry-level job was always partly a training subsidy — you paid someone to do the checkable version of the work while they learned the unwritten version. When the checkable version gets cheap, that arrangement is the first thing to come under pressure.</p>
<p class="text-muted-foreground leading-7 my-4">Two more facts to file next to it.</p>
<p class="text-muted-foreground leading-7 my-4"><strong class="text-foreground font-bold">It is showing up in headcount, not in pay.</strong> Compensation trends barely diverge by age or exposure; employment trends diverge a lot. The market is repricing <em class="font-mono font-normal text-primary/80 print:text-primary">whether the role exists</em>, not what it pays. (The paper is careful here too — it offers offsetting wage effects or simple short-run wage stickiness as competing explanations.)</p>
<p class="text-muted-foreground leading-7 my-4"><strong class="text-foreground font-bold">It is probably hiring, not firing — and that part is a hypothesis, not a measurement.</strong> The authors suggest reduced hiring is &quot;the lowest-friction adjustment margin,&quot; and that firms &quot;may primarily shrink junior inflows rather than displace incumbents.&quot; Note every &quot;may.&quot; They did not measure inflows against separations; they proposed a mechanism that fits the shape of the data. Treat it as the best current guess, not a result.</p>
<p class="text-muted-foreground leading-7 my-4">Which, if it holds, means the adjustment is nearly invisible from inside a company. Nobody announces it. There is no layoff. A role just quietly stops being opened.</p>
<h2 id="so-what-do-you-do-with-it" class="font-display text-2xl lg:text-3xl font-bold text-foreground mt-12 mb-6 border-l-2 border-primary pl-4"><a href="#so-what-do-you-do-with-it" class="heading-permalink text-inherit no-underline">So What Do You Do With It</a></h2>
<p class="text-muted-foreground leading-7 my-4">I run a company where agents do a large share of the checkable work. So I have to take this seriously rather than argue with it, and here is where I have landed.</p>
<p class="text-muted-foreground leading-7 my-4"><strong class="text-foreground font-bold">Audit your roles by task, not by title.</strong> The unit of exposure in this data is the task, not the job. Take each role you were planning to hire and split it: which parts are codified and checkable, which need judgment, context or a relationship. If a role is mostly the first kind, you are not hiring a person, you are buying a workflow — and you should ask whether you want it as a workflow.</p>
<p class="text-muted-foreground leading-7 my-4"><strong class="text-foreground font-bold">Do not read &quot;automate&quot; as &quot;eliminate.&quot;</strong> The augmentation quintiles grew fastest. The instruction the data actually gives is to hire <em class="font-mono font-normal text-primary/80 print:text-primary">into</em> the augmented shape: fewer people doing more leveraged work, earlier. That is a different plan from a hiring freeze, and it is a better one.</p>
<p class="text-muted-foreground leading-7 my-4"><strong class="text-foreground font-bold">Notice that the cheap-to-check work was also the training ladder.</strong> If you automate every task a junior used to learn on, you have optimised this year and mortgaged your senior bench. Somebody has to still be growing into the judgment work, and it will not happen by accident. I wrote about the individual side of this in <a href="/blog/own-your-career/" class="text-primary underline underline-offset-2 decoration-primary/40 print:decoration-primary hover:decoration-primary transition-colors">Own Your Career</a>; the company side is the same problem viewed from the other end.</p>
<p class="text-muted-foreground leading-7 my-4"><strong class="text-foreground font-bold">Stop measuring people by output volume.</strong> When the checkable output is nearly free, counting it tells you nothing about who is valuable. That was already true before this paper — it is the whole argument in <a href="/blog/10x-engineer-myth/" class="text-primary underline underline-offset-2 decoration-primary/40 print:decoration-primary hover:decoration-primary transition-colors">The 10x Engineer Myth</a> — and this data makes it expensive to keep getting wrong.</p>
<h2 id="the-canary-is-a-warning-not-a-verdict" class="font-display text-2xl lg:text-3xl font-bold text-foreground mt-12 mb-6 border-l-2 border-primary pl-4"><a href="#the-canary-is-a-warning-not-a-verdict" class="heading-permalink text-inherit no-underline">The Canary Is a Warning, Not a Verdict</a></h2>
<p class="text-muted-foreground leading-7 my-4">The paper&#x27;s title is the right metaphor and it is worth taking literally. A canary is an early indicator in a confined space. It tells you something about the air before you can feel it yourself. It does not tell you the mine has collapsed, and it does not tell you why.</p>
<p class="text-muted-foreground leading-7 my-4">That is roughly the epistemic state we are in. Something real is happening at the entry level of AI-exposed work. It is showing up in headcount rather than wages, in automated tasks rather than augmented ones, and most likely through hiring that quietly does not happen. Whether generative AI is the cause is not established, and the honest researchers on this are the ones saying so.</p>
<p class="text-muted-foreground leading-7 my-4">You do not need causality to act on it. You need to know which of your tasks are checkable, and what you plan to do with the people who used to check them.</p>]]></content:encoded>
      <pubDate>Tue, 18 Aug 2026 12:00:00 GMT</pubDate>
    </item>
    <item>
      <title>Your Eval Suite Measures the Wrong Thing</title>
      <link>https://pratik.pa.tel/blog/your-eval-suite-measures-the-wrong-thing/</link>
      <guid isPermaLink="true">https://pratik.pa.tel/blog/your-eval-suite-measures-the-wrong-thing/</guid>
      <description>Nearly a quarter of what goes wrong in multi-agent systems is a verification failure. The most common one is a check that ran, and said yes.</description>
      <dc:creator>Pratik Patel</dc:creator>
      <category>ai</category>
      <category>agents</category>
      <category>evals</category>
      <category>reliability</category>
      <content:encoded><![CDATA[<p class="text-muted-foreground leading-7 my-4">Everyone measures whether the agent works. Almost nobody measures whether they&#x27;d know if it stopped.</p>
<p class="text-muted-foreground leading-7 my-4">That distinction used to be academic. It isn&#x27;t anymore. In LangChain&#x27;s <a href="https://www.langchain.com/state-of-agent-engineering" class="text-primary underline underline-offset-2 decoration-primary/40 print:decoration-primary hover:decoration-primary transition-colors">State of Agent Engineering</a> survey, <strong class="text-foreground font-bold">57.3% of respondents already had agents running in production</strong>, with another 30.4% actively building toward deployment. Whatever you believe about the hype cycle, the deployment happened. The measurement did not keep pace.</p>
<p class="text-muted-foreground leading-7 my-4">I&#x27;ve been reading the <a href="https://arxiv.org/abs/2503.13657" class="text-primary underline underline-offset-2 decoration-primary/40 print:decoration-primary hover:decoration-primary transition-colors">MAST paper</a> — a failure taxonomy built from 1,642 annotated execution traces across seven multi-agent frameworks — and one number in it reorganized how I think about evals.</p>
<p class="text-muted-foreground leading-7 my-4">Failures sort into three categories. System design issues are the largest. Inter-agent misalignment is second, and <a href="/blog/the-handoff-is-where-agents-break/" class="text-primary underline underline-offset-2 decoration-primary/40 print:decoration-primary hover:decoration-primary transition-colors">that&#x27;s the seam I wrote about last week</a>. The third category is <strong class="text-foreground font-bold">task verification, at 23.5% of everything observed</strong>.</p>
<p class="text-muted-foreground leading-7 my-4">Not &quot;the agent was wrong.&quot; The process that was supposed to catch the agent being wrong didn&#x27;t catch it.</p>
<p class="text-muted-foreground leading-7 my-4">Nearly a quarter of observed failures are failures of the checking layer. Which is to say: of the thing you built to tell you about the other failures.</p>
<h2 id="the-three-ways-checking-fails" class="font-display text-2xl lg:text-3xl font-bold text-foreground mt-12 mb-6 border-l-2 border-primary pl-4"><a href="#the-three-ways-checking-fails" class="heading-permalink text-inherit no-underline">The Three Ways Checking Fails</a></h2>
<p class="text-muted-foreground leading-7 my-4">The 23.5% breaks into three named modes, and the breakdown is more useful than the total, because each one is a different thing your eval suite isn&#x27;t doing:</p>
<p class="text-muted-foreground leading-7 my-4"><strong class="text-foreground font-bold">Incorrect verification — 9.10%.</strong> Something was checked. The check was wrong.</p>
<p class="text-muted-foreground leading-7 my-4"><strong class="text-foreground font-bold">No or incomplete verification — 8.20%.</strong> Nothing was checked, or only part of it was.</p>
<p class="text-muted-foreground leading-7 my-4"><strong class="text-foreground font-bold">Premature termination — 6.20%.</strong> The system stopped before the work was actually done.</p>
<p class="text-muted-foreground leading-7 my-4">Sit with the ordering for a second. The <strong class="text-foreground font-bold">most common</strong> verification failure is not the missing check. It&#x27;s the check that ran and returned a pass.</p>
<p class="text-muted-foreground leading-7 my-4">That inverts how most teams think about eval coverage. The instinct is that reliability is a coverage problem — we haven&#x27;t written enough tests yet, we&#x27;ll get to it. But the largest single bucket here is code that executed, evaluated the output, and concluded it was fine. More coverage of that kind adds more surface for the same failure.</p>
<h2 id="what-a-superficial-check-looks-like" class="font-display text-2xl lg:text-3xl font-bold text-foreground mt-12 mb-6 border-l-2 border-primary pl-4"><a href="#what-a-superficial-check-looks-like" class="heading-permalink text-inherit no-underline">What a Superficial Check Looks Like</a></h2>
<p class="text-muted-foreground leading-7 my-4">The paper is specific about why verifiers fail, and the answer is unglamorous. From the authors&#x27; analysis of the traces:</p>
<blockquote class="my-6 border-l-2 border-primary/40 print:border-primary pl-6">
<p class="text-muted-foreground leading-7 my-4">&quot;We find that many existing verifiers perform only superficial checks, despite being prompted to perform thorough verification, such as checking if the code compiles or if there are leftover TODO comments.&quot;</p>
</blockquote>
<p class="text-muted-foreground leading-7 my-4">Their worked example is a ChatDev-generated chess program. It compiles. It passes review. It has runtime bugs, because nothing validated it against the actual rules of chess. The output is unusable, and every check in the pipeline said it was fine.</p>
<p class="text-muted-foreground leading-7 my-4">Read that list of superficial checks again — does it compile, are there leftover TODOs — and ask how much of your own agent eval suite it describes. Did the code parse. Did the JSON validate. Did the response come back non-empty. Did it mention the keyword. These are all real checks. They are all checks on the <em class="font-mono font-normal text-primary/80 print:text-primary">shape</em> of the output rather than its <em class="font-mono font-normal text-primary/80 print:text-primary">correctness</em>, and shape is exactly what a fluent model gets right while getting the substance wrong.</p>
<p class="text-muted-foreground leading-7 my-4">The paper&#x27;s own conclusion on this is blunt:</p>
<blockquote class="my-6 border-l-2 border-primary/40 print:border-primary pl-6">
<p class="text-muted-foreground leading-7 my-4">&quot;Current verifier implementations are often insufficient; sole reliance on final-stage, low-level checks is inadequate.&quot;</p>
</blockquote>
<p class="text-muted-foreground leading-7 my-4">Final-stage and low-level. That&#x27;s a description of the median eval suite: one assertion, at the end, on the surface of the answer.</p>
<h2 id="the-part-that-should-actually-worry-you" class="font-display text-2xl lg:text-3xl font-bold text-foreground mt-12 mb-6 border-l-2 border-primary pl-4"><a href="#the-part-that-should-actually-worry-you" class="heading-permalink text-inherit no-underline">The Part That Should Actually Worry You</a></h2>
<p class="text-muted-foreground leading-7 my-4">Here is the finding that changed my mind about something.</p>
<p class="text-muted-foreground leading-7 my-4">The authors split their traces by outcome — failures observed in runs that <em class="font-mono font-normal text-primary/80 print:text-primary">succeeded</em> versus runs that <em class="font-mono font-normal text-primary/80 print:text-primary">failed</em> — and the split is not uniform. Some failure modes appear almost exclusively in failed runs. Unaware of termination conditions, information withholding: when those show up, the task is usually going down with them.</p>
<p class="text-muted-foreground leading-7 my-4">Verification failures behave differently. Verbatim:</p>
<blockquote class="my-6 border-l-2 border-primary/40 print:border-primary pl-6">
<p class="text-muted-foreground leading-7 my-4">&quot;In contrast, verification-related failures like 3.2 No or Incomplete Verification and 3.3 Incorrect Verification appear frequently even in successful runs. This suggests that while these systems can complete some tasks, their verification process still contains flaws.&quot;</p>
</blockquote>
<p class="text-muted-foreground leading-7 my-4">Your broken verifier does not announce itself by breaking the task. It sits inside runs that came out fine.</p>
<p class="text-muted-foreground leading-7 my-4">This is the measured version of something I&#x27;ve argued from instinct before — that <a href="/blog/agents-fail-quietly/" class="text-primary underline underline-offset-2 decoration-primary/40 print:decoration-primary hover:decoration-primary transition-colors">agents fail quietly</a>, in ways that pass the checks you remembered to write. What I hadn&#x27;t appreciated is that the quiet failure is often <em class="font-mono font-normal text-primary/80 print:text-primary">in the checking layer itself</em>, and that it persists happily through green runs. The successful run and the broken verifier coexist, and one of them is visible to you.</p>
<p class="text-muted-foreground leading-7 my-4">So the green dashboard is not evidence the verification works. It&#x27;s evidence the task succeeded, which is a different claim, and on this data those two things come apart routinely.</p>
<p class="text-muted-foreground leading-7 my-4">One honesty note: this success/failure split is a smaller, per-system analysis than the 1,642-trace headline, and the rounding in the published rates suggests roughly ten to a dozen successful traces per system. Treat the direction as the finding, not the magnitude.</p>
<h2 id="what-actually-moved-the-number" class="font-display text-2xl lg:text-3xl font-bold text-foreground mt-12 mb-6 border-l-2 border-primary pl-4"><a href="#what-actually-moved-the-number" class="heading-permalink text-inherit no-underline">What Actually Moved the Number</a></h2>
<p class="text-muted-foreground leading-7 my-4">The paper doesn&#x27;t stop at taxonomy. The authors ran interventions on ChatDev and measured them, and the verification-flavored one is the strongest result in the paper.</p>
<p class="text-muted-foreground leading-7 my-4">They changed the system&#x27;s shape so that it terminates <strong class="text-foreground font-bold">only when a designated agent confirms all reviews are properly satisfied</strong>, with an iteration cutoff to prevent infinite loops. Before, the pipeline ended when it reached the end of the pipeline. After, it ended when an explicit objective check passed.</p>
<p class="text-muted-foreground leading-7 my-4">On a custom 32-task benchmark the authors call ProgramDev-v0, success went from <strong class="text-foreground font-bold">25.0 to 40.6 points</strong>.</p>
<p class="text-muted-foreground leading-7 my-4">Three caveats, and I want to give them properly because this figure travels badly.</p>
<p class="text-muted-foreground leading-7 my-4">First, those are percentage points off a very low base. You will see this result quoted as &quot;+15.6% improvement,&quot; which reads as relative and overstates it. The system went from failing three quarters of the time to failing three fifths of the time.</p>
<p class="text-muted-foreground leading-7 my-4">Second, the benchmark is 32 custom tasks — the paper explicitly distinguishes ProgramDev-v0 from the ProgramDev dataset it discusses elsewhere. On the second benchmark, HumanEval, the identical change moved success from <strong class="text-foreground font-bold">89.6 to 91.5</strong>. That&#x27;s under two points, on a benchmark already near ceiling. This is not a general result and shouldn&#x27;t be sold as one.</p>
<p class="text-muted-foreground leading-7 my-4">Third, the authors themselves decline to oversell it:</p>
<blockquote class="my-6 border-l-2 border-primary/40 print:border-primary pl-6">
<p class="text-muted-foreground leading-7 my-4">&quot;Even though our interventions are successful in improving the performance of the framework in different tasks, they do not constitute substantial improvements.&quot;</p>
</blockquote>
<p class="text-muted-foreground leading-7 my-4">I still think it&#x27;s the most instructive number here, precisely because of how unspectacular it is. The mechanism is what matters: making termination conditional on an explicit objective check, rather than on having reached the end of the workflow. That is a one-line change in what &quot;done&quot; means, and on the benchmark where there was room to move, it moved more than anything else they tried.</p>
<h2 id="four-things-to-measure-instead" class="font-display text-2xl lg:text-3xl font-bold text-foreground mt-12 mb-6 border-l-2 border-primary pl-4"><a href="#four-things-to-measure-instead" class="heading-permalink text-inherit no-underline">Four Things to Measure Instead</a></h2>
<p class="text-muted-foreground leading-7 my-4">I don&#x27;t want to end on a checklist, so let me end on the smallest set of questions your current setup probably can&#x27;t answer.</p>
<p class="text-muted-foreground leading-7 my-4"><strong class="text-foreground font-bold">Can you make your verifier fail?</strong> Take a known-bad output — genuinely wrong, not malformed — and run it through your eval suite. If it passes, you&#x27;ve measured the thing you needed to measure. Mutation testing is old, unglamorous, and almost nobody applies it to agent evals, which is odd given that 9.10% is the single largest verification failure mode.</p>
<p class="text-muted-foreground leading-7 my-4"><strong class="text-foreground font-bold">Do you distinguish &quot;finished&quot; from &quot;stopped&quot;?</strong> Premature termination is 6.20% on its own, and it&#x27;s the mode with no artifact to inspect — there&#x27;s no wrong answer sitting there, just less right answer than there should have been. If your telemetry records completion but not <em class="font-mono font-normal text-primary/80 print:text-primary">why</em> the loop exited — objective satisfied, step budget exhausted, tool error swallowed — you cannot see this class at all.</p>
<p class="text-muted-foreground leading-7 my-4"><strong class="text-foreground font-bold">Does anything check the coverage of the check?</strong> No-or-incomplete verification is 8.20%, and &quot;incomplete&quot; is the operative word. A verifier that inspects three of the seven things that had to be true will report success four sevenths of the time it shouldn&#x27;t.</p>
<p class="text-muted-foreground leading-7 my-4"><strong class="text-foreground font-bold">Are you looking at the successful runs?</strong> This is the one that follows directly from the split above, and it&#x27;s the cheapest. Sample runs that passed. Read the traces. Verification flaws live there, and by definition no alert will ever route you to them.</p>
<h2 id="the-honest-caveat" class="font-display text-2xl lg:text-3xl font-bold text-foreground mt-12 mb-6 border-l-2 border-primary pl-4"><a href="#the-honest-caveat" class="heading-permalink text-inherit no-underline">The Honest Caveat</a></h2>
<p class="text-muted-foreground leading-7 my-4">MAST is a taxonomy paper built on annotated traces, and the annotation is human judgment applied at scale — the authors report strong inter-annotator agreement, but on a small labelled sample that was then extended by an LLM judge. The intervention results are one system, two benchmarks, one of which is 32 tasks. And the LangChain production figure comes from a self-selected community survey — 63% technology industry, roughly half at companies under a hundred people — so read it as &quot;people who build agents&quot; rather than &quot;enterprises.&quot;</p>
<p class="text-muted-foreground leading-7 my-4">None of that changes the shape of the argument, because the argument doesn&#x27;t need the numbers to be precise. It needs them to be the right <em class="font-mono font-normal text-primary/80 print:text-primary">kind</em> of number, and they are: an independent look at what actually goes wrong, which found that the checking layer is a first-class source of failure rather than the neutral instrument everyone treats it as.</p>
<p class="text-muted-foreground leading-7 my-4">Your eval suite is not a measuring device pointed at your agent. It is another component in the system, written with the same care as the rest of it, failing in the same ways, and — unlike everything else you built — there is nothing downstream of it that would notice if it broke.</p>
<p class="text-muted-foreground leading-7 my-4">Everyone measures whether the agent works.</p>
<p class="text-muted-foreground leading-7 my-4">Start measuring whether you&#x27;d know if it stopped.</p>]]></content:encoded>
      <pubDate>Tue, 11 Aug 2026 12:00:00 GMT</pubDate>
    </item>
    <item>
      <title>The Handoff Is Where Agents Break</title>
      <link>https://pratik.pa.tel/blog/the-handoff-is-where-agents-break/</link>
      <guid isPermaLink="true">https://pratik.pa.tel/blog/the-handoff-is-where-agents-break/</guid>
      <description>You debug the agent that produced the bad output. The bug was in the message it got handed.</description>
      <dc:creator>Pratik Patel</dc:creator>
      <category>ai</category>
      <category>agents</category>
      <category>reliability</category>
      <category>multi-agent</category>
      <content:encoded><![CDATA[<p class="text-muted-foreground leading-7 my-4">Something goes wrong in your multi-agent system. The final output is confidently, specifically wrong.</p>
<p class="text-muted-foreground leading-7 my-4">So you open the last agent&#x27;s trace. You read its prompt, its reasoning, its tool calls. And the frustrating part is that none of it looks broken. Given what it was told, it did something reasonable. You swap in a better model. Same class of failure. You rewrite the prompt. Same class of failure.</p>
<p class="text-muted-foreground leading-7 my-4">The agent was fine. The brief it received wasn&#x27;t.</p>
<p class="text-muted-foreground leading-7 my-4">This is the thing that surprises people about running more than one agent: the failures stop living inside the agents and start living in the space between them. We <a href="/blog/your-second-agent-is-the-hard-one/" class="text-primary underline underline-offset-2 decoration-primary/40 print:decoration-primary hover:decoration-primary transition-colors">wrote about that shift</a> when you go from one agent to two. This post is about the part that comes after the realization — actually engineering that space.</p>
<h2 id="the-failure-rates-are-worse-than-the-demos-suggest" class="font-display text-2xl lg:text-3xl font-bold text-foreground mt-12 mb-6 border-l-2 border-primary pl-4"><a href="#the-failure-rates-are-worse-than-the-demos-suggest" class="heading-permalink text-inherit no-underline">The Failure Rates Are Worse Than the Demos Suggest</a></h2>
<p class="text-muted-foreground leading-7 my-4">There is now real data on this, and it is bracing.</p>
<p class="text-muted-foreground leading-7 my-4">Researchers behind the <a href="https://arxiv.org/abs/2503.13657" class="text-primary underline underline-offset-2 decoration-primary/40 print:decoration-primary hover:decoration-primary transition-colors">MAST taxonomy</a> annotated 1,642 execution traces across seven state-of-the-art open-source multi-agent systems. It is worth knowing exactly how, because the how is what makes it usable: six human experts read 150+ traces to build the taxonomy, three of them then labeled independently until they reached 0.88 Cohen&#x27;s kappa, and the full 1,642-trace dataset was labeled by an LLM-as-judge pipeline calibrated against those human annotations (94% accuracy, 0.77 kappa).</p>
<p class="text-muted-foreground leading-7 my-4">So it is not 1,642 hand-reads, and anyone citing it should say so. What it is, is a human-built and human-validated taxonomy applied at a scale humans could not reach — the closest thing the field has to a failure epidemiology.</p>
<p class="text-muted-foreground leading-7 my-4">Their headline: <strong class="text-foreground font-bold">41% to 86.7% failure rates.</strong> The six systems the paper breaks out individually: AppWorld failed 86.7% of the time, HyperAgent 74.7%, ChatDev 66.7%, Magentic-One 62.0%, MetaGPT 60.0%, AG2 41.0%.</p>
<p class="text-muted-foreground leading-7 my-4">One honest caveat before anyone quotes that range at a standup, and it is the paper&#x27;s own: these systems were measured on <em class="font-mono font-normal text-primary/80 print:text-primary">different benchmarks</em>, so the numbers are not directly comparable to each other. It is not a leaderboard. What it is, is a floor-level observation that serious multi-agent systems built by serious people fail a lot, and the good ones still fail more than you would guess from a demo video.</p>
<h2 id="a-third-of-it-is-the-seam" class="font-display text-2xl lg:text-3xl font-bold text-foreground mt-12 mb-6 border-l-2 border-primary pl-4"><a href="#a-third-of-it-is-the-seam" class="heading-permalink text-inherit no-underline">A Third of It Is the Seam</a></h2>
<p class="text-muted-foreground leading-7 my-4">Here is the number this post is actually about.</p>
<p class="text-muted-foreground leading-7 my-4">MAST sorts every failure into three categories. <strong class="text-foreground font-bold">Inter-agent misalignment accounts for 32.3% of them.</strong> Not the model being incapable. Not the task being impossible. Roughly a third of failures are agents failing to communicate correctly with each other.</p>
<p class="text-muted-foreground leading-7 my-4">Break that category open — all six sub-modes, so the arithmetic is checkable — and they are painfully recognizable:</p>
<ul class="space-y-2 my-6 ml-4">
<li class="flex gap-3 text-muted-foreground leading-7"><span class="text-primary shrink-0 mt-1.5">▸</span><span class="min-w-0"><strong class="text-foreground font-bold">Reasoning-action mismatch — 13.2%.</strong> The agent&#x27;s stated reasoning and its actual action diverge. It says it will check the schema, then doesn&#x27;t.</span></li>
<li class="flex gap-3 text-muted-foreground leading-7"><span class="text-primary shrink-0 mt-1.5">▸</span><span class="min-w-0"><strong class="text-foreground font-bold">Task derailment — 7.4%.</strong> The work drifts from what was asked. Each handoff nudges it slightly, and no single step looks wrong.</span></li>
<li class="flex gap-3 text-muted-foreground leading-7"><span class="text-primary shrink-0 mt-1.5">▸</span><span class="min-w-0"><strong class="text-foreground font-bold">Failure to ask for clarification — 6.8%.</strong> The agent receives an ambiguous brief and resolves the ambiguity by guessing, silently.</span></li>
<li class="flex gap-3 text-muted-foreground leading-7"><span class="text-primary shrink-0 mt-1.5">▸</span><span class="min-w-0"><strong class="text-foreground font-bold">Conversation reset — 2.2%.</strong> The thread restarts and everything established up to that point is gone.</span></li>
<li class="flex gap-3 text-muted-foreground leading-7"><span class="text-primary shrink-0 mt-1.5">▸</span><span class="min-w-0"><strong class="text-foreground font-bold">Ignoring another agent&#x27;s input — 1.9%.</strong> The information arrived. It was not used.</span></li>
<li class="flex gap-3 text-muted-foreground leading-7"><span class="text-primary shrink-0 mt-1.5">▸</span><span class="min-w-0"><strong class="text-foreground font-bold">Information withholding — 0.85%.</strong> An agent knows something relevant and does not pass it on.</span></li>
</ul>
<p class="text-muted-foreground leading-7 my-4">Look at that list as a whole and a pattern falls out. Almost none of these are intelligence failures. They are <strong class="text-foreground font-bold">protocol</strong> failures. The receiving agent was never told what it needed, never told what it was allowed to assume, and never given a way to say &quot;this brief is underspecified.&quot;</p>
<p class="text-muted-foreground leading-7 my-4">One thing worth being precise about, because the sloppy version of this argument is everywhere: the 41–86.7% range is an <em class="font-mono font-normal text-primary/80 print:text-primary">overall</em> failure rate, not a coordination-failure rate. The paper does not claim coordination causes most failures. It measures that inter-agent misalignment is 32.3% of them. The argument that this is the most <em class="font-mono font-normal text-primary/80 print:text-primary">fixable</em> third is mine, not theirs.</p>
<h2 id="why-the-seam-is-invisible" class="font-display text-2xl lg:text-3xl font-bold text-foreground mt-12 mb-6 border-l-2 border-primary pl-4"><a href="#why-the-seam-is-invisible" class="heading-permalink text-inherit no-underline">Why the Seam Is Invisible</a></h2>
<p class="text-muted-foreground leading-7 my-4">Every agent in the chain looks locally correct. That is the whole problem.</p>
<p class="text-muted-foreground leading-7 my-4">When you debug a single agent, you have a clean pair: the input you gave it, the output it produced. When you debug a chain, the input to agent three is an artifact produced by agent two, which was itself working from agent one&#x27;s artifact. By the time you inspect the failure, the original intent has been paraphrased twice.</p>
<p class="text-muted-foreground leading-7 my-4">Paraphrasing is lossy. Not dramatically — that would be easy to catch. It is lossy at the edges, in exactly the places that turn out to matter: the constraint that was mentioned once, the exception the user flagged in passing, the &quot;don&#x27;t touch the billing table&quot; that agent one understood as context and agent two reasonably summarized away.</p>
<p class="text-muted-foreground leading-7 my-4">This is why &quot;just use a better model&quot; doesn&#x27;t fix it. A better model paraphrases more fluently. It does not know which edge case you were going to care about, because that information was destroyed one hop upstream.</p>
<h2 id="engineer-the-seam-like-an-api" class="font-display text-2xl lg:text-3xl font-bold text-foreground mt-12 mb-6 border-l-2 border-primary pl-4"><a href="#engineer-the-seam-like-an-api" class="heading-permalink text-inherit no-underline">Engineer the Seam Like an API</a></h2>
<p class="text-muted-foreground leading-7 my-4">The fix is not more intelligence. It is treating the handoff as an interface with a contract, the same way you would treat a service boundary between two teams.</p>
<p class="text-muted-foreground leading-7 my-4"><strong class="text-foreground font-bold">Pass structure, not prose.</strong> The single highest-leverage change is to stop letting agents hand each other paragraphs. A paragraph invites paraphrase; a schema doesn&#x27;t. If agent two must produce <code class="font-mono rounded bg-muted px-1.5 py-0.5 text-foreground print:border print:border-border">{ file, change, constraints[], unresolved[] }</code>, then <code class="font-mono rounded bg-muted px-1.5 py-0.5 text-foreground print:border print:border-border">constraints</code> cannot get quietly dropped in a rewrite — a missing field is a visible, checkable error rather than an absence nobody notices.</p>
<p class="text-muted-foreground leading-7 my-4"><strong class="text-foreground font-bold">Make &quot;I don&#x27;t know&quot; a first-class value.</strong> The 6.8% clarification-failure mode exists because the schema had no slot for uncertainty. Give every handoff an explicit <code class="font-mono rounded bg-muted px-1.5 py-0.5 text-foreground print:border print:border-border">unresolved</code> list, and make it legal — expected, even — for an agent to hand back work with three open questions instead of a confidently guessed answer. An agent that guesses silently is not being helpful; it is <a href="/blog/agents-fail-quietly/" class="text-primary underline underline-offset-2 decoration-primary/40 print:decoration-primary hover:decoration-primary transition-colors">failing quietly</a>, and quiet failure is the expensive kind.</p>
<p class="text-muted-foreground leading-7 my-4"><strong class="text-foreground font-bold">Validate at the boundary, not at the end.</strong> Check the handoff artifact the moment it is produced, against the contract, before the next agent starts. A malformed brief caught at the seam costs one retry. The same brief caught after three more agents have built on it costs the whole run — and by then the evidence of what went wrong has been paraphrased out of existence.</p>
<p class="text-muted-foreground leading-7 my-4"><strong class="text-foreground font-bold">Carry provenance forward.</strong> Every constraint in a handoff should say where it came from: the user, an upstream agent, or the agent&#x27;s own inference. This one line of metadata is what lets a downstream agent — or you, reading the trace — tell the difference between a hard requirement and a guess that has been repeated so many times it looks like one.</p>
<p class="text-muted-foreground leading-7 my-4"><strong class="text-foreground font-bold">Keep the original brief reachable.</strong> Don&#x27;t let the user&#x27;s actual words disappear behind two layers of summary. Pass a pointer to the source alongside the summary, so any agent in the chain can go check the original instead of trusting the paraphrase. This is the cheapest possible fix and it is skipped constantly.</p>
<h2 id="the-boring-discipline" class="font-display text-2xl lg:text-3xl font-bold text-foreground mt-12 mb-6 border-l-2 border-primary pl-4"><a href="#the-boring-discipline" class="heading-permalink text-inherit no-underline">The Boring Discipline</a></h2>
<p class="text-muted-foreground leading-7 my-4">None of this is exotic. It is schema validation, explicit nullability, provenance, and error handling — the same things we learned to do at service boundaries twenty years ago, applied to a boundary we have been pretending is a conversation.</p>
<p class="text-muted-foreground leading-7 my-4">That pretense is the root of it. We describe agents as &quot;talking to each other,&quot; and that framing quietly imports every assumption we make about human conversation: that context is shared, that ambiguity gets resolved by asking, that the listener will flag something that sounds off. None of that is true by default here. It is true only if you build it.</p>
<p class="text-muted-foreground leading-7 my-4">The good news is that this is the tractable third. You cannot make the model smarter this quarter. You can absolutely make the message between agent two and agent three a validated object with a slot for &quot;here is what I wasn&#x27;t sure about.&quot;</p>
<p class="text-muted-foreground leading-7 my-4">Start there. The next time an agent produces something confidently wrong, before you touch its prompt, go read what it was handed. The bug is upstream more often than it has any right to be.</p>]]></content:encoded>
      <pubDate>Tue, 04 Aug 2026 12:00:00 GMT</pubDate>
    </item>
    <item>
      <title>Your Second Agent Is the Hard One</title>
      <link>https://pratik.pa.tel/blog/your-second-agent-is-the-hard-one/</link>
      <guid isPermaLink="true">https://pratik.pa.tel/blog/your-second-agent-is-the-hard-one/</guid>
      <description>Going from one agent to many does not multiply what you can do. It moves every failure into the space between them.</description>
      <dc:creator>Pratik Patel</dc:creator>
      <category>ai</category>
      <category>agents</category>
      <category>reliability</category>
      <category>engineering</category>
      <content:encoded><![CDATA[<p class="text-muted-foreground leading-7 my-4">Two of the most-shared pieces of engineering writing about agents say the opposite thing.</p>
<p class="text-muted-foreground leading-7 my-4">Cognition published <a href="https://cognition.com/blog/dont-build-multi-agents" class="text-primary underline underline-offset-2 decoration-primary/40 print:decoration-primary hover:decoration-primary transition-colors">Don&#x27;t Build Multi-Agents</a>, arguing that splitting work across parallel agents is fragile by construction. Anthropic published its <a href="https://www.anthropic.com/engineering/built-multi-agent-research-system" class="text-primary underline underline-offset-2 decoration-primary/40 print:decoration-primary hover:decoration-primary transition-colors">multi-agent research system</a>, reporting that an orchestrator with parallel subagents beat a single strong model by 90.2 percent on their internal research eval.</p>
<p class="text-muted-foreground leading-7 my-4">Both teams are serious. Both are describing real systems. And most people reading them treat the disagreement as a matter of taste, or of who had the better framework.</p>
<p class="text-muted-foreground leading-7 my-4">It isn&#x27;t. They&#x27;re describing two different kinds of work, and the difference between those two kinds is the single most useful thing I know about scaling from one agent to a fleet.</p>
<h2 id="the-first-agent-is-an-engineering-problem" class="font-display text-2xl lg:text-3xl font-bold text-foreground mt-12 mb-6 border-l-2 border-primary pl-4"><a href="#the-first-agent-is-an-engineering-problem" class="heading-permalink text-inherit no-underline">The First Agent Is an Engineering Problem</a></h2>
<p class="text-muted-foreground leading-7 my-4">Everything I&#x27;ve written about agent reliability so far assumes a boundary around one agent. Decide <a href="/blog/agent-permissions-are-product-design/" class="text-primary underline underline-offset-2 decoration-primary/40 print:decoration-primary hover:decoration-primary transition-colors">what it&#x27;s allowed to touch</a>. Give it <a href="/blog/agent-runbooks-beat-better-prompts/" class="text-primary underline underline-offset-2 decoration-primary/40 print:decoration-primary hover:decoration-primary transition-colors">a runbook instead of a better prompt</a>. Notice <a href="/blog/agents-fail-quietly/" class="text-primary underline underline-offset-2 decoration-primary/40 print:decoration-primary hover:decoration-primary transition-colors">when it fails quietly</a>. Make its mistakes <a href="/blog/give-your-agent-an-undo-button/" class="text-primary underline underline-offset-2 decoration-primary/40 print:decoration-primary hover:decoration-primary transition-colors">cheap to take back</a>. Teach it <a href="/blog/teach-your-agent-to-ask-for-help/" class="text-primary underline underline-offset-2 decoration-primary/40 print:decoration-primary hover:decoration-primary transition-colors">when to stop and ask</a>. Keep <a href="/blog/trust-comes-from-the-trace/" class="text-primary underline underline-offset-2 decoration-primary/40 print:decoration-primary hover:decoration-primary transition-colors">a trace of what it did</a>.</p>
<p class="text-muted-foreground leading-7 my-4">Every one of those controls is scoped to a single actor. The permission set belongs to an agent. The undo button unwinds one agent&#x27;s actions. The escalation gate pauses one agent&#x27;s decision. The trace reconstructs one agent&#x27;s run.</p>
<p class="text-muted-foreground leading-7 my-4">Add a second agent and none of those controls break in an obvious way. They just stop covering the interesting part, because the interesting part is now between the two of them.</p>
<p class="text-muted-foreground leading-7 my-4">Who owns the shared file when both agents want to write it? When agent A assumes the API returns cents and agent B assumes dollars, whose trace shows the bug? When the run goes wrong at step nine of a forty-step plan spread across five agents, which agent gets the pause? Nothing in a single-agent design answers these. Each one is a question about authority: who owns what, and who decides when two agents disagree.</p>
<p class="text-muted-foreground leading-7 my-4">That&#x27;s the actual shape of the jump. The first agent is a systems problem you can solve with better tools. The second agent is a coordination problem you solve with better contracts, and coordination problems have never been solved by making the individual participants smarter.</p>
<h2 id="the-failures-are-boring-which-is-the-bad-news" class="font-display text-2xl lg:text-3xl font-bold text-foreground mt-12 mb-6 border-l-2 border-primary pl-4"><a href="#the-failures-are-boring-which-is-the-bad-news" class="heading-permalink text-inherit no-underline">The Failures Are Boring, Which Is the Bad News</a></h2>
<p class="text-muted-foreground leading-7 my-4">There&#x27;s now real data on how these systems fail. A team from Berkeley and collaborators read <a href="https://arxiv.org/abs/2503.13657" class="text-primary underline underline-offset-2 decoration-primary/40 print:decoration-primary hover:decoration-primary transition-colors">150 execution traces</a> closely enough, and agreed with each other closely enough (Cohen&#x27;s kappa of 0.88), to build a taxonomy of what actually goes wrong. It has fourteen distinct failure modes in three buckets: system design issues, inter-agent misalignment, and task verification. They then built an LLM annotator against that taxonomy and ran it over more than 1,600 traces from seven popular multi-agent frameworks, which is how we know the taxonomy holds beyond the traces that produced it.</p>
<p class="text-muted-foreground leading-7 my-4">Read that list again and notice what isn&#x27;t on it. Not &quot;the model wasn&#x27;t capable enough.&quot; Not &quot;reasoning was too shallow.&quot; The failures are specification, coordination, and checking the work. If you replaced every agent in those traces with a smarter model, most of those traces would still fail, which is roughly what the authors concluded when they said the failures they found demand more sophisticated solutions than surface-level fixes.</p>
<p class="text-muted-foreground leading-7 my-4">I find this genuinely clarifying, and a little deflating. We&#x27;re rediscovering, at great expense, that a team of capable individuals with unclear ownership and no shared definition of done produces worse work than one competent person. Gartner&#x27;s forecast that <a href="https://www.gartner.com/en/newsroom/press-releases/2025-06-25-gartner-predicts-over-40-percent-of-agentic-ai-projects-will-be-canceled-by-end-of-2027" class="text-primary underline underline-offset-2 decoration-primary/40 print:decoration-primary hover:decoration-primary transition-colors">over 40 percent of agentic AI projects will be canceled by the end of 2027</a> blames escalating cost, unclear value, and inadequate risk controls. All three are management failures.</p>
<p class="text-muted-foreground leading-7 my-4">Cognition names the mechanism precisely. Actions carry implicit decisions, and conflicting decisions carry bad results. Their example is a Flappy Bird clone split between two subagents: one renders a background in the style of Super Mario Bros., the other builds a bird that doesn&#x27;t match anything, and the orchestrator is left holding two incompatible halves. Neither subagent did anything wrong by its own lights. They just each made a silent decision the other never saw.</p>
<h2 id="parallelize-reading-serialize-writing" class="font-display text-2xl lg:text-3xl font-bold text-foreground mt-12 mb-6 border-l-2 border-primary pl-4"><a href="#parallelize-reading-serialize-writing" class="heading-permalink text-inherit no-underline">Parallelize Reading. Serialize Writing.</a></h2>
<p class="text-muted-foreground leading-7 my-4">So why did Anthropic&#x27;s fleet work?</p>
<p class="text-muted-foreground leading-7 my-4">Look at what those subagents were doing. Searching. Reading. Exploring independent branches of a research question, then handing back findings that get synthesized by a lead agent. Nothing any subagent does changes what another subagent sees. The context each one gathers is additive. When two agents turn up the same source, the worst thing that happens is you paid twice for it. And the cost of that architecture was roughly fifteen times the tokens of a normal chat, which is a fine trade when the alternative is a question that doesn&#x27;t fit in one context window at all.</p>
<p class="text-muted-foreground leading-7 my-4">Now look at the Flappy Bird case. Both subagents were writing to one artifact. Every choice one made silently constrained the other. Their outputs collide.</p>
<p class="text-muted-foreground leading-7 my-4">That&#x27;s the variable. It isn&#x27;t framework, orchestration topology, or prompt quality:</p>
<p class="text-muted-foreground leading-7 my-4"><strong class="text-foreground font-bold">Parallelize reading. Serialize writing.</strong></p>
<p class="text-muted-foreground leading-7 my-4">Fan out as widely as you like across work that only reads. Search, retrieval, analysis, exploring five approaches to see which is viable, reviewing a diff from six angles. The results combine cleanly because nothing contended for shared state. This is exactly the workload where a fleet pays for itself, because you&#x27;re buying breadth that a single context window can&#x27;t hold.</p>
<p class="text-muted-foreground leading-7 my-4">Then narrow to one writer for anything that mutates shared state. One agent holds the pen for a given artifact. Not one agent for the whole system. One agent per thing that can be written.</p>
<p class="text-muted-foreground leading-7 my-4">Most multi-agent designs I see get this backwards. They parallelize the writes, because that&#x27;s where the wall-clock time is, and then spend the savings debugging conflicts that never had to exist.</p>
<h2 id="contracts-not-conversations" class="font-display text-2xl lg:text-3xl font-bold text-foreground mt-12 mb-6 border-l-2 border-primary pl-4"><a href="#contracts-not-conversations" class="heading-permalink text-inherit no-underline">Contracts, Not Conversations</a></h2>
<p class="text-muted-foreground leading-7 my-4">Once you accept that constraint, the rest of a fleet design falls out of it.</p>
<p class="text-muted-foreground leading-7 my-4"><strong class="text-foreground font-bold">Make handoffs artifacts, not chat.</strong> Agents passing messages recreates every ambiguity of a hallway conversation, and current models are bad at the social reasoning that makes hallway conversations work. Have agents produce structured, inspectable outputs that the next agent consumes as input. A file, a diff, a typed result. Something you can look at afterward and say: this was correct when it left agent A, and agent B mishandled it. Chat transcripts don&#x27;t let you say that.</p>
<p class="text-muted-foreground leading-7 my-4"><strong class="text-foreground font-bold">Pass the full trace, not the summary.</strong> Cognition&#x27;s first principle is to share complete agent traces rather than individual messages, and it&#x27;s the right call. A subagent that receives a one-line task description will invent everything the description left out. The invented parts are where conflicts come from.</p>
<p class="text-muted-foreground leading-7 my-4"><strong class="text-foreground font-bold">Give shared state an owner.</strong> For every resource two agents can touch, name the one that writes. The others read. If two agents genuinely must write the same thing, that&#x27;s not a coordination problem to solve with better prompts. That&#x27;s a design smell telling you the work was split along the wrong seam.</p>
<p class="text-muted-foreground leading-7 my-4"><strong class="text-foreground font-bold">Trace across the fleet, not per agent.</strong> A per-agent trace tells you what each one did. It won&#x27;t tell you that agent B&#x27;s correct action was based on agent A&#x27;s wrong assumption. You need spans that link, so the causal chain survives the hop between agents. Reconstruction was already <a href="/blog/trust-comes-from-the-trace/" class="text-primary underline underline-offset-2 decoration-primary/40 print:decoration-primary hover:decoration-primary transition-colors">the foundation of trust in a single agent</a>. In a fleet it&#x27;s the only thing standing between you and a bug that no single trace contains.</p>
<p class="text-muted-foreground leading-7 my-4"><strong class="text-foreground font-bold">Escalate at the seams.</strong> The most valuable place for a human pause isn&#x27;t inside an agent&#x27;s reasoning. It&#x27;s at the handoff, where one agent&#x27;s output becomes another&#x27;s assumption, and where a wrong assumption is still cheap to catch.</p>
<h2 id="ask-what-theyll-contend-over" class="font-display text-2xl lg:text-3xl font-bold text-foreground mt-12 mb-6 border-l-2 border-primary pl-4"><a href="#ask-what-theyll-contend-over" class="heading-permalink text-inherit no-underline">Ask What They&#x27;ll Contend Over</a></h2>
<p class="text-muted-foreground leading-7 my-4">The question to ask before adding an agent isn&#x27;t &quot;could a second agent do this in parallel.&quot; It almost always could. The question is what the two of them will contend over, and who decides when they disagree.</p>
<p class="text-muted-foreground leading-7 my-4">If the answer is nothing, they contend over nothing, then fan out. Run ten. The read-only fleet is one of the genuinely good deals in this technology, and it&#x27;s why the research-agent results are real.</p>
<p class="text-muted-foreground leading-7 my-4">If the answer involves shared state, a shared artifact, or a decision that has to be consistent across both, then a second agent doesn&#x27;t split the work. It splits the ownership, and hands you a distributed-systems problem you now have to solve in natural language, with participants that don&#x27;t reliably remember what they agreed to.</p>
<p class="text-muted-foreground leading-7 my-4">One agent that reads widely and writes carefully will beat five that all think they&#x27;re holding the pen.</p>
<p class="text-muted-foreground leading-7 my-4">Your second agent isn&#x27;t a second worker. It&#x27;s the first day your agents needed a manager.</p>]]></content:encoded>
      <pubDate>Tue, 28 Jul 2026 12:00:00 GMT</pubDate>
    </item>
    <item>
      <title>Trust Comes From the Trace</title>
      <link>https://pratik.pa.tel/blog/trust-comes-from-the-trace/</link>
      <guid isPermaLink="true">https://pratik.pa.tel/blog/trust-comes-from-the-trace/</guid>
      <description>You bounded what your agent can touch, gave it an undo button, and taught it when to ask. But when it does something surprising, can you reconstruct exactly why? That is the observability gap.</description>
      <dc:creator>Pratik Patel</dc:creator>
      <category>ai</category>
      <category>agents</category>
      <category>reliability</category>
      <category>observability</category>
      <category>engineering</category>
      <content:encoded><![CDATA[<p class="text-muted-foreground leading-7 my-4">An agent closes a customer ticket it should have escalated. You go looking for what happened. The request logged a clean 200. The final message reads fine. Latency was normal. Every line in the log says the run was healthy.</p>
<p class="text-muted-foreground leading-7 my-4">And you still have no idea why it did what it did.</p>
<p class="text-muted-foreground leading-7 my-4">That is the gap I want to talk about. Not &quot;did the agent fail,&quot; which you can often catch. Something harder. &quot;The agent did something surprising, and I cannot reconstruct the path it took to get there.&quot;</p>
<p class="text-muted-foreground leading-7 my-4">I have written about bounding what an agent can touch, measuring when it fails quietly, giving it an undo button, and teaching it when to ask a human. Each of those makes the agent&#x27;s mistakes smaller or more survivable. None of them tells you <em class="font-mono font-normal text-primary/80 print:text-primary">why</em> a specific run went the way it did. For that you need to be able to replay the decision. And most teams cannot.</p>
<h2 id="a-log-says-what-happened-it-does-not-say-why" class="font-display text-2xl lg:text-3xl font-bold text-foreground mt-12 mb-6 border-l-2 border-primary pl-4"><a href="#a-log-says-what-happened-it-does-not-say-why" class="heading-permalink text-inherit no-underline">A Log Says What Happened. It Does Not Say Why.</a></h2>
<p class="text-muted-foreground leading-7 my-4">Traditional monitoring was built for deterministic services. The same input produces the same output, every request follows a known code path, and a 200 is a strong signal that things went right. Agents break all three assumptions. The same prompt can produce different tool calls on different runs, the path branches on model output, and <a href="https://www.braintrust.dev/articles/agent-tracing-debug-ai-agents-production" class="text-primary underline underline-offset-2 decoration-primary/40 print:decoration-primary hover:decoration-primary transition-colors">a 200 can wrap a confidently wrong answer</a>.</p>
<p class="text-muted-foreground leading-7 my-4">So the usual stack shows you the outside of the run and none of the inside. It can tell you the request returned successfully. It cannot tell you that the agent <a href="https://www.braintrust.dev/articles/agent-tracing-debug-ai-agents-production" class="text-primary underline underline-offset-2 decoration-primary/40 print:decoration-primary hover:decoration-primary transition-colors">looped twice, called the wrong tool, and hallucinated a billing policy</a> on the way to that successful-looking answer. The most dangerous agent failures are the ones that look like success: well-formed but wrong outputs, redundant tool calls, semantically invalid actions. A binary up-or-down check waves all of those through.</p>
<p class="text-muted-foreground leading-7 my-4">The uncomfortable part is how few teams have anything better. Gartner puts spending on LLM observability at <a href="https://www.gartner.com/en/newsroom/press-releases/2026-03-30-gartner-predicts-by-2028-explainable-ai-will-drive-llm-observability-investments-to-50-percent-for-secure-genai-deployment" class="text-primary underline underline-offset-2 decoration-primary/40 print:decoration-primary hover:decoration-primary transition-colors">15% of GenAI deployments today, and expects it to reach 50% by 2028</a>. The direction of travel is obvious, but the baseline is the part worth sitting with: most agents in production right now are black boxes to the people who run them.</p>
<h2 id="a-trace-is-the-inside-of-the-decision" class="font-display text-2xl lg:text-3xl font-bold text-foreground mt-12 mb-6 border-l-2 border-primary pl-4"><a href="#a-trace-is-the-inside-of-the-decision" class="heading-permalink text-inherit no-underline">A Trace Is the Inside of the Decision</a></h2>
<p class="text-muted-foreground leading-7 my-4">The fix is not more logs. It is a different shape of record: a trace.</p>
<p class="text-muted-foreground leading-7 my-4">A log is a flat list of events. A trace is the causal tree. Every reasoning step, every tool call, every retrieval, every decision branch becomes <a href="https://greptime.com/blogs/2026-05-09-opentelemetry-genai-semantic-conventions" class="text-primary underline underline-offset-2 decoration-primary/40 print:decoration-primary hover:decoration-primary transition-colors">a nested span linked to the step that caused it</a>. Open one and you can walk the whole path from the initial request to the final action: what the agent saw, what it decided, which tool it reached for, what came back, and what it did next.</p>
<p class="text-muted-foreground leading-7 my-4">That is the difference between &quot;the ticket got closed&quot; and &quot;the agent read the ticket, retrieved the wrong policy doc, concluded the issue was resolved, and closed it without checking the account flag.&quot; The first is a log line. The second is a trace, and it is the only one of the two you can actually learn from.</p>
<p class="text-muted-foreground leading-7 my-4">Concretely, a trace worth keeping captures three things for every run:</p>
<ul class="space-y-2 my-6 ml-4">
<li class="flex gap-3 text-muted-foreground leading-7"><span class="text-primary shrink-0 mt-1.5">▸</span><span class="min-w-0"><strong class="text-foreground font-bold">Every model call</strong>, with the prompt, the context it was given, and the raw response. Not a summary. The actual input and output, because that is where wrong decisions are born.</span></li>
<li class="flex gap-3 text-muted-foreground leading-7"><span class="text-primary shrink-0 mt-1.5">▸</span><span class="min-w-0"><strong class="text-foreground font-bold">Every tool call and retrieval</strong>, with arguments and results, linked to the reasoning step that triggered it. This is where you catch the wrong tool, the malformed argument, the stale document.</span></li>
<li class="flex gap-3 text-muted-foreground leading-7"><span class="text-primary shrink-0 mt-1.5">▸</span><span class="min-w-0"><strong class="text-foreground font-bold">Every decision branch</strong>, so a non-deterministic run is still reconstructable after the fact. When the same prompt could have gone three ways, you want to know which way this run went and what tipped it.</span></li>
</ul>
<p class="text-muted-foreground leading-7 my-4">Get those and a surprising run stops being a mystery. It becomes a recording you can scrub through.</p>
<h2 id="you-do-not-have-to-build-this-from-scratch" class="font-display text-2xl lg:text-3xl font-bold text-foreground mt-12 mb-6 border-l-2 border-primary pl-4"><a href="#you-do-not-have-to-build-this-from-scratch" class="heading-permalink text-inherit no-underline">You Do Not Have to Build This From Scratch</a></h2>
<p class="text-muted-foreground leading-7 my-4">The good news is that this is standardizing. OpenTelemetry, the same tracing standard most backends already use, <a href="https://greptime.com/blogs/2026-05-09-opentelemetry-genai-semantic-conventions" class="text-primary underline underline-offset-2 decoration-primary/40 print:decoration-primary hover:decoration-primary transition-colors">formed a GenAI group to define spans for LLM calls, agent orchestration, and MCP tool calls</a>. The conventions are still moving, so treat attribute names as not-yet-frozen, but the shape is real and you can adopt it today. Most agent frameworks and observability platforms can <a href="https://www.langchain.com/resources/agent-observability" class="text-primary underline underline-offset-2 decoration-primary/40 print:decoration-primary hover:decoration-primary transition-colors">emit and read agent traces</a> without you hand-rolling the plumbing.</p>
<p class="text-muted-foreground leading-7 my-4">If you want a place to start, start small and specific:</p>
<p class="text-muted-foreground leading-7 my-4">Instrument the model call and the tool call first. Those two spans explain the majority of surprising behavior, because that is where the agent&#x27;s intent turns into an action against the real world.</p>
<p class="text-muted-foreground leading-7 my-4">Keep the raw inputs and outputs, not summaries. The moment you compress a trace down to &quot;called billing tool, got result,&quot; you have thrown away the exact detail you will need at 2am.</p>
<p class="text-muted-foreground leading-7 my-4">Make one real run replayable end to end before you scale. If you can open a single production run and narrate every step out loud from the trace alone, the instrumentation is working. If there is a gap you have to guess across, fix that gap before you add more agents.</p>
<h2 id="observability-is-how-trust-gets-earned" class="font-display text-2xl lg:text-3xl font-bold text-foreground mt-12 mb-6 border-l-2 border-primary pl-4"><a href="#observability-is-how-trust-gets-earned" class="heading-permalink text-inherit no-underline">Observability Is How Trust Gets Earned</a></h2>
<p class="text-muted-foreground leading-7 my-4">There is a strategic reason this matters beyond debugging. The thing keeping most agent pilots from graduating to production is not capability. It is trust. Leaders will not hand real authority to a system whose decisions they cannot inspect. Gartner frames the same dynamic from the other side: as enterprises scale GenAI, <a href="https://www.gartner.com/en/newsroom/press-releases/2026-03-30-gartner-predicts-by-2028-explainable-ai-will-drive-llm-observability-investments-to-50-percent-for-secure-genai-deployment" class="text-primary underline underline-offset-2 decoration-primary/40 print:decoration-primary hover:decoration-primary transition-colors">&quot;the trust requirement grows faster than the technology itself&quot;</a> — which is exactly why observability spend is forecast to triple as a share of deployments.</p>
<p class="text-muted-foreground leading-7 my-4">That is the quiet payoff. A trace is not only how you debug a bad run. It is how you show a skeptical stakeholder exactly what the agent did and why, how you satisfy an auditor, how you prove after an incident that you can explain the machine you are running. An agent you can replay is an agent you can defend. An agent you cannot replay is one you are asking everyone to take on faith, and faith does not survive the first surprising ticket.</p>
<h2 id="the-bottom-line" class="font-display text-2xl lg:text-3xl font-bold text-foreground mt-12 mb-6 border-l-2 border-primary pl-4"><a href="#the-bottom-line" class="heading-permalink text-inherit no-underline">The Bottom Line</a></h2>
<p class="text-muted-foreground leading-7 my-4">We spend enormous effort trying to make agents make the right decision. We spend almost none making sure we can see the decision after the fact. That is backwards, because the agent will occasionally be surprising no matter how good it gets, and the only question that matters in that moment is whether you can reconstruct what it did.</p>
<p class="text-muted-foreground leading-7 my-4">Bounding limits the damage. Detection catches the miss. An undo button takes it back. Escalation routes the hard calls to a human. But all four assume you can eventually understand what happened. Observability is the layer that makes that assumption true.</p>
<p class="text-muted-foreground leading-7 my-4">So before you give your agent more to do, ask the plain question. When this does something I did not expect, can I replay exactly how it got there? If the answer is no, the problem is not that the agent needs a better prompt.</p>
<p class="text-muted-foreground leading-7 my-4">It is that you cannot see it think.</p>]]></content:encoded>
      <pubDate>Tue, 21 Jul 2026 12:00:00 GMT</pubDate>
    </item>
    <item>
      <title>Teach Your Agent to Ask for Help</title>
      <link>https://pratik.pa.tel/blog/teach-your-agent-to-ask-for-help/</link>
      <guid isPermaLink="true">https://pratik.pa.tel/blog/teach-your-agent-to-ask-for-help/</guid>
      <description>Full autonomy is the wrong goal. The best agents know exactly when to stop and hand the decision back to you.</description>
      <dc:creator>Pratik Patel</dc:creator>
      <category>ai</category>
      <category>agents</category>
      <category>reliability</category>
      <category>human-in-the-loop</category>
      <content:encoded><![CDATA[<p class="text-muted-foreground leading-7 my-4">We keep grading agents on the wrong thing.</p>
<p class="text-muted-foreground leading-7 my-4">The demos that get shared are the ones where the agent does everything itself. No hand-offs, no pauses, no human touching the keyboard. Full autonomy, start to finish. It looks like the future.</p>
<p class="text-muted-foreground leading-7 my-4">Then you put that same agent on real work, and the trait you were cheering for becomes the thing that scares you. An agent that never stops is an agent that will confidently do the one thing you would have told it not to.</p>
<p class="text-muted-foreground leading-7 my-4">The skill that actually matters in production is the opposite of what the demos reward. It&#x27;s not &quot;how much can the agent do alone.&quot; It&#x27;s &quot;does the agent know when to stop and ask.&quot;</p>
<h2 id="autonomy-is-not-the-prize" class="font-display text-2xl lg:text-3xl font-bold text-foreground mt-12 mb-6 border-l-2 border-primary pl-4"><a href="#autonomy-is-not-the-prize" class="heading-permalink text-inherit no-underline">Autonomy Is Not the Prize</a></h2>
<p class="text-muted-foreground leading-7 my-4">There is a quiet assumption baked into most agent projects: more autonomy is better, and any moment where a human has to step in is a failure we will eventually engineer away.</p>
<p class="text-muted-foreground leading-7 my-4">The teams running agents on things that matter have stopped believing that. In Okta&#x27;s 2026 survey of the agentic enterprise, <a href="https://www.okta.com/newsroom/articles/ai-agents-at-work-2026-agentic-enterprise-security/" class="text-primary underline underline-offset-2 decoration-primary/40 print:decoration-primary hover:decoration-primary transition-colors">95 percent of executives said they were confident</a> in their organization&#x27;s ability to detect an agent acting outside its intended scope. Read that again. The headline number is about <em class="font-mono font-normal text-primary/80 print:text-primary">detection</em>, not prevention. Everyone is investing in catching the agent after it steps out of bounds. Far fewer are designing the moment where the agent stops itself before it does.</p>
<p class="text-muted-foreground leading-7 my-4">That gap is the whole game. Detection tells you something went wrong. A well-placed pause stops it from going wrong in the first place.</p>
<p class="text-muted-foreground leading-7 my-4">Regulators are already treating the pause as mandatory. The EU AI Act&#x27;s oversight requirements, in force from August 2026, make demonstrable human intervention points a legal condition for high-risk autonomous systems, not a nice-to-have. &quot;The model was very capable&quot; isn&#x27;t a defense. &quot;A human approved the consequential action&quot; is.</p>
<h2 id="not-every-action-deserves-the-same-trust" class="font-display text-2xl lg:text-3xl font-bold text-foreground mt-12 mb-6 border-l-2 border-primary pl-4"><a href="#not-every-action-deserves-the-same-trust" class="heading-permalink text-inherit no-underline">Not Every Action Deserves the Same Trust</a></h2>
<p class="text-muted-foreground leading-7 my-4">The mistake is treating autonomy as a single dial you turn up or down for the whole agent. It&#x27;s not one dial. It&#x27;s a decision you make per action.</p>
<p class="text-muted-foreground leading-7 my-4">The way we sort it: every action an agent can take falls into one of four tiers, ranked by reversibility and blast radius.</p>
<p class="text-muted-foreground leading-7 my-4"><strong class="text-foreground font-bold">Tier 1. Read-only.</strong> Queries, searches, analysis. Nothing changes in the outside world. Let the agent run. Gating these just manufactures friction and trains your reviewers to rubber-stamp.</p>
<p class="text-muted-foreground leading-7 my-4"><strong class="text-foreground font-bold">Tier 2. Reversible.</strong> Drafts, internal state, anything that can be cleanly undone. Let the agent act, but log everything so you can walk it back. This is where an <a href="/blog/give-your-agent-an-undo-button/" class="text-primary underline underline-offset-2 decoration-primary/40 print:decoration-primary hover:decoration-primary transition-colors">undo button</a> earns its keep.</p>
<p class="text-muted-foreground leading-7 my-4"><strong class="text-foreground font-bold">Tier 3. External or third-party.</strong> Actions that touch systems you don&#x27;t fully control. Route these to a review queue instead of firing them off. The cost of being wrong just left your building.</p>
<p class="text-muted-foreground leading-7 my-4"><strong class="text-foreground font-bold">Tier 4. High-risk and irreversible.</strong> Production deploys, financial transactions, deleting data, changing permissions, sending external messages. These require a human approval, <em class="font-mono font-normal text-primary/80 print:text-primary">regardless of how confident the agent claims to be.</em></p>
<p class="text-muted-foreground leading-7 my-4">The last clause is the important one. The agent&#x27;s confidence isn&#x27;t a signal you can trust at Tier 4, because confidence and correctness come apart exactly when the stakes are highest.</p>
<h2 id="the-agents-confidence-is-not-evidence" class="font-display text-2xl lg:text-3xl font-bold text-foreground mt-12 mb-6 border-l-2 border-primary pl-4"><a href="#the-agents-confidence-is-not-evidence" class="heading-permalink text-inherit no-underline">The Agent&#x27;s Confidence Is Not Evidence</a></h2>
<p class="text-muted-foreground leading-7 my-4">Here is the arithmetic that should end the &quot;just let it decide&quot; argument. Chain three agents together, each reporting 90 percent confidence but actually running around 75 percent accurate. Every step has to land for the chain to land, so the accuracies multiply: 0.75 × 0.75 × 0.75, which is just over 0.42. Three steps in, real end-to-end reliability is roughly 42 percent.</p>
<p class="text-muted-foreground leading-7 my-4">The agent will tell you it&#x27;s 90 percent sure. The system is a coin flip. And it will report that 90 percent with exactly the same fluent, self-assured tone whether it&#x27;s right or catastrophically wrong.</p>
<p class="text-muted-foreground leading-7 my-4">This is why &quot;ask the agent how confident it is and gate on that&quot; quietly fails. You&#x27;re gating on a number the agent isn&#x27;t qualified to produce. The trigger for a pause can&#x27;t be the agent&#x27;s <em class="font-mono font-normal text-primary/80 print:text-primary">feeling</em> about the action. It has to be the <em class="font-mono font-normal text-primary/80 print:text-primary">category</em> of the action. Deleting a database is Tier 4 whether the agent is nervous or serene about it.</p>
<h2 id="design-the-pause-like-a-feature" class="font-display text-2xl lg:text-3xl font-bold text-foreground mt-12 mb-6 border-l-2 border-primary pl-4"><a href="#design-the-pause-like-a-feature" class="heading-permalink text-inherit no-underline">Design the Pause Like a Feature</a></h2>
<p class="text-muted-foreground leading-7 my-4">The good news is that a pause isn&#x27;t a design failure to apologize for. It&#x27;s a feature to build well. The teams doing this treat the hand-off as a first-class part of the system, not an error path.</p>
<p class="text-muted-foreground leading-7 my-4">A few principles that hold up:</p>
<p class="text-muted-foreground leading-7 my-4"><strong class="text-foreground font-bold">Gate on categories, not vibes.</strong> Pick the handful of action types that are genuinely irreversible or high-blast-radius and require a human every time. A widely used default: pause for production deploys, external communications, financial transactions above a set threshold (often as low as $100), data deletion, and privilege changes. Everything else runs.</p>
<p class="text-muted-foreground leading-7 my-4"><strong class="text-foreground font-bold">Make the pause cheap to answer.</strong> An approval request that dumps raw JSON on a reviewer is a bottleneck, not oversight. Give the human a plain-language summary: what the agent wants to do, what changes, whether it&#x27;s reversible, and the estimated impact. The goal is a confident decision in seconds, not an archaeology project.</p>
<p class="text-muted-foreground leading-7 my-4"><strong class="text-foreground font-bold">Freeze the state, don&#x27;t restart it.</strong> When the agent pauses, it should serialize its state and resume cleanly on approval, not re-run from the top. Hash the proposed action at the moment of the pause and re-check it before executing, so nothing drifts during the approval window.</p>
<p class="text-muted-foreground leading-7 my-4"><strong class="text-foreground font-bold">Set a timeout, and make it a kill switch.</strong> An approval request that hangs forever is its own failure. Give it a window. Thirty minutes is a common default. After that the action is cancelled, not silently executed. The safe default when a human never answers is <em class="font-mono font-normal text-primary/80 print:text-primary">stop</em>, not <em class="font-mono font-normal text-primary/80 print:text-primary">proceed</em>.</p>
<h2 id="asking-for-help-is-the-advanced-move" class="font-display text-2xl lg:text-3xl font-bold text-foreground mt-12 mb-6 border-l-2 border-primary pl-4"><a href="#asking-for-help-is-the-advanced-move" class="heading-permalink text-inherit no-underline">Asking for Help Is the Advanced Move</a></h2>
<p class="text-muted-foreground leading-7 my-4">There is a version of this that reads as a limitation: the agent isn&#x27;t good enough to be trusted, so we bolt a human onto the risky parts. That framing is backwards.</p>
<p class="text-muted-foreground leading-7 my-4">An agent that barrels through every decision isn&#x27;t more advanced. It&#x27;s less aware. The genuinely capable behavior is the thing we struggle to get juniors to do: recognize the edge of your own competence and escalate <em class="font-mono font-normal text-primary/80 print:text-primary">before</em> you cross it, not after. An agent that stops at the right moment and says &quot;this one is above my pay grade, confirm before I proceed&quot; is showing more judgment, not less.</p>
<p class="text-muted-foreground leading-7 my-4">The counterargument is worth taking seriously. In April 2026, MIT Technology Review argued that human-in-the-loop oversight has quietly become an illusion in some systems, because a reviewer can&#x27;t actually verify what a model reasoned about internally before it acted. That&#x27;s a real risk, and it&#x27;s a warning about <em class="font-mono font-normal text-primary/80 print:text-primary">lazy</em> oversight: the rubber-stamp, the ten-thousand-alert queue nobody reads. It isn&#x27;t an argument against oversight. It&#x27;s an argument for the kind you can actually perform: few gates, high stakes, clear summaries, real decisions.</p>
<h2 id="the-bottom-line" class="font-display text-2xl lg:text-3xl font-bold text-foreground mt-12 mb-6 border-l-2 border-primary pl-4"><a href="#the-bottom-line" class="heading-permalink text-inherit no-underline">The Bottom Line</a></h2>
<p class="text-muted-foreground leading-7 my-4">Stop optimizing for how much your agent can do without you. Start optimizing for whether it stops in the right places.</p>
<p class="text-muted-foreground leading-7 my-4">The reliability arc bends the same way every time. You limit what the agent <em class="font-mono font-normal text-primary/80 print:text-primary">can</em> do. You give it repeatable procedures. You measure whether it actually worked. You make its mistakes reversible. And then you draw the last line: the short list of actions where no amount of agent confidence is enough, and a human has to say yes.</p>
<p class="text-muted-foreground leading-7 my-4">Full autonomy makes a better demo. Knowing when to ask for help makes an agent you can leave running.</p>
<p class="text-muted-foreground leading-7 my-4">Build the pause. Make it fast. And put it exactly where being wrong is expensive.</p>]]></content:encoded>
      <pubDate>Tue, 14 Jul 2026 12:00:00 GMT</pubDate>
    </item>
    <item>
      <title>Give Your Agent an Undo Button</title>
      <link>https://pratik.pa.tel/blog/give-your-agent-an-undo-button/</link>
      <guid isPermaLink="true">https://pratik.pa.tel/blog/give-your-agent-an-undo-button/</guid>
      <description>You can bound what an agent touches and measure when it fails. Neither one saves you if the mistake it makes cannot be taken back.</description>
      <dc:creator>Pratik Patel</dc:creator>
      <category>ai</category>
      <category>agents</category>
      <category>reliability</category>
      <category>engineering</category>
      <content:encoded><![CDATA[<p class="text-muted-foreground leading-7 my-4">In April 2026, an AI coding agent <a href="https://www.euronews.com/next/2026/04/28/an-ai-agent-deleted-a-companys-entire-database-in-9-seconds-then-wrote-an-apology" class="text-primary underline underline-offset-2 decoration-primary/40 print:decoration-primary hover:decoration-primary transition-colors">deleted a company&#x27;s production database in nine seconds</a>, then wrote an apology.</p>
<p class="text-muted-foreground leading-7 my-4">The part that stuck with me was not the deletion. It was what happened next. The backups were gone too. The most recent thing left to restore from was three months old. Three months of customer data, wiped by a single confident action, with nothing behind it.</p>
<p class="text-muted-foreground leading-7 my-4">The agent did not get hacked. It used <a href="https://www.eon.io/blog/ai-agent-data-loss" class="text-primary underline underline-offset-2 decoration-primary/40 print:decoration-primary hover:decoration-primary transition-colors">valid credentials, a valid endpoint, and a permitted operation</a>. Every individual step was allowed. The problem was that the most destructive thing it could do was also one of the easiest, and there was no way to take it back.</p>
<p class="text-muted-foreground leading-7 my-4">That is the failure mode I keep seeing under all the others. Not &quot;the agent did something it was not allowed to do.&quot; Something quieter. &quot;The agent did something allowed, and it could not be undone.&quot;</p>
<h2 id="prevention-and-detection-are-not-enough" class="font-display text-2xl lg:text-3xl font-bold text-foreground mt-12 mb-6 border-l-2 border-primary pl-4"><a href="#prevention-and-detection-are-not-enough" class="heading-permalink text-inherit no-underline">Prevention and Detection Are Not Enough</a></h2>
<p class="text-muted-foreground leading-7 my-4">I have written before about the two obvious levers. Decide what the agent is allowed to touch. Measure whether its work is actually correct. Both matter. Both are necessary.</p>
<p class="text-muted-foreground leading-7 my-4">Neither one helps you at 2am when a permitted action has already landed and it was wrong.</p>
<p class="text-muted-foreground leading-7 my-4">Prevention narrows the set of things that can go wrong. Detection tells you when one of them did. But there is always a gap between the moment a bad action executes and the moment you notice. In that gap, the only thing that saves you is whether the action can be reversed.</p>
<p class="text-muted-foreground leading-7 my-4">Reliability people have a name for the worst version of this: the blast radius. How much damage can one wrong move cause before anything stops it? For traditional software, a bad deploy has a blast radius you can reason about. For an autonomous agent making its own decisions about what to do next, <a href="https://tianpan.co/blog/2026-05-05-agent-blast-radius-bounding-worst-case-impact-production" class="text-primary underline underline-offset-2 decoration-primary/40 print:decoration-primary hover:decoration-primary transition-colors">the blast radius gets wider and the rollback gets harder</a>, because the agent can chain several allowed actions into one unplanned outcome faster than a human can react.</p>
<p class="text-muted-foreground leading-7 my-4">So the third lever is recovery. Design the system so that when the agent is wrong, and it will be, the mistake is cheap to take back.</p>
<h2 id="reversibility-is-a-property-you-design" class="font-display text-2xl lg:text-3xl font-bold text-foreground mt-12 mb-6 border-l-2 border-primary pl-4"><a href="#reversibility-is-a-property-you-design" class="heading-permalink text-inherit no-underline">Reversibility Is a Property You Design</a></h2>
<p class="text-muted-foreground leading-7 my-4">The instinct is to treat reversibility as a backup problem. Take snapshots, keep copies, restore if something breaks. That is the exact assumption that failed in the nine-second story. The backups shared the same credentials and the same control plane as the thing they were protecting, so the same action that killed production killed the safety net.</p>
<p class="text-muted-foreground leading-7 my-4">Reversibility is not a copy you keep. It is a property of each action the agent can take.</p>
<p class="text-muted-foreground leading-7 my-4">The most useful thing I have started doing is boring: classify every tool the agent has by how hard it is to undo.</p>
<p class="text-muted-foreground leading-7 my-4">Green actions are reversible and cheap. Reading data, running a query, drafting a message, opening a pull request. If the agent gets these wrong, you close the tab and move on. Let them run.</p>
<p class="text-muted-foreground leading-7 my-4">Yellow actions are reversible but not free. Writing to a database with a soft delete, sending an internal notification, changing a config that has a known rollback. These are fine to automate if you have actually tested the path back, not just assumed it exists.</p>
<p class="text-muted-foreground leading-7 my-4">Red actions are irreversible or externally visible. Hard-deleting data, moving money, sending an email to a customer, deleting the backups. These are the ones that should never be one confident sentence away from execution.</p>
<p class="text-muted-foreground leading-7 my-4">Once you sort tools this way, a lot of design decisions make themselves. The stronger the approval, the smaller the blast radius, and the clearer the evidence you require, the more irreversible the action is. That single rule of thumb would have stopped the nine-second incident cold.</p>
<h2 id="the-undo-button-in-practice" class="font-display text-2xl lg:text-3xl font-bold text-foreground mt-12 mb-6 border-l-2 border-primary pl-4"><a href="#the-undo-button-in-practice" class="heading-permalink text-inherit no-underline">The Undo Button, In Practice</a></h2>
<p class="text-muted-foreground leading-7 my-4">Giving an agent an undo button is less exotic than it sounds. It is mostly patterns engineers already know, applied one layer earlier.</p>
<p class="text-muted-foreground leading-7 my-4">Prefer soft deletes over hard deletes. A row marked deleted is an undo. A row that is gone is an incident. The agent should have to work much harder to do the second one.</p>
<p class="text-muted-foreground leading-7 my-4">Make destructive actions two-step. The agent proposes, something else confirms. That something can be a human for the truly irreversible stuff, or a separate check for the merely expensive stuff. The point is that no single call both decides and destroys.</p>
<p class="text-muted-foreground leading-7 my-4">Run in a copy first. Staging, a sandbox, a dry-run mode that shows the diff before it is applied. If the agent can rehearse the change against something that is not production, most bad plans reveal themselves before they touch anything real.</p>
<p class="text-muted-foreground leading-7 my-4">Keep the recovery path off the same switch. Backups that share credentials and storage with production are not backups. They are a second copy of the thing you are about to lose. Immutability and a separate control plane are the difference between a bad hour and a dead company.</p>
<p class="text-muted-foreground leading-7 my-4">Checkpoint the work. For multi-step workflows, save state between steps so you can roll back to the last good point instead of unwinding the entire run by hand. This is the recovery version of measuring the middle, and it is why teams increasingly treat <a href="https://medium.com/@raktims2210/the-enterprise-ai-control-plane-why-reversible-autonomy-is-the-missing-layer-for-scalable-ai-8dd1edef2ab5" class="text-primary underline underline-offset-2 decoration-primary/40 print:decoration-primary hover:decoration-primary transition-colors">reversible autonomy as its own layer</a> of the stack rather than an afterthought.</p>
<p class="text-muted-foreground leading-7 my-4">None of this makes the agent smarter. It makes the agent&#x27;s mistakes survivable, which is a different and more useful goal.</p>
<h2 id="save-the-human-for-the-irreversible" class="font-display text-2xl lg:text-3xl font-bold text-foreground mt-12 mb-6 border-l-2 border-primary pl-4"><a href="#save-the-human-for-the-irreversible" class="heading-permalink text-inherit no-underline">Save The Human For The Irreversible</a></h2>
<p class="text-muted-foreground leading-7 my-4">There is a real cost to a confirmation step. Ask a person to approve everything and you have not built an agent, you have built a slow form. So spend the human attention where it actually pays off.</p>
<p class="text-muted-foreground leading-7 my-4">The trigger is irreversibility, not importance. A reversible action can be big and still run on its own, because if it is wrong you fix it. An irreversible action can be small and still deserve a human, because if it is wrong you cannot. <a href="https://www.microsoft.com/en-us/security/blog/2026/05/14/defense-in-depth-autonomous-ai-agents/" class="text-primary underline underline-offset-2 decoration-primary/40 print:decoration-primary hover:decoration-primary transition-colors">Microsoft&#x27;s guidance on autonomous agents</a> lands in the same place. Let the low-risk, bounded, reversible work flow, and gate the actions whose worst case you could not walk back.</p>
<p class="text-muted-foreground leading-7 my-4">Done right, the human is not in the loop for everything. They are in the loop for exactly the handful of moves that cannot be undone. That is a load a person can actually carry, and it is where their judgment is worth the interruption.</p>
<h2 id="the-bottom-line" class="font-display text-2xl lg:text-3xl font-bold text-foreground mt-12 mb-6 border-l-2 border-primary pl-4"><a href="#the-bottom-line" class="heading-permalink text-inherit no-underline">The Bottom Line</a></h2>
<p class="text-muted-foreground leading-7 my-4">We spend most of our agent-building energy trying to make the agent right. That is the wrong thing to optimize alone, because it will still be wrong sometimes, and the interesting question is what happens when it is.</p>
<p class="text-muted-foreground leading-7 my-4">An agent whose every action is reversible can be wrong all day and cost you a few minutes. An agent with one irreversible action wired to a confident decision can be right for months and then end you in nine seconds.</p>
<p class="text-muted-foreground leading-7 my-4">So before you widen what your agent can do, ask the unglamorous question. If this goes wrong, can I take it back? If the answer is no, that action does not need a better prompt.</p>
<p class="text-muted-foreground leading-7 my-4">It needs an undo button.</p>]]></content:encoded>
      <pubDate>Tue, 07 Jul 2026 12:00:00 GMT</pubDate>
    </item>
    <item>
      <title>Agents Fail Quietly</title>
      <link>https://pratik.pa.tel/blog/agents-fail-quietly/</link>
      <guid isPermaLink="true">https://pratik.pa.tel/blog/agents-fail-quietly/</guid>
      <description>A passing test and a working workflow are not the same thing. The gap between them is where production breaks.</description>
      <dc:creator>Pratik Patel</dc:creator>
      <category>ai</category>
      <category>agents</category>
      <category>evals</category>
      <category>reliability</category>
      <content:encoded><![CDATA[<p class="text-muted-foreground leading-7 my-4">Most AI agents look flawless right up until they don&#x27;t.</p>
<p class="text-muted-foreground leading-7 my-4">The demo runs clean. The test suite is green. The final answer reads well. And then the same agent, pointed at real work, quietly produces something wrong.</p>
<p class="text-muted-foreground leading-7 my-4">Not wrong in an obvious way. Wrong in a way that passes every check you thought to write.</p>
<p class="text-muted-foreground leading-7 my-4">This is the part of building with agents that nobody puts in the launch video.</p>
<h2 id="the-demo-was-the-easy-part" class="font-display text-2xl lg:text-3xl font-bold text-foreground mt-12 mb-6 border-l-2 border-primary pl-4"><a href="#the-demo-was-the-easy-part" class="heading-permalink text-inherit no-underline">The Demo Was the Easy Part</a></h2>
<p class="text-muted-foreground leading-7 my-4">Every agent demo is built on the same foundation: clean inputs, a cooperative user, a defined scenario, and an environment where the agent&#x27;s strengths are on display and its failure modes are conveniently out of frame.</p>
<p class="text-muted-foreground leading-7 my-4">Production is the opposite. Inputs are messy. Users ask for things the agent was never shaped to do. The scenario drifts halfway through. And the failure modes you kept off-screen in the demo are now the main event.</p>
<p class="text-muted-foreground leading-7 my-4">The numbers make this concrete. A <a href="https://www.inovabeing.com/blog/ai-agent-reliability-production-failure-2026" class="text-primary underline underline-offset-2 decoration-primary/40 print:decoration-primary hover:decoration-primary transition-colors">March 2026 reliability report</a> analyzed roughly 4.5 million tests across more than 6,000 production agents across ten regions. The aggregate end-to-end success rate was 56.6 percent.</p>
<p class="text-muted-foreground leading-7 my-4">These were not toy projects. They were real systems handling customer service, document processing, internal tooling, and workflow automation. Just over half of the runs actually worked.</p>
<p class="text-muted-foreground leading-7 my-4">That gap, between the demo and the deployment, is not a model problem you can prompt your way out of.</p>
<p class="text-muted-foreground leading-7 my-4">It is a measurement problem.</p>
<h2 id="a-green-test-is-not-a-working-agent" class="font-display text-2xl lg:text-3xl font-bold text-foreground mt-12 mb-6 border-l-2 border-primary pl-4"><a href="#a-green-test-is-not-a-working-agent" class="heading-permalink text-inherit no-underline">A Green Test Is Not a Working Agent</a></h2>
<p class="text-muted-foreground leading-7 my-4">Traditional software has a comforting property: if the inputs are the same, the outputs are the same. A passing test means the code did what you asked.</p>
<p class="text-muted-foreground leading-7 my-4">Agents break that assumption.</p>
<p class="text-muted-foreground leading-7 my-4">An agent can pass a unit test and still fail the workflow. It can select the wrong tool. It can pass a malformed argument. It can hand off to the wrong sub-agent. It can take an unsafe path and then arrive at a plausible final answer that looks completely fine.</p>
<p class="text-muted-foreground leading-7 my-4">Here is the trap. Most checks look at the final state. The agent&#x27;s mistakes happen in the middle.</p>
<p class="text-muted-foreground leading-7 my-4">Picture a research agent. It correctly retrieves competitor information in step one. It misattributes a feature to the wrong company in step three. It builds the rest of its analysis on that mistake. The final summary is well written, confident, and wrong, and it passes a surface-level check because the format is right and the conclusion sounds reasonable.</p>
<p class="text-muted-foreground leading-7 my-4">The error did not show up at the end.</p>
<p class="text-muted-foreground leading-7 my-4">It propagated from the middle and never announced itself.</p>
<p class="text-muted-foreground leading-7 my-4">This is what I mean by failing quietly. The agent does not crash. It does not throw an error. It hands you a clean, finished, incorrect result.</p>
<h2 id="the-math-is-against-long-workflows" class="font-display text-2xl lg:text-3xl font-bold text-foreground mt-12 mb-6 border-l-2 border-primary pl-4"><a href="#the-math-is-against-long-workflows" class="heading-permalink text-inherit no-underline">The Math Is Against Long Workflows</a></h2>
<p class="text-muted-foreground leading-7 my-4">The reason this gets worse with ambition is simple arithmetic.</p>
<p class="text-muted-foreground leading-7 my-4">Say your agent is 85 percent reliable at each individual step. That sounds strong. Most people would ship it.</p>
<p class="text-muted-foreground leading-7 my-4">Now chain ten steps together. The end-to-end success rate is not 85 percent. It is 0.85 to the tenth power, which is roughly 20 percent.</p>
<p class="text-muted-foreground leading-7 my-4">Eighty-five percent per step. Twenty percent overall.</p>
<p class="text-muted-foreground leading-7 my-4">And real production workflows are often longer than ten steps. Every tool call, every handoff, every retrieval is another place for a small error to slip in and compound. Strong step-level performance can still produce cascading, end-to-end failure.</p>
<p class="text-muted-foreground leading-7 my-4">This is why &quot;the model got better&quot; rarely fixes a flaky agent. A better model raises per-step reliability a few points. The compounding math eats most of the gain. The fix is fewer steps, tighter scope, and a way to catch errors the moment they happen rather than at the finish line.</p>
<p class="text-muted-foreground leading-7 my-4">You cannot do any of that without measuring the middle.</p>
<h2 id="benchmarks-will-lie-to-you-too" class="font-display text-2xl lg:text-3xl font-bold text-foreground mt-12 mb-6 border-l-2 border-primary pl-4"><a href="#benchmarks-will-lie-to-you-too" class="heading-permalink text-inherit no-underline">Benchmarks Will Lie to You Too</a></h2>
<p class="text-muted-foreground leading-7 my-4">The natural instinct is to reach for a benchmark. Run the agent against a standard suite, get a score, ship if the score is high.</p>
<p class="text-muted-foreground leading-7 my-4">Be careful there as well.</p>
<p class="text-muted-foreground leading-7 my-4">In April 2026, <a href="https://insights.reinventing.ai/articles/ai-agents-evaluation-production-reliability-2026-04-27" class="text-primary underline underline-offset-2 decoration-primary/40 print:decoration-primary hover:decoration-primary transition-colors">UC Berkeley researchers showed</a> that every major AI agent benchmark could be exploited to reach near-perfect scores without actually solving a single task. The benchmark measured something. It just was not the thing you cared about.</p>
<p class="text-muted-foreground leading-7 my-4">A high benchmark score and a reliable production agent are not the same claim. One is a number on a leaderboard. The other is a system that behaves on the inputs your users actually send.</p>
<p class="text-muted-foreground leading-7 my-4">The benchmark is a starting point, not a verdict.</p>
<h2 id="build-evals-like-you-mean-it" class="font-display text-2xl lg:text-3xl font-bold text-foreground mt-12 mb-6 border-l-2 border-primary pl-4"><a href="#build-evals-like-you-mean-it" class="heading-permalink text-inherit no-underline">Build Evals Like You Mean It</a></h2>
<p class="text-muted-foreground leading-7 my-4">So what actually works? Less hope, more measurement.</p>
<p class="text-muted-foreground leading-7 my-4">The teams shipping reliable agents in 2026 treat evaluation as core infrastructure, not an afterthought. A few principles I keep coming back to.</p>
<p class="text-muted-foreground leading-7 my-4">Measure steps, not just outcomes. Check the tool calls, the arguments, and the handoffs, not only the final answer. A workflow that gets the right answer for the wrong reason will eventually get the wrong answer.</p>
<p class="text-muted-foreground leading-7 my-4">Build evals from real failures. Every production incident is a test case you did not have yet. When an agent fails quietly, capture the trace and turn it into a permanent check. Your eval set should grow every week.</p>
<p class="text-muted-foreground leading-7 my-4">Run evals in CI. Treat an agent regression like a code regression. The point of <a href="https://www.confident-ai.com/knowledge-base/compare/best-ci-cd-tools-testing-ai-agents-before-production-2026" class="text-primary underline underline-offset-2 decoration-primary/40 print:decoration-primary hover:decoration-primary transition-colors">CI tooling for agents</a> is to catch the drop before users do, not to explain it afterward.</p>
<p class="text-muted-foreground leading-7 my-4">Score reliability, not vibes. &quot;It feels smarter&quot; is not a metric. End-to-end success rate, step-level accuracy, and tool-selection precision are. If you cannot put a number on it, you cannot tell whether your last change helped or hurt.</p>
<p class="text-muted-foreground leading-7 my-4">None of this is glamorous. It is the agent equivalent of writing tests and reading logs. But it is the difference between an agent that demos well and an agent you can actually leave running.</p>
<h2 id="quiet-failure-is-a-trust-problem" class="font-display text-2xl lg:text-3xl font-bold text-foreground mt-12 mb-6 border-l-2 border-primary pl-4"><a href="#quiet-failure-is-a-trust-problem" class="heading-permalink text-inherit no-underline">Quiet Failure Is a Trust Problem</a></h2>
<p class="text-muted-foreground leading-7 my-4">The deeper issue is what quiet failure does to trust.</p>
<p class="text-muted-foreground leading-7 my-4">A loud failure is almost a gift. The agent crashes, you see the error, you fix it, you move on. You know exactly where you stand.</p>
<p class="text-muted-foreground leading-7 my-4">A quiet failure erodes something harder to rebuild. The agent hands you a confident, finished, wrong result, and you only find out later, after you have acted on it. Do that a few times and the user stops trusting any output, even the correct ones. They start re-checking everything by hand, which defeats the entire purpose of the agent.</p>
<p class="text-muted-foreground leading-7 my-4">An agent you have to fully re-verify is not saving you work. It is adding a step.</p>
<p class="text-muted-foreground leading-7 my-4">The whole value of an autonomous workflow is that you can trust the parts you did not watch. That trust is not earned by a good demo. It is earned by a system that catches its own mistakes and tells you when it is unsure.</p>
<h2 id="the-bottom-line" class="font-display text-2xl lg:text-3xl font-bold text-foreground mt-12 mb-6 border-l-2 border-primary pl-4"><a href="#the-bottom-line" class="heading-permalink text-inherit no-underline">The Bottom Line</a></h2>
<p class="text-muted-foreground leading-7 my-4">The flashy part of agent building is the capability. Watch it browse, write code, call APIs, chain tools, finish the job.</p>
<p class="text-muted-foreground leading-7 my-4">The part that decides whether it survives contact with real work is quieter. It is the measurement underneath: the evals, the traces, the step-level checks, the honest reliability numbers.</p>
<p class="text-muted-foreground leading-7 my-4">Agents do not usually fail loudly. They fail in the middle, with a clean face, in a way that passes the checks you remembered to write.</p>
<p class="text-muted-foreground leading-7 my-4">So write the checks you would rather not think about. Measure the steps, not just the endings. Treat every quiet failure as the next test case.</p>
<p class="text-muted-foreground leading-7 my-4">The agents that win will not be the ones that demo best.</p>
<p class="text-muted-foreground leading-7 my-4">They will be the ones you can prove are right.</p>]]></content:encoded>
      <pubDate>Tue, 30 Jun 2026 12:00:00 GMT</pubDate>
    </item>
    <item>
      <title>Agent Permissions Are Product Design</title>
      <link>https://pratik.pa.tel/blog/agent-permissions-are-product-design/</link>
      <guid isPermaLink="true">https://pratik.pa.tel/blog/agent-permissions-are-product-design/</guid>
      <description>The next reliable AI workflow starts with deciding what the agent is allowed to touch.</description>
      <dc:creator>Pratik Patel</dc:creator>
      <category>ai</category>
      <category>agents</category>
      <category>security</category>
      <category>product</category>
      <content:encoded><![CDATA[<p class="text-muted-foreground leading-7 my-4">The most important setting in an AI agent is not the model.</p>
<p class="text-muted-foreground leading-7 my-4">It is the boundary.</p>
<p class="text-muted-foreground leading-7 my-4">What can it read? What can it change? Which tools load by default? Which actions need approval? Which systems are simply not in scope?</p>
<p class="text-muted-foreground leading-7 my-4">That used to sound like security plumbing. Necessary, but boring. Something you handled after the product worked.</p>
<p class="text-muted-foreground leading-7 my-4">I think that is backwards now.</p>
<p class="text-muted-foreground leading-7 my-4">Agent permissions are becoming product design.</p>
<p class="text-muted-foreground leading-7 my-4">When a tool can write code, browse the web, operate a browser, run a terminal, call APIs, remember context, trigger workflows, and delegate work to other agents, the permission surface is no longer a back-office detail. It is the shape of the experience. It determines how much the user can trust the agent, how much the agent can do without asking, and how easy it is to understand what happened after the work is done.</p>
<p class="text-muted-foreground leading-7 my-4">This is why I found the recent <a href="https://hermes-agent.nousresearch.com/docs/getting-started/quickstart" class="text-primary underline underline-offset-2 decoration-primary/40 print:decoration-primary hover:decoration-primary transition-colors">Hermes Agent Blank Slate setup</a> interesting. The feature starts from a minimal agent, then asks you to opt into capabilities instead of quietly loading everything. No web, browser, code execution, vision, memory, delegation, cron, skills, plugins, or MCP servers unless you choose them. The details are specific to Hermes, but the pattern is bigger than one tool.</p>
<p class="text-muted-foreground leading-7 my-4">The default future should be opt-in capability.</p>
<p class="text-muted-foreground leading-7 my-4">Not because agents are useless without access.</p>
<p class="text-muted-foreground leading-7 my-4">Because access is the product.</p>
<h2 id="every-tool-changes-the-job" class="font-display text-2xl lg:text-3xl font-bold text-foreground mt-12 mb-6 border-l-2 border-primary pl-4"><a href="#every-tool-changes-the-job" class="heading-permalink text-inherit no-underline">Every Tool Changes the Job</a></h2>
<p class="text-muted-foreground leading-7 my-4">It is tempting to describe an agent by its intelligence.</p>
<p class="text-muted-foreground leading-7 my-4">This model is better at reasoning. That one is better at code. This setup has a longer context window. That one has lower latency.</p>
<p class="text-muted-foreground leading-7 my-4">All of that matters, but it misses the operational truth. An agent with no tools is mostly a thinking partner. An agent with file access is a reviewer. An agent with terminal access is a builder. An agent with browser control is a tester. An agent with production API credentials is an operator. An agent with scheduling, memory, and delegation is a process.</p>
<p class="text-muted-foreground leading-7 my-4">Each new permission changes the job you are assigning.</p>
<p class="text-muted-foreground leading-7 my-4">That means each permission should change the product surface too.</p>
<p class="text-muted-foreground leading-7 my-4">If the agent can only read files, the interface can be lightweight. If it can delete data, the interface needs stronger consent. If it can call payment APIs, the workflow needs clear policy and audit trails. If it can install MCP servers, the product needs to explain what those servers expose and what downstream systems they can reach.</p>
<p class="text-muted-foreground leading-7 my-4">The user should not have to reverse engineer the blast radius.</p>
<p class="text-muted-foreground leading-7 my-4">Good products make the boundary visible.</p>
<h2 id="defaults-teach-users-what-is-safe" class="font-display text-2xl lg:text-3xl font-bold text-foreground mt-12 mb-6 border-l-2 border-primary pl-4"><a href="#defaults-teach-users-what-is-safe" class="heading-permalink text-inherit no-underline">Defaults Teach Users What Is Safe</a></h2>
<p class="text-muted-foreground leading-7 my-4">Defaults are not neutral.</p>
<p class="text-muted-foreground leading-7 my-4">When an agent ships with every tool enabled, it teaches the user that broad access is normal. When it starts locked down and asks for capabilities as the task needs them, it teaches a different habit: grant the smallest useful permission, then expand only when the work proves it needs more.</p>
<p class="text-muted-foreground leading-7 my-4">That is the AI version of progressive disclosure.</p>
<p class="text-muted-foreground leading-7 my-4">The product does not need to dump a security lecture on the user. It just needs to make the next safe step obvious.</p>
<p class="text-muted-foreground leading-7 my-4">Want the agent to summarize a repo? Read access is enough.</p>
<p class="text-muted-foreground leading-7 my-4">Want it to fix a bug? It needs write access and a test command.</p>
<p class="text-muted-foreground leading-7 my-4">Want it to verify a UI? It needs a browser and maybe a local server.</p>
<p class="text-muted-foreground leading-7 my-4">Want it to publish? That is a different level of trust.</p>
<p class="text-muted-foreground leading-7 my-4">Each step should feel like a deliberate escalation, not a hidden side effect of installing the tool.</p>
<p class="text-muted-foreground leading-7 my-4">This is also where the &quot;power user&quot; argument gets weak. Power users do want speed, but they also want predictability. They do not want to inspect a settings file after every update to see whether new tools appeared. They do not want an agent to gain browser, memory, or plugin access because a vendor decided the new default would be more impressive in a demo.</p>
<p class="text-muted-foreground leading-7 my-4">Fast is good.</p>
<p class="text-muted-foreground leading-7 my-4">Predictable is better.</p>
<h2 id="mcp-makes-this-more-urgent" class="font-display text-2xl lg:text-3xl font-bold text-foreground mt-12 mb-6 border-l-2 border-primary pl-4"><a href="#mcp-makes-this-more-urgent" class="heading-permalink text-inherit no-underline">MCP Makes This More Urgent</a></h2>
<p class="text-muted-foreground leading-7 my-4">The Model Context Protocol made agent tooling feel composable. That is useful. It also means capability can spread quickly.</p>
<p class="text-muted-foreground leading-7 my-4">An MCP server can expose data, tools, prompts, and workflows through a standard interface. That makes it easier to connect agents to real systems. It also creates a new class of product question: when the agent connects to a server, what exactly did it gain?</p>
<p class="text-muted-foreground leading-7 my-4">The official <a href="https://modelcontextprotocol.io/docs/tutorials/security/security_best_practices" class="text-primary underline underline-offset-2 decoration-primary/40 print:decoration-primary hover:decoration-primary transition-colors">MCP security guidance</a> focuses on things like consent, token handling, confused deputy risks, session security, and scope minimization. The <a href="https://modelcontextprotocol.io/specification/draft/basic/authorization" class="text-primary underline underline-offset-2 decoration-primary/40 print:decoration-primary hover:decoration-primary transition-colors">authorization draft</a> also pushes clients toward least-privilege scope selection and step-up authorization when more access is needed.</p>
<p class="text-muted-foreground leading-7 my-4">That language sounds technical because the implementation is technical.</p>
<p class="text-muted-foreground leading-7 my-4">But the user experience problem is plain.</p>
<p class="text-muted-foreground leading-7 my-4">Do not make users approve a black box.</p>
<p class="text-muted-foreground leading-7 my-4">If an agent is connecting to a GitHub MCP server, the user should know whether the agent can read public repos, read private repos, open issues, create branches, push code, or trigger workflow runs. Those are different permissions. They deserve different consent, different defaults, and different review paths.</p>
<p class="text-muted-foreground leading-7 my-4">The <a href="https://www.nsa.gov/Portals/75/documents/Cybersecurity/CSI_MCP_SECURITY.pdf?ver=bmgiSbNQLP6Z_GiWtRt6bg%3D%3D" class="text-primary underline underline-offset-2 decoration-primary/40 print:decoration-primary hover:decoration-primary transition-colors">NSA&#x27;s May 2026 MCP security guidance</a> makes the broader point: as MCP adoption spreads into production workflows, the security model needs implementation rigor and clearer boundaries. That is not just a warning for security teams. It is a design brief for everyone building agent products.</p>
<p class="text-muted-foreground leading-7 my-4">The product has to translate capability into judgment.</p>
<h2 id="the-best-permission-ui-is-operational" class="font-display text-2xl lg:text-3xl font-bold text-foreground mt-12 mb-6 border-l-2 border-primary pl-4"><a href="#the-best-permission-ui-is-operational" class="heading-permalink text-inherit no-underline">The Best Permission UI Is Operational</a></h2>
<p class="text-muted-foreground leading-7 my-4">Most permission screens are bad because they are abstract.</p>
<p class="text-muted-foreground leading-7 my-4">They ask for access to &quot;files&quot; or &quot;tools&quot; or &quot;workspace resources&quot; without explaining what the user is actually trying to do. The result is a familiar consent problem: users click approve because the product blocks progress, not because the decision is meaningful.</p>
<p class="text-muted-foreground leading-7 my-4">Agent products can do better because the agent usually has a task.</p>
<p class="text-muted-foreground leading-7 my-4">Tie permissions to the task.</p>
<p class="text-muted-foreground leading-7 my-4">&quot;To update this blog post, I need write access to src/data/blog-posts and permission to run the focused data test.&quot;</p>
<p class="text-muted-foreground leading-7 my-4">That is useful.</p>
<p class="text-muted-foreground leading-7 my-4">&quot;To debug this failed checkout flow, I need browser control, local server access, and permission to edit files under src/pages and src/components.&quot;</p>
<p class="text-muted-foreground leading-7 my-4">That is useful too.</p>
<p class="text-muted-foreground leading-7 my-4">The permission request should name the outcome, the resources, the action level, and the proof the agent will leave behind. If the agent asks for more access later, it should explain what changed.</p>
<p class="text-muted-foreground leading-7 my-4">This turns permissioning from a modal into part of the workflow.</p>
<p class="text-muted-foreground leading-7 my-4">The boundary becomes inspectable.</p>
<h2 id="capability-should-expire" class="font-display text-2xl lg:text-3xl font-bold text-foreground mt-12 mb-6 border-l-2 border-primary pl-4"><a href="#capability-should-expire" class="heading-permalink text-inherit no-underline">Capability Should Expire</a></h2>
<p class="text-muted-foreground leading-7 my-4">One subtle mistake in agent design is treating permissions as permanent identity.</p>
<p class="text-muted-foreground leading-7 my-4">Humans log into a tool and keep their access. That works because organizations have identity systems, managers, offboarding processes, and audit policies. Even then, stale access is a constant problem.</p>
<p class="text-muted-foreground leading-7 my-4">Agents make stale access worse because their work is often task-shaped.</p>
<p class="text-muted-foreground leading-7 my-4">An agent may need write access to a repo for one bug fix. It may need a browser for one QA pass. It may need a calendar integration for one scheduling task. That does not mean it should keep those capabilities forever.</p>
<p class="text-muted-foreground leading-7 my-4">The better default is task-scoped access.</p>
<p class="text-muted-foreground leading-7 my-4">Grant the capability for this run. Record why it was granted. Revoke it when the task completes. Ask again if the next task needs it.</p>
<p class="text-muted-foreground leading-7 my-4">That may sound slower, but it creates a cleaner mental model. The user stops thinking, &quot;This agent has my environment.&quot; They start thinking, &quot;This task has these capabilities.&quot;</p>
<p class="text-muted-foreground leading-7 my-4">That is a safer product.</p>
<p class="text-muted-foreground leading-7 my-4">It is also a clearer one.</p>
<h2 id="trust-comes-from-boring-boundaries" class="font-display text-2xl lg:text-3xl font-bold text-foreground mt-12 mb-6 border-l-2 border-primary pl-4"><a href="#trust-comes-from-boring-boundaries" class="heading-permalink text-inherit no-underline">Trust Comes From Boring Boundaries</a></h2>
<p class="text-muted-foreground leading-7 my-4">The AI demos that get attention usually show an agent doing a lot.</p>
<p class="text-muted-foreground leading-7 my-4">It opens tools. It clicks around. It writes code. It sends messages. It books things. It looks alive.</p>
<p class="text-muted-foreground leading-7 my-4">The workflows that survive in real teams will be more boring.</p>
<p class="text-muted-foreground leading-7 my-4">They will show exactly what the agent can touch. They will keep destructive actions behind approval. They will separate read from write. They will make capability changes visible in review. They will log the actual tool calls. They will make it easy to replay why the agent did something.</p>
<p class="text-muted-foreground leading-7 my-4">That is not anti-autonomy.</p>
<p class="text-muted-foreground leading-7 my-4">That is how autonomy becomes usable.</p>
<p class="text-muted-foreground leading-7 my-4">If I cannot tell what an agent was allowed to do, I cannot trust what it did. If I cannot see when its permissions changed, I cannot review the work responsibly. If every task starts with broad access, I have to treat every task as high risk.</p>
<p class="text-muted-foreground leading-7 my-4">Good boundaries lower the review burden.</p>
<p class="text-muted-foreground leading-7 my-4">They make the agent&#x27;s work smaller, clearer, and easier to approve.</p>
<h2 id="builders-should-design-the-permission-ladder" class="font-display text-2xl lg:text-3xl font-bold text-foreground mt-12 mb-6 border-l-2 border-primary pl-4"><a href="#builders-should-design-the-permission-ladder" class="heading-permalink text-inherit no-underline">Builders Should Design the Permission Ladder</a></h2>
<p class="text-muted-foreground leading-7 my-4">If you are building agent workflows, I would start with a simple exercise.</p>
<p class="text-muted-foreground leading-7 my-4">Write down the permission ladder for your product.</p>
<p class="text-muted-foreground leading-7 my-4">Not the settings page. The ladder.</p>
<p class="text-muted-foreground leading-7 my-4">What can the agent do with no access? What can it do with read access? What can it do with write access? What requires browser control? What requires network access? What requires external credentials? What requires human approval every time?</p>
<p class="text-muted-foreground leading-7 my-4">Then design the product around those steps.</p>
<p class="text-muted-foreground leading-7 my-4">Make the first step useful. Make escalation contextual. Make dangerous capabilities rare and visible. Make revocation normal. Make the audit trail part of the happy path.</p>
<p class="text-muted-foreground leading-7 my-4">The goal is not to scare users away from powerful agents.</p>
<p class="text-muted-foreground leading-7 my-4">The goal is to make power legible.</p>
<p class="text-muted-foreground leading-7 my-4">The best agent products will not be the ones that ask for everything up front. They will be the ones that earn access as the work demands it.</p>
<h2 id="the-bottom-line" class="font-display text-2xl lg:text-3xl font-bold text-foreground mt-12 mb-6 border-l-2 border-primary pl-4"><a href="#the-bottom-line" class="heading-permalink text-inherit no-underline">The Bottom Line</a></h2>
<p class="text-muted-foreground leading-7 my-4">Agents are becoming capable enough that permissions can no longer be treated as a checkbox in settings.</p>
<p class="text-muted-foreground leading-7 my-4">Permissions define the work. They shape the user&#x27;s trust. They decide whether an agent feels like a helpful operator or an unpredictable process with too much reach.</p>
<p class="text-muted-foreground leading-7 my-4">The next generation of AI workflow tools will compete on models, speed, integrations, and UX polish. But underneath all of that, the real product question will be simple:</p>
<p class="text-muted-foreground leading-7 my-4">What can the agent touch, and why?</p>
<p class="text-muted-foreground leading-7 my-4">Answer that well, and the agent feels powerful.</p>
<p class="text-muted-foreground leading-7 my-4">Answer it poorly, and every new capability becomes another reason to hesitate.</p>]]></content:encoded>
      <pubDate>Tue, 23 Jun 2026 12:00:00 GMT</pubDate>
    </item>
    <item>
      <title>Agent Runbooks Beat Better Prompts</title>
      <link>https://pratik.pa.tel/blog/agent-runbooks-beat-better-prompts/</link>
      <guid isPermaLink="true">https://pratik.pa.tel/blog/agent-runbooks-beat-better-prompts/</guid>
      <description>The best AI work happens when delegation is repeatable, visible, and bounded.</description>
      <dc:creator>Pratik Patel</dc:creator>
      <category>ai</category>
      <category>agents</category>
      <category>engineering</category>
      <category>productivity</category>
      <content:encoded><![CDATA[<p class="text-muted-foreground leading-7 my-4">I started writing tiny runbooks for AI agent tasks, and the quality of the work changed almost immediately.</p>
<p class="text-muted-foreground leading-7 my-4">Not because the model got smarter.</p>
<p class="text-muted-foreground leading-7 my-4">Because the work got less ambiguous.</p>
<p class="text-muted-foreground leading-7 my-4">Most people still treat agent delegation like prompt craft. They keep trying to find the perfect sentence, the magic wording, the clever instruction that makes the model behave. I get the instinct. When the interface is a text box, it is natural to believe the answer is a better text box input.</p>
<p class="text-muted-foreground leading-7 my-4">But that is not how real delegated work gets better.</p>
<p class="text-muted-foreground leading-7 my-4">If a human teammate kept making inconsistent decisions, you would not solve it by giving them a prettier paragraph every morning. You would give them context. You would show them the expected path. You would name the edge cases. You would define when to stop and ask. You would make the work inspectable.</p>
<p class="text-muted-foreground leading-7 my-4">That is a runbook.</p>
<p class="text-muted-foreground leading-7 my-4">And for agent workflows, runbooks are starting to matter more than prompts.</p>
<h2 id="prompts-are-not-enough" class="font-display text-2xl lg:text-3xl font-bold text-foreground mt-12 mb-6 border-l-2 border-primary pl-4"><a href="#prompts-are-not-enough" class="heading-permalink text-inherit no-underline">Prompts Are Not Enough</a></h2>
<p class="text-muted-foreground leading-7 my-4">A prompt describes what you want right now.</p>
<p class="text-muted-foreground leading-7 my-4">A runbook describes how the work should be done every time.</p>
<p class="text-muted-foreground leading-7 my-4">That distinction matters because the biggest agent failures I see are not caused by a lack of raw intelligence. They are caused by missing operating context.</p>
<p class="text-muted-foreground leading-7 my-4">The agent changes the right file but verifies the wrong behavior. It fixes the visible bug but misses the product constraint. It keeps digging after the task is already complete. It treats a flaky test as a code problem. It stops at a plan when the task clearly needed implementation. It implements the request but forgets to leave a useful handoff.</p>
<p class="text-muted-foreground leading-7 my-4">These are not prompt wording problems.</p>
<p class="text-muted-foreground leading-7 my-4">They are workflow design problems.</p>
<p class="text-muted-foreground leading-7 my-4">The model needs to know more than the goal. It needs to know the local rules of the system it is operating inside. Which commands prove success. Which files are dangerous. Which tests are worth running. Which changes should stay out of scope. Which blocker is real enough to stop work.</p>
<p class="text-muted-foreground leading-7 my-4">That information does not belong in a one-off prompt.</p>
<p class="text-muted-foreground leading-7 my-4">It belongs in a reusable operating guide.</p>
<h2 id="the-runbook-is-the-interface" class="font-display text-2xl lg:text-3xl font-bold text-foreground mt-12 mb-6 border-l-2 border-primary pl-4"><a href="#the-runbook-is-the-interface" class="heading-permalink text-inherit no-underline">The Runbook Is the Interface</a></h2>
<p class="text-muted-foreground leading-7 my-4">The more agents operate software environments, the more the runbook becomes the actual interface between human intent and machine work.</p>
<p class="text-muted-foreground leading-7 my-4">The codebase is not enough. The ticket is not enough. The chat history is not enough. Each one has pieces of the truth, but none of them reliably tells the agent how to move through the work.</p>
<p class="text-muted-foreground leading-7 my-4">A good runbook does.</p>
<p class="text-muted-foreground leading-7 my-4">It turns vague delegation into a bounded loop:</p>
<ul class="space-y-2 my-6 ml-4">
<li class="flex gap-3 text-muted-foreground leading-7"><span class="text-primary shrink-0 mt-1.5">▸</span><span class="min-w-0">What is the outcome?</span></li>
<li class="flex gap-3 text-muted-foreground leading-7"><span class="text-primary shrink-0 mt-1.5">▸</span><span class="min-w-0">What context should be read first?</span></li>
<li class="flex gap-3 text-muted-foreground leading-7"><span class="text-primary shrink-0 mt-1.5">▸</span><span class="min-w-0">What is explicitly out of scope?</span></li>
<li class="flex gap-3 text-muted-foreground leading-7"><span class="text-primary shrink-0 mt-1.5">▸</span><span class="min-w-0">What is the smallest useful verification?</span></li>
<li class="flex gap-3 text-muted-foreground leading-7"><span class="text-primary shrink-0 mt-1.5">▸</span><span class="min-w-0">What counts as a blocker?</span></li>
<li class="flex gap-3 text-muted-foreground leading-7"><span class="text-primary shrink-0 mt-1.5">▸</span><span class="min-w-0">What evidence should be left behind?</span></li>
<li class="flex gap-3 text-muted-foreground leading-7"><span class="text-primary shrink-0 mt-1.5">▸</span><span class="min-w-0">Who owns the next step?</span></li>
</ul>
<p class="text-muted-foreground leading-7 my-4">That sounds simple, but it changes the shape of the work.</p>
<p class="text-muted-foreground leading-7 my-4">Without a runbook, the agent has to infer the workflow from scattered clues. Sometimes it guesses well. Sometimes it confidently follows the wrong path.</p>
<p class="text-muted-foreground leading-7 my-4">With a runbook, the agent has a track to run on. It can still make mistakes, but the mistakes become easier to spot because you can compare what happened against an expected process.</p>
<p class="text-muted-foreground leading-7 my-4">That is the beginning of trust.</p>
<h2 id="my-smallest-useful-runbook" class="font-display text-2xl lg:text-3xl font-bold text-foreground mt-12 mb-6 border-l-2 border-primary pl-4"><a href="#my-smallest-useful-runbook" class="heading-permalink text-inherit no-underline">My Smallest Useful Runbook</a></h2>
<p class="text-muted-foreground leading-7 my-4">The most useful runbooks I write are not long documents. They are usually small, sharp, and boring.</p>
<p class="text-muted-foreground leading-7 my-4">For a coding task, the skeleton looks something like this:</p>
<ol class="space-y-2 my-6 ml-4">
<li class="flex gap-3 text-muted-foreground leading-7"><span class="text-primary shrink-0 tabular-nums">1.</span><span class="min-w-0">Read the issue and the latest comment first.</span></li>
<li class="flex gap-3 text-muted-foreground leading-7"><span class="text-primary shrink-0 tabular-nums">2.</span><span class="min-w-0">Inspect the existing code before proposing changes.</span></li>
<li class="flex gap-3 text-muted-foreground leading-7"><span class="text-primary shrink-0 tabular-nums">3.</span><span class="min-w-0">Keep the edit scoped to the requested behavior.</span></li>
<li class="flex gap-3 text-muted-foreground leading-7"><span class="text-primary shrink-0 tabular-nums">4.</span><span class="min-w-0">Use the repo&#x27;s existing patterns unless there is a clear reason not to.</span></li>
<li class="flex gap-3 text-muted-foreground leading-7"><span class="text-primary shrink-0 tabular-nums">5.</span><span class="min-w-0">Run the smallest verification that proves the change.</span></li>
<li class="flex gap-3 text-muted-foreground leading-7"><span class="text-primary shrink-0 tabular-nums">6.</span><span class="min-w-0">Do not revert unrelated work.</span></li>
<li class="flex gap-3 text-muted-foreground leading-7"><span class="text-primary shrink-0 tabular-nums">7.</span><span class="min-w-0">Leave a handoff with changed files, verification, and remaining risk.</span></li>
</ol>
<p class="text-muted-foreground leading-7 my-4">That is not a prompt trick. It is an operating contract.</p>
<p class="text-muted-foreground leading-7 my-4">The exact details change by project. A frontend task might require screenshots. A database migration might require rollback notes. A security fix might require a test that proves the boundary fails closed. A content task might require checking the publish date field instead of trusting the PR description.</p>
<p class="text-muted-foreground leading-7 my-4">The point is not to write one universal agent manual.</p>
<p class="text-muted-foreground leading-7 my-4">The point is to make the repeatable parts of the work explicit.</p>
<h2 id="runbooks-create-better-stops" class="font-display text-2xl lg:text-3xl font-bold text-foreground mt-12 mb-6 border-l-2 border-primary pl-4"><a href="#runbooks-create-better-stops" class="heading-permalink text-inherit no-underline">Runbooks Create Better Stops</a></h2>
<p class="text-muted-foreground leading-7 my-4">One of the underrated parts of delegation is knowing when work should stop.</p>
<p class="text-muted-foreground leading-7 my-4">Agents are good at continuing. That is useful until it is not.</p>
<p class="text-muted-foreground leading-7 my-4">An agent can keep refactoring because it sees adjacent cleanup. It can keep trying tests because the failure looks solvable. It can keep changing copy because there is always a smoother sentence. It can keep exploring because the repo has more context to read.</p>
<p class="text-muted-foreground leading-7 my-4">Humans do the same thing, but humans usually have more ambient judgment about when the extra motion is no longer worth it.</p>
<p class="text-muted-foreground leading-7 my-4">Runbooks give agents better stop conditions.</p>
<p class="text-muted-foreground leading-7 my-4">Stop when the focused test passes and the change is narrow. Stop when the blocker is outside this environment. Stop when the next action belongs to another agent. Stop when the task asks for a review and no code change is needed. Stop when the open PR already satisfies the issue and the remaining work is a reviewer decision.</p>
<p class="text-muted-foreground leading-7 my-4">This matters because productivity is not just about making agents move faster.</p>
<p class="text-muted-foreground leading-7 my-4">It is about making sure they stop in the right place.</p>
<h2 id="the-proof-matters-more-than-the-claim" class="font-display text-2xl lg:text-3xl font-bold text-foreground mt-12 mb-6 border-l-2 border-primary pl-4"><a href="#the-proof-matters-more-than-the-claim" class="heading-permalink text-inherit no-underline">The Proof Matters More Than the Claim</a></h2>
<p class="text-muted-foreground leading-7 my-4">The best runbooks also define what proof looks like.</p>
<p class="text-muted-foreground leading-7 my-4">&quot;I fixed it&quot; is not proof.</p>
<p class="text-muted-foreground leading-7 my-4">&quot;The blog post is added&quot; is not proof.</p>
<p class="text-muted-foreground leading-7 my-4">&quot;The PR is open&quot; is closer, but still incomplete if the post was not wired into the index, the date is wrong, or the route is missing from prerendering.</p>
<p class="text-muted-foreground leading-7 my-4">Useful proof is specific:</p>
<ul class="space-y-2 my-6 ml-4">
<li class="flex gap-3 text-muted-foreground leading-7"><span class="text-primary shrink-0 mt-1.5">▸</span><span class="min-w-0">The new file exists at the expected path.</span></li>
<li class="flex gap-3 text-muted-foreground leading-7"><span class="text-primary shrink-0 mt-1.5">▸</span><span class="min-w-0">The slug is exported in the post registry.</span></li>
<li class="flex gap-3 text-muted-foreground leading-7"><span class="text-primary shrink-0 mt-1.5">▸</span><span class="min-w-0">The dateISO is the next Tuesday.</span></li>
<li class="flex gap-3 text-muted-foreground leading-7"><span class="text-primary shrink-0 mt-1.5">▸</span><span class="min-w-0">The focused data test passes.</span></li>
<li class="flex gap-3 text-muted-foreground leading-7"><span class="text-primary shrink-0 mt-1.5">▸</span><span class="min-w-0">The PR URL is recorded.</span></li>
<li class="flex gap-3 text-muted-foreground leading-7"><span class="text-primary shrink-0 mt-1.5">▸</span><span class="min-w-0">The review owner has a real next action.</span></li>
</ul>
<p class="text-muted-foreground leading-7 my-4">The proof changes by task, but the principle does not.</p>
<p class="text-muted-foreground leading-7 my-4">If you want agent work to become reliable, do not just ask for the outcome. Ask for the evidence that the outcome is real.</p>
<p class="text-muted-foreground leading-7 my-4">This is where a lot of AI workflows quietly fail. The agent produces plausible completion language, and the human has to reconstruct whether anything actually worked. That burns the time the agent was supposed to save.</p>
<p class="text-muted-foreground leading-7 my-4">Runbooks make the verification trail part of the work, not an optional afterthought.</p>
<h2 id="repetition-is-where-agents-get-useful" class="font-display text-2xl lg:text-3xl font-bold text-foreground mt-12 mb-6 border-l-2 border-primary pl-4"><a href="#repetition-is-where-agents-get-useful" class="heading-permalink text-inherit no-underline">Repetition Is Where Agents Get Useful</a></h2>
<p class="text-muted-foreground leading-7 my-4">The first time you write a runbook, it can feel slower than just asking the agent to do the task.</p>
<p class="text-muted-foreground leading-7 my-4">That is true if the task will never happen again.</p>
<p class="text-muted-foreground leading-7 my-4">But most valuable work is repetitive. Not identical, but similar enough that the same operating shape appears over and over.</p>
<p class="text-muted-foreground leading-7 my-4">Write a weekly blog post. Review a PR. Fix a failing check. Generate a social calendar. Audit a permission boundary. Triage a bug report. Test a mobile flow. Prepare a release note. Verify a customer issue.</p>
<p class="text-muted-foreground leading-7 my-4">These are not one-off miracles. They are loops.</p>
<p class="text-muted-foreground leading-7 my-4">Loops deserve runbooks.</p>
<p class="text-muted-foreground leading-7 my-4">Once the loop is written down, every future task starts with better defaults. The agent reads less random context. It makes fewer avoidable decisions. It leaves a better handoff. You spend less time correcting process and more time reviewing judgment.</p>
<p class="text-muted-foreground leading-7 my-4">That is the compounding effect.</p>
<p class="text-muted-foreground leading-7 my-4">A prompt helps once.</p>
<p class="text-muted-foreground leading-7 my-4">A runbook improves the next hundred runs.</p>
<h2 id="the-human-job-moves-upstream" class="font-display text-2xl lg:text-3xl font-bold text-foreground mt-12 mb-6 border-l-2 border-primary pl-4"><a href="#the-human-job-moves-upstream" class="heading-permalink text-inherit no-underline">The Human Job Moves Upstream</a></h2>
<p class="text-muted-foreground leading-7 my-4">The objection I hear is that this sounds like process overhead.</p>
<p class="text-muted-foreground leading-7 my-4">It can be, if you turn every agent task into a ceremony.</p>
<p class="text-muted-foreground leading-7 my-4">But the good version is lightweight. A useful runbook is not bureaucracy. It is compressed judgment.</p>
<p class="text-muted-foreground leading-7 my-4">You are taking the lessons you already learned the hard way and putting them where the agent can use them. Do not touch these files. Always check this field. Use this command first. Ask for approval before this class of action. Prefer a child issue over polling. Keep the final answer short but include verification.</p>
<p class="text-muted-foreground leading-7 my-4">That is not paperwork.</p>
<p class="text-muted-foreground leading-7 my-4">That is operational taste.</p>
<p class="text-muted-foreground leading-7 my-4">And it is one of the human skills that gets more important as agents get more capable. The better the agent is at moving, the more valuable it becomes to define the lane, the guardrails, and the finish line.</p>
<p class="text-muted-foreground leading-7 my-4">The future builder is not just a prompt writer.</p>
<p class="text-muted-foreground leading-7 my-4">The future builder is a workflow designer.</p>
<h2 id="what-i-would-start-with" class="font-display text-2xl lg:text-3xl font-bold text-foreground mt-12 mb-6 border-l-2 border-primary pl-4"><a href="#what-i-would-start-with" class="heading-permalink text-inherit no-underline">What I Would Start With</a></h2>
<p class="text-muted-foreground leading-7 my-4">If you are using agents today, do not try to build a huge operating manual.</p>
<p class="text-muted-foreground leading-7 my-4">Start with the task that repeats and still annoys you.</p>
<p class="text-muted-foreground leading-7 my-4">Pick one. Write the tiny runbook. Include the context to read, the boundaries, the verification, and the stop conditions. Use it twice. After each run, add the one instruction that would have prevented the most recent mistake.</p>
<p class="text-muted-foreground leading-7 my-4">That is enough.</p>
<p class="text-muted-foreground leading-7 my-4">The document will get better because the work will teach you what it needs.</p>
<p class="text-muted-foreground leading-7 my-4">The agent will get better because the work will stop depending on hidden context in your head.</p>
<p class="text-muted-foreground leading-7 my-4">And you will get better because you will be forced to separate the outcome you want from the process that reliably produces it.</p>
<p class="text-muted-foreground leading-7 my-4">That separation is the whole game.</p>
<h2 id="the-bottom-line" class="font-display text-2xl lg:text-3xl font-bold text-foreground mt-12 mb-6 border-l-2 border-primary pl-4"><a href="#the-bottom-line" class="heading-permalink text-inherit no-underline">The Bottom Line</a></h2>
<p class="text-muted-foreground leading-7 my-4">Better prompts can improve an agent task.</p>
<p class="text-muted-foreground leading-7 my-4">Better runbooks improve the agent system.</p>
<p class="text-muted-foreground leading-7 my-4">That is the shift I care about. AI agents are not just text generators anymore. They are software operators moving through repos, terminals, browsers, queues, and review paths. If we want that work to be useful, we have to give them more than a clever request.</p>
<p class="text-muted-foreground leading-7 my-4">We have to give them a way to work.</p>
<p class="text-muted-foreground leading-7 my-4">Tiny runbooks are how that starts.</p>]]></content:encoded>
      <pubDate>Tue, 16 Jun 2026 12:00:00 GMT</pubDate>
    </item>
    <item>
      <title>AI Made Bugs Cheap to Find</title>
      <link>https://pratik.pa.tel/blog/ai-made-bugs-cheap-to-find/</link>
      <guid isPermaLink="true">https://pratik.pa.tel/blog/ai-made-bugs-cheap-to-find/</guid>
      <description>The new security bottleneck is triage, patching, and judgment.</description>
      <dc:creator>Pratik Patel</dc:creator>
      <category>ai</category>
      <category>security</category>
      <category>engineering</category>
      <category>building</category>
      <content:encoded><![CDATA[<p class="text-muted-foreground leading-7 my-4">The most important AI security story right now is not that models can find bugs.</p>
<p class="text-muted-foreground leading-7 my-4">It is that models can find more bugs than humans can responsibly process.</p>
<p class="text-muted-foreground leading-7 my-4">That is the part that changes how builders should think about software. For years, security work was constrained by discovery. Could someone find the vulnerability? Could they reproduce it? Could they build an exploit? Could a small team afford enough expert review to catch the important issues before attackers did?</p>
<p class="text-muted-foreground leading-7 my-4">Now that bottleneck is moving.</p>
<p class="text-muted-foreground leading-7 my-4">Anthropic&#x27;s recent <a href="https://www.anthropic.com/research/glasswing-initial-update" class="text-primary underline underline-offset-2 decoration-primary/40 print:decoration-primary hover:decoration-primary transition-colors">Project Glasswing update</a> is the clearest signal yet. The company says Claude Mythos Preview and its partners found more than 10,000 high- or critical-severity vulnerabilities across major software systems. In open source alone, Anthropic says it scanned more than 1,000 projects and surfaced thousands of serious findings, with human triage becoming the slow part.</p>
<p class="text-muted-foreground leading-7 my-4">You do not have to take every number at face value to see the shape of the shift.</p>
<p class="text-muted-foreground leading-7 my-4">AI is making vulnerability discovery cheaper. That sounds like good news, and it is. But it also means every software team is about to face a harder question:</p>
<p class="text-muted-foreground leading-7 my-4"><strong class="text-foreground font-bold">What happens when the scanner is faster than the organization?</strong></p>
<h2 id="the-patch-window-is-the-product-now" class="font-display text-2xl lg:text-3xl font-bold text-foreground mt-12 mb-6 border-l-2 border-primary pl-4"><a href="#the-patch-window-is-the-product-now" class="heading-permalink text-inherit no-underline">The Patch Window Is the Product Now</a></h2>
<p class="text-muted-foreground leading-7 my-4">Security used to have a familiar rhythm. A bug was found. A report was filed. A team reproduced it. Someone argued about severity. Someone wrote a patch. Users eventually upgraded.</p>
<p class="text-muted-foreground leading-7 my-4">That process was never fast enough, but it mostly matched the speed of human discovery.</p>
<p class="text-muted-foreground leading-7 my-4">AI breaks that balance.</p>
<p class="text-muted-foreground leading-7 my-4">If models can search codebases, reason about exploit paths, generate reports, and repeat that work across thousands of projects, then finding bugs stops being the scarce skill. The scarce skill becomes the system around the finding:</p>
<ul class="space-y-2 my-6 ml-4">
<li class="flex gap-3 text-muted-foreground leading-7"><span class="text-primary shrink-0 mt-1.5">▸</span><span class="min-w-0">Can you tell which reports are real?</span></li>
<li class="flex gap-3 text-muted-foreground leading-7"><span class="text-primary shrink-0 mt-1.5">▸</span><span class="min-w-0">Can you prioritize the ones that actually matter?</span></li>
<li class="flex gap-3 text-muted-foreground leading-7"><span class="text-primary shrink-0 mt-1.5">▸</span><span class="min-w-0">Can you patch without breaking production?</span></li>
<li class="flex gap-3 text-muted-foreground leading-7"><span class="text-primary shrink-0 mt-1.5">▸</span><span class="min-w-0">Can you ship fixes before attackers learn the same thing?</span></li>
<li class="flex gap-3 text-muted-foreground leading-7"><span class="text-primary shrink-0 mt-1.5">▸</span><span class="min-w-0">Can you keep maintainers from drowning in low-quality reports?</span></li>
</ul>
<p class="text-muted-foreground leading-7 my-4">That is the new security stack. Not just detection. Response capacity.</p>
<p class="text-muted-foreground leading-7 my-4">A vulnerability that sits untriaged for three weeks is not meaningfully safer because an AI found it. In some cases, it is riskier, because the same class of model may soon make the path to exploitation easier for everyone else.</p>
<p class="text-muted-foreground leading-7 my-4">The patch window is not an operational detail anymore. It is part of the product.</p>
<h2 id="ai-does-not-remove-security-work" class="font-display text-2xl lg:text-3xl font-bold text-foreground mt-12 mb-6 border-l-2 border-primary pl-4"><a href="#ai-does-not-remove-security-work" class="heading-permalink text-inherit no-underline">AI Does Not Remove Security Work</a></h2>
<p class="text-muted-foreground leading-7 my-4">There is a lazy version of the story where AI agents make security easy.</p>
<p class="text-muted-foreground leading-7 my-4">Run the model. Get the report. Apply the patch. Done.</p>
<p class="text-muted-foreground leading-7 my-4">That is not how real systems work.</p>
<p class="text-muted-foreground leading-7 my-4">Real systems are full of tradeoffs. The obvious fix can break an integration. The technically correct fix can create a migration problem. The most severe-looking vulnerability might be unreachable in production, while the boring one in a forgotten admin path is exposed to the internet.</p>
<p class="text-muted-foreground leading-7 my-4">AI can help find and explain these problems. It can write a first patch. It can generate a regression test. It can compare similar code paths and look for variants.</p>
<p class="text-muted-foreground leading-7 my-4">But someone still has to own the decision.</p>
<p class="text-muted-foreground leading-7 my-4">That is the same pattern I keep seeing across AI-assisted building. The manual work shrinks, but the judgment work expands. The operator has to decide what deserves attention, what can wait, what risk is acceptable, and what needs a deeper human review.</p>
<p class="text-muted-foreground leading-7 my-4">Security is just where this becomes impossible to ignore.</p>
<h2 id="the-dangerous-middle" class="font-display text-2xl lg:text-3xl font-bold text-foreground mt-12 mb-6 border-l-2 border-primary pl-4"><a href="#the-dangerous-middle" class="heading-permalink text-inherit no-underline">The Dangerous Middle</a></h2>
<p class="text-muted-foreground leading-7 my-4">There is a period we are entering that feels especially unstable.</p>
<p class="text-muted-foreground leading-7 my-4">Eventually, AI should make software much safer. Every serious codebase should have agents continuously searching for vulnerabilities, proposing patches, generating tests, and checking whether fixes actually landed. That world is better than the one we have now.</p>
<p class="text-muted-foreground leading-7 my-4">But the transition is messy.</p>
<p class="text-muted-foreground leading-7 my-4">The discovery side is improving faster than the response side. That creates a gap. More findings, more reports, more possible attack paths, more pressure on teams that already do not have enough security capacity.</p>
<p class="text-muted-foreground leading-7 my-4">This is especially painful for open source.</p>
<p class="text-muted-foreground leading-7 my-4">A large company can assign security engineers, rotate incident response, and fund dedicated tooling. A maintainer with a popular library might be doing all of this after work, for free, while also reviewing feature requests and answering issue comments. Dumping hundreds of AI-generated reports into that maintainer&#x27;s inbox does not automatically make the ecosystem safer.</p>
<p class="text-muted-foreground leading-7 my-4">It might make it worse unless the reports are high quality, reproducible, prioritized, and paired with patches that are easy to review.</p>
<p class="text-muted-foreground leading-7 my-4">AI security only works if it respects the human throughput on the other side.</p>
<h2 id="what-builders-should-change-now" class="font-display text-2xl lg:text-3xl font-bold text-foreground mt-12 mb-6 border-l-2 border-primary pl-4"><a href="#what-builders-should-change-now" class="heading-permalink text-inherit no-underline">What Builders Should Change Now</a></h2>
<p class="text-muted-foreground leading-7 my-4">If you are building with AI agents, this is not just a cybersecurity industry story. It changes the default operating model for anyone shipping software.</p>
<p class="text-muted-foreground leading-7 my-4">The old advice was &quot;move fast and break things.&quot; The AI-era version needs an asterisk:</p>
<p class="text-muted-foreground leading-7 my-4">Move fast, but build a system that can notice what broke.</p>
<p class="text-muted-foreground leading-7 my-4">That means security cannot be a quarterly cleanup pass. It has to live inside the same loop as product work.</p>
<h3 id="1-treat-every-ai-generated-change-as-reviewable-work" class="font-display text-xl font-bold text-foreground mt-10 mb-4"><a href="#1-treat-every-ai-generated-change-as-reviewable-work" class="heading-permalink text-inherit no-underline">1. Treat every AI-generated change as reviewable work</a></h3>
<p class="text-muted-foreground leading-7 my-4">AI code should not feel like magic output. It should feel like a pull request from a very fast junior-to-mid-level engineer who sometimes has excellent instincts and sometimes misses the reason the system is shaped the way it is.</p>
<p class="text-muted-foreground leading-7 my-4">Review it. Ask what changed. Ask what assumptions it made. Ask what surfaces it touched. If a change affects auth, payments, permissions, file handling, secrets, user data, or external integrations, slow down.</p>
<p class="text-muted-foreground leading-7 my-4">Fast does not mean casual.</p>
<h3 id="2-make-tests-prove-the-risky-behavior" class="font-display text-xl font-bold text-foreground mt-10 mb-4"><a href="#2-make-tests-prove-the-risky-behavior" class="heading-permalink text-inherit no-underline">2. Make tests prove the risky behavior</a></h3>
<p class="text-muted-foreground leading-7 my-4">AI is good at producing tests that increase coverage and bad at knowing which behavior deserves proof unless you tell it.</p>
<p class="text-muted-foreground leading-7 my-4">For security-sensitive changes, generic tests are not enough. Ask for tests that prove the boundary:</p>
<ul class="space-y-2 my-6 ml-4">
<li class="flex gap-3 text-muted-foreground leading-7"><span class="text-primary shrink-0 mt-1.5">▸</span><span class="min-w-0">A user cannot access another user&#x27;s data</span></li>
<li class="flex gap-3 text-muted-foreground leading-7"><span class="text-primary shrink-0 mt-1.5">▸</span><span class="min-w-0">A disabled feature cannot be invoked through an API path</span></li>
<li class="flex gap-3 text-muted-foreground leading-7"><span class="text-primary shrink-0 mt-1.5">▸</span><span class="min-w-0">A webhook cannot be replayed without detection</span></li>
<li class="flex gap-3 text-muted-foreground leading-7"><span class="text-primary shrink-0 mt-1.5">▸</span><span class="min-w-0">A malformed upload cannot escape its expected directory</span></li>
<li class="flex gap-3 text-muted-foreground leading-7"><span class="text-primary shrink-0 mt-1.5">▸</span><span class="min-w-0">A permission check fails closed, not open</span></li>
</ul>
<p class="text-muted-foreground leading-7 my-4">These tests do more than catch regressions. They teach the agent what matters next time.</p>
<h3 id="3-keep-a-patch-lane-open" class="font-display text-xl font-bold text-foreground mt-10 mb-4"><a href="#3-keep-a-patch-lane-open" class="heading-permalink text-inherit no-underline">3. Keep a patch lane open</a></h3>
<p class="text-muted-foreground leading-7 my-4">The teams that handle AI-era security well will not be the teams with the most findings. They will be the teams with the shortest path from confirmed issue to shipped fix.</p>
<p class="text-muted-foreground leading-7 my-4">That means knowing who can approve a security patch. Knowing which tests must run. Knowing how to ship a small hotfix without dragging in unrelated product work. Knowing how to communicate a change if users need to update.</p>
<p class="text-muted-foreground leading-7 my-4">If your process requires three meetings to patch a serious bug, AI did not solve your security problem. It just made the backlog visible.</p>
<h3 id="4-build-less-surface-area" class="font-display text-xl font-bold text-foreground mt-10 mb-4"><a href="#4-build-less-surface-area" class="heading-permalink text-inherit no-underline">4. Build less surface area</a></h3>
<p class="text-muted-foreground leading-7 my-4">The most underrated security feature is not having the feature.</p>
<p class="text-muted-foreground leading-7 my-4">Every integration, admin panel, file parser, OAuth scope, background job, and public endpoint becomes another place where a model can find something interesting. AI makes it easier to build all of that. It also makes it easier to discover what you accidentally exposed.</p>
<p class="text-muted-foreground leading-7 my-4">This is another reason taste matters. Good builders cut surface area. They do not add settings because settings are easy. They do not expose APIs because the model can scaffold them. They ask whether the capability deserves to exist.</p>
<p class="text-muted-foreground leading-7 my-4">The safest code is still the code you never had to ship.</p>
<h2 id="this-is-why-human-judgment-gets-more-valuable" class="font-display text-2xl lg:text-3xl font-bold text-foreground mt-12 mb-6 border-l-2 border-primary pl-4"><a href="#this-is-why-human-judgment-gets-more-valuable" class="heading-permalink text-inherit no-underline">This Is Why Human Judgment Gets More Valuable</a></h2>
<p class="text-muted-foreground leading-7 my-4">Every time AI makes a technical task cheaper, people assume the human role shrinks.</p>
<p class="text-muted-foreground leading-7 my-4">I think the opposite keeps happening.</p>
<p class="text-muted-foreground leading-7 my-4">When code generation gets cheaper, product judgment matters more. When content generation gets cheaper, taste and distribution matter more. When vulnerability discovery gets cheaper, triage and patch judgment matter more.</p>
<p class="text-muted-foreground leading-7 my-4">The bottleneck moves up the stack.</p>
<p class="text-muted-foreground leading-7 my-4">That is the lesson for builders. Do not measure your AI workflow by how much code it can produce or how many issues it can find. Measure it by how quickly it helps you make good decisions and ship the right fixes.</p>
<p class="text-muted-foreground leading-7 my-4">The future is not &quot;AI finds every bug, so security is solved.&quot;</p>
<p class="text-muted-foreground leading-7 my-4">The future is closer to this:</p>
<p class="text-muted-foreground leading-7 my-4">AI finds more than you can handle. The winners are the teams that built the judgment, process, and restraint to handle the right things first.</p>
<h2 id="the-bottom-line" class="font-display text-2xl lg:text-3xl font-bold text-foreground mt-12 mb-6 border-l-2 border-primary pl-4"><a href="#the-bottom-line" class="heading-permalink text-inherit no-underline">The Bottom Line</a></h2>
<p class="text-muted-foreground leading-7 my-4">Project Glasswing is a preview of a much larger shift. Software is entering an era where the cost of finding flaws drops dramatically, while the cost of responsibly fixing them remains stubbornly human.</p>
<p class="text-muted-foreground leading-7 my-4">That is uncomfortable, but it is also useful clarity.</p>
<p class="text-muted-foreground leading-7 my-4">If you are building with AI, do not wait for a security crisis to design your response loop. Review agent-written code like it matters. Test the boundaries. Keep patches small. Reduce surface area. Build a process that can absorb uncomfortable findings without freezing.</p>
<p class="text-muted-foreground leading-7 my-4">AI made bugs cheap to find.</p>
<p class="text-muted-foreground leading-7 my-4">Now the hard part is proving you can fix them.</p>]]></content:encoded>
      <pubDate>Tue, 09 Jun 2026 12:00:00 GMT</pubDate>
    </item>
    <item>
      <title>The Agent Left the IDE</title>
      <link>https://pratik.pa.tel/blog/the-agent-left-the-ide/</link>
      <guid isPermaLink="true">https://pratik.pa.tel/blog/the-agent-left-the-ide/</guid>
      <description>Computer use turns AI coding from text generation into software operation.</description>
      <dc:creator>Pratik Patel</dc:creator>
      <category>ai</category>
      <category>engineering</category>
      <category>codex</category>
      <category>productivity</category>
      <content:encoded><![CDATA[<p class="text-muted-foreground leading-7 my-4">The most interesting thing about AI coding agents right now is not that they can write code.</p>
<p class="text-muted-foreground leading-7 my-4">It is that they are starting to operate computers.</p>
<p class="text-muted-foreground leading-7 my-4">That sounds like a small distinction until you feel it in the workflow. A code generator lives inside a text box. It waits for a prompt, returns a patch, and leaves the rest of the job to you. A software operator can inspect the app, click through the broken flow, read the console, run the server, reproduce the issue, change the code, and check whether the thing actually works.</p>
<p class="text-muted-foreground leading-7 my-4">That is a different kind of tool.</p>
<p class="text-muted-foreground leading-7 my-4">OpenAI&#x27;s May 29 <a href="https://help.openai.com/en/articles/6825453-chatgpt-release-notes" class="text-primary underline underline-offset-2 decoration-primary/40 print:decoration-primary hover:decoration-primary transition-colors">Codex update</a> points in that direction. Codex now supports computer use on Windows in the Codex app for eligible users, so it can see, click, and type in Windows applications while testing and refining software. The same release also expands remote control, letting a user steer work from ChatGPT on mobile or Codex on Mac while the Windows machine remains the host for the project files, shell, app server, and local context.</p>
<p class="text-muted-foreground leading-7 my-4">I do not think the important part is Windows support by itself.</p>
<p class="text-muted-foreground leading-7 my-4">The important part is the new shape of work.</p>
<h2 id="coding-was-never-just-typing" class="font-display text-2xl lg:text-3xl font-bold text-foreground mt-12 mb-6 border-l-2 border-primary pl-4"><a href="#coding-was-never-just-typing" class="heading-permalink text-inherit no-underline">Coding Was Never Just Typing</a></h2>
<p class="text-muted-foreground leading-7 my-4">For a while, the AI coding story was mostly about generation. Could the model write a component? Could it scaffold an API route? Could it refactor a file without losing the plot?</p>
<p class="text-muted-foreground leading-7 my-4">Useful, but narrow.</p>
<p class="text-muted-foreground leading-7 my-4">Real software work has always been messier than text generation. You open the app. You notice the layout is wrong. You click a button. Nothing happens. You check the terminal. The dev server crashed. You restart it. The page loads, but the empty state is off. You resize the browser. The mobile nav breaks. You skim the network tab. The request is fine, but the UI state is stale.</p>
<p class="text-muted-foreground leading-7 my-4">None of that is &quot;write code&quot; in the pure sense.</p>
<p class="text-muted-foreground leading-7 my-4">It is operating the system around the code.</p>
<p class="text-muted-foreground leading-7 my-4">That is why computer use matters. It gives the agent access to the loop that human engineers actually live in: observe, diagnose, change, verify. The text editor is only one stop in that loop.</p>
<h2 id="the-ide-is-too-small" class="font-display text-2xl lg:text-3xl font-bold text-foreground mt-12 mb-6 border-l-2 border-primary pl-4"><a href="#the-ide-is-too-small" class="heading-permalink text-inherit no-underline">The IDE Is Too Small</a></h2>
<p class="text-muted-foreground leading-7 my-4">The IDE was a natural starting point for AI coding tools because code is text. Put the model near the text and it can help.</p>
<p class="text-muted-foreground leading-7 my-4">But the product surface of software is not the IDE. It is the browser, the terminal, the database, the logs, the design tool, the cloud dashboard, the test runner, the email preview, the mobile simulator, and sometimes a random desktop app that only exists because some enterprise workflow depends on it.</p>
<p class="text-muted-foreground leading-7 my-4">If the agent can only see the repository, it is always working from a partial truth.</p>
<p class="text-muted-foreground leading-7 my-4">It can infer what should happen. It can read tests. It can inspect types. It can even run commands if the environment allows it. But it cannot fully understand the gap between the code and the experience unless it can look at the experience.</p>
<p class="text-muted-foreground leading-7 my-4">This is why frontend work has been such a revealing test. A model can produce valid React and still ship an interface that feels wrong. It can pass tests and still overlap text on mobile. It can implement the requested behavior and miss that the loading state jumps the layout.</p>
<p class="text-muted-foreground leading-7 my-4">The browser catches what the diff cannot.</p>
<p class="text-muted-foreground leading-7 my-4">An agent that can look, click, and iterate has a better shot at closing that gap.</p>
<h2 id="remote-control-changes-the-cadence" class="font-display text-2xl lg:text-3xl font-bold text-foreground mt-12 mb-6 border-l-2 border-primary pl-4"><a href="#remote-control-changes-the-cadence" class="heading-permalink text-inherit no-underline">Remote Control Changes the Cadence</a></h2>
<p class="text-muted-foreground leading-7 my-4">The remote-control part may end up being just as important as computer use.</p>
<p class="text-muted-foreground leading-7 my-4">When an agent can keep working on the host machine while you check in from somewhere else, the job starts to feel less like a chat session and more like delegated work. You do not need to sit there watching every command. You can let the agent run until it hits a decision point, then answer the question, redirect it, or approve the next step.</p>
<p class="text-muted-foreground leading-7 my-4">That changes the cadence of engineering.</p>
<p class="text-muted-foreground leading-7 my-4">The old cadence was synchronous. You were either coding or you were not. If you stepped away, the work stopped.</p>
<p class="text-muted-foreground leading-7 my-4">The new cadence is supervisory. You define the goal, give the agent enough context, and let it move through the loop. Your job is to keep the judgment layer alive. Is this still the right approach? Did it choose the right tradeoff? Is the patch too broad? Did it verify the thing that matters?</p>
<p class="text-muted-foreground leading-7 my-4">That is closer to managing a capable junior engineer than using autocomplete.</p>
<p class="text-muted-foreground leading-7 my-4">And like managing a junior engineer, the value depends on the quality of your delegation.</p>
<h2 id="the-risk-moves-too" class="font-display text-2xl lg:text-3xl font-bold text-foreground mt-12 mb-6 border-l-2 border-primary pl-4"><a href="#the-risk-moves-too" class="heading-permalink text-inherit no-underline">The Risk Moves Too</a></h2>
<p class="text-muted-foreground leading-7 my-4">There is a tempting version of this story where more agent autonomy simply means more productivity.</p>
<p class="text-muted-foreground leading-7 my-4">That is not the full picture.</p>
<p class="text-muted-foreground leading-7 my-4">An agent with computer use has a wider action surface. It can click the wrong thing. It can misunderstand a modal. It can test against the wrong environment. It can mistake a locally cached state for a real fix. It can spend time polishing the visible symptom while missing the deeper bug.</p>
<p class="text-muted-foreground leading-7 my-4">More access is only useful when the workflow has boundaries.</p>
<p class="text-muted-foreground leading-7 my-4">That means you still need clear permissions, disposable environments, human approval for risky actions, and a review process that treats agent work like real work. Especially when the agent is touching systems outside the editor.</p>
<p class="text-muted-foreground leading-7 my-4">The mistake is assuming that because the agent can operate more of the computer, it should be allowed to operate everything.</p>
<p class="text-muted-foreground leading-7 my-4">Good delegation is scoped. Give the agent a sandbox. Give it the app server, the browser, the test suite, and enough project context to make progress. Keep production credentials, irreversible actions, billing changes, and sensitive user data behind a stronger gate.</p>
<p class="text-muted-foreground leading-7 my-4">The point is not to make the agent fearless.</p>
<p class="text-muted-foreground leading-7 my-4">The point is to make it useful without making it dangerous.</p>
<h2 id="the-best-workflows-will-be-visible" class="font-display text-2xl lg:text-3xl font-bold text-foreground mt-12 mb-6 border-l-2 border-primary pl-4"><a href="#the-best-workflows-will-be-visible" class="heading-permalink text-inherit no-underline">The Best Workflows Will Be Visible</a></h2>
<p class="text-muted-foreground leading-7 my-4">As agents become more operational, the winning workflows will be the ones that make the agent&#x27;s work easy to inspect.</p>
<p class="text-muted-foreground leading-7 my-4">I want to see what it tried. I want screenshots when the UI changes. I want terminal output when a test fails. I want a short explanation of why it chose one fix over another. I want the final diff to be boring and the verification trail to be clear.</p>
<p class="text-muted-foreground leading-7 my-4">This is the difference between autonomy and trust.</p>
<p class="text-muted-foreground leading-7 my-4">Autonomy means the agent can move. Trust means I can understand what happened after it moved.</p>
<p class="text-muted-foreground leading-7 my-4">That is also where many teams will get the first productivity gains. Not from letting agents do huge open-ended tasks, but from handing them bounded loops:</p>
<ul class="space-y-2 my-6 ml-4">
<li class="flex gap-3 text-muted-foreground leading-7"><span class="text-primary shrink-0 mt-1.5">▸</span><span class="min-w-0">Reproduce this bug and propose the smallest fix</span></li>
<li class="flex gap-3 text-muted-foreground leading-7"><span class="text-primary shrink-0 mt-1.5">▸</span><span class="min-w-0">Run the app and check the empty states</span></li>
<li class="flex gap-3 text-muted-foreground leading-7"><span class="text-primary shrink-0 mt-1.5">▸</span><span class="min-w-0">Add the test that proves this permission boundary</span></li>
<li class="flex gap-3 text-muted-foreground leading-7"><span class="text-primary shrink-0 mt-1.5">▸</span><span class="min-w-0">Verify this onboarding flow on mobile</span></li>
<li class="flex gap-3 text-muted-foreground leading-7"><span class="text-primary shrink-0 mt-1.5">▸</span><span class="min-w-0">Compare the implementation to the design and list mismatches</span></li>
</ul>
<p class="text-muted-foreground leading-7 my-4">Those are not glamorous tasks. They are exactly the tasks that slow teams down every day.</p>
<p class="text-muted-foreground leading-7 my-4">Computer use makes them more delegable.</p>
<h2 id="what-builders-should-do-now" class="font-display text-2xl lg:text-3xl font-bold text-foreground mt-12 mb-6 border-l-2 border-primary pl-4"><a href="#what-builders-should-do-now" class="heading-permalink text-inherit no-underline">What Builders Should Do Now</a></h2>
<p class="text-muted-foreground leading-7 my-4">If you are building with AI agents, I would not wait for the perfect tool before changing your habits.</p>
<p class="text-muted-foreground leading-7 my-4">Start by making your work easier for an agent to operate.</p>
<p class="text-muted-foreground leading-7 my-4">Keep local setup simple. Document the command that runs the app. Make tests deterministic. Write down the flows that matter. Keep secrets out of default environments. Add screenshots or acceptance criteria when the task is visual. Ask the agent to verify behavior, not just change files.</p>
<p class="text-muted-foreground leading-7 my-4">Most of this is just good engineering hygiene.</p>
<p class="text-muted-foreground leading-7 my-4">That is the recurring pattern with AI tools. The better your system is for humans, the better it tends to be for agents. Clear docs, clear tests, clear boundaries, clear review paths. AI does not remove the need for that discipline. It makes the payoff more obvious.</p>
<p class="text-muted-foreground leading-7 my-4">The agent leaving the IDE does not mean engineers leave the process.</p>
<p class="text-muted-foreground leading-7 my-4">It means the process needs to be legible enough that an agent can participate in it.</p>
<h2 id="the-bottom-line" class="font-display text-2xl lg:text-3xl font-bold text-foreground mt-12 mb-6 border-l-2 border-primary pl-4"><a href="#the-bottom-line" class="heading-permalink text-inherit no-underline">The Bottom Line</a></h2>
<p class="text-muted-foreground leading-7 my-4">The next phase of AI coding is not about prettier autocomplete.</p>
<p class="text-muted-foreground leading-7 my-4">It is about agents that can operate the software environment around the code. They will run apps, inspect interfaces, respond to prompts, test changes, and keep moving while humans supervise from the judgment layer.</p>
<p class="text-muted-foreground leading-7 my-4">That is a big shift.</p>
<p class="text-muted-foreground leading-7 my-4">The IDE was where AI coding started because it was the easiest surface to understand. But software does not live in the IDE. It lives in the messy loop between code, runtime, product, and user experience.</p>
<p class="text-muted-foreground leading-7 my-4">Now the agents are entering that loop.</p>
<p class="text-muted-foreground leading-7 my-4">The builders who benefit most will not be the ones who hand over everything. They will be the ones who design tight, visible, reviewable workflows where agents can do real operating work and humans still own the decisions that matter.</p>]]></content:encoded>
      <pubDate>Tue, 02 Jun 2026 12:00:00 GMT</pubDate>
    </item>
    <item>
      <title>I Haven&apos;t Touched Code in One Month</title>
      <link>https://pratik.pa.tel/blog/i-have-not-touched-code-in-one-month/</link>
      <guid isPermaLink="true">https://pratik.pa.tel/blog/i-have-not-touched-code-in-one-month/</guid>
      <description>AI did not make me obsolete. It moved my job up the stack.</description>
      <dc:creator>Pratik Patel</dc:creator>
      <category>ai</category>
      <category>engineering</category>
      <category>productivity</category>
      <category>building</category>
      <content:encoded><![CDATA[<p class="text-muted-foreground leading-7 my-4">I realized something uncomfortable this week: I have not really touched code in a full month.</p>
<p class="text-muted-foreground leading-7 my-4">Not &quot;I stopped building.&quot; The opposite. More shipped. More got reviewed. More ideas made it from vague note to working software. But the actual implementation work - the typing, wiring, scaffolding, renaming, fixing imports, chasing test failures - increasingly happened somewhere else.</p>
<p class="text-muted-foreground leading-7 my-4">AI agents handled it.</p>
<p class="text-muted-foreground leading-7 my-4">The honest reaction is not pure excitement. It is stranger than that. Part of me feels more effective than ever. Part of me wonders if I am getting worse at the craft I spent years sharpening. If you do not use the coding muscle, does it atrophy? Probably. But I also think that is the wrong question.</p>
<p class="text-muted-foreground leading-7 my-4">The better question is: <strong class="text-foreground font-bold">what skill is replacing it?</strong></p>
<h2 id="the-work-did-not-disappear" class="font-display text-2xl lg:text-3xl font-bold text-foreground mt-12 mb-6 border-l-2 border-primary pl-4"><a href="#the-work-did-not-disappear" class="heading-permalink text-inherit no-underline">The Work Did Not Disappear</a></h2>
<p class="text-muted-foreground leading-7 my-4">When people talk about AI coding, they usually frame it as automation. The machine writes code, so the human does less work. That is technically true and practically misleading.</p>
<p class="text-muted-foreground leading-7 my-4">The work did not disappear. It changed shape.</p>
<p class="text-muted-foreground leading-7 my-4">I still have to know what good looks like. I still have to understand the system well enough to spot a bad abstraction, a leaky permission model, or a feature that technically works but should not exist. I still have to decide which problems are worth solving and which ones are distractions wearing a product costume.</p>
<p class="text-muted-foreground leading-7 my-4">What I am not doing as much is the mechanical part:</p>
<ul class="space-y-2 my-6 ml-4">
<li class="flex gap-3 text-muted-foreground leading-7"><span class="text-primary shrink-0 mt-1.5">▸</span><span class="min-w-0">Creating the fifth version of the same CRUD flow</span></li>
<li class="flex gap-3 text-muted-foreground leading-7"><span class="text-primary shrink-0 mt-1.5">▸</span><span class="min-w-0">Moving state between components because the first pass guessed wrong</span></li>
<li class="flex gap-3 text-muted-foreground leading-7"><span class="text-primary shrink-0 mt-1.5">▸</span><span class="min-w-0">Writing boilerplate tests for behavior that is already clear</span></li>
<li class="flex gap-3 text-muted-foreground leading-7"><span class="text-primary shrink-0 mt-1.5">▸</span><span class="min-w-0">Updating copy across three files</span></li>
<li class="flex gap-3 text-muted-foreground leading-7"><span class="text-primary shrink-0 mt-1.5">▸</span><span class="min-w-0">Chasing TypeScript errors caused by a renamed prop</span></li>
<li class="flex gap-3 text-muted-foreground leading-7"><span class="text-primary shrink-0 mt-1.5">▸</span><span class="min-w-0">Reading docs for the third-party integration and translating them into glue code</span></li>
</ul>
<p class="text-muted-foreground leading-7 my-4">That work still matters. It is just no longer the highest-leverage place for me to spend my attention.</p>
<h2 id="i-am-not-coding-less-i-am-delegating-more" class="font-display text-2xl lg:text-3xl font-bold text-foreground mt-12 mb-6 border-l-2 border-primary pl-4"><a href="#i-am-not-coding-less-i-am-delegating-more" class="heading-permalink text-inherit no-underline">I Am Not Coding Less. I Am Delegating More.</a></h2>
<p class="text-muted-foreground leading-7 my-4">The mental model that finally clicked for me is delegation.</p>
<p class="text-muted-foreground leading-7 my-4">For most of my career, &quot;building software&quot; meant converting intent into code with my own hands. I would hold the product shape in my head, make a series of implementation decisions, and type the thing into existence.</p>
<p class="text-muted-foreground leading-7 my-4">Now the workflow feels closer to leading a small implementation team:</p>
<ol class="space-y-2 my-6 ml-4">
<li class="flex gap-3 text-muted-foreground leading-7"><span class="text-primary shrink-0 tabular-nums">1.</span><span class="min-w-0">Define the outcome clearly.</span></li>
<li class="flex gap-3 text-muted-foreground leading-7"><span class="text-primary shrink-0 tabular-nums">2.</span><span class="min-w-0">Point the agent at the relevant context.</span></li>
<li class="flex gap-3 text-muted-foreground leading-7"><span class="text-primary shrink-0 tabular-nums">3.</span><span class="min-w-0">Review the plan before it writes code.</span></li>
<li class="flex gap-3 text-muted-foreground leading-7"><span class="text-primary shrink-0 tabular-nums">4.</span><span class="min-w-0">Let it implement.</span></li>
<li class="flex gap-3 text-muted-foreground leading-7"><span class="text-primary shrink-0 tabular-nums">5.</span><span class="min-w-0">Review the diff, the tests, and the behavior.</span></li>
<li class="flex gap-3 text-muted-foreground leading-7"><span class="text-primary shrink-0 tabular-nums">6.</span><span class="min-w-0">Push back where the work misses the intent.</span></li>
</ol>
<p class="text-muted-foreground leading-7 my-4">That is not passive. It is a different kind of active.</p>
<p class="text-muted-foreground leading-7 my-4">The quality of the output is still a reflection of the quality of the input. If I give an agent a lazy task description, I get lazy work back. If I give it the actual constraints, the edge cases, the product intent, and the shape of the existing system, the result is usually good enough to review like a normal pull request.</p>
<p class="text-muted-foreground leading-7 my-4">The job moved from &quot;write the code&quot; to &quot;create the conditions where good code is likely to emerge.&quot;</p>
<h2 id="the-weird-skill-loss-is-real" class="font-display text-2xl lg:text-3xl font-bold text-foreground mt-12 mb-6 border-l-2 border-primary pl-4"><a href="#the-weird-skill-loss-is-real" class="heading-permalink text-inherit no-underline">The Weird Skill Loss Is Real</a></h2>
<p class="text-muted-foreground leading-7 my-4">There is a cost here, and I do not want to hand-wave it away.</p>
<p class="text-muted-foreground leading-7 my-4">I am probably worse at remembering exact APIs than I was a year ago. I am less practiced at grinding through implementation details from a blank file. I feel the friction when I do drop back into the code editor and have to rehydrate the local context myself.</p>
<p class="text-muted-foreground leading-7 my-4">That is real.</p>
<p class="text-muted-foreground leading-7 my-4">But it is not the same as getting worse at engineering. It is closer to what happens when a senior engineer stops being the person who personally writes every line and starts being the person who sets direction, reviews work, and keeps the system coherent.</p>
<p class="text-muted-foreground leading-7 my-4">Some skills fade. Other skills compound.</p>
<p class="text-muted-foreground leading-7 my-4">The skills that feel more important now:</p>
<ul class="space-y-2 my-6 ml-4">
<li class="flex gap-3 text-muted-foreground leading-7"><span class="text-primary shrink-0 mt-1.5">▸</span><span class="min-w-0"><strong class="text-foreground font-bold">Taste:</strong> knowing which version of the product should exist</span></li>
<li class="flex gap-3 text-muted-foreground leading-7"><span class="text-primary shrink-0 mt-1.5">▸</span><span class="min-w-0"><strong class="text-foreground font-bold">Context design:</strong> giving AI the right boundaries, examples, and constraints</span></li>
<li class="flex gap-3 text-muted-foreground leading-7"><span class="text-primary shrink-0 mt-1.5">▸</span><span class="min-w-0"><strong class="text-foreground font-bold">Review judgment:</strong> spotting the subtle wrongness in code that passes tests</span></li>
<li class="flex gap-3 text-muted-foreground leading-7"><span class="text-primary shrink-0 mt-1.5">▸</span><span class="min-w-0"><strong class="text-foreground font-bold">Systems thinking:</strong> understanding where a change belongs before anyone implements it</span></li>
<li class="flex gap-3 text-muted-foreground leading-7"><span class="text-primary shrink-0 mt-1.5">▸</span><span class="min-w-0"><strong class="text-foreground font-bold">Risk calibration:</strong> knowing when AI output is fine and when it needs deep scrutiny</span></li>
</ul>
<p class="text-muted-foreground leading-7 my-4">Those are not softer skills. They are the skills that decide whether AI-generated work becomes leverage or liability.</p>
<h2 id="the-dangerous-part-is-feeling-fast" class="font-display text-2xl lg:text-3xl font-bold text-foreground mt-12 mb-6 border-l-2 border-primary pl-4"><a href="#the-dangerous-part-is-feeling-fast" class="heading-permalink text-inherit no-underline">The Dangerous Part Is Feeling Fast</a></h2>
<p class="text-muted-foreground leading-7 my-4">The biggest risk in this workflow is not that AI writes bad code. Bad code can be reviewed, tested, and fixed.</p>
<p class="text-muted-foreground leading-7 my-4">The bigger risk is that AI makes bad decisions feel cheap.</p>
<p class="text-muted-foreground leading-7 my-4">When implementation is expensive, you naturally hesitate. You ask whether the feature is worth it. You think about maintenance. You negotiate scope because every extra requirement costs real time.</p>
<p class="text-muted-foreground leading-7 my-4">When implementation feels almost free, the discipline has to come from somewhere else. You need stronger taste, not weaker. You need to say no more often, not less. Otherwise you end up with a product full of features nobody needed, all shipped efficiently.</p>
<p class="text-muted-foreground leading-7 my-4">That is the operator lesson I keep coming back to: <strong class="text-foreground font-bold">AI removes implementation friction, so judgment becomes the bottleneck.</strong></p>
<p class="text-muted-foreground leading-7 my-4">This matters a lot for what we are building at <a href="https://poof.new" class="text-primary underline underline-offset-2 decoration-primary/40 print:decoration-primary hover:decoration-primary transition-colors">poof.new</a>. If someone can describe an app and get working software back, the scarce skill is no longer syntax. It is intent. The builder has to know what they want, why it matters, and how to tell whether the result is actually good.</p>
<p class="text-muted-foreground leading-7 my-4">AI can make software real. It cannot tell you whether that software deserves to exist.</p>
<h2 id="what-my-day-looks-like-now" class="font-display text-2xl lg:text-3xl font-bold text-foreground mt-12 mb-6 border-l-2 border-primary pl-4"><a href="#what-my-day-looks-like-now" class="heading-permalink text-inherit no-underline">What My Day Looks Like Now</a></h2>
<p class="text-muted-foreground leading-7 my-4">A typical building day used to start with me opening the editor and finding the first file to change.</p>
<p class="text-muted-foreground leading-7 my-4">Now it starts with writing a better task.</p>
<p class="text-muted-foreground leading-7 my-4">I spend more time turning fuzzy ideas into crisp implementation briefs. I describe the user outcome, the current behavior, the desired behavior, the files that probably matter, the constraints that cannot be violated, and the tests that should pass when the work is done.</p>
<p class="text-muted-foreground leading-7 my-4">Then I review the plan.</p>
<p class="text-muted-foreground leading-7 my-4">This is the step I used to undervalue. Plan review is where most AI work succeeds or fails. If the plan reveals a wrong assumption, I would much rather catch it before the agent rewrites half the feature. A five-minute correction at the plan stage can save an hour of cleanup later.</p>
<p class="text-muted-foreground leading-7 my-4">Once the PR exists, I review it in layers:</p>
<ul class="space-y-2 my-6 ml-4">
<li class="flex gap-3 text-muted-foreground leading-7"><span class="text-primary shrink-0 mt-1.5">▸</span><span class="min-w-0">Does the behavior match the intent?</span></li>
<li class="flex gap-3 text-muted-foreground leading-7"><span class="text-primary shrink-0 mt-1.5">▸</span><span class="min-w-0">Is the implementation consistent with the existing system?</span></li>
<li class="flex gap-3 text-muted-foreground leading-7"><span class="text-primary shrink-0 mt-1.5">▸</span><span class="min-w-0">Did it add unnecessary abstraction?</span></li>
<li class="flex gap-3 text-muted-foreground leading-7"><span class="text-primary shrink-0 mt-1.5">▸</span><span class="min-w-0">Are the tests proving the right thing?</span></li>
<li class="flex gap-3 text-muted-foreground leading-7"><span class="text-primary shrink-0 mt-1.5">▸</span><span class="min-w-0">Is there any security or data exposure risk?</span></li>
<li class="flex gap-3 text-muted-foreground leading-7"><span class="text-primary shrink-0 mt-1.5">▸</span><span class="min-w-0">Would I be comfortable owning this code six months from now?</span></li>
</ul>
<p class="text-muted-foreground leading-7 my-4">That last question is the anchor. AI can write the code, but I still own the consequences.</p>
<h2 id="the-new-builder-is-more-editor-than-typist" class="font-display text-2xl lg:text-3xl font-bold text-foreground mt-12 mb-6 border-l-2 border-primary pl-4"><a href="#the-new-builder-is-more-editor-than-typist" class="heading-permalink text-inherit no-underline">The New Builder Is More Editor Than Typist</a></h2>
<p class="text-muted-foreground leading-7 my-4">There is an old romantic version of programming where the builder sits alone, enters flow state, and produces perfect software through direct contact with the machine.</p>
<p class="text-muted-foreground leading-7 my-4">I still love that feeling. I do not think it is going away entirely. There will always be moments where the fastest path is to open the file and make the change yourself.</p>
<p class="text-muted-foreground leading-7 my-4">But for a growing slice of product work, the highest-leverage builder looks less like a typist and more like an editor:</p>
<p class="text-muted-foreground leading-7 my-4">They know what to ask for. They know what to cut. They know when the first draft is good enough and when it is structurally wrong. They can look at a pile of generated work and see the one decision that needs to change.</p>
<p class="text-muted-foreground leading-7 my-4">That is not less creative. It is more like the creativity moved earlier in the process.</p>
<p class="text-muted-foreground leading-7 my-4">The blank page is not the code editor anymore. The blank page is the instruction.</p>
<h2 id="the-bottom-line" class="font-display text-2xl lg:text-3xl font-bold text-foreground mt-12 mb-6 border-l-2 border-primary pl-4"><a href="#the-bottom-line" class="heading-permalink text-inherit no-underline">The Bottom Line</a></h2>
<p class="text-muted-foreground leading-7 my-4">I have not touched much code in a month, and I am trying to be honest about both sides of that.</p>
<p class="text-muted-foreground leading-7 my-4">Yes, some hands-on sharpness fades when AI handles the implementation reps. If I had to sit down tomorrow and rebuild a feature entirely from scratch, I might feel slower than I used to.</p>
<p class="text-muted-foreground leading-7 my-4">But I am also shipping more, thinking more clearly, and spending more of my time on the parts of building that actually determine whether the work matters.</p>
<p class="text-muted-foreground leading-7 my-4">So no, I do not think AI made me dumber.</p>
<p class="text-muted-foreground leading-7 my-4">It made me more aware of which parts of my intelligence were being spent on low-leverage work.</p>
<p class="text-muted-foreground leading-7 my-4">The goal is not to never code again. The goal is to make code the last mile, not the whole journey. And once you experience that shift, it is hard to go back.</p>]]></content:encoded>
      <pubDate>Tue, 26 May 2026 12:00:00 GMT</pubDate>
    </item>
    <item>
      <title>I Replaced My Team With AI. Here&apos;s What I Miss.</title>
      <link>https://pratik.pa.tel/blog/what-i-miss-about-having-a-team/</link>
      <guid isPermaLink="true">https://pratik.pa.tel/blog/what-i-miss-about-having-a-team/</guid>
      <description>Solo building is a superpower. But superpowers have side effects.</description>
      <dc:creator>Pratik Patel</dc:creator>
      <category>ai</category>
      <category>building</category>
      <category>teams</category>
      <category>reflection</category>
      <content:encoded><![CDATA[<p class="text-muted-foreground leading-7 my-4">I&#x27;ve spent the last few months writing about how AI makes it possible to build alone. Ship it yourself. The $0 startup. Taste as a moat. Distribution as the new code. All of it true. All of it real. I believe every word I wrote.</p>
<p class="text-muted-foreground leading-7 my-4">But I haven&#x27;t been fully honest.</p>
<p class="text-muted-foreground leading-7 my-4">There&#x27;s a version of this story I&#x27;ve been leaving out. The version where it&#x27;s 11 PM, you&#x27;ve shipped something you&#x27;re proud of, and there&#x27;s nobody to high-five. The version where you make a decision that feels right but you&#x27;re not 100% sure, and there&#x27;s no one to push back. The version where the silence of solo building starts to feel less like freedom and more like isolation.</p>
<p class="text-muted-foreground leading-7 my-4">This is that version.</p>
<h2 id="whats-genuinely-better" class="font-display text-2xl lg:text-3xl font-bold text-foreground mt-12 mb-6 border-l-2 border-primary pl-4"><a href="#whats-genuinely-better" class="heading-permalink text-inherit no-underline">What&#x27;s Genuinely Better</a></h2>
<p class="text-muted-foreground leading-7 my-4">Let me be clear: I&#x27;m not writing a &quot;going back to the office&quot; piece. Working with AI agents is objectively better in a dozen ways.</p>
<p class="text-muted-foreground leading-7 my-4"><strong class="text-foreground font-bold">Speed.</strong> Decisions that used to require three meetings and a Slack thread now happen in my head. I think it, the agent builds it, I ship it. The feedback loop is measured in minutes, not sprints.</p>
<p class="text-muted-foreground leading-7 my-4"><strong class="text-foreground font-bold">No consensus tax.</strong> Every team I&#x27;ve been on has a hidden cost: the energy spent getting everyone aligned. Debating naming conventions. Arguing about architecture decisions that don&#x27;t actually matter. When you&#x27;re solo, that tax drops to zero. You just decide and move.</p>
<p class="text-muted-foreground leading-7 my-4"><strong class="text-foreground font-bold">Deep focus.</strong> I haven&#x27;t been interrupted mid-thought in months. No standup pulling me out of flow state. No &quot;quick question&quot; that takes 45 minutes. My calendar is empty and my output has never been higher.</p>
<p class="text-muted-foreground leading-7 my-4"><strong class="text-foreground font-bold">Full context.</strong> I know every line of code, every design decision, every tradeoff. There&#x27;s no knowledge transfer problem because there&#x27;s no one to transfer to. The entire system lives in my head, and the AI agents have access to all of it.</p>
<p class="text-muted-foreground leading-7 my-4">These are real gains. I&#x27;m more productive than I&#x27;ve ever been. But productivity isn&#x27;t the whole picture.</p>
<h2 id="the-are-you-sure-moment" class="font-display text-2xl lg:text-3xl font-bold text-foreground mt-12 mb-6 border-l-2 border-primary pl-4"><a href="#the-are-you-sure-moment" class="heading-permalink text-inherit no-underline">The &quot;Are You Sure?&quot; Moment</a></h2>
<p class="text-muted-foreground leading-7 my-4">The thing I miss most is the challenge.</p>
<p class="text-muted-foreground leading-7 my-4">Not conflict. Not arguing. The specific moment when a smart colleague looks at your work and says, &quot;Have you thought about this from the user&#x27;s perspective?&quot; or &quot;What happens when this breaks at scale?&quot; or simply, &quot;Are you sure?&quot;</p>
<p class="text-muted-foreground leading-7 my-4">AI agents don&#x27;t do this. They execute. They&#x27;re remarkably good at building what you ask for. But they almost never question whether you should be asking for it. They won&#x27;t tell you your priority is wrong. They won&#x27;t push back on a design because it feels off, even if it technically meets the spec.</p>
<p class="text-muted-foreground leading-7 my-4">I used to find that friction annoying. Now I realize it was the most valuable part of working with people. The best ideas I&#x27;ve ever shipped got better because someone challenged them. And the worst ideas I&#x27;ve ever had died because someone had the courage to say, &quot;I don&#x27;t think this is right.&quot;</p>
<p class="text-muted-foreground leading-7 my-4">When you&#x27;re solo, that safety net disappears. Every bad idea has a clear path to production.</p>
<h2 id="complementary-taste" class="font-display text-2xl lg:text-3xl font-bold text-foreground mt-12 mb-6 border-l-2 border-primary pl-4"><a href="#complementary-taste" class="heading-permalink text-inherit no-underline">Complementary Taste</a></h2>
<p class="text-muted-foreground leading-7 my-4">I wrote about taste being a moat. I still believe that. But here&#x27;s what I didn&#x27;t say: your taste has blind spots.</p>
<p class="text-muted-foreground leading-7 my-4">Everyone&#x27;s does. You gravitate toward certain aesthetics, certain patterns, certain types of solutions. When you work with people who have different taste, different backgrounds, different instincts, the output is richer than anything one person can produce alone.</p>
<p class="text-muted-foreground leading-7 my-4">My AI agents share my taste because they learned it from my instructions and my feedback. They&#x27;re an amplifier, not a counterweight. And sometimes what you need isn&#x27;t amplification. It&#x27;s someone who sees the thing you&#x27;re missing.</p>
<h2 id="the-accountability-gap" class="font-display text-2xl lg:text-3xl font-bold text-foreground mt-12 mb-6 border-l-2 border-primary pl-4"><a href="#the-accountability-gap" class="heading-permalink text-inherit no-underline">The Accountability Gap</a></h2>
<p class="text-muted-foreground leading-7 my-4">When I had a team, there was a natural rhythm of accountability. Someone was waiting on my work. Someone would notice if I went down a rabbit hole for three days. Someone would flag if the project was drifting off course.</p>
<p class="text-muted-foreground leading-7 my-4">Now? I&#x27;m accountable to myself. Which sounds empowering until you realize that humans are terrible at holding themselves accountable. I&#x27;ve spent entire weeks perfecting features that didn&#x27;t matter, because nobody was there to tap me on the shoulder and say, &quot;Is this really the highest-impact thing you could be doing right now?&quot;</p>
<p class="text-muted-foreground leading-7 my-4">AI agents will happily help you polish something irrelevant. They don&#x27;t have the judgment to say, &quot;Stop. Step back. Look at the big picture.&quot; That judgment is a uniquely human contribution, and it&#x27;s one I took for granted.</p>
<h2 id="the-energy-problem" class="font-display text-2xl lg:text-3xl font-bold text-foreground mt-12 mb-6 border-l-2 border-primary pl-4"><a href="#the-energy-problem" class="heading-permalink text-inherit no-underline">The Energy Problem</a></h2>
<p class="text-muted-foreground leading-7 my-4">This one surprised me.</p>
<p class="text-muted-foreground leading-7 my-4">I expected to love the quiet. And I do, sometimes. But building is an emotional activity, not just an intellectual one. There&#x27;s an energy you get from working alongside people who care about the same thing. The laugh when something finally works after hours of debugging. The shared frustration when a deploy goes sideways. The momentum that comes from knowing someone else is counting on you to show up.</p>
<p class="text-muted-foreground leading-7 my-4">AI agents don&#x27;t generate that energy. They&#x27;re tireless but not inspiring. They&#x27;ll work at 3 AM without complaint, but they won&#x27;t text you at 3 AM with an idea they can&#x27;t stop thinking about. The human spark, the irrational enthusiasm that makes hard work feel meaningful, you can&#x27;t prompt your way to that.</p>
<h2 id="what-ive-learned" class="font-display text-2xl lg:text-3xl font-bold text-foreground mt-12 mb-6 border-l-2 border-primary pl-4"><a href="#what-ive-learned" class="heading-permalink text-inherit no-underline">What I&#x27;ve Learned</a></h2>
<p class="text-muted-foreground leading-7 my-4">I&#x27;m not going back. The productivity gains are too real, and the model works too well to abandon. But I&#x27;ve stopped pretending that solo AI building is a pure upgrade with no downsides.</p>
<p class="text-muted-foreground leading-7 my-4">What I&#x27;ve started doing instead:</p>
<p class="text-muted-foreground leading-7 my-4"><strong class="text-foreground font-bold">Finding deliberate friction.</strong> I share work with a small group of people I trust before shipping anything important. Not for approval. For challenge. I specifically ask them to find problems, to question assumptions, to tell me what feels wrong. This replaces the organic friction that teams provide naturally.</p>
<p class="text-muted-foreground leading-7 my-4"><strong class="text-foreground font-bold">Seeking complementary taste.</strong> I follow builders whose instincts are different from mine. When I&#x27;m stuck on a design decision, I look at how they&#x27;d approach it. Not to copy, but to see what I&#x27;m not seeing.</p>
<p class="text-muted-foreground leading-7 my-4"><strong class="text-foreground font-bold">Creating external accountability.</strong> Building in public is part of this. When I commit to a shipping date on Twitter, I&#x27;ve created accountability that didn&#x27;t exist before. The audience becomes the team, in a sense.</p>
<p class="text-muted-foreground leading-7 my-4"><strong class="text-foreground font-bold">Protecting against isolation.</strong> I schedule regular calls with other solo builders. Not networking. Not masterminds. Just human conversation with people who understand the specific loneliness of building alone. It&#x27;s the cheapest investment with the highest return.</p>
<h2 id="build-solo-but-dont-build-alone" class="font-display text-2xl lg:text-3xl font-bold text-foreground mt-12 mb-6 border-l-2 border-primary pl-4"><a href="#build-solo-but-dont-build-alone" class="heading-permalink text-inherit no-underline">Build Solo, But Don&#x27;t Build Alone</a></h2>
<p class="text-muted-foreground leading-7 my-4">The future of building is smaller teams, more AI, more individual leverage. I&#x27;m convinced of that. But &quot;smaller teams&quot; is not the same as &quot;no team.&quot; And &quot;more AI&quot; is not the same as &quot;no humans.&quot;</p>
<p class="text-muted-foreground leading-7 my-4">The best version of this new model isn&#x27;t a person alone with their agents. It&#x27;s a person with agents AND a small, intentional network of humans who provide the things AI can&#x27;t: challenge, complementary taste, accountability, and energy.</p>
<p class="text-muted-foreground leading-7 my-4">You don&#x27;t need those people on your payroll. You don&#x27;t need them in your Slack. But you need them in your life. Because the most dangerous thing about building alone isn&#x27;t that you&#x27;ll build something bad. It&#x27;s that you&#x27;ll build something good enough, and never know how much better it could have been if someone had pushed you.</p>
<p class="text-muted-foreground leading-7 my-4">Ship it yourself. But find your people first.</p>]]></content:encoded>
      <pubDate>Tue, 19 May 2026 12:00:00 GMT</pubDate>
    </item>
    <item>
      <title>Distribution Is the New Code</title>
      <link>https://pratik.pa.tel/blog/distribution-is-the-new-code/</link>
      <guid isPermaLink="true">https://pratik.pa.tel/blog/distribution-is-the-new-code/</guid>
      <description>AI can build anything. The hard part is getting anyone to care.</description>
      <dc:creator>Pratik Patel</dc:creator>
      <category>ai</category>
      <category>distribution</category>
      <category>startups</category>
      <category>marketing</category>
      <content:encoded><![CDATA[<p class="text-muted-foreground leading-7 my-4">There&#x27;s a pattern I keep seeing in my DMs. Someone builds something genuinely good (clean product, solid code, real utility) and then wonders why nobody&#x27;s using it. They share a link on Twitter, post it to a couple of subreddits, and wait. Nothing happens. A week later they&#x27;re demoralized, convinced the product wasn&#x27;t good enough.</p>
<p class="text-muted-foreground leading-7 my-4">It was good enough. The product wasn&#x27;t the problem. <strong class="text-foreground font-bold">Getting it in front of people</strong> was the problem.</p>
<p class="text-muted-foreground leading-7 my-4">We&#x27;ve spent the last two years celebrating how AI collapsed the cost of building. I wrote about it. The $0 startup is real, shipping solo is real, taste as a moat is real. But nobody wants to talk about what comes next: <strong class="text-foreground font-bold">building was never the hardest part</strong>. Distribution was. And it still is.</p>
<h2 id="the-biggest-lie-in-tech" class="font-display text-2xl lg:text-3xl font-bold text-foreground mt-12 mb-6 border-l-2 border-primary pl-4"><a href="#the-biggest-lie-in-tech" class="heading-permalink text-inherit no-underline">The Biggest Lie in Tech</a></h2>
<p class="text-muted-foreground leading-7 my-4">Silicon Valley has a mythology problem. The story goes like this: build something great, and the world will beat a path to your door. &quot;If you build it, they will come.&quot; It&#x27;s the foundational myth of the startup world, and it has always been a lie.</p>
<p class="text-muted-foreground leading-7 my-4">Google wasn&#x27;t the best search engine when it launched. It had better distribution through academic networks. Facebook wasn&#x27;t the best social network. It had exclusivity and campus-by-campus rollout. Slack wasn&#x27;t the best chat app. It had a viral loop baked into its team onboarding. The winners didn&#x27;t just build better products. They <strong class="text-foreground font-bold">built better distribution</strong>.</p>
<p class="text-muted-foreground leading-7 my-4">This was already true before AI. Now it&#x27;s the <em class="font-mono font-normal text-primary/80 print:text-primary">only</em> truth that matters.</p>
<h2 id="building-isnt-the-bottleneck-anymore" class="font-display text-2xl lg:text-3xl font-bold text-foreground mt-12 mb-6 border-l-2 border-primary pl-4"><a href="#building-isnt-the-bottleneck-anymore" class="heading-permalink text-inherit no-underline">Building Isn&#x27;t the Bottleneck Anymore</a></h2>
<p class="text-muted-foreground leading-7 my-4">The math is simple. In 2024, maybe 500,000 people worldwide could build a production-quality web app from scratch. Designers, engineers, full-stack developers, the usual suspects. In 2026, that number is closer to 50 million. AI agents turned &quot;non-technical&quot; people into builders overnight.</p>
<p class="text-muted-foreground leading-7 my-4">That&#x27;s a 100x increase in supply. What happens when supply explodes? The thing that was scarce stops being scarce. <strong class="text-foreground font-bold">The ability to build is no longer a competitive advantage.</strong> It&#x27;s table stakes.</p>
<p class="text-muted-foreground leading-7 my-4">So what&#x27;s still scarce? The ability to reach people. Attention. Trust. An audience that listens when you have something to say. That&#x27;s distribution, and no AI agent is going to build it for you while you sleep.</p>
<h2 id="what-distribution-actually-means" class="font-display text-2xl lg:text-3xl font-bold text-foreground mt-12 mb-6 border-l-2 border-primary pl-4"><a href="#what-distribution-actually-means" class="heading-permalink text-inherit no-underline">What Distribution Actually Means</a></h2>
<p class="text-muted-foreground leading-7 my-4">Let me be specific, because &quot;distribution&quot; gets thrown around as a buzzword. It&#x27;s not one thing. It&#x27;s a stack:</p>
<p class="text-muted-foreground leading-7 my-4"><strong class="text-foreground font-bold">Audience.</strong> People who already know you exist and have some reason to pay attention. This could be a Twitter following, a newsletter list, a YouTube channel, a podcast. The format doesn&#x27;t matter. What matters is that when you say &quot;I built a thing,&quot; there are humans on the other end who will actually look at it.</p>
<p class="text-muted-foreground leading-7 my-4"><strong class="text-foreground font-bold">Trust.</strong> An audience that doesn&#x27;t trust you is just a number. Trust comes from consistency, from showing up, sharing real work, being honest about what works and what doesn&#x27;t. The people who do best at distribution aren&#x27;t the loudest. They&#x27;re the most consistent.</p>
<p class="text-muted-foreground leading-7 my-4"><strong class="text-foreground font-bold">Channel fit.</strong> Every product has natural channels where it spreads, and forcing it into the wrong one is a waste of time. A developer tool spreads through GitHub stars and technical blog posts. A consumer app spreads through TikTok and word of mouth. A B2B product spreads through case studies and LinkedIn. If you&#x27;re posting your B2B SaaS on Reddit&#x27;s r/sideproject, you&#x27;re putting premium gas in a bicycle.</p>
<p class="text-muted-foreground leading-7 my-4"><strong class="text-foreground font-bold">Story.</strong> People don&#x27;t share products. They share stories. &quot;I built an app&quot; is not a story. &quot;I quit my job, built a tool that replaced my old team&#x27;s entire workflow, and now it&#x27;s making $10K/month&quot; is a story. The narrative around your product matters as much as the product itself. Maybe more.</p>
<h2 id="the-distribution-divide" class="font-display text-2xl lg:text-3xl font-bold text-foreground mt-12 mb-6 border-l-2 border-primary pl-4"><a href="#the-distribution-divide" class="heading-permalink text-inherit no-underline">The Distribution Divide</a></h2>
<p class="text-muted-foreground leading-7 my-4">I&#x27;m watching this play out in real time: the gap between builders and distributors is becoming the defining divide in tech.</p>
<p class="text-muted-foreground leading-7 my-4">On one side, you have incredible builders, technical people shipping polished products every week, who can&#x27;t get past 50 users. They keep iterating on features, convinced that the <em class="font-mono font-normal text-primary/80 print:text-primary">next</em> feature will be the one that unlocks growth. It never is.</p>
<p class="text-muted-foreground leading-7 my-4">On the other side, you have people with strong audiences and distribution channels who build something mediocre and immediately get traction. Their v1 is worse, but it doesn&#x27;t matter because 10,000 people tried it on day one. And with real user feedback flowing in from day one, their v2 is better than the first group&#x27;s v5.</p>
<p class="text-muted-foreground leading-7 my-4">This isn&#x27;t fair. But it&#x27;s how markets work. <strong class="text-foreground font-bold">The product with distribution beats the better product without it, almost every time.</strong></p>
<h2 id="building-distribution-before-you-build-the-product" class="font-display text-2xl lg:text-3xl font-bold text-foreground mt-12 mb-6 border-l-2 border-primary pl-4"><a href="#building-distribution-before-you-build-the-product" class="heading-permalink text-inherit no-underline">Building Distribution Before You Build the Product</a></h2>
<p class="text-muted-foreground leading-7 my-4">The single best piece of advice I can give to anyone planning to build something: start your distribution engine <em class="font-mono font-normal text-primary/80 print:text-primary">before</em> you write a line of code.</p>
<h3 id="1-build-in-public" class="font-display text-xl font-bold text-foreground mt-10 mb-4"><a href="#1-build-in-public" class="heading-permalink text-inherit no-underline">1. Build in Public</a></h3>
<p class="text-muted-foreground leading-7 my-4">Share your process, not just your product. People love watching things get built. Tweet your progress. Write about your decisions. Show the messy middle. By the time you launch, you&#x27;ll have an audience that feels invested in your success. They watched it happen.</p>
<h3 id="2-own-a-narrative" class="font-display text-xl font-bold text-foreground mt-10 mb-4"><a href="#2-own-a-narrative" class="heading-permalink text-inherit no-underline">2. Own a Narrative</a></h3>
<p class="text-muted-foreground leading-7 my-4">Pick a point of view and commit to it. &quot;AI is changing how we work&quot; is too generic. &quot;Solo founders with AI teams will outperform funded startups within 5 years&quot; is a narrative. When you own a specific, debatable take, people remember you. They share your posts because they either strongly agree or strongly disagree. Both are good for distribution.</p>
<h3 id="3-give-away-your-best-ideas" class="font-display text-xl font-bold text-foreground mt-10 mb-4"><a href="#3-give-away-your-best-ideas" class="heading-permalink text-inherit no-underline">3. Give Away Your Best Ideas</a></h3>
<p class="text-muted-foreground leading-7 my-4">Counterintuitive, but the most effective distribution strategy I&#x27;ve seen is radical generosity. Share your frameworks. Publish your playbooks. Give away the thinking behind your product for free. The people who consume that content self-select as your ideal users. When you launch, they already understand the problem and trust your approach to solving it.</p>
<h3 id="4-choose-one-channel-and-go-deep" class="font-display text-xl font-bold text-foreground mt-10 mb-4"><a href="#4-choose-one-channel-and-go-deep" class="heading-permalink text-inherit no-underline">4. Choose One Channel and Go Deep</a></h3>
<p class="text-muted-foreground leading-7 my-4">Don&#x27;t try to be everywhere. Pick the one channel where your target users already spend time, and become impossible to ignore on that channel. One great Twitter presence beats a mediocre presence on Twitter, LinkedIn, YouTube, TikTok, and a blog combined. Depth beats breadth in distribution.</p>
<h2 id="ai-cant-distribute-for-you" class="font-display text-2xl lg:text-3xl font-bold text-foreground mt-12 mb-6 border-l-2 border-primary pl-4"><a href="#ai-cant-distribute-for-you" class="heading-permalink text-inherit no-underline">AI Can&#x27;t Distribute For You</a></h2>
<p class="text-muted-foreground leading-7 my-4">The uncomfortable truth is one the AI hype cycle doesn&#x27;t want to acknowledge: AI is great at building things, but it&#x27;s mediocre at distributing them.</p>
<p class="text-muted-foreground leading-7 my-4">Yes, AI can write social media posts and schedule them. It can generate content at scale. But the output feels like what it is: machine-generated filler. And people are getting better at pattern-matching on AI slop every day.</p>
<p class="text-muted-foreground leading-7 my-4">Distribution that works is built on <strong class="text-foreground font-bold">authenticity</strong>. It&#x27;s your real story, your real struggle, your real perspective. The irony is rich: in a world where AI can fake almost anything, the unfakeable stuff, genuine human experience and earned trust, has become the most valuable asset of all.</p>
<h2 id="the-new-stack" class="font-display text-2xl lg:text-3xl font-bold text-foreground mt-12 mb-6 border-l-2 border-primary pl-4"><a href="#the-new-stack" class="heading-permalink text-inherit no-underline">The New Stack</a></h2>
<p class="text-muted-foreground leading-7 my-4">We used to say &quot;code is the new literacy.&quot; Everyone should learn to code, we said, because code is how you build the future. That was true for a while. It&#x27;s not anymore.</p>
<p class="text-muted-foreground leading-7 my-4">In 2026, the stack looks different:</p>
<ul class="space-y-2 my-6 ml-4">
<li class="flex gap-3 text-muted-foreground leading-7"><span class="text-primary shrink-0 mt-1.5">▸</span><span class="min-w-0"><strong class="text-foreground font-bold">Taste</strong> decides what to build (I covered this in &quot;Your Taste Is Your Moat&quot;)</span></li>
<li class="flex gap-3 text-muted-foreground leading-7"><span class="text-primary shrink-0 mt-1.5">▸</span><span class="min-w-0"><strong class="text-foreground font-bold">AI</strong> handles the building</span></li>
<li class="flex gap-3 text-muted-foreground leading-7"><span class="text-primary shrink-0 mt-1.5">▸</span><span class="min-w-0"><strong class="text-foreground font-bold">Distribution</strong> determines whether anyone ever uses it</span></li>
</ul>
<p class="text-muted-foreground leading-7 my-4">Two out of three are in your control. And the one most builders ignore is the one that requires the most sustained effort — distribution.</p>
<h2 id="the-question-you-should-be-asking" class="font-display text-2xl lg:text-3xl font-bold text-foreground mt-12 mb-6 border-l-2 border-primary pl-4"><a href="#the-question-you-should-be-asking" class="heading-permalink text-inherit no-underline">The Question You Should Be Asking</a></h2>
<p class="text-muted-foreground leading-7 my-4">If you&#x27;re building something right now, I have a simple question: <strong class="text-foreground font-bold">do you have a plan to get it in front of 1,000 people on launch day?</strong> Not a vague hope. A plan. Names of communities, channels, people who will amplify it.</p>
<p class="text-muted-foreground leading-7 my-4">If the answer is no, stop adding features. Close your code editor. Open a blank document and write your distribution plan instead.</p>
<p class="text-muted-foreground leading-7 my-4">The products that win in 2026 won&#x27;t be the best-built. They&#x27;ll be the best-distributed. And the founders who figure that out early will run circles around the ones who keep polishing features in the dark.</p>
<p class="text-muted-foreground leading-7 my-4">Code used to be the bottleneck. Now it&#x27;s free. Distribution is the new code. And unlike code, you can&#x27;t outsource it to an agent.</p>
<p class="text-muted-foreground leading-7 my-4">Start building your audience today. Ship the product tomorrow. That&#x27;s the new order of operations.</p>]]></content:encoded>
      <pubDate>Tue, 12 May 2026 12:00:00 GMT</pubDate>
    </item>
    <item>
      <title>Your Taste Is Your Moat</title>
      <link>https://pratik.pa.tel/blog/taste-is-your-moat/</link>
      <guid isPermaLink="true">https://pratik.pa.tel/blog/taste-is-your-moat/</guid>
      <description>AI made execution cheap. The scarce resource now is knowing what to build.</description>
      <dc:creator>Pratik Patel</dc:creator>
      <category>ai</category>
      <category>engineering</category>
      <category>leadership</category>
      <category>product</category>
      <content:encoded><![CDATA[<p class="text-muted-foreground leading-7 my-4">Last month I watched someone build a full SaaS app in 45 minutes using AI tools. Working backend, auth, payments, dashboard. It was genuinely impressive.</p>
<p class="text-muted-foreground leading-7 my-4">It was also completely useless. Nobody wanted it.</p>
<p class="text-muted-foreground leading-7 my-4">The app worked perfectly. It just solved a problem that didn&#x27;t exist, in a way that nobody would choose to use, with a UI that felt like it was designed by committee. Every technical box was checked. Every product instinct was wrong.</p>
<p class="text-muted-foreground leading-7 my-4">This is the new failure mode. And I&#x27;m seeing it everywhere.</p>
<h2 id="execution-is-no-longer-the-hard-part" class="font-display text-2xl lg:text-3xl font-bold text-foreground mt-12 mb-6 border-l-2 border-primary pl-4"><a href="#execution-is-no-longer-the-hard-part" class="heading-permalink text-inherit no-underline">Execution Is No Longer the Hard Part</a></h2>
<p class="text-muted-foreground leading-7 my-4">For twenty years, the bottleneck in software was building the thing. You had an idea, and the gap between that idea and a working product was months of engineering, tens of thousands of dollars, and a team of specialists.</p>
<p class="text-muted-foreground leading-7 my-4">AI collapsed that gap to almost nothing. I wrote about this in &quot;The $0 Startup.&quot; The tools exist. The cost is near zero. Anyone can ship.</p>
<p class="text-muted-foreground leading-7 my-4">Which means shipping isn&#x27;t the differentiator anymore. If everyone can build, the question stops being &quot;can you build it?&quot; and becomes &quot;should you build it?&quot; And more importantly: &quot;should you build it <em class="font-mono font-normal text-primary/80 print:text-primary">this way</em>?&quot;</p>
<p class="text-muted-foreground leading-7 my-4">That&#x27;s taste. And right now, it&#x27;s the scarcest resource in tech.</p>
<h2 id="what-taste-actually-means" class="font-display text-2xl lg:text-3xl font-bold text-foreground mt-12 mb-6 border-l-2 border-primary pl-4"><a href="#what-taste-actually-means" class="heading-permalink text-inherit no-underline">What Taste Actually Means</a></h2>
<p class="text-muted-foreground leading-7 my-4">Taste isn&#x27;t about aesthetics. It&#x27;s not &quot;make it look pretty.&quot; It&#x27;s the ability to make a thousand small decisions correctly without having to reason through each one from first principles.</p>
<p class="text-muted-foreground leading-7 my-4">Should this button be here or there? Should this feature exist at all? Should we ship this now or wait until we&#x27;ve talked to ten more users? Should this error message be technical or friendly? Should this API return one object or a list?</p>
<p class="text-muted-foreground leading-7 my-4">Each decision is tiny. Each one barely matters on its own. But they compound. A product built by someone with good taste feels <em class="font-mono font-normal text-primary/80 print:text-primary">right</em> in a way that&#x27;s hard to articulate but impossible to miss. A product built without it feels off, even when nothing is technically broken.</p>
<p class="text-muted-foreground leading-7 my-4">Steve Jobs talked about this constantly. So did Dieter Rams. But you don&#x27;t have to be a design legend to have taste. You just have to care about the details that most people skip.</p>
<h2 id="the-ai-taste-gap" class="font-display text-2xl lg:text-3xl font-bold text-foreground mt-12 mb-6 border-l-2 border-primary pl-4"><a href="#the-ai-taste-gap" class="heading-permalink text-inherit no-underline">The AI Taste Gap</a></h2>
<p class="text-muted-foreground leading-7 my-4">AI tools are fantastic at execution. Give Claude or Cursor a well-defined task and it&#x27;ll produce clean, working code faster than most engineers. But ask it to decide <em class="font-mono font-normal text-primary/80 print:text-primary">what</em> to build? That&#x27;s where things fall apart.</p>
<p class="text-muted-foreground leading-7 my-4">I&#x27;ve tested this repeatedly. When I give an AI agent a specific, well-scoped task with clear context, the output is great. When I give it an open-ended problem, &quot;build something that helps developers manage their dotfiles,&quot; the result is technically competent and creatively dead. It builds the obvious thing. The thing that already exists twelve times on GitHub.</p>
<p class="text-muted-foreground leading-7 my-4">AI doesn&#x27;t have taste because taste comes from experience, opinions, and the willingness to say no to things that technically work. AI is a people-pleaser. It&#x27;ll build whatever you ask for. It won&#x27;t push back and say &quot;that feature is a bad idea&quot; or &quot;your users don&#x27;t actually want that.&quot;</p>
<p class="text-muted-foreground leading-7 my-4">That pushback is where taste lives.</p>
<h2 id="three-things-ive-learned-about-taste" class="font-display text-2xl lg:text-3xl font-bold text-foreground mt-12 mb-6 border-l-2 border-primary pl-4"><a href="#three-things-ive-learned-about-taste" class="heading-permalink text-inherit no-underline">Three Things I&#x27;ve Learned About Taste</a></h2>
<h3 id="1-taste-is-mostly-about-removal" class="font-display text-xl font-bold text-foreground mt-10 mb-4"><a href="#1-taste-is-mostly-about-removal" class="heading-permalink text-inherit no-underline">1. Taste is mostly about removal</a></h3>
<p class="text-muted-foreground leading-7 my-4">The best product decisions I&#x27;ve made weren&#x27;t about what to add. They were about what to cut. The feature that seemed important but would confuse the core flow. The settings page that gave users control they&#x27;d never use. The onboarding step that felt necessary but killed activation.</p>
<p class="text-muted-foreground leading-7 my-4">Junior engineers add. Senior engineers remove. The willingness to kill something that took a week to build, because it makes the overall product worse, is one of the clearest signals of taste I know.</p>
<h3 id="2-taste-requires-contact-with-real-users" class="font-display text-xl font-bold text-foreground mt-10 mb-4"><a href="#2-taste-requires-contact-with-real-users" class="heading-permalink text-inherit no-underline">2. Taste requires contact with real users</a></h3>
<p class="text-muted-foreground leading-7 my-4">You can&#x27;t develop taste in isolation. Every product instinct I trust was built by watching someone struggle with software I built. Not reading analytics dashboards. Not reviewing survey results. Sitting next to someone and watching them try to complete a task.</p>
<p class="text-muted-foreground leading-7 my-4">The founders I know who ship great products all do some version of this. They talk to users constantly. Not through feedback forms. Through actual conversations where they shut up and watch.</p>
<p class="text-muted-foreground leading-7 my-4">AI can&#x27;t do this. It can analyze usage data. It can summarize feedback. But it can&#x27;t sit in a room and notice that a user paused for three seconds before clicking a button, and understand what that pause means.</p>
<h3 id="3-taste-compounds-like-interest" class="font-display text-xl font-bold text-foreground mt-10 mb-4"><a href="#3-taste-compounds-like-interest" class="heading-permalink text-inherit no-underline">3. Taste compounds like interest</a></h3>
<p class="text-muted-foreground leading-7 my-4">Every product you ship, every user interaction you observe, every decision you make and see the result of, it all accumulates. Five years of building products gives you intuitions that no amount of reading or theorizing can replicate.</p>
<p class="text-muted-foreground leading-7 my-4">This is why experienced product builders are more valuable now than ever. Not less. The execution layer got automated. The judgment layer didn&#x27;t. A founder with ten years of product sense and AI tools is a force multiplier. A founder with AI tools and no product sense is just building faster in the wrong direction.</p>
<h2 id="what-this-means-for-engineers" class="font-display text-2xl lg:text-3xl font-bold text-foreground mt-12 mb-6 border-l-2 border-primary pl-4"><a href="#what-this-means-for-engineers" class="heading-permalink text-inherit no-underline">What This Means for Engineers</a></h2>
<p class="text-muted-foreground leading-7 my-4">If you&#x27;re an engineer reading this, the implication is clear: technical skill alone isn&#x27;t enough anymore. It never was, really, but now the gap is obvious.</p>
<p class="text-muted-foreground leading-7 my-4">The engineers I want to work with aren&#x27;t just good coders. They&#x27;re people who ask &quot;why are we building this?&quot; before they ask &quot;how should we build this?&quot; They&#x27;re the ones who push back on specs that don&#x27;t make sense. Who prototype three different approaches and pick the one that <em class="font-mono font-normal text-primary/80 print:text-primary">feels</em> right, not just the one that&#x27;s technically cleanest.</p>
<p class="text-muted-foreground leading-7 my-4">You build taste by building things. Ship side projects. Use your own products. Pay attention to what annoys you about software you use every day. Develop opinions. Strong ones. Be willing to be wrong, but always have a point of view.</p>
<h2 id="the-new-competitive-advantage" class="font-display text-2xl lg:text-3xl font-bold text-foreground mt-12 mb-6 border-l-2 border-primary pl-4"><a href="#the-new-competitive-advantage" class="heading-permalink text-inherit no-underline">The New Competitive Advantage</a></h2>
<p class="text-muted-foreground leading-7 my-4">The next decade of tech belongs to people with good taste and access to AI tools. Not to the people with the best AI tools and no taste.</p>
<p class="text-muted-foreground leading-7 my-4">Execution is a commodity now. Taste isn&#x27;t. If you&#x27;re wondering what to invest in, invest in your judgment. Talk to users. Ship things. Develop opinions about what good software feels like.</p>
<p class="text-muted-foreground leading-7 my-4">The moat isn&#x27;t your code. It&#x27;s your ability to decide what code should exist.</p>]]></content:encoded>
      <pubDate>Tue, 05 May 2026 12:00:00 GMT</pubDate>
    </item>
    <item>
      <title>The $0 Startup: Why Your Next Company Should Cost Almost Nothing to Build</title>
      <link>https://pratik.pa.tel/blog/the-zero-dollar-startup/</link>
      <guid isPermaLink="true">https://pratik.pa.tel/blog/the-zero-dollar-startup/</guid>
      <description>The most expensive part of building used to be people. AI just zeroed out that line item.</description>
      <dc:creator>Pratik Patel</dc:creator>
      <category>ai</category>
      <category>startups</category>
      <category>building</category>
      <category>economics</category>
      <content:encoded><![CDATA[<p class="text-muted-foreground leading-7 my-4">There&#x27;s a number that used to haunt every first-time founder: the cost of getting to v1.</p>
<p class="text-muted-foreground leading-7 my-4">Five years ago, the math looked something like this. You needed a designer ($8-15K for a freelancer, more for an agency). You needed a developer, or more likely two ($15-30K each, if you were lucky). You needed hosting, a domain, maybe some SaaS subscriptions for analytics, email, and payments. By the time you had something real enough to put in front of customers, you were $50-100K deep — and that was the <em class="font-mono font-normal text-primary/80 print:text-primary">lean</em> version.</p>
<p class="text-muted-foreground leading-7 my-4">That math broke most ideas before they started. Not because the ideas were bad, but because the price of finding out was too high. How many great products never existed because the founder looked at a $75K price tag and said &quot;maybe next year&quot;?</p>
<p class="text-muted-foreground leading-7 my-4">I&#x27;ll tell you: almost all of them.</p>
<h2 id="the-new-math" class="font-display text-2xl lg:text-3xl font-bold text-foreground mt-12 mb-6 border-l-2 border-primary pl-4"><a href="#the-new-math" class="heading-permalink text-inherit no-underline">The New Math</a></h2>
<p class="text-muted-foreground leading-7 my-4">Here&#x27;s what building a product costs in 2026:</p>
<p class="text-muted-foreground leading-7 my-4"><strong class="text-foreground font-bold">Design:</strong> $0. AI design tools generate production-ready interfaces from a text description. Not wireframes. Not mockups. Actual, deployable component code with responsive layouts, consistent design systems, and accessibility built in. I covered this in &quot;No More Ugly Websites&quot; — the design barrier is gone.</p>
<p class="text-muted-foreground leading-7 my-4"><strong class="text-foreground font-bold">Development:</strong> $0 (or close to it). AI agents write backend services, API endpoints, database migrations, and test suites. They don&#x27;t write <em class="font-mono font-normal text-primary/80 print:text-primary">perfect</em> code, but they write code that works, passes tests, and ships. The gap between &quot;AI-generated&quot; and &quot;production-ready&quot; has collapsed to a few hours of review and iteration.</p>
<p class="text-muted-foreground leading-7 my-4"><strong class="text-foreground font-bold">Infrastructure:</strong> $0 to start. Serverless platforms, free-tier databases, and edge hosting mean you can serve thousands of users before you pay a single dollar for infrastructure. The days of provisioning servers before you had customers are over.</p>
<p class="text-muted-foreground leading-7 my-4"><strong class="text-foreground font-bold">Content and copy:</strong> $0. AI writes marketing copy, blog posts, documentation, and email sequences. Again, not perfect — you&#x27;ll want to edit for voice and accuracy — but the first draft is free and usually 80% of the way there.</p>
<p class="text-muted-foreground leading-7 my-4"><strong class="text-foreground font-bold">Total cost to get a real product in front of real users:</strong> Your time, a laptop, and maybe $20/month in API costs.</p>
<p class="text-muted-foreground leading-7 my-4">This isn&#x27;t theoretical. This is how I build. This is how a growing number of founders are building. And the implications are enormous.</p>
<h2 id="why-this-changes-everything" class="font-display text-2xl lg:text-3xl font-bold text-foreground mt-12 mb-6 border-l-2 border-primary pl-4"><a href="#why-this-changes-everything" class="heading-permalink text-inherit no-underline">Why This Changes Everything</a></h2>
<p class="text-muted-foreground leading-7 my-4">The obvious takeaway is &quot;building is cheaper now.&quot; But that understates what&#x27;s actually happening. Cheap building doesn&#x27;t just mean more products get built. It means the <em class="font-mono font-normal text-primary/80 print:text-primary">entire startup model</em> changes.</p>
<h3 id="the-end-of-fundraising-as-a-prerequisite" class="font-display text-xl font-bold text-foreground mt-10 mb-4"><a href="#the-end-of-fundraising-as-a-prerequisite" class="heading-permalink text-inherit no-underline">The End of Fundraising as a Prerequisite</a></h3>
<p class="text-muted-foreground leading-7 my-4">The traditional startup path was: have an idea, raise money to build it, build it, and then find out if anyone wants it. The fundraising step wasn&#x27;t just about money. It was a filter. You had to convince investors that your idea was worth building before you could build it. That filter was imperfect — it selected for charisma and credentials as much as for good ideas.</p>
<p class="text-muted-foreground leading-7 my-4">When building costs nothing, you skip the filter entirely. Build first, then decide if you need money to scale. The question changes from &quot;can I convince someone to fund this?&quot; to &quot;can I convince someone to use this?&quot; That&#x27;s a much better question.</p>
<h3 id="the-death-of-the-mvp" class="font-display text-xl font-bold text-foreground mt-10 mb-4"><a href="#the-death-of-the-mvp" class="heading-permalink text-inherit no-underline">The Death of the MVP</a></h3>
<p class="text-muted-foreground leading-7 my-4">The Minimum Viable Product was a response to high build costs. Strip everything down to the absolute minimum, ship that, and iterate. It was a good framework for a world where every feature cost real money to build.</p>
<p class="text-muted-foreground leading-7 my-4">But when building is nearly free, the concept of &quot;minimum&quot; changes. Your v1 doesn&#x27;t have to be a stripped-down embarrassment. It can be genuinely good. It can have polish, it can have features that delight users, it can have the kind of fit and finish that used to require months of work. The &quot;minimum&quot; in MVP just got a lot more viable.</p>
<h3 id="speed-as-the-only-moat" class="font-display text-xl font-bold text-foreground mt-10 mb-4"><a href="#speed-as-the-only-moat" class="heading-permalink text-inherit no-underline">Speed as the Only Moat</a></h3>
<p class="text-muted-foreground leading-7 my-4">When anyone can build anything for free, the only competitive advantage is speed. Not speed of coding — AI handles that. Speed of <em class="font-mono font-normal text-primary/80 print:text-primary">insight</em>. How fast can you identify what users actually want? How fast can you iterate on their feedback? How fast can you go from &quot;I think this might work&quot; to &quot;I know this works because 500 people are using it&quot;?</p>
<p class="text-muted-foreground leading-7 my-4">The winners in this new landscape won&#x27;t be the best-funded teams. They&#x27;ll be the fastest learners.</p>
<h2 id="what-actually-costs-money-now" class="font-display text-2xl lg:text-3xl font-bold text-foreground mt-12 mb-6 border-l-2 border-primary pl-4"><a href="#what-actually-costs-money-now" class="heading-permalink text-inherit no-underline">What Actually Costs Money Now</a></h2>
<p class="text-muted-foreground leading-7 my-4">If building is free, where does the money go? This is where the startup model gets interesting:</p>
<p class="text-muted-foreground leading-7 my-4"><strong class="text-foreground font-bold">Distribution.</strong> Building the product is the easy part. Getting it in front of the right people still costs time, effort, and sometimes money. SEO, content marketing, paid acquisition, partnerships — the go-to-market machine is where the real investment happens.</p>
<p class="text-muted-foreground leading-7 my-4"><strong class="text-foreground font-bold">Taste.</strong> AI can generate a hundred UI variations in an hour. Knowing which one is right? That&#x27;s human judgment. The scarcest resource in a $0-build world isn&#x27;t engineering talent — it&#x27;s product taste. The ability to look at ten options and pick the one that resonates. That can&#x27;t be automated, and it&#x27;s worth more than ever.</p>
<p class="text-muted-foreground leading-7 my-4"><strong class="text-foreground font-bold">Support and trust.</strong> Customers still want to know there&#x27;s a real human behind the product. Response time, reliability, and genuine care — these are the things that turn a side project into a business. They cost attention, not dollars, but they&#x27;re non-negotiable.</p>
<p class="text-muted-foreground leading-7 my-4"><strong class="text-foreground font-bold">Scale.</strong> Eventually, if your product works, you&#x27;ll outgrow the free tiers. Infrastructure costs kick in. You might need to hire humans for customer support, partnerships, or sales. But by that point, you have revenue, users, and data — which means you can either self-fund or raise money from a position of strength instead of desperation.</p>
<h2 id="the-playbook" class="font-display text-2xl lg:text-3xl font-bold text-foreground mt-12 mb-6 border-l-2 border-primary pl-4"><a href="#the-playbook" class="heading-permalink text-inherit no-underline">The Playbook</a></h2>
<p class="text-muted-foreground leading-7 my-4">If you&#x27;re starting something in 2026, here&#x27;s how I&#x27;d think about money:</p>
<p class="text-muted-foreground leading-7 my-4"><strong class="text-foreground font-bold">Spend $0 on v1.</strong> Use AI agents for design, development, and content. Use free-tier infrastructure. Get something real in front of real people without spending a dollar. This isn&#x27;t cutting corners — this is the new standard.</p>
<p class="text-muted-foreground leading-7 my-4"><strong class="text-foreground font-bold">Spend your time on taste and distribution.</strong> These are the two things that actually matter now. What should your product feel like? Who needs it? How do they find it? If you&#x27;re spending your days writing code instead of answering these questions, you&#x27;re doing it wrong.</p>
<p class="text-muted-foreground leading-7 my-4"><strong class="text-foreground font-bold">Don&#x27;t raise money until you have signal.</strong> Revenue, active users, organic growth — any of these. Going to investors with &quot;I built this for $0 and 200 people are paying for it&quot; is a fundamentally different conversation than &quot;I have an idea and I need $500K to find out if it works.&quot;</p>
<p class="text-muted-foreground leading-7 my-4"><strong class="text-foreground font-bold">Reinvest revenue before outside capital.</strong> When money does start coming in, put it back into the things that got you here: faster iteration, better distribution, deeper understanding of your users. The compounding effects of reinvestment in a $0-cost structure are absurd. Your margins are effectively 100% until you choose to spend.</p>
<h2 id="the-uncomfortable-truth" class="font-display text-2xl lg:text-3xl font-bold text-foreground mt-12 mb-6 border-l-2 border-primary pl-4"><a href="#the-uncomfortable-truth" class="heading-permalink text-inherit no-underline">The Uncomfortable Truth</a></h2>
<p class="text-muted-foreground leading-7 my-4">Here&#x27;s what nobody in the startup world wants to say out loud: <strong class="text-foreground font-bold">most venture-backed startups were always solving an artificial problem.</strong> They raised money to hire engineers to build something that, in many cases, one focused person with the right tools could build in a week. The complexity was the product of the tooling, not the problem.</p>
<p class="text-muted-foreground leading-7 my-4">AI didn&#x27;t just reduce costs. It exposed how much of the startup ecosystem was a tax on building. Accelerators, talent recruiters, office space, team retreats, standup meetings — an entire industry existed to manage the complexity of building with humans. When AI replaces that complexity, the industry around it loses its reason to exist.</p>
<p class="text-muted-foreground leading-7 my-4">I&#x27;m not saying every startup should be one person with a laptop. Some problems genuinely require teams, capital, and coordination. But a lot more problems than we thought can be solved by one person who cares deeply about the outcome and has AI agents to handle the execution.</p>
<h2 id="what-this-means-for-you" class="font-display text-2xl lg:text-3xl font-bold text-foreground mt-12 mb-6 border-l-2 border-primary pl-4"><a href="#what-this-means-for-you" class="heading-permalink text-inherit no-underline">What This Means for You</a></h2>
<p class="text-muted-foreground leading-7 my-4">If you&#x27;ve been waiting for the right time to start something, consider this: the financial excuse is officially dead. You don&#x27;t need savings. You don&#x27;t need investors. You don&#x27;t need a co-founder with a trust fund. You need an idea, a laptop, and the willingness to ship.</p>
<p class="text-muted-foreground leading-7 my-4">The $0 startup isn&#x27;t a gimmick. It&#x27;s the new default. And the founders who figure this out first will build the next decade&#x27;s most interesting companies — not because they raised the most money, but because they needed the least.</p>
<p class="text-muted-foreground leading-7 my-4">Stop budgeting. Start building.</p>]]></content:encoded>
      <pubDate>Tue, 28 Apr 2026 12:00:00 GMT</pubDate>
    </item>
    <item>
      <title>Security Incidents on the Rise: Is Vibe Coding the Common Thread?</title>
      <link>https://pratik.pa.tel/blog/security-incidents-on-the-rise/</link>
      <guid isPermaLink="true">https://pratik.pa.tel/blog/security-incidents-on-the-rise/</guid>
      <description>As AI-generated code floods production, the security bills are starting to come due</description>
      <dc:creator>Pratik Patel</dc:creator>
      <category>security</category>
      <category>ai</category>
      <category>engineering</category>
      <content:encoded><![CDATA[<p class="text-muted-foreground leading-7 my-4">April 2026 has been a brutal month for cybersecurity. Vercel confirmed a breach tied to a compromised AI tool. Drift Protocol lost $285 million in twelve minutes. Kelp DAO was exploited for $292 million, leaving Aave with over $200 million in bad debt. And those are just the headlines.</p>
<p class="text-muted-foreground leading-7 my-4">Something feels different about this wave of incidents. Not just the scale — we&#x27;ve seen big numbers before — but the <strong class="text-foreground font-bold">pattern</strong>. A growing number of breaches trace back to code that was shipped fast, reviewed lightly, and built with AI assistance. The security community is starting to ask an uncomfortable question: is the vibe coding revolution creating a generation of applications that are fundamentally less secure?</p>
<p class="text-muted-foreground leading-7 my-4">I&#x27;ve been thinking about this a lot. As someone who builds with AI tools daily and writes about the experience, I can&#x27;t ignore the data anymore.</p>
<h2 id="the-numbers-are-stark" class="font-display text-2xl lg:text-3xl font-bold text-foreground mt-12 mb-6 border-l-2 border-primary pl-4"><a href="#the-numbers-are-stark" class="heading-permalink text-inherit no-underline">The Numbers Are Stark 📊</a></h2>
<p class="text-muted-foreground leading-7 my-4">Let&#x27;s start with what we know. According to recent research, AI code generators produce vulnerabilities at roughly <strong class="text-foreground font-bold">2x the rate</strong> of human-written code. A Veracode analysis of 4 million code scans found that AI-generated code contained security flaws <strong class="text-foreground font-bold">45% of the time</strong>. The Cloud Security Alliance puts that number even higher — 62% in their study.</p>
<p class="text-muted-foreground leading-7 my-4">And the trend is accelerating. In January 2026, six new CVE entries were directly attributed to AI-generated code. By February, it was fifteen. By March, <strong class="text-foreground font-bold">thirty-five</strong>. Georgia Tech researchers estimate the real number could be five to ten times what&#x27;s currently being detected — roughly 400 to 700 cases across the open-source ecosystem.</p>
<p class="text-muted-foreground leading-7 my-4">Meanwhile, 46% of all code on GitHub is now AI-generated. The vibe coding market hit $4.7 billion in 2026. We&#x27;re shipping more AI-written code into production than ever before, and the vulnerability surface is expanding with it.</p>
<h2 id="anatomy-of-aprils-worst-month" class="font-display text-2xl lg:text-3xl font-bold text-foreground mt-12 mb-6 border-l-2 border-primary pl-4"><a href="#anatomy-of-aprils-worst-month" class="heading-permalink text-inherit no-underline">Anatomy of April&#x27;s Worst Month 💥</a></h2>
<p class="text-muted-foreground leading-7 my-4">Let&#x27;s look at what actually happened this month, because the details matter more than the dollar figures.</p>
<h3 id="vercel-when-your-ai-tool-becomes-the-attack-vector" class="font-display text-xl font-bold text-foreground mt-10 mb-4"><a href="#vercel-when-your-ai-tool-becomes-the-attack-vector" class="heading-permalink text-inherit no-underline">Vercel: When Your AI Tool Becomes the Attack Vector</a></h3>
<p class="text-muted-foreground leading-7 my-4">Vercel&#x27;s breach didn&#x27;t come from a zero-day or a sophisticated protocol exploit. It came from a <strong class="text-foreground font-bold">third-party AI tool</strong>. An employee signed up for an AI productivity suite called Context.ai using their Vercel enterprise account and granted it &quot;Allow All&quot; OAuth permissions. When Context.ai was compromised, the attackers walked right into Vercel&#x27;s Google Workspace through that OAuth token.</p>
<p class="text-muted-foreground leading-7 my-4">This is the new attack surface that nobody&#x27;s talking about enough. Engineers are adopting AI tools at breakneck speed — browser extensions, coding assistants, AI office suites — and each one is a potential entry point. The Vercel breach wasn&#x27;t about bad code. It was about the <strong class="text-foreground font-bold">toolchain sprawl</strong> that comes with an AI-everything culture.</p>
<h3 id="drift-protocol-285-million-in-twelve-minutes" class="font-display text-xl font-bold text-foreground mt-10 mb-4"><a href="#drift-protocol-285-million-in-twelve-minutes" class="heading-permalink text-inherit no-underline">Drift Protocol: $285 Million in Twelve Minutes</a></h3>
<p class="text-muted-foreground leading-7 my-4">On April 1st — yes, April Fool&#x27;s Day — attackers drained $285 million from Drift Protocol on Solana. The method was audacious: they manufactured a completely fictitious token called CarbonVote, seeded it with a few thousand dollars in fake liquidity, and Drift&#x27;s oracles treated it as legitimate collateral worth hundreds of millions.</p>
<p class="text-muted-foreground leading-7 my-4">The staging began weeks earlier. On-chain forensics traced the initial funding to a Tornado Cash withdrawal on March 11th, with movement patterns consistent with DPRK-attributed operations. The attack executed in roughly twelve minutes, with most stolen funds bridged to Ethereum within hours.</p>
<p class="text-muted-foreground leading-7 my-4">The deeper question: how did a fabricated token bypass validation? The answer likely involves the same pattern we see across the industry — systems built for speed, with security assumptions that went unquestioned because the code &quot;worked&quot; and the tests passed.</p>
<h3 id="kelp-dao-and-aave-the-292-million-cascade" class="font-display text-xl font-bold text-foreground mt-10 mb-4"><a href="#kelp-dao-and-aave-the-292-million-cascade" class="heading-permalink text-inherit no-underline">Kelp DAO and Aave: The $292 Million Cascade</a></h3>
<p class="text-muted-foreground leading-7 my-4">On April 18th, an attacker exploited Kelp DAO&#x27;s bridge infrastructure to release 116,500 unbacked rsETH tokens — about 18% of the token&#x27;s circulating supply. These phantom tokens were immediately deposited into Aave as collateral to borrow real assets.</p>
<p class="text-muted-foreground leading-7 my-4">The cascade was devastating. Aave&#x27;s total value locked plunged by $6.6 billion. The AAVE token dropped 16%. Whales pulled more than $6 billion in 24 hours, pushing major lending pools to 100% utilization and effectively trapping remaining depositors.</p>
<p class="text-muted-foreground leading-7 my-4">Again, Aave&#x27;s own contracts weren&#x27;t compromised. The vulnerability existed in the <strong class="text-foreground font-bold">integration layer</strong> — the assumptions about what constitutes valid collateral. These are exactly the kinds of assumptions that get lost when code is generated fast and reviewed at the surface level.</p>
<h2 id="the-vibe-coding-problem" class="font-display text-2xl lg:text-3xl font-bold text-foreground mt-12 mb-6 border-l-2 border-primary pl-4"><a href="#the-vibe-coding-problem" class="heading-permalink text-inherit no-underline">The Vibe Coding Problem 🎯</a></h2>
<p class="text-muted-foreground leading-7 my-4">Here&#x27;s the uncomfortable truth about vibe coding: it fundamentally breaks traditional application security models.</p>
<p class="text-muted-foreground leading-7 my-4">The term &quot;vibe coding,&quot; coined by Andrej Karpathy, describes a development approach where you describe what you want and let AI generate the implementation. The philosophy prioritizes speed and rapid iteration. Ship fast, fix later. The vibes are good. The code compiles. The tests pass. Deploy.</p>
<p class="text-muted-foreground leading-7 my-4">But security isn&#x27;t about whether code compiles. It&#x27;s about whether code <strong class="text-foreground font-bold">fails safely</strong> under adversarial conditions. And that requires a kind of paranoid, defensive thinking that AI code generators simply don&#x27;t exhibit by default.</p>
<p class="text-muted-foreground leading-7 my-4">The most common vulnerabilities in vibe-coded applications are telling:</p>
<ul class="space-y-2 my-6 ml-4">
<li class="flex gap-3 text-muted-foreground leading-7"><span class="text-primary shrink-0 mt-1.5">▸</span><span class="min-w-0"><strong class="text-foreground font-bold">Disabled row-level security</strong> — found in roughly 70% of apps built with AI-first platforms</span></li>
<li class="flex gap-3 text-muted-foreground leading-7"><span class="text-primary shrink-0 mt-1.5">▸</span><span class="min-w-0"><strong class="text-foreground font-bold">Leaked secrets</strong> — API keys and credentials hardcoded or exposed in client bundles</span></li>
<li class="flex gap-3 text-muted-foreground leading-7"><span class="text-primary shrink-0 mt-1.5">▸</span><span class="min-w-0"><strong class="text-foreground font-bold">Missing webhook verification</strong> — endpoints that accept any payload without signature checks</span></li>
<li class="flex gap-3 text-muted-foreground leading-7"><span class="text-primary shrink-0 mt-1.5">▸</span><span class="min-w-0"><strong class="text-foreground font-bold">Absent authorization checks</strong> — the classic CWE-862, where endpoints work but don&#x27;t verify who&#x27;s asking</span></li>
</ul>
<p class="text-muted-foreground leading-7 my-4">These aren&#x27;t exotic attack vectors. They&#x27;re <strong class="text-foreground font-bold">Security 101 failures</strong> — the kind that a human developer with a few years of experience would catch instinctively, but that an AI code generator will happily produce because the code technically fulfills the functional requirement.</p>
<h2 id="the-moltbook-cautionary-tale" class="font-display text-2xl lg:text-3xl font-bold text-foreground mt-12 mb-6 border-l-2 border-primary pl-4"><a href="#the-moltbook-cautionary-tale" class="heading-permalink text-inherit no-underline">The Moltbook Cautionary Tale 🚨</a></h2>
<p class="text-muted-foreground leading-7 my-4">Perhaps no incident better illustrates the risk than Moltbook, an AI social network whose founder publicly stated he &quot;didn&#x27;t write a single line of code.&quot; The entire application was vibe-coded.</p>
<p class="text-muted-foreground leading-7 my-4">Within three days of launch, security researchers discovered the application had exposed its <strong class="text-foreground font-bold">entire production database</strong> — 1.5 million API authentication tokens, 35,000 email addresses, and private messages. All publicly accessible. No authentication required.</p>
<p class="text-muted-foreground leading-7 my-4">Moltbook is what happens when the entire security posture of an application depends on an AI code generator that optimizes for functionality, not defense. The app worked. Users could sign up, post, and interact. The vibes were immaculate. The security was nonexistent.</p>
<h2 id="speed-vs-insight-the-real-trade-off" class="font-display text-2xl lg:text-3xl font-bold text-foreground mt-12 mb-6 border-l-2 border-primary pl-4"><a href="#speed-vs-insight-the-real-trade-off" class="heading-permalink text-inherit no-underline">Speed vs. Insight: The Real Trade-Off ⚖️</a></h2>
<p class="text-muted-foreground leading-7 my-4">I want to be clear: I&#x27;m not anti-AI coding. I use AI tools every day and I&#x27;ve written about how they&#x27;ve made me more productive. The issue isn&#x27;t AI assistance itself — it&#x27;s the <strong class="text-foreground font-bold">absence of human security judgment</strong> in the loop.</p>
<p class="text-muted-foreground leading-7 my-4">When an experienced engineer writes code, they bring accumulated knowledge about failure modes. They know that an API endpoint needs rate limiting because they&#x27;ve seen what happens without it. They add input validation not because the spec says to, but because they&#x27;ve been burned by SQL injection before. They check authorization on every endpoint because they understand that &quot;the frontend handles it&quot; is not a security strategy.</p>
<p class="text-muted-foreground leading-7 my-4">AI code generators don&#x27;t have this scar tissue. They produce code that matches the pattern of what was requested, but they don&#x27;t anticipate how that code might be <strong class="text-foreground font-bold">abused</strong>. And when developers accept that code without applying their own security judgment — when they vibe with it instead of scrutinizing it — the defensive layer disappears entirely.</p>
<p class="text-muted-foreground leading-7 my-4">The trade-off isn&#x27;t speed vs. security. It&#x27;s <strong class="text-foreground font-bold">speed vs. insight</strong>. And right now, the industry is overwhelmingly choosing speed.</p>
<h2 id="what-needs-to-change" class="font-display text-2xl lg:text-3xl font-bold text-foreground mt-12 mb-6 border-l-2 border-primary pl-4"><a href="#what-needs-to-change" class="heading-permalink text-inherit no-underline">What Needs to Change 🔧</a></h2>
<p class="text-muted-foreground leading-7 my-4">The answer isn&#x27;t to stop using AI coding tools. That ship has sailed — 46% of GitHub is already AI-generated. The answer is to build security back into the workflow in ways that work alongside AI-assisted development.</p>
<p class="text-muted-foreground leading-7 my-4"><strong class="text-foreground font-bold">First, treat AI-generated code as untrusted input.</strong> Every line should pass through the same scrutiny you&#x27;d give to a dependency from an unknown source. Static analysis, secret scanning, and authorization audits should run automatically on every commit, not as a manual step that gets skipped when velocity is the priority.</p>
<p class="text-muted-foreground leading-7 my-4"><strong class="text-foreground font-bold">Second, keep a human in the security loop.</strong> Code review for AI-generated code should specifically focus on security assumptions — authentication, authorization, input validation, error handling, data exposure. These are the areas where AI consistently underperforms.</p>
<p class="text-muted-foreground leading-7 my-4"><strong class="text-foreground font-bold">Third, audit your AI toolchain.</strong> The Vercel breach wasn&#x27;t about code — it was about the tools around the code. Every AI tool your team adopts is a potential attack surface. OAuth permissions should be reviewed. Third-party integrations should be inventoried. The convenience of &quot;Allow All&quot; permissions is the enemy of security.</p>
<p class="text-muted-foreground leading-7 my-4"><strong class="text-foreground font-bold">Fourth, invest in security education that acknowledges the AI reality.</strong> Developers need to understand not just how to use AI tools, but where those tools have systematic blind spots. Security training needs to evolve from &quot;how to write secure code&quot; to &quot;how to verify that AI-generated code is secure.&quot;</p>
<h2 id="the-stakes-are-only-getting-higher" class="font-display text-2xl lg:text-3xl font-bold text-foreground mt-12 mb-6 border-l-2 border-primary pl-4"><a href="#the-stakes-are-only-getting-higher" class="heading-permalink text-inherit no-underline">The Stakes Are Only Getting Higher 📈</a></h2>
<p class="text-muted-foreground leading-7 my-4">We&#x27;re at an inflection point. The volume of AI-generated code in production is growing exponentially. The sophistication of attackers — including nation-state actors like the DPRK group behind the Drift exploit — isn&#x27;t slowing down. And the gap between &quot;code that works&quot; and &quot;code that&#x27;s secure&quot; is widening as AI makes it easier than ever to ship the former without achieving the latter.</p>
<p class="text-muted-foreground leading-7 my-4">April 2026 should be a wake-up call. Not because AI coding tools are inherently dangerous, but because we&#x27;re adopting them faster than we&#x27;re adapting our security practices to account for their limitations. The vibes might be good, but the threat model doesn&#x27;t care about vibes.</p>
<p class="text-muted-foreground leading-7 my-4">The question isn&#x27;t whether AI-assisted development will continue — it will. The question is whether we&#x27;ll build the security culture and tooling to match the pace of adoption. Right now, the scoreboard says we&#x27;re losing.</p>]]></content:encoded>
      <pubDate>Sat, 25 Apr 2026 12:00:00 GMT</pubDate>
    </item>
    <item>
      <title>The 10x Engineer Is a Myth</title>
      <link>https://pratik.pa.tel/blog/10x-engineer-myth/</link>
      <guid isPermaLink="true">https://pratik.pa.tel/blog/10x-engineer-myth/</guid>
      <description>I&apos;ve worked with hundreds of engineers and never met a 10x engineer. I&apos;ve met plenty with 10x impact, and they do something completely different.</description>
      <dc:creator>Pratik Patel</dc:creator>
      <category>engineering</category>
      <category>leadership</category>
      <category>career</category>
      <category>teams</category>
      <content:encoded><![CDATA[<p class="text-muted-foreground leading-7 my-4">The &quot;10x engineer&quot; is one of tech&#x27;s most persistent myths. You know the archetype: the lone genius who cranks out code at superhuman speed, headphones on, hooded up, fueled by caffeine and pure talent.</p>
<p class="text-muted-foreground leading-7 my-4">I&#x27;ve worked with hundreds of engineers across startups and big tech. I&#x27;ve never met one.</p>
<p class="text-muted-foreground leading-7 my-4">But I&#x27;ve met plenty of people with 10x impact. And they do something completely different.</p>
<h2 id="the-myth-more-code-more-value" class="font-display text-2xl lg:text-3xl font-bold text-foreground mt-12 mb-6 border-l-2 border-primary pl-4"><a href="#the-myth-more-code-more-value" class="heading-permalink text-inherit no-underline">The Myth: More Code = More Value</a></h2>
<p class="text-muted-foreground leading-7 my-4">The 10x engineer myth is built on a flawed assumption: that an engineer&#x27;s value is measured by their individual output.</p>
<p class="text-muted-foreground leading-7 my-4">Write more code. Ship more features. Close more tickets. If one person does 10x the tickets, they&#x27;re 10x the engineer. Simple math.</p>
<p class="text-muted-foreground leading-7 my-4">Except code isn&#x27;t an asset. Code is a liability. Every line you write is a line someone has to maintain, debug, and eventually rewrite. More code doesn&#x27;t mean more value. It often means more complexity, more bugs, and more surface area for things to go wrong.</p>
<p class="text-muted-foreground leading-7 my-4">The engineer who writes 10x the code might also be creating 10x the maintenance burden. That&#x27;s not a 10x engineer. That&#x27;s a 10x cost center.</p>
<h2 id="what-10x-impact-actually-looks-like" class="font-display text-2xl lg:text-3xl font-bold text-foreground mt-12 mb-6 border-l-2 border-primary pl-4"><a href="#what-10x-impact-actually-looks-like" class="heading-permalink text-inherit no-underline">What 10x Impact Actually Looks Like</a></h2>
<p class="text-muted-foreground leading-7 my-4">The people I&#x27;ve seen with genuine 10x impact don&#x27;t produce 10x the output. They multiply everyone else&#x27;s.</p>
<h3 id="code-reviews-that-teach" class="font-display text-xl font-bold text-foreground mt-10 mb-4"><a href="#code-reviews-that-teach" class="heading-permalink text-inherit no-underline">Code reviews that teach</a></h3>
<p class="text-muted-foreground leading-7 my-4">There&#x27;s a difference between a code review that says &quot;approved&quot; and one that says &quot;this works, but pattern X would handle the edge case on line 47 better&quot; with a link to a post explaining the tradeoff.</p>
<p class="text-muted-foreground leading-7 my-4">The first review gets the PR merged. The second review gets the PR merged <em class="font-mono font-normal text-primary/80 print:text-primary">and</em> makes the author a better engineer. Multiply that across hundreds of reviews a year, and you&#x27;ve raised the quality of every PR the team ships.</p>
<p class="text-muted-foreground leading-7 my-4">The engineers with 10x impact treat code review as mentoring, not gatekeeping.</p>
<h3 id="documentation-that-saves-hours" class="font-display text-xl font-bold text-foreground mt-10 mb-4"><a href="#documentation-that-saves-hours" class="heading-permalink text-inherit no-underline">Documentation that saves hours</a></h3>
<p class="text-muted-foreground leading-7 my-4">I&#x27;ve seen a single well-written architecture doc save an entire team weeks of confusion. A runbook that prevents 3 AM page escalations. An onboarding guide that cuts ramp-up time from two months to two weeks.</p>
<p class="text-muted-foreground leading-7 my-4">Nobody gets promoted for writing docs. But the engineers who write them anyway, who explain how something works so everyone else doesn&#x27;t have to reverse-engineer it, have an outsized impact that never shows up in ticket counts.</p>
<h3 id="mentoring-that-compounds" class="font-display text-xl font-bold text-foreground mt-10 mb-4"><a href="#mentoring-that-compounds" class="heading-permalink text-inherit no-underline">Mentoring that compounds</a></h3>
<p class="text-muted-foreground leading-7 my-4">The math is simple: if you make 5 people 20% better at their jobs, that&#x27;s the equivalent of adding a full engineer to the team. If you do that consistently over a year, you&#x27;ve created more value than any individual contributor could.</p>
<p class="text-muted-foreground leading-7 my-4">The best multipliers I&#x27;ve worked with do this naturally. They pair-program when someone&#x27;s stuck. They explain the &quot;why&quot; behind technical decisions, not just the &quot;what.&quot; They create an environment where asking questions is easier than guessing.</p>
<p class="text-muted-foreground leading-7 my-4">This compounds. The person you mentored mentors someone else. The patterns you taught become team standards. The documentation culture you started outlasts your tenure.</p>
<h3 id="the-question-that-saves-a-month" class="font-display text-xl font-bold text-foreground mt-10 mb-4"><a href="#the-question-that-saves-a-month" class="heading-permalink text-inherit no-underline">The question that saves a month</a></h3>
<p class="text-muted-foreground leading-7 my-4">Every engineering team has had this meeting: someone is 30 minutes into presenting a plan, and one person raises their hand and asks a simple question that reveals the entire approach is solving the wrong problem.</p>
<p class="text-muted-foreground leading-7 my-4">That question, the one that redirects a month of misguided work, is worth more than any amount of code. But it requires two things most &quot;10x coders&quot; don&#x27;t have: deep understanding of the business context, and the courage to speak up.</p>
<h2 id="the-multiplier-framework" class="font-display text-2xl lg:text-3xl font-bold text-foreground mt-12 mb-6 border-l-2 border-primary pl-4"><a href="#the-multiplier-framework" class="heading-permalink text-inherit no-underline">The Multiplier Framework</a></h2>
<p class="text-muted-foreground leading-7 my-4">A simple way to think about it:</p>
<p class="text-muted-foreground leading-7 my-4"><strong class="text-foreground font-bold">Individual output</strong> = what you ship yourself.</p>
<p class="text-muted-foreground leading-7 my-4"><strong class="text-foreground font-bold">Multiplier effect</strong> = how much better you make everyone around you.</p>
<p class="text-muted-foreground leading-7 my-4">A team of 10 engineers where one person has a 2x multiplier effect is more productive than a team of 10 where one person writes 2x the code. Because the multiplier raises everyone. The individual just raises themselves.</p>
<p class="text-muted-foreground leading-7 my-4">Most engineering cultures reward the individual. Promotions go to the person who shipped the Big Feature. Performance reviews measure tickets closed, lines written, projects delivered.</p>
<p class="text-muted-foreground leading-7 my-4">But the people who quietly make everyone around them more effective? They&#x27;re the actual force multipliers. And they&#x27;re chronically under-recognized.</p>
<h2 id="how-to-become-a-multiplier" class="font-display text-2xl lg:text-3xl font-bold text-foreground mt-12 mb-6 border-l-2 border-primary pl-4"><a href="#how-to-become-a-multiplier" class="heading-permalink text-inherit no-underline">How to Become a Multiplier</a></h2>
<p class="text-muted-foreground leading-7 my-4">You don&#x27;t need to be a senior staff engineer to start. Three things you can do this week:</p>
<h3 id="1-make-your-code-reviews-useful" class="font-display text-xl font-bold text-foreground mt-10 mb-4"><a href="#1-make-your-code-reviews-useful" class="heading-permalink text-inherit no-underline">1. Make your code reviews useful</a></h3>
<p class="text-muted-foreground leading-7 my-4">Stop rubber-stamping. When you review code, leave at least one comment that teaches something: a better pattern, a potential edge case, a relevant resource. Takes 5 extra minutes per review and compounds over months.</p>
<h3 id="2-write-the-doc-nobody-asked-for" class="font-display text-xl font-bold text-foreground mt-10 mb-4"><a href="#2-write-the-doc-nobody-asked-for" class="heading-permalink text-inherit no-underline">2. Write the doc nobody asked for</a></h3>
<p class="text-muted-foreground leading-7 my-4">You know that thing on your team that everyone asks about and nobody writes down? Write it down. It doesn&#x27;t have to be perfect. A mediocre doc that exists is infinitely more useful than a perfect doc that doesn&#x27;t.</p>
<h3 id="3-share-context-not-just-answers" class="font-display text-xl font-bold text-foreground mt-10 mb-4"><a href="#3-share-context-not-just-answers" class="heading-permalink text-inherit no-underline">3. Share context, not just answers</a></h3>
<p class="text-muted-foreground leading-7 my-4">When someone asks you a question, don&#x27;t just give the answer. Explain how you found it. &quot;I checked the logs in CloudWatch, filtered by this query, and found the error here&quot; teaches them to fish. &quot;The bug is on line 47&quot; gives them a fish.</p>
<h2 id="the-takeaway" class="font-display text-2xl lg:text-3xl font-bold text-foreground mt-12 mb-6 border-l-2 border-primary pl-4"><a href="#the-takeaway" class="heading-permalink text-inherit no-underline">The Takeaway</a></h2>
<p class="text-muted-foreground leading-7 my-4">Stop trying to be the fastest coder in the room. Be the person your team can&#x27;t function without, not because you hoard knowledge or write all the critical code, but because everyone around you does better work when you&#x27;re there.</p>
<p class="text-muted-foreground leading-7 my-4">That&#x27;s not a myth. That&#x27;s 10x impact.</p>]]></content:encoded>
      <pubDate>Mon, 20 Apr 2026 12:00:00 GMT</pubDate>
    </item>
    <item>
      <title>No More Ugly Websites: AI Killed Every Excuse for Bad Design</title>
      <link>https://pratik.pa.tel/blog/no-more-ugly-websites/</link>
      <guid isPermaLink="true">https://pratik.pa.tel/blog/no-more-ugly-websites/</guid>
      <description>Billion-dollar companies still ship interfaces from 2003. The tools to fix that now cost $0 and take minutes.</description>
      <dc:creator>Pratik Patel</dc:creator>
      <category>ai</category>
      <category>design</category>
      <category>web</category>
      <content:encoded><![CDATA[<p class="text-muted-foreground leading-7 my-4">Open Craigslist in 2026 and you&#x27;re looking at the same HTML table layout from 1995. Try Namecheap&#x27;s dashboard and you&#x27;re fighting a wall of cluttered panels that haven&#x27;t aged well since the mid-2010s. Visit berkshirehathaway.com, the website of an $800 billion company, and you&#x27;ll find a single page of unstyled hyperlinks on a white background. It looks like a professor&#x27;s personal homepage from the Geocities era.</p>
<p class="text-muted-foreground leading-7 my-4">These aren&#x27;t edge cases. Ugly, dated web design is <em class="font-mono font-normal text-primary/80 print:text-primary">everywhere</em>, and it costs real money. <strong class="text-foreground font-bold">75% of users judge a company&#x27;s credibility based on its website design</strong> (Stanford Web Credibility Project). <strong class="text-foreground font-bold">38% of visitors will stop engaging entirely if the layout is unattractive</strong> (Adobe). First impressions are 94% design-related, and they form in 0.05 seconds.</p>
<p class="text-muted-foreground leading-7 my-4">It&#x27;s 2026. None of these sites need to look like this anymore.</p>
<h2 id="the-old-excuse-died-this-year" class="font-display text-2xl lg:text-3xl font-bold text-foreground mt-12 mb-6 border-l-2 border-primary pl-4"><a href="#the-old-excuse-died-this-year" class="heading-permalink text-inherit no-underline">The Old Excuse Died This Year</a></h2>
<p class="text-muted-foreground leading-7 my-4">For decades, the excuse was legitimate. A proper redesign meant hiring a designer, a frontend engineer (or three), and committing to months of work. For Craigslist, which earns hundreds of millions from classified ads, the calculus was simple: the ugly design <em class="font-mono font-normal text-primary/80 print:text-primary">works</em>, and a redesign is expensive, risky, and probably not worth the ROI.</p>
<p class="text-muted-foreground leading-7 my-4">That calculus just broke.</p>
<p class="text-muted-foreground leading-7 my-4">AI design tools have collapsed the cost and timeline of a website redesign from months and six figures to <strong class="text-foreground font-bold">minutes and zero dollars</strong>. Not hype. The new baseline.</p>
<h2 id="whats-actually-out-there-now" class="font-display text-2xl lg:text-3xl font-bold text-foreground mt-12 mb-6 border-l-2 border-primary pl-4"><a href="#whats-actually-out-there-now" class="heading-permalink text-inherit no-underline">What&#x27;s Actually Out There Now</a></h2>
<p class="text-muted-foreground leading-7 my-4">The AI design tool space in 2026 is wild. I&#x27;ve been watching a few closely:</p>
<p class="text-muted-foreground leading-7 my-4"><strong class="text-foreground font-bold">v0.dev</strong> (by Vercel) lets you describe a UI in plain English and get production-ready React components back. You can screenshot an ugly site, paste it in, and ask for a modern version. It outputs clean code with Tailwind CSS and proper component architecture.</p>
<p class="text-muted-foreground leading-7 my-4"><strong class="text-foreground font-bold">Bolt.new</strong> gives you a full-stack web app from a text prompt, running entirely in the browser. Describe what you want, and it scaffolds a modern app with your choice of framework. No local setup, no deployment pipeline. Just an idea to a live site.</p>
<p class="text-muted-foreground leading-7 my-4"><strong class="text-foreground font-bold">Lovable</strong> (formerly GPT Engineer) takes natural language descriptions and generates full-stack applications with polished design out of the box. It&#x27;s aimed squarely at people who have a vision but not a design team.</p>
<p class="text-muted-foreground leading-7 my-4"><strong class="text-foreground font-bold">Uizard</strong> can take a <em class="font-mono font-normal text-primary/80 print:text-primary">screenshot</em> of your existing legacy site and convert it into an editable, modernized mockup. The &quot;before and after&quot; workflow is built right in.</p>
<p class="text-muted-foreground leading-7 my-4">Then there&#x27;s <strong class="text-foreground font-bold">Framer AI</strong> generating publishable websites from descriptions, <strong class="text-foreground font-bold">Figma</strong> with AI plugins (Musho, Relume) that generate complete page designs from prompts, and <strong class="text-foreground font-bold">Wix</strong> and <strong class="text-foreground font-bold">Hostinger</strong> with AI builders that create responsive sites from a sentence about your business.</p>
<p class="text-muted-foreground leading-7 my-4">The tools aren&#x27;t making design faster. They&#x27;re making <strong class="text-foreground font-bold">the absence of design a deliberate choice</strong>.</p>
<h2 id="the-ugly-works-myth" class="font-display text-2xl lg:text-3xl font-bold text-foreground mt-12 mb-6 border-l-2 border-primary pl-4"><a href="#the-ugly-works-myth" class="heading-permalink text-inherit no-underline">The &quot;Ugly Works&quot; Myth</a></h2>
<p class="text-muted-foreground leading-7 my-4">Defenders of dated design often argue that sites like Craigslist and Hacker News prove ugly works. There&#x27;s a kernel of truth there. Both sites have massive, loyal user bases that value function over form.</p>
<p class="text-muted-foreground leading-7 my-4">But this argument confuses <strong class="text-foreground font-bold">tolerance</strong> with <strong class="text-foreground font-bold">preference</strong>. Users tolerate Craigslist&#x27;s design because the utility is irreplaceable, not because the interface is good. Craigslist succeeds <em class="font-mono font-normal text-primary/80 print:text-primary">despite</em> its design, not because of it. And the data on what happens when you actually improve UX is unambiguous:</p>
<ul class="space-y-2 my-6 ml-4">
<li class="flex gap-3 text-muted-foreground leading-7"><span class="text-primary shrink-0 mt-1.5">▸</span><span class="min-w-0">A well-designed UI can raise conversion rates by <strong class="text-foreground font-bold">up to 200%</strong> (Forrester Research)</span></li>
<li class="flex gap-3 text-muted-foreground leading-7"><span class="text-primary shrink-0 mt-1.5">▸</span><span class="min-w-0">Better UX design yields conversion rates <strong class="text-foreground font-bold">up to 400%</strong> higher</span></li>
<li class="flex gap-3 text-muted-foreground leading-7"><span class="text-primary shrink-0 mt-1.5">▸</span><span class="min-w-0">Every $1 invested in UX returns roughly <strong class="text-foreground font-bold">$100</strong> in value</span></li>
</ul>
<p class="text-muted-foreground leading-7 my-4">The &quot;ugly works&quot; argument is really saying: &quot;We&#x27;re leaving money on the table, but we&#x27;re making enough that we don&#x27;t care.&quot; That&#x27;s a valid business decision. But it&#x27;s not an argument that the design is <em class="font-mono font-normal text-primary/80 print:text-primary">good</em>, and it&#x27;s definitely not an argument that it&#x27;s <em class="font-mono font-normal text-primary/80 print:text-primary">necessary</em> anymore.</p>
<h2 id="now-anyone-can-do-it" class="font-display text-2xl lg:text-3xl font-bold text-foreground mt-12 mb-6 border-l-2 border-primary pl-4"><a href="#now-anyone-can-do-it" class="heading-permalink text-inherit no-underline">Now Anyone Can Do It</a></h2>
<p class="text-muted-foreground leading-7 my-4">The bigger shift is <strong class="text-foreground font-bold">who can use these tools</strong>. Modernizing a website used to be an engineering task. You needed someone who knew HTML, CSS, JavaScript, responsive design, accessibility, and deployment.</p>
<p class="text-muted-foreground leading-7 my-4">Now, a marketing manager can redesign a landing page during lunch. A founder can go from &quot;our site looks outdated&quot; to &quot;here&#x27;s the new version&quot; in an afternoon. A small business owner who&#x27;s been embarrassed by their website for years can finally fix it without hiring an agency.</p>
<p class="text-muted-foreground leading-7 my-4">Squarespace and Wix started this shift with templates, but AI tools finish it by removing the template constraint entirely. You&#x27;re not picking from a menu anymore. You describe what you want and get something custom.</p>
<h2 id="vibe-coding-and-what-comes-next" class="font-display text-2xl lg:text-3xl font-bold text-foreground mt-12 mb-6 border-l-2 border-primary pl-4"><a href="#vibe-coding-and-what-comes-next" class="heading-permalink text-inherit no-underline">Vibe Coding and What Comes Next</a></h2>
<p class="text-muted-foreground leading-7 my-4">Some people call this broader movement <strong class="text-foreground font-bold">vibe coding</strong>: describe the <em class="font-mono font-normal text-primary/80 print:text-primary">vibe</em> of what you want and let AI figure out the implementation. Not about writing code. About expressing intent.</p>
<p class="text-muted-foreground leading-7 my-4">At <a href="https://poof.new" class="text-primary underline underline-offset-2 decoration-primary/40 print:decoration-primary hover:decoration-primary transition-colors">Tarobase (poof.new)</a>, where I work as Chief Architect, we&#x27;re building tools around this exact thesis. The web should be a place where ideas become reality without requiring a computer science degree. When the barrier between imagination and implementation drops to near-zero, the entire economics of web development changes.</p>
<p class="text-muted-foreground leading-7 my-4">The ugly website problem was never about technology. It&#x27;s <strong class="text-foreground font-bold">inertia</strong>. Companies kept dated designs because redesigning was hard. That excuse just expired.</p>
<h2 id="so-why-do-sites-still-look-terrible" class="font-display text-2xl lg:text-3xl font-bold text-foreground mt-12 mb-6 border-l-2 border-primary pl-4"><a href="#so-why-do-sites-still-look-terrible" class="heading-permalink text-inherit no-underline">So Why Do Sites Still Look Terrible?</a></h2>
<p class="text-muted-foreground leading-7 my-4">If AI tools can already redesign a website in minutes, why hasn&#x27;t it happened?</p>
<p class="text-muted-foreground leading-7 my-4">The answer is organizational, not technical. Large companies have entrenched codebases, bureaucratic approval processes, and teams optimized for maintaining the status quo. Craigslist doesn&#x27;t look the way it does because no one knows how to make it better. It looks that way because no one with the authority to change it has prioritized doing so.</p>
<p class="text-muted-foreground leading-7 my-4">But that&#x27;s changing. As AI design tools go mainstream, the social pressure mounts. When your competitor can ship a gorgeous, modern experience with a fraction of the effort, &quot;our site has always looked like this&quot; stops being defensible.</p>
<p class="text-muted-foreground leading-7 my-4">The companies that move first will set new baselines for their industries. The ones that don&#x27;t will look like relics. Not because they lack resources, but because they lack urgency.</p>
<h2 id="the-excuse-is-dead" class="font-display text-2xl lg:text-3xl font-bold text-foreground mt-12 mb-6 border-l-2 border-primary pl-4"><a href="#the-excuse-is-dead" class="heading-permalink text-inherit no-underline">The Excuse Is Dead</a></h2>
<p class="text-muted-foreground leading-7 my-4">Good-looking, functional web design doesn&#x27;t require a team of specialists anymore. Text prompt. Five minutes.</p>
<p class="text-muted-foreground leading-7 my-4">The last excuse for ugly websites is dead. How long will companies keep pretending it&#x27;s still alive?</p>]]></content:encoded>
      <pubDate>Thu, 16 Apr 2026 12:00:00 GMT</pubDate>
    </item>
    <item>
      <title>Ship It Yourself: Why the Best Time to Build Is Right Now</title>
      <link>https://pratik.pa.tel/blog/ship-it-yourself/</link>
      <guid isPermaLink="true">https://pratik.pa.tel/blog/ship-it-yourself/</guid>
      <description>AI didn&apos;t just lower the barrier to entry. It removed it entirely.</description>
      <dc:creator>Pratik Patel</dc:creator>
      <category>ai</category>
      <category>startups</category>
      <category>building</category>
      <content:encoded><![CDATA[<p class="text-muted-foreground leading-7 my-4">Five years ago, if you wanted to launch a product, you needed a team. A designer to make it look right. A frontend engineer to build the interface. A backend engineer to wire the logic. A DevOps person to deploy it. Maybe a copywriter to make the landing page not sound like it was written by a robot. The minimum viable <em class="font-mono font-normal text-primary/80 print:text-primary">team</em> was five people before you even had a minimum viable <em class="font-mono font-normal text-primary/80 print:text-primary">product</em>.</p>
<p class="text-muted-foreground leading-7 my-4">That world is gone.</p>
<p class="text-muted-foreground leading-7 my-4">In 2026, a single person with a laptop and a clear idea can ship a product that looks, works, and scales like it was built by a funded startup. I know this because I&#x27;m living it. And if you&#x27;re sitting on an idea right now, waiting for the &quot;right time&quot; or the &quot;right team,&quot; I&#x27;m here to tell you: <strong class="text-foreground font-bold">the right time is now, and the right team is already on your machine</strong>.</p>
<h2 id="the-excuses-are-dead" class="font-display text-2xl lg:text-3xl font-bold text-foreground mt-12 mb-6 border-l-2 border-primary pl-4"><a href="#the-excuses-are-dead" class="heading-permalink text-inherit no-underline">The Excuses Are Dead</a></h2>
<p class="text-muted-foreground leading-7 my-4">Let&#x27;s run through the greatest hits of reasons people don&#x27;t build:</p>
<p class="text-muted-foreground leading-7 my-4"><strong class="text-foreground font-bold">&quot;I can&#x27;t design.&quot;</strong> AI design tools generate production-ready interfaces from a text description. Entire component libraries, color systems, and responsive layouts — built in minutes. I wrote about this in my last post: there is genuinely no excuse for an ugly website anymore. That same logic applies to your product.</p>
<p class="text-muted-foreground leading-7 my-4"><strong class="text-foreground font-bold">&quot;I can&#x27;t code the backend.&quot;</strong> AI agents write, test, and deploy backend services. Describe your data model and business logic, and an agent will scaffold the API, write the tests, handle the migrations, and open a PR for your review. You&#x27;re the architect, not the bricklayer.</p>
<p class="text-muted-foreground leading-7 my-4"><strong class="text-foreground font-bold">&quot;I don&#x27;t have time.&quot;</strong> This one used to be real. Building something meaningful on nights and weekends was genuinely brutal. But when AI agents handle 60-70% of the implementation work, your time equation changes dramatically. What used to take a weekend now takes an evening. What used to take a month now takes a week.</p>
<p class="text-muted-foreground leading-7 my-4"><strong class="text-foreground font-bold">&quot;I don&#x27;t have a team.&quot;</strong> You don&#x27;t need one. Not the way you used to. AI agents can fill the roles that previously required hiring: content creation, code review, testing, deployment, even basic project management. You&#x27;re not a solo founder anymore. You&#x27;re a founder with a tireless, always-available team.</p>
<p class="text-muted-foreground leading-7 my-4"><strong class="text-foreground font-bold">&quot;I don&#x27;t have funding.&quot;</strong> Most of the tools that make this possible cost less than your monthly coffee budget. The expensive part of building used to be <em class="font-mono font-normal text-primary/80 print:text-primary">people</em>. When AI handles the work that people used to do, the cost structure collapses. You can build and launch a real product for nearly zero dollars.</p>
<p class="text-muted-foreground leading-7 my-4">Every single excuse has an AI-shaped hole in it.</p>
<h2 id="what-actually-changed" class="font-display text-2xl lg:text-3xl font-bold text-foreground mt-12 mb-6 border-l-2 border-primary pl-4"><a href="#what-actually-changed" class="heading-permalink text-inherit no-underline">What Actually Changed</a></h2>
<p class="text-muted-foreground leading-7 my-4">It&#x27;s easy to wave your hands and say &quot;AI makes everything easier.&quot; But the specific changes matter, because they&#x27;re what make this moment different from every other &quot;democratization of technology&quot; wave.</p>
<h3 id="1-ai-got-good-enough-to-ship" class="font-display text-xl font-bold text-foreground mt-10 mb-4"><a href="#1-ai-got-good-enough-to-ship" class="heading-permalink text-inherit no-underline">1. AI Got Good Enough to Ship</a></h3>
<p class="text-muted-foreground leading-7 my-4">The gap between &quot;AI demo&quot; and &quot;production-ready&quot; used to be enormous. AI could generate impressive-looking code that fell apart under real usage. That&#x27;s no longer the case. The current generation of AI agents produces code that passes tests, handles edge cases, and follows established patterns. Is it perfect? No. But it&#x27;s <strong class="text-foreground font-bold">good enough to ship</strong>, and shipping beats perfection every single time.</p>
<h3 id="2-the-full-stack-collapsed" class="font-display text-xl font-bold text-foreground mt-10 mb-4"><a href="#2-the-full-stack-collapsed" class="heading-permalink text-inherit no-underline">2. The Full Stack Collapsed</a></h3>
<p class="text-muted-foreground leading-7 my-4">You used to need different specialists for different layers. Now, the same set of AI tools can handle frontend, backend, infrastructure, and content. The &quot;full stack&quot; isn&#x27;t a rare skillset anymore — it&#x27;s the default mode of AI-assisted development. One person can operate across the entire stack because the AI fills in the gaps in their expertise.</p>
<h3 id="3-iteration-got-radically-faster" class="font-display text-xl font-bold text-foreground mt-10 mb-4"><a href="#3-iteration-got-radically-faster" class="heading-permalink text-inherit no-underline">3. Iteration Got Radically Faster</a></h3>
<p class="text-muted-foreground leading-7 my-4">The most underrated change isn&#x27;t the first version — it&#x27;s the second, third, and tenth version. AI makes iteration almost free. Don&#x27;t like the UI? Regenerate it. Need to pivot the data model? Let the agent handle the migration. Want to A/B test a new approach? Spin up a variant in an hour. When iteration is cheap, you can experiment fearlessly. And fearless experimentation is how good products are born.</p>
<h2 id="the-builders-playbook-for-2026" class="font-display text-2xl lg:text-3xl font-bold text-foreground mt-12 mb-6 border-l-2 border-primary pl-4"><a href="#the-builders-playbook-for-2026" class="heading-permalink text-inherit no-underline">The Builder&#x27;s Playbook for 2026</a></h2>
<p class="text-muted-foreground leading-7 my-4">If you&#x27;re convinced but not sure where to start, here&#x27;s the playbook I&#x27;d recommend:</p>
<h3 id="start-with-the-problem-not-the-tech" class="font-display text-xl font-bold text-foreground mt-10 mb-4"><a href="#start-with-the-problem-not-the-tech" class="heading-permalink text-inherit no-underline">Start With the Problem, Not the Tech</a></h3>
<p class="text-muted-foreground leading-7 my-4">The biggest trap I see new builders fall into is leading with the technology. &quot;I want to build something with AI&quot; is not a starting point. <strong class="text-foreground font-bold">&quot;I&#x27;m frustrated that X is broken and I think Y would fix it&quot;</strong> is a starting point. AI is the engine, but you still need to point the car somewhere worth driving.</p>
<h3 id="ship-in-days-not-months" class="font-display text-xl font-bold text-foreground mt-10 mb-4"><a href="#ship-in-days-not-months" class="heading-permalink text-inherit no-underline">Ship in Days, Not Months</a></h3>
<p class="text-muted-foreground leading-7 my-4">The old startup playbook said: spend months building, then launch. The new playbook says: <strong class="text-foreground font-bold">ship the smallest possible version this week</strong>. AI makes this feasible because the cost of building v1 is so low. Get it in front of people. Learn what&#x27;s wrong. Fix it. Repeat. The feedback loop is where all the value lives, and the sooner you enter it, the faster you learn.</p>
<h3 id="use-ai-as-a-team-not-a-tool" class="font-display text-xl font-bold text-foreground mt-10 mb-4"><a href="#use-ai-as-a-team-not-a-tool" class="heading-permalink text-inherit no-underline">Use AI as a Team, Not a Tool</a></h3>
<p class="text-muted-foreground leading-7 my-4">Stop thinking of AI as a code generator. Start thinking of it as a <strong class="text-foreground font-bold">team you manage</strong>. Assign tasks. Review output. Set standards. Give feedback. The mental model shift from &quot;AI writes my code&quot; to &quot;I lead a team of AI agents&quot; is the single biggest unlock for solo builders. You&#x27;re not doing less work — you&#x27;re doing <em class="font-mono font-normal text-primary/80 print:text-primary">different</em> work. Higher-leverage work.</p>
<h3 id="dont-polish-ship" class="font-display text-xl font-bold text-foreground mt-10 mb-4"><a href="#dont-polish-ship" class="heading-permalink text-inherit no-underline">Don&#x27;t Polish, Ship</a></h3>
<p class="text-muted-foreground leading-7 my-4">Perfectionism kills more projects than competition ever will. Your AI-generated UI doesn&#x27;t need to be pixel-perfect before launch. Your API doesn&#x27;t need 100% test coverage on day one. Your copy doesn&#x27;t need to win a Pulitzer. It needs to <strong class="text-foreground font-bold">exist</strong> and <strong class="text-foreground font-bold">work</strong> and be <strong class="text-foreground font-bold">in front of real users</strong>. Polish comes after validation, not before.</p>
<h2 id="the-real-unlock" class="font-display text-2xl lg:text-3xl font-bold text-foreground mt-12 mb-6 border-l-2 border-primary pl-4"><a href="#the-real-unlock" class="heading-permalink text-inherit no-underline">The Real Unlock</a></h2>
<p class="text-muted-foreground leading-7 my-4">Here&#x27;s what I&#x27;ve realized after months of building this way: the technology isn&#x27;t the breakthrough. The breakthrough is <strong class="text-foreground font-bold">permission</strong>.</p>
<p class="text-muted-foreground leading-7 my-4">For years, most of us told ourselves stories about why we couldn&#x27;t build. We didn&#x27;t have the skills, the time, the team, the money. Those stories felt true because they <em class="font-mono font-normal text-primary/80 print:text-primary">were</em> true — in a world where building required all of those things.</p>
<p class="text-muted-foreground leading-7 my-4">AI didn&#x27;t just give us new tools. It <strong class="text-foreground font-bold">invalidated our excuses</strong>. And when your excuses go away, the only thing left is the question you&#x27;ve been avoiding: <em class="font-mono font-normal text-primary/80 print:text-primary">do you actually want to build this, or were the excuses more comfortable?</em></p>
<p class="text-muted-foreground leading-7 my-4">That&#x27;s a harder question than any technical challenge. But if your answer is yes — if there&#x27;s something you&#x27;ve been wanting to create, a problem you&#x27;ve been wanting to solve, an idea that won&#x27;t leave you alone — then you&#x27;re out of reasons to wait.</p>
<h2 id="the-bottom-line" class="font-display text-2xl lg:text-3xl font-bold text-foreground mt-12 mb-6 border-l-2 border-primary pl-4"><a href="#the-bottom-line" class="heading-permalink text-inherit no-underline">The Bottom Line</a></h2>
<p class="text-muted-foreground leading-7 my-4">The gap between &quot;idea&quot; and &quot;product&quot; has never been smaller. The cost has never been lower. The tools have never been better. And the window won&#x27;t stay this wide forever — as more people realize what&#x27;s possible, the advantage of being early shrinks.</p>
<p class="text-muted-foreground leading-7 my-4">So stop planning. Stop researching. Stop waiting for the perfect moment or the perfect co-founder or the perfect market conditions.</p>
<p class="text-muted-foreground leading-7 my-4">Open your laptop. Describe what you want to build. And ship it yourself.</p>
<p class="text-muted-foreground leading-7 my-4">The world doesn&#x27;t need another pitch deck. It needs another product. And you&#x27;re the one who can build it.</p>]]></content:encoded>
      <pubDate>Sun, 12 Apr 2026 12:00:00 GMT</pubDate>
    </item>
    <item>
      <title>From Co-Pilots to Colleagues: How AI Agents Changed My Engineering Workflow</title>
      <link>https://pratik.pa.tel/blog/from-copilots-to-colleagues/</link>
      <guid isPermaLink="true">https://pratik.pa.tel/blog/from-copilots-to-colleagues/</guid>
      <description>A year of working alongside AI teammates reshaped how I think about building software</description>
      <dc:creator>Pratik Patel</dc:creator>
      <category>ai</category>
      <category>engineering</category>
      <category>productivity</category>
      <content:encoded><![CDATA[<p class="text-muted-foreground leading-7 my-4">A little over a year ago, I wrote about my first experience using <strong class="text-foreground font-bold">Devin AI</strong> as a coding co-pilot. The takeaway was clear: AI wasn&#x27;t replacing engineers, but it was becoming a surprisingly capable junior teammate. Fast forward to today, and that framing already feels quaint. The AI agents I work with now aren&#x27;t co-pilots. They&#x27;re closer to <strong class="text-foreground font-bold">colleagues</strong>.</p>
<p class="text-muted-foreground leading-7 my-4">Here&#x27;s what changed, what I got wrong, and what I&#x27;ve learned about building software alongside AI agents in 2026.</p>
<h2 id="the-shift-from-autocomplete-to-autonomy" class="font-display text-2xl lg:text-3xl font-bold text-foreground mt-12 mb-6 border-l-2 border-primary pl-4"><a href="#the-shift-from-autocomplete-to-autonomy" class="heading-permalink text-inherit no-underline">The Shift: From Autocomplete to Autonomy 🔄</a></h2>
<p class="text-muted-foreground leading-7 my-4">When I first started using AI coding tools, the mental model was simple: <strong class="text-foreground font-bold">I think, it types</strong>. Copilot-style tools predicted the next line. I was still the driver. The AI was a fancy autocomplete engine that occasionally read my mind.</p>
<p class="text-muted-foreground leading-7 my-4">The agents I use today operate differently. I describe a problem, point them at the relevant code, and they go figure it out. They read documentation, explore the codebase, draft a plan, write the implementation, run the tests, and open a PR. Sometimes they even catch edge cases I didn&#x27;t think of.</p>
<p class="text-muted-foreground leading-7 my-4">The biggest mental shift wasn&#x27;t learning new tools. It was learning to <strong class="text-foreground font-bold">delegate</strong>. And delegation, it turns out, is a skill that most engineers never had to practice with machines before.</p>
<h2 id="what-i-got-wrong-last-year" class="font-display text-2xl lg:text-3xl font-bold text-foreground mt-12 mb-6 border-l-2 border-primary pl-4"><a href="#what-i-got-wrong-last-year" class="heading-permalink text-inherit no-underline">What I Got Wrong Last Year 🤔</a></h2>
<p class="text-muted-foreground leading-7 my-4">In my Devin AI post, I framed the value proposition as a time trade-off: 15 minutes of my own coding vs. 1 hour with Devin. That math was real, but the conclusion I drew was too narrow. I was measuring the wrong thing.</p>
<p class="text-muted-foreground leading-7 my-4">The real value isn&#x27;t &quot;did this specific task get done faster?&quot; It&#x27;s <strong class="text-foreground font-bold">&quot;what did I do with the time I didn&#x27;t spend on it?&quot;</strong> When I stopped measuring AI by how fast it could do <em class="font-mono font-normal text-primary/80 print:text-primary">my</em> tasks and started measuring it by how much it expanded <em class="font-mono font-normal text-primary/80 print:text-primary">my capacity</em>, the picture changed dramatically.</p>
<p class="text-muted-foreground leading-7 my-4">These days, I routinely have two or three agent sessions running in parallel while I focus on architecture decisions, stakeholder conversations, or code review. My throughput hasn&#x27;t just increased — it&#x27;s <strong class="text-foreground font-bold">qualitatively different</strong>. I spend more time on the problems that actually need a human brain, and less time on the ones that don&#x27;t.</p>
<h2 id="the-trust-calibration-problem" class="font-display text-2xl lg:text-3xl font-bold text-foreground mt-12 mb-6 border-l-2 border-primary pl-4"><a href="#the-trust-calibration-problem" class="heading-permalink text-inherit no-underline">The Trust Calibration Problem ⚖️</a></h2>
<p class="text-muted-foreground leading-7 my-4">Here&#x27;s the thing nobody warns you about when working with AI agents: <strong class="text-foreground font-bold">trust is harder than prompting</strong>.</p>
<p class="text-muted-foreground leading-7 my-4">Early on, I over-trusted. I&#x27;d skim an AI-generated PR, approve it, and move on. Then I&#x27;d find a subtle bug two days later — something that passed tests but violated an unwritten assumption about how our system handles state. The AI didn&#x27;t know our system&#x27;s history. It only knew the code as it existed on disk.</p>
<p class="text-muted-foreground leading-7 my-4">Then I over-corrected. I reviewed AI PRs with more scrutiny than I&#x27;d give a senior engineer&#x27;s code. That defeated the entire purpose. I was spending <em class="font-mono font-normal text-primary/80 print:text-primary">more</em> time reviewing than I would have spent just writing the code myself.</p>
<p class="text-muted-foreground leading-7 my-4">The sweet spot — and I think every engineer working with agents has to find their own — is what I call <strong class="text-foreground font-bold">calibrated trust</strong>:</p>
<ul class="space-y-2 my-6 ml-4">
<li class="flex gap-3 text-muted-foreground leading-7"><span class="text-primary shrink-0 mt-1.5">▸</span><span class="min-w-0"><strong class="text-foreground font-bold">High trust</strong> for well-defined, well-tested tasks: CRUD endpoints, data transformations, boilerplate setup, migrations with clear schemas</span></li>
<li class="flex gap-3 text-muted-foreground leading-7"><span class="text-primary shrink-0 mt-1.5">▸</span><span class="min-w-0"><strong class="text-foreground font-bold">Medium trust</strong> for tasks that require domain context: business logic, API integrations, anything touching auth or payments</span></li>
<li class="flex gap-3 text-muted-foreground leading-7"><span class="text-primary shrink-0 mt-1.5">▸</span><span class="min-w-0"><strong class="text-foreground font-bold">Low trust</strong> for tasks involving system design, performance-sensitive code, or subtle correctness requirements</span></li>
</ul>
<p class="text-muted-foreground leading-7 my-4">This isn&#x27;t that different from how you&#x27;d calibrate trust with a human teammate. The difference is that AI agents are <strong class="text-foreground font-bold">consistently good at their strengths and consistently blind to their weaknesses</strong>. Humans are more variable but also more self-aware. Once you internalize that pattern, the collaboration gets much smoother.</p>
<h2 id="three-habits-that-made-the-difference" class="font-display text-2xl lg:text-3xl font-bold text-foreground mt-12 mb-6 border-l-2 border-primary pl-4"><a href="#three-habits-that-made-the-difference" class="heading-permalink text-inherit no-underline">Three Habits That Made the Difference 🛠️</a></h2>
<p class="text-muted-foreground leading-7 my-4">After a year of iteration, three practices made my AI-augmented workflow actually <em class="font-mono font-normal text-primary/80 print:text-primary">work</em>:</p>
<h3 id="1-write-better-context-not-better-prompts" class="font-display text-xl font-bold text-foreground mt-10 mb-4"><a href="#1-write-better-context-not-better-prompts" class="heading-permalink text-inherit no-underline">1. Write Better Context, Not Better Prompts</a></h3>
<p class="text-muted-foreground leading-7 my-4">The prompt engineering hype was overblown. What actually matters is <strong class="text-foreground font-bold">context</strong>. AI agents do better work when they have access to clear documentation, well-named functions, and explicit conventions. Every time I improved our codebase&#x27;s readability for humans, the AI agents got better too.</p>
<p class="text-muted-foreground leading-7 my-4">The irony isn&#x27;t lost on me: the best way to make AI productive is to make your codebase better for <em class="font-mono font-normal text-primary/80 print:text-primary">everyone</em>.</p>
<h3 id="2-review-the-plan-not-just-the-code" class="font-display text-xl font-bold text-foreground mt-10 mb-4"><a href="#2-review-the-plan-not-just-the-code" class="heading-permalink text-inherit no-underline">2. Review the Plan, Not Just the Code</a></h3>
<p class="text-muted-foreground leading-7 my-4">Most AI agent tools now show you a plan before they start coding. I used to skip this step. Now it&#x27;s the most valuable part of the process. Catching a wrong assumption at the plan stage saves 10x the time compared to catching it in code review.</p>
<p class="text-muted-foreground leading-7 my-4">When I review an agent&#x27;s plan, I&#x27;m asking: <em class="font-mono font-normal text-primary/80 print:text-primary">Does this agent understand the problem the way I do?</em> If the answer is no, I course-correct before a single line of code gets written.</p>
<h3 id="3-keep-a-human-in-the-architecture-loop" class="font-display text-xl font-bold text-foreground mt-10 mb-4"><a href="#3-keep-a-human-in-the-architecture-loop" class="heading-permalink text-inherit no-underline">3. Keep a Human in the Architecture Loop</a></h3>
<p class="text-muted-foreground leading-7 my-4">AI agents are great at implementing within a well-defined boundary. They&#x27;re not great at deciding where the boundary should be. Architectural decisions — where does this logic live, how do these services communicate, what are the failure modes — still need human judgment.</p>
<p class="text-muted-foreground leading-7 my-4">I&#x27;ve settled into a rhythm: I make the structural decisions, the agents fill in the implementation, and I review the result. It&#x27;s not unlike being a tech lead, except my &quot;team&quot; never gets tired and never has opinions about tabs vs. spaces.</p>
<h2 id="what-im-watching-next" class="font-display text-2xl lg:text-3xl font-bold text-foreground mt-12 mb-6 border-l-2 border-primary pl-4"><a href="#what-im-watching-next" class="heading-permalink text-inherit no-underline">What I&#x27;m Watching Next 🔮</a></h2>
<p class="text-muted-foreground leading-7 my-4">The pace of improvement in AI agents is staggering. A year ago, getting an agent to handle a multi-file refactor reliably felt like a stretch. Now it&#x27;s routine. The frontier is moving toward agents that can maintain context across longer arcs of work — understanding not just the current task but the <em class="font-mono font-normal text-primary/80 print:text-primary">project trajectory</em>.</p>
<p class="text-muted-foreground leading-7 my-4">I&#x27;m also seeing more teams adopt agents not as individual tools but as <strong class="text-foreground font-bold">team members with defined roles</strong>: one agent handles test coverage, another manages dependency updates, another writes documentation. The multi-agent workflow is still early, but the pattern is emerging.</p>
<p class="text-muted-foreground leading-7 my-4">The engineers who will thrive in this landscape aren&#x27;t the ones who write the fastest code. They&#x27;re the ones who can <strong class="text-foreground font-bold">orchestrate, review, and architect</strong> — the skills that have always defined senior engineering, now amplified by a new kind of teammate.</p>
<h2 id="the-bottom-line" class="font-display text-2xl lg:text-3xl font-bold text-foreground mt-12 mb-6 border-l-2 border-primary pl-4"><a href="#the-bottom-line" class="heading-permalink text-inherit no-underline">The Bottom Line 🎯</a></h2>
<p class="text-muted-foreground leading-7 my-4">Working with AI agents for the past year taught me something I didn&#x27;t expect: it made me a <strong class="text-foreground font-bold">better engineer</strong>, not because the AI wrote my code, but because it forced me to think more clearly about what I actually wanted built. You can&#x27;t delegate effectively if you don&#x27;t understand the problem deeply yourself.</p>
<p class="text-muted-foreground leading-7 my-4">AI agents aren&#x27;t replacing engineers. They&#x27;re raising the bar for what &quot;engineering&quot; means. Less time typing, more time thinking. Less time on the routine, more time on the remarkable.</p>
<p class="text-muted-foreground leading-7 my-4">And honestly? I wouldn&#x27;t go back. The way I work now — with AI colleagues running alongside me — feels like the way software was always meant to be built. We just didn&#x27;t have the teammates for it until now.</p>]]></content:encoded>
      <pubDate>Sun, 05 Apr 2026 12:00:00 GMT</pubDate>
    </item>
    <item>
      <title>A Real-World Coding Story: Devin AI as My Co-Pilot</title>
      <link>https://pratik.pa.tel/blog/devin-ai-as-my-co-pilot/</link>
      <guid isPermaLink="true">https://pratik.pa.tel/blog/devin-ai-as-my-co-pilot/</guid>
      <description>What happens when you hand off a real task to an AI teammate</description>
      <dc:creator>Pratik Patel</dc:creator>
      <category>ai</category>
      <category>engineering</category>
      <category>productivity</category>
      <content:encoded><![CDATA[<p class="text-muted-foreground leading-7 my-4">I recently took <strong class="text-foreground font-bold">Devin AI</strong> for a spin on a real development task, and the experience felt like something between magic and mentorship. I had a straightforward job: <strong class="text-foreground font-bold">implement an API endpoint to generate user access tokens for a third-party service</strong>. Normally I&#x27;d crank this out in ~15 minutes of coding. This time, I decided to hand it off to my new &quot;AI teammate&quot; and see what happened. Here&#x27;s how it went, step by step.</p>
<h2 id="the-task-initial-plan" class="font-display text-2xl lg:text-3xl font-bold text-foreground mt-12 mb-6 border-l-2 border-primary pl-4"><a href="#the-task-initial-plan" class="heading-permalink text-inherit no-underline">The Task &amp; Initial Plan 🚀</a></h2>
<p class="text-muted-foreground leading-7 my-4">My instruction to Devin was simple:</p>
<blockquote class="my-6 border-l-2 border-primary/40 print:border-primary pl-6">
<p class="text-muted-foreground leading-7 my-4">Implement the backend API based on this documentation [link to 3P website]. Add it to [microservice API name]. Utilize the existing CDK secrets logic to store the 3P API Key.</p>
</blockquote>
<p class="text-muted-foreground leading-7 my-4">Devin&#x27;s response was almost immediate:</p>
<ul class="space-y-2 my-6 ml-4">
<li class="flex gap-3 text-muted-foreground leading-7"><span class="text-primary shrink-0 mt-1.5">▸</span><span class="min-w-0"><strong class="text-foreground font-bold">Understanding the goal:</strong> It parsed my request and the service docs, and quickly outlined a plan to create the token-generation endpoint.</span></li>
<li class="flex gap-3 text-muted-foreground leading-7"><span class="text-primary shrink-0 mt-1.5">▸</span><span class="min-w-0"><strong class="text-foreground font-bold">Collecting details:</strong> Devin identified required endpoints and data (thanks to the link I gave) and noted it would need to handle authentication, token storage, etc.</span></li>
<li class="flex gap-3 text-muted-foreground leading-7"><span class="text-primary shrink-0 mt-1.5">▸</span><span class="min-w-0"><strong class="text-foreground font-bold">Generating a plan:</strong> Before writing code, Devin presented a high-level game plan. This included creating a new route in our backend, calling the third-party API for the token, and returning the result to our app. I was impressed – the plan was thorough and made sense given the task.</span></li>
</ul>
<p class="text-muted-foreground leading-7 my-4">In essence, Devin did a great job figuring out <em class="font-mono font-normal text-primary/80 print:text-primary">what</em> needed to be done without me hand-holding the requirements. It felt like I was working with an autonomous engineer who eagerly drafts a design spec after a short request.</p>
<h2 id="from-prompt-to-pull-request" class="font-display text-2xl lg:text-3xl font-bold text-foreground mt-12 mb-6 border-l-2 border-primary pl-4"><a href="#from-prompt-to-pull-request" class="heading-permalink text-inherit no-underline">From Prompt to Pull Request 💻</a></h2>
<p class="text-muted-foreground leading-7 my-4">With the plan in place, I gave Devin the green light. It jumped into coding mode. Within minutes, <strong class="text-foreground font-bold">Devin had opened a GitHub pull request</strong> on our repository with the new API implementation. I could follow its progress in real-time through Devin&#x27;s interface. (The UI actually provides a step-by-step log of what the AI is doing – very cool!).</p>
<p class="text-muted-foreground leading-7 my-4">Watching Devin work was surreal. The <strong class="text-foreground font-bold">UI/UX</strong> of the tool is <strong class="text-foreground font-bold">amazing</strong> – it felt seamless and intuitive to use. As one early user noted, <em class="font-mono font-normal text-primary/80 print:text-primary">&quot;Devin feels UI/UX first, not GenAI first,&quot;</em> emphasizing that the surrounding experience is the star. I have to agree. In my case, giving instructions felt as easy as chatting with a colleague, and I could see Devin&#x27;s thought process and actions clearly. The combination of a chat interface, an embedded code editor, and live updates made it <strong class="text-foreground font-bold">feel like pair-programming with a supercharged junior dev</strong>.</p>
<p class="text-muted-foreground leading-7 my-4">Soon, the PR was ready for review. The code Devin produced was surprisingly solid for a first pass. It had set up the new API endpoint, made calls to the third-party service, and hooked everything into our backend. All of this happened while I was hands-off.</p>
<h2 id="hiccups-and-iterations" class="font-display text-2xl lg:text-3xl font-bold text-foreground mt-12 mb-6 border-l-2 border-primary pl-4"><a href="#hiccups-and-iterations" class="heading-permalink text-inherit no-underline">Hiccups and Iterations 🛠️</a></h2>
<p class="text-muted-foreground leading-7 my-4">Of course, it wasn&#x27;t perfect out of the box. Upon reviewing the pull request, I spotted a few issues that needed addressing before this code could go to production. Notably:</p>
<ul class="space-y-2 my-6 ml-4">
<li class="flex gap-3 text-muted-foreground leading-7"><span class="text-primary shrink-0 mt-1.5">▸</span><span class="min-w-0"><strong class="text-foreground font-bold">Missing API key logic:</strong> Devin forgot to include the authentication API key when calling the third-party service. A human engineer knows that&#x27;s a must for the request to succeed, but the AI overlooked this detail on the first try.</span></li>
<li class="flex gap-3 text-muted-foreground leading-7"><span class="text-primary shrink-0 mt-1.5">▸</span><span class="min-w-0"><strong class="text-foreground font-bold">Code style differences:</strong> Some of the naming conventions and formatting didn&#x27;t match our project&#x27;s style guidelines. (Minor issue, but something we&#x27;d fix in any code review – even with human contributors.)</span></li>
<li class="flex gap-3 text-muted-foreground leading-7"><span class="text-primary shrink-0 mt-1.5">▸</span><span class="min-w-0"><strong class="text-foreground font-bold">Best practice tweaks:</strong> I noticed that the API logic for calling the 3P service should have been abstracted into our &quot;clients&quot; library. While Devin&#x27;s solution was functional, integrating this logic into the shared library would improve maintainability and align with our standards.</span></li>
</ul>
<p class="text-muted-foreground leading-7 my-4">The great thing was that, <strong class="text-foreground font-bold">addressing these issues felt very natural.</strong> I went into the PR on GitHub and dropped comments exactly like I would for a human colleague. For example, I pointed out where the API key should be injected, suggested where to move code around, and reminded to clean up some unneeded additions.</p>
<p class="text-muted-foreground leading-7 my-4">Devin took this feedback in stride. It truly felt like collaborating with a keen junior developer: I&#x27;d leave a note, and Devin would go off to fix it. Devin was <strong class="text-foreground font-bold">eager to improve</strong> and quickly pushed new commits to the PR, incorporating my suggestions.</p>
<p class="text-muted-foreground leading-7 my-4">We went through about <strong class="text-foreground font-bold">2-3 iterations</strong> like this. Each cycle, I&#x27;d review the updates, find fewer things to tweak, and comment on the remaining issues. Devin would promptly address them. After this iterative back-and-forth, the API code was <strong class="text-foreground font-bold">production-ready</strong>. All tests passed, the style was consistent, and the integration worked flawlessly with the third-party service.</p>
<h2 id="time-trade-off-15-minutes-vs-1-hour" class="font-display text-2xl lg:text-3xl font-bold text-foreground mt-12 mb-6 border-l-2 border-primary pl-4"><a href="#time-trade-off-15-minutes-vs-1-hour" class="heading-permalink text-inherit no-underline">Time Trade-Off: 15 Minutes vs 1 Hour ⏱️</a></h2>
<p class="text-muted-foreground leading-7 my-4">You might be wondering: <em class="font-mono font-normal text-primary/80 print:text-primary">Was using Devin worth it, time-wise?</em> By the numbers, <strong class="text-foreground font-bold">doing it myself would have been faster</strong> – roughly 15 minutes of coding versus about <strong class="text-foreground font-bold">1 hour</strong> to get it done via Devin (including the initial setup, waiting for the AI to do its work, reviewing the PR, and guiding the fixes). That&#x27;s a 4x increase in wall-clock time for the task, which sounds like a loss in efficiency.</p>
<p class="text-muted-foreground leading-7 my-4">In fact, others have noted this current limitation of AI coding agents. The waiting and iterative feedback loop can indeed make the process longer than just writing the code yourself, especially for a simple task.</p>
<p class="text-muted-foreground leading-7 my-4">However, here&#x27;s the catch: <strong class="text-foreground font-bold">while Devin was working, I wasn&#x27;t stuck waiting idly.</strong> During that 1 hour I was free to focus on other work. I answered a couple of messages, reviewed a different PR from a teammate, and even started brainstorming a design for an upcoming feature – all while Devin handled the heavy lifting for this task in the background.</p>
<p class="text-muted-foreground leading-7 my-4">In essence, that hour wasn&#x27;t me twiddling my thumbs; it was more like delegating to a capable assistant. Yes, the <strong class="text-foreground font-bold">calendar time</strong> was longer, but my <strong class="text-foreground font-bold">personal time investment</strong> was much less than an hour of active coding. I probably spent only a few minutes giving instructions and about 5 minutes total reviewing and commenting. The rest of the time, Devin was on the job autonomously.</p>
<p class="text-muted-foreground leading-7 my-4">It&#x27;s a trade-off: <strong class="text-foreground font-bold">faster solo vs. parallelized teamwork.</strong> If I had 10 such small tasks in a sprint, I could theoretically assign them all to Devin and attend to bigger challenges, checking in occasionally for reviews. That ability to multitask is where a tool like this shows its value.</p>
<h2 id="the-ui-ux-a-smooth-ride" class="font-display text-2xl lg:text-3xl font-bold text-foreground mt-12 mb-6 border-l-2 border-primary pl-4"><a href="#the-ui-ux-a-smooth-ride" class="heading-permalink text-inherit no-underline">The UI/UX: A Smooth Ride 🎨</a></h2>
<p class="text-muted-foreground leading-7 my-4">I have to come back to <strong class="text-foreground font-bold">Devin&#x27;s user experience</strong>, because it really enhanced the whole process. The interface made it super easy to interact with the AI:</p>
<ul class="space-y-2 my-6 ml-4">
<li class="flex gap-3 text-muted-foreground leading-7"><span class="text-primary shrink-0 mt-1.5">▸</span><span class="min-w-0">I gave my instructions in a chat-like format (in our case, through Devin&#x27;s Slack integration and then via GitHub PR comments). No complex setup or configuration; it was like talking to a teammate.</span></li>
<li class="flex gap-3 text-muted-foreground leading-7"><span class="text-primary shrink-0 mt-1.5">▸</span><span class="min-w-0">Devin kept me updated with a live log and even a &quot;plan file&quot; of notes. I could literally see what it was thinking – the steps it was taking, the commands it ran, files it created or modified, etc. This transparency is <strong class="text-foreground font-bold">huge</strong> for trust when an AI is writing your code.</span></li>
<li class="flex gap-3 text-muted-foreground leading-7"><span class="text-primary shrink-0 mt-1.5">▸</span><span class="min-w-0">The pull request it opened was clear and well-structured. It included a description of what the change was and even referenced the task (just as a diligent dev would do).</span></li>
<li class="flex gap-3 text-muted-foreground leading-7"><span class="text-primary shrink-0 mt-1.5">▸</span><span class="min-w-0">When I left feedback, the UI (and GitHub integration) notified Devin immediately. It felt like the system was built around a smooth feedback loop, which is crucial. I commented and within a minute Devin&#x27;s next update had the fix implemented.</span></li>
</ul>
<p class="text-muted-foreground leading-7 my-4">The polish and thought put into Devin&#x27;s UX did not go unnoticed. It didn&#x27;t feel like using a clunky experimental tool; it felt like working in an environment <strong class="text-foreground font-bold">built for developers&#x27; comfort</strong>. This level of refinement in developer tools is refreshing – it let me focus on the results rather than wrestling with the tool itself.</p>
<h2 id="final-thoughts-a-promising-co-pilot-not-a-replacement" class="font-display text-2xl lg:text-3xl font-bold text-foreground mt-12 mb-6 border-l-2 border-primary pl-4"><a href="#final-thoughts-a-promising-co-pilot-not-a-replacement" class="heading-permalink text-inherit no-underline">Final Thoughts: A Promising Co-Pilot, Not a Replacement 🚧</a></h2>
<p class="text-muted-foreground leading-7 my-4">After this trial run, here&#x27;s my reflection: <strong class="text-foreground font-bold">Devin AI is not replacing engineers anytime soon, but it&#x27;s certainly a promising co-pilot</strong>. The experience was akin to working with a supercharged junior engineer who can execute tasks and learn from feedback. It had its blind spots and needed guidance, but ultimately it delivered real value.</p>
<p class="text-muted-foreground leading-7 my-4">The limitations I encountered (missing a key detail, needing adjustments, slower turnaround) underscore that human expertise is still crucial. In real-world development, context and subtle requirements matter – things an experienced human developer intuitively catches, but an AI might miss without a proper prompt. I had to be the quality control, just like I would with a less-experienced team member. <strong class="text-foreground font-bold">Devin isn&#x27;t about to take over my job</strong>, and I wouldn&#x27;t trust it to run completely unsupervised on anything critical just yet.</p>
<p class="text-muted-foreground leading-7 my-4">That said, the benefits were significant. By offloading a chunk of work to the AI, I freed up mental space and time. I found myself thinking more about <em class="font-mono font-normal text-primary/80 print:text-primary">what</em> needed to be done, and less about the minute details of <em class="font-mono font-normal text-primary/80 print:text-primary">how</em> to code it in the moment. It&#x27;s a different way of working – more high-level orchestration, less in-the-trenches coding for certain tasks. As the creators of Devin intended, it&#x27;s meant to be a collaborative helper rather than a threat.</p>
<p class="text-muted-foreground leading-7 my-4">For a first-gen AI developer agent, Devin exceeded my expectations in UX and autonomy. It felt like I had an eager intern who works blindingly fast and never gets tired, but occasionally needs me to double-check the work. I can live with that! The technology will only get better from here. With more polish and learning from each interaction, I imagine tools like Devin will handle bigger chunks of the development workload, and do so more reliably.</p>
<p class="text-muted-foreground leading-7 my-4"><strong class="text-foreground font-bold">Bottom line:</strong> Devin AI isn&#x27;t a replacement for an engineer – it&#x27;s a new kind of teammate. Using it won&#x27;t instantly halve your development time (yet), but it <em class="font-mono font-normal text-primary/80 print:text-primary">will</em> change how you can allocate your time. I was able to focus on other priorities while it cranked out code. That kind of <strong class="text-foreground font-bold">parallel productivity</strong> is game-changing if used right.</p>
<p class="text-muted-foreground leading-7 my-4">I&#x27;m excited to continue using Devin as a co-pilot for future projects. It&#x27;s like having a junior dev in the background who speeds through the boring stuff and lets me concentrate on the fun and hard parts of engineering. <strong class="text-foreground font-bold">Not perfect, but very promising.</strong> The future of development might not be &quot;AI vs Human&quot;, but <strong class="text-foreground font-bold">AI + Human, working in tandem</strong> – and my experience with Devin AI this week certainly makes me feel that way. 🚀👏</p>]]></content:encoded>
      <pubDate>Thu, 20 Feb 2025 12:00:00 GMT</pubDate>
    </item>
    <item>
      <title>Own Your Career</title>
      <link>https://pratik.pa.tel/blog/own-your-career/</link>
      <guid isPermaLink="true">https://pratik.pa.tel/blog/own-your-career/</guid>
      <description>No one will fight harder for your career growth than you. Five lessons on showing your value, creating your own opportunities, and earning the promotion.</description>
      <dc:creator>Pratik Patel</dc:creator>
      <category>growth</category>
      <category>engineering</category>
      <content:encoded><![CDATA[<p class="text-muted-foreground leading-7 my-4">Early in my career, I made a huge mistake. I thought if I worked hard and hit all my deadlines, my manager would naturally recognize my efforts and drive my career forward. Promotions, salary bumps, new opportunities—they&#x27;d all come, right?</p>
<p class="text-muted-foreground leading-7 my-4">Wrong. 🚫</p>
<p class="text-muted-foreground leading-7 my-4">After waiting (and waiting… and waiting) for someone else to advocate for me, I realized I was approaching promotions all wrong. The truth is, no one will fight harder for your career growth than YOU. It&#x27;s your responsibility to show your value, take ownership of your path, and create opportunities for yourself.</p>
<p class="text-muted-foreground leading-7 my-4">Here are the five biggest lessons I&#x27;ve learned about driving my career and earning my promotions.</p>
<h2 id="1-you-own-your-career-path" class="font-display text-2xl lg:text-3xl font-bold text-foreground mt-12 mb-6 border-l-2 border-primary pl-4"><a href="#1-you-own-your-career-path" class="heading-permalink text-inherit no-underline">1. You Own Your Career Path</a></h2>
<p class="text-muted-foreground leading-7 my-4">It sounds obvious, but many professionals (my younger self included) mistakenly assume that their manager or team lead is responsible for charting their career progression. The truth? Nope. That&#x27;s all on you. Your growth, development, and career path are in your hands. While managers can provide guidance, feedback, and opportunities, it&#x27;s ultimately up to you to set goals, seek out learning experiences, and take the steps needed to advance.</p>
<p class="text-muted-foreground leading-7 my-4">That&#x27;s not to say that some managers don&#x27;t genuinely want to see you succeed—they absolutely do. But keep in mind, most managers are juggling their own set of responsibilities, goals, and ambitions. They may be focused on hitting deadlines, managing team performance, or achieving personal career milestones. As a result, even the most well-intentioned managers may not always have the time, resources, or bandwidth to give your career growth the dedicated attention it deserves. This is why it&#x27;s crucial to take ownership of your own development.</p>
<p class="text-muted-foreground leading-7 my-4">Think of your career as a long-term project. You&#x27;re the project manager. Success depends on setting goals, creating a roadmap, and consistently reviewing your progress. Want to move upward in your career? Ask yourself questions like:</p>
<ul class="space-y-2 my-6 ml-4">
<li class="flex gap-3 text-muted-foreground leading-7"><span class="text-primary shrink-0 mt-1.5">▸</span><span class="min-w-0">What roles interest me in the next 2 or 5 years?</span></li>
<li class="flex gap-3 text-muted-foreground leading-7"><span class="text-primary shrink-0 mt-1.5">▸</span><span class="min-w-0">What skills do I need to develop to land those roles?</span></li>
<li class="flex gap-3 text-muted-foreground leading-7"><span class="text-primary shrink-0 mt-1.5">▸</span><span class="min-w-0">Who can help me achieve these goals?</span></li>
</ul>
<p class="text-muted-foreground leading-7 my-4">The moment you adopt a proactive mindset, you start taking control of your trajectory. Don&#x27;t wait for someone to hand you a blueprint. Draft one for yourself and take the lead.</p>
<h2 id="2-document-your-wins-religiously" class="font-display text-2xl lg:text-3xl font-bold text-foreground mt-12 mb-6 border-l-2 border-primary pl-4"><a href="#2-document-your-wins-religiously" class="heading-permalink text-inherit no-underline">2. Document Your Wins Religiously</a></h2>
<p class="text-muted-foreground leading-7 my-4">People love a good story, and your career is no different—it&#x27;s up to you to craft a compelling narrative. That narrative begins with documentation. Every single achievement, no matter how small, deserves a record.</p>
<p class="text-muted-foreground leading-7 my-4">Implemented a new software feature? Write it down. Resolved a high-stakes bug right before launch? Document it. Improved a process that saved your team time or reduced costs? Log it.</p>
<p class="text-muted-foreground leading-7 my-4">Why is this important? When promotion discussions happen (or interview panels for a new job roll around), you&#x27;ll have concrete examples to prove your value:</p>
<ul class="space-y-2 my-6 ml-4">
<li class="flex gap-3 text-muted-foreground leading-7"><span class="text-primary shrink-0 mt-1.5">▸</span><span class="min-w-0">Outline the <strong class="text-foreground font-bold">problem</strong> you solved.</span></li>
<li class="flex gap-3 text-muted-foreground leading-7"><span class="text-primary shrink-0 mt-1.5">▸</span><span class="min-w-0">Highlight your <strong class="text-foreground font-bold">actions</strong>.</span></li>
<li class="flex gap-3 text-muted-foreground leading-7"><span class="text-primary shrink-0 mt-1.5">▸</span><span class="min-w-0">Summarize the <strong class="text-foreground font-bold">results</strong> (bonus points for measurable outcomes like percentage improvement or dollars saved).</span></li>
</ul>
<p class="text-muted-foreground leading-7 my-4">Your wins demonstrate your growth and impact. They&#x27;re your ultimate receipts. And when you bring them up in conversations, your contributions go from &quot;implied&quot; to undeniable.</p>
<h2 id="3-seek-out-high-impact-projects" class="font-display text-2xl lg:text-3xl font-bold text-foreground mt-12 mb-6 border-l-2 border-primary pl-4"><a href="#3-seek-out-high-impact-projects" class="heading-permalink text-inherit no-underline">3. Seek Out High-Impact Projects</a></h2>
<p class="text-muted-foreground leading-7 my-4">Here&#x27;s a truth bomb I wish someone had dropped on me earlier in my career: the work you do matters, but <em class="font-mono font-normal text-primary/80 print:text-primary">where</em> you focus your energy matters even more. High-impact projects are your ticket to visibility and leadership opportunities.</p>
<p class="text-muted-foreground leading-7 my-4">What makes a project high-impact?</p>
<ul class="space-y-2 my-6 ml-4">
<li class="flex gap-3 text-muted-foreground leading-7"><span class="text-primary shrink-0 mt-1.5">▸</span><span class="min-w-0">It solves a major pain point for your team or company.</span></li>
<li class="flex gap-3 text-muted-foreground leading-7"><span class="text-primary shrink-0 mt-1.5">▸</span><span class="min-w-0">It affects a large number of people.</span></li>
<li class="flex gap-3 text-muted-foreground leading-7"><span class="text-primary shrink-0 mt-1.5">▸</span><span class="min-w-0">It drives measurable results, like saving time, generating revenue, or improving efficiency.</span></li>
</ul>
<p class="text-muted-foreground leading-7 my-4">Don&#x27;t wait for these projects to come to you—volunteer for them. Offer to tackle that backlog no one wants to touch, spearhead a cross-functional initiative, or run an experiment that aligns with your company&#x27;s top priorities.</p>
<p class="text-muted-foreground leading-7 my-4">Pro tip: High-impact projects also showcase your ability to handle bigger responsibilities, making it easier for decision-makers to envision you in advanced roles.</p>
<p class="text-muted-foreground leading-7 my-4"><strong class="text-foreground font-bold">Relevant Aside</strong>: While high-impact projects often take center stage, it&#x27;s important to recognize that in large companies, not all initiatives need to yield direct, measurable results to be valuable. A prime example of this lies in projects like redesigning or rewriting applications. These efforts, while not always critical in terms of immediate business outcomes, often become what&#x27;s colloquially known as &quot;promo projects.&quot; They give employees the opportunity to showcase their technical abilities, lead teams, or experiment with modern frameworks and tools — all of which can serve as stepping stones in career advancement.</p>
<p class="text-muted-foreground leading-7 my-4">The reality is, sometimes you have to play the game—it&#x27;s about visibility and shaping the narrative. While redesigns or rewrites may not always yield clear, measurable ROI compared to maintenance or incremental updates, they capture attention. Being involved in these high-profile projects can position you as a forward-thinker, someone driving innovation or tackling significant challenges, even if the actual impact is more subtle. The key is finding a balance: embrace these promotional opportunities while also delivering meaningful, substantive work.</p>
<h2 id="4-build-relationships-across-teams" class="font-display text-2xl lg:text-3xl font-bold text-foreground mt-12 mb-6 border-l-2 border-primary pl-4"><a href="#4-build-relationships-across-teams" class="heading-permalink text-inherit no-underline">4. Build Relationships Across Teams</a></h2>
<p class="text-muted-foreground leading-7 my-4">We&#x27;ve all heard the saying, &quot;It&#x27;s not what you know, but who you know.&quot; While skills and performance are crucial, relationships can play a big role in career growth.</p>
<p class="text-muted-foreground leading-7 my-4">Building relationships isn&#x27;t about networking in the stereotypical sense (no one likes forced LinkedIn messages). It&#x27;s about creating genuine connections with colleagues, managers, and leaders—within <strong class="text-foreground font-bold">and beyond</strong> your immediate team.</p>
<p class="text-muted-foreground leading-7 my-4">Why does this matter?</p>
<ul class="space-y-2 my-6 ml-4">
<li class="flex gap-3 text-muted-foreground leading-7"><span class="text-primary shrink-0 mt-1.5">▸</span><span class="min-w-0"><strong class="text-foreground font-bold">Advocates and allies:</strong> People who see your potential can vouch for you during promotion or hiring discussions.</span></li>
<li class="flex gap-3 text-muted-foreground leading-7"><span class="text-primary shrink-0 mt-1.5">▸</span><span class="min-w-0"><strong class="text-foreground font-bold">Opportunities:</strong> Opportunities often come from unexpected places, like collaborations or referrals from someone in another department.</span></li>
<li class="flex gap-3 text-muted-foreground leading-7"><span class="text-primary shrink-0 mt-1.5">▸</span><span class="min-w-0"><strong class="text-foreground font-bold">Broader perspective:</strong> Understanding the challenges and goals of other teams enhances your ability to make a bigger impact in your role.</span></li>
<li class="flex gap-3 text-muted-foreground leading-7"><span class="text-primary shrink-0 mt-1.5">▸</span><span class="min-w-0"><strong class="text-foreground font-bold">Promotion support during panels:</strong> Building relationships across teams can be crucial when promotion panels or reviews occur. Colleagues and leaders from other areas of the organization who know your work and value your contributions can provide key insights and endorsements.</span></li>
</ul>
<p class="text-muted-foreground leading-7 my-4">Start simple. Attend cross-team meetings, schedule a coffee chat with someone who inspires you, or offer to help another department with your unique skillset. Relationships are the bridge between where you are and where you want to go.</p>
<h2 id="5-ask-for-regular-feedback" class="font-display text-2xl lg:text-3xl font-bold text-foreground mt-12 mb-6 border-l-2 border-primary pl-4"><a href="#5-ask-for-regular-feedback" class="heading-permalink text-inherit no-underline">5. Ask for Regular Feedback</a></h2>
<p class="text-muted-foreground leading-7 my-4">Here&#x27;s something I didn&#x27;t realize early enough: Feedback isn&#x27;t criticism; it&#x27;s guidance. Regular feedback from managers and peers can help you refine your skills, avoid blind spots, and identify areas to improve before they become obstacles.</p>
<p class="text-muted-foreground leading-7 my-4">Don&#x27;t wait for your annual performance review to ask, &quot;How am I doing?&quot; Instead, actively seek feedback throughout the year:</p>
<ul class="space-y-2 my-6 ml-4">
<li class="flex gap-3 text-muted-foreground leading-7"><span class="text-primary shrink-0 mt-1.5">▸</span><span class="min-w-0">After delivering a major project or presentation, ask what went well and what could have been improved.</span></li>
<li class="flex gap-3 text-muted-foreground leading-7"><span class="text-primary shrink-0 mt-1.5">▸</span><span class="min-w-0">During 1-on-1 meetings, invite input on your performance and discuss how you can position yourself for growth.</span></li>
<li class="flex gap-3 text-muted-foreground leading-7"><span class="text-primary shrink-0 mt-1.5">▸</span><span class="min-w-0">Reach out to trusted colleagues for peer feedback on how you collaborate and contribute.</span></li>
</ul>
<p class="text-muted-foreground leading-7 my-4">Remember, feedback is a dialogue, not a one-way street. Use it as a tool to learn, adapt, and grow continuously. Your willingness to evolve demonstrates maturity and leadership potential.</p>
<h2 id="bonus-know-your-value-and-advocate-for-fair-compensation" class="font-display text-2xl lg:text-3xl font-bold text-foreground mt-12 mb-6 border-l-2 border-primary pl-4"><a href="#bonus-know-your-value-and-advocate-for-fair-compensation" class="heading-permalink text-inherit no-underline">Bonus: Know Your Value and Advocate for Fair Compensation</a></h2>
<p class="text-muted-foreground leading-7 my-4">Understanding your worth in the market is just as vital as pursuing career growth. Your skills, expertise, and contributions have tangible value, and recognizing this is the first step toward advocating for yourself. Research industry standards, salaries within your field, and the pay scales for your role to arm yourself with the knowledge needed to have open and constructive conversations about compensation.</p>
<p class="text-muted-foreground leading-7 my-4">While these discussions can feel daunting, they are necessary. Approach them with confidence and professionalism, outlining the impact you&#x27;ve made and why it merits fair recompense. Remember, advocating for your worth is not selfish—it&#x27;s about ensuring a balanced exchange for your work and dedication.</p>
<h2 id="dont-wait-own-your-narrative" class="font-display text-2xl lg:text-3xl font-bold text-foreground mt-12 mb-6 border-l-2 border-primary pl-4"><a href="#dont-wait-own-your-narrative" class="heading-permalink text-inherit no-underline">Don&#x27;t Wait—Own Your Narrative</a></h2>
<p class="text-muted-foreground leading-7 my-4">Here&#x27;s the truth no one (except maybe your most honest mentor) tells you: Career growth isn&#x27;t just about being good at what you do. It&#x27;s about showing your value, proving your contributions, and taking ownership of your narrative.</p>
<ul class="space-y-2 my-6 ml-4">
<li class="flex gap-3 text-muted-foreground leading-7"><span class="text-primary shrink-0 mt-1.5">▸</span><span class="min-w-0">Advocate for yourself.</span></li>
<li class="flex gap-3 text-muted-foreground leading-7"><span class="text-primary shrink-0 mt-1.5">▸</span><span class="min-w-0">Showcase your skills and initiative.</span></li>
<li class="flex gap-3 text-muted-foreground leading-7"><span class="text-primary shrink-0 mt-1.5">▸</span><span class="min-w-0">Build relationships that bolster your career.</span></li>
<li class="flex gap-3 text-muted-foreground leading-7"><span class="text-primary shrink-0 mt-1.5">▸</span><span class="min-w-0">Seek opportunities to learn, improve, and grow.</span></li>
</ul>
<p class="text-muted-foreground leading-7 my-4">Your next promotion? It&#x27;s not a milestone in someone else&#x27;s timeline. It&#x27;s your achievement to claim. The moment you decide to take control of your career, you level up—a mindset, a skill, a role at a time.</p>
<p class="text-muted-foreground leading-7 my-4">Start now. Decide what&#x27;s next for you, map your path, and show the world what you&#x27;re capable of.</p>]]></content:encoded>
      <pubDate>Sat, 15 Feb 2025 12:00:00 GMT</pubDate>
    </item>
    <item>
      <title>The Power of Saying No</title>
      <link>https://pratik.pa.tel/blog/the-power-of-saying-no/</link>
      <guid isPermaLink="true">https://pratik.pa.tel/blog/the-power-of-saying-no/</guid>
      <description>Saying yes to everything doesn&apos;t make you a team player. It makes you the bottleneck. How I learned to say no, and the diplomatic versions that work better.</description>
      <dc:creator>Pratik Patel</dc:creator>
      <category>leadership</category>
      <category>career</category>
      <content:encoded><![CDATA[<p class="text-muted-foreground leading-7 my-4">Early in my career, I was <em class="font-mono font-normal text-primary/80 print:text-primary">that</em> engineer. The &quot;Yes Person™.&quot;</p>
<p class="text-muted-foreground leading-7 my-4">&quot;Can you take on this extra feature?&quot; <em class="font-mono font-normal text-primary/80 print:text-primary">Yes!</em></p>
<p class="text-muted-foreground leading-7 my-4">&quot;Can we launch a week early?&quot; <em class="font-mono font-normal text-primary/80 print:text-primary">Sure! Why not!</em></p>
<p class="text-muted-foreground leading-7 my-4">&quot;Can you hop on a quick call at 9 PM?&quot; <em class="font-mono font-normal text-primary/80 print:text-primary">Of course!</em></p>
<p class="text-muted-foreground leading-7 my-4">I thought saying yes to everything made me a team player—someone reliable, indispensable, and on-track for all the accolades. But more often than not, those yeses led me to <strong class="text-foreground font-bold">tight deadlines, late nights, and some… creative technical workarounds</strong> (you know, the kind that makes future-you weep).</p>
<p class="text-muted-foreground leading-7 my-4">The turning point? Learning to say <em class="font-mono font-normal text-primary/80 print:text-primary">no</em>. Or more often, the more diplomatic cousin of no:</p>
<ul class="space-y-2 my-6 ml-4">
<li class="flex gap-3 text-muted-foreground leading-7"><span class="text-primary shrink-0 mt-1.5">▸</span><span class="min-w-0">&quot;Yes, but we&#x27;ll need to adjust the timeline.&quot;</span></li>
<li class="flex gap-3 text-muted-foreground leading-7"><span class="text-primary shrink-0 mt-1.5">▸</span><span class="min-w-0">&quot;Yes, if we drop another lower-priority task.&quot;</span></li>
<li class="flex gap-3 text-muted-foreground leading-7"><span class="text-primary shrink-0 mt-1.5">▸</span><span class="min-w-0">&quot;No, because that would introduce tech debt that will haunt us forever.&quot;</span></li>
</ul>
<p class="text-muted-foreground leading-7 my-4">Turns out, strategic no&#x27;s transformed me into a <strong class="text-foreground font-bold">better engineer, a better teammate, and a much more effective professional overall</strong>. Here&#x27;s why—and how you can harness the power of no to elevate your own work and sanity.</p>
<h2 id="the-problem-with-always-saying-yes" class="font-display text-2xl lg:text-3xl font-bold text-foreground mt-12 mb-6 border-l-2 border-primary pl-4"><a href="#the-problem-with-always-saying-yes" class="heading-permalink text-inherit no-underline">The Problem with Always Saying Yes</a></h2>
<p class="text-muted-foreground leading-7 my-4">The tech world often glorifies the Yes Person. They&#x27;re seen as flexible, eager, and a team player. But saying yes to everything has some hidden costs, particularly in fields like engineering and project management, where the stakes are consistently high.</p>
<p class="text-muted-foreground leading-7 my-4">Here&#x27;s what unchecked yeses can lead to:</p>
<ul class="space-y-2 my-6 ml-4">
<li class="flex gap-3 text-muted-foreground leading-7"><span class="text-primary shrink-0 mt-1.5">▸</span><span class="min-w-0"><strong class="text-foreground font-bold">Burnout:</strong> Saying yes to more commitments than you have time for inevitably results in exhaustion. Constant late nights and rushed work don&#x27;t lead to personal or professional growth—they just empty your tank.</span></li>
<li class="flex gap-3 text-muted-foreground leading-7"><span class="text-primary shrink-0 mt-1.5">▸</span><span class="min-w-0"><strong class="text-foreground font-bold">Technical Debt:</strong> Agreeing to overly ambitious deadlines often means cutting corners. And what looks like success in the short term will leave your team paying interest on those decisions for years.</span></li>
<li class="flex gap-3 text-muted-foreground leading-7"><span class="text-primary shrink-0 mt-1.5">▸</span><span class="min-w-0"><strong class="text-foreground font-bold">Chaos over Quality:</strong> When everything is a priority, <em class="font-mono font-normal text-primary/80 print:text-primary">nothing</em> is a priority. Projects finished under the weight of &quot;yes to everything&quot; tend to lack the structure and finesse that truly stand out.</span></li>
</ul>
<p class="text-muted-foreground leading-7 my-4">Ultimately, saying yes to everything doesn&#x27;t make you a hero; it makes you a hazard—to yourself and your team.</p>
<h2 id="why-saying-no-is-a-superpower" class="font-display text-2xl lg:text-3xl font-bold text-foreground mt-12 mb-6 border-l-2 border-primary pl-4"><a href="#why-saying-no-is-a-superpower" class="heading-permalink text-inherit no-underline">Why Saying No is a Superpower</a></h2>
<p class="text-muted-foreground leading-7 my-4">Saying no strategically isn&#x27;t just about protecting your sanity (although, that&#x27;s important too). It&#x27;s about ensuring long-term success—for yourself, your projects, and your team.</p>
<p class="text-muted-foreground leading-7 my-4">Here&#x27;s what saying no can do for you:</p>
<ul class="space-y-2 my-6 ml-4">
<li class="flex gap-3 text-muted-foreground leading-7"><span class="text-primary shrink-0 mt-1.5">▸</span><span class="min-w-0"><strong class="text-foreground font-bold">Prioritization:</strong> Establishing boundaries allows you to focus on the tasks that truly matter. Delivering one product milestone well beats delivering five poorly.</span></li>
<li class="flex gap-3 text-muted-foreground leading-7"><span class="text-primary shrink-0 mt-1.5">▸</span><span class="min-w-0"><strong class="text-foreground font-bold">Preserving Quality:</strong> Fewer commitments mean more time to work thoughtfully and effectively on the things that count. Instead of sweating over quick fixes, you can build something to be proud of.</span></li>
<li class="flex gap-3 text-muted-foreground leading-7"><span class="text-primary shrink-0 mt-1.5">▸</span><span class="min-w-0"><strong class="text-foreground font-bold">Realistic Planning:</strong> Pushing back helps prevent impossible timelines and stops burnout culture in its tracks. Deadlines grounded in reality are better for you, your team, and—surprise—for your leadership&#x27;s trust in you.</span></li>
</ul>
<p class="text-muted-foreground leading-7 my-4">The key lesson? Saying no isn&#x27;t about being difficult or resistant; it&#x27;s about ensuring decisions align with real constraints, shared goals, and long-term success.</p>
<h2 id="how-to-say-no-without-burning-bridges" class="font-display text-2xl lg:text-3xl font-bold text-foreground mt-12 mb-6 border-l-2 border-primary pl-4"><a href="#how-to-say-no-without-burning-bridges" class="heading-permalink text-inherit no-underline">How to Say No Without Burning Bridges</a></h2>
<p class="text-muted-foreground leading-7 my-4">If the thought of saying no sends a wave of anxiety through your body, don&#x27;t worry—you&#x27;re not alone. Engineers, tech leads, and project managers alike often worry about how saying no might damage their reputation or relationships. But saying no <strong class="text-foreground font-bold">doesn&#x27;t have to come across as negative</strong>.</p>
<p class="text-muted-foreground leading-7 my-4">Here are practical ways you can say no while still being collaborative and solution-driven:</p>
<h3 id="1-frame-it-as-a-yes-but" class="font-display text-xl font-bold text-foreground mt-10 mb-4"><a href="#1-frame-it-as-a-yes-but" class="heading-permalink text-inherit no-underline">1. Frame It as a Yes-But</a></h3>
<p class="text-muted-foreground leading-7 my-4">This is one of my go-to strategies. When a request lands on your plate, consider responding with a conditional yes. For example:</p>
<ul class="space-y-2 my-6 ml-4">
<li class="flex gap-3 text-muted-foreground leading-7"><span class="text-primary shrink-0 mt-1.5">▸</span><span class="min-w-0">&quot;Yes, I can take this on, but we&#x27;ll need to extend the launch date by two weeks.&quot;</span></li>
<li class="flex gap-3 text-muted-foreground leading-7"><span class="text-primary shrink-0 mt-1.5">▸</span><span class="min-w-0">&quot;Yes, but for this to be delivered on time, we&#x27;ll have to postpone feature XYZ.&quot;</span></li>
</ul>
<p class="text-muted-foreground leading-7 my-4">This shows you&#x27;re being thoughtful and realistic while presenting solutions instead of obstacles.</p>
<h3 id="2-lean-on-data" class="font-display text-xl font-bold text-foreground mt-10 mb-4"><a href="#2-lean-on-data" class="heading-permalink text-inherit no-underline">2. Lean on Data</a></h3>
<p class="text-muted-foreground leading-7 my-4">Pushback becomes a lot easier when it&#x27;s backed by facts. For example:</p>
<ul class="space-y-2 my-6 ml-4">
<li class="flex gap-3 text-muted-foreground leading-7"><span class="text-primary shrink-0 mt-1.5">▸</span><span class="min-w-0">&quot;No, because adding this feature would introduce latency that exceeds our performance benchmarks.&quot;</span></li>
<li class="flex gap-3 text-muted-foreground leading-7"><span class="text-primary shrink-0 mt-1.5">▸</span><span class="min-w-0">&quot;No, because we&#x27;ve already committed 12 developer hours to the sprint, and squeezing this in would put us over capacity.&quot;</span></li>
</ul>
<p class="text-muted-foreground leading-7 my-4">Data makes the conversation less personal and more objective, which helps teammates and stakeholders understand your reasoning.</p>
<h3 id="3-prioritize-tech-debt" class="font-display text-xl font-bold text-foreground mt-10 mb-4"><a href="#3-prioritize-tech-debt" class="heading-permalink text-inherit no-underline">3. Prioritize Tech Debt</a></h3>
<p class="text-muted-foreground leading-7 my-4">Get comfortable with saying things like:</p>
<ul class="space-y-2 my-6 ml-4">
<li class="flex gap-3 text-muted-foreground leading-7"><span class="text-primary shrink-0 mt-1.5">▸</span><span class="min-w-0">&quot;No, because rushing this will create tech debt that will haunt us down the line.&quot;</span></li>
<li class="flex gap-3 text-muted-foreground leading-7"><span class="text-primary shrink-0 mt-1.5">▸</span><span class="min-w-0">&quot;No, because this violates our codebase standard and will slow development in future sprints.&quot;</span></li>
</ul>
<p class="text-muted-foreground leading-7 my-4">Every tech lead or product owner encountering these perfectly logical reasons will breathe a secret sigh of relief (whether they admit it or not).</p>
<h3 id="4-practice-transparency" class="font-display text-xl font-bold text-foreground mt-10 mb-4"><a href="#4-practice-transparency" class="heading-permalink text-inherit no-underline">4. Practice Transparency</a></h3>
<p class="text-muted-foreground leading-7 my-4">Don&#x27;t just say no; explain the &quot;why.&quot; Transparency builds trust. For instance:</p>
<ul class="space-y-2 my-6 ml-4">
<li class="flex gap-3 text-muted-foreground leading-7"><span class="text-primary shrink-0 mt-1.5">▸</span><span class="min-w-0">&quot;No, because I&#x27;m already loaded with task A and task B, and compromising quality isn&#x27;t something I&#x27;m comfortable with.&quot;</span></li>
<li class="flex gap-3 text-muted-foreground leading-7"><span class="text-primary shrink-0 mt-1.5">▸</span><span class="min-w-0">&quot;No, because this request requires resources we don&#x27;t currently have allocated.&quot;</span></li>
</ul>
<p class="text-muted-foreground leading-7 my-4">When people understand your reason, they&#x27;re more likely to agree with your decision.</p>
<h3 id="5-default-to-diplomacy" class="font-display text-xl font-bold text-foreground mt-10 mb-4"><a href="#5-default-to-diplomacy" class="heading-permalink text-inherit no-underline">5. Default to Diplomacy</a></h3>
<p class="text-muted-foreground leading-7 my-4">Sometimes, the classic &quot;no&quot; can feel abrupt. Instead, try softer alternatives that maintain rapport:</p>
<ul class="space-y-2 my-6 ml-4">
<li class="flex gap-3 text-muted-foreground leading-7"><span class="text-primary shrink-0 mt-1.5">▸</span><span class="min-w-0">&quot;That&#x27;s a great idea, but I think we should revisit it after we finish XYZ.&quot;</span></li>
<li class="flex gap-3 text-muted-foreground leading-7"><span class="text-primary shrink-0 mt-1.5">▸</span><span class="min-w-0">&quot;I wish I could help, but unfortunately, I can&#x27;t commit right now.&quot;</span></li>
</ul>
<p class="text-muted-foreground leading-7 my-4">You&#x27;re upholding boundaries without creating friction.</p>
<h2 id="saying-no-is-saying-yes-to-better-things" class="font-display text-2xl lg:text-3xl font-bold text-foreground mt-12 mb-6 border-l-2 border-primary pl-4"><a href="#saying-no-is-saying-yes-to-better-things" class="heading-permalink text-inherit no-underline">Saying No is Saying Yes (to Better Things)</a></h2>
<p class="text-muted-foreground leading-7 my-4">Every time you say no, you&#x27;re simultaneously saying yes—to focus, quality, precision, and meaningful work. Balancing commitments isn&#x27;t just healthier for you; it also strengthens your team dynamic and ensures your engineering output always shines.</p>
<p class="text-muted-foreground leading-7 my-4">The ironic truth? Learning to say no strategically makes you far more valuable—far more indispensable—than constant agreement ever could.</p>
<p class="text-muted-foreground leading-7 my-4">If you&#x27;re reading this and thinking, &quot;I <em class="font-mono font-normal text-primary/80 print:text-primary">am</em> the Yes Person,&quot; don&#x27;t worry. We&#x27;ve all been there. The beauty of this lesson is that it&#x27;s never too late to start sprinkling in some thoughtful no&#x27;s. The difference will amaze you.</p>
<h2 id="final-thoughts" class="font-display text-2xl lg:text-3xl font-bold text-foreground mt-12 mb-6 border-l-2 border-primary pl-4"><a href="#final-thoughts" class="heading-permalink text-inherit no-underline">Final Thoughts</a></h2>
<p class="text-muted-foreground leading-7 my-4">Whether you&#x27;re an engineer, a tech lead, or a project manager, mastering the art of saying no is one of the most powerful skills you can develop. Start small—be strategic—and watch how prioritizing quality over chaos transforms your career.</p>
<p class="text-muted-foreground leading-7 my-4">Remember, saying no isn&#x27;t about shutting doors; it&#x27;s about opening the right ones. Your best work—your truly impactful work—awaits on the other side.</p>]]></content:encoded>
      <pubDate>Mon, 10 Feb 2025 12:00:00 GMT</pubDate>
    </item>
  </channel>
</rss>
