<?xml version="1.0" encoding="utf-8"?><?xml-stylesheet type="text/xsl" href="atom.xsl"?>
<feed xmlns="http://www.w3.org/2005/Atom">
    <id>https://hassant.github.io</id>
    <title>/dev/hassan Blog</title>
    <updated>2026-08-04T00:00:00.000Z</updated>
    <generator>https://github.com/jpmonette/feed</generator>
    <link rel="alternate" href="https://hassant.github.io"/>
    <subtitle>/dev/hassan Blog</subtitle>
    <icon>https://hassant.github.io/img/favicon.ico</icon>
    <entry>
        <title type="html"><![CDATA[A Human Checkpoint Is Not an AI Safety System]]></title>
        <id>https://hassant.github.io/human-checkpoint-is-not-a-safety-system</id>
        <link href="https://hassant.github.io/human-checkpoint-is-not-a-safety-system"/>
        <updated>2026-08-04T00:00:00.000Z</updated>
        <summary type="html"><![CDATA[A friend in a regulated industry recently described their AI governance with complete confidence: "A human approves every output."]]></summary>
        <content type="html"><![CDATA[<p>A friend in a regulated industry recently described their AI governance with complete confidence: "A human approves every output."</p>
<p>I asked how many outputs a day. About four thousand. I asked how long each approval takes. A pause. "A few seconds."</p>
<p>That is not a safety system. It is a queue with a person sitting in front of it.</p>
<!-- -->
<p>"Human in the loop" has become the industry's comfort phrase. It appears in responsible AI decks, enterprise sales pitches, and governance plans because it sounds like control. Put a person between the model and the action, and somebody is accountable.</p>
<p>The trouble is that the human is often right where the system is weakest: watching a mostly-correct machine for long enough that attention becomes ceremonial.</p>
<p>The person is still in the process. Their judgment is no longer carrying the weight we pretend it is.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="why-rubber-stamp-oversight-fails">Why rubber-stamp oversight fails<a href="https://hassant.github.io/human-checkpoint-is-not-a-safety-system#why-rubber-stamp-oversight-fails" class="hash-link" aria-label="Direct link to Why rubber-stamp oversight fails" title="Direct link to Why rubber-stamp oversight fails" translate="no">​</a></h2>
<p>The problem is not that people are careless. The job is badly designed.</p>
<p><strong>Automation bias</strong> makes people over-trust automated recommendations, even when contradictory evidence is available. In a <a href="https://doi.org/10.1109/HRI.2016.7451740" target="_blank" rel="noopener noreferrer" class="">2016 emergency evacuation study</a>, participants followed a robot toward the wrong exit despite having seen it behave unreliably.</p>
<p><strong>Vigilance decrement</strong> makes passive supervision worse over time. A system that is correct often enough trains the operator to stop expecting the rare failure. When that failure arrives, the person has the least context and the least practice at exactly the moment intervention matters.</p>
<p><strong>The moral crumple zone</strong>, Madeleine Clare Elish's phrase, describes what happens after failure. The human operator absorbs the legal, moral, and reputational impact even when the system and the institution gave them little practical ability to prevent it. The person becomes a liability sponge.</p>
<p><strong>The scale math fails.</strong> Four thousand decisions multiplied by a few seconds each does not produce four thousand acts of judgment. It produces a habit.</p>
<p>The <a href="https://eur-lex.europa.eu/eli/reg/2024/1689/oj/eng" target="_blank" rel="noopener noreferrer" class="">EU AI Act's Article 14</a> is more nuanced than the slogan. For high-risk AI systems, it requires effective human oversight measures and explicitly calls out automation bias. Overseers need enough competence, authority, and system support to understand limitations, monitor operation, override output, and stop the system. It does not generally require a person to approve every output, although it does impose stricter confirmation requirements for a narrow class of remote biometric identification systems.</p>
<p>A person with no time, weak context, and no real stop authority is not meaningful oversight.</p>
<p>They are evidence that a checkbox was filled.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="oversight-is-a-risk-decision-not-a-ladder">Oversight is a risk decision, not a ladder<a href="https://hassant.github.io/human-checkpoint-is-not-a-safety-system#oversight-is-a-risk-decision-not-a-ladder" class="hash-link" aria-label="Direct link to Oversight is a risk decision, not a ladder" title="Direct link to Oversight is a risk decision, not a ladder" translate="no">​</a></h2>
<p>It is tempting to draw a neat progression:</p>
<ol>
<li class="">human in the loop;</li>
<li class="">human on the loop;</li>
<li class="">human above the loop.</li>
</ol>
<p>The third phrase is the label I use for the person who designs the gates, owns the evidence, and has authority to stop the system. It is useful, but it is not a universal replacement for the other two.</p>
<p>Different actions need different controls.</p>
<ul>
<li class="">A reversible draft can be generated automatically and reviewed later.</li>
<li class="">A production configuration change may need monitoring, a rollback, and a named owner.</li>
<li class="">A payment, dismissal, medical decision, or destructive action may still require explicit approval before execution.</li>
<li class="">A high-volume, low-risk classification task may be safer with automated thresholds and sampled review than with thousands of rushed clicks.</li>
</ul>
<p>The right question is not "where is the human?" It is "where does human judgment change the outcome, and does that person have the evidence, time, and authority to act?"</p>
<!-- -->
<p>This is the difference between putting a human inside every decision and making humans responsible for the design of the decision system.</p>
<p>The first consumes attention. The second compounds judgment.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="what-the-frontier-frameworks-actually-teach">What the frontier frameworks actually teach<a href="https://hassant.github.io/human-checkpoint-is-not-a-safety-system#what-the-frontier-frameworks-actually-teach" class="hash-link" aria-label="Direct link to What the frontier frameworks actually teach" title="Direct link to What the frontier frameworks actually teach" translate="no">​</a></h2>
<p>Frontier model safety frameworks are not the same thing as enterprise agent governance. One deals with catastrophic model capabilities and deployment thresholds; the other often deals with tools, data, workflows, and business actions.</p>
<p>They are still useful analogies because the control shape is similar: measure capability, define thresholds, assign owners, bind thresholds to mitigations, keep evidence, and rehearse escalation.</p>
<p>Anthropic's <a href="https://www.anthropic.com/responsible-scaling-policy/rsp-v3-0" target="_blank" rel="noopener noreferrer" class="">Responsible Scaling Policy v3.0</a> uses capability thresholds, containment measures, a Responsible Scaling Officer, and escalation processes. It is important not to romanticize it. The <a href="https://www.governance.ai/analysis/anthropics-rsp-v3-0-how-it-works-whats-changed-and-some-reflections" target="_blank" rel="noopener noreferrer" class="">GovAI analysis</a> notes that v3.0 dropped the earlier general pause commitment and reframed some of the strongest security measures as industry-wide recommendations.</p>
<p>OpenAI's <a href="https://openai.com/index/updating-our-preparedness-framework/" target="_blank" rel="noopener noreferrer" class="">Preparedness Framework</a> similarly separates capability assessment from safeguards and routes decisions through a formal review structure.</p>
<p>Google DeepMind's <a href="https://deepmind.google/discover/blog/introducing-the-frontier-safety-framework/" target="_blank" rel="noopener noreferrer" class="">Frontier Safety Framework</a> defines capability levels and mitigations for severe risks.</p>
<p>An <a href="https://arxiv.org/abs/2512.01166" target="_blank" rel="noopener noreferrer" class="">independent evaluation of twelve providers</a> is the useful cold shower. It found weak commitments across both organizational governance and hard technical assurance. Internal audit, named risk ownership, risk tolerances, escalation, and assurance against loss-of-control risks remain uneven or underdeveloped.</p>
<p>The lesson is not that the labs solved governance. They did not.</p>
<p>The lesson is that serious safety work looks like thresholds, evidence, owners, mitigations, and escalation. It does not look like a tired person approving four thousand outputs.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="identity-is-where-responsibility-becomes-architecture">Identity is where responsibility becomes architecture<a href="https://hassant.github.io/human-checkpoint-is-not-a-safety-system#identity-is-where-responsibility-becomes-architecture" class="hash-link" aria-label="Direct link to Identity is where responsibility becomes architecture" title="Direct link to Identity is where responsibility becomes architecture" translate="no">​</a></h2>
<p>Every control above assumes you can answer a simple question:</p>
<p><strong>Who did that?</strong></p>
<p>In many enterprises, the honest answer is still "a shared service account created years ago, used by several pipelines, carrying a credential nobody wants to rotate."</p>
<p>That is not just an identity hygiene problem. It breaks governance.</p>
<p>You cannot reliably:</p>
<ul>
<li class="">attribute an action to a specific agent run;</li>
<li class="">separate the agent's authority from the user's authority;</li>
<li class="">revoke one compromised actor without breaking unrelated systems;</li>
<li class="">prove which policy and permission set applied;</li>
<li class="">route an incident to a named organizational owner.</li>
</ul>
<p>Identity does not create responsibility by itself. An organization remains accountable for a system even when the logs are poor. But a first-class agent identity gives responsibility somewhere concrete to attach.</p>
<p>Microsoft's <a href="https://learn.microsoft.com/en-us/entra/agent-id/what-are-agent-identities" target="_blank" rel="noopener noreferrer" class="">Entra Agent ID</a> distinguishes first-class agent identities from agents built on ordinary application identities and service principals. The important idea is broader than one product: an autonomous agent should not quietly borrow a human session or share a long-lived credential with unrelated workloads.</p>
<p>The audit record should be able to say:</p>
<ul>
<li class="">agent X acted;</li>
<li class="">for user or process Y;</li>
<li class="">during run Z;</li>
<li class="">with this bounded authority;</li>
<li class="">under this policy version;</li>
<li class="">and this named owner was accountable for the system.</li>
</ul>
<!-- -->
<p>This is where the phrase <strong>least agency</strong> becomes useful. Least privilege limits what an identity can access. Least agency also limits how freely the system may act with that access.</p>
<p>An agent may be allowed to read a repository but not push. It may draft a payment but not submit it. It may open a pull request but not merge it. It may inspect production telemetry but need approval before changing production.</p>
<p>Autonomy should be granted in bounded slices, with evidence and rollback attached.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="the-threat-model-is-already-here">The threat model is already here<a href="https://hassant.github.io/human-checkpoint-is-not-a-safety-system#the-threat-model-is-already-here" class="hash-link" aria-label="Direct link to The threat model is already here" title="Direct link to The threat model is already here" translate="no">​</a></h2>
<p>The <a href="https://genai.owasp.org/resource/owasp-top-10-for-agentic-applications-for-2026/" target="_blank" rel="noopener noreferrer" class="">OWASP Top 10 for Agentic Applications 2026</a>, published in December 2025, includes <strong>ASI03: Identity and Privilege Abuse</strong>. The reason is straightforward: excessive or ambiguous authority turns every other failure into a larger incident.</p>
<p>EchoLeak, <a href="https://msrc.microsoft.com/update-guide/vulnerability/CVE-2025-32711" target="_blank" rel="noopener noreferrer" class="">CVE-2025-32711</a>, demonstrated a zero-click indirect prompt injection against Microsoft 365 Copilot. A crafted email could plant instructions that the system later retrieved as context.</p>
<p>The <a href="https://aws.amazon.com/security/security-bulletins/AWS-2025-015/" target="_blank" rel="noopener noreferrer" class="">Amazon Q Developer extension incident</a> showed a different path. An over-scoped GitHub token let an attacker commit agent-directed instructions into the extension's repository. Those instructions shipped in version 1.84.0 and told the agent to damage the local system. The payload did not execute because of a syntax error, and AWS reported no customer impact.</p>
<p>That syntax error was luck, not a control.</p>
<p>In both cases, the potential blast radius was defined by the tools, credentials, and data available to the agent. Prompt injection is dangerous. Prompt injection connected to broad, long-lived authority is dangerous at machine speed.</p>
<p>The standards work is moving in the same direction. The IETF's current <a href="https://datatracker.ietf.org/doc/draft-ietf-oauth-transaction-tokens/" target="_blank" rel="noopener noreferrer" class="">OAuth Transaction Tokens draft</a> explores short-lived, purpose-bound tokens for transactions. An individual Internet-Draft, <a href="https://datatracker.ietf.org/doc/draft-klrc-aiagent-auth/" target="_blank" rel="noopener noreferrer" class="">AI Agent Authentication and Authorization</a>, composes workload identity and OAuth patterns into an agent-focused architecture. These are works in progress, not settled standards, but the direction is sensible: replace static shared secrets with attributable, short-lived, attenuated authority.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="the-person-needs-teeth">The person needs teeth<a href="https://hassant.github.io/human-checkpoint-is-not-a-safety-system#the-person-needs-teeth" class="hash-link" aria-label="Direct link to The person needs teeth" title="Direct link to The person needs teeth" translate="no">​</a></h2>
<p>There is a fear underneath a lot of responsible AI language that the trajectory is to make humans decorative.</p>
<p>First the human writes the work. Then the human reviews it. Then the human rubber-stamps it. Finally, when the system fails, the human becomes the moral crumple zone.</p>
<p>That is not meaningful human oversight.</p>
<p>Meaningful oversight gives specific people:</p>
<ul>
<li class="">authority to stop or constrain the system;</li>
<li class="">competence to understand its capabilities and failure modes;</li>
<li class="">evidence that does not depend on the model's own summary;</li>
<li class="">time to make the decisions that genuinely require judgment;</li>
<li class="">an organizational environment that supports intervention rather than punishing delay.</li>
</ul>
<p>The human should not be the fuse inside the system. The organization should build controls that make human judgment load-bearing where it matters.</p>
<p>That includes explicit approval for irreversible actions. It also includes policy engines, evals, runtime monitors, bounded identities, audit trails, rollback, and named ownership.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="the-field-map">The field map<a href="https://hassant.github.io/human-checkpoint-is-not-a-safety-system#the-field-map" class="hash-link" aria-label="Direct link to The field map" title="Direct link to The field map" translate="no">​</a></h2>
<ul>
<li class=""><strong>A checkpoint is not automatically oversight.</strong> If the reviewer has seconds, weak context, or no stop authority, the approval is theatre.</li>
<li class=""><strong>Put humans where judgment changes the outcome.</strong> Risk and reversibility should decide whether an action is automatic, monitored, sampled, or explicitly approved.</li>
<li class=""><strong>Identity enables attribution, not absolution.</strong> A distinct agent identity makes actions traceable; a named human and organization still own the system.</li>
<li class=""><strong>Least privilege is not enough.</strong> Limit both what the agent can access and how freely it can act with that access.</li>
<li class=""><strong>Safety is a system of controls.</strong> Thresholds, evals, permissions, monitoring, rollback, audit, and escalation work together. No single tired reviewer can replace them.</li>
</ul>
<p>The loop was never the safety net.</p>
<p>The safety net is the system of evidence, boundaries, and authority that gives a person the power to make it stop. 🛑</p>
<h3 class="anchor anchorTargetStickyNavbar_Vzrq" id="sources-and-further-reading">Sources and further reading<a href="https://hassant.github.io/human-checkpoint-is-not-a-safety-system#sources-and-further-reading" class="hash-link" aria-label="Direct link to Sources and further reading" title="Direct link to Sources and further reading" translate="no">​</a></h3>
<ul>
<li class=""><a href="https://eur-lex.europa.eu/eli/reg/2024/1689/oj/eng" target="_blank" rel="noopener noreferrer" class="">EU AI Act, official text</a></li>
<li class=""><a href="https://genai.owasp.org/resource/owasp-top-10-for-agentic-applications-for-2026/" target="_blank" rel="noopener noreferrer" class="">OWASP Top 10 for Agentic Applications 2026</a></li>
<li class=""><a href="https://learn.microsoft.com/en-us/entra/agent-id/what-are-agent-identities" target="_blank" rel="noopener noreferrer" class="">Microsoft Entra Agent ID</a></li>
</ul>]]></content>
        <author>
            <name>Hassan Tariq</name>
            <uri>https://github.com/hassant</uri>
        </author>
        <category label="ai-agents" term="ai-agents"/>
        <category label="governance" term="governance"/>
        <category label="identity" term="identity"/>
        <category label="responsible-ai" term="responsible-ai"/>
        <category label="security" term="security"/>
    </entry>
    <entry>
        <title type="html"><![CDATA[When Software Stopped Being a Recipe]]></title>
        <id>https://hassant.github.io/when-software-stopped-being-a-recipe</id>
        <link href="https://hassant.github.io/when-software-stopped-being-a-recipe"/>
        <updated>2026-08-03T00:00:00.000Z</updated>
        <summary type="html"><![CDATA[The problem, stated plainly: software engineering trained us to make behaviour repeatable, then handed us a component whose useful output has to be judged across a distribution.]]></summary>
        <content type="html"><![CDATA[<p>The problem, stated plainly: software engineering trained us to make behaviour repeatable, then handed us a component whose useful output has to be judged across a distribution.</p>
<p>That is not traditional development with a faster autocomplete. It changes what the important artifact is, what verification means, and what an engineer has to be good at.</p>
<!-- -->
<p>Stack Overflow's write-up of the 2025 Developer Survey names this directly: ask the same question twice and you can get two plausible answers with different structures and trade-offs. I felt the difference most clearly while reviewing an agent pipeline. Nothing had crashed. The same goal, tools, and context were still there. But one run chose a different route and produced a different implementation from the run before it.</p>
<p>Both answers looked reasonable.</p>
<p>That was the uncomfortable part.</p>
<p>With conventional code, we aim for repeatable behaviour under controlled state. Production systems have never been perfectly deterministic. They have clocks, networks, concurrency, randomness, and dependencies that change underneath them. But the implementation still gives us a concrete thing to inspect. We can trace a branch, reproduce a state, and assert an invariant.</p>
<p>LLM sampling adds a more visible kind of variability. Hosted model updates add a separate versioning risk. The same intent can produce several plausible outputs, and correctness is no longer just "did this line execute as expected?" It is also "does this result satisfy the goal, the constraints, and the standard we agreed on?"</p>
<p>The numbers reflect that tension. Stack Overflow's write-up of its <a href="https://stackoverflow.blog/2026/02/18/closing-the-developer-ai-trust-gap/" target="_blank" rel="noopener noreferrer" class="">2025 Developer Survey</a> says 84% of respondents use or plan to use AI tools, up from 76% the year before, while trust in AI accuracy fell to 29%. Their sharpest line is still the simplest:</p>
<blockquote>
<p>"It's what makes the field software engineering rather than software hoping-for-the-best."</p>
</blockquote>
<p>Adoption is rising. Trust is falling. That is not a contradiction. It is a profession discovering that a useful new component plays by different rules.</p>
<!-- -->
<p>Three shifts inside that diagram changed how I build, and one thing carries all three.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="the-scarce-artifact-moved">The scarce artifact moved<a href="https://hassant.github.io/when-software-stopped-being-a-recipe#the-scarce-artifact-moved" class="hash-link" aria-label="Direct link to The scarce artifact moved" title="Direct link to The scarce artifact moved" translate="no">​</a></h2>
<p>When code generation gets cheap, describing the right system becomes the expensive part.</p>
<p>I used to think of a prompt as the thing that kicked off the work. Now I think of the specification as the thing that makes the work governable. A useful spec has a goal, boundaries, examples, non-goals, acceptance evidence, and a stopping condition. It says what the agent may touch, what it must not touch, and what would prove the task is complete.</p>
<p>That sounds like requirements engineering because it is requirements engineering. The difference is that the spec is no longer a document people read before writing code. It is executable context for a system that can turn intent into implementation.</p>
<p>The tooling is catching up. <a href="https://github.com/github/spec-kit" target="_blank" rel="noopener noreferrer" class="">GitHub Spec Kit</a> and <a href="https://github.com/Fission-AI/OpenSpec" target="_blank" rel="noopener noreferrer" class="">OpenSpec</a> move work from a loose prompt toward durable specifications, plans, tasks, and implementation. <a href="https://tessl.io/" target="_blank" rel="noopener noreferrer" class="">Tessl</a> approaches the same pressure differently, through versioned agent skills, evals, governance, and evidence.</p>
<p>I learned this the boring way while building Agentry. The useful part of an agent contract was never the sentence describing the goal. It was everything around it:</p>
<ul>
<li class="">the boundary that prevented a helpful agent from touching the wrong files;</li>
<li class="">the evidence requirement that stopped a confident summary from standing in for a review;</li>
<li class="">the budget and stopping condition that turned a runaway loop into a bounded loss;</li>
<li class="">the identity and approval rules that made every action attributable.</li>
</ul>
<p>Each field felt like paperwork until a run needed it. Then it became the receipt.</p>
<p>I wrote about the same lesson from the autonomy side in <a class="" href="https://hassant.github.io/evidence-picks-the-level">You Don't Pick the Autonomy Level. Your Evidence Does.</a>. The short version is that reach is earned by proof. A larger model, a longer context window, or more agents do not change that.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="tests-did-not-disappear">Tests did not disappear<a href="https://hassant.github.io/when-software-stopped-being-a-recipe#tests-did-not-disappear" class="hash-link" aria-label="Direct link to Tests did not disappear" title="Direct link to Tests did not disappear" translate="no">​</a></h2>
<p>The lazy version of this argument says tests are obsolete and evals replace them.</p>
<p>That is wrong.</p>
<p>Tests assert specific invariants. They are still the best tool for questions such as:</p>
<ul>
<li class="">Did the parser reject malformed input?</li>
<li class="">Did the permission check deny an unauthorized call?</li>
<li class="">Did the API preserve the contract?</li>
<li class="">Did the retry stop at the configured limit?</li>
</ul>
<p>Evals answer a different class of question:</p>
<ul>
<li class="">Did the answer satisfy the user's intent?</li>
<li class="">Did it stay grounded in the supplied evidence?</li>
<li class="">Did it choose the right tool?</li>
<li class="">Did it avoid a known failure mode across a representative set of cases?</li>
<li class="">How often did quality fall below an acceptable threshold?</li>
</ul>
<p>A test can tell me the citation field exists. An eval has to tell me whether the citation supports the sentence.</p>
<p>That distinction matters because a green suite can still be looking in the wrong place. I have shipped code where every test passed because the doubles never had the shape of the real primitive. I wrote about that in <a class="" href="https://hassant.github.io/green-reviewed-still-wrong">Green, Reviewed, and Still Wrong</a>. The lesson carries directly into AI systems: verification is only as good as the cases, evidence, and failure shapes it covers.</p>
<!-- -->
<p>The hard part is that evals do not settle the question forever. The task distribution changes. Users find new edge cases. Grounding sources move. Providers update hosted models. A threshold that looked sensible in a lab can fail under production pressure.</p>
<p>So the eval suite is not a graduation exam. It is a living description of what failure looks like.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="the-engineer-became-the-director-and-the-verifier">The engineer became the director and the verifier<a href="https://hassant.github.io/when-software-stopped-being-a-recipe#the-engineer-became-the-director-and-the-verifier" class="hash-link" aria-label="Direct link to The engineer became the director and the verifier" title="Direct link to The engineer became the director and the verifier" translate="no">​</a></h2>
<p>Microsoft's <a href="https://www.microsoft.com/en-us/worklab/work-trend-index/agents-human-agency-and-the-opportunity-for-every-organization" target="_blank" rel="noopener noreferrer" class="">2026 Work Trend Index</a> frames the shift as expanded human agency: as agents take on more execution, the premium moves toward judgment, quality control, and redesigning work across people and agents. Anthropic's usage data shows a similar direction from the model side. In its <a href="https://www.anthropic.com/research/anthropic-economic-index-september-2025-report" target="_blank" rel="noopener noreferrer" class="">September 2025 Economic Index</a>, directive conversations on Claude.ai rose from 27% to 39% in eight months, while 77% of sampled first-party API business usage showed automation patterns.</p>
<p>The numbers are less interesting than the direction. People are delegating larger units of work.</p>
<p>That can look like the engineer is disappearing from the process. In practice, the job moves to the places where judgment is hardest to automate:</p>
<ol>
<li class="">deciding what deserves delegation;</li>
<li class="">writing the contract the system has to satisfy;</li>
<li class="">choosing evidence that does not depend on the agent's own confidence;</li>
<li class="">designing the rollback before increasing autonomy;</li>
<li class="">recognizing when a plausible result is still the wrong result.</li>
</ol>
<p>The last one is where most of my mistakes have lived.</p>
<p>In <a class="" href="https://hassant.github.io/certainty-is-the-bug">Certainty Is the Bug</a>, I described a review finding that sounded precise and urgent. It was also wrong about the failure it claimed. Then two of my own fixes were wrong in different ways. The review was useful, the tests were green, and none of that made the claims true until I reproduced them.</p>
<p>That is the new centre of the craft for me. Not generating more. Building a system in which claims have to arrive with evidence.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="the-harness-is-the-product">The harness is the product<a href="https://hassant.github.io/when-software-stopped-being-a-recipe#the-harness-is-the-product" class="hash-link" aria-label="Direct link to The harness is the product" title="Direct link to The harness is the product" translate="no">​</a></h2>
<p>The model gets the attention because it produces the visible result. The reliability lives around it.</p>
<p>The harness decides:</p>
<ul>
<li class="">what context enters the run;</li>
<li class="">which tools are available;</li>
<li class="">what identity and permissions the run receives;</li>
<li class="">which deterministic checks must pass;</li>
<li class="">which evals measure quality;</li>
<li class="">which actions require approval;</li>
<li class="">what gets logged;</li>
<li class="">when the system stops;</li>
<li class="">how a failed action is reversed.</li>
</ul>
<p>Prompt engineering matters inside that system, but it is only one control surface. A beautifully phrased prompt cannot compensate for a missing permission boundary, an eval set that rewards the wrong thing, or an audit trail that records only the final summary.</p>
<p>This is why "coding is dead" and "nothing changed, it is just autocomplete" are both unhelpful. The craft did not vanish, and it did not stay put. It moved up a layer.</p>
<p>We still need people who understand code. We also need people who can specify intent, model risk, design evidence, and direct systems that will not produce the same implementation every time.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="the-field-map">The field map<a href="https://hassant.github.io/when-software-stopped-being-a-recipe#the-field-map" class="hash-link" aria-label="Direct link to The field map" title="Direct link to The field map" translate="no">​</a></h2>
<ul>
<li class=""><strong>The spec is part of the runtime.</strong> Goal, boundaries, examples, non-goals, evidence, and stopping conditions are not planning decoration. They shape what the agent can produce.</li>
<li class=""><strong>Tests and evals answer different questions.</strong> Tests assert invariants. Evals estimate quality across representative cases. AI systems need both.</li>
<li class=""><strong>A green result is evidence, not a verdict.</strong> If the cases miss the real failure shape, the system can be confidently wrong.</li>
<li class=""><strong>The harness carries reliability.</strong> Context, permissions, checks, telemetry, rollback, and audit matter more than one clever prompt.</li>
<li class=""><strong>The engineer's job moved toward judgment.</strong> Delegation increases the value of deciding what correct looks like and proving that the system met it.</li>
</ul>
<p>Software did not stop needing engineers when it stopped behaving like a recipe.</p>
<p>It started asking us to become better at the parts we used to leave implicit. 🧭</p>
<h3 class="anchor anchorTargetStickyNavbar_Vzrq" id="sources-and-further-reading">Sources and further reading<a href="https://hassant.github.io/when-software-stopped-being-a-recipe#sources-and-further-reading" class="hash-link" aria-label="Direct link to Sources and further reading" title="Direct link to Sources and further reading" translate="no">​</a></h3>
<ul>
<li class=""><a href="https://stackoverflow.blog/2026/02/18/closing-the-developer-ai-trust-gap/" target="_blank" rel="noopener noreferrer" class="">Stack Overflow: Closing the developer AI trust gap</a></li>
<li class=""><a href="https://www.microsoft.com/en-us/worklab/work-trend-index/agents-human-agency-and-the-opportunity-for-every-organization" target="_blank" rel="noopener noreferrer" class="">Microsoft 2026 Work Trend Index</a></li>
<li class=""><a href="https://www.anthropic.com/research/anthropic-economic-index-september-2025-report" target="_blank" rel="noopener noreferrer" class="">Anthropic Economic Index, September 2025</a></li>
</ul>]]></content>
        <author>
            <name>Hassan Tariq</name>
            <uri>https://github.com/hassant</uri>
        </author>
        <category label="ai-agents" term="ai-agents"/>
        <category label="agentic" term="agentic"/>
        <category label="evals" term="evals"/>
        <category label="software-engineering" term="software-engineering"/>
        <category label="devtools" term="devtools"/>
    </entry>
    <entry>
        <title type="html"><![CDATA[Certainty Is the Bug: A Review, a Guard, and a Docstring That All Lied]]></title>
        <id>https://hassant.github.io/certainty-is-the-bug</id>
        <link href="https://hassant.github.io/certainty-is-the-bug"/>
        <updated>2026-07-05T00:00:00.000Z</updated>
        <summary type="html"><![CDATA[I put a side project through a hard review — a written list of findings, each one a specific claim about a specific line. Good review. It caught real things. It was also, in three places, confidently wrong. And when I sat down to fix the things it got right, two of my fixes were wrong too.]]></summary>
        <content type="html"><![CDATA[<p>I put a side project through a hard review — a written list of findings, each one a specific claim about a specific line. Good review. It caught real things. It was also, in three places, confidently wrong. And when I sat down to fix the things it got <em>right</em>, two of my fixes were wrong too.</p>
<p>The through-line wasn't sloppiness. It was <strong>certainty</strong> — the review's, and then mine.</p>
<!-- -->
<p>Last time I wrote that <a class="" href="https://hassant.github.io/green-reviewed-still-wrong">green doesn't mean done, even after a review</a>: a passing suite plus a careful read still shipped bugs that hid behind the shapes my tests never took. This is the sequel from the other side of the desk — what happens when you <em>do</em> get the review, take it seriously, and discover that a finding is not a fact, a guard is not a guarantee, and the word "never" in your own docstring is just a promise you haven't checked yet.</p>
<p>Three of them. None exotic. All the kind you nod along to and ship.</p>
<hr>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="the-finding-that-wasnt-a-crash">The finding that wasn't a crash<a href="https://hassant.github.io/certainty-is-the-bug#the-finding-that-wasnt-a-crash" class="hash-link" aria-label="Direct link to The finding that wasn't a crash" title="Direct link to The finding that wasn't a crash" translate="no">​</a></h2>
<p>Near the top of the review: <em>"This workshop slide tells people to use a grounding source that doesn't exist. With the new fail-closed behavior, copying it verbatim now raises an error and the agent won't run — a regression."</em></p>
<p>That reads as a hair-on-fire bug. A first lesson in a tutorial that hard-crashes is about the worst learner experience there is. I was one keystroke from writing the fix and moving on.</p>
<p>Instead I ran the thing the finding described:</p>
<div class="language-python codeBlockContainer_Ckt0 theme-code-block" style="--prism-color:#393A34;--prism-background-color:#f6f8fa"><div class="codeBlockContent_QJqH"><pre tabindex="0" class="prism-code language-python codeBlock_bY9V thin-scrollbar" style="color:#393A34;background-color:#f6f8fa"><code class="codeBlockLines_e6Vv"><div class="token-line" style="color:#393A34"><span class="token plain">cfg </span><span class="token operator" style="color:#393A34">=</span><span class="token plain"> </span><span class="token punctuation" style="color:#393A34">{</span><span class="token string" style="color:#e3116c">"grounding"</span><span class="token punctuation" style="color:#393A34">:</span><span class="token plain"> </span><span class="token punctuation" style="color:#393A34">{</span><span class="token string" style="color:#e3116c">"sources"</span><span class="token punctuation" style="color:#393A34">:</span><span class="token plain"> </span><span class="token punctuation" style="color:#393A34">[</span><span class="token string" style="color:#e3116c">"ghost-source"</span><span class="token punctuation" style="color:#393A34">]</span><span class="token punctuation" style="color:#393A34">,</span><span class="token plain"> </span><span class="token string" style="color:#e3116c">"require_packet"</span><span class="token punctuation" style="color:#393A34">:</span><span class="token plain"> </span><span class="token boolean" style="color:#36acaa">True</span><span class="token punctuation" style="color:#393A34">}</span><span class="token punctuation" style="color:#393A34">}</span><span class="token plain"></span><br></div><div class="token-line" style="color:#393A34"><span class="token plain">packet </span><span class="token operator" style="color:#393A34">=</span><span class="token plain"> build_packet</span><span class="token punctuation" style="color:#393A34">(</span><span class="token plain">cfg</span><span class="token punctuation" style="color:#393A34">,</span><span class="token plain"> repo_root</span><span class="token punctuation" style="color:#393A34">,</span><span class="token plain"> required</span><span class="token operator" style="color:#393A34">=</span><span class="token boolean" style="color:#36acaa">True</span><span class="token punctuation" style="color:#393A34">)</span><span class="token plain"></span><br></div><div class="token-line" style="color:#393A34"><span class="token plain"></span><span class="token comment" style="color:#999988;font-style:italic"># claimed:  raises, agent won't run</span><span class="token plain"></span><br></div><div class="token-line" style="color:#393A34"><span class="token plain"></span><span class="token comment" style="color:#999988;font-style:italic"># actual:   {'topic': ..., 'grounded': [], 'ungrounded': ['ghost-source'], 'chunks': []}</span><br></div></code></pre></div></div>
<p>No exception. The unresolved source doesn't crash the run — it comes back <strong>flagged <code>ungrounded</code></strong>, which is the system working exactly as designed: <em>cite what you can, and be honest about what you couldn't.</em> The hard error only fires when the citation index itself can't be found at all — a different failure, on a different line, that this config never touches.</p>
<p>The finding was still worth acting on — a tutorial shouldn't teach a source ID that isn't real. But the <em>reason</em> was fiction, and the reason changes the fix. "It crashes" means stop-the-world. "It shows up ungrounded" means correct the ID and tighten the copy. If I'd inherited the reviewer's certainty, I'd have written a scarier changelog than the truth, and learned the wrong thing about my own system.</p>
<p><strong>A finding is a claim, not a fact.</strong> The confidence in a review is the author's, not the code's. The five-line repro is cheap; the wrong mental model you adopt by skipping it is not.</p>
<hr>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="the-guard-with-a-hole-in-it">The guard with a hole in it<a href="https://hassant.github.io/certainty-is-the-bug#the-guard-with-a-hole-in-it" class="hash-link" aria-label="Direct link to The guard with a hole in it" title="Direct link to The guard with a hole in it" translate="no">​</a></h2>
<p>Now the part where <em>my</em> certainty was the bug.</p>
<p>There's a redaction path that runs a config-supplied regex over model output. A pathological pattern like <code>`(a+)+$`</code> backtracks for <em>seconds</em> on a short string — a classic denial-of-service footgun. I'd flagged this exact risk in the last post and, at the time, argued the tempting fix (a watchdog thread) was theatre, because the regex engine holds the interpreter lock and a watchdog can never get a turn to fire. So this round I built the honest version: <strong>analyze the pattern before running it</strong> and refuse the dangerous shape.</p>
<p>My first cut was a tidy little regex that recognized "a quantifier wrapped around a quantified atom" — the <code>(a+)+</code> shape. It caught the textbook case. Tests green. I felt done.</p>
<p>Then I threw one more string at it:</p>
<div class="language-python codeBlockContainer_Ckt0 theme-code-block" style="--prism-color:#393A34;--prism-background-color:#f6f8fa"><div class="codeBlockContent_QJqH"><pre tabindex="0" class="prism-code language-python codeBlock_bY9V thin-scrollbar" style="color:#393A34;background-color:#f6f8fa"><code class="codeBlockLines_e6Vv"><div class="token-line" style="color:#393A34"><span class="token plain">is_dangerous</span><span class="token punctuation" style="color:#393A34">(</span><span class="token string" style="color:#e3116c">r"(a+)+$"</span><span class="token punctuation" style="color:#393A34">)</span><span class="token plain">     </span><span class="token comment" style="color:#999988;font-style:italic"># True   — caught</span><span class="token plain"></span><br></div><div class="token-line" style="color:#393A34"><span class="token plain">is_dangerous</span><span class="token punctuation" style="color:#393A34">(</span><span class="token string" style="color:#e3116c">r"((a+))+$"</span><span class="token punctuation" style="color:#393A34">)</span><span class="token plain">   </span><span class="token comment" style="color:#999988;font-style:italic"># False  — sailed straight through</span><br></div></code></pre></div></div>
<p>Same catastrophic pattern, one redundant pair of parentheses, and my detector shrugged. Feed <code>((a+))+$</code> to the real matcher and it grinds for <strong>about ninety seconds</strong> on a thirty-character input. My "guard" waved it through. And it got worse: the check only sat on the redaction path — I'd left the <em>same</em> config regexes unguarded on two other call sites that also run <code>re.search</code> directly. Three doors, one lock.</p>
<!-- -->
<p>A pattern-matcher that recognizes <em>one spelling</em> of a dangerous thing isn't a guard; it's a suggestion. The fix was to stop pattern-matching the surface and start reducing the structure — peel off redundant wrapping, handle the lazy <code>`+?`</code> spelling, and only then ask "is this, underneath, a quantifier over a quantifier?" — and to put that one check on <strong>every</strong> site that runs an untrusted pattern, not just the one that prompted it.</p>
<p><strong>A security check with a trivial bypass is worse than none — it's a bypass with a false sense of safety stapled to it.</strong> If you're writing one, your job isn't to catch the example in the ticket. It's to catch the example <em>plus the same thing with a paren, a lazy quantifier, and a nested group</em> — because whoever trips it next won't use your spelling.</p>
<hr>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="i-wrote-never-in-my-own-docstring">I wrote "never" in my own docstring<a href="https://hassant.github.io/certainty-is-the-bug#i-wrote-never-in-my-own-docstring" class="hash-link" aria-label="Direct link to I wrote &quot;never&quot; in my own docstring" title="Direct link to I wrote &quot;never&quot; in my own docstring" translate="no">​</a></h2>
<p>The last one is the smallest and the most embarrassing, because I typed the overclaim with my own hands.</p>
<p>There's an audit log that hashes a snapshot of each decision. It used to serialize with <code>default=str</code>, which quietly stringifies anything weird — including an object's <strong>memory address</strong>, which changes every run, which silently breaks the tamper-evident chain across processes. Real bug. I replaced it with a deterministic encoder and, pleased with myself, wrote the docstring:</p>
<blockquote>
<p><em>…is deterministic and <strong>never raises</strong>, so a malformed snapshot can't crash a governed turn.</em></p>
</blockquote>
<p>Deterministic: true, and a good fix. <strong>Never raises:</strong> a wish wearing a fact's clothing. Because the serializer only gets consulted for values the JSON encoder already accepted — and the encoder throws <em>before</em> it ever calls me:</p>
<div class="language-python codeBlockContainer_Ckt0 theme-code-block" style="--prism-color:#393A34;--prism-background-color:#f6f8fa"><div class="codeBlockContent_QJqH"><pre tabindex="0" class="prism-code language-python codeBlock_bY9V thin-scrollbar" style="color:#393A34;background-color:#f6f8fa"><code class="codeBlockLines_e6Vv"><div class="token-line" style="color:#393A34"><span class="token plain">cyclic </span><span class="token operator" style="color:#393A34">=</span><span class="token plain"> </span><span class="token punctuation" style="color:#393A34">[</span><span class="token punctuation" style="color:#393A34">]</span><span class="token punctuation" style="color:#393A34">;</span><span class="token plain"> cyclic</span><span class="token punctuation" style="color:#393A34">.</span><span class="token plain">append</span><span class="token punctuation" style="color:#393A34">(</span><span class="token plain">cyclic</span><span class="token punctuation" style="color:#393A34">)</span><span class="token plain"></span><br></div><div class="token-line" style="color:#393A34"><span class="token plain">audit</span><span class="token punctuation" style="color:#393A34">.</span><span class="token plain">record</span><span class="token punctuation" style="color:#393A34">(</span><span class="token string" style="color:#e3116c">"input"</span><span class="token punctuation" style="color:#393A34">,</span><span class="token plain"> </span><span class="token punctuation" style="color:#393A34">{</span><span class="token string" style="color:#e3116c">"x"</span><span class="token punctuation" style="color:#393A34">:</span><span class="token plain"> cyclic</span><span class="token punctuation" style="color:#393A34">}</span><span class="token punctuation" style="color:#393A34">,</span><span class="token plain"> verdict</span><span class="token punctuation" style="color:#393A34">)</span><span class="token plain">   </span><span class="token comment" style="color:#999988;font-style:italic"># ValueError: Circular reference detected</span><span class="token plain"></span><br></div><div class="token-line" style="color:#393A34"><span class="token plain">audit</span><span class="token punctuation" style="color:#393A34">.</span><span class="token plain">record</span><span class="token punctuation" style="color:#393A34">(</span><span class="token string" style="color:#e3116c">"input"</span><span class="token punctuation" style="color:#393A34">,</span><span class="token plain"> </span><span class="token punctuation" style="color:#393A34">{</span><span class="token punctuation" style="color:#393A34">(</span><span class="token string" style="color:#e3116c">"a"</span><span class="token punctuation" style="color:#393A34">,</span><span class="token string" style="color:#e3116c">"b"</span><span class="token punctuation" style="color:#393A34">)</span><span class="token punctuation" style="color:#393A34">:</span><span class="token plain"> </span><span class="token number" style="color:#36acaa">1</span><span class="token punctuation" style="color:#393A34">}</span><span class="token punctuation" style="color:#393A34">,</span><span class="token plain"> verdict</span><span class="token punctuation" style="color:#393A34">)</span><span class="token plain">  </span><span class="token comment" style="color:#999988;font-style:italic"># TypeError: keys must be str...</span><br></div></code></pre></div></div>
<p>Both blow up. And the call site that records the audit entry has no <code>try</code> around it — so my "can't crash a governed turn" crashes the governed turn, on exactly the malformed input the fix was supposed to make safe. I'd fixed the determinism and left the crash, then <em>documented</em> the crash out of existence.</p>
<p>The repair was to make the sentence true: wrap the serialization so any encoder failure falls back to a stable, deterministic sentinel instead of propagating — genuinely fail-safe — and narrow the claim to what the code actually guarantees. Same amount of confidence in the docstring. Backed, this time, by a line that catches the exception.</p>
<p><strong>"Never" is a proof obligation, not an adjective.</strong> If a hot path calls you and can't handle your exception, either you actually never raise — because you wrapped the thing that does — or you don't write "never." The docstring is a claim like any other, and the reader will trust it more than they trust the code.</p>
<hr>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="the-field-map">The field map<a href="https://hassant.github.io/certainty-is-the-bug#the-field-map" class="hash-link" aria-label="Direct link to The field map" title="Direct link to The field map" translate="no">​</a></h2>
<ul>
<li class=""><strong>A finding is a claim.</strong> A review's confidence belongs to its author, not to your code. Reproduce the specific claim before you adopt its mental model — the wrong model is more expensive than the bug.</li>
<li class=""><strong>A guard that catches one spelling is a suggestion.</strong> For anything security-shaped, test the bypass on purpose: the extra paren, the lazy quantifier, the nested group. Then put the check on <em>every</em> door, not the one that filed the ticket.</li>
<li class=""><strong>"Never" is a proof, not an adjective.</strong> If an unguarded caller depends on your promise, make it true or stop making it. An overclaim in a docstring outlives the code that broke it.</li>
<li class=""><strong>Determinism and safety are different fixes.</strong> I fixed one and wrote the docstring as if I'd fixed both. Check that the sentence and the code agree.</li>
<li class=""><strong>The most dangerous line is the one you're sure about</strong> — the review you trusted, the guard you tested once, the docstring you typed on a roll. Certainty is where verification stops, which is exactly where the bug moves in.</li>
</ul>
<p>The review was right that these lines needed changing. It was wrong about why one of them did, and I was wrong twice about how. Green didn't lie last time; the review didn't lie this time. They were both just <em>confident</em> — and confidence is not the same as having run it.</p>
<p>Run it. Then believe it. 🔎</p>]]></content>
        <author>
            <name>Hassan Tariq</name>
            <uri>https://github.com/hassant</uri>
        </author>
        <category label="ai-agents" term="ai-agents"/>
        <category label="agentic" term="agentic"/>
        <category label="code-review" term="code-review"/>
        <category label="security" term="security"/>
        <category label="devtools" term="devtools"/>
    </entry>
    <entry>
        <title type="html"><![CDATA[You Don't Pick the Autonomy Level. Your Evidence Does.]]></title>
        <id>https://hassant.github.io/evidence-picks-the-level</id>
        <link href="https://hassant.github.io/evidence-picks-the-level"/>
        <updated>2026-07-05T00:00:00.000Z</updated>
        <summary type="html"><![CDATA[The interesting question in agentic engineering stopped being what do I prompt months ago. These days it's a quieter, more nervous one: how much rope do I hand this thing, and what do I need back before I can walk away from it?]]></summary>
        <content type="html"><![CDATA[<p>The interesting question in agentic engineering stopped being <em>what do I prompt</em> months ago. These days it's a quieter, more nervous one: <strong>how much rope do I hand this thing, and what do I need back before I can walk away from it?</strong></p>
<p>I've been keeping a rough ladder in my head for that — low end, it just suggests and I decide; high end, it runs on its own and only surfaces the decisions that actually need me. It's a useful ladder. It's also been quietly lying to me, and it took a bad week to see how.</p>
<!-- -->
<p>Here's the lie, stated plainly: I kept treating the rung I <em>reached</em> as the rung I'd <em>earned</em>. They are not the same thing, and the gap between them is where every agentic disaster I've had actually lived. I've written before that <a class="" href="https://hassant.github.io/green-doesnt-mean-done">green doesn't mean done</a> — a passing build tells you no check failed, not that the work is right. This is the same disease one level up. <strong>The autonomy you can defend is exactly the autonomy your evidence can carry. Everything above that line is a story you're telling yourself with more agents running.</strong></p>
<p>You don't pick the level. Your proof does. "Add more verification" is the kind of advice that's true but useless until it has edges — so let me show you mine.</p>
<hr>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="two-skills-i-kept-scoring-as-one">Two skills I kept scoring as one<a href="https://hassant.github.io/evidence-picks-the-level#two-skills-i-kept-scoring-as-one" class="hash-link" aria-label="Direct link to Two skills I kept scoring as one" title="Direct link to Two skills I kept scoring as one" translate="no">​</a></h2>
<p>For months I bragged about the same two things as if they were one thing. "I let this agent run for twenty minutes untouched." "I had eight agents going at once." Both feel like <em>look how hands-off I've become</em> — and they are completely different skills that break in opposite directions.</p>
<p>One is about <strong>reach</strong>: how far from my own judgment I let a single agent wander before it has to come back and show me something. The other is about <strong>coordination</strong>: whether I can carve the work into slices that don't collide, so I only step in when one of them face-plants.</p>
<p>You can be strong at one and a menace at the other. The trap I kept walking into was running "several agents at once" while personally holding every dependency in my head — deciding what each one touched, merging their output by hand, arbitrating conflicts on the fly. That's not orchestration. That's me being the orchestrator, manually, at higher speed. It <em>looked</em> like leverage. It was just faster context-switching wearing leverage's jacket. <strong>Running many agents is not the same as coordinating them, and running one agent far is not the same as running it safely.</strong> Score them separately or you'll congratulate yourself for the wrong one.</p>
<hr>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="the-contract-is-the-receipt">The contract is the receipt<a href="https://hassant.github.io/evidence-picks-the-level#the-contract-is-the-receipt" class="hash-link" aria-label="Direct link to The contract is the receipt" title="Direct link to The contract is the receipt" translate="no">​</a></h2>
<p>Here's the thing that actually moved my autonomy up a rung, and it's boring: I stopped letting an agent start until I'd written down what it was for.</p>
<p>Not a prompt. A contract. On agentry — the governed-agent tooling I've been building — that's a single declarative file: the goal, the boundaries, the identity it runs as, what it's allowed to touch, how it's grounded, who has to approve what, how it's checked, when it stops, and what it costs. The first few times it felt like paperwork written to feel responsible. It isn't. <strong>Every field on that thing is a specific failure I'm pre-paying:</strong></p>
<ul>
<li class="">A <strong>stopping condition</strong> is the entire difference between "chase this goal" and an agent that hits your time-to-interactive target by deleting the page. "Do whatever it takes" is only safe when <em>whatever it takes</em> has a floor it isn't allowed to cross — and the floor has to be measurable, or it's decoration.</li>
<li class=""><strong>Evidence</strong> is the field that decides whether I'm delegating or just hoping. Not "it says the tests pass" — the diff, the failing-then-passing run, the repro, the logs. Something a <em>different</em> process can check without asking the agent that produced it.</li>
<li class=""><strong>Escalation and budget</strong> are the two I used to skip and the two that turn a runaway into a bounded loss. An agent that can retry forever on my tokens isn't autonomous. It's a leak with initiative.</li>
</ul>
<p>The contract isn't ceremony. It's the receipt I'll want when a run goes sideways and someone — usually future-me at an ugly hour — asks what this thing was even permitted to do. No contract, no receipt. No receipt, and "high autonomy" just means I was lucky, and luck has never once shipped a stopping condition.</p>
<hr>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="three-things-i-build-backwards-from">Three things I build backwards from<a href="https://hassant.github.io/evidence-picks-the-level#three-things-i-build-backwards-from" class="hash-link" aria-label="Direct link to Three things I build backwards from" title="Direct link to Three things I build backwards from" translate="no">​</a></h2>
<p>Once I started treating autonomy as something you <em>earn with evidence</em>, three pieces of machinery stopped being optional — the things I now build <em>before</em> I let the rope out, each one answering a question I'll be asking at the worst possible moment:</p>
<p><strong>A tripwire, not a transcript</strong> — for <em>how fast will I know it went wrong?</em> Every governed decision writes a tamper-evident trail, so catching a bad turn is a query, not an archaeology dig. If finding out I was wrong means re-reading a chat log, I don't have autonomy, I have a liability with good manners.</p>
<p><strong>A cheap undo</strong> — for <em>how cleanly can I roll it back?</em> Fail-closed defaults and a real revert path aren't paranoia, they're the toll for the next rung. Reversibility is what lets you run hotter: the refactor behind strong tests and a clean revert can take far more rope than a task with no source of truth to revert <em>to</em>.</p>
<p><strong>Evidence that travels</strong> — for <em>what would prove I was right?</em> The proof rides <em>with</em> the work — a diff, the tests, the logs, the findings — instead of collapsing into a confident paragraph at the end.</p>
<p>If my honest answers are <em>slowly, painfully, and "I trust the summary,"</em> I'm not operating at high autonomy. I'm exposed, and calling the exposure a feature.</p>
<p>Which is exactly the wall I hit last week.</p>
<hr>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="the-week-i-almost-trusted-the-summary">The week I almost trusted the summary<a href="https://hassant.github.io/evidence-picks-the-level#the-week-i-almost-trusted-the-summary" class="hash-link" aria-label="Direct link to The week I almost trusted the summary" title="Direct link to The week I almost trusted the summary" translate="no">​</a></h2>
<p>The most dangerous version of this failure is the tidiest one: you let the agent's own write-up stand in for the review. The summary is clean, the tone is confident, and reading it <em>feels</em> like doing the work. It isn't. It's surrender with good formatting.</p>
<p>I ran straight at it a few days ago. I had a manager agent taking a list of findings, handing slices out to a small fleet of reviewers, collecting what they found, and returning only the decisions that needed me. A small version of the shape I actually trust — one coordinator, separate checkers, me on the decisions. The efficient move was to skim the manager's summary and ship.</p>
<p>I didn't — and the only reason it held is that I made the reviewers argue with the <em>code</em>, not with each other's summaries. They earned it twice: they caught claims in the original findings that turned out to be confidently wrong, and then caught blind spots in <em>my own fixes</em> that my green tests had waved straight through. I wrote up the wreckage in <a class="" href="https://hassant.github.io/certainty-is-the-bug">certainty is the bug</a>, but the autonomy lesson is sharper than the debugging one: <strong>a fleet without independent verification isn't leverage. It's faster wrongness, delivered with more confidence and a nicer diff.</strong></p>
<p>The fix was never "trust it harder." It was the unglamorous stuff — separate the thing that does the work from the thing that checks it, make the evidence independent of whatever generated it, keep the scope small enough that a rollback is cheap. The summary is not the review. The number of agents is not the leverage. The evidence is the only part of any of this that actually scales.</p>
<hr>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="the-field-map">The field map<a href="https://hassant.github.io/evidence-picks-the-level#the-field-map" class="hash-link" aria-label="Direct link to The field map" title="Direct link to The field map" translate="no">​</a></h2>
<ul>
<li class=""><strong>Reach and coordination are different skills.</strong> "How far one agent runs" and "how many I run at once" fail in opposite ways. Score them apart, or you'll mistake fast context-switching for orchestration.</li>
<li class=""><strong>The level follows the verification, not the task's vibe.</strong> A payment refactor behind strong tests and a clean revert can run hotter than a docs task with no canonical truth. Risk and reversibility set the ceiling.</li>
<li class=""><strong>Write the contract before the agent starts.</strong> Goal, boundaries, stopping condition, evidence, escalation, budget. Each field is a failure you're pre-paying; skipping it doesn't remove the failure, it just makes it a surprise.</li>
<li class=""><strong>Build backwards from three questions.</strong> Know-it's-wrong-fast, undo-cleanly, prove-it's-right aren't a post-hoc quiz — they're why you build the audit trail, the fail-closed default, and the evidence that travels with the work.</li>
<li class=""><strong>Independent verification is the only thing that scales.</strong> A summary is not a review; a fleet is not leverage. Separate who builds from who checks, or high autonomy is just faster wrongness.</li>
</ul>
<p>Verification was always going to be the ceiling. The mature posture isn't the highest rung you can climb to — it's the highest one you can currently <em>defend</em>, and the honest ones move up slowly, one axis at a time, only as the evidence for the next rung actually accumulates.</p>
<p>So no, you don't get to pick the autonomy level. You write the contract, you build the proof, and the proof tells you which rung you've earned. Everything above it is theatre.</p>
<p>Pick the level your evidence can defend. Then, and only then, climb. 🪜</p>]]></content>
        <author>
            <name>Hassan Tariq</name>
            <uri>https://github.com/hassant</uri>
        </author>
        <category label="ai-agents" term="ai-agents"/>
        <category label="agentic" term="agentic"/>
        <category label="autonomy" term="autonomy"/>
        <category label="governance" term="governance"/>
        <category label="devtools" term="devtools"/>
    </entry>
    <entry>
        <title type="html"><![CDATA[Green, Reviewed, and Still Wrong: The Bugs Behind Your Test Doubles]]></title>
        <id>https://hassant.github.io/green-reviewed-still-wrong</id>
        <link href="https://hassant.github.io/green-reviewed-still-wrong"/>
        <updated>2026-06-26T00:00:00.000Z</updated>
        <summary type="html"><![CDATA[I shipped a security-hardening pass on a side project talk to a real pipe.]]></summary>
        <content type="html"><![CDATA[<p>I shipped a security-hardening pass on a side project: the suite green, the validator clean, the release tagged. A day later I found a bug in it that would have broken <strong>every</strong> run — not by re-running the tests, but by writing twenty lines that did one thing the whole suite never did: talk to a real pipe.</p>
<p>Green wasn't lying. It just wasn't looking where the bug was.</p>
<!-- -->
<p>I've argued before that <a class="" href="https://hassant.github.io/green-doesnt-mean-done">green doesn't mean done</a> — a passing build says <em>no check failed</em>, not <em>the code is correct</em>. This is the follow-up with receipts: three bugs that sailed through a green suite and a careful read, because each one hid in a place verification almost never looks — <strong>behind a test double, inside the fix itself, and in the direction a fallback falls.</strong> None of them are exotic. That's the point. They're the kind you ship on a good day.</p>
<hr>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="the-mock-that-hid-the-hang">The mock that hid the hang<a href="https://hassant.github.io/green-reviewed-still-wrong#the-mock-that-hid-the-hang" class="hash-link" aria-label="Direct link to The mock that hid the hang" title="Direct link to The mock that hid the hang" translate="no">​</a></h2>
<p>I'd just rewritten the part of a harness that reads JSON-RPC from a subprocess over a pipe. The old version could hang if the other end went quiet; the new one added a deadline. <code>select()</code> for readability, then read a line — textbook. Tests passed.</p>
<p>Here's the repro I wrote anyway, because "added a deadline" is a claim, not a fact:</p>
<div class="language-python codeBlockContainer_Ckt0 theme-code-block" style="--prism-color:#393A34;--prism-background-color:#f6f8fa"><div class="codeBlockContent_QJqH"><pre tabindex="0" class="prism-code language-python codeBlock_bY9V thin-scrollbar" style="color:#393A34;background-color:#f6f8fa"><code class="codeBlockLines_e6Vv"><div class="token-line" style="color:#393A34"><span class="token plain">read_fd</span><span class="token punctuation" style="color:#393A34">,</span><span class="token plain"> write_fd </span><span class="token operator" style="color:#393A34">=</span><span class="token plain"> os</span><span class="token punctuation" style="color:#393A34">.</span><span class="token plain">pipe</span><span class="token punctuation" style="color:#393A34">(</span><span class="token punctuation" style="color:#393A34">)</span><span class="token plain"></span><br></div><div class="token-line" style="color:#393A34"><span class="token plain">os</span><span class="token punctuation" style="color:#393A34">.</span><span class="token plain">write</span><span class="token punctuation" style="color:#393A34">(</span><span class="token plain">write_fd</span><span class="token punctuation" style="color:#393A34">,</span><span class="token plain"> </span><span class="token string" style="color:#e3116c">b'{"jsonrpc":"2.0","id":1,"result":{"ok":true}}\n'</span><span class="token punctuation" style="color:#393A34">)</span><span class="token plain">  </span><span class="token comment" style="color:#999988;font-style:italic"># one whole line, atomically</span><span class="token plain"></span><br></div><div class="token-line" style="color:#393A34"><span class="token plain"></span><span class="token comment" style="color:#999988;font-style:italic"># read through os.fdopen(read_fd, "r") — exactly what a real subprocess hands you</span><span class="token plain"></span><br></div><div class="token-line" style="color:#393A34"><span class="token plain">resp </span><span class="token operator" style="color:#393A34">=</span><span class="token plain"> read_response</span><span class="token punctuation" style="color:#393A34">(</span><span class="token plain">read_fd</span><span class="token punctuation" style="color:#393A34">,</span><span class="token plain"> timeout</span><span class="token operator" style="color:#393A34">=</span><span class="token number" style="color:#36acaa">2.0</span><span class="token punctuation" style="color:#393A34">)</span><span class="token plain"></span><br></div><div class="token-line" style="color:#393A34"><span class="token plain"></span><span class="token comment" style="color:#999988;font-style:italic"># expected: the response.   actual: TimeoutError after 2.0s.</span><br></div></code></pre></div></div>
<p>A complete, valid response — and it <strong>timed out</strong>. Every real call would have stalled for the full timeout and then died.</p>
<p>The cause is a lovely little trap. <code>select()</code> only sees the operating-system file descriptor. It knows nothing about Python's <em>own</em> buffer. My first <code>read(1)</code> asked a buffered text stream for one character; the stream went to the kernel, slurped the <strong>whole</strong> line into its internal buffer, and handed me back one char. Next loop iteration, <code>select()</code> checked the fd again — now empty, because the data was sitting in Python's buffer the entire time. So <code>select()</code> reported "nothing here," and I timed out on a response I'd already received.</p>
<!-- -->
<p>Now look at <em>why</em> the suite was green. One test fed the reader a <code>StringIO</code> — which has no file descriptor, so it quietly skipped the entire <code>select()</code> codepath where the bug lives. Another used a real pipe that <em>never sent anything</em> — perfect for proving the timeout fires, useless for proving a normal read works. Between "no real fd" and "no data," not one test ever put a real, complete line into a real pipe. The mocks weren't wrong; their <strong>shape</strong> was wrong, and the bug lived exactly in the shape they didn't have.</p>
<p>The fix was to stop reading through the buffered wrapper and pull raw bytes off the descriptor with <code>os.read</code>, so the thing <code>select()</code> watches and the thing I read from are the same object. But the fix isn't the lesson. The lesson is that I only knew to write it because I'd reproduced the failure against the <strong>real primitive</strong> — a real fd, a real pipe — instead of the convenient stand-in. Test doubles are how you go fast. They are also where integration bugs go to hide.</p>
<hr>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="a-fix-worse-than-the-bug">A fix worse than the bug<a href="https://hassant.github.io/green-reviewed-still-wrong#a-fix-worse-than-the-bug" class="hash-link" aria-label="Direct link to A fix worse than the bug" title="Direct link to A fix worse than the bug" translate="no">​</a></h2>
<p>The next one I'd written myself, recently, with tests, feeling good about it. A careful read of the diff is what turned it up.</p>
<p>The job was redaction: take a pattern and a replacement from config and scrub matches out of some text. Mine did the obvious thing:</p>
<div class="language-python codeBlockContainer_Ckt0 theme-code-block" style="--prism-color:#393A34;--prism-background-color:#f6f8fa"><div class="codeBlockContent_QJqH"><pre tabindex="0" class="prism-code language-python codeBlock_bY9V thin-scrollbar" style="color:#393A34;background-color:#f6f8fa"><code class="codeBlockLines_e6Vv"><div class="token-line" style="color:#393A34"><span class="token plain">re</span><span class="token punctuation" style="color:#393A34">.</span><span class="token plain">subn</span><span class="token punctuation" style="color:#393A34">(</span><span class="token plain">pattern</span><span class="token punctuation" style="color:#393A34">,</span><span class="token plain"> replacement</span><span class="token punctuation" style="color:#393A34">,</span><span class="token plain"> text</span><span class="token punctuation" style="color:#393A34">)</span><span class="token plain">   </span><span class="token comment" style="color:#999988;font-style:italic"># replacement is... a regex template, not literal text</span><br></div></code></pre></div></div>
<p><code>re.subn</code> treats the replacement as a <strong>template</strong>, not literal text — which means it can contain a backreference. Watch what a replacement of <code>`\g&lt;0&gt;`</code> does to a secret:</p>
<div class="language-python codeBlockContainer_Ckt0 theme-code-block" style="--prism-color:#393A34;--prism-background-color:#f6f8fa"><div class="codeBlockContent_QJqH"><pre tabindex="0" class="prism-code language-python codeBlock_bY9V thin-scrollbar" style="color:#393A34;background-color:#f6f8fa"><code class="codeBlockLines_e6Vv"><div class="token-line" style="color:#393A34"><span class="token plain">re</span><span class="token punctuation" style="color:#393A34">.</span><span class="token plain">subn</span><span class="token punctuation" style="color:#393A34">(</span><span class="token string" style="color:#e3116c">r"\d{3}-\d{2}-\d{4}"</span><span class="token punctuation" style="color:#393A34">,</span><span class="token plain"> </span><span class="token string" style="color:#e3116c">r"\g&lt;0&gt;_SEEN"</span><span class="token punctuation" style="color:#393A34">,</span><span class="token plain"> </span><span class="token string" style="color:#e3116c">"ssn 123-45-6789"</span><span class="token punctuation" style="color:#393A34">)</span><span class="token plain"></span><br></div><div class="token-line" style="color:#393A34"><span class="token plain"></span><span class="token comment" style="color:#999988;font-style:italic"># -&gt; ('ssn 123-45-6789_SEEN', 1)     the secret is right there. redaction re-emitted it.</span><br></div></code></pre></div></div>
<p>The feature whose entire job is to <em>remove</em> the secret can be made to <strong>re-print it</strong> — and a replacement ending in a stray backslash throws an uncaught error that kills the turn instead. The repair is one line: substitute with a function so the replacement is treated literally, and fail closed if the pattern is malformed.</p>
<p>But the one line isn't the lesson either. The lesson is <em>where</em> this lived: in the security-critical path, in code I was confident about, behind tests that all used boring, well-behaved replacement strings. No green test ever tried a hostile replacement, because the person who wrote the redaction also wrote the tests, and we shared the assumption that the replacement was friendly. <strong>A fix can quietly invert the very guarantee it's supposed to provide — and the more sure you are of a piece of code, the less precisely your own tests will probe it.</strong></p>
<hr>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="fail-closed-in-the-wrong-direction">Fail-closed, in the wrong direction<a href="https://hassant.github.io/green-reviewed-still-wrong#fail-closed-in-the-wrong-direction" class="hash-link" aria-label="Direct link to Fail-closed, in the wrong direction" title="Direct link to Fail-closed, in the wrong direction" translate="no">​</a></h2>
<p>The smallest bug taught the sharpest lesson, and I caught it the lowest-tech way there is: reading my own diff out loud.</p>
<p>A rule can say <em>deny when <code>amount &gt;= 1000</code></em>. Earlier in the same pass I'd hardened the number parser to reject nonsense like <code>inf</code> and <code>nan</code>. Obviously a safety improvement.</p>
<p>My first version returned "not a number" for those. <em>Feels</em> fail-closed. It's the exact opposite. The rule only <em>fires</em> when the comparison is true; <code>float("nan") &gt;= 1000</code> is <code>False</code>, and "not a number" also makes the rule not fire — so a <code>nan</code> amount strolls straight past a deny-the-big-ones rule. My "safety" fix <strong>failed open</strong> on the one path that mattered.</p>
<p>The correct version <em>raises</em>, so the evaluator's catch-all turns it into a hard deny. Same intent, opposite outcome, and the entire difference is which way the fallback falls. "Reject the bad input" is not a direction. <strong>In a guard, every fallback has a polarity — and you have to trace, for the specific rule, whether your safe-looking default lands on allow or deny.</strong> Tests wouldn't have caught this; the rule still "worked" on every finite number anyone thought to type.</p>
<hr>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="the-hardening-that-would-have-been-theatre">The hardening that would have been theatre<a href="https://hassant.github.io/green-reviewed-still-wrong#the-hardening-that-would-have-been-theatre" class="hash-link" aria-label="Direct link to The hardening that would have been theatre" title="Direct link to The hardening that would have been theatre" translate="no">​</a></h2>
<p>One more, because it's the inverse mistake — not a bug I almost shipped, but a "fix" I almost built.</p>
<p>That same redaction path runs a config-supplied regex on the hot path, and a pathological pattern like <code>`(a+)+$`</code> can backtrack for <em>seconds</em> on a short string. The tempting hardening is to run the match inside a timeout — spawn a thread, kill it if it overruns. I started reaching for it.</p>
<p>Then I remembered how the runtime actually works: CPython's regex engine holds the interpreter lock while it matches. A watchdog thread can't preempt it, because the runaway match never lets go of the lock for the watchdog to <em>run</em>. The timeout would be pure ceremony — a guard that can't fire. The honest fix is smaller and less satisfying: treat config patterns as trusted code, document the footgun, and fail safe on a malformed one. <strong>Know your runtime before you "harden" against it — half of hardening is not building the protection that can't work.</strong></p>
<hr>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="the-field-map">The field map<a href="https://hassant.github.io/green-reviewed-still-wrong#the-field-map" class="hash-link" aria-label="Direct link to The field map" title="Direct link to The field map" translate="no">​</a></h2>
<ul>
<li class=""><strong>Mock the primitive, mock the bug away.</strong> A double's <em>shape</em> decides which codepaths it touches. Reproduce the failure against the real fd, socket, or clock at least once, or you've only tested the stand-in.</li>
<li class=""><strong>A fix can invert the guarantee.</strong> Especially in a security path, especially in code you're sure about — which is exactly the code your own tests probe least precisely.</li>
<li class=""><strong>Every fallback has a polarity.</strong> "Reject the bad input" lands on <em>allow</em> or <em>deny</em> depending on the rule. Trace it for the specific case; "safe-looking" is not the same as fail-closed.</li>
<li class=""><strong>Know the runtime before you harden.</strong> The protection that can't fire is worse than none — it's a false sense of one. Half of good hardening is deleting the fix that doesn't work.</li>
<li class=""><strong>Reproduce, don't re-run.</strong> Re-running your own green suite just re-confirms your blind spots. The thing that moves is a fresh script pointed at reality.</li>
<li class=""><strong>You are the last reviewer.</strong> Tools and second reads are worth it — but the bug that would have broken every run was caught by twenty lines I wrote against a real pipe, after the suite went green.</li>
</ul>
<p>Green told me no check failed, and that was true. The bug that mattered most was hiding behind a <code>StringIO</code> the whole time, in the one shape my tests never took.</p>
<p>Don't ship the green. Ship the thing you reproduced against the real world. 🔌</p>]]></content>
        <author>
            <name>Hassan Tariq</name>
            <uri>https://github.com/hassant</uri>
        </author>
        <category label="ai-agents" term="ai-agents"/>
        <category label="agentic" term="agentic"/>
        <category label="code-review" term="code-review"/>
        <category label="testing" term="testing"/>
        <category label="devtools" term="devtools"/>
    </entry>
    <entry>
        <title type="html"><![CDATA[I Read 31 Agent Loops — Most Stop When They're Tired, Not When They're Right]]></title>
        <id>https://hassant.github.io/stop-when-right-not-tired</id>
        <link href="https://hassant.github.io/stop-when-right-not-tired"/>
        <updated>2026-06-21T00:00:00.000Z</updated>
        <summary type="html"><![CDATA[The problem, stated plainly: most agent loops don't know when to stop. They stop when they run out of turns, when they hit a wall they can't get past, or when a human finally looks — and then they call whatever state they're in "done." Repeating is the easy part. Knowing the latest pass is actually right — and not just the pass you happened to be on when the budget ran out — is the whole game.]]></summary>
        <content type="html"><![CDATA[<p><strong>The problem, stated plainly: most agent loops don't know when to stop. They stop when they run out of turns, when they hit a wall they can't get past, or when a human finally looks — and then they call whatever state they're in "done."</strong> Repeating is the easy part. Knowing the <em>latest</em> pass is actually right — and not just the pass you happened to be on when the budget ran out — is the whole game.</p>
<p>This week the loop became a <strong>named, shareable artifact</strong> — a library you can browse, copy, and submit to. But the names aren't the interesting part. The exits are. Here's what changed my mind, the four ways a loop can stop ranked by how much trust each one earns, and why — now that generation is nearly free — the exit is the only part left worth engineering.</p>
<!-- -->
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="i-read-31-loops-to-feel-smug-it-backfired">I read 31 loops to feel smug. It backfired.<a href="https://hassant.github.io/stop-when-right-not-tired#i-read-31-loops-to-feel-smug-it-backfired" class="hash-link" aria-label="Direct link to I read 31 loops to feel smug. It backfired." title="Direct link to I read 31 loops to feel smug. It backfired." translate="no">​</a></h2>
<p>Forward Future just published the <a href="https://signals.forwardfuture.ai/loop-library/" target="_blank" rel="noopener noreferrer" class="">Loop Library</a> — 31 named agent loops (a docs sweep, a production-error sweep, a builder–reviewer "autonomy" loop, a versioned-experiment loop) with a clean thesis: a good loop answers four questions — what's the goal, how does it check the last attempt, what does it do with what it learned, and <em>when does it stop</em>.</p>
<p>I started this post as a tour of all 31. I deleted that draft — a catalog isn't an argument. I went back in looking for "vibes with a <code>while</code> loop" so I could feel clever about my own setup. Mostly I found the opposite, and it indicted <em>my</em> loops, not theirs.</p>
<p>Because my throwaway loops look like this: <em>"keep fixing the flaky tests until they pass."</em> It runs, something goes green on pass nine, <code>max_turns</code> is close, it stops, and I get a cheerful <strong>✅ all tests passing</strong>. The suite is still flaky. The loop didn't decide it was right. It decided it was tired. Call it <strong>stop-on-exhaustion</strong> — and once you have the phrase, you start seeing it everywhere.</p>
<p>The good library loops don't do that. They decide what would make them stop <em>before</em> the first pass. That's the tell of a real one.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="why-this-matters-now-generation-got-cheap">Why this matters now: generation got cheap<a href="https://hassant.github.io/stop-when-right-not-tired#why-this-matters-now-generation-got-cheap" class="hash-link" aria-label="Direct link to Why this matters now: generation got cheap" title="Direct link to Why this matters now: generation got cheap" translate="no">​</a></h2>
<p>A year ago, an iteration was expensive enough that you babied each one. That's over. Charity Majors put it bluntly this month — <a href="https://charitydotwtf.substack.com/p/ai-demands-more-engineering-discipline" target="_blank" rel="noopener noreferrer" class="">"AI demands more engineering discipline. Not less."</a> — because the economics of code production flipped: lines went from treasured and carefully curated to <strong>disposable and regenerable</strong>, practically overnight. The same week, <a href="https://simonwillison.net/2026/Jun/17/glm-52/" target="_blank" rel="noopener noreferrer" class="">GLM-5.2 shipped as open weights</a> with a million-token context. On models like that, a short coding pass is now near-zero marginal cost.</p>
<p>Here's the consequence nobody likes saying out loud: <strong>when generation is a commodity, the code the loop emits is the cheap part.</strong> The scarce resource is the judgment about which pass to <em>keep</em>. The asset isn't the output — it's the evaluator that judges it. A frontier model wired to a "did it crash?" check is a Ferrari behind a turnstile. (I called that <em>rigor cosplay</em> in <a class="" href="https://hassant.github.io/agent-engineering-stack">the stack post</a>; cheap iteration just made it more expensive to fake.)</p>
<p>So the question stops being <em>"can the agent generate a fix?"</em> — of course it can, ten times before lunch — and becomes <em>"can the loop tell a good pass from a lucky one, and stop on the right one?"</em></p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="hope-loops-vs-prune-loops">Hope loops vs prune loops<a href="https://hassant.github.io/stop-when-right-not-tired#hope-loops-vs-prune-loops" class="hash-link" aria-label="Direct link to Hope loops vs prune loops" title="Direct link to Hope loops vs prune loops" translate="no">​</a></h2>
<p>Watch two loops chase the same goal and the difference jumps out.</p>
<!-- -->
<p>A <strong>hope loop</strong> regenerates the same doomed approach until the budget dies, then ships whatever was on screen. A <strong>prune loop</strong> spends the now-cheap generations on <em>variety</em>, kills the regressions fast, and iterates only the survivor.</p>
<p>Which means the metric I actually care about isn't a loop's success rate — it's its <strong>discard ratio</strong>: how quickly it throws away a path that isn't working. A loop that takes ten passes to abandon a dead end is worse than one that abandons it in two, even if both eventually pass. Cheap generation didn't make iteration the skill. It made <em>discarding</em> the skill.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="the-four-ways-a-loop-can-stop">The four ways a loop can stop<a href="https://hassant.github.io/stop-when-right-not-tired#the-four-ways-a-loop-can-stop" class="hash-link" aria-label="Direct link to The four ways a loop can stop" title="Direct link to The four ways a loop can stop" translate="no">​</a></h2>
<p>Strip any loop down and its exit falls into one of four buckets. They are not equal.</p>
<!-- -->
<p>Trust climbs as you go down the list. <strong>Exhaustion</strong> is the bottom — it's not a stop condition at all, it's a loop fainting and you reading meaning into where it fell. <strong>Surrender</strong> is honest and often correct: the loop did its mechanical work and routes the judgment to a person. <strong>Consensus</strong> (a second model signs off) is stronger, but agreement and correctness are different things — two models can be confidently wrong together. <strong>Provable Truth</strong> — a test, an audit, a compile, a check a reviewer can re-run and watch turn red — is the only exit that proves the property you were actually after.</p>
<p>The honest finding from the ones I read closely: the loops worth naming anchor on Provable Truth or a deliberate human gate. None of <em>those</em> stop because they got tired.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="the-teardown">The teardown<a href="https://hassant.github.io/stop-when-right-not-tired#the-teardown" class="hash-link" aria-label="Direct link to The teardown" title="Direct link to The teardown" translate="no">​</a></h2>
<p>Three specimens, read for their exits rather than their actions:</p>
<table><thead><tr><th>Loop (from the library)</th><th>The check it runs</th><th>How it stops</th><th>What that earns</th></tr></thead><tbody><tr><td><strong>Revolve versioned-experiment</strong></td><td>frozen tests + scoring against a recorded baseline; keep only a regression-free win</td><td><strong>Provable Truth</strong> — plus no-progress / budget as backstops</td><td>a kept change is <em>provably</em> better than the last checkpoint, with a rollback</td></tr><tr><td><strong>The docs sweep</strong></td><td>verifies commands, links, and examples against the repo, then opens a PR</td><td>mechanical proof <strong>+ Surrender</strong> to human review</td><td>docs match the code — <em>if</em> someone actually reviews the PR</td></tr><tr><td><strong>The production data cleanup</strong></td><td>audits records against explicit rules — sampled and queried — with reversible ops; holds uncertain rows back</td><td><strong>Provable Truth + a human gate</strong> on the irreversible step</td><td>it deletes bad prod data without <em>being</em> how you delete prod data</td></tr></tbody></table>
<p>That last row is the one that humbled me. A loop that mutates production is exactly where a lazy <code>while</code> would be catastrophic — and the published version refuses to be lazy: explicit inclusion/exclusion rules written <em>before</em> any data changes, an audit that must pass against those rules, reversible operations, and uncertain records parked for a human instead of deleted. That isn't a loop that stops when it's tired. It's a loop engineered to stop only when it's <em>right</em>, with a person on the one action that has no undo.</p>
<p>Notice what's missing from all three: nobody stops on exhaustion. That's not a coincidence — it's the difference between a loop somebody published under their name and the "keep improving X" one-liners I write for myself and never inspect.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="the-exit-is-a-spec">The exit is a spec<a href="https://hassant.github.io/stop-when-right-not-tired#the-exit-is-a-spec" class="hash-link" aria-label="Direct link to The exit is a spec" title="Direct link to The exit is a spec" translate="no">​</a></h2>
<p><a class="" href="https://hassant.github.io/agent-engineering-stack">The stack post</a> put loop engineering on my map as a <em>layer</em>. Here's the part I underplayed: <strong>the stop condition deserves the same up-front design as the goal — because it <em>is</em> the definition of done.</strong> A practical test, learned the hard way this week: if you can't name the exit contract <em>before</em> the first pass, you don't have a loop. You have motion that happens to halt.</p>
<p>That also sharpens <a class="" href="https://hassant.github.io/green-doesnt-mean-done"><em>green doesn't mean done</em></a>: green is only a verdict if the loop stopped <em>because</em> the check passed — not because the check happened to be green when the turns ran out.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="the-field-map">The field map<a href="https://hassant.github.io/stop-when-right-not-tired#the-field-map" class="hash-link" aria-label="Direct link to The field map" title="Direct link to The field map" translate="no">​</a></h2>
<ul>
<li class=""><strong>A loop's value is its exit, not its iteration.</strong> Anyone can write <code>while not done</code>. The engineering is <code>done</code>.</li>
<li class=""><strong>Exhaustion is a failure mode, not a stop condition.</strong> If the honest answer to "why did you stop?" is "I ran out of turns," the loop didn't finish — it fainted.</li>
<li class=""><strong>Optimize the discard, not the generator.</strong> Measure how fast a loop kills dead ends. Now that passes are cheap, the discard is the scarce skill.</li>
<li class=""><strong>The evaluator is the asset.</strong> Code is regenerable; the rubric that judges it is what's worth keeping. A frontier model behind a "did it crash?" check is horsepower wired to nothing.</li>
<li class=""><strong>Irreversible steps get a human gate, by design.</strong> The good prod loops pair a provable audit with a person on the delete. Copy the stop, not just the action.</li>
<li class=""><strong>Naming a loop forces the exit into the open.</strong> That — not the catalog — is why a library of loops is worth anything.</li>
</ul>
<p>So before you write another <code>while not done</code>, write <code>done</code> first. The iterations are free now — the only scarce thing left is a loop that can prove <em>why</em> it stopped, instead of reporting <strong>✅ done</strong> because it got tired.</p>]]></content>
        <author>
            <name>Hassan Tariq</name>
            <uri>https://github.com/hassant</uri>
        </author>
        <category label="ai-agents" term="ai-agents"/>
        <category label="agentic" term="agentic"/>
        <category label="loop-engineering" term="loop-engineering"/>
        <category label="evals" term="evals"/>
        <category label="devtools" term="devtools"/>
    </entry>
    <entry>
        <title type="html"><![CDATA[Inside Agentry: How a Governed Agent Harness Actually Works]]></title>
        <id>https://hassant.github.io/inside-agentry</id>
        <link href="https://hassant.github.io/inside-agentry"/>
        <updated>2026-06-20T00:00:00.000Z</updated>
        <summary type="html"><![CDATA[The problem, stated plainly: you cannot trust an agent you cannot constrain, ground, measure, or reconstruct — and "trust me" is not an architecture. A model emits tokens. Everything that decides whether those tokens are safe to act on in production lives in the code around it.]]></summary>
        <content type="html"><![CDATA[<p><strong>The problem, stated plainly: you cannot trust an agent you cannot constrain, ground, measure, or reconstruct — and "trust me" is not an architecture.</strong> A model emits tokens. Everything that decides whether those tokens are <em>safe to act on in production</em> lives in the code around it.</p>
<p><a class="" href="https://hassant.github.io/agentry-governed-agents">The last Agentry post</a> made that pitch. This one opens the hood — every pillar, in the weeds: the loop, the policy engine, identity, the audit chain, grounding, fleets, evals, and the one-dependency bet that holds it all together.</p>
<!-- -->
<p>One orientation note, then the schematics: every section starts where it should — with the problem that part exists to solve. The full thesis lives in the <a class="" href="https://hassant.github.io/agentry-governed-agents">previous post</a>; here we stay in the weeds.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="the-unit-agentyaml-and-the-runtime-that-wires-it">The unit: <code>agent.yaml</code>, and the runtime that wires it<a href="https://hassant.github.io/inside-agentry#the-unit-agentyaml-and-the-runtime-that-wires-it" class="hash-link" aria-label="Direct link to the-unit-agentyaml-and-the-runtime-that-wires-it" title="Direct link to the-unit-agentyaml-and-the-runtime-that-wires-it" translate="no">​</a></h2>
<p><strong>Problem: agents in the wild sprawl across a prompt, some glue code, a vector store, and a half-documented deploy script. There is no single object you can review, diff, or govern.</strong></p>
<p>Agentry collapses that into one file — the <a class="" href="https://hassant.github.io/agentry-governed-agents">previous post showed its shape</a>. What matters in the weeds is what <code>runtime.run_agent()</code> does with it: resolve the provider, build the tool table, construct the policy engine, attach an identity and an audit log, and hand the lot to a <code>Harness</code> — or, when the file declares an <code>agents:</code> list, to the orchestrator instead.</p>
<p>The seam that makes the whole thing testable is the <strong>model function</strong>. The harness never calls an LLM itself — you inject <code>model_fn(context: list[dict]) -&gt; dict</code>, returning either <code>{"tool_call": {"name", "args"}}</code> or <code>{"done": str}</code>. In production that closure wraps a provider; in a test it returns a canned dict. That one indirection is why the entire suite runs offline with no network and no mocked SDK.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="the-governed-loop-eight-interception-points-fail-closed">The governed loop: eight interception points, fail-closed<a href="https://hassant.github.io/inside-agentry#the-governed-loop-eight-interception-points-fail-closed" class="hash-link" aria-label="Direct link to The governed loop: eight interception points, fail-closed" title="Direct link to The governed loop: eight interception points, fail-closed" translate="no">​</a></h2>
<p><strong>Problem: a control that runs <em>after</em> the model has already decided is theater. By the time you're reading the output, the tool already fired.</strong></p>
<p>So Agentry evaluates policy at <strong>eight lifecycle points</strong> wrapped around every step of the loop — not one "is this output okay?" check at the end:</p>
<div class="language-text codeBlockContainer_Ckt0 theme-code-block" style="--prism-color:#393A34;--prism-background-color:#f6f8fa"><div class="codeBlockContent_QJqH"><pre tabindex="0" class="prism-code language-text codeBlock_bY9V thin-scrollbar" style="color:#393A34;background-color:#f6f8fa"><code class="codeBlockLines_e6Vv"><div class="token-line" style="color:#393A34"><span class="token plain">agent_startup → input → pre_model_call → post_model_call</span><br></div><div class="token-line" style="color:#393A34"><span class="token plain">             → pre_tool_call → post_tool_call → output → agent_shutdown</span><br></div></code></pre></div></div>
<!-- -->
<p>Each point builds a small snapshot (the input text, the drafted response, the pending tool call and its args) and asks the engine for a <code>Verdict</code>. <code>Harness.run()</code> is the sync path, <code>arun()</code> the async one, and they share identical governance semantics. The loop is bounded by <code>max_turns</code> so a model that loops forever returns an honest <code>{"error": "max turns exceeded"}</code> rather than spinning. The return value is small and inspectable: <code>{"answer", "trace"}</code> on success, <code>{"blocked", "verdict", "trace"}</code> when a verdict stopped it, each step recorded.</p>
<p>One honest note, because <a class="" href="https://hassant.github.io/green-doesnt-mean-done">an earlier review caught exactly this</a>: there was a moment when only four of those eight points were actually wired into the loop, and a <code>redact</code> verdict fired but never masked anything. A second-opinion agent on a different model reproduced both with file-and-line receipts. They're wired and enforced now — which is the point of keeping the receipts.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="the-policy-engine-a-swappable-seam-not-a-prompt">The policy engine: a swappable seam, not a prompt<a href="https://hassant.github.io/inside-agentry#the-policy-engine-a-swappable-seam-not-a-prompt" class="hash-link" aria-label="Direct link to The policy engine: a swappable seam, not a prompt" title="Direct link to The policy engine: a swappable seam, not a prompt" translate="no">​</a></h2>
<p><strong>Problem: "please don't email external recipients" in a system prompt is a wish. Enforcement has to be deterministic, inspectable, and runnable outside the model.</strong></p>
<p><code>PolicyEngine.evaluate(point, snapshot)</code> returns one of five verdicts:</p>
<table><thead><tr><th>Verdict</th><th>Effect</th></tr></thead><tbody><tr><td><code>allow</code></td><td>proceed</td></tr><tr><td><code>warn</code></td><td>proceed, record a warning</td></tr><tr><td><code>deny</code></td><td>stop the step with a reason</td></tr><tr><td><code>escalate</code></td><td>block pending a human sign-off</td></tr><tr><td><code>redact</code></td><td>mask the matched content (the rule <strong>must</strong> carry a pattern)</td></tr></tbody></table>
<p>The evaluator is a <strong>seam</strong>, not a hardcode. <code>NativeEvaluator</code> is the deterministic default; <code>OPAEvaluator</code> ships the same decision to an Open Policy Agent server over standard-library HTTP, so the policy that runs in the teaching engine is the policy you can run in prod:</p>
<div class="language-python codeBlockContainer_Ckt0 theme-code-block" style="--prism-color:#393A34;--prism-background-color:#f6f8fa"><div class="codeBlockContent_QJqH"><pre tabindex="0" class="prism-code language-python codeBlock_bY9V thin-scrollbar" style="color:#393A34;background-color:#f6f8fa"><code class="codeBlockLines_e6Vv"><div class="token-line" style="color:#393A34"><span class="token keyword" style="color:#00009f">from</span><span class="token plain"> agentry</span><span class="token punctuation" style="color:#393A34">.</span><span class="token plain">governance </span><span class="token keyword" style="color:#00009f">import</span><span class="token plain"> OPAEvaluator</span><span class="token punctuation" style="color:#393A34">,</span><span class="token plain"> PolicyEngine</span><br></div><div class="token-line" style="color:#393A34"><span class="token plain" style="display:inline-block"></span><br></div><div class="token-line" style="color:#393A34"><span class="token plain">engine </span><span class="token operator" style="color:#393A34">=</span><span class="token plain"> PolicyEngine</span><span class="token punctuation" style="color:#393A34">.</span><span class="token plain">from_path</span><span class="token punctuation" style="color:#393A34">(</span><span class="token plain"></span><br></div><div class="token-line" style="color:#393A34"><span class="token plain">    </span><span class="token string" style="color:#e3116c">"governance/policy/manifest.yaml"</span><span class="token punctuation" style="color:#393A34">,</span><span class="token plain"></span><br></div><div class="token-line" style="color:#393A34"><span class="token plain">    evaluator</span><span class="token operator" style="color:#393A34">=</span><span class="token plain">OPAEvaluator</span><span class="token punctuation" style="color:#393A34">(</span><span class="token string" style="color:#e3116c">"http://127.0.0.1:8181"</span><span class="token punctuation" style="color:#393A34">,</span><span class="token plain"> package</span><span class="token operator" style="color:#393A34">=</span><span class="token string" style="color:#e3116c">"agentry/policy"</span><span class="token punctuation" style="color:#393A34">)</span><span class="token punctuation" style="color:#393A34">,</span><span class="token plain"></span><br></div><div class="token-line" style="color:#393A34"><span class="token plain"></span><span class="token punctuation" style="color:#393A34">)</span><span class="token plain"></span><br></div><div class="token-line" style="color:#393A34"><span class="token plain">verdict </span><span class="token operator" style="color:#393A34">=</span><span class="token plain"> engine</span><span class="token punctuation" style="color:#393A34">.</span><span class="token plain">evaluate</span><span class="token punctuation" style="color:#393A34">(</span><span class="token string" style="color:#e3116c">"pre_tool_call"</span><span class="token punctuation" style="color:#393A34">,</span><span class="token plain"> </span><span class="token punctuation" style="color:#393A34">{</span><span class="token string" style="color:#e3116c">"tool_call"</span><span class="token punctuation" style="color:#393A34">:</span><span class="token plain"> </span><span class="token punctuation" style="color:#393A34">{</span><span class="token string" style="color:#e3116c">"name"</span><span class="token punctuation" style="color:#393A34">:</span><span class="token plain"> </span><span class="token string" style="color:#e3116c">"send_email"</span><span class="token punctuation" style="color:#393A34">}</span><span class="token punctuation" style="color:#393A34">}</span><span class="token punctuation" style="color:#393A34">)</span><br></div></code></pre></div></div>
<p>Both evaluators <strong>fail closed</strong>: an unknown interception point denies, a transport error or malformed response denies, a <code>redact</code> rule with no pattern denies, and <code>from_path(..., profile=...)</code> raises rather than silently running ungoverned if a profile can't load. Rules read declaratively — a <code>when:</code> clause matches on <code>tool</code>, on an <code>arg</code> with <code>matches</code>/<code>not_matches</code> or numeric <code>gte|gt|lte|lt|eq</code>, or on <code>target_matches</code> — and the manifest is mirrored by a <code>policy.rego</code> kept in parity. Three profiles (<code>advisory</code>, <code>balanced</code>, <code>strict</code>) layer hard tool policies on top. The verdict is never the model's discretion; it's a check a reviewer can re-run and watch turn red.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="prove-it-then-contain-it-identity-audit-approval-sandbox">Prove it, then contain it: identity, audit, approval, sandbox<a href="https://hassant.github.io/inside-agentry#prove-it-then-contain-it-identity-audit-approval-sandbox" class="hash-link" aria-label="Direct link to Prove it, then contain it: identity, audit, approval, sandbox" title="Direct link to Prove it, then contain it: identity, audit, approval, sandbox" translate="no">​</a></h2>
<p><strong>Problem: "an agent did something" is not an incident response. And a sub-agent that can hand itself more authority than its parent is a privilege-escalation bug wearing a helpful face.</strong></p>
<p>Four mechanisms, each opt-in and <code>None</code> by default so they never change behavior unless you ask for them:</p>
<ul>
<li class=""><strong>Identity.</strong> Every agent carries a name, a clearance, and a set of scopes. <code>Identity.delegate(...)</code> can only <strong>narrow</strong> — scopes are intersected, clearance is clamped down. A specialist can never out-scope its orchestrator, by construction.</li>
<li class=""><strong>Audit.</strong> Every verdict is appended to a <strong>hash-chained, tamper-evident log</strong>: each record's hash covers the previous one, so altering, reordering, or dropping any entry breaks the chain and a one-line <code>verify()</code> catches it. Pass <code>AuditLog(key=...)</code> and the chain upgrades to HMAC-SHA256.</li>
<li class=""><strong>Approval.</strong> Irreversible writes (a refund, an outbound email, a <code>pull_request</code>) go through a <strong>default-cancel</strong> <code>propose → confirm</code> rail with <strong>one-time tokens that are never persisted</strong>. The model proposes; a human confirms; nothing fires in between.</li>
<li class=""><strong>Sandbox.</strong> Tool execution can run <code>InProcessSandbox</code>, <code>TimeoutSandbox</code>, or <code>SubprocessSandbox</code>. Honest limit, documented rather than hidden: the subprocess isolator can't process-isolate a tool loaded dynamically from a file at runtime — it returns a clear "not isolatable" sentinel instead of pretending.</li>
</ul>
<p>That audit chain is the line I care about most. "Trust me" becomes "here is the signed sequence of what policy was active, what was requested, and why it was allowed or denied."</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="grounded-with-a-contract-that-forbids-orphans">Grounded, with a contract that forbids orphans<a href="https://hassant.github.io/inside-agentry#grounded-with-a-contract-that-forbids-orphans" class="hash-link" aria-label="Direct link to Grounded, with a contract that forbids orphans" title="Direct link to Grounded, with a contract that forbids orphans" translate="no">​</a></h2>
<p><strong>Problem: a model that recalls instead of citing will confidently invent a source, and a docs tree that drifts from the code will confidently lie about it.</strong></p>
<p>Grounding is "cite, don't recall." Before generating, an agent emits a <strong>Learning Packet</strong> naming the approved canon entries that ground a claim, and flags anything it can't ground as <code>UNGROUNDED</code> — out loud — instead of smoothing the gap. Retrieval, when enabled, returns <code>Chunk</code> objects carrying <code>(text, score, provenance, source_id)</code>; a <code>type: vectorstore</code> tool is confined to the agent's own directory and answers with source-and-SHA provenance on every hit, all through a dependency-free keyword retriever.</p>
<p>The same honesty is enforced on the repo itself by a <strong>registration contract</strong>: a construct, pattern, level, or lab <em>exists</em> only if it is (a) a real file, (b) registered in <code>manifest.yml</code>, and (c) cites only canon ids that resolve in the source index. <code>agentry check</code> fails on any dangling citation or orphan. This is not bureaucracy for its own sake — it's the mechanism that makes "all claims are sourced" a build gate instead of a promise.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="fleets-multiplying-the-work-without-multiplying-the-trust-boundary">Fleets: multiplying the work without multiplying the trust boundary<a href="https://hassant.github.io/inside-agentry#fleets-multiplying-the-work-without-multiplying-the-trust-boundary" class="hash-link" aria-label="Direct link to Fleets: multiplying the work without multiplying the trust boundary" title="Direct link to Fleets: multiplying the work without multiplying the trust boundary" translate="no">​</a></h2>
<p><strong>Problem: every time you add an agent, you add a handoff — and a handoff is a place where authority and policy can quietly leak.</strong></p>
<p>When <code>agent.yaml</code> declares a top-level <code>agents:</code> list, the runtime dispatches to <code>orchestrator.run_workflow()</code>, which runs the specialists sequentially under <strong>one shared policy, one audit log, and one tool table</strong>, each with a delegated identity that can only narrow:</p>
<!-- -->
<p>This week I shipped the lab that stress-tests the idea: <strong>L5 — an agent fleet with shared, governed memory.</strong> Five specialists share a small persistent store (decisions, code patterns, tests, design tokens) so the fleet reuses prior work and gets faster as the corpus grows. The sharp part is <em>how</em> they touch it: not raw SQL, but two governed tools — <code>query_memory</code> is read-only, restricted to an allow-list of tables, caps <code>LIMIT</code>, validates filter columns as identifiers, and parameterizes every value; <code>store_memory</code> validates required fields and enforces a payload cap. It defaults to standard-library <code>sqlite3</code>, so the whole lab runs offline with zero setup, with PostgreSQL as an import-guarded option. An agent fleet with institutional memory is a great idea right up until someone hands it a database connection; the lab's entire teaching point is the layer that doesn't.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="evals-that-refuse-to-lie">Evals that refuse to lie<a href="https://hassant.github.io/inside-agentry#evals-that-refuse-to-lie" class="hash-link" aria-label="Direct link to Evals that refuse to lie" title="Direct link to Evals that refuse to lie" translate="no">​</a></h2>
<p><strong>Problem: "it's better now" without a measured delta is a vibe, and an eval that fabricates a score is worse than no eval at all.</strong></p>
<p>Agentry grades with deterministic checks (tool selection, exact match, F1/BLEU/ROUGE/Levenshtein/Jaccard), an optional LLM judge, and RAG metrics — and it is structurally honest:</p>
<ul>
<li class="">A case with <strong>no offline-verifiable expectation</strong> is reported <code>skipped</code>, never faked.</li>
<li class="">An LLM judge with <strong>no model</strong> returns a skip, not an invented number.</li>
<li class="">A run where <strong>everything</strong> skipped exits non-zero, unless you explicitly pass <code>--allow-skips</code>.</li>
</ul>
<p>You measure change over time, too: <code>agentry eval --store runs.jsonl</code> appends each run as JSONL, <code>--compare</code> flags <code>pass → fail</code> regressions against the previous run, and <code>agentry report</code> renders the whole history — plus the verifiable audit trail — into a single self-contained HTML page built from <code>html.escape</code> and string templates. No dashboard framework: the runs are already open JSONL and the CLI already prints the regressions, so a web stack would have added moving parts, not signal.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="the-one-dependency-bet-and-the-surfaces-around-it">The one-dependency bet, and the surfaces around it<a href="https://hassant.github.io/inside-agentry#the-one-dependency-bet-and-the-surfaces-around-it" class="hash-link" aria-label="Direct link to The one-dependency bet, and the surfaces around it" title="Direct link to The one-dependency bet, and the surfaces around it" translate="no">​</a></h2>
<p><strong>Problem: a toolkit whose pitch is <em>auditability</em> has no business being too large to audit — and every vendor SDK is supply-chain surface and lock-in you'll regret.</strong></p>
<p>The whole core runs on <strong>one runtime dependency: <code>pyyaml</code>.</strong> Real LLM execution against four providers (<code>ollama</code>, <code>openai</code>, <code>anthropic</code>, <code>azure_openai</code>) goes through standard-library <code>urllib</code>; each provider is <code>@register</code>-ed, normalizes its response to the same <code>{"tool_call"}</code>/<code>{"done"}</code> contract, exposes typed errors with retry and backoff, imports its SDK lazily if at all, and offers an async fallback. The HTTP governor is <code>http.server</code> with bearer-token auth, a request-body cap, and sanitized errors; a raw ASGI adapter exposes the same routes; an MCP layer speaks stdlib JSON-RPC both ways — consuming external MCP tools <em>governed at the same <code>pre</code>/<code>post_tool_call</code> points as native ones</em>, and exposing Agentry itself as a governed MCP server. Observability is an opt-in OpenTelemetry span (<code>agentry.gov</code>, tagged with <code>point</code>, <code>decision</code>, <code>rule_id</code>) that is a <strong>no-op</strong> when the extra isn't installed. Everything heavy — a JSON-Schema validator, hosted-provider keys, <code>numpy</code>, <code>uvicorn</code> — is an import-guarded extra, never the price of entry.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="the-meta-loop-govern-the-contributions-not-just-the-agents">The meta-loop: govern the contributions, not just the agents<a href="https://hassant.github.io/inside-agentry#the-meta-loop-govern-the-contributions-not-just-the-agents" class="hash-link" aria-label="Direct link to The meta-loop: govern the contributions, not just the agents" title="Direct link to The meta-loop: govern the contributions, not just the agents" translate="no">​</a></h2>
<p><strong>Problem: the same green-checkmark trap that fools you about your own code fools you about a contribution — especially one an AI helped write.</strong></p>
<p>A contributed PR landed this week adding that L5 fleet lab. CI was green across four Python versions. It was also, underneath: leaking internal infrastructure URLs in shipped code and live-network tests, calling methods that didn't exist, using a YAML schema the runtime doesn't understand — and <strong>unregistered</strong>, so <code>agentry check</code> never actually validated the lab it claimed to add. Green meant "no check failed," not "done." Again.</p>
<p>What caught it was the same discipline the platform preaches, turned inward: three reviewers on three different models, each with a hostile brief and read-only access, cross-checked the diff against the registration contract and the invariants. They converged on <em>do-not-merge</em> with file-and-line receipts — the internal-URL leak first. The fix wasn't to reject the idea; it was to strip the proprietary and broken parts, make the lab actually runnable and offline, register it so <code>agentry check</code> could see it, ground it in the real canon, and re-validate. The teaching intent survived; the leak and the lies didn't.</p>
<p>That is the recursion I find most useful: a governed harness, reviewed by governed reviewers, contributing to a repo whose own contract forbids the orphan and the dangling citation. Governance isn't a feature you point at the agent. It's how the whole system — code, docs, contributions, and the AI writing them — keeps itself honest.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="the-field-map">The field map<a href="https://hassant.github.io/inside-agentry#the-field-map" class="hash-link" aria-label="Direct link to The field map" title="Direct link to The field map" translate="no">​</a></h2>
<ul>
<li class=""><strong>Controls run before the action, at eight points, fail-closed.</strong> That's the gap between "the prompt asked the model nicely" and "a verdict a reviewer re-runs and watches turn red."</li>
<li class=""><strong>Authority only narrows; every verdict rides a hash chain.</strong> Delegation can't escalate, and the log can't be quietly edited — containment <em>and</em> proof, not one or the other.</li>
<li class=""><strong>The fleet shares memory through governed tools, never a raw handle.</strong> Read-only, allow-listed, parameterized: institutional memory without the database footgun.</li>
<li class=""><strong>The contract forbids the orphan; the reviewers run on other models.</strong> <code>agentry check</code> won't see an unregistered lab, and three hostile briefs caught the green-but-broken PR. Govern the contributions, not just the agents.</li>
</ul>
<p>The pitch was <em>governed, grounded, evaluated.</em> The schematics are how it earns each of those words — and keeps the receipts. 🧾</p>]]></content>
        <author>
            <name>Hassan Tariq</name>
            <uri>https://github.com/hassant</uri>
        </author>
        <category label="ai-agents" term="ai-agents"/>
        <category label="agentic" term="agentic"/>
        <category label="governance" term="governance"/>
        <category label="evals" term="evals"/>
        <category label="architecture" term="architecture"/>
        <category label="devtools" term="devtools"/>
    </entry>
    <entry>
        <title type="html"><![CDATA[Agentry: Governed Agents That Can Prove What Happened]]></title>
        <id>https://hassant.github.io/agentry-governed-agents</id>
        <link href="https://hassant.github.io/agentry-governed-agents"/>
        <updated>2026-06-16T00:00:00.000Z</updated>
        <summary type="html"><![CDATA[In the last post I reviewed "a toolkit I'd been building." This is that toolkit — and the opinions baked into it.]]></summary>
        <content type="html"><![CDATA[<p>In <a class="" href="https://hassant.github.io/green-doesnt-mean-done">the last post</a> I reviewed "a toolkit I'd been building." This is that toolkit — and the opinions baked into it.</p>
<p>The pitch in one line: a model is one component; everything that makes it <em>safe and useful in production</em> lives in the harness around it. <strong>Agentry</strong> is that harness — governed, grounded, evaluated — with one property I care about a lot: it can <strong>prove what it did.</strong></p>
<!-- -->
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="the-bet-the-harness-is-the-product">The bet: the harness is the product<a href="https://hassant.github.io/agentry-governed-agents#the-bet-the-harness-is-the-product" class="hash-link" aria-label="Direct link to The bet: the harness is the product" title="Direct link to The bet: the harness is the product" translate="no">​</a></h2>
<p><a class="" href="https://hassant.github.io/agent-engineering-stack">I've argued before</a> that the model is the least interesting layer of a real agent. The interesting parts are the controls: what the agent is <em>allowed</em> to do, what it <em>grounds</em> its answers in, how you <em>measure</em> whether it works, and whether you can <em>reconstruct</em> what happened afterward. Bolt those on as an afterthought and you get a demo. Build them in and you get something you can run in front of a customer.</p>
<!-- -->
<p>So agentry's unit isn't a prompt — it's one declarative <code>agent.yaml</code> that carries the model, its tools, its <strong>policy</strong>, its <strong>grounding</strong>, and its <strong>evals</strong> together:</p>
<div class="language-yaml codeBlockContainer_Ckt0 theme-code-block" style="--prism-color:#393A34;--prism-background-color:#f6f8fa"><div class="codeBlockContent_QJqH"><pre tabindex="0" class="prism-code language-yaml codeBlock_bY9V thin-scrollbar" style="color:#393A34;background-color:#f6f8fa"><code class="codeBlockLines_e6Vv"><div class="token-line" style="color:#393A34"><span class="token key atrule" style="color:#00a4db">name</span><span class="token punctuation" style="color:#393A34">:</span><span class="token plain"> support</span><span class="token punctuation" style="color:#393A34">-</span><span class="token plain">triage</span><br></div><div class="token-line" style="color:#393A34"><span class="token plain"></span><span class="token key atrule" style="color:#00a4db">model</span><span class="token punctuation" style="color:#393A34">:</span><span class="token plain"> </span><span class="token punctuation" style="color:#393A34">{</span><span class="token plain"> </span><span class="token key atrule" style="color:#00a4db">provider</span><span class="token punctuation" style="color:#393A34">:</span><span class="token plain"> ollama</span><span class="token punctuation" style="color:#393A34">,</span><span class="token plain"> </span><span class="token key atrule" style="color:#00a4db">name</span><span class="token punctuation" style="color:#393A34">:</span><span class="token plain"> llama3.2</span><span class="token punctuation" style="color:#393A34">:</span><span class="token plain">latest </span><span class="token punctuation" style="color:#393A34">}</span><span class="token plain">   </span><span class="token comment" style="color:#999988;font-style:italic"># or openai / anthropic / azure</span><span class="token plain"></span><br></div><div class="token-line" style="color:#393A34"><span class="token plain"></span><span class="token key atrule" style="color:#00a4db">instructions</span><span class="token punctuation" style="color:#393A34">:</span><span class="token plain"> </span><span class="token punctuation" style="color:#393A34">{</span><span class="token plain"> </span><span class="token key atrule" style="color:#00a4db">file</span><span class="token punctuation" style="color:#393A34">:</span><span class="token plain"> instructions/system.md </span><span class="token punctuation" style="color:#393A34">}</span><span class="token plain"></span><br></div><div class="token-line" style="color:#393A34"><span class="token plain" style="display:inline-block"></span><br></div><div class="token-line" style="color:#393A34"><span class="token plain"></span><span class="token key atrule" style="color:#00a4db">grounding</span><span class="token punctuation" style="color:#393A34">:</span><span class="token plain">                 </span><span class="token comment" style="color:#999988;font-style:italic"># cite, don't recall</span><span class="token plain"></span><br></div><div class="token-line" style="color:#393A34"><span class="token plain">  </span><span class="token key atrule" style="color:#00a4db">sources</span><span class="token punctuation" style="color:#393A34">:</span><span class="token plain"> </span><span class="token punctuation" style="color:#393A34">[</span><span class="token plain">support</span><span class="token punctuation" style="color:#393A34">-</span><span class="token plain">kb</span><span class="token punctuation" style="color:#393A34">]</span><span class="token plain"></span><br></div><div class="token-line" style="color:#393A34"><span class="token plain">  </span><span class="token key atrule" style="color:#00a4db">require_packet</span><span class="token punctuation" style="color:#393A34">:</span><span class="token plain"> </span><span class="token boolean important" style="color:#36acaa">true</span><span class="token plain"></span><br></div><div class="token-line" style="color:#393A34"><span class="token plain" style="display:inline-block"></span><br></div><div class="token-line" style="color:#393A34"><span class="token plain"></span><span class="token key atrule" style="color:#00a4db">governance</span><span class="token punctuation" style="color:#393A34">:</span><span class="token plain">                </span><span class="token comment" style="color:#999988;font-style:italic"># policy as code — or it doesn't ship</span><span class="token plain"></span><br></div><div class="token-line" style="color:#393A34"><span class="token plain">  </span><span class="token key atrule" style="color:#00a4db">profile</span><span class="token punctuation" style="color:#393A34">:</span><span class="token plain"> balanced</span><br></div><div class="token-line" style="color:#393A34"><span class="token plain">  </span><span class="token key atrule" style="color:#00a4db">policy</span><span class="token punctuation" style="color:#393A34">:</span><span class="token plain"> governance/policy/manifest.yaml</span><br></div><div class="token-line" style="color:#393A34"><span class="token plain" style="display:inline-block"></span><br></div><div class="token-line" style="color:#393A34"><span class="token plain"></span><span class="token key atrule" style="color:#00a4db">evaluations</span><span class="token punctuation" style="color:#393A34">:</span><span class="token plain"></span><br></div><div class="token-line" style="color:#393A34"><span class="token plain">  </span><span class="token key atrule" style="color:#00a4db">metrics</span><span class="token punctuation" style="color:#393A34">:</span><span class="token plain"> </span><span class="token punctuation" style="color:#393A34">[</span><span class="token punctuation" style="color:#393A34">{</span><span class="token plain"> </span><span class="token key atrule" style="color:#00a4db">type</span><span class="token punctuation" style="color:#393A34">:</span><span class="token plain"> tool_selection</span><span class="token punctuation" style="color:#393A34">,</span><span class="token plain"> </span><span class="token key atrule" style="color:#00a4db">threshold</span><span class="token punctuation" style="color:#393A34">:</span><span class="token plain"> </span><span class="token number" style="color:#36acaa">1.0</span><span class="token plain"> </span><span class="token punctuation" style="color:#393A34">}</span><span class="token punctuation" style="color:#393A34">]</span><br></div></code></pre></div></div>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="governed-by-default--policy-is-code-not-a-paragraph">Governed by default — policy is code, not a paragraph<a href="https://hassant.github.io/agentry-governed-agents#governed-by-default--policy-is-code-not-a-paragraph" class="hash-link" aria-label="Direct link to Governed by default — policy is code, not a paragraph" title="Direct link to Governed by default — policy is code, not a paragraph" translate="no">​</a></h2>
<p>Here's the opinion that drives the whole design: <strong>governance has to be deterministic and run <em>before</em> the action</strong> — not a paragraph in the system prompt that asks the model nicely to behave.</p>
<p>Agentry evaluates policy at eight lifecycle interception points around every step of the loop — startup, input, before/after the model call, before/after each tool call, output, shutdown — and each check returns one of <code>allow · warn · deny · escalate · redact</code>. It is <strong>fail-closed</strong>: an unknown point or a broken rule denies. The same rules live as a manifest <em>and</em> mirror to OPA/Rego, so the policy that runs in the teaching engine is the policy you can run in production.</p>
<!-- -->
<p>A deny stops the loop. A <code>redact</code> rule actually masks the matched content (an SSN-shaped string never leaves the harness). An <code>escalate</code> is where a human signs off. The point is that none of this is the model's discretion — it's a check a reviewer can re-run and watch turn red.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="prove-what-happened">Prove what happened<a href="https://hassant.github.io/agentry-governed-agents#prove-what-happened" class="hash-link" aria-label="Direct link to Prove what happened" title="Direct link to Prove what happened" translate="no">​</a></h2>
<p>The part I'm most attached to: every verdict is appended to a <strong>hash-chained, tamper-evident audit log</strong>. Each record's hash covers the previous one, so altering, reordering, or dropping any entry breaks the chain — and a one-line <code>verify()</code> catches it. "An agent did something" is not an incident response; "here is the signed sequence of what policy was active, what was requested, and why it was allowed or denied" is.</p>
<p>Agent <strong>identity</strong> rides alongside it: every agent carries a name, a clearance, and a set of scopes, and delegation to a sub-agent can only <em>narrow</em> authority — never escalate it.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="grounded-and-honest-about-evals">Grounded, and honest about evals<a href="https://hassant.github.io/agentry-governed-agents#grounded-and-honest-about-evals" class="hash-link" aria-label="Direct link to Grounded, and honest about evals" title="Direct link to Grounded, and honest about evals" translate="no">​</a></h2>
<p>Two more non-negotiables, both about honesty:</p>
<ul>
<li class=""><strong>Cite, don't recall.</strong> Before generating, an agent emits a Learning Packet naming the approved sources that ground a claim, and flags anything it can't ground as <code>UNGROUNDED</code> — out loud — instead of smoothing over the gap.</li>
<li class=""><strong>Evals never lie.</strong> A check with no offline-verifiable expectation is reported <code>skipped</code>, not faked. An LLM-judge with no model returns a skip, not an invented score. A run where <em>everything</em> skipped exits non-zero. Better to admit than to lie.</li>
</ul>
<p>You measure change, too: <code>agentry eval --store</code> records runs, <code>--compare</code> flags pass→fail regressions, and <code>agentry report</code> renders the whole history — plus the verifiable audit trail — into a single HTML page.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="the-contrarian-bet-one-dependency">The contrarian bet: one dependency<a href="https://hassant.github.io/agentry-governed-agents#the-contrarian-bet-one-dependency" class="hash-link" aria-label="Direct link to The contrarian bet: one dependency" title="Direct link to The contrarian bet: one dependency" translate="no">​</a></h2>
<p>Here's the engineering stance that surprised even me by holding up: <strong>the whole platform runs on one runtime dependency — <code>pyyaml</code>.</strong></p>
<p>Real LLM execution against four providers? Standard-library <code>urllib</code>, no vendor SDKs. The HTTP governance service? <code>http.server</code>. The eval metrics (F1, BLEU, ROUGE)? Hand-rolled, deterministic. The report? <code>html.escape</code> plus string templates plus inline CSS — a self-contained page, no dashboard framework.</p>
<!-- -->
<p>Why bother? A dependency-light core installs anywhere, has almost no supply-chain surface, and is small enough to actually audit — which is a funny thing to skimp on in a project whose whole pitch is <em>auditability</em>. The heavy, optional things (a JSON-Schema validator, hosted-provider keys) are opt-in extras, never the price of entry.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="what-i-deliberately-left-out">What I deliberately left out<a href="https://hassant.github.io/agentry-governed-agents#what-i-deliberately-left-out" class="hash-link" aria-label="Direct link to What I deliberately left out" title="Direct link to What I deliberately left out" translate="no">​</a></h2>
<p>The honest part. Knowing what <em>not</em> to build is most of the work:</p>
<ul>
<li class=""><strong>No dashboard.</strong> I wanted run-history charts, and I nearly reached for a dashboard framework. Then I remembered the data is already open JSONL and a CLI already prints regressions. A web stack would have been rigor cosplay. So <code>agentry report</code> is a read-only, stdlib HTML page — and that's it.</li>
<li class=""><strong>The subprocess sandbox has a real limit.</strong> It isolates importable tools with a timeout and crash-containment, but it <em>can't</em> process-isolate tools loaded dynamically from a file at runtime — it returns a clear "not isolatable" sentinel instead. I documented the caveat rather than hiding it behind green tests.</li>
<li class=""><strong>Scaffolds are labeled scaffolds.</strong> Only the first lab is fully worked; the language clients are preview; tracing is declarative metadata, not emitted spans. The docs say so.</li>
<li class=""><strong>I said no to a memory layer and an autonomous task queue.</strong> They're great — for a <em>different</em> product (an always-on orchestrator). Absorbing them would have been scope creep that broke the dependency bet. Right-size the harness; don't grow it for the sake of growing it.</li>
</ul>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="the-field-map">The field map<a href="https://hassant.github.io/agentry-governed-agents#the-field-map" class="hash-link" aria-label="Direct link to The field map" title="Direct link to The field map" translate="no">​</a></h2>
<ul>
<li class=""><strong>The harness is the product.</strong> The model is the commodity; the controls around it are the engineering.</li>
<li class=""><strong>Governance is code, not a prompt.</strong> Deterministic, fail-closed, before the action — and mirrored to a policy engine you can run in prod.</li>
<li class=""><strong>Prove what happened.</strong> A tamper-evident audit log turns "trust me" into "verify me."</li>
<li class=""><strong>Ground it, and don't fake the evals.</strong> <code>UNGROUNDED</code> and <code>skipped</code> are features. Honesty is how the thing stays trustworthy as it grows.</li>
<li class=""><strong>One dependency.</strong> If the pitch is auditability, the toolkit should be small enough to audit.</li>
<li class=""><strong>Delete the ceremony you don't need.</strong> No dashboard, no memory layer, no orchestrator. Right-size to the failure mode.</li>
</ul>
<p>The agents that behave like senior engineers aren't the ones with the cleverest prompts. They're the ones wrapped in honest controls — and able to show their work afterward.</p>
<p>Govern it. Ground it. Evaluate it. And keep the receipts. 🧾</p>]]></content>
        <author>
            <name>Hassan Tariq</name>
            <uri>https://github.com/hassant</uri>
        </author>
        <category label="ai-agents" term="ai-agents"/>
        <category label="agentic" term="agentic"/>
        <category label="governance" term="governance"/>
        <category label="evals" term="evals"/>
        <category label="devtools" term="devtools"/>
    </entry>
    <entry>
        <title type="html"><![CDATA[Green Doesn't Mean Done: Reviewing an Agent Toolkit with a Sub-Agent]]></title>
        <id>https://hassant.github.io/green-doesnt-mean-done</id>
        <link href="https://hassant.github.io/green-doesnt-mean-done"/>
        <updated>2026-06-15T00:00:00.000Z</updated>
        <summary type="html"><![CDATA[I pointed my coding agent at a toolkit I'd been building — an agent harness with governance, grounding, and evals baked in — and gave it one job: prove it's complete. The validator was green. Fourteen unit tests passed. Every command in the README's quickstart did exactly what the README said.]]></summary>
        <content type="html"><![CDATA[<p>I pointed my coding agent at a toolkit I'd been building — an agent harness with governance, grounding, and evals baked in — and gave it one job: <em>prove it's complete.</em> The validator was green. Fourteen unit tests passed. Every command in the README's quickstart did exactly what the README said.</p>
<p>So the agent did the one thing that actually earns trust: <strong>it refused to believe itself.</strong> It spun up a <em>second</em> agent, on a different model, and told it to try to break every claim the toolkit made. Ten minutes later I had a punch-list of four places where my toolkit was quietly lying to me — with file-and-line receipts for each.</p>
<!-- -->
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="green-is-a-hypothesis-not-a-verdict">Green is a hypothesis, not a verdict<a href="https://hassant.github.io/green-doesnt-mean-done#green-is-a-hypothesis-not-a-verdict" class="hash-link" aria-label="Direct link to Green is a hypothesis, not a verdict" title="Direct link to Green is a hypothesis, not a verdict" translate="no">​</a></h2>
<p>Here's the uncomfortable thing about a passing build: it tells you <em>no check failed</em>. It does <strong>not</strong> tell you the thing is done. Those are different statements, and the gap between them is where production incidents live.</p>
<!-- -->
<p>A green check covers the cases someone <em>thought to write down</em>. The interesting failures are the ones nobody encoded — the promise in the prose that no test ever asserts. You don't find those by re-running your own checks. You find them by handing the work to a reader who doesn't share your blind spots.</p>
<hr>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="dont-grade-your-own-homework">Don't grade your own homework<a href="https://hassant.github.io/green-doesnt-mean-done#dont-grade-your-own-homework" class="hash-link" aria-label="Direct link to Don't grade your own homework" title="Direct link to Don't grade your own homework" translate="no">​</a></h2>
<p>The author of a thing — human or agent — is the worst possible reviewer of it, because the assumptions that produced the bug also produce the test that misses the bug. Self-review re-runs your own mental model. What you need is an <strong>adversarial second reader with different priors.</strong></p>
<p>For an agent, "different priors" has a literal, cheap implementation: <em>a different model.</em> My orchestrator runs on one model; the reviewer it dispatched runs on another (GPT-5.5), with a deliberately hostile brief — "assume the README is marketing; make the code prove every sentence."</p>
<!-- -->
<p>Two things matter in that picture. The reviewer is a <strong>separate context on a separate model</strong>, so it isn't anchored to my framing. And the orchestrator doesn't just wait — it runs its <em>own</em> verification in parallel (the quickstart, the test suite, a few scans) so the duck's claims land against independent evidence instead of replacing one opinion with another.</p>
<hr>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="what-the-duck-found">What the duck found<a href="https://hassant.github.io/green-doesnt-mean-done#what-the-duck-found" class="hash-link" aria-label="Direct link to What the duck found" title="Direct link to What the duck found" translate="no">​</a></h2>
<p>This is the honest part. Every row is a sentence the toolkit advertises, next to what the code actually does when you run it.</p>
<table><thead><tr><th>The promise</th><th>What actually happens</th><th>Where</th></tr></thead><tbody><tr><td>Governance at <strong>8 lifecycle points</strong></td><td>only <strong>4</strong> are wired into the loop; <code>post_tool_call</code> is even evaluated, then its verdict is thrown away</td><td><code>harness.py</code></td></tr><tr><td>A <code>redact</code> verdict scrubs PII at output</td><td><code>redact</code> isn't treated as "blocked", so an SSN flows through the response <strong>unchanged</strong></td><td><code>harness.py</code> + <code>governance.py</code></td></tr><tr><td><code>agentry check</code> refuses a lab with <strong>no policy</strong></td><td>it only checks the <code>governance:</code> <em>key</em> exists — <code>profile: balanced</code> with no policy file sails through</td><td><code>validate.py</code></td></tr><tr><td>Grounding flags <code>UNGROUNDED</code> gaps honestly</td><td><code>"banana governance"</code> comes back fully "grounded", with no gap reported for <em>banana</em></td><td><code>grounding.py</code></td></tr><tr><td>Evals fail loudly</td><td>an all-skipped run prints <code>0/1 passed</code> and exits <strong>0</strong> (success)</td><td><code>cli.py</code></td></tr><tr><td>Block refunds of $1000 or more</td><td><code>$5000</code> slips past the rule's <code>`^[0-9]{4,}`</code> regex and returns <strong>allow</strong></td><td><code>manifest.yaml</code></td></tr></tbody></table>
<p>The reviewer didn't <em>assert</em> these — it <strong>reproduced</strong> them. The redaction one, for example, came back with the actual return value: a harness call whose output still contained <code>123-45-6789</code> and a trace that proudly recorded <code>('output', 'redact')</code> right before handing the secret back. The verdict fired. Nothing acted on it.</p>
<p>To be fair to the toolkit: it's young — a Phase-A MVP — and several of these are the seam between an <em>honest scaffold</em> and an <em>over-eager README</em>. That's exactly why the exercise is worth it. The sub-agent's whole value is finding the gap between what the docs promise and what the code enforces, fast, with receipts — before a user does.</p>
<h3 class="anchor anchorTargetStickyNavbar_Vzrq" id="the-discipline-that-made-the-findings-trustworthy">The discipline that made the findings trustworthy<a href="https://hassant.github.io/green-doesnt-mean-done#the-discipline-that-made-the-findings-trustworthy" class="hash-link" aria-label="Direct link to The discipline that made the findings trustworthy" title="Direct link to The discipline that made the findings trustworthy" translate="no">​</a></h3>
<p>The brief banned vibes. Every finding had to arrive as <strong>claim → reproduction → <code>file:line</code> → fix.</strong> A reviewer that says "the governance feels incomplete" is noise. A reviewer that says "points 5–8 are declared in <code>governance.py</code> and never called in <code>harness.py</code>, here's the deny rule that did nothing" is a work item.</p>
<!-- -->
<p>This is just <a class="" href="https://hassant.github.io/agent-engineering-stack">harness thinking</a> pointed at a review: a finding only counts when it's a check a human can re-run and watch turn red. Everything else is ceremony.</p>
<hr>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="building-on-it-the-sub-agent-is-a-loop-not-a-one-shot">Building on it: the sub-agent is a loop, not a one-shot<a href="https://hassant.github.io/green-doesnt-mean-done#building-on-it-the-sub-agent-is-a-loop-not-a-one-shot" class="hash-link" aria-label="Direct link to Building on it: the sub-agent is a loop, not a one-shot" title="Direct link to Building on it: the sub-agent is a loop, not a one-shot" translate="no">​</a></h2>
<p>The part I want to highlight — because it's where sub-agents stop being a gimmick — is that I kept <em>steering</em> the reviewer while it worked. It wasn't "fire prompt, read essay." It was a conversation with a worker that holds its own context.</p>
<p>Mid-review I realised the scope was bigger than the code: the repo also ships a two-day workshop. So I sent the running agent a follow-up — <em>the workshop is a first-class deliverable too; judge whether its promises match what the toolkit can actually do.</em> It folded that in and immediately found the next class of gap: slides that promise PII redaction and tracing the runtime doesn't implement yet.</p>
<!-- -->
<p>A sub-agent you can talk to mid-flight is a different tool from a prompt you fire and forget. You refine the scope, add a constraint, push back on a weak finding — and it carries the whole prior context forward instead of starting cold.</p>
<hr>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="govern-your-reviewers-too">Govern your reviewers, too<a href="https://hassant.github.io/green-doesnt-mean-done#govern-your-reviewers-too" class="hash-link" aria-label="Direct link to Govern your reviewers, too" title="Direct link to Govern your reviewers, too" translate="no">​</a></h2>
<p>One more lesson, and it's the one people skip. A powerful reviewer is <em>eager</em> — it will happily recommend things your project explicitly forbids: pull in a new dependency, expand the scope, "just add attribution to that source," rewrite a module it was only asked to inspect. Helpful, and wrong.</p>
<p>So the orchestrator governs the reviewer the same way a <a class="" href="https://hassant.github.io/building-baymax-agentic-os">governed harness governs a tool call</a>: the sub-agent <em>proposes</em>, the manager <em>filters against the project's constraints</em>, and only verified, in-bounds findings survive into the report. A second opinion you can't constrain isn't a reviewer — it's a second source of scope creep. The manager's job is to keep the duck honest <em>and</em> on-leash.</p>
<hr>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="the-field-map">The field map<a href="https://hassant.github.io/green-doesnt-mean-done#the-field-map" class="hash-link" aria-label="Direct link to The field map" title="Direct link to The field map" translate="no">​</a></h2>
<ul>
<li class=""><strong>Green is a hypothesis.</strong> A passing build means "no check failed," not "the work is done." Treat it as the start of review, not the end.</li>
<li class=""><strong>Don't grade your own homework.</strong> The thing that built the bug also wrote the test that misses it. Bring a reader with different priors.</li>
<li class=""><strong>Different model, different blind spots.</strong> The cheapest way to get an adversarial reviewer is to run it on another model with a hostile brief.</li>
<li class=""><strong>Demand evidence, not opinions.</strong> <code>file:line</code> plus a reproduction, or it didn't happen. "Feels incomplete" is noise; a re-runnable red check is a work item.</li>
<li class=""><strong>Verify in parallel.</strong> While the reviewer reasons, run the quickstart and the tests yourself, so its claims land against independent facts.</li>
<li class=""><strong>Steer it — it's a loop.</strong> A sub-agent that holds context lets you refine scope and push back mid-flight. Use that.</li>
<li class=""><strong>Govern the reviewer.</strong> A strong second opinion will cheerfully suggest things your constraints forbid. The manager filters.</li>
<li class=""><strong>A human still reads the diff.</strong> The loop produces a trustworthy punch-list; a person decides what ships.</li>
</ul>
<p>The agents that behave like senior engineers aren't the ones that declare victory when the build goes green. They're the ones that get suspicious <em>because</em> it went green — and call in a second pair of eyes that doesn't share their assumptions.</p>
<p>Ship the report, not the vibe. Green is where the review starts. 🦆</p>]]></content>
        <author>
            <name>Hassan Tariq</name>
            <uri>https://github.com/hassant</uri>
        </author>
        <category label="ai-agents" term="ai-agents"/>
        <category label="agentic" term="agentic"/>
        <category label="code-review" term="code-review"/>
        <category label="evals" term="evals"/>
        <category label="devtools" term="devtools"/>
    </entry>
    <entry>
        <title type="html"><![CDATA[Prompt, Context, Harness, Loop: The Real Agent Engineering Stack]]></title>
        <id>https://hassant.github.io/agent-engineering-stack</id>
        <link href="https://hassant.github.io/agent-engineering-stack"/>
        <updated>2026-06-12T00:00:00.000Z</updated>
        <summary type="html"><![CDATA[Last week I asked my coding agent to delete an unused workflow harness, keep the docs honest, and ship a release. It reviewed the change, bumped the version, opened a PR — tripped three CI gates, from four self-inflicted root causes — fixed every one, went green, and shipped. A human signed exactly one thing: the irreversible step.]]></summary>
        <content type="html"><![CDATA[<p>Last week I asked my coding agent to delete an unused workflow harness, keep the docs honest, and ship a release. It reviewed the change, bumped the version, opened a PR — <strong>tripped three CI gates, from four self-inflicted root causes</strong> — fixed every one, went green, and shipped. A human signed exactly one thing: the irreversible step.</p>
<p>That small episode is the whole argument of this post. The thing that made the agent <em>trustworthy</em> wasn't a clever prompt. It was the layers around the prompt.</p>
<!-- -->
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="five-words-everyones-arguing-about">Five words everyone's arguing about<a href="https://hassant.github.io/agent-engineering-stack#five-words-everyones-arguing-about" class="hash-link" aria-label="Direct link to Five words everyone's arguing about" title="Direct link to Five words everyone's arguing about" translate="no">​</a></h2>
<p>"Prompt engineering." "Context engineering." "Harness engineering." "Loop engineering." "Spec-driven development" — people say <em>"speckit."</em> The discourse treats them like competing fads, each year's skill obsoleting the last. On real enterprise agents they're nothing of the sort. <strong>They're layers in a stack</strong> — and most failures I see come from over-investing in one layer while ignoring the one that would actually have caught the bug.</p>
<p>Here's the map I use.</p>
<table><thead><tr><th>Discipline</th><th>The unit</th><th>What it optimizes</th><th>The failure it prevents</th><th>When it's <em>enough</em></th></tr></thead><tbody><tr><td><strong>Prompt engineering</strong></td><td>one instruction</td><td>clarity of a single ask</td><td>vague, ambiguous output</td><td>one-shot, low-stakes tasks</td></tr><tr><td><strong>Context engineering</strong></td><td>the window's contents</td><td><em>what the model sees</em> — sources, memory, tool results</td><td>confident hallucination</td><td>most enterprise reliability</td></tr><tr><td><strong>Harness engineering</strong></td><td>gates and validators</td><td>deterministic enforcement</td><td>mistakes reaching production</td><td>when "looks right" isn't proof</td></tr><tr><td><strong>Loop engineering</strong></td><td>the act → verify cycle</td><td>self-correction</td><td>one bad step becoming a bad outcome</td><td>autonomous, multi-step work</td></tr><tr><td><strong>Spec-driven dev</strong></td><td>the spec</td><td>shared intent before code</td><td>building the wrong thing well</td><td>non-trivial, multi-file changes</td></tr></tbody></table>
<p>Look at the last column. Each discipline has a narrow regime where it is <em>sufficient</em> — and a much larger regime where it is necessary but nowhere near enough.</p>
<hr>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="they-stack-they-dont-compete">They stack, they don't compete<a href="https://hassant.github.io/agent-engineering-stack#they-stack-they-dont-compete" class="hash-link" aria-label="Direct link to They stack, they don't compete" title="Direct link to They stack, they don't compete" translate="no">​</a></h2>
<!-- -->
<p>Read it from the inside out. A <strong>prompt</strong> wraps a model call. <strong>Context</strong> decides what that prompt can even see. A <strong>harness</strong> is the deterministic shell that refuses to let bad output through. The <strong>loop</strong> drives the whole thing — act, observe, verify, repeat — and an explicit <strong>spec</strong> sits on top as the definition of "done."</p>
<p>You can be a world-class prompt engineer and still ship nonsense if the outer layers are missing. The reverse is rarely true: with strong context, a real harness, and a tight loop, even an unremarkable prompt has a real shot at a useful, checkable answer.</p>
<hr>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="context-engineering-the-biggest-lever-in-the-enterprise">Context engineering: the biggest lever in the enterprise<a href="https://hassant.github.io/agent-engineering-stack#context-engineering-the-biggest-lever-in-the-enterprise" class="hash-link" aria-label="Direct link to Context engineering: the biggest lever in the enterprise" title="Direct link to Context engineering: the biggest lever in the enterprise" translate="no">​</a></h2>
<p>If you upgrade only one layer, upgrade this one.</p>
<p>In an enterprise, the model's parametric memory is the <em>least</em> trustworthy thing in the room. Your policies, your architecture standards, your approved patterns, last quarter's decision — none of that is in the weights. The job of context engineering is to put the <strong>right, current, authoritative</strong> material in front of the model <em>before</em> it generates, and to make ungrounded output visible instead of plausible.</p>
<p>The pattern I keep coming back to is <strong>ground before you generate</strong>:</p>
<!-- -->
<p>The branch that makes it enterprise-grade is <strong>"cite or flag."</strong> If there is no approved source, the agent doesn't smooth over the gap — it labels the output ungrounded and says so. A confident <em>"I can't confirm this is standard"</em> beats a fluent fabrication every time. (I wrote about the same instinct for <em>code</em> in <a class="" href="https://hassant.github.io/building-baymax-agentic-os">the Baymax post</a>: don't guess an API, read the source.)</p>
<p>Prompt engineering tweaks the question. Context engineering changes the evidence. At scale, evidence wins.</p>
<hr>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="harness-engineering-deterministic-gates-beat-ceremony">Harness engineering: deterministic gates beat ceremony<a href="https://hassant.github.io/agent-engineering-stack#harness-engineering-deterministic-gates-beat-ceremony" class="hash-link" aria-label="Direct link to Harness engineering: deterministic gates beat ceremony" title="Direct link to Harness engineering: deterministic gates beat ceremony" translate="no">​</a></h2>
<p>A harness is everything around the agent that <strong>fails closed</strong> — schema validators, a build that must stay green, a registration contract, an eval suite, a policy check before a tool fires.</p>
<!-- -->
<p>The mistake is to confuse <em>ceremony</em> with <em>rigor</em>. Ceremony is paperwork that makes a process feel serious. Rigor is a check a reviewer can re-run that turns red when something is actually wrong. They are not the same — and only one of them is load-bearing.</p>
<p>Which brings me to the harness I deleted.</p>
<hr>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="the-speckit-lesson-i-removed-a-harness-and-shipped-faster">The "speckit" lesson: I removed a harness and shipped faster<a href="https://hassant.github.io/agent-engineering-stack#the-speckit-lesson-i-removed-a-harness-and-shipped-faster" class="hash-link" aria-label="Direct link to The &quot;speckit&quot; lesson: I removed a harness and shipped faster" title="Direct link to The &quot;speckit&quot; lesson: I removed a harness and shipped faster" translate="no">​</a></h2>
<p>The change my agent was shipping last week <em>was itself</em> a harness removal — which makes it a clean illustration.</p>
<p>We had vendored the full <strong>GitHub Spec Kit</strong> harness: a <code>/speckit.specify → /speckit.plan → /speckit.tasks → /speckit.implement</code> pipeline (plus a constitution and optional clarify/analyze steps), around fifteen generated agents, and a pile of scripts and templates. It is a genuinely good methodology. But in our repo it had <strong>never been run.</strong> Every guarantee it was meant to provide — grounding, an artifact-registration contract, a green build before merge — was <em>already</em> enforced by two things that run on every change: a one-page spec reviewed in the PR, and a deterministic <code>build.py --check</code> gate.</p>
<!-- -->
<p>So we deleted the harness and kept the gate. The rigor was never in the ceremony — it was in the gate.</p>
<p>The principle generalizes: <strong>right-size the harness to the failure mode.</strong> A heavyweight pipeline nobody exercises isn't rigor — it's rigor <em>cosplay</em>. The deterministic gate a reviewer can rerun in five seconds is the real thing. If your ceremony and your gate enforce the same property, delete the ceremony.</p>
<p>(For the full "build it, test it, delete it" instinct — including a whole execution plane I tore down — see <a class="" href="https://hassant.github.io/building-baymax-agentic-os">the Baymax post</a>. Knowing what to remove is part of the engineering.)</p>
<hr>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="loop-engineering-the-layer-everyone-underestimates">Loop engineering: the layer everyone underestimates<a href="https://hassant.github.io/agent-engineering-stack#loop-engineering-the-layer-everyone-underestimates" class="hash-link" aria-label="Direct link to Loop engineering: the layer everyone underestimates" title="Direct link to Loop engineering: the layer everyone underestimates" translate="no">​</a></h2>
<p>Here's the part of last week's episode I find most instructive.</p>
<!-- -->
<p>The agent did not get it right the first time. It removed <code>speckit</code> from the spell-check dictionary while a changelog still used it; it tripped two <em>separate</em> markdown-lint rules — a blockquote it stacked wrong, and a generated file whose lint-disable header it dropped; and it left two links pointing at a repo that doesn't exist. <strong>Three red checks, four root causes.</strong> Then it read the failures, traced each to a cause it had introduced, fixed all of them, and only <em>then</em> asked a human to merge.</p>
<p>That cycle — <em>act → observe → verify → correct → repeat</em> — is what converts a capable model into a dependable colleague. And notice the dependency: <strong>the loop only has teeth because the harness gave it something real to fail against.</strong> Loop and harness are symbiotic. A loop with no gates is just an agent confidently walking off a cliff in a very tight circle.</p>
<p>This is also where the human belongs — not babysitting every token, but signing the <strong>irreversible</strong> step. The agent ran the whole review-fix-verify loop on its own; a person approved exactly one thing: the merge that cut a public release.</p>
<hr>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="so-how-much-ceremony-does-this-change-need">So how much ceremony does <em>this</em> change need?<a href="https://hassant.github.io/agent-engineering-stack#so-how-much-ceremony-does-this-change-need" class="hash-link" aria-label="Direct link to so-how-much-ceremony-does-this-change-need" title="Direct link to so-how-much-ceremony-does-this-change-need" translate="no">​</a></h2>
<p>The map collapses into a single decision I make dozens of times a week:</p>
<!-- -->
<ul>
<li class=""><strong>Trivial</strong> (typo, dependency bump, doc tweak) → ship the diff. A reviewer can fully judge it.</li>
<li class=""><strong>Non-trivial</strong> (new behavior, multi-file) → a one-page spec for shared intent, then let the gate enforce correctness.</li>
<li class=""><strong>Irreversible / high blast radius</strong> (publish, migrate, delete data) → spec <em>plus</em> a threat model, <em>plus</em> a human on the irreversible action.</li>
</ul>
<p>Match the weight to the risk. Most teams either gate everything (and grind to a halt) or gate nothing (and page themselves at 2am). The skill is calibration.</p>
<hr>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="the-field-map">The field map<a href="https://hassant.github.io/agent-engineering-stack#the-field-map" class="hash-link" aria-label="Direct link to The field map" title="Direct link to The field map" translate="no">​</a></h2>
<ul>
<li class=""><strong>Prompt is table stakes.</strong> Necessary, lowest leverage. Stop polishing it past "clear."</li>
<li class=""><strong>Context is the lever.</strong> Feed the agent real, current, cited source. Make ungrounded output <em>visible</em>, not plausible.</li>
<li class=""><strong>Harness is the safety net.</strong> Deterministic gates that fail closed. A green check beats a confident sentence.</li>
<li class=""><strong>Loop is the trust.</strong> Act → verify → correct. It earns autonomy — but only because the harness gives it something honest to fail against.</li>
<li class=""><strong>Spec is the steering.</strong> A one-page statement of intent stops you building the wrong thing well. Anything heavier has to earn its keep against the failure mode.</li>
<li class=""><strong>Delete the ceremony the gate already covers.</strong> Rigor cosplay is worse than nothing — it costs time <em>and</em> manufactures false confidence.</li>
<li class=""><strong>Humans sign the irreversible step.</strong> Everything else, let the loop run.</li>
</ul>
<p>The agents that behave like senior engineers aren't the ones with the cleverest prompts. They're the ones wrapped in real context, honest gates, and a loop that refuses to declare victory until the checks are green.</p>
<p>Ship the gate. Delete the cosplay. Let the loop run — with a human on the one button that has no undo.</p>]]></content>
        <author>
            <name>Hassan Tariq</name>
            <uri>https://github.com/hassant</uri>
        </author>
        <category label="ai-agents" term="ai-agents"/>
        <category label="agentic" term="agentic"/>
        <category label="context-engineering" term="context-engineering"/>
        <category label="devtools" term="devtools"/>
        <category label="enterprise" term="enterprise"/>
    </entry>
    <entry>
        <title type="html"><![CDATA[Building Baymax: A Personal Agentic OS on Microsoft Scout]]></title>
        <id>https://hassant.github.io/building-baymax-agentic-os</id>
        <link href="https://hassant.github.io/building-baymax-agentic-os"/>
        <updated>2026-06-08T00:00:00.000Z</updated>
        <summary type="html"><![CDATA[How I turned an off-the-shelf AI agent into a disciplined, always-on engineering companion —]]></summary>
        <content type="html"><![CDATA[<p>How I turned an off-the-shelf AI agent into a disciplined, always-on engineering companion —
with skills, memory, isolation, governance, and evals. And the one piece I built, tested, and
deliberately threw away.</p>
<!-- -->
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="the-premise">The premise<a href="https://hassant.github.io/building-baymax-agentic-os#the-premise" class="hash-link" aria-label="Direct link to The premise" title="Direct link to The premise" translate="no">​</a></h2>
<p>Most AI setups stop at "ask a question, get an answer." The unlock is <strong>follow-through</strong>: an agent that holds your priorities, does the mechanical work, and stays honest about it — while a human still signs off on what matters.</p>
<p>I didn't want to build an agent <em>runtime</em>. <strong>Microsoft Scout</strong> already ships one (it's an always-on "Autopilot" with its own Entra identity, native Teams/Outlook/Work IQ, and memory). What was missing was a <strong>methodology</strong> layered on top: how the agent searches for truth, contains risk, decides what needs approval, and proves its work.</p>
<p>I called the result <strong>Baymax</strong> — after the calm, protective companion from <em>Big Hero 6</em>. The metaphor turned out to be the architecture:</p>
<table><thead><tr><th>Baymax the character</th><th>Baymax the agent</th></tr></thead><tbody><tr><td>"Your personal healthcare companion"</td><td>your personal <strong>engineering</strong> companion</td></tr><tr><td>Scans before treating</td><td>search real source before coding</td></tr><tr><td>First, do no harm</td><td>run risky code in a disposable sandbox</td></tr><tr><td>"I cannot deactivate until you are satisfied"</td><td>won't ship without your sign-off</td></tr><tr><td>Keeps a care record</td><td>an append-only audit log</td></tr></tbody></table>
<hr>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="the-architecture-at-a-glance">The architecture at a glance<a href="https://hassant.github.io/building-baymax-agentic-os#the-architecture-at-a-glance" class="hash-link" aria-label="Direct link to The architecture at a glance" title="Direct link to The architecture at a glance" translate="no">​</a></h2>
<p>Scout runs in <strong>two tiers</strong>, and knowing which one a task needs is the whole game.</p>
<!-- -->
<p><strong>Rule of thumb:</strong> M365 coordination → the always-on cloud Autopilot. Local / code / sandbox work → the desktop tier.</p>
<hr>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="the-layers">The layers<a href="https://hassant.github.io/building-baymax-agentic-os#the-layers" class="hash-link" aria-label="Direct link to The layers" title="Direct link to The layers" translate="no">​</a></h2>
<p>Baymax is a thin methodology over Scout, expressed as <strong>composable skills</strong> plus a few conventions.</p>
<!-- -->
<h3 class="anchor anchorTargetStickyNavbar_Vzrq" id="1-orchestrator--scout">1. Orchestrator — Scout<a href="https://hassant.github.io/building-baymax-agentic-os#1-orchestrator--scout" class="hash-link" aria-label="Direct link to 1. Orchestrator — Scout" title="Direct link to 1. Orchestrator — Scout" translate="no">​</a></h3>
<p>The harness matters more than the model. Scout brings native Teams/Work IQ, an Entra identity, and an always-on cloud tier. Baymax adds discipline, not plumbing.</p>
<h3 class="anchor anchorTargetStickyNavbar_Vzrq" id="2-memory--durable--synced">2. Memory — durable &amp; synced<a href="https://hassant.github.io/building-baymax-agentic-os#2-memory--durable--synced" class="hash-link" aria-label="Direct link to 2. Memory — durable &amp; synced" title="Direct link to 2. Memory — durable &amp; synced" translate="no">​</a></h3>
<p>Long-term memory is Scout's native, OneDrive-synced store. Write durable facts and decisions; recall before non-trivial work. Never store secrets.</p>
<h3 class="anchor anchorTargetStickyNavbar_Vzrq" id="3-skills--modular-capability-the-modern-unit">3. Skills — modular capability (the modern unit)<a href="https://hassant.github.io/building-baymax-agentic-os#3-skills--modular-capability-the-modern-unit" class="hash-link" aria-label="Direct link to 3. Skills — modular capability (the modern unit)" title="Direct link to 3. Skills — modular capability (the modern unit)" translate="no">​</a></h3>
<p>Ten Anthropic-style <code>SKILL.md</code> files, loaded by description (no slash commands). Saying <strong>"baymax"</strong> activates the constitution + router, which hands off to the right skill.</p>
<table><thead><tr><th>Skill</th><th>What it does</th></tr></thead><tbody><tr><td><code>agentic-os</code></td><td>constitution + router + identity</td></tr><tr><td><code>governance</code></td><td>budget caps, approval gate, audit log</td></tr><tr><td><code>heartbeat</code></td><td>wake → claim → work → verify → done</td></tr><tr><td><code>daytona-sandbox</code></td><td>disposable, isolated code execution</td></tr><tr><td><code>source-code-context</code></td><td>read real SDK source before coding</td></tr><tr><td><code>teams-workiq-brief</code></td><td>pull Teams/mail/calendar context</td></tr><tr><td><code>agentic-engineering-workflow</code></td><td>the per-feature build loop</td></tr><tr><td><code>code-structure-cleanup</code></td><td>behavior-preserving refactor pass</td></tr><tr><td><code>grep-loop-review-workflow</code></td><td>review-fix loop to merge-ready</td></tr><tr><td><code>eval-harness</code></td><td>turn intent into objective checks</td></tr></tbody></table>
<h3 class="anchor anchorTargetStickyNavbar_Vzrq" id="4-isolation--daytona">4. Isolation — Daytona<a href="https://hassant.github.io/building-baymax-agentic-os#4-isolation--daytona" class="hash-link" aria-label="Direct link to 4. Isolation — Daytona" title="Direct link to 4. Isolation — Daytona" translate="no">​</a></h3>
<p>Blast radius is the bottleneck. LLM-generated, risky, or parallel code runs in a <strong>disposable Daytona sandbox</strong> (a full VM that boots in ~1s and is deleted after), never on the host.</p>
<h3 class="anchor anchorTargetStickyNavbar_Vzrq" id="5-governance--a-person-still-signs-off">5. Governance — a person still signs off<a href="https://hassant.github.io/building-baymax-agentic-os#5-governance--a-person-still-signs-off" class="hash-link" aria-label="Direct link to 5. Governance — a person still signs off" title="Direct link to 5. Governance — a person still signs off" translate="no">​</a></h3>
<p>Ported from the "control plane" idea at personal scale:</p>
<!-- -->
<h3 class="anchor anchorTargetStickyNavbar_Vzrq" id="6-autonomy--the-heartbeat-loop">6. Autonomy — the heartbeat loop<a href="https://hassant.github.io/building-baymax-agentic-os#6-autonomy--the-heartbeat-loop" class="hash-link" aria-label="Direct link to 6. Autonomy — the heartbeat loop" title="Direct link to 6. Autonomy — the heartbeat loop" translate="no">​</a></h3>
<p>A file-per-task queue with <strong>atomic <code>mv</code> checkout</strong> (POSIX rename can't double-grab), so multiple workers never collide. Scout wakes on a schedule, claims the top task, works it, verifies, and files the result.</p>
<!-- -->
<h3 class="anchor anchorTargetStickyNavbar_Vzrq" id="7-evals--prove-it-dont-claim-it">7. Evals — prove it, don't claim it<a href="https://hassant.github.io/building-baymax-agentic-os#7-evals--prove-it-dont-claim-it" class="hash-link" aria-label="Direct link to 7. Evals — prove it, don't claim it" title="Direct link to 7. Evals — prove it, don't claim it" translate="no">​</a></h3>
<p>The modern alternative to prompt/persona theater: encode <em>what "done" means</em> as checks, then verify.</p>
<!-- -->
<p>Deterministic checks (<code>command</code>, <code>file_exists</code>, <code>grep</code>) are run by a tiny runner; qualitative <code>llm_judge</code> checks are <strong>never auto-passed</strong> — the agent grades them against a rubric, with evidence. <em>Better to admit than to lie.</em></p>
<hr>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="the-soul">The soul<a href="https://hassant.github.io/building-baymax-agentic-os#the-soul" class="hash-link" aria-label="Direct link to The soul" title="Direct link to The soul" translate="no">​</a></h2>
<p>Personas as roleplay are obsolete; persona as a <strong>behavioral contract</strong> is not. Baymax's <code>SOUL.md</code> is lean:</p>
<ul>
<li class=""><strong>Voice:</strong> crisp, professional, calm under pressure, lightly deadpan.</li>
<li class=""><strong>Values (ranked):</strong> honesty (check sources, admit unknowns) → thoroughness → the user's wellbeing.</li>
<li class=""><strong>Behavior:</strong> proactive within governance, but <strong>always confirm before anything outbound.</strong></li>
<li class=""><strong>Signatures:</strong> introduces itself, says <em>"Scanning…"</em> before diagnosing, closes with <em>"Are you satisfied with this work?"</em>, and a <em>ba-la-la-la</em> fist-bump on a win.</li>
</ul>
<hr>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="what-i-built-tested-and-deleted">What I built, tested, and deleted<a href="https://hassant.github.io/building-baymax-agentic-os#what-i-built-tested-and-deleted" class="hash-link" aria-label="Direct link to What I built, tested, and deleted" title="Direct link to What I built, tested, and deleted" translate="no">​</a></h2>
<p>The most useful part of the project was an experiment I <strong>killed on purpose</strong>.</p>
<p>I wanted true 24/7 — Teams/email watched even with my Mac off — so I built a second "execution plane": an always-on VPS running the heartbeat, fed by <strong>Power Automate → a webhook</strong> (because a server can't hold M365 tokens). I got it fully working: containerized webhook behind Traefik with TLS, atomic queue handoff over git, the lot.</p>
<p>Then two facts ended it:</p>
<ol>
<li class=""><strong>Conditional Access</strong> flatly blocked any Microsoft Graph token issued to the VPS (<code>AADSTS53003</code>) — so the server could never read my mail directly anyway.</li>
<li class="">The Microsoft blog confirmed Scout's <strong>cloud Autopilot is already always-on</strong> with its own identity — making the entire webhook detour <strong>redundant</strong> for inbox-watching.</li>
</ol>
<!-- -->
<p>So I tore it down — container, systemd units, deploy key, DNS, the lot — and wrote the reason into the audit log. <strong>Knowing what to remove is part of the engineering.</strong></p>
<hr>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="lessons-the-modern-stack">Lessons (the modern stack)<a href="https://hassant.github.io/building-baymax-agentic-os#lessons-the-modern-stack" class="hash-link" aria-label="Direct link to Lessons (the modern stack)" title="Direct link to Lessons (the modern stack)" translate="no">​</a></h2>
<ul>
<li class=""><strong>Context engineering &gt; prompt engineering.</strong> The win is feeding the agent real source, memory, and tool results — not crafting clever prompts. <code>source-code-context</code> is the whole philosophy: <em>don't guess APIs, read the source.</em></li>
<li class=""><strong>Skills are the unit of capability.</strong> Modular, description-triggered, composable.</li>
<li class=""><strong>Evals are the new spec.</strong> Thin specs that generate checks beat big design docs that rot.</li>
<li class=""><strong>Persona is a contract, not a costume.</strong> Encode honesty and guardrails; skip the roleplay.</li>
<li class=""><strong>Governance keeps autonomy safe.</strong> Budget caps, an approval gate, and an audit log are what let you turn the agent loose.</li>
<li class=""><strong>Delete boldly.</strong> The VPS detour taught me more dead than alive.</li>
</ul>
<hr>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="how-i-use-it-day-to-day">How I use it day to day<a href="https://hassant.github.io/building-baymax-agentic-os#how-i-use-it-day-to-day" class="hash-link" aria-label="Direct link to How I use it day to day" title="Direct link to How I use it day to day" translate="no">​</a></h2>
<ol>
<li class="">Say <strong>"baymax"</strong> in Scout → the constitution + persona activate.</li>
<li class="">M365 coordination runs 24/7 on the cloud Autopilot.</li>
<li class="">Code/local work routes through skills; risky code is sandboxed; costly or outbound actions stop and ask.</li>
<li class="">Autonomous work flows through the queue; every task is verified by evals and recorded in the audit log.</li>
</ol>
<blockquote>
<p>Runs with the lights dimmed, not off — a person still signs off.</p>
</blockquote>
<p><em>Ba-la-la-la.</em> 👊</p>]]></content>
        <author>
            <name>Hassan Tariq</name>
            <uri>https://github.com/hassant</uri>
        </author>
        <category label="ai-agents" term="ai-agents"/>
        <category label="agentic" term="agentic"/>
        <category label="scout" term="scout"/>
        <category label="automation" term="automation"/>
        <category label="devtools" term="devtools"/>
    </entry>
</feed>