Skip to main content
On this page

Separate the Role from the Runtime

Tomohiro Iida · Published August 5, 2026 · Updated September 20, 2026

Separate the Role from the Runtime

In summer 2026, part of Netsujo’s development used a fixed sequence: ChatGPT organised the purpose and acceptance criteria, Claude Code implemented in the real repository, Codex reviewed the near-final diff, and a human made the release decision. That split was useful because implementation and evaluation stopped being the same job. The next bottleneck was runtime capacity: quota, concurrent sessions, guard health, and the need to preserve scarce reviewer capacity. We now keep the roles but do not bind each role permanently to one product. Role and Runtime are separate pieces of authority.

Key takeaways

  • Keep design, implementation, deterministic checks, and independent review as distinct roles. Product names are not the authority boundary.
  • Select an execution Runtime only after checking capability, authority, fresh quota telemetry, live capacity, and guard health.
  • Run format, lint, type checks, tests, and builds before LLM review whenever those checks can decide the question deterministically.
  • Independent review remains required when the preferred reviewer Runtime is unavailable. Use another eligible reviewer Runtime or block; do not silently drop the review.
  • Reserve scarce model capacity for higher-value strategic work and cross-runtime quality checks, and keep the final release decision with a human.

The first step was fixing roles to three AIs

In a human development team, a planner organises the purpose, an engineer implements, and a different engineer reviews the code. When one person or a very small team owns the work, that separation weakens. The person who defined the requirement also directs the implementation and judges completion, so the original assumption can survive untested all the way to the end.

RoleInitial runtimeMain workAuthority
Business and designChatGPTPut customer value, purpose, requirements, non-goals, and acceptance criteria into words.Design information and reference material.
ImplementationClaude CodeInspect the repository, change code, run tests and builds, fix.Edit code on the working branch.
Independent auditCodexCompare the diff, the specification, and test results, and report serious defects.Read-only as a rule.
Integration and releaseHumanDecide which findings to accept, the priority, the risk tolerance, and whether to publish.Final decision and responsibility.

Different products do not create independence on their own. Independence comes from separate implementation and reviewer assignments, an exact reviewed HEAD, restricted reviewer authority, explicit evaluation criteria, and an exit condition. The original three-product mapping was an implementation of that idea, not the permanent definition of the roles.

The next bottleneck was allocating scarce runtime capacity

Once the roles were separated, the next constraint was no longer only code quality. A runtime can have quota remaining and still be unable to take a new job because another live session occupies shared capacity. It can also be excluded before quota matters because its execution guard is not healthy. Netsujo therefore treats eligibility, quota, live capacity, and guard health as separate observations.

ObservationQuestion
EligibilityDoes this runtime have the required role capability, authority, and safety state?
QuotaHow much of the relevant usage window remains?
Live capacityCan this runtime accept another job right now?
Guard healthIs the budget and execution boundary demonstrably protected?

On September 19, 2026, Netsujo observed a concrete version of this distinction: Claude Code still had substantial five-hour headroom, but an unrelated legitimate interactive session occupied the shared slot, so a new independent QC run was not started. At the same time Codex was excluded because the budget-guard monitor could not prove a protected state. Remaining tokens alone are not dispatch authority.

The current Netsujo source policy reserves part of Codex capacity for strategic interactive work and cross-runtime QC. The exact thresholds are internal operating values, not universal recommendations.

Why ChatGPT handles the design

My work starts from business development, so the first thing I think about is not code but customers, revenue, and operations. A request such as “let customers enter their goal on the analysis screen” hides several questions: who enters it, for what purpose, one goal or several, free text or a selection, how the input affects the analysis logic, which parts of the existing screen stay untouched, what is explicitly out of scope this time, and what counts as done. Left unstated, an AI fills the missing assumptions by inference. The faster the implementation, the faster a vague assumption becomes code. So before any code, the request is converted into purpose, current state, the problem to solve, target users, mandatory requirements, conditions that must not change, non-goals, acceptance criteria, expected risks, verification method, and open questions.

The quality of a design document is not decided by its length. It is decided by whether the purpose, constraints, acceptance criteria, and non-goals are correct. Specifying file names and implementation details before the repository has been inspected only freezes a wrong assumption in place.

What the implementation role needs from a runtime

A correct design document still meets a repository with existing conventions, dependencies, tests, CI, and prior design decisions. Claude Code was initially the main implementation runtime because it could inspect and modify that real environment. The requirement is now expressed as an implementation role: the assigned runtime must have the required repository capability and mutation authority, must inspect current state before changing it, and must report contradictions between the Task Contract and the repository. Claude Code and Codex can be implementation candidates when the canonical policy says they are eligible; the product name itself does not grant the role.

Independent review is runtime-independent

Codex was the initial independent-review runtime. The current rule is expressed at the role level instead: the reviewer assignment must be independent of the implementation assignment, must review the exact HEAD, and must not gain mutation authority merely because the preferred reviewer is unavailable. If the first reviewer runtime cannot run, another eligible reviewer runtime may be selected. If none is available, the work blocks. A human approval marker or an implementer self-check does not become independent-review evidence.

What the split delivered

  • Vague requirements surface before implementation, while the cost of fixing them is still small.
  • What was asked for, which assumption changed during implementation, and which checks ran all stay in Markdown and the Git diff rather than in someone’s memory.
  • Self-approval is harder to reach. Design, implementation, and evaluation in one continuous conversation tend to evaluate against the AI’s own chosen approach.
  • A human can manage quality through acceptance criteria, forbidden changes, risk, and accept-or-reject decisions without hand-writing every line.
  • A good instruction document becomes an asset. Failure causes, correct constraints, verification commands, and preventive measures move into project rules, CI, and quality gates.
  • There is one deliberate pause before release, which keeps “it works” separate from “it may be published”.

What it cost

The serial arrangement carries a handoff tax. ChatGPT, Claude Code, and Codex each re-read the same background, and the longer the design document grows, the more effort the implementer and the auditor spend locating what they need. Using several AIs is not by itself an efficiency gain.

  • A ChatGPT design can drift from the repository. Specifying file layout and function names without having inspected the code pulls Claude Code toward a wrong design.
  • Introducing independent review only just before completion means architecture and data-model problems are found last, when the fix is widest.
  • Sending problems that format, lint, type checking, unit tests, or a build can decide to an LLM raises cost and dilutes attention on the serious findings.
  • A false sense of safety appears. Several AIs reaching the same conclusion does not make it correct; given the same vague requirement, different AIs can err in the same direction.
  • When both Claude Code and Codex edit the same branch, ownership of a change becomes ambiguous, and fixes get reverted, unrelated diffs get mixed in, and the same problem is investigated twice.
  • AI keeps producing improvement candidates, so “just one more fix” blurs the line between a release blocker and a post-release improvement.

Review reliability is decided by clear ground truth, the latest diff, findings with a stated basis, separated authority, and an exit condition. It is not decided by how many AIs looked at it.

What actually happened at Netsujo

A request to add a customer goal field to an existing analysis screen grew, through successive improvement instructions, into a redesign of the whole screen. The cause was less the capability of the AI than our failure to write down that the existing screen structure was to be preserved and that only the input field and its connection to the analysis were in scope. A vague request was implemented at high speed. Non-goals and invariants became mandatory fields in the Task Contract after that. Separately, letting both Claude Code and Codex make changes on GitHub made the state of the main branch and the uncommitted changes hard to track, because the owned area, working branch, source of truth, and integration owner were undefined. The standard now is one task per branch, no direct changes to main, Codex read-only as a rule, implementation fixes returned to Claude Code, merge, push, and deploy only on an explicit human decision, and no double review of the same head SHA.

The defect that passed every check

The customer-facing analysis screen for Netsujo SIGNAL lets a user pick a period and a comparison condition, including “no comparison”. The expected behaviour is unambiguous: when “no comparison” is selected, no section uses previous-period data. After the implementation, comparing the Task Contract against the diff with Codex surfaced the real state. Some sections honoured the setting. Others kept receiving the previous-period specification. The screen rendered, the type check passed, the existing tests passed, and the build succeeded. The code was not broken. The meaning of some results no longer matched the condition the user had chosen. This class of defect is hard to catch with type checking: if the previous-period value has the right type and the function accepts it, TypeScript has nothing to say. If the existing test only asserts that data can be fetched, a contract violation like “pass no previous period in any section when comparison is off” goes straight through. What a diff review examines is not syntax but the flow of data. The fix did not stop at the missing application sites. We added a regression test for the no-comparison case, a script that checks for unapplied analysis conditions, and wiring for both into the package scripts and GitHub Actions CI. What mattered was not that Codex found the problem once. It was converting the finding into a check that stops the same defect mechanically next time. Fixing only what an AI review comment mentions lets the same defect reappear in another section.

A check you wrote has the holes you cannot imagine

Moving a finding into a check is not the end. After adding the check script above, a later diff review reported a defect in the check itself: the comment-stripping logic treated the double slash inside a URL in a string literal as the start of a comment and silently skipped the rest of that line. The check appeared to work and quietly missed lines written a particular way, and the miss never appeared in its output. Whoever writes a check is poorly placed to imagine what it cannot catch, because the tests and the check script both come from the same person, and no check gets written where the blind spot is. So check scripts are review targets too. A green CI run is not evidence that a check worked as intended. Every time a check is added, a second pair of eyes has to establish what it fails to catch.

Assign work by permissions, unit of work, and acceptance criteria

Tool names are a poor way to define roles over time: tools change, and the same tool can do different things under different permissions. We therefore define an assignment by three axes.

AxisWhat to defineWhat happens if it is undefined
PermissionsWhether the role is read-only, may change files, may commit and push, or may release to production.A role intended only to inspect can make changes, and ownership becomes untraceable.
Unit of workThe task, branch, or diff that bounds the assignment.Several roles edit the same surface and create integration debt.
Acceptance criteriaMachine-check results, the required evidence, and human approval.A completion report proves nothing concrete.
RolePermissionsUnit of workAcceptance criteria
Design and requirementsRead and draft documents; no repository changes.One contract per task.Scope, non-goals, and completion conditions are explicit.
ImplementationChange repository files and commit and push to a branch.One task and one branch.Lint, types, tests, and build pass.
Independent reviewRead-only; no automatic fixes, commits, or pushes.One diff-limited bundle.Findings stay within the supplied diff.
Machine checksReturn a decision and change nothing.One diff.The result is deterministic.
Release decisionAuthority to execute a production release.One release.A human has approved the subject, impact, and release time.

The important boundary is that the independent reviewer cannot modify the subject it reviews. An implementer still needs to investigate the cause, make the change, and test their own work. That self-check must not be counted as the independent review, because a mistaken implementation premise can otherwise survive under the same premise.

This is Netsujo’s current operating model, not a universal allocation for every development team. The appropriate boundaries vary with organisation size, data sensitivity, and release frequency.

The review requirement survives runtime changes

The reasonable order is: define the purpose and risk; write a Task Contract; assign an eligible implementation runtime; inspect the repository before mutation; work within one bounded assignment; run deterministic checks; fix the candidate; bind the candidate to an exact HEAD; run one independent review on that subject; and let the Controller or human integrate the evidence before release.

Risk classification is not used to delete the review requirement. It is used to select reasoning strength, scope, and any additional security review.

We stopped using risk classification to decide whether to review

Previously, low-risk changes skipped the external review and medium-risk changes were reviewed only when a condition matched. The idea was to allocate reasoning budget in proportion to risk.

That arrangement had a hole: when the classification itself is wrong, nothing is detected. A diff classified as low was merged without review, and defects that presented customers with a state contrary to fact were found afterwards. Improving classification accuracy is one response, but a classification can still be wrong, and a design whose detection drops to zero when it is wrong cannot be repaired by better classification.

So we went back to putting every diff that passes the machine checks through one external review, regardless of its content or its size. This also avoids the inversion in which larger diffs receive less review.

Today there are only three cases in which no review is called.

  • There are zero lines of reviewable text in the diff, for example an image-only replacement.
  • The machine checks failed, or were not run. Fix that first; not run is not the same as passed.
  • The automatic review attempts for that diff are used up, after which a human decides.

Risk classification selects the reasoning effort

RiskExample changeWhat the classification is used for
HighAuthentication, authorization, billing, database migrations, customer data, public APIs, diagnosis and scoring logic, production environment and deploy settings.Review at high reasoning effort, and confirm design questions before implementation.
Not highEverything else, including wording, articles, CSS, and documentation.Review at medium reasoning effort, once, after the machine checks pass.

A file that matches no rule is treated as the higher category rather than dropped into the lower one. Not being able to judge something is not the same as it being safe.

This allocates the strength of the reasoning budget by risk, not the number of AIs and not the size of the change. Whether a review happens at all is not part of that allocation. A very large change is still not handed over in one piece; split the pull request or the unit of change instead.

From a detailed order to a verifiable contract

FieldWhat it records
goalWhat is being achieved.
user_valueWho benefits and how.
scopeWhat changes this time.
non_goalsWhat does not change this time.
invariantsConditions that must never break.
acceptance_criteriaConditions that decide completion.
risk_levelLow, medium, or high.
test_planThe checks that will run.
evidenceThe evidence required in the completion report.
stop_conditionsWhen to stop auto-fixing and return to a human.

Non-goals, invariants, and stop conditions matter most. If adding a goal field must not replace the existing analysis screen, write that down. If three rounds of fixes produce the same failure, hand it back to a human instead of continuing. Weakening tests, type settings, or security settings in order to obtain a PASS is prohibited. Raising an AI’s self-correction ability depends more on clear pass and stop conditions than on more iterations.

From full re-scan to exact-HEAD diff review

The independent reviewer receives a bounded subject: the Task Contract, base and head SHAs, the target diff, the minimum related code, a summary of machine checks, the risks to focus on, and the instruction to return actionable findings only. The reviewer runtime may change; the exact subject and review contract do not.

  • No unbounded re-exploration of the whole repository.
  • No long successful logs.
  • No aimless full text of lockfiles or generated output.
  • No refactoring proposals outside the diff.
  • No duplicate review of the same head SHA.
  • No automatic commit, push, merge, or deploy by the reviewer.

The response format is specified too: per finding, the severity, the file and location, the contract that was broken, the impact on users or operations, the basis, the minimal fix, and the missing regression test. Where nothing applies, the reviewer writes “no applicable findings” and is explicitly not allowed to write “this is safe”. Those are not the same statement. Nothing was found within the scope reviewed, the evidence supplied, and the criteria used. Release still needs machine checks, authority design, operational confirmation, and human approval. AI review is one layer of quality assurance, not a replacement for tests, CI, branch protection, or human approval.

Do not confuse internal judges with independent review

An implementation runtime can contain useful internal roles: a Builder, deterministic checks, a self-check, and a manager for stop conditions. Those checks speed the implementation loop, but they are not promoted into independent-review evidence when they share the implementation assignment. Independent review uses a separate assignment and exact subject. What changes dynamically is the eligible runtime that performs the role.

Measure whether it actually works

A feeling that several AIs raise quality does not justify the operating model. We need lead time, first-pass success, fix rounds, reviewer runtime, reasoning strength, input and output volume, duplicate reviews of the same SHA, serious findings, findings converted into regression checks, incorrect findings, post-release defects, rollbacks, and human integration time. Compare those over fixed windows. If quality is unchanged and inference cost rises, narrow the review input or lower reasoning strength where appropriate; do not silently remove the review requirement. OpenAI’s January 2026 Datadog case study remains a useful external example of a review layer catching system-level interactions, but it is a single company case study rather than a general detection rate.

More AI does not distribute human responsibility

Lining up ChatGPT, Claude Code, and Codex does not produce a team of three engineers. Each is a probabilistic system that carries no responsibility. Given a wrong goal, it will refine the wrong goal. Given weak acceptance criteria, it will call a presentable but unfinished state a PASS. With nobody to stop the release, it will keep producing improvement candidates. What stays with the human is deciding what to build, what not to change, which risks to accept, whether a finding is a specification change or a defect, where to stop fixing, and who answers for it after release. A development lead in the agent era writes less code and designs more roles, authority, evaluation criteria, and stop conditions.

Conclusion

  • Structure the instruction document as a Task Contract.
  • Separate Role from Runtime and assign only an eligible execution surface.
  • Run deterministic checks before LLM review.
  • Keep independent review separate from the implementation assignment and bind it to the exact HEAD.
  • Allocate scarce runtime capacity explicitly and keep integration responsibility with the Controller or human.

Products and models will change. What remains is the operating design: which role receives what, what it is not given, on what basis something passes, and where it stops.

Frequently asked questions

Are ChatGPT, Claude Code, and Codex still fixed to one role each?
No. The roles stay distinct, but the runtime is selected from eligible execution surfaces. ChatGPT is often useful for business and requirements work; Claude Code and Codex can be implementation or reviewer candidates when capability, authority, quota, and live-capacity policy allows it.
Should every change use the same reviewer runtime?
No. Reviewable text diffs keep the independent-review requirement, while risk and complexity select reasoning strength and the eligible reviewer runtime. If the preferred runtime is unavailable, use another eligible reviewer or block; do not turn unavailability into a review waiver.
Does using different AIs make the review independent?
Different products do not guarantee independence. It takes a separate reviewer assignment, an exact reviewed HEAD, explicit evaluation criteria, bounded authority, findings with a stated basis, and a defined stop condition. An implementer self-check is still a self-check even when the same product can act as a reviewer elsewhere.
Should a human still read the code?
The need to read every line by hand goes down, but purpose, risk, acceptance criteria, accepting or rejecting findings, and the release decision stay with the human. For high-risk changes, a human also reads the important diffs and test results.
Does letting AIs auto-fix each other raise quality?
Increasing iterations alone raises cost, wrong fixes, and infinite loops. Define the maximum number of fix rounds, the stop rule for a repeating failure, the budget ceiling, and the escalation condition first.

AI agent development and operations incident logThe series hub lists every incident in this record and the order to read them in.

A repeated instruction is a system bugThe separate article covers moving repeated human attention into a Controller, Evidence Ledger, and Gates.

The same work item, different chat namesThe separate article covers how work is identified when it moves between tools.

We design roles, authority, evaluation criteria, and stop conditions alongside the implementation itself.

Talk to Netsujo about AI implementation

Continue the series

Next: #01 Give AI review a budget and an exit condition

Back to all 12 episodes of the AI Agent Development Incident Log

What Broke When We Put AI Agents to Work #11

We delegated the work and still worked from morning to midnight

A poorly designed Executor turns the human into a command proxy

Agents repeatedly returned operations they could not execute to the human. Operator Protection requires capability-aware planning, bounded execution envelopes, leases, session TTLs, WIP limits, and a rule that humans decide exceptions rather than execute raw commands.

Incident Card

The promise was that agents would absorb implementation detail while the human focused on important decisions. In practice, the operator kept receiving requests for one more command, one merge, one branch operation, or one recovery action.

The agents were working, but the human had become the fallback executor for every missing capability.

FieldObserved condition
SymptomAgents delegated operations they could not perform back to the human.
Direct impactThe human executed branch, worktree, and pull-request operations across tasks.
Hidden impactOld sessions and branch contamination created additional recovery labor.
Incorrect premiseA little manual completion at the end still counts as automation.

One command is not a free interruption

A thirty-second command can require several minutes of cognitive recovery. The operator must identify the task, reconstruct state, verify branch and worktree, assess side effects, execute the operation, return the result, and then recover the previous context.

Automation quality therefore needs to measure human interruptions and operator minutes, not only commands executed by agents.

Return decisions to humans, not operations

Human-in-the-loop should mean human judgment at authority or risk boundaries. It should not mean using a person as an authenticated shell runner.

A Human Gate should contain the decision, reason, target artifact, expected head, risk, available options, recommendation, and the automated action that follows approval.

Do not start a plan with no capable executor

If a plan contains a merge step but the assigned executor cannot merge, the plan is invalid before execution starts. Capability gaps should be resolved during planning, not discovered after every earlier step has completed.

  • Route the step to a capable executor.
  • Require human approval and then let an authorized executor perform the mutation.
  • Return PLAN_INVALID when no legitimate execution authority exists.

The Executor operates inside an Execution Envelope

Executors change the world: they edit files, run commands, commit, push, and call mutation APIs. Strong capability needs a narrow envelope rather than broad discretion.

  • Task ID, plan version, and step ID
  • Expected input SHA
  • Allowed branch and worktree
  • Allowed write scope
  • Allowed tools and forbidden actions
  • Stop conditions
  • Lease and fencing token

Branch contamination is an Operator Protection failure

When an unrelated session can write commits onto the canonical task branch, the problem is not merely an untidy Git history. The ownership and mutation boundary has failed.

Before any mutation, verify canonical session, claimed worktree, claimed branch, expected head, allowed path scope, and active lease. If contamination is discovered, freeze mutation and enter a recovery task rather than performing an automatic reset, revert, or force push.

Old sessions need TTL and lease expiry

Long-lived sessions retain old heads, plans, policies, and ownership assumptions. A session whose heartbeat stops should lose write authority. Resume should hydrate current state from the system and issue a fresh execution envelope.

WIP limits protect the human decision queue

Agents can start new work more cheaply than humans can review it. Without WIP limits, the machine produces pull requests, reports, approvals, and collisions faster than the operator can resolve them.

Limit mutating tasks, unresolved HUMAN_REQUIRED states, and unreviewed pull requests. When a limit is reached, finish, retire, or replan existing work before starting more mutation.

Make session lifecycle explicit

  • ACTIVE: executing an authorized step.
  • FROZEN: mutation disabled while state is preserved.
  • WAITING: blocked on an external condition.
  • HANDOFF_PENDING: preparing transfer to another executor.
  • RECOVERY: auditing contamination or state inconsistency.
  • TERMINATED: lease and mutation authority revoked.

Rule

Return decisions, not raw operations, to the human. An Executor may perform only actions for which capability, scope, authority, and lease are valid, and an infeasible plan must not be completed through manual command delegation.

Guardrail

  • Check executor capability before plan execution.
  • Do not delegate raw commands to the human operator.
  • Attach risk, evidence, and exact artifact identity to Human Gates.
  • Perform approved operations through an authorized executor.
  • Bind branch, worktree, write scope, and lease to the task.
  • Check expected head before every mutation.
  • Require session TTL and heartbeat.
  • Stop new mutation when WIP limits are exceeded.
  • Implement FREEZE, RESUME, HANDOFF, RECOVERY, and TERMINATE states.

Evidence

  • Raw commands delegated to humans per task
  • Human interruption count
  • Time from Human Gate creation to decision
  • Mid-plan stops caused by capability gaps
  • Average and maximum session lifetime
  • Mutation attempts from expired leases
  • Branch contamination incidents
  • Human minutes spent on recovery
  • WIP-limit violations
  • Human operation time required to reach DONE

Remaining Risk

Over-protection can create unnecessary stalls if every ambiguity becomes a human gate. The controller should route, retry, re-observe, or replan ordinary uncertainty before escalating.

Human attention remains necessary for risk acceptance, goal changes, and external authority. Operator Protection exists to reserve human effort for those decisions instead of routine execution.