Separate the Role from the Runtime
Tomohiro Iida · Published August 5, 2026 · Updated September 20, 2026
In summer 2026, part of Netsujo’s development used a fixed sequence: ChatGPT organised the purpose and acceptance criteria, Claude Code implemented in the real repository, Codex reviewed the near-final diff, and a human made the release decision. That split was useful because implementation and evaluation stopped being the same job. The next bottleneck was runtime capacity: quota, concurrent sessions, guard health, and the need to preserve scarce reviewer capacity. We now keep the roles but do not bind each role permanently to one product. Role and Runtime are separate pieces of authority.
Key takeaways
- Keep design, implementation, deterministic checks, and independent review as distinct roles. Product names are not the authority boundary.
- Select an execution Runtime only after checking capability, authority, fresh quota telemetry, live capacity, and guard health.
- Run format, lint, type checks, tests, and builds before LLM review whenever those checks can decide the question deterministically.
- Independent review remains required when the preferred reviewer Runtime is unavailable. Use another eligible reviewer Runtime or block; do not silently drop the review.
- Reserve scarce model capacity for higher-value strategic work and cross-runtime quality checks, and keep the final release decision with a human.
The first step was fixing roles to three AIs
In a human development team, a planner organises the purpose, an engineer implements, and a different engineer reviews the code. When one person or a very small team owns the work, that separation weakens. The person who defined the requirement also directs the implementation and judges completion, so the original assumption can survive untested all the way to the end.
| Role | Initial runtime | Main work | Authority |
|---|---|---|---|
| Business and design | ChatGPT | Put customer value, purpose, requirements, non-goals, and acceptance criteria into words. | Design information and reference material. |
| Implementation | Claude Code | Inspect the repository, change code, run tests and builds, fix. | Edit code on the working branch. |
| Independent audit | Codex | Compare the diff, the specification, and test results, and report serious defects. | Read-only as a rule. |
| Integration and release | Human | Decide which findings to accept, the priority, the risk tolerance, and whether to publish. | Final decision and responsibility. |
Different products do not create independence on their own. Independence comes from separate implementation and reviewer assignments, an exact reviewed HEAD, restricted reviewer authority, explicit evaluation criteria, and an exit condition. The original three-product mapping was an implementation of that idea, not the permanent definition of the roles.
The next bottleneck was allocating scarce runtime capacity
Once the roles were separated, the next constraint was no longer only code quality. A runtime can have quota remaining and still be unable to take a new job because another live session occupies shared capacity. It can also be excluded before quota matters because its execution guard is not healthy. Netsujo therefore treats eligibility, quota, live capacity, and guard health as separate observations.
| Observation | Question |
|---|---|
| Eligibility | Does this runtime have the required role capability, authority, and safety state? |
| Quota | How much of the relevant usage window remains? |
| Live capacity | Can this runtime accept another job right now? |
| Guard health | Is the budget and execution boundary demonstrably protected? |
On September 19, 2026, Netsujo observed a concrete version of this distinction: Claude Code still had substantial five-hour headroom, but an unrelated legitimate interactive session occupied the shared slot, so a new independent QC run was not started. At the same time Codex was excluded because the budget-guard monitor could not prove a protected state. Remaining tokens alone are not dispatch authority.
The current Netsujo source policy reserves part of Codex capacity for strategic interactive work and cross-runtime QC. The exact thresholds are internal operating values, not universal recommendations.
Why ChatGPT handles the design
My work starts from business development, so the first thing I think about is not code but customers, revenue, and operations. A request such as “let customers enter their goal on the analysis screen” hides several questions: who enters it, for what purpose, one goal or several, free text or a selection, how the input affects the analysis logic, which parts of the existing screen stay untouched, what is explicitly out of scope this time, and what counts as done. Left unstated, an AI fills the missing assumptions by inference. The faster the implementation, the faster a vague assumption becomes code. So before any code, the request is converted into purpose, current state, the problem to solve, target users, mandatory requirements, conditions that must not change, non-goals, acceptance criteria, expected risks, verification method, and open questions.
The quality of a design document is not decided by its length. It is decided by whether the purpose, constraints, acceptance criteria, and non-goals are correct. Specifying file names and implementation details before the repository has been inspected only freezes a wrong assumption in place.
What the implementation role needs from a runtime
A correct design document still meets a repository with existing conventions, dependencies, tests, CI, and prior design decisions. Claude Code was initially the main implementation runtime because it could inspect and modify that real environment. The requirement is now expressed as an implementation role: the assigned runtime must have the required repository capability and mutation authority, must inspect current state before changing it, and must report contradictions between the Task Contract and the repository. Claude Code and Codex can be implementation candidates when the canonical policy says they are eligible; the product name itself does not grant the role.
Independent review is runtime-independent
Codex was the initial independent-review runtime. The current rule is expressed at the role level instead: the reviewer assignment must be independent of the implementation assignment, must review the exact HEAD, and must not gain mutation authority merely because the preferred reviewer is unavailable. If the first reviewer runtime cannot run, another eligible reviewer runtime may be selected. If none is available, the work blocks. A human approval marker or an implementer self-check does not become independent-review evidence.
What the split delivered
- Vague requirements surface before implementation, while the cost of fixing them is still small.
- What was asked for, which assumption changed during implementation, and which checks ran all stay in Markdown and the Git diff rather than in someone’s memory.
- Self-approval is harder to reach. Design, implementation, and evaluation in one continuous conversation tend to evaluate against the AI’s own chosen approach.
- A human can manage quality through acceptance criteria, forbidden changes, risk, and accept-or-reject decisions without hand-writing every line.
- A good instruction document becomes an asset. Failure causes, correct constraints, verification commands, and preventive measures move into project rules, CI, and quality gates.
- There is one deliberate pause before release, which keeps “it works” separate from “it may be published”.
What it cost
The serial arrangement carries a handoff tax. ChatGPT, Claude Code, and Codex each re-read the same background, and the longer the design document grows, the more effort the implementer and the auditor spend locating what they need. Using several AIs is not by itself an efficiency gain.
- A ChatGPT design can drift from the repository. Specifying file layout and function names without having inspected the code pulls Claude Code toward a wrong design.
- Introducing independent review only just before completion means architecture and data-model problems are found last, when the fix is widest.
- Sending problems that format, lint, type checking, unit tests, or a build can decide to an LLM raises cost and dilutes attention on the serious findings.
- A false sense of safety appears. Several AIs reaching the same conclusion does not make it correct; given the same vague requirement, different AIs can err in the same direction.
- When both Claude Code and Codex edit the same branch, ownership of a change becomes ambiguous, and fixes get reverted, unrelated diffs get mixed in, and the same problem is investigated twice.
- AI keeps producing improvement candidates, so “just one more fix” blurs the line between a release blocker and a post-release improvement.
Review reliability is decided by clear ground truth, the latest diff, findings with a stated basis, separated authority, and an exit condition. It is not decided by how many AIs looked at it.
What actually happened at Netsujo
A request to add a customer goal field to an existing analysis screen grew, through successive improvement instructions, into a redesign of the whole screen. The cause was less the capability of the AI than our failure to write down that the existing screen structure was to be preserved and that only the input field and its connection to the analysis were in scope. A vague request was implemented at high speed. Non-goals and invariants became mandatory fields in the Task Contract after that. Separately, letting both Claude Code and Codex make changes on GitHub made the state of the main branch and the uncommitted changes hard to track, because the owned area, working branch, source of truth, and integration owner were undefined. The standard now is one task per branch, no direct changes to main, Codex read-only as a rule, implementation fixes returned to Claude Code, merge, push, and deploy only on an explicit human decision, and no double review of the same head SHA.
The defect that passed every check
The customer-facing analysis screen for Netsujo SIGNAL lets a user pick a period and a comparison condition, including “no comparison”. The expected behaviour is unambiguous: when “no comparison” is selected, no section uses previous-period data. After the implementation, comparing the Task Contract against the diff with Codex surfaced the real state. Some sections honoured the setting. Others kept receiving the previous-period specification. The screen rendered, the type check passed, the existing tests passed, and the build succeeded. The code was not broken. The meaning of some results no longer matched the condition the user had chosen. This class of defect is hard to catch with type checking: if the previous-period value has the right type and the function accepts it, TypeScript has nothing to say. If the existing test only asserts that data can be fetched, a contract violation like “pass no previous period in any section when comparison is off” goes straight through. What a diff review examines is not syntax but the flow of data. The fix did not stop at the missing application sites. We added a regression test for the no-comparison case, a script that checks for unapplied analysis conditions, and wiring for both into the package scripts and GitHub Actions CI. What mattered was not that Codex found the problem once. It was converting the finding into a check that stops the same defect mechanically next time. Fixing only what an AI review comment mentions lets the same defect reappear in another section.
A check you wrote has the holes you cannot imagine
Moving a finding into a check is not the end. After adding the check script above, a later diff review reported a defect in the check itself: the comment-stripping logic treated the double slash inside a URL in a string literal as the start of a comment and silently skipped the rest of that line. The check appeared to work and quietly missed lines written a particular way, and the miss never appeared in its output. Whoever writes a check is poorly placed to imagine what it cannot catch, because the tests and the check script both come from the same person, and no check gets written where the blind spot is. So check scripts are review targets too. A green CI run is not evidence that a check worked as intended. Every time a check is added, a second pair of eyes has to establish what it fails to catch.
Assign work by permissions, unit of work, and acceptance criteria
Tool names are a poor way to define roles over time: tools change, and the same tool can do different things under different permissions. We therefore define an assignment by three axes.
| Axis | What to define | What happens if it is undefined |
|---|---|---|
| Permissions | Whether the role is read-only, may change files, may commit and push, or may release to production. | A role intended only to inspect can make changes, and ownership becomes untraceable. |
| Unit of work | The task, branch, or diff that bounds the assignment. | Several roles edit the same surface and create integration debt. |
| Acceptance criteria | Machine-check results, the required evidence, and human approval. | A completion report proves nothing concrete. |
| Role | Permissions | Unit of work | Acceptance criteria |
|---|---|---|---|
| Design and requirements | Read and draft documents; no repository changes. | One contract per task. | Scope, non-goals, and completion conditions are explicit. |
| Implementation | Change repository files and commit and push to a branch. | One task and one branch. | Lint, types, tests, and build pass. |
| Independent review | Read-only; no automatic fixes, commits, or pushes. | One diff-limited bundle. | Findings stay within the supplied diff. |
| Machine checks | Return a decision and change nothing. | One diff. | The result is deterministic. |
| Release decision | Authority to execute a production release. | One release. | A human has approved the subject, impact, and release time. |
The important boundary is that the independent reviewer cannot modify the subject it reviews. An implementer still needs to investigate the cause, make the change, and test their own work. That self-check must not be counted as the independent review, because a mistaken implementation premise can otherwise survive under the same premise.
This is Netsujo’s current operating model, not a universal allocation for every development team. The appropriate boundaries vary with organisation size, data sensitivity, and release frequency.
The review requirement survives runtime changes
The reasonable order is: define the purpose and risk; write a Task Contract; assign an eligible implementation runtime; inspect the repository before mutation; work within one bounded assignment; run deterministic checks; fix the candidate; bind the candidate to an exact HEAD; run one independent review on that subject; and let the Controller or human integrate the evidence before release.
Risk classification is not used to delete the review requirement. It is used to select reasoning strength, scope, and any additional security review.
We stopped using risk classification to decide whether to review
Previously, low-risk changes skipped the external review and medium-risk changes were reviewed only when a condition matched. The idea was to allocate reasoning budget in proportion to risk.
That arrangement had a hole: when the classification itself is wrong, nothing is detected. A diff classified as low was merged without review, and defects that presented customers with a state contrary to fact were found afterwards. Improving classification accuracy is one response, but a classification can still be wrong, and a design whose detection drops to zero when it is wrong cannot be repaired by better classification.
So we went back to putting every diff that passes the machine checks through one external review, regardless of its content or its size. This also avoids the inversion in which larger diffs receive less review.
Today there are only three cases in which no review is called.
- There are zero lines of reviewable text in the diff, for example an image-only replacement.
- The machine checks failed, or were not run. Fix that first; not run is not the same as passed.
- The automatic review attempts for that diff are used up, after which a human decides.
Risk classification selects the reasoning effort
| Risk | Example change | What the classification is used for |
|---|---|---|
| High | Authentication, authorization, billing, database migrations, customer data, public APIs, diagnosis and scoring logic, production environment and deploy settings. | Review at high reasoning effort, and confirm design questions before implementation. |
| Not high | Everything else, including wording, articles, CSS, and documentation. | Review at medium reasoning effort, once, after the machine checks pass. |
A file that matches no rule is treated as the higher category rather than dropped into the lower one. Not being able to judge something is not the same as it being safe.
This allocates the strength of the reasoning budget by risk, not the number of AIs and not the size of the change. Whether a review happens at all is not part of that allocation. A very large change is still not handed over in one piece; split the pull request or the unit of change instead.
From a detailed order to a verifiable contract
| Field | What it records |
|---|---|
| goal | What is being achieved. |
| user_value | Who benefits and how. |
| scope | What changes this time. |
| non_goals | What does not change this time. |
| invariants | Conditions that must never break. |
| acceptance_criteria | Conditions that decide completion. |
| risk_level | Low, medium, or high. |
| test_plan | The checks that will run. |
| evidence | The evidence required in the completion report. |
| stop_conditions | When to stop auto-fixing and return to a human. |
Non-goals, invariants, and stop conditions matter most. If adding a goal field must not replace the existing analysis screen, write that down. If three rounds of fixes produce the same failure, hand it back to a human instead of continuing. Weakening tests, type settings, or security settings in order to obtain a PASS is prohibited. Raising an AI’s self-correction ability depends more on clear pass and stop conditions than on more iterations.
From full re-scan to exact-HEAD diff review
The independent reviewer receives a bounded subject: the Task Contract, base and head SHAs, the target diff, the minimum related code, a summary of machine checks, the risks to focus on, and the instruction to return actionable findings only. The reviewer runtime may change; the exact subject and review contract do not.
- No unbounded re-exploration of the whole repository.
- No long successful logs.
- No aimless full text of lockfiles or generated output.
- No refactoring proposals outside the diff.
- No duplicate review of the same head SHA.
- No automatic commit, push, merge, or deploy by the reviewer.
The response format is specified too: per finding, the severity, the file and location, the contract that was broken, the impact on users or operations, the basis, the minimal fix, and the missing regression test. Where nothing applies, the reviewer writes “no applicable findings” and is explicitly not allowed to write “this is safe”. Those are not the same statement. Nothing was found within the scope reviewed, the evidence supplied, and the criteria used. Release still needs machine checks, authority design, operational confirmation, and human approval. AI review is one layer of quality assurance, not a replacement for tests, CI, branch protection, or human approval.
Do not confuse internal judges with independent review
An implementation runtime can contain useful internal roles: a Builder, deterministic checks, a self-check, and a manager for stop conditions. Those checks speed the implementation loop, but they are not promoted into independent-review evidence when they share the implementation assignment. Independent review uses a separate assignment and exact subject. What changes dynamically is the eligible runtime that performs the role.
Measure whether it actually works
A feeling that several AIs raise quality does not justify the operating model. We need lead time, first-pass success, fix rounds, reviewer runtime, reasoning strength, input and output volume, duplicate reviews of the same SHA, serious findings, findings converted into regression checks, incorrect findings, post-release defects, rollbacks, and human integration time. Compare those over fixed windows. If quality is unchanged and inference cost rises, narrow the review input or lower reasoning strength where appropriate; do not silently remove the review requirement. OpenAI’s January 2026 Datadog case study remains a useful external example of a review layer catching system-level interactions, but it is a single company case study rather than a general detection rate.
More AI does not distribute human responsibility
Lining up ChatGPT, Claude Code, and Codex does not produce a team of three engineers. Each is a probabilistic system that carries no responsibility. Given a wrong goal, it will refine the wrong goal. Given weak acceptance criteria, it will call a presentable but unfinished state a PASS. With nobody to stop the release, it will keep producing improvement candidates. What stays with the human is deciding what to build, what not to change, which risks to accept, whether a finding is a specification change or a defect, where to stop fixing, and who answers for it after release. A development lead in the agent era writes less code and designs more roles, authority, evaluation criteria, and stop conditions.
Conclusion
- Structure the instruction document as a Task Contract.
- Separate Role from Runtime and assign only an eligible execution surface.
- Run deterministic checks before LLM review.
- Keep independent review separate from the implementation assignment and bind it to the exact HEAD.
- Allocate scarce runtime capacity explicitly and keep integration responsibility with the Controller or human.
Products and models will change. What remains is the operating design: which role receives what, what it is not given, on what basis something passes, and where it stops.
Frequently asked questions
- Are ChatGPT, Claude Code, and Codex still fixed to one role each?
- No. The roles stay distinct, but the runtime is selected from eligible execution surfaces. ChatGPT is often useful for business and requirements work; Claude Code and Codex can be implementation or reviewer candidates when capability, authority, quota, and live-capacity policy allows it.
- Should every change use the same reviewer runtime?
- No. Reviewable text diffs keep the independent-review requirement, while risk and complexity select reasoning strength and the eligible reviewer runtime. If the preferred runtime is unavailable, use another eligible reviewer or block; do not turn unavailability into a review waiver.
- Does using different AIs make the review independent?
- Different products do not guarantee independence. It takes a separate reviewer assignment, an exact reviewed HEAD, explicit evaluation criteria, bounded authority, findings with a stated basis, and a defined stop condition. An implementer self-check is still a self-check even when the same product can act as a reviewer elsewhere.
- Should a human still read the code?
- The need to read every line by hand goes down, but purpose, risk, acceptance criteria, accepting or rejecting findings, and the release decision stay with the human. For high-risk changes, a human also reads the important diffs and test results.
- Does letting AIs auto-fix each other raise quality?
- Increasing iterations alone raises cost, wrong fixes, and infinite loops. Define the maximum number of fix rounds, the stop rule for a repeating failure, the budget ceiling, and the escalation condition first.
AI agent development and operations incident logThe series hub lists every incident in this record and the order to read them in.
A repeated instruction is a system bugThe separate article covers moving repeated human attention into a Controller, Evidence Ledger, and Gates.
The same work item, different chat namesThe separate article covers how work is identified when it moves between tools.
We design roles, authority, evaluation criteria, and stop conditions alongside the implementation itself.
Talk to Netsujo about AI implementationContinue the series
Next: #01 Give AI review a budget and an exit condition
Back to all 12 episodes of the AI Agent Development Incident Log