Skip to content
Amila

Palahepitiya Gamage Amila · Weekly Learnings

The operating model behind a multi-agent engineering team

· Palahepitiya Gamage Amila

Pipeline diagram: Route, Plan, Challenge, Human gate, Execute, UI QA, Review, with Verification on demand and a cost meter.

I spent the last stretch of September running a multi-agent setup on real engineering work: planning, challenging, executing, and reviewing changes against an approved plan. This was production-shaped work with cost ceilings, handoffs, and people who had to say yes before anything moved.

Some context. I work as an AI-first fractional CTO. Agents and specialists collect facts and write drafts. I decide. Anything that leaves the building waits for approval. That same discipline is what I have been testing on multi-agent delivery itself.

The easy part was adding agents. The hard part was giving each one a narrow job, a clear handoff, a human gate, and a way to see cost and failure. Without that, capable models still drift, duplicate work, and spend without a ceiling.

Treat the work as a pipeline of artifacts

Most multi-agent demos look like a chat room. Everyone talks. Memory is shared and fuzzy. Nobody can point to the contract that was supposed to constrain the next step.

What worked better was a controlled pipeline of durable artifacts:

request → route → plan → challenge → human approval → execute → UI QA → review → final approval

Each arrow is a handoff. Each handoff carries a status, an artifact, any findings, a next owner, and a stop condition. Hidden memory is a weak transport. If the plan changed, you see a new version. If the challenger blocked it, you see findings rather than a quiet rewrite.

That sounds bureaucratic until you watch two agents invent different scopes for the same ticket. Then the paper trail stops being ceremony and starts being the only way to know what was approved.

I also stopped treating “the conversation” as the source of truth. The source of truth is the latest approved plan, plus the findings that forced a re-plan, plus the review that compared the result to the contract. Chat can help. It should not be the ledger.

Roles in plain words

The stack is not “more models.” It is a small set of roles with written boundaries. Model choice sits underneath, in a registry. The harness that runs the code is an implementation detail and can change by job.

Orchestrator. One shallow classification and a route. It does not design the solution. Cheap routing, no long reasoning. If the orchestrator starts planning, you have already lost the cost model.

Planner. Turns the request and repository context into a plan as a contract. Scope, acceptance criteria, assumptions, estimated tokens and time, cost ceiling, explicit stop conditions. The planner never executes its own plan. That separation matters more than people expect. When the same agent writes the plan and then runs it, challenge arrives too late, if it arrives at all.

Challenger. Attacks the plan, not the people. Returns findings only. A critical finding blocks until the plan is revised. The challenger does not patch the plan into an ad hoc fix and wave it through. Findings go back to planning as inputs for a new version.

Human gate. Approves the plan before execution, and later approves the result. Can send work back for a clean re-plan. Gates do not silently expire. Downstream work does not start until the approval is clear. If a human is away, the queue shows blocked work. It does not invent consent.

Executor. Follows the approved steps. Logs transitions. Fails fast. Stops at the cost ceiling rather than improvising. Scope creep is a re-plan request, even when the model sounds confident.

UI QA. Checks what was built against the design, not only against the plan text. Screenshots (or the live UI) are compared to the intended screens so visual drift does not hide behind a green unit check. Findings are concrete: wrong state, missing empty state, layout that does not match the design. Critical UI findings block the same way critical plan findings do.

Reviewer. Compares the result with the plan. Read-only findings. Critical divergence goes back to planning, rather than into a quiet second pass that changes the deal.

Verification (on demand). Not a default stage. When you need facts checked, a side skill returns bounded, checked facts only. It sits beside Execute, not in the main line. It does not make product or delivery decisions. Facts in, facts out.

None of this requires one particular product. It requires that each role has a one-page profile: mission, inputs, outputs, forbidden actions, escalation rule, and a success test you can actually check.

A roughly three-day cost check

I looked at a gateway export from a roughly three-day run: about 1,000 recorded requests, about 126.8 million tokens, across three model families (GLM, DeepSeek, and Qwen). It cost under $40 for that window, roughly.

That is a useful reminder: routing and role boundaries can matter more than always using one expensive model for every step. Put a heavier model on planning or challenge when you need depth. Put a lighter model on high-volume execution when the steps are already approved and narrow. Keep a single model registry so those choices are not buried inside prompts and scripts.

Call the infrastructure what it is for this post: a model gateway. The vendor name is less important than whether you can see spend before the next step runs, and stop when you hit the ceiling. If you cannot answer “what did the last approved plan cost, and what would the next step cost?” you do not have cost visibility. You have a bill after the fact.

How a project actually moves

Intake is classified once. The planner gets the request plus the relevant role profile and skills, and writes a versioned plan. The challenger reviews that plan independently. Findings cause a re-plan. They are not converted into a side fix by the challenger.

After the human gate, the executor works step by step. A watcher or step log makes transitions visible. A spend meter checks the next step before it runs and halts near or over the ceiling. UI QA checks screenshots or the live UI against the design. Then the reviewer compares output to the contract. Critical mismatch loops back to planning. Otherwise the final human gate decides whether the work is done.

Repeated work is a deliberate re-plan. Keeping the orchestrator shallow also keeps orchestration from becoming the dominant spend. Routing should be cheap. Thinking should happen where it is named.

When something fails, the question is usually which boundary broke. Did routing get ambitious? Did the planner skip acceptance criteria? Did the challenger get ignored? Did the executor invent steps past the ceiling? Did review check the result against the plan, or only against “does it look plausible?” Fix the boundary. Do not paper over it with a longer prompt.

What you can copy

You do not need my stack. You need the operating discipline. Start here:

  1. Written boundaries. Route. Plan. Challenge. Execute. UI QA. Review. Write them down. If a role crosses a boundary, stop and redesign the profile.
  2. One-page role profiles. Mission, inputs, outputs, forbidden actions, escalation rule, success test. One page is enough. Three pages usually means you are still negotiating with yourself.
  3. One model registry. One place that maps role to model family. Do not hard-code model names through prompts or one-off scripts.
  4. Plan as a durable artifact. Version history, acceptance criteria, cost ceiling, named human gate. If it is not written, it was not approved.
  5. A small handoff schema. Status, artifact, findings, next owner, stop condition. Enough to debug a stuck run without archaeology.
  6. The smallest useful observability. Step timestamps, plan diffs, a token and spend meter, and a visible queue for blocked approvals.

Pilot on a bounded feature or migration. Measure re-plans, blocked findings, review defects, time to approval, and spend. Lines of code are a weak scoreboard for this. A clean re-plan after a critical finding is success, not delay. An executor that “finished” while ignoring the cost ceiling is failure, even if the diff looks impressive.

What lasts

Model families will change. Harnesses will change. The gateway you use this quarter may not be the one you use next year.

What lasts is narrower jobs, clearer handoffs, human gates that actually gate, and cost you can see before the next step starts. Adding agents without that mostly adds variance.

If you are experimenting with multi-agent delivery, I would be interested in comparing notes: which role boundary, handoff, or cost guardrail made the biggest difference in your own projects?

Questions this note answers

What worked better than a shared multi-agent chat?

A controlled pipeline of durable artifacts: request, route, plan, challenge, human approval, execute, UI QA, review, and final approval. Each handoff carries a status, an artifact, any findings, a next owner, and a stop condition. The source of truth is the latest approved plan, plus the findings that forced a re-plan, plus the review that compared the result to the contract.

Which roles does the pipeline use?

Orchestrator, planner, challenger, human gate, executor, UI QA, and reviewer. Verification is on demand: it sits beside Execute, returns bounded checked facts, and does not make product or delivery decisions. The orchestrator only classifies and routes. The planner writes the plan and does not execute it. The challenger returns findings only. Human gates approve the plan before execution and the result afterwards. They do not silently expire.

What did the roughly three-day run cost?

A gateway export for that window showed about 1,000 recorded requests and about 126.8 million tokens, across GLM, DeepSeek, and Qwen. It cost under $40, roughly. Heavier models can sit on planning or challenge. A lighter model can sit on high-volume execution once the steps are approved and narrow. The check that matters is whether you can see spend before the next step runs, and stop at the ceiling.

How does a project move through the model?

Intake is classified once. The planner writes a versioned plan. The challenger reviews it independently, and findings cause a re-plan rather than a side fix. After the human gate, the executor follows the approved steps, a spend meter checks the next step before it runs, UI QA checks the design, and the reviewer compares the result to the plan. Critical mismatch goes back to planning. The final human gate decides whether the work is done.