Palahepitiya Gamage Amila · Weekly Learnings
Two harnesses, seven agents, one feature: what the brief changed
· Palahepitiya Gamage Amila

This post is for founders and engineering leads who are trying agents on real delivery work. It covers how I run one workflow across two agent tools, what one real feature cost, and the lesson that mattered most: how I wrote the brief decided the outcome more than which tool or model I picked.
Early on 24 September I typed "let's start phase 1" into my orchestrator agent. Phase 1 was a new frame and menu for a client's legacy web app. In less than a day the plan was written, challenged, rewritten, built, fact-checked, checked against the design and code reviewed. I approved it the next day and it was pushed. It cost $19.19 against a $60 cap.
No engineering team worked on it. It was me and seven agents.
Why I use two harnesses
Some context. I work as a fractional CTO, and for the last few weeks I've used agents for real delivery on a client project. I kept hitting the same question: which tool?
Claude models work best in Claude Code. The other models I wanted to use run in Pi, a different agent harness, through a model gateway. I didn't want to pick one. So I built one workflow that runs across both, inside herdr, a terminal multiplexer, with each agent in its own pane.
The setup is simple to describe. One registry file says which model each role uses. When a stage starts, a script reads that file. If the model is a Claude model, the pane opens in Claude Code. Anything else opens in Pi. The roles, rules and handoff files are the same either way, so an agent doesn't need to know which harness the one before it ran in.
On this project, Claude models ran the orchestrator, planner and code reviewer. Pi ran the challenger, the executor, verification and UI QA on other model families. I picked those by testing, not by reputation. My first challenger came back clean 6 times out of 10 through the gateway. Qwen came back clean 21 out of 21, so it got the job.
I wrote up the roles, handoffs and approval gates in an earlier post, so I won't repeat them here. The short version: each agent has one narrow job, agents hand off through small checked files, and nothing runs or ships until I approve it.
What the second harness bought me
Mixing model families paid off in what each one caught in the others. Qwen blocked a Claude plan that used a git command the machine's git version couldn't run. Verification caught the executor reporting a diff size it had measured before its own last edit. The Claude code reviewer refused to carry on when my orchestrator tried to route past a failed verification.
Same-family review was softer. After five rounds of challenger findings on one plan, I asked the orchestrator whether the planner was the problem. Most of its answer was fair, and the fix it suggested was good. But it was a Claude model defending a Claude model's work. So now the reviewer comes from a different family from the author, and I stay at the gate.
The second harness also changed the bill. The Claude models were the expensive part. The other models did a lot of the heavy work for very little. On one plan, the planner cost $13.13 by the time the executor handed off. The executor, which did the actual work, cost $0.23. The challenger cost $0.19. Planning is the cost I watch now.
The brief mattered more than the tool
What surprised me most was how little the choice of harness mattered. The agents work hard on whatever you give them. When my brief was vague, a lot of that effort went to the wrong places.
My first big plan on this project started from one sentence: review the progress of the new UI and "get a plan to complete it fully". Before planning started, the orchestrator flagged four open questions, including what "fully" meant. There was no done-line, so the plan had to invent one as it went.
That plan went through seven versions in four days. The first was 35 KB. The last was 228 KB. By the fifth review round, 12 of the challenger's 19 findings were the same kind of fault: a fact written in several places in the plan and updated in only one. After six rounds of planning and review, the test suite was still at 16 passing out of 38. All the actual build work in that stretch cost less than a dollar. The money and time went on a document that kept growing.
So I wrote the next brief differently. It said, in bold, "Phase it. Do not plan the whole UI here." It listed what was in scope and what was out. It set a $60 cap. The plan opened with one plain line: done means 39 of 43 tests pass, the same 4 old failures and no others, and zero build errors.
That plan stayed between 16 and 21 KB across five versions. It went from brief to final code review in less than a day. When a real problem came up (the menu icons needed a script I hadn't approved), the icons moved to their own later phase. The plan didn't grow to absorb them.
I don't want to oversell this. It's two plans on one project. The second had a narrower job, better tools (a plan checker that came out of the first plan's mess) and a pinned copy of the live design to work from. But it matches the rest of my notes. Long plans drift. A small piece with a stated done-line gets to real code sooner, and real code is where the real problems show up. Three plan reviews missed that the design's icons would break every page in the app's frame. Loading real pages found it.
Where I stayed in the loop
The night phase 1 started, I went to bed and left the orchestrator running it. I gave it one narrow exception. It could approve the plan gate for me, but only if the challenger's latest verdict had no hard block, nothing needed a change to the client's source code and spend stayed under the cap. On anything else it had to stop. It could never push, and the final approval stayed with me.
It read those limits literally. The icons needed a script, so it held them back. When the executor landed one test short of the done-line, it didn't move the line. When UI QA flagged a menu item with colour contrast of 4.3 to 1 against a 4.5 to 1 standard, everything stopped and waited for me. The night cost $8.54, and I made the calls in the morning.
That contrast issue is worth a note. The design had the same fault, so checking the build against the design alone would never have caught it.
The limits I still have
Some gaps are still open, and I'd rather name them:
- There's no cap yet on how many times a plan can bounce between planner and challenger.
- The loop prints "needs me" instead of stopping on it.
- Review roles are told not to write, but that's a rule in a text file. They still have write access. One night a challenger's "harmless" syntax check changed two database permissions. It told me straight away and it was a two-minute fix, but removing that access is still on my list.
What to take from this
If you're trying agents on real work, spend your effort on the brief before you spend it on the tooling. Cut the work into a phase. Say what's in and what's out. Set a cost cap. Write one line that says what done means, in terms a test can check. Then let the agents work, and keep the approvals with a person.
Running two harnesses works fine once the roles and handoffs don't depend on either one. Pick models per role by testing them on your own work, and have a different model family review what another one wrote.
If you're running agents on delivery work, I'd like to compare notes. How are you splitting work between tools, and how small do you cut a piece of work before you hand it to an agent? If you'd rather talk it through, you can book a discovery call or email hello@pgamila.com.
Questions this note answers
Why use two harnesses?
Claude models work best in Claude Code. The other models I wanted to use run in Pi, a different agent harness, through a model gateway. One registry file says which model each role uses. If the model is a Claude model, the pane opens in Claude Code. Anything else opens in Pi. The roles, rules and handoff files are the same either way, so an agent doesn't need to know which harness the one before it ran in.
What did phase 1 cost?
Phase 1 was a new frame and menu for a client's legacy web app. In less than a day the plan was written, challenged, rewritten, built, fact-checked, checked against the design and code reviewed. I approved it the next day and it was pushed. It cost $19.19 against a $60 cap.
What changed when the brief was rewritten?
The first plan started from one sentence: review the progress of the new UI and get a plan to complete it fully. It went through seven versions in four days, from 35 KB to 228 KB. After six rounds of planning and review, the test suite was still at 16 passing out of 38. The next brief said to phase the work, listed what was in and out, and set a $60 cap. Done meant 39 of 43 tests pass, the same 4 old failures and no others, and zero build errors. That plan stayed between 16 and 21 KB and went from brief to final code review in less than a day.