Writing / ai engineering
One Harness, Any Model: Why I Use Oh My Pi
omp is the open-source terminal harness I run my coding agents in. Here is what it does when a model goes down mid-task, how subagents, the advisor, and memory work, and the small config I actually use.
You know this moment.
The agent is forty minutes into a refactor. It has read half the repository, it has a plan, it has started editing. Then the provider returns a 429, or a 529, or just stops answering.
The tool prints a retry countdown. Then another one. Then a longer one.
You now have a choice. Wait and hope the provider recovers before the context goes stale. Or kill the session, open a different tool with a different model, and explain the whole task again from zero.
Neither option is engineering. Both are you paying for a vendor’s bad afternoon with your own context.
That moment is the reason I moved my terminal workflow onto Oh My Pi, which installs as the omp command. The harness stays. The model is a setting.
What omp is, in one paragraph
omp is an open-source coding-agent harness for the terminal. It is not a model and it is not tied to one vendor. You point it at Anthropic, OpenAI, Google, OpenRouter, a local Ollama, or whatever else you have credentials for, and you get the same tools, the same session format, and the same YAML config regardless of which model is answering.
It reads the instruction files you already have (AGENTS.md, CLAUDE.md, GEMINI.md, .github/copilot-instructions.md) and it can import an existing Claude Code or Codex session with --from-claude or --from-codex. So switching did not mean throwing away the workflow from my previous article. It meant the loop got one layer underneath it.
you
-> omp (tools, session, approvals, config)
-> model role: default | smol | slow | plan | advisor
-> provider A ... falls back to provider B
The rest of this article is the four parts that changed how I work, plus the honest version of my own config.
1. Models are roles, and roles have fallback chains
In omp you do not configure “the model”. You configure roles:
modelRoles:
default: anthropic/claude-sonnet-4-5
smol: openai/gpt-4.1-mini
slow: anthropic/claude-opus-4-5:high
plan: anthropic/claude-opus-4-5
advisor: anthropic/claude-sonnet-4-5:medium
default does the normal work. smol is the cheap, fast one for mechanical tasks and background jobs. slow is the expensive one for the hard problem. The :high suffix is the thinking level. Ctrl+P cycles through the roles in cycleOrder mid-session, and --model @slow starts a session on one.
That alone is useful. The part that fixes the opening of this article is retry.fallbackChains:
retry:
enabled: true
modelFallback: true
fallbackRevertPolicy: cooldown-expiry
fallbackChains:
default:
- anthropic/claude-opus-4-5
- openai/gpt-5.5
- google/gemini-3-pro
smol:
- openai/gpt-5.5-mini
- anthropic/claude-haiku-4-5
google-antigravity/*:
- google/*
- google-vertex/*
When the active model keeps failing with rate limits or outages, omp switches to the next selector in the chain that owns the failing model, skips anything still cooling down, and finishes the turn on the fallback. With cooldown-expiry it comes back to the primary once the suppression window ends. The provider/* form keeps the model id and swaps the provider, which is the right shape when the same model is served through two endpoints.
The context does not go stale. You do not re-explain the task. The retry countdown becomes a one-line notice that the model changed.
The trade-off
A fallback model is a different model. If your primary is a top-tier reasoning model and the fallback is a mid-tier one, the quality of the next few tool calls drops, and you should know that when you review the diff. I keep the chain short and put the closest-quality model first for exactly that reason.
Two things to know before you copy that block:
- Roles without their own chain inherit
default. That is convenient until your cheapsmolrole falls back to your most expensive model on a bad day. Givesmolits own chain. - Arrays in omp config replace, they do not append. A project-level
.omp/config.ymlthat setsfallbackChains.defaultbecomes the entire chain for that project. Same fordisabledProvidersandenabledModels.
2. Subagents and vibe mode
If you read the Projectinator article, you know I like the idea of one model coordinating and other models doing the work. omp has that built in, in two shapes.
task: fan out, then inspect
The task tool spawns subagents. You give it shared context once and a batch of independent jobs, and each job runs as its own session with its own model:
scout read-only research, fast model
reviewer code review
security-reviewer read-only vulnerability pass
task general implementation
sonic low-reasoning mechanical edits
You can also define your own agents in .omp/agents/, and per-agent model overrides live in settings, so a scout can run on smol while a reviewer runs on slow.
The part I actually care about is what happens after. Every subagent leaves two things behind:
agent://<id>is its final output.history://<id>is its transcript: every tool call, every file it read, every command it ran.
When a subagent claims “done” and the result looks wrong, I read history://<id> before I read the diff. That is where you find the moment it stopped reading the repository and started guessing.
Finished subagents do not vanish either. They go idle, then park after a timeout, and you can message a parked one through the hub tool to ask a follow-up. It still has its context. That is usually cheaper and better than spawning a fresh one and re-explaining.
Two settings worth knowing on day one:
task:
maxConcurrency: 4 # bounds parallel subagents
isolation:
enabled: true # each subagent works in its own worktree, returns a patch
Isolation is the answer to two subagents editing the same file at the same time. Each one works on a clone and hands back a patch or a branch; you merge deliberately instead of hoping.
/vibe: the director pattern
/vibe flips the interactive session into a director. The director loses edit, bash, and every other mutating tool. It keeps read, todo, and five worker-control tools. Workers are real, persistent subagents in two tiers:
| Tier | Agent | Role | Use it for |
|---|---|---|---|
fast |
sonic |
@smol |
mechanical execution, drafts, volume |
good |
task |
@task |
design decisions, judgment, reviewing fast |
Because the director cannot edit, it has to verify worker claims by reading the touched files. That constraint is the whole point. It is the same separation I built by hand in Projectinator, without the orchestration code.
Vibe mode is not for small tasks. It earns its overhead when a change touches several independent areas, and the worker transcripts stay inspectable afterwards.
3. The advisor: a second model that reviews every turn
This is the feature that sounds like overhead until you see what it catches.
The advisor is a second model attached to the session. After each primary turn it receives the transcript delta, including the primary’s reasoning and tool calls, and it can inspect the workspace with read, grep, and glob. If it has something to say, it uses one tool, advise, with a severity:
| Severity | What happens |
|---|---|
nit |
batched into the transcript at the next step boundary, no interruption |
concern |
steers the live turn: material risk, wrong direction, hallucinated API |
blocker |
interrupts even a finished answer: continuing would waste work |
Enabling it is two lines plus a role:
modelRoles:
advisor: anthropic/claude-sonnet-4-5:medium
advisor:
enabled: true
Or /advisor on for one session, or omp -p --advisor "..." for a headless run.
What makes it usable rather than annoying is the throttling. One accepted note per update. Exact-text duplicates are dropped. After an interrupting concern or blocker, further ones are downgraded to asides for the next three turns (advisor.immuneTurns). Content-free notes like “looks good” never reach you.
The trade-off is cost. Every turn now has a second model reading it. I turn the advisor on for work where a wrong direction is expensive, migrations, auth changes, anything touching data, and leave it off for a blog post edit.
For a team, WATCHDOG.yml lets you define a roster of named advisors with their own models and tool grants (/advisor configure opens the editor). An “Architecture” advisor and a “Security” advisor with different prompts, both watching the same session, is a very cheap version of a review board.
4. Memory, skills, and the files that travel with the repo
The last article ended with “save what you learned in AGENTS.md”. omp gives that idea three layers.
Context files
omp discovers instruction files before the session starts. AGENTS.md and CLAUDE.md walk up from the current directory to the repository root; .omp/AGENTS.md is the native format. You never tell the agent to go read them.
The one I did not have before is .omp/RULES.md. It is a sticky rule: short, hard requirements that get re-attached near the current turn so they survive a long conversation. AGENTS.md is background. RULES.md is “never commit, never push, run the build before you say done”.
Skills
A skill is a directory with a SKILL.md. The agent sees only the name and description in its prompt, and reads the full content on demand through skill://<name>. That keeps the system prompt small while still giving the agent a playbook for “how we deploy this service” or “how to add an app page to this site”.
Skills are discovered one level deep under skills/. A nested team/internal/SKILL.md is silently ignored, which is the first thing to check when a skill does not show up.
Memory
Memory is off by default. The local backend is the one I would start with:
memory:
backend: local
autolearn:
enabled: true
With local, a background pipeline reads past sessions for the project and consolidates them into a MEMORY.md and a compact summary that is injected at session start. autolearn exposes a learn tool so the agent can record an explicit lesson to learned.md when something is worth keeping. /memory view shows exactly what is being injected, which matters, because memory is heuristic context, not truth about the current repo state.
My rule: repository files win over memory, memory wins over the model’s guess. When the three disagree, the memory is stale.
My actual config
Here is the honest part. My global ~/.omp/agent/config.yml is nine lines:
modelRoles:
default: anthropic/claude-fable-5-1
symbolPreset: unicode
composer:
shape: band
theme:
dark: titanium
setupVersion: 2
defaultThinkingLevel: auto
That is it. One model role, a theme, and automatic thinking-level selection. Everything else in this article is a default or a per-session toggle.
I am not showing you that to be modest. I am showing it because the value of a harness is mostly in the defaults being right, and in the fact that every one of them is a plain YAML key I can read with omp config list and change with omp config set.
What I would add next, in order:
- A
fallbackChains.defaultentry, because the opening of this article is not hypothetical. smolandslowroles, so subagents and scouts stop running on the expensive model.tools.approvalMode: writewith a couple ofbash.patterns, to get the human boundary from my last article into config instead of into every prompt.memory.backend: localon the repositories I return to most often.
What I give up
omp is a large tool. omp config list is long, the docs are long, and there is more surface than any one person uses. If you want a two-flag CLI, this is not that.
It is also not a vendor product. There is no support contract behind it. When something breaks, you read the source or the docs, and you should be comfortable with that before you put it in front of a team.
And a second model watching a first model is still two models. The advisor catches wrong directions. It does not catch wrong requirements. That is still my job.
The simple version
- The harness is the constant. The model is a setting.
- Give every role a fallback chain before you need one.
- Fan out with subagents, then read
history://before you trust the result. - Turn the advisor on when a wrong direction costs more than the tokens.
- Put hard rules in
RULES.md, background inAGENTS.md, playbooks in skills, and let memory be heuristic.
The repository is on GitHub. Start with omp config list, change one thing, and see what the next bad provider afternoon feels like.
