Openrig operator
Skill mvschwarz/openrig/packages/daemon/assets/plugins/openrig-core/skills/openrig-operator
Use when debugging host-side OpenRig runtime issues: daemon reachability, Codex permission or writable-root failures, command approval friction, rate limits/account switches, helper cleanup, or topology health confusion. NOT for ordinary CLI operation (use openrig-user) or for changing OpenRig itself (work in the openrig product repo).From its SKILL.md
npx -y skills add mvschwarz/openrig --skill openrig-operatorAssembled from the repository path, not quoted from the project. Check it against their README if it does not work.
SKILL.md
10.9 KB, ~2.3k tokens by cl100k_base, as published. Nobody here has run it
OpenRig Operator
Overview
This skill covers host/runtime/operator triage around OpenRig itself. Use it when the problem may be the daemon, the shell/runtime surface, or stale helper processes rather than the product workflow you are trying to run.
CANONICAL SURFACE NOTE (2026-05-11) — for durable queue routing while doing operator triage, use
rig queue(daemon-backed SQLite).rigx queueis recovery-only fallback; qitems written there are invisible to daemon-backed reads and should not be used for new substantive work.
When to Use
Use this skill when you see:
rig whoami --jsonreturning partial identityrig ps --nodes --jsonfailing while some otherrigcommands still workSent to ...plusVerified: no- repeated unified-exec-process warnings
- suspicion that stale helper processes are accumulating
- Codex seats hit
Operation not permitted, command approval friction, or stale writable roots - Codex seats report usage-limit/rate-limit or need a ChatGPT account switch
Do not use this skill for normal product workflow routing, queue handling, or ordinary peer communication. Use openrig-user for that.
Do not use this skill to decide how to change OpenRig behavior. Changes to the OpenRig product (CLI, daemon, UI, specs) happen in the openrig product repo — operate that as a normal software project, not via this troubleshooting skill.
First Checks
Start with the minimum truthful operator read:
rig whoami --json
rig daemon status
rig ps --nodes --json
Interpret them together, not in isolation:
- partial
whoamican mean identity is inferable while daemon-backed surfaces are degraded daemon statustells you whether the host daemon is up, not whether every seat can reach it cleanlyps --nodes --jsonis the best machine-readable topology check when it works
Verification Drift Vs Send Failure
For rig send:
Sent to ...+Verified: yes= strong positive delivery evidenceSent to ...+Verified: no= ambiguous delivery, not automatic failure- no
Sent to ...line or a hard error = send failure
When verification is ambiguous, check:
- direct reply
rig capture <session>- transcript evidence
- queue/outbox state if the message asked for a durable handoff
Do not blindly retry until you have checked one of those.
Unified Exec Warning
If you see:
Warning: The maximum number of unified exec processes you can keep open is 60 ...
treat it first as a host/tooling-layer warning, not as automatic proof that the OpenRig topology is unhealthy.
This warning can coexist with a healthy live topology.
Safe Process Triage
Inspect the process surface first:
ps -axo pid,ppid,command | rg 'tmux send-keys|rig queue create|tmux attach|codex|claude'
Think in layers:
- host/tooling layer: stale one-shot wrappers, session bookkeeping, helper shells
- topology layer: live
tmux attachseats, livecodex/clauderuntimes, daemon health
Do not diagnose topology failure from tooling-layer warnings alone.
Safe Cleanup Boundary
Usually safe to reap when clearly orphaned / one-shot:
tmux send-keys ...- short-lived shell wrappers created only to enqueue or send one message
Do not mass-kill:
tmux attach ...codex ...claude ...- other long-lived daemon/runtime processes
The point is to remove garbage, not workers.
Terminal Node Escalation
Use a terminal node when evidence shows a seat-level sandbox/profile cannot perform required host work, but another approved operator surface can. This is an explicit operator lane, not a silent permission bypass.
Good fits:
- Codex/Claude seat cannot access Tart, SSH, tmux, queue directories, or host files needed for a proof
- a VM/test-proof or current-host-operator task needs real host capability
- another seat or terminal/sysadmin node has approved access and can return command evidence
Do not use this lane to bypass repo safety, dirty-worktree boundaries, review gates, or destructive-action approval. If the command would stop live rigs, delete data, reset git state, or mutate VM state beyond the accepted plan, require explicit authorization and record it.
Protocol:
- Classify the original failure as seat/profile-specific if possible, not product failure.
- Write or cite a packet with objective, lane, exact commands, cwd, expected outputs, stop conditions, and artifact path.
- Route to an existing sysadmin/infrastructure/terminal node when available; otherwise provision a named terminal node through the topology/config layer instead of using a hidden ad hoc shell.
- Terminal node runs only the packeted commands and returns command log, exit codes, cwd/env notes, and artifacts.
- Original agent keeps task ownership and proof classification; terminal node supplies host capability.
Field lesson: if tester sees Operation not permitted for Tart/SSH while driver or another approved seat can reach the VM, treat it as permission/profile variance and route a terminal-node/operator remediation before changing product code.
Codex Permission Policy
Use the security-and-consequence-boundary-policy skill for the security model.
OpenRig gates consequence boundaries, not ordinary work inside the intended boundary.
Treat Codex permissions as two independent layers:
- command approval rules decide which shell commands can run outside the sandbox
- filesystem writable roots decide which paths a
workspace-writeseat can mutate
Codex auto_review, --full-auto, and approval-bypass modes are not normal fleet defaults. They
can burn quota catastrophically and do not widen filesystem roots for an already-running seat. Full
access scope with approvals_reviewer = "user" is the intended Codex default here.
Claude Code auto permissions are different and are the preferred Claude default on this host when the consequence-boundary rules still apply.
Current host policy:
- default profile is
fleet fleetusessandbox_mode = "danger-full-access",approval_policy = "on-request",approvals_reviewer = "user"- top-level
[sandbox_workspace_write].writable_rootsintentionally includes~/codeand tool dotdirs:.codex,.claude,.openrig,.agents,.config,.cache,.local,.docker,.npm,.nvm,.pnpm-store permissions.fleet.filesystemmirrors those writes and denies.ssh,.gnupg, and project env filesdefault.rulesbroadly allowsrigandrigx;rigxis fully permitted for now because it is the fast-moving config-layer overlaydefault.rulesstill prompts for destructive/publishing surfaces such asrm,mv,chmod,git push, PR mutation,sudo, process kills, daemon lifecycle, and destructive Docker/Brew/rig commands
Codex Account / Usage-Limit Refresh
When Codex seats hit usage limits or need a ChatGPT/OpenAI account rotation, use the focused
codex-seat-auth-refresh skill instead of improvising.
Key reminder: host auth can switch while already-running Codex TUIs keep the old account. Refresh only the scoped seats, preserve stable seat names, and record auth-seat registry rows sequentially.
Known failure modes this policy prevents:
- Codex seats assuming a task is impossible when the real issue is stale launch roots
- agents creating workaround slices/features for what is actually a sandbox/config problem
- writing load-bearing canon into
state/,/tmp, or a nearby writable folder because the intended durable path was blocked - trusting labels like
full access; verify actual roots and command policy instead - changing
config.tomland expecting already-running seats to pick it up without restart
When Operation not permitted appears:
- Identify whether it is command approval, filesystem root, macOS privacy, or stale session.
- Check effective roots with
codex -p fleet debug prompt-input <probe-name>. - Verify with a tiny direct write probe in an allowed target and a negative probe in a protected target such as
.ssh. - If
config.tomlchanged, restart one seat and re-run the probes before fleet rollout. - If the target should be durable and is not writable, stop and escalate; do not invent fallback storage.
Field note: as of Codex CLI 0.125.0, top-level [sandbox_workspace_write] was the shape reflected by debug prompt-input; profile-scoped writable roots did not show in the effective prompt. Re-test this if Codex changes.
Durable Write Escalation
If the task requires writing load-bearing knowledge or behavior and the intended target is not writable, stop and escalate. Do not silently write to state/, /tmp, or a nearby writable folder.
Use the intended durable home for the kind of content:
skills/(e.g.~/.claude/skills/,~/.agents/skills/, or your team's shared skill folder) for operating rules and refocus behavior- Your team's workstream notebook (field notes, lab experiments, mission packets, etc.) for durable observations and PM canon
- The product repo for shipped OpenRig daemon/CLI/config/spec/test behavior
A runtime mirror under a rig state/ path may be used only as a temporary live patch, must be labeled as non-canonical, and must have a canonical sync follow-up.
Common Mistakes
- treating
Verified: noas if it proves the message did not land - treating the unified-exec warning as if it proves the rig is overloaded
- killing live seats when only stale helper wrappers needed cleanup
- concluding "daemon down" from one seat's failure without checking host-level daemon status
- assuming Codex config changes apply to already-running seats without a restart/probe
- confusing command approval with filesystem write permission
- calling a VM or product path broken before comparing another approved seat or terminal-node probe
- assuming a label like
full accessproves effective write/network capability; verify command approval and filesystem writable-root coverage separately
Practical Rule
Clean the smallest safe surface that matches the evidence.
If the warning or failure remains after stale-wrapper cleanup, re-check:
rig daemon status
rig ps --nodes --json
If those remain healthy, the residual issue may still be in the host/tool/session layer rather than in OpenRig topology state.