Sortie Dogs
Bounded agent harness and validated orchestration loop plugin for OpenCode
0
13,250
3.8k in 7 days
45.3
Multi-signal model
5 hours ago
2026-10-05
Install and configure
opencode.jsonWrites to this project's opencode.json — applies to this repository only.
opencode.json
{
"$schema": "https://opencode.ai/config.json",
"plugin": ["sortie-dogs@0.13.8"]
}Writes to ~/.config/opencode/opencode.json — applies to every project.
~/.config/opencode/opencode.json
{
"$schema": "https://opencode.ai/config.json",
"plugin": ["sortie-dogs@0.13.8"]
}If you want to modify the plugin locally, install it into the project and reference the local path.
shell
pnpm add -D sortie-dogsOpenCode loads npm dependencies through its embedded runtime on startup and caches them locally — no manual global install needed.
A goal-preserving, adaptive execution harness for OpenCode that optimizes cost, time, and proof without taking your setup over.
Use OpenCode normally. Invoke Sortie only when you want scoped investigation, implementation, validation, review, and model routing.
- Goal invariance: accepted outcomes and proof requirements survive delegation, continuation, remediation, and restart.
- Adaptive execution: small work stays small; additional agents and stronger models are used only when task shape or risk justifies them.
- Clear Worker instructions: give Luna a concise goal, explicit constraints and completion criteria; keep orchestration bookkeeping in the harness. See the instruction design principles.
- Coexistence: Sortie activates only when selected and preserves normal OpenCode agents, settings, and user-owned files.
- Cost, time, and proof: the objective is a verified result at the lowest practical cost and wall time, not the largest agent count.

Guides: 日本語 · 简体中文 · Testing · CLI testing
Current release: v0.13.8
(release notes). The default Mission runtime retains the v010
profile, command and configuration names for compatibility; these names do not mean v0.10 is installed.
The current asset marker is 0.13.8-native-binding-v1.
SWE-bench Lite: 170/300 (56.67%)
The fixed Sortie-dogs v0.12.24 harness resolved 170 of 300 SWE-bench Lite test issues in one pass@1 campaign, with 9 empty patches and no official evaluation errors. Every instance has a frozen prediction and an inference-time trajectory. The task Workers ran openai/gpt-6-luna-fast#max; operator, coordinator and review roles ran openai/gpt-6-sol#xhigh. This is a system result, not a Luna-only model comparison or a Verified/full SWE-bench score.
Technical report and per-repository results · Public predictions, logs and trajectories
The single official 300-instance report and frozen predictions are hash-bound in the report. Confirmed inference expense was $162.99; a separate $34.60 of usage has unknown pricing and is held against the campaign cap, not counted as known expense. Leaderboard registration and maintainer acceptance are separate from this official local evaluation.
Historical scores below belong to their fixed candidates, not v0.13.8. SWE-bench is a separate, optional measurement rather than a mandatory release gate.
Beta: v0.13.x is still stabilizing. Runtime behavior, configuration, and generated assets may still change before 1.0.
Quick start
Requirements: Node.js 22.6 or newer, npm, and OpenCode V2.
Run these commands in the target project:
npm install --save-dev sortie-dogs@latest
npx sortie-dogs init .
init defaults to the v010 Mission profile and registers the OpenCode V2 plugin in
.opencode/opencode.json(c), preserving unrelated settings. It sets subagent depth to at least two:
{
"plugins": ["sortie-dogs"],
"experimental": { "subagent_depth": 2 }
}
If an existing local bridge already imports sortie-dogs/server, init reuses it instead of adding
a duplicate package entry. A larger existing subagent depth is retained.
Completely restart OpenCode, then run:
/sortie-v010 <task>
Selecting dog-operator directly starts the same workflow. dog-operator is the
user-facing entry point for the Mission profile. dogs-coordinator and every *-v010 role are
internal children and must not be selected as task entry points.
init installs runtime assets and merges the required OpenCode settings; the package entry or
existing local bridge loads enforcement and model routing. OpenCode can reload watched configuration,
but replacing an installed dependency may require a full restart. A new chat session alone does not
prove the newly installed plugin is loaded.
v0.13.8 runtime updates
PR #161 captures native validation bindings through the readable resolved target of a Windows
directory junction instead of enumerating its failing logical alias. Logical evidence labels,
link identity and complete target bytes remain bound; unchanged readable junctions retain their
evidence identity. External-link/cycle handling, actual source/output freshness and Review remain
unchanged. No new approval, permissions or model changes. The PR's isolated native Windows fixture
reached independent Review PASS and accepted completion after resuming the same Mission, without
rerunning its successful check; this does not claim original product recovery or uninterrupted
single-turn completion. Release receipts: _testenv/releases/0.13.8/;
see release notes.
v0.13.7 runtime updates (retained)
PRs #158/#159 prevent late write-only reports from falsely invalidating bound non-generating
native checks, while preserving actual source/output freshness, failed checks and legacy recipes.
Whole-root grants without a concrete inventory keep full-candidate freshness. Non-Git Mission roots
can prepare independent Review from declared nested repositories/worktrees/plain outputs, including
committed content; corrupt Git/HEAD errors remain visible. Operator performs authorized local recovery
in the same turn without another approval, changed requirements or reset spending. Ubuntu's Git
filesystem-boundary diagnostic is recognized alongside the standard non-repository error.
The PR native observation reached Reviewer startup, not final acceptance. Release receipts:
_testenv/releases/0.13.7/; see release notes.
v0.13.6 runtime updates (retained)
PR #156 references recorded formal PASS evidence in acceptance summaries instead of fetching
the same successful unit's native history again for display. Missing/failed proof retains history
diagnostics; historical evidence is not current freshness or acceptance. Identical external artifact
inventories are shared only inside one snapshot refresh, never across later calls. Operator guidance
proceeds to acceptance when current evidence covers the request and no concrete gap remains.
Existing freshness/review guards, provider/model selection and cache-prefix machinery remain unchanged.
No end-to-end speedup or live token/cache-hit improvement was measured. Release receipts:
_testenv/releases/0.13.6/; see release notes.
v0.13.5 runtime updates (retained)
PR #154 preserves stable model instructions and appends changed host state at native history boundaries, with complete current-state reconstruction after compaction. Genuine Reviewer tools have deterministic read-only-first ordering; the real correction-permission transition still remains. The final candidate restores v3 behavioral review and combined assignment/findings while retaining exact instruction discovery, inherited Reviewer formal-command delivery, literal local shell-file scope reconciliation and explicit repository-root read/write scope support. New whole-project captures avoid bookkeeping-only invalidation; legacy evidence and real source/artifact freshness remain.
The final pre-release cycle 11 Anko sample passed official local score 1 and selected public probes 9/9
in 26m1.866s at estimated $1.41968316. It was faster but 15.762% costlier than the earlier v3 sample,
not combined cost-preserving optimization or a general quality/speed guarantee. Release receipts:
_testenv/releases/0.13.5/. See release notes and
candidate tradeoffs/failures. Native Worker startup is not completion.
v0.13.4 runtime updates (retained)
PR #152 restores same-session Coordinator/Operator implementation and formal validation through
plan_units(executor="self"), start_direct_unit and finish_direct_unit. Known single-unit work can
combine Mission start and planning; Luna Fast/max Worker routing remains the default. Full generated
contracts remain visible when an explicit Read line range covers the file. Native background
responsiveness remains: a launch acknowledgement or idle root is not Mission completion.
The initial independent Reviewer can investigate, correct, formally validate, deliver and self-recheck
continuously in its original native Task. Later Major/Medium findings accumulate in that same correction
context. Author self-recheck remains self-rechecked, independent=false, never independent PASS.
A different Reviewer is conditional on concrete residual Major risk; unresolved Major/Medium findings
still block acceptance. Operator owns final comparison and receipt.
Inherited compiler scratch no longer falsely invalidates broad-scope formal proof. Host-observed Git delivery, caller-setting review and test-composition guidance reduce avoidable detours. Saved host completion cards remain in tool history/UI, while outgoing V2 model/compaction context omits only their presentation body, retaining receipt and evidence identities.
Two runs of the same fixed pre-release v8 package completed the original Anko task with official local
score 1 (F2P 9/9, P2P 94/94) and selected public probes 9/9. Times were 26m49s and 25m55s, costs
$1.47908048 and $1.54610816; each saved more than seven minutes versus the recorded v5 sample.
These limited same-task observations do not establish general speedup or a new SWE-bench score.
Release preflight, full tests, fixed-commit Windows CI and native Worker-start receipts are retained in
_testenv/releases/0.13.4/; startup/model identity is not task completion. See the
release notes and quality-loop evidence.
Mission workflow
Use Operator → Worker when one useful unit and its meaningful formal check are known; use Operator → Coordinator → Worker for actual discovery or decomposition:
dog-operatorstates a few requirements/negative constraints and owns user decisions and final acceptance. The host saves the original user message verbatim.- Hidden
dogs-coordinatorowns investigation, unit declarations, Worker/Scout/Advisor/Reviewer dispatch, in-request write-scope extensions, and corrections. It can read/search and run confirmation shell commands; it can implement and formally validate directly in its own session, or delegate a unit to Worker. dog-worker-v010implements a host-generated unit within its file/directory write scopes. Investigation commands need no pre-registration; formal checks retain real host-recorded results.- High-risk changes require an initial independent Reviewer with read/search access. Low-risk skips are explicit and recorded. Reviewer-owned corrections follow the self-recheck policy above.
- Fast-lane can include high-risk single-unit work; it never implies a review skip.
- Investigation, edits, formal checks and requested Git delivery stay in the same implementing child. Explicit user ordering is retained; no routine plan-approval or commit-only handoff is needed.
- Unit progress appears on the running Task without stopping Coordinator or prompting Operator.
The v010 Mission profile is serial by design; background responsiveness does not add parallel writers.
The stable profile's Luna fabric and parallel integration path are not exposed in this profile. More agents are not a
goal; preserving quality while reducing unnecessary expensive work is.
SWE-bench evaluation
Official SWE-bench Lite dev results on the same 23 public instances:
| Candidate / benchmark adapter | Official resolved | Empty patches | Run conditions | Details |
|---|---|---|---|---|
v0.10.6 candidate (859c396) |
4 / 23 (17.4%) | 5 | Four inference slots | Handoff |
| Frozen v0.10.14 build | 6 / 23 (26.1%) | — | Four slots; $1.50/task | Per-task results |
v0.12.8 (e0f8cef adapter) |
5 / 23 (21.7%) | 9 | Four slots; budget amended across two batches; 30-minute timeout | Campaign |
v0.12.8 (84ccdf1 main adapter) |
6 / 23 (26.1%) | 2 | Fresh 23-task run; four slots; 30-minute timeout | Main-integrated run |
v0.12.15 (0c9690d) |
5 / 23 (21.7%) | 3 | Fresh 23-task ext4 retry; four slots; 40-minute timeout | Campaign |
v0.12.16 (9b05a34 release; b1a6c0e runner) |
4 / 23 (17.4%) | 5 | Fresh 23-task run; eight slots; 40-minute timeout | Campaign |
v0.12.19 (24f5386 release; matched rerun) |
8 / 23 (34.8%) | 0 | Eight slots; effective $2/instance; 40-minute timeout; one inference timeout | Official result and provenance |
v0.12.20 (628eb81 release) |
7 / 23 (30.4%) | 0 | Eight slots; effective $2/instance; 40-minute timeout | Official result and caveat |
v0.12.25 (49eb1e4 release) |
7 / 23 (30.4%) | 0 | Eight slots; $2/instance; $30 total cap; 20-minute progress check / 40-minute hard maximum | Comparison baseline |
v0.13.1 (d19e8be release; 2026-09-30) |
8 / 23 (34.8%) | 0 | Eight slots; $2/instance; $46 total cap; 20-minute progress check / 40-minute hard maximum; GPT-6.1 Sol + Luna Fast | Run summary |
Every row has 23 submitted official predictions; an empty patch counts against
the score, not as a missing evaluation. The v0.10.14 report does not separately
summarize empty patches. The v0.12.15 row is the separately approved fresh
run after an initial /tmp quota failure, not an additional score for that
failed attempt. Inference completion is not official resolution.
Budgets, runtime/adapter versions, and execution conditions changed between
campaigns, so this table is a history of observed results, not a controlled
head-to-head comparison or a general success-rate claim. The v0.10.14 run
estimated $15.75 in model cost and a 15.2-minute median agent runtime.
The v0.12.16 run is the first eight-slot inference score in this table;
five runners timed out, and this does not establish a model-quality regression
against runs with different concurrency and conditions.
The v0.12.19 row is the corrected run with an effective $2 per-instance cap;
an earlier v0.12.19 run scored 7/23 but had no effective per-instance cap and
is not a same-condition comparison. The v0.12.20 run lost the
sqlfluff__sqlfluff-2419 resolution relative to the corrected run; this
single run-to-run difference does not establish causation.
Historical qualification references remain in benchmark reference.
v0.13.1 dev23 (2026-09-30)
One fresh pass@1 run and one official SWE-bench harness evaluation resolved
8/23, versus 7/23 for v0.12.25. The new resolution was
pylint-dev__astroid-1333; all seven previously resolved IDs were retained.
Resolved by repository: marshmallow 2/2, pvlib 0/5, pydicom 2/5,
astroid 3/5, pyvista 0/1, sqlfluff 1/5.
- The scored row is the user-requested fresh run after a host restart. The interrupted initial run is excluded from this score; the fresh run made one attempt per instance with no inference retry.
- The dataset revision (
6ec7bb89b9342f664a54a6e0a6ea6501d3437cc2), public rows, and all 23 official evaluation image IDs match the v0.12.25 run. Both usedofficial-image-testbedand sequential official scoring. - Operator/Coordinator/Reviewer/Advisor defaults changed to
openai/gpt-6.1-sol#xhigh. Actual task Workers remainedopenai/gpt-6-luna-fast#max, observed across all 23 instances. The harness and total budget also changed, so the extra resolution cannot be attributed to the model change alone. - Inference ended with 20 normal completions, two timeouts
(
pvlib__pvlib-python-1154,sqlfluff__sqlfluff-1763) and one agent failure (pvlib__pvlib-python-1854). All patches, including stopped attempts, were officially scored: 23 completed evaluations, zero empty patches and zero official evaluation errors or infrastructure failures. - Known estimated inference cost: $17.73; separate unknown-usage hold: $1.98, not counted as known expense. Inference wall time was about 81 minutes, followed by 7.2 minutes of official scoring.
- Fixed release commit:
d19e8be0d21180cc23ad2ae4b853d846a18e77bc; package SHA-256:99300ceec0c3eee4fa1f984fed50d15d58b2ff4455df0041850ecd63514b0a51. OpenCode 2.0.20, official harness 5.0.2. Local evidence is retained in_testenv/swebench-v0131-dev23-20260930-r2/result-summary.json; generated predictions, databases and raw logs are not committed.
Mission tools
start_mission: Operator supplies concise requirements; the host saves original messages. A known single unit can includeunitto combine start/planning and return its configured Worker task.plan_units: Operator or Coordinator supplies title, objective, file/directory scopes and formal checks. The host generates IDs, handoff, manifest, proof mapping and the ready Worker task.executor="self"keeps execution in the same controller session;start_direct_unit/finish_direct_unitretain observed formal-check freshness.operator_next: advance serial units.expand_unitreconciles required in-request outputs while preserving the same Task. A reasonedplan_unitscorrection orretry_mission_unithandles ordinary unit recovery under the original requirements and cumulative budget.review_mission: generate the independent review packet from source, requirements and observed checks; dispatch its Reviewer task for high-risk changes or record a low-risk skip.repair_review: record findings and continue correction in the running original Reviewer Task, or resume that same native session. Later findings accumulate; run inherited checks/requested Git delivery and explicitly reportSELF_RECHECKEDin that same Task. LegacyCORRECTION_READYalone requires a same-author read-only fallback throughreview_mission.submit_mission: Coordinator returns a completion candidate, user-only decision, or proven external/scope/budget blocker.complete_mission: Operator compares the original request, source and evidence, then explicitly accepts. Only a succeeded receipt authorizes DONE and the measured 🐾 return report.
All tool names use the sortie_v010_ prefix. Prior proposal/plan-repair tools remain in the compatibility
implementation but are hidden from the normal Mission tool list. Operator/Coordinator handle in-request
path reconciliation without a new user approval. Changes beyond the original requirements or cumulative
budget return to Operator/user. EVIDENCE_GAPS is an advisory limitation, not Review PASS or an
automatic extra review; failed or missing required checks still prevent acceptance.
Durable profile state and hash-bound task references support restart and
compaction recovery without reconstructing criteria from summary prose. Stale,
foreign-root, or changed references are rejected. Requested git add <paths> and git commit -m ...
use the actual source write scope, not a fabricated .git/** scope. An optional host-managed Git
lifecycle also retains its branch, commit and post-commit boundaries. Neither mode grants arbitrary
Git, force push, release or publication authority.
Progress and acceptance evidence
sortie_v010_operator_status keeps original requests, formal command/exit/timing observations,
review disposition and recorded delivery in a compact Mission view. { "view": "progress" } exposes
the current unit, completed/total units, host budget and next action; { "view": "full" } or
details_ref provides full snapshot diagnostics. Unknown clean state or a failed commit is not delivery
success. Reading progress does not dispatch, retry or accept work; use native completion notifications
instead of polling. Existing status reconciliation can recover a missed child terminal event.
Configuration
Profile files and precedence
The default package entry is the v010 Mission profile:
- Command:
/sortie-v010 - Primary agent:
dog-operator - Project settings:
.opencode/sortie-dogs-v010.json - Global settings:
~/.config/opencode/sortie-dogs-v010.json - JSON environment override:
SORTIE_DOGS_V010_CONFIG - Runtime state:
.sortie-dogs-v010/ - Installed asset marker:
.opencode/sortie-dogs-v010.version
For a global install, the marker is <OpenCode config root>/sortie-dogs-v010.version.
Precedence is built-in defaults, global file, project file, environment JSON,
then plugin factory options. Unknown properties or invalid types are rejected.
Use external v010 role names such as dog-operator, dogs-coordinator, and
dog-reviewer-v010 in modelRouting; do not also declare their stable aliases.
Example .opencode/sortie-dogs-v010.json:
{
"validationProfile": "balanced",
"readOnlyTools": ["my_mcp_search"],
"freeTierFallbackModels": ["opencode/deepseek-v4-flash-free"],
"modelRouting": {
"dog-operator": {
"preferred": { "model": "provider/model", "variant": "high" }
},
"dogs-coordinator": {
"preferred": { "model": "provider/model", "variant": "deep" }
}
},
"modelCatalog": {
"project": [
{ "model": "provider/model", "variants": ["high", "deep"] }
]
},
"continuation": {
"enabled": true,
"taskWatchdogMilliseconds": 300000
}
}
Declare only models and named variants the host actually provides. Sortie does not invent, probe, or translate variant names.
Settings reference
readOnlyTools: additional host-specific tools known not to mutate project files. Values accumulate across configuration layers. Unknown tools are denied in a bound worker session.modelRouting: preferred and ordered fallback targets by external profile role.modelCatalog: availableprojectandglobalmodel/variant declarations.freeTierFallbackModels: ordered global last-resort model IDs. Default:opencode/deepseek-v4-flash-free;[]disables this fallback.dedicatedWorkerModel: canonical stable serial target, defaultopenai/gpt-6.1-sol/medium. The Mission profile supplies its explicit role routes below; do not infer its Worker route from this stable setting.consultation.strategy: fixed advisor identity, optionalrequired, and positivemaxCallsPerCandidate; default one call and not required.consultation.sourceReview: risk-based review withmaxCallsPerCandidatedefault1andmaxArtifactBytesdefault/maximum30720. Unavailable review blocks only when review is required.continuation.enabled: defaulttrue.- Automatic continuation has no turn-count ceiling. The legacy positive-integer
continuation.maxAutoContinuessetting is accepted but ignored. continuation.taskWatchdogMilliseconds: root inactivity while an implementation Task is outstanding; default300000, valid range10..1800000.continuation.summarizeModel: optional explicit compaction model; omission reuses the latest observed root model.validationProfile:fast,balanced, orassurance; defaultbalanced.reflection: enabled by default forrun,projectandgloballayers, with at most three entries / 500 estimated tokens injected. Root Operator can usesortie_v010_reflectionto retain verified process causes/preventions for later turns and sessions; this is not model training. Storage and managed blocks are separate from stable. Setreflection.enabledtofalseto disable.
The Mission host owns handoff and manifest controls under
.sortie-dogs-v010/contracts/. Do not create a legacy root
operation-manifest.json for this profile and do not edit generated controls.
Delete .sortie-dogs-v010/ only when no Sortie run is active.
Validation policy
validationProfile chooses supplementary non-canonical depth:
fast: static checksbalanced: targeted checksassurance: related checks
It does not replace meaningful declared formal checks or user/project-required broad validation.
Batch related edits, run focused checks, then execute every declared formal check in order on the
stable candidate. The implementing Worker or admitted correcting Reviewer runs those commands;
canonical/full-suite owner=coordinator is evidence accounting, not a requirement for root execution.
Keep required broad checks for the final integrated candidate. Reuse valid unchanged evidence only when the contract permits; identity includes candidate, command, environment, scope and owner. Required repeated occurrences retain their own execution identities and cannot be skipped as duplicates. Native commands, actual working directory, exit, duration and saved source bindings establish freshness. Diagnostics do not substitute for formal proof. Repeat checks when changes, failures or freshness require it, rather than solely because a Worker changed or documentation was edited.
Default routes
dog-operator:openai/gpt-6.1-sol/xhighdogs-coordinator:openai/gpt-6.1-sol/xhighdog-worker-v010:openai/gpt-6-luna-fast/maxdog-scout-v010:openai/gpt-6-luna-fast/maxdog-reviewer-v010:openai/gpt-6.1-sol/xhighdog-advisor-v010:openai/gpt-6.1-sol/xhigh
An explicit model and variant selected in OpenCode remains authoritative for that session. Child role defaults fill absent native settings and may be overridden by valid profile routing. Review never silently inherits the implementation model.
Stable compatibility profile
The earlier parallel-capable runtime remains available explicitly:
npx sortie-dogs init . --profile stable
Load it through a project bridge instead of the default package plugin entry:
export { SortieDogsPlugin } from "sortie-dogs/plugin/stable";
The stable profile uses /sortie, dog-coordinator,
.opencode/sortie-dogs.json, SORTIE_DOGS_CONFIG, and .sortie-dogs/. Do not
register stable and v010 from the same package installation path in one host.
Global availability
Project-local installation is recommended. To expose the current Mission assets globally:
npm install --global sortie-dogs@0.13.8
sortie-dogs init --global --profile v010
Global initialization registers the package or reuses an existing local V2 bridge, and sets subagent depth to at least two, preserving unrelated settings and the default agent.
An existing <OpenCode config root>/plugins/sortie-dogs/index.js bridge importing sortie-dogs/server
can resolve a separate dependency under that config root. Updating npm-global alone does not update
it. For that layout, also install the same release at the actual config root, then rerun global init:
npm install --prefix "$HOME/.config/opencode" sortie-dogs@0.13.8
sortie-dogs init --global --profile v010
The command shows the default config root; use your actual root if overridden. For a configured npm
package entry, OpenCode V2 also provides opencode plugin list / opencode plugin update; exact
version pins require an explicit version change. Completely restart OpenCode after updating, then
check the installed package, asset marker and loaded plugin version.
Updates and removal
For project-local updates, replace the dependency, rerun initialization and completely restart OpenCode:
npm install --save-dev sortie-dogs@latest
npx sortie-dogs init .
init is idempotent. It updates recognized Sortie-owned assets, records the asset
version, preserves user configuration, and stops safely on unknown ownership or
conflicting files.
Align any exact version pin or separate bridge dependency with the intended release too. An installed
marker of 0.13.8-native-binding-v1 identifies the assets; it does not prove an already-running
OpenCode process has reloaded the plugin.
There is no supported uninstall command. Remove the npm dependency separately,
then follow the safe manual removal guide. Delete only known
Sortie-owned paths; never remove the whole .opencode directory or use broad
wildcards.
Maintainers: the release batch guide covers fixed-tarball validation, global application, GitHub publication, and manual npm publication.
Similar plugins
Dynamic Workflows
@malhashemi/opencode-dynamic-workflows
Deterministic multi-agent Workflows for OpenCode: TypeScript scripts that fan work out to subagents, with typed results, a live TUI panel, a web app and a versioned protocol.
Herdr
opencode-herdr
Route OpenCode agents through Herdr panes as herdr/<adapter>/<model> (Cursor, Claude, Codex, OpenCode)
Workflows
@rphang/opencode-workflows
Claude Code-style dynamic workflows for opencode v2: the model writes a JS orchestration script that fans work out to parallel subagents (agent, parallel, pipeline, budget, resume)