Vision Delegate
Visual bridge for opencode. Registers the native vision_analyze tool (model configured by the user via the opencode agent model override on vision-agent) with the vision-agent subagent as fallback, and teaches a text-only orchestrator to extract visual in
0
361
217 in 7 days
35.7
Multi-signal model
7 days ago
2026-09-27
Install and configure
opencode.jsonWrites to this project's opencode.json — applies to this repository only.
opencode.json
{
"$schema": "https://opencode.ai/config.json",
"plugin": ["opencode-vision-delegate@0.2.0"]
}Writes to ~/.config/opencode/opencode.json — applies to every project.
~/.config/opencode/opencode.json
{
"$schema": "https://opencode.ai/config.json",
"plugin": ["opencode-vision-delegate@0.2.0"]
}If you want to modify the plugin locally, install it into the project and reference the local path.
shell
pnpm add -D opencode-vision-delegateOpenCode loads npm dependencies through its embedded runtime on startup and caches them locally — no manual global install needed.
Disclaimer: OpenCode Vision Delegate is an independent, community-built project. It is not built by, endorsed by, or affiliated with the opencode team. It is a port of kilo-vision-bridge (itself based on wezzard/opencode-vision, MIT) back to the opencode plugin SDK, and builds on their design.
Introduction
Give text-only opencode orchestrators (GLM, DeepSeek, and similar models) eyes by delegating visual tasks to a vision-capable model through a dynamically registered vision subagent.
When the orchestrator model is text-only and a task needs pixels — not just accessibility metadata — the plugin's vision skill detects the visual intent, extracts a task-specific JSON response template, delegates the task, and parses the structured findings back into the conversation.
Tool-first architecture. Delegation targets a native plugin tool, vision_analyze, registered through the @opencode-ai/plugin tool hook. The tool runs the visual judgment in-process — it reads the listed image files and calls the configured vision model directly, so no subagent nesting is required and the tool works from any session, including subagent sessions. The skill calls vision_analyze first and falls back to spawning the vision-agent subagent only when the tool is unavailable in the session or the call fails with a provider/protocol/HTTP error. The vision-agent subagent also has an advanced role: an orchestrator may delegate a whole investigation (multi-image sweeps, deep zoom chains) to it — the subagent can call vision_analyze, including region crops, in its own loop before returning its single final JSON. With the tool path, permission.task configuration is not needed for visual delegation.
About the v2 plugin architecture
opencode is migrating to a new ("v2") internal plugin architecture. As of opencode 1.18, the v2 plugin surface covers agents, skills, catalog, commands, and integrations — but not custom tools, chat message/system transforms, or permission hooks, which are exactly what this plugin's core features need. The official plugin docs still document the hooks API ("v1") as the way to write plugins, and the opencode plugin installer loads it.
This package therefore ships both entries in one default export:
server(hooks API) — full features:vision_analyzetool, per-model vision routing, image materialization, permission handling. This is whatopencode pluginloads today.setup(v2 API) — forward compatibility: registers thevision-agentsubagent and the vision skill through the v2agent/skilldomains. A v2 host that loadssetupgets subagent + skill delegation (the tool-dependent features stay on the hooks entry until upstream v2 gains tool registration).
The two entries write disjoint registration systems, so there is no double-registration conflict if both loaders ever see the package.
Requirements
- opencode 1.18+
- At least one configured provider with an image-capable model (
enabled_providersand/orproviderentries in opencode config). The plugin discovers models from your configured providers and opencode's cached model catalog (~/.cache/opencode/models.json) — it does not ship a fixed model list.
Installation
# install globally (available to all projects) — installs the package and patches the config
opencode plugin opencode-vision-delegate --global
# or install for the current project only
opencode plugin opencode-vision-delegate
# or install straight from GitHub (git source, works the same)
opencode plugin github:ChengZiiii/opencode-vision-delegate --global
Pin a specific version with opencode plugin opencode-vision-delegate@<version> (or a branch/tag with github:ChengZiiii/opencode-vision-delegate#<ref>). The command adds the package to the plugin array in your opencode config (global: ~/.config/opencode/opencode.json; project: .opencode/opencode.json) and manages the package under opencode's package store (~/.cache/opencode/packages/). Equivalent manual config:
{
"plugin": ["opencode-vision-delegate"]
}
After install, restart opencode. The plugin registers the vision-agent subagent and the vision_analyze tool on launch; the vision skill is discovered straight from the installed package directory via skills.paths.
Renamed from
opencode-vision-bridge(the npm name belongs to an unrelated pre-captioning plugin by martinmose). If you installed this plugin before the rename via the git specgithub:ChengZiiii/opencode-vision-bridge: switch thepluginentry in your opencode config togithub:ChengZiiii/opencode-vision-delegate, runopencode plugin github:ChengZiiii/opencode-vision-delegate --global --force, and delete the old store dir~/.cache/opencode/packages/github_ChengZiiii/opencode-vision-bridge— opencode keys its package store on the install spec, so the renamed spec installs to a new dir and the old one is dead weight. (The old GitHub URL redirects, but keeping the old spec in your config would keep installing under the old store key.)
Why this package ships no
build/postinstallscripts: opencode's bundled installer runs npm's git-dependency preparation whenever an installed-from-git package declares any ofpreinstall/install/postinstall/prepack/prepare/build(or aworkspacesfield), and that preparation fails inside the compiled opencode binary — see opencode issue #49704. This package commits a pre-builtdist/index.jsand keepsscriptsfree of those names, so GitHub installs work on any machine.
Alternative installs
- Local checkout:
"plugin": ["file:///<repo-absolute-path>"]— the skill is discovered straight from the package directory viaskills.paths. - Single file: copy
dist/index.jsto~/.config/opencode/plugin/vision.jsand copySKILL.mdto~/.config/opencode/skills/vision/SKILL.mdmanually (there is no installer script).
Do not mix install methods for the same plugin id (vision) — they would double-register.
Updating / uninstalling
To upgrade to a newly published version, --force alone is NOT enough: the package store pins the exact version at first install and a forced reinstall reuses it (verified on 0.1.0 → 0.2.0). Delete the store directory (or directories) first, then reinstall:
# remove every store dir for the npm spec — note there may be TWO:
# ~/.cache/opencode/packages/opencode-vision-delegate AND .../opencode-vision-delegate@latest
# (the runtime loads the one your `opencode agent list` output references)
opencode plugin opencode-vision-delegate --global
Then restart opencode and verify with opencode agent list.
opencode 1.18 has no built-in plugin uninstall command. To uninstall manually (verified working):
- Remove the entry from the
pluginarray in your opencode config (~/.config/opencode/opencode.jsonfor global,.opencode/opencode.jsonfor project). - Delete the package from opencode's store:
~/.cache/opencode/packages/<sanitized-spec>/(e.g.opencode-vision-delegateorgithub_ChengZiiii/opencode-vision-delegate). - Delete
~/.config/opencode/skills/vision/if it exists (a leftover from a single-file install's manual copy). - Optionally remove the
agent["vision-agent"]model knob — otherwiseopencode agent listkeeps showing the name.
Restart opencode and the vision-agent subagent, the vision_analyze tool, and the vision skill are gone.
Quick start
- Install with
opencode plugin opencode-vision-delegate --globaland restart opencode. - Set the vision model (see below):
agent["vision-agent"].model = "<provider-id>/<model-id>"— use any vision-capable model from your configured providers. - Drag an image into the opencode input (or reference an image path) and ask a visual question. The orchestrator detects the visual intent, delegates to
vision_analyze, and the vision model returns structured JSON matching the template.
Usage
The vision model knob
The plugin registers exactly one subagent, vision-agent, WITHOUT a default model — the plugin never writes model. The vision model is set by you through the agent model override on vision-agent:
{
"agent": {
"vision-agent": {
"model": "<provider-id>/<model-id>" // any vision-capable model from your configured providers
}
}
}
If no override is set, opencode falls back to the default model.
Free models: zero-setup vision
opencode's Zen gateway (provider id opencode) serves free models that need no provider authentication — e.g. opencode/space-bunny-free (image input, 1M context). The plugin treats this gateway as keyless: with no auth.json entry and no OPENCODE_API_KEY, the vision_analyze tool simply sends the request without an auth header. Zero-setup usage:
{
"agent": {
"vision-agent": { "model": "opencode/space-bunny-free" }
}
}
Every other provider keeps the strict contract: a missing API key is a provider error naming the fix, and no request is made. If Zen ever starts requiring auth for a model, the HTTP 401 surfaces as a normal provider error (the skill's subagent fallback applies).
This override is the single vision model knob for both delegation paths: the vision_analyze tool's model source is exactly this override, and the vision-agent subagent fallback uses it too. The plugin never writes the field, so your override is always preserved.
Request tuning knobs (tool path)
The vision_analyze tool calls the provider directly, so knobs opencode's runtime would normally inject (model variants, thinking control) do not apply to it. Three knobs cover the gap — all optional, set on the same vision-agent entry (or via env):
{
"agent": {
"vision-agent": {
"model": "minimax-cn-coding-plan/MiniMax-M3",
// 1. Native opencode agent option, passed through to the request body
// (default 0.1 when unset).
"temperature": 0.1,
// 2. Generic body passthrough: deep-merged into the final request
// body — e.g. disable thinking on GLM endpoints.
"extraBody": { "thinking": { "type": "disabled" } }
}
}
}
# 3. Timeout ceiling for one vision call (default 300000ms; <=0 disables
# the timeout; invalid values fall back to the default).
OPENCODE_VISION_TIMEOUT_MS=300000
Body semantics: OpenAI-shaped requests set no max_tokens (server default applies); Anthropic-shaped requests keep the protocol-required max_tokens (default 8192) — override it through extraBody if needed. On timeout the error reads timeout after <N>ms posting <url> instead of a generic network error, so the skill's subagent fallback routing still applies.
Dual request shapes
The tool selects the HTTP request shape from the resolved endpoint URL: endpoints containing /anthropic get Anthropic /messages bodies (base64 image blocks, x-api-key); everything else gets OpenAI /chat/completions bodies (image_url data: URLs, Bearer auth). This matters in practice: some providers' OpenAI-compatible endpoints silently drop data: images (measured on api.minimaxi.com), while their Anthropic-style endpoint delivers them — the built-in endpoint map points minimax-family providers at the Anthropic-style URL.
Per-model vision routing
The plugin routes images based on the handling model of each request, not a single global toggle. The vision-agent subagent is always registered — regardless of the top-level model — so a text-only agent in a mixed config can always delegate.
- Multimodal (vision-capable) model. Image
FileParts pass through untouched — the model sees images natively. A system transform injects a[vision:native]instruction telling the model to inspect images directly and NOT to delegate plain reading to the vision skill or avision-*subagent. One narrow exception: when small text or fine detail is beyond native resolution, the model MAY callvision_analyzeWITH aregionto zoom into that area of an image file on disk (the assist needs a file path — natively attached dropped images are not materialized, so it applies to tool-output screenshots and user-provided paths). - Text-only model. Image
FileParts are materialized under the plugin's temp dir (<system-tmp>/opencode-vision-delegate/) and rewritten to[vision:dropped-image]markers carrying the resulting path. The orchestrator then delegates via thevision_analyzetool, falling back to thevision-agentsubagent only when the tool is unavailable or errors with a provider/protocol/HTTP failure.
Capability is resolved per request: the messages transform checks the message's info.model first, then the agent's configured model, then the top-level config model as a final fallback. Provider/model ids match case-insensitively. To bypass the skill per task on a text-only model, prepend this to your prompt:
You MUST not use the vision skill.
The vision_analyze tool and permissions
The tool's arguments are images ([{id, path, region?}] — short contract ids plus local image paths, each with an optional region), question (the exact visual question), response_template (JSON string defining the required response shape), and optional response_rules. The tool reads the images, calls the configured vision model, and returns exactly one JSON object matching the template.
Region crop (zoom). Each images entry may carry region: [x1, y1, x2, y2] — integer pixel coordinates in the ORIGINAL image (x2/y2 exclusive, clamped to bounds). PNG only, detected by file signature (a PNG named .jpg still crops; cropped entries are submitted as image/png). The crop happens in memory — no files are written — and keeps the region at full source resolution, so provider-side downsampling never applies. The tool appends a coordinate-mapping note to the request for cropped images, and the vision model reports coordinates in original-image pixel space. The zoom workflow: call on the full image first, then re-call with a region around whatever needs a closer look. Non-PNG images with a region, unsupported PNG variants (16-bit, interlaced), and empty regions return a crop error naming the image and the fix (re-call without region, or convert to PNG); the skill neither retries nor falls back on it.
vision_analyze is auto-allowed: sessions — including subagent sessions — call it without permission prompts. An explicit user deny wins: set permission.vision_analyze = "deny" and the plugin never upgrades it.
Disabling vision delegation
To stop delegation entirely, set disable: true on agent["vision-agent"]. This disables both paths: the agent does not appear in opencode agent list and cannot be delegated to, and the vision_analyze tool is not registered in any session. The plugin never writes disable, so your setting stays effective.
Troubleshooting
- Plugin not loading: run opencode with
opencode --print-logsand check the output for plugin load errors. - Stale plugin cache: reset the package store under
~/.cache/opencode/packages/and restart. - Missing
vision-agentsubagent /vision_analyzetool: no model is pre-configured — set one via the agent model override onvision-agent. Confirm the override is set and that~/.cache/opencode/models.jsoncontains the provider. The tool is also absent whendisable: trueis set onvision-agent. vision_analyzereturns "model not configured": setagent["vision-agent"].modelto a vision-capable provider/model from your configured providers; the skill does not fall back on this error.vision_analyzereturns a provider error: check the API key (opencode auth login <provider>/auth.jsonentry, or the provider's*_API_KEYenv var) and the endpoint (provider.<id>.options.baseURLin config, else the catalog/built-in endpoint). The skill automatically falls back to thevision-agentsubagent on provider/protocol/HTTP errors.vision_analyzereturns a crop error:regionwas used on an image that is not a PNG (by signature), an unsupported PNG variant (16-bit, interlaced, >40 MP), or an empty region. The error names the image and the fix: re-call withoutregion, or convert the image to a standard 8-bit non-interlaced PNG. This category is deterministic and local — it never triggers the skill's retry or subagent fallback.- A multimodal model still delegates to the vision skill /
vision_analyze: update the plugin. Before the keyless-provider fix, capability resolution only consulted providers that passed availability gating (config blocks / env keys /auth.json), so keyless providers like opencode Zen (opencode/space-bunny-free,opencode/big-pickle, …) were misjudged text-only and their images were rewritten to[vision:dropped-image]markers. Capability is now resolved from the full cached model catalog (~/.cache/opencode/models.json) plus config provider model overrides, independent of availability.
License
MIT — see LICENSE.
Upstream design and implementation: wezzard/opencode-vision, MIT, and kilo-vision-bridge. See also I Gave GLM-5.2 Eyes for the design rationale.
Similar plugins
Vision
opencode-vision
Dynamic visual-response skill for opencode. Registers vision subagents from OpenCode's configured image-capable models and teaches a text-only orchestrator to extract visual intent, design a task-specific JSON response template, and delegate to a vision s
Vision Paste
opencode-vision-paste
OpenCode plugin: intercept pasted images → local VL API analysis → replace with text
Subagent Magazine
opencode-subagent-magazine
OpenCode TUI plugin monitoring sub-agent invocation status in real time