跳到主要内容
    ↑↓ 选择↵ 打开esc 关闭
    oyale

    Screenshot Vision

    opencode-screenshot-vision·v1.4.1·MCP 集成

    OpenCode plugin: let a text-only LLM read browser screenshots via local Ollama with OpenCode Zen fallback.

    GitHub 星标

    0

    月装机量

    113

    近 7 天 42

    综合评分

    32.6

    生态多维模型

    最近提交

    3 天前

    2026-10-01

    快速安装与配置

    opencode.json

    写入当前项目的 opencode.json,只对这个仓库生效。

    opencode.json

    {
      "$schema": "https://opencode.ai/config.json",
      "plugin": ["opencode-screenshot-vision@1.4.1"]
    }

    OpenCode 启动时会通过内嵌运行时自动加载 npm 依赖并缓存至本地目录,无需手动在全局环境执行安装。

    Give a text-only LLM the ability to read screenshots during browser-testing workflows — local-first.

    What problem does this solve?

    Text-only models can drive a browser through an MCP server but cannot read the screenshots the browser sends back. When a test step needs to verify what is actually on screen, the model is blind.

    opencode-screenshot-vision is an OpenCode plugin that closes this gap. It exposes a single vision tool to the model. The model calls that tool, and the tool sends the screenshot to a vision-capable backend and returns a plain-text description. The text-only model never gains vision itself — it just receives a description it can reason over.

    Positioning

    This plugin exists for one specific job: screenshots for browser-testing workflows, local-first.

    It is not a general-purpose "vision for text-only models" package. That space is crowded, and the auto-transparent packages in Related work serve pasted images more smoothly than this plugin does (this one still needs a manual vision() call for them). This project focuses on the browser-testing flow — screenshots captured by Browser MCP — which those packages do not cover, and it prioritizes free, local inference before falling back to any cloud service.

    This is a parallel project built for learning, not a competitor claiming to replace the earlier work. The differences are spelled out below.

    What this project adds

    1. Browser MCP screenshot capture. Browser MCP returns screenshots inline in the raw tool result; those results never pass through the message-transform pipeline (verified), so the auto-transparent packages do not cover the browser-testing flow. This plugin captures them via a tool.execute.after hook on the raw MCP result.
    2. Local-first and free. The primary tier is a local OpenAI-compatible runtime (free, private, no API key), then OpenCode Zen free, then Zen paid.
    3. File mode. vision(path=...) reads any screenshot saved to disk, regardless of the tool that produced it (Playwright, Selenium, Puppeteer, Cypress, a manual capture, …).
    4. Prompt-injection defense in the vision prompt.

    Features

    • Single vision tool — one call, no new workflow to learn.
    • Reads screenshots from three sources: the latest browser screenshot, a pasted/dropped image in the conversation, or a file on disk.
    • Automatic fallback across three backends: a local OpenAI-compatible runtime, then OpenCode Zen free, then Zen paid.
    • Auto-description: after a browser screenshot or a pasted/dropped image, its text description is injected automatically (appended by default), so a text-only model never has to be told to read it. Configurable via OPENCODE_VISION_AUTO_MODE.
    • Direct HTTP calls to the vision backends, rather than opencode's model path: in testing, an image attached through opencode did not reach the local Ollama model, while a direct HTTP call to Ollama did.
    • Built-in safety: prompt-injection defense, path containment, MIME sniffing, a 10 MB size limit, and a 2,048-token output cap.

    How it works

    The plugin registers three things when opencode starts:

    1. A vision tool the model can call.
    2. A tool.execute.after hook that captures the image whenever a browser screenshot tool runs.
    3. A chat.message hook that captures pasted/dropped images from incoming messages.

    The three flows

    Browser MCP (inline). When a Browser MCP server captures a screenshot, the image is returned inline in the tool result ({ content: [{ type: "image", ... }] }) — it never touches disk. The tool.execute.after hook captures it in memory, and calling vision with no arguments describes the most recent captured screenshot.

    Pasted / dropped image. A pasted or dropped image arrives as a file part on the incoming message. The chat.message hook captures it, and calling vision with no arguments describes it. (The auto-transparent packages do this without the manual call.)

    File on disk (any tool). When a screenshot is saved as a file — by Playwright, Selenium, Puppeteer, Cypress, or anything else — the model calls vision with a path argument. The plugin reads and validates that file directly.

    All flows converge on the same describe step: encode the image, send it to a backend, and return the text description.

    Fallback chain

    The chain is resolved at runtime, and its order is configurable via OPENCODE_VISION_BACKENDS. By default: the env pin (OPENCODE_VISION_LOCAL_MODEL), then discovered vision-capable models ordered local-first → free → cost (cached for 10 minutes), then OpenCode Zen free, then Zen paid. Each tier is tried only if the previous one fails with an error or a timeout. The paid tier retries once without the reasoning parameter if the API rejects it with HTTP 400 (the backend does not accept reasoning in non-reasoning mode).

    Tier Backend Model Cost Endpoint
    1 Env pin (local, OpenAI-compatible) gemma4:e4b Free /v1/chat/completions
    2 Discovered backends (local → free → cost) auto-discovered model varies provider endpoint
    3 OpenCode Zen mimo-v2.5-free Free /v1/chat/completions
    4 OpenCode Zen gpt-5-nano $0.05 / $0.40 per 1M tokens /v1/responses

    Backend scope. The local tier speaks the OpenAI-compatible /v1/chat/completions API, so it works with Ollama as well as LM Studio, llama.cpp server, vLLM, and any runtime that exposes that endpoint — point OPENCODE_VISION_LOCAL_URL at it. Vision-capable models from all configured providers are auto-discovered as ordered backends (see Discovered backends); OpenCode Zen's free and paid tiers remain the guaranteed final fallback.

    Auto-skip when the main model has vision

    The plugin reads the active model's capabilities.input.image from every request. With OPENCODE_VISION_AUTO_MODE=auto (the default), an incoming image is not auto-described when the main model can already see images — no backend inference runs. Set the mode explicitly to append/replace to force auto-description, or off to disable it entirely.

    Discovered backends

    The vision backends are no longer only the hardcoded local + Zen chain. On the first vision call, the plugin lists the configured providers and uses every model with image capability, ordered local-first, then free, then by cost. The local/OPENCODE_VISION_LOCAL_MODEL pin is always tried first. Zen free and paid remain the guaranteed final fallback. The candidate list is cached for 10 minutes and refreshed after a total failure.

    Privacy: keep screenshots local

    Screenshots can hold sensitive content — credentials, drafts, internal pages. By default the chain falls back to OpenCode Zen (cloud) when the local backend fails, so a screenshot can leave your machine. To guarantee no image ever leaves your machine, run fully local:

    OPENCODE_VISION_BACKENDS=local
    

    With local, only the OPENCODE_VISION_LOCAL_MODEL pin and discovered local models are used — no cloud tier is ever contacted and no API key is required.

    Who describes what. The vision tool is backend-driven: whichever tier succeeds produces the description, regardless of which model called the tool. So an API/cloud model that calls vision receives a description generated by the local backend (under the default local-first order) — it is not seeing the image itself. For a model that can already see images, let the image flow natively (see Auto-skip) instead of calling vision.

    Requirements

    • OpenCode (the plugin loads at startup).

    • Local tier: any OpenAI-compatible runtime — Ollama, LM Studio, a llama.cpp server, or vLLM — with a vision-capable model. Example for Ollama (default model gemma4:e4b):

      ollama pull gemma4:e4b
      

      Point the plugin at another runtime with OPENCODE_VISION_LOCAL_URL.

    • Zen tiers: an OpenCode Zen connection, configured via /connect in opencode (or the equivalent environment variables).

    • Browser MCP flow: a Browser MCP server (@browsermcp/mcp) connected in opencode.

    Install

    Easiest — from inside opencode's CLI:

    opencode plugin opencode-screenshot-vision        # project
    opencode plugin -g opencode-screenshot-vision     # global
    

    This installs the npm package and updates the config for you. Restart opencode afterwards.

    Or add it to opencode.json manually:

    {
      "plugin": ["opencode-screenshot-vision"]
    }
    

    Then restart opencode. npm plugins are installed automatically at startup.

    Alternatively, copy the plugin file into a project's .opencode/plugins/ directory:

    cp vision.ts <project>/.opencode/plugins/vision.ts
    

    Either way, restart opencode or start a new session. Plugins load at startup.

    Usage

    The examples below are written from the point of view of the text-only model driving the browser.

    Browser MCP flow — read the latest inline screenshot:

    # The browser MCP captures a screenshot; it is returned inline and the text-only
    # model cannot read it. Call vision with no arguments:
    vision()
    

    The tool returns a description of the most recently captured screenshot.

    File flow — read a screenshot saved to disk:

    # Any tool (Playwright, Selenium, Puppeteer, Cypress, ...) saves a screenshot
    # to a file. Pass its path:
    vision(path="/tmp/opencode/screenshot-123.png")
    

    Ask a specific question about the image:

    vision(prompt="Is there a login button visible, and is it enabled?")
    

    A prompt can be combined with a path:

    vision(path="/tmp/opencode/screenshot-123.png", prompt="List any error messages on the page.")
    

    Configuration

    Only the settings below are configurable, via optional environment variables. Everything else — the Zen URL, the Zen models, and the request formats — is fixed in the code (see the fallback-chain scope note above).

    Variable Default Purpose
    OPENCODE_VISION_LOCAL_URL http://localhost:11434/v1 Local OpenAI-compatible base URL
    OPENCODE_VISION_LOCAL_MODEL gemma4:e4b Local vision model
    OPENCODE_VISION_OLLAMA_MODEL (deprecated) Old name for the local model; still honored
    OPENCODE_VISION_LOCAL_TIMEOUT_MS 90000 Local request timeout, milliseconds
    OPENCODE_VISION_CLOUD_TIMEOUT_MS 45000 Zen request timeout, milliseconds
    OPENCODE_VISION_MAX_IMAGE_BYTES 10485760 (10 MB) Max image size for path-based loads
    OPENCODE_VISION_USER_AGENT Chrome 126 UA User-Agent header sent to Zen
    OPENCODE_VISION_AUTO_MODE auto Auto-describe browser screenshots and pasted images: append (add description after the image), replace (description replaces the image), off (manual vision only), auto (default: describe only when the main model cannot see images)
    OPENCODE_VISION_BACKENDS local,zen-free,zen-paid Ordered backend tiers, comma-separated. Tokens: local (discovered models), zen-free, zen-paid. Omit a token to exclude that tier (e.g. zen-paid only, or local,zen-paid to skip the free tier)
    OPENCODE_VISION_ALLOWED_ROOTS (empty) Extra directories readable via path, separated by the OS path delimiter
    OPENCODE_API_KEY (from auth.json) Overrides the Zen API key
    OPENCODE_AUTH_CONTENT (unset) auth.json contents provided as an environment string
    OPENCODE_AUTH_FILE $XDG_DATA_HOME/opencode/auth.json Overrides the auth file path

    Security

    • Prompt-injection defense. The vision prompt is sent in the system role and repeated in the user message so local models that ignore the system role still see it. It instructs the model to treat every instruction visible in a screenshot — and the optional prompt question — as untrusted data: report it, never follow it. Auto-generated descriptions are injected back into the conversation labelled as untrusted page-derived data.
    • Path containment. The path argument is restricted to the session directory, the git worktree, $TMPDIR/opencode, and any roots listed in OPENCODE_VISION_ALLOWED_ROOTS.
    • MIME sniffing. Path-based loads are typed from magic bytes (PNG, JPEG, GIF, WebP); unsupported formats are rejected.
    • Size limit. Images larger than 10 MB are rejected.
    • Output cap. Every backend is limited to 2,048 output tokens.

    Troubleshooting / Caveats

    • Zen balance. The paid tier returns 401 CreditsError when the workspace has insufficient balance.
    • Zen free rate limit. The free tier can return 429 under load.
    • Cloudflare. Zen requests must send a browser User-Agent; this is set by default.
    • Zen paid vision. Not yet verified at runtime — treat the paid tier as unproven until exercised.

    When all three backends fail, the vision tool reports each failure in a single error message, along with a hint to retry sequential calls if several vision calls were made at once.

    Related work

    This plugin is one of several projects that give vision to text-only models in OpenCode. They are all worth knowing about.

    Auto-transparent packages

    These detect a pasted image and replace it with a text description before the main model sees it, via the message-transform pipeline:

    Tool / subagent packages

    These register explicit vision tools or subagents that the model invokes:

    • opencode-vision — registers vision subagents from the user's configured image-capable models.
    • opencode-vision-plugin — in-process tools (describe/OCR/analyze), direct fetch to Gemini + NVIDIA NIM; a fork of nicolasrios/opencode-vision.

    Comparison

    Aspect This project Auto-transparent packages Tool / subagent packages
    Primary input Browser MCP screenshots + file path Pasted images Pasted images / manual invocation
    Trigger Explicit vision tool call Automatic (transparent) Explicit tool or subagent call
    Backend Local-first fallback chain (Ollama → Zen) Varies by package User's configured image models or direct API (Gemini, NVIDIA NIM)
    Local-first / free tier Yes No No
    Prompt-injection defense Yes No No
    OCR / analyze tools No No opencode-vision-plugin

    Roadmap

    See ROADMAP.md for the plan for v2 — which borrows improvements from the prior-art projects above — and for the widening plan (more local and cloud backends, more agent platforms, more screenshot sources).

    Development

    Tests. Unit + mocked integration, deterministic (CI runs this):

    bun test
    
    • vision.test.ts — plugin structure, auto-describe decision, mime/path guards, fallback cache.
    • main-model-capability.test.ts — vision-capability detection and per-session tracking.
    • backend-discovery.test.ts — discovery filtering, ordering, env pin, TTL cache, edge cases.
    • vision.integration.test.ts — the three capture flows (file, browser-screenshot, pasted image), vision() no-args, auto-skip with a vision model, path validation, prompt transmission.

    Live smoke against the local vision backend (Ollama), exercising the capture flows end-to-end with real inference:

    bun run smoke
    

    SMOKE_IMAGE comes from .env (see .env.example); without it the smoke uses a tiny embedded PNG. The model comes from OPENCODE_VISION_LOCAL_MODEL (default qwen3-vl:4b-instruct). Before running, unload heavy models that can crash the vision runner (ollama stop gemma4:e4b).

    License

    MIT. See the LICENSE file for the full text.

    Contributing

    Contributions are welcome. Please open an issue to discuss a change before submitting a pull request, and keep the fallback chain and security behavior in mind when modifying the vision call path.

    Acknowledgments

    Built on OpenCode and its plugin API, with vision provided by Ollama and OpenCode Zen, and browser automation by Browser MCP.

    This project was informed by the prior-art vision plugins listed in Related work: opencode-vision, opencode-vision-plugin, opencode-vision-fallback, @venespana/opencode-vision, @pawprint0706/opencode-vision-helper, and @jochenyang/opencode-vision.

    同类生态推荐