跳到主要内容
    ↑↓ 选择↵ 打开esc 关闭
    toninho09

    Cache Scope

    @toninho09/opencode-cache-scope·v0.3.0·可观测与分析

    opencode TUI sidebar panel with prompt cache hit rate and per-category cost for the current session

    GitHub 星标

    0

    月装机量

    2,277

    近 7 天 22

    综合评分

    39.4

    生态多维模型

    最近提交

    29 天前

    2026-09-06

    快速安装与配置

    opencode.json

    写入当前项目的 opencode.json,只对这个仓库生效。

    opencode.json

    {
      "$schema": "https://opencode.ai/config.json",
      "plugin": ["@toninho09/opencode-cache-scope@0.3.0"]
    }

    OpenCode 启动时会通过内嵌运行时自动加载 npm 依赖并缓存至本地目录,无需手动在全局环境执行安装。

    A sidebar panel for the opencode TUI showing how much of your session context is being reused from prompt cache, and what each part costs.

    LLM providers charge very different rates for fresh tokens versus tokens replayed from cache. Without visibility you only find out on the invoice. This panel gives the feedback while you work.

    It never touches your sessions: no alerts, no config, no interference with any other opencode flow. The only thing it writes is whether you left the subagent list expanded.

    Install

    opencode plugin add @toninho09/opencode-cache-scope
    

    Or add it manually to tui.json (~/.config/opencode/tui.json or .opencode/tui.json):

    {
      "$schema": "https://opencode.ai/tui.json",
      "plugin": ["@toninho09/opencode-cache-scope"]
    }
    

    TUI plugins go in tui.json, not opencode.json. Restart opencode afterwards.

    Requires opencode >= 1.18.0.

    What it shows

    The panel stays hidden until the session proves it uses caching (at least one turn with cache read or cache write above zero), so models without cache support never add noise to the sidebar.

    Usage
    hit rate: 90%
    cache read: 45.0k ($0.0135)
    cache write: 3.0k ($0.0113)
    input: 2.0k ($0.0060)
    output: 1.2k ($0.0180)
    total: $0.0488
    subagents: 3 ($0.0072) +
    
    • hit rate — share of the processed context that came from cache.
    • cache read — tokens replayed from cache, plus their cost.
    • cache write — tokens written into the cache on their first pass, plus cost.
    • input — tokens that went through neither cache path, plus cost.
    • output — tokens generated by the model, plus cost. Includes reasoning tokens, which several providers report separately but opencode bills at the output rate.
    • total — everything the session has been billed so far.
    • subagents — how many child sessions the task tool spawned, and their share of the total.

    The total is not a per-turn figure and it is not only what you typed: it is the whole session, subagents included. If your main agent runs on a subscription model priced at $0 and your subagents run on a metered one, the total will equal the subagent line. That is arithmetic, not a bug.

    Token counts are abbreviated (12.3k, 1.2m). Costs use 4 decimals because a single turn is usually a fraction of a cent.

    Subagents

    The task tool runs each subagent in its own child session. Their tokens are counted in every line above, so the panel reflects what the session actually burned rather than only the turns you typed. This matters: on a session driven by subagents they can account for over 90% of the cache traffic.

    Click the subagents line to break it down into one row per child, showing its agent name, context tokens and cost:

    total: $0.0488
    subagents: 3 ($0.0072) -
      explore: 162.5k ($0.0031)
      explore: 44.3k ($0.0025)
      general: 25.5k ($0.0016)
    

    The line only appears once the session has spawned a subagent. The expanded/collapsed choice is remembered across restarts.

    How the numbers are computed

    Token counts and the total cost come from the session record the opencode server maintains, which accumulates every turn of the session plus, here, every child session spawned by the task tool.

    Deriving these from the message list would be wrong: the TUI keeps only the last 100 messages of a session in memory and discards the rest, so a long session would report a sliding window instead of its lifetime total.

    Why the per-category costs can disappear

    The server tracks tokens per category but only one cost figure for the session, so the split across cache read / cache write / input / output has to be derived locally, by pricing each category's tokens with the session's model.

    That derivation has no way to know a session switched models halfway through — a session planned on a paid model and executed on a subscription one would price everything at the subscription's $0 and show four zeroes against a real bill.

    So the split has to earn its place: if the four costs do not add up to the billed total within 5%, they are dropped and those lines show tokens only. The total is always exact.

    cache read: 22.3m          <- split did not reconcile, cost withheld
    cache write: 338.5k
    input: 491.5k
    output: 177.0k
    total: $4.4493             <- always the billed amount
    

    Hit rate

    hit rate = cache read / (cache read + cache write + input)
    

    The denominator is all input context the model had to process, split three ways: replayed from cache, newly written to cache, and never cached. Output tokens are excluded — they are generated, not context that can be cached.

    If the denominator is zero the hit rate is reported as 0%. The result is multiplied by 100 and rounded to an integer.

    Example: 45,000 / (45,000 + 3,000 + 2,000) = 0.90 → 90%. Nine out of ten context tokens in that session were reused instead of reprocessed at full price.

    Because the value is cumulative, one bad turn will not move it much in a long session. It shows the overall cache health of the session, not a snapshot of the last turn.

    Known limitations

    Some providers apply different rates once the context passes a large threshold (cost.experimentalOver200K in opencode's model data). The per-category split does not reproduce those tiers; the total is unaffected, since it comes from the server.

    Development

    bun install
    bun run typecheck
    bun run build
    

    src/tui.tsx is compiled by scripts/build.mjs into dist/tui.js using babel-preset-solid with the @opentui/solid universal runtime.

    This build step is mandatory, not cosmetic. opencode applies its own Solid JSX transform only to files outside node_modules. A published package lives inside node_modules, so shipping raw .tsx would fall back to Bun's default JSX transform and produce a non-reactive component: the panel would render once and never update.

    License

    MIT

    同类生态推荐