跳到主要内容
    ↑↓ 选择↵ 打开esc 关闭
    TejasS1233

    Parser

    opencode-parser·v2.0.1·代码智能

    Parse any file in opencode. Supports PDF, DOCX, XLSX, PPTX, images with OCR, EPUB, HTML, XML, Markdown, Jupyter notebooks, CSV, TSV, ZIP, RAR, 7z, TAR, GZip, JSON, YAML, TOML, INI, and plain text. Detects format by magic bytes. Extracts text, tables, meta

    GitHub 星标

    61

    近 30 天 +6

    月装机量

    1,744

    近 7 天 707

    综合评分

    59.3

    生态多维模型

    最近提交

    6 天前

    2026-09-29

    快速安装与配置

    opencode.json

    写入当前项目的 opencode.json,只对这个仓库生效。

    opencode.json

    {
      "$schema": "https://opencode.ai/config.json",
      "plugin": ["opencode-parser@2.0.1"]
    }

    OpenCode 启动时会通过内嵌运行时自动加载 npm 依赖并缓存至本地目录,无需手动在全局环境执行安装。

    An opencode plugin that parses any file into structured text the LLM can work with.

    This plugin supports both OpenCode V1 and V2 from the same package. V1 calls server(), V2 calls setup() (requires OpenCode V1 >= 1.18.29 for the object entrypoint).

    Demo video

    Install

    OpenCode V2

    {
      "plugins": ["opencode-parser"]
    }
    

    OpenCode V1

    {
      "plugin": ["opencode-parser"]
    }
    

    Or install via CLI: opencode plugin opencode-parser -g

    Or copy src/ into .opencode/tools/ for a local zero-config setup.

    Supported formats

    Format Extensions Extracted
    PDF .pdf Text, metadata, pages
    Word .docx Text, tables, metadata
    Excel .xlsx, .xls, .csv, .tsv Text, tables, sheet names
    PowerPoint .pptx, .ppt Slide text, speaker notes
    Images .png, .jpg, .jpeg, .webp, .gif, .bmp, .tiff OCR text (opt-in)
    EPUB .epub Full text with heading structure
    HTML .html, .htm Body text, headings
    XML .xml Stripped text content
    Markdown .md Raw text
    Jupyter .ipynb Code, markdown, outputs
    ZIP .zip File listing with sizes
    Archives .rar, .7z, .tar, .gz Listing (extraction notes)
    Plain text .txt, .json, .yaml, .toml, .ini Raw content

    Usage

    Parse @report.pdf and give me a summary
    
    parse the spreadsheet at @data.xlsx but only the first 3 sheets
    
    parse @report.pdf and save the full output
    

    Options

    Option Default Description
    filePath — Path to the file (required)
    maxChars 50000 Limit output chars (-1 for unlimited). Pass -1 to get the full document.
    extractTables true Extract tables from docs/spreadsheets
    extractImages false Enable OCR for images
    ocrLang "eng" OCR language for tesseract.js (e.g. "eng", "fra", "ara")
    maxPages varies Limit pages/slides/sheets processed
    save false Save the full parsed output as a .md file alongside the original (no truncation)
    outputPath — Custom path for the Markdown export (overrides save path)

    How it works

    1. File is verified by magic bytes, not just extension
    2. Type detection dispatches to the right parser
    3. Metadata is extracted (author, pages, sheet count, etc.)
    4. Tables become readable markdown
    5. Large content is truncated gracefully with a note to the LLM

    All 15+ format handlers return the same output structure, so the LLM gets consistent results regardless of file type.

    Development

    npm install
    npm run typecheck
    

    License

    MIT

    同类生态推荐