跳到主要内容
    ↑↓ 选择↵ 打开esc 关闭
    中文English
    TejasS1233

    Parser

    v1.1.0模型接入
    opencode-parser

    Parse any file in opencode — PDF, DOCX, XLSX, PPTX, images (OCR), EPUB, HTML, IPYNB, archives, and plain text

    GitHub 星标

    50

    月装机量

    1,058

    近 7 天 239

    综合评分SCORE

    48.6

    生态多维模型

    最近提交

    2 个月前

    2026-06-05

    快速安装与配置

    opencode.json

    写入当前项目的 opencode.json,只对这个仓库生效。

    opencode.json

    {
      "$schema": "https://opencode.ai/config.json",
      "plugin": ["opencode-parser@1.1.0"]
    }

    opencode 启动时会通过内嵌运行时自动加载 npm 依赖并缓存至本地目录,无需手动在全局环境执行安装。

    An opencode plugin that parses any file into structured text the LLM can work with.

    https://github.com/user-attachments/assets/ed9d6ee7-d30b-43d5-83e0-4e09dafaa422

    Install

    {
      "plugin": ["opencode-parser"]
    }
    

    Or install via CLI: opencode plugin opencode-parser -g

    Or copy src/ into .opencode/tools/ for a local zero-config setup.

    Supported formats

    Format Extensions Extracted
    PDF .pdf Text, metadata, pages
    Word .docx Text, tables, metadata
    Excel .xlsx, .xls, .csv, .tsv Text, tables, sheet names
    PowerPoint .pptx, .ppt Slide text, speaker notes
    Images .png, .jpg, .jpeg, .webp, .gif, .bmp, .tiff OCR text (opt-in)
    EPUB .epub Full text with heading structure
    HTML .html, .htm Body text, headings
    XML .xml Stripped text content
    Markdown .md Raw text
    Jupyter .ipynb Code, markdown, outputs
    ZIP .zip File listing with sizes
    Archives .rar, .7z, .tar, .gz Listing (extraction notes)
    Plain text .txt, .json, .yaml, .toml, .ini Raw content

    Usage

    Parse @report.pdf and give me a summary
    
    parse the spreadsheet at @data.xlsx but only the first 3 sheets
    
    parse @report.pdf and save the full output
    

    Options

    Option Default Description
    filePath Path to the file (required)
    maxChars 50000 Limit output chars (-1 for unlimited). Pass -1 to get the full document.
    extractTables true Extract tables from docs/spreadsheets
    extractImages false Enable OCR for images
    ocrLang "eng" OCR language for tesseract.js (e.g. "eng", "fra", "ara")
    maxPages varies Limit pages/slides/sheets processed
    save false Save the full parsed output as a .md file alongside the original (no truncation)
    outputPath Custom path for the Markdown export (overrides save path)

    How it works

    1. File is verified by magic bytes, not just extension
    2. Type detection dispatches to the right parser
    3. Metadata is extracted (author, pages, sheet count, etc.)
    4. Tables become readable markdown
    5. Large content is truncated gracefully with a note to the LLM

    All 15+ format handlers return the same output structure, so the LLM gets consistent results regardless of file type.

    Development

    npm install
    npm run typecheck
    

    License

    MIT