Parser
Parse any file in opencode. Supports PDF, DOCX, XLSX, PPTX, images with OCR, EPUB, HTML, XML, Markdown, Jupyter notebooks, CSV, TSV, ZIP, RAR, 7z, TAR, GZip, JSON, YAML, TOML, INI, and plain text. Detects format by magic bytes. Extracts text, tables, meta
61
+6 in 30 days
1,744
707 in 7 days
59.3
Multi-signal model
6 days ago
2026-09-29
Install and configure
opencode.jsonWrites to this project's opencode.json — applies to this repository only.
opencode.json
{
"$schema": "https://opencode.ai/config.json",
"plugin": ["opencode-parser@2.0.1"]
}Writes to ~/.config/opencode/opencode.json — applies to every project.
~/.config/opencode/opencode.json
{
"$schema": "https://opencode.ai/config.json",
"plugin": ["opencode-parser@2.0.1"]
}If you want to modify the plugin locally, install it into the project and reference the local path.
shell
pnpm add -D opencode-parserOpenCode loads npm dependencies through its embedded runtime on startup and caches them locally — no manual global install needed.
An opencode plugin that parses any file into structured text the LLM can work with.
This plugin supports both OpenCode V1 and V2 from the same package.
V1 calls server(), V2 calls setup() (requires OpenCode V1 >= 1.18.29 for the object entrypoint).
Install
OpenCode V2
{
"plugins": ["opencode-parser"]
}
OpenCode V1
{
"plugin": ["opencode-parser"]
}
Or install via CLI: opencode plugin opencode-parser -g
Or copy src/ into .opencode/tools/ for a local zero-config setup.
Supported formats
| Format | Extensions | Extracted |
|---|---|---|
.pdf |
Text, metadata, pages | |
| Word | .docx |
Text, tables, metadata |
| Excel | .xlsx, .xls, .csv, .tsv |
Text, tables, sheet names |
| PowerPoint | .pptx, .ppt |
Slide text, speaker notes |
| Images | .png, .jpg, .jpeg, .webp, .gif, .bmp, .tiff |
OCR text (opt-in) |
| EPUB | .epub |
Full text with heading structure |
| HTML | .html, .htm |
Body text, headings |
| XML | .xml |
Stripped text content |
| Markdown | .md |
Raw text |
| Jupyter | .ipynb |
Code, markdown, outputs |
| ZIP | .zip |
File listing with sizes |
| Archives | .rar, .7z, .tar, .gz |
Listing (extraction notes) |
| Plain text | .txt, .json, .yaml, .toml, .ini |
Raw content |
Usage
Parse @report.pdf and give me a summary
parse the spreadsheet at @data.xlsx but only the first 3 sheets
parse @report.pdf and save the full output
Options
| Option | Default | Description |
|---|---|---|
filePath |
— | Path to the file (required) |
maxChars |
50000 | Limit output chars (-1 for unlimited). Pass -1 to get the full document. |
extractTables |
true | Extract tables from docs/spreadsheets |
extractImages |
false | Enable OCR for images |
ocrLang |
"eng" | OCR language for tesseract.js (e.g. "eng", "fra", "ara") |
maxPages |
varies | Limit pages/slides/sheets processed |
save |
false | Save the full parsed output as a .md file alongside the original (no truncation) |
outputPath |
— | Custom path for the Markdown export (overrides save path) |
How it works
- File is verified by magic bytes, not just extension
- Type detection dispatches to the right parser
- Metadata is extracted (author, pages, sheet count, etc.)
- Tables become readable markdown
- Large content is truncated gracefully with a note to the LLM
All 15+ format handlers return the same output structure, so the LLM gets consistent results regardless of file type.
Development
npm install
npm run typecheck
License
MIT
Similar plugins
Openclaude Memory
@openlines/openclaude-memory
OPEN. CONFIGURABLE. Global persistent memory for opencode sessions via MEMORY.md injection (like our mother Claude Code).
Permission Reviewer
opencode-permission-reviewer
Policy-aware permission reviewer for OpenCode V1 and V2
Jev Router
@robertn702/opencode-jev-router
Adaptive reasoning effort for OpenCode V2 with request-local GPT-6 model selection