MCPHub LabRegistrykreuzberg-dev/kreuzberg
kreuzberg-dev

kreuzberg dev/kreuzberg

Built by kreuzberg-dev β€’ 7,139 stars

What is kreuzberg dev/kreuzberg?

A polyglot document intelligence framework with a Rust core. Extract text, metadata, and structured information from PDFs, Office documents, images, and 88+ formats. Available for Rust, Python, Ruby,

How to use kreuzberg dev/kreuzberg?

1. Install a compatible MCP client (like Claude Desktop). 2. Open your configuration settings. 3. Add kreuzberg dev/kreuzberg using the following command: npx @modelcontextprotocol/kreuzberg-dev-kreuzberg 4. Restart the client and verify the new tools are active.
πŸ›‘οΈ Scoped (Restricted)
npx @modelcontextprotocol/kreuzberg-dev-kreuzberg --scope restricted
πŸ”“ Unrestricted Access
npx @modelcontextprotocol/kreuzberg-dev-kreuzberg

Key Features

Native MCP Protocol Support
Real-time Tool Activation & Execution
Verified High-performance Implementation
Secure Resource & Context Handling

Optimized Use Cases

Extending AI models with custom local capabilities
Automating system workflows via natural language
Connecting external data sources to LLM context windows

kreuzberg dev/kreuzberg FAQ

Q

Is kreuzberg dev/kreuzberg safe?

Yes, kreuzberg dev/kreuzberg follows the standardized Model Context Protocol security patterns and only executes tools with explicit user-granted permissions.

Q

Is kreuzberg dev/kreuzberg up to date?

kreuzberg dev/kreuzberg is currently active in the registry with 7,139 stars on GitHub, indicating its reliability and community support.

Q

Are there any limits for kreuzberg dev/kreuzberg?

Usage limits depend on the specific implementation of the MCP server and your system resources. Refer to the official documentation below for technical details.

Official Documentation

View on GitHub

Xberg

<div align="center" style="display: flex; flex-wrap: wrap; gap: 8px; justify-content: center; margin: 20px 0;"> <a href="https://github.com/xberg-io/alef"> <img src="https://img.shields.io/badge/Bindings-alef%20%D7%90-007ec6" alt="Bindings"> </a> <!-- Language Bindings --> <a href="https://crates.io/crates/xberg"> <img src="https://img.shields.io/crates/v/xberg?label=Rust&color=007ec6" alt="Rust"> </a> <a href="https://pypi.org/project/xberg/"> <img src="https://img.shields.io/pypi/v/xberg?label=Python&color=007ec6" alt="Python"> </a> <a href="https://www.npmjs.com/package/@xberg-io/xberg"> <img src="https://img.shields.io/npm/v/@xberg-io/xberg?label=Node.js&color=007ec6" alt="Node.js"> </a> <a href="https://www.npmjs.com/package/@xberg-io/xberg-wasm"> <img src="https://img.shields.io/npm/v/@xberg-io/xberg-wasm?label=WASM&color=007ec6" alt="WASM"> </a> <a href="https://central.sonatype.com/artifact/io.xberg/xberg"> <img src="https://img.shields.io/maven-central/v/io.xberg/xberg?label=Java&color=007ec6" alt="Java"> </a> <a href="https://github.com/xberg-io/xberg/tree/main/packages/go"> <img src="https://img.shields.io/github/v/tag/xberg-io/xberg?label=Go&color=007ec6&filter=v1*" alt="Go"> </a> <a href="https://www.nuget.org/packages/Xberg/"> <img src="https://img.shields.io/nuget/v/Xberg?label=C%23&color=007ec6" alt="C#"> </a> <a href="https://packagist.org/packages/xberg-io/xberg"> <img src="https://img.shields.io/packagist/v/xberg-io/xberg?label=PHP&color=007ec6" alt="PHP"> </a> <a href="https://rubygems.org/gems/xberg"> <img src="https://img.shields.io/gem/v/xberg?label=Ruby&color=007ec6" alt="Ruby"> </a> <a href="https://hex.pm/packages/xberg"> <img src="https://img.shields.io/hexpm/v/xberg?label=Elixir&color=007ec6" alt="Elixir"> </a> <a href="https://pub.dev/packages/xberg"> <img src="https://img.shields.io/pub/v/xberg?label=Dart&color=007ec6" alt="Dart"> </a> <a href="https://central.sonatype.com/artifact/io.xberg/xberg-android"> <img src="https://img.shields.io/maven-central/v/io.xberg/xberg-android?label=Kotlin&color=007ec6" alt="Kotlin"> </a> <a href="https://github.com/xberg-io/xberg/tree/main/packages/swift"> <img src="https://img.shields.io/badge/Swift-SPM-007ec6" alt="Swift"> </a> <a href="https://github.com/xberg-io/xberg/tree/main/packages/zig"> <img src="https://img.shields.io/badge/Zig-package-007ec6" alt="Zig"> </a> <a href="https://github.com/xberg-io/xberg/releases"> <img src="https://img.shields.io/badge/C-FFI-007ec6" alt="C FFI"> </a> <a href="https://github.com/xberg-io/xberg/pkgs/container/xberg"> <img src="https://img.shields.io/badge/Docker-ghcr.io-007ec6?logo=docker&logoColor=white" alt="Docker"> </a> <!-- Project Info --> <a href="https://github.com/xberg-io/xberg/blob/main/LICENSE"> <img src="https://img.shields.io/badge/License-MIT-007ec6" alt="License"> </a> <a href="https://docs.xberg.io"> <img src="https://img.shields.io/badge/Docs-xberg-007ec6" alt="Documentation"> </a> <a href="https://huggingface.co/xberg-io"> <img src="https://img.shields.io/badge/Hugging%20Face-Xberg-007ec6" alt="Hugging Face"> </a> </div> <div align="center" style="display: flex; flex-wrap: wrap; gap: 12px; justify-content: center; margin: 28px 0 24px;"> <a href="https://discord.gg/xt9WY3GnKR"> <img height="22" src="https://img.shields.io/badge/Discord-Chat-007ec6?logo=discord&logoColor=white" alt="Join Discord"> </a> <a href="https://docs.xberg.io/demo.html"> <img height="22" src="https://img.shields.io/badge/Live%20Demo-Open-007ec6?logo=webassembly&logoColor=white" alt="Live Demo"> </a> <a href="https://github.com/xberg-io/xberg/stargazers"> <img height="22" src="https://img.shields.io/github/stars/xberg-io/xberg?style=social" alt="GitHub Stars"> </a> </div>

Extract clean text, tables, and structured data from documents and code β€” no format detection, no OCR setup, no stitched-together libraries. One engine, 15 language bindings, runs anywhere.

Xberg is the next iteration of Kreuzberg. Same document-intelligence engine, rebuilt and rebranded under a fresh v1 line.

<div align="center">

Feed documents β†’ get clean text, tables, metadata, transcripts, code intelligence Β· Run it library, CLI, REST API, or MCP server Β· No GPU needed Β· Stream multi-GB files Β· Cache results.

Documents Β· Images Β· Spreadsheets Β· Email Β· Archives Β· Code Β· Audio Β· Video

crates.io npm PyPI License: MIT

Quick start Β· What you get Β· Capabilities Β· CLI Β· Docs

</div>
<!-- markdownlint-disable MD013 --> <p align="center"><img src="docs/assets/demos/extract.gif" alt="Extracting clean Markdown from a PDF in the CLI" width="820"></p> <p align="center"><em>Feed any documentβ€”get structured text. Extract, batch, stream, or crawl.</em></p> <!-- markdownlint-enable MD013 --> <div align="center"><sub><a href="#demos">See more ↓</a></sub></div>

What you get

Point Xberg at anything β€” a PDF, a spreadsheet, a scanned image, an audio file, a source tree β€” and get back clean, structured content you can use right away. One core does the format detection, reading, and extraction, so you don't assemble a pipeline yourself. Call it from Rust, Python, Node.js, Go, Java, C#, Ruby, PHP, Elixir, Dart, Swift, Zig, WASM, Kotlin, or C FFI, and run it as a library, CLI tool, REST API, or MCP server.

What it doesHow
Extract from 96 formatsPDFs, Office, images, HTML, email, archives, scientific publications, and code β€” intelligent MIME detection, streaming for large files.
6 output formatsPlain text, Markdown, Djot, HTML, JSON tree structure, or Structured (JSON with OCR metadata and bounding boxes).
Code intelligenceFunctions, classes, imports, symbols, docstrings from 306 programming languages. Syntax-aware chunking for RAG pipelines.
Crawl & recurseFollow URLs, extract documents from within documents (nested archives, embedded PDFs). Auto/Document/Crawl modes.
OCR on demandTesseract, PaddleOCR, Candle, or VLM backends β€” fallback chains, extensible via plugins. Confidence scores. Language auto-detection.
TranscriptionWhisper ONNX for audio/video tracks (MP3, M4A, WAV, WebM, MP4).
Embeddings & searchLocal (ONNX models) or provider-hosted (OpenAI, Anthropic, Google, 143 providers via liter-llm). Reranking.
Structured outputsLLM-powered extraction β€” local (Ollama, LM Studio, vLLM) or remote (OpenAI, Anthropic, Google).
EnrichmentNER, redaction, summarization, translation, QR code detection, page classification, keyword extraction (YAKE/RAKE), language detection, layout detection, table extraction, token reduction (TOON).
Batch & parallelProcess 100s of documents in parallel. Per-file timeouts. Configurable batch concurrency (max_concurrent_extractions).
CachingContent-hash cache keys β€” skip re-extraction when the file and config are unchanged.
DeploymentLibrary, CLI (12 commands), REST API (xberg serve), MCP server (9 tools, 3 prompts, 4 resources), Docker.

Demos

<!-- markdownlint-disable MD013 --> <p align="center"><img src="docs/assets/demos/cli.gif" alt="Xberg CLI: extract, batch, detect, formats, cache, serve, mcp" width="760"></p> <p align="center"><em>The CLI: 12 commands for extraction, caching, serving, and MCP.</em></p> <p align="center"><img src="docs/assets/demos/ocr.gif" alt="OCR from a scanned image with confidence scores and bounding boxes" width="820"></p> <p align="center"><em>OCR with confidence scores and bounding boxes. Switch backends without code changes.</em></p> <p align="center"><img src="docs/assets/demos/crawl.gif" alt="Crawling a website and extracting all linked documents" width="820"></p> <p align="center"><em>Web crawl: fetch a page, follow links, extract all documents recursively.</em></p> <p align="center"><img src="docs/assets/demos/mcp.gif" alt="MCP server integration with Claude Desktop showing extraction tools and prompts" width="820"></p> <p align="center"><em>MCP server: AI agents extract documents, detect formats, warm models, manage cache.</em></p> <p align="center"><img src="docs/assets/demos/serve.gif" alt="REST API: POST a document, get JSON extraction results with streaming support" width="820"></p> <p align="center"><em>REST API: stream large files, get JSON or Markdown, one endpoint for all formats.</em></p> <!-- markdownlint-enable MD013 -->

Installation

Language Packages

<details open> <summary><strong>Python</strong></summary>
pip install xberg

See Python README for full documentation.

</details> <details> <summary><strong>Node.js / TypeScript</strong></summary>
npm install @xberg-io/xberg

See Node.js README for full documentation.

</details> <details> <summary><strong>Rust</strong></summary>
cargo add xberg

See Rust README for full documentation.

</details> <details> <summary><strong>Go</strong></summary>
go get github.com/xberg-io/xberg

See Go README for full documentation.

</details> <details> <summary><strong>Java</strong></summary>

Available on Maven Central as io.xberg:xberg. See Java README for the dependency snippet.

</details> <details> <summary><strong>C#</strong></summary>
dotnet add package Xberg

See C# README for full documentation.

</details> <details> <summary><strong>Ruby</strong></summary>
gem install xberg

See Ruby README for full documentation.

</details> <details> <summary><strong>PHP</strong></summary>
composer require xberg-io/xberg

See PHP README for full documentation.

</details> <details> <summary><strong>Elixir</strong></summary>

Add {:xberg, "~> 1.0"} to your mix.exs dependencies. See Elixir README for full documentation.

</details> <details> <summary><strong>WebAssembly</strong></summary>
npm install @xberg-io/xberg-wasm

See WebAssembly README for full documentation.

</details> <details> <summary><strong>Kotlin (Android)</strong></summary>

Available on Maven Central as io.xberg:xberg-android. See Kotlin README for the dependency snippet.

</details> <details> <summary><strong>Swift</strong></summary>

Add via Swift Package Manager. See Swift README for full documentation.

</details> <details> <summary><strong>Dart / Flutter</strong></summary>
dart pub add xberg

See Dart README for full documentation.

</details> <details> <summary><strong>Zig</strong></summary>

Add via zig fetch. See Zig README for full documentation.

</details> <details> <summary><strong>C/C++ (FFI)</strong></summary>

Build from source as part of this workspace. See C (FFI) README for full documentation.

</details>

CLI & Deployment

<details> <summary><strong>CLI Tool</strong></summary>
brew install xberg-io/tap/xberg

12 commands: extract, batch, detect, formats, version, cache (stats/clear/manifest/warm), serve, mcp, api, embed, chunk, completions.

See CLI usage guide for detailed documentation.

</details> <details> <summary><strong>Docker</strong></summary>
docker pull ghcr.io/xberg-io/xberg:latest

Run in API, CLI, or MCP modes. See Docker guide for examples.

</details> <details> <summary><strong>REST API Server</strong></summary>
xberg serve --host 0.0.0.0 --port 8000

One POST endpoint handles all formats. Returns JSON or Markdown. Stream large files. See API server guide.

</details> <details> <summary><strong>MCP Server</strong></summary>
xberg mcp --transport stdio

9 tools (extract, extract_batch, detect_mime_type, cache_stats, list_formats, cache_clear, get_version, cache_manifest, cache_warm). 3 prompts (extract_document, extract_with_ocr, semantic_search). 4 resources (formats, models, OCR languages, embedding presets).

Add to Claude Desktop or Cursor:

{
  "mcpServers": {
    "xberg": { "command": "xberg", "args": ["mcp"] }
  }
}

See MCP integration guide.

</details>

AI Coding Assistants

Install the Xberg plugin from xberg-io/plugins. Ships extraction APIs, OCR backends, configuration, and language conventions.

<details open> <summary><strong>Claude Code</strong></summary>
/plugin marketplace add xberg-io/plugins
/plugin install xberg@xberg
</details> <details> <summary><strong>Codex CLI</strong></summary>
/plugins add https://github.com/xberg-io/plugins

Search for xberg and select Install Plugin.

</details> <details> <summary><strong>Cursor</strong></summary>

Settings β†’ Plugins β†’ Add from URL β†’ https://github.com/xberg-io/plugins, then select xberg.

</details> <details> <summary><strong>Gemini CLI</strong></summary>
gemini extensions install https://github.com/xberg-io/plugins
</details> <details> <summary><strong>Factory Droid</strong></summary>
droid plugin marketplace add https://github.com/xberg-io/plugins
droid plugin install xberg@xberg
</details> <details> <summary><strong>GitHub Copilot CLI</strong></summary>
copilot plugin marketplace add https://github.com/xberg-io/plugins
copilot plugin install xberg@xberg
</details> <details> <summary><strong>opencode</strong></summary>

Add to opencode.json:

{
  "$schema": "https://opencode.ai/config.json",
  "plugin": ["@xberg-io/opencode-xberg"]
}
</details>

Quick Start

Extract text from a document:

use xberg::{extract, ExtractInput, ExtractionConfig};

#[tokio::main]
async fn main() -> xberg::Result<()> {
    let config = ExtractionConfig::default();
    let output = extract(
        ExtractInput::from_uri("document.pdf"),
        &config
    ).await?;

    println!("{}", output.results[0].content);
    Ok(())
}

Common use cases β€” see Quick start guide for language-specific examples, OCR, batch processing, and API configuration.


Capabilities

<details> <summary><strong>Full feature list</strong></summary>

Supported File Formats (96)

96 file formats across 8 major categories with intelligent format detection and comprehensive metadata extraction.

Office Documents

CategoryFormatsCapabilities
Word Processing.docx, .docm, .doc, .dotx, .dotm, .dot, .odt, .pagesFull text, tables, images, metadata, styles
Spreadsheets.xlsx, .xlsm, .xlsb, .xls, .xla, .xlam, .xltm, .xltx, .xlt, .ods, .numbersSheet data, formulas, cell metadata, charts
Presentations.pptx, .pptm, .ppt, .ppsx, .potx, .potm, .pot, .keySlides, speaker notes, images, metadata
PDF.pdfText, tables, images, metadata, OCR support
eBooks.epub, .fb2Chapters, metadata, embedded resources
Database.dbfTable data extraction, field type support
Hangul.hwp, .hwpxKorean document format, text extraction

Images (OCR-Enabled)

CategoryFormatsFeatures
Raster.png, .jpg, .jpeg, .gif, .webp, .bmp, .tiff, .tifOCR, table detection, EXIF metadata, dimensions, color space
Advanced.jp2, .jpx, .jpm, .mj2, .jbig2, .jb2, .pnm, .pbm, .pgm, .ppmOCR via pure-Rust JPEG2000 decoder, JBIG2 support, table detection
HEIC family.heic, .heics, .heif, .avif, .avcsEXIF metadata, optional pixel decoding
Vector.svgDOM parsing, embedded text, graphics metadata

Audio & Video

CategoryFormatsFeatures
Audio.mp3, .mpga, .m4a, .wav, .webmWhisper transcription
Video audio track.mp4, .mpeg, .webmAudio-track transcription only

Web & Data

CategoryFormatsFeatures
Markup.html, .htm, .xhtml, .xml, .svgDOM parsing, metadata (Open Graph, Twitter Card), link extraction
Structured Data.json, .yaml, .yml, .toml, .csv, .tsvSchema detection, nested structures, validation
Text & Markdown.txt, .md, .markdown, .djot, .mdx, .rst, .org, .rtfCommonMark, GFM, Djot, MDX, reStructuredText, Org Mode

Email & Archives

CategoryFormatsFeatures
Email.eml, .msg, .pstHeaders, body (HTML/plain), attachments, threading
Archives.zip, .tar, .tgz, .gz, .7zFile listing, nested archives, metadata, recursive extraction

Academic & Scientific

CategoryFormatsFeatures
Citations.bib, .ris, .nbib, .enwStructured parsing: RIS, PubMed/MEDLINE, EndNote XML, BibTeX/BibLaTeX
Scientific.tex, .latex, .typ, .typst, .jats, .ipynbLaTeX, Typst, Jupyter notebooks, PubMed JATS
Publishing.fb2, .docbook, .dbk, .docbook4, .docbook5, .opmlFictionBook, DocBook XML, OPML outlines

Code Intelligence (306 Languages)

Extract structure from 306 programming languages via tree-sitter:

FeatureDescription
Structure ExtractionFunctions, classes, methods, structs, interfaces, enums
Import/Export AnalysisModule dependencies, re-exports, wildcard imports
Symbol ExtractionVariables, constants, type aliases, properties
Docstring ParsingGoogle, NumPy, Sphinx, JSDoc, RustDoc, and 10+ formats
Syntax-Aware ChunkingSplit code by semantic boundaries for RAG pipelines
DiagnosticsParse errors with line/column positions

Powered by tree-sitter-language-pack.

Output Formats (6)

FormatUse caseExample
PlainRaw text, no markup"Chapter 1\nIntroduction"
MarkdownReadable, structured, RAG-friendly"# Chapter 1\n## Introduction"
DjotModern lightweight markupSimilar to Markdown but stricter
HTMLStyled, browser-ready<h1>Chapter 1</h1>
JSONMachine-readable tree structureHierarchical sections with heading levels
StructuredOCR metadata, bounding boxesJSON with elements[] containing {text, bbox, confidence}

Deployment Modes

ModeCommandTransportUse case
Libraryxberg::extract()Async functionsEmbed in your application
CLIxberg extract document.pdf12 commandsScripts, batch jobs, CI/CD
REST APIxberg serveHTTP POSTMicroservice, serverless deployment
MCP Serverxberg mcpstdio or HTTPClaude, Cursor, IDE agents
Dockerdocker run ghcr.io/xberg-io/xbergAll modesContainer deployment

OCR Backends

  • Tesseract β€” Native C FFI (Linux/macOS/Windows) and WASM (browser)
  • PaddleOCR β€” ONNX Runtime, mobile-optimized models
  • Candle β€” Pure Rust, CPU-only, lightweight
  • VLM β€” GPT-4 Vision, Claude Vision, Gemini Vision, or 143 providers via liter-llm

Fallback chains. Extensible via plugin system.

Embeddings

Local (ONNX Runtime):

  • Preset models: fast, balanced (default), quality, multilingual
  • Dimensions: 384, 768, 1024

Provider-hosted:

  • OpenAI, Anthropic, Google, Hugging Face, Mistral, Cohere, and 143 providers total
  • Via liter-llm integration

Reranking:

  • Local ONNX rerankers (cross-encoder models)
  • Provider-hosted: Cohere Rerank, others

Structured LLM Extraction

Local engines: Ollama, LM Studio, vLLM

Remote: OpenAI, Anthropic, Google, Mistral, Cohere, and 143 providers via liter-llm

Schema validation. Temperature, top-p, frequency penalty tuning.

Enrichment

  • NER β€” GLiNER or LLM-based entity recognition
  • Redaction β€” Mask PII (phone, email, SSN, credit card, addresses)
  • Summarization β€” Document and section summaries via LLM
  • Translation β€” Multi-language via LLM
  • Page Classification β€” Tag document pages (cover, toc, content, etc.)
  • QR Code Detection β€” Extract and decode QR codes from images
  • Keyword Extraction β€” YAKE or RAKE algorithms
  • Language Detection β€” Detect document language
  • Layout Detection β€” RT-DETR + TATR models for document structure
  • Table Extraction β€” Cell-level structure and content
  • Token Reduction β€” TOON wire format (~30–50% fewer tokens than JSON)
</details>

CLI Reference

<details> <summary><strong>All 12 commands</strong></summary>
CommandSubcommandsPurpose
extractβ€”Extract text from a single document (path, URL, or stdin)
batchβ€”Extract from multiple documents in parallel
detectβ€”Identify MIME type of a file
formatsβ€”List all 96 supported formats and MIME types
versionβ€”Show Xberg version
cachestats, clear, manifest, warmManage extraction cache and models
serveβ€”Start REST API server (default: http://127.0.0.1:8000)
mcpβ€”Start MCP server (stdio or HTTP transport)
apischemaOutput OpenAPI 3.1 specification
embedβ€”Generate embeddings for text (local or provider-hosted)
chunkβ€”Split text into chunks (text, markdown, YAML, or semantic)
completionsβ€”Generate shell completion scripts

Run xberg --help or xberg <command> --help for detailed options.

</details>

Documentation

Full guides, API references for every binding, format reference, and configuration docs live at xberg.io.


Contributing

Contributions are welcome! See CONTRIBUTING.md for guidelines.

Join our Discord community for questions and discussion.


Part of Xberg.dev

Xberg is one of six open-source projects from Kreuzberg, Inc.:

  • Xberg β€” document intelligence: text, tables, metadata from 91+ formats with optional OCR.
  • Xberg Enterprise β€” managed extraction API with SDKs, dashboards, and observability.
  • crawlberg β€” web crawling and scraping with HTMLβ†’Markdown and headless-Chrome fallback.
  • html-to-markdown β€” fast, lossless HTMLβ†’Markdown engine.
  • liter-llm β€” universal LLM API client with native bindings for 14 languages and 143 providers.
  • tree-sitter-language-pack β€” tree-sitter grammars and code-intelligence primitives.
  • alef β€” the polyglot binding generator that produces every per-language binding across the 5 polyglot repos.

License

MIT License (MIT) β€” see LICENSE for details.

Global Ranking

8.5
Trust ScoreMCPHub Index

Based on codebase health & activity.

Manual Config

{ "mcpServers": { "kreuzberg-dev-kreuzberg": { "command": "npx", "args": ["kreuzberg-dev-kreuzberg"] } } }