lemonade-sdk

lemonade

Built by lemonade-sdk 2,605 stars

What is lemonade?

Lemonade helps users discover and run local AI apps by serving optimized LLMs right from their own GPUs and NPUs. Join our discord: https://discord.gg/5xXzkMu8Zk

How to use lemonade?

1. Install a compatible MCP client (like Claude Desktop). 2. Open your configuration settings. 3. Add lemonade using the following command: npx @modelcontextprotocol/lemonade 4. Restart the client and verify the new tools are active.
🛡️ Scoped (Restricted)
npx @modelcontextprotocol/lemonade --scope restricted
🔓 Unrestricted Access
npx @modelcontextprotocol/lemonade

Key Features

Native MCP Protocol Support
Real-time Tool Activation & Execution
Verified High-performance Implementation
Secure Resource & Context Handling

Optimized Use Cases

Extending AI models with custom local capabilities
Automating system workflows via natural language
Connecting external data sources to LLM context windows

lemonade FAQ

Q

Is lemonade safe?

Yes, lemonade follows the standardized Model Context Protocol security patterns and only executes tools with explicit user-granted permissions.

Q

Is lemonade up to date?

lemonade is currently active in the registry with 2,605 stars on GitHub, indicating its reliability and community support.

Q

Are there any limits for lemonade?

Usage limits depend on the specific implementation of the MCP server and your system resources. Refer to the official documentation below for technical details.

Official Documentation

View on GitHub

🍋 Lemonade: Refreshingly fast local AI

<p align="center"> <a href="https://discord.gg/5xXzkMu8Zk"> <img src="https://img.shields.io/badge/Discord-7289DA?logo=discord&logoColor=white" alt="Discord" /></a> <a href="https://github.com/lemonade-sdk/lemonade/blob/main/docs/dev/contribute.md" title="Contribution Guide"> <img src="https://img.shields.io/badge/PRs-welcome-brightgreen.svg" alt="PRs Welcome" /></a> <a href="https://github.com/lemonade-sdk/lemonade/releases/latest" title="Download the latest release"> <img src="https://img.shields.io/github/v/release/lemonade-sdk/lemonade?include_prereleases" alt="Latest Release" /></a> <a href="https://tooomm.github.io/github-release-stats/?username=lemonade-sdk&repository=lemonade"> <img src="https://img.shields.io/github/downloads/lemonade-sdk/lemonade/total.svg" alt="GitHub downloads" /></a> <a href="https://github.com/lemonade-sdk/lemonade/issues"> <img src="https://img.shields.io/github/issues/lemonade-sdk/lemonade" alt="GitHub issues" /></a> <a href="https://github.com/lemonade-sdk/lemonade/blob/main/LICENSE"> <img src="https://img.shields.io/badge/License-Apache-yellow.svg" alt="License: Apache" /></a> <a href="https://star-history.com/#lemonade-sdk/lemonade"> <img src="https://img.shields.io/badge/Star%20History-View-brightgreen" alt="Star History Chart" /></a> </p> <p align="center"> <img src="https://github.com/lemonade-sdk/assets/blob/main/docs/banner_02.png?raw=true" alt="Lemonade Banner" /> </p> <h3 align="center"> <a href="https://lemonade-server.ai/install_options.html">Download</a> | <a href="https://lemonade-server.ai/docs/">Documentation</a> | <a href="https://discord.gg/5xXzkMu8Zk">Discord</a> </h3>

Lemonade is the local AI server that gives you the same capabilities as cloud APIs, except 100% free and private. Use the latest models for chat, coding, speech, and image generation on your own NPU and GPU.

Lemonade comes in two flavors:

  • Lemonade Server installs a service you can connect to hundreds of great apps using standard OpenAI, Anthropic, and Ollama APIs.
  • Embeddable Lemonade is a portable binary you can package into your own application to give it multi-modal local AI that auto-optimizes for your user’s PC.

This project is built by the community for every PC, with optimizations by AMD engineers to get the most from Ryzen AI, Radeon, and Strix Halo PCs.

Getting Started

  1. Install: Windows · Linux · macOS · Docker · Source
  2. Get Models: Browse and download with the Model Manager
  3. Generate: Try models with the built-in interfaces for chat, image gen, speech gen, and more
  4. Mobile: Take your lemonade to go: iOS · Android · Source
  5. Connect: Use Lemonade with your favorite apps:
<!-- MARKETPLACE_START --> <p align="center"> <a href="https://lemonade-server.ai/docs/server/apps/claude-code/" title="Claude Code"><img src="https://raw.githubusercontent.com/lemonade-sdk/marketplace/main/apps/claude-code/logo.png" alt="Claude Code" width="60" /></a>&nbsp;&nbsp;<a href="https://quickthoughts.ca/posts/firefox-chatback-lemonade-sdk/" title="Firefox Chatbot"><img src="https://raw.githubusercontent.com/lemonade-sdk/marketplace/main/apps/fx-chatbot/logo.png" alt="Firefox Chatbot" width="60" /></a>&nbsp;&nbsp;<a href="https://lemonade-server.ai/docs/server/apps/anythingLLM/" title="AnythingLLM"><img src="https://raw.githubusercontent.com/lemonade-sdk/marketplace/main/apps/anythingllm/logo.png" alt="AnythingLLM" width="60" /></a>&nbsp;&nbsp;<a href="https://marketplace.dify.ai/plugins/langgenius/lemonade" title="Dify"><img src="https://raw.githubusercontent.com/lemonade-sdk/marketplace/main/apps/dify/logo.png" alt="Dify" width="60" /></a>&nbsp;&nbsp;<a href="https://github.com/amd/gaia?tab=readme-ov-file#getting-started-guide" title="GAIA"><img src="https://raw.githubusercontent.com/lemonade-sdk/marketplace/main/apps/gaia/logo.png" alt="GAIA" width="60" /></a>&nbsp;&nbsp;<a href="https://admcpr.com/local-github-copilot-with-lemonade-server-on-windows" title="GitHub Copilot"><img src="https://raw.githubusercontent.com/lemonade-sdk/marketplace/main/apps/github-copilot/logo.png" alt="GitHub Copilot" width="60" /></a>&nbsp;&nbsp;<a href="https://github.com/lemonade-sdk/infinity-arcade" title="Infinity Arcade"><img src="https://raw.githubusercontent.com/lemonade-sdk/marketplace/main/apps/infinity-arcade/logo.png" alt="Infinity Arcade" width="60" /></a>&nbsp;&nbsp;<a href="https://n8n.io/integrations/lemonade-model/" title="n8n"><img src="https://raw.githubusercontent.com/lemonade-sdk/marketplace/main/apps/n8n/logo.png" alt="n8n" width="60" /></a>&nbsp;&nbsp;<a href="https://lemonade-server.ai/docs/server/apps/open-webui/" title="Open WebUI"><img src="https://raw.githubusercontent.com/lemonade-sdk/marketplace/main/apps/open-webui/logo.png" alt="Open WebUI" width="60" /></a>&nbsp;&nbsp;<a href="https://lemonade-server.ai/docs/server/apps/open-hands/" title="OpenHands"><img src="https://raw.githubusercontent.com/lemonade-sdk/marketplace/main/apps/openhands/logo.png" alt="OpenHands" width="60" /></a> </p> <p align="center"><em>Want your app featured here? <a href="https://github.com/lemonade-sdk/marketplace">Just submit a marketplace PR!</a></em></p> <!-- MARKETPLACE_END -->

Supported Platforms

PlatformBuild
Arch LinuxBuild on Arch
Debian Trixie+Build on Debian
DockerBuild Container Image
Fedora 43+Build .rpm
macOSBuild .pkg
SnapBuild Snap
Ubuntu 24.04+Build Launchpad PPA
Windows 11Build .msi

Using the CLI

To run and chat with Gemma:

lemonade run Gemma-4-E2B-it-GGUF

To code with Lemonade models:

lemonade launch claude

Multi-modality:

# image gen
lemonade run SDXL-Turbo

# speech gen
lemonade run kokoro-v1

# transcription
lemonade run Whisper-Large-v3-Turbo

To see available models and download them:

lemonade list

lemonade pull Gemma-4-E2B-it-GGUF

To see the backends available on your PC:

lemonade backends

For hybrid setups, Lemonade can also route to any OpenAI-compatible cloud provider (Fireworks, OpenAI, OpenRouter, Together, …) alongside local models — see Cloud Offload. (Experimental.)

Model Library

<img align="right" src="https://github.com/lemonade-sdk/assets/blob/main/docs/model_manager_02.png?raw=true" alt="Model Manager" width="280" />

Lemonade supports a wide variety of LLMs (GGUF, FLM, and ONNX), whisper, stable diffusion, etc. models across CPU, GPU, and NPU.

Use lemonade pull or the built-in Model Manager to download models. Custom GGUF/ONNX models can be pulled from Hugging Face or ModelScope, with their source retained for future updates.

Browse all built-in models →

<br clear="right"/>

Supported Configurations

Lemonade supports multiple inference engines for LLM, speech, TTS, and image generation, and each has its own backend and hardware requirements.

<!-- BEGIN GENERATED: backends-matrix --> <table> <thead> <tr> <th>Modality</th> <th>Engine</th> <th>Backend</th> <th>Device</th> <th>OS</th> </tr> </thead> <tbody> <tr> <td rowspan="9"><strong>Text generation</strong></td> <td rowspan="6"><code>llamacpp</code></td> <td><code>system</code></td> <td><code>x86_64</code>/ARM64 CPU, GPU</td> <td>Linux</td> </tr> <tr> <td><code>metal</code></td> <td>Apple Silicon GPU</td> <td>macOS</td> </tr> <tr> <td><code>cuda</code></td> <td>NVIDIA GPUs (Turing or newer)**</td> <td>Windows, Linux</td> </tr> <tr> <td><code>vulkan</code></td> <td><code>x86_64</code> CPU, AMD iGPU, AMD dGPU; ARM64 CPU/GPU (Linux)</td> <td>Windows, Linux</td> </tr> <tr> <td><code>rocm</code></td> <td>Supported AMD ROCm iGPU/dGPU families, incl. AMD Instinct MI300X (gfx942) and MI350X (gfx950, Linux + stable only)*</td> <td>Windows, Linux</td> </tr> <tr> <td><code>cpu</code></td> <td><code>x86_64</code> CPU; ARM64 CPU (Linux)</td> <td>Windows, Linux</td> </tr> <tr> <td rowspan="1"><code>flm</code></td> <td><code>npu</code></td> <td>XDNA2 NPU</td> <td>Windows, Linux</td> </tr> <tr> <td rowspan="1"><code>ryzenai-llm</code></td> <td><code>npu</code></td> <td>XDNA2 NPU</td> <td>Windows</td> </tr> <tr> <td rowspan="1"><code>vllm</code> (experimental)</td> <td><code>rocm</code></td> <td>Strix Halo iGPU (gfx1151)</td> <td>Linux</td> </tr> <tr> <td rowspan="6"><strong>Speech-to-text</strong></td> <td rowspan="5"><code>whispercpp</code></td> <td><code>npu</code></td> <td>XDNA2 NPU</td> <td>Windows</td> </tr> <tr> <td><code>rocm</code></td> <td>Supported AMD ROCm iGPU/dGPU families*</td> <td>Windows, Linux</td> </tr> <tr> <td><code>vulkan</code></td> <td><code>x86_64</code> CPU</td> <td>Windows, Linux</td> </tr> <tr> <td><code>cpu</code></td> <td><code>x86_64</code> CPU</td> <td>Windows, Linux</td> </tr> <tr> <td><code>metal</code></td> <td>Apple Silicon GPU</td> <td>macOS</td> </tr> <tr> <td rowspan="1"><code>moonshine</code></td> <td><code>cpu</code></td> <td><code>x86_64</code>/<code>arm64</code> CPU</td> <td>Windows, Linux, macOS</td> </tr> <tr> <td rowspan="5"><strong>Text-to-speech</strong></td> <td rowspan="2"><code>kokoro</code></td> <td><code>cpu</code></td> <td><code>x86_64</code> CPU</td> <td>Windows, Linux</td> </tr> <tr> <td><code>metal</code></td> <td>Apple Silicon GPU</td> <td>macOS</td> </tr> <tr> <td rowspan="3"><code>openmoss</code> (experimental)</td> <td><code>rocm</code></td> <td>AMD GPUs (ROCm via TheRock)</td> <td>Windows, Linux</td> </tr> <tr> <td><code>cuda</code></td> <td>NVIDIA GPUs</td> <td>Windows, Linux</td> </tr> <tr> <td><code>vulkan</code></td> <td>Vulkan-capable GPUs</td> <td>Windows, Linux</td> </tr> <tr> <td rowspan="6"><strong>Audio generation</strong></td> <td rowspan="3"><code>thinksound</code> (experimental)</td> <td><code>rocm</code></td> <td>Supported AMD ROCm iGPU/dGPU families (ROCm via TheRock)</td> <td>Windows, Linux</td> </tr> <tr> <td><code>cuda</code></td> <td>NVIDIA GPUs</td> <td>Windows, Linux</td> </tr> <tr> <td><code>vulkan</code></td> <td>Vulkan-capable GPUs</td> <td>Windows, Linux</td> </tr> <tr> <td rowspan="3"><code>acestep</code> (experimental)</td> <td><code>rocm</code></td> <td>Supported AMD ROCm iGPU/dGPU families (ROCm via TheRock)</td> <td>Windows, Linux</td> </tr> <tr> <td><code>cuda</code></td> <td>NVIDIA GPUs</td> <td>Windows, Linux</td> </tr> <tr> <td><code>vulkan</code></td> <td>Vulkan-capable GPUs</td> <td>Windows, Linux</td> </tr> <tr> <td rowspan="5"><strong>Image generation</strong></td> <td rowspan="5"><code>sd-cpp</code></td> <td><code>rocm</code></td> <td>Supported AMD ROCm iGPU/dGPU families*</td> <td>Windows, Linux</td> </tr> <tr> <td><code>cuda</code></td> <td>NVIDIA GPUs (Turing or newer)**</td> <td>Windows, Linux</td> </tr> <tr> <td><code>vulkan</code></td> <td>Vulkan-capable GPUs</td> <td>Windows, Linux</td> </tr> <tr> <td><code>cpu</code></td> <td><code>x86_64</code> CPU</td> <td>Windows, Linux</td> </tr> <tr> <td><code>metal</code></td> <td>Apple Silicon GPU</td> <td>macOS</td> </tr> <tr> <td rowspan="3"><strong>3D generation</strong></td> <td rowspan="3"><code>trellis</code> (experimental)</td> <td><code>rocm</code></td> <td>Supported AMD ROCm iGPU/dGPU families (ROCm via TheRock)</td> <td>Windows, Linux</td> </tr> <tr> <td><code>cuda</code></td> <td>NVIDIA GPUs</td> <td>Windows, Linux</td> </tr> <tr> <td><code>vulkan</code></td> <td>Vulkan-capable GPUs</td> <td>Windows, Linux</td> </tr> <tr> <td rowspan="3"><strong>Text classification</strong></td> <td rowspan="3"><code>onnxruntime</code> (experimental)</td> <td><code>cpu</code></td> <td><code>x86_64</code> CPU</td> <td>Windows</td> </tr> <tr> <td><code>cpu</code></td> <td><code>x86_64</code>/<code>arm64</code> CPU</td> <td>Linux</td> </tr> <tr> <td><code>cpu</code></td> <td><code>arm64</code> CPU</td> <td>macOS</td> </tr> </tbody> </table> <!-- END GENERATED: backends-matrix -->

To check exactly which recipes/backends are supported on your own machine, run:

lemonade backends
<details> <summary><small><i>* See supported AMD ROCm platforms</i></small></summary> <br> <table> <thead> <tr> <th>Architecture</th> <th>Platform Support</th> <th>GPU Models</th> </tr> </thead> <tbody> <tr> <td><b>gfx1151</b> (STX Halo)</td> <td>Windows, Ubuntu</td> <td>Ryzen AI MAX+ Pro 395</td> </tr> <tr> <td><b>gfx120X</b> (RDNA4)</td> <td>Windows, Ubuntu</td> <td>Radeon AI PRO R9700, RX 9070 XT/GRE/9070, RX 9060 XT</td> </tr> <tr> <td><b>gfx110X</b> (RDNA3)</td> <td>Windows, Ubuntu</td> <td>Radeon PRO W7900/W7800/W7700/V710, RX 7900 XTX/XT/GRE, RX 7800 XT, RX 7700 XT</td> </tr> </tbody> </table> </details> <details> <summary><small><i>** See supported NVIDIA CUDA platforms</i></small></summary> <br> <table> <thead> <tr> <th>Compute Capability</th> <th>Architecture</th> <th>GPU Models</th> </tr> </thead> <tbody> <tr> <td><b>sm_75</b></td> <td>Turing</td> <td>RTX 20-series, GTX 16-series, T4</td> </tr> <tr> <td><b>sm_80</b> / <b>sm_86</b></td> <td>Ampere</td> <td>RTX 30-series, A100, A40</td> </tr> <tr> <td><b>sm_89</b></td> <td>Ada Lovelace</td> <td>RTX 40-series, L40, L4</td> </tr> <tr> <td><b>sm_90</b></td> <td>Hopper</td> <td>H100, H200</td> </tr> <tr> <td><b>sm_100</b> / <b>sm_120</b></td> <td>Blackwell</td> <td>RTX 50-series, B100, B200</td> </tr> </tbody> </table> </details>

Project Roadmap

Lemonade's roadmap is defined by a set of working groups. Visit the landing page here to learn each group's goal and roadmap.

Integrate Embeddable Lemonade in Your Application

Embeddable Lemonade is a binary version of Lemonade that you can bundle into your own app to give it a portable, auto-optimizing, multi-modal local AI stack. This lets users focus on your app, with zero Lemonade installers, branding, or telemetry.

Check out the Embeddable Lemonade guide.

Connect Lemonade Server to Your Application

You can use any OpenAI-compatible client library by configuring it to use http://localhost:13305/v1 as the base URL. A table containing official and popular OpenAI clients on different languages is shown below.

Feel free to pick and choose your preferred language.

PythonC++JavaC#Node.jsGoRubyRustPHP
openai-pythonopenai-cppopenai-javaopenai-dotnetopenai-nodego-openairuby-openaiasync-openaiopenai-php

Python Client Example

from openai import OpenAI

# Initialize the client to use Lemonade Server
client = OpenAI(
    base_url="http://localhost:13305/api/v1",
    api_key="lemonade"  # required but unused
)

# Create a chat completion
completion = client.chat.completions.create(
    model="Gemma-4-E2B-it-GGUF",  # or any other available model
    messages=[
        {"role": "user", "content": "What is the capital of France?"}
    ]
)

# Print the response
print(completion.choices[0].message.content)

Click to learn more about the available APIs and how to embed Lemonade in your own application.

FAQ

To read our frequently asked questions, see our FAQ Guide

Contributing

Lemonade is built by the local AI community! If you would like to contribute to this project, please check out our contribution guide.

Maintainers

This is a community project maintained by @amd-pworfolk @bitgamma @danielholanda @jeremyfowers @kenvandine @Geramy @ramkrishna2910 @sawansri @siavashhub @sofiageo @superm1 @vgodsoe, and sponsored by AMD. You can reach us by filing an issue, emailing lemonade@amd.com, or joining our Discord.

Code Signing Policy

Free code signing provided by SignPath.io, certificate by SignPath Foundation.

Privacy policy: This program will not transfer any information to other networked systems unless specifically requested by the user or the person installing or operating it. When the user requests a model download or registry lookup, Lemonade may contact Hugging Face Hub (see their privacy policy) or ModelScope, according to the model source selected by the user or packager.

License and Attribution

This project is:

<!--This file was originally licensed under Apache 2.0. It has been modified. Modifications Copyright (c) 2025 AMD-->

Global Ranking

-
Trust ScoreMCPHub Index

Based on codebase health & activity.

Manual Config

{ "mcpServers": { "lemonade": { "command": "npx", "args": ["lemonade"] } } }