Vibeleaderboard
Index / app
Visit github.com
Category
Developer Tools
Rank
Pricing
Open Source
Type
APP
Use case
Models: Train & Run
Interfaces
Desktop · Web · API
Latest release
v4.9
Date

About

An open-source desktop application for running large language models locally, offering a chat and notebook UI, multimodal (image/file) input, LoRA fine-tuning, and image generation via diffusers models. It exposes an OpenAI/Anthropic-compatible API with tool-calling and supports multiple inference backends (llama.cpp, Transformers, ExLlamaV3, TensorRT-LLM) with zero telemetry.

What it does

TextGen runs language models entirely on your own machine and puts a browser interface in front of them: a conversational chat screen, a free-write tab for generating text outside of a dialogue, and support for dropping in images or documents so the model can react to their contents. It can start several different engines to actually run the model depending on the file format you have, switching between them without restarting the app. The same conversation is also reachable through a network endpoint shaped like the interfaces ChatGPT and Claude use, so other software can talk to it the same way it would talk to a hosted service.

Why it's ranked here

The privacy claim in the README checks out in the code: nothing here phones home, and the app runs fully offline with no external resource calls. The real strength is breadth without a rewrite for each backend: one loader map drives five separate inference engines including llama.cpp, Transformers, ExLlamaV3, and TensorRT-LLM, and the same server exposes a chat-compatible API other tools can call like a hosted model. Getting there is not one command, though: the full install branches across hardware-specific requirement files and a multi-step Conda setup, which slows down anyone past the portable build.

What's good

The tool-calling flow is built with real caution: when a model tries to invoke a function, generation blocks and waits for an explicit approval before the call runs, with an option to remember that decision for the rest of the session. Prompt formatting for different model families runs through a sandboxed templating engine instead of special-cased string building, so a new model's chat format is one template away, not a code change. An extension system exposes hooks into input, output, conversation state, and even the token sampler, letting add-ons change behavior without touching the core.

Tradeoffs

The command-line surface is large: dozens of flags across cache types, speculative decoding, GPU splitting, and per-backend quantization, which is powerful but demands real configuration knowledge to get a given model running well. The full installation path pulls in PyTorch and roughly 10GB of dependencies, and you pick the matching requirements file for your OS and GPU vendor by hand. Generation is serialized behind a single lock for most backends, so one conversation blocks another unless parallel request slots are explicitly enabled.

How to use it well

This fits someone who wants a private chat interface backed by their own downloaded model, plus a drop-in replacement for a hosted LLM API inside their own scripts or agents. Grab the portable build if GGUF models through the default engine are enough; move to the full installation only when training, image generation, or one of the other backends is actually needed, since those pull in the heavier dependency set. Multi-user mode turns off saved chat history, so treat it as a shared, stateless deployment rather than a personal one.

Technical notes+

server.py is the entry point: create_interface() builds the Gradio app and starts the OpenAI/Anthropic-compatible API server when --api is set. modules/models.py routes a chosen backend name through a load_func_map dictionary to one of five loader functions (llama_cpp_server_loader, transformers_loader, ExLlamav3_HF_loader, ExLlamav3_loader, TensorRT_LLM_loader), and modules/loaders.py defines which parameters each backend accepts. modules/chat.py implements request_tool_approval, a threading.Event-based block that holds generation until a human approves, rejects, or allows a tool call for the rest of the session, and renders prompts through an ImmutableSandboxedEnvironment Jinja2 instance rather than hand-written formatting per model. modules/extensions.py defines EXTENSION_MAP, the fixed set of hook points (input, output, state, history, tokenizer, logits_processor, custom_generate_reply) that third-party extensions can attach to.

Observed

Language
Python throughout the server and module code
Interface
Browser-based UI built on Gradio, launched from a local server process
API
OpenAI/Anthropic-compatible chat and completions endpoints, enabled with an --api flag
Tool integration
Custom Python-file tools plus MCP stdio server support
Inference backends
llama.cpp, Transformers, ExLlamaV3 (HF and native), and TensorRT-LLM, selected through a single loader map
Packaging
Portable prebuilt archives, manual pip/venv install, Conda environment setup, and Docker Compose files for NVIDIA, AMD, Intel, and CPU-only
Platform support
Linux, Windows, and macOS, with CUDA, Vulkan, ROCm, and CPU-only build variants
Extensibility
A fixed extension hook map covering input, output, state, history, tokenizer, and logits-processor modification points
Telemetry
README states no analytics, external resource calls, or remote update requests

Read from README.md, server.py, modules/text_generation.py, modules/models.py, modules/loaders.py, modules/chat.py, modules/extensions.py, modules/ui.py, modules/shared.py.

What it can do

  • Run large language models locally for chat conversations

    text prompt → chat response

  • Process multimodal inputs including images and files

    image or file → model response

  • Fine-tune language models using LoRA

    training data → fine-tuned LoRA model

  • Generate images using diffusers models

    text prompt → generated image

  • Expose an OpenAI/Anthropic-compatible API with tool-calling support

    API request → model response with tool calls

  • Run inference using multiple backend engines (llama.cpp, Transformers, ExLlamaV3, TensorRT-LLM)

    model files → generated text

Tags

local-llmllama.cppself-hostedchatbotopen-sourceopenai-apilora-traininggguf

Comments (0)

No comments yet

Editorially curated, with community endorsements as a secondary signal. Corrections welcome.