Vibeleaderboard
Index / tool
Visit github.com
Category
Developer Tools
Rank

Previous survey · No. 805 ·

Pricing
Open Source
Type
TOOL
Use case
Models: Train & Run · Model & Agent Evaluation
Interfaces
CLI
Latest release
v0.75.1
Date

About

Soup is a CLI for fine-tuning and post-training LLMs from a single YAML config, using layer streaming to keep the frozen base model out of VRAM so an 8B model can train on a 4GB laptop GPU. It also includes a 'soup ship' regression gate that benchmarks model checkpoints against suites like tool-calling, MMLU, and refusal-rate to catch regressions before deployment.

What it does

Soup is a command line tool that turns one configuration file into a complete training run: it chooses batch size and precision automatically, drives the job, and can hand the result off to export or a local server afterward. Its standout feature moves the unchanging half of a large model off the graphics card during training, feeding it back one section at a time, which is how a full-size model trains on hardware that would normally only fit something much smaller. A separate command grades a finished checkpoint against a battery of tests before anyone decides to release it.

Why it's ranked here

Engineering care carries this more than any single novel idea: the same tool handles data cleaning, quantized training, evaluation gating, multi-format export and deployment, with every heavy capability split into its own optional install so a casual user never pulls down a full GPU toolchain by accident. The dependency manifest reads like it was written by people who got burned before, with real compatibility breakages named rather than guessed at. The headline memory-saving trick is honestly labeled experimental and opt-in rather than sold as the default, which is the kind of restraint that is easy to skip and rare to see.

What's good

The install layering is thoughtful: the base package carries no heavy machine-learning dependencies at all, and training, serving, the web dashboard and the agent-facing server are each their own add-on. That agent-facing server takes security seriously for something this niche: it checks its access token in constant time, blocks browser-driven requests to the local port with strict host checking, and scrubs every error message so a failure never leaks a stack trace or a file path back to the caller. The changelog's fixes read like they came from real, reproduced failures rather than hypothetical ones.

Tradeoffs

The command surface is enormous, spanning training, evaluation, deployment and governance, which makes it hard for an outside user to judge how deeply tested any one corner is compared with the well documented core loop. The project's own release notes admit a history of settings that were accepted, validated and then quietly ignored on one backend instead of erroring, a pattern only recently closed off. The headline low-memory training path is still marked experimental, and the maintainers themselves note that its last published benchmark predates a later correctness fix and has not been repeated since.

How to use it well

Reach for this when you want one tool to carry a model from a bare config file through training, evaluation and export without stitching together separate scripts, especially if you are experimenting on a small consumer graphics card and want the memory-saving path (treat it as experimental and confirm the numbers on your own hardware rather than trusting a published figure). Install only the light core if you need its data or configuration tooling and nothing more. Read the release notes before upgrading: this project ships real breaking changes to its configuration format, not just additive ones.

Technical notes+

The console entry point is registered in pyproject.toml as soup = soup_cli.cli:run, and src/soup_cli/cli.py wires dozens of typer subcommands and sub-apps (train, data, eval, serve, deploy, registry, mcp, and more) onto one Typer application. The Model Context Protocol integration lives entirely in src/soup_cli/mcp_server/server.py, the only module that imports the mcp SDK; build_server() introspects Server.__init__ for an on_list_tools parameter to branch between the SDK's decorator-based and callback-based APIs rather than parsing a version string. Its network transports run behind a bearer-auth ASGI middleware that compares the Authorization header with secrets.compare_digest, and behind an allowed_hosts_for()/_security_settings() pair that builds a DNS-rebinding host allowlist, expanding loopback names to their standard aliases and bracketing IPv6 literals for the Host header. _dispatch_tool() catches every handler exception and rewraps it as a generic type(exc).__name__ message, so no stack trace or filesystem path reaches a client. pyproject.toml keeps the base package free of the training stack, pins Python support to 3.10 through 3.12, and organizes optional capabilities (train, mlx, ui, serve, mcp, and others) as separate extras.

Observed

License
Apache-2.0 license, declared in the project manifest.
Packaging
Distributed as a Python package with a console-script entry point, installable via pip or pipx.
Install extras
The core install has no PyTorch dependency; the full training stack sits behind a separate optional extra.
Interfaces
Ships three distinct interfaces: a command-line tool, a local web dashboard, and a Model Context Protocol server.
MCP server
The Model Context Protocol server supports stdio transport locally and network transports (SSE, streamable HTTP); the network transports require bearer-token authentication and DNS-rebinding host checks.
Python support
Declared Python support is 3.10 through 3.12; newer Python majors are not declared as supported in the package manifest.
Install extras
Optional capabilities (training backends, serving, web UI, MCP server, Apple-Silicon support, evaluation tooling, and more) are each packaged as separate install extras.

Read from README.md, pyproject.toml, src/soup_cli/__init__.py, src/soup_cli/__main__.py, src/soup_cli/cli.py, src/soup_cli/mcp_server/server.py, src/soup_cli/mcp_server/__init__.py, docs/commands.md, docs/README.md, docs/backends-and-ops.md.

What it can do

  • Fine-tune and post-train LLMs from a YAML config

    YAML config file → Fine-tuned LLM

  • Stream model layers to keep frozen base model out of VRAM during training

    Base LLM weights → Reduced VRAM usage during training

  • Benchmark model checkpoints against regression test suites (tool-calling, MMLU, refusal-rate) before deployment

    Model checkpoint → Regression benchmark report

Tags

llmfine-tuningqloraloraclimlgpumlops

Tech Stack

PythonDocker

Comments (0)

No comments yet

Editorially curated, with community endorsements as a secondary signal. Corrections welcome.