Vibeleaderboard
Index / tool
Visit github.com
Category
AI Tools
Rank
No. 1984Tools index

Previous survey · No. 1989 ·

Pricing
Open Source
Type
TOOL
Use case
Models: Train & Run · Model & Agent Evaluation
Interfaces
CLI
Latest release
v1.1.16
Date

About

A terminal tool that works out which local LLMs will actually run well on your machine. It detects CPU, GPU and RAM, then scores each model on quality, speed, fit and context, handling multi-GPU setups, mixture-of-experts architectures and quantization choices, and reads from local runtimes including Ollama, llama.cpp, MLX, LM Studio and Docker Model Runner. It also measures real tokens per second on your hardware and lets you contribute those numbers back, so estimates get replaced by measurements from identical machines.

What it does

Llmfit answers a narrow question before you download forty gigabytes: will this model actually load and run at a usable speed here? It reads your memory, cores and graphics cards, then walks each catalogued model down a ladder of compression levels until one fits, halving the context window if nothing does. Every candidate gets one of four fit verdicts based on how full the memory pool would be, plus a speed estimate derived from the card's memory bandwidth. Your own benchmark runs, and matching community ones, replace the formula guesses when they exist.

Why it's ranked here

The estimates show their working. Each speed figure carries a confidence label saying whether it was measured on this machine, measured by others on matching hardware, calibrated, or pure formula, and the inputs behind it are exposed so a number can be checked. The fit bands are explicit (60, 85 and 98 percent pool utilization), mixture-of-experts models are sized by their active parameters rather than their total, and paths that spill off the graphics card can never earn the top verdict. Add an MIT license, a single Rust binary, and installs through Homebrew, Scoop, MacPorts, pip and Docker, and it is an easy tool to trust and to adopt.

What's good

Hardware detection is broad and specific about its limits. It covers NVIDIA, AMD, Intel Arc, Apple Silicon and Ascend accelerators, sums memory across multiple cards, and corrects for AMD laptop chips whose BIOS hides part of the RAM from the operating system. It talks to five local runtimes (Ollama, llama.cpp, MLX, Docker Model Runner, LM Studio), marks what you already have installed, and can pull a model from the terminal interface. Machine-readable JSON output on every command, documented exit codes and an HTTP endpoint make it usable from scripts and agents, not only by hand. A profile flag scores models against hardware you do not own.

Tradeoffs

Most numbers are estimates until someone benchmarks. Speed comes from a bandwidth formula with a fixed 0.55 efficiency factor, and prompt processing time is left blank unless the card's compute throughput is known. The model catalogue is compiled into the binary, so new releases arrive only when you upgrade the tool. The quality score leans partly on a model family's reputation. Detection is thinner off Linux and Apple Silicon: Windows and Intel Macs see NVIDIA cards only through its own utility, and Android phones get no GPU detection at all. Windows release binaries can ship unsigned if the signing job fails.

How to use it well

Run it before choosing a model for a laptop, workstation or home server, and treat the Too Tight and Marginal verdicts as the useful output: they save a doomed download. Pick the use case (coding, reasoning, chat) so the weighting matches your work, and run its benchmark command once a model is serving so your real throughput replaces the estimate. On a machine whose detection misreads memory, override it by hand. It is a sizing tool, not an inference server or an evaluation harness: it will not tell you whether a model answers your questions well, only whether it will run.

Technical notes+

Workspace of three crates (llmfit-core, llmfit-tui, llmfit-desktop) per Cargo.toml, MIT, edition 2024. llmfit-core/src/fit.rs defines CalcConfig with efficiency 0.55, a DEFAULT_ESTIMATION_CTX of 8192 tokens for KV-cache sizing, run-mode speed factors (GPU 1.0, tensor parallel 0.9, MoE offload 0.8, CPU offload 0.5, CPU only 0.3) and per-use-case weights, e.g. Reasoning quality 0.55, Chat speed 0.35. EstimateConfidence orders measured_local, measured_community, calibrated, estimated, unsupported. llmfit-core/src/models.rs holds QUANT_HIERARCHY from Q8_0 to Q2_K, plus MXFP4 handling for gpt-oss and ternary detection for BitNet; unknown quant labels fall back to 0.58 bytes per parameter. llmfit-core/src/hardware.rs does non-short-circuiting multi-vendor detection and queries Metal's recommendedMaxWorkingSetSize on macOS. llmfit-tui/src/main.rs declares an mcp_server module and serve_api.

Observed

License
MIT
Language
Rust (Cargo workspace, edition 2024)
Interfaces
Terminal UI (default), CLI with JSON and CSV output, web dashboard, HTTP JSON API
Install surface
Homebrew, Scoop, MacPorts, uv/pip, install script, multi-arch container image, cargo build from source
Platform support
macOS (Apple Silicon and Intel), Linux (x86_64 and ARM64), Windows (x86_64)
GPU detection
NVIDIA, AMD, Intel Arc, Apple Silicon, Ascend; no GPU detection on Android/Termux
Runtime providers
Ollama, llama.cpp, MLX, Docker Model Runner, LM Studio
Model catalogue
Embedded in the binary at compile time; updated by upgrading the tool
Speed estimate provenance
Every figure labelled measured locally, measured by community, calibrated, estimated or unsupported

Read from README.md, Cargo.toml, LICENSE, docs/how-it-works.md, llmfit-core/src/fit.rs, llmfit-core/src/hardware.rs, llmfit-core/src/models.rs, llmfit-core/Cargo.toml, docs/providers.md, docs/platform-support.md, llmfit-tui/src/main.rs.

Tags

ggufllmlocalaimlxskillunsloth

Tech Stack

RustDocker

Comments (0)

No comments yet

Editorially curated, with community endorsements as a secondary signal. Corrections welcome.