- Category
- Developer Tools
- Rank
- No. 724Tools index
- Listed in
- #15 Run models locally
- Pricing
- Open Source
- Type
- TOOL
- Use case
- Models: Train & Run
- Interfaces
- CLI · Desktop
- GitHub
- 26.2k stars
- Latest release
- 0.10.6
- Date
About
llamafile packages an LLM's weights together with llama.cpp and Cosmopolitan Libc into a single cross-platform executable that runs locally with no installation, across macOS, Windows, Linux, and BSD variants on multiple CPU architectures. It also bundles whisperfile, a single-file speech-to-text and translation tool built on whisper.cpp.
What it does
This tool turns a language model into one downloadable file you make executable and run directly, no interpreter, no dependency install, no build step. The same file boots on Mac, Windows, and Linux machines and opens into a terminal chat, a local web interface, or an API server that speaks the same protocol OpenAI and Anthropic clients already use. Companion single-file builds handle speech transcription and image generation using the same trick.
Why it's ranked here
The default run sits behind a syscall sandbox that blocks outbound network and file writes, an unusual posture for something that loads model files and serves an HTTP API. That protection disappears the moment a GPU backend loads, which the project's own docs call the common production case, and the default combined chat-plus-server mode also runs unsandboxed. Apache 2.0 licensed, with dedicated unit and integration tests just for the sandbox behavior. Worth trusting more on CPU than on GPU.
What's good
Distribution could not be simpler: one file, no package manager, no container. It exposes a server endpoint compatible with both OpenAI's and Anthropic's chat APIs, so client code written against either often works unchanged once pointed at a local port. GPU support is compiled on the fly against whatever toolchain is already on the host, avoiding separate binaries per graphics card. The security posture is unusually explicit, with default restrictions on outbound network and writes, plus an optional stricter mode that hides the rest of the filesystem.
Tradeoffs
Windows executables cap out at four gigabytes, so larger models need their weights supplied as a separate file rather than baked into the binary. The optional filesystem confinement locks down whole directories, not individual files, so anything sharing a folder with the model weights stays readable too. Sandboxing is self-imposed: a build from an untrusted source could simply have the protection removed, so it defends against bugs in this code, not against a hostile one.
How to use it well
Good fit for local, offline experimentation, or for shipping a demo a non-technical person can run without installing anything: download, make executable, and go. For anything exposed to a network or run with GPU acceleration, treat the sandbox as absent, since both GPU mode and the default combined mode skip it, and add your own isolation. If existing code already talks to OpenAI's or Anthropic's client libraries, point it at the local server instead of rewriting the integration.
Technical notes+
llamafile/sandbox.c implements the pledge() and unveil() policy: pledge() becomes a SECCOMP BPF filter on Linux or native pledge(2) on OpenBSD, while unveil() uses the Landlock LSM (kernel 5.13+) for the opt-in read confinement flag; docs/security.md lists the exact permission strings used for server, cli, and chat modes and notes that GPU mode, remote-procedure features, and the combined default mode relax or skip the sandbox entirely. Model loading in llamafile/llamafile.c memory-maps the GGUF weights straight out of the executable's own zip central directory, with explicit bounds checks against the archive's declared offsets and sizes to guard against a malformed file. docs/technical_details.md explains the packaging trick itself: llama.cpp is built twice, once per architecture, wrapped in a shell script that runs natively on Windows and through a small loader on Linux, with GPU backend source compiled against whatever host toolchain is present at runtime instead of being statically linked.
Observed
- License
- Apache 2.0 for the project itself; its changes to llama.cpp and whisper.cpp are kept MIT-licensed to stay upstreamable
- Distribution
- Single downloadable executable, no package manager or installer required
- Platform support
- macOS, Windows, Linux, and BSD variants, across multiple CPU architectures, from one binary
- Interfaces
- Terminal chat UI, single-prompt CLI mode, and an HTTP server exposing OpenAI- and Anthropic-compatible chat completion endpoints
- Companion tools
- Separate single-file builds for speech-to-text and translation, transcription, and image generation, built with the same packaging approach
- Sandboxing
- pledge() syscall restriction active by default (SECCOMP filter on Linux, native pledge on OpenBSD); optional unveil()-based filesystem confinement behind a flag; both are no-ops on macOS and Windows
- GPU support
- Backend libraries compiled at runtime against whatever GPU toolchain is present on the host, rather than statically linked into the binary
- Windows constraint
- Executables over 4GB cannot run on Windows; larger models require supplying weights as a separate external file
- Testing
- Repository includes dedicated unit and integration test suites specifically for sandbox behavior
- Build system
- GNU Make-based build pulling in llama.cpp, whisper.cpp, stable-diffusion.cpp, and transcribe.cpp as git submodules with project-specific patches applied on setup
Read from README.md, Makefile, llamafile/main.cpp, llamafile/args.cpp, llamafile/llamafile.c, llamafile/chatbot_main.cpp, llamafile/sandbox.c, llamafile/compute.cpp, docs/quickstart.md, docs/technical_details.md, docs/security.md.
What it can do
Run a large language model locally from a single executable file
Model weights bundled in executable → Generated text
Transcribe speech to text
Audio file → Text transcript
Translate spoken audio
Audio file → Translated text
Tags
Media
Comments (0)
No comments yet
Editorially curated, with community endorsements as a secondary signal. Corrections welcome.
