Modular
github.com/modular/modular- Category
- Developer Tools
- Rank
- No. 1985Tools index
- Type
- APP
- Use case
- Models: Train & Run
- Interfaces
- API
- GitHub
- 29.9k stars
- Latest release
- max/v26.6.0
- Date
About
The open-source home of the Modular Platform, spanning the Mojo language and the MAX framework for AI inference. Includes the Mojo compiler and standard library, MAX accelerator kernels, Python-based model pipelines, and an OpenAI-compatible serving endpoint.
What it does
Two linked products live in one repository. Mojo is a programming language whose code can import Python packages, call Python functions and convert values in both directions. MAX takes a Hugging Face model name, downloads and compiles the weights, and runs a local server that answers chat requests from the stock OpenAI Python client with only the base address changed. It targets NVIDIA and AMD datacenter GPUs and also runs, more slowly and with fewer models, on consumer machines including Macs.
Why it's ranked here
Model and hardware breadth, with a clear license split. The supported-models table covers DeepSeek V2 through V4, Gemma 3 and 4, FLUX.2 image generation and sentence embedding models, many marked multi-GPU and offered in 8-bit and 4-bit float weight formats. The same server ships as Docker images for NVIDIA, AMD or both. The repository code is Apache 2.0 with LLVM exceptions, while using and distributing MAX itself falls under the separate Modular Community License, so read both before production use.
What's good
Getting a model served takes three steps in the quickstart: export a Hugging Face token, accept the model license, run one serve command. The container is described as a wrapper around that same command, so a laptop trial and a Kubernetes deployment behave alike. Images come in full and slim flavors per vendor, and the hardware-agnostic one bundles Ubuntu, Python, CUDA, ROCm and PyTorch. The contribution rules are concrete: label AI-assisted work, keep a human reviewing it, aim for pull requests under 100 lines, and expect a first maintainer response within three weeks.
Tradeoffs
The docs strongly recommend datacenter GPUs such as NVIDIA B200, H200 and H100 or AMD MI355X, MI325X and MI300X. The default quickstart model needs more than 96 GiB of GPU memory. The container does not run on macOS. Gated models require a Hugging Face token and license acceptance per model. The compiler does not accept outside contributions, and accepted changes are synced to an internal repository first, reaching the public branch with a nightly build a day or two after merge.
How to use it well
It fits teams self-hosting open-weight models on their own GPUs who want to keep existing OpenAI client code: point the client at the local server and pass any placeholder key. Try it first with Llama 3.1 8B, which the docs say needs about 15 GiB of RAM and works on a Mac. Mount the Hugging Face and MAX cache directories into containers so downloads are reused. Lower the maximum sequence length when a large model does not fit in memory. It is not a hosted API.
Technical notes+
pyproject.toml declares the project as modular version 0 with requires-python >= 3.10, configures black to format .mojo files at 80 columns, enables ruff with Google-style pydocstyle enforced only on the public MAX Python API, and runs mypy with the pydantic plugin; it lists files vendored verbatim from xgrammar and a DeepSeek-V4 checkpoint as excluded from formatting. docs/max/container.mdx names five images (max-full, max-amd, max-amd-base, max-nvidia-full, max-nvidia-base) and documents the --devices gpu:0,1,2,3 / gpu:all / cpu selector and --max-length. Mojo/stdlib/std/python/__init__.mojo exports Python, PythonObject, ConvertibleFromPython and ConvertibleToPython. Mojo/examples/operators/main.mojo demonstrates operator overloading and indexing on a Complex struct. AI_TOOL_POLICY.md asks for an Assisted-by: AI commit trailer.
Observed
- License
- Apache License v2.0 with LLVM Exceptions for the repository
- Runtime license
- MAX usage and distribution under the Modular Community License
- Languages
- Mojo and Python, Python 3.10 or newer
- Interfaces
- max CLI, OpenAI-compatible REST API, Docker images, Mojo standard library with Python interop
- Install surface
- pixi (recommended), uv, pip or conda; Docker images for NVIDIA, AMD or both
- Platform support
- NVIDIA and AMD GPUs recommended; CPU and Mac with fewer models; container is Linux only
- Build tooling
- Bazel wrapper with a format command, plus pre-commit via pixi
- Contribution policy
- AI-assisted work must be labelled and human-reviewed; compiler closed to contributions
Read from README.md, pyproject.toml, LICENSE, AI_TOOL_POLICY.md, CONTRIBUTING.md, docs/max/get-started.mdx, docs/max/models.mdx, docs/max/container.mdx, Mojo/examples/operators/main.mojo, Mojo/stdlib/std/__init__.mojo, Mojo/stdlib/std/python/__init__.mojo.
Intel on Modular
- Select, Don't Train: The Benefits of Modular Entity Disambiguation with LLM-Based Selection
- Modular Manifolds
- Modular Cognitive Architecture Emerges in Large Language Models
- Modular vs end-to-end voice agents, and what the text bottleneck costs
- When Policies Change Probabilities: Modular Decision-Making for LLM Code Review
Tags
Tech Stack
Comments (0)
No comments yet
Editorially curated, with community endorsements as a secondary signal. Corrections welcome.