Vibeleaderboard
Index / tool
Visit mini-swe-agent.com
Category
AI Agents
Rank

Previous survey · No. 812 ·

Pricing
Open Source
Type
TOOL
Use case
Coding
Interfaces
CLI · SDK
Builder
swe-agent
Latest release
v2.4.6
Date

About

A minimal open-source AI coding agent built from roughly 100 lines of Python that only uses bash (no tool-calling APIs) yet scores over 74% on SWE-bench Verified. Created by the Princeton/Stanford team behind SWE-bench and SWE-agent, it supports local, Docker, and sandboxed execution and is used internally by companies like Meta, NVIDIA, and IBM.

What it does

This is a coding agent that talks to a language model and lets it solve programming tasks purely through shell commands. Each turn sends the running conversation to the model, executes whatever command it proposes as a fresh, independent process, and appends the result back into the same linear transcript. There is no custom tool-calling protocol and no persistent shell session: every command runs on its own, which makes it straightforward to redirect execution into a container or remote sandbox instead of the local machine. It ships as both a command line program and a set of importable Python classes for scripting.

Why it's ranked here

The lineage matters: the same Princeton and Stanford researchers who built the original benchmark and its first, tool-heavy agent turned around and cut that agent down to almost nothing, betting that a strong model with plain shell access beats a scaffold full of custom tools. That bet mostly paid off. The result is now widely used as a baseline across the field precisely because it is small enough to read start to finish and reason about with confidence. For engineers who want to see exactly what an agent does at every step, rather than trust a black box, that transparency is the whole argument for using it.

What's good

The core control loop is compact: query the model, run the command it returns, feed the output back, repeat, with no branching state machine to trace. Every action executes as an independent process rather than inside a kept-open shell, so swapping a local run for a containerized or remote one is a small change rather than a rewrite. Model access, execution environment, and even the agent implementation itself are each selected by a short name and resolved on the fly, so trying a different backend or sandbox does not require forking the project.

Tradeoffs

The minimalism is also the limit: there are no built-in tools beyond a shell, so any task-specific capability, like opening a pull request safely, depends entirely on the model choosing the right shell command rather than on a purpose-built interface. Running each action as an independent process means no shell state carries over between steps, so a multi-step workflow that depends on something like an activated virtual environment has to be re-established inside a single command every time it is needed.

How to use it well

It fits best as a baseline or a starting scaffold: point it at a repository with a plain natural-language task, let it iterate entirely through shell commands, and read the resulting transcript to see exactly what it tried and why. Because the agent, model backend, and execution environment are all swappable independently, it also works well as a testbed for comparing models or sandboxes without touching the core loop. It is a weaker fit for teams that want a polished, tool-rich assistant with built-in guardrails rather than raw shell access to a machine.

Technical notes+

The package entry points are defined in pyproject.toml: several console scripts all map to functions inside minisweagent's run package. src/minisweagent/__init__.py defines the Agent, Model, and Environment types as typing.Protocol interfaces rather than base classes, so a new implementation only needs matching method signatures, not inheritance. src/minisweagent/agents/__init__.py and src/minisweagent/environments/__init__.py both resolve an implementation from a short string key through a lookup dictionary and importlib, with docker, singularity, local, bubblewrap, and contree listed among the built-in environments. src/minisweagent/config/__init__.py resolves YAML config files by checking, in order, an explicit path, an environment-variable override, and two built-in config subdirectories. src/minisweagent/__main__.py is a two-line wrapper that calls the same CLI entry point as the console scripts, so python -m minisweagent and the installed commands run identical code.

Observed

License
MIT
Language
Python, requires Python 3.10 or newer
Interfaces
Command-line interface and a Python library API
Packaging
Published to the Python Package Index; installable via pip, uv, or pipx, or from source
Execution environments
Local process, Docker or Podman, Singularity or Apptainer, and other sandboxed backends, selectable by name
Configuration
YAML configuration files resolved from a project path, an environment-variable override, or built-in defaults
Extensibility
Agent, model, and environment implementations are resolved dynamically from short string identifiers rather than hardcoded imports

Read from README.md, pyproject.toml, src/minisweagent/__init__.py, src/minisweagent/__main__.py, src/minisweagent/agents/__init__.py, src/minisweagent/run/__init__.py, src/minisweagent/environments/__init__.py, src/minisweagent/config/__init__.py, docs/quickstart.md, docs/advanced/control_flow.md.

What it can do

  • Autonomously resolve software engineering tasks using bash commands

    Coding task/issue description → Code changes/patch

  • Execute agent workflows in local, Docker, or sandboxed environments

    Execution environment configuration → Isolated task execution

  • Interact with a coding environment using only bash commands without tool-calling APIs

    Bash commands → Command execution results

Tags

coding-agentswe-benchai-agentscli-toolpythonbenchmarkopen-sourcelitellm

Tech Stack

Python

Comments (0)

No comments yet

Editorially curated, with community endorsements as a secondary signal. Corrections welcome.