Agentless
github.com/openautocoder/agentless- Category
- Developer Tools
- Rank
- No. 1050Tools index
- Pricing
- Open Source
- Type
- TOOL
- Use case
- Software Testing & Quality
- Interfaces
- CLI
- Builder
- openautocoder
- GitHub
- 2.1k stars
- Latest release
- v1.5.0
- Date
About
Agentless is a research framework that automatically fixes software issues without using an autonomous agent loop, instead following a fixed three-phase pipeline of fault localization, patch generation, and validation. It achieved 27.3% resolve rate on SWE-bench Lite at roughly $0.34 per issue, outperforming several agent-based baselines, and later reached 40.7-50.8% when paired with Claude 3.5 Sonnet.
What it does
Agentless hands a bug report to a language model in fixed stages instead of letting the model choose its own next step. First it narrows the codebase: a handful of suspect files, then the classes and functions inside them, then the exact lines to change. Next it asks the model for dozens of candidate fixes. Finally it has the model write a small script meant to reproduce the bug, runs that script and the project's existing tests against every candidate, and submits the fix that the most samples agree on among those that pass.
Why it's ranked here
A published result backs the design. The README reports 82 SWE-bench Lite issues fixed (27.3%) at an average of 34 cents each on release, then 40.7% on Lite and 50.8% on Verified once paired with Claude 3.5 Sonnet. The argument behind those numbers matters more than the numbers: a scripted pipeline with heavy sampling and test-based voting outscored the open-source agent loops it was compared against, while staying cheaper and easier to debug. The code is MIT licensed and the complete benchmark runs are downloadable, so the claims can be checked rather than taken on faith.
What's good
Every stage leaves a record. Each phase writes its results and the full model exchange, token counts included, to line-per-record JSON files, so a bad fix can be traced to the stage that went wrong, and a rerun skips issues already finished. File search combines two signals: the model's own pick of up to five files and an embedding search over the folders the model did not rule out. Candidate fixes are normalized before the vote so equivalent changes count together, and a bundled script prices a run from the logged token counts.
Tradeoffs
It is built around the benchmark, not your repository. Every script loads issues from the SWE-bench dataset by ID and checks out the benchmark's own snapshot of the code, so pointing it at a live bug means writing that glue yourself. Search looks only at Python files and skips test files. The dependency list carries no version numbers and pulls the benchmark harness straight from its git repository. One repair step turns the model's written file name into a string by running it as Python code, which executes model output directly. The cost script hard-codes a single per-token price.
How to use it well
Best fit: researchers and teams building their own issue-fixing systems who want a strong, inspectable baseline to beat or borrow from. Run the stages one command at a time as the SWE-bench guide lays out, start with a single issue ID to see what each stage produces, and set the thread count to fit your API rate limits. The expensive knobs are the sample counts: four location sets times ten repair samples yields forty candidate patches per bug before any test runs. Nothing in the pipeline opens a pull request; its final output is a predictions file for the benchmark's evaluator.
Technical notes+
File-level localization in agentless/fl/FL.py asks for at most 5 files with max_tokens 300 at temperature 0. agentless/fl/localize.py runs filter_none_python and filter_out_test_files over the repo structure from get_repo_structure (checked out under a playground directory at the benchmark base_commit), compresses candidate files to skeletons with get_skeleton, and retries related-element localization up to MAX_RETRIES = 5, raising temperature from 0 to 1.0 when no valid location comes back. Per README_swebench.md, edit locations are sampled 4 times at temperature 0.8, retrieval uses OpenAI text-embedding-3-small via llama-index, and agentless/repair/repair.py draws 10 samples per location set (1 greedy, 9 at temperature) with a 10-line context window in SEARCH/REPLACE format. _post_process_multifile_repair calls eval(edited_file_key) on the model-derived file key. agentless/test/generate_reproduction_tests.py prompts for a test that prints Issue reproduced, Issue resolved or Other issues. agentless/repair/rerank.py keeps patches with the minimum regression result count that also pass reproduction, falls back to regression-only, then majority-votes on normalized_patch with first appearance as the tiebreak. agentless/util/model.py defines OpenAI, Anthropic and DeepSeek decoders; the Anthropic path adds a str_replace_editor tool loop capped at MAX_CODEGEN_ITERATIONS = 10. dev/util/cost.py prices tokens at $5 and $15 per million input and output, and $0.02 per million embedding tokens.
Observed
- License
- MIT
- Language
- Python, with a Python 3.11 conda environment in the setup instructions
- Install surface
- Clone, install requirements.txt, export PYTHONPATH; no package install step documented
- Dependencies
- Nine entries in requirements.txt with no version pins; SWE-bench installed from its git repository
- Interface
- Command-line Python scripts, one per pipeline stage
- Model backends
- OpenAI, Anthropic and DeepSeek decoder classes
- Target benchmark
- SWE-bench Lite and SWE-bench Verified datasets, selected by a dataset flag
- Localization scope
- Non-Python files and test files are filtered out before file search
Read from README.md, requirements.txt, LICENSE, README_swebench.md, agentless/fl/localize.py, agentless/fl/FL.py, agentless/repair/repair.py, agentless/repair/rerank.py, agentless/test/generate_reproduction_tests.py, agentless/util/model.py, dev/util/cost.py.
What it can do
Automatically fix software issues via a fixed localization-repair-validation pipeline
Software issue/bug report and codebase → Patch that resolves the issue
Localize the fault in a codebase related to a reported issue
Codebase and issue description → Identified faulty code location
Generate a candidate patch for an identified fault
Localized fault information → Generated code patch
Validate generated patches
Candidate patch → Validation result (pass/fail)
Tags
Tech Stack
Comments (0)
No comments yet
Editorially curated, with community endorsements as a secondary signal. Corrections welcome.