Claude Engineer v3
github.com/doriandarko/claude-engineer- Category
- AI Agents
- Rank
- No. 1786Tools index
Previous survey · No. 1793 ·
- Pricing
- Open Source
- Type
- AGENT
- Use case
- Coding
- Interfaces
- CLI · Web
- Builder
- doriandarko
- GitHub
- 11.2k stars
- Date
About
A self-improving AI coding assistant framework that uses Claude 3.5 Sonnet to autonomously identify capability gaps, then design and implement new tools for itself during a conversation. It ships with a CLI and a web interface, plus built-in tools for file editing, code execution (via E2B), web search/scraping, and linting.
What it does
A Python chat agent that talks to Anthropic's API and hands the model a folder of plugins it can call. The twist: one plugin asks the model to write a brand new plugin as source code, saves it into that folder, and after you type refresh the agent re-imports everything and can use it. You drive it from a terminal session with rich text output or from a small local browser page that also accepts image uploads.
Why it's ranked here
The idea is the draw, and the code shows it plainly: a four-method plugin contract, a loader that re-imports the folder on demand, and a generator that is a single prompt plus a file write. That makes it a readable teaching example of an agent extending itself. It is not a hardened product. Generated code is saved and imported with no review or test step, the model name is pinned in two places, and the packaging metadata still carries template placeholders.
What's good
The plugin contract is small enough to learn in a minute: a name, a description, an input schema and one execute method. When a plugin fails to import because a package is missing, the loader names the package and offers to install it through uv rather than crashing. Tool call inputs and results are printed in panels with large base64 image data stripped out, so debugging stays readable. Code execution is pushed into a remote E2B sandbox rather than your own machine.
Tradeoffs
Model-written plugins land on disk and get imported into the running process with no validation, so a bad or harmful generation runs with your permissions. The web server keeps one assistant object for every request, so all browser tabs share a single conversation. Errors come back as successful responses. The system prompt advertises shell, git and explorer tools that the README's built-in list does not include. Packaging metadata points at a module and project URLs that do not match the repository layout, and the plain requirements file omits the E2B client.
How to use it well
Treat it as a sandbox for studying how an agent can grow its own toolset, not as a daily coding assistant. Run it in a throwaway virtual environment or container, read every generated plugin before typing refresh, and keep the web server bound to your own machine since it has one shared session. Install from the project file rather than the requirements file if you want sandboxed code execution. For editing a real codebase with review gates, look elsewhere.
Technical notes+
tools/base.py defines BaseTool as an ABC with name, description, input_schema and execute. ce3.py's Assistant._load_tools deletes cached tools.* entries from sys.modules and re-imports each module found by pkgutil.iter_modules; an ImportError triggers an interactive y/n prompt that routes through the uvpackagemanager tool. tools/toolcreator.py calls claude-3-5-sonnet-20241022 at temperature 0 with max_tokens 4000, extracts the name by regex and writes the returned text straight to disk; _validate_tool_name is defined but not called in execute. config.py sets MAX_TOKENS 8000, MAX_CONVERSATION_TOKENS 200000, temperature 0.7. app.py instantiates a module-level Assistant, caps uploads at 16MB, hardcodes image/jpeg in /chat, and returns HTTP 200 on exceptions. tools/e2bcodetool.py creates a new Sandbox per call. pyproject.toml points its version path and console script at a cev3 package with placeholder author and URLs; requirements.txt lacks e2b-code-interpreter.
Observed
- License
- MIT (pyproject.toml and readme.md)
- Language
- Python, requires-python >=3.9 in pyproject.toml
- Interfaces
- Terminal chat (ce3.py) and local Flask web app (app.py)
- Model provider
- Anthropic API only; model id hardcoded in config.py and tools/toolcreator.py
- Plugin contract
- Abstract base class with name, description, input_schema, execute (tools/base.py)
- Generated tool safety
- Generated code written to disk and imported without validation or tests (tools/toolcreator.py)
- Code execution
- Remote E2B sandbox, requires E2B API key (tools/e2bcodetool.py)
- Web session model
- Single module-level assistant shared by all requests (app.py)
- Packaging
- pyproject.toml build targets and URLs do not match repository layout; requirements.txt omits e2b-code-interpreter
Read from readme.md, pyproject.toml, requirements.txt, ce3.py, app.py, config.py, tools/base.py, tools/toolcreator.py, tools/e2bcodetool.py, prompts/system_prompts.py.
What it can do
Autonomously design and implement new tools during a conversation to fill capability gaps
Conversation context → New tool
Edit files
File → Edited file
Execute code
Code snippet → Execution result
Search the web
Search query → Search results
Scrape web pages
URL → Scraped content
Lint code
Code → Lint report
Tags
Tech Stack
Comments (0)
No comments yet
Editorially curated, with community endorsements as a secondary signal. Corrections welcome.