Vibeleaderboard
Index / agent
Visit github.com
Category
AI Agents
Rank
No. 1786Tools index

Previous survey · No. 1793 ·

Pricing
Open Source
Type
AGENT
Use case
Coding
Interfaces
CLI · Web
Date

About

A self-improving AI coding assistant framework that uses Claude 3.5 Sonnet to autonomously identify capability gaps, then design and implement new tools for itself during a conversation. It ships with a CLI and a web interface, plus built-in tools for file editing, code execution (via E2B), web search/scraping, and linting.

What it does

A Python chat agent that talks to Anthropic's API and hands the model a folder of plugins it can call. The twist: one plugin asks the model to write a brand new plugin as source code, saves it into that folder, and after you type refresh the agent re-imports everything and can use it. You drive it from a terminal session with rich text output or from a small local browser page that also accepts image uploads.

Why it's ranked here

The idea is the draw, and the code shows it plainly: a four-method plugin contract, a loader that re-imports the folder on demand, and a generator that is a single prompt plus a file write. That makes it a readable teaching example of an agent extending itself. It is not a hardened product. Generated code is saved and imported with no review or test step, the model name is pinned in two places, and the packaging metadata still carries template placeholders.

What's good

The plugin contract is small enough to learn in a minute: a name, a description, an input schema and one execute method. When a plugin fails to import because a package is missing, the loader names the package and offers to install it through uv rather than crashing. Tool call inputs and results are printed in panels with large base64 image data stripped out, so debugging stays readable. Code execution is pushed into a remote E2B sandbox rather than your own machine.

Tradeoffs

Model-written plugins land on disk and get imported into the running process with no validation, so a bad or harmful generation runs with your permissions. The web server keeps one assistant object for every request, so all browser tabs share a single conversation. Errors come back as successful responses. The system prompt advertises shell, git and explorer tools that the README's built-in list does not include. Packaging metadata points at a module and project URLs that do not match the repository layout, and the plain requirements file omits the E2B client.

How to use it well

Treat it as a sandbox for studying how an agent can grow its own toolset, not as a daily coding assistant. Run it in a throwaway virtual environment or container, read every generated plugin before typing refresh, and keep the web server bound to your own machine since it has one shared session. Install from the project file rather than the requirements file if you want sandboxed code execution. For editing a real codebase with review gates, look elsewhere.

Technical notes+

tools/base.py defines BaseTool as an ABC with name, description, input_schema and execute. ce3.py's Assistant._load_tools deletes cached tools.* entries from sys.modules and re-imports each module found by pkgutil.iter_modules; an ImportError triggers an interactive y/n prompt that routes through the uvpackagemanager tool. tools/toolcreator.py calls claude-3-5-sonnet-20241022 at temperature 0 with max_tokens 4000, extracts the name by regex and writes the returned text straight to disk; _validate_tool_name is defined but not called in execute. config.py sets MAX_TOKENS 8000, MAX_CONVERSATION_TOKENS 200000, temperature 0.7. app.py instantiates a module-level Assistant, caps uploads at 16MB, hardcodes image/jpeg in /chat, and returns HTTP 200 on exceptions. tools/e2bcodetool.py creates a new Sandbox per call. pyproject.toml points its version path and console script at a cev3 package with placeholder author and URLs; requirements.txt lacks e2b-code-interpreter.

Observed

License
MIT (pyproject.toml and readme.md)
Language
Python, requires-python >=3.9 in pyproject.toml
Interfaces
Terminal chat (ce3.py) and local Flask web app (app.py)
Model provider
Anthropic API only; model id hardcoded in config.py and tools/toolcreator.py
Plugin contract
Abstract base class with name, description, input_schema, execute (tools/base.py)
Generated tool safety
Generated code written to disk and imported without validation or tests (tools/toolcreator.py)
Code execution
Remote E2B sandbox, requires E2B API key (tools/e2bcodetool.py)
Web session model
Single module-level assistant shared by all requests (app.py)
Packaging
pyproject.toml build targets and URLs do not match repository layout; requirements.txt omits e2b-code-interpreter

Read from readme.md, pyproject.toml, requirements.txt, ce3.py, app.py, config.py, tools/base.py, tools/toolcreator.py, tools/e2bcodetool.py, prompts/system_prompts.py.

What it can do

  • Autonomously design and implement new tools during a conversation to fill capability gaps

    Conversation context → New tool

  • Edit files

    File → Edited file

  • Execute code

    Code snippet → Execution result

  • Search the web

    Search query → Search results

  • Scrape web pages

    URL → Scraped content

  • Lint code

    Code → Lint report

Tags

claudeai-agentscoding-assistantself-improvingclideveloper-toolsllm

Tech Stack

Python

Comments (0)

No comments yet

Editorially curated, with community endorsements as a secondary signal. Corrections welcome.