Vibeleaderboard
Index / tool
Visit github.com
Category
Developer Tools
Pricing
Open Source
Type
TOOL
GitHub
37 stars
Added
Jun 18, 2026

About

A specialized benchmarking tool for evaluating Large Language Models on Apple's MLX framework knowledge and coding tasks. It includes 441 questions across different categories and difficulty levels, supports multiple LLM providers (local and cloud), and generates detailed performance reports.

Why it made the leaderboard

Benchmarks how well an LLM actually knows Apple's MLX framework: 441 questions across QA, coding, and debugging, runnable against local Ollama models or OpenAI, Anthropic, and Groq, with detailed reports. Useful for picking which model to trust with MLX work.

Tags

mlxbenchmarkllmapplemachine-learningclievaluation

Tech Stack

Python

Comments (0)

No comments yet

Indexed by a proprietary survey. Corrections welcome.