
VisionClaw
github.com/intent-lab/visionclaw- Category
- AI Agents
- Rank
- No. 956Tools index
Previous survey · No. 940 ·
- Pricing
- Open Source
- Type
- APP
- Builder
- @rohanpaul_ai
- GitHub
- 2.5k stars
- Date
About
A real-time AI assistant that works with Meta Ray-Ban smart glasses to see what you see, hear what you say, and take actions on your behalf through voice commands. Combines live vision, voice interaction, and app integration through Gemini Live API.
What it can do
Analyze visual scenes in real-time
Live camera footage from Meta Ray-Ban glasses → Visual scene description and understanding
Process voice commands
Spoken voice commands through glasses microphone → Executed actions or verbal responses
Stream live video to AI processing
Real-time camera feed at 1fps → Continuous visual analysis via Gemini API
Provide bidirectional audio interaction
User voice input and AI responses → Two-way voice conversation through glasses speakers
Integrate with external applications
Voice commands for app actions → Executed tasks in connected applications
Why it made the leaderboard
Turns Meta Ray-Ban glasses into a real-time AI assistant: your camera streams to Gemini Live at 1fps while bidirectional audio carries voice commands, so it can see what you see and act on it. A working, hackable step toward the sci-fi assistant, built on WebRTC with OpenClaw integration.
Tags
Comments (0)
No comments yet
Editorially curated, with community endorsements as a secondary signal. Corrections welcome.