- Category
- Developer Tools
- Rank
- No. 2260Tools index
- Type
- TOOL
- GitHub
- 4 stars
- Date
About
vla-edge provides optimized inference runtimes for three vision-language-action models (ABC-VLA, MolmoAct2, and pi-0.5) running on NVIDIA Jetson Thor hardware. Using an agentic search process to propose custom fused CUDA kernels and runtime glue, alongside CUDA graphs, TensorRT, and a lossless weight decoder, it cuts MolmoAct2 inference from 611ms to 113ms and ABC-VLA inference to 32.4ms versus 63.4ms for a plain TensorRT conversion, without distillation, pruning, or sub-fp16 quantization.
Tags
vlajetson-thorinference-optimizationtensorrtcudarobotics
Tech Stack
Python
Comments (0)
No comments yet
Editorially curated, with community endorsements as a secondary signal. Corrections welcome.
