Introducing AMem NCCL-Plugin, the 2nd #OpenSource component to ASystem! This library solves the critical problem of inefficient NCCL memory offload in RL workflows: 🪄Significant VRAM Savings: free up 10GB+ on a single Hopper-architecture GPU 🚀Extreme Efficiency: switching time is optimized from the typical minutes to $<1$ second Verified by Ring-1T large-scale RL training, try it today!
Welcome to the ASystem's community! 📦 GitHub for AMem NCCL-Plugin: https://t.co/sdYML0ipfX Last week's AWEX release: https://t.co/gwotBgBIk6 For the awesome Ling/Ring models trained by ASystem: 🤗 Hugging Face: https://t.co/ivGFnm3sM8 🤖 ModelScope Community: https://t.co/qaUr3bKcDw
Core Features 1/2 💾 Transparent Memory Management: Resolving the NCCL cross-rank cross-reference challenge to enable transparent offloading. The resulting ability to free over 10 GB of GPU memory per Hopper card while maintaining the communication state offers substantial resource optimization during training/inference phase transitions. ⏳ Performance Optimization: Reducing the transition latency from the typical minutes required for connection rebuilds down to under 1 second. This gain, achieved by exclusively restoring metadata and preserving the communication group, will accelerate iterative RL development.
Core Features 2/2 🧩 Pluggable Architecture: Seamless integration through a lightweight plugin mechanism with the “ncclPause()” and “ncclResume()” memory management APIs. This provides a flexible, non-intrusive approach to enhancing GPU memory utilization within existing RL workflows. 📷 Validation at Scale: The efficacy and robustness were rigorously validated in the RL training environment supporting the Ring-1T trillion-parameter inference model, confirming reliable performance at a massive scale.
RL workflows lose minutes rebuilding NCCL communicators on every train-to- switch. AMem preserves the communication group and restores only metadata, freeing over 10GB per Hopper GPU and making the switch sub-second.
Checking sign-in…
Loading comments…