RT-Swap
Real-time GPU memory management using CPU memory as swap space.
Addressing GPU memory bottlenecks for real-time multi-DNN inference
The growing memory requirements of deep neural networks can prevent sophisticated models from running together on memory-constrained GPUs. RT-Swap is a real-time memory-management framework that transparently extends GPU capacity with CPU memory without compromising timing guarantees.
RT-Swap coordinates memory-object swapping with DNN execution, enabling multi-DNN task sets to remain schedulable even when their combined memory demand exceeds physical GPU capacity.
Evaluation on representative machine-learning frameworks showed that RT-Swap:
- Improved task-set schedulability by at least 72% over existing approaches.
- Supported workloads demanding up to 96.2% more memory than the GPU’s physical capacity.
- Preserved the timing guarantees required by real-time inference workloads.
RT-Swap was published at IEEE RTAS 2024.
References
2024
- In 2024 IEEE 30th Real-Time and Embedded Technology and Applications Symposium (RTAS), 2024