RT-Swap

Real-time GPU memory management using CPU memory as swap space.

Addressing GPU memory bottlenecks for real-time multi-DNN inference

The growing memory requirements of deep neural networks can prevent sophisticated models from running together on memory-constrained GPUs. RT-Swap is a real-time memory-management framework that transparently extends GPU capacity with CPU memory without compromising timing guarantees.

RT-Swap coordinates memory-object swapping with DNN execution, enabling multi-DNN task sets to remain schedulable even when their combined memory demand exceeds physical GPU capacity.

RT-Swap architecture showing the preloaded library, scheduler, GPU and host memory, and memory-object swap-out and swap-in operations
RT-Swap system architecture and memory-object swap-in and swap-out process.

Evaluation on representative machine-learning frameworks showed that RT-Swap:

  • Improved task-set schedulability by at least 72% over existing approaches.
  • Supported workloads demanding up to 96.2% more memory than the GPU’s physical capacity.
  • Preserved the timing guarantees required by real-time inference workloads.

RT-Swap was published at IEEE RTAS 2024.

References

2024

  1. Woosung Kang, Jinkyu Lee, Youngmoon Lee, Sangeun Oh, Kilho Lee, and Hoon Sung Chwa
    In 2024 IEEE 30th Real-Time and Embedded Technology and Applications Symposium (RTAS), 2024