LaLaRAND

Fine-grained CPU/GPU scheduling for real-time DNN layers.

Flexible layer-by-layer CPU/GPU scheduling for real-time DNN tasks

Conventional machine-learning frameworks commonly assign each DNN task to a single processor, limiting their ability to exploit heterogeneous embedded CPU/GPU platforms. LaLaRAND introduces fine-grained, layer-level resource allocation for real-time DNN execution.

The framework combines CPU-friendly quantization with schedulability-aware CPU/GPU allocation while mitigating inference-accuracy loss. Its system-wide scheduler makes runtime decisions using DNN profile data and current resource availability.

LaLaRAND architecture showing layer-level CPU and GPU kernel execution coordinated by a system-wide runtime scheduler
LaLaRAND system overview: layer-level CPU/GPU execution coordinated by a system-wide scheduler.

Compared with prior approaches, LaLaRAND:

  • Improved schedulability by 56% over an existing approach.
  • Improved schedulability by 80% over vanilla PyTorch.
  • Limited the inference-accuracy difference to at most 0.4%.

LaLaRAND was published at IEEE RTSS 2021.

References

2021

  1. Woosung Kang, Kilho Lee, Jinkyu Lee, Insik Shin, and Hoon Sung Chwa
    In 2021 IEEE Real-Time Systems Symposium (RTSS), 2021