CV

Curriculum vitae of Woosung Kang.

Contact Information

Name Woosung Kang
Email woosungkang@dgist.ac.kr

Experience

Profile

  • Postdoctoral researcher in real-time and on-device AI systems, focusing on resource-constrained AI, deadline-aware multi-DNN scheduling, timing guarantees under resource contention, and system-level optimization for edge AI.

Academic Interests

On-device AI Systems: Resource-constrained AI, Real-time DNN inference, System-level optimization for edge AI
Real-Time Scheduling: Deadline-aware multi-DNN scheduling, Timing guarantees under resource contention

Education

  • 2019 - 2026

    Daegu, South Korea

    Ph.D.
    Daegu Gyeongbuk Institute of Science and Technology (DGIST)
    Computer Science, Real-Time Computing Lab
    • Advisor: Prof. Hoon Sung Chwa
  • 2014 - 2019

    Daegu, South Korea

    B.S.
    DGIST, School of Undergraduate Studies
    Computer Science

Recognition

  • Best Paper Award, IEEE RTAS 2026 - May 2026
  • DGIST Kyu-Young Whang Outstanding Research Award - October 2025

Publications

  • 2021
    LaLaRAND: Flexible Layer-by-Layer CPU/GPU Scheduling for Real-Time DNN Tasks
    IEEE Real-Time Systems Symposium (RTSS'21), Dortmund, Germany

    W. Kang, K. Lee, J. Lee, I. Shin, and H. S. Chwa

  • 2022
    DNN-SAM: Split-and-Merge DNN Execution for Real-Time Object Detection
    IEEE Real-Time and Embedded Technology and Applications Symposium (RTAS'22), Milano, Italy

    W. Kang, S. Chung, J. Y. Kim, Y. Lee, K. Lee, J. Lee, K. G. Shin, and H. S. Chwa

  • 2024
    RT-Swap: Addressing GPU Memory Bottlenecks for Real-Time Multi-DNN Inference
    IEEE Real-Time and Embedded Technology and Applications Symposium (RTAS'24), Hong Kong, China

    W. Kang, J. Lee, Y. Lee, S. Oh, K. Lee, and H. S. Chwa

  • 2025
    Mitigating Resource Contention for Responsive On-Device Machine Learning Inferences
    IEEE/ACM International Conference on Computer-Aided Design (ICCAD'25), Munich, Germany

    M. Kim, J. Lee, S. Chou, W. Chung, I. Kim, W. Kang, H. Kim, S. Oh, H. S. Chwa, and K. Lee

  • 2025
    Timing Guarantees for Inference of AI Models in Embedded Systems
    Real-Time Systems Journal

    S. Lee, W. Kang, M. Bertogna, H. S. Chwa, and J. Lee

  • 2026
    ZeroSwap: Minimizing Swap Overhead for Real-Time Multi-DNN Inference via SSD-based GPU Memory Extension
    IEEE Real-Time and Embedded Technology and Applications Symposium (RTAS'26), Saint-Malo, France

    Best Paper Award. W. Kang, F. Muzzini, G. Brilli, J.-C. Kim, J. Lee, and H. S. Chwa

  • 2023
    Paste-and-Cut: Collective Image Localization and Classification for Real-Time Multi-Camera Object Detection
    International Conference on Information and Communication Technology Convergence (ICTC), Jeju, Korea

    Y. E. Kang, W. Kang, T. Lee, and H. S. Chwa

Patents

  • Real-Time Object Detection System and Method Using Dynamic Image Merging and Patching Technique - Republic of Korea, granted. Patent No. 10-2988385.
    Inventors: Hoon Sung Chwa, Woosung Kang, Youngeun Kang, Taehoon Lee
  • Memory Management System and Method Using Memory Virtualization and Real-Time Scheduling - Republic of Korea, published. Application No. 10-2024-0038877.
    Inventors: Hoon Sung Chwa, Woosung Kang
  • Memory Management System and Method Using Memory Virtualization and Real-Time Scheduling - PCT, published. Application No. PCT/KR2024/011386.
    Inventors: Hoon Sung Chwa, Woosung Kang

Professional Service

  • Program Committee, ACM S3 Workshop (with MobiCom’25)
  • Program Committee, Junior Researcher Workshop on Real-Time Computing (with RTNS’24)
  • Secondary Reviewer, IEEE Real-Time Systems Symposium (RTSS), 2022-2024

Skills

Programming: C++, Python, CUDA
ML Systems: ML system design, GPU/CPU memory management, Concurrency, Efficient I/O, Runtime design, Scheduling, Memory optimization, Framework internals
Frameworks: PyTorch, TensorFlow Lite, Darknet
Platforms: NVIDIA Jetson (Orin, Nano), Server-class GPUs, Linux, Docker, Git

Languages

Korean : Native
English : Fluent

Projects

  • ZeroSwap: Minimizing Swap Overhead for Real-Time Multi-DNN Inference via SSD-based GPU Memory Extension

    A memory swapping framework that minimizes swap overhead for real-time multi-DNN inference on embedded systems with integrated GPUs.

    • Published in RTAS’26 (first author)
    • Semantic-aware selective swapping minimizes swap volume.
    • Shared pinned allocation eliminates runtime allocation overhead through physical memory sharing.
    • Segment-level overlapping hides PCIe transfer latency by overlapping swapping with computation.
    • Improved DNN task-set schedulability by up to 101.7% and response time by up to 3.2x over existing approaches.
  • RT-Swap: Addressing GPU Memory Bottlenecks for Real-Time Multi-DNN Inference

    A real-time memory-management framework that extends GPU memory with CPU memory while preserving timing guarantees for multi-DNN inference.

    • Published in RTAS’24 (first author)
    • Improved task-set schedulability by at least 72% over existing approaches.
    • Supported task sets requiring up to 96.2% more memory than the GPU’s physical capacity.
  • DNN-SAM: Split-and-Merge DNN Execution for Real-Time Object Detection

    A split-and-merge execution and scheduling framework that provides different accuracy and timeliness levels to image regions with different criticality.

    • Published in RTAS’22 (first author)
    • Improved safety-critical-region detection accuracy by 2.0-3.7x.
    • Reduced average inference latency by 4.8-9.7x without violating timing constraints.
  • LaLaRAND: Flexible Layer-by-Layer CPU/GPU Scheduling for Real-Time DNN Tasks

    A real-time layer-level DNN scheduling framework that combines CPU-friendly quantization with fine-grained CPU/GPU allocation.

    • Published in RTSS’21 (first author)
    • Improved schedulability by 56% over an existing approach and 80% over vanilla PyTorch.
    • Limited inference-accuracy difference to at most 0.4%.