CV
Curriculum vitae of Woosung Kang.
Contact Information
| Name | Woosung Kang |
| woosungkang@dgist.ac.kr |
Experience
-
2026 - present Daegu, South Korea
Profile
- Postdoctoral researcher in real-time and on-device AI systems, focusing on resource-constrained AI, deadline-aware multi-DNN scheduling, timing guarantees under resource contention, and system-level optimization for edge AI.
Academic Interests
Education
Recognition
- Best Paper Award, IEEE RTAS 2026 - May 2026
- DGIST Kyu-Young Whang Outstanding Research Award - October 2025
Publications
-
2021 LaLaRAND: Flexible Layer-by-Layer CPU/GPU Scheduling for Real-Time DNN Tasks
IEEE Real-Time Systems Symposium (RTSS'21), Dortmund, Germany
W. Kang, K. Lee, J. Lee, I. Shin, and H. S. Chwa
-
2022 DNN-SAM: Split-and-Merge DNN Execution for Real-Time Object Detection
IEEE Real-Time and Embedded Technology and Applications Symposium (RTAS'22), Milano, Italy
W. Kang, S. Chung, J. Y. Kim, Y. Lee, K. Lee, J. Lee, K. G. Shin, and H. S. Chwa
-
2024 RT-Swap: Addressing GPU Memory Bottlenecks for Real-Time Multi-DNN Inference
IEEE Real-Time and Embedded Technology and Applications Symposium (RTAS'24), Hong Kong, China
W. Kang, J. Lee, Y. Lee, S. Oh, K. Lee, and H. S. Chwa
-
2025 Mitigating Resource Contention for Responsive On-Device Machine Learning Inferences
IEEE/ACM International Conference on Computer-Aided Design (ICCAD'25), Munich, Germany
M. Kim, J. Lee, S. Chou, W. Chung, I. Kim, W. Kang, H. Kim, S. Oh, H. S. Chwa, and K. Lee
-
2025 Timing Guarantees for Inference of AI Models in Embedded Systems
Real-Time Systems Journal
S. Lee, W. Kang, M. Bertogna, H. S. Chwa, and J. Lee
-
2026 ZeroSwap: Minimizing Swap Overhead for Real-Time Multi-DNN Inference via SSD-based GPU Memory Extension
IEEE Real-Time and Embedded Technology and Applications Symposium (RTAS'26), Saint-Malo, France
Best Paper Award. W. Kang, F. Muzzini, G. Brilli, J.-C. Kim, J. Lee, and H. S. Chwa
-
2023 Paste-and-Cut: Collective Image Localization and Classification for Real-Time Multi-Camera Object Detection
International Conference on Information and Communication Technology Convergence (ICTC), Jeju, Korea
Y. E. Kang, W. Kang, T. Lee, and H. S. Chwa
Patents
- Real-Time Object Detection System and Method Using Dynamic Image Merging and Patching Technique - Republic of Korea, granted. Patent No. 10-2988385.
Inventors: Hoon Sung Chwa, Woosung Kang, Youngeun Kang, Taehoon Lee
- Memory Management System and Method Using Memory Virtualization and Real-Time Scheduling - Republic of Korea, published. Application No. 10-2024-0038877.
Inventors: Hoon Sung Chwa, Woosung Kang
- Memory Management System and Method Using Memory Virtualization and Real-Time Scheduling - PCT, published. Application No. PCT/KR2024/011386.
Inventors: Hoon Sung Chwa, Woosung Kang
Professional Service
- Program Committee, ACM S3 Workshop (with MobiCom’25)
- Program Committee, Junior Researcher Workshop on Real-Time Computing (with RTNS’24)
- Secondary Reviewer, IEEE Real-Time Systems Symposium (RTSS), 2022-2024
Skills
Languages
Projects
-
ZeroSwap: Minimizing Swap Overhead for Real-Time Multi-DNN Inference via SSD-based GPU Memory Extension
A memory swapping framework that minimizes swap overhead for real-time multi-DNN inference on embedded systems with integrated GPUs.
- Published in RTAS’26 (first author)
- Semantic-aware selective swapping minimizes swap volume.
- Shared pinned allocation eliminates runtime allocation overhead through physical memory sharing.
- Segment-level overlapping hides PCIe transfer latency by overlapping swapping with computation.
- Improved DNN task-set schedulability by up to 101.7% and response time by up to 3.2x over existing approaches.
-
RT-Swap: Addressing GPU Memory Bottlenecks for Real-Time Multi-DNN Inference
A real-time memory-management framework that extends GPU memory with CPU memory while preserving timing guarantees for multi-DNN inference.
- Published in RTAS’24 (first author)
- Improved task-set schedulability by at least 72% over existing approaches.
- Supported task sets requiring up to 96.2% more memory than the GPU’s physical capacity.
-
DNN-SAM: Split-and-Merge DNN Execution for Real-Time Object Detection
A split-and-merge execution and scheduling framework that provides different accuracy and timeliness levels to image regions with different criticality.
- Published in RTAS’22 (first author)
- Improved safety-critical-region detection accuracy by 2.0-3.7x.
- Reduced average inference latency by 4.8-9.7x without violating timing constraints.
-
LaLaRAND: Flexible Layer-by-Layer CPU/GPU Scheduling for Real-Time DNN Tasks
A real-time layer-level DNN scheduling framework that combines CPU-friendly quantization with fine-grained CPU/GPU allocation.
- Published in RTSS’21 (first author)
- Improved schedulability by 56% over an existing approach and 80% over vanilla PyTorch.
- Limited inference-accuracy difference to at most 0.4%.