About
I am an Associate Professor in the Department of Computer Science and Engineering at Ewha Womans University, where I direct the IP-CAL (Intelligent system and Parallel Computer Architecture) Lab, which I founded in March 2021. My research focuses on the architecture of modern parallel computing systems — including GPU micro-architecture, memory systems such as unified virtual memory, machine learning and ray tracing accelerators, and general-purpose parallel processors — with the goal of improving the performance, efficiency, and programmability of next-generation hardware. My work has appeared at venues including ISCA, MICRO, HPCA, and PACT. Before joining Ewha, I worked on graphics and GPU software at Samsung Electronics. I received my Ph.D. in Electrical and Electronic Engineering from Yonsei University in 2018 and my B.S. in Computer Engineering and Computational Mathematics from Washington State University (WSU) in 2011.
Research Interests
- Graphics Processing Unit (GPU)
- Ray Tracing Accelerator (RTA)
- Machine Learning Accelerator
- Unified Virtual Memory (UVM)
- Parallel Programming
- Computer Architecture
Education
- 2018 Ph.D. in Electrical and Electronic Engineering, Yonsei University, Korea
- 2011 B.S. in Computer Engineering and Computational Math, Washington State University (WSU), USA
Work Experience
-
Mar. 2026 – Present
Associate Professor, Ewha Womans University
Department of Computer Science and Engineering -
Mar. 2021 – Feb. 2026
Assistant Professor, Ewha Womans University
Department of Computer Science and Engineering -
Jun. 2020 – Feb. 2021
Network Business, Samsung Electronics
Advanced S/W Lab -
Mar. 2018 – May 2020
Mobile Communications Business, Samsung Electronics
Graphic R&D Group
Academic Service
Organizing Committee
- Registration Chair, 2026 IEEE International Symposium on Performance Analysis of Systems and Software (ISPASS 2026)
- Local Arrangements Co-Chairs, 58th IEEE/ACM International Symposium on Microarchitecture (MICRO 2025)
Program Committee
- 33rd IEEE International Symposium on High-Performance Computer Architecture (HPCA 2027)
Reviewer
- ACM Transactions on Architecture and Code Optimization (TACO), 2026, 2025, 2024
- IEEE Transactions on Computers (TC), 2026, 2025, 2021
- IEEE Computer Architecture Letters (CAL), 2026
- ETRI Journal, 2026, 2024
- The Journal of Supercomputing, 2025, 2024, 2023
- Microprocessors and Microsystems, 2024, 2022
- IEEE Journal on Emerging and Selected Topics in Circuits and Systems (JETCAS), 2024
- IEEE Transactions on Emerging Topics in Computing, 2023, 2021
- ACM Transactions on Design Automation of Electronic Systems, 2023
- IEEE Transactions on Parallel and Distributed Systems (TPDS), 2021
Publications
-
ConfMake Every Batch Count: Fault Entry Merging for Efficient Batching in Unified Virtual Memory
59th IEEE/ACM International Symposium on Microarchitecture (MICRO 2026), Athens, Greece, Oct. 31 – Nov. 4, 2026
-
ConfComplex Tensor Core: Software-Hardware Co-Design for Accelerating Complex-Valued Neural Networks on GPUs
59th IEEE/ACM International Symposium on Microarchitecture (MICRO 2026), Athens, Greece, Oct. 31 – Nov. 4, 2026
-
ConfUnderstanding and Exploiting Cache Asymmetry in Chiplet Processors via Profile-Predict-Guard Allocation
IEEE International Symposium on Workload Characterization (IISWC 2026), Boulder, USA, Sep. 27 – 29, 2026
-
ConfCharacterizing Performance Bottleneck of Distributed Shared Memory in Modern GPUs
IEEE International Symposium on Workload Characterization (IISWC 2026), Boulder, USA, Sep. 27 – 29, 2026
-
ConfAccelerating Vision Transformer Inference via Non-GEMM Kernel Fusion on Edge GPUs
2026 Summer Annual Conference of IEIE, Jeju, Korea, June 23 – 26, 2026
-
JournalBiKD: Bidirectional Kernel Decomposition for Large-Scale GCNs on GPU
IEEE Access, Vol. 14, pp. 96754-96770, June 2026
-
JournalRestructuring the Implicit GEMM Workflow for Complex-Valued Convolution
IEEE Access, Vol. 14, pp. 61010-61024, Apr. 2026
-
JournalCommunication-Optimized Tensor Parallelism for Efficient Multi-GPU Training of Complex-Valued CNNs
Journal of the Institute of Electronics and Information Engineers (IEIE), Vol. 63, No. 4, pp. 53-66, Apr. 2026
-
ConfCharacterizing Cache-Asymmetric CPU Topologies on AMD 3D V-Cache ProcessorsPoster Abstract
2026 IEEE International Symposium on Performance Analysis of Systems and Software (ISPASS 2026), Seoul, Korea, April 26 – 28, 2026
-
ConfFINEA: An Efficient Neural Network Accelerator Exploiting Factorized Input Features
IEEE International Conference on Computer Design (ICCD 2025), Dallas, USA, Nov. 10 – 12, 2025
-
JournalTM-Training: An Energy-Efficient Tiered Memory System for Deep Learning Training in NPUs
ACM Transactions on Storage (TOS), Vol. 21, Issue 4, Article No. 32, pp. 1-26, Nov. 2025
-
ConfWINS: Winograd Structured Pruning for Fast Winograd Convolution
IEEE/CVF International Conference on Computer Vision (ICCV 2025), Honolulu, Hawaii, USA, Oct. 19 – 23, 2025
-
ConfUnderstanding Distributed Training of Large Language Models with Unified Virtual Memory
IEEE International Symposium on Workload Characterization (IISWC 2025), Irvine, USA, Oct. 12 – 14, 2025
-
ConfHALO: Hybrid Systolic Arrays via Logical Partitioning for Acceleration of Complex-Valued Neural Networks
IEEE International Symposium on Workload Characterization (IISWC 2025), Irvine, USA, Oct. 12 – 14, 2025
-
ConfEnergy-Efficient Systolic Array for Complex-Valued Convolutional Neural Networks
40th International Technical Conference on Circuits/Systems, Computers, and Communications (ITC-CSCC 2025), Seoul, Korea, July 07 – 10, 2025
-
JournalMOST: Memory Oversubscription-aware Scheduling for Tensor Migration on GPU Unified Storage
IEEE Computer Architecture Letters (CAL), Vol. 24, Issue 2, pp. 213-216, July 2025
-
ConfAvant-Garde: Empowering GPUs with Scaled Numeric Formats
52nd IEEE/ACM International Symposium on Computer Architecture (ISCA 2025), Tokyo, Japan, June 21 – 25, 2025
-
ConfSSFFT: Energy-Efficient Selective Scaling for Fast Fourier Transform in Embedded GPUs
26th ACM SIGPLAN/SIGBED International Conference on Languages, Compilers, and Tools for Embedded Systems (LCTES 2025), Seoul, Korea, June 15 – 16, 2025
-
JournalTLP Balancer: Predictive Thread Allocation for Multi-Tenant Inference in Embedded GPUs
IEEE Embedded Systems Letters (ESL), Vol. 17, Issue 3, pp. 180-183, June 2025
-
ConfHierarchical Traversal Stack Design Using Shared Memory for GPU Ray TracingBest Paper Nominee
2025 IEEE International Symposium on Performance Analysis of Systems and Software (ISPASS 2025), Ghent, Belgium, May 11 – 13, 2025
-
JournalBeyond VABlock: Improving Transformer Workloads through Aggressive Prefetching
Journal of Systems Architecture (JSA), Vol. 162, pp. 103389, May 2025
-
ConfHyMM: A Hybrid Sparse-Dense Matrix Multiplication Accelerator for GCNs
Design, Automation and Test in Europe Conference (DATE 2025), Lyon, France, Mar. 31 – Apr. 2, 2025
-
ConfWarped-Compaction: Maximizing GPU Register File Bandwidth Utilization via Operand Compaction
31st IEEE International Symposium on High Performance Computer Architecture (HPCA 2025), Las Vegas, USA, Mar. 1 – 5, 2025
-
ConfMarching Page Walks: Batching and Concurrent Page Table Walks for Enhancing GPU Throughput
31st IEEE International Symposium on High Performance Computer Architecture (HPCA 2025), Las Vegas, USA, Mar. 1 – 5, 2025
-
ConfDEPrune: Depth-wise Separable Convolution Pruning for Maximizing GPU Parallelism
38th Annual Conference on Neural Information Processing Systems (NeurIPS 2024), Vancouver, Canada, Dec. 9 – 15, 2024
-
JournalAdaptive Kernel Merge and Fusion for Multi-Tenant Inference in Embedded GPUs
IEEE Embedded Systems Letters (ESL), Vol. 16, Issue 4, pp. 421-424, Dec. 2024
-
ConfPerformance Comparison of CNN Pruning Techniques Using NanoSAM Model on Jetson Orin Nano
2024 Autumn Annual Conference of IEIE, Jeongseon, Gangwon, Korea, Nov. 22 – 23, 2024
-
JournalAdvancements in GPUs for Maximizing AI Application Performance and Research Trends
Communications of the Korean Institute of Information Scientists and Engineers, Vol. 42, Issue 9, pp. 8-13, Sep. 2024
-
ConfVitBit: Enhancing Embedded GPU Performance for AI Workloads through Register Operand Packing
53rd International Conference on Parallel Processing (ICPP 2024), Gotland, Sweden, Aug. 12 – 15, 2024
-
ConfTwisted Bank Arbitrator for Balanced Register Bank Accesses on Graphics Processing Units
2024 Summer Annual Conference of IEIE, Jeju, Korea, June 26 – 28, 2024
-
JournalTriple-A: Early Operand Collector Allocation for Maximizing GPU Register Bank Utilization
IEEE Embedded Systems Letters (ESL), Vol. 16, Issue 2, pp. 206-209, June 2024
-
JournalConflict-Aware Compiler for Hierarchical Register File on GPUs
Journal of Systems Architecture (JSA), Vol. 149, pp. 103099, Apr. 2024
-
JournalSAVector: Vectored Systolic Arrays
IEEE Access, Vol. 12, pp. 44446-44461, Mar. 2024
-
ConfINTERPRET: Inter-Warp Register Reuse for GPU Tensor Core
32nd International Conference on Parallel Architectures and Compilation Techniques (PACT 2023), Vienna, Austria, Oct. 21 – 25, 2023
-
ConfWarped-MC: An Efficient Memory Controller Scheme for Massively Parallel ProcessorsBest Paper Award
52nd International Conference on Parallel Processing (ICPP 2023), Salt Lake City, Utah, USA, Aug. 7 – 10, 2023
-
JournalPerformance Analysis of Neural Processing Units with Emerging Memory Technologies
Journal of the Institute of Electronics and Information Engineers (IEIE), Vol. 60, No. 7, pp. 30-39, July 2023
-
ConfPreloading Architecture for Graphics Processing UnitPaper Award
2023 Summer Annual Conference of IEIE, Jeju, Korea, June 28 – 30, 2023
-
ConfReduced Precision Floating Point for Ray Tracing
2023 Summer Annual Conference of IEIE, Jeju, Korea, June 28 – 30, 2023
-
ConfEarly-Adaptor: An Adaptive Framework For Proactive UVM Memory Management
2023 IEEE International Symposium on Performance Analysis of Systems and Software (ISPASS 2023), Raleigh, NC, USA, Apr. 23 – 25, 2023
-
JournalFairness Analysis of Multi-Tenant Applications on Multi-Instance GPUs
Journal of the Institute of Electronics and Information Engineers (IEIE), Vol. 60, No. 4, pp. 11-23, Apr. 2023
-
ConfBalanced Column-Wise Block Pruning for Maximizing GPU Parallelism
37th Association for the Advancement of Artificial Intelligence (AAAI-23), Washington DC, USA, Feb. 07 – 14, 2023
-
JournalCASH-RF: A Compiler-Assisted Hierarchical Register File in GPUs
IEEE Embedded Systems Letters (ESL), Vol. 14, Issue 4, pp. 187-190, Dec. 2022
-
ConfReconstructing Out-of-Order Issue Queue
55th IEEE/ACM International Symposium on Microarchitecture (MICRO 2022), Chicago, Illinois, USA, Oct. 01 – 05, 2022
-
JournalAnalyzing GCN Aggregation on GPU
IEEE Access, Vol. 10, pp. 113046-113060, Oct. 2022
-
JournalGhostLeg: Selective Memory Coalescing for Secure GPU Architecture
IEEE Access, Vol. 10, pp. 111449-111462, Oct. 2022
-
JournalTEA-RC: Thread Context-Aware Register Cache for GPUs
IEEE Access, Vol. 10, pp. 82049-82062, Aug. 2022
-
ConfCompiler-Assisted GPU Register File Power Management Technique
2022 International Conference on Electronics, Information, and Communication (ICEIC 2022), Jeju, Korea, Feb. 06 – 09, 2022
-
ConfAnalyzing Characteristics of Memory Side-Channels in GPU
2021 Korea Software Congress (KSC 2021), Pyeongchang, Korea, Dec. 20 – 22, 2021
-
JournalREACT: Scalable and High-Performance Regular Expression Pattern Matching Accelerator for In-Storage Processing
IEEE Transactions on Parallel and Distributed Systems (TPDS), Vol. 31, Issue 5, pp. 1137-1151, May 2020
-
JournalAdaptive Cooperation of Prefetching and Warp Scheduling on GPUs
IEEE Transactions on Computers (TC), Vol. 68, No. 4, pp. 609-616, Apr. 2019
-
ConfFineReg: Fine-Grained Register File Management for Augmenting GPU Throughput
51st IEEE/ACM International Symposium on Microarchitecture (MICRO 2018), Fukuoka, Japan, Oct. 20 – 24, 2018
-
JournalWASP: Selective Data Prefetching with Monitoring Runtime Warp Progress on GPUs
IEEE Transactions on Computers (TC), Vol. 67, No. 9, pp. 1366-1373, Sep. 2018
-
JournalDynamic Resizing on Active Warps Scheduler to Hide Operation Stalls on GPUs
IEEE Transactions on Parallel and Distributed Systems (TPDS), Vol. 28, No. 11, pp. 3142-3156, Nov. 2017
-
ConfOptimizing Intersection and Reflection Step of Geometrical Optics using GPUs
16th International Conference on Electronics, Information and Communication (ICEIC 2017), Phuket, Thailand, Jan. 11 – 14, 2017
-
ConfVirtual Thread: Maximizing Thread-Level Parallelism beyond GPU Scheduling Limit
43rd ACM/IEEE International Symposium on Computer Architecture (ISCA 2016), Seoul, Korea, Jun. 18 – 22, 2016
-
ConfAPRES: Improving Cache Efficiency by Exploiting Load Characteristics on GPUs
43rd ACM/IEEE International Symposium on Computer Architecture (ISCA 2016), Seoul, Korea, Jun. 18 – 22, 2016
-
ConfWarped-Preexecution: A GPU Pre-execution Approach for Improving Latency Hiding
22nd International IEEE Symposium on High Performance Computer Architecture (HPCA 2016), Barcelona, Spain, Mar. 12 – 16, 2016
-
ConfDRAW: Investigating Benefits of Adaptive Fetch Group Size on GPU
2015 IEEE International Symposium on Performance Analysis of Systems and Software (ISPASS 2015), Philadelphia, PA, USA, Mar. 29 – 31, 2015
-
JournalIntroduction to Researches on Performance Bottlenecks of Many-Core GPU Architectures
Communications of KIISE, Vol. 32, No. 5, May 2014
-
JournalA Distributed Signature Detection Method for Detecting Intrusions in Sensor Systems
Sensors, Vol. 13, No. 4, pp. 3998-4016, Mar. 2013
-
ConfDirectory Centralized Ring-based Interconnection for Multi-Core Systems
12th International Conference on Electronics, Information and Communication (ICEIC 2013), Bali, Indonesia, Jan. 30 – Feb. 2, 2013
Projects
Ongoing
-
Basic Research Laboratory for Energy-Efficient, General-Purpose Multi-Modal AI with Heterogeneous Computing Accelerators
National Research Foundation of Korea, June 2025 – May 2028
-
A Study on Massively Parallel Processing Architectures for Maximizing Computational Performance and Energy Efficiency of Complex-Valued Neural Network
National Research Foundation of Korea, Mar. 2025 – Feb. 2028
-
Development of 5G-A vRAN Research Platform
Institute for Information & Communication Technology Planning & Evaluation, Apr. 2024 – Dec. 2028
Completed
-
Artificial Intelligence Innovation Hub
Institute for Information & Communication Technology Planning & Evaluation, July 2021 – Dec. 2025
-
Fine-grained Power Management Technique for Large-Scale Parallel Processors
National Research Foundation of Korea, Sep. 2021 – Feb. 2024
-
Developing General Purpose Processing Unit for Machine Learning Algorithms
Ewha Womans University (Ewha Start-up Fund), Mar. 2021 – Feb. 2023
-
Development of Multi-GPU Based High Speed Ray-Tracing Engine
Samsung Electronics, 2017 – 2018
-
Development of Ray Tracing Simulator for Wireless Communication on GPUs
Samsung Electronics, 2016 – 2017
-
GPU Architectures for Unstructured and Irregular Parallel Programs
National Research Foundation of Korea, 2015 – 2018
-
Development of Regular Expression Accelerator IP Embedded on SSD
Samsung Electronics, 2015 – 2016
-
Developing SSD-based MapReduce Acceleration Technology for Efficient Analysis of Big Data
Samsung Electronics, 2014 – 2015
-
Development of Power-awareness Android Binder Monitoring and Enhancement Techniques
LG Electronics, 2013 – 2014
-
Development of Low Power Cache Coherence Protocol and Interconnection Network
LG Electronics, 2012 – 2013