PhD Research Projects

My doctoral research is supervised by Prof. J.P. Shen and Prof. Shawn Blanton through CMU ACTL and CMU NCAL, with project collaborations involving UCF UNARY and NEXUS.

Mugi — Value-Level Parallelism for Efficient LLMs

ASPLOS 2026

Co-developed Mugi, a technique that exploits value-level parallelism in transformer nonlinear operations such as softmax, SiLU, and GELU, as well as small-batch GEMMs with asymmetric inputs used by weight-only and KV-cache quantization.

Key results: Up to 45× throughput & 668× energy efficiency for softmax; 2.07× LLM throughput and 3.11× energy efficiency end-to-end; 1.45× reduction in operational carbon and 1.48× embodied carbon. Outperforms existing nonlinear approximations in accuracy, performance, and efficiency.

LLM Inference Value Parallelism Transformers ASPLOS 2026

Tempus Core — Temporal-Unary Convolution Core for Edge DLAs

DATE 2025

Architected an INT8 temporal-unary convolution core for NVDLA, targeting low-precision edge deep learning inference. The paper reports post-synthesis comparisons and post-place-and-route analysis in 45nm CMOS.

Key results: For the evaluated INT4 16×4 arrays, post-place-and-route results show 53% lower area and 44% lower power than the NVDLA CMAC baseline. Separately, the post-synthesis 16×16 INT8 comparison reports 5× higher iso-area throughput.

NVDLA INT8 Post-Place-and-Route DATE 2025

TNNGen — Automated Neuromorphic SPU Design Framework

ISCAS 2024 · TCAS-II 2024

Developed TNNGen, an automation framework that compiles PyTorch Temporal Neural Network (TNN) models through PyTorch-to-RTL and RTL-to-layout stages. The paper evaluates seven time-series clustering designs across different sensory modalities.

Key results: 32% lower place-and-route runtime on average and nearly 47% lower full hardware-flow runtime for the largest evaluated column. Selected for journal publication in IEEE TCAS-II 2024.

TNN RTL Automation PyTorch Post-Layout Evaluation ISCAS 2024 TCAS-II 2024

TNN7 — Custom 7nm PDK Extension for Neuromorphic TNNs

ISVLSI 2022

Devised TNN7: a custom predictive 7nm open-source PDK extension (ASAP7) comprising 9 custom hard macros for Temporal Neural Networks.

Key results: 14% power, 16% delay, 28% area, 45% EDP reduction over baseline ASAP7 designs.

ASAP7 Custom PDK Hard Macros TNN ISVLSI 2022

Industry Research Contributions

These projects were developed through research internships and industry collaborations, with an emphasis on deployment-facing performance characterization and low-power accelerator design.

LLM Inference Profiling on CPU-GPU Coupled Architectures

ISPASS 2025 · Samsung-Funded

Built SKIP — a PyTorch-based profiling tool for operator-kernel dynamics in LLM inference. The paper uses fine-grained operator-to-kernel traces and Total Kernel Launch and Queuing Time (TKLQT) to compare PCIe A100/H100 systems with GH200 Grace Hopper.

Key results: For Llama-3.2-1B at large batch sizes, GH200 achieves 1.9×–2.7× lower prefill latency than the loosely coupled systems. The study also finds that GH200 remains CPU-bound up to 4× larger batch sizes, exposing a low-batch host-side bottleneck.

LLM Profiling H100 / GH200 KV Cache PyTorch CUDA Kernels ISPASS 2025

tubGEMM — Temporal-Unary-Binary GEMM Unit

ISVLSI 2023 · MediaTek

Devised an ultra-low-power hybrid temporal-unary-binary GEMM unit for edge AI. Designed and evaluated on TSMC N5 process technology during an internship at MediaTek USA.

Key results: Compared with uGEMM, the paper reports 89% lower area, 87% lower power, and 50% lower energy. For the evaluated 128×128 INT8 design in TSMC N5, workload sparsity in MobileNetV2 and ResNet-50 reduces energy by more than 3×.

GEMM Unary Computing TSMC N5 Edge AI ISVLSI 2023

OzMAC — Zero-Skipping MAC for Bit-Sparse DL Inference

VLSI-SoC 2024 · MediaTek

Co-developed OzMAC, a modified implementation of Bit-Pragmatic that dynamically skips zero bits. The paper evaluates OzMAC against a binary MAC across multiple precisions and clock frequencies using post-synthesis results in TSMC N5.

Key results: For the evaluated INT8 configuration, the paper reports 21% lower area, 70% lower power, and 28% lower energy. At throughput-normalized frequency, power and energy are each 30% lower than the binary MAC baseline.

MAC Sparsity TSMC N5 VLSI-SoC 2024

Talks & Presentations

Invited Talk May 20, 2025
"Characterizing and Optimizing LLM Inference Workloads on CPU-GPU Coupled Architectures"
Jülich Supercomputing Center, Forschungszentrum Jülich, Germany (Remote)
Conference May 12, 2025
"Characterizing and Optimizing LLM Inference Workloads on CPU-GPU Coupled Architectures"
ISPASS 2025, Ghent, Belgium
Conference April 1, 2025
"Tempus Core: Area-Power Efficient Temporal-Unary Convolution Core for Low-Precision Edge DLAs"
DATE 2025, Lyon, France
Conference October 7, 2024
"Commercial Evaluation of Zero-Skipping MAC Design for Bit Sparsity Exploitation in DL Inference"
VLSI-SoC 2024, Tangier, Morocco
Conference July 2, 2024
"Exploration of Unary Arithmetic-Based Matrix Multiply Units for Low Precision DL Accelerators"
ISVLSI 2024, Knoxville, TN
Conference May 21, 2024
"TNNGen: Automated Design of Neuromorphic Sensory Processing Units for Time-Series Clustering"
ISCAS 2024, Singapore

Fellowships & Awards

ISVLSI 2024 Travel Grant

IEEE ISVLSI 2024

CMU GSA Conference Grant

Carnegie Mellon University Graduate Student Assembly

DAC Young Fellow

Design Automation Conference, 2022

ASPLOS Young Architect

ASPLOS 2022

Professional Service

Peer Reviewer: IEEE Transactions on VLSI Systems (TVLSI) · IEEE Journal of Exploratory Solid-State Computational Devices and Circuits (JXCDC)

Memberships: IEEE-Eta Kappa Nu (HKN) · Sigma Xi Scientific Research Honor Society