Research
LLM inference optimization, deep learning accelerators, neuromorphic computing, and unary arithmetic for efficient hardware systems.
Advisers' groups CMU NCAL · CMU ACTL
Collaborating groups UCF UNARY · NEXUS
PhD Research Projects
My doctoral research is supervised by Prof. J.P. Shen and Prof. Shawn Blanton through CMU ACTL and CMU NCAL, with project collaborations involving UCF UNARY and NEXUS.
Mugi — Value-Level Parallelism for Efficient LLMs
ASPLOS 2026Co-developed Mugi, a technique that exploits value-level parallelism in transformer nonlinear operations such as softmax, SiLU, and GELU, as well as small-batch GEMMs with asymmetric inputs used by weight-only and KV-cache quantization.
Key results: Up to 45× throughput & 668× energy efficiency for softmax; 2.07× LLM throughput and 3.11× energy efficiency end-to-end; 1.45× reduction in operational carbon and 1.48× embodied carbon. Outperforms existing nonlinear approximations in accuracy, performance, and efficiency.
Tempus Core — Temporal-Unary Convolution Core for Edge DLAs
DATE 2025Architected an INT8 temporal-unary convolution core for NVDLA, targeting low-precision edge deep learning inference. The paper reports post-synthesis comparisons and post-place-and-route analysis in 45nm CMOS.
Key results: For the evaluated INT4 16×4 arrays, post-place-and-route results show 53% lower area and 44% lower power than the NVDLA CMAC baseline. Separately, the post-synthesis 16×16 INT8 comparison reports 5× higher iso-area throughput.
TNNGen — Automated Neuromorphic SPU Design Framework
ISCAS 2024 · TCAS-II 2024Developed TNNGen, an automation framework that compiles PyTorch Temporal Neural Network (TNN) models through PyTorch-to-RTL and RTL-to-layout stages. The paper evaluates seven time-series clustering designs across different sensory modalities.
Key results: 32% lower place-and-route runtime on average and nearly 47% lower full hardware-flow runtime for the largest evaluated column. Selected for journal publication in IEEE TCAS-II 2024.
TNN7 — Custom 7nm PDK Extension for Neuromorphic TNNs
ISVLSI 2022Devised TNN7: a custom predictive 7nm open-source PDK extension (ASAP7) comprising 9 custom hard macros for Temporal Neural Networks.
Key results: 14% power, 16% delay, 28% area, 45% EDP reduction over baseline ASAP7 designs.
Industry Research Contributions
These projects were developed through research internships and industry collaborations, with an emphasis on deployment-facing performance characterization and low-power accelerator design.
LLM Inference Profiling on CPU-GPU Coupled Architectures
ISPASS 2025 · Samsung-FundedBuilt SKIP — a PyTorch-based profiling tool for operator-kernel dynamics in LLM inference. The paper uses fine-grained operator-to-kernel traces and Total Kernel Launch and Queuing Time (TKLQT) to compare PCIe A100/H100 systems with GH200 Grace Hopper.
Key results: For Llama-3.2-1B at large batch sizes, GH200 achieves 1.9×–2.7× lower prefill latency than the loosely coupled systems. The study also finds that GH200 remains CPU-bound up to 4× larger batch sizes, exposing a low-batch host-side bottleneck.
tubGEMM — Temporal-Unary-Binary GEMM Unit
ISVLSI 2023 · MediaTekDevised an ultra-low-power hybrid temporal-unary-binary GEMM unit for edge AI. Designed and evaluated on TSMC N5 process technology during an internship at MediaTek USA.
Key results: Compared with uGEMM, the paper reports 89% lower area, 87% lower power, and 50% lower energy. For the evaluated 128×128 INT8 design in TSMC N5, workload sparsity in MobileNetV2 and ResNet-50 reduces energy by more than 3×.
OzMAC — Zero-Skipping MAC for Bit-Sparse DL Inference
VLSI-SoC 2024 · MediaTekCo-developed OzMAC, a modified implementation of Bit-Pragmatic that dynamically skips zero bits. The paper evaluates OzMAC against a binary MAC across multiple precisions and clock frequencies using post-synthesis results in TSMC N5.
Key results: For the evaluated INT8 configuration, the paper reports 21% lower area, 70% lower power, and 28% lower energy. At throughput-normalized frequency, power and energy are each 30% lower than the binary MAC baseline.
Talks & Presentations
Fellowships & Awards
Amar Mukherjee Best Paper Award
ISVLSI 2025
Qualcomm Innovation Fellowship
North America, 2023
CIT Dean's Fellowship
Carnegie Mellon University doctoral fellowship
ISVLSI 2024 Travel Grant
IEEE ISVLSI 2024
CMU GSA Conference Grant
Carnegie Mellon University Graduate Student Assembly
DAC Young Fellow
Design Automation Conference, 2022
ASPLOS Young Architect
ASPLOS 2022
Professional Service
Peer Reviewer: IEEE Transactions on VLSI Systems (TVLSI) · IEEE Journal of Exploratory Solid-State Computational Devices and Circuits (JXCDC)
Memberships: IEEE-Eta Kappa Nu (HKN) · Sigma Xi Scientific Research Honor Society