About
I am a Silicon CoDesign Engineer at NVIDIA and a PhD candidate in Electrical & Computer Engineering at Carnegie Mellon University, co-advised by Prof. John Paul Shen and Prof. Shawn Blanton. I joined NVIDIA full-time in July 2026 after completing a Silicon Solution Engineering Internship there from March to June 2026. I remain a PhD candidate at CMU, with expected completion in January 2027.
My research spans LLM inference optimization on CPU-GPU coupled architectures, energy-efficient deep learning accelerators, and neuromorphic computing with Temporal Neural Networks. I received the 2023 Qualcomm Innovation Fellowship and the CIT Dean's Fellowship.
Research Focus
LLM systems CPU-GPU coupled inference, orchestration overhead, workload characterization, and performance optimization.
AI accelerators Energy-efficient GEMM, convolution, and MAC architectures for low-precision deep learning.
Neuromorphic hardware Temporal Neural Networks, unary arithmetic, automated PyTorch-to-layout design flows, and custom PDK extensions.
News
All news- Joined NVIDIA full-time as a Silicon CoDesign Engineer in Santa Clara, California.
- TaxBreak published at IEEE ISPASS 2026 and presented on April 27, 2026.
- Started a Silicon Solution Engineering Internship at NVIDIA in Santa Clara, California.
- Mugi accepted at ACM ASPLOS 2026.
- Received the Amar Mukherjee Best Paper Award at IEEE ISVLSI 2025 for the Catwalk paper.
Selected Publications
All publications- TaxBreak: Unmasking the Hidden Costs of LLM Inference Through Overhead Decomposition IEEE ISPASS 2026
- Mugi: Value Level Parallelism For Efficient LLMs ACM ASPLOS 2026
- Characterizing and Optimizing LLM Inference Workloads on CPU-GPU Coupled Architectures IEEE ISPASS 2025
- Tempus Core: Area-Power Efficient Temporal-Unary Convolution Core for Low-Precision Edge DLAs IEEE DATE 2025
- Catwalk: Unary Top-K for Efficient Ramp-No-Leak Neuron Design for Temporal Neural Networks IEEE ISVLSI 2025 ยท Amar Mukherjee Best Paper Award