Carnegie Mellon University campus

About

I am a Silicon CoDesign Engineer at NVIDIA and a PhD candidate in Electrical & Computer Engineering at Carnegie Mellon University, co-advised by Prof. John Paul Shen and Prof. Shawn Blanton. I joined NVIDIA full-time in July 2026 after completing a Silicon Solution Engineering Internship there from March to June 2026. I remain a PhD candidate at CMU, with expected completion in January 2027.

My research spans LLM inference optimization on CPU-GPU coupled architectures, energy-efficient deep learning accelerators, and neuromorphic computing with Temporal Neural Networks. I received the 2023 Qualcomm Innovation Fellowship and the CIT Dean's Fellowship.

Research Focus

LLM systems CPU-GPU coupled inference, orchestration overhead, workload characterization, and performance optimization.

AI accelerators Energy-efficient GEMM, convolution, and MAC architectures for low-precision deep learning.

Neuromorphic hardware Temporal Neural Networks, unary arithmetic, automated PyTorch-to-layout design flows, and custom PDK extensions.

News

All news
  • Joined NVIDIA full-time as a Silicon CoDesign Engineer in Santa Clara, California.
  • TaxBreak published at IEEE ISPASS 2026 and presented on April 27, 2026.
  • Started a Silicon Solution Engineering Internship at NVIDIA in Santa Clara, California.
  • Mugi accepted at ACM ASPLOS 2026.
  • Received the Amar Mukherjee Best Paper Award at IEEE ISVLSI 2025 for the Catwalk paper.

Selected Publications

All publications
  1. TaxBreak: Unmasking the Hidden Costs of LLM Inference Through Overhead Decomposition P. Vellaisamy, S. Tripathi, V. Natarajan, S.S. Thenarasu, R.D.S. Blanton, J.P. Shen IEEE ISPASS 2026
  2. Mugi: Value Level Parallelism For Efficient LLMs D. Price, P. Vellaisamy, J.P. Shen, D. Wu ACM ASPLOS 2026
  3. Characterizing and Optimizing LLM Inference Workloads on CPU-GPU Coupled Architectures P. Vellaisamy, T. Labonte, S. Chakraborty, M. Turner, S. Sury, J.P. Shen IEEE ISPASS 2025
  4. Tempus Core: Area-Power Efficient Temporal-Unary Convolution Core for Low-Precision Edge DLAs P. Vellaisamy, H. Nair, T. Kang, Y. Ni, H. Fan, B. Qi, J. Chen, R.D.S. Blanton, J.P. Shen IEEE DATE 2025
  5. Catwalk: Unary Top-K for Efficient Ramp-No-Leak Neuron Design for Temporal Neural Networks D. Lister, P. Vellaisamy, J.P. Shen, D. Wu IEEE ISVLSI 2025 ยท Amar Mukherjee Best Paper Award