Education

Sep 2021 – Jan 2027 (Expected)
Doctor of Philosophy, Electrical & Computer Engineering
Carnegie Mellon University — Pittsburgh, PA
Co-Advisors: Prof. J.P. Shen & Prof. Shawn Blanton · CMU NCAL, CMU ACTL, UCF UNARY Research Groups
CIT Dean's Fellowship · 2023 Qualcomm Innovation Fellowship
Jan 2020 – May 2021
Master of Science, Electrical & Computer Engineering
Carnegie Mellon University — Pittsburgh, PA
Jun 2014 – Jul 2018
Bachelor of Technology, Electrical & Electronics Engineering
SRM Institute of Science and Technology — Chennai, India

Research Interests

LLM inference optimization Deep learning accelerators Neuromorphic computing Unary computing VLSI / ASIC design Hardware-software co-design CPU-GPU coupled architectures Edge AI

Academic and Industry Experience

Jul 2026 – Present
Silicon CoDesign Engineer
NVIDIA Corporation — Santa Clara, CA
Characterizing LLM inference and agentic AI workloads on NVIDIA platforms and contributing to an internal end-to-end application simulator.
Mar 2026 – Jun 2026
Silicon Solution Engineering Intern
NVIDIA Corporation — Santa Clara, CA
Jun 2024 – Aug 2024
AI Characterization & Tight Coupling Analysis Intern
Samsung Semiconductor Inc. — San Jose, CA
  • Built SKIP, a PyTorch profiling framework used for fine-grained operator-to-kernel trace analysis. The ISPASS 2025 paper reports that GH200 improves large-batch prefill latency for Llama-3.2-1B while remaining CPU-bound through batch sizes up to 4× larger than the evaluated loosely coupled systems.
  • Spearheaded a 5-person CMU-Samsung research collaboration from concept to publication; first-authored paper accepted at ISPASS 2025.
Jun 2022 – Dec 2022
AI Architecture & Algorithm Intern
MediaTek USA Inc. — San Jose, CA (Full-time Jun–Sep; Part-time Remote Aug–Dec)
  • Developed tubGEMM (ISVLSI 2023) and co-developed OzMAC (VLSI-SoC 2024), with both designs evaluated using the TSMC N5 process. The papers report post-synthesis area, power, and energy comparisons against their respective baselines.

PhD Research Highlights

2021 – Present
Doctoral Researcher
CMU NCAL / CMU ACTL — Co-Advisors: Prof. J.P. Shen & Prof. Shawn Blanton
  • Collaborated with research teams across 4 research groups (CMUNCAL, CMU-ACTL, UCF-UNARY, NEXUS), delivering 13 peer-reviewed publications and mentoring graduate and undergraduate students.
  • Characterized performance bottlenecks in LLM inference on GH200 through systematic profiling of model configurations; research funded by Samsung Semiconductor.
  • Created Tempus Core, a temporal-unary convolution engine for NVDLA. The paper reports 53% lower area and 44% lower power for the post-place-and-route INT4 16×4 comparison, and 5× higher iso-area throughput for the post-synthesis INT8 16×16 comparison.
  • Created TNNGen, an automation framework that compiles PyTorch models through PyTorch-to-RTL and RTL-to-layout stages and was evaluated using seven time-series clustering designs across diverse sensory modalities.
  • Developed TNN7, a set of 9 macros for a 7nm PDK extension to ASAP7, reducing energy-delay product (EDP) by 45% against the baseline design.

Recent Publications

  1. Vellaisamy, P., Lam, V., Blanton, S., Shen, J.P. “Characterization of Request and Token Energy Costs for LLM Inference Workloads on GPU Platforms.” IISWC 2026 [Accepted].
  2. Vellaisamy, P., Tripathi, S., Natarajan, V., Thenarasu, S., Blanton, S., Shen, J.P. “TaxBreak: Unmasking the Hidden Costs of LLM Inference Through Overhead Decomposition.” [ISPASS 2026].

View all publications

Presentations

  • Invited Talk. “Characterizing and Optimizing LLM Inference Workloads on CPU-GPU Coupled Architectures.” Jülich Supercomputing Center, Jülich, Germany (Remote), May 20, 2025.
  • Conference Presentation. “Characterizing and Optimizing LLM Inference Workloads on CPU-GPU Coupled Architectures.” IEEE ISPASS 2025, Ghent, Belgium, May 12, 2025.
  • Conference Presentation. “Tempus Core: Area-Power Efficient Temporal-Unary Convolution Core for Low-Precision Edge DLAs.” IEEE DATE 2025, Lyon, France, April 1, 2025.
  • Conference Presentation. “Commercial Evaluation of Zero-Skipping MAC Design for Bit Sparsity Exploitation in DL Inference.” IEEE VLSI-SoC 2024, Tangier, Morocco, October 7, 2024.
  • Conference Presentation. “Exploration of Unary Arithmetic-Based Matrix Multiply Units for Low Precision DL Accelerators.” IEEE ISVLSI 2024, Knoxville, TN, July 2, 2024.
  • Conference Presentation. “TNNGen: Automated Design of Neuromorphic Sensory Processing Units for Time-Series Clustering.” IEEE ISCAS 2024, Singapore, May 21, 2024.

Fellowships, Awards, and Honors

  • Amar Mukherjee Best Paper Award, ISVLSI 2025.
  • ISVLSI 2024 Travel Grant.
  • CMU GSA Conference Grant.
  • Qualcomm Innovation Fellowship 2023.
  • DAC 2022 Young Fellow.
  • ASPLOS Young Architect 2022.
  • Carnegie Institute of Technology Dean's Fellowship.

Teaching Experience

Carnegie Mellon University — Department of Electrical and Computer Engineering

Course Role Semesters
18-340/640: Hardware Arithmetic for Machine Learning Teaching Instructor 4 semesters (approximately 50 students per semester)
18-743: Neuromorphic Computer Architecture & Processor Design Teaching Instructor 5 semesters (approximately 20 graduate students per semester)
18-740: Modern Computer Architecture Teaching Instructor 1 semester (approximately 100 students)

Technical Skills

Tools
vLLM TensorRT NVIDIA Nsight Systems Nsight Compute nvprof Synopsys Design Compiler Synopsys VCS Cadence Genus Cadence Xcelium Cadence Innovus AMD Vivado Intel Quartus Prime
Programming
Python PyTorch SystemVerilog Verilog C++ Tcl
Languages
English Hindi Tamil Japanese

Relevant Coursework

Large Language Models: Methods and Applications Neuromorphic Computer Architecture Modern Computer Architecture Introduction to Machine Learning Hardware Arithmetic for Machine Learning Introduction to Embedded Deep Learning Advanced Digital Integrated Circuit Design Applied Cryptography Fundamentals of Computational Biology

Professional Service

  • IEEE Transactions on Very Large-Scale Integration (VLSI) Systems (IEEE TVLSI) reviewer.
  • IEEE Journal of Exploratory Solid-State Computational Devices and Circuits (IEEE JXCDC) reviewer.

Professional Memberships and Honor Societies

  • IEEE-Eta Kappa Nu (HKN).
  • Sigma Xi Scientific Research Honor Society.