Characterizing and Optimizing LLM Inference Workloads on CPU-GPU Coupled Architectures
This paper uses SKIP operator-to-kernel traces and Total Kernel Launch and Queuing Time to compare LLM inference on PCIe A100/H100 and GH200 systems. GH200 improves large-batch prefill latency but remains CPU-bound through batch sizes up to four times larger than the loosely coupled systems.