Characterization of Request and Token Energy Costs for LLM Inference Workloads on GPU Platforms
This paper decomposes GPU energy for LLM inference into request-level and token-level costs across H100 and H200 platforms, showing how batching, context length, output length, and model architecture affect energy efficiency.