tuGEMM: Area-Power-Efficient Temporal Unary GEMM Architecture for Low-Precision Edge AI
tuGEMM performs exact temporal-coded matrix multiplication using serial and parallel variants with different area/power-latency tradeoffs. The paper reports post-synthesis results in 45 nm CMOS for 2-, 4-, and 8-bit configurations.