KEY RESPONSIBILITIES: Design, research, implement, and rigorously optimize high-performance matmul, attention (flash, paged, grouped-query), MoE, and fully fused transformer kernels using Triton, targeting large-scale LLM and multimodal workloads. Drive deep kernel-level optimizations across the AMD memory hierarchy (LDS, L2, HBM), wavefront execution (wave32/wave64), vectorization, MFMA utilization, occupancy tuning, and instruction scheduling to maximize hardware efficiency.