Experience with Triton, compiler optimization, advanced GPU profiling tools, multiple accelerator generations, long-sequence or multimodal models, neural time series, or large-scale model training is a plus. Optimize tensor layouts, precision, memory allocation, activation checkpointing, operator fusion, and execution graphs to fit larger or longer-context models within available resources.