Invent, evaluate, and productionize methods for quantization, QAT, distillation, speculative decoding, KV-cache reuse, KV-cache compression, long-context inference, MoE routing, and model/runtime co-optimization. Build high-quality prototypes in PyTorch, Triton, CUDA-adjacent tooling, or inference-serving frameworks, then work with MLEs and platform engineers to productionize them.