Your work will span model internals, quantization and model compression, speculative decoding, KV-cache optimization, inference engines, serving architecture, and benchmarking to improve latency, throughput, memory efficiency, GPU utilization, and cost per token while preserving model quality and reliability. Practical knowledge of at least one modern inference stack, such as vLLM, SGLang, TensorRT-LLM, Triton Inference Server, NVIDIA Dynamo, Ray Serve, KServe, or an equivalent internal system.