Working knowledge of GPU acceleration and the software layers that shape local AI performance, such as inference frameworks (vLLM, SGLang, Llama.cpp, Ollama), model formats, quantization, memory management, and hardware-aware optimization. This role will work directly with the companies and collaborative software projects building the models, runtimes, tools, agent platforms, and applications that bring AI onto PCs, workstations, and personal AI.