Familiar with mainstream model compilation stacks (such as TVM, MLIR, XLA, etc.), with relevant experience in development, and optimization; Proficient in C/C++ development, familiar with assembly, CPU/GPU architecture, and cache mechanism, with practical experience in high-performance kernel development; Familiar with the underlying principles of deep learning frameworks (such as TensorFlow, PyTorch, OneFlow, etc.), understand the computational graph structure, inference/training execution process, and have experience in implementing model graph optimization and compilation optimization; Possess strong abilities in independent thinking, problem decomposition, performance troubleshooting, and practical optimization, and be able to independently overcome complex performance bottlenecks; Preferred Qualifications: Have experience in joint hardware and software design, and possess experience in heterogeneous computing projects; Have experience in contributing to or developing open-source deep learning kernel libraries, compilers, or inference engines; Have in-depth research experience on the underlying architecture and mechanisms of at least one machine learning framework (TensorFlow / PyTorch / MxNet or other self-developed frameworks). Collaborate with the business and algorithm teams to identify performance issues, provide full stack performance analysis, bottleneck diagnosis, and optimization solutions, consolidate general-purpose performance optimization components, toolchains, and platform capabilities, and empower multiple internal business.