It covers a variety of technical fields, including deep learning, computer vision, graphics, speech, recording and editing, special effects, client and server engineering, and provides cutting-edge content understanding, content creation, interactive experience, and consumption capabilities and industry solutions to other business lines within the company and external partners in various forms. Masters or PhD in computer science, mathematics, engineering engineering with at least 5 years of research and practical experience in one or more areas of computer vision, including but not limited to: Experience in multimodal understanding, such as video highlight detection and slicing, audio/music understanding, etc.