The TikTok Agentic Arch team is building AI-native engineering infrastructure that transforms how software is designed, built, and operated at TikTok scale. We research and deploy reliable long-horizon agents that can reason over large codebases, interact with complex tools and environments, and complete consequential engineering work across distributed production systems.
Our environment provides a unique opportunity to advance agent research through real-world execution feedback. You will develop frontier methods, evaluate them rigorously, and bring them into production through platforms used at large scale-creating measurable improvements in engineering productivity, software quality, and operational effectiveness.
We are looking for talented individuals to join our team. As a graduate, you will get opportunities to pursue bold ideas, tackle complex challenges, and unlock limitless growth.
Successful candidates must be able to commit to an onboarding date by the end of the year. Please state your availability and graduation date clearly in your resume.
Responsibilities
Conduct frontier research on reliable long-horizon agents for complex, multi-step software engineering and operations tasks.
Advance agent capabilities in areas such as hierarchical planning and reasoning, memory and context management, tool use, environment interaction, and multi-agent coordination.
Co-design agents and models through post-training, reinforcement learning, learning from execution feedback, search, and test-time scaling to improve performance on real-world tasks.
Build scalable agent infrastructure for orchestration, evaluation, observability, and reliable execution across large codebases, engineering toolchains, and distributed production environments.
Apply and validate new methods in representative workflows such as cross-repository software changes, testing and verification, large-scale migrations, deployment, and incident diagnosis and remediation.
Translate research into production systems, define rigorous evaluation methodologies, and measure impact through task success, software quality, engineering efficiency, and system performance.
Collaborate with researchers, infrastructure teams, developer-platform teams, and product engineers to deploy solutions at scale and produce publishable research and broader scientific insights where appropriate. Minimum Qualifications
Individuals who are completing or have recently completed a PhD degree in Computer Science or a related discipline.
Demonstrated research or engineering experience in AI agents or closely related areas such as large language model reasoning, reinforcement learning, program synthesis, or AI for code.
Strong understanding of one or more relevant areas, including agent planning and reasoning, model post-training, reinforcement learning, memory and context systems, tool learning, multi-agent systems, or agent evaluation.
Strong programming and systems-building ability in at least one language such as Python, C++, Go, or Java, with the ability to turn research ideas into robust implementations.
Ability to formulate ambiguous real-world problems, design rigorous experiments and evaluations, analyze results, and iterate from evidence.
Strong communication and collaboration skills, with the ability to work across research, infrastructure, platform, and product teams.
Preferred Qualifications
Evidence of research excellence through influential publications, open-source work, deployed systems, or other significant contributions. Relevant venues include NeurIPS, ICML, ICLR, ACL, MLSys, OSDI, SOSP, NSDI, ICSE, and FSE.
Experience building or deploying LLM agents, AI developer tools, or agent platforms in complex or large-scale environments.•Hands-on experience with model post-training, reinforcement learning, execution-feedback loops, tool-using agents, distributed agent runtimes, or scalable evaluation systems.
Experience working with large codebases, distributed systems, CI/CD and testing infrastructure, developer platforms, or production operations
A track record of translating research into measurable production impact.
Numbers & Facts
Location
San Jose, CA
Skills
Analysis Skillsunmatched
Artificial Intelligence (AI)unmatched
Artificial Intelligence (AI) Agentsunmatched
C++ Programming Languageunmatched
Communication Skillsunmatched
Computer Programmingunmatched
Computer Scienceunmatched
Continuous Deployment/Deliveryunmatched
Continuous Integrationunmatched
Distributed Computingunmatched
Engineeringunmatched
Experiment Designunmatched
Go Programming Language (Golang)unmatched
Javaunmatched
Memory Hardwareunmatched
Memory Managementunmatched
Modeling Languagesunmatched
Onboardingunmatched
Open Sourceunmatched
Operations Processesunmatched
Performance Managementunmatched
Product Engineeringunmatched
Production Systemsunmatched
Productivity Managementunmatched
Programming Toolsunmatched
Publicationsunmatched
Python Programming/Scripting Languageunmatched
Quality Engineeringunmatched
Reinforcement Learningunmatched
Research Skillsunmatched
Scalable System Developmentunmatched
Scientific Researchunmatched
Software Designunmatched
Software Engineeringunmatched
Software Testingunmatched
Systems Scalabilityunmatched
Team Playerunmatched
Test Plan/Scheduleunmatched
🎯
Be found by employers
5,500+ employers search our resume database daily. Add yours to get found by recruiters looking for candidates like you.
Level up your application
Professional resume templates
Browse dozens of recruiter approved resume templates, layouts and formats. Choose your favorite and make it your own in minutes.