Site Reliability Engineer (SRE) Intern — AI Infrastructure

Tencent
  • Palo Alto, California
    8 days ago

    Job Description

    About the Hiring Team

    What the Role Entails

    Role Summary
    We are seeking a motivated Site Reliability Engineer (SRE) Intern to join our AI Compute team, supporting the daily operations of AI infrastructure. In this role, you will work closely with internal business teams and external engineering partners to build and operate AI infrastructure. This is a hands-on opportunity to gain knowledge and experience in cutting-edge AI infrastructure.

    Key Responsibilities

    • Support the deployment, configuration, and maintenance of high-performance AI infrastructure servers, storage servers, networking equipment, and software components in secure environments.
    • Assist with hardware diagnostics, system functionality checks, and firmware updates as required.
    • Collaborate with engineering teams to help deliver tailored customer environments (e.g., bare-metal systems, Kubernetes, Slurm, etc.).
    • Provide first-line engineering support for onsite operational issues, including troubleshooting hardware, network, and software problems, and firmware compliance.
    • Document incident details, resolutions, and lessons learned to improve future problem-solving.
    • Maintain clear, accurate, and up-to-date documentation to support knowledge sharing across the team.
    • Participate in team meetings and knowledge-sharing sessions to foster collaboration and continuous learning.

    Who We Look For

    Qualifications & Requirements

    • Currently pursuing or recently completed a Bachelor's or Master's degree in computer engineering, computer science, or a related technical field.
    • Basic understanding of server hardware, firmware lifecycle, and Linux environments, with an awareness of physical and system-level security standards.
    • Exposure to scripting languages such as Bash or Python.
    • Familiarity with — or strong interest in — configuration management, CI/CD tools, workload managers, and cluster software (e.g., Slurm, Kubernetes), and observability tools (e.g., Prometheus, Grafana, ELK).
    • Strong problem-solving and analytical skills.
    • Ability to work both independently and as part of a team.
    • Professional fluency in English and Mandarin is highly preferred

    Preferred Qualifications

    • Coursework, projects, or hands-on experience related to AI Infrastructure, distributed systems, or cloud infrastructure.
    • Familiarity with networking fundamentals and Linux system administration.
    • Genuine interest in AI/ML infrastructure.

    Location State(s)

    US-California-Palo Alto

    The expected base pay range for this position in the location(s) listed above is $27.12 to $51.93 per hour. Actual pay may vary depending on job-related knowledge, skills, and experience. This position will be eligible for 1 hour of paid sick leave for every 30 hours worked and up to 13 paid holidays throughout the calendar year. Subject to the terms and conditions of the applicable plans then in effect, full-time interns are also eligible to enroll in the Company-sponsored medical plan.

    Equal Employment Opportunity at Tencent

    As an equal opportunity employer, we firmly believe that diverse voices fuel our innovation and allow us to better serve our users and the community. We foster an environment where every employee of Tencent feels supported and inspired to achieve individual and common goals.

    Numbers & Facts

    LocationPalo Alto, California

    Skills

    • 1st Level Supportunmatched
    • Analysis Skillsunmatched
    • Artificial Intelligence (AI)unmatched
    • Bash Scriptingunmatched
    • Cloud Computingunmatched
    • Computer Engineeringunmatched
    • Computer Firmwareunmatched
    • Computer Scienceunmatched
    • Configuration Managementunmatched
    • Continuous Deployment/Deliveryunmatched
    • Continuous Integrationunmatched
    • Distributed Computingunmatched
    • English Languageunmatched
    • Health Planunmatched
    • Linux Administrationunmatched
    • Linux Operating Systemunmatched
    • Mandarin Chinese Languageunmatched
    • Network Administration/Managementunmatched
    • Network Softwareunmatched
    • Network System Hardwareunmatched
    • Operational Supportunmatched
    • Operationsunmatched
    • Problem Solving Skillsunmatched
    • Python Programming/Scripting Languageunmatched
    • Reliability Engineeringunmatched
    • Scripting (Scripting Languages)unmatched
    • Server Hardwareunmatched
    • Support Documentationunmatched

    Be found by employers

    5,500+ employers search our resume database daily. Add yours to get found by recruiters looking for candidates like you.

    Level up your application

    Professional resume templates

    Browse dozens of recruiter approved resume templates, layouts and formats. Choose your favorite and make it your own in minutes.

    Free resume templates

    Free resume builder

    Improve your existing resume or start from scratch and create a standout, ATS-friendly resume. Add job-specific content, download and apply.

    Free resume builder