AI Reliability Operations Engineer

Lenovo Group Ltd
  • Chicago, IL
    14 days ago

    Job Description

    General Information

    Req #

    100017497

    Career area:

    Engineering

    Country/Region:

    United States of America

    State:

    Illinois

    City:

    Chicago

    Date:

    Tuesday, August 18, 2026

    Additional Locations:

    • United States of America - Illinois - Chicago

    Why Work at Lenovo

    We are Lenovo. We do what we say. We own what we do. We WOW our customers.

    Lenovo is a US$83 billion revenue global technology powerhouse, ranked #153 in the Fortune Global 500, and serving millions of customers every day in 180 markets. Focused on a bold vision to deliver Smarter Technology for All, Lenovo has built on its success as the world's largest PC company with a full-stack portfolio of AI-enabled, AI-ready, and AI-optimized devices (PCs, workstations, smartphones, tablets), infrastructure (server, storage, edge, high performance computing and software defined infrastructure), software, solutions, and services. Lenovo's continued investment in world-changing innovation is building a more equitable, trustworthy, and smarter future for everyone, everywhere. Lenovo is listed on the Hong Kong stock exchange under Lenovo Group Limited (HKSE: 992) (ADR: LNVGY).

    To find out more visit www.lenovo.com and read about the latest news via our StoryHub.

    Description and Requirements

    About Our Team

    We are looking for an AI Reliability Operations Engineer to support the operational health of Qira''s production and non-production systems. Qira is Lenovo's cross-device Personal AI that works across phones, PCs, and other Lenovo and Motorola products. This role spans system monitoring, alert response, incident response, and observability across the full AI stack, including model performance, inference pipelines, and cloud services. You will also have visibility into SDLC

    operations across staging and pre-production environments, helping ensure that releases and configuration changes land cleanly. This is a foundational role in keeping Qira stable and available for users around the world.

    Location: Onsite in Chicago, IL (Hybrid, 3 days onsite, 2 days remote)

    What You''ll Do

    Operations Domain

    • Perform incident response: contain issues and work cross-functionally with dev teams to

    fully resolve them.

    • Monitor system health via Grafana dashboards, catching issues early and verifying resolution.
    • Serve as point of contact for change requests (CRs), triaging bug tickets from internal testers

    to the correct dev group.

    • Keep incident and CR records clear and accurate in ticketing systems.

    Monitoring and Observability

    • Monitor production and non-production systems using observability dashboards,

    alerting tools, and AI-specific signals, including model performance, inference latency,

    and data pipeline health.

    • Watch proactively for early warning signals across Qira''s cloud services, device

    integrations, and AI components, not just respond to alerts after they fire.

    • Review alert thresholds, update runbooks, and flag procedural gaps to the shift lead or

    SRE. This role is expected to improve the process, not just execute it.

    Build Tooling and Documentation

    • Build and improve tooling and scripts to reduce manual, repetitive work across incident

    and CR handling.

    • Maintain documentation standards: keep runbooks, CR records, and process docs

    accurate and usable by anyone on the team.

    Release Support

    • Observe and report on SDLC operations across staging and pre-production

    environments, flagging anomalies and supporting engineering teams during releases

    and configuration changes.

    • Verify system health before and after deployments and configuration changes, and

    assist engineering with deployment checks.

    Basic Qualifications

    • Direct experience in incident response or an SRE-adjacent role, not just monitoring or

    support.

    • Experience with observability tools such as Grafana, Datadog, or cloud-native

    dashboards.

    • Experience with alerting tools such as PagerDuty or OpsGenie, and ticketing systems

    such as Jira or ServiceNow.

    • Experience Troubleshooting: can isolate where a problem actually lives (which

    service, which layer) without being handed the answer.

    • Experience in Azure, including core cloud concepts and how services in an Azure

    environment are monitored.

    Preferred Qualifications

    • Experience in a technical operations, SRE, or production support environment.
    • Exposure to AI or ML systems, including awareness of how model quality and data pipelines

    are monitored.

    • Basic scripting ability or comfort reading and adapting existing scripts and runbook

    commands.

    • Experience working across time zones in a globally distributed team.
    • Working knowledge of SRE concepts: P50/P95/P99 latency, MTTA/MTTM/MTTR (MTTx),

    the four golden signals (latency, traffic, errors, saturation), and how to apply them to

    triage.

    • Clear, precise written communication in English, including accurate incident updates

    under pressure.

    • Ability to work assigned on-call coverage

    What Success Looks Like

    A successful AI Reliability Operations Engineer detects issues early, responds to alerts quickly,

    performs accurate initial triage, and keeps clear and complete records across production and

    non-production environments. Their work directly supports system uptime and ensures that Qira

    delivers a reliable, consistent experience for users at all times

    The base salary budgeted range for this position is $80K - $90K. Individuals may also be considered for bonus and/or commission.

    Lenovo's various benefits can be found on www.lenovobenefits.com.

    We are an Equal Opportunity Employer and do not discriminate against any employee or applicant for employment because of race, color, sex, age, religion, sexual orientation, gender identity, national origin, status as a veteran, and basis of disability or any federal, state, or local protected class.

    Additional Locations:

    • United States of America - Illinois - Chicago
    • United States of America
    • United States of America - Illinois
    • United States of America - Illinois - Chicago

    Numbers & Facts

    LocationChicago, IL

    Skills

    • Artificial Intelligence (AI)unmatched
    • Bug Tracking/Defect Managementunmatched
    • Change Requests/Ordersunmatched
    • Cloud Computingunmatched
    • Computer Softwareunmatched
    • Cross-Functionalunmatched
    • Data Managementunmatched
    • Data Modelingunmatched
    • Documentationunmatched
    • English Languageunmatched
    • Incident Responseunmatched
    • Infrastructure Softwareunmatched
    • Machine Toolunmatched
    • Operational Supportunmatched
    • Patient Assessmentunmatched
    • Performance Modelingunmatched
    • Process Improvementunmatched
    • Production Controlunmatched
    • Production Supportunmatched
    • Production Systemsunmatched
    • Reliability Engineeringunmatched
    • Reporting Dashboardsunmatched
    • Scripting (Scripting Languages)unmatched
    • Smartphonesunmatched
    • Software Development Lifecycle (SDLC)unmatched
    • Stock Marketunmatched
    • Systems Administration/Managementunmatched
    • Technical Deliveryunmatched
    • Technical Operationsunmatched
    • Technical Supportunmatched

    Be found by employers

    5,500+ employers search our resume database daily. Add yours to get found by recruiters looking for candidates like you.

    Level up your application

    Professional resume templates

    Browse dozens of recruiter approved resume templates, layouts and formats. Choose your favorite and make it your own in minutes.

    Free resume templates

    Free resume builder

    Improve your existing resume or start from scratch and create a standout, ATS-friendly resume. Add job-specific content, download and apply.

    Free resume builder