Machine Learning Operations (MLOps) Engineer

Openlane Inc
  • Carmel, IN
    1 day ago

    Job Description

    Who We Are:

    At OPENLANE we make wholesale easy so our customers can be more successful.

    We're a technology company building the world's most advanced-and uncomplicated-digital marketplace for used vehicles.

    We're a data company helping customers buy and sell smarter with clear, actionable insights they can understand and use.

    And we're an innovation company accelerating the future of wholesale remarketing through curiosity, collaboration, and an entrepreneurial spirit.

    Our Values:

    Driven Waybuilders. We pursue challenges that inspire us to build, create, and innovate.

    Relentless Curiosity. We seek to understand and improve our customers' experience.

    Smart Risk-Taking. We transform risk into progress through data, experience, and intuition.

    Fearless Ownership. We deliver what we promise and learn along the way.

    We're Looking For

    An Machine Learning Operations (MLOps) Engineer who builds and operates the pipelines, platforms, and observability that carry machine learning models from experimentation into reliable, cost-efficient production. You will join a new MLOps team within Site Reliability Engineering and help define how OPENLANE trains, ships, monitors, and safely updates models at scale. You will independently design automated ML delivery workflows, reduce operational toil, and ensure models are reproducible, monitored, and serving predictions safely in production.

    You Are

    Analytical, pragmatic, proactive, collaborative, quality-focused, and AI-forward. You treat AI as a first-class engineering tool, critically evaluate its output, and apply strong engineering judgment to systems that must stay healthy long after the first deployment.

    You Bring

    • Core knowledge of system architecture, distributed systems, cloud-native design, networking, and agile development methodologies (Scrum, Kanban).
    • Functional expertise in designing end-to-end ML pipelines (data ingestion, training, validation, deployment), including requirements for latency, redundancy, scalability, and error handling.
    • Technical expertise in CI/CD for ML, infrastructure-as-code, containerization and orchestration, and metrics, monitoring, and alerting toolsets.
    • Working understanding of the model lifecycle: experiment tracking and lineage, model registries, model drift, data quality, and retraining triggers, with enough ML fluency to diagnose model issues alongside data scientists.
    • Strong Python skills and the ability to write clean, well-tested, production-ready code.

    You Own

    • Design, build, and operation of automated ML pipelines covering data ingestion, training, validation, deployment, and rollback.
    • Model CI/CD, including automated testing and canary and rollback strategies for safe releases.
    • Monitoring and observability for model performance, data drift, inference latency, and availability, with actionable alerting and incident response.
    • Reproducibility of model workloads through versioned code, data, features, and environments.
    • Identification and reduction of operational toil through automation, standardization, and reusable ML platform components.
    • Documentation, testing, and debugging of pipelines and serving infrastructure, maintaining high code and design standards.

    You Drive

    Execution: Translate data science and product needs into scalable, automated ML delivery designs, and implement technical roadmap projects for the MLOps platform.

    Results: Shorten time-to-production for models while improving reliability, minimizing failure risk, and optimizing the cost, scaling, and utilization of training and inference workloads in partnership with Infrastructure teams.

    Team: Support hiring, knowledge sharing, and mentorship as the MLOps practice grows.

    How You'll Use AI in This Role

    • AI-assisted coding, refactoring, and test generation for pipelines and platform tooling.
    • AI-supported debugging and root-cause analysis of pipeline failures and model degradation.
    • AI-assisted documentation and runbooks to reduce repetitive work and increase engineering leverage.

    Who You Will Work With

    Reporting to the Director, Site Reliability Engineering, this role collaborates daily with Data Scientists, Data Engineers, Software Engineers, Infrastructure, Security, Product Owners, and SRE teams across the organization.

    Where You Work

    Your work is performed as a Remote employee, with occasional team alignment meetings as needed.

    Must Have's

    • 3+ years of experience in MLOps, ML engineering, Site Reliability Engineering, DevOps, or Software Engineering, including experience building or operating ML pipelines in production.
    • Bachelor's degree preferred or equivalent practical experience.
    • Strong Python skills and hands-on experience with a major cloud platform (AWS preferred).
    • Experience with workflow orchestration (e.g., Airflow, Kubeflow, Step Functions), infrastructure-as-code (e.g., Terraform), and CI/CD systems.
    • Experience with containers and Kubernetes (Docker, EKS preferred).
    • Familiarity with model serving frameworks (e.g., TorchServe, TF Serving, Triton, BentoML, SageMaker endpoints) and monitoring tools (e.g., Prometheus, Grafana, Evidently AI).
    • Hands-on experience with system debugging, observability, and incident response in production environments.
    • Actively uses AI development tools and can critically evaluate AI outputs.

    Nice to Have's

    • Experience with ML platforms and tooling such as SageMaker, MLflow, Databricks, or Vertex AI (GCP).
    • Experience with feature stores, model registries, or LLM/GenAI deployment and evaluation (LLMOps).
    • Experience with GPU workload scheduling and inference cost optimization.
    • Experience interviewing candidates, conducting technical debriefs, and mentoring junior engineers or interns.
    • Proven ability to acquire deep domain and architectural knowledge within 6 months of joining a team.

    What We Offer:

    • Competitive pay

    • Medical, dental, and vision benefits with employer HSA contributions (US) and FSA options (US)

    • Immediately vested 401K (US) or RRSP (Canada) with company match

    • Paid Vacation, Personal, and Sick Time

    • Paid maternity and paternity leave (US)

    • Employer-paid short-term disability, long-term disability, life insurance, and AD&D (US)

    • Robust Employee Assistance Program

    • Employer paid Leap into Service Day to volunteer

    • Tuition Reimbursement for eligible programs

    • Opportunities to expand your skill set and share your knowledge across a publicly traded, global organization

    • Company culture of internal promotions, diverse career paths, and meaningful advancement

    Sound like a match? Apply Now - We can't wait to hear from you!

    Numbers & Facts

    LocationCarmel, IN

    Skills

    • Accidental Death and Dismemberment (AD&D)unmatched
    • Agile Programming Methodologiesunmatched
    • Amazon Web Services (AWS)unmatched
    • Analysis Skillsunmatched
    • Architectural Servicesunmatched
    • Artificial Intelligence (AI)unmatched
    • Automationunmatched
    • Cloud Computingunmatched
    • Continuous Deployment/Deliveryunmatched
    • Continuous Integrationunmatched
    • Cost Controlunmatched
    • Customer Experienceunmatched
    • Customer Support/Serviceunmatched
    • Data Modelingunmatched
    • Data Qualityunmatched
    • Data Scienceunmatched
    • Debugging Skillsunmatched
    • DevOpsunmatched
    • Disability Insuranceunmatched
    • Distributed Computingunmatched
    • Dockerunmatched
    • Documentationunmatched
    • Entrepreneurshipunmatched
    • Error Handlingunmatched
    • Failure Analysisunmatched
    • Flexible Spending Accountsunmatched
    • GCP (Good Clinical Practices)unmatched
    • GPU (Graphics Processing Unit)unmatched
    • Identify Issuesunmatched
    • Incident Responseunmatched
    • Interviewing Skillsunmatched
    • Kanbanunmatched
    • Life Insuranceunmatched
    • Machine Learningunmatched
    • Machine Toolunmatched
    • Machining Operationsunmatched
    • Mentoringunmatched
    • Metricsunmatched
    • Network Designunmatched
    • Performance Modelingunmatched
    • Production Costingunmatched
    • Production Systemsunmatched
    • Programming Toolsunmatched
    • Python Programming/Scripting Languageunmatched
    • Refactoringunmatched
    • Reliability Engineeringunmatched
    • Riskunmatched
    • Risk Managementunmatched
    • Root Cause Analysisunmatched
    • Salesunmatched
    • Scrum Project Management and Software Developmentunmatched
    • Software Engineeringunmatched
    • System Architectureunmatched
    • Systems Engineeringunmatched
    • Team Playerunmatched
    • Technical/Engineering Designunmatched
    • Test Automationunmatched
    • Testingunmatched
    • Wholesale Industryunmatched
    • Writing Skillsunmatched

    Be found by employers

    5,500+ employers search our resume database daily. Add yours to get found by recruiters looking for candidates like you.

    Level up your application

    Professional resume templates

    Browse dozens of recruiter approved resume templates, layouts and formats. Choose your favorite and make it your own in minutes.

    Free resume templates

    Free resume builder

    Improve your existing resume or start from scratch and create a standout, ATS-friendly resume. Add job-specific content, download and apply.

    Free resume builder