NTT DATA logo

Real-Time Inference Engineering Lead

NTT DATA
  • Charlotte, NC
  • $70–$80 Per Hour
  • Quick Apply
5 days ago

Job Description

Req ID: 388176
NTT DATA strives to hire exceptional, innovative and passionate individuals who want to grow with us. If you want to be part of an inclusive, adaptable, and forward-thinking organization, apply now.

We are currently seeking a Real-Time Inference Engineering Lead to join our team in Charlotte, North Carolina (US-NC), United States (US).

Job Description:
Platform Context
The Cortex Predictive AI Platform accelerates predictive AI modernization and enterprise adoption across the full model lifecycle: governed data and features; model build, training, and validation; deployment and inference; and ongoing monitoring and operations. The Real-Time Services portfolio provides standardized, scalable model-serving capabilities for applications requiring reliable, performant online inference.

Position Summary
The Real-Time Inference Engineering Lead will design, build, and industrialize low-latency, resilient model-serving services for real-time predictive AI use cases. This role provides technical leadership for online inference architecture, deployment patterns, API services, capacity controls, observability, and operational practices across public cloud and on-premises environments.
The successful candidate will establish reusable patterns that enable application, data science, and ML engineering teams to deploy and operate predictive models safely and efficiently at scale. This is a hands-on engineering role requiring strong experience with model serving, Kubernetes, APIs, performance optimization, reliability engineering, CI/CD, and production operations.

Key Responsibilities
  • Define the target architecture and engineering standards for real-time predictive model-serving services across cloud and on-premises environments.
  • Design, build, test, deploy, and operate scalable online inference services that meet latency, throughput, availability, resiliency, and security requirements.
  • Establish reusable model-serving patterns for synchronous APIs, asynchronous inference, batch-adjacent processing, and event-driven real-time use cases where appropriate.
  • Build standardized deployment approaches for predictive models, including model packaging, versioning, release promotion, canary deployment, rollback, and retirement.
  • Design and implement secure API patterns for inference services, including authentication, authorization, traffic management, rate limiting*** auditability, and integration with enterprise systems.
  • Engineer Kubernetes-based serving platforms using GKE, OpenShift, and related container orchestration capabilities.
  • Implement autoscaling, resource allocation, quota management, capacity planning, and workload-isolation controls for variable inference demand.
  • Conduct performance engineering, load testing, stress testing, and failure testing to validate service behavior under expected and peak production workloads.
  • Identify and implement latency-optimization opportunities across model initialization, feature retrieval, network paths, API handling, runtime configuration, and infrastructure utilization.
  • Define and implement monitoring, telemetry, dashboards, alerts, SLIs, SLOs, and error-budget practices for real-time inference services.
  • Partner with ML platform, data engineering, application engineering, security, and operations teams to integrate model services with governed data, feature, network, and identity capabilities.
  • Implement CI/CD and automated validation for model-serving services, infrastructure configuration, APIs, performance benchmarks, and release-readiness checks.
  • Build operational runbooks, incident-response procedures, support models, and production-readiness artifacts for real-time services.
  • Drive reliability improvements through root-cause analysis, capacity reviews, resiliency testing, disaster-recovery planning, and continuous operational improvement.
  • Mentor engineers and establish reusable technical documentation, reference implementations, and knowledge-transfer materials for real-time inference capabilities.
Required Qualifications
  • 8+ years of software engineering, platform engineering, cloud engineering, SRE, or infrastructure engineering experience.
  • 4+ years of experience designing, building, or operating production APIs, distributed systems, platform services, or real-time data and ML workloads.
  • Demonstrated experience leading technical design and engineering delivery for highly available, performance-sensitive production services.
  • Strong experience with online inference architecture, model-serving frameworks, or predictive-model deployment patterns.
  • Hands-on experience designing and operating RESTful, gRPC, or event-driven APIs.
  • 4+ years satrong experience with Kubernetes and container platforms in production, including GKE, OpenShift, or comparable environments.
  • Experience with autoscaling, resource management, capacity planning, performance testing, and load testing for distributed services.
  • Experience implementing observability, monitoring, dashboards, alerts, SLIs, SLOs, and incident-management practices.
  • Experience with CI/CD, Git-based development, automated testing, deployment automation, and production-release practices.
  • Strong understanding of resiliency, high availability, fault tolerance, disaster recovery, and operational support for critical services.
  • Ability to work effectively with data science, ML engineering, platform engineering, application teams, security, and business stakeholders.
Required Skills / Knowledge
  • Online inference and low-latency model-serving architecture.
  • Model deployment, versioning, routing, rollout, rollback, and lifecycle management.
  • REST APIs, gRPC, API gateways, authentication, authorization, traffic management, and API observability.
  • Kubernetes, GKE, OpenShift, containers, service meshes, ingress, workload scheduling, and autoscaling.
  • Performance engineering, load testing, stress testing, benchmarking, profiling, and latency optimization.
  • Monitoring, telemetry, distributed tracing, dashboards, alerting, SLIs, SLOs, and error budgets.
  • CI/CD, automated testing, deployment automation, infrastructure-as-code, and release controls.
  • Resiliency engineering, high availability, capacity controls, incident response, root-cause analysis, and operational runbooks.
  • Cloud and on-premises platform operations, networking, identity, data protection, and secure production delivery.
Preferred Qualifications
  • Experience with Vertex AI endpoints, KServe, Seldon, NVIDIA Triton Inference Server, MLflow deployments, or comparable model-serving technologies.
  • Experience deploying and operating models on GCP, Azure, AWS, private cloud, or hybrid-cloud environments.
  • Experience with service mesh, API gateway, traffic-routing, or edge-serving technologies.
  • Experience serving high-volume, customer-facing, fraud, risk, personalization, decisioning, or other latency-sensitive predictive models.
  • Experience with feature-serving, online feature stores, caching, streaming platforms, or real-time data enrichment.
  • Experience with Terraform, Helm, Argo CD, Jenkins, GitHub Actions, GitLab CI, or similar automation tooling.
  • Experience in banking, financial services, healthcare, insurance, or another regulated enterprise environment.
  • Experience participating in a 24x7 operational support model for high-priority production services.
  • Expected Outcomes
  • Standardized, production-ready real-time inference architecture and reusable model-serving patterns.
  • Reliable online inference services that meet defined latency, throughput, availability, and resiliency objectives.
  • Automated deployment, testing, monitoring, capacity-management, and rollback capabilities for predictive models.
  • Clear operational dashboards, SLOs, alerts, runbooks, and readiness evidence for real-time services.
  • Improved engineering productivity and faster adoption of secure, scalable real-time predictive AI capabilities across the Cortex portfolio.
Where required by law, NTT DATA provides a reasonable range of compensation for specific roles. The starting hourly range for this remote role is ($70-$80/hour). This range reflects the minimum and maximum target compensation for the position across all US locations. Actual compensation will depend on several factors, including the candidate's actual work location, relevant experience, technical skills, and other qualifications. This position may also be eligible for incentive compensation based on individual and/or company performance.

This position is eligible for company benefits that will depend on the nature of the role offered. Company benefits may include medical, dental, and vision insurance, flexible spending or health savings account, life, and AD&D insurance, short-and long-term disability coverage, paid time off, employee assistance, participation in a 401k program with company match, and additional voluntary or legally required benefits .




About NTT DATA:

NTT DATA is a $30 billion trusted global innovator of business and technology services. We serve 75% of the Fortune Global 100 and are committed to helping clients innovate, optimize and transform for long term success. As a Global Top Employer, we have diverse experts in more than 50 countries and a robust partner ecosystem of established and start-up companies. Our services include business and technology consulting, data and artificial intelligence, industry solutions, as well as the development, implementation and management of applications, infrastructure and connectivity. We are one of the leading providers of digital and AI infrastructure in the world. NTT DATA is a part of NTT Group, which invests over $3.6 billion each year in R&D to help organizations and society move confidently and sustainably into the digital future. Visit us at us.nttdata.com

NTT DATA endeavors to make https://us.nttdata.com accessible to any and all users. If you would like to contact us regarding the accessibility of our website or need assistance completing the application process, please contact us at https://us.nttdata.com/en/contact-us. This contact information is for accommodation requests only and cannot be used to inquire about the status of applications. NTT DATA is an equal opportunity employer. Qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, disability or protected veteran status. For our EEO Policy Statement, please click here. If you'd like more information on your EEO rights under the law, please click here. For Pay Transparency information, please click here.

Numbers & Facts

LocationCharlotte, NC
IndustryManagement Consulting Services
Salary$70–$80 Per Hour
Company Size10,000 employees or more
Year Founded1967
Websitehttp://americas.nttdata.com/Careers/Careers.aspx

Benefits

Retirement / Pension Plans

About Company

NTT DATA means Business

NTT DATA is your Innovation Partner anywhere around the world. With business operations in more than 50 countries, we put emphasis on long-term commitment and combine global reach and local intimacy to provide premier professional services from consulting, system development, business process and IT outsourcing, to cloud-based solutions.

NTT DATA Americas' Strategic Staffing group provides our clients with top notch technical talent to augment their core IT staff. Our approach is customer centric, partnering to assist you in achieving your strategic goals and IT initiatives. We get to know your company’s culture and the type of technical staff that thrive within your organization. We understand your specific technical and business requirements, timing, and budget.

  • NTT DATA is part of the NTT Group – a Fortune 31 Global IT & Telecom services company.
  • NTT Group one of the largest Telecommunications Companies in the world.
  • NTT DATA is ranked in the top 10 largest global IT services provider in the world.
Collectively, the integrated company generates $16B in annual revenues with over 130,000 employees across 50 countries Visit www.nttdata.com/americas to learn how our consultants, projects, managed services, and outsourcing engagements deliver value for a range of businesses and government agencies.

Skills

  • Amazon Web Services (AWS)unmatched
  • Application Programming Interface (API)unmatched
  • Applications Securityunmatched
  • Artificial Intelligence (AI)unmatched
  • Authenticationunmatched
  • Automationunmatched
  • Autoscalingunmatched
  • Banking Servicesunmatched
  • Benchmarkingunmatched
  • Budgetingunmatched
  • Business Operationsunmatched
  • Business Servicesunmatched
  • Cachingunmatched
  • Capacity Managementunmatched
  • Capacity and Performance Managementunmatched
  • Cloud Computingunmatched
  • Consultingunmatched
  • Continuous Deployment/Deliveryunmatched
  • Continuous Integrationunmatched
  • Customer Relationsunmatched
  • Data Scienceunmatched
  • Disaster Recoveryunmatched
  • Distributed Computingunmatched
  • Ecosystemsunmatched
  • Engineeringunmatched
  • Failure Analysisunmatched
  • Financial Servicesunmatched
  • GCP (Good Clinical Practices)unmatched
  • Gitunmatched
  • GitHubunmatched
  • Health Insuranceunmatched
  • High Availabilityunmatched
  • Hybrid Cloudunmatched
  • Incident Managementunmatched
  • Incident Responseunmatched
  • Information/Data Security (InfoSec)unmatched
  • International Businessunmatched
  • Internet/Online Serviceunmatched
  • Jenkinsunmatched
  • Knowledge Transferunmatched
  • Leadershipunmatched
  • Legalunmatched
  • Load Testingunmatched
  • Machine Toolunmatched
  • Mentoringunmatched
  • Microsoft Windows Azureunmatched
  • Model Validationunmatched
  • Operational Auditunmatched
  • Operational Improvementunmatched
  • Operational Supportunmatched
  • Operations Planningunmatched
  • Performance Engineeringunmatched
  • Performance Testingunmatched
  • Performance Tuning/Optimizationunmatched
  • Predictive Modelingunmatched
  • Private Cloudunmatched
  • Process Improvementunmatched
  • Production Supportunmatched
  • Productivity Managementunmatched
  • Public Cloudunmatched
  • REST (Representational State Transfer)unmatched
  • Reliability Engineeringunmatched
  • Reporting Dashboardsunmatched
  • Research & Development (R&D)unmatched
  • Resource Managementunmatched
  • Riskunmatched
  • Root Cause Analysisunmatched
  • Software Engineeringunmatched
  • Startupunmatched
  • Stress Testingunmatched
  • Technical Leadershipunmatched
  • Technical Writingunmatched
  • Technical/Engineering Designunmatched
  • Telemetryunmatched
  • Test Automationunmatched
  • Testingunmatched
  • Traffic Shapingunmatched
  • Use Casesunmatched

Be found by employers

5,500+ employers search our resume database daily. Add yours to get found by recruiters looking for candidates like you.

Level up your application

Professional resume templates

Browse dozens of recruiter approved resume templates, layouts and formats. Choose your favorite and make it your own in minutes.

Free resume templates

Free resume builder

Improve your existing resume or start from scratch and create a standout, ATS-friendly resume. Add job-specific content, download and apply.

Free resume builder