Real-Time Inference Engineering Lead (FTE / Hybrid)

NTT DATA, Inc.

Charlotte (NC)

Hybrid

USD 84,000 - 125,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Medical, dental, and vision insurance
401k with company match
Paid time off
Employee assistance program
Flexible spending or health savings

Job summary

NTT DATA Services in Charlotte, NC seeks a Real-Time Inference Engineering Lead (FTE / Hybrid) to design, build, and industrialize low-latency, resilient model-serving services for real-time predictive AI use cases. The role combines hands-on engineering with technical leadership across online inference architecture, deployment patterns, APIs, and operations.

The successful candidate will establish reusable patterns enabling multiple teams to deploy and operate predictive models safely and at

Qualifications

  • 8+ years of software engineering, platform engineering, or related infra experience.
  • 4+ years of building or operating production APIs, distributed systems, or real-time ML workloads.
  • Experience delivering highly available, performance-sensitive services.

Responsibilities

  • Define target architecture and standards for real-time model-serving across cloud and on-premises.
  • Design, build, test, deploy, and operate scalable online inference services with low latency.
  • Establish reusable model-serving patterns for synchronous and asynchronous inference.
  • Develop deployment approaches for model packaging, versioning, and canary deployments.
  • Design secure API patterns for inference services, including auth, authz, and rate limiting.
  • Lead Kubernetes-based serving platforms using GKE and OpenShift.
  • Implement autoscaling, capacity planning, and workload isolation for variable demand.
  • Establish monitoring, dashboards, SLIs/SLOs, and incident-management practices.

Skills

Online inference architecture
Model-serving patterns
REST APIs
Kubernetes
CI/CD
Observability
Security patterns
Performance optimization
Distributed systems

Tools

GKE
OpenShift

Job description

Real-Time Inference Engineering Lead (FTE / Hybrid)

Date: Aug 28, 2026

Location: Charlotte, NC, US

Company: NTT DATA Services

NTT DATA strives to hire exceptional, innovative and passionate individuals who want to grow with us. If you want to be part of an inclusive, adaptable, and forward-thinking organization,

We are currently seeking a Real-Time Inference Engineering Lead (FTE / Hybrid) to join our team in Charlotte, North Carolina (US-NC), United States (US).

Job Duties and Responsibilities:

Platform Context

The Cortex Predictive AI Platform accelerates predictive AI modernization and enterprise adoption across the full model lifecycle: governed data and features; model build, training, and validation; deployment and inference; and ongoing monitoring and operations. The Real-Time Services portfolio provides standardized, scalable model-serving capabilities for applications requiring reliable, performant online inference.

Position Summary

The Real-Time Inference Engineering Lead will design, build, and industrialize low-latency, resilient model-serving services for real-time predictive AI use cases. This role provides technical leadership for online inference architecture, deployment patterns, API services, capacity controls, observability, and operational practices across public cloud and on-premises environments.

The successful candidate will establish reusable patterns that enable application, data science, and ML engineering teams to deploy and operate predictive models safely and efficiently at scale. This is a hands-on engineering role requiring strong experience with model serving, Kubernetes, APIs, performance optimization, reliability engineering, CI/CD, and production operations.

Key Responsibilities

  • Define the target architecture and engineering standards for real-time predictive model-serving services across cloud and on-premises environments.
  • Design, build, test, deploy, and operate scalable online inference services that meet latency, throughput, availability, resiliency, and security requirements.
  • Establish reusable model-serving patterns for synchronous APIs, asynchronous inference, batch-adjacent processing, and event-driven real-time use cases where appropriate.
  • Build standardized deployment approaches for predictive models, including model packaging, versioning, release promotion, canary deployment, rollback, and retirement.
  • Design and implement secure API patterns for inference services, including authentication, authorization, traffic management, rate limiting, auditability, and integration with enterprise systems.
  • Engineer Kubernetes-based serving platforms using GKE, OpenShift, and related container orchestration capabilities.
  • Implement autoscaling, resource allocation, quota management, capacity planning, and workload-isolation controls for variable inference demand.
  • Conduct performance engineering, load testing, stress testing, and failure testing to validate service behavior under expected and peak production workloads.
  • Identify and implement latency-optimization opportunities across model initialization, feature retrieval, network paths, API handling, runtime configuration, and infrastructure utilization.
  • Define and implement monitoring, telemetry, dashboards, alerts, SLIs, SLOs, and error-budget practices for real-time inference services.
  • Partner with ML platform, data engineering, application engineering, security, and operations teams to integrate model services with governed data, feature, network, and identity capabilities.
  • Implement CI/CD and automated validation for model-serving services, infrastructure configuration, APIs, performance benchmarks, and release-readiness checks.
  • Build operational runbooks, incident-response procedures, support models, and production-readiness artifacts for real-time services.
  • Drive reliability improvements through root-cause analysis, capacity reviews, resiliency testing, disaster-recovery planning, and continuous operational improvement.
  • Mentor engineers and establish reusable technical documentation, reference implementations, and knowledge-transfer materials for real-time inference capabilities.

Required Qualifications

  • 8+ years of software engineering, platform engineering, cloud engineering, SRE, or infrastructure engineering experience.
  • 4+ years of experience designing, building, or operating production APIs, distributed systems, platform services, or real-time data and ML workloads.
  • Demonstrated experience leading technical design and engineering delivery for highly available, performance-sensitive production services.
  • Strong experience with online inference architecture, model-serving frameworks, or predictive-model deployment patterns.
  • Hands-on experience designing and operating RESTful, gRPC, or event-driven APIs.
  • 4+ years satrong experience with Kubernetes and container platforms in production, including GKE, OpenShift, or comparable environments.
  • Experience with autoscaling, resource management, capacity planning, performance testing, and load testing for distributed services.
  • Experience implementing observability, monitoring, dashboards, alerts, SLIs, SLOs, and incident-management practices.
  • Experience with CI/CD, Git-based development, automated testing, deployment automation, and production-release practices.
  • Strong understanding of resiliency, high availability, fault tolerance, disaster recovery, and operational support for critical services.
  • Ability to work effectively with data science, ML engineering, platform engineering, application teams, security, and business stakeholders.

Required Skills / Knowledge

  • Online inference and low-latency model-serving architecture.
  • Model deployment, versioning, routing, rollout, rollback, and lifecycle management.
  • REST APIs, gRPC, API gateways, authentication, authorization, traffic management, and API observability.
  • Kubernetes, GKE, OpenShift, containers, service meshes, ingress, workload scheduling, and autoscaling.
  • Performance engineering, load testing, stress testing, benchmarking, profiling, and latency optimization.
  • Monitoring, telemetry, distributed tracing, dashboards, alerting, SLIs, SLOs, and error budgets.
  • CI/CD, automated testing, deployment automation, infrastructure-as-code, and release controls.
  • Resiliency engineering, high availability, capacity controls, incident response, root-cause analysis, and operational runbooks.
  • Cloud and on-premises platform operations, networking, identity, data protection, and secure production delivery.

Preferred Qualifications

  • Experience with Vertex AI endpoints, KServe, Seldon, NVIDIA Triton Inference Server, MLflow deployments, or comparable model-serving technologies.
  • Experience deploying and operating models on GCP, Azure, AWS, private cloud, or hybrid-cloud environments.
  • Experience with service mesh, API gateway, traffic-routing, or edge-serving technologies.
  • Experience serving high-volume, customer-facing, fraud, risk, personalization, decisioning, or other latency-sensitive predictive models.
  • Experience with feature-serving, online feature stores, caching, streaming platforms, or real-time data enrichment.
  • Experience with Terraform, Helm, Argo CD, Jenkins, GitHub Actions, GitLab CI, or similar automation tooling.
  • Experience in banking, financial services, healthcare, insurance, or another regulated enterprise environment.
  • Experience participating in a 24x7 operational support model for high-priority production services.

Expected Outcomes

  • Standardized, production-ready real-time inference architecture and reusable model-serving patterns.
  • Reliable online inference services that meet defined latency, throughput, availability, and resiliency objectives.
  • Automated deployment, testing, monitoring, capacity-management, and rollback capabilities for predictive models.
  • Clear operational dashboards, SLOs, alerts, runbooks, and readiness evidence for real-time services.
  • Improved engineering productivity and faster adoption of secure, scalable real-time predictive AI capabilities across the Cortex portfolio.

#LI-NorthAmerica

NTT DATA provides a reasonable range of compensation for U.S.-based positions. The starting pay range for this role is $83,520.00 - $125,280.00. Actual compensation will depend on a number of factors, including the candidate’s relevant experience, technical skills, and other qualifications.

This position may also be eligible for incentive compensation based on individual and/or company performance.

This position is eligible for company benefits including medical, dental, and vision insurance with an employer contribution, flexible spending or health savings account, lifeand AD&D insurance, short and long term disability coverage, paid time off, employee assistance, participation in a 401k program with company match, and additional voluntary or legally-required benefits.

About NTT DATA

NTT DATA is a $30 billion business and technology services leader, serving 75% of the Fortune Global 100. We are committed to accelerating client success and positively impacting society through responsible innovation. We are one of the world's leading AI and digital infrastructure providers, with unmatched capabilities in enterprise-scale AI, cloud, security, connectivity, data centers and application services. our consulting and Industry solutions help organizations and society move confidently and sustainably into the digital future. As a Global Top Employer, we have experts in more than 50 countries. We also offer clients access to a robust ecosystem of innovation centers as well as established and start-up partners. NTT DATA is a part of NTT Group, which invests over $3 billion each year in R&D.


Nearest Major Market: Charlotte
Job Segment: Testing, Cloud, Consulting, Technology

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Real-Time Inference Engineering Lead (FTE / Hybrid)
Real-Time Inference Engineering Lead (FTE / Hybrid)

NTT DATA North America • Charlotte (NC)

Hybrid
USD 84,000 - 125,000
Medical, dental, vision insurance
401k with company match
Paid time off
Lead ML Platform Engineer
Lead ML Platform Engineer

Creative Solutions Services, LLC • Charlotte (NC)

On-site
USD 80,000 - 89,000
Medical insurance
Dental insurance
Vision insurance
+3
Lead ML Platform Engineer (SRE / FTE / Onsite)
Lead ML Platform Engineer (SRE / FTE / Onsite)

NTT DATA, Inc. • Charlotte (NC)

On-site
USD 84,000 - 125,000
Medical Insurance
Dental Insurance
Vision Insurance
+6
Lead ML Platform Engineer (SRE / FTE / Onsite)
Lead ML Platform Engineer (SRE / FTE / Onsite)

NTT DATA North America • Charlotte (NC)

On-site
USD 84,000 - 125,000
Medical, dental, and vision insurance
401(k) with company match
Paid time off
Lead ML Platform Engineer
Lead ML Platform Engineer

NTT DATA North America • Charlotte (NC)

On-site
USD 150,000 - 230,000
Real-Time Inference Platform Lead
Real-Time Inference Platform Lead

Creative Solutions Services, LLC • Charlotte (NC)

On-site
USD 80,000 - 89,000
Medical insurance
Dental insurance
Vision insurance
+3
Real-Time Inference Platform Lead
Real-Time Inference Platform Lead

NTT DATA North America • Charlotte (NC)

Hybrid
USD 84,000 - 125,000
Medical, dental, vision insurance
401k with company match
Paid time off
Platform Analyst
Platform Analyst

NTT DATA North America • Charlotte (NC)

Hybrid
USD 94,000 - 109,000
Medical, dental, vision
401(k) plan
AD&D insurance
+3
Principal Engineer
Principal Engineer

NTT DATA North America • Charlotte (NC)

Hybrid
USD 96,000 - 110,000
Medical, dental, vision insurance
401(k) program
AD&D insurance
Java and Python Developer (FTE / Hybrid)
Java and Python Developer (FTE / Hybrid)

NTT DATA, Inc. • Charlotte (NC)

Hybrid
USD 97,000 - 145,000
Medical, dental, and vision insurance
HSA/FSAs
Life & AD&D insurance
+4