Real-Time Inference Platform Lead

NTT DATA North America

Charlotte (NC)

Hybrid

USD 84,000 - 125,000

Full time

9 days ago
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Medical, dental, vision insurance
401k with company match
Paid time off

Job summary

NTT DATA North America in Charlotte, NC seeks a Real-Time Inference Engineering Lead to design, build, and operate low-latency model-serving services across cloud and on-prem environments. You will lead online inference architecture, deployment patterns, APIs, observability, and production operations in a hands-on role.

The role requires strong experience with model serving, Kubernetes, REST/gRPC APIs, performance optimization, and CI/CD.

Qualifications

  • 8+ years of software engineering, platform engineering, cloud engineering, SRE, or infrastructure engineering experience.
  • 4+ years of experience designing, building, or operating production APIs, distributed systems, platform services, or real-time data and ML workloads.
  • Demonstrated experience leading technical design and engineering delivery for highly available, performance-sensitive production services.
  • Strong experience with online inference architecture, model-serving frameworks, or predictive-model deployment patterns.
  • Hands-on experience designing and operating RESTful, gRPC, or event-driven APIs.
  • 4+ years satrong experience with Kubernetes and container platforms in production, including GKE, OpenShift, or comparable environments.
  • Experience with autoscaling, resource management, capacity planning, performance testing, and load testing for distributed services.
  • Experience implementing observability, monitoring, dashboards, alerts, SLIs, SLOs, and incident-management practices.
  • Experience with CI/CD, Git-based development, automated testing, deployment automation, and production-release practices.
  • Strong understanding of resiliency, high availability, fault tolerance, disaster recovery, and operational support for critical services.
  • Ability to work effectively with data science, ML engineering, platform engineering, application teams, security, and business stakeholders.

Responsibilities

  • Define the target architecture and engineering standards for real-time predictive model-serving services across cloud and on-premises environments.
  • Design, build, test, deploy, and operate scalable online inference services that meet latency, throughput, availability, resiliency, and security requirements.
  • Establish reusable model-serving patterns for synchronous APIs, asynchronous inference, batch-adjacent processing, and event-driven real-time use cases where appropriate.
  • Build standardized deployment approaches for predictive models, including model packaging, versioning, release promotion, canary deployment, rollback, and retirement.
  • Design and implement secure API patterns for inference services, including authentication, authorization, traffic management, rate limiting, auditability, and integration with enterprise systems.
  • Engineer Kubernetes-based serving platforms using GKE, OpenShift, and related container orchestration capabilities.
  • Implement autoscaling, resource allocation, quota management, capacity planning, and workload-isolation controls for variable inference demand.
  • Conduct performance engineering, load testing, stress testing, and failure testing to validate service behavior under expected and peak production workloads.
  • Identify and implement latency-optimization opportunities across model initialization, feature retrieval, network paths, API handling, runtime configuration, and infrastructure utilization.
  • Define and implement monitoring, telemetry, dashboards, alerts, SLIs, SLOs, and error-budget practices for real-time inference services.
  • Partner with ML platform, data engineering, application engineering, security, and operations teams to integrate model services with governed data, feature, network, and identity capabilities.
  • Implement CI/CD and automated validation for model-serving services, infrastructure configuration, APIs, performance benchmarks, and release-readiness checks.
  • Build operational runbooks, incident-response procedures, support models, and production-readiness artifacts for real-time services.
  • Drive reliability improvements through root-cause analysis, capacity reviews, resiliency testing, disaster-recovery planning, and continuous operational improvement.
  • Mentor engineers and establish reusable technical documentation, reference implementations, and knowledge-transfer materials for real-time inference capabilities.

Skills

Online inference architecture
Model-serving
Kubernetes
REST/gRPC APIs
CI/CD
Observability
Performance optimization
Cloud/on-prem deployment

Tools

GKE
OpenShift

Job description

NTT DATA North America in Charlotte, NC seeks a Real-Time Inference Engineering Lead to design, build, and operate low-latency model-serving services across cloud and on-prem environments. You will lead online inference architecture, deployment patterns, APIs, observability, and production operations in a hands-on role.

The role requires strong experience with model serving, Kubernetes, REST/gRPC APIs, performance optimization, and CI/CD.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Real-Time Inference Platform Lead
Real-Time Inference Platform Lead

Creative Solutions Services, LLC • Charlotte (NC)

On-site
USD 80,000 - 89,000
Medical insurance
Dental insurance
Vision insurance
+3
Real-Time Inference Platform Lead
Real-Time Inference Platform Lead

NTT DATA, Inc. • Charlotte (NC)

Hybrid
USD 84,000 - 125,000
Medical, dental, and vision insurance
401k with company match
Paid time off
+2
Inference Lead
Inference Lead

VDart Inc • Charlotte (NC)

Hybrid
USD 140,000 - 190,000
Lead ML Platform Engineer
Lead ML Platform Engineer

Creative Solutions Services, LLC • Charlotte (NC)

On-site
USD 80,000 - 89,000
Medical insurance
Dental insurance
Vision insurance
+3
Inference Lead
Inference Lead

Stellent IT LLC • Charlotte (NC)

Hybrid
USD 96,000 - 138,000
Real-Time Inference Engineering Lead (FTE / Hybrid)
Real-Time Inference Engineering Lead (FTE / Hybrid)

NTT DATA, Inc. • Charlotte (NC)

Hybrid
USD 84,000 - 125,000
Medical, dental, and vision insurance
401k with company match
Paid time off
+2
Real-Time AI Inference Engineering Lead
Real-Time AI Inference Engineering Lead

VDart Inc • Charlotte (NC)

Hybrid
USD 140,000 - 190,000
Real-Time Inference Architect & Platform Lead
Real-Time Inference Architect & Platform Lead

Compunnel, Inc. • Charlotte (NC), Northern (KY)

Hybrid
USD 170,000 - 210,000
Inference Lead
Inference Lead

Compunnel, Inc. • Charlotte (NC), Northern (KY)

Hybrid
USD 170,000 - 210,000
Real-Time ML Inference Lead — Latency & Platform Architect
Real-Time ML Inference Lead — Latency & Platform Architect

Stellent IT LLC • Charlotte (NC)

Hybrid
USD 96,000 - 138,000