Lead ML Platform Engineer

NTT DATA North America

Charlotte (NC)

On-site

USD 150,000 - 230,000

Full time

6 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

NTT DATA North America is seeking a Lead ML Platform Engineer in Charlotte, NC to drive architecture and hands-on leadership for the Cortex Predictive AI Platform. You will enable data scientists, ML engineers, and application teams to build, validate, deploy, and operate predictive models efficiently and reliably.

You will lead design decisions across the ML platform lifecycle, mentor teams, and ensure secure, scalable platform operations with incident response and runbooks.

Qualifications

  • 8+ years of experience in platform engineering, cloud engineering, SRE, MLOps, or related roles.
  • 4+ years designing, building, or operating enterprise AI/ML or data platforms.
  • Experience leading architecture and engineering delivery for complex, production-grade cloud and/or on-premises platforms.
  • Strong hands-on experience with GCP and multi-cloud or hybrid-cloud architecture.
  • Experience with Kubernetes-based platforms (GKE, OpenShift) in production.
  • Experience implementing MLOps capabilities, model lifecycle workflows, or ML platform services.
  • Proficiency in Python for automation, tooling, or ML workflows.
  • Experience with CI/CD, Git-based development, automated testing, and IaC practices.
  • Strong understanding of enterprise security, data protection, IAM, encryption, audit logging, and secure software delivery.
  • Experience observability, monitoring, dashboards, SLI/SLOs, incident response, and runbooks.
  • Experience mentoring engineers and communicating architecture decisions to stakeholders.

Responsibilities

  • Define and lead target architecture for predictive AI and ML platform capabilities across cloud and on-premises.
  • Design, build, and operate reusable platform services for end-to-end ML lifecycle.
  • Establish scalable reference architectures, engineering standards, and templates for Cortex platform workloads.
  • Lead platform engineering for GCP and multi-cloud environments with secure connectivity and governance.
  • Design and operate Kubernetes-based ML platforms using GKE and OpenShift.
  • Implement MLOps capabilities including experiment tracking, model packaging, validation, and deployment gates.
  • Build CI/CD pipelines and infrastructure automation for platform services and ML workflows.
  • Enable model migration from legacy environments into Cortex platform patterns.
  • Engineer real-time and batch inference with scalable serving patterns and observability.
  • Collaborate with data governance, security, and application teams to embed data protection.

Skills

Platform engineering
Cloud engineering
MLOps
GCP
Kubernetes
Python
CI/CD
Security
Observability
Leadership

Tools

GKE
OpenShift

Job description

Company Overview

Req ID: 388174
NTT DATA strives to hire exceptional, innovative and passionate individuals who want to grow with us. If you want to be part of an inclusive, adaptable, and forward-thinking organization,
We are currently seeking a Lead ML Platform Engineer to join our team in Charlotte, North Carolina (US-NC), United States (US).

Job Description

Position Summary
The Lead ML Platform Engineer provides architecture and hands-on engineering leadership for the Cortex Predictive AI Platform across cloud and on-premises environments. This role will establish and implement reusable, secure, scalable standards that enable data scientists, ML engineers, and application teams to build, validate, deploy, monitor, and operate predictive models efficiently and reliably.
The successful candidate will lead technical design and engineering decisions across the ML platform lifecycle, including governed data and feature access, model development environments, training and validation workflows, model delivery pipelines, real-time and batch inference, observability, reliability, and operational readiness. This role will also mentor engineering teams and transfer knowledge to support sustainable platform operations and adoption.

Key Responsibilities
  • Define and lead the target architecture for predictive AI and ML platform capabilities spanning public cloud and on-premises environments.
  • Design, build, and operate reusable platform services supporting the end-to-end ML lifecycle: governed data and features, model development, training, validation, deployment, inference, monitoring, and operations.
  • Establish scalable reference architectures, engineering standards, reusable templates, and implementation patterns for ML workloads across the Cortex portfolio.
  • Lead platform engineering for GCP and multi-cloud environments, including secure connectivity, identity, network controls, compute, storage, and managed AI/ML services where applicable.
  • Design and operate Kubernetes-based ML platforms using GKE, OpenShift, and associated container, workload orchestration, and resource-management capabilities.
  • Implement and improve MLOps capabilities for experiment tracking, model packaging, validation, approval gates, model registry integration, deployment automation, rollback, and lifecycle management.
  • Build CI/CD pipelines and infrastructure automation for platform services, ML workflows, model delivery, and environment provisioning.
  • Enable model migration from legacy environments into standardized Cortex platform patterns, minimizing delivery risk and operational disruption.
  • Engineer production-grade real-time and batch inference capabilities, including API-based serving, scalable runtime patterns, resiliency, performance, and operational support.
  • Partner with data engineering, data governance, security, privacy, risk, model validation, and application teams to ensure data protection and control requirements are embedded into platform design.
  • Implement platform observability, including logs, metrics, traces, dashboards, alerts, service-level indicators, service-level objectives, and operational runbooks.
  • Drive reliability engineering practices for ML platform services, including capacity planning, high availability, disaster recovery, incident management, root-cause analysis, and continuous improvement.
  • Ensure platform designs meet enterprise security requirements for authentication, authorization, secrets management, encryption, data access, auditability, and environment isolation.
  • Provide technical leadership, architecture reviews, code reviews, design guidance, and mentoring to ML platform engineers and adjacent delivery teams.
  • Produce clear technical documentation, reference implementations, operational procedures, and knowledge-transfer materials to enable self-service adoption and long-term support.
Required Qualifications
  • 8+ years of experience in platform engineering, cloud engineering, infrastructure engineering, SRE, MLOps, or related technical roles.
  • 4+ years of experience designing, building, or operating enterprise AI/ML or data platforms.
  • Demonstrated experience leading architecture and engineering delivery for complex, production-grade cloud and/or on-premises platforms.
  • Strong hands-on experience with GCP and working knowledge of multi-cloud or hybrid-cloud architecture.
  • Experience with Kubernetes-based platforms, including GKE and OpenShift, in production environments.
  • Strong experience implementing MLOps capabilities, model lifecycle workflows, or ML platform services.
  • Proficiency in Python for platform automation, integration, operational tooling, or ML workflow development.
  • Experience with CI/CD, Git-based development, automated testing, deployment automation, and infrastructure-as-code practices.
  • Strong understanding of enterprise security, data protection, identity and access management, secrets management, encryption, audit logging, and secure software delivery.
  • Experience implementing observability, monitoring, alerting, dashboards, SLOs, incident response, and operational runbooks.
  • Experience mentoring engineers and communicating technical architecture decisions to engineering, product, security, data, and executive stakeholders.
Required Skills / Knowledge
  • Enterprise ML platform architecture and end-to-end predictive model lifecycle management.
  • GCP, hybrid cloud, multi-cloud, on-premises platform, networking, identity, and security concepts.
  • Kubernetes, GKE, OpenShift, containers, workload orchestration, and scalable compute platforms.
  • MLOps, model development environments, model registries, validation workflows, model deployment, and model monitoring.
  • Python, CI/CD, Git, automated testing, infrastructure automation, and API-based integration.
  • Real-time and batch inference architecture, model-serving patterns, performance optimization, and operational support.
  • Data protection, governance, access controls, encryption, auditability, and regulated-platform design.
  • Observability, telemetry, dashboards, alerting, SLI/SLO design, reliability engineering, and production troubleshooting.
  • Technical leadership, reusable pattern development, engineering documentation, and knowledge transfer.
Preferred Qualifications
  • Experience with Vertex AI or comparable cloud ML platform services.
  • Experience designing or operating on-premises AI/ML platforms, private cloud, or hybrid ML workloads.
  • Experience with feature stores, model registries, experiment tracking, data lineage, model governance, or model risk-management processes.
  • Experience supporting model migration, platform modernization, or transition from legacy data science and ML environments.
  • Experience with real-time, low-latency model-serving systems and event-driven inference architectures.
  • Experience with Terraform, Helm, Argo CD, Jenkins, GitHub Actions, GitLab CI, or similar automation and deployment tooling.
  • Experience in banking, financial services, healthcare, insurance, or another regulated enterprise environment.
  • Experience establishing self-service platform capabilities for data scientists, ML engineers, and application teams.
Expected Outcomes
  • A secure, scalable, and reusable Cortex ML platform architecture spanning public cloud and on-premises environments.
  • Standardized MLOps, CI/CD, and model-delivery patterns that reduce time to train, validate, deploy, and operate predictive models.
  • Reliable platform capabilities for governed data and features, model migration, batch and real-time inference, and production operations.
  • Improved observability, resiliency, service-level management, and operational readiness for ML platform services and models.
  • Reusable engineering standards, reference implementations, documentation, and knowledge-transfer assets that enable self-service adoption and sustainable platform support.
About NTT DATA

NTT DATA is a $30 billion trusted global innovator of business and technology services. We serve 75% of the Fortune Global 100 and are committed to helping clients innovate, optimize and transform for long term success. As a Global Top Employer, we have diverse experts in more than 50 countries and a robust partner ecosystem of established and start-up companies. Our services include business and technology consulting, data and artificial intelligence, industry solutions, as well as the development, implementation and management of applications, infrastructure and connectivity. We are one of the leading providers of digital and AI infrastructure in the world. NTT DATA is a part of NTT Group, which invests over $3.6 billion each year in R&D to help organizations and society move confidently and sustainably into the digital future. Visit us at us.nttdata.com

NTT DATA endeavors to make https://us.nttdata.com accessible to any and all users. If you would like to contact us regarding the accessibility of our website or need assistance completing the application process, please contact us at https://us.nttdata.com/en/contact-us. This contact information is for accommodation requests only and cannot be used to inquire about the status of applications. NTT DATA is an equal opportunity employer. Qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, disability or protected veteran status. For our EEO Policy Statement, please click here. If you’d like more information on your EEE rights under the law, please click here. For Pay Transparency information, please click here.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Lead ML Platform Engineer (SRE / FTE / Onsite)
Lead ML Platform Engineer (SRE / FTE / Onsite)

NTT DATA North America • Charlotte (NC)

On-site
USD 84,000 - 125,000
Medical, dental, and vision insurance
401(k) with company match
Paid time off
Lead ML Platform Engineer (SRE / FTE / Onsite)
Lead ML Platform Engineer (SRE / FTE / Onsite)

NTT DATA, Inc. • Charlotte (NC)

On-site
USD 84,000 - 125,000
Medical Insurance
Dental Insurance
Vision Insurance
+6
Real-Time Inference Engineering Lead (FTE / Hybrid)
Real-Time Inference Engineering Lead (FTE / Hybrid)

NTT DATA North America • Charlotte (NC)

Hybrid
USD 84,000 - 125,000
Medical, dental, vision insurance
401k with company match
Paid time off
Senior AI DevOps Engineer (AI Ops / Platform Engineering) (Hybrid)
Senior AI DevOps Engineer (AI Ops / Platform Engineering) (Hybrid)

NTT DATA North America • Atlanta (GA)

Hybrid
USD 87,000 - 163,000
Cloud Agentic AI Platform Engineer / Architect
Cloud Agentic AI Platform Engineer / Architect

NTT DATA North America • Charlotte (NC)

Hybrid
USD 96,000 - 110,000
Azure AI Platform Engineer
Azure AI Platform Engineer

NTT Data Americas, Inc. • Irving (TX)

Hybrid
USD 140,000 - 210,000
AI Security Architect
AI Security Architect

Creative Solutions Services, LLC • Boston (MA)

Hybrid
USD 314,093,000 - 358,176,000
Real-Time Inference Engineering Lead (FTE / Hybrid)
Real-Time Inference Engineering Lead (FTE / Hybrid)

NTT DATA, Inc. • Charlotte (NC)

Hybrid
USD 84,000 - 125,000
Medical, dental, and vision insurance
401k with company match
Paid time off
+2
Data Scientist
Data Scientist

NTT DATA North America • Camden (NJ)

On-site
USD 95,000 - 120,000
Inclusive workplace environment
AI Architect
AI Architect

NTT DATA North America • Jersey City (NJ)

On-site
Health insurance
Vision insurance
Dental insurance
+5