SVP, Lead AI/MLOps Infrastructure Engineer - Full Time - Hybrid - 5540

Benchmark IT - Technology Talent

New York (NY)

Hybrid

USD 200,000 - 230,000

Full time

7 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Retirement match
Unlimited PTO

Job summary

Benchmark IT - Technology Talent is partnering with a fintech client in New York to hire a Senior SVP, Lead AI/MLOps Infrastructure Engineer. The hybrid role leads end-to-end AI/ML platform stack, sets MLOps standards, and drives cloud-native AWS infra with Terraform and Kubernetes.

Expect a strategic, hands-on leader shaping GenAI workloads and cost optimization. The role emphasizes production ML/GenAI deployments, governance, and cross-team collaboration, with strong compensation and a

Qualifications

  • 15+ years of experience in DevOps, SRE, or platform engineering with AWS as primary cloud
  • Hands-on experience building and operating MLOps pipelines in production
  • Experience with MLOps tooling including model registries, experiment tracking, and feature stores
  • Exposure to Generative AI / LLM workloads, including AWS Bedrock
  • Strong IaC and scripting skills (Terraform; Python or similar)
  • Solid Linux, systems, and troubleshooting fundamentals
  • Excellent communicator, comfortable collaborating across teams

Responsibilities

  • Own the end-to-end AI/ML platform stack: orchestration, compute (including GPU), storage, and model serving
  • Build and operate MLOps pipelines across the full model lifecycle: training, validation, versioning, deployment
  • Productionize AI/ML and GenAI workloads in partnership with ML engineers and data scientists
  • Design and manage cloud-native AWS infrastructure using Kubernetes, and own Infrastructure as Code standards (Terraform)
  • Own SLAs/SLOs for model serving and inference; lead monitoring, drift detection, and incident response
  • Build internal tooling and standardized environments to boost ML engineer productivity (MLflow, Kubeflow, Weights & Biases, Ray)
  • Drive data governance, privacy compliance, and cost optimization across training and inference workloads
  • Set the MLOps roadmap and mentor engineers on infrastructure and MLOps best practices

Skills

AWS
MLOps
Python
Linux
DevOps
SRE
Cloud architecture
ML platform

Tools

Terraform
Kubernetes
MLflow
Kubeflow
Weights & Biases
Ray
Bedrock

Job description

SVP, Lead AI/MLOps Infrastructure Engineer – Full Time – Hybrid

We’re partnering with our client, a fast-growing fintech firm, on a senior-level hire to lead and own the infrastructure behind their AI and machine learning platforms.

This is a highly visible, hands-on leadership role where you’ll own the end-to-end AI/ML platform stack, from training and inference infrastructure to model serving, reliability, and cost, while setting the MLOps roadmap and standards for the team.

The priority here is MLOps and AI infrastructure first, built on deep AWS, Terraform, and production ML / GenAI experience.

What You’ll Be Doing
  • Own the end-to-end AI/ML platform stack: orchestration, compute (including GPU), storage, and model serving
  • Build and operate MLOps pipelines across the full model lifecycle: training, validation, versioning, and deployment
  • Productionize AI/ML and GenAI (LLM) workloads in partnership with ML engineers and data scientists
  • Design and manage cloud-native AWS infrastructure using Kubernetes, and own Infrastructure as Code standards (Terraform)
  • Own SLAs/SLOs for model serving and inference; lead monitoring, drift detection, and incident response
  • Build internal tooling and standardized environments that boost ML engineer productivity (e.g., MLflow, Kubeflow, Weights & Biases, Ray)
  • Drive data governance, privacy compliance, and cost optimization across training and inference workloads
  • Set the MLOps roadmap and mentor engineers on infrastructure and MLOps best practices
What They’re Looking For
  • 15+ years of experience in DevOps, SRE, or platform engineering, with AWS as primary cloud
  • Proven, hands-on experience building and operating MLOps pipelines in production (key priority)
  • Experience with MLOps tooling, including model registries, experiment tracking, and feature stores
  • Exposure to Generative AI / LLM workloads, including AWS Bedrock
  • Strong Infrastructure as Code (Terraform) and scripting skills (Python or similar)
  • Solid Linux, systems, and troubleshooting fundamentals
  • Excellent communicator, comfortable collaborating across teams
Nice to Have
  • Hands-on Kubernetes, containerized workloads, and cloud networking
  • Experience in regulated or fintech environments
  • Background optimizing costs for compute-intensive (GPU) workloads
Why This Role
  • Own and shape the AI platform at a growing fintech
  • High-impact leadership role setting MLOps strategy and standards
  • Strong compensation: $200K–$230K base + bonus + equity
  • Comprehensive benefits, including retirement match and unlimited PTO
  • Hybrid model: 4 days onsite / 1 day remote (NYC area)
Get your free, confidential resume review.

or drag and drop your file here.