Senior MLOps Engineer: GPU AI Infra & Production

Blue Signal Search

Santa Clara (CA)

On-site

USD 140,000 - 190,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Advanced GPU infra exposure
Collaborative engineering culture
Open source AI frameworks access
Competitive compensation
Career growth in AI infra

Job summary

Blue Signal Search is seeking an experienced MLOps Engineer to design, automate, and scale production AI infrastructure in Santa Clara, California. You will join a team building distributed GPU computing backbones for large-scale machine learning workloads and enterprise AI applications.

You will implement ML pipelines with Kubeflow and Airflow, orchestrate on Kubernetes and SLURM, and develop Python automation to improve reliability, performance, and operational efficiency across the platform.

Qualifications

  • Five+ years in Infrastructure, DevOps, Platform Engineering, or ML Ops.
  • Two+ years supporting production GPU/AI infrastructure.
  • Production experience orchestrating distributed GPU workloads with Kubernetes, SLURM, Ray, or similar.
  • Production ML pipelines with Kubeflow or Airflow.
  • Deploying open source foundation models like Llama, Qwen, DeepSeek or similar.
  • Strong Python programming and automation.
  • Excellent cross-functional communication.
  • Experience in AI-focused organizations or GPU cloud providers.

Responsibilities

  • Design, maintain, and improve production infrastructure for distributed GPU ML workloads.
  • Build and optimize automated ML workflows with Kubeflow, Airflow, or similar.
  • Deploy and maintain open source LLMs in production ensuring performance and scalability.
  • Orchestrate AI workloads across Kubernetes, SLURM, Ray, and similar platforms.
  • Develop Python automation to streamline infra management and deployment processes.
  • Monitor GPU resource utilization and improve scheduling efficiency.
  • Collaborate with software engineers, AI researchers, and infra teams to productionize models.
  • Create operational docs, automation tools, and deployment standards.
  • Provide guidance on AI infra capabilities to internal and customer teams.

Skills

Python
MLOps
Kubernetes
Kubeflow
Airflow
SLURM
Ray
Automation
GPU infra

Tools

Kubeflow
Airflow
Kubernetes
Ray
SLURM

Job description

Blue Signal Search is seeking an experienced MLOps Engineer to design, automate, and scale production AI infrastructure in Santa Clara, California. You will join a team building distributed GPU computing backbones for large-scale machine learning workloads and enterprise AI applications.

You will implement ML pipelines with Kubeflow and Airflow, orchestrate on Kubernetes and SLURM, and develop Python automation to improve reliability, performance, and operational efficiency across the platform.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

MLOps Engineer
MLOps Engineer

Blue Signal Search • Santa Clara (CA)

On-site
USD 140,000 - 190,000
Advanced GPU infra exposure
Collaborative engineering culture
Open source AI frameworks access
+2
AI Infra Engineer — GPU Cloud, Kubernetes/Slurm
AI Infra Engineer — GPU Cloud, Kubernetes/Slurm

Blue Signal Search • San Francisco (CA)

On-site
USD 180,000 - 240,000
Annual bonus
Equity participation
Comprehensive benefits
+1
Senior MLOps Engineer — Enterprise AI Infrastructure Lead
Senior MLOps Engineer — Enterprise AI Infrastructure Lead

Nace AI • Palo Alto (CA)

On-site
USD 180,000 - 240,000
Equity
Premium benefits
AI Infra/HPC Engineer
AI Infra/HPC Engineer

Blue Signal Search • San Francisco (CA)

On-site
USD 180,000 - 240,000
Annual bonus
Equity participation
Comprehensive benefits
+1
Senior GPU Compute Cluster Architect
Senior GPU Compute Cluster Architect

Blue Signal Search • San Francisco (CA)

On-site
USD 150,000 - 230,000
Senior MLOps Engineer, AI Infrastructure & Open-Source
Senior MLOps Engineer, AI Infrastructure & Open-Source

NVIDIA Corporation • Santa Clara (CA), Northern (KY)

Hybrid
USD 184,000 - 357,000
Equity
Benefits
AI Kernel & GPU HPC Cluster Engineer
AI Kernel & GPU HPC Cluster Engineer

Blue Signal Search • Santa Clara (CA)

On-site
USD 150,000 - 210,000
Senior GPU Network Architect for AI Clusters
Senior GPU Network Architect for AI Clusters

Blue Signal Search • Santa Clara (CA)

On-site
USD <240,000
Senior AI Infra Engineer: LLMOps, GPU, Kubernetes
Senior AI Infra Engineer: LLMOps, GPU, Kubernetes

eloquentai • San Francisco (CA)

On-site
USD 130,000 - 160,000
Senior MLOps Engineer: AI Infrastructure & Scale Expert
Senior MLOps Engineer: AI Infrastructure & Scale Expert

NVIDIA • Seattle (WA)

On-site
USD 184,000 - 357,000