Senior MLOps Engineer: GPU AI Infra & Production

Blue Signal Search

Santa Clara (CA)

On-site

USD 140,000 - 190,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Advanced GPU infra exposure
Collaborative engineering culture
Open source AI frameworks access
Competitive compensation
Career growth in AI infra

Job summary

Blue Signal Search is seeking an experienced MLOps Engineer to design, automate, and scale production AI infrastructure in Santa Clara, California. You will join a team building distributed GPU computing backbones for large-scale machine learning workloads and enterprise AI applications.

You will implement ML pipelines with Kubeflow and Airflow, orchestrate on Kubernetes and SLURM, and develop Python automation to improve reliability, performance, and operational efficiency across the platform.

Qualifications

  • Five+ years in Infrastructure, DevOps, Platform Engineering, or ML Ops.
  • Two+ years supporting production GPU/AI infrastructure.
  • Production experience orchestrating distributed GPU workloads with Kubernetes, SLURM, Ray, or similar.
  • Production ML pipelines with Kubeflow or Airflow.
  • Deploying open source foundation models like Llama, Qwen, DeepSeek or similar.
  • Strong Python programming and automation.
  • Excellent cross-functional communication.
  • Experience in AI-focused organizations or GPU cloud providers.

Responsibilities

  • Design, maintain, and improve production infrastructure for distributed GPU ML workloads.
  • Build and optimize automated ML workflows with Kubeflow, Airflow, or similar.
  • Deploy and maintain open source LLMs in production ensuring performance and scalability.
  • Orchestrate AI workloads across Kubernetes, SLURM, Ray, and similar platforms.
  • Develop Python automation to streamline infra management and deployment processes.
  • Monitor GPU resource utilization and improve scheduling efficiency.
  • Collaborate with software engineers, AI researchers, and infra teams to productionize models.
  • Create operational docs, automation tools, and deployment standards.
  • Provide guidance on AI infra capabilities to internal and customer teams.

Skills

Python
MLOps
Kubernetes
Kubeflow
Airflow
SLURM
Ray
Automation
GPU infra

Tools

Kubeflow
Airflow
Kubernetes
Ray
SLURM

Job description

Blue Signal Search is seeking an experienced MLOps Engineer to design, automate, and scale production AI infrastructure in Santa Clara, California. You will join a team building distributed GPU computing backbones for large-scale machine learning workloads and enterprise AI applications.

You will implement ML pipelines with Kubeflow and Airflow, orchestrate on Kubernetes and SLURM, and develop Python automation to improve reliability, performance, and operational efficiency across the platform.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

MLOps Engineer
MLOps Engineer

Blue Signal Search • Santa Clara (CA)

On-site
USD 140,000 - 190,000
Advanced GPU infra exposure
Collaborative engineering culture
Open source AI frameworks access
+2
AI Infra Engineer — GPU Cloud, Kubernetes/Slurm
AI Infra Engineer — GPU Cloud, Kubernetes/Slurm

Blue Signal Search • San Francisco (CA)

On-site
USD 180,000 - 240,000
Annual bonus
Equity participation
Comprehensive benefits
+1
AI Infra/HPC Engineer
AI Infra/HPC Engineer

Blue Signal Search • San Francisco (CA)

On-site
USD 180,000 - 240,000
Annual bonus
Equity participation
Comprehensive benefits
+1
AI Kernel & GPU HPC Cluster Engineer
AI Kernel & GPU HPC Cluster Engineer

Blue Signal Search • Santa Clara (CA)

On-site
USD 150,000 - 210,000
Senior GPU Compute Cluster Architect
Senior GPU Compute Cluster Architect

Blue Signal Search • San Francisco (CA)

On-site
USD 150,000 - 230,000
ML Ops Engineer
ML Ops Engineer

Rise Technical Recruitment Limited • San Francisco (CA)

Hybrid
USD 140,000 - 180,000
Equity
Healthcare
401(k)
+1
Senior GPU Network Architect for AI Clusters
Senior GPU Network Architect for AI Clusters

Blue Signal Search • Santa Clara (CA)

On-site
USD 180,000 - 240,000
Senior MLOps Engineer: AI Infrastructure & Scale Expert
Senior MLOps Engineer: AI Infrastructure & Scale Expert

NVIDIA • Seattle (WA)

On-site
USD 184,000 - 357,000
Senior ML Infrastructure Engineer - GPU Training & MLOps
Senior ML Infrastructure Engineer - GPU Training & MLOps

Atoms • San Francisco (CA)

On-site
USD 224,000 - 280,000
Medical, Dental, Vision, Disability, and Life Insurance
Flexible Spending Account / Health Savings Account Options
401(k)
+2
Senior MLOps Engineer: Scalable Pipelines & GPUs
Senior MLOps Engineer: Scalable Pipelines & GPUs

Boston Dynamics • Waltham (MA)

On-site
USD 150,000 - 180,000
Medical, dental, vision
401(k)
Paid time off
+1