Senior AI Infra Engineer: Scale & Secure ML on GPUs

Dover

Santa Clara, Northern (CA, KY)

Hybrid

USD 150,000 - 170,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Unlimited PTO
Generous parental leave
Stock Purchase Program (ESPP)
Annual bonus program
Medical, Dental and Vision coverage
Mental health support

Job summary

Dover is seeking a Software Engineer specialized in AI/ML applications to independently drive end-to-end lifecycle management of AI-powered systems, blending advanced AI development with robust software engineering and automation.

You will deploy and scale open-source models on distributed GPU infrastructure, design automated testing frameworks, and implement secure CI/CD pipelines that ensure quality on every merge.

Qualifications

  • 6+ years of production-grade, asynchronous Python development with clean architecture.

Responsibilities

  • Architect and scale AI infrastructure across multi-node clusters using Kubernetes, Ray, or Slurm.
  • Design, evaluate, and deploy production-grade AI models and agents with benchmarks.
  • Conduct comprehensive model benchmarks and dashboards to communicate performance to stakeholders.
  • Develop extensive automated test suites for end-to-end applications and stochastic AI outputs.
  • Automate DevSecOps with vulnerability scanning in Merge Requests and fast remediation pipelines.
  • Own features end-to-end from ideation to production, including repository synchronization and OSS interactions.

Skills

Python programming
AI systems design
Data analysis (pandas/NumPy)
GitLab pipelines
PyTest testing
Advanced Git workflows

Education

Bachelor's or Master's in CS/Engineering

Tools

Kubernetes
Ray
Slurm
LangChain
Hugging Face
MLOps
vLLM
SGLang
GitLab
PyTest

Job description

Dover is seeking a Software Engineer specialized in AI/ML applications to independently drive end-to-end lifecycle management of AI-powered systems, blending advanced AI development with robust software engineering and automation.

You will deploy and scale open-source models on distributed GPU infrastructure, design automated testing frameworks, and implement secure CI/CD pipelines that ensure quality on every merge.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior ML Infra Engineer - Scale GPU Clusters, Remote
Senior ML Infra Engineer - Scale GPU Clusters, Remote

AI Breaking Wire • San Francisco (CA), Northern (KY)

Hybrid
USD 320,000 - 500,000
Equity
Medical/Dental/Vision coverage
Unlimited PTO
+1
Senior AI Infra Engineer - Scale ML Platforms
Senior AI Infra Engineer - Scale ML Platforms

NVIDIA • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Equity
Benefits
Remote AI Infrastructure Engineer — Scale ML & Inference
Remote AI Infrastructure Engineer — Scale ML & Inference

Vantaca, LLC • Redwood City (CA), Northern (KY)

Hybrid
USD 150,000 - 190,000
Medical, Dental, Vision
Senior AI Infrastructure Engineer – Edge & Cloud ML
Senior AI Infrastructure Engineer – Edge & Cloud ML

Anduril • Costa Mesa (CA)

On-site
USD 150,000 - 230,000
Senior AI/ML Engineer - Scale & Deploy Intelligent Models
Senior AI/ML Engineer - Scale & Deploy Intelligent Models

HP • Spring (TX)

On-site
USD 147,000 - 231,000
Health insurance
Dental insurance
Vision insurance
+5
Senior ML Engineer — Scale Training Infra & AI Deployments
Senior ML Engineer — Scale Training Infra & AI Deployments

Best AI Tools Wiki • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 350,000
Equity package
Health insurance
Unlimited PTO
+3
Senior AI Infrastructure Engineer — Scale GPU Clusters Remote
Senior AI Infrastructure Engineer — Scale GPU Clusters Remote

AI Breaking Wire • San Francisco (CA), Northern (KY)

Hybrid
USD 280,000 - 420,000
Equity options
Health, vision, dental benefits
Unlimited PTO
+2
AI Infrastructure Engineer — Scale ML Training & Inference
AI Infrastructure Engineer — Scale ML Training & Inference

Triwill Group • San Francisco (CA), Northern (KY)

Hybrid
USD 140,000 - 190,000
Remote AI Research Clusters Engineer - ML Infra & GPU
Remote AI Research Clusters Engineer - ML Infra & GPU

NEPSE Trading • Northern (KY)

Hybrid
USD 124,000 - 196,000
Senior AI Infrastructure Engineer — Edge-Cloud ML + Equity
Senior AI Infrastructure Engineer — Edge-Cloud ML + Equity

Anduril Industries • Washington

On-site
USD 191,000 - 253,000