Sr. Platform Engineer, ML Infrastructure

Insilico Search Partners

Cambridge (MA)

On-site

USD 140,000 - 210,000

Full time

35 hours ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Insilico Search Partners is seeking an experienced infrastructure/ML Ops engineer to own production-grade Kubernetes clusters and data pipelines. The role focuses on scaling ML workloads, implementing CI/CD and GitOps, and ensuring reliable, observable infrastructure across on-prem and cloud.

You will collaborate with ML engineers and researchers, manage GPU workloads, and drive cross-functional platform initiatives to reduce toil and improve repeatability.

Qualifications

  • 5+ years in production infrastructure, platform engineering, DevOps, SRE, or MLOps.
  • Hands-on Kubernetes and infrastructure-as-code experience.
  • Experience with major cloud platforms, CI/CD pipelines and containerization.
  • Familiarity with ML workflows and model-serving platforms.

Responsibilities

  • Manage infrastructure as code (Terraform) and production Kubernetes clusters.
  • Build and maintain CI/CD and GitOps workflows; secure and optimize container environments.
  • Operate workflow orchestration and storage for large ML datasets.
  • Manage GPU-enabled infrastructure: provisioning, drivers, utilization.
  • Establish observability (metrics, logs, dashboards, alerting) and runbooks.
  • Deploy and operate ML training/inference infrastructure and tooling for experiment tracking.
  • Partner with ML Engineers to move model workloads onto production-ready infrastructure.

Skills

Kubernetes
Terraform
CI/CD
GitOps
Workflow orchestration
Cloud platforms

Tools

Kubeflow
Argo CD
Nextflow
Ray
Vertex AI
Argo Workflows
Helm
CircleCI
GitHub Actions
Seqera

Job description

About the Company

Our client is a venture-backed biotech company applying AI to drug discovery, using a proprietary platform to identify novel drug targets and therapeutics from complex biological data.


Own and evolve the shared infrastructure — Kubernetes, Terraform, CI/CD, orchestration, storage, security, observability, and GPU systems — behind the company's ML and scientific workloads, across on-prem and cloud. You'll work closely with Machine Learning Engineers and researchers to keep training, evaluation, and inference reliable and scalable. Infrastructure-first role, with meaningful ML systems ownership (~60/40 split).


Key Responsibilities


  • Manage infrastructure as code (Terraform) and operate production Kubernetes clusters — networking, IAM/RBAC, security, resource allocation, reliability.

  • Build and maintain CI/CD and GitOps workflows (GitHub Actions, CircleCI, Helm, Argo CD); secure and optimize container environments.

  • Operate workflow orchestration (Kubeflow, Argo Workflows, Nextflow, or Seqera) and storage infrastructure for large scientific/ML datasets.

  • Manage GPU-enabled infrastructure: provisioning, scheduling, drivers, utilization, troubleshooting.

  • Establish observability (metrics, logs, dashboards, alerting) and document architecture/runbooks so operations don't depend on one person.

  • Deploy and operate ML training/inference infrastructure (Anyscale, Ray, Vertex AI, Kubernetes); own shared tooling for experiment tracking and model/data versioning.

  • Partner with ML Engineers to move model workloads onto shared, production-ready infrastructure.

  • Lead cross-functional infrastructure initiatives and mentor engineers/scientists on the platform.


Qualifications


  • 5+ years in production infrastructure, platform engineering, DevOps, SRE, or MLOps.

  • Strong hands-on Kubernetes and infrastructure-as-code (Terraform) experience.

  • Experience with major cloud platforms, CI/CD, containerization, and GPU/compute-intensive workloads.

  • Workflow orchestration and GitOps tooling (Kubeflow, Argo, Nextflow, Helm, Kueue).

  • Familiarity with PyTorch and distributed ML execution; experience with model-serving platforms (Ray, Vertex AI).

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior ML Infra Engineer
Senior ML Infra Engineer

Maxinsights Corporation • Santa Clara (CA)

On-site
USD 180,000 - 240,000
Staff Platform Machine Learning Engineer – Engine
Staff Platform Machine Learning Engineer – Engine

Jobtailor • New York (NY)

On-site
USD 140,000 - 190,000
Platform Engineer
Platform Engineer

Harrison Clarke • San Francisco (CA)

On-site
USD 120,000 - 160,000
MLOps Engineer MLOps Engineer
MLOps Engineer MLOps Engineer

Kurai • Austin (TX)

On-site
USD 140,000 - 190,000
Platform Engineer - AI/ML Infrastructure (Kubernetes & Terraform)
Platform Engineer - AI/ML Infrastructure (Kubernetes & Terraform)

Madrona Venture Labs • United States

Hybrid
USD 180,000 - 260,000
Member of Technical Staff (AI Infrastructure Engineer)
Member of Technical Staff (AI Infrastructure Engineer)

Perplexity • San Francisco (CA)

On-site
USD 120,000 - 150,000
Principal Platform Engineer
Principal Platform Engineer

European Recruitment BV • United States

On-site
USD 150,000 - 200,000
ML Infrastructure Engineer
ML Infrastructure Engineer

Lattice, Inc. • San Francisco (CA)

Hybrid
USD 200,000 - 280,000
Competitive salary
Premium health, dental, and vision insurance
Unlimited PTO
+2
Software Engineer, AI Infrastructure
Software Engineer, AI Infrastructure

Harell Data • Palo Alto (CA)

On-site
USD 180,000 - 260,000
AI/ML Platform Engineer
AI/ML Platform Engineer

Planet Pharma • Indianapolis (IN)

On-site
USD 130,000 - 170,000