ML Infrastructure Engineer: Build Scalable GPU Clusters

Cursor

California (MO)

On-site

USD 140,000 - 185,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Cursor is seeking an experienced ML infrastructure engineer to help build and optimize high-performance GPU infrastructure that supports large-scale RL workloads. You will work closely with ML researchers and engineers to enhance training frameworks, reliability, and developer experience.

You will contribute to automation, monitoring, and scalable cluster management, leveraging Python, Typescript, Rust, and Golang across Linux-based environments and Kubernetes.

Qualifications

  • Strong background in systems and infrastructure-focused software engineering.
  • Proficient in Python, Typescript, Rust and Golang.
  • Experience with distributed storage and networking.
  • Experience with Kubernetes and Linux-based deployments.

Responsibilities

  • Collaborate with ML researchers and engineers to improve throughput and reliability of training.
  • Plan and build cutting-edge GPU infrastructure with OEMs and cloud providers.
  • Improve density and scalability of compute environments for large RL workloads.
  • Create software to automate building, monitoring, and running GPU clusters.
  • Develop workload scheduling and data movement systems for Cursor’s training footprint.

Skills

Systems engineering
Python
Typescript
Rust
Golang
Distributed systems

Tools

Kubernetes
Linux

Job description

Cursor is seeking an experienced ML infrastructure engineer to help build and optimize high-performance GPU infrastructure that supports large-scale RL workloads. You will work closely with ML researchers and engineers to enhance training frameworks, reliability, and developer experience.

You will contribute to automation, monitoring, and scalable cluster management, leveraging Python, Typescript, Rust, and Golang across Linux-based environments and Kubernetes.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

ML Infrastructure Engineer: Build Scalable GPU Clusters
ML Infrastructure Engineer: Build Scalable GPU Clusters

cursor • New York (NY), San Francisco (CA)

On-site
USD 120,000 - 150,000
Software Engineer, ML Infrastructure
Software Engineer, ML Infrastructure

Cursor • California (MO)

On-site
USD 140,000 - 185,000
Software Engineer, ML Infrastructure
Software Engineer, ML Infrastructure

Cursor • New York (NY), San Francisco (CA)

On-site
USD 120,000 - 150,000
Senior ML Infra Engineer - Scale GPU Clusters, Remote
Senior ML Infra Engineer - Scale GPU Clusters, Remote

AI Breaking Wire • San Francisco (CA), Northern (KY)

Hybrid
USD 320,000 - 500,000
Equity
Medical/Dental/Vision coverage
Unlimited PTO
+1
Senior ML Infra Engineer: GPU-Optimized Kubernetes Platform
Senior ML Infra Engineer: GPU-Optimized Kubernetes Platform

Hamilton Barnes Associates Limited • San Francisco (CA)

On-site
USD 225,000 - 275,000
Stock options
ML Platform Engineer — Infra for Research on GPU Fleets
ML Platform Engineer — Infra for Research on GPU Fleets

cursor • New York (NY), San Francisco (CA)

On-site
USD 120,000 - 180,000
ML Infra Engineer — GPU Clusters & Distributed Systems
ML Infra Engineer — GPU Clusters & Distributed Systems

AI Breaking Wire • San Francisco (CA), Northern (KY)

Hybrid
USD 170,000 - 250,000
Industry-leading compensation and/or:?
Unlimited PTO
Top-tier medical, dental, and vision
+1
AI/ML Infra Engineer - Hosting
AI/ML Infra Engineer - Hosting

Hamilton Barnes Associates Limited • San Francisco (CA)

On-site
USD 225,000 - 275,000
Stock options
ML Infra Engineer: GPU Orchestration & Observability
ML Infra Engineer: GPU Orchestration & Observability

Autolab • San Francisco (CA)

On-site
USD 150,000 - 210,000
ML Platform Engineer — Scalable GPU & Kubernetes
ML Platform Engineer — Scalable GPU & Kubernetes

Socket.dev • Palo Alto (CA)

On-site
USD 180,000 - 280,000
Healthcare coverage
Relocation support
Retirement plans
+2