Staff ML Platform Engineer: Scale GPU Pipelines

SimplyHired

United States

Remote

USD 180,000 - 240,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Arriver System Software S.R.L. is seeking a Staff Software Engineer to join the ML Platform team.

You will design, build, and optimize large-scale ML infrastructure across on-prem NVIDIA DGX clusters and AWS Cloud, enabling training pipelines and production-grade ML solutions. You will write high-quality, maintainable code in multiple languages, lead complex projects, and collaborate with data science and engineering teams to deliver scalable data workflows and reliable platform services across

Qualifications

  • Bachelor's degree in Engineering, Information Systems, Computer Science, or related field and 4+ years of Software Engineering or related work experience.
  • Master's degree in Engineering, Information Systems, Computer Science, or related field and 3+ years of Software Engineering or related work experience.
  • PhD in Engineering, Information Systems, Computer Science, or related field and 2+ years of Software Engineering or related work experience.
  • 2+ years of work experience with programming languages such as C, C++, Java, Python, etc.

Responsibilities

  • Architect and develop core components of the ML platform and data infrastructure for training, inference, and large-scale data processing.
  • Design and implement scalable solutions for GPU clusters and distributed data pipelines on-prem and in AWS.
  • Lead project workstreams, ensuring timely delivery and alignment with platform roadmap; operate independently and drive outcomes end-to-end.
  • Build and optimize data pipelines for ingestion, transformation, storage, and retrieval supporting ML workflows.
  • Write clean, efficient, and maintainable code (services, operators, automation, tooling) in multiple programming languages.
  • Collaborate with data science and engineering teams to integrate ML and data workflows seamlessly.
  • Implement CI/CD for ML and data workflows using Argo Workflows, ArgoCD, GitHub Actions; champion testability and reproducibility.
  • Maintain observability (Prometheus, Grafana) and logging (AWS CloudWatch, ELK/Opensearch); drive SLOs, tracing, and cost-awareness.
  • Operate AWS services (EKS, EC2, VPC, IAM, S3, EFS, Batch) across hybrid environments; contribute to security and compliance controls.
  • Continuously improve platform reliability, performance (GPU utilization, throughput), and developer experience; stay current on modern MLOps/Data engineering practices.

Skills

C
C++
Java
Python

Education

Bachelor's degree
Master's degree
PhD

Job description

Arriver System Software S.R.L. is seeking a Staff Software Engineer to join the ML Platform team.

You will design, build, and optimize large-scale ML infrastructure across on-prem NVIDIA DGX clusters and AWS Cloud, enabling training pipelines and production-grade ML solutions. You will write high-quality, maintainable code in multiple languages, lead complex projects, and collaborate with data science and engineering teams to deliver scalable data workflows and reliable platform services across

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

ML Platform Software Engineer - Timisoara, Romania
ML Platform Software Engineer - Timisoara, Romania

SimplyHired • United States

On-site
USD 180,000 - 240,000
Staff ML Platform Engineer: Scale Production ML Pipelines
Staff ML Platform Engineer: Scale Production ML Pipelines

Datavant • Salt Lake City (UT)

On-site
USD 224,000 - 280,000
Senior ML Platform Engineer: Scale Production Pipelines
Senior ML Platform Engineer: Scale Production Pipelines

Datavant • Montpelier (VT)

On-site
USD 224,000 - 280,000
Senior GPU HPC Infrastructure Engineer for ML Pipelines
Senior GPU HPC Infrastructure Engineer for ML Pipelines

Jaide Health • United States

Hybrid
USD 180,000 - 250,000
Weekly lunch stipend
Health and dental benefits
Parental leave
+3
Senior ML Platform & Infra Engineer - Scale AI Pipelines
Senior ML Platform & Infra Engineer - Scale AI Pipelines

Monograph • United States

Hybrid
USD 160,000 - 240,000
Competitive base pay
Equity (RSUs)
Benefits
Staff ML Platform Engineer: Scale Production ML
Staff ML Platform Engineer: Scale Production ML

Datavant • Boise (ID)

On-site
USD 224,000 - 280,000
ML Platform Engineer — Infra for Research on GPU Fleets
ML Platform Engineer — Infra for Research on GPU Fleets

cursor • New York (NY), San Francisco (CA)

On-site
USD 120,000 - 180,000
Staff ML Platform Engineer: Scale Production AI & LLM
Staff ML Platform Engineer: Scale Production AI & LLM

Datavant • Helena (MT)

On-site
USD 224,000 - 280,000
Staff ML Platform Engineer: Scale Graph ML & MLOps
Staff ML Platform Engineer: Scale Graph ML & MLOps

Reddit, Inc. • San Francisco (CA)

On-site
USD 230,000 - 322,000
Equity in RSUs
Medical, dental, and vision insurance
401(k) employer match
+1
ML Platform Engineer — Scale AI Pipelines
ML Platform Engineer — Scale AI Pipelines

Gusto, Inc. • Denver (CO)

Hybrid
USD 160,000 - 200,000
Equity (RSUs)
Benefits