GPU & ML Infrastructure Engineer, Data-Quality Focus

Aether Biomedical

United States

Remote

USD 104,000 - 150,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Svitla Systems Inc. is seeking a GPU & ML Infrastructure Engineer for a full-time position targeting Europe. You will own end-to-end data generation for GPU and compute health experiments, deploying workloads, collecting telemetry, and ensuring reproducible runs across NVIDIA datacenter GPUs and edge platforms.

This role emphasizes automated deployment, data quality, and robust logging. The ideal candidate has hands-on experience with LLM inference/training on GPUs, time-series telemetry, and

Qualifications

  • Experience deploying LLM inference and training workloads on GPUs, including quantized models.
  • Ability to diagnose sensor and sampling problems in time-series hardware data.
  • Understand reproducibility of GPU workloads and nondeterminism sources.
  • Strong Linux systems knowledge, GPU driver stacks, and process orchestration.
  • Experience collecting hardware telemetry programmatically via NVML and DCGM and out-of-band interfaces.
  • Familiarity with GPU benchmarking and stress tooling.
  • Knowledge of GPU power and thermal management and multi-GPU scaling.

Responsibilities

  • Port the test procedure to new hardware and document changes.
  • Build and maintain the benchmark workload suite for GPUs and edge devices.
  • Own logger correctness and ensure timestamp consistency across data sources.
  • Automate deployment of runs and clean teardown with reproducible configurations.
  • Create data-quality checks to catch missing or misaligned samples.
  • Oversee storage handoff with full run metadata.
  • Document each hardware target.

Skills

LLM workloads on GPUs
Time-series telemetry analysis
Linux systems
GPU driver stacks
Benchmark and stress tooling
Data collection pipelines

Tools

NVML
DCGM
Redfish
IPMI

Job description

Svitla Systems Inc. is seeking a GPU & ML Infrastructure Engineer for a full-time position targeting Europe. You will own end-to-end data generation for GPU and compute health experiments, deploying workloads, collecting telemetry, and ensuring reproducible runs across NVIDIA datacenter GPUs and edge platforms.

This role emphasizes automated deployment, data quality, and robust logging. The ideal candidate has hands-on experience with LLM inference/training on GPUs, time-series telemetry, and

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

GPU & ML Infrastructure Engineer — Remote
GPU & ML Infrastructure Engineer — Remote

Svitla Systems, Inc. • United States

Hybrid
USD 120,000 - 180,000
Remote work option
GPU & ML Infrastructure Engineer
GPU & ML Infrastructure Engineer

Svitla Systems, Inc. • United States

On-site
USD 120,000 - 180,000
Remote work option
GPU & ML Infrastructure Engineer
GPU & ML Infrastructure Engineer

ConsultBae India Private limited • United States

Remote
USD 150,000 - 210,000
Remote GPU & ML Infrastructure Engineer
Remote GPU & ML Infrastructure Engineer

ConsultBae India Private limited • United States

Remote
USD 150,000 - 210,000
ML Infrastructure Engineer – GPU Compute Platform
ML Infrastructure Engineer – GPU Compute Platform

Remanence • Paris (TX)

Hybrid
USD 124,000 - 186,000
Visa sponsorship
Relocation support
Hybrid work setup
+1
Principal ML Infrastructure Engineer (Relocation Available)
Principal ML Infrastructure Engineer (Relocation Available)

Franklin Fitch • Dallas (TX)

On-site
USD 100,000 - 140,000
Lead GPU Inference & ML Infrastructure Engineer
Lead GPU Inference & ML Infrastructure Engineer

Avride Inc. • Austin (TX)

On-site
USD 140,000 - 190,000
Software Engineer: ML Infra
Software Engineer: ML Infra

Generalist • Somerville (MA), San Mateo (CA)

On-site
USD 120,000 - 160,000
ML Platform Engineer — Infra for Research on GPU Fleets
ML Platform Engineer — Infra for Research on GPU Fleets

cursor • New York (NY), San Francisco (CA)

On-site
USD 120,000 - 180,000
Senior ML Infra Architect - Data Pipelines & GPUs
Senior ML Infra Architect - Data Pipelines & GPUs

Engg • New York (NY)

On-site
USD 180,000 - 250,000