GPU & ML Infrastructure Engineer

Svitla Systems, Inc.

United States

Hybrid

USD 120,000 - 180,000

Full time

11 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Remote work option

Job summary

Svitla Systems, Inc. is seeking a GPU & ML Infrastructure Engineer for a full-time position in the USA. The role leads the end-to-end data generation process for GPU systems, including benchmarks, automated deployment, hardware telemetry collection, data quality validation, and dataset delivery.

The successful candidate will work with multiple NVIDIA data-center GPU generations and edge platforms, overlapping with the client’s team until 11:00 a.m. PST.

Qualifications

  • :

Responsibilities

  • Port existing test procedures to new data-center GPUs and edge devices.
  • Build and maintain GPU benchmark workloads, including synthetic kernels and LLM inference and training.
  • Ensure telemetry collection is accurate and consistent across platforms and data sources.
  • Investigate sampling issues, timestamp inconsistencies, missing sensor data, and logging anomalies.
  • Automate workload deployment, execution, data collection, and environment cleanup.
  • Build automated checks for missing samples, irregular intervals, clock mismatches, and invalid telemetry.
  • Deliver datasets in a consistent and documented format with complete run metadata.
  • Document hardware targets, configurations, procedures, driver versions, and firmware versions.

Skills

LLM workloads on GPUs
Linux systems
Python automation
Hardware telemetry
GPU driver stacks
NVML/DCGM/IPMI/Redfish telemetry
Benchmark workloads

Tools

NVML
DCGM
BMC
IPMI
Redfish

Job description

Svitla Systems Inc. is looking for a GPU & ML Infrastructure Engineer for a full-time position (40 hours per week) in the USA. Our client is a stealth startup. The successful candidate will own the end-to-end data generation process for GPU systems, including benchmark workloads, automated deployment, hardware telemetry collection, data quality validation, and dataset delivery. The role requires working with multiple NVIDIA data-center GPU generations and embedded or edge platforms.

Requirements
  • Strong experience deploying LLM inference and training workloads on GPUs, including quantized models.
  • Experience diagnosing sensor, logging, and sampling issues in time-series hardware data.
  • Ability to build reproducible GPU workloads and control sources of run-to-run variation.
  • Strong Linux systems knowledge, including GPU driver stacks, process orchestration, scheduling, and timing.
  • Experience collecting hardware telemetry programmatically using NVML, DCGM, BMC, IPMI, or Redfish.
  • Strong Python skills for automation, telemetry collection, and data processing.
  • Experience automating workload deployment and data collection across different hardware platforms.
  • Ability to work independently and take ownership of technical processes.
  • Availability to overlap with the client’s team until 11:00 a.m. PST.
Nice to have
  • Experience building data collection pipelines for hardware testing or systems research.
  • Familiarity with GPU benchmarking, stress testing, and benchmark methodology.
  • Knowledge of GPU power, thermal management, multi-GPU scaling, and NCCL.
  • Experience building automated data quality checks for time-series or sensor data.
Responsibilities
  • Port existing test procedures to new data-center GPUs and edge devices.
  • Build and maintain GPU benchmark workloads, including synthetic kernels and LLM inference and training.
  • Ensure telemetry collection is accurate and consistent across platforms and data sources.
  • Investigate sampling issues, timestamp inconsistencies, missing sensor data, and logging anomalies.
  • Automate workload deployment, execution, data collection, and environment cleanup.
  • Build automated checks for missing samples, irregular intervals, clock mismatches, and invalid telemetry.
  • Deliver datasets in a consistent and documented format with complete run metadata.
  • Document hardware targets, configurations, procedures, driver versions, and firmware versions.
We offer
  • US and EU projects based on advanced technologies.
  • Competitive compensation based on skills and experience.
  • Flexibility in workspace, either remote or our welcoming office.
  • Bonuses for article writing, public talks, and other activities.
  • Free tech webinars and meetups organized by Svitla.
  • Regular corporate online activities.
  • Awesome team, friendly and supportive community!
Why join Svitla

Svitla Systems is a global digital solutions company headquartered in the U.S. and operating across the Americas, Europe, Asia, and APAC. Since 2003, we have served a wide range of clients - from innovative start-ups to Fortune 500 companies. Our success is built on partnership. By integrating seamlessly with clients' teams, we create lasting collaborations that drive real results. We are strong advocates of workplace flexibility, remote culture, individual approach to professional and personal growth.

Our global mission is to build a business that contributes to wellbeing of our partners, personnel, and their families, improves our communities, and makes a lasting difference in the world. Together, we are coding a brighter tomorrow - and living it.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

GPU & ML Infrastructure Engineer — Remote
GPU & ML Infrastructure Engineer — Remote

Svitla Systems, Inc. • United States

Hybrid
USD 120,000 - 180,000
Remote work option
Infrastructure Engineer (GPU & Compute)
Infrastructure Engineer (GPU & Compute)

Lightning-Ai • New York (NY)

Hybrid
USD 180,000 - 200,000
Medical coverage (US)
Dental coverage (US)
Vision coverage (US)
+6
Infrastructure Engineer (GPU & Compute)
Infrastructure Engineer (GPU & Compute)

Lightning AI • San Francisco (CA)

On-site
USD 180,000 - 200,000
Comprehensive medical, dental, and vision coverage
Generous paid time off
Flexible work environment
Senior Software Engineer, AI Inference Systems
Senior Software Engineer, AI Inference Systems

NVIDIA • Santa Clara (CA)

Hybrid
USD 184,000 - 357,000
Equity
Benefits
AI Practice Lead
AI Practice Lead

Svitla Systems, Inc. • New York (NY)

Hybrid
USD 180,000 - 240,000
Remote or office work
Competitive compensation
Travel opportunities
+1
Infrastructure Engineer, LLM Inference Optimization
Infrastructure Engineer, LLM Inference Optimization

GMI Cloud • Mountain View (CA)

On-site
USD 170,000 - 230,000
Senior Software Engineer - GPU Local AI Platforms
Senior Software Engineer - GPU Local AI Platforms

NVIDIA AI • Seattle (WA)

On-site
USD 224,000 - 431,250
Equity
Benefits
Senior Software Engineer - GPU Local AI Platforms
Senior Software Engineer - GPU Local AI Platforms

NVIDIA • Westford (MA)

On-site
USD 224,000 - 432,000
Senior Software Engineer - GPU Local AI Platforms
Senior Software Engineer - GPU Local AI Platforms

NVIDIA • Durham (NC)

On-site
USD 224,000 - 357,000
Equity
Benefits
GPU Performance / Kernel Engineer
GPU Performance / Kernel Engineer

Designworks Talent • Bellevue (WA)

Hybrid
USD 180,000 - 240,000
Health insurance
401(k) plan
Annual bonus
+1