GPU & ML Infrastructure Engineer

ConsultBae India Private limited

United States

Remote

USD 150,000 - 210,000

Full time

8 days ago
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

ConsultBae India Private Limited is seeking a seasoned GPU & ML Infrastructure Engineer to build, test, monitor, and maintain GPU-based benchmarking and ML workloads. You will own end-to-end data generation, telemetry collection, and reproducible results.

Experience with NVIDIA data-center GPUs, Linux systems, and automation is essential to deploy workloads, validate data, and document platform configurations across edge and data-center platforms.

Qualifications

  • 10 years of relevant experience preferred; exceptional candidates considered.
  • Hands-on experience deploying LLM workloads on GPUs.
  • Experience with quantized and full-precision ML models.
  • Strong Linux system administration and troubleshooting skills.
  • Experience collecting and analyzing hardware telemetry programmatically.
  • Telemetry technologies such as NVML, DCGM, BMC, IPMI, Redfish.

Responsibilities

  • Port and adapt GPU test procedures to NVIDIA data-center and edge platforms.
  • Analyze driver differences, power/thermal limits, and telemetry sources.
  • Build and maintain benchmarking, stress-testing, and workload suites.
  • Deploy and execute LLM inference/training workloads, including Llama-family models.
  • Ensure reproducibility and consistency across runs.
  • Automate deployment, logging, data collection, and cleanup.
  • Build data-quality checks for time-series and hardware telemetry data.
  • Document platform configurations, limitations, and hardware behavior.

Skills

LLM inference
LLM training
NVIDIA GPU environments
Linux administration
Telemetry programmatically
Automation scripting
Multi-GPU scaling
NCCL
Tensor parallelism
Pipeline parallelism
Data-quality validation

Tools

NVML
DCGM
BMC
IPMI
Redfish

Job description

GPU & ML Infrastructure Engineer

Experience: 10 Years

Location: Remote

About the Role

We are looking for an experienced GPU & ML Infrastructure Engineer to build, test, monitor, and maintain reliable infrastructure for GPU-based benchmarking, stress testing, telemetry collection, and machine learning workloads.

The role involves working with NVIDIA data-center GPUs and embedded/edge platforms, deploying ML workloads, collecting hardware telemetry, automating testing processes, and ensuring that the resulting data is accurate, consistent, and reproducible.

The engineer will own the end-to-end testing and data-generation process, from preparing new GPU platforms and deploying workloads to collecting telemetry, validating data, and documenting results.

Key Responsibilities
  • Port and adapt existing GPU test procedures to new NVIDIA data-center and edge/embedded platforms.

  • Analyze differences in GPU drivers, power/thermal limits, hardware sensors, and telemetry sources across platforms.

  • Build and maintain GPU benchmarking, stress-testing, and workload suites.

  • Deploy and execute LLM inference and training workloads, including Llama-family models.

  • Work with both quantized and full-precision ML models.

  • Develop and execute synthetic workloads such as:

  • GEMM

  • Convolution

  • Compute workloads

  • Memory-bandwidth workloads

  • Multi-GPU workloads

  • Vision models for edge platforms

  • Ensure GPU workloads are reproducible and consistent across repeated runs.

  • Collect and validate hardware telemetry using:

  • NVML

  • DCGM

  • BMC

  • IPMI

  • Redfish

  • Other external/lab-grade measurement equipment

  • Ensure consistent sampling rates, timestamps, field names, and clock alignment across telemetry sources.

  • Identify and troubleshoot missing, irregular, inconsistent, or inaccurate sensor data.

  • Automate infrastructure deployment, workload execution, logging, data collection, and cleanup.

  • Build automated data-quality checks for time-series and hardware telemetry data.

  • Detect missing samples, clock mismatches, workload/telemetry misalignment, and invalid sensor data.

  • Maintain consistent and well-documented datasets.

  • Maintain complete run metadata, including:

  • Hardware model

  • Driver version

  • Firmware version

  • Procedure version

  • Execution schedule

  • Document platform configurations, limitations, and hardware behavior.

Required Skills & Experience
  • 10 years of relevant experience preferred ; exceptional candidates with 8 years may be considered.

  • Hands-on experience deploying LLM inference and training workloads on GPUs.

  • Experience with quantized and full-precision ML models.

  • Strong hands-on experience with NVIDIA GPU environments and infrastructure.

  • Strong Linux system administration and troubleshooting skills.

  • Good understanding of:

  • GPU driver stacks

  • Linux processes

  • Process orchestration

  • Scheduling

  • Timing behavior

  • Hands-on experience collecting and analyzing hardware telemetry programmatically.

  • Experience with telemetry technologies such as NVML, DCGM, BMC, IPMI, and/or Redfish.

  • Ability to troubleshoot sensor and sampling issues in hardware time-series data.

  • Strong understanding of workload reproducibility and non-determinism.

  • Experience with automation and scripting for infrastructure deployment and workload execution.

Good to Have
  • Experience developing data-collection pipelines for:

  • Hardware testing

  • Hardware qualification

  • Systems research

  • Performance benchmarking

  • Experience with GPU benchmarking and stress-testing tools.

  • Understanding of benchmark methodology, including:

  • Warm-up

  • Steady-state execution

  • Run-to-run variance

  • Performance consistency

  • Understanding of GPU power and thermal management.

  • Familiarity with GPU clock-throttling reasons and related telemetry.

  • Experience with multi-GPU scaling.

  • Knowledge of NCCL.

  • Knowledge of tensor parallelism and pipeline parallelism.

  • Experience developing automated data-quality validation for time-series or sensor data.

Key Technical Skills

GPU: NVIDIA Data Center GPUs, Embedded/Edge GPUs

ML Workloads: LLM Inference, LLM Training, Llama, Vision Models

GPU Telemetry: NVML, DCGM

Hardware Management: BMC, IPMI, Redfish

Systems: Linux, GPU Drivers, Process Orchestration

Benchmarking: GEMM, Convolution, Compute, Memory Bandwidth, Stress Testing

Multi-GPU: NCCL, Tensor Parallelism, Pipeline Parallelism

Data: Time-Series Telemetry, Data Validation, Dataset Generation

Automation: Scripting, Infrastructure Deployment, Workload Automation

The candidate should be comfortable owning the complete workflow---from preparing a new GPU platform and deploying workloads to collecting reliable telemetry and producing reproducible, well-documented datasets.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Remote GPU & ML Infrastructure Engineer
Remote GPU & ML Infrastructure Engineer

ConsultBae India Private limited • United States

Remote
USD 150,000 - 210,000
GPU Systems Infrastructure Engineer
GPU Systems Infrastructure Engineer

Blue Signal Search • Fremont (CA)

On-site
USD 120,000 - 170,000
Cluster Engineer
Cluster Engineer

STN Inc • San Francisco (CA)

On-site
USD 180,000 - 240,000
Site Reliability Engineer
Site Reliability Engineer

GMI Cloud • United States

On-site
USD 110,000 - 170,000
Principal Infrastructure Engineer, AI Cluster Performance & Validation
Principal Infrastructure Engineer, AI Cluster Performance & Validation

Nscale • New York (NY), San Francisco (CA), Seattle (WA)

On-site
USD 180,000 - 240,000
Principal ML Infrastructure Engineer (Relocation Available)
Principal ML Infrastructure Engineer (Relocation Available)

Franklin Fitch • Dallas (TX)

On-site
USD 100,000 - 140,000
GPU Architect
GPU Architect

EngineersOfAI • Milpitas (CA)

On-site
USD 140,000 - 190,000
GPU Systems Engineer
GPU Systems Engineer

Career Techniques • New York (NY)

On-site
USD 200,000 - 300,000
AI Senior Engineer - Fulltime
AI Senior Engineer - Fulltime

IMR Soft Llc • Plano (TX)

On-site
USD 140,000 - 190,000
Member of Technical Staff - GPU Infrastructure
Member of Technical Staff - GPU Infrastructure

Hyperbolic • San Francisco (CA)

On-site
USD 180,000 - 260,000