A complete application in a minute — tailored resume and cover letter, ready to send.
ConsultBae India Private Limited is seeking a seasoned GPU & ML Infrastructure Engineer to build, test, monitor, and maintain GPU-based benchmarking and ML workloads. You will own end-to-end data generation, telemetry collection, and reproducible results.
Experience with NVIDIA data-center GPUs, Linux systems, and automation is essential to deploy workloads, validate data, and document platform configurations across edge and data-center platforms.
GPU & ML Infrastructure Engineer
Experience: 10 Years
Location: Remote
We are looking for an experienced GPU & ML Infrastructure Engineer to build, test, monitor, and maintain reliable infrastructure for GPU-based benchmarking, stress testing, telemetry collection, and machine learning workloads.
The role involves working with NVIDIA data-center GPUs and embedded/edge platforms, deploying ML workloads, collecting hardware telemetry, automating testing processes, and ensuring that the resulting data is accurate, consistent, and reproducible.
The engineer will own the end-to-end testing and data-generation process, from preparing new GPU platforms and deploying workloads to collecting telemetry, validating data, and documenting results.
Port and adapt existing GPU test procedures to new NVIDIA data-center and edge/embedded platforms.
Analyze differences in GPU drivers, power/thermal limits, hardware sensors, and telemetry sources across platforms.
Build and maintain GPU benchmarking, stress-testing, and workload suites.
Deploy and execute LLM inference and training workloads, including Llama-family models.
Work with both quantized and full-precision ML models.
Develop and execute synthetic workloads such as:
GEMM
Convolution
Compute workloads
Memory-bandwidth workloads
Multi-GPU workloads
Vision models for edge platforms
Ensure GPU workloads are reproducible and consistent across repeated runs.
Collect and validate hardware telemetry using:
NVML
DCGM
BMC
IPMI
Redfish
Other external/lab-grade measurement equipment
Ensure consistent sampling rates, timestamps, field names, and clock alignment across telemetry sources.
Identify and troubleshoot missing, irregular, inconsistent, or inaccurate sensor data.
Automate infrastructure deployment, workload execution, logging, data collection, and cleanup.
Build automated data-quality checks for time-series and hardware telemetry data.
Detect missing samples, clock mismatches, workload/telemetry misalignment, and invalid sensor data.
Maintain consistent and well-documented datasets.
Maintain complete run metadata, including:
Hardware model
Driver version
Firmware version
Procedure version
Execution schedule
Document platform configurations, limitations, and hardware behavior.
10 years of relevant experience preferred ; exceptional candidates with 8 years may be considered.
Hands-on experience deploying LLM inference and training workloads on GPUs.
Experience with quantized and full-precision ML models.
Strong hands-on experience with NVIDIA GPU environments and infrastructure.
Strong Linux system administration and troubleshooting skills.
Good understanding of:
GPU driver stacks
Linux processes
Process orchestration
Scheduling
Timing behavior
Hands-on experience collecting and analyzing hardware telemetry programmatically.
Experience with telemetry technologies such as NVML, DCGM, BMC, IPMI, and/or Redfish.
Ability to troubleshoot sensor and sampling issues in hardware time-series data.
Strong understanding of workload reproducibility and non-determinism.
Experience with automation and scripting for infrastructure deployment and workload execution.
Experience developing data-collection pipelines for:
Hardware testing
Hardware qualification
Systems research
Performance benchmarking
Experience with GPU benchmarking and stress-testing tools.
Understanding of benchmark methodology, including:
Warm-up
Steady-state execution
Run-to-run variance
Performance consistency
Understanding of GPU power and thermal management.
Familiarity with GPU clock-throttling reasons and related telemetry.
Experience with multi-GPU scaling.
Knowledge of NCCL.
Knowledge of tensor parallelism and pipeline parallelism.
Experience developing automated data-quality validation for time-series or sensor data.
GPU: NVIDIA Data Center GPUs, Embedded/Edge GPUs
ML Workloads: LLM Inference, LLM Training, Llama, Vision Models
GPU Telemetry: NVML, DCGM
Hardware Management: BMC, IPMI, Redfish
Systems: Linux, GPU Drivers, Process Orchestration
Benchmarking: GEMM, Convolution, Compute, Memory Bandwidth, Stress Testing
Multi-GPU: NCCL, Tensor Parallelism, Pipeline Parallelism
Data: Time-Series Telemetry, Data Validation, Dataset Generation
Automation: Scripting, Infrastructure Deployment, Workload Automation
The candidate should be comfortable owning the complete workflow---from preparing a new GPU platform and deploying workloads to collecting reliable telemetry and producing reproducible, well-documented datasets.