Senior AI Benchmark Engineer - GPU Systems & Inference

IMR Soft Llc

Plano (TX)

On-site

USD 140,000 - 190,000

Full time

12 days ago
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

IMR Soft Llc in Plano, TX is seeking an AI Senior Engineer to build and validate GPU infrastructure for distributed training and production inference. You will own the validation suite, optimize benchmarks, and deliver reports that help customers verify performance against reference architecture.

The role requires 5+ years of ML engineering with production GPU workloads, hands-on multi-node distributed training, and expertise in PyTorch DDP/FSDP, DeepSpeed, Triton, and vLLM.

Qualifications

  • 5+ years ML engineering, MLOps, or performance engineering with production GPU workloads.
  • Hands-on distributed training on multi-node GPU clusters not single-GPU or notebook-scale work.
  • Production LLM inference serving: Triton, vLLM, or TensorRT-LLM.
  • Rigorous benchmarking discipline you can profile a system, explain where GPU time actually goes, and defend a number.
  • Strong Python; containers and Kubernetes-native workflows (Kubeflow, Ray, Argo Workflows).
  • Clear technical writing. A validation report is a deliverable a customer pays for.

Responsibilities

  • Enable and tune distributed training on multi-node GPU clusters.
  • Architect inference serving with Triton, NVIDIA NIM, vLLM, TensorRT-LLM.
  • Size inference platforms by customer SLOs: latency, throughput, concurrency.
  • Support GenAI patterns: RAG, fine-tuning, agentic pipelines.
  • Build and own GPU infrastructure validation suite: NCCL scaling, MFU baselines.
  • Deliver acceptance validation engagements with signed reports for customers.
  • Design the continuous-validation service: post-downtime revalidation, driver/firmware certification.
  • Produce published benchmark methodology and reports to establish credibility.
  • Own PoC workload design and benchmark evidence to close engagements.

Skills

Distributed training
PyTorch DDP/FSDP
DeepSpeed/ZeRO
NCCL tuning
Python
Kubernetes
Kubeflow
vLLM
Triton
TensorRT-LLM
Benchmarking
Technical writing

Job description

IMR Soft Llc in Plano, TX is seeking an AI Senior Engineer to build and validate GPU infrastructure for distributed training and production inference. You will own the validation suite, optimize benchmarks, and deliver reports that help customers verify performance against reference architecture.

The role requires 5+ years of ML engineering with production GPU workloads, hands-on multi-node distributed training, and expertise in PyTorch DDP/FSDP, DeepSpeed, Triton, and vLLM.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

AI Senior Engineer - Fulltime
AI Senior Engineer - Fulltime

IMR Soft Llc • Plano (TX)

On-site
USD 140,000 - 190,000
Senior AI Infra Engineer — Large-Scale GPU Cloud Equity
Senior AI Infra Engineer — Large-Scale GPU Cloud Equity

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Remote Senior AI Inference Optimization Engineer
Remote Senior AI Inference Optimization Engineer

DigitalOcean • San Francisco (CA)

On-site
USD 191,200 - 239,000
Equity compensation
Remote work
Senior DL Inference Engineer — GPU-Accelerated AI, Equity
Senior DL Inference Engineer — GPU-Accelerated AI, Equity

NVIDIA Gruppe • California (MO)

On-site
USD 152,000 - 288,000
Senior AI Inference Performance Engineer
Senior AI Inference Performance Engineer

NVIDIA • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Lead AI Cloud Architect: Scalable GPU Clusters
Lead AI Cloud Architect: Scalable GPU Clusters

IREN • United States

On-site
USD 180,000 - 260,000
Medical Insurance
Dental Insurance
Vision Insurance
+8
Senior DL Inference Engineer — GPU-Optimized LLMs (Remote)
Senior DL Inference Engineer — GPU-Optimized LLMs (Remote)

NVIDIA Corporation • Northern (KY)

Hybrid
USD 152,000 - 288,000
Senior AI Infrastructure Lead: GPU Clusters & LLMs
Senior AI Infrastructure Lead: GPU Clusters & LLMs

Cadence Design Systems • San Jose (CA)

On-site
USD 137,000 - 254,000
Senior AI Infra Engineer-Distributed GPU Clusters (Equity)
Senior AI Infra Engineer-Distributed GPU Clusters (Equity)

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 184,000 - 356,500
AI/ML Engineer – (Next-Generation AI Platforms & Workloads)
AI/ML Engineer – (Next-Generation AI Platforms & Workloads)

VeeAR Projects Inc. • Sunnyvale (CA)

On-site
USD 140,000 - 210,000