AI Senior Engineer - Fulltime

IMR Soft Llc

Plano (TX)

On-site

USD 140,000 - 190,000

Full time

12 days ago
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

IMR Soft Llc in Plano, TX is seeking an AI Senior Engineer to build and validate GPU infrastructure for distributed training and production inference. You will own the validation suite, optimize benchmarks, and deliver reports that help customers verify performance against reference architecture.

The role requires 5+ years of ML engineering with production GPU workloads, hands-on multi-node distributed training, and expertise in PyTorch DDP/FSDP, DeepSpeed, Triton, and vLLM.

Qualifications

  • 5+ years ML engineering, MLOps, or performance engineering with production GPU workloads.
  • Hands-on distributed training on multi-node GPU clusters not single-GPU or notebook-scale work.
  • Production LLM inference serving: Triton, vLLM, or TensorRT-LLM.
  • Rigorous benchmarking discipline you can profile a system, explain where GPU time actually goes, and defend a number.
  • Strong Python; containers and Kubernetes-native workflows (Kubeflow, Ray, Argo Workflows).
  • Clear technical writing. A validation report is a deliverable a customer pays for.

Responsibilities

  • Enable and tune distributed training on multi-node GPU clusters.
  • Architect inference serving with Triton, NVIDIA NIM, vLLM, TensorRT-LLM.
  • Size inference platforms by customer SLOs: latency, throughput, concurrency.
  • Support GenAI patterns: RAG, fine-tuning, agentic pipelines.
  • Build and own GPU infrastructure validation suite: NCCL scaling, MFU baselines.
  • Deliver acceptance validation engagements with signed reports for customers.
  • Design the continuous-validation service: post-downtime revalidation, driver/firmware certification.
  • Produce published benchmark methodology and reports to establish credibility.
  • Own PoC workload design and benchmark evidence to close engagements.

Skills

Distributed training
PyTorch DDP/FSDP
DeepSpeed/ZeRO
NCCL tuning
Python
Kubernetes
Kubeflow
vLLM
Triton
TensorRT-LLM
Benchmarking
Technical writing

Job description

AI Senior Engineer

Plano, TX

Fulltime permanent

About the role

Two halves, and both matter. You make the platform useful enabling distributed training and production inference on the clusters we build. And you own our infrastructure validation offering: the independent, evidence-based assessment that tells a customer whether the eight-figure GPU estate they just bought actually performs the way the reference architecture promised.

If you like building the benchmark that settles the argument, this is your role.

What you ll do
  • Workload engineering - Enable and tune distributed training: PyTorch DDP and FSDP, DeepSpeed/ZeRO, tensor and pipeline parallelism, NCCL tuning, MFU measurement and improvement.
  • Architect inference serving Triton, NVIDIA NIM, vLLM, TensorRT-LLM with quantization, continuous batching, KV-cache and paged attention, and autoscaling to latency SLOs.
  • Size inference platforms backward from customer SLOs: time-to-first-token, inter-token latency, concurrency, context length.
  • Support the GenAI patterns customers actually ask for RAG, fine-tuning, agentic pipelines, vector database integration.
  • Build and own our GPU infrastructure validation suite: fabric validation, NCCL scaling curves, GPU burn and thermal/power soak, storage throughput, GPUDirect verification, reference-workload MFU baselines.
  • Deliver acceptance validation engagements and author the signed acceptance reports customers use to hold vendors to their commitments.
  • Design the continuous-validation service: post-downtime revalidation, driver and firmware matrix certification, performance drift detection against baseline.
  • Produce published benchmark methodology and reports that establish our technical credibility in the market.
  • Own PoC workload design and the benchmark evidence that closes engagements.
What you need
  • 5+ years ML engineering, MLOps, or performance engineering with production GPU workloads.
  • Hands-on distributed training on multi-node GPU clusters not single-GPU or notebook-scale work.
  • Production LLM inference serving: Triton, vLLM, or TensorRT-LLM.
  • Rigorous benchmarking discipline you can profile a system, explain where GPU time actually goes, and defend a number.
  • Strong Python; containers and Kubernetes-native workflows (Kubeflow, Ray, Argo Workflows).
  • Clear technical writing. A validation report is a deliverable a customer pays for.
Nice to have
  • NVIDIA NIM and NeMo; MLPerf or formal benchmark program experience; test-and-validation engineering background; fine-tuning at scale (LoRA/QLoRA); customer-facing PoC delivery.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Product Manager - AI Inference Performance
Senior Product Manager - AI Inference Performance

NVIDIA • United States

On-site
USD 180,000 - 260,000
Principal Infrastructure Engineer, AI Cluster Performance & Validation
Principal Infrastructure Engineer, AI Cluster Performance & Validation

Nscale • New York (NY), San Francisco (CA), Seattle (WA)

On-site
USD 180,000 - 240,000
Cluster Engineer
Cluster Engineer

STN Inc • San Francisco (CA)

On-site
USD 180,000 - 240,000
Senior AI Benchmark Engineer - GPU Systems & Inference
Senior AI Benchmark Engineer - GPU Systems & Inference

IMR Soft Llc • Plano (TX)

On-site
USD 140,000 - 190,000
Member of Technical Staff (AI Inference Engineer)
Member of Technical Staff (AI Inference Engineer)

Perplexity • San Francisco (CA)

On-site
USD 140,000 - 210,000
Staff Applied AI Inference Engineer
Staff Applied AI Inference Engineer

Crusoe Energy Systems • San Francisco (CA)

On-site
USD 180,000 - 240,000
Health benefits
401(k) match
Paid time off
+1
Member of Technical Staff - GPU Infrastructure
Member of Technical Staff - GPU Infrastructure

Prime Intellect • San Francisco (CA), Northern (KY)

On-site
USD 150,000 - 300,000
Equity incentives
Senior AI Performance Engineer
Senior AI Performance Engineer

Brillfy Technology Inc • United States

On-site
USD 150,000 - 210,000
Member of Technical Staff, AI Compute & Data Infrastructure Vinci
Member of Technical Staff, AI Compute & Data Infrastructure Vinci

CDFAM - Computational Design Symposium • Palo Alto (CA), Northern (KY)

On-site
USD 180,000 - 240,000
AI/ML Engineer – (Next-Generation AI Platforms & Workloads)
AI/ML Engineer – (Next-Generation AI Platforms & Workloads)

VeeAR Projects Inc. • Sunnyvale (CA)

On-site
USD 140,000 - 210,000