Principal ML Infra Engineer - GPU Inference & C++ Systems

Franklin Fitch

Dallas (TX)

On-site

USD 100,000 - 140,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

A technology solutions provider in Dallas is seeking an AI Infrastructure Engineer with expertise in C++ and CUDA. The role involves designing and optimizing GPU-accelerated systems for deploying machine learning models in production. Ideal candidates will have a Master's or PhD, along with strong experience in GPU optimization and high-performance system design. Responsibilities include building inference pipelines, supporting model deployment, and working closely with ML researchers to ensure models are production-ready.

Qualifications

  • Strong C++ expertise with experience in production-grade systems.
  • Hands-on experience with CUDA programming and GPU optimization.
  • Solid understanding of GPU architectures and memory management.

Responsibilities

  • Design and maintain GPU-accelerated infrastructure for ML models.
  • Build and optimize inference pipelines for low-latency performance.
  • Support model conversion and deployment using inference runtimes.
  • Partner with ML researchers to transition models to production.

Skills

C++ expertise
CUDA programming
GPU optimization
Linux

Education

Masters or PhD

Tools

TensorRT

Job description

A technology solutions provider in Dallas is seeking an AI Infrastructure Engineer with expertise in C++ and CUDA. The role involves designing and optimizing GPU-accelerated systems for deploying machine learning models in production. Ideal candidates will have a Master's or PhD, along with strong experience in GPU optimization and high-performance system design. Responsibilities include building inference pipelines, supporting model deployment, and working closely with ML researchers to ensure models are production-ready.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Principal ML Infrastructure Engineer (Relocation Available)
Principal ML Infrastructure Engineer (Relocation Available)

Franklin Fitch • Dallas (TX)

On-site
USD 100,000 - 140,000
Senior GPU ML Infra Engineer — Mid-Training & Inference
Senior GPU ML Infra Engineer — Mid-Training & Inference

Reflection AI • San Francisco (CA)

On-site
USD 120,000 - 160,000
Staff ML Performance Engineer — Scalable Inference & CUDA
Staff ML Performance Engineer — Scalable Inference & CUDA

Modal • New York (NY)

On-site
USD 120,000 - 160,000
Senior Inference Performance Engineer - GPU & CUDA
Senior Inference Performance Engineer - GPU & CUDA

inference.net • San Francisco (CA)

Hybrid
USD 220,000 - 320,000
Engineering Manager, GPU-Accelerated LLM Inference
Engineering Manager, GPU-Accelerated LLM Inference

NVIDIA • California (MO)

On-site
USD 184,000 - 357,000
Senior ML Engineer - Low-Latency Inference & Systems
Senior ML Engineer - Low-Latency Inference & Systems

Inworld • Germany (OH)

Hybrid
USD 120,000 - 160,000
Senior Inference Performance Engineer — GPU & CUDA
Senior Inference Performance Engineer — GPU & CUDA

Inference • San Francisco (CA)

Hybrid
USD 220,000 - 320,000
Competitive compensation
Equity in a high-growth startup
Comprehensive benefits
Backend / Infra Engineer for Cloud GPU Inference
Backend / Infra Engineer for Cloud GPU Inference

Pear VC • Austin (TX), California (MO)

On-site
USD 120,000 - 150,000
Senior AI Inference Engineer — GPU-Accelerated DL Systems
Senior AI Inference Engineer — GPU-Accelerated DL Systems

NVIDIA Corporation • California (MO)

On-site
USD 152,000 - 287,500
Equity
Benefits
Senior CI Architect — GPU Inference & Open-Source Infra
Senior CI Architect — GPU Inference & Open-Source Infra

RadixArk • Palo Alto (CA)

On-site
USD 120,000 - 150,000
Equity opportunities
Comprehensive health benefits
Flexible work arrangements