Principal ML Infrastructure Engineer (Relocation Available)

Franklin Fitch

Dallas (TX)

On-site

USD 100,000 - 140,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

A technology solutions provider in Dallas is seeking an AI Infrastructure Engineer with expertise in C++ and CUDA. The role involves designing and optimizing GPU-accelerated systems for deploying machine learning models in production. Ideal candidates will have a Master's or PhD, along with strong experience in GPU optimization and high-performance system design. Responsibilities include building inference pipelines, supporting model deployment, and working closely with ML researchers to ensure models are production-ready.

Qualifications

  • Strong C++ expertise with experience in production-grade systems.
  • Hands-on experience with CUDA programming and GPU optimization.
  • Solid understanding of GPU architectures and memory management.

Responsibilities

  • Design and maintain GPU-accelerated infrastructure for ML models.
  • Build and optimize inference pipelines for low-latency performance.
  • Support model conversion and deployment using inference runtimes.
  • Partner with ML researchers to transition models to production.

Skills

C++ expertise
CUDA programming
GPU optimization
Linux

Education

Masters or PhD

Tools

TensorRT

Job description

Overview

AI Infrastructure Engineer (GPU Systems & Model Deployment) (Principal and Entry level available)

We are seeking an AI Infrastructure Engineer to design and optimize high-performance systems that enable machine learning models to run reliably and efficiently in production environments. This role is focused on GPU-accelerated inference, low-latency model serving, and bridging the gap between research models and real-world deployment. You will work closely with ML researchers and software engineers to ensure models are production-ready, scalable, and performant.

This is a hands-on systems role with a strong emphasis on C++, CUDA, and GPU inference optimisation.

Core Responsibilities
  • Design and maintain GPU-accelerated infrastructure for deploying machine learning models in production
  • Build and optimize high-throughput, low-latency inference pipelines
  • Develop and maintain performance-critical components in C++
  • Optimize GPU utilization through CUDA programming and kernel tuning
  • Support model conversion, optimization, and deployment using inference runtimes
  • Partner with ML researchers to transition models from experimentation to production
  • Diagnose and improve system performance relative to baseline benchmarks
  • Ensure deployed systems are reliable, observable, and maintainable in production environments
Required Qualifications
  • Masters or PhD required
  • Strong C++ expertise with experience writing and optimizing production-grade systems
  • Hands-on CUDA programming experience and GPU performance optimization
  • Solid understanding of GPU architectures and memory management
Preferred / Nice-to-Have Qualifications
  • Experience with TensorRT or similar GPU inference runtimes
  • 1–7 years of experience as a Software Development Engineer supporting production model deployment
  • Experience with model optimization, quantization, or runtime acceleration techniques
  • Exposure to ML frameworks (e.g., PyTorch, TensorFlow) from a systems or deployment perspective
  • Experience working with containerized environments and CI/CD pipelines
Tech Environment (Representative, Not Exhaustive)
  • C++, CUDA
  • GPU inference runtimes (e.g., TensorRT)
  • Linux, containers, cloud or on-prem GPU systems
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Machine Learning Engineer
Senior Machine Learning Engineer

Acceler8 Talent • San Francisco (CA), Northern (KY)

Hybrid
USD 170,000 - 240,000
Principal ML Infra Engineer - GPU Inference & C++ Systems
Principal ML Infra Engineer - GPU Inference & C++ Systems

Franklin Fitch • Dallas (TX)

On-site
USD 100,000 - 140,000
AI/ML Infra Engineer - Hosting
AI/ML Infra Engineer - Hosting

Hamilton Barnes Associates Limited • San Francisco (CA)

On-site
USD 225,000 - 275,000
Stock options
Senior ML Performance Engineer
Senior ML Performance Engineer

well-funded deeptech startup • California (MO)

On-site
USD 200,000 - 250,000
AI/ML Engineer – (Next-Generation AI Platforms & Workloads)
AI/ML Engineer – (Next-Generation AI Platforms & Workloads)

VeeAR Projects Inc. • Sunnyvale (CA)

On-site
USD 140,000 - 210,000
Software Engineer: ML Infra
Software Engineer: ML Infra

Generalist • Somerville (MA), San Mateo (CA)

On-site
USD 120,000 - 160,000
Senior ML Infrastructure Engineer - GPU Training & Inference
Senior ML Infrastructure Engineer - GPU Training & Inference

TensorWave • Las Vegas (NM)

On-site
Senior Software Engineer, Infrastructure Software for AI (Centralized AI Data Centers & Distrib[...]
Senior Software Engineer, Infrastructure Software for AI (Centralized AI Data Centers & Distrib[...]

Intelliswift - An LTTS Company • Sunnyvale (CA)

On-site
USD 120,000 - 150,000
Competitive salary
Health insurance
Flexible work hours
ML Inference Engineer San Francisco · Engineering · Full Time →
ML Inference Engineer San Francisco · Engineering · Full Time →

Reactor • San Francisco (CA)

On-site
USD 120,000 - 160,000
Visa sponsorship
Relocation support
Generous health, dental, and vision coverage
MLOps Engineer
MLOps Engineer

Blue Signal Search • Santa Clara (CA)

On-site
USD 140,000 - 190,000
Advanced GPU infra exposure
Collaborative engineering culture
Open source AI frameworks access
+2