Principal ML Infra Engineer - GPU Inference & C++ Systems

Franklin Fitch

Dallas (TX)

On-site

USD 100,000 - 140,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

A technology solutions provider in Dallas is seeking an AI Infrastructure Engineer with expertise in C++ and CUDA. The role involves designing and optimizing GPU-accelerated systems for deploying machine learning models in production. Ideal candidates will have a Master's or PhD, along with strong experience in GPU optimization and high-performance system design. Responsibilities include building inference pipelines, supporting model deployment, and working closely with ML researchers to ensure models are production-ready.

Qualifications

  • Strong C++ expertise with experience in production-grade systems.
  • Hands-on experience with CUDA programming and GPU optimization.
  • Solid understanding of GPU architectures and memory management.

Responsibilities

  • Design and maintain GPU-accelerated infrastructure for ML models.
  • Build and optimize inference pipelines for low-latency performance.
  • Support model conversion and deployment using inference runtimes.
  • Partner with ML researchers to transition models to production.

Skills

C++ expertise
CUDA programming
GPU optimization
Linux

Education

Masters or PhD

Tools

TensorRT

Job description

A technology solutions provider in Dallas is seeking an AI Infrastructure Engineer with expertise in C++ and CUDA. The role involves designing and optimizing GPU-accelerated systems for deploying machine learning models in production. Ideal candidates will have a Master's or PhD, along with strong experience in GPU optimization and high-performance system design. Responsibilities include building inference pipelines, supporting model deployment, and working closely with ML researchers to ensure models are production-ready.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Principal ML Infrastructure Engineer (Relocation Available)
Principal ML Infrastructure Engineer (Relocation Available)

Franklin Fitch • Dallas (TX)

On-site
USD 100,000 - 140,000
Senior Inference Performance Engineer - GPU & CUDA
Senior Inference Performance Engineer - GPU & CUDA

inference.net • San Francisco (CA)

Hybrid
USD 220,000 - 320,000
Equity in a high-growth startup
Comprehensive benefits
Senior Inference Performance Engineer — GPU & CUDA
Senior Inference Performance Engineer — GPU & CUDA

Inference • San Francisco (CA)

Hybrid
USD 220,000 - 320,000
Competitive compensation
Equity in a high-growth startup
Comprehensive benefits
ML Infra Engineer: Scale GPU Training & Inference
ML Infra Engineer: Scale GPU Training & Inference

Reducto • San Francisco (CA)

On-site
USD 120,000 - 160,000
Unlimited PTO
Free lunch
Reimbursed transportation
+3
ML Infrastructure Engineer: GPU Training & Serving
ML Infrastructure Engineer: GPU Training & Serving

Character.AI • San Francisco (CA)

On-site
USD 130,000 - 207,000
Health insurance
Flexible working hours
Professional development opportunities
Senior ML Systems Engineer - Model Inference & Efficiency
Senior ML Systems Engineer - Model Inference & Efficiency

Cohere • New York (NY)

Hybrid
USD 100,000 - 150,000
Inclusive culture and work environment
Weekly lunch stipend, in-office lunches & snacks
Full health and dental benefits
+4
Software Engineer: ML Infra
Software Engineer: ML Infra

Generalist • Somerville (MA), San Mateo (CA)

On-site
USD 120,000 - 160,000
On-Prem LLM Inference Engineer: GPU & AI Infra
On-Prem LLM Inference Engineer: GPU & AI Infra

Compunnel, Inc. • Charlotte (NC)

On-site
USD 120,000 - 150,000
Senior ML Infrastructure Engineer - GPU & Scale
Senior ML Infrastructure Engineer - GPU & Scale

TensorWave • Las Vegas (NV)

On-site
USD 120,000 - 150,000
Competitive Salary
Stock Options
100% paid Medical, Dental, and Vision insurance
+8
Senior ML Training Systems Engineer - Distributed GPU Infra
Senior ML Training Systems Engineer - Distributed GPU Infra

Baseten • San Francisco (CA)

On-site
USD 150,000 - 200,000
Competitive compensation, including equity
100% coverage of medical, dental, and vision insurance
Generous PTO policy
+2