Senior Inference Systems Engineer (GPU/On-Device)

Genesis AI

Northern (KY)

Hybrid

USD 150,000 - 210,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Genesis AI is seeking an experienced ML infrastructure engineer to build low-latency inference pipelines and optimize distributed systems across GPU clusters for real-time robotics applications.

You will implement CUDA, Triton, and custom kernels, and integrate them into high-level frameworks, while developing monitoring tools to ensure reliability, determinism, and rapid diagnosis of regressions across both stacks.

Qualifications

  • 8+ years of experience in distributed systems or ML infrastructure.
  • Production-grade Python programming and experience with systems languages (C++, Rust, Go).
  • Low-level performance mastery: CUDA, Triton, kernel optimization and memory management.

Responsibilities

  • Build low-latency inference pipelines for on-device deployment and real-time control loops in robotics.
  • Design and optimize distributed inference systems on GPU clusters for high throughput and efficient resource utilization.
  • Implement efficient low-level code (CUDA, Triton, custom kernels) and integrate into high-level frameworks.
  • Optimize workloads for throughput and latency (batching, scheduling, quantization, caching, memory management).
  • Develop monitoring and debugging tools to guarantee reliability, determinism, and rapid diagnosis of regressions across stacks.

Skills

Distributed systems
ML infrastructure
High-performance serving
Python
C++/Rust/Go

Tools

CUDA
Triton
Kernel optimization
Quantization
Memory scheduling
GPU clusters

Job description

Genesis AI is seeking an experienced ML infrastructure engineer to build low-latency inference pipelines and optimize distributed systems across GPU clusters for real-time robotics applications.

You will implement CUDA, Triton, and custom kernels, and integrate them into high-level frameworks, while developing monitoring tools to ensure reliability, determinism, and rapid diagnosis of regressions across both stacks.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Inference Systems Engineer - Low-Latency & On-Device
Senior Inference Systems Engineer - Low-Latency & On-Device

Genesis AI • United States

On-site
USD 180,000 - 240,000
Senior ML Training Systems Engineer - Distributed CUDA
Senior ML Training Systems Engineer - Distributed CUDA

Genesis AI • San Francisco (CA)

On-site
USD 180,000 - 260,000
Senior GPU AI Inference Systems Engineer
Senior GPU AI Inference Systems Engineer

NVIDIA • California (MO)

On-site
USD 196,000 - 288,000
Equity
Comprehensive benefits
Inference
Inference

Genesis AI • Northern (KY)

Hybrid
USD 150,000 - 210,000
Inference
Inference

Genesis AI • United States

On-site
USD 180,000 - 240,000
Senior AI Inference Systems Engineer (GPU, Open Source)
Senior AI Inference Systems Engineer (GPU, Open Source)

NVIDIA Corporation • Santa Clara (CA)

Hybrid
USD 224,000 - 431,000
Equity
Employee benefits
Senior ML Infra Engineer for Distributed GPU Training
Senior ML Infra Engineer for Distributed GPU Training

Genesis Molecular AI • City of Utica (NY)

On-site
USD 150,000 - 190,000
Competitive compensation with salary +
Senior System Software Engineer, Dynamo-Triton Inference
Senior System Software Engineer, Dynamo-Triton Inference

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 224,000 - 357,000
Senior DL Inference Engineer — GPU-Optimized LLMs (Remote)
Senior DL Inference Engineer — GPU-Optimized LLMs (Remote)

NVIDIA Corporation • Northern (KY)

Hybrid
USD 152,000 - 288,000
Remote Senior AI Inference Optimization Engineer
Remote Senior AI Inference Optimization Engineer

DigitalOcean • San Francisco (CA)

On-site
USD 191,200 - 239,000
Equity compensation
Remote work