Senior Inference Systems Engineer - Low-Latency & On-Device

Genesis AI

United States

On-site

USD 180,000 - 240,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Genesis AI is seeking a senior engineer to build and optimize real-time inference pipelines for on-device and cluster deployment. You will design GPU-accelerated systems, implement CUDA and Triton kernels, and ensure efficient memory and compute scheduling across heterogeneous stacks.

You’ll work on throughput- and latency-optimized workloads, with monitoring and debugging tools to diagnose regressions rapidly.

Qualifications

  • 8+ years in distributed systems, ML infrastructure, or high-performance serving.
  • Production-grade Python with strong C++/Rust/Go systems background.
  • Deep CUDA, Triton, kernel optimization, memory management expertise.
  • Experience scaling inference workloads for clusters and on-device deployments.
  • System-level mindset for hardware–software tuning and efficiency.

Responsibilities

  • Build low-latency inference pipelines for on-device deployment.
  • Design and optimize distributed inference systems on GPU clusters.
  • Implement low-level code (CUDA, Triton, custom kernels) into high-level frameworks.
  • Optimize workloads for throughput and latency (batching, quantization, caching).
  • Develop monitoring and debugging tools for reliability and fast regression diagnosis.

Skills

Distributed systems
ML infrastructure
Python
C++/Rust/Go
CUDA
Triton
Kernel optimization
Quantization
Memory scheduling

Tools

CUDA
Triton

Job description

Genesis AI is seeking a senior engineer to build and optimize real-time inference pipelines for on-device and cluster deployment. You will design GPU-accelerated systems, implement CUDA and Triton kernels, and ensure efficient memory and compute scheduling across heterogeneous stacks.

You’ll work on throughput- and latency-optimized workloads, with monitoring and debugging tools to diagnose regressions rapidly.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Inference Systems Engineer (GPU/On-Device)
Senior Inference Systems Engineer (GPU/On-Device)

Genesis AI • Northern (KY)

Hybrid
USD 150,000 - 210,000
Senior GPU AI Inference Systems Engineer
Senior GPU AI Inference Systems Engineer

NVIDIA • California (MO)

On-site
USD 196,000 - 288,000
Equity
Comprehensive benefits
Inference
Inference

Genesis AI • Northern (KY)

Hybrid
USD 150,000 - 210,000
Senior ML Training Systems Engineer - Distributed CUDA
Senior ML Training Systems Engineer - Distributed CUDA

Genesis AI • San Francisco (CA)

On-site
USD 180,000 - 260,000
Remote Senior AI Inference Optimization Engineer
Remote Senior AI Inference Optimization Engineer

DigitalOcean • San Francisco (CA)

On-site
USD 191,200 - 239,000
Equity compensation
Remote work
Inference
Inference

Genesis AI • United States

On-site
USD 180,000 - 240,000
Senior System Software Engineer, Dynamo-Triton Inference
Senior System Software Engineer, Dynamo-Triton Inference

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 224,000 - 357,000
Inference Runtime Engineer - On-Device & Cloud AI, Flexible WFH
Inference Runtime Engineer - On-Device & Cloud AI, Flexible WFH

EngRadar • New York (NY)

On-site
USD 150,000 - 230,000
Equity grants
Medical plan
Vision plan
+5
Senior AI Inference Optimization Engineer
Senior AI Inference Optimization Engineer

DigitalOcean • Austin (TX)

On-site
USD 191,200 - 239,000
On-Device AI Inference Engineer — Ultra-Low Latency
On-Device AI Inference Engineer — Ultra-Low Latency

Hark • San Jose (CA), Northern (KY)

Hybrid
USD 200,000 - 450,000