Real-Time Inference Architect for On-Device & GPU Clusters

Genesis

United States

Remote

USD 140,000 - 210,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Genesis is seeking a senior engineering expert to build low-latency inference pipelines for on-device deployment and real-time control loops in robotics. You will design and optimize distributed inference on GPU clusters, push throughput with large-batch serving, and implement efficient low-level code integrated into high-level frameworks.

The role emphasizes a system-level mindset, with deep experience in CUDA/Triton kernel optimization and a proven track record scaling inference workloads in

Qualifications

  • 8+ years of experience in distributed systems, ML infrastructure, or high-performance serving.
  • Production-grade Python with strong background in systems languages (C++/Rust/Go).
  • Experience optimizing CUDA/Triton kernels and memory scheduling.

Responsibilities

  • Build low-latency inference pipelines for on-device deployment enabling real-time next-token and diffusion-based control loops in robotics.
  • Design and optimize distributed inference systems on GPU clusters for high-throughput serving.
  • Implement efficient low-level code (CUDA, Triton, custom kernels) and integrate into high-level frameworks.
  • Optimize workloads for throughput and latency including caching and memory management.
  • Develop monitoring and debugging tools to ensure reliability, determinism, and rapid diagnosis of regressions.

Skills

Distributed systems
ML infrastructure
High-performance serving
Python programming
Systems languages (C++/Rust/Go)

Tools

CUDA
Triton
GPU kernel optimization

Job description

Genesis is seeking a senior engineering expert to build low-latency inference pipelines for on-device deployment and real-time control loops in robotics. You will design and optimize distributed inference on GPU clusters, push throughput with large-batch serving, and implement efficient low-level code integrated into high-level frameworks.

The role emphasizes a system-level mindset, with deep experience in CUDA/Triton kernel optimization and a proven track record scaling inference workloads in

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Inference Systems Engineer (GPU/On-Device)
Senior Inference Systems Engineer (GPU/On-Device)

Genesis AI • Northern (KY)

Hybrid
USD 150,000 - 210,000
Senior Inference Systems Engineer - Low-Latency & On-Device
Senior Inference Systems Engineer - Low-Latency & On-Device

Genesis AI • United States

On-site
USD 180,000 - 240,000
Inference
Inference

Genesis AI • Northern (KY)

Hybrid
USD 150,000 - 210,000
Inference
Inference

Genesis AI • United States

On-site
USD 180,000 - 240,000
Inference
Inference

Genesis • United States

Remote
USD 140,000 - 210,000
Real-Time GPU Optimization Engineer - Inference
Real-Time GPU Optimization Engineer - Inference

techire ai • San Francisco (CA)

On-site
USD 230,000 - 300,000
Senior GPU AI Inference Systems Engineer
Senior GPU AI Inference Systems Engineer

NVIDIA • California (MO)

On-site
USD 196,000 - 288,000
Equity
Comprehensive benefits
Remote Senior AI Inference Optimization Engineer
Remote Senior AI Inference Optimization Engineer

DigitalOcean • San Francisco (CA)

On-site
USD 191,200 - 239,000
Equity compensation
Remote work
Senior ML Training Systems Engineer - Distributed CUDA
Senior ML Training Systems Engineer - Distributed CUDA

Genesis AI • San Francisco (CA)

On-site
USD 180,000 - 260,000
Senior System Software Engineer, Dynamo-Triton Inference
Senior System Software Engineer, Dynamo-Triton Inference

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 224,000 - 357,000