Inference

Genesis AI

Northern (KY)

Hybrid

USD 150,000 - 210,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Genesis AI is seeking an experienced ML infrastructure engineer to build low-latency inference pipelines and optimize distributed systems across GPU clusters for real-time robotics applications.

You will implement CUDA, Triton, and custom kernels, and integrate them into high-level frameworks, while developing monitoring tools to ensure reliability, determinism, and rapid diagnosis of regressions across both stacks.

Qualifications

  • 8+ years of experience in distributed systems or ML infrastructure.
  • Production-grade Python programming and experience with systems languages (C++, Rust, Go).
  • Low-level performance mastery: CUDA, Triton, kernel optimization and memory management.

Responsibilities

  • Build low-latency inference pipelines for on-device deployment and real-time control loops in robotics.
  • Design and optimize distributed inference systems on GPU clusters for high throughput and efficient resource utilization.
  • Implement efficient low-level code (CUDA, Triton, custom kernels) and integrate into high-level frameworks.
  • Optimize workloads for throughput and latency (batching, scheduling, quantization, caching, memory management).
  • Develop monitoring and debugging tools to guarantee reliability, determinism, and rapid diagnosis of regressions across stacks.

Skills

Distributed systems
ML infrastructure
High-performance serving
Python
C++/Rust/Go

Tools

CUDA
Triton
Kernel optimization
Quantization
Memory scheduling
GPU clusters

Job description

What You’ll Do
  • Build low-latency inference pipelines for on-device deployment, enabling real-time next-token and diffusion-based control loops in robotics

  • Design and optimize distributed inference systems on GPU clusters, pushing throughput with large-batch serving and efficient resource utilization

  • Implement efficient low-level code (CUDA, Triton, custom kernels) and integrate it seamlessly into high-level frameworks

  • Optimize workloads for both throughput (batching, scheduling, quantization) and latency (caching, memory management, graph compilation)

  • Develop monitoring and debugging tools to guarantee reliability, determinism, and rapid diagnosis of regressions across both stacks

What You’ll Bring
  • Deep experience in distributed systems, ML infrastructure, or high-performance serving (8+ years)

  • Production-grade expertise in Python, with strong background in systems languages (C++/Rust/Go)

  • Low-level performance mastery: CUDA, Triton, kernel optimization, quantization, memory and compute scheduling

  • Proven track record scaling inference workloads in both throughput-oriented cluster environments and latency-critical on-device deployments

  • System-level mindset with a history of tuning hardware–software interactions for maximum efficiency, throughput, and responsiveness

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Inference
Inference

Genesis • United States

Remote
USD 140,000 - 210,000
Inference
Inference

Genesis AI • United States

On-site
USD 180,000 - 240,000
Machine Learning Engineer- Inference Optimization | Experienced Hire
Machine Learning Engineer- Inference Optimization | Experienced Hire

Susquehanna International Group, LLP • Bala Cynwyd (PA)

On-site
USD 110,000 - 150,000
Member of Technical Staff (AI Inference Engineer)
Member of Technical Staff (AI Inference Engineer)

Perplexity • San Francisco (CA)

On-site
USD 140,000 - 210,000
Training / AI Infrastructure
Training / AI Infrastructure

Genesis AI • San Francisco (CA)

On-site
USD 180,000 - 260,000
Real-Time Inference Architect for On-Device & GPU Clusters
Real-Time Inference Architect for On-Device & GPU Clusters

Genesis • United States

Remote
USD 140,000 - 210,000
Member of Technical Staff, Inference
Member of Technical Staff, Inference

Mount Thor • San Francisco (CA)

On-site
USD 240,000 - 320,000
Senior Inference Systems Engineer (GPU/On-Device)
Senior Inference Systems Engineer (GPU/On-Device)

Genesis AI • Northern (KY)

Hybrid
USD 150,000 - 210,000
Member of Technical Staff (AI Inference Engineer)
Member of Technical Staff (AI Inference Engineer)

Perplexity • New York (NY)

On-site
USD 220,000 - 485,000
Member of Technical Staff (AI Inference Engineer)
Member of Technical Staff (AI Inference Engineer)

Perplexity • Palo Alto (CA)

On-site
USD 220,000 - 485,000