Inference

Genesis

United States

Remote

USD 140,000 - 210,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Genesis is seeking a senior engineering expert to build low-latency inference pipelines for on-device deployment and real-time control loops in robotics. You will design and optimize distributed inference on GPU clusters, push throughput with large-batch serving, and implement efficient low-level code integrated into high-level frameworks.

The role emphasizes a system-level mindset, with deep experience in CUDA/Triton kernel optimization and a proven track record scaling inference workloads in

Qualifications

  • 8+ years of experience in distributed systems, ML infrastructure, or high-performance serving.
  • Production-grade Python with strong background in systems languages (C++/Rust/Go).
  • Experience optimizing CUDA/Triton kernels and memory scheduling.

Responsibilities

  • Build low-latency inference pipelines for on-device deployment enabling real-time next-token and diffusion-based control loops in robotics.
  • Design and optimize distributed inference systems on GPU clusters for high-throughput serving.
  • Implement efficient low-level code (CUDA, Triton, custom kernels) and integrate into high-level frameworks.
  • Optimize workloads for throughput and latency including caching and memory management.
  • Develop monitoring and debugging tools to ensure reliability, determinism, and rapid diagnosis of regressions.

Skills

Distributed systems
ML infrastructure
High-performance serving
Python programming
Systems languages (C++/Rust/Go)

Tools

CUDA
Triton
GPU kernel optimization

Job description

What You’ll Do
  • Build low-latency inference pipelines for on-device deployment, enabling real-time next-token and diffusion-based control loops in robotics

  • Design and optimize distributed inference systems on GPU clusters, pushing throughput with large-batch serving and efficient resource utilization

  • Implement efficient low-level code (CUDA, Triton, custom kernels) and integrate it seamlessly into high-level frameworks

  • Optimize workloads for both throughput (batching, scheduling, quantization) and latency (caching, memory management, graph compilation)

  • Develop monitoring and debugging tools to guarantee reliability, determinism, and rapid diagnosis of regressions across both stacks

What You’ll Bring
  • Deep experience in distributed systems, ML infrastructure, or high-performance serving (8+ years)

  • Production-grade expertise in Python, with strong background in systems languages (C++/Rust/Go)

  • Low-level performance mastery: CUDA, Triton, kernel optimization, quantization, memory and compute scheduling

  • Proven track record scaling inference workloads in both throughput-oriented cluster environments and latency-critical on-device deployments

  • System-level mindset with a history of tuning hardware–software interactions for maximum efficiency, throughput, and responsiveness

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Inference
Inference

Genesis AI • United States

On-site
USD 180,000 - 240,000
Inference
Inference

Genesis AI • Northern (KY)

Hybrid
USD 150,000 - 210,000
Machine Learning Engineer- Inference Optimization | Experienced Hire
Machine Learning Engineer- Inference Optimization | Experienced Hire

Susquehanna International Group, LLP • Bala Cynwyd (PA)

On-site
USD 110,000 - 150,000
Member of Technical Staff (AI Inference Engineer)
Member of Technical Staff (AI Inference Engineer)

Perplexity • San Francisco (CA)

On-site
USD 140,000 - 210,000
Training / AI Infrastructure
Training / AI Infrastructure

Genesis AI • San Francisco (CA)

On-site
USD 180,000 - 260,000
Member of Technical Staff, Inference
Member of Technical Staff, Inference

Mount Thor • San Francisco (CA)

On-site
USD 240,000 - 320,000
Member of Technical Staff (AI Inference Engineer)
Member of Technical Staff (AI Inference Engineer)

Perplexity • New York (NY)

On-site
USD 220,000 - 485,000
Real-Time Inference Architect for On-Device & GPU Clusters
Real-Time Inference Architect for On-Device & GPU Clusters

Genesis • United States

Remote
USD 140,000 - 210,000
Member of Technical Staff (AI Inference Engineer)
Member of Technical Staff (AI Inference Engineer)

Perplexity • Palo Alto (CA)

On-site
USD 220,000 - 485,000
Senior Inference Systems Engineer (GPU/On-Device)
Senior Inference Systems Engineer (GPU/On-Device)

Genesis AI • Northern (KY)

Hybrid
USD 150,000 - 210,000