Inference

Genesis AI

United States

On-site

USD 180,000 - 240,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Genesis AI is seeking a senior engineer to build and optimize real-time inference pipelines for on-device and cluster deployment. You will design GPU-accelerated systems, implement CUDA and Triton kernels, and ensure efficient memory and compute scheduling across heterogeneous stacks.

You’ll work on throughput- and latency-optimized workloads, with monitoring and debugging tools to diagnose regressions rapidly.

Qualifications

  • 8+ years in distributed systems, ML infrastructure, or high-performance serving.
  • Production-grade Python with strong C++/Rust/Go systems background.
  • Deep CUDA, Triton, kernel optimization, memory management expertise.
  • Experience scaling inference workloads for clusters and on-device deployments.
  • System-level mindset for hardware–software tuning and efficiency.

Responsibilities

  • Build low-latency inference pipelines for on-device deployment.
  • Design and optimize distributed inference systems on GPU clusters.
  • Implement low-level code (CUDA, Triton, custom kernels) into high-level frameworks.
  • Optimize workloads for throughput and latency (batching, quantization, caching).
  • Develop monitoring and debugging tools for reliability and fast regression diagnosis.

Skills

Distributed systems
ML infrastructure
Python
C++/Rust/Go
CUDA
Triton
Kernel optimization
Quantization
Memory scheduling

Tools

CUDA
Triton

Job description

What You’ll Do
  • Build low-latency inference pipelines for on-device deployment, enabling real-time next-token and diffusion-based control loops in robotics

  • Design and optimize distributed inference systems on GPU clusters, pushing throughput with large-batch serving and efficient resource utilization

  • Implement efficient low-level code (CUDA, Triton, custom kernels) and integrate it seamlessly into high-level frameworks

  • Optimize workloads for both throughput (batching, scheduling, quantization) and latency (caching, memory management, graph compilation)

  • Develop monitoring and debugging tools to guarantee reliability, determinism, and rapid diagnosis of regressions across both stacks

What You’ll Bring
  • Deep experience in distributed systems, ML infrastructure, or high-performance serving (8+ years)

  • Production-grade expertise in Python, with strong background in systems languages (C++/Rust/Go)

  • Low-level performance mastery: CUDA, Triton, kernel optimization, quantization, memory and compute scheduling

  • Proven track record scaling inference workloads in both throughput-oriented cluster environments and latency-critical on-device deployments

  • System-level mindset with a history of tuning hardware–software interactions for maximum efficiency, throughput, and responsiveness

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Inference
Inference

Genesis AI • Northern (KY)

Hybrid
USD 150,000 - 210,000
Machine Learning Engineer- Inference Optimization | Experienced Hire
Machine Learning Engineer- Inference Optimization | Experienced Hire

Susquehanna International Group, LLP • Bala Cynwyd (PA)

On-site
USD 110,000 - 150,000
Member of Technical Staff (AI Inference Engineer)
Member of Technical Staff (AI Inference Engineer)

Perplexity • San Francisco (CA)

On-site
USD 140,000 - 210,000
Training / AI Infrastructure
Training / AI Infrastructure

Genesis AI • San Francisco (CA)

On-site
USD 180,000 - 260,000
Member of Technical Staff, Inference
Member of Technical Staff, Inference

Mount Thor • San Francisco (CA)

On-site
USD 240,000 - 320,000
Member of Technical Staff (AI Inference Engineer)
Member of Technical Staff (AI Inference Engineer)

Perplexity • Palo Alto (CA)

On-site
USD 220,000 - 485,000
Senior Inference Systems Engineer (GPU/On-Device)
Senior Inference Systems Engineer (GPU/On-Device)

Genesis AI • Northern (KY)

Hybrid
USD 150,000 - 210,000
Member of Technical Staff (AI Inference Engineer)
Member of Technical Staff (AI Inference Engineer)

Perplexity • New York (NY)

On-site
USD 220,000 - 485,000
Staff Machine Learning Engineer
Staff Machine Learning Engineer

Unity • San Francisco (CA)

On-site
USD 180,000 - 240,000
Staff Machine Learning Engineer
Staff Machine Learning Engineer

Unity • Mountain View (CA)

On-site
USD 180,000 - 280,000