Senior ML Engineer - GPU Inference & Large-Scale AI Cloud

Nebius

United States

Remote

USD 180,000 - 250,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Nebius is building a high-performance inference and fine-tuning platform in a leading GPU cloud. We aim to push foundation models to hardware limits with maximum throughput, minimal latency, and optimised cost-per-token across thousands of GPUs.

Join Token Factory within Nebius Cloud to work on inference optimization, low-precision training, and novel architectures, contributing to open-source and in-house components. A strong ML/engineering background is essential.

Qualifications

  • Strong foundations in machine learning and transformer architecture.
  • Experience profiling GPU workloads with profiling tools.
  • Proficiency in Python and modern deep learning frameworks.

Responsibilities

  • Develop and optimize inference and fine-tuning pipelines at scale.
  • Profile and improve GPU utilization, latency, and throughput.

Skills

Transformer architectures
Gpu profiling
Python
CI/CD
Deep learning frameworks
Communication

Tools

Nsight
PyTorch
Triton
CUDA

Job description

Nebius is building a high-performance inference and fine-tuning platform in a leading GPU cloud. We aim to push foundation models to hardware limits with maximum throughput, minimal latency, and optimised cost-per-token across thousands of GPUs.

Join Token Factory within Nebius Cloud to work on inference optimization, low-precision training, and novel architectures, contributing to open-source and in-house components. A strong ML/engineering background is essential.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

GPU Systems Engineer for AI Inference & Performance
GPU Systems Engineer for AI Inference & Performance

Nebius • United States

Remote
USD 180,000 - 260,000
Competitive compensation
Career growth
Flexible ownership & autonomy
+3
GPU ML Benchmarking Engineer for Next-Gen AI Infra
GPU ML Benchmarking Engineer for Next-Gen AI Infra

Nebius • Amsterdam (VA)

On-site
USD 130,000 - 190,000
Competitive compensation
Career growth and learning
Flexibility and ownership
+2
GPU Benchmark Engineer for AI Cloud Infra
GPU Benchmark Engineer for AI Cloud Infra

Nebius • United States

Remote
USD 150,000 - 190,000
Senior ML Engineer: AI Inference & Performance Optimizer
Senior ML Engineer: AI Inference & Performance Optimizer

Nebius • Palo Alto (CA)

Hybrid
USD 195,000 - 263,000
Health insurance
401(k) plan
Parental leave
+2
ML Infrastructure Engineer
ML Infrastructure Engineer

Nebius • Amsterdam (VA)

On-site
USD 130,000 - 190,000
Competitive compensation
Career growth and learning
Flexibility and ownership
+2
ML Systems Engineer: Distributed GPU Training & RL Pipelines
ML Systems Engineer: Distributed GPU Training & RL Pipelines

Nebius B.V. • Palo Alto (CA)

On-site
USD 180,000 - 260,000
Senior AI Cloud Infrastructure Engineer
Senior AI Cloud Infrastructure Engineer

Nebius • United States

Remote
USD 150,000 - 190,000
Competitive compensation
Career growth and learning
Flexible ownership and culture
+3
Senior HPC Systems Engineer: GPU Clusters & AI Infra
Senior HPC Systems Engineer: GPU Clusters & AI Infra

Nebius • United States

Remote
USD 180,000 - 240,000
Competitive pay
Career growth
Flexibility and ownership
+3
Senior ML Solutions Architect – Remote Token Platform (LLM)
Senior ML Solutions Architect – Remote Token Platform (LLM)

Nebius • United States

On-site
USD 210,000 - 260,000
Health Insurance
401(k) Plan
Parental Leave
+2
Senior Systems Software Engineer, GPU Compute
Senior Systems Software Engineer, GPU Compute

Nebius • United States

On-site
USD 170,000 - 300,000
Competitive compensation
Career growth
Flexible and ownership culture
+1