ML Infrastructure Engineer

Nebius B.V.

Deutschland

Remote

EUR 90.000 - 120.000

Vollzeit

Vor 3 Tagen
Sei unter den ersten Bewerbenden

Erhalte mehr Antworten von Arbeitgebern

Versende in nur wenigen Minuten einen passgenauen Lebenslauf.

Benefits dieser Stelle

Competitive compensation
Career growth and learning
Flexibility and ownership
Collaborative and innovative culture
Impactful AI projects
International environment and talented

Zusammenfassung

Nebius B.V. is seeking a highly skilled ML/AI Engineer to lead benchmarking of GPU platforms for ML/AI workloads. You will profile GPU performance at system and kernel levels, compare across platforms, and optimize workloads to remove bottlenecks.

You will build tools and dashboards to visualise performance trends and contribute to internal ML benchmarking frameworks. Strong ML foundations and experience with PyTorch/JAX and modern LLM stacks are essential.

Qualifikationen

  • Proven ML/AI knowledge and foundations.
  • Experience optimizing large neural networks for training and inference.
  • Strong familiarity with PyTorch, JAX, Megatron-LM, Tensort-LLM.
  • Solid grasp of CUDA, NCCL, drivers and libraries.
  • Experience in containerized environments (Docker, Kubernetes).
  • Excellent communication and independent work capability.

Aufgaben

  • Profile and analyze GPU performance at system and kernel level.
  • Evaluate and compare GPU performance across platforms, architectures and stacks (CUDA, ROCm).
  • Debug and optimise ML workloads to reduce bottlenecks on GPU hardware.
  • Perform acceptance testing for new GPU clusters ensuring performance and compatibility.
  • Run experiments across GPU configurations to study interconnects and system optimisations.
  • Develop tools and dashboards to visualise performance metrics and trends.
  • Contribute to internal tooling, frameworks and best practices.

Kenntnisse

GPU benchmarking
Performance profiling
CUDA programming
LLM inference frameworks
System/kernel profiling
Python
Visualization dashboards
TensorRT
PyTorch
JAX
Megatron-LM
Nsight

Tools

Docker
Kubernetes
CUDA
NCCL
Nsight
nvprof
Perf
PyTorch
JAX
Megatron-LM
TensorRT

Jobbeschreibung

The role

We are seeking a highly skilled ML/AI Engineer to join our team to lead and support benchmarking of GPU platforms benchmarking of GPU platforms for machine learning and AI workloads. You will play a critical role in evaluating the performance of GPU-based hardware for various deep learning and AI frameworks, enabling data-driven decisions for platform optimisation and next-generation hardware development.

Your responsibilities will include:
  • Work closely with hardware, development teams to profile and analyse GPU performance at the system and kernel level.
  • Evaluate and compare GPU performance across different platforms, architectures, and software stacks (e.g.,CUDA, ROCm).
  • Debug and optimise ML workloads to run efficiently on GPU hardware, identifying and resolving performance bottlenecks.
  • Perform acceptance testing acceptance testing for new GPU clusters, ensuring hardware and software meet performance, stability, and compatibility requirements for AI workloads.
  • Perform experiments across diverse GPU system configurations to assess the impact of varying interconnect strategies and system-level optimisations on performance and scalability.
  • Develop tools and dashboards to visualise performance metrics visualise performance metrics, bottlenecks, and trends.
  • Contribute to internal tooling, frameworks, and best practices
We expect you to have:
  • A profound understanding of theoretical foundations of machine learning
  • Deep understanding of performance aspects of large neural networks training and inference (data/tensor/context/expert parallelism, offloading, custom kernels, hardware features, attention optimisations, dynamic batching etc.)
  • Deep experience with modern deep learning frameworks (PyTorch, JAX, Megatron-LM, Tensort-LLM)
  • Good understanding of the GPU stack: CUDA,NCCL, drivers, and relevant libraries
  • Familiarity with containerized environments (e.g., Docker, Kubernetes).
  • Strong communication and ability to work independently
Ways to stand out from the crowd:
  • Familiarity with modern LLM inference frameworks (vLLM, SGLang, TensorRT)
  • Experience in Python and performance profiling tools (e.g., Nsight, nvprof, perf).
  • Familiarity with cloud ML platforms like AWS, GCP, Azure ML
  • Contributions to open-source ML benchmarking tools
Benefits & Perks:
  • Competitive compensation
  • Career growth and learning opportunities
  • Flexibility and ownership
  • Collaborative and innovative culture
  • Opportunity to work on impactful AI projects
  • International environment and talented teams
What's it like to work at Nebius:

Fast moving - Bold thinking - Constant growth - Meaningful impact - Trust and real ownership - Opportunity to shape the future of AI

Equal Opportunity Statement:

Nebius is an equal opportunity employer. We are committed to fostering an inclusive and diverse workplace and to providing equal employment opportunities in all aspects of employment. We do not discriminate on the basis of race, color, religion, sex (including pregnancy), national origin, ancestry, age, disability, genetic information, marital status, veteran status, sexual orientation, gender identity or expression, or any other characteristic protected by applicable law.

Applicants must be authorized to work in the country in which they apply and will be required to provide proof of employment eligibility as a condition of hire.

If you need accommodations during the application process, please let us know.

Hol dir deinen kostenlosen, vertraulichen Lebenslauf-Check.
oder ziehe deine Datei hierhin.
Similar jobs

Ähnliche Jobs, die dir auch gefallen könnten

Customer Engineer
Customer Engineer

Nebius B.V. • Deutschland

Remote
EUR 155.000 - 193.000
Health Insurance
401(k) Plan
Parental Leave
+2
Cloud Solution Architect
Cloud Solution Architect

Nebius B.V. • Deutschland

Remote
EUR 90.000 - 130.000
Competitive compensation
Career growth and learning
Flexibility and ownership
+3
Technical Project Manager (Hardware)
Technical Project Manager (Hardware)

Nebius B.V. • Deutschland

Hybrid
EUR 90.000 - 130.000
Competitive compensation
Career growth opportunities
Flexibility and ownership
+2
AI/ML Specialist Solutions Architect
AI/ML Specialist Solutions Architect

Nebius B.V. • Deutschland

Remote
EUR 90.000 - 130.000
Flexible working options
EU-wide remote work
Career growth opportunities
+1
Technical Account Manager
Technical Account Manager

Nebius B.V. • Deutschland

Remote
EUR 90.000 - 130.000
Competitive pay
Career growth
Flexibility
+3
System Engineer (Compute Node)
System Engineer (Compute Node)

Nebius B.V. • Deutschland

Remote
EUR 90.000 - 130.000
Competitive compensation
Career growth and learning opportunit​
Flexibility and ownership
+3
Sr. Solution Engineer (GPU Infrastructure)
Sr. Solution Engineer (GPU Infrastructure)

Nebius B.V. • Deutschland

Remote
EUR 155.000 - 190.000
Health insurance
401(k)
Parental leave
+3
Field CTO
Field CTO

Nebius B.V. • Deutschland

Remote
EUR 172.000 - 211.000
Health Insurance
401(k) Plan
Parental Leave
+2
Senior Software Engineer
Senior Software Engineer

Nebius B.V. • Deutschland

Hybrid
EUR 112.000 - 146.000
Health insurance
401(k) plan
Parental leave
+2
Senior Applied ML Engineer (Agentic Search)
Senior Applied ML Engineer (Agentic Search)

Nebius B.V. • Deutschland

Remote
EUR 110.000 - 160.000