AI and ML Infra Software Engineer, GPU Clusters

Jobtailor

California (MO)

On-site

USD 120,000 - 190,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Jobtailor is seeking a recent graduate with an MS/PhD in Computer Science to join our AI/ML infrastructure team in the US. You will support HPC workloads, optimize performance, and help scale distributed training using PyTorch, NeMo, or JAX.

You will work with researchers, data engineers, and DevOps to implement accelerated computing, storage, and networking solutions across AWS, GCP, and Azure. Strong scripting and cloud familiarity are essential.

Qualifications

  • Recent graduate with an MS, PhD or equivalent experience in Computer Science or related field.
  • Proven experience in AI/ML and HPC workloads and infrastructure.
  • Hands-on experience in using or operating High Performance Computing (HPC) grade infrastructure.
  • In-depth knowledge of accelerated computing (e.g., GPU, custom silicon).
  • Storage (e.g., Lustre, GPFS, BeeGFS).
  • Scheduling & orchestration (e.g., Slurm, Kubernetes, LSF).
  • High-speed networking (e.g., Infiniband, RoCE, Amazon EFA).
  • Containers technologies (Docker, Enroot).
  • Expertise in running and optimizing large-scale distributed training workloads using PyTorch (DDP, FSDP), NeMo, or JAX.
  • Deep understanding of AI/ML workflows, encompassing data processing, model training, and inference pipelines.
  • Proficiency in programming & scripting languages such as Python, Go, Bash.
  • Familiarity with cloud computing platforms (e.g., AWS, GCP, Azure).
  • Experience with parallel computing frameworks and paradigms.
  • Passion for continual learning and keeping abreast of new technologies and effective approaches in the AI/ML infrastructure field.
  • Excellent communication and collaboration skills

Responsibilities

  • Collaborate closely with our AI and ML research teams to understand infrastructure needs and obstacles.
  • Monitor and optimize the performance of our infrastructure ensuring high availability, scalability, and efficient resource utilization.
  • Help define and improve important measures of AI researcher efficiency, ensuring measurable results.
  • Collaborate with researchers, data engineers, and DevOps professionals to build a seamless AI/ML infrastructure ecosystem
  • Stay on top of the latest advancements in AI/ML technologies, frameworks, and effective strategies, and promote their implementation within the company.

Skills

AI/ML Infrastructure
HPC
Distributed Training
Python
Go
Bash
Docker
Enroot
Kubernetes
Slurm
LSF
Lustre
GPFS
BeeGFS
Infiniband
RoCE
AWS
GCP
Azure
NeMo
PyTorch
JAX
Data processing
Model training
Inference pipelines

Education

MS or PhD in Computer Science or related field

Tools

Lustre
GPFS
BeeGFS
Slurm
Kubernetes
LSF
Docker
Enroot

Job description

  • Collaborate closely with our AI and ML research teams to understand their infrastructure needs and obstacles
  • Monitor and optimize the performance of our infrastructure ensuring high availability, scalability, and efficient resource utilization
  • Help define and improve important measures of AI researcher efficiency, ensuring that our actions are in line with measurable results
  • Collaborate with diverse teams, including researchers, data engineers, and DevOps professionals, to build a seamless and coordinated AI/ML infrastructure ecosystem
  • Stay on top of the latest advancements in AI/ML technologies, frameworks, and effective strategies, and promote their implementation within the company
Requirements
  • Recent graduate with a MS, PhD or equivalent experience in Computer Science or related field
  • Proven experience in AI/ML and HPC workloads and infrastructure
  • Hands-on experience in using or operating High Performance Computing (HPC) grade infrastructure
  • In-depth knowledge of accelerated computing (e.g., GPU, custom silicon)
  • Storage (e.g., Lustre, GPFS, BeeGFS)
  • Scheduling & orchestration (e.g., Slurm, Kubernetes, LSF)
  • High-speed networking (e.g., Infiniband, RoCE, Amazon EFA)
  • Containers technologies (Docker, Enroot)
  • Expertise in running and optimizing large-scale distributed training workloads using PyTorch (DDP, FSDP), NeMo, or JAX
  • Deep understanding of AI/ML workflows, encompassing data processing, model training, and inference pipelines
  • Proficiency in programming & scripting languages such as Python, Go, Bash
  • Familiarity with cloud computing platforms (e.g., AWS, GCP, Azure)
  • Experience with parallel computing frameworks and paradigms.
  • Passion for continual learning and keeping abreast of new technologies and effective approaches in the AI/ML infrastructure field.
  • Excellent communication and collaboration skills
Core Competencies

Demonstrates expertise in AI/ML infrastructure, including High Performance Computing (HPC) and accelerated computing technologies. Proficient in optimizing large-scale distributed training workloads and collaborating with diverse teams to enhance AI researcher efficiency.

Highest-signal resume keywords
  • AI/ML Infrastructure
  • High Performance Computing (HPC)
  • Distributed Training Workloads Optimization
  • Programming Languages (Python, Go, Bash)
  • Cloud Computing Platforms (AWS, GCP, Azure)
ATS Optimization Keywords
Hard Skills
  • AI/ML Workloads
  • Accelerated Computing
  • Storage Technologies (Lustre, GPFS, BeeGFS)
  • Scheduling & Orchestration (Slurm, Kubernetes, LSF)
  • High-Speed Networking (Infiniband, RoCE, Amazon EFA)
  • Container Technologies (Docker, Enroot)
  • Parallel Computing Frameworks
  • Data Processing
  • Model Training
  • Inference Pipelines
Soft Skills
  • Excellent Communication
  • Collaboration Skills
  • Passion for Learning
Certifications & Qualifications
  • MS or PhD in Computer Science or Related Field
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Member of Technical Staff, ML Engineer
Member of Technical Staff, ML Engineer

Jobtailor • Boston (MA)

On-site
USD 120,000 - 160,000
Staff Engineer, Senior Manager
Staff Engineer, Senior Manager

Jobtailor • Connecticut

On-site
USD 140,000 - 190,000
Senior AI Systems and Algorithms Engineer
Senior AI Systems and Algorithms Engineer

Jobtailor • California (MO)

On-site
USD 180,000 - 260,000
Software Engineer – AI & Cloud Engineering
Software Engineer – AI & Cloud Engineering

Jobtailor • Massachusetts

On-site
USD 110,000 - 170,000
Principal AI/ML Engineer
Principal AI/ML Engineer

Jobtailor • United States

On-site
USD 180,000 - 240,000
Senior Software Engineer – Local AI
Senior Software Engineer – Local AI

Jobtailor • California (MO)

On-site
USD 140,000 - 210,000
Member of Technical Staff – AI Cloud Infrastructure
Member of Technical Staff – AI Cloud Infrastructure

Jobtailor • California (MO)

On-site
USD 150,000 - 190,000
Director of Machine Learning
Director of Machine Learning

Jobtailor • Washington

On-site
USD 180,000 - 260,000
Principal Cloud Engineer – AI
Principal Cloud Engineer – AI

Jobtailor • West Chester

On-site
USD 150,000 - 210,000
System Software Engineer – AI
System Software Engineer – AI

Jobtailor • California (MO)

On-site
USD 120,000 - 170,000