Senior HPC Cloud Engineer — GPU/InfiniBand AI Infra

Jobgether

Germany (OH)

On-site

USD 81,272 - 104,493

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Career development opportunities
Flexible working arrangements
Collaborative engineering environment
Opportunities to work on AI infrastructure

Job summary

Jobgether is seeking a Senior HPC Cluster Engineer based in Germany to work on optimizing high-performance computing systems. This role involves deep system-level engineering and performance tuning to improve AI cloud infrastructure. You will collaborate with experts to enhance compute platforms used for data-intensive applications.

The ideal candidate has 5+ years of experience in system-level software engineering, proficiency in programming languages like C and Python, and familiarity with advanced AI infrastructures. The position offers a competitive compensation package and flexible working arrangements.

Qualifications

  • 5+ years of experience in system-level software engineering focused on performance.
  • 3+ years of hands-on experience with Linux systems administration.
  • Strong understanding of server and hardware architecture.
  • Experience with GPU clusters and InfiniBand networking is highly desirable.

Responsibilities

  • Tune and optimize GPU cluster performance and InfiniBand fabric.
  • Diagnose, troubleshoot, and resolve complex system-level issues.
  • Integrate and validate new hardware components into the HPC infrastructure.
  • Develop automation for monitoring and proactive remediation.

Skills

System-level software engineering
Performance tuning
Linux systems administration
C, C++, Go, or Python
Analytical and problem-solving skills

Tools

KVM/QEMU
Kubernetes
MPI or NCCL

Job description

Jobgether is seeking a Senior HPC Cluster Engineer based in Germany to work on optimizing high-performance computing systems. This role involves deep system-level engineering and performance tuning to improve AI cloud infrastructure. You will collaborate with experts to enhance compute platforms used for data-intensive applications.

The ideal candidate has 5+ years of experience in system-level software engineering, proficiency in programming languages like C and Python, and familiarity with advanced AI infrastructures. The position offers a competitive compensation package and flexible working arrangements.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior HPC Systems Engineer: GPU Clusters & AI Infra
Senior HPC Systems Engineer: GPU Clusters & AI Infra

Nebius • United States

Remote
USD 180,000 - 240,000
Competitive pay
Career growth
Flexibility and ownership
+3
Senior HPC Engineer - GPU Compute & InfiniBand
Senior HPC Engineer - GPU Compute & InfiniBand

Nebius • United States

Remote
USD 150,000 - 230,000
Staff Compute Infra Engineer - GPU & AI Systems
Staff Compute Infra Engineer - GPU & AI Systems

xAI • Palo Alto (CA)

On-site
USD 180,000 - 440,000
Senior HPC AI Cluster Architect — Equity Eligible
Senior HPC AI Cluster Architect — Equity Eligible

NVIDIA Corporation • Santa Clara (CA), Northern (KY)

Hybrid
USD 176,000 - 334,000
Head of AI Data Center Infrastructure Platforms and Software
Head of AI Data Center Infrastructure Platforms and Software

Summit Group Solutions, LLC • United States

On-site
USD 150,000 - 350,000
Global HPC Network Engineer for AI Infra
Global HPC Network Engineer for AI Infra

Together • San Francisco (CA)

On-site
USD 190,000 - 280,000
Startup equity
Health insurance
Competitive benefits
Senior Network Engineer — AI Infra & HPC Fabric Expert
Senior Network Engineer — AI Infra & HPC Fabric Expert

Nscale • Houston (TX)

On-site
USD 150,000 - 210,000
Competitive benefits package
Flexible paid time off
Parental leave
+1
Senior HPC AI Cluster Engineer
Senior HPC AI Cluster Engineer

NVIDIA • California (MO)

On-site
USD 176,000 - 334,000
Senior HPC Systems Architect: Liquid-Cooled GPU AI Infra
Senior HPC Systems Architect: Liquid-Cooled GPU AI Infra

Lambda • San Jose (CA)

Hybrid
USD 180,000 - 280,000
401k Plan
Health, dental, and vision coverage
Wellness stipend
+2
Senior AI Infrastructure Engineer — Scale GPU Clusters Remote
Senior AI Infrastructure Engineer — Scale GPU Clusters Remote

AI Breaking Wire • San Francisco (CA), Northern (KY)

Hybrid
USD 280,000 - 420,000
Equity options
Health, vision, dental benefits
Unlimited PTO
+2