Member of Technical Staff - Research Infrastructure Engineer

Black Forest Labs

Freiburg im Breisgau

Vor Ort

EUR 100.000 - 230.000

Vollzeit

14 Tage+
Bewerbungsgenerator

Eine komplette Bewerbung in einer Minute — maßgeschneiderter Lebenslauf und Anschreiben, versandbereit.

Schaffe es an den ATS-Filtern vorbei

Zusammenfassung

Black Forest Labs is looking for engineers in Freiburg to build and maintain the infrastructure supporting visual intelligence. Responsibilities include optimizing performance of large-scale infrastructure and collaborating with research teams.

The ideal candidate should have experience with distributed systems and Kubernetes, alongside strong problem-solving skills. This role offers a hybrid work model with reasonable travel costs covered for in-office presence.

Qualifikationen

  • Experience building or operating large-scale training platforms.
  • Strong problem-solving skills and ability to work independently.
  • Deep knowledge of modern cloud infrastructure including Kubernetes, Infrastructure as Code, AWS, and GCP.

Aufgaben

  • Maintain research infrastructure and optimize components for peak performance.
  • Scale infrastructure to meet growing research demands while ensuring reliability.
  • Collaborate with research teams to understand infrastructure needs and design solutions.

Kenntnisse

Distributed systems
Infrastructure reliability
Kubernetes
Performance optimization

Tools

Python
Bash
Go
SLURM

Jobbeschreibung

Why This Role

We're looking for engineers to build and maintain the engine that powers our mission to develop visual intelligence. From maintaining and scaling clusters, to building research platforms to accelerate the rate of innovation, this team operates with large breadth and depth. We build the systems to make multi-week/month long training possible, to orchestrate resources at scale, and at the same time efficiently, enabling the next breakthrough model. If you’re obsessed with distributed systems at scale, infrastructure reliability, scalability, security, and continuous improvement, this team would be perfect for you.

What You’ll Work On
  • Maintain research infrastructure, ensuring health, and optimizing components to extract peak performance from the system (both on application, and infrastructure side)
  • Scale infrastructure to meet growing research demands while maintaining reliability and performance
  • Collaborate with research teams to deeply understand their infrastructure needs, and design solutions that balance performance with cost efficiency.
  • Identify and resolve performance bottlenecks and capacity hotspots through deep analysis of distributed systems at scale.
  • Build and evolve telemetry and monitoring systems to provide deep visibility into infrastructure performance, utilization, and costs across our cloud and datacenter fleets.
  • Participate in on-call rotations and incident response to maintain system reliability
Technical Focus
  • Python, Bash, Go
  • Kubernetes
  • Nvidia GPU drivers, and operators
  • OTel, Prometheus
What We’re Looking For
  • Experience building or operating large-scale training platforms
  • Worked with large scale compute clusters (GPUs)
  • Proven ability to debug performance and reliability issues across large distributed fleets
  • Strong problem-solving skills and ability to work independently
  • Strong communication skills and the ability to work effectively with both internal and external partners
  • Deep knowledge of modern cloud infrastructure including Kubernetes, Infrastructure as Code, AWS, and GCP
  • Experience with SLURM
  • Experience building or operating large-scale training platforms
How We Work Together

We’re a distributed team with real offices that people actually use. Depending on your role, you’ll either join us in Freiburg or SF at least 2 days a week (or one full week every other week), or work remotely with a monthly in-person week to stay connected. We’ll cover reasonable travel costs to make this possible. We think in-person time matters, and we’ve structured things to make it accessible to all. We’ll discuss what this will look like for the role during our interview process.

Base Annual Salary

EU €100,000 - €230,000 + Equity

US $150,000 - $300,000 + Equity

This role is based in our Freiburg / San Francisco office. We operate a hybrid model and cover reasonable travel costs — relocation is encouraged but not required. We do expect a meaningful in-person presence, and we'll discuss what that looks like for your situation during the process.

Hol dir deinen kostenlosen, vertraulichen Lebenslauf-Check.
oder ziehe deine Datei hierhin.
Similar jobs

Ähnliche Jobs, die dir auch gefallen könnten

Member of Technical Staff - Research Infrastructure Engineer
Member of Technical Staff - Research Infrastructure Engineer

Black Forest Labs Inc. • Deutschland

Hybrid
EUR 100.000 - 230.000
Hybrid work model
Equity
Travel costs covered
Infrastructure Operations Engineer
Infrastructure Operations Engineer

lightningai • Deutschland

Hybrid
EUR 138.000 - 172.000
Discretionary bonus
Equity
401(k) matching
+1
Senior Site Reliability Engineer — Token Factory (Inference Platform)
Senior Site Reliability Engineer — Token Factory (Inference Platform)

Jobgether • Deutschland

Vor Ort
EUR 120.000 - 160.000
Competitive compensation
Learning opportunities
Ownership of projects
+4
(Senior) Distributed Systems Engineer - m/f/d
(Senior) Distributed Systems Engineer - m/f/d

DUDE CHEM • Berlin

Vor Ort
EUR 90.000 - 140.000
Forward Deployed Engineer
Forward Deployed Engineer

turbalance • Heidelberg

Hybrid
EUR 60.000 - 80.000
Competitive compensation
Performance-based incentives
Subsidized Deutschlandticket
+2
Member of Technical Staff - Research Engineer
Member of Technical Staff - Research Engineer

Black Forest Labs • Freiburg im Breisgau

Hybrid
EUR 130.000 - 240.000
(Senior) Distributed Systems Engineer - m/f/d
(Senior) Distributed Systems Engineer - m/f/d

Langdock • Berlin

Vor Ort
EUR 90.000 - 140.000
Equity
Senior Software Engineer, GPU Cluster Infrastructure
Senior Software Engineer, GPU Cluster Infrastructure

Socket.dev • Deutschland

Hybrid
EUR 129.000 - 198.000
Health Insurance
Retirement - 401(k)
PTO - 25 days per year
+3
Infrastructure Engineer (GPU & Compute)
Infrastructure Engineer (GPU & Compute)

lightningai • Deutschland

Hybrid
EUR 155.000 - 189.000
Discretionary bonus
Equity
Comprehensive medical coverage
+3
Senior HPC Cluster Administrator - Deep Learning Frameworks Infrastructure
Senior HPC Cluster Administrator - Deep Learning Frameworks Infrastructure

NVIDIA • Deutschland

Vor Ort
USD 58.306 - 101.063