Infrastructure/Systems Engineer

DUDE CHEM

Berlin

Hybrid

EUR 90.000 - 130.000

Vollzeit

14 Tage+

Erhalte mehr Antworten von Arbeitgebern

Versende in nur wenigen Minuten einen passgenauen Lebenslauf.

Benefits dieser Stelle

Remote/Hybrid Berlin
Home office budget
30 days vacation
Direct impact on missions

Zusammenfassung

Orcrist is building a next-generation data intelligence platform with a Kubernetes-based stack, delivering B2B SaaS and self-hosted options including air-gapped deployments. You will own the underlying bare-metal servers, OS, data-center networking and storage—partnering with SRE, MLOps, and ML teams to ensure GPU inference runs reliably at scale.

Some work happens at customer sites. We seek a strong infrastructure engineer with hands-on experience in Linux, automation, and networking.

Qualifikationen

  • 5+ years in bare-metal or data-center infrastructure engineering.
  • Strong Linux administration and kernel/storage tuning experience.
  • Proficient in infrastructure automation and scripting (Python/Bash/Go).
  • Solid networking fundamentals including L2/L3 and BGP.
  • Experience with air-gapped/on-prem deployments is a plus.

Aufgaben

  • Design, size, provision and operate bare-metal server fleets across on-prem and air-gapped environments.
  • Automate provisioning with PXE/iPXE and tooling (Ansible, Terraform or equivalent).
  • Build and run data-center networking and storage with resilience and encryption.
  • Collaborate with SRE, MLOps and ML teams to support GPU inference workloads.
  • Plan on-site builds including rack integration, power, cooling and UPS sizing.

Kenntnisse

Bare-metal infra
Linux admin
Networking fundamentals
Automation tooling
Scripting languages

Tools

Ansible
Terraform
Git
CI/CD
PXE/iPXE
Redfish/IPMI

Jobbeschreibung

Infrastructure Engineer – Platform
Company

Orcrist is building a next generation data intelligence platform using cutting-edge technologies. We’re handling petabyte-scale data with sub-second queries. Our product is a Kubernetes-based platform delivered as B2B SaaS or as a self-hosted on-prem solution, including air-gapped deployments. We enable customers across defense, law enforcement, and enterprise to turn mission-critical data into actionable intelligence. Our Platform team owns the infrastructure that powers every deployment, from the metal up.

Role

You'll own the layer everything else runs on: bare-metal servers, operating systems, data-center networking, and storage across on-prem and fully air-gapped sites — the physical and infrastructure foundation our platform and GPU fleets are built on. You design, build, and operate server fleets with a strong automation and DevOps mindset, then partner with our SRE, MLOps, and ML teams to ensure everything running above the metal — including GPU inference — performs reliably at scale. Some of this work is hands‑on at customer sites, where you size, rack, and commission self-contained server environments with no internet uplink.

We weight depth in modern data‑center infrastructure, networking, and automation more heavily than GPU‑specific experience. A strong infrastructure and network engineer with a genuine automation mindset — even without prior GPU exposure — is a better fit for this role than a candidate with GPU expertise whose networking background is rooted in legacy, corporate-style L2 designs.

What you’ll do
  • Design, size, provision, and operate bare-metal server fleets across on-prem and air-gapped environments (firmware/BIOS/UEFI, BMC via Redfish/IPMI, OS, RAID, kernel and storage tuning) using zero-touch provisioning (PXE/iPXE, MAAS/Metal3/Tinkerbell/Ironic) and automation (Ansible, Terraform, or equivalent — the tooling matters less than the automation mindset).
  • Build and run modern data-center networking: L2/L3 design, IP Fabric, BGP and switching, and RDMA fabrics (RoCE/InfiniBand) sized to scale without ripping out the core.
  • Engineer resilient, highly available storage (Ceph/Rook, NVMe) with capacity planning and encryption at rest.
  • Operate confidently in air-gapped and on-prem environments: offline mirrors and registries, signed artifacts, firmware/driver lifecycle without internet access, and system hardening.
  • Support MLOps and inference-serving fundamentals — GPU model serving (Triton/KServe/vLLM), GPU scheduling and sharing, and throughput/latency optimization — in partnership with our SRE and ML teams.
  • Plan and run on-site build-outs: rack integration, power budgets, thermal/cooling and UPS sizing, commissioning, capacity planning, runbooks, and operator handover, with SWaP awareness for field sites.
About You
  • 5+ years in bare-metal, data-center, or systems infrastructure engineering, with hands‑on ownership of physical and compute infrastructure at scale.
  • Strong bare-metal Linux (Ubuntu, RKE2, Talos, or similar): provisioning, firmware/BMC management, PXE/iPXE, kernel and storage tuning, systemd, RAID.
  • Real experience with infrastructure automation (Ansible, Terraform, or equivalent), Git and CI/CD, and scripting in Python, Bash, or Go.
  • Solid, current data-center networking fundamentals: L2/L3, IP Fabric, BGP, and switching — this is a hard requirement, not a nice‑to‑have. RDMA (RoCE/InfiniBand) experience is a strong plus.
  • Comfortable operating in air-gapped or on-prem environments and traveling to customer sites for builds and deployments.
  • Practical hardware sizing literacy: power budgets, thermal/cooling, UPS sizing, and rack integration.
  • Documentation-focused, methodical, and calm during hardware incidents. Eligible to work in Germany.
Nice‑to‑haves
  • German language (B1+); exposure to regulated or security‑critical environments (e.g. BSI C5, ISO 27001, or defense-sector delivery) is a plus.
  • NVIDIA GPU stack knowledge (drivers, CUDA, GPU Operator, MIG, DCGM) and cross‑node GPU interconnect experience (NVLink, InfiniBand, NCCL).
  • Kubernetes bare-metal fundamentals — how cluster bring‑up and GPU device plugins interact with the underlying hardware, not day‑to‑day cluster operation.
  • Inference optimization (vLLM, TensorRT-LLM, quantization) and familiarity with switch NOS (SONiC/Cumulus).
  • Relevant certifications (NVIDIA, Red Hat, CKA/CKS) or field/forward‑deployed engineering experience.
What We Offer
  • Modern architecture & stack.
  • Remote/Hybrid setup in Berlin with occasional team events in Berlin.
  • Home office budget and great equipment.
  • 30 days vacation.
  • Direct impact on critical missions across private and public‑sector customers.
Hol dir deinen kostenlosen, vertraulichen Lebenslauf-Check.
oder ziehe deine Datei hierhin.
Similar jobs

Ähnliche Jobs, die dir auch gefallen könnten

Senior Technical Operations & Deployment Engineer (GPU Cloud Infrastructure)
Senior Technical Operations & Deployment Engineer (GPU Cloud Infrastructure)

Jobgether • Deutschland

Vor Ort
EUR 90.000 - 140.000
Head of Compute Engineering
Head of Compute Engineering

Impossible Cloud GmbH • Hamburg

Vor Ort
EUR 80.000 - 100.000
Competitive salary
ESOP
Subsidized gym membership
+1
Member of Technical Staff - Research Infrastructure Engineer
Member of Technical Staff - Research Infrastructure Engineer

Black Forest Labs • Freiburg im Breisgau

Vor Ort
EUR 100.000 - 230.000
Forward Deployed Engineer
Forward Deployed Engineer

turbalance • Heidelberg

Hybrid
EUR 60.000 - 80.000
Competitive compensation
Performance-based incentives
Subsidized Deutschlandticket
+2
Senior System Engineer (Munich, Germany)
Senior System Engineer (Munich, Germany)

Remotestar • München

Hybrid
EUR 80.000 - 110.000
Indefinite contract
Equal pay guaranteed
Variable performance bonus
+8
Senior Data Center Infrastructure Engineer
Senior Data Center Infrastructure Engineer

NVIDIA • Deutschland

Remote
EUR 145.000 - 230.000
Equity
Benefits
Principal Software Engineer - Compute Infrastructure
Principal Software Engineer - Compute Infrastructure

NVIDIA • Deutschland

Hybrid
EUR 214.000 - 337.000
Equity
Hybrid work model
Senior Infrastructure & Security Engineer
Senior Infrastructure & Security Engineer

European Tech Recruit • Frankfurt

Vor Ort
EUR 110.000 - 140.000
Senior Observability & Telemetry Engineer - Radian Arc
Senior Observability & Telemetry Engineer - Radian Arc

Jobgether • Deutschland

Hybrid
EUR 110.000 - 170.000
Remote work model (EMEA)
Hybrid-friendly environment
International collaboration
+2
Senior Solutions Architect, HPC and AI
Senior Solutions Architect, HPC and AI

NVIDIA • Berlin

Vor Ort
EUR 110.000 - 170.000