Infrastructure/Systems Engineer

Remotely

Berlin

Remote

EUR 90.000 - 140.000

Vollzeit

14 Tage+
Bewerbungsgenerator

A complete application in a minute — tailored resume and cover letter, ready to send.

Schaffe es an den ATS-Filtern vorbei

Benefits dieser Stelle

Remote-first, Germany-wide
Flexible hours
Your setup, your way
30 days vacation
Keep growing
Get rewarded for impact
A warm welcome
Good people, good times
A mission that matters

Zusammenfassung

Orcrist is building a next generation data intelligence platform using cutting-edge technologies. We’re handling petabyte-scale data with sub-second queries.

Our product is a Kubernetes-based platform delivered as B2B SaaS or as a self-hosted on-prem solution, including air-gapped deployments. You’ll own the layer everything else runs on: bare-metal servers, operating systems, data-center networking, and storage across on-prem and fully air-gapped sites — the physical and infrastructure

Qualifikationen

  • 5+ years in bare-metal, data-center, or systems infrastructure engineering.
  • Strong bare-metal Linux (Ubuntu, RKE2, Talos, or similar) provisioning, firmware/BMC management, PXE/iPXE, kernel and storage tuning.
  • Infrastructure automation experience with Ansible, Terraform, or equivalent; Git and CI/CD; scripting in Python, Bash, or Go.
  • Solid data-center networking fundamentals: L2/L3, IP Fabric, BGP, and switching; RDMA (RoCE/InfiniBand) a strong plus.

Aufgaben

  • Design, size, provision, and operate bare-metal server fleets across on-prem and air-gapped environments using zero-touch provisioning and automation.
  • Build and run modern data-center networking: L2/L3 design, IP Fabric, BGP and switching.
  • Engineer resilient, highly available storage (Ceph/Rook, NVMe) with capacity planning and encryption at rest.
  • Operate offline mirrors and registries, signed artifacts, firmware/driver lifecycle without internet, and system hardening.
  • Support MLOps and inference-serving fundamentals — GPU model serving, GPU scheduling and sharing, throughput/latency optimization.
  • Plan and run on-site build-outs: rack integration, power budgets, thermal/cooling, UPS sizing, capacity planning, runbooks.

Jobbeschreibung

Orcrist is building a next generation data intelligence platform using cutting‑edge technologies. We're handling petabyte‑scale data with sub‑second queries. Our product is a Kubernetes‑based platform delivered as B2B SaaS or as a self‑hosted on‑prem solution, including air‑gapped deployments. We enable customers across defense, law enforcement, and enterprise to turn mission‑critical data into actionable intelligence.

Role

You'll own the layer everything else runs on: bare‑metal servers, operating systems, data‑center networking, and storage across on‑prem and fully air‑gapped sites — the physical and infrastructure foundation our platform and GPU fleets are built on. You design, build, and operate server fleets with a strong automation and DevOps mindset, then partner with our SRE, MLOps, and ML teams to ensure everything running above the metal — including GPU inference — performs reliably at scale. Some of this work is hands‑on at customer sites, where you size, rack, and commission self‑contained server environments with no internet uplink.

We weight depth in modern data‑center infrastructure, networking, and automation more heavily than GPU‑specific experience. A strong infrastructure and network engineer with a genuine automation mindset — even without prior GPU exposure — is a better fit for this role than a candidate with GPU expertise whose networking background is rooted in legacy, corporate‑style L2 designs.

What you’ll do
  • Design, size, provision, and operate bare‑metal server fleets across on‑prem and air‑gapped environments (firmware/BIOS/UEFI, BMC via Redfish/IPMI, OS, RAID, kernel and storage tuning) using zero‑touch provisioning (PXE/iPXE, MAAS/Metal3/Tinkerbell/Ironic) and automation (Ansible, Terraform, or equivalent — the tooling matters less than the automation mindset).
  • Build and run modern data‑center networking: L2/L3 design, IP Fabric, BGP and switching, and RDMA fabrics (RoCE/InfiniBand) sized to scale without ripping out the core.
  • Engineer resilient, highly available storage (Ceph/Rook, NVMe) with capacity planning and encryption at rest.
  • Operate confidently in air‑gapped and on‑prem environments: offline mirrors and registries, signed artifacts, firmware/driver lifecycle without internet access, and system hardening.
  • Support MLOps and inference‑serving fundamentals — GPU model serving (Triton/KServe/vLLM), GPU scheduling and sharing, and throughput/latency optimization — in partnership with our SRE and ML teams.
  • Plan and run on‑site build‑outs: rack integration, power budgets, thermal/cooling and UPS sizing, commissioning, capacity planning, runbooks, and operator handover, with SWaP awareness for field sites.
About You
  • 5+ years in bare‑metal, data‑center, or systems infrastructure engineering, with hands‑on ownership of physical and compute infrastructure at scale.
  • Strong bare‑metal Linux (Ubuntu, RKE2, Talos, or similar): provisioning, firmware/BMC management, PXE/iPXE, kernel and storage tuning, systemd, RAID.
  • Real experience with infrastructure automation (Ansible, Terraform, or equivalent), Git and CI/CD, and scripting in Python, Bash, or Go.
  • Solid, current data‑center networking fundamentals: L2/L3, IP Fabric, BGP, and switching — this is a hard requirement, not a nice‑to‑have. RDMA (RoCE/InfiniBand) experience is a strong plus.
  • Comfortable operating in air‑gapped or on‑prem environments and traveling to customer sites for builds and deployments.
  • Practical hardware sizing literacy: power budgets, thermal/cooling, UPS sizing, and rack integration.
  • Documentation‑focused, methodical, and calm during hardware incidents. Eligible to work in Germany.
Nice‑to‑haves
  • German language (B1+); exposure to regulated or security‑critical environments (e.g. BSI C5, ISO 27001, or defense‑sector delivery) is a plus.
  • NVIDIA GPU stack knowledge (drivers, CUDA, GPU Operator, MIG, DCGM) and cross‑node GPU interconnect experience (NVLink, InfiniBand, NCCL).
  • Kubernetes bare‑metal fundamentals — how cluster bring‑up and GPU device plugins interact with the underlying hardware, not day‑to‑day cluster operation.
  • Inference optimization (vLLM, TensorRT‑LLM, quantization) and familiarity with switch NOS (SONiC/Cumulus).
  • Relevant certifications (NVIDIA, Red Hat, CKA/CKS) or field/forward‑deployed engineering experience.
What we offer
  • Remote‑first, Germany‑wide:Work from wherever you do your best work, with regular team gatherings in Berlin and other off‑site locations.
  • Flexibility by default: Flexible working hours help you make work fit your life.
  • Your setup, your way: Get a personal home‑office equipment budget to create a workspace that works for you.
  • 30 days of vacation:Take the time you need to recharge and come back with fresh energy.
  • Keep growing: We invest in your personal and professional development.
  • Get rewarded for impact: Performance bonuses are tied to agreed objectives and key results.
  • A warm welcome: Every new team member gets a welcome goodie bag.
  • Good people, good times: From summer and Christmas parties to regular team gatherings, we make time to celebrate together.
  • A mission that matters: Work on challenges with tangible impact on public safety and national security.
Hol dir deinen kostenlosen, vertraulichen Lebenslauf-Check.

oder ziehe deine Datei hierhin.

Similar jobs

Ähnliche Jobs, die dir auch gefallen könnten

System Software Engineer (gn) Cloud & Simulation
System Software Engineer (gn) Cloud & Simulation

Personio • Dresden

Vor Ort
EUR 90.000 - 130.000
Network Engineer (m/f/d)
Network Engineer (m/f/d)

www.polarise.eu • Düsseldorf

Vor Ort
EUR 90.000 - 130.000
Remote work model
Structured onboarding
Hardware provided (MacBook Pro/Think?)
Infrastructure Support Engineer
Infrastructure Support Engineer

Lyceum • Berlin

Vor Ort
EUR 60.000 - 85.000
Senior Forward Deployed Platform Engineer (m/w/d)
Senior Forward Deployed Platform Engineer (m/w/d)

Recare • Deutschland

Hybrid
EUR 90.000 - 120.000
Edenred card
Extra vacation day
Remote-friendly with flexible hours
Senior Forward Deployed Platform Engineer (m/w/d)
Senior Forward Deployed Platform Engineer (m/w/d)

Recare Deutschland GmbH • Berlin

Remote
EUR 95.000 - 130.000
Edenred card
Extra vacation day
Remote-friendly & flexible hours
GPU Cluster Engineer - Scalable AI Infrastructure
GPU Cluster Engineer - Scalable AI Infrastructure

NEURA Robotics • Deutschland

Remote
USD 140.000 - 195.000
Senior Infrastructure Support Engineer
Senior Infrastructure Support Engineer

Nscale • Deutschland

Vor Ort
EUR 103.000 - 147.000
Base salary + equity
Remote-first team
Annual reviews and progression plan
+1
Head of Compute Engineering
Head of Compute Engineering

Impossible Cloud GmbH • Hamburg

Vor Ort
EUR 80.000 - 100.000
Competitive salary
ESOP
Subsidized gym membership
+1
Senior System Engineer (Munich, Germany)
Senior System Engineer (Munich, Germany)

Remotestar • München

Vor Ort
EUR 80.000 - 110.000
Indefinite contract
Equal pay guaranteed
Variable performance bonus
+8
GPU Cluster Engineer (human)
GPU Cluster Engineer (human)

NEURA Robotics • Deutschland

Vor Ort
USD 140.000 - 195.000