Site Reliability Engineer II - AI & Infrastructure (f/m/d)

Meyandy LLC

Berlin

Vor Ort

EUR 85.000 - 120.000

Vollzeit

Vor 11 Tagen
Bewerbungsgenerator

Eine zielgenaue Bewerbung für diesen Job — ein maßgeschneiderter Lebenslauf und ein Anschreiben, die genau zur Stellenanzeige passen.

Schaffe es an den ATS-Filtern vorbei

Zusammenfassung

Focused Energy in Berlin is seeking a Site Reliability Engineer II to build reliable, secure, observable systems for its internal applications. You’ll own day-to-day reliability, deployments, incident response, monitoring, backup, recovery, and rollback across a broad technical environment spanning cloud services, databases, identity, networking, secrets, and CI/CD.

This role also supports AI workflows and Langdock when needed.

Qualifikationen

  • Strong experience with Linux, networking, structured troubleshooting, and cloud hosting concepts.
  • Experience operating internal applications and deployment platforms in Azure or a comparable cloud environment.
  • Practical knowledge of relational databases, particularly PostgreSQL or an equivalent platform, including safe migration and recovery practices.
  • Experience with infrastructure as code and CI/CD tools such as OpenTofu, Terraform, GitLab CI, GitHub Actions, or Azure DevOps.
  • Knowledge of monitoring, alerting, backup, recovery, rollback automation, incident response, and service reliability practices.
  • Experience with horizontal scaling, stateless system design, migrations, and reversible change management.

Aufgaben

  • Own the reliability and day-to-day operation of multiple internal applications, deployment platforms, and supporting services.
  • Lead the investigation and resolution of incidents and deployment failures across cloud services, databases, identity, networking, secrets, and CI/CD systems.
  • Drive automation improvements across build, release, deployment, monitoring, alerting, backup, recovery, rollback, and runbook processes.
  • Implement safe, well-tested code, configuration, infrastructure, and database fixes to restore or improve service reliability.
  • Develop proposals for cloud, hosting, database, and runtime migrations, as well as horizontal scaling approaches, supported by appropriate testing and rollback plans.
  • Build and maintain clear operational documentation, including runbooks, recovery procedures, incident fixes, and system knowledge that can be reused by other engineers.
  • Partner with software engineering, IT, security, identity, and external platform-support stakeholders to plan and execute infrastructure changes safely.
  • Provide overflow diagnostic and troubleshooting support for AI workflows and tools, including Langdock, without losing focus on core reliability priorities.

Kenntnisse

Linux
Networking
Troubleshooting
Cloud hosting concepts
Azure
PostgreSQL
Monitoring
Alerting
Backup
Recovery
Rollback
Incident response
SRE practices
Horizontal scaling
Stateless design
Migrations
Reversible changes

Tools

OpenTofu
Terraform
GitLab CI
GitHub Actions
Azure DevOps

Jobbeschreibung

Focused Energy is a pioneering international deep-tech company with locations in Germany and the US, dedicated to commercializing laser-driven nuclear fusion. Our mission is to deliver clean, safe, and virtually limitless energy to the world. As we rapidly scale, we seek talented individuals who thrive on bringing clarity and structure to fast-growing environments.

About the Role

Focused Energy is looking for a Site Reliability Engineer II to help build reliable, secure, observable, and easy-to-operate systems for its internal applications. You’ll own day-to-day reliability, deployments, incident response, monitoring, backup, recovery, and rollback processes across a broad technical environment spanning cloud services, databases, identity, networking, secrets, and CI/CD. This role is ideal for an engineer who enjoys solving complex operational problems and turning recurring incidents into lasting improvements.

Alongside your core infrastructure responsibilities, you’ll provide overflow support for AI-enablement workflows and tools when the dedicated AI Tech Enabler needs additional support.

What You’ll Do
  • Own the reliability and day-to-day operation of multiple internal applications, deployment platforms, and supporting services.
  • Lead the investigation and resolution of incidents and deployment failures across cloud services, databases, identity, networking, secrets, and CI/CD systems.
  • Drive automation improvements across build, release, deployment, monitoring, alerting, backup, recovery, rollback, and runbook processes.
  • Implement safe, well-tested code, configuration, infrastructure, and database fixes to restore or improve service reliability.
  • Develop proposals for cloud, hosting, database, and runtime migrations, as well as horizontal scaling approaches, supported by appropriate testing and rollback plans.
  • Build and maintain clear operational documentation, including runbooks, recovery procedures, incident fixes, and system knowledge that can be reused by other engineers.
  • Partner with software engineering, IT, security, identity, and external platform-support stakeholders to plan and execute infrastructure changes safely.
  • Provide overflow diagnostic and troubleshooting support for AI workflows and tools, including Langdock, without losing focus on core reliability priorities.
Who You Are
  • You take ownership of well-scoped technical work from investigation through testing, deployment, and follow-up.
  • You are calm and methodical when responding to live incidents and can diagnose problems across multiple interconnected systems.
  • You focus on root-cause resolution rather than repeatedly applying temporary fixes.
  • You communicate incidents, trade-offs, risks, and recovery plans clearly to both technical and non-technical stakeholders.
  • You are proactive about identifying reliability risks, automation opportunities, recurring issues, and operational improvements.
  • You make independent decisions on tactical fixes and safe configuration changes while seeking appropriate sign-off for architecture, migration, and scaling decisions.
  • You document changes and operational knowledge so that other engineers can support systems effectively.
  • You ask for help early when risks or dependencies are unclear and keep the Head of IT informed about significant issues and progress.
Desirable Skills & Knowledge
  • Strong experience with Linux, networking, structured troubleshooting, and cloud hosting concepts.
  • Experience operating internal applications and deployment platforms in Azure or a comparable cloud environment.
  • Practical knowledge of relational databases, particularly PostgreSQL or an equivalent platform, including safe migration and recovery practices.
  • Experience with infrastructure as code and CI/CD tools such as OpenTofu, Terraform, GitLab CI, GitHub Actions, or Azure DevOps.
  • Knowledge of monitoring, alerting, backup, recovery, rollback automation, incident response, and service reliability practices.
  • Experience with horizontal scaling, stateless system design, migrations, and reversible change ...

Site Reliability Engineer II - AI & Infrastructure (f/m/d) - Focused, Berlin.

Hol dir deinen kostenlosen, vertraulichen Lebenslauf-Check.
oder ziehe deine Datei hierhin.
Similar jobs

Ähnliche Jobs, die dir auch gefallen könnten

Site Reliability Engineer II - AI & Infrastructure (f/m/d)
Site Reliability Engineer II - AI & Infrastructure (f/m/d)

Fusion Energy Base • Berlin

Vor Ort
EUR 70.000 - 110.000
Site Reliability Engineer II - AI & Infrastructure (f/m/d)
Site Reliability Engineer II - AI & Infrastructure (f/m/d)

Linuxconfig • Berlin

Vor Ort
EUR 85.000 - 120.000
Site Reliability Engineer II - AI & Infrastructure (f/m/d)
Site Reliability Engineer II - AI & Infrastructure (f/m/d)

DUDE CHEM • Berlin

Vor Ort
EUR 85.000 - 110.000
IT Systems Engineer (f/m/d)
IT Systems Engineer (f/m/d)

Meyandy LLC • Darmstadt

Vor Ort
EUR 70.000 - 100.000
IT Operations Administrator (f/m/d)
IT Operations Administrator (f/m/d)

Meyandy LLC • Berlin

Vor Ort
EUR 60.000 - 90.000
IT Systems Engineer (f/m/d)
IT Systems Engineer (f/m/d)

Focused • Darmstadt

Vor Ort
EUR 65.000 - 85.000
IT Systems Engineer (f/m/d)
IT Systems Engineer (f/m/d)

Focused Energy • Darmstadt

Vor Ort
EUR 70.000 - 95.000
IT Operations Administrator (f/m/d)
IT Operations Administrator (f/m/d)

Focused Energy • Berlin

Hybrid
EUR 60.000 - 90.000
IT Operations Administrator (f/m/d)
IT Operations Administrator (f/m/d)

Focused • Berlin

Vor Ort
EUR 55.000 - 75.000
Infrastructure Engineer - Deployments (d/f/m)
Infrastructure Engineer - Deployments (d/f/m)

Meyandy LLC • München

Vor Ort
EUR 70.000 - 110.000
EGYM Wellpass
language courses
Jobrad
+1