Erhalte mehr Antworten von Arbeitgebern
Versende in nur wenigen Minuten einen passgenauen Lebenslauf.
CloudFactory in Berlin is seeking a Senior Site Reliability Engineer to keep production systems running smoothly, focusing on reliability, security, and automation across ML/LLM workloads.
You will own observability, platform health, and the end-to-end software delivery lifecycle, partnering with engineers to deploy securely to multiple public clouds and maintain high availability.
At CloudFactory, we are a mission-driven team passionate about unlocking the potential of AI to transform the world. By combining advanced technology with a global network of talented people, we make unusable data usable, driving real-world impact at scale. More than just a workplace, we’re a global community founded on strong relationships and the belief that meaningful work transforms lives. Our commitment to earning, learning, and serving fuels everything we do as we strive to connect one million people to meaningful work and build leaders worth following.
At CloudFactory, we believe in building a workplace where everyone feels empowered, valued, and inspired to bring their authentic selves to work. We are: Mission-Driven: We focus on creating economic and social impact. People-Centric: We care deeply about our team’s growth, well-being, and sense of belonging. Innovative: We embrace change and find better ways to do things together. Globally Connected: We foster collaboration between diverse cultures and perspectives. If you’re passionate about innovation, collaboration, and making a real impact, we’d love to have you on board!
As a Site Reliability Engineer, you will play a key role in keeping all production systems running smoothly. You will work closely with other engineers and operators to fuse engineering principles, operational knowledge, security, and automation to work towards platform/service production excellence from an angle of infrastructure, reliability, and security. The SRE team owns the foundation of AI Platform’s Core platform - the services and infrastructure that let us deploy to a multitude of public cloud providers and that powers many ML and LLM powered features. We give every other engineering team a reliable base to build on, and we own the software delivery lifecycle end to end: the tooling, patterns, and automation that reduce friction for the whole org. This is an exciting opportunity to grow professionally while contributing to a mission-driven organization.
Who you are (must-haves)