YOUR TASKS
- Embed as Site Reliability Engineer within a product team, taking ownership of reliability and driving resiliency and scalability of their services
- Design, implement, and maintain robust infrastructure on Google Cloud Platform (GCP), ensuring services achieve reliability targets defined by business-aligned SLOs
- Act as a subject matter expert, championing "You Build It, You Run It" principles, modern Reliability Engineering, and DevOps best practices through hands‑on implementation and roadmap consulting
- Participate in Incident Management, on‑call rotation, and drive learning from failure as well as safe preparation of changes
- Enhance and maintain CI/CD pipelines (e.g., deployment, canaries, rollbacks) and observability tooling (monitoring, logging, metrics, and tracing), with a senior‑level focus on creating shared, reusable components and establishing appropriate guardrails
- Take an active part in evolving our system architecture and non‑functional requirements to foster a modern, resilient, and observable application ecosystem
YOUR PROFILE
- Proven experience in a Site Reliability, DevOps, or equivalent role, with demonstrated expertise in designing and operating services on a major public cloud platform (GCP preferred). Relevant certifications are a plus
- A growth mindset and a passion for continuous improvement, with an interest in industry trends like Platform Engineering and the application of AI in operations
- Know‑how for translating business requirements into technical reliability targets (SLOs/SLIs), and a solid cross‑functional collaboration with Product Owners, Architects, Software and QA Engineers
Proficiency in infrastructure
- Serverless Computing: Experience with serverless platforms like Cloud Run or AWS Fargate
- Containerization & Orchestration: Deploying and managing applications with Docker and
- Infrastructure as Code (IaC): Strong command of Terraform and CI/CD tools like GitHub Actions for automating infrastructure provisioning and application deployments
- A working understanding of at least some of networking (VPCs, firewalls, DNS), load balancing, and databases
- Experience in Software engineering, with proficiency in at least one of: Python, Java, or Groovy, Go (with and without AI‑powered coding assistants)
- Experience in working with observability tools like Datadog, Prometheus, Grafana
- Sound communication and collaboration skills, with a commitment to mentoring others and thriving in an international team
- Willingness to participate in a sustainable on‑call rotation to ensure service availability
- Professional fluency in English. German is optional but welcome
OUR BENEFITS
- Tasty Breaks: Delicious meals in our canteen – including vegetarian or vegan options, and many are even free of charge
- Move it, move it: 24/7 access to our gym – work out whenever you like, we cover the costs for you
- Do it your way: Flexible working hours, 30 days of vacation, up to 2 days of home office per week
- Easy Going: Germany ticket included
- Get Together: Team events, social days and more
- Grow like a Pro: Individual internal trainings and a wide range of development opportunities
- Good Vibes: You can expect an open, modern company culture with definitely no dress code, and plenty of room for new ideas and initiative
PAYBACK values diversity - come as you are