- Refactor existing solutions using appropriate design patterns.
We are looking for a highly skilled Site Reliability Engineer II (SRE II) to join a technology‑driven environment focused on reliability, scalability, performance, and automation. This role is ideal for an engineer who treats operations as a software problem and enjoys building solutions that eliminate operational toil while improving system resilience.
You will work closely with software engineers, infrastructure teams, and product stakeholders to ensure systems remain highly available, observable, secure, and scalable. The position combines software engineering, production operations, incident management, and cloud/infrastructure automation.
The environment is heavily focused on AWS, Kubernetes, CI/CD automation, observability, and reliability engineering, Java.
Key Responsibilities
Reliability & Operations
- Own services and applications throughout their lifecycle.
- Monitor application health, reliability, and performance.
- Define, maintain, and improve operational metrics and SLIs/SLOs.
- Participate in incident response and on‑call activities.
- Conduct root cause analyses and contribute to postmortem activities.
Automation & Engineering
- Develop automation solutions to reduce manual operational work.
- Build software and tooling that improve scalability, availability, and efficiency.
- Reduce technical debt and improve platform maintainability.Optimize infrastructure and application performance.
Software Development
- Build and maintain production‑grade software applications.
- Write clean, reusable, and maintainable code.
- Refactor existing solutions using appropriate design patterns.
- Ensure code quality through automated testing and engineering best practices.
Architecture & System Design
- Contribute to solution design discussions.
- Evaluate architectural approaches based on business and technical requirements.Support future scalability and reliability initiatives.
- Advise development teams on observability, reliability, and operational excellence.
Observability & Monitoring
- Improve monitoring, logging, tracing, and alerting capabilities.
- Define meaningful operational dashboards and KPIs.
- Support capacity planning and performance optimization initiatives.
Required Skills & Experience
- 3‑5 years of relevant experience in Site Reliability Engineering, DevOps, Platform Engineering, or Software Engineering.
- Strong software development background.
- Experience with automation and infrastructure management.
- Knowledge of observability, monitoring, and alerting frameworks.
- Experience managing production systems and cloud‑based environments.
- Strong incident management and troubleshooting skills.
- Understanding of system design, scalability, reliability, and performance engineering.
- Experience with CI/CD and modern deployment practices.
- Strong analytical and problem‑solving capabilities.
- Excellent communication and stakeholder management skills.
Details
- Start: Asap
- Duration: 12 months (+)
- Location: Amsterdam, hybrid