N’envoyez pas un CV générique — générez un CV et une lettre de motivation adaptés à ce poste précis.
Swan's Core Infrastructure team seeks an experienced Site Reliability Engineer to own production systems, improve observability, automate tasks, and collaborate with development, product, and security teams.
You will respond to incidents, contribute to runbooks, and help design reliable, scalable services with strong security and compliance awareness in a financial services context.
Working within Swan's Core Infrastructure team, you will help ensure the reliability, scalability, security, and performance of the platforms that support our financial services. You will take ownership of well-scoped services and operational incidents, improve observability, automate repetitive tasks, and collaborate closely with development, product, and security teams.
This is an independent engineering role for someone who has developed solid operational foundations and is ready to take greater ownership of production systems. You will contribute to incident response, infrastructure improvements, service design reviews, and the continuous improvement of our reliability practices.
On a daily basis, you will:
Core Infrastructure is responsible for building and operating the foundations that enable Swan's products to remain reliable as the business grows. We work closely with development and other technical teams to improve system resilience, operational efficiency, and customer experience.
We value ownership, pragmatism, knowledge sharing, and open communication. Engineers are encouraged to challenge ideas constructively, document what they learn, and continuously improve the way we build and operate services. You will work in a supportive environment where reliability is a shared responsibility and where operational excellence is developed through collaboration.
Together alongside Engineering Productivity, our squad constitutes the broader Platform Engineering team.
You have typically 2 to 4 years of experience in Site Reliability Engineering, DevOps, platform engineering, infrastructure engineering, software engineering, or a related field. You have hands-on experience supporting production services and participating in an on-call rotation. You can independently respond to well-understood incidents, follow escalation procedures, and contribute to postmortems and runbook improvements. You are comfortable working with logs, metrics, dashboards, alerting, and basic distributed tracing. You understand the Four Golden Signals: latency, traffic, errors, and saturation. You have experience creating dashboards and meaningful alerts, and understand the fundamentals of SLIs and SLOs. You have practical experience with cloud infrastructure, ideally AWS, including compute, storage, networking, and managed services. You have experience with Infrastructure as Code, particularly Terraform, CloudFormation, or equivalent tools. You can write automation scripts in one or more languages such as Bash, Python, or Go. You understand CI/CD practices and have contributed to deployment or infrastructure automation. You have a working understanding of high availability, fault tolerance, redundancy, health checks, retries, circuit breakers, and failure recovery. You are familiar with event-driven architectures and message queuing technologies such as Kafka or AWS SQS. You understand the operational implications of distributed systems, including consistency, availability, and partition tolerance. You have an awareness of security and compliance requirements relevant to financial services, such as PCI DSS, ISO 27001, encryption, least privilege, and data classification. You are comfortable monitoring resource usage, thinking about capacity, and identifying basic cloud cost optimisation opportunities. You communicate clearly, document your work thoroughly, and collaborate effectively across technical and non-technical teams.
Nice to have AWS certification, such as AWS Certified Solutions Architect, Associate level. Experience with Kubernetes and container orchestration. Experience with observability platforms such as Grafana, Datadog, or similar. Experience operating services that process financial transactions or other highly sensitive data. Experience improving service-level objectives, performance, or capacity for production systems. Familiarity with Go or another programming language used for internal tooling.
Empathetic. Skilled. Frank. We love to challenge each other, and we leave our egos at the door.
It's okay if you don't tick all the boxes - don't let imposter syndrome prevent you from applying!
Swan is committed to providing a caring work environment for all employees, regardless of age, sex, disability, sexual orientation, race, religion, or belief. When it comes to recruitment, we're interested in your work experience, skills, and overall personality. Because diversity makes the workplace stronger and is necessary for Swan's success, we are intensifying efforts to incorporate concrete actions to help us improve in this area.