Confirmed Site Reliability Engineer

swan.io

France

Sur place

EUR 70 000 - 110 000

Plein temps

Il y a 13 jours
Générateur de candidature

N’envoyez pas un CV générique — générez un CV et une lettre de motivation adaptés à ce poste précis.

Passez les filtres ATS

Résumé du poste

Swan's Core Infrastructure team seeks an experienced Site Reliability Engineer to own production systems, improve observability, automate tasks, and collaborate with development, product, and security teams.

You will respond to incidents, contribute to runbooks, and help design reliable, scalable services with strong security and compliance awareness in a financial services context.

Qualifications

  • 2 to 4 years of experience in Site Reliability Engineering, DevOps, platform engineering, infrastructure engineering, software engineering, or a related field.
  • Hands-on experience supporting production services and participating in an on-call rotation.
  • Experience with logs, metrics, dashboards, alerting, and basic distributed tracing.
  • Understanding of the Four Golden Signals: latency, traffic, errors, and saturation.
  • Practical experience with cloud infrastructure, ideally AWS, including compute, storage, networking, and managed services.
  • Experience with Infrastructure as Code, particularly Terraform, CloudFormation, or equivalent tools.
  • Ability to write automation scripts in Bash, Python, or Go.
  • Familiarity with CI/CD practices and deployment or infrastructure automation.
  • Knowledge of high availability, fault tolerance, redundancy, health checks, retries, circuit breakers, and failure recovery.
  • Familiarity with event-driven architectures and message queuing technologies like Kafka or AWS SQS.
  • Understanding of security/compliance requirements relevant to financial services (PCI DSS, ISO 27001, encryption, least privilege, data classification).
  • Awareness of capacity planning and basic cloud cost optimization.
  • Strong communication and collaboration across technical and non-technical teams.

Responsabilités

  • Act as a primary responder for production incidents and participate in the on-call rotation.
  • Assess incident impact across transaction volume, revenue, and data integrity.
  • Investigate issues using logs, metrics, dashboards, and tracing; document updates for stakeholders.
  • Contribute to postmortems, runbooks, and find recurring incident patterns.
  • Create and maintain dashboards, alerts, and SLIs/SLOs for supported services.
  • Tune alert thresholds to reduce noise and improve signal quality.
  • Participate in system design reviews focusing on reliability and production readiness.
  • Implement reliability improvements such as health checks, retries with backoff, circuit breakers, and monitoring.
  • Manage cloud resources and contribute to Infrastructure as Code (Terraform, CloudFormation, etc.).
  • Write automation scripts/tools in Bash, Python, or Go to reduce toil.
  • Contribute to CI/CD pipelines and automate routine maintenance tasks (backups, cert renewals, log management).
  • Review infrastructure code to maintain safety and repeatability of changes.
  • Assist security/compliance activities including PCI DSS and ISO 27001 controls.

Connaissances

Site Reliability Engineering
DevOps
Cloud infrastructure
IaC
Automation scripts
CI/CD
Observability
Distributed systems

Outils

Terraform
CloudFormation
Grafana
Datadog
Kafka
AWS SQS
Kubernetes

Description du poste

Working within Swan's Core Infrastructure team, you will help ensure the reliability, scalability, security, and performance of the platforms that support our financial services. You will take ownership of well-scoped services and operational incidents, improve observability, automate repetitive tasks, and collaborate closely with development, product, and security teams.

This is an independent engineering role for someone who has developed solid operational foundations and is ready to take greater ownership of production systems. You will contribute to incident response, infrastructure improvements, service design reviews, and the continuous improvement of our reliability practices.

Main responsibilities

On a daily basis, you will:

  • Act as a primary responder for well-understood production incidents and participate independently in the on-call rotation.
  • Assess the impact of incidents, including transaction volume affected, potential revenue impact, and implications for data integrity.
  • Investigate operational issues using logs, metrics, dashboards, and distributed tracing, then document clear incident updates for stakeholders.
  • Contribute to postmortems, update runbooks, identify recurring incident patterns, and suggest preventive measures.
  • Create and maintain dashboards, alerts, and basic service-level indicators for the services you support.
  • Tune alert thresholds to reduce noise and improve the quality of operational signals.
  • Participate in system design reviews, with a particular focus on reliability, operability, failure modes, and production readiness.
  • Implement reliability improvements such as health checks, retries with exponential backoff, circuit breakers, and appropriate monitoring.
  • Manage cloud resources and contribute to Infrastructure as Code using tools such as Terraform.
  • Write automation scripts and small internal tools in Bash, Python, or Go to reduce manual toil and improve operational efficiency.
  • Contribute to CI/CD pipelines and automate routine maintenance tasks such as backup verification, certificate renewal, and log management.
  • Participate in infrastructure code reviews and help maintain high standards for safe, repeatable changes.
  • Support security and compliance activities, including PCI DSS controls, ISO 27001 initiatives, security remediation, data classification, and encryption requirements.
  • Monitor resource utilisation, provide basic capacity forecasts, and implement practical cost optimisation measures such as rightsizing resources and removing unused infrastructure.
  • Collaborate with development, product, and security teams to improve the resilience and operability of services.
  • Provide clear handovers, maintain high-quality documentation, and communicate technical topics effectively to both technical and non-technical stakeholders.
  • Use approved AI tools responsibly to support tasks such as code generation, documentation and log analysis, while validating outputs and protecting sensitive information.
Your team

Core Infrastructure is responsible for building and operating the foundations that enable Swan's products to remain reliable as the business grows. We work closely with development and other technical teams to improve system resilience, operational efficiency, and customer experience.

We value ownership, pragmatism, knowledge sharing, and open communication. Engineers are encouraged to challenge ideas constructively, document what they learn, and continuously improve the way we build and operate services. You will work in a supportive environment where reliability is a shared responsibility and where operational excellence is developed through collaboration.

Together alongside Engineering Productivity, our squad constitutes the broader Platform Engineering team.

You're a great match if:

You have typically 2 to 4 years of experience in Site Reliability Engineering, DevOps, platform engineering, infrastructure engineering, software engineering, or a related field. You have hands-on experience supporting production services and participating in an on-call rotation. You can independently respond to well-understood incidents, follow escalation procedures, and contribute to postmortems and runbook improvements. You are comfortable working with logs, metrics, dashboards, alerting, and basic distributed tracing. You understand the Four Golden Signals: latency, traffic, errors, and saturation. You have experience creating dashboards and meaningful alerts, and understand the fundamentals of SLIs and SLOs. You have practical experience with cloud infrastructure, ideally AWS, including compute, storage, networking, and managed services. You have experience with Infrastructure as Code, particularly Terraform, CloudFormation, or equivalent tools. You can write automation scripts in one or more languages such as Bash, Python, or Go. You understand CI/CD practices and have contributed to deployment or infrastructure automation. You have a working understanding of high availability, fault tolerance, redundancy, health checks, retries, circuit breakers, and failure recovery. You are familiar with event-driven architectures and message queuing technologies such as Kafka or AWS SQS. You understand the operational implications of distributed systems, including consistency, availability, and partition tolerance. You have an awareness of security and compliance requirements relevant to financial services, such as PCI DSS, ISO 27001, encryption, least privilege, and data classification. You are comfortable monitoring resource usage, thinking about capacity, and identifying basic cloud cost optimisation opportunities. You communicate clearly, document your work thoroughly, and collaborate effectively across technical and non-technical teams.

Nice to have AWS certification, such as AWS Certified Solutions Architect, Associate level. Experience with Kubernetes and container orchestration. Experience with observability platforms such as Grafana, Datadog, or similar. Experience operating services that process financial transactions or other highly sensitive data. Experience improving service-level objectives, performance, or capacity for production systems. Familiarity with Go or another programming language used for internal tooling.

Our ideal teammate

Empathetic. Skilled. Frank. We love to challenge each other, and we leave our egos at the door.

It's okay if you don't tick all the boxes - don't let imposter syndrome prevent you from applying!

Swan is committed to providing a caring work environment for all employees, regardless of age, sex, disability, sexual orientation, race, religion, or belief. When it comes to recruitment, we're interested in your work experience, skills, and overall personality. Because diversity makes the workplace stronger and is necessary for Swan's success, we are intensifying efforts to incorporate concrete actions to help us improve in this area.

Obtenez votre examen gratuit et confidentiel de votre CV.

ou faites glisser et déposez votre fichier ici.

Similar jobs

Postes similaires à comparer

Confirmed Site Reliability Engineer
Confirmed Site Reliability Engineer

Swan • Paris

Hybride
EUR 80 000 - 110 000
Meal vouchers
Transport package
Holidays 25 days + RTT
+6
Confirmed Site Reliability Engineer
Confirmed Site Reliability Engineer

swan.io • Paris

Hybride
EUR 75 000 - 110 000
Meal vouchers
Transport package
Holidays: 25 days + RTT
+3
Senior Site Reliability Engineer (SRE)
Senior Site Reliability Engineer (SRE)

Swile (ex Lunchr) • Montpellier

Sur place
EUR 65 000 - 85 000
Competitive salary and benefits package
Professional development opportunities
Collaborative work environment
Software Engineer (backend Typescript)
Software Engineer (backend Typescript)

swan.io • Paris, Bordeaux

Sur place
EUR 53 000 - 70 000
Software Engineer (backend Typescript)
Software Engineer (backend Typescript)

swan.io • France

Hybride
EUR 53 000 - 70 000
Software Engineer (backend Typescript)
Software Engineer (backend Typescript)

swanio • Paris

Sur place
EUR 53 000 - 70 000
Hybrid remote policy
Relocation package to Paris
Health insurance (mutuelle)
Site Reliability Engineer — On-Call & Automation
Site Reliability Engineer — On-Call & Automation

Swan • Paris

Hybride
EUR 80 000 - 110 000
Meal vouchers
Transport package
Holidays 25 days + RTT
+6
Senior Site Reliability Engineer (SRE)
Senior Site Reliability Engineer (SRE)

Swile • Montpellier

Sur place
EUR 50 000 - 80 000
Competitive salary and benefits
Professional development opportunities
Collaborative work environment
Confirmed TypeScript Software Engineer
Confirmed TypeScript Software Engineer

swanio • Paris

Sur place
EUR 53 000 - 63 000
Hybrid remote policy
Relocation package (Paris)
25 days holidays + RTT
+5
Software Engineer - Banking Experience Team (backend TypeScript)
Software Engineer - Banking Experience Team (backend TypeScript)

Swan • Paris

Hybride
EUR 53 000 - 70 000
Hybrid remote policy
Relocation package
25 days + RTT holidays
+6