- Ensure the reliability, scalability, security, and performance of platforms supporting Swan's financial services
- Own well-scoped services and operational incidents
- Participate independently in the on-call rotation as a primary responder for well-understood production incidents
- Assess incident impact, including transaction volume, revenue impact, and data-integrity implications
- Investigate operational issues using logs, metrics, dashboards, and distributed tracing
- Document incident updates for stakeholders
- Contribute to postmortems, update runbooks, identify recurring incident patterns, and suggest preventive measures
- Create and maintain dashboards, alerts, and basic service-level indicators
- Tune alert thresholds to reduce noise and improve operational signals
- Participate in system design reviews focused on reliability, operability, failure modes, and production readiness
- Implement health checks, exponential-backoff retries, circuit breakers, and monitoring
- Manage cloud resources and contribute to Infrastructure as Code using Terraform
- Write Bash, Python, or Go automation scripts and internal tools
- Contribute to CI/CD pipelines and automate backup verification, certificate renewal, and log management
- Participate in infrastructure code reviews
- Support PCI DSS, ISO 27001, security remediation, data classification, and encryption activities
- Monitor resource utilisation, provide capacity forecasts, and optimise cloud costs
- Collaborate with development, product, and security teams to improve service resilience and operability
- Maintain documentation, provide handovers, and communicate technical topics to technical and non-technical stakeholders
- Use approved AI tools responsibly for code generation, documentation, and log analysis while protecting sensitive information
Requirements
- Typically 2 to 4 years of experience in Site Reliability Engineering, DevOps, platform engineering, infrastructure engineering, software engineering, or a related field
- Hands-on experience supporting production services and participating in an on-call rotation
- Ability to independently respond to well-understood incidents, follow escalation procedures, and contribute to postmortems and runbook improvements
- Comfortable working with logs, metrics, dashboards, alerting, and basic distributed tracing
- Understanding of the Four Golden Signals: latency, traffic, errors, and saturation
- Experience creating dashboards and meaningful alerts, with understanding of SLIs and SLOs
- Practical experience with cloud infrastructure, ideally AWS
- Experience with Infrastructure as Code, particularly Terraform, CloudFormation, or equivalent tools
- Ability to write automation scripts in Bash, Python, or Go
- Understanding of CI/CD practices and deployment or infrastructure automation
- Working understanding of high availability, fault tolerance, redundancy, health checks, retries, circuit breakers, and failure recovery
- Familiarity with event-driven architectures and message queuing technologies such as Kafka or AWS SQS
- Understanding of distributed-systems operational implications, including consistency, availability, and partition tolerance
- Awareness of financial-services security and compliance requirements, including PCI DSS, ISO 27001, encryption, least privilege, and data classification
- Comfortable monitoring resource usage, capacity planning, and basic cloud cost optimisation
- Clear communication, thorough documentation, and effective collaboration across technical and non-technical teams
- AWS certification, such as AWS Certified Solutions Architect Associate, is nice to have
- Experience with Kubernetes and container orchestration is nice to have
- Experience with observability platforms such as Grafana or Datadog is nice to have
- Experience operating services processing financial transactions or highly sensitive data is nice to have
- Experience improving production-system service-level objectives, performance, or capacity is nice to have
- Familiarity with Go or another internal-tooling language is nice to have
Core Competencies
Demonstrates expertise in Site Reliability Engineering and DevOps practices, with a strong focus on cloud infrastructure management, incident response, and automation scripting. Proficient in creating dashboards, monitoring systems, and ensuring compliance with financial services security standards.
Highest-signal resume keywords
- Site Reliability Engineering
- Infrastructure as Code
- Cloud Infrastructure Management
- Automation Scripting
- Incident Response
ATS Optimization Keywords
Hard Skills
- Bash Scripting
- Python Scripting
- Go Scripting
- Terraform
- AWS
- CI/CD Practices
- Monitoring Systems
- Distributed Tracing
- Service-Level Indicators
- High Availability
Soft Skills
- Clear Communication
- Effective Collaboration
- Thorough Documentation
Certifications & Qualifications
- AWS Certified Solutions Architect Associate
Industry Keywords
- PCI DSS
- ISO 27001
- Data Classification
- Encryption
- Financial Services
Tools & Technologies
- Grafana
- Datadog
- Kubernetes
- AWS SQS
- Kafka