A global online fashion retailer is seeking a Senior Site Reliability Engineer to ensure the reliability of mission-critical systems. This role requires maintaining high availability services, collaborating with cross-functional teams, and managing infrastructure using tools such as Kubernetes and Nginx. Ideal candidates have deep experience in operating large-scale systems, a solid background in Linux and networking, and a passion for problem-solving. Employees enjoy a range of benefits, including health accounts, 401(k) plans, and more.
Qualifications
3+ years of experience with high-traffic production systems.
Hands-on experience with troubleshooting distributed systems.
Strong communication skills for collaboration across teams.
Responsibilities
Maintain mission-critical production systems 24/7.
Triage production incidents and analyze root causes.
Monitor capacity planning and resource utilization.
Skills
Linux expertise
Networking knowledge
Software engineering in Python or Go
Incident response
Performance optimization
Observability systems (Prometheus, Grafana)
Education
Bachelor’s degree in Computer Science or related field
Tools
Kubernetes
Redis
Kafka
APISIX
Elasticsearch
Job description
A global online fashion retailer is seeking a Senior Site Reliability Engineer to ensure the reliability of mission-critical systems. This role requires maintaining high availability services, collaborating with cross-functional teams, and managing infrastructure using tools such as Kubernetes and Nginx. Ideal candidates have deep experience in operating large-scale systems, a solid background in Linux and networking, and a passion for problem-solving. Employees enjoy a range of benefits, including health accounts, 401(k) plans, and more.