Senior Site Reliability Engineer - Build SRE Ops (Hybrid)
Mission Staffing
New York (NY)
Hybrid
USD 140,000 - 200,000
Full time
14 days+
Application generator
Turn this role into an interview — a resume and cover letter built around what this employer wants.
Get past ATS filters
Job summary
A technology-driven investment firm is seeking a Senior Site Reliability Engineer to establish innovative SRE practices across its infrastructure. This key position in a hybrid setting involves collaborating with engineering teams to enhance service reliability and performance while maintaining high availability for critical production systems. The ideal candidate will have over 8 years of experience and expertise in observability tools like Prometheus and Grafana, ensuring robust operational performance in a complex tech ecosystem.
Qualifications
8+ years of experience in Site Reliability Engineering, Platform Engineering, or Infrastructure Engineering roles.
Experience operating large-scale distributed systems in production environments.
Strong expertise with observability and monitoring platforms.
Responsibilities
Establish and evolve Site Reliability Engineering practices and standards.
Design and scale observability and monitoring platforms.
Define reliability standards for applications running in Kubernetes environments.
Skills
Site Reliability Engineering
Observability and monitoring platforms
Containerization and orchestration
Scripting and automation with Python/Bash/Go
Tools
Prometheus
Grafana
Loki
Docker
Kubernetes
Job description
A technology-driven investment firm is seeking a Senior Site Reliability Engineer to establish innovative SRE practices across its infrastructure. This key position in a hybrid setting involves collaborating with engineering teams to enhance service reliability and performance while maintaining high availability for critical production systems. The ideal candidate will have over 8 years of experience and expertise in observability tools like Prometheus and Grafana, ensuring robust operational performance in a complex tech ecosystem.