Senior Site Reliability Engineer: Observability & Cloud
VBeyond Corporation
Jersey City (NJ)
On-site
USD 100,000 - 260,000
Full time
14 days+
Application generator
Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.
Get past ATS filters
Job summary
A technology firm is seeking a Site Reliability Engineer for a full-time position. The role involves focusing on observability, Kubernetes, and cloud infrastructure. Responsibilities include managing an observability stack, improving cluster reliability, and developing Terraform modules. Ideal candidates will have 4-8 years of SRE experience and strong automation skills in Python or Go. Compensation is competitive, reflecting seniority and expertise in the field.
Qualifications
4–8 years of experience in SRE, infrastructure, or Kubernetes operations.
Strong expertise in observability tools, Terraform, automation, CI/CD, and cloud networking.
Responsibilities
Ownership of observability stack: Prometheus, Grafana, OpenTelemetry.
Build and maintain reliable monitoring pipelines for metrics and alerts.
Implement AI-assisted diagnostics for anomaly detection.
Lead SLO reporting, incident management, and root cause analysis.
Skills
Kubernetes
Observability tools
Terraform
Automation (Python/Go)
CI/CD
Cloud networking
Job description
A technology firm is seeking a Site Reliability Engineer for a full-time position. The role involves focusing on observability, Kubernetes, and cloud infrastructure. Responsibilities include managing an observability stack, improving cluster reliability, and developing Terraform modules. Ideal candidates will have 4-8 years of SRE experience and strong automation skills in Python or Go. Compensation is competitive, reflecting seniority and expertise in the field.