Site Reliability Engineer - Fulltime Only
Direct message the job poster from VBeyond Corporation
- SRE role focused on observability, Kubernetes, and cloud infrastructure (AWS/GCP/EKS)
- Ownership of observability stack: Prometheus, Grafana, OpenTelemetry, ELK/Loki/Splunk, Jaeger, Alertmanager, SLOs
- Build and maintain reliable monitoring pipelines for metrics, logs, tracing, dashboards, and alerts
- Develop Terraform modules for observability infrastructure, Kubernetes components, and cluster add-ons
- Improve cluster reliability through automation, performance tuning, capacity planning, and remediation
- Implement AI-assisted diagnostics for anomaly detection, alert tuning, and noise reduction
- Collaborate with Platform Engineering on Istio/service mesh telemetry and platform health
- Lead SLO reporting, incident management, and root cause analysis
- 4–8 years of experience in SRE, infrastructure, or Kubernetes operations
- Strong expertise in observability tools, Terraform, automation (Python/Go), CI/CD, and cloud networking
Note – VBeyond is fully committed to Diversity and Equal Employment Opportunity.
Seniority level
Mid-Senior level
Employment type
Full-time
New York, NY $100,000.00-$260,000.00 4 days ago