Site Reliability Engineering (SRE)
Location: New York City, NY
Salary Range: $150,000 – $250,000 (Base)
About the Role
We are seeking an experienced Site Reliability Engineer (SRE) to join our engineering team in New York City. In this role, you will be responsible for building and maintaining highly available, scalable, and reliable systems. You will work across infrastructure, cloud platforms, automation, observability, and deployment processes to ensure our applications operate efficiently and reliably at scale.
You will collaborate closely with software engineers, DevOps teams, and infrastructure stakeholders to improve system performance, automate operational processes, strengthen reliability, and minimize downtime.
This role is ideal for engineers who enjoy solving complex infrastructure challenges, automating repetitive processes, and building resilient distributed systems.
Required Qualifications
- Bachelor's degree in Computer Science, Engineering, Information Technology, or a related technical field (or equivalent experience)
- 3+ years of professional experience in Site Reliability Engineering, DevOps, Infrastructure Engineering, or a related field
- Strong experience with cloud platforms such as AWS, Azure, or Google Cloud Platform
- Hands‑on experience with Infrastructure as Code (IaC) using Terraform, CloudFormation, or similar tools
- Experience designing and maintaining CI/CD pipelines using Jenkins, GitHub Actions, GitLab CI, or similar technologies
- Strong knowledge of Docker, Kubernetes, and container orchestration
- Experience administering Linux‑based systems and writing automation scripts using Bash, Python, or PowerShell
- Experience with monitoring, observability, and logging tools such as Prometheus, Grafana, ELK, Datadog, or similar platforms
- Strong understanding of system reliability, availability, scalability, and performance
- Experience troubleshooting production systems and resolving complex infrastructure issues
- Strong automation and problem‑solving skills
- Ability to work collaboratively with software engineering and infrastructure teams
Preferred Qualifications
- Hands‑on experience operating Kubernetes in production environments
- Experience with Site Reliability Engineering practices, including SLIs, SLOs, SLAs, and error budgets
- Knowledge of networking, cloud security, IAM, and infrastructure best practices
- Experience implementing GitOps workflows using tools such as Argo CD or Flux
- Familiarity with microservices and distributed systems architecture
- Experience with incident response, root cause analysis, and production troubleshooting
- Experience building scalable and highly available systems in cloud environments
- Familiarity with capacity planning, disaster recovery, and business continuity strategies
- Experience with observability, alerting, and performance monitoring at scale
Compensation & Benefits
- Base Salary: $150,000 – $250,000, depending on experience and qualifications
- Competitive equity or bonus opportunities
- Comprehensive health, dental, and vision benefits
- Generous paid time off and holidays
- Professional development and learning opportunities
- Opportunity to work with modern cloud and infrastructure technologies
- Collaborative and innovative engineering environment in New York City