Blitzy is seeking a Senior Site Reliability Engineer for its Cambridge, MA headquarters. The role centers on ensuring reliability and operational excellence for our AI software development platform. Candidates should have 5–8 years of experience with Kubernetes expertise and proficiency in cloud environments, particularly AWS. The position offers a high-impact opportunity to influence systems architecture while providing competitive compensation and equity in a rapidly growing company.
Qualifications
5–8 years of experience in Site Reliability Engineering, DevOps, or Platform Engineering.
Deep expertise in Kubernetes.
Strong proficiency with cloud platforms like AWS.
Hands-on Terraform experience for infrastructure provisioning.
Solid scripting and automation skills in Python, Go, or Bash.
Responsibilities
Design, build, and operate fault-tolerant infrastructure across cloud environments.
Define and own SLOs, SLAs, and error budgets for critical services.
Build and maintain CI/CD pipelines and deployment infrastructure.
Manage Kubernetes clusters for AI workloads.
Drive infrastructure-as-code practices using Terraform.
Skills
Kubernetes expertise
Cloud platform proficiency (AWS preferred)
Terraform experience
Scripting and automation skills
Tools
Prometheus
Grafana
Datadog
OpenTelemetry
Job description
Blitzy is seeking a Senior Site Reliability Engineer for its Cambridge, MA headquarters. The role centers on ensuring reliability and operational excellence for our AI software development platform. Candidates should have 5–8 years of experience with Kubernetes expertise and proficiency in cloud environments, particularly AWS. The position offers a high-impact opportunity to influence systems architecture while providing competitive compensation and equity in a rapidly growing company.