Senior SRE: Scale, Observability & Automation Leader

Alembic Technologies

Dunwoody (GA)

On-site

USD 200,000 - 225,000

Full time

37 hours ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Alembic Technologies is seeking an experienced Site Reliability Engineer to scale our platform with reliability, observability, and operational excellence. This onsite role in Dunwoody, GA partners with engineers and data scientists to automate and maintain infrastructure powering data pipelines, ML workloads, and real-time analytics systems.

The role is hands-on and high impact, offering visibility across the stack and opportunity to shape our infrastructure and operations.

Qualifications

  • 8+ years of experience in SRE, DevOps, or infrastructure engineering roles (highly experienced).
  • 5+ years of datacenter operations and/or system and network administration.
  • Experience with containerization (Docker) and orchestration (Kubernetes).
  • Strong knowledge of Linux systems, networking, and performance tuning; proficient with SSH and Linux CLI.
  • Infrastructure-as-code experience (Terraform, Ansible).
  • Good programming skills and ability to apply coding principles to IaC and scripting (Python, Bash).
  • Experience with monitoring/observability stacks (Prometheus, Grafana, Datadog, ELK, OpenTelemetry) and CI/CD tools (GitHub Actions, ArgoCD).
  • Ability to debug complex systems and automate solutions in scripting languages.

Responsibilities

  • Design, build, and maintain scalable infrastructure to support real‑time analytics and ML workloads.
  • Improve reliability and performance through automation, observability, and capacity planning.
  • Own and evolve CI/CD pipelines, deployment automation, rollback mechanisms, and config management.
  • Implement and maintain monitoring, alerting, and incident response processes (SLOs, runbooks, on‑call rotations).
  • Collaborate across engineering and data science to drive performance and reliability.
  • Ensure security, compliance, and operational readiness across cloud infrastructure.
  • Drive post‑incident analysis and continuous improvement initiatives.

Skills

SRE/DevOps experience
Linux administration
Networking basics
CI/CD tooling
Scripting (Python/Bash)
Infrastructure as Code
Monitoring & observability
Cloud concepts

Tools

Docker
Kubernetes
GitHub Actions
ArgoCD
Prometheus
Grafana
Datadog
ELK
OpenTelemetry
Terraform
Ansible

Job description

Alembic Technologies is seeking an experienced Site Reliability Engineer to scale our platform with reliability, observability, and operational excellence. This onsite role in Dunwoody, GA partners with engineers and data scientists to automate and maintain infrastructure powering data pipelines, ML workloads, and real-time analytics systems.

The role is hands-on and high impact, offering visibility across the stack and opportunity to shape our infrastructure and operations.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer
Senior Site Reliability Engineer

Alembic Technologies • Dunwoody (GA)

On-site
USD 200,000 - 225,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Alembic • Dunwoody (GA)

On-site
USD 150,000 - 190,000
Senior SRE: Observability, Automation & Scalable Systems
Senior SRE: Observability, Automation & Scalable Systems

Replit • Northern (KY)

Hybrid
USD 140,000 - 190,000
Competitive Salary & Equity
401(k) 4% match (US)
Health, Dental, Vision & Life
+7
Senior SRE: Scale Infra, Automate, Elevate Reliability
Senior SRE: Scale Infra, Automate, Elevate Reliability

Fathom.ai • United States

On-site
USD 100,000 - 130,000
Competitive compensation
Supportive environment for personal growth
Dynamic and collaborative team
Senior SRE: Scale Reliability, Observability & Resilience
Senior SRE: Scale Reliability, Observability & Resilience

Early Warning Services LLC • Scottsdale (AZ)

Hybrid
USD 106,000 - 130,000
Healthcare Coverage
401(k) Retirement Plan
Paid Time Off
+2
Senior SRE: Reliability & Observability Lead
Senior SRE: Reliability & Observability Lead

Inspire • Atlanta (GA)

On-site
USD 140,000 - 200,000
Senior SRE: Scale Reliability & Observability
Senior SRE: Scale Reliability & Observability

Megaport • Abbeyville (CO)

On-site
USD 130,000 - 190,000
Contractor (PJ)
Paid Time Off
Competitive Compensation
+4
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Alembic Technologies • San Francisco (CA)

On-site
USD 210,000 - 240,000
Ownership of critical infrastructure
High-performance engineering culture
Influence platform scaling
Senior SRE: AI-Driven Platform Reliability & Scale
Senior SRE: AI-Driven Platform Reliability & Scale

Medallia • McLean (VA)

On-site
USD 129,000 - 190,000
Health benefits
401(k) matching
Paid parental leave
+1
Senior Site Reliability Engineer – Scale & Observability
Senior Site Reliability Engineer – Scale & Observability

Inspire Brands, Inc. • Atlanta (GA)

On-site
USD 120,000 - 180,000