SRE Architect: Reliability, Platform & Automation

MACHINE LEARNING TECHNOLOGIES LLC

Atlanta (GA)

On-site

USD 110,000 - 193,000

Full time

39 hours ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

MACHINE LEARNING TECHNOLOGIES LLC seeks an experienced SRE Architect in Atlanta to lead reliability and scalability strategy for critical services. You will define standards, blueprints, and frameworks enabling teams to build and operate highly resilient systems.

You will mentor SREs, drive architectural reviews, and champion observability, incident readiness, and postmortems to improve overall platform health and efficiency.

Qualifications

  • Proven experience in an architectural role, designing solutions for reliability, scalability, and performance.
  • Deep understanding of SRE principles (SLIs/SLOs, error budgets, toil reduction, automation, incident management, postmortems).
  • Expertise in AWS cloud, networking, and security services.
  • Strong experience with Kubernetes, Docker, and serverless computing.
  • Solid experience designing and implementing observability solutions (Dynatrace, Prometheus, Grafana, ELK/EFK Stack, Jaeger, OpenTelemetry).
  • Strong programming/scripting skills (Python, Go, Bash) for automation and tool development.
  • Excellent analytical, problem-solving, and strategic thinking skills.
  • Strong communication, collaboration, and leadership skills with the ability to influence technical direction across teams.

Responsibilities

  • Architect and design highly available, scalable, secure, and cost-effective infrastructure and application patterns on AWS.
  • Define and evangelize SRE best practices, standards, and blueprints for service design, deployment, monitoring, and operational readiness.
  • Review observability implementation to identify gaps and define steps to reach higher maturity of observability setup.
  • Lead the definition and implementation strategy for SLIs, SLOs, and error budgets for critical services.
  • Design solutions to reduce operational toil through automation and improved system design.
  • Evaluate SRE tools and automation frameworks (CI/CD pipelines, IaC, automated incident remediation, chaos engineering) and suggest enhancements.
  • Prototype and recommend new technologies, tools, and methodologies to enhance reliability and developer productivity.
  • Act as a senior technical advisor on reliability, scalability, and performance for development teams.

Skills

Python
Go
Bash
SRE principles

Tools

AWS
Kubernetes
Docker
Serverless
Dynatrace
Prometheus
Grafana
ELK/EFK Stack
Jaeger
OpenTelemetry
CI/CD pipelines
Infrastructure as Code

Job description

MACHINE LEARNING TECHNOLOGIES LLC seeks an experienced SRE Architect in Atlanta to lead reliability and scalability strategy for critical services. You will define standards, blueprints, and frameworks enabling teams to build and operate highly resilient systems.

You will mentor SREs, drive architectural reviews, and champion observability, incident readiness, and postmortems to improve overall platform health and efficiency.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Site Reliability Engineering (SRE) Architect
Site Reliability Engineering (SRE) Architect

MACHINE LEARNING TECHNOLOGIES LLC • Atlanta (GA)

On-site
USD 110,000 - 193,000
Hybrid SRE Architect: Reliability, AWS & Observability
Hybrid SRE Architect: Reliability, AWS & Observability

STAFFWORXS • Atlanta (GA)

Hybrid
USD 96,432 - 103,320
Senior Enterprise SRE: Reliability & Observability
Senior Enterprise SRE: Reliability & Observability

Trigint Solutions LLC • Atlanta (GA)

Hybrid
USD 124,000 - 220,000
SRE Manager: Reliability, Automation & Platform Ops
SRE Manager: Reliability, Automation & Platform Ops

mtb • Buffalo (NY)

Hybrid
USD 150,000 - 230,000
SRE Lead: Reliability & Automation in AWS | Hybrid
SRE Lead: Reliability & Automation in AWS | Hybrid

Seek Now • Atlanta (GA)

Hybrid
USD 120,000 - 180,000
Competitive salary
Health, dental, and vision coverage
401(k) with company match
+1
Site Reliability Engineering (SRE) Architect
Site Reliability Engineering (SRE) Architect

STAFFWORXS • Atlanta (GA)

On-site
USD 96,432 - 103,320
Site Reliability Engineering (SRE) Architect
Site Reliability Engineering (SRE) Architect

Ethereum Technologies LLC • Atlanta (GA)

On-site
USD 150,000 - 190,000
Remote SRE Lead: Cloud Modernization & Reliability
Remote SRE Lead: Cloud Modernization & Reliability

LexisNexis Special Services Inc. • Alpharetta (GA), Northern (KY)

Hybrid
USD 118,000 - 264,000
Annual incentive bonus
Senior SRE: Architect of Reliable, Scalable Systems
Senior SRE: Architect of Reliable, Scalable Systems

Practice By Numbers, Inc. • Bellevue (WA), Northern (KY)

Hybrid
USD 140,000 - 200,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Alembic • Dunwoody (GA)

On-site
USD 150,000 - 190,000