SRE Manager / SRE Architect

Olik Global

South Carolina

Hybrid

USD 130,000 - 160,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Olik Global is seeking a hands-on SRE Manager/Architect located in South Carolina. This role involves leading reliability initiatives, managing release processes, and optimizing cloud infrastructure across enterprise applications.

The ideal candidate will have 15+ years of IT experience with a focus on Site Reliability Engineering and strong expertise in AWS, Azure, DevOps, and Automation. The position also offers a hybrid work model allowing flexibility.

Qualifications

  • 15+ years of IT experience with focus on SRE, DevOps, or Production Support.
  • Hands-on experience implementing SRE practices in enterprise environments.
  • Strong understanding of SLI/SLO/Error Budgets and Incident Management.

Responsibilities

  • Design and implement SRE best practices for reliability and scalability.
  • Drive end-to-end release management processes across multiple environments.
  • Automate infrastructure provisioning and deployment workflows.

Skills

Site Reliability Engineering
DevOps
Cloud Engineering
Release Management
Automation

Tools

AWS
Azure
Kubernetes
Docker
Terraform
GitHub Actions
Jenkins

Job description

Job Description - SRE Manager / SRE Architect (Hands-on)

Location: New York City, NY / Fort Mill, SC (Hybrid)

Employment Type: Full‑Time / Contract

Industry: Financial Services

Position Overview

We are seeking a highly experienced and hands‑on Site Reliability Engineering (SRE) Manager / SRE Architect to lead reliability, availability, performance, and release management initiatives across enterprise‑scale applications and platforms. This role requires a strong blend of SRE, DevOps, Release Management, Cloud Engineering, Automation, and Production Operations expertise.

The ideal candidate will be deeply involved in designing and implementing reliability strategies, driving release governance, improving deployment processes, and ensuring operational excellence across cloud‑native environments.

Key Responsibilities

LaunchDarkly experience is highly preferred but not mandatory.

Site Reliability Engineering (SRE)
  • Design and implement SRE best practices focused on reliability, scalability, performance, and availability.
  • Define and monitor SLIs, SLOs, and error budgets across critical applications and services.
  • Drive proactive monitoring, alerting, observability, and incident management processes.
  • Lead root cause analysis (RCA) efforts and implement preventive measures.
  • Improve system resiliency through automation, self‑healing capabilities, and operational excellence.
  • Establish reliability standards across distributed systems and cloud platforms.
Release Management
  • Own and drive end‑to‑end release management processes across multiple environments.
  • Coordinate application releases across development, QA, UAT, staging, and production environments.
  • Develop release governance, release calendars, deployment strategies, rollback procedures, and change management processes.
  • Partner with development, QA, infrastructure, and business teams to ensure smooth production deployments.
  • Identify and mitigate release risks while minimizing downtime and business impact.
  • Implement deployment automation and continuous delivery best practices.
DevOps & Automation
  • Design and maintain CI/CD pipelines using modern DevOps tools.
  • Automate infrastructure provisioning, deployment, monitoring, and operational workflows.
  • Drive Infrastructure as Code (IaC) adoption using Terraform or similar technologies.
  • Support cloud‑native architectures and containerised application deployments.
  • Partner with engineering teams to improve developer productivity and deployment velocity.
Cloud & Platform Engineering
  • Manage and optimise cloud infrastructure on AWS and/or Azure.
  • Support Kubernetes, container orchestration, and cloud‑native application platforms.
  • Ensure platform scalability, security, compliance, and operational readiness.
  • Drive platform modernization initiatives and operational transformation efforts.
Required Skills & Experience
Core SRE Skills
  • 15+ years of IT experience with a strong focus on SRE, DevOps, Platform Engineering, or Production Support.
  • Extensive hands‑on experience implementing SRE practices in enterprise environments.
  • Strong understanding of:
  • SLI/SLO/Error Budgets
  • Incident Management
  • Problem Management
  • Capacity Planning
  • Reliability Engineering
  • Observability & Monitoring
Release Management
  • Proven experience managing large‑scale production releases.
  • Strong expertise in:
  • Release Planning
  • Release Governance
  • Change Management
  • Deployment Automation
  • Rollback Strategies
  • Production Readiness Reviews
DevOps & Cloud
  • Hands‑on experience with:
  • AWS and/or Azure
  • Kubernetes (EKS, AKS, OpenShift preferred)
  • Docker
  • Terraform
  • GitHub Actions, Jenkins, Azure DevOps, GitLab CI/CD
  • Experience building and maintaining CI/CD pipelines.
Monitoring & Observability
  • Strong experience with:
  • Dynatrace
  • Datadog
  • Splunk
  • Prometheus
  • Grafana
  • ELK Stack
  • CloudWatch
Scripting & Automation
  • Experience with Python, Bash, PowerShell, or similar scripting languages.
  • Strong automation mindset with focus on operational efficiency.
Nice to Have
  • LaunchDarkly end‑to‑end implementation experience.
  • Feature flag management and progressive delivery strategies.
  • Financial Services, Banking, or Wealth Management domain experience.
  • Experience leading SRE or DevOps transformation initiatives.
  • Cloud certifications (AWS, Azure, Kubernetes).
Preferred Candidate Profile
  • Strong hands‑on SRE leader, not just a people manager.
  • Deep expertise in Release Management and Production Support.
  • Proven background in DevOps, Cloud Engineering, and Platform Reliability.
  • Ability to work with development, infrastructure, security, and business teams.
Keywords

SRE, Site Reliability Engineering, Release Management, DevOps, Terraform, AWS, Azure, Kubernetes, Dynatrace, CI/CD, LaunchDarkly, Production Support, Incident Management, Reliability Engineering, Observability, Platform Engineering, Infrastructure Automation.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Lead Site Reliability Engineer (SRE) / Principal Site Reliability Engineer (SRE)
Lead Site Reliability Engineer (SRE) / Principal Site Reliability Engineer (SRE)

Mindlance • Irving (TX)

Hybrid
USD 120,000 - 160,000
Site Reliability Engineer
Site Reliability Engineer

Jobtailor • New Jersey

On-site
USD 120,000 - 180,000
Site Reliability Engineer
Site Reliability Engineer

TalentDome Staffing • United States

On-site
USD 140,000 - 210,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Mission Staffing • New York (NY)

Hybrid
USD 140,000 - 200,000
Site Reliability Engineering (SRE)
Site Reliability Engineering (SRE)

Weekday (YC W21) • New York (NY)

On-site
USD 150,000 - 250,000
Health, dental, vision insurance
Generous PTO
Learning & development
+2
Site Reliability Engineer -Jersey City, NJ & Dallas, TX
Site Reliability Engineer -Jersey City, NJ & Dallas, TX

StradIT • Jersey City (NJ)

Hybrid
USD 120,000 - 160,000
CloudDevs: Senior Site Reliability Engineer (SRE)
CloudDevs: Senior Site Reliability Engineer (SRE)

Breakout Tools • San Francisco (CA)

On-site
USD 120,000 - 160,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

The ReWork Group • New York (NY)

On-site
USD 120,000 - 160,000
Staff Site Reliability Engineer, SRE
Staff Site Reliability Engineer, SRE

Jobtailor • California (MO)

On-site
USD 120,000 - 210,000
Lead, Site Reliability Engineer
Lead, Site Reliability Engineer

CardWorks • Pittsburgh

Hybrid
USD 146,000 - 163,000
Competitive Pay
Medical, Dental, and Vision Benefits
401(k) Plan with Company Match
+1