SRE Operations

SFE

Charlotte (NC)

On-site

USD 140,000 - 180,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

SFE is seeking an SRE Operations Lead to spearhead 24x7 production reliability for AWS, GitLab and AI-driven operations in Charlotte, onsite. You will own reliability metrics, drive observability and automation, and coordinate major incident response with cross-functional teams.

The role demands 5+ years in SRE leadership, strong CI/CD/DevSecOps experience, and hands-on cloud, Python, and AI tooling expertise.

Qualifications

  • 5+ years in SRE / application reliability leadership.
  • Experience leading 24x7 SRE operations and incident response.
  • Experience with CI/CD, DevSecOps workflows.

Responsibilities

  • Lead 24x7 SRE operations and coordinate P1/P2 incident response through restoration and follow-up.
  • Own reliability measures: SLIs/SLOs/SLAs and MTTR improvements.
  • Drive end-to-end observability using Dynatrace, Grafana, CloudWatch, and Splunk.
  • Lead automation, AIOps, self-healing, event correlation, and AI-driven operations initiatives.
  • Oversee AWS platform operations, batch processing, and related environments.
  • Integrate Claude AI on AWS Bedrock with GitLab via APIs and custom workflows.
  • Develop AI-driven analysis of GitLab project data, vulnerabilities, pipelines and security findings.
  • Design GitLab API automation, DevSecOps controls, and CI/CD pipeline improvements.
  • Build and maintain Python-based GitLab integrations and REST API solutions.
  • Support vulnerability remediation and onboarding/configuration of security scanning tools.
  • Manage AWS Lambda, ECR, and Bedrock for deployment and automation; optimize Lambda settings.
  • Design and support resilient multi-region AWS architectures and containerized deployments with Docker.

Skills

SRE/ARE
Claude AI
GitLab access
Python/Shell scripting
AWS
Incident Management
SRE leadership

Tools

GitLab
Dynatrace
Grafana
CloudWatch
Splunk
AWS Bedrock
Control-M
Docker
ECR

Job description

Role: SRE Operations Lead - AWS, GitLab & AIOps
Location: Charlotte, NC - Onsite
Job Type: Fulltime
Job Description
Must Have Technical/Functional Skills:
  • SRE/ARE + 24x7 Production Support (CI/CD, DevSecOps)
  • Claude Code to be used for troubleshooting and some level of automation
  • Accessing repositories (Gitlab etc), awareness of branching strategies
  • Claude to be used for troubleshooting and some level of automation
  • Python - Shell scripting/ intermediate
  • AWS (CloudWatch, Route53, S3, CloudFront, ECR , EC2) (some, not all)
  • Incident Management + Automation/AIOps
  • 5+ years in SRE / application reliability leadership
Roles & Responsibilities:
  • Lead 24x7 SRE operations and coordinate P1/P2 major incident response through service restoration and follow-up.
  • Own reliability measures including SLIs, SLOs, SLAs, MTTR improvement, and service reliability KPIs.
  • Drive end-to-end observability using Dynatrace, Grafana, CloudWatch, and Splunk.
  • Lead automation, AIOps, self-healing, event correlation, and AI-driven operations initiatives.
  • Oversee AWS platform operations, batch processing, and Control-M environments.
  • Integrate Claude AI on AWS Bedrock with GitLab using APIs, PATs, and custom workflows.
  • Develop AI-driven analysis of GitLab project data, vulnerabilities, pipelines, and security findings.
  • Design GitLab API automation, custom workflows, DevSecOps controls, and CI/CD pipeline improvements.
  • Build and maintain Python-based GitLab integrations and REST API solutions.
  • Support vulnerability remediation and onboarding/configuration of security scanning tools.
  • Manage AWS Lambda, ECR, and Bedrock for deployment and automation; optimize Lambda configuration, concurrency, and scaling.
  • Design and support resilient multi-region AWS architectures and containerized deployments using Docker and Amazon ECR.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

SRE Operations Lead - AWS, GitLab & AIOps
SRE Operations Lead - AWS, GitLab & AIOps

Tata Consultancy Services • Charlotte (NC)

On-site
USD 70,000 - 120,000
SRE Operations Lead — AI-Driven Cloud & CI/CD
SRE Operations Lead — AI-Driven Cloud & CI/CD

SFE • Charlotte (NC)

On-site
USD 140,000 - 180,000
Sr SRE Automation Engineer
Sr SRE Automation Engineer

Compunnel, Inc. • Austin (TX), Northern (KY)

On-site
USD 130,000 - 180,000
Lead Site Reliability Engineer (SRE) / Principal Site Reliability Engineer (SRE)
Lead Site Reliability Engineer (SRE) / Principal Site Reliability Engineer (SRE)

Mindlance • Irving (TX)

On-site
USD 120,000 - 160,000
SRE Engineer
SRE Engineer

Programmers.io • Austin (TX)

Hybrid
USD 120,000 - 180,000
SRE Lead
SRE Lead

TechDigital Group • Woonsocket (RI)

On-site
USD 140,000 - 190,000
SRE Engineer
SRE Engineer

Tata Consultancy Services • Englewood Cliffs (NJ)

On-site
USD 110,000 - 125,000
Lead SRE
Lead SRE

JPMorgan Chase & Co. • Plano (TX)

On-site
USD 150,000 - 190,000
SRE AWS DevOps
SRE AWS DevOps

SFE • Malvern

On-site
USD 150,000 - 190,000
Site Reliability Engineer Lead
Site Reliability Engineer Lead

Good co India • United States

Remote
USD 120,000 - 160,000