SRE AWS DevOps

SFE

Malvern (Chester County)

On-site

USD 150,000 - 190,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

SFE is seeking an SRE AWS DevOps Lead (MUST HAVE AWS expertise) to drive building and maintaining scalable, secure AWS cloud platforms from Malvern, PA. This role emphasises reliability, IaC, and modern CI/CD practices across data, ML, and production teams.

The ideal candidate has 8+ years in SRE/DevOps, strong experience with Kubernetes, Terraform, CI/CD tooling, and AI observability. You will lead cross-functional incident response and ensure high availability of critical AI workloads.

Qualifications

  • Knowledge of AWS Cloud, DevOps, Reliability Engineering, Monitoring, MLOps, and AI Observability.

Responsibilities

  • SRE Lead with 8+ years of experience in building and managing scalable, secure, and highly available AWS cloud platforms leveraging DevOps, Kubernetes, and IaC practices.
  • Experience in implementing CI/CD pipelines, driving SRE best practices, and ensuring platform reliability through proactive monitoring, incident management, and performance optimization using CloudWatch, Prometheus, Grafana, and Arize.
  • Strong collaboration with engineering, data, and ML teams, enabling MLOps, AI model observability, drift detection, and reliable deployment of production-grade AI/GenAI solutions.
  • Generic Managerial Skills, If any
  • Strong leadership and stakeholder management skills, with experience leading cross-functional teams and driving delivery excellence.
  • Effective communication with business

Skills

AWS
DevOps
Reliability Engineering
Monitoring
MLOps
AI Observability
Python
Bash
PowerShell
SQL
Security
DevSecOps
Kubernetes
Docker
Helm
Terraform
AWS CloudFormation
Ansible
Jenkins
GitHub Actions
GitLab CI/CD
AWS CodePipeline
Prometheus
Grafana
Datadog
ELK Stack
OpenTelemetry
CloudWatch
Incident Management
Root Cause Analysis
Communication

Tools

Terraform
AWS CloudFormation
Ansible
Jenkins
GitHub Actions
GitLab CI/CD
AWS CodePipeline
Docker
Kubernetes
Helm
Prometheus
Grafana
Datadog
OpenTelemetry
CloudWatch
Terraform

Job description

Role: SRE AWS DevOps (LEAD AWS IS MUST )
Location: Malvern PA
Term: Full Time

Job Description
Must Have Technical/Functional Skills

For an SRE / AWS DevOps Engineer with Arize Observability, the candidate should have knowledge on AWS Cloud, DevOps, Reliability Engineering, Monitoring, MLOps, and AI Observability.

AWS Cloud Services EC2, EKS, ECS, Lambda, S3, RDS, Redshift, CloudWatch, IAM, VPC.

Infrastructure as Code (IaC) Terraform, AWS CloudFormation, Ansible.

CI/CD Automation Jenkins, GitHub Actions, GitLab CI/CD, AWS CodePipeline.

Containerization & Orchestration Docker, Kubernetes (EKS), Helm.

Site Reliability Engineering (SRE) SLI/SLO/SLA management, incident response, root cause analysis (RCA), reliability engineering.

Monitoring & Observability Arize AI, Prometheus, Grafana, Datadog, ELK Stack, OpenTelemetry, CloudWatch.

MLOps & AI Observability Arize platform, model monitoring, drift detection, model performance tracking, data quality monitoring, LLM observability.

Programming & Scripting Python, Bash, PowerShell, SQL.

Security & DevSecOps IAM, Secrets Manager, AWS Security Hub, vulnerability scanning, policy enforcement.

Collaboration & Agile Delivery Scrum, Jira, stakeholder communication, cross-functional incident management, technical documentation.

Roles & Responsibilities

SRE Lead with 8+ years of experience in building and managing scalable, secure, and highly available AWS cloud platforms leveraging DevOps, Kubernetes, and Infrastructure-as-Code practices.

Experienced in implementing CI/CD pipelines, driving SRE best practices, and ensuring platform reliability through proactive monitoring, incident management, and performance optimization using CloudWatch, Prometheus, Grafana, and Arize.

Strong collaborator with engineering, data, and ML teams, enabling MLOps, AI model observability, drift detection, and reliable deployment of production-grade AI/GenAI solutions.

Generic Managerial Skills, If any

Strong leadership and stakeholder management skills, with experience leading cross-functional teams and driving delivery excellence.

Effective communication with business

Get your free, confidential resume review.

or drag and drop your file here.