Turn this role into an interview — a resume and cover letter built around what this employer wants.
SFE is seeking an SRE AWS DevOps Lead (MUST HAVE AWS expertise) to drive building and maintaining scalable, secure AWS cloud platforms from Malvern, PA. This role emphasises reliability, IaC, and modern CI/CD practices across data, ML, and production teams.
The ideal candidate has 8+ years in SRE/DevOps, strong experience with Kubernetes, Terraform, CI/CD tooling, and AI observability. You will lead cross-functional incident response and ensure high availability of critical AI workloads.
Role: SRE AWS DevOps (LEAD AWS IS MUST )
Location: Malvern PA
Term: Full Time
For an SRE / AWS DevOps Engineer with Arize Observability, the candidate should have knowledge on AWS Cloud, DevOps, Reliability Engineering, Monitoring, MLOps, and AI Observability.
AWS Cloud Services EC2, EKS, ECS, Lambda, S3, RDS, Redshift, CloudWatch, IAM, VPC.
Infrastructure as Code (IaC) Terraform, AWS CloudFormation, Ansible.
CI/CD Automation Jenkins, GitHub Actions, GitLab CI/CD, AWS CodePipeline.
Containerization & Orchestration Docker, Kubernetes (EKS), Helm.
Site Reliability Engineering (SRE) SLI/SLO/SLA management, incident response, root cause analysis (RCA), reliability engineering.
Monitoring & Observability Arize AI, Prometheus, Grafana, Datadog, ELK Stack, OpenTelemetry, CloudWatch.
MLOps & AI Observability Arize platform, model monitoring, drift detection, model performance tracking, data quality monitoring, LLM observability.
Programming & Scripting Python, Bash, PowerShell, SQL.
Security & DevSecOps IAM, Secrets Manager, AWS Security Hub, vulnerability scanning, policy enforcement.
Collaboration & Agile Delivery Scrum, Jira, stakeholder communication, cross-functional incident management, technical documentation.
SRE Lead with 8+ years of experience in building and managing scalable, secure, and highly available AWS cloud platforms leveraging DevOps, Kubernetes, and Infrastructure-as-Code practices.
Experienced in implementing CI/CD pipelines, driving SRE best practices, and ensuring platform reliability through proactive monitoring, incident management, and performance optimization using CloudWatch, Prometheus, Grafana, and Arize.
Strong collaborator with engineering, data, and ML teams, enabling MLOps, AI model observability, drift detection, and reliable deployment of production-grade AI/GenAI solutions.
Generic Managerial Skills, If any
Strong leadership and stakeholder management skills, with experience leading cross-functional teams and driving delivery excellence.
Effective communication with business