AWS SRE – AIOps & Full Stack Engineer

Savvyan Technologies

Columbus (OH)

On-site

USD 140,000 - 190,000

Full time

22 hours ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Savvyan Technologies seeks an AWS Site Reliability Engineer (SRE) with strong AIOps and full-stack development experience to design, build, automate, and operate highly available cloud applications and platforms.

The role spans cloud infrastructure, CI/CD, observability, AI-driven automation, incident management, and performance optimization across backend and frontend services. You will collaborate across teams to implement scalable solutions and leverage AI/ML for proactive remediation.

Qualifications

  • Bachelor's degree in Computer Science, Engineering, Information Technology, or related field.
  • 7+ years of experience in software engineering, DevOps, Cloud Engineering, or SRE.
  • Hands-on AWS experience with production systems and cloud infrastructure.

Responsibilities

  • Design, deploy, and maintain highly available apps and infrastructure on AWS.
  • Implement SRE principles including SLOs, SLAs, and capacity planning.
  • Develop automation to reduce manual operational tasks and implement IaC.
  • Build and maintain CI/CD pipelines (Jenkins, GitHub Actions, GitLab CI, CodePipeline).
  • Implement AIOps and AI-driven operations for proactive remediation and observability.
  • Establish dashboards, metrics, logs, and tracing for cloud and applications.

Skills

AWS
SRE
DevOps
CI/CD
Terraform
Docker
Kubernetes
OpenTelemetry
Datadog
Python

Education

Bachelor's degree in Computer Science or related field

Tools

Jenkins
GitHub Actions
GitLab CI
AWS CodePipeline
CloudFormation
AWS CDK

Job description

We are seeking a highly skilled AWS Site Reliability Engineer (SRE) with strong AIOps and Full Stack development experience to design, build, automate, and operate highly available cloud applications and platforms.

The ideal candidate will have a strong background in AWS infrastructure, Site Reliability Engineering, DevOps, observability, AIOps, automation, and full-stack application development. This role requires someone who can work across the entire technology stack—from cloud infrastructure and CI/CD pipelines to backend services, APIs, frontend applications, monitoring, and AI-driven operational automation.

The engineer will be responsible for improving system reliability, scalability, performance, and operational efficiency while leveraging AI/ML and AIOps capabilities to proactively identify incidents, predict failures, automate remediation, and improve application observability.

Key Responsibilities
AWS & Site Reliability Engineering
  • Design, deploy, and maintain highly available, scalable, and fault-tolerant applications and infrastructure on AWS.
  • Implement SRE principles including SLIs, SLOs, SLAs, error budgets, availability, reliability, and capacity planning.
  • Manage AWS services including EC2, EKS, ECS, Lambda, S3, RDS, DynamoDB, API Gateway, CloudFront, Route 53, IAM, VPC, CloudWatch, SNS, SQS, and EventBridge.
  • Troubleshoot complex production issues involving application, infrastructure, networking, database, and cloud components.
  • Participate in incident management, root-cause analysis, problem management, and post-incident reviews.
  • Develop automation to reduce manual operational activities and eliminate repetitive tasks.
  • Perform capacity planning, performance tuning, disaster recovery, and business continuity activities.
DevOps & Infrastructure Automation
  • Build and maintain CI/CD pipelines using tools such as Jenkins, GitHub Actions, GitLab CI/CD, AWS CodePipeline, or Azure DevOps.
  • Implement Infrastructure as Code using Terraform, CloudFormation, or AWS CDK.
  • Automate infrastructure provisioning, configuration management, deployments, and operational processes.
  • Implement containerized workloads using Docker and Kubernetes/Amazon EKS.
  • Develop automated deployment strategies including blue-green, canary, and rolling deployments.
  • Integrate security, compliance, testing, and quality checks into CI/CD pipelines.
AIOps & AI-Driven Operations
  • Implement AIOps solutions to improve monitoring, incident detection, event correlation, root-cause analysis, and automated remediation.
  • Leverage AI/ML and Generative AI capabilities to analyze logs, metrics, traces, alerts, and operational data.
  • Develop intelligent alerting and anomaly-detection mechanisms to identify potential production issues before they impact customers.
  • Build AI-assisted incident investigation and troubleshooting workflows.
  • Integrate LLM/GenAI capabilities into SRE and DevOps workflows for automated log analysis, incident summarization, knowledge retrieval, and remediation recommendations.
  • Develop automated runbooks and self-healing mechanisms using event-driven AWS services and AI-assisted decision making.
  • Integrate AIOps platforms and observability tools such as Datadog, Dynatrace, New Relic, Splunk, CloudWatch, Grafana, and Prometheus.
  • Develop or integrate AI agents/workflows that can assist with incident response, operational diagnostics, and infrastructure management.
  • Monitor AIOps/AI solutions for accuracy, reliability, security, and operational effectiveness.
Observability & Monitoring
  • Implement comprehensive metrics, logs, traces, dashboards, and alerting across cloud and application environments.
  • Work with Prometheus, Grafana, CloudWatch, OpenTelemetry, Datadog, Dynatrace, Splunk, or similar observability platforms.
  • Establish meaningful service-level indicators and operational dashboards.
  • Implement distributed tracing and application performance monitoring.
  • Tune alerts to reduce false positives and alert fatigue.
  • Build proactive monitoring and predictive health checks.
Full Stack Development
  • Develop and maintain scalable backend services, APIs, and web applications.
  • Build RESTful APIs and microservices using technologies such as Java/Spring Boot, Python/FastAPI, Node.js, or similar.
  • Develop responsive frontend applications using React, Angular, TypeScript, JavaScript, HTML, and CSS.
  • Integrate frontend applications with REST/GraphQL APIs and cloud-native backend services.
  • Design and optimize database interactions using PostgreSQL, MySQL, MongoDB, DynamoDB, or similar databases.
  • Implement authentication and authorization using OAuth 2.0, OpenID Connect, JWT, AWS IAM, or similar technologies.
  • Develop automated unit, integration, API, and end-to-end tests.
  • Troubleshoot application performance and scalability issues across frontend, backend, database, and infrastructure layers.
Required Technical Skills
Cloud
  • AWS
  • EC2, S3, VPC, IAM, RDS, DynamoDB
  • EKS/ECS, Lambda
  • CloudWatch, Route 53, API Gateway
  • SQS, SNS, EventBridge
  • AWS networking and security
SRE / DevOps
  • Site Reliability Engineering principles
  • Incident Management & Root Cause Analysis
  • CI/CD
  • Terraform / CloudFormation / AWS CDK
  • Docker
  • Kubernetes / EKS
  • Jenkins / GitHub Actions / GitLab CI
  • Linux administration and troubleshooting
  • Bash/Shell scripting
  • Python
AIOps / AI
  • AIOps and intelligent event management
  • Generative AI / LLM concepts
  • AI-assisted incident management
  • Anomaly detection and predictive monitoring
  • Automated remediation / self-healing systems
  • LLM APIs and AI agent workflows
  • RAG/vector database concepts are a plus
  • Experience integrating AI with DevOps/SRE workflows
Observability
  • Datadog / Dynatrace / New Relic
  • Prometheus
  • Grafana
  • Splunk
  • AWS CloudWatch
  • OpenTelemetry
  • Application Performance Monitoring
Full Stack
  • React / Angular
  • JavaScript / TypeScript
  • HTML / CSS
  • Node.js / Python / Java
  • REST APIs / GraphQL
  • Microservices
  • SQL / NoSQL databases
Preferred Qualifications
  • Bachelor's degree in Computer Science, Engineering, Information Technology, or a related field.
  • 7+ years of experience in software engineering, DevOps, Cloud Engineering, or SRE.
  • 4+ years of hands-on AWS experience.
  • Strong experience supporting production environments and distributed systems.
  • Experience implementing AIOps or AI-powered operational solutions.
  • Experience with Kubernetes and cloud-native architectures.
  • Experience developing full-stack applications.
  • Strong programming and scripting skills in Python, Java, Node.js, or similar languages.
  • Experience working in Agile/Scrum environments.
  • Strong troubleshooting, analytical, communication, and problem-solving skills.
Key Competencies
  • Cloud-native architecture
  • Site Reliability Engineering
  • AIOps & Generative AI
  • Infrastructure automation
  • Full-stack development
  • Observability & monitoring
  • Production support
  • Incident response
  • Automation & self-healing
  • Performance optimization
  • Security and reliability
  • Continuous improvement
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Sr SRE Automation Engineer
Sr SRE Automation Engineer

Compunnel, Inc. • Austin (TX), Northern (KY)

On-site
USD 130,000 - 180,000
Full Stack Engineer
Full Stack Engineer

Quadrant IQ Solutions LLC • New Jersey

On-site
USD 140,000 - 190,000
Site Reliability Engineer
Site Reliability Engineer

TalentDome Staffing • United States

On-site
USD 140,000 - 210,000
Sr. Site Reliability Engineer(Local to Atlanta GA Only)
Sr. Site Reliability Engineer(Local to Atlanta GA Only)

Trigint Solutions LLC • Atlanta (GA)

Hybrid
USD 124,000 - 220,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Spectraforce Technologies • Austin (TX)

Hybrid
USD 130,000 - 170,000
Site Reliability Engineer
Site Reliability Engineer

Compunnel, Inc. • Greenwood Village (CO)

On-site
USD 120,000 - 150,000
Senior Lead Site Reliability Engineer
Senior Lead Site Reliability Engineer

JPMorgan Chase & Co. • Jersey City (NJ)

On-site
USD 150,000 - 210,000
Senior Forward Deployed Engineer (DevOps/SRE)
Senior Forward Deployed Engineer (DevOps/SRE)

LeoForce • Pleasanton (CA)

On-site
USD 300,000 - 350,000
Medical benefits
401(k) plan
Free meals and snacks
+2
Site Reliability Engineer
Site Reliability Engineer

Harvey Nash • United States

Remote
USD 120,000 - 150,000
Lead Site Reliability Engineer (SRE) / Principal Site Reliability Engineer (SRE)
Lead Site Reliability Engineer (SRE) / Principal Site Reliability Engineer (SRE)

Mindlance • Irving (TX)

On-site
USD 120,000 - 160,000