AWS DevOps / SRE Engineer

TechDigital Group

Eagan (MN)

On-site

USD 90,000 - 130,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

TechDigital Group in Minnesota is seeking an AWS Cloud Operations / Infrastructure Support Engineer to monitor, operate, and support AWS cloud infrastructure across production and non-production environments. You will perform health checks, monitor with CloudWatch and Dynatrace, investigate incidents, validate backups, manage EBS snapshots, and coordinate logs and access within the LZ2 landing zone.

You will configure monitoring, respond to alerts, assist with deployments using runbooks, and

Qualifications

  • 5+ years of IT infrastructure, cloud operations, production support, or systems engineering experience.
  • 3+ years of hands-on AWS operations experience preferred.
  • Experience with monitoring, incident response, backup validation, and runbooks.

Responsibilities

  • Perform daily AWS cloud infrastructure health checks, including disk space validation, SSM agent status, EC2/RDS availability, and operational readiness checks.
  • Configure, manage, and enhance cloud infrastructure monitoring using AWS CloudWatch and Dynatrace, including dashboards, metrics, alarms, and notification workflows.
  • Investigate monitoring alerts from CloudWatch, Dynatrace, and related tools for outages, performance degradation, system issues, and infrastructure anomalies.
  • Configure CloudWatch monitoring and alert notifications to automatically notify the on-call team for critical events and production-impacting incidents.
  • Manage production EBS snapshot activities, EBS disk sizing requests, backup validation for EC2 and RDS instances, and associated operational reporting.
  • Manage VPC Flow Logs, EC2 logs, log bucket access, S3 bucket access policies, and AWS account access within the LZ2 landing zone environment.
  • Troubleshoot cloud infrastructure service outages, EC2/container restarts, network issues, DNS private hosted zones, firewall change requests, and access-related issues.
  • Support AWS account creation, AMI and S3 object sharing across AWS accounts, EC2 instance restarts, container restarts, SSL certificate requests, and infrastructure resizing guided by FME.
  • Support application deployments by executing approved deployment runbooks, work instructions, validation steps, and weekend service downtime communications.
  • Monitor Subversion code repositories, maintain infrastructure documentation, and support TriZetto Facets suite installation activities.

Skills

AWS operations
Cloud monitoring
Production support

Tools

Dynatrace
CloudWatch
SSM

Job description

Top Skills Required
  • AWS Cloud Operations (EC2, EBS, RDS, S3, VPC, Route 53, SSM, IAM)
  • Cloud Monitoring & Observability (CloudWatch, Dynatrace, Incident Management)
  • Production Support & Infrastructure Troubleshooting
Job Description/Responsibilities

AWS Cloud Operations / Infrastructure Support Engineer to monitor, operate, and support AWS cloud infrastructure services across production and non-production environments. The role is responsible for daily health checks, CloudWatch and Dynatrace monitoring, incident investigation, EC2/RDS backup validation, EBS snapshot and disk management, VPC and EC2 log operations, DNS/firewall coordination, deployment runbook execution, documentation, and TriZetto Facets suite installation support.

Key Responsibilities
  • Perform daily AWS cloud infrastructure health checks, including disk space validation, SSM agent status, EC2/RDS availability, and operational readiness checks.
  • Configure, manage, and enhance cloud infrastructure monitoring using AWS CloudWatch and Dynatrace, including dashboards, metrics, alarms, and notification workflows.
  • Investigate monitoring alerts from CloudWatch, Dynatrace, and related tools for outages, performance degradation, system issues, and infrastructure anomalies.
  • Configure CloudWatch monitoring and alert notifications to automatically notify the on-call team for critical events and production-impacting incidents.
  • Manage production EBS snapshot activities, EBS disk sizing requests, backup validation for EC2 and RDS instances, and associated operational reporting.
  • Manage VPC Flow Logs, EC2 logs, log bucket access, S3 bucket access policies, and AWS account access within the LZ2 landing zone environment.
  • Troubleshoot cloud infrastructure service outages, EC2/container restarts, network issues, DNS private hosted zones, firewall change requests, and access-related issues.
  • Support AWS account creation, AMI and S3 object sharing across AWS accounts, EC2 instance restarts, container restarts, SSL certificate requests, and infrastructure resizing guided by FME.
  • Support application deployments by executing approved deployment runbooks, work instructions, validation steps, and weekend service downtime communications.
  • Monitor Subversion code repositories, maintain infrastructure documentation, and support TriZetto Facets suite installation activities.
Required Technical Skills
  • Hands-on AWS operations experience across EC2, EBS, RDS, S3, VPC, Route 53 Private Hosted Zones, CloudWatch, SSM, AMIs, IAM/access, and AWS account operations.
  • Strong monitoring and observability experience using AWS CloudWatch and Dynatrace, including alert investigation, dashboard creation, one-agent installation, and alert routing.
  • Experience with production support, incident triage, outage investigation, infrastructure troubleshooting, log analysis, backup validation, and operational reporting.
  • Working knowledge of VPC Flow Logs, EC2 logs, S3/log bucket access, access policies, DNS, firewall change coordination, SSL certificate requests, and cross-account sharing.
  • Ability to follow approved runbooks and work instructions for application deployment support, infrastructure resizing, instance/container restarts, and weekend communication support.
  • Familiarity with Subversion repositories, infrastructure documentation practices, and healthcare/payer application platforms such as TriZetto Facets is preferred.
Qualifications
  • 5+ years of IT infrastructure, cloud operations, production support, or systems engineering experience; 3+ years of hands‑on AWS operations experience preferred.
  • Experience supporting production AWS workloads, monitoring tools, incident response processes, backup validation, and operational runbooks in enterprise environments.
  • Strong communication, documentation, ownership, troubleshooting, collaboration, and weekend/on‑call support readiness.
  • Preferred certifications: AWS SysOps Administrator – Associate, AWS Solutions Architect – Associate, AWS Cloud Practitioner, Dynatrace Associate/Professional, ITIL Foundation.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AWS Devops Engineer
AWS Devops Engineer

Compunnel, Inc. • Charlotte (NC)

On-site
USD 120,000 - 180,000
AWS Cloud Operations & Platform Readiness Engineer
AWS Cloud Operations & Platform Readiness Engineer

Apexon • New Jersey

On-site
USD 120,000 - 160,000
Senior AWS Devops Engineer
Senior AWS Devops Engineer

RxCloud • Oregon (WI)

Hybrid
USD 120,000 - 150,000
AWS DevOps / SRE Engineer
AWS DevOps / SRE Engineer

TechDigital Group • Norfolk (VA)

On-site
USD 140,000 - 165,000
AWS DevOps Engineer
AWS DevOps Engineer

koundinyasa Technology Services • Indiana (PA)

On-site
USD 110,000 - 160,000
AWS Devops Engineer
AWS Devops Engineer

Compunnel, Inc. • Atlanta (GA)

On-site
USD 100,000 - 130,000
AWS DevOps / SRE Engineer — Cloud Reliability & Incidents
AWS DevOps / SRE Engineer — Cloud Reliability & Incidents

TechDigital Group • Eagan (MN)

On-site
USD 90,000 - 130,000
SRE Engineer
SRE Engineer

Blue Ribbon Global Technologies • Atlanta (GA)

On-site
USD 90,000 - 120,000
Operations Engineer (Cloud)
Operations Engineer (Cloud)

Stifel Financial Corp. • Memphis (TN)

On-site
USD 90,000 - 120,000
SRE Engineer
SRE Engineer

SOMERSET STAFFING • Atlanta (GA)

On-site
USD 110,000 - 170,000