Site Reliability Engineer

Resolve Tech Solutions

Irving (TX)

On-site

USD 120,000 - 160,000

Full time

3 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Resolve Tech Solutions is seeking a senior SRE/DevOps professional to lead incident response, triage, and rapid resolution for our production systems in Irving, TX. You will drive RCA, implement long-term corrective actions, and maintain runbooks and escalation paths.

You will optimize AWS services, monitor health across cloud environments, and collaborate with engineering, QA, and product teams to prevent recurrence of issues.

Qualifications

  • Minimum 5+ years in SRE, DevOps, Production Support, or similar operational roles.

Responsibilities

  • Serve as a primary responder for production incidents, ensuring rapid triage, mitigation, and resolution.

Skills

Incident response
RCA
On-call rotations
SRE
DevOps
Production Support
Cross-functional collaboration
SLOs
Linux
Networking
Distributed systems
Scripting

Education

AWS Solutions Architect
AWS SysOps
MongoDB DBA

Tools

AWS EC2
AWS ECS
AWS EKS
AWS Lambda
AWS S3
CloudWatch
IAM
RDS
VPC networking
MongoDB
New Relic
Postman
Intune
Firebase
ServiceNow
Jenkins
GitHub Actions
GitLab CI
Kubernetes
Kafka
SQS
RabbitMQ
Docker
Linux

Job description

“MUST HAVE” SPECIFIC KNOWLEDGE AND SKILLS
  • Serve as a primary responder for production incidents, ensuring rapid triage, mitigation, and resolution.
  • Lead root cause analysis (RCA) and drive long?term corrective actions.
  • Maintain and improve incident response processes, runbooks, and escalation paths.
  • Collaborate with engineering, QA, and product teams to prevent recurrence of issues.
  • AWS Infrastructure OperationsSupport and optimize AWS services such as EC2, ECS/EKS, Lambda, S3, CloudWatch, IAM, RDS, and VPC networking.
  • Monitor system health, performance, and capacity across cloud environments.
  • Implement infrastructure best practices around reliability, scalability, and cost efficiency.
  • Assist with deployments, environment configuration, and CI/CD pipelines.
  • Database & Storage SupportManage and troubleshoot MongoDB clusters, including performance tuning, replication, backups, and failover.
  • Diagnose query performance issues and collaborate with developers on schema optimization.
  • Ensure data integrity, availability, and recovery readiness.
  • Monitoring, Observability & AlertingUse New Relic, CloudWatch, and other observability tools to monitor application and infrastructure performance.
  • Build dashboards, alerts, and telemetry that provide actionable insights.
  • Continuously refine monitoring thresholds to reduce noise and improve signal quality.
  • Experience with on-call rotations and 24/7 production environments.
  • Work cross-functionally with the various teams in the organization and help establish SLOs and achieve those SLOs.
  • 5+ years of experience in SRE, DevOps, Production Support, or similar operational roles.
  • Strong hands-on experience with AWS services and cloud-native architectures.
  • Proficiency with MongoDB administration and troubleshooting.
  • Experience with New Relic or similar APM/observability platforms.
  • Experience using additional tools like Postman, Intune, and Firebase, Service Now, Cloudwatch.
  • Strong understanding of Linux systems, networking, and distributed systems.
  • Solid scripting skills (Python, Bash, or similar).
  • 5+ years Monitoring and Alarming in all environments and familiar with tools like Mongo Charts, New Relic, Cloudwatch, Service Now.
  • Proven experience managing high-severity incidents and driving RCA processes.
  • Familiarity with CI/CD tools (Jenkins, GitHub Actions, GitLab CI, etc.).
ADDITIONAL SKILLS AND OTHER REQUIREMENTS
  • Experience with container orchestration (ECS, EKS, Kubernetes).
  • Knowledge of message queues (Kafka, SQS, RabbitMQ).
  • Exposure to microservices architectures.
  • Certifications such as AWS Solutions Architect, AWS SysOps, or MongoDB DBA.
  • Working experience with IoT devices, and Microsoft Intune.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Veriipro • Atlanta (GA)

On-site
USD 110,000 - 160,000
Site Reliability Engineer
Site Reliability Engineer

TalentDome Staffing • United States

On-site
USD 140,000 - 210,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

MeridianLink, Inc. • Northern (KY)

Hybrid
USD 120,000 - 170,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Apply • Northern (KY)

Hybrid
USD 120,000 - 150,000
Site Reliability Engineer/Production Support
Site Reliability Engineer/Production Support

Smart IT Frame LLC • New York (NY)

On-site
USD 120,000 - 160,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

MeridianLink • United States

Remote
USD 140,000 - 180,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

The ReWork Group • New York (NY)

On-site
USD 120,000 - 160,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Knack Solutions • Reston (VA)

On-site
USD 120,000 - 160,000
Site Reliability Engineer
Site Reliability Engineer

Evlo AI • San Francisco (CA)

On-site
USD 140,000 - 200,000
Site Reliability Engineer
Site Reliability Engineer

Compunnel, Inc. • Greenwood Village (CO)

On-site
USD 120,000 - 150,000