Site Reliability Engineer

SCIGON

Naperville (IL)

Hybrid

USD 110,000 - 170,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

SCIGON is seeking a Site Reliability Engineer (SRE) – Release & Operations to bridge software development and technology operations in a hybrid role. You will manage the release lifecycle, drive change governance, and foster cross-team communication.

The role requires hands-on automation development to reduce operational overhead, production support, and a focus on cloud-based infrastructure reliability, security, and scalability.

Qualifications

  • Experience in Site Reliability Engineering (SRE), DevOps, cloud operations, or related roles.
  • Demonstrated experience managing software releases and coordinating deployments within change management processes.
  • Knowledge of cloud infrastructure platforms and services.
  • Experience building, deploying, and maintaining scalable cloud-based systems.
  • Solid understanding of Infrastructure-as-Code (IaC) principles and tools.
  • Strong scripting and automation skills using Python, Bash, PowerShell, or similar.
  • Experience supporting production systems and incident management.

Responsibilities

  • Lead end-to-end software release lifecycle including planning, scheduling, deployment, validation, and post-release activities.
  • Participate in change management processes to ensure governance and compliance prior to production deployment.
  • Serve as a primary contact for stakeholders regarding deployment schedules, status, risks, and rollback strategies.
  • Automate manual activities into repeatable CI/CD workflows and controls.
  • Provide hands-on production support, incident response, on-call participation, and root cause analyses.
  • Design and maintain IaC solutions for consistent deployments.
  • Configure and optimize monitoring, logging, and alerting to ensure application health.
  • Collaborate across engineering, platform, security, and operations teams to improve reliability and performance.

Skills

Strong communication
Systems-thinking
Detail-oriented
Analytical problem solving

Tools

Python
Bash
PowerShell
Shell scripting
Terraform
CloudFormation
Ansible
GitHub Actions
Jenkins
Grafana
ECS
EKS

Job description

The Site Reliability Engineer (SRE) – Release & Operations is a hybrid technical and operational role responsible for bridging software development and technology operations. This position focuses on managing and optimizing the software release lifecycle, driving change management governance, and facilitating communication across technical and business teams.

In addition to release engineering responsibilities, this role requires a hands-on technical professional who can develop automation solutions to reduce operational overhead, actively support production environments, and help ensure the availability, security, scalability, and reliability of cloud-based infrastructure and applications.

Responsibilities
Release Management & Change Governance
  • Lead and coordinate the end-to-end software release lifecycle, including planning, scheduling, staging, deployment, validation, and post-release activities.
  • Participate in and facilitate change management processes, evaluating release readiness, assessing risks, and ensuring governance and compliance requirements are met prior to production deployment.
  • Serve as a primary point of contact for engineering, quality assurance, product, operations, and business stakeholders regarding deployment schedules, release status, risk assessments, and rollback strategies.
  • Continuously improve release management processes by transitioning manual activities into automated, repeatable workflows and CI/CD controls.
  • Drive release standardization and promote best practices across development and operations teams.
Site Reliability & Production Operations
  • Provide hands-on production support, ensuring operational stability and participating in incident response and on-call support activities.
  • Design, develop, and maintain automation tools, scripts, and workflows to reduce operational effort and improve system reliability.
  • Respond to service interruptions, outages, and operational incidents, performing root cause analysis and implementing long-term corrective actions.
  • Implement and maintain Infrastructure-as-Code (IaC) solutions to ensure consistent, scalable, and repeatable infrastructure deployments.
  • Configure, manage, and optimize monitoring, logging, and alerting systems to provide visibility into application and infrastructure health.
  • Identify opportunities to improve system reliability, scalability, performance, and operational efficiency.
Security, Compliance & Collaboration
  • Ensure deployment and operational processes adhere to applicable security, governance, compliance, and risk management requirements.
  • Collaborate closely with software engineering, quality assurance, platform, infrastructure, and security teams to establish reliable deployment and operational practices.
  • Maintain accurate and audit-ready documentation, including operational procedures, deployment records, runbooks, incident reports, and change logs.
  • Support continuous improvement initiatives related to operational excellence, reliability engineering, and service delivery.
Required Skills & Experience
  • Experience in Site Reliability Engineering (SRE), DevOps, system administration, cloud operations, release engineering, or related technical roles.
  • Demonstrated experience managing software releases and coordinating deployments within structured change management processes.
  • Knowledge of cloud infrastructure platforms and services.
  • Experience building, deploying, and maintaining scalable cloud-based systems and applications.
  • Solid understanding of Infrastructure-as-Code (IaC) principles and tools.
  • Strong scripting and automation skills using languages such as Python, Bash, PowerShell, or similar technologies.
  • Experience supporting production systems and participating in incident management and root cause analysis activities.
  • Understanding of monitoring, observability, alerting, and operational support practices.
  • Strong communication and stakeholder management skills, with the ability to collaborate across technical and non-technical teams.
  • Ability to balance operational stability with delivery speed and business objectives.
Preferred Qualifications
  • Experience with change management, IT service management, or governance frameworks.
  • Professional certifications related to cloud platforms, DevOps, Site Reliability Engineering, infrastructure, or IT service management.
  • Experience designing and maintaining CI/CD pipelines and deployment automation frameworks.
  • Familiarity with modern monitoring, observability, and application performance management solutions.
  • Experience with containerization and container orchestration technologies.
  • Exposure to highly regulated or compliance-driven environments.
  • Experience supporting audit activities, change controls, and operational governance processes.
  • Familiarity with cloud-native architectures and distributed systems.
Technology Environment

Experience with technologies similar to the following is beneficial:

  • AWS, Azure, Google Cloud Platform (GCP), or comparable cloud providers
Automation & Scripting
  • Python
  • Bash
  • PowerShell
  • Shell scripting and automation frameworks
Infrastructure as Code
  • Terraform
  • CloudFormation
  • Ansible
CI/CD & Release Automation
  • GitHub Actions
  • Jenkins
  • Other modern deployment automation platforms
Monitoring & Observability
  • Grafana
  • Cloud-native monitoring solutions
  • Log aggregation and alerting platforms
Containers & Orchestration
  • Docker
  • ECS
  • EKS
  • Other container orchestration platforms
  • Strong collaboration and communication skills with the ability to align technical and business stakeholders.
  • Systems-thinking mindset with a focus on solving root causes through automation and process improvement.
  • Detail-oriented approach to change management, risk mitigation, and operational excellence.
  • Strong analytical and troubleshooting skills in complex production environments.
  • Ability to identify opportunities to improve reliability, scalability, and operational efficiency.
  • Commitment to continuous learning and adoption of modern cloud, DevOps, and Site Reliability Engineering practices.
  • Proactive ownership mentality with the ability to operate effectively in fast-paced and evolving environments.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Cloud Solutions Engineer
Cloud Solutions Engineer

Tyler Technologies • Lakewood (CO)

On-site
USD 120,000 - 150,000
Site Reliability Engineer – Lead
Site Reliability Engineer – Lead

Jobtailor • Arizona

On-site
USD 140,000 - 230,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Virtual Tech Gurus • Puerto Rico

On-site
USD 140,000 - 210,000
Senior SRE Engineer
Senior SRE Engineer

Compunnel, Inc. • Alpharetta (GA)

On-site
USD 140,000 - 190,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Veriipro • Washington

On-site
USD 120,000 - 180,000
Cloud Solutions Engineer
Cloud Solutions Engineer

Tyler Technologies, Inc. • Plano (TX), Latham (NY), Lubbock (TX), Lakewood (CO)

On-site
USD 93,547 - 150,000
Cloud Solutions Engineer
Cloud Solutions Engineer

Tyler-Technologies-29572f8 • Lakewood (CO)

On-site
USD 93,547 - 150,000
Site Reliability Engineer
Site Reliability Engineer

TalentDome Staffing • United States

On-site
USD 140,000 - 210,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Veloc Inc • Coppell (TX)

On-site
USD 140,000 - 190,000
Site Reliability Engineer
Site Reliability Engineer

SRE • Puerto Rico

Hybrid
USD 120,000 - 180,000