DevOps / Site Reliability Engineer

AgileEngine, LLC

Brazil, Northern (IN, KY)

Hybrid

USD 120,000 - 160,000

Full time

6 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Remote work 100%
Flexible hours
Annual learning budget

Job summary

AgileEngine, LLC seeks a DevOps / Site Reliability Engineer to maintain 24/7 stability for a multi-cloud enterprise security program. You will act as Incident Commander for major incidents, own IaC, CI/CD pipelines, and CSPM telemetry using Terraform and Wiz. Requires 5+ years of SRE experience in a 24/7 financial services environment.

You will drive major-incident calls, coordinate cross-functional responders, and own post-incident remediation and playbooks to prevent reoccurrence.

Qualifications

  • 5+ years of SRE experience in a 24/7 production environment.
  • Hands-on incident command during major/critical incidents.
  • Strong IaC practice and multi-cloud exposure.
  • Experience with Kubernetes, Terraform, CI/CD tools, and scripting (Python/Go).
  • Ability to draft post-incident communications and runbooks.

Responsibilities

  • Scale and maintain operational stability across multi-cloud environments (Azure, AWS, GCP).
  • Engineer IaC and baselines to prevent misconfigurations using Terraform/Wiz.
  • Design, maintain, and optimize CI/CD pipelines for deployment and security data ingestion.
  • Lead incident bridges during major incidents and coordinate cross-functional responders.
  • Draft and circulate incident notifications and status updates to all stakeholders.
  • Develop divisional incident-management playbooks and escalation procedures.

Skills

Kubernetes
Terraform
CI/CD orchestration
Python/Go scripting
Incident command

Tools

Wiz
PagerDuty
ServiceNow

Job description

We are looking for a DevOps / Site Reliability Engineer to maintain operational resilience and 24/7 stability for a multi-cloud enterprise security program, serving as Incident Commander on major and critical incidents while also owning IaC, CI/CD pipelines, and CSPM telemetry. You will drive major-incident calls, own post-incident remediation follow-through, draft stakeholder communications, and develop divisional incident-management playbooks alongside multi-cloud security guardrails using Terraform and Wiz. The role requires 5+ years of SRE experience with hands-on incident command in a 24/7 financial services environment.

What you will do
  • Scale and maintain the ability to drive operational stability across multi-cloud environments (Azure, AWS, GCP).
  • Engineer unified security policies and configuration baselines using IaC (Terraform) to prevent misconfigurations.
  • Design, maintain, and optimize enterprise CI/CD pipelines to support continuous ASPM ingestion and deployment.
  • Act on continuous monitoring alerts, utilizing Cloud Security Posture Management (CSPM) tools like Wiz to secure workloads.
  • Serve as Incident Commander on major and critical incidents - running the bridge, directing technical workstreams, making time-critical decisions, and coordinating cross-functional responders under pressure.
  • Own the post-incident loop - track remediation items to closure, hold owning teams accountable to timelines, and drive systemic fixes and preventative actions across groups.
  • Draft and send clear, accurate, audience-appropriate incident notifications and status updates to technical teams, management, and stakeholders throughout the incident lifecycle.
  • Develop, maintain, and socialize divisional / group-level incident-management playbooks, runbooks, and escalation procedures that standardize response and reduce time-to-resolution.
Must haves
  • 5+ years of experience.
  • In-depth architectural expertise in multi-cloud defense, federated IAM, and zero-trust principles.
  • Strong practical experience with Kubernetes, Terraform, CI/CD orchestration, and Python/Go scripting.
  • Senior-level, hands-on incident-command experience driving major/critical incident calls to resolution in a 24x7 production environment.
  • Proven track record of remediation follow-up - coordinating with teams and holding owners accountable until issues are fully closed.
  • Demonstrated skill drafting and issuing incident notification communications to both technical and executive audiences.
  • Direct experience authoring divisional/group incident-management playbooks and escalation procedures.
  • Fully autonomous.
  • Drives the architecture of complex automated runbooks and mentors Middle-level SREs.
  • Extensive experience deploying and tuning APIs from modern CNAPP/CSPM platforms, ideally Wiz.
  • Prior experience building platforms subject to strict financial compliance standards (PCI-DSS, SOC2).
Nice to haves
  • PagerDuty - hands-on experience with on-call scheduling, alert routing, and incident orchestration.
  • ServiceNow - familiarity with incident, problem, and change management workflows and reporting.
  • Growth without limits: build your skills through mentorship, internal TechTalks, challenging projects, and a dedicated annual learning budget
  • Competitive compensation: get recognition that reflects your skills and impact, with regular performance and compensation reviews
  • Flexibility: work 100% remotely with flexible hours that support focus, autonomy, and a healthy work rhythm
  • Meaningful, modern projects: build impactful products using modern technologies alongside global teams and leading brands
  • Collaborative culture: join a supportive environment with zero micromanagement where ideas are welcomed and contributions are recognized
  • Well-being & support: access local well-being programs and people-focused support tailored to your location
FAQ
Have any questions?
What is the work format andschedule?
What level of English isrequired?
Does AgileEngine provide workequipment?
What opportunities for professional growth do youoffer?
What does a typical team look like, and what tools do youuse?
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

DevOps / Site Reliability Engineer ID70127
DevOps / Site Reliability Engineer ID70127

AgileEngine, LLC. • Irving (TX)

On-site
USD 140,000 - 190,000
Growth without limits
Competitive compensation
Flexibility: remote work with flexible
+3
DevOps / Site Reliability Engineer ID70127
DevOps / Site Reliability Engineer ID70127

AgileEngine, LLC. • Dallas (TX)

On-site
USD 150,000 - 190,000
Growth opportunities
Competitive compensation
Remote work options
+3
DevOps / Site Reliability Engineer ID70127
DevOps / Site Reliability Engineer ID70127

AgileEngine • Chicago (IL)

Hybrid
USD 120,000 - 150,000
Professional growth
Competitive compensation
Exciting projects
+1
DevOps / Site Reliability Engineer ID70127
DevOps / Site Reliability Engineer ID70127

AgileEngine • Atlanta (GA)

Hybrid
USD 100,000 - 130,000
Professional growth
Competitive compensation
Exciting projects
+1
Remote DevOps & SRE Lead: Multi-Cloud & Incidents
Remote DevOps & SRE Lead: Multi-Cloud & Incidents

AgileEngine, LLC. • Boston (MA)

On-site
USD 150,000 - 190,000
Growth without limits
Competitive compensation
Flexibility: 100% remote with flexible
+3
Senior Multi-Cloud SRE & Incident Commander (Remote)
Senior Multi-Cloud SRE & Incident Commander (Remote)

AgileEngine, LLC. • New York (NY)

On-site
USD 140,000 - 200,000
Growth without limits
Competitive compensation
Flexibility: 100% remote with flexible
+3
Remote SRE - DevOps, Incident Commander & CI/CD
Remote SRE - DevOps, Incident Commander & CI/CD

AgileEngine, LLC. • Blacksburg (VA)

On-site
USD 140,000 - 180,000
Growth without limits
Competitive compensation
Flexibility: 100% remote
+3
Remote Senior DevOps & SRE — Multi-Cloud Incident Commander
Remote Senior DevOps & SRE — Multi-Cloud Incident Commander

AgileEngine, LLC. • Texas City (TX)

On-site
USD 140,000 - 190,000
Remote work
Annual learning budget
Competitive compensation reviews
+3
Remote SRE Lead: Multi-Cloud Incident Command & CI/CD
Remote SRE Lead: Multi-Cloud Incident Command & CI/CD

AgileEngine, LLC • Brazil (IN), Northern (KY)

Hybrid
USD 120,000 - 160,000
Remote work 100%
Flexible hours
Annual learning budget
Senior DevOps & SRE: Multi-Cloud Incident Lead (Remote)
Senior DevOps & SRE: Multi-Cloud Incident Lead (Remote)

AgileEngine, LLC. • Baltimore (MD)

On-site
USD 140,000 - 190,000
Growth without limits
Competitive compensation
100% remote with flexible hours
+3