Manager - Software Engineer DevOps

Request Technology, LLC

Chicago (IL)

Hybrid

USD 140,000 - 200,000

Full time

2 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Request Technology, LLC seeks a Manager of DevOps - Infrastructure and Applications Production Support in Chicago, IL. This hybrid role leads a 6–10 person team delivering 24/7 L1/L2 support for deployments, middleware, and platform operations.

You will coordinate incident response, triage, and RCA documentation, while driving automation and monitoring improvements across multiple tech stacks. You will manage cross-functional incident resolution, engage Platform/S&I/Security teams, and own the

Qualifications

  • Minimum 5 years in environment operations, production support, or infrastructure operations.
  • Experience across at least four areas: deployment pipelines, container platforms, middleware, incident management, monitoring, configuration, release engineering, scripting or platform operations.
  • Proven ability to define and enforce SLA frameworks in a tiered support model (L1/L2/L3).
  • Familiarity with regulated environments and change governance is a plus.
  • Industry knowledge of environment operations, platform reliability, and automation practices.
  • Willingness to shift work and be on-call (24x7) as needed.
  • Previous people management or team lead experience is required.

Responsibilities

  • Lead L1 and L2 engineers in incident response, triage, coordination, resolution, and post-incident reporting.
  • Oversee analysis of incidents across deployments, middleware, and platform layers, coordinating with internal teams.
  • Serve as Tier 3 escalation for complex incidents, directing resolution across Platform, S&I, Security, and App Dev teams.
  • Own the full incident lifecycle from alert to RCA and permanent fix or workaround.
  • Drive post-incident reviews for P1/P2 incidents ensuring root causes are identified and acted upon.

Skills

Leadership
Incident management
SLA frameworks
Team mentoring
On-call readiness
Cross-functional collaboration

Tools

Jenkins
Kubernetes
Kafka
HashiCorp Vault
Splunk
Dynatrace
Datadog
ServiceNow
Terraform (TFE)
Harness

Job description

Manager, DevOps - Infrastructure - Applications Production Support

LOCATION: Chicago, IL

Hybrid 3 days onsite and 2 days remote

You will lead a team of 6-10, 24/7 support, L1, L2 support. Application deployment container platform ops middleware support incident management monitoring observability configuration release engineering platform operations scripting automation SLA frameworks harness Jenkins Kubernetes apache kakfa splunk Dynatrace datadog.

  • Lead L1 and L2 support engineers in all incident response activities including triage, investigation, coordination, resolution, closure, and post-incident reporting.
  • Oversee technical analysis of environment incidents across application deployments, middleware, and platform layers while coordinating response activities with internal engineering, platform, and application development teams.
  • Serve as Tier 3 escalation point for complex incidents beyond L2 capability — triaging, directing, and driving resolution across Platform (k8s, Kafka, TFE), S&I (deployment, middleware, storage, network), Security (Vault, certs, secrets), and App Dev teams.
  • Own the full incident lifecycle — from first alert through to RCA documentation and permanent fix or accepted workaround.
  • Drive post-incident reviews for all P1 and P2 incidents, ensuring root cause is identified, documented, and actioned — not filed.
Technical Skills:
  • Deployment & Pipeline tooling: Harness (continuous delivery pipelines, deployment verification, rollback automation), Jenkins (CI/CD pipeline management, job configuration, build troubleshooting), GitHub (branching strategies, pull request workflows, pipeline integration).
  • Container & orchestration platforms: Kubernetes (k8s) pod lifecycle management, namespace operations, log retrieval, resource troubleshooting, and coordination with Platform teams on cluster-level issues.
  • Messaging & streaming platforms: Apache Kafka topic management, consumer group monitoring, lag analysis, and escalation to Platform for broker-level issues.
  • Secrets & configuration management: HashiCorp Vault — secrets retrieval, token/lease troubleshooting, policy review, and escalation to Security teams for certificate and secrets rotation.
  • Monitoring & observability: Proficiency in at least two production monitoring toolsets (e.g. Splunk, Dynatrace, Datadog, AppDynamics, Prometheus/Grafana) alert triage, dashboard interpretation, log analysis, and tuning requests.
  • Middleware platforms: Working knowledge of middleware infrastructure including application servers, messaging brokers, storage integrations, and network-layer dependencies sufficient to triage, gather diagnostics, and route correctly to L3.
  • Incident and ticketing platforms: ServiceNow or equivalent ITSM tooling incident creation, SLA tracking, problem record management, and reporting.
  • MTTR and operational metrics: Ability to build and maintain operational dashboards and reports covering MTTR, SLA compliance, alert-to-incident ratio, repeat incident rate, and deployment success rate.
Education and/or Experience:
  • Minimum 5 years of hands‑on environment operations, production support, or infrastructure operations experience, including interdisciplinary experience across four or more of the following: application deployment pipelines, container platform operations, middleware support, incident management, monitoring and observability, configuration management, release engineering, platform operations, or scripting and automation.
  • Technical experience and comprehensive knowledge of production environment failure modes — including deployment failures, configuration drift, platform instability, and integration breakdowns — and the methodologies used to diagnose and resolve them.
  • Demonstrated experience defining and enforcing SLA frameworks in a tiered support model (L1/L2/L3 or equivalent).
  • Familiarity with financial services or other regulated‑industry production environments is a strong advantage — understanding of change governance, audit requirements, and production access controls.
  • Industry knowledge of current and emerging practices in environment operations, platform reliability, and support automation.
  • Shift work and on-call availability required — including 24×7 on-call response capacity and availability during planned and emergency maintenance windows.
  • Previous people management or team lead experience required; formal people management experience
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Manager - Software Engineering DevOps
Manager - Software Engineering DevOps

Request Technology, LLC • Chicago (IL)

Hybrid
USD 170,000 - 200,000
Manager of DevOps Operational Support
Manager of DevOps Operational Support

Request Technology, LLC • Chicago (IL)

On-site
USD 140,000 - 190,000
Bonus eligible
Manager, DevSecOps Production Support
Manager, DevSecOps Production Support

Request Technology, LLC • Chicago (IL)

On-site
USD 120,000 - 180,000
Manager, Software Engineering DevOps
Manager, Software Engineering DevOps

The Options Clearing Corporation • Chicago (IL)

Hybrid
USD 120,000 - 160,000
Hybrid work up to 2 days remote
Tuition Reimbursement
Student Loan Repayment Assistance
+4
Hybrid DevOps Manager: Incident Lead & Platform Reliability
Hybrid DevOps Manager: Incident Lead & Platform Reliability

Request Technology, LLC • Chicago (IL)

Hybrid
USD 140,000 - 200,000
DevOps Engineer
DevOps Engineer

Stelvio Inc. • Frisco (TX)

On-site
USD 110,000 - 160,000
Paid vacation
Sick leave
Bereavement leave
+2
DevOps Platform Manager — Lead 24/7 Incident Response
DevOps Platform Manager — Lead 24/7 Incident Response

Request Technology, LLC • Chicago (IL)

Hybrid
USD 170,000 - 200,000
IT CONSULTANT SR
IT CONSULTANT SR

First Horizon Corp. • Memphis (TN), Northern (KY)

Hybrid
USD 120,000 - 160,000
Software Engineering Group Manager – Site Reliability
Software Engineering Group Manager – Site Reliability

Jobtailor • Alabama

On-site
USD 120,000 - 180,000
IT CONSULTANT SR
IT CONSULTANT SR

First Horizon Bank • Memphis (TN)

On-site
USD 120,000 - 180,000