Senior Systems Engineer

Jobtailor

Atlanta (GA)

On-site

USD 120,000 - 180,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Jobtailor in Atlanta, GA seeks a senior infrastructure operations engineer responsible for the health of critical domains and for leading major incidents and DR exercises. You will own patching, capacity planning, and automation efforts while improving monitoring and cross‑team collaboration.

Ideal candidates have 5+ years in systems engineering or senior IC roles, strong written and verbal communication, and the ability to guide audits with clear documentation.

Qualifications

  • Senior individual contributor or technical lead capacity.
  • Experience with multiple infrastructure domains (servers, virtualization, cloud, networking, identity, backup).
  • Proven track record in leading major incidents and DR exercises.
  • Strong ability to document architectures and runbooks for audit and review.

Responsibilities

  • Own operational health of one or two infrastructure domains to keep them healthy and improving.
  • Lead major incident responses and drive resolution, coordinating responders and stakeholders.
  • Own patching programs and disaster‑recovery execution with evidence of testing and results.
  • Drive capacity planning, refresh cycles, and decommissioning of legacy systems at scale.
  • Operate and improve monitoring and observability; tune alerts and build dashboards.
  • Lead cross‑team initiatives with Platform Engineering, Cybersecurity, Networking, and ITSM.
  • Define and document operational patterns, runbooks, and standards for audits.
  • Mentor engineers and review quality of output, supporting less senior teammates.
  • Serve as senior escalation in on‑call rotations and after‑hours incidents.

Skills

Senior IC / Tech Lead
Cloud infrastructure
Incident management
Disaster recovery planning
Automation / IaC
Monitoring & observability

Job description

Responsibilities
  • Domain ownership: Own the operational health of one or two infrastructure domains (e.g., server platforms, virtualization, cloud infrastructure, identity, backup and recovery, specialty business systems). Keep them measurably healthy and improving.
  • Major incident leadership: Lead major incident response, drive technical resolution, coordinate responders, communicate to stakeholders, and own the post‑incident review and corrective actions.
  • Patching and DR programs: Own the patching program for assigned domains—cadence, exception handling, reporting, continuous improvement. Own disaster recovery execution: maintain and exercise DR runbooks, coordinate tests, and ensure recovery objectives are met and evidenced.
  • Server lifecycle at scale: Drive capacity planning, refresh cycles, configuration baselines, and decommissioning of legacy systems.
  • Monitoring and observability: Operate and improve monitoring and observability across managed systems—tune alerting, eliminate noise, build dashboards, and contribute to anomaly‑detection and AIOps initiatives.
  • Cross‑team initiatives: Lead initiatives that span Platform Engineering, Cybersecurity, Networking, and application teams—controlled rollouts, hardening efforts, platform migrations. Land them without breaking production.
  • Standards and patterns: Define and document operational patterns, runbooks, and standards the team executes against and the auditors review.
  • Mentorship: Pair with Systems Engineers, run technical reviews, give substantive feedback, and grow the next tier. Quality of output from less senior engineers is part of your scope.
  • Operational partnership: Be the senior partner with Platform Engineering, Cybersecurity, Networking, and IT Service Management when they need operational input. Solve problems with them, not at them.
  • Specialty systems: Provide operational ownership of specialty business systems and legacy platforms supporting business‑critical applications.
  • Security and Audit: Apply and validate security baselines, lead remediation of high‑severity findings, and keep your domains' evidence map current.
  • Automation: Push toward repeatable, codified operations (IaC, automated evidence collection, scripted runbooks) instead of one‑off manual work.
  • On‑call: Participate in and serve as senior escalation for the on‑call rotation, including after‑hours support for high‑severity incidents, change windows, and disaster recovery events.
Requirements
  • 5+ years in systems engineering, systems administration, or infrastructure operations, including time in a senior individual contributor or technical lead capacity.
  • Strong fundamentals across multiple infrastructure domains (server platforms, virtualization, cloud infrastructure, networking, identity, backup and recovery).
  • Experience operating production workloads in at least one major cloud platform.
  • Demonstrated experience leading major incidents and disaster recovery exercises.
  • Ability to produce clear architecture, operational, and decision documentation that holds up under audit and peer review.
  • Excellent written and verbal communication; able to explain trade‑offs across technical and business audiences in plain language.
  • Comfortable mentoring less senior engineers and owning quality‑of‑output for one or more domains.
  • Comfortable serving as senior escalation in an on‑call rotation.
Core Competencies

Demonstrates expertise in operational health management of infrastructure domains, including server platforms and cloud infrastructure, while leading major incident responses and disaster recovery initiatives. Proficient in mentoring engineers and producing clear documentation for operational standards and audits.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Systems Engineer
Senior Systems Engineer

NextGenEnergyJobs • Atlanta (GA)

On-site
USD 120,000 - 180,000
Systems Operations and Engineering Manager
Systems Operations and Engineering Manager

LOOP • Greenville (SC), Spartanburg (SC), Anderson (SC)

On-site
USD 120,000 - 160,000
Senior System Engineer
Senior System Engineer

Confidential • Lansing (MI)

On-site
USD 90,000 - 120,000
Vice President, Lead Engineer – Infrastructure Operations
Vice President, Lead Engineer – Infrastructure Operations

Jobtailor • Illinois

On-site
USD 130,000 - 170,000
Systems Engineer
Systems Engineer

VT Group (VTG) • McLean (VA)

On-site
USD 125,000 - 170,000
Senior System Engineer
Senior System Engineer

Hollingsworth & Vose • East Walpole (MA)

On-site
USD 120,000 - 180,000
Systems Engineer
Systems Engineer

Vosper Thornycroft Group • McLean (VA)

On-site
USD 120,000 - 180,000
Sr Systems Engineer
Sr Systems Engineer

Insight Global • New York (NY)

On-site
USD 120,000 - 150,000
Software Engineering Manager - Reliability Engineering, Store Systems (Remote)
Software Engineering Manager - Reliability Engineering, Store Systems (Remote)

The Home Depot • Tallahassee (FL)

Hybrid
USD 150,000 - 230,000
Senior Lead System Engineer
Senior Lead System Engineer

Infinite Computer Solutions • Campus (IL)

On-site
USD 120,000 - 180,000