Site Reliability Engineer

Gridiron IT

Arlington (VA)

On-site

USD 120,000 - 210,000

Full time

8 days ago
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Medical insurance
Dental insurance
Vision insurance
HSAs/FSAs
401(k)
Disability & ADD insurance
Life and pet insurance

Job summary

Gridiron IT is seeking a DevOps Site Reliability Engineer (SRE) to ensure continuous availability, performance, and health of the enterprise platform across multi-cloud environments.

You will implement monitoring, alerting, and automation, handle on-call incidents within 1 hour, and collaborate with CSPs to maintain 99.9% uptime. This role requires strong SRE skills, scripting in Python/Bash, and a calm, methodical approach to outages.

Qualifications

  • Strong background in Site Reliability Engineering principles and practices.
  • Hands-on experience with Kubernetes deployments and orchestration.
  • Proven experience managing and troubleshooting multi-cloud environments (GCP, Azure, AWS).
  • Expertise in setting up and managing monitoring, logging, and automated alerting systems.
  • Proficiency in scripting and automation (e.g., Python, Bash) for system stabilization and diagnostic tasks.

Responsibilities

  • Continuous Monitoring: Provide 24/7 infrastructure health monitoring and automated telemetry tracking to maintain the mandated 99.9% core platform availability.
  • Incident Response: Serve on-call to respond to major incidents or platform downtime within 1 hour of notification, executing rapid platform downtime response and system stabilization maneuvers.
  • Automation: Develop and maintain a library of scripts and internal tools to streamline and automate repetitive diagnostic tasks, health checks, and data-gathering procedures.
  • Dashboarding & Telemetry: Create and maintain automated dashboards for uptime, incident status, API latency, and other critical support metrics to ensure platform health visibility.
  • Vendor Coordination: Lead direct engineering-level coordination with Cloud Service Providers (CSPs) during outages or underlying infrastructure issues.
  • Reliability Engineering: Continuously evaluate, monitor, and provide recommended improvements to logging, system metrics, and architecture to ensure the platform remains rapidly scalable and highly available.

Skills

Site Reliability Engineering
Kubernetes
Multi-cloud (GCP/Azure/AWS)
Monitoring & alerting
Scripting (Python/Bash)

Tools

Kubernetes

Job description

DevOps Site Reliability Engineer (SRE) essential to ensuring the continuous availability, performance, and health of an enterprise platform. This role focuses on maintaining reliable and stable software deployments in multi-cloud environments, establishing comprehensive monitoring and alerting systems, and providing rapid-reaction troubleshooting. You will bridge the gap between development and operations to uphold the mandated 99.9% uptime requirement for a critical defense enterprise platform.


Desired Qualifications


  • Strong background in Site Reliability Engineering principles and practices.

  • Hands-on experience with Kubernetes deployments and orchestration.

  • Proven experience managing and troubleshooting multi-cloud environments (GCP, Azure, AWS).

  • Expertise in setting up and managing monitoring, logging, and automated alerting systems.

  • Proficiency in scripting and automation (e.g., Python, Bash) for system stabilization and diagnostic tasks.


What we are looking for in a candidate


  • A strong interest in maintaining highly reliable software deployments under stringent uptime requirements.

  • Ability to remain calm and systematic during high-pressure downtime incidents.

  • Excellent diagnostic and troubleshooting skills with an eye for rapid resolution.

  • A \"build-to-manage\" mindset, focusing on automating repetitive tasks to reduce operational toil.


Key Responsibilities


  • Continuous Monitoring: Provide 24/7 infrastructure health monitoring and automated telemetry tracking to maintain the mandated 99.9% core platform availability.

  • Incident Response: Serve on-call to respond to major incidents or platform downtime within 1 hour of notification, executing rapid platform downtime response and system stabilization maneuvers.

  • Automation: Develop and maintain a library of scripts and internal tools to streamline and automate repetitive diagnostic tasks, health checks, and data-gathering procedures.

  • Dashboarding & Telemetry: Create and maintain automated dashboards for uptime, incident status, API latency, and other critical support metrics to ensure platform health visibility.

  • Vendor Coordination: Lead direct engineering-level coordination with Cloud Service Providers (CSPs) during outages or underlying infrastructure issues.

  • Reliability Engineering: Continuously evaluate, monitor, and provide recommended improvements to logging, system metrics, and architecture to ensure the platform remains rapidly scalable and highly available.


Clearance

Applicants selected will be subject to a security investigation and may need to meet eligibility requirements for access to classified information. TS/SCI required.


Compensation and Benefits

Salary Range $120,000 - $210,000/YR (Compensation is determined by various factors, including but not limited to location, work experience, skills, education, certifications, seniority, and business needs. This range may be modified in the future.)


Benefits: Gridiron offers a comprehensive benefits package including medical, dental, vision insurance, HSA, FSA, 401(k), disability & ADD insurance, life and pet insurance to eligible employees. Full-time and part-time employees working at least 30 hours per week on a regular basis are eligible to participate in Gridiron’s benefits programs.


Gridiron IT Solutions is an Equal Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, pregnancy, sexual orientation, gender identity, national origin, age, protected veteran status or disability status.


Gridiron IT is a Women Owned Small Business (WOSB) headquartered in the Washington, D.C. area that supports our clients' missions throughout the United States. Gridiron IT specializes in providing comprehensive IT services tailored to meet the needs of federal agencies. Our capabilities include IT Infrastructure & Cloud Services, Cyber Security, Software Integration & Development, Data Solution & AI, and Enterprise Applications. These capabilities are backed by Gridiron IT's experienced workforce and our commitment to ensuring we meet and exceed our clients' expectations.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

DevOps Site Reliability Engineer (SRE)
DevOps Site Reliability Engineer (SRE)

IT Veterans • Washington

On-site
USD 120,000 - 180,000
Site Reliability Engineer SRE SecOps
Site Reliability Engineer SRE SecOps

Arkenstone • Menlo Park (CA)

On-site
USD 140,000 - 190,000
Competitive Salary
Health & Wellness Programs
401(k) Plan
+3
Release and Deployment Engineer
Release and Deployment Engineer

Gridiron IT • Fort Meade (MD)

On-site
USD 75,000 - 125,000
Medical
Dental
Vision
+5
Senior DevOps/SRE Engineer
Senior DevOps/SRE Engineer

VITG • Ellicott City (MD)

Hybrid
USD 90,000 - 120,000
401(k) with employer contribution
Medical/Dental/Vision insurance
Paid vacation (PTO)
K8 platform engineer
K8 platform engineer

Seneca Resources Company, LLC • Washington

On-site
USD 185,000 - 230,000
Performance bonuses
Company-paid training/certifications
Referral bonuses
+1
Lead Site Reliability Engineer (SRE)
Lead Site Reliability Engineer (SRE)

IP Secure, LLC • San Antonio (TX)

Hybrid
USD 140,000 - 190,000
Medical
Dental
Vision
+3
Site Reliability Engineering (SRE)
Site Reliability Engineering (SRE)

Weekday (YC W21) • New York (NY)

On-site
USD 150,000 - 250,000
Health, dental, vision insurance
Generous PTO
Learning & development
+2
Site Reliability Engineer
Site Reliability Engineer

Brooksource • San Antonio (TX)

On-site
USD 80,000 - 120,000
SRE: 99.9% Uptime in Multi-Cloud & Kubernetes
SRE: 99.9% Uptime in Multi-Cloud & Kubernetes

Gridiron IT • Arlington (VA)

On-site
USD 120,000 - 210,000
Medical insurance
Dental insurance
Vision insurance
+4
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Cross River • United States

On-site
USD 160,000 - 200,000