Site Reliability Engineer

Gridiron IT

Arlington (VA)

On-site

USD 120,000 - 210,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Medical Insurance
401(k) Plan
Disability & ADD Insurance
Life & Pet Insurance

Job summary

Gridiron IT Solutions seeks a DevOps Site Reliability Engineer to ensure continuous availability and performance of an enterprise platform across multi-cloud environments. You will implement monitoring, alerting, and rapid response strategies to meet a 99.9% uptime requirement and bridge development with operations for stable deployments.

The role requires hands-on Kubernetes management, incident response, and automation to reduce toil, with security clearance considerations and a strong

Qualifications

  • Strong background in Site Reliability Engineering principles and practices.
  • Hands-on experience with Kubernetes deployments and orchestration.
  • Proven experience managing and troubleshooting multi-cloud environments (GCP, Azure, AWS).
  • Expertise in setting up and managing monitoring, logging, and automated alerting systems.
  • Proficiency in scripting and automation for system stabilization and diagnostic tasks.

Responsibilities

  • Continuous Monitoring: 24/7 infrastructure health monitoring and automated telemetry tracking to maintain 99.9% availability.
  • Incident Response: on-call to respond to major incidents within 1 hour of notification and perform rapid stabilization.
  • Automation: develop and maintain scripts and tools to automate diagnostic tasks and health checks.
  • Dashboarding & Telemetry: create dashboards for uptime, incident status, API latency, and health metrics.
  • Vendor Coordination: lead engineering-level coordination with CSPs during outages or infra issues.
  • Reliability Engineering: ongoing evaluation and improvements to logging, metrics, and architecture for scalability and availability.

Skills

SRE principles
Kubernetes
Multi-cloud environments
Monitoring & alerting
Scripting (Python/Bash)

Tools

Kubernetes

Job description

DevOps Site Reliability Engineer (SRE) essential to ensuring the continuous availability, performance, and health of an enterprise platform. This role focuses on maintaining reliable and stable software deployments in multi-cloud environments, establishing comprehensive monitoring and alerting systems, and providing rapid-reaction troubleshooting. You will bridge the gap between development and operations to uphold the mandated 99.9% uptime requirement for a critical defense enterprise platform.

Desired Qualifications
  • - Strong background in Site Reliability Engineering principles and practices.
  • - Hands-on experience with Kubernetes deployments and orchestration.
  • - Proven experience managing and troubleshooting multi-cloud environments (GCP, Azure, AWS).
  • - Expertise in setting up and managing monitoring, logging, and automated alerting systems.
  • - Proficiency in scripting and automation (e.g., Python, Bash) for system stabilization and diagnostic tasks.
What we are looking for in a candidate
  • - A strong interest in maintaining highly reliable software deployments under stringent uptime requirements.
  • - Ability to remain calm and systematic during high-pressure downtime incidents.
  • - Excellent diagnostic and troubleshooting skills with an eye for rapid resolution.
  • - A "build-to-manage" mindset, focusing on automating repetitive tasks to reduce operational toil.
Key Responsibilities
  • - Continuous Monitoring: Provide 24/7 infrastructure health monitoring and automated telemetry tracking to maintain the mandated 99.9% core platform availability.
  • - Incident Response: Serve on-call to respond to major incidents or platform downtime within 1 hour of notification, executing rapid platform downtime response and system stabilization maneuvers.
  • - Automation: Develop and maintain a library of scripts and internal tools to streamline and automate repetitive diagnostic tasks, health checks, and data-gathering procedures.
  • - Dashboarding & Telemetry: Create and maintain automated dashboards for uptime, incident status, API latency, and other critical support metrics to ensure platform health visibility.
  • - Vendor Coordination: Lead direct engineering-level coordination with Cloud Service Providers (CSPs) during outages or underlying infrastructure issues.
  • - Reliability Engineering: Continuously evaluate, monitor, and provide recommended improvements to logging, system metrics, and architecture to ensure the platform remains rapidly scalable and highly available.
Clearance

Applicants selected will be subject to a security investigation and may need to meet eligibility requirements for access to classified information. TS/SCI required.

Compensation and Benefits

Salary Range $120,000 - $210,000/YR (Compensation is determined by various factors, including but not limited to location, work experience, skills, education, certifications, seniority, and business needs. This range may be modified in the future.)

Benefits: Gridiron offers a comprehensive benefits package including medical, dental, vision insurance, HSA, FSA, 401(k), disability & ADD insurance, life and pet insurance to eligible employees. Full-time and part-time employees working at least 30 hours per week on a regular basis are eligible to participate in Gridiron's benefits programs.

Gridiron IT Solutions is an Equal Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, pregnancy, sexual orientation, gender identity, national origin, age, protected veteran status or disability status.

Gridiron IT is a Women Owned Small Business (WOSB) headquartered in the Washington, D.C. area that supports our clients' missions throughout the United States. Gridiron IT specializes in providing comprehensive IT services tailored to meet the needs of federal agencies. Our capabilities include IT Infrastructure & Cloud Services, Cyber Security, Software Integration & Development, Data Solution & AI, and Enterprise Applications. These capabilities are backed by Gridiron IT's experienced workforce and our commitment to ensuring we meet and exceed our clients' expectations.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

DevOps Site Reliability Engineer (SRE)
DevOps Site Reliability Engineer (SRE)

IT Veterans • Washington

On-site
USD 120,000 - 180,000
Site Reliability Engineer
Site Reliability Engineer

SRE • Puerto Rico

Hybrid
USD 120,000 - 180,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

GovCIO • Arlington (VA)

Hybrid
USD 210,000 - 230,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Govcio LLC • United States

Hybrid
USD 210,000 - 230,000
Senior DevOps/SRE Engineer
Senior DevOps/SRE Engineer

VITG • Ellicott City (MD)

Hybrid
USD 90,000 - 120,000
401(k) with employer contribution
Medical/Dental/Vision insurance
Paid vacation (PTO)
Site Reliability Engineer
Site Reliability Engineer

Experis • Charlotte (NC)

Hybrid
USD 96,000 - 103,000
Site Reliability Engineer
Site Reliability Engineer

TalentDome Staffing • United States

On-site
USD 140,000 - 210,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Virtual Tech Gurus • Puerto Rico

On-site
USD 140,000 - 210,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

OutSolve • Mission (KS)

Remote
USD 90,000 - 130,000
100% remote work environment
Competitive compensation
Professional development opportunities
+1
Site Reliability Engineer
Site Reliability Engineer

Govcio LLC • Arlington (TX)

Hybrid
USD 230,000 - 250,000