Senior Site Reliability Engineer

ecsfederal

Virginia (MN)

Ibrido

USD 118.000 - 177.000

Tempo pieno

2 giorni fa
Candidati tra i primi
Generatore di candidature

Una candidatura completa in un minuto — curriculum e lettera di presentazione personalizzati, pronti da inviare.

Supera i filtri ATS

Descrizione del lavoro

Everforth ECS is seeking a Senior Site Reliability Engineer to work remotely on the CDM data services program. You will define and grow the SRE practice, ensuring reliability, availability, and performance across critical production environments.

You will implement logging, monitoring, and alerting with Elastic stack and related tools, lead incident response, and perform root-cause analyses to prevent recurrence. Collaboration with multiple teams is essential.

Competenze

  • US citizenship with ability to obtain Public Trust.
  • 6+ years as a Site Reliability Engineer (SRE) or equivalent.
  • 6+ years of experience designing, implementing, and maintaining observability solutions (logging, monitoring, alerting).
  • 3+ years defining and measuring SLOs and SLIs.
  • 3+ years of hands-on experience with cloud platforms (AWS GovCloud preferred).
  • 3+ years of hands-on programming or scripting (e.g., Python, Bash).
  • Strong knowledge of microservices, containerization, and orchestration tools (Docker, Kubernetes).
  • Experience collaborating with cross-functional teams to integrate reliability and observability into the software development lifecycle.
  • Proficiency in developing Synthetic monitoring scripts using TypeScript.

Mansioni

  • Define, implement, and grow our SRE practice to ensure reliability, availability, and performance of our production environments.
  • Set up comprehensive logging, monitoring, and alerting solutions using Elastic stack and other tools; respond to incidents and perform root cause analyses.
  • Collaborate with developers, testers, infrastructure and DevOps to embed reliability and observability into the SDLC.

Conoscenze

SRE experience
Observability
Cloud platforms
Programming scripting
Microservices containers
Cross-functional collaboration
Synthetic monitoring

Strumenti

Elastic
Prometheus
Grafana
Splunk

Descrizione del lavoro

Everforth ECS is seeking a Senior Site Reliability Engineer to work remotely.

Everforth ECS is seeking talented professionals to join our successful and growing team in building the next-generation Continuous Diagnostics and Mitigation (CDM) Cyber data solution. The CDM Program is the Cybersecurity and Infrastructure Security Agency’s (CISA) dynamic approach to strengthening the cybersecurity of Federal networks and systems through better awareness and visibility into their security posture and cyber threats. ECS is responsible for designing, building, deploying, operating, and maintaining a complete ‘Data Services’ solution which includes the collection, normalization, visualization, and sharing of cyber data from more than 100 Federal agencies. The CDM Data Services product is an integrated suite of multiple Commercial Off the Shelf (COTS) products, software configuration packages, and custom code which work together to operate as an integrated solution tailored to meet Department of Homeland Security (DHS) requirements.

We are seeking professionals who thrive in a dynamic, fast-paced, and highly collaborative environment where problem-solving, critical thinking, and a holistic approach to serving the mission are key. Our program operates within the Scaled Agile Framework (SAFe). An aptitude and enthusiasm for continuous learning, improvement, and cyber security is a must!

Role & Responsibilities

ECS is seeking a talented Senior Site Reliability Engineer (SRE) to play a key role in defining, implementing, and growing our SRE practice to ensure the reliability, availability, and performance of our critical production environments.

The Senior SRE will contribute to a culture of continuous improvement, identifying areas for enhancement, and driving initiatives to improve system reliability, scalability, and efficiency.

The successful candidate will have demonstrated hands-on experience designing, implementing, and maintaining solutions to ensure that systems, including infrastructure and applications, are resilient, highly available, and performant. The Senior SRE will also play a critical role in defining and measuring the Service Level Objectives (SLOs) and Service Level Indicators (SLIs) for our solution.

The Senior SRE will be responsible for setting up comprehensive logging, monitoring, and alerting solutions using the Elastic stack and other tools as necessary to ensure the continuous performance of services. Additionally, they will respond to incidents, perform root cause analyses, and implement solutions to prevent reoccurrence. The Senior SRE will work in close collaboration with other SRE team members, developers, testers, infrastructure engineers, DevOps engineers, and other stakeholders to integrate reliability and observability into the software development lifecycle.

Salary Range: $118,000 - $177,000

General Description of Benefits

  • Must be a US citizen with the ability to obtain Public Trust Suitability.
  • 6+ years of experience as a Site Reliability Engineer (SRE) or equivalent
  • 6+ years of demonstrated experience designing, implementing, and maintaining observability solutions to include logging, monitoring, and alerting
  • 6+ years of hands-on experience with SRE tools (e.g., Elastic, Prometheus, Grafana, Splunk, etc.)
  • 3+ years defining and measuring SLOs and SLIs
  • 3+ years of relevant experience using cloud platforms (AWS GovCloud preferred)
  • 3+ years of hands-on programming or scripting (e.g., Python, Bash, etc.)
  • Strong knowledge of microservices, containerization, and orchestration tools (Docker, Kubernetes)
  • Proven ability to collaborate with cross-functional teams (development, testing, and product) to integrate reliability and observability into the software development lifecycle
  • Strong problem-solving and analytical skills
  • Proactive, detail-oriented approach to identifying inefficiencies and implementing improvements
  • Proficient in developing Synthetic monitoring scripts using typescript.
Ottieni la revisione del curriculum gratis e riservata.

o trascina qui il file.

Similar jobs

Offerte di lavoro simili che vale la pena confrontare

Senior DevOps/SRE Engineer
Senior DevOps/SRE Engineer

VITG • Ellicott City (MD)

In loco
USD 90.000 - 120.000
401(k) with employer contribution
Medical/Dental/Vision insurance
Paid vacation (PTO)
Cloud Site Reliability Engineer (SRE)
Cloud Site Reliability Engineer (SRE)

ECS • Arlington (VA)

Ibrido
USD 130.000 - 180.000
Site Reliability Engineer
Site Reliability Engineer

TalentDome Staffing • Stati Uniti

In loco
USD 140.000 - 210.000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Storm2 • Scottsdale (AZ)

In loco
USD 140.000 - 150.000
Competitive healthcare, dental, and vision coverage
401(k) with company match
Generous PTO and paid holidays
+1
Site Reliability Engineer (SRE) - 175164
Site Reliability Engineer (SRE) - 175164

Piper Companies • Stati Uniti

Remoto
USD 120.000 - 145.000
Health insurance
Vision insurance
Dental insurance
+3
Senior Site Reliability Engineer
Senior Site Reliability Engineer

GovCIO • Arlington (VA)

In loco
USD 210.000 - 230.000
Lead Site Reliability Engineer (SRE) / Principal Site Reliability Engineer (SRE)
Lead Site Reliability Engineer (SRE) / Principal Site Reliability Engineer (SRE)

Mindlance • Irving (TX)

In loco
USD 120.000 - 160.000
Site Reliability Engineer (Secret Clearance)
Site Reliability Engineer (Secret Clearance)

ROI Services LLC • Huntsville (AL)

In loco
USD 110.000 - 150.000
SRE - Site Reliability Engineer - Senior
SRE - Site Reliability Engineer - Senior

ManpowerGroup Global, Inc. • Austin (TX)

In loco
USD 66.000 - 90.000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

The ReWork Group • New York (NY)

In loco
USD 120.000 - 160.000