Senior Site Reliability Engineer

Valid8 Financial, Inc.

Washington (District of Columbia)

Hybrid

USD 120,000 - 180,000

Full time

3 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Health Insurance
Unlimited Vacation
Training
Mentoring/Coaching

Job summary

Elevate Government Solutions is seeking a Senior Site Reliability Engineer to join a Federal client project. The role is remote for US citizens with a SECRET clearance, focused on maintaining production infrastructure and improving uptime.

The candidate will work with Python, Linux environments, Terraform, Ansible, Docker, and CI/CD pipelines, leveraging AWS services like EKS, S3, and EMR, plus Spark/JupyterHub/Hue for data processing tasks.

Qualifications

  • Proactive problem solving with ability to diagnose and resolve complex operational issues.
  • Experience with Python in operational/support environments.
  • Professional experience with Terraform, Ansible, Docker and CI/CD pipelines.
  • Strong Git version control skills.
  • Comfort with Linux command line and operational tasks.
  • Ability to learn new technologies quickly.
  • Experience supporting production infrastructure in a collaborative team.
  • Excellent incident response and communication skills.
  • Ability to work East Coast hours and hybrid in Vienna, VA or remotely.

Responsibilities

  • Maintain, monitor, and troubleshoot production environments for uptime and performance.
  • Operate and manage infrastructure using Terraform, Ansible, and Docker.
  • Oversee CI/CD pipelines and automate workflows with Git tools.
  • Ensure reliability of systems running on AWS services including EKS, S3, EMR.
  • Support Spark, JupyterHub, and Hue in operational contexts.
  • Diagnose and resolve issues, perform root cause analyses, drive resolution.
  • Implement and monitor infrastructure alerting and diagnostics.
  • Collaborate with engineering, product, and client staff on deployments.

Skills

Python
Terraform
Ansible
Docker
CI/CD pipelines
Git
Linux
Incident response
Communication
Learning agility
Hybrid work

Education

Bachelor's Degree in CS/Engineering or related

Tools

AWS (EKS, S3, EMR)
JupyterHub
Hue
Spark
EMR

Job description

This position requires an ACTIVE SECRET security clearance. Your application will not be reviewed if you do not have an ACTIVE security clearance at SECRET or above.

Elevate Government Solutions is a high growth Technology Services company focused on driving technological change in the government space. Our teams engage with various government agencies to develop and deploy emerging technology solutions using a tailored Agile methodology.

We are seeking a highly motivated and intellectually curious Senior Site Reliability Engineer to join our team working with a Federal client. The position will be a remote role open to US citizens residing in the United States with a Secret security clearance. The Senior Site Reliability Engineer will oversee the efficient, reliable, and secure operation of critical application infrastructure for a federal financial agency. In this operationally focused role, you will be responsible for maintaining production environments, troubleshooting incidents, performing monitoring and diagnostics, supporting deployments, and continually improving system uptime and reliability.

The ideal candidate brings Python experience, deep operational expertise with Linux environments, and hands‑on experience managing infrastructure with tools like Terraform, Ansible, and Docker. Experience with CI/CD pipelines, containerization, Git‑based workflows, and monitoring/alerting systems are essential. Familiarity with managing AWS services—particularly EKS, S3, and EMR (Elastic Map Reduce)—is highly valued. In addition, familiarity with Spark, JupyterHub, and Hue is preferred. Effective communication and a collaborative mindset are critical to success in this client‑facing environment.

Responsibilities and Duties
  • Maintain, monitor, and troubleshoot production environments to ensure optimal uptime and performance
  • Operate and manage infrastructure using Terraform, Ansible, and Docker
  • Oversee and support CI/CD pipelines and automate operational workflows using Git and related tooling
  • Ensure ongoing reliability and operational excellence for critical systems running on AWS services, including EKS, S3, and EMR (Elastic Map Reduce)
  • Support operational use of Spark, JupyterHub, and Hue
  • Diagnose and resolve operational issues, perform root cause analysis, and drive problem resolution
  • Implement, refine, and monitor infrastructure and application alerting and diagnostics
  • Optimize data flows and storage integrations
  • Collaborate with engineering, product, and client stakeholders to communicate issues, coordinate maintenance, and support deployments
  • Contribute to continual improvement of platform operational processes, documentation, and best practices
Required Experience, Skills and Qualifications
  • Proactive, creative problem‑solving mindset with a demonstrated ability to anticipate, diagnose, and resolve complex operational challenges
  • Experience with Python, specifically in operational and support environments
  • Professional experience with Terraform, Ansible, Docker, and CI/CD pipeline operations
  • Strong proficiency with Git and version control workflows
  • Comfortable working extensively on the Linux command line and performing operational tasks
  • Demonstrated ability to learn new technologies quickly
  • Experience monitoring, maintaining, and supporting production infrastructure in a collaborative environment
  • Strong incident response and communication skills
  • Ability to work East Coast hours
  • Ability to work hybrid in Vienna, VA, or remotely
Preferred Qualifications
  • Experience operating AWS or other cloud platforms
  • Familiarity with Spark, JupyterHub, and Hue in an operational context
  • Hands‑on experience with EMR (Elastic Map Reduce)
  • Experience with Databricks
  • Experience supporting federal, regulated, or client‑facing environments
Education Requirement
  • 4 year Bachelor's Degree in Computer Science/Eng or related (highly preferred)
  • Health Insurance
  • Unlimited Vacation
  • Training
  • Mentoring/Coaching
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Platform Operations Engineer
Senior Platform Operations Engineer

Valid8 Financial, Inc. • Washington

Hybrid
USD 120,000 - 180,000
Health Insurance
Unlimited Vacation
Training
+1
Remote Senior Site Reliability Engineer - Secret Clearance
Remote Senior Site Reliability Engineer - Secret Clearance

Valid8 Financial, Inc. • Washington

Hybrid
USD 120,000 - 180,000
Health Insurance
Unlimited Vacation
Training
+1
Senior Site Reliability Engineer
Senior Site Reliability Engineer

ClearanceJobs • Springfield (VA)

On-site
USD 130,000 - 190,000
Site Reliability Engineer
Site Reliability Engineer

GovCIO • Arlington (VA)

Hybrid
USD 230,000 - 250,000
Lead Site Reliability Engineer (SRE)
Lead Site Reliability Engineer (SRE)

IP Secure, LLC • San Antonio (TX)

Hybrid
USD 140,000 - 190,000
Medical
Dental
Vision
+3
Cloud/Network Infrastructure Engineer, Senior
Cloud/Network Infrastructure Engineer, Senior

ECS Corporate Services • Arlington (VA)

On-site
USD 123,000 - 184,000
DevSecOps Engineer
DevSecOps Engineer

Global Alliant Inc • Rockville (MD)

On-site
USD 90,000 - 135,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Jobgether • United States

Hybrid
USD 120,000 - 160,000
Competitive compensation package
Flexible work arrangements
Professional development opportunities
+2
Staff Site Reliability Engineer (Government)
Staff Site Reliability Engineer (Government)

SentinelOne • New York (NY)

On-site
USD 180,000 - 260,000
Sr Software Development Engineer, SRE (US Federal)
Sr Software Development Engineer, SRE (US Federal)

Workday, Inc. • Reston (VA)

On-site
USD 140,000 - 170,000