Senior SRE: Cloud, Kubernetes & Linux (TS/SCI)

Momentum Engineering, Inc

Annapolis (MD)

On-site

USD 83,000 - 124,000

Full time

5 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

11 paid holidays
3 weeks PTO
Group medical plan
Dental coverage
Vision
Life insurance
STD/LTD

Job summary

Momentum Engineering, Inc. is seeking a Site Reliability Engineer to support a cloud-based platform built with Java and FOSS technologies including Kubernetes, Hadoop, and Accumulo.

You will be on-call and provide Tier 1–3 support to ensure reliability and performance of distributed systems. The ideal candidate thrives in a fast-paced team, has strong Linux troubleshooting skills, and can develop automation with Python/Bash, using Prometheus and Grafana for observability.

Qualifications

  • Must have active Top Secret/SCI clearance with NSA Full Scope Polygraph.
  • Bachelor's degree in Computer Science or related field is highly desired, may be equivalent to 2 years' experience.
  • Master's degree may be considered equivalent to 4 years of experience.
  • Degrees in Mathematics, Information Systems, Engineering, or similar disciplines are technical degrees.
  • Fourteen (14) years of relevant technical experience.
  • Strong experience troubleshooting operational issues in Linux environments.
  • DoD 8570 IAT Level I certification or higher.
  • Ability to provide Tier 1 through Tier 3 support in a mission-critical environment.
  • Candidates must possess at least one of the following certifications:
  • AWS Certified Solutions Architect - Associate
  • AWS Certified Solutions Architect - Professional
  • AWS Certified SysOps Administrator - Associate
  • Elastic Certified Engineer
  • Elastic Certified Observability Engineer

Responsibilities

  • Support the operation, administration, reliability, and availability of cloud-based infrastructure and platform services.
  • Provide Tier 1 through Tier 3 technical support for operational issues affecting users, applications, and infrastructure.
  • Troubleshoot and resolve complex system, application, networking, and infrastructure issues within Linux environments.
  • Monitor system health, availability, performance, and operational status.
  • Support distributed computing technologies including Kubernetes, Hadoop, and Accumulo.
  • Support containerized applications and services using Docker and Kubernetes.
  • Develop and maintain automation and administrative scripts using Python, Bash, or similar scripting languages.
  • Use monitoring and observability tools such as Prometheus and Grafana to identify and resolve operational issues.
  • Support configuration management and automation tools such as Salt and Ansible.
  • Participate in incident response, root cause analysis, corrective actions, and continuous improvement activities.
  • Support virtualization, cloud, and hybrid infrastructure environments.
  • Document troubleshooting procedures, operational processes, system changes, and recurring issues.
  • Collaborate with developers, system administrators, engineers, and mission stakeholders to maintain reliable and scalable platform operations.
  • Participate in on-call support and respond to operational issues as required.

Skills

Active TS/SCI clearance
On-call support
Linux troubleshooting
Problem solving
Team collaboration

Education

Bachelor's degree in CS or related field
Master's degree; equivalence possible
Engineering/Math/Info Systems disciplines

Tools

Docker
Kubernetes
Prometheus
Grafana
Salt
Ansible
AWS
OpenStack
Hadoop
HDFS
Python

Job description

Momentum Engineering, Inc. is seeking a Site Reliability Engineer to support a cloud-based platform built with Java and FOSS technologies including Kubernetes, Hadoop, and Accumulo.

You will be on-call and provide Tier 1–3 support to ensure reliability and performance of distributed systems. The ideal candidate thrives in a fast-paced team, has strong Linux troubleshooting skills, and can develop automation with Python/Bash, using Prometheus and Grafana for observability.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer - Cloud, Kubernetes & Linux
Senior Site Reliability Engineer - Cloud, Kubernetes & Linux

Momentum Engineering • Maryland

On-site
USD 165,000 - 230,000
11 paid holidays
Minimum 3 weeks PTO
Group medical plan
+1
Site Reliability Engineer III — Mission-Critical Cloud & Kubernetes
Site Reliability Engineer III — Mission-Critical Cloud & Kubernetes

Momentumcareers • Corridor North (MD)

On-site
USD 150,000 - 190,000
11 paid holidays
3 weeks PTO
Medical, dental, vision plans
Onsite SRE — TS/SCI Clearance | Linux & Kubernetes
Onsite SRE — TS/SCI Clearance | Linux & Kubernetes

CyberPoint International • Maryland

Hybrid
USD 120,000 - 170,000
Cloud Systems Engineer — Linux, Kubernetes & TS/SCI Clearance
Cloud Systems Engineer — Linux, Kubernetes & TS/SCI Clearance

Momentumcareers • Corridor North (MD)

On-site
USD 120,000 - 160,000
11 paid holidays
3 weeks PTO
Group medical plan
+4
Senior SRE: Cloud, Kubernetes & Terraform (Remote)
Senior SRE: Cloud, Kubernetes & Terraform (Remote)

Motion Recruitment • Mount Laurel Township (NJ)

Remote
USD 130,000 - 180,000
Medical, dental, and vision benefits
Equity / Stock Options
Remote equipment stipend
+3
ME00680-Site Reliability Engineer 3
ME00680-Site Reliability Engineer 3

Momentum Engineering, Inc • Annapolis (MD)

On-site
USD 83,000 - 124,000
11 paid holidays
3 weeks PTO
Group medical plan
+4
Cloud Platform Administrator: TS/SCI, Kubernetes & Hadoop
Cloud Platform Administrator: TS/SCI, Kubernetes & Hadoop

Momentum Engineering, Inc • Annapolis (MD)

Hybrid
USD 120,000 - 160,000
11 paid holidays
Minimum 3 weeks PTO
Company-sponsored group medical plan
+4
ME00680-Site Reliability Engineer 3
ME00680-Site Reliability Engineer 3

Momentumcareers • Corridor North (MD)

On-site
USD 150,000 - 190,000
11 paid holidays
3 weeks PTO
Medical, dental, vision plans
Senior SRE - Cloud Infra, AWS & Kubernetes
Senior SRE - Cloud Infra, AWS & Kubernetes

Kids for the Future • United States

Hybrid
USD 130,000 - 195,000
Healthcare benefits
401(k) matching
Paid time off
Senior SRE: Cloud Reliability & AI-Driven Incident Triage
Senior SRE: Cloud Reliability & AI-Driven Incident Triage

Clearwater Analytics • Boise (ID)

On-site
USD 130,000 - 170,000