Site Reliability Engineer (req-307)

CATHEXIS

Tysons (VA)

On-site

USD 100,000 - 160,000

Full time

13 days ago
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Performance Bonuses
Medical Insurance
Dental Insurance
Vision Insurance
401(k) Plan (Traditional and ROTH)
Life Insurance (Basic, Voluntary & AD&
Paid Time Off
11 Federal Holidays
Parental Leave
Commute Benefits
Disability (Short/Long Term)

Job summary

CATHEXIS is seeking a Site Reliability Engineer (SRE) with a Top Secret clearance to join our team in Tysons, VA. You will manage Kubernetes clusters, monitor cloud infrastructure, and drive automation using Terraform, Helm, and IaC practices.

The role requires hands-on experience with AWS/Azure, Linux, Docker, and a passion for scalable AI-enabled solutions, working in an agile environment to enhance reliability and security for client environments.

Qualifications

  • Top Secret clearance required.
  • Bachelor’s degree in Computer Science or related field.
  • 2+ years of experience with on-prem and cloud environments.
  • Experience with AWS and/or Azure.
  • Hands-on with Linux, Docker, Kubernetes, Terraform, Helm, PostgreSQL, and similar tech.
  • Proficiency in at least one high-level language (Python/Java/C/C++/Ruby/JavaScript).
  • Experience with distributed storage (NFS/HDFS/Ceph/S3).
  • Ability to work in Agile/Scrum and lead initiatives.
  • Strong collaborative and communication skills.

Responsibilities

  • Monitor and manage Kubernetes clusters for stability, health, and scalability.
  • Deploy, monitor, and scale applications; maintain Helm charts and resource allocation.
  • Design and maintain Docker-based microservices and reproducible deployments.
  • Work with AWS/Azure/GCP to configure infrastructure using Terraform/CloudFormation.
  • Set up monitoring, define alerts, and manage incident response for Jenkins/Kubernetes.
  • Build automation for scaling and infrastructure using IaC and related tools.
  • Collaborate with development, services, and operations teams for seamless integration.
  • Ensure security/compliance with RBAC, encryption and vulnerability scanning.

Skills

Kubernetes
Cloud infrastructure
IaC (Terraform, CloudFormation)
Docker
Linux
Python/Java/C/C++
Jenkins
Security & compliance
Agile/Scrum
Open-source tooling

Education

Bachelor’s degree in Computer Science or related field

Tools

Docker
Kubernetes
Terraform
Helm
PostgreSQL
Linux
CloudFormation

Job description

Team CATHEXIS elevates the government contracting experience through rapid response, deep skill, and thoughtful problem-solving and communication. Our core capabilities are our top-tier program and project management, data analytics, and audit services, the backbone of which is our integrated approach to operational excellence.

You worked hard to get to where you are. You strive to make every day better than the day before. So do we. Team CATHEXIS operates with an all-in mindset. We are working together to create a company that supports our shared values and individual goals. Our values are centered around leading with integrity, owning the outcome, growing together, and moving with purpose in everything we do for our employees, customers, partners, and communities. We believe success is best when we listen and lead with empathy; model high standards of ethics to provide a rewarding candidate experience; work hard, have fun, and appreciate the strengths we all bring to the team; and empower our employees to create innovative and trusted results.

We are looking for a dynamic Site Reliability Engineer (SRE) with a Top Secret clearance to join our team! The Site Reliability Engineer (SRE) will manage, monitor, and optimize clusters on Kubernetes. Together, we’re accelerating our clients’ digital transformation through the building and deployment of data-driven, scalable AI solutions. The ideal candidate will have a deep understanding of Kubernetes, Cloud Infrastructure, and Infrastructure as Code (IaC) practices. You will be responsible for ensuring the reliability and scalability of our clients’ Kubernetes clusters and Cloud Infrastructure.

Responsibilities

The responsibilities include, but are not limited to:

  • Monitor and Manage Kubernetes Clusters: Ensure the stability, health, and scalability of Kubernetes Clusters, deploying applications and services on Kubernetes
  • Kubernetes Management: Deploy, monitor, and scale applications on Kubernetes clusters. Maintain Helm charts, manage services, and ensure resource allocation for optimal cluster performance
  • Containerization & Deployment: Design and maintain Docker-based microservices architecture, ensuring consistent and reproducible deployments across staging, QA, and production environments
  • Cloud Infrastructure Management: Work with leading Cloud Platforms (AWS, Azure and/or GCP) to set up, configure, and manage infrastructure resources using Infrastructure as Code (Terraform, CloudFormation, etc.)
  • Monitoring & Incident Response: Set up monitoring solutions, define alerts, an manage the incident response process for any issues related to Jenkins or Kubernetes clusters
  • Automate Infrastructure Processes: Build automation tools for scaling, monitoring, and maintaining infrastructure using modern tools like Terraform, Ansible, Linux, or equivalent
  • Collaborate Across Teams: Work closely with development, services, and operations teams to ensure a seamless integration between application development, deployment, and infrastructure
  • Security & Compliance: Ensure all systems follow best practices in terms of security and compliance with relevant regulations. This includes role-based access, encryption, and automated vulnerability scanning
  • Active TOP SECRET clearance or higher is required
  • Bachelor’s degree in Computer Science or related field
  • A minimum of two (2) years of experience working with on-premise and off-premise cloud environments
  • Experience with AWS and/or Azure
  • Hands-on experience with a range of open-source technologies, such as Linux, Docker, Kubernetes, K8s, Terraform, Helm, PostgreSQL, or similar technologies
  • Ability to program (structured and OOP) using one or more high-level languages, such as Python, Java, C/C++, Ruby, and JavaScript
  • Experience with distributed storage technologies such as NFS, HDFS, Ceph, and Amazon S3, as well as dynamic resource management frameworks (Apache Mesos, Kubernetes, Yarn)
  • Proactive approach to identifying problems, performance bottlenecks, and areas for improvement
  • Ability to lead and work independently in an Agile/Scrum environment
  • Real passion for developing team-oriented solutions to complex engineering problems
  • Thrive in an autonomous, empowering and exciting environment
  • Great verbal and written communication skills to collaborate multi-functionally and improve scalability
  • Interest in committing to a fun, friendly, expansive, and intellectually stimulating environment
Desired Skills
  • Hands-on experience deploying and operating applications using IaaS and PaaS on major cloud providers, such as Amazon AWS, Microsoft Azure, or Google Cloud Services
  • Experience with deep learning, natural language processing, computer vision, or reinforcement learning
  • Conveys highly technical concepts and information in written form to technical and non-technical audiences
  • The ability to work on multiple concurrent projects is essential. Strong self-motivation and the ability to work with minimal supervision
  • Must be a team-oriented individual, energetic, result & delivery oriented, with a keen interest on quality and the ability to meet deadlines

CATHEXIS offers competitive compensation packages to all eligible employees. Our goal is to provide a compensation package that reflects the value you bring to our team, is competitive with national average market rates, and promotes your financial security and personal well-being. The annual salary range for this role is $100,000 - $160,000. Please note that the salary information provided is a general guideline. CATHEXIS considers various factors in its final offer, including location, qualifications, experience, and skills.

  • Performance Bonuses
  • Medical Insurance
  • Dental Insurance
  • Vision Insurance
  • 401(k) Plan (Traditional and ROTH)
  • Life Insurance (Basic, Voluntary & AD&D)
  • Paid Time Off
  • 11 Federal Holidays
  • Parental Leave
  • Commute Benefits
  • Short Term & Long Term Disability
  • Training & Development
  • Wellness Program
  • Community Outreach Initiatives

CATHEXIS is an Equal Employment Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, or protected veteran status and will not be discriminated againston the basis ofdisability EEO IS THE LAW.If you are an individual with a disability and would like to request a reasonable accommodation as part of the employment selection process, please contact the Recruiting DepartmentRecruitingTeam@cathexiscorp.com

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Software Engineer (req-291)
Software Engineer (req-291)

Cathexisfederal • Redwood City (CA)

On-site
USD 137,000 - 167,000
Performance Bonuses
Medical Insurance
Dental Insurance
+11
Software Engineer (req-292)
Software Engineer (req-292)

Cathexisfederal • Tysons (VA), Northern (KY)

On-site
USD 123,000 - 150,000
Medical Insurance
Dental Insurance
Vision Insurance
+11
Senior IT Talent Acquisition Specialist (req-304)
Senior IT Talent Acquisition Specialist (req-304)

CATHEXIS • Tysons (VA)

On-site
USD 115,000 - 125,000
Performance Bonuses
Medical Insurance
Dental Insurance
+11
Software Engineer (req-291)
Software Engineer (req-291)

CATHEXIS • Redwood City (CA)

On-site
USD 137,000 - 167,000
Performance Bonuses
Medical Insurance
Dental Insurance
+11
Senior Talent Acquisition Specialist (req-304)
Senior Talent Acquisition Specialist (req-304)

CATHEXIS • Tysons (VA)

Hybrid
USD 115,000 - 125,000
Performance Bonuses
Medical Insurance
Dental Insurance
+11
Site Reliability Engineer
Site Reliability Engineer

CyberPoint International • Maryland

Hybrid
USD 120,000 - 170,000
Senior AI/ML Software Developer - Top Secret (req-318)
Senior AI/ML Software Developer - Top Secret (req-318)

Cathexisfederal • Tysons (VA)

Hybrid
USD 148,000 - 179,000
Performance Bonuses
Medical Insurance
Dental Insurance
+11
SRE/Devops Engineer
SRE/Devops Engineer

INSPYR Solutions • Sunnyvale (CA)

On-site
USD 120,000 - 180,000
Work-life balance
No on-call requirements
Standard business hours
Senior AI/ML Software Developer - Top Secret (req-318)
Senior AI/ML Software Developer - Top Secret (req-318)

CATHEXIS • Tysons (VA)

On-site
USD 148,000 - 179,000
Performance Bonuses
Medical Insurance
Dental Insurance
+11
Site Reliability Engineer - 174242 - 174721 - 175240
Site Reliability Engineer - 174242 - 174721 - 175240

ZP Group • United States

Remote
USD 140,000 - 165,000
Medical, Dental, Vision
401(k) with match