Site Reliability Engineer

Jobtailor

Rockville (MD)

On-site

USD 120,000 - 180,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Jobtailor is seeking a Site Reliability Engineer to join a team modernizing CMS enterprise data platforms in Rockville, MD. You will operate AWS environments, build observability, and drive reliable, automated deployments using Terraform, Ansible, Jenkins, and Docker.

You will ensure performance, security, and cost efficiency while supporting security/compliance initiatives and disaster recovery planning. Strong communication and problem‑solving skills are essential for this fast‑paced,

Qualifications

  • Bachelor's degree in computer science, engineering, or related field or equivalent hands‑on experience.
  • 3–5 years in site reliability, systems, or cloud engineering with AWS experience.
  • Solid working knowledge of core AWS services and best practices.
  • Hands‑on experience with infrastructure‑as‑code tools (Terraform, Ansible, CloudFormation).
  • Good understanding of CI/CD pipelines and automation tools (Jenkins, GitLab CI).
  • Comfort scripting and automating in Python.
  • Familiarity with monitoring/observability tools (CloudWatch, New Relic, Splunk).
  • Strong problem‑solving and calm under pressure; clear communication skills.

Responsibilities

  • Join CMS modernization effort to unify knowledge and data platforms.
  • Operate and tune AWS environments to meet availability SLAs during transition.
  • Build observability with dashboards and alerts; establish performance baselines.
  • Automate toil with Terraform/Ansible and support CI/CD pipelines and Docker workloads.
  • Define and track SLIs/SLOs; produce performance and bottleneck reports.
  • Optimize performance, security, and cost using AWS tools and best practices.
  • Support security/compliance modernization and risk scans within RMF/IS2P2.
  • Design and maintain disaster recovery and COOP continuity.
  • Own incidents end to end; drive blameless post-mortems and preventative fixes.

Skills

Site Reliability Engineering
AWS
AWS core services
Python scripting
Problem solving
Communication

Education

Bachelor's degree in computer science, engineering, or related field

Tools

Terraform
Ansible
CloudFormation
Jenkins
GitLab CI
CloudWatch
New Relic
Splunk

Job description

Responsibilities
  • Join the team supporting the Centers for Medicare & Medicaid Services (CMS) as it merges and modernizes its enterprise knowledge and data systems into a single, AI-driven platform, reducing manual effort, improving data accuracy, and enhancing transparency for stakeholders.
  • Keep the systems up and the users happy. Operate and tune AWS environments to meet infrastructure and application availability SLAs, even during transition and change.
  • Build observability that actually informs. Implement continuous monitoring, alerting, and dashboards using tools like AWS CloudWatch, New Relic, and Splunk, and establish performance baselines so you can spot degradation before users do.
  • Automate the toil. Write infrastructure-as-code (Terraform, Ansible) and support CI/CD pipelines (Jenkins) and containerized workloads (Docker) for repeatable, reliable deployments.
  • Define and track the numbers that matter. Set and monitor SLIs and SLOs, and produce performance, load/stress, and bottleneck reports that drive smarter decisions.
  • Optimize for performance, security, and cost. Use tools like AWS Trusted Advisor to find and act on improvement opportunities.
  • Support security and compliance modernization. Partner with the Security & Compliance SME to review vulnerability and security scans, feed continuous monitoring, and help advance the move toward a Continuous ATO (cATO) within a FISMA Moderate boundary (RMF, ARS, IS2P2).
  • Strengthen resilience. Help design and maintain disaster recovery and COOP continuity so the systems hold up against outages, incidents, and the unexpected.
  • Own incidents end to end. Drive response, run blameless post-mortems, and implement the preventative fixes that keep the same thing from happening twice.
Requirements
  • A bachelor’s degree in computer science, engineering, or a related field (or equivalent hands‑on experience).
  • 3–5 years of experience in site reliability, systems, or cloud engineering, with meaningful time spent in AWS environments.
  • Solid working knowledge of core AWS services, architecture, and best practices.
  • Hands‑on experience with infrastructure-as-code tools (Terraform, Ansible, or CloudFormation).
  • A good understanding of CI/CD pipelines and automation tools (Jenkins, GitLab CI, or similar).
  • Comfort scripting and automating in Python.
  • Familiarity with monitoring and observability tooling (CloudWatch, New Relic, Splunk, or comparable).
  • Strong problem‑solving instincts and the composure to work calmly under pressure.
  • Clear communication skills, with the ability to make complex technical concepts understandable.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

Skyward • Rockville (MD)

Hybrid
USD 112,000 - 150,000
Fully paid medical, dental, and vision insurance
401(k) with 4% employer contribution
Up to 4 weeks paid paternity and maternity leave
+2
Site Reliability Engineer
Site Reliability Engineer

TalentDome Staffing • United States

On-site
USD 140,000 - 210,000
Lead Site Reliability Engineer
Lead Site Reliability Engineer

Selby Jennings • Wilmington (NC)

On-site
USD 140,000 - 200,000
Site Reliability Engineer
Site Reliability Engineer

TechDigital Group • Houston (TX), Juno Beach (FL)

On-site
USD 120,000 - 180,000
Site Reliability Engineer
Site Reliability Engineer

Shrive Technologies LLC • Schaumburg (IL)

On-site
USD 110,000 - 170,000
Site Reliability Engineer
Site Reliability Engineer

Shrive Technologies • Schaumburg (IL)

On-site
USD 120,000 - 180,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Veloc Inc • Coppell (TX)

On-site
USD 140,000 - 190,000
Site Reliability Engineer
Site Reliability Engineer

TechDigital Group • Englewood Cliffs (NJ)

On-site
USD 110,000 - 130,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Jobtailor • Arlington (VA)

On-site
USD 140,000 - 200,000
Site Reliability Engineer
Site Reliability Engineer

Govcio LLC • Arlington (TX)

Hybrid
USD 230,000 - 250,000