Site Reliability Engineer

Skyward

Rockville (MD)

On-site

USD 110,000 - 170,000

Full time

10 hours ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Medical, dental, vision insurance
401K with 4% employer contribution
Company provided laptop
Paid time off and holidays
Professional development budget

Job summary

Skyward in Rockville, MD is seeking a Site Reliability Engineer to join a mission-driven team modernizing CMS data platforms with AI-driven solutions.

You will manage AWS environments, build observability, and implement IaC with Terraform/Ansible, Jenkins, and Docker, ensuring reliability and cost efficiency. We value collaboration, continuous learning, and a flexible work approach with remote options.

Qualifications

  • A bachelor’s degree in computer science, engineering, or a related field (or equivalent hands-on experience).
  • 3-5 years of experience in site reliability, systems, or cloud engineering, with meaningful time spent in AWS environments.
  • Solid working knowledge of core AWS services, architecture, and best practices.
  • Hands-on experience with infrastructure-as-code tools (Terraform, Ansible, or CloudFormation).
  • A good understanding of CI/CD pipelines and automation tools (Jenkins, GitLab CI, or similar).
  • Comfort scripting and automating in Python.
  • Familiarity with monitoring and observability tooling (CloudWatch, New Relic, Splunk, or comparable).

Responsibilities

  • Join CMS team as it merges and modernizes enterprise knowledge and data systems into a single AI-driven platform.
  • Operate and tune AWS environments to meet infrastructure and application availability SLAs, even during transition.
  • Build observability with dashboards and alerts; establish performance baselines to spot degradation early.
  • Write infrastructure-as-code and support CI/CD pipelines and Docker workloads for repeatable deployments.
  • Define and track SLIs/SLOs, and produce performance/load/bottleneck reports for smarter decisions.
  • Optimize for performance, security, and cost using AWS Trusted Advisor.
  • Support security/compliance modernization and move toward continuous ATO within RMF-boundary.
  • Strengthen resilience with disaster recovery and COOP planning.
  • Own incidents end to end; drive blameless post-mortems and preventative fixes.

Skills

Strong problem-solving
Clear communication
Adaptability

Education

Bachelor's degree in CS/Engineering/related field

Tools

Terraform
Ansible
CloudFormation
Jenkins
GitLab CI
Python
Docker
CloudWatch
New Relic
Splunk

Job description

We are Skyward.

That is, a love for people, for improvement, for human advancement through information technology. We are a people-centered business with a desire to serve others. We are diverse and unified; creative and collaborative; a collection of complementary, not competing talents. And though on the surface we remain relaxed, beneath, a torrent of energy links us to our civic tech mission.

We stand by our values, and we won’t compromise on any of them.

Integrity:

We’re conscientious, intentional, and empathetic. Our words and actions align. That’s our character. Please don’t ask us to play another part, we’re poor actors.

Compassionate:

If we may borrow a quote from Theodore Roosevelt: “No one cares how much you know until they know how much you care.” Because our team is thoughtful and supportive, caring deeply for each other, our clients, and our work, this comes naturally.

Inquisitive:

We remain students by failing openly and turning lessons into solutions.

Unconventional:

For us, life isn’t what happens outside of work. Work happens inside of life and our culture erases the line often dividing the two.

Authentic:

Made possible only because we embody the values listed above. We’re relaxed and fun yet intensely curious and driven. Team members are placed with thought, care, and precision to ensure that Trust, Truth, and Transparency continue to represent our brand.

Because of that, we continue Onward, Upward, and Skyward.

We need an SRE.

Do you have a real feel for how distributed systems behave, and a knack for tracking down the network, infrastructure, or pipeline issue everyone else gave up on? Are you comfortable in the cloud, fluent in CI/CD, and the type who believes an alert should mean something and a dashboard should tell a story? If you love keeping complex systems healthy, fast, and quietly reliable, then

Come join us if you're motivated to learn from others, to learn from mistakes, to be part of a future-looking and growth-oriented team.

Let's go Skyward together.

What you’ll do:
  • Join the team supporting the Centers for Medicare & Medicaid Services (CMS) as it merges and modernizes its enterprise knowledge and data systems into a single, AI-driven platform, reducing manual effort, improving data accuracy, and enhancing transparency for stakeholders
  • Keep the systems up and the users happy. Operate and tune AWS environments to meet infrastructure and application availability SLAs, even during transition and change
  • Build observability that actually informs. Implement continuous monitoring, alerting, and dashboards using tools like AWS CloudWatch, New Relic, and Splunk, and establish performance baselines so you can spot degradation before users do
  • Automate the toil. Write infrastructure-as-code (Terraform, Ansible) and support CI/CD pipelines (Jenkins) and containerized workloads (Docker) for repeatable, reliable deployments
  • Define and track the numbers that matter. Set and monitor SLIs and SLOs, and produce performance, load/stress, and bottleneck reports that drive smarter decisions
  • Optimize for performance, security, and cost. Use tools like AWS Trusted Advisor to find and act on improvement opportunities
  • Support security and compliance modernization. Partner with the Security & Compliance SME to review vulnerability and security scans, feed continuous monitoring, and help advance the move toward a Continuous ATO (cATO) within a FISMA Moderate boundary (RMF, ARS, IS2P2)
  • Strengthen resilience. Help design and maintain disaster recovery and COOP continuity so the systems hold up against outages, incidents, and the unexpected
  • Own incidents end to end. Drive response, run blameless post-mortems, and implement the preventative fixes that keep the same thing from happening twice
What we’d like you to have:
  • A bachelor’s degree in computer science, engineering, or a related field (or equivalent hands-on experience)
  • 3-5 years of experience in site reliability, systems, or cloud engineering, with meaningful time spent in AWS environments
  • Solid working knowledge of core AWS services, architecture, and best practices
  • Hands-on experience with infrastructure-as-code tools (Terraform, Ansible, or CloudFormation)
  • A good understanding of CI/CD pipelines and automation tools (Jenkins, GitLab CI, or similar)
  • Comfort scripting and automating in Python
  • Familiarity with monitoring and observability tooling (CloudWatch, New Relic, Splunk, or comparable)
  • Strong problem-solving instincts and the composure to work calmly under pressure
  • Clear communication skills, with the ability to make complex technical concepts understandable
What would blow us away:
  • You’ve previously worked with CMS
  • You have experience working in AI, NLP, or LLM-driven environments
  • You have all the AWS certifications and the real-world scars that come with them
What we offer you:
  • Medical, dental, vision insurance
  • 15 days of paid leave
  • 7 days of sick leave
  • 2 days bereavement leave
  • 11 paid Federal holidays
  • Up to 40 hours for jury duty
  • 401K with 4% employer contribution (and no vesting period)
  • Up to 4 weeks of paid paternity and maternity leave
  • Company provided laptop
  • $5,000 per year for professional development
  • $600 per year for technical supplies and equipment
  • $2,000 referral bonus
  • Life and disability insurance
  • HSA and FSA
  • Legal Shield and ID Shield Voluntary Benefits
  • Opportunity to work in a collaborative, motivated team focused on modernizing government services with cutting-edge technology and innovative solutions. Who says government work can't be exciting!

At Skyward, we support flexible working hours and remote opportunities to help maintain a healthy work-life balance for all employees.

Offers of employment with Skyward are contingent upon acceptable results of a background investigation.

Applicants must have the ability to obtain and maintain a Public Trust security clearance due to the nature of our work as a government contractor.

We may use artificial intelligence (AI) tools to support parts of the hiring process, such as reviewing applications, analyzing resumes, or assessing responses and identifying potential inconsistencies or verification signals in application materials based on available information. These tools assist our recruitment team but do not replace human judgment. Final hiring decisions are ultimately made by humans. If you would like more information about how your data is processed, please contact us.

We believe great work deserves great pay. That’s why we ensure our compensation is not only competitive but also fair and transparent, as required by Maryland law. Expect a salary that matches your skills, experience, and the value you bring to the table — because you’re worth it!

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

Skyward IT Solutions, LLC • Rockville (MD)

Hybrid
USD 112,000 - 150,000
Medical, dental, vision insurance (fully paid for employees)
401(k) with 4% employer contribution
Up to 4 weeks of paid parental leave
+1
Technical Analyst
Technical Analyst

Skyward • United States

On-site
USD 65,000 - 105,000
Medical, dental, vision insurance
401K with employer contribution
Company provided laptop
+1
Solutions Architect
Solutions Architect

Skyward IT Solutions, LLC • Rockville (MD)

Hybrid
USD 150,000 - 190,000
Medical, dental, vision insurance (fully paid for employees)
401K with 4% employer contribution
Up to 40 hours for jury duty
+2
Cloud Engineer
Cloud Engineer

Skyward IT Solutions, LLC • Rockville (MD)

Hybrid
USD 115,000 - 145,000
Medical, dental, vision insurance
15 days paid leave
7 days sick leave
+3
Reporting & Technical Writer Skyward IT Solutions, LLC · USD 62k-75k/yr Rockville, MD, US Workplace 12 hours ago
Reporting & Technical Writer Skyward IT Solutions, LLC · USD 62k-75k/yr Rockville, MD, US Workplace 12 hours ago

Content Creators • Rockville (MD), Northern (KY)

Hybrid
USD 70,000 - 100,000
Medical, dental, vision insurance
15 days paid leave
7 days sick leave
+2
Data Integrity Specialist
Data Integrity Specialist

Sky Solutions • Tysons (VA)

On-site
USD 90,000 - 150,000
Medical coverage
Time Off & Work-Life Balance
Education stipend
+2
Sr Site Reliability Engineer (US Federal)
Sr Site Reliability Engineer (US Federal)

Workday • Reston (VA)

On-site
USD 147,400 - 221,200
Flexible work schedule
Health benefits
Employee Referral program
Senior Software Development Engineer (US Federal)
Senior Software Development Engineer (US Federal)

Workday • Reston (VA)

Hybrid
USD 164,000 - 264,000
Staff Site Reliability Engineer, Government
Staff Site Reliability Engineer, Government

Talanto • Northern (KY)

Hybrid
USD 13,000 - 17,000
RSUs
ESPP
Flexible time off
+3
Application Security Engineer
Application Security Engineer

WellSky • Overland Park (KS)

On-site
USD 70,000 - 110,000
Excellent medical benefits
Mental Health support through EAP
Generous PTO + 13 paid holidays
+2