Site Reliability Engineer

Koitecc Solutions

Baltimore, Northern (MD, KY)

Hybrid

USD 108,000 - 195,000

Full time

2 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

FedRAMP security cognizance
Career development

Job summary

Leidos is seeking a Site Reliability Engineer (SRE) / Senior Cloud Engineer to join our Baltimore, Maryland office to modernize CMS Contact Center CRM within the CMS AWS Enclave and Pega Cloud for Government.

You will design, build, and operate highly available AWS infrastructure, architect secure interconnects, and lead incident response with robust observability and automation. The role requires SRE discipline, cloud security, and cross-team collaboration.

Qualifications

  • Bachelor’s degree with 6–8 years of cloud/SRE/infrastructure experience.
  • Experience building highly available AWS infrastructure per AWS Well-Architected Framework.
  • Experience with Infrastructure as Code, automation, and configuration management for cloud resources.
  • Ability to obtain Public Trust.

Responsibilities

  • Design, build, and operate highly available AWS infrastructure within CMS Enclave.
  • Architect secure multi-account VPC interconnects (CMS Enclave to Pega Cloud).
  • Support container and serverless architectures (Lambda, Glue) for data integration and API layers.
  • Apply SRE practices: SLOs, error budgets, reliability metrics aligned to SLAs.
  • Build observability with centralized logging, monitoring, and alerting.
  • Lead incident response and root-cause analysis for production issues.
  • Automate provisioning with Terraform, CloudFormation, Ansible.
  • Test Disaster Recovery, including multi-AZ/region failover.
  • Support capacity planning for 20,000 concurrent CSR sessions.
  • Ensure security monitoring and Zero Trust alignment across environments.
  • Develop CI/CD pipelines with automated testing and security scans.
  • Coordinate cloud connectivity for data migration pipelines (AWS Glue, S3, RDS).
  • Collaborate with Hybrid Cloud Team on provisioning and lifecycle management.
  • Minimize operational toil through automation and log analysis.
  • Document architecture and runbooks; mentor other engineers.
  • Communicate risks and dependencies to CMS leadership.

Skills

AWS infrastructure
Infrastructure as Code
Disaster Recovery design
Multi-Account VPC
FedRAMP security
GovCloud familiarity
Container/Serverless (Docker/Lambda)
Observability tooling

Education

Bachelor’s degree or higher

Tools

Terraform
CloudFormation
Ansible
PrivateLink
Direct Connect
API Gateway
AWS CloudWatch / Splunk / New Relic

Job description

The Digital Sector at Leidos currently has an opening for a Site Reliability Engineer (SRE) / Senior Cloud Engineer to work in our Baltimore, Maryland office. This is an exciting opportunity to use your experience helping the Center for Medicare and Medicaid Services (CMS) modernize its legacy Contact Center CRM platform within the CMS AWS Enclave and Pega Cloud for Government.

Primary Responsibilities
  • Design, build, and operate highly available AWS infrastructure within the CMS AWS Enclave (FedRAMP Moderate), applying AWS Well-Architected Framework best practices.
  • Architect secure, scalable multi-account VPC interconnectivity between the CMS AWS Enclave and Pega Cloud for Government (PCFG) in AWS GovCloud US-West, including PrivateLink, API Gateway, and Direct Connect.
  • Support container and serverless architectures (e.g., AWS Lambda, Glue) for data integration, batch processing, and API layers supporting the modernized CRM.
  • Apply Site Reliability Engineering (SRE) practices, including defining and tracking Service Level Objectives (SLOs), error budgets, and reliability metrics aligned to contract SLAs (e.g., ≥99.9% availability).
  • Build and maintain observability across the AWS and Pega environments, including centralized logging, monitoring, and alerting, to enable proactive detection of performance and availability issues.
  • Lead incident response and root-cause analysis for production issues, and long-term reliability improvements.
  • Automate infrastructure provisioning, configuration, and environment build-out using Infrastructure as Code (e.g., Terraform, CloudFormation, Ansible).
  • Design and test Disaster Recovery capability for cloud-based workloads, including backup, failover, and Multi-AZ/Multi-Region resilience.
  • Support performance testing and capacity planning to validate the platform's ability to scale to 20,000 concurrent CSR sessions and peak Open Enrollment Period (OEP) volumes.
  • Support continuous security monitoring, vulnerability remediation, and Zero Trust alignment across the AWS and Pega environments.
  • Partner with the DevOps Lead/Configuration Manager to build and maintain CI/CD pipelines, ensuring automated testing, security scanning, and deployment across all SDLC environments.
  • Support cloud connectivity and data movement for the AWS-based data migration pipeline (e.g., AWS Glue, S3, RDS/Aurora PostgreSQL) between legacy Siebel and the modernized CRM.
  • Coordinate with the CMS Hybrid Cloud Team on cloud environment provisioning, patching, and lifecycle management activities.
  • Manage release coordination and change windows in support of OEP blackout periods and other critical operational periods, minimizing risk of service disruption.
  • Collaborate with the Release Train Engineer, Solution Architect, and Agile delivery teams to align infrastructure readiness with sprint and PI planning.
  • Support integration of Genesys Cloud CX infrastructure and telephony/chat channels with the modernized CRM environment.
  • Continuously identify opportunities to reduce operational toil through automation of repetitive tasks, log analysis, and routine operational activities.
  • Document cloud architecture, operational runbooks, and disaster recovery procedures to support the Transition-Out Plan and audit readiness.
  • Provide technical mentoring and knowledge-sharing to other engineers on cloud architecture, automation, and reliability engineering best practices.
  • Communicate technical status, risks, and dependencies to CMS leadership, and Leidos management.
  • Support requirements traceability and technical documentation related to infrastructure and integration architecture.
  • Actively participate in planning sessions, requirements gathering activities, design sessions, Agile sessions, and other events supporting the CRM modernization effort.
Required Qualifications:
  • Bachelor’s degree and a minimum of 6-8 years of relevant experience in cloud engineering, site reliability engineering, or infrastructure, or an equivalent combination of education and experience
  • Experience building highly available AWS infrastructure based on industry best practices and the AWS Well-Architected Framework
  • Experience with Infrastructure as Code, automation, and configuration management of cloud-based resources
  • Experience designing Disaster Recovery for cloud-based workloads, including Multi-AZ/Multi-Region resilience
  • Ability to obtain Public Trust
Preferred Qualifications:
  • AWS Solutions Architect or SysOps Administrator Certification
  • Experience with AWS GovCloud and FedRAMP-authorized cloud environments
  • Experience supporting Pega Cloud for Government (PCFG) or similar SaaS platform connectivity (e.g., AWS PrivateLink, VPC peering)
  • Familiarity with container (Docker, ECS, EKS) and serverless (Lambda) architectures
  • Experience with observability/monitoring tooling (e.g., Splunk, New Relic, CloudWatch) and incident response practices
  • Experience supporting federal contact center or other 24x7 mission-critical government systems
  • Agile delivery experience

If you're looking for comfort, keep scrolling. At Leidos, we outthink, outbuild, and outpace the status quo - because the mission demands it. We're not hiring followers. We're recruiting the ones who disrupt, provoke, and refuse to fail. Step 10 is ancient history. We're already at step 30 - and moving faster than anyone else dares.

Pay Range

Pay Range $107,900.00 - $195,050.00

The Leidos pay range for this job level is a general guideline only and not a guarantee of compensation or salary. Additional factors considered in extending an offer include (but are not limited to) responsibilities of the job, education, experience, knowledge, skills, and abilities, as well as internal equity, alignment with market data, applicable bargaining agreement (if any), or other law.

About Leidos

Leidos is an industry and technology leader serving government and commercial customers with smarter, more efficient digital and mission innovations. Headquartered in Reston, Virginia, with 47,000 global employees, Leidos reported annual revenues of approximately $16.7 billion for the fiscal year ended January 3, 2025. For more information, visit www.Leidos.com.

Pay and Benefits

Pay and benefits are fundamental to any career decision. That's why we craft compensation packages that reflect the importance of the work we do for our customers. Employment benefits include competitive compensation, Health and Wellness programs, Income Protection, Paid Leave and Retirement. More details are available at www.leidos.com/careers/pay-benefits.

Commitment to Non-Discrimination

All qualified applicants will receive consideration for employment without regard to sex, race, ethnicity, age, national origin, citizenship, religion, physical or mental disability, medical condition, genetic information, pregnancy, family structure, marital status, ancestry, domestic partner status, sexual orientation, gender identity or expression, veteran or military status, or any other basis prohibited by law. Leidos will also consider for employment qualified applicants with criminal histories consistent with relevant laws.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

Leidos Inc • Baltimore (MD)

On-site
USD 108,000 - 195,000
Site Reliability Engineer
Site Reliability Engineer

Leidos • Baltimore (MD)

On-site
USD 108,000 - 195,000
DevOps Lead/Configuration Manager
DevOps Lead/Configuration Manager

Koitecc Solutions • Baltimore (MD), Northern (KY)

Hybrid
USD 108,000 - 195,000
DevOps Lead/Configuration Manager
DevOps Lead/Configuration Manager

Leidos Inc • Baltimore (MD)

On-site
USD 108,000 - 195,000
Senior Database Administrator (DBA)
Senior Database Administrator (DBA)

Leidos • Baltimore (MD)

On-site
USD 108,000 - 195,000
Senior Database Administrator (DBA)
Senior Database Administrator (DBA)

Careerwebsite • Baltimore (MD), Northern (KY)

Hybrid
USD 108,000 - 195,000
Senior Database Administrator (DBA)
Senior Database Administrator (DBA)

Leidos Inc • Baltimore (MD)

On-site
USD 108,000 - 195,000
DevOps Lead/Configuration Manager
DevOps Lead/Configuration Manager

Leidos • Baltimore (MD)

On-site
USD 108,000 - 195,000
Pega Lead System Architect
Pega Lead System Architect

Leidos Inc • Baltimore (MD)

On-site
USD 131,000 - 237,000
SAFe Release Train Engineer
SAFe Release Train Engineer

Leidos Inc • Baltimore (MD)

On-site
USD 108,000 - 195,000