Senior Software Engineer- Site Reliability Engineering (SRE)

Noctua Technology

United States

Remote

USD 149,000 - 202,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Noctua Technology is seeking an experienced Senior Site Reliability Engineer (SRE) to lead reliability strategy for large-scale cloud-native systems. You will define SLOs, architect IaC, and drive automation to reduce toil while ensuring security and high availability.

In this role, you will champion modern cloud architectures, mentor teams, and coordinate incident response with blameless postmortems. Remote-friendly with US-based candidates eligible for clearance.

Qualifications

  • 5+ years of SRE/Cloud engineering experience.
  • Strong software engineering skills for automation.
  • Deep experience with IaC (Terraform, CloudFormation).
  • Experience with Docker and Kubernetes.
  • Networking and cloud security knowledge.
  • Scripting in Python, Bash, or Go.
  • Experience with CI/CD and DevOps.
  • Ability to influence decisions across teams.

Responsibilities

  • Drive SLIs and SLOs across services and platforms.
  • Design IaC solutions for large-scale environments.
  • Implement containerized and serverless architectures.
  • Build reliable CI/CD pipelines for deployments.
  • Monitor, alert, and log to detect performance issues.
  • Lead toil reduction and automation projects.
  • Coordinate high-severity incident response and postmortems.
  • Collaborate with development teams on reliability.

Skills

SRE/Cloud Engineering
Docker
Kubernetes
Python/Bash/Go
CI/CD
Observability/Monitoring
SLOs/SLAs
Automation/Infra as Code

Education

Bachelor’s or advanced degree in Computer Science or related field

Tools

Terraform
CloudFormation

Job description

The Site Reliability Engineering discipline at Noctua Technology, LLC is a strategic force driving digital transformation. We treat operations as a software engineering challenge, focusing on the seamless integration, scalability, and long-term reliability of cloud native systems. Our SREs don’t just manage infrastructure; they build it using Infrastructure as Code (IaC), monitor it through advanced observability stacks, and protect it by engineering for failure. We work closely with clients to bridge the gap between development and operations.

We are seeking a highly experienced and autonomous Senior Site Reliability Engineer (SRE) to join our dynamic team. As a technical leader, you will define the strategy and apply advanced software engineering principles to operations, focusing on the architecture, reliability, and long-term performance of large-scale production systems. You will play a crucial role in reducing toil through automation, defining and monitoring Service Level Objectives (SLOs), and implementing best practices for system stability and incident response. This role requires working with modern cloud technologies to ensure the high availability and efficiency of applications and infrastructure.

  • Location : Primarily Remote. Candidates must be based in CA or DC Metro Area for proximity to project and client teams.
  • Security Clearance Requirement: Applicants must be US citizens and eligible to obtain and maintain an active Secret security clearance or above.
Key Responsibilities
Site Reliability Engineering
  • Drive the definition and adoption of SLIs and SLOs across multiple services or entire platforms, ensuring alignment with business goals.
  • Design and architect Infrastructure as Code (IaC) solutions for large-scale, complex environments, establishing standards and best practices.
  • Implement and manage containerized and serverless architectures using Docker, Kubernetes, and cloud-native services, focusing on performance and error budgets.
  • Build and maintain reliable and self-healing CI/CD pipelines to automate deployments and improve development workflows.
Toil Reduction and Incident Management
  • Implement and refine comprehensive monitoring, alerting, and logging to detect and address performance and availability issues proactively.
  • Lead the strategic effort to eliminate toil, identifying and championing major automation projects that deliver significant organizational efficiency.
  • Lead high-severity incident response and coordinate blameless postmortems for major outages, driving the resulting remediation and systemic improvements.
Testing and Service Resiliency
  • Implement cloud security best practices, including identity and access management (IAM), encryption, and compliance controls.
  • Proactively identify and address system weaknesses and ensure performance under stress.
  • Support disaster recovery and high availability strategies through backup and failover planning.
Collaboration and Knowledge Sharing
  • Serve as a primary SRE liaison for development teams, influencing application architecture and design to meet reliability and scalability targets from inception.
  • Create and maintain documentation for cloud architectures, deployment processes, and best practices.
  • Contribute to internal knowledge-sharing initiatives, ensuring continuous learning within the team.
Stakeholder Communication
  • Act as a subject matter expert and trusted advisor to clients and internal leadership on cloud infrastructure, reliability strategy, and Service Level Agreement (SLA) negotiations.
  • Act on client feedback to refine and enhance cloud solutions.
  • Conduct training and knowledge-sharing sessions to help clients manage their cloud environments effectively.
Continuous Learning and Innovation
  • Stay updated on the latest developments in cloud infrastructure and technology trends.
  • Drive innovation by proposing and implementing new techniques and technologies.
Qualifications
  • 5+ years of experience in site reliability engineering, cloud engineering, or related fields.
  • Strong software engineering skills with an emphasis on writing clean, modular, and maintainable code, specifically for automation and system management.
  • Deep experience with Infrastructure as Code (IaC) tools like Terraform or CloudFormation.
  • Deep experience with containerization and orchestration tools like Docker and Kubernetes.
  • Deep knowledge of networking concepts, cloud security best practices, and identity management.
  • Experience with programming or scripting languages such as Python, Bash, or Go.
  • Experience with CI/CD pipelines and DevOps methodologies.
  • Strong problem-solving skills and the ability to troubleshoot complex cloud environments.
  • Demonstrated ability to influence technical decision-making across organizational boundaries
Preferred qualifications:
  • Bachelor’s or advanced degree in Computer Science or a related field.
  • Any of the below cloud certifications:
    • Google Cloud Professional Cloud Architect
    • Google Cloud Professional Cloud DevOps Engineer
    • AWS Certified Solutions Architect
    • AWS Certified Developer
    • AWS Certified SysOps Administrator
  • CompTIA Security+ certification or an equivalent DoD 8140/8570 IAT Level II baseline certification.
Salary Range : $149,400 - $202,000
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Software Engineer- Site Reliability Engineering (SRE)
Senior Software Engineer- Site Reliability Engineering (SRE)

Noctua Technology • Virginia (MN), California (MO), Washington

On-site
USD 149,400 - 202,000
Software Engineer- Site Reliability Engineering (SRE)
Software Engineer- Site Reliability Engineering (SRE)

Noctua Technology • United States

Remote
USD 107,000 - 178,000
Software Engineer- Site Reliability Engineering (SRE)
Software Engineer- Site Reliability Engineering (SRE)

Noctua Technology • Virginia (MN), California (MO), Washington

On-site
USD 106,500 - 177,500
Senior Site Reliability Engineer
Senior Site Reliability Engineer

GovCIO • Arlington (VA)

On-site
USD 210,000 - 230,000
Remote Senior SRE: Cloud Reliability & Automation Leader
Remote Senior SRE: Cloud Reliability & Automation Leader

Noctua Technology • United States

Remote
USD 149,000 - 202,000
Site Reliability Engineer
Site Reliability Engineer

TalentDome Staffing • United States

On-site
USD 140,000 - 210,000
Senior SRE Engineer: Remote Cloud Reliability & Automation
Senior SRE Engineer: Remote Cloud Reliability & Automation

Noctua Technology • Virginia (MN), California (MO), Washington

Remote
USD 149,400 - 202,000
Remote SRE Engineer: Cloud Reliability & Automation
Remote SRE Engineer: Cloud Reliability & Automation

Noctua Technology • United States

Remote
USD 107,000 - 178,000
Senior DevOps/SRE Engineer
Senior DevOps/SRE Engineer

SEI • Chicago (IL)

On-site
USD 140,000 - 170,000
Comprehensive healthcare benefits
401(k) match
Paid Time Off (PTO)
+2
Senior DevOps/SRE Engineer
Senior DevOps/SRE Engineer

VITG • Ellicott City (MD)

On-site
USD 90,000 - 120,000
401(k) with employer contribution
Medical/Dental/Vision insurance
Paid vacation (PTO)