Software Engineer- Site Reliability Engineering (SRE)

Noctua Technology

Virginia, California, Washington (MN, MO, District of Columbia)

Remote

USD 106,500 - 177,500

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Noctua Technology is seeking a Software Engineer specializing in Site Reliability Engineering (SRE). This role focuses on driving digital transformation while ensuring the reliability, scalability, and performance of production systems. As part of a dynamic team, you will automate processes and enhance system reliability using cutting-edge cloud technologies.

The position is primarily remote, requiring candidates to be based in California or the DC Metro Area, and offers a salary range between $106,500 and $177,500.

Qualifications

  • 1-5 years of experience in site reliability engineering, cloud engineering, or related fields.
  • Experience with programming or scripting languages such as Python, Bash, or Go.
  • Familiarity with DevOps methodologies.

Responsibilities

  • Define, measure, and report on Service Level Indicators (SLIs) and Service Level Objectives (SLOs).
  • Develop and deploy Infrastructure as Code (IaC) using Terraform or CloudFormation.
  • Implement and manage containerized architectures using Docker and Kubernetes.

Skills

Strong software engineering skills
Proficiency in Infrastructure as Code (IaC) tools
Experience with Docker and Kubernetes
Knowledge of cloud security best practices
Programming languages like Python, Bash, or Go
Problem-solving skills
Effective communication skills

Education

Bachelor's or advanced degree in Computer Science or a related field

Tools

Terraform
CloudFormation
CI/CD pipelines

Job description

Software Engineer- Site Reliability Engineering (SRE)

DC, MD, VA, CA

The Site Reliability Engineering discipline at Noctua Technology, LLC is a strategic force driving digital transformation. We treat operations as a software engineering challenge, focusing on the seamless integration, scalability, and long-term reliability of cloud native systems. Our SREs don’t just manage infrastructure; they build it using Infrastructure as Code (IaC), monitor it through advanced observability stacks, and protect it by engineering for failure. We work closely with clients to bridge the gap between development and operations.

We are seeking a motivated Site Reliability Engineer (SRE) to join our dynamic team. As a key contributor, you will apply software engineering principles to operations, focusing on the reliability, scalability, and performance of production systems. You will play a crucial role in reducing toil through automation, defining and monitoring Service Level Objectives (SLOs), and implementing best practices for system stability and incident response. This role requires working with modern cloud technologies to ensure the high availability and efficiency of applications and infrastructure.

  • Location: Primarily Remote. Candidates must be based in CA or DC Metro Area for proximity to project and client teams.
  • Security Clearance Requirement: Applicants must be US citizens and eligible to obtain and maintain an active Secret security clearance or above.
Key Responsibilities
Site Reliability Engineering
  • Define, measure, and report on Service Level Indicators (SLIs) and Service Level Objectives (SLOs) to ensure system reliability and uptime.
  • Develop and deploy Infrastructure as Code (IaC) using Terraform, CloudFormation, or similar tools, with an emphasis on repeatability and change management.
  • Implement and manage containerized and serverless architectures using Docker, Kubernetes, and cloud-native services, focusing on performance and error budgets.
  • Build and maintain reliable and self-healing CI/CD pipelines to automate deployments and improve development workflows.
Toil Reduction and Incident Management
  • Implement and refine comprehensive monitoring, alerting, and logging to detect and address performance and availability issues proactively.
  • Eliminate toil by extensively automating operational tasks, including provisioning, patching, and deployments, using scripting and configuration management tools such as Python, Bash, or Ansible.
  • Conduct post-incident reviews (blameless postmortems) to drive continuous improvement in system reliability and operational processes.
Testing and Service Resiliency
  • Implement cloud security best practices, including identity and access management (IAM), encryption, and compliance controls.
  • Proactively identify and address system weaknesses and ensure performance under stress.
  • Support disaster recovery and high availability strategies through backup and failover planning.
Collaboration and Knowledge Sharing
  • Collaborate with development teams to improve the operability and production readiness of applications from design through deployment.
  • Create and maintain documentation for cloud architectures, deployment processes, and best practices.
  • Contribute to internal knowledge-sharing initiatives, ensuring continuous learning within the team.
Stakeholder Communication
  • Provide technical guidance and support to clients and internal teams on cloud infrastructure and reliability best practices, with a focus on defining Service Level Agreements (SLAs).
  • Act on client feedback to refine and enhance cloud solutions.
  • Conduct training and knowledge-sharing sessions to help clients manage their cloud environments effectively.
Continuous Learning and Innovation
  • Stay updated on the latest developments in cloud infrastructure and technology trends.
  • Drive innovation by proposing and implementing new techniques and technologies.
Qualifications
  • 1-5 years of experience in site reliability engineering, cloud engineering, or related fields.
  • Strong software engineering skills with an emphasis on writing clean, modular, and maintainable code, specifically for automation and system management.
  • Proficiency in Infrastructure as Code (IaC) tools like Terraform or CloudFormation.
  • Experience with containerization and orchestration tools like Docker and Kubernetes.
  • Knowledge of networking concepts, cloud security best practices, and identity management.
  • Experience with programming or scripting languages such as Python, Bash, or Go.
  • Familiarity with CI/CD pipelines and DevOps methodologies.
  • Strong problem-solving skills and the ability to troubleshoot complex cloud environments.
  • Effective communication skills and a willingness to learn and collaborate.
Preferred Qualifications
  • Bachelor's or advanced degree in Computer Science or a related field.
  • Any of the below cloud certifications:
    • Google Cloud Professional Cloud Architect
    • Google Cloud Professional Cloud DevOps Engineer
    • AWS Certified Solutions Architect
    • AWS Certified Developer
    • AWS Certified SysOps Administrator
    • Azure Solutions Architect Expert
  • CompTIA Security+ certification or an equivalent DoD 8140/8570 IAT Level II baseline certification.

Salary Range: $106,500 - $177,500

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Software Engineer- Site Reliability Engineering (SRE)
Senior Software Engineer- Site Reliability Engineering (SRE)

Noctua Technology • Virginia (MN), California (MO), Washington

Remote
USD 149,000 - 202,000
Site Reliability Engineer
Site Reliability Engineer

TalentDome Staffing • United States

On-site
USD 140,000 - 210,000
Senior SRE Engineer: Remote Cloud Reliability & Automation
Senior SRE Engineer: Remote Cloud Reliability & Automation

Noctua Technology • Virginia (MN), California (MO), Washington

Remote
USD 149,000 - 202,000
Cloud Infrastructure Site Reliability Engineer
Cloud Infrastructure Site Reliability Engineer

Robotics Prcocess Automation, LLC • Berkeley Heights (NJ)

On-site
Site Reliability Engineer
Site Reliability Engineer

SRE • Puerto Rico

Hybrid
USD 120,000 - 180,000
Senior DevOps/SRE Engineer
Senior DevOps/SRE Engineer

VITG • Ellicott City (MD)

Hybrid
USD 90,000 - 120,000
401(k) with employer contribution
Medical/Dental/Vision insurance
Paid vacation (PTO)
Site Reliability Engineer
Site Reliability Engineer

Govcio LLC • Arlington (TX)

Hybrid
USD 230,000 - 250,000
Sr. Cloud Operations Reliability Engineer (SRE)
Sr. Cloud Operations Reliability Engineer (SRE)

NextGen Healthcare • Georgia

On-site
USD 140,000 - 210,000
Site Reliability Engineer
Site Reliability Engineer

Tata Consultancy Services • Scottsdale (AZ)

On-site
USD 100,000 - 110,000
Discretionary Annual Incentive
Comprehensive Medical Coverage
Family Support: Parental Leaves
+4
Remote SRE Engineer — Cloud Reliability & Automation
Remote SRE Engineer — Cloud Reliability & Automation

Noctua Technology • Virginia (MN), California (MO), Washington

Remote
USD 106,500 - 177,500