Site Reliability Engineer SRE SecOps

Arkenstone

Menlo Park (CA)

On-site

USD 150,000 - 210,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Competitive Salary
Health and Wellness
401(k) Plan
Paid Time Off
Employee Assistance Program
Professional Development

Job summary

Arkenstone is seeking a highly capable Systems Reliability Engineer (SRE) to lead operational excellence across AWS, Azure, and GCP. You will own uptime, observability, and system resilience for critical services, collaborating with product owners, developers, and security teams.

Responsibilities include defining SLOs/SLAs, driving automation, and mentoring junior engineers while advocating ownership and accountability across the organization.

Qualifications

  • 5+ years in SRE, DevOps, or infrastructure engineering.
  • Hands-on multi-cloud experience (AWS and GCP) required.
  • Experience with Prometheus, Grafana, Datadog, or ELK.
  • Strong scripting and infrastructure automation skills.

Responsibilities

  • Design and own infrastructure reliability strategy across AWS, Azure, and GCP.
  • Develop logging, monitoring, and alerting systems.
  • Define and enforce SLOs/SLAs for critical systems.
  • Lead performance tuning, capacity planning, and DR.
  • Own incident management lifecycle from detection to root cause.
  • Automate deployment, scaling, and recovery workflows.
  • Contribute to infrastructure as code using Terraform, ARM, CloudFormation.
  • Mentor junior engineers and cross-functional partners.
  • Foster accountability and continuous improvement.

Skills

Python scripting
Bash scripting
Kubernetes
CI/CD pipelines
Observability tooling

Tools

Terraform
ARM templates
CloudFormation

Job description

About Us

At Arkenstone Defense, we empower defense tech startups with the tools, infrastructure, and compliance solutions they need to become successful prime contractors. Our mission is to remove barriers and help innovators grow - from day one to becoming a trusted prime for the U.S. Government.

We're early, we're lean, and we're building something that actually matters. The people who do well here aren't waiting to be told what to do; they see a gap and fill it.

Overview

We are seeking a highly motivated Systems Reliability Engineer (SRE) to lead the design and implementation of operational excellence across our multi-cloud environments. This role is central to ensuring the scalability, reliability, and performance of our products running in AWS, Azure, and GCP infrastructure.

As the lead SRE, you will own the uptime, observability, and system resilience for our critical services. This includes driving architecture decisions, automation practices, and incident response strategies - working closely with the product owner(s), developer teams, and security operations teams.

What You’ll Do
  • Design, implement, and own the infrastructure reliability strategy across AWS, Azure, and GCP
  • Champion observability by developing and maintaining effective logging, monitoring, and alerting systems
  • Define and enforce SLOs/SLAs for critical systems and services
  • Lead efforts in performance tuning, system hardening, capacity planning, and disaster recovery
  • Own the incident management lifecycle: from detection to postmortem and root cause analysis
  • Automate deployment, scaling, and recovery workflows to reduce manual toil
  • Contribute to infrastructure as code (Terraform, ARM templates, CloudFormation, etc.)
  • Act as a mentor and technical leader to junior engineers and cross-functional partners
  • Drive a culture of accountability, ownership, and continuous improvement
  • Perform any other related duties as required or assigned.
Requirements
  • 5+ years of experience in SRE, DevOps, or infrastructure engineering roles
  • Proven track record of operating large-scale systems in multi-cloud environments, with hands-on expertise in AWS and GCP
  • Strong knowledge of cloud-native architecture, container orchestration with Kubernetes, and CI/CD pipelines
  • Proficient in scripting (Python, Bash, etc.) and infrastructure automation tools (e.g., Terraform)
  • Experience with monitoring and observability platforms (e.g., Prometheus, Grafana, Datadog, ELK)
  • Excellent problem-solving skills with the ability to manage incidents and make sound decisions under pressure
  • Clear communicator capable of translating technical concepts to mixed audiences and participating in customer discussions
Who You Are
  • The Security Builder: You don't just consume security tools - you extend and improve them. You're energized by the opportunity to make analysts faster and compliance more automated.
  • Quality Over Speed: You write code that lasts. You push for clean interfaces, good documentation, and tests - even in a fast-moving environment.
  • Security-Minded Developer: You treat security as a first-class requirement, not an afterthought. You're comfortable reading CVEs, threat models, and compliance controls.
  • Cross-Functional Partner: You can work fluidly with security analysts, engineers, and compliance professionals, translating needs into reliable software.
Mission Alignment

We are a Defense-focused company supporting sensitive and cleared workforces. The Site Reliability Engineer (SRE) - SecOps will embrace our commitment to operational excellence, compliance rigor, and a world-class employee experience.

Physical Requirements
  • Prolonged periods of sitting at a desk and working on a computer
  • Must be able to lift up to 15 pounds at times
  • May require occasional travel to office locations or client sites
  • Ability to communicate effectively in written and verbal form
Benefits for working with us!

We are committed to supporting our employees both professionally and personally. Our robust benefits package is designed to promote your well-being, growth, and work-life balance

  • Competitive Salary: Recognizing your hard work with attractive compensation and rewarding excellence.
  • Health and Wellness Programs: Including medical, dental, & vision insurance options, along with mental health support & wellness initiatives.
  • Retirement Planning: Secure your future with our flexible 401(k) plan and matching company contributions.
  • Paid Time Off & Holidays: Generous PTO, sick leave, and holiday pay to help you recharge and enjoy life outside of work.
  • Employee Assistance Program: Confidential resources for personal and professional support.
  • Professional Development: Access to training, certifications, and continuing education to foster your career growth.

We are an Equal Opportunity Employer. We celebrate diversity and are committed to creating an inclusive environment for all employees. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex (including pregnancy, gender identity, and sexual orientation), national origin, age, disability, genetic information, veteran status, or any other characteristic protected under applicable law.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

VantageScore® • San Francisco (CA)

On-site
USD 150,000
Medical insurance
Dental insurance
401(k) plan
+1
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

OutSolve • Mission (KS)

Remote
USD 90,000 - 130,000
100% remote work environment
Competitive compensation
Professional development opportunities
+1
Senior Site Reliability Engineer (SRE)
Senior Site Reliability Engineer (SRE)

PrincePerelson and Associates • Salt Lake City (UT)

Hybrid
USD 120,000 - 190,000
Hybrid work schedule (4 days in-office
Medical, dental, vision
Retirement plan
+3
Staff Software Engineer (SWE) - SecOps
Staff Software Engineer (SWE) - SecOps

Arkenstone Defense • Menlo Park (CA)

On-site
USD 150,000 - 210,000
Competitive Salary
Health and Wellness Programs
Retirement Planning
+3
Senior DevOps/SRE Engineer
Senior DevOps/SRE Engineer

SEI • Chicago (IL)

Hybrid
USD 140,000 - 170,000
Comprehensive healthcare benefits
401(k) match
Paid Time Off (PTO)
+2
Site Reliability Engineer
Site Reliability Engineer

TalentDome Staffing • United States

On-site
USD 140,000 - 210,000
Sr SRE Automation Engineer
Sr SRE Automation Engineer

Compunnel, Inc. • Austin (TX), Northern (KY)

Hybrid
USD 130,000 - 180,000
Senior DevOps/SRE Engineer
Senior DevOps/SRE Engineer

VITG • Ellicott City (MD)

Hybrid
USD 90,000 - 120,000
401(k) with employer contribution
Medical/Dental/Vision insurance
Paid vacation (PTO)
Site Reliability Engineer
Site Reliability Engineer

Skyhigh Security • Frisco (TX)

Hybrid
USD 110,000 - 140,000
Retirement Plans
Medical, Dental and Vision Coverage
Paid Time Off
+2
Site Reliability Engineer
Site Reliability Engineer

Fortress Information Security, LLC • Patuxent Highland (MD)

Hybrid
USD 160,000 - 180,000
Medical, dental, and vision plans
401(k) match
Flexible Paid Time Off
+1