Senior Site Reliability Engineer -

United States Digital Space LLC

Bellevue (WA)

On-site

USD 147,000 - 202,000

Full time

3 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

United States Digital Space LLC is seeking a highly skilled Senior Site Reliability Engineer to join our team. This role blends software engineering and systems administration to build and maintain reliable, scalable, and secure infrastructure for our security SaaS offerings.

You will automate toil, participate in on-call, and collaborate with security, data, and development teams. In-person onboarding at our Toronto office during the first week is required, with travel as needed.

Qualifications

  • Strong coding skills and production-grade coding ability.
  • IaC experience with Terraform for provisioning cloud infra.
  • Familiarity with modern CI/CD practices, especially Spinnaker.
  • Containerization expertise and managing large-scale clusters with Kubernetes.
  • Experience with database migrations using Flyway.

Responsibilities

  • Platform & Reliability: Design, build, and maintain core infrastructure ensuring high availability and scalability for security SaaS offerings.
  • Automation: Develop robust automation to eliminate toil across environments.
  • Security & Compliance: Embed security-first mindset and ensure compliance with standards.
  • Incident Response: Participate in on-call rotations and lead root-cause analysis.
  • Collaboration: Work with development, data science, and security teams on architectural decisions.

Skills

Strong coding skills
IaC Terraform
CI/CD Spinnaker
Containerization Kubernetes
Database migrations Flyway
Snowflake data systems
AI/ML experience
Problem solving

Tools

Terraform
Spinnaker
Kubernetes
Flyway
Snowflake

Job description

**Secure Every Identity, from AI to Human

**Identity is the key to unlocking the potential of AI. the company secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence.

This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk.

Senior Site Reliability Engineer (SRE) - Security and Data Systems

Our company is seeking a highly skilled Senior Site Reliability Engineer to join our team. We are a SaaS company specializing in securing large-scale systems. This role is a blend of software engineering and systems administration, where you'll be responsible for building and maintaining highly reliable, scalable, and secure infrastructure. You will be a key contributor, applying your expertise to automate manual processes and proactively solve complex problems before they become incidents, handling incidents, and includes on-call shifts.

Responsibilities
  • Platform & Reliability: Design, build, and maintain the core infrastructure that underpins our security SaaS offerings, ensuring high availability, performance, and scalability. This includes building and operating the tooling for our Snowflake data systems.
  • Automation: Develop robust automation using code to eliminate toil and ensure consistency across our environments. You'll be a key driver in automating everything from infrastructure provisioning to application deployment and incident response.
  • Security & Compliance: Work closely with our security teams to embed a security-first mindset into all our processes and infrastructure. You will be responsible for ensuring our systems and data platforms are compliant with industry standards.
  • Incident Response: Participate in on-call rotations and be a primary responder for critical incidents, leading root cause analysis and implementing preventative measures to ensure issues don't recur.
  • Collaboration: Partner with development, data science, and security teams to provide expert guidance on architectural decisions, best practices, and the implementation of new services.
Key Skills & Qualifications
  • Strong Coding Skills: You are a developer at heart and are comfortable writing production-level code to solve complex operational challenges.
  • Infrastructure as Code (IaC): Deep experience with Terraform for provisioning and managing cloud infrastructure and services.
  • Continuous Delivery: Familiarity with modern CI/CD practices and tools, particularly Spinnaker, to automate and standardize our release pipelines.
  • Containerization & Orchestration: Expertise in container technologies and hands-on experience managing large-scale, production-ready clusters with Kubernetes.
  • Database Migrations: Experience with database schema management tools like Flyway for safely and reliably handling database changes.
  • Data Systems: Direct experience with large-scale data systems, specifically with the Snowflake platform.
  • AI/ML Experience (a plus): Experience or a strong interest in AI/ML, particularly how these technologies can be applied to improve reliability, security, and operational efficiency (e.g., AIOps, predictive analysis).
  • Problem-Solving: Excellent analytical and problem-solving skills with a proactive approach to identifying and addressing potential issues.

*This role requires in-person onboarding and travel to our Toronto, Canada, Office during the first week of employment.*

Below is the annual base salary range for candidates located in California (excluding San Francisco Bay Area), Colorado, Illinois, New York and Washington. Your actual base salary will depend on factors such as your skills, qualifications, experience, and work location. In addition, the company offers equity (where applicable), bonus, and benefits, including health, dental and vision insurance, 401(k), flexible spending account, and paid leave (including PTO and parental leave) in accordance with our applicable plans and policies. To learn more about our Total Rewards program please visit: https://rewards.the company.com/us.

The annual base salary range for this position for candidates located in California (excluding San Francisco Bay Area), Colorado, Illinois, New York, and Washington is between:

$147,000—$202,400 USD

  • Supporting Your Well-Being
  • Driving Social Impact
  • Developing Talent and Fostering Connection + Community

We are intentional about connection. Our global community, spanning over 20 offices worldwide, is united by a drive to innovate. Your journey begins with an immersive, in-person onboarding experience designed to accelerate your impact and connect you to our mission and team from day one.

the company is an Equal Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, ancestry, marital status, age, physical or mental disability, or status as a protected veteran. We also consider for employment qualified applicants with arrest and convictions records, consistent with applicable laws.

If reasonable accommodation is needed to complete any part of the job application, interview process, or onboarding pleaseuse this Form to request an accommodation.

Notice for New York City Applicants & Employees: the company may use Automated Employment Decision Tools (AEDT), as defined by New York City Local Law 144, that use artificial intelligence, machine learning, or other automated processes to assist in our recruitment and hiring process. In accordance with NYC Local Law 144, if you are an applicant or employee residing in New York City, pleaseclick here to view our full NYC AEDT Notice.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer - Security and Data Systems (Federal)
Senior Site Reliability Engineer - Security and Data Systems (Federal)

United States Digital Space LLC • Bellevue (WA)

On-site
USD 147,000 - 202,000
Staff Site Reliability Engineer (FedRAMP)
Staff Site Reliability Engineer (FedRAMP)

United States Digital Space LLC • Washington

On-site
USD 174,000 - 239,000
Health insurance
Dental insurance
Vision insurance
+2
Manager, Site Reliability Engineering
Manager, Site Reliability Engineering

United States Digital Space LLC • San Francisco (CA)

Hybrid
USD 204,000 - 306,000
Equity
Bonus
Health insurance
+6
Senior Manager, Site Reliability Engineering - Infrastructure Platform
Senior Manager, Site Reliability Engineering - Infrastructure Platform

United States Digital Space LLC • San Francisco (CA)

Hybrid
USD 232,000 - 319,000
Equity
Bonus
Health insurance
+2
Staff TDI Site Reliability Engineer, Okta Federal
Staff TDI Site Reliability Engineer, Okta Federal

United States Digital Space LLC • Washington

Hybrid
USD 174,000 - 239,000
Equity
Bonus
Health insurance
+5
Staff TDI Site Reliability Engineer, Okta Federal
Staff TDI Site Reliability Engineer, Okta Federal

United States Digital Space LLC • San Francisco (CA)

Hybrid
USD 174,000 - 239,000
Staff Software Engineer, Core Infrastructure
Staff Software Engineer, Core Infrastructure

United States Digital Space LLC • San Francisco (CA)

On-site
USD 194,000 - 267,000
Equity
Bonus
Health insurance
+6
Senior Site Reliability Engineer
Senior Site Reliability Engineer

United States Digital Space LLC • Bellevue (CA)

On-site
USD 180,000 - 230,000
Amazing Benefits
Making Social Impact
Fostering Diversity, Equity, Inclusion
Staff Platform Engineer (FedRAMP)
Staff Platform Engineer (FedRAMP)

United States Digital Space LLC • Washington

Hybrid
USD 174,000 - 239,000
Equity
Bonus
Health insurance
+3
Senior Manager, Software Engineering- Core (FED)
Senior Manager, Software Engineering- Core (FED)

United States Digital Space LLC • San Francisco (CA)

Hybrid
USD 232,000 - 290,000
Equity
Bonus eligible
Health, dental, and vision insurance
+2