Manager, Site Reliability Engineering

United States Digital Space LLC

San Francisco (CA)

Hybrid

USD 204,000 - 306,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Equity
Bonus
Health insurance
Dental insurance
Vision insurance
401(k)
Flexible spending account
Paid leave
Parental leave

Job summary

the company in San Francisco is seeking a Manager of Infrastructure Platform and Shared Services within the IDaaS SRE group. You will oversee Edge networking, Kubernetes platform, CI/CD, Observability, and tooling.

You will lead a multi-team, drive DevOps maturity, build self-service tooling, mentor engineers, and ensure on-time delivery within budget. Requires 3+ years of technical leadership and strong AWS/Kubernetes/IaC experience; SF office 2 days/week.

Qualifications

  • 3+ years of experience in technical leadership & people management.
  • Extensive experience using Agile and DevOps methodologies to build product infrastructure and shared service at scale.
  • Experience running large-scale infrastructure platforms supporting a SaaS/Cloud service in a public Cloud, preferably AWS.
  • Strong expertise in cloud-native architectures, containerization (Kubernetes), IaC (Terraform), and CI/CD pipelines.
  • Strong background and hands-on experience in SW development, PaaS and automation.
  • Deep experience with building and operating observability platforms and monitoring tools (Grafana, Splunk, APM etc.) in a large scale environment.
  • Effective verbal, written communication and interpersonal skills.
  • Computer Science Degree or related degree or equivalent experience.

Responsibilities

  • Manage a team of SREs supporting various workloads and teams that support our IDaaS platform.
  • Drive the microservice journey, DevOps maturity, and workload reliability in tandem with architects and teams across the organization.
  • Accelerate the velocity of SRE and product engineering by developing powerful tooling, intuitive self-service capabilities, and robust self-healing patterns.
  • Lead, mentor, and grow a high-performing team of engineers and managers across platform, infrastructure, and shared services domains.
  • Perform engineering design evaluations and ensure the completion of projects within resource, budget, and scheduling constraints.
  • Improve SDLC processes for Cloud infrastructure as a code, including the maturity of CI/CD pipelines, change and release management
  • Manage service and business expectations and prioritize resource allocation
  • Maintain a deep knowledge of industry best practices, evolving trends, and technologies

Skills

Technical leadership
Agile & DevOps
AWS & multi-cloud
Kubernetes
Terraform IaC
CI/CD pipelines
Observability tools
Communication skills
CS Degree or related

Education

CS Degree or related degree

Tools

Grafana
Splunk
APM tools
CI/CD tooling

Job description

**Secure Every Identity, from AI to Human

**Identity is the key to unlocking the potential of AI. the company secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence.

This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk.

Manager, Site Reliability Engineering

San Francisco, California

Secure Every Identity, from AI to Human

**Identity is the key to unlocking the potential of AI. the company secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence.

This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk.

**This position requires 2 days a week in our San Francisco Office.

The IDaaS Site Reliability Engineering Group

the company authenticates, authorizes and provisions millions of users a day. The service is hosted on Amazon Web Services (AWS) across multiple availability zones and geographically separated regions. The service is designed for high throughput and 99.999 availability. We're looking for a technical leader to help us continue to scale the service with great people and reliable, cost-effective, and efficient infrastructure, processes, and tooling.

As the Manager of Infrastructure Platform and Shared Services, you will oversee multiple teams focused on Edge networking, K8s platform, CI/CD, Observability, automation platform & tooling.

What you’ll be doing

  • Managing a team of SRE’s supporting various workloads and teams that support our IDaaS platform.
  • Drive the microservice journey, DevOps maturity, and workload reliability in tandem with architects and teams across the organization.
  • Accelerate the velocity of SRE and product engineering by developing powerful tooling, intuitive self-service capabilities, and robust self-healing patterns.
  • Lead, mentor, and grow a high-performing team of engineers and managers across platform, infrastructure, and shared services domains.
  • Perform engineering design evaluations and ensure the completion of projects within resource, budget, and scheduling constraints.
  • Improve SDLC processes for Cloud infrastructure as a code, including the maturity of CI/CD pipelines, change and release management
  • Manage service and business expectations and prioritize resource allocation
  • Maintain a deep knowledge of industry best practices, evolving trends, and technologies

What you’ll bring to the role

  • 3+ years of experience in technical leadership & people management
  • Extensive experience using Agile and DevOps methodologies to build product infrastructure and shared service at scale
  • Experience running large-scale infrastructure platforms supporting a SaaS/Cloud service in a public Cloud, preferably AWS. Experience supporting a multi-Cloud environment will be a plus.
  • Strong expertise in cloud-native architectures, containerization (Kubernetes), IaC (Terraform), and CI/CD pipelines
  • Strong background and hands-on experience in SW development, PaaS and automation
  • Deep experience with building and operating observability platforms and monitoring tools (Grafana, Splunk, APM etc.) in a large scale environment.
  • Effective verbal, written communication and interpersonal skills
  • Computer Science Degree or related degree or equivalent experience

Additional requirements:

  • This position requires the ability to access federal environments and/or have access to protected federal data. As a condition of employment for this position, the successful candidate must be able to submit documentation establishing U.S. Person status (e.g. a U.S. Citizen, National, Lawful Permanent Resident, Refugee, or Asylee. 22 CFR 120.15) upon hire.

#LI-Hybrid

Below is the annual base salary range for candidates located in San Francisco Bay Area. Your actual base salary will depend on factors such as your skills, qualifications, experience, and work location. In addition, the company offers equity (where applicable), bonus, and benefits, including health, dental and vision insurance, 401(k), flexible spending account, and paid leave (including PTO and parental leave) in accordance with our applicable plans and policies. To learn more about our Total Rewards program please visit: https://rewards.the company.com/us.

The annual base salary range for this position for candidates located in the San Francisco Bay area is between:

$204,000—$306,000 USD

The the company Experience
  • Supporting Your Well-Being
  • Driving Social Impact
  • Developing Talent and Fostering Connection + Community

We are intentional about connection. Our global community, spanning over 20 offices worldwide, is united by a drive to innovate. Your journey begins with an immersive, in-person onboarding experience designed to accelerate your impact and connect you to our mission and team from day one.

the company is an Equal Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, ancestry, marital status, age, physical or mental disability, or status as a protected veteran. We also consider for employment qualified applicants with arrest and convictions records, consistent with applicable laws.

If reasonable accommodation is needed to complete any part of the job application, interview process, or onboarding pleaseuse this Form to request an accommodation.

Notice for New York City Applicants & Employees: the company may use Automated Employment Decision Tools (AEDT), as defined by New York City Local Law 144, that use artificial intelligence, machine learning, or other automated processes to assist in our recruitment and hiring process. In accordance with NYC Local Law 144, if you are an applicant or employee residing in New York City, pleaseclick here to view our full NYC AEDT Notice.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Manager, Site Reliability Engineering - Infrastructure Platform
Senior Manager, Site Reliability Engineering - Infrastructure Platform

United States Digital Space LLC • San Francisco (CA)

Hybrid
USD 232,000 - 319,000
Equity
Bonus
Health insurance
+2
Staff Site Reliability Engineer (FedRAMP)
Staff Site Reliability Engineer (FedRAMP)

United States Digital Space LLC • Washington

On-site
USD 174,000 - 239,000
Health insurance
Dental insurance
Vision insurance
+2
Staff Software Engineer, Core Infrastructure
Staff Software Engineer, Core Infrastructure

United States Digital Space LLC • San Francisco (CA)

On-site
USD 194,000 - 267,000
Equity
Bonus
Health insurance
+6
Staff TDI Site Reliability Engineer, Okta Federal
Staff TDI Site Reliability Engineer, Okta Federal

United States Digital Space LLC • Washington

Hybrid
USD 174,000 - 239,000
Equity
Bonus
Health insurance
+5
Manager, Site Reliability Engineering
Manager, Site Reliability Engineering

Okta • Chicago (IL)

Hybrid
USD 204,000 - 306,000
Equity
Bonus
Health benefits
Manager, Site Reliability Engineering
Manager, Site Reliability Engineering

Okta • Bellevue (WA)

Hybrid
USD 204,000 - 306,000
Equity
Bonus
Health insurance
+4
Senior Manager, Software Engineering- Core (FED)
Senior Manager, Software Engineering- Core (FED)

United States Digital Space LLC • San Francisco (CA)

Hybrid
USD 232,000 - 290,000
Equity
Bonus eligible
Health, dental, and vision insurance
+2
Manager, Site Reliability Engineering
Manager, Site Reliability Engineering

Okta • San Francisco (CA)

On-site
USD 204,000 - 281,000
Health, dental, and vision insurance
401(k) plan
Paid leave including PTO and parental leave
Staff TDI Site Reliability Engineer, Okta Federal
Staff TDI Site Reliability Engineer, Okta Federal

United States Digital Space LLC • San Francisco (CA)

Hybrid
USD 174,000 - 239,000
Senior Manager, Site Reliability Engineering - Infrastructure Platform
Senior Manager, Site Reliability Engineering - Infrastructure Platform

Okta • Washington

On-site
USD 232,000 - 319,000
Health insurance
Dental insurance
Vision insurance
+2