Member of Technical Staff - Site Reliability

Runlayer

San Francisco (CA)

On-site

USD 150,000 - 190,000

Full time

43 hours ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Salary & equity
Paid time off
Professional development
Equipment
Health benefits
Customer exposure

Job summary

Runlayer is seeking a Site Reliability Engineer to own reliability, performance, and scalability of our cloud infrastructure as we scale to serve enterprise customers across multi-tenant SaaS, single-tenant SaaS, and BYOC.

You will collaborate with founders and a senior engineering team, manage AWS/GCP workloads, Kubernetes, CI/CD, and incident response, delivering resilient systems for enterprise clients.

Qualifications

  • Experience deploying and supporting on-prem / BYOC environments.
  • Background at a B2B company serving enterprise customers.
  • Strong AWS experience, ECS, Aurora, Kinesis.
  • Networking: VPC peering, Transit Gateway, PrivateLink, DNS, TLS termination.
  • CI/CD pipeline ownership and incident response experience.

Responsibilities

  • Own reliability and performance of cloud infrastructure across AWS (ECS, Aurora, CloudWatch) and GCP.
  • Manage and optimize Kubernetes clusters and container orchestration.
  • Drive database reliability engineering, including performance tuning and scaling.
  • Build and maintain CI/CD pipelines for rapid, safe deployments.
  • Run incident response and on-call rotations.
  • Partner with product engineers to design scalable, resilient systems.

Skills

AWS experience
Kubernetes
CI/CD ownership
BYOC operations

Job description

About Runlayer

AI is transforming how every company operates, but most enterprises are stuck. They want to move fast with AI Agents, tools, and workflows, but they can't do it safely. We're fixing that.

Our team built AI Actions for OpenAI, shipped Zapier Agents to millions of users, and launched the first remote MCP server with Anthropic. We helped establish the protocol, and now we're building the platform enterprises need to actually put AI to work.

Runlayer is one platform for MCPs, Skills, and Agents: purpose-built security, fine-grained governance, and complete observability so organizations can go all-in on AI across the entire company without the risk. We just raised a $30M Series A led by Felicis, with participation from Khosla Ventures, bringing our total raised to $42M. Already trusted by Gusto, Instacart, Opendoor, dbt Labs, and Decagon.

About Runlayer

AI is transforming how every company operates, but most enterprises are stuck. They want to move fast with AI Agents, tools, and workflows, but they can't do it safely. We're fixing that.

Our team built AI Actions for OpenAI, shipped Zapier Agents to millions of users, and launched the first remote MCP server with Anthropic. We helped establish the protocol, and now we're building the platform enterprises need to actually put AI to work.

Runlayer is one platform for MCPs, Skills, and Agents: purpose-built security, fine-grained governance, and complete observability so organizations can go all-in on AI across the entire company without the risk. We just raised a $30M Series A led by Felicis, with participation from Khosla Ventures, bringing our total raised to $42M. Already trusted by Gusto, Instacart, Opendoor, dbt Labs, and Decagon.

About The Role

As our Site Reliability Engineer, you'll own the reliability, performance, and scalability of Runlayer's infrastructure as we grow to serve enterprise customers across multiple deployment models: multi-tenant SaaS, single-tenant SaaS, and BYOC.

Why You'll Thrive Here
  • Impact: Build the infrastructure foundation for the enterprise AI platform, directly enabling AI adoption at scale
  • Excellence: Work closely with founders and a small, senior engineering team shipping fast in a high-growth environment
  • Ownership: Own reliability end-to-end, from database performance to incident response to CI/CD pipelines
What You'll Do
  • Own reliability and performance of our cloud infrastructure across AWS (ECS, Aurora, CloudWatch) and GCP
  • Manage and optimize Kubernetes clusters and container orchestration
  • Drive database reliability engineering, including performance tuning and scaling
  • Build and maintain CI/CD pipelines for rapid, safe deployments
  • Run incident response and on-call rotations
  • Partner with product engineers to design scalable, resilient systems
What We're Looking For
  • Experience deploying and supporting on-prem / BYOC environments. We are looking for engineers who have experienced the challenges of scaling BYOC operations, and know what they would/wouldn’t do again
  • Background at a B2B company serving enterprise customers, ideally building infrastructure/platform products
  • Strong AWS experience, particularly ECS, Aurora, Kinesis
  • Networking: VPC peering, Transit Gateway, PrivateLink, security groups, DNS, TLS termination
  • CI/CD pipeline ownership and incident response experience
Bonus Qualifications
  • GCP experience as we expand cross-cloud
  • Python/Go familiarity
  • Experience at an early-stage or high-growth company
What We Offer

We provide a competitive package designed to attract and retain top talent who can work effectively with enterprise customers.

  • Competitive salary and equity — compensation that reflects your expertise and customer-facing responsibilities.
  • Paid time off — paid vacation, paid sick leave, and paid parental leave.
  • Professional development — budget for conferences, courses, and certifications in AI, enterprise software, and customer success.
  • Top-tier equipment — your choice of laptop and accessories to create your ideal work environment.
  • Health benefits — comprehensive health, dental, and vision coverage.
  • Customer interaction opportunities — work directly with innovative companies and see the immediate impact of your work.

Not quite the right fit? Reach out to careers@runlayer.com with details about your experience and interests.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Member of Technical Staff - Infrastructure
Member of Technical Staff - Infrastructure

Runlayer • San Francisco (CA)

On-site
USD 140,000 - 200,000
Equity
PTO
Development budget
+3
Member of Technical Staff - Infrastructure
Member of Technical Staff - Infrastructure

Runlayer • New York (NY)

On-site
USD 120,000 - 160,000
Competitive salary and equity
Paid time off
Professional development
+3
Account Executive
Account Executive

Runlayer • San Francisco (CA)

Hybrid
USD 90,000 - 150,000
Competitive salary and equity
Paid time off
Professional development
+3
Founding Product Manager, Control Plane
Founding Product Manager, Control Plane

Runlayer • New York (NY)

On-site
USD 180,000 - 240,000
Competitive salary
Equity
Paid time off
+4
Member of Technical Staff - Security
Member of Technical Staff - Security

Runlayer • New York (NY)

On-site
USD 150,000 - 210,000
Competitive salary and equity
Paid time off
Professional development
+3
Head of Field Engineering
Head of Field Engineering

Runlayer • New York (NY)

On-site
USD 150,000 - 200,000
4 weeks paid vacation
Paid sick leave
Paid parental leave
+2
Regional Sales Director – East
Regional Sales Director – East

Runlayer • New York (NY)

On-site
USD 150,000 - 210,000
Competitive salary and equity
Paid time off
Professional development
+3
Integrations Engineer
Integrations Engineer

Runlayer • New York (NY)

On-site
USD 120,000 - 190,000
Competitive salary and equity
Paid time off
Professional development
+3
East Regional Sales Director for AI Security Growth
East Regional Sales Director for AI Security Growth

Runlayer • New York (NY)

On-site
USD 150,000 - 210,000
Competitive salary and equity
Paid time off
Professional development
+3
Founding Product Manager, Agents
Founding Product Manager, Agents

Runlayer • New York (NY)

On-site
USD 150,000 - 210,000
Competitive salary and equity
Paid time off
Professional development
+3