Sr. Principal Engineer — Platform

Rafay Systems

Sunnyvale (CA)

On-site

USD 260,000 - 360,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Rafay Systems in Sunnyvale, CA is seeking a Sr. Principal Engineer to provide technical leadership and hands-on architecture for Rafay’s multi-tenant cloud and AI infrastructure platform.

You will guide architecture across cloud infrastructure, Kubernetes, observability, security, and automation while mentoring teams and driving reliable production systems. The role blends architecture with hands-on implementation, focusing on scalable, secure platforms suitable for enterprise and regulated

Qualifications

  • 12+ years designing large-scale enterprise software platforms.
  • Senior architect or principal engineer responsible for platform decisions.
  • Deep understanding of distributed systems: scalability, HA, resiliency.
  • Expert in Golang or Python; strong cloud and Kubernetes experience.
  • Microservices, multi-tenant control planes, and security-focused design.
  • Experience with public clouds (AWS/Azure/GCP) and networking basics.
  • Observability, telemetry, and proactive reliability practices.

Responsibilities

  • Design core architectural components for a large-scale multi-tenant platform.
  • Set technical vision and align major platform areas across teams.
  • Build modular, scalable, secure, observable distributed services.
  • Lead platform security architecture incl. IAM, secrets, TLS, and isolation.
  • Drive observability via metrics, logs, traces, and health frameworks.
  • Mentor engineers and establish scalable engineering practices.

Skills

Golang
Python
Kubernetes
Distributed systems
Security
Observability
Multi-tenant
Cloud infrastructure
Networking
Architecture leadership

Tools

OpenTelemetry
Prometheus
Grafana
AWS
GCP

Job description

We are looking for a Sr. Principal Engineer to provide technical leadership and make significant contributions to the architecture and development of Rafay's multi-tenant cloud and AI infrastructure platform.

Rafay operates at the intersection of distributed systems, Kubernetes, virtualization, infrastructure automation, observability, and security. This role provides an opportunity to build foundational technologies used to operate complex cloud and AI infrastructure environments at scale.

As a Sr. Principal Engineer, you will serve as a technical leader and hands-on architect, setting technical direction for critical platform capabilities and multiplying the effectiveness of engineering teams across the organization.

You will be expected to move comfortably between architecture and implementation, reason about complex distributed systems and failure modes, and design platforms that are scalable, highly available, observable, secure, and suitable for enterprise and regulated environments.

Responsibilities
  • Design and implement core architectural components for critical services within a large-scale, multi-tenant distributed platform.
  • Set the technical vision and long-term architectural direction for major platform areas and drive alignment across engineering teams.
  • Design highly modular, scalable, resilient, secure, and maintainable distributed services.
  • Architect platform capabilities spanning cloud infrastructure, Kubernetes, virtualization, observability, security, and automation.
  • Drive the design of modern observability capabilities covering metrics, logs, traces, events, infrastructure telemetry, and service health.
  • Develop approaches for correlating information across multiple infrastructure and application layers to improve troubleshooting and operational reliability.
  • Design service health monitoring, alerting, SLI/SLO frameworks, synthetic monitoring, and proactive validation capabilities.
  • Develop capabilities that improve incident detection, troubleshooting, root-cause analysis, and operational automation.
  • Help establish architectures for safe, controlled, and auditable automation of infrastructure operations.
  • Define and drive platform security architecture, including identity, authorization, secrets management, workload isolation, network security, secure APIs, and privileged operations.
  • Design systems suitable for high-security and regulated environments, including government and enterprise deployments.
  • Partner with security and compliance teams to translate compliance requirements into scalable platform capabilities.
  • Contribute to architectures involving confidential computing, trusted execution environments, hardware-backed security, attestation, and protection of sensitive workloads and data.
  • Participate in security architecture reviews and threat modeling for distributed and multi-tenant systems.
  • Ensure systems provide strong auditability for administrative actions, configuration changes, and automated operations.
  • Perform R&D and feasibility analysis on emerging technologies related to cloud infrastructure, AI infrastructure, observability, distributed systems, and security.
  • Assist engineering and operations teams with diagnosing complex production issues and drive systemic improvements based on lessons learned.
  • Lead architecture reviews, design reviews, and code reviews and help establish engineering standards across teams.
  • Mentor engineers and raise the technical bar across the engineering organization.
  • Collaborate with engineering, product management, security, SRE, QA, and customer-facing organizations.
  • Champion a customer-focused engineering culture where production experience informs improvements in reliability, usability, security, and automation.
Skills and Qualifications
  • 12+ years of experience designing, building, and delivering large-scale enterprise software platforms.
  • Demonstrated experience operating as a senior technical architect or Principal-level engineer responsible for significant platform architecture decisions.
  • Deep understanding of distributed systems fundamentals, including:
    • Scalability
    • High availability
    • Resiliency
    • Distributed state
    • Concurrency
    • Failure handling
    • Performance
  • Expert knowledge of one or more programming languages, preferably:
    • Golang
    • Python
  • Strong experience designing and developing microservices and distributed control-plane systems.
  • Excellent troubleshooting and debugging skills across complex production environments.
  • Hands-on experience building services on public cloud platforms such as AWS, Azure, or GCP.
  • Strong practical knowledge of networking fundamentals and protocols including TCP/IP, HTTP/HTTPS, DNS, TLS, load balancing, and network security.
  • Strong experience with Kubernetes and cloud-native architectures.
  • Experience with container orchestration, distributed control planes, APIs, lifecycle management, and multi-tenancy.
Observability & Reliability
  • Deep experience designing or operating large-scale observability and telemetry systems.
  • Experience with metrics, logging, distributed tracing, event processing, and infrastructure monitoring.
  • Hands-on experience with technologies such as OpenTelemetry, Prometheus, Grafana, or comparable platforms.
  • Experience defining and operating:
    • SLIs
    • SLOs
    • Service-health models
    • Alerting strategies
    • Capacity and performance monitoring
  • Experience designing systems for proactive problem detection, telemetry correlation, and production troubleshooting.
  • Experience with synthetic monitoring, anomaly detection, or automated service validation is highly desirable.
  • Experience applying AI or automation to improve operational diagnosis and reliability is a plus.
Security
  • Strong understanding of cloud-native and distributed-systems security.
  • Hands-on experience with areas such as:
    • IAM and RBAC
    • Least privilege
    • Workload identity
    • PKI and certificate lifecycle
    • Secrets management
    • TLS/mTLS
    • Network security and segmentation
    • API security
    • Secure software supply chains
    • Vulnerability management
    • Security logging and auditing
  • Experience designing secure multi-tenant platforms with strong workload and tenant isolation.
  • Experience performing threat modeling and security architecture reviews.
  • Familiarity with policy-as-code, automated security controls, and security telemetry is desirable.
FedRAMP & Regulated Environments
  • Experience designing, implementing, or operating software platforms in FedRAMP, U.S. government, or similarly regulated environments is highly desirable.
  • Experience translating regulatory and security requirements into technical architecture and software controls.
  • Experience with continuous compliance, security monitoring, audit evidence, vulnerability management, and configuration governance is desirable.
Confidential Computing
  • Understanding of confidential computing and trusted execution environments.
  • Familiarity with hardware-backed security technologies available across modern CPU, GPU, virtualization, and cloud platforms.
  • Understanding of concepts such as:
    • Hardware roots of trust
    • Secure and measured boot
    • Attestation
    • Trusted execution environments
    • Memory encryption
    • Data-in-use protection
  • Experience integrating confidential-computing technologies into Kubernetes, virtualization, or cloud platforms is highly desirable.
Technical Leadership
  • Proven ability to make architecture decisions involving substantial scale, reliability, security, and operational complexity.
  • Demonstrated ability to influence technical direction and build consensus across multiple teams.
  • Ability to identify architectural limitations and drive longer-term improvements across multiple releases.
  • Experience solving difficult customer production problems and converting those learnings into platform improvements.
  • Strong ability to mentor senior engineers and establish engineering practices that scale beyond an individual team.
Preferred Background

A particularly strong candidate will have experience across several of the following areas:

Distributed Systems | Kubernetes | Cloud Infrastructure | Observability | Security | GPU/AI Infrastructure | Virtualization | FedRAMP | Confidential Computing

Candidates are not expected to be experts in every area. We are looking for individuals with deep expertise in several domains and sufficient architectural breadth to reason across the entire platform.

WHY JOIN RAFAY

Rafay is building foundational infrastructure for modern cloud and AI environments. Engineers at Rafay work on challenging problems involving distributed systems, Kubernetes, virtualization, observability, security, and infrastructure automation.

We offer an environment where senior engineers can influence platform architecture, work on technically demanding infrastructure problems, and help define the next generation of cloud and AI infrastructure management.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Principal Solutions Architect
Principal Solutions Architect

Rafay • United States

Hybrid
USD 180,000 - 240,000
Lead Technical Support Engineer
Lead Technical Support Engineer

Rafay Systems • Northern (KY)

On-site
USD 120,000 - 160,000
Competitive compensation
Comprehensive benefits
Stock options
Senior Solutions Engineer
Senior Solutions Engineer

Rafay Systems • United States

On-site
USD 140,000 - 170,000
Sr. Solutions Engineer
Sr. Solutions Engineer

Rafay • United States

On-site
USD 140,000 - 180,000
Competitive salary
Robust benefits
Stock options
Sr. Implementation Engineer (Kubernetes/AI)
Sr. Implementation Engineer (Kubernetes/AI)

Rafay • United States

On-site
USD 120,000 - 150,000
Competitive salary
Robust benefits
Attractive stock options
Technical Solutions Architect (GPU Platform)
Technical Solutions Architect (GPU Platform)

Rafay • United States

On-site
USD 180,000 - 240,000
Lead Technical Support Engineer
Lead Technical Support Engineer

Rafay • San Francisco (CA)

On-site
USD 140,000 - 180,000
Stock options
Comprehensive benefits
Career mentorship
Solutions Architect
Solutions Architect

Rafay Systems • United States

On-site
USD 140,000 - 210,000
Principal Solutions Architect
Principal Solutions Architect

Rafay Systems • United States

Hybrid
USD 190,000 - 280,000
Stock options
Senior Engineering Program Manager
Senior Engineering Program Manager

Rafay Systems • Northern (KY)

On-site
USD 150,000 - 190,000
a fun and dynamic work environment
a competitive salary
robust benefits
+1