Site Reliability Engineer/Cloud Platform Engineer - Operations (PST Timezone)

Skyflow

United States

Remote

USD 150,000 - 210,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Work from home expense
Excellent Health Insurance Options
Very generous PTO
Flexible Hours
Generous Equity

Job summary

Skyflow is seeking a Platform Engineer to design and run the infrastructure for a multi-tenant, security-sensitive B2B cloud product used by enterprises worldwide.

The role emphasizes automation, reliability, and operational ownership, with heavy Go and Python development, Kubernetes operations, and IaC tooling. You will work across AWS and GCP in multi-tenant and BYOC deployments.

Qualifications

  • Hands-on platform/infra engineering with production ownership.
  • Strong coding in Go/Python with tested, maintainable code.
  • Experience with Kubernetes in production and IaC tooling.

Responsibilities

  • Provide operational support aligned with US time zones.
  • Design and build automation for multi-cloud, multi-tenant environments.
  • Write production-grade Go and Python services and CLIs for self-service platforms.
  • Own and evolve the IaC stack (Terraform/OpenTofu, Helm, GitOps).
  • Operate and scale core platform services (Kubernetes, Istio, databases, queues).
  • Improve observability and alerting to catch issues before customers notice.
  • Participate in on-call rotations and drive root-cause analyses.

Skills

Go
Python
Kubernetes
Terraform
OpenTofu
ArgoCD
Istio
AWS/GCP

Tools

Pulumi
Terraform/OpenTofu
Helm
GitOps

Job description

Skyflow secures the flow of data across datastores, models, and agents. Enterprises turn to Skyflow as their runtime AI data control layer to protect sensitive data, enable safe AI deployment, and unlock full value from their applications, data platforms, and AI systems. Skyflow is trusted by Fortune 500 enterprises, and leading SaaS companies in financial services, healthcare, retail, travel and hospitality.

Skyflow is headquartered in Palo Alto, California and was founded in 2019. For more information, www.skyflow.com or follow on X and LinkedIn.

About the role

We're looking for a Platform Engineer to help build and run the infrastructure that powers a multi-tenant, security-sensitive B2B cloud product used by enterprise customers around the world. Our platform spans multi-tenant environments and single-tenant (BYOC) deployments into customer cloud accounts, so the bar for automation, consistency, and reliability is high — every environment needs to be provisioned, upgraded, monitored, and recovered the same way, at scale, with minimal manual intervention.

This is a hands-on engineering role, not a ticket-queue ops role. You'll spend most of your time writing Go and Python to turn repeatable infrastructure work into platform capabilities — provisioning pipelines, internal CLIs, self-service tooling, and automation that lets the rest of engineering ship without waiting on infrastructure as a bottleneck. You'll also carry deep operational ownership: on-call, incident response, and root-cause analysis on the systems you build.

*must be able to work PST hours

You have:
  • 4+ years of experience in platform engineering, infrastructure engineering, DevOps, or SRE roles, with real ownership of production cloud infrastructure.
  • Strong software engineering skills in Go and/or Python — you write tested, maintainable code and think of infrastructure automation as software, not scripting.
  • Deep hands-on experience with Kubernetes in production (workload scheduling, networking, autoscaling, upgrades) and Infrastructure-as-Code tooling such as Terraform, Pulumi, or OpenTofu.
  • Solid experience with at least one major public cloud (AWS or GCP or Azure); experience operating in all the three is a strong plus given our multi-cloud footprint.
  • Track record of building tools or platforms that other engineers use — internal CLIs, provisioning frameworks, self-service portals, or automation pipelines — not just maintaining existing infrastructure.
  • Comfort operating in a B2B environment with enterprise customers, where infrastructure changes carry compliance, security, and contractual weight (e.g. dedicated/BYOC deployments, uptime SLAs, audit requirements).
  • Strong incident-response instincts: you can debug distributed systems under pressure, drive a root-cause analysis to a real fix, and communicate clearly during and after an incident.
  • A bias toward root-causing and automating away recurring problems over repeatedly firefighting the same issue.
You will:
  • Provide operational support aligned with US time zones, ensuring system reliability and availability.
  • Design and build automation that provisions and manages cloud infrastructure end-to-end — new environment onboarding, upgrades, scaling, and decommissioning — across multiple cloud providers (AWS, GCP) and multiple deployment models (multi-tenant and dedicated/BYOC).
  • Write production-grade Go and Python services and CLIs that turn infrastructure operations into self-service platform capabilities for other engineering teams, rather than one-off scripts or manual runbooks.
  • Own and evolve the Infrastructure-as-Code stack (Terraform/OpenTofu, Helm, GitOps/ArgoCD) that defines every environment, and drive migrations across the fleet (Kubernetes version upgrades, node pool migrations, service mesh changes) with minimal customer impact.
  • Operate and scale core platform services — Kubernetes clusters, service mesh (Istio), data stores (Aerospike, PostgreSQL), messaging (Kafka), and GPU-backed inference workloads — with a focus on capacity planning, cost efficiency, and right-sizing.
  • Build and improve observability and alerting (metrics, logs, synthetic monitoring) so that failures are caught before customers notice, and drive the automation that turns repeat incidents into permanent fixes.
  • Participate in an on-call rotation, lead incident response for the systems you own, and write root-cause analyses that result in concrete corrective action — not just documentation.
  • Partner with security and compliance stakeholders to build guardrails (secrets management, access control, network policy, audit logging) directly into the platform, so secure-by-default is the path of least resistance for every team.
  • Continuously identify manual, repetitive, or error-prone infrastructure work and eliminate it — the measure of success in this role is less manual toil across the org, not more tickets closed.
Nice to have
  • Experience with service mesh (Istio/Envoy), GitOps workflows (ArgoCD/Flux), or policy-as-code (OPA).
  • Experience operating stateful systems in production — distributed databases (Aerospike, PostgreSQL, Cassandra), message queues (Kafka), or GPU-backed inference workloads.
  • Experience with security-sensitive or regulated environments (PCI, SOC 2, HIPAA) and building compliance controls into infrastructure by default.
  • Experience with cost optimization and capacity planning at fleet scale (right-sizing, autoscaling policy design, spot/preemptible usage).
  • Contributions to open-source infrastructure tooling, or experience building an internal developer platform (IDP).
Our stack (roughly)

Go, Python · Kubernetes, Helm, ArgoCD · Terraform/OpenTofu · AWS, GCP · Istio · Aerospike, PostgreSQL, Kafka · Coralogix/Prometheus-style observability · GitHub Actions and Spacelift-style CI/CD for infrastructure changes.

Benefits:
  • Work from home expense
  • Excellent Health Insurance Options
  • Very generous PTO
  • Flexible Hours
  • Generous Equity

At Skyflow, we believe that diverse teams are the strongest teams. We invite applicants of all genders, races, ethnicities, nationalities, ages, religions, sexual orientations, disability statuses, educational experiences, family situations, and socio-economic backgrounds.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Product Marketing Manager (all levels)
Product Marketing Manager (all levels)

Skyflow • United States

On-site
USD 140,000 - 210,000
Work from home stipend
Health/Dental/Vision insurance
401k
+2
Product Marketing Manager (all levels)
Product Marketing Manager (all levels)

Skyflow • United States

Remote
USD 120,000 - 180,000
Work from home expense
Health insurance
Dental insurance
+4
OEM Data Privacy Partnerships Executive
OEM Data Privacy Partnerships Executive

Foundation Capital • United States

On-site
USD 180,000 - 240,000
Work from home expense
Excellent Health Insurance Options
Very generous PTO
+2
AI Product Lead
AI Product Lead

Skyflow • United States

On-site
USD 120,000 - 170,000
Work-from-home reimbursement
Health/Dental/Vision insurance
401k Vanguard
+2
ISV Account Executive - United States
ISV Account Executive - United States

Skyflow • United States

On-site
USD 180,000 - 240,000
Work from home expense
Excellent Health Insurance Options
Very generous PTO
+2
Strategic Enterprise AI & Privacy SaaS Account Executive
Strategic Enterprise AI & Privacy SaaS Account Executive

Skyflow • United States

Hybrid
USD 120,000 - 200,000
Work from home stipend (US)
Health, Dental, Vision Insurance
401k
+2
Account Executive, AI & Digital Natives
Account Executive, AI & Digital Natives

Skyflow • United States

On-site
USD 120,000 - 200,000
Work from home stipend (US)
Health, Dental, Vision Insurance
401k
+2
Software Engineer, Ontology
Software Engineer, Ontology

Fluidstack • New York (NY)

On-site
USD 140,000 - 180,000
Equity
Health insurance
Retirement plan
+1
Enterprise Solutions Architect
Enterprise Solutions Architect

Skyflow • United States

On-site
USD 90,000 - 130,000
Home Office Expense
Health, Vision, Dental
Unlimited PTO
+2
Senior Software Engineer, Platform
Senior Software Engineer, Platform

Astronomer • London (KY)

On-site
USD 150,000 - 210,000