Site Reliability Engineer/Cloud Platform Engineer - Operations (PST Timezone)

skyflow

India

On-site

INR 2,500,000 - 3,500,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Skyflow in India is seeking a Platform Engineer to build and run scalable, multi-tenant infrastructure for a security-sensitive cloud product used by enterprise customers worldwide. You will write Go and Python to turn repeated infrastructure work into platform capabilities, contribute to provisioning pipelines, CLIs, and self-service tooling, and own on-call incident response.

This hands-on role requires deep production cloud experience, strong IaC skills, and a bias for automation.

Qualifications

  • 4+ years in platform/infrastructure engineering with prod cloud infra ownership.
  • Proficient in Go and Python; treats infra as software.
  • Hands-on Kubernetes production experience and IaC tooling (Terraform, Pulumi, OpenTofu).
  • Experience with at least one major cloud (AWS/GCP/Azure) and multi-cloud exposure.
  • Track record of building internal tooling or self-service platforms.
  • Comfort in B2B enterprise environments with security/compliance requirements.
  • Strong incident response and root-cause analysis skills.

Responsibilities

  • Provide operational support aligned with US time zones.
  • Design and build end-to-end automation for cloud infra across multi-tenant and BYOC models.
  • Develop production-grade Go and Python services and CLIs for self-service platforms.
  • Own and evolve the IaC stack with migrations with minimal customer impact.
  • Operate and scale core platform services (Kubernetes, Istio, data stores, messaging) with cost efficiency.
  • Improve observability and alerting to catch failures before customers notice.
  • Participate in on-call rotations and write RCAs.
  • Partner with security/compliance to bake guardrails into the platform.
  • Drive automation to reduce manual toil.

Skills

Go
Python
Kubernetes
Terraform
Pulumi
OpenTofu
Cloud (AWS/GCP/Azure)

Tools

ArgoCD
Istio
Helm
GitOps

Job description

Skyflow secures the flow of data across datastores, models, and agents. Enterprises turn to Skyflow as their runtime AI data control layer to protect sensitive data, enable safe AI deployment, and unlock full value from their applications, data platforms, and AI systems.

Skyflow is trusted by Fortune 500 enterprises, and leading SaaS companies in financial services, healthcare, retail, travel and hospitality.

Skyflow is headquartered in Palo Alto, California and was founded in 2019. For more information, visit www.skyflow.com or follow on X and LinkedIn.

About the role

We're looking for a Platform Engineer to help build and run the infrastructure that powers a multi-tenant, security-sensitive B2B cloud product used by enterprise customers around the world. Our platform spans multi-tenant environments and single-tenant (BYOC) deployments into customer cloud accounts, so the bar for automation, consistency, and reliability is high - every environment needs to be provisioned, upgraded, monitored, and recovered the same way, at scale, with minimal manual intervention.

This is a hands-on engineering role, not a ticket-queue ops role. You'll spend most of your time writing Go and Python to turn repeatable infrastructure work into platform capabilities - provisioning pipelines, internal CLIs, self-service tooling, and automation that lets the rest of engineering ship without waiting on infrastructure as a bottleneck. You'll also carry deep operational ownership: on-call, incident response, and root-cause analysis on the systems you build.

You have:
  • 4+ years of experience in platform engineering, infrastructure engineering, DevOps, or SRE roles, with real ownership of production cloud infrastructure.
  • Strong software engineering skills in Go and/or Python - you write tested, maintainable code and think of infrastructure automation as software, not scripting.
  • Deep hands-on experience with Kubernetes in production (workload scheduling, networking, autoscaling, upgrades) and Infrastructure-as-Code tooling such as Terraform, Pulumi, or OpenTofu.
  • Solid experience with at least one major public cloud (AWS or GCP or Azure); experience operating in all the three is a strong plus given our multi-cloud footprint.
  • Track record of building tools or platforms that other engineers use - internal CLIs, provisioning frameworks, self-service portals, or automation pipelines - not just maintaining existing infrastructure.
  • Comfort operating in a B2B environment with enterprise customers, where infrastructure changes carry compliance, security, and contractual weight (e.g. dedicated/BYOC deployments, uptime SLAs, audit requirements).
  • Strong incident-response instincts: you can debug distributed systems under pressure, drive a root-cause analysis to a real fix, and communicate clearly during and after an incident.
  • A bias toward root-causing and automating away recurring problems over repeatedly firefighting the same issue.
You will:
  • Provide operational support aligned with US time zones, ensuring system reliability and availability.
  • Design and build automation that provisions and manages cloud infrastructure end-to-end - new environment onboarding, upgrades, scaling, and decommissioning - across multiple cloud providers (AWS, GCP) and multiple deployment models (multi-tenant and dedicated/BYOC).
  • Write production-grade Go and Python services and CLIs that turn infrastructure operations into self-service platform capabilities for other engineering teams, rather than one-off scripts or manual runbooks.
  • Own and evolve the Infrastructure-as-Code stack (Terraform/OpenTofu, Helm, GitOps/ArgoCD) that defines every environment, and drive migrations across the fleet (Kubernetes version upgrades, node pool migrations, service mesh changes) with minimal customer impact.
  • Operate and scale core platform services - Kubernetes clusters, service mesh (Istio), data stores (Aerospike, PostgreSQL), messaging (Kafka), and GPU-backed inference workloads - with a focus on capacity planning, cost efficiency, and right-sizing.
  • Build and improve observability and alerting (metrics, logs, synthetic monitoring) so that failures are caught before customers notice, and drive the automation that turns repeat incidents into permanent fixes.
  • Participate in an on-call rotation, lead incident response for the systems you own, and write root-cause analyses that result in concrete corrective action - not just documentation.
  • Partner with security and compliance stakeholders to build guardrails (secrets management, access control, network policy, audit logging) directly into the platform, so secure-by-default is the path of least resistance for every team.
  • Continuously identify manual, repetitive, or error-prone infrastructure work and eliminate it - the measure of success in this role is less manual toil across the org, not more tickets closed.
Nice to have
  • Experience with service mesh (Istio/Envoy), GitOps workflows (ArgoCD/Flux), or policy-as-cod
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer/Cloud Platform Engineer
Senior Site Reliability Engineer/Cloud Platform Engineer

skyflow • India

On-site
INR 4,000,000 - 7,000,000
Infra Team Manager
Infra Team Manager

KRAFTON INDIA • Bengaluru

On-site
INR 4,000,000 - 7,000,000
Infra Team Manager
Infra Team Manager

Krafton • Bengaluru

On-site
INR 1,500,000 - 2,000,000
Senior Platform Engineer
Senior Platform Engineer

LE300 Optiva (India) Technologies Pvt. Ltd. • Hyderabad

On-site
INR 1,800,000 - 2,500,000
Tazapay - Staff DevOps Engineer
Tazapay - Staff DevOps Engineer

Tazapay • Chennai District

On-site
INR 3,500,000 - 6,000,000
Cloud Operations Lead – SRE / DevOps / Platform Engineering
Cloud Operations Lead – SRE / DevOps / Platform Engineering

PeoplePilot • Pune District

On-site
INR 2,600,000 - 5,200,000
Platform Engineer
Platform Engineer

LE300 Optiva (India) Technologies Pvt. Ltd. • Hyderabad

On-site
INR 800,000 - 1,500,000
Sr. Infrastructure Engineer – Data Platforms (Python + Cloud)
Sr. Infrastructure Engineer – Data Platforms (Python + Cloud)

General Mills • Pune District

On-site
INR 1,800,000 - 3,000,000
Lead DevOps Engineer
Lead DevOps Engineer

Lenskart • Gurugram District

On-site
INR 1,200,000 - 2,400,000
Lead SDE - DevOps
Lead SDE - DevOps

Flourish Ventures • Chennai District

On-site
INR 2,000,000 - 3,000,000
Inclusive and people-first culture
Health & wellness programs
Comprehensive medical insurance
+2