Senior Site Reliability Engineer/Cloud Platform Engineer

skyflow

India

On-site

INR 4,000,000 - 7,000,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Skyflow invites an experienced Platform Engineer to design and run scalable, security-conscious cloud infrastructure for a multi-tenant B2B product used by enterprises worldwide.

You will own automation, IaC, and platform services (Kubernetes, Istio, data stores) while mentoring peers and ensuring uptime and audit readiness in a fast-paced environment.

Qualifications

  • 8+ years of experience in platform engineering, infrastructure engineering, DevOps, or SRE roles with ownership of production cloud infrastructure.
  • Strong software engineering skills in Go and/or Python; treat infrastructure automation as software.
  • Hands-on production experience with Kubernetes (workloads, networking, autoscaling, upgrades) and IaC tooling like Terraform, Pulumi or OpenTofu.
  • Experience with at least one major public cloud (AWS, GCP, or Azure) and multi-cloud environments.
  • Track record of building tools or platforms used by other engineers (internal CLIs, provisioning frameworks, self-service portals).
  • Comfort working in a B2B, enterprise environment with compliance, security, and SLAs.
  • Incident-response mindset with root-cause analysis and clear post-incident communication.
  • Ability to lead and mentor a team of platform engineers.

Responsibilities

  • Design and build automation that provisions and manages cloud infrastructure end-to-end across multi-tenant and BYOC deployments.
  • Develop production-grade Go and Python services and CLIs to enable self-service platform capabilities.
  • Own the IaC stack (Terraform/OpenTofu, Helm, GitOps/ArgoCD) and drive fleet migrations with minimal customer impact.
  • Operate and scale core platform services (Kubernetes, Istio, Aerospike, PostgreSQL, Kafka, GPU workloads) with focus on capacity and cost efficiency.
  • Improve observability, alerts, and automation to reduce manual toil and incidents.
  • Participate in on-call rotations and write RCA that leads to concrete actions.
  • Collaborate with security/compliance to bake guardrails into the platform for secure-by-default usage.
  • Mentor junior engineers and drive engineering best practices across the team.

Skills

Go
Python
Kubernetes
Terraform
Pulumi
OpenTofu
AWS
GCP
Azure
Internal tools
SRE practices

Tools

Helm
GitOps
ArgoCD
Istio
Envoy
Terraform/OpenTofu

Job description

Skyflow secures the flow of data across datastores, models, and agents. Enterprises turn to Skyflow as their runtime AI data control layer to protect sensitive data, enable safe AI deployment, and unlock full value from their applications, data platforms, and AI systems.

Skyflow is trusted by Fortune 500 enterprises, and leading SaaS companies in financial services, healthcare, retail, travel and hospitality.

Skyflow is headquartered in Palo Alto, California and was founded in 2019. For more information, visit www.skyflow.com or follow on X and LinkedIn.

About the role

We're looking for a Platform Engineer to help build and run the infrastructure that powers a multi-tenant, security-sensitive B2B cloud product used by enterprise customers around the world. Our platform spans multi-tenant environments and single-tenant (BYOC) deployments into customer cloud accounts, so the bar for automation, consistency, and reliability is high - every environment needs to be provisioned, upgraded, monitored, and recovered the same way, at scale, with minimal manual intervention.

This is a hands-on engineering role, not a ticket-queue ops role. You'll spend most of your time writing Go and Python to turn repeatable infrastructure work into platform capabilities - provisioning pipelines, internal CLIs, self-service tooling, and automation that lets the rest of engineering ship without waiting on infrastructure as a bottleneck. You'll also carry deep operational ownership: on-call, incident response, and root-cause analysis on the systems you build.

You have:
  • 8+ years of experience in platform engineering, infrastructure engineering, DevOps, or SRE roles, with real ownership of production cloud infrastructure.
  • Strong software engineering skills in Go and/or Python - you write tested, maintainable code and think of infrastructure automation as software, not scripting.
  • Deep hands-on experience with Kubernetes in production (workload scheduling, networking, autoscaling, upgrades) and Infrastructure-as-Code tooling such as Terraform, Pulumi, or OpenTofu.
  • Solid experience with at least one major public cloud (AWS or GCP or Azure); experience operating in all the three is a strong plus given our multi-cloud footprint.
  • Track record of building tools or platforms that other engineers use - internal CLIs, provisioning frameworks, self-service portals, or automation pipelines - not just maintaining existing infrastructure.
  • Comfort operating in a B2B environment with enterprise customers, where infrastructure changes carry compliance, security, and contractual weight (e.g. dedicated/BYOC deployments, uptime SLAs, audit requirements).
  • Strong incident-response instincts: you can debug distributed systems under pressure, drive a root-cause analysis to a real fix, and communicate clearly during and after an incident.
  • A bias toward root-causing and automating away recurring problems over repeatedly firefighting the same issue.
  • Lead, mentor, and technically guide a team of junior/early-career platform engineers.
You will:
  • Design and build automation that provisions and manages cloud infrastructure end-to-end - new environment onboarding, upgrades, scaling, and decommissioning - across multiple cloud providers (AWS, GCP) and multiple deployment models (multi-tenant and dedicated/BYOC).
  • Write production-grade Go and Python services and CLIs that turn infrastructure operations into self-service platform capabilities for other engineering teams, rather than one-off scripts or manual runbooks.
  • Own and evolve the Infrastructure-as-Code stack (Terraform/OpenTofu, Helm, GitOps/ArgoCD) that defines every environment, and drive migrations across the fleet (Kubernetes version upgrades, node pool migrations, service mesh changes) with minimal customer impact.
  • Operate and scale core platform services - Kubernetes clusters, service mesh (Istio), data stores (Aerospike, PostgreSQL), messaging (Kafka), and GPU-backed inference workloads - with a focus on capacity planning, cost efficiency, and right-sizing.
  • Build and improve observability and alerting (metrics, logs, synthetic monitoring) so that failures are caught before customers notice, and drive the automation that turns repeat incidents into permanent fixes.
  • Participate in an on-call rotation, lead incident response for the systems you own, and write root-cause analyses that result in concrete corrective action - not just documentation.
  • Partner with security and compliance stakeholders to build guardrails (secrets management, access control, network policy, audit logging) directly into the platform, so secure-by-default is the path of least resistance for every team.
  • Continuously identify manual, repetitive, or error-prone infrastructure work and eliminate it - the measure of success in this role is less manual toil across the org, not more tickets closed.
Nice to have
  • Experience with service mesh (Istio/Envoy), GitOps workflows (ArgoCD/Flux), or policy-as-code (OPA).
  • Experi
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer/Cloud Platform Engineer - Operations (PST Timezone)
Site Reliability Engineer/Cloud Platform Engineer - Operations (PST Timezone)

skyflow • India

On-site
INR 2,500,000 - 3,500,000
Infra Team Manager
Infra Team Manager

KRAFTON INDIA • Bengaluru

On-site
INR 4,000,000 - 7,000,000
Senior Platform Engineer
Senior Platform Engineer

LE300 Optiva (India) Technologies Pvt. Ltd. • Hyderabad

On-site
INR 1,800,000 - 2,500,000
Infra Team Manager
Infra Team Manager

Krafton • Bengaluru

On-site
INR 1,500,000 - 2,000,000
Platform Lead
Platform Lead

Wenger & Watson • Bengaluru

On-site
INR 1,200,000 - 1,800,000
Tazapay - Staff DevOps Engineer
Tazapay - Staff DevOps Engineer

Tazapay • Chennai District

On-site
INR 3,500,000 - 6,000,000
Lead DevOps Engineer
Lead DevOps Engineer

Lenskart • Gurugram District

On-site
INR 1,200,000 - 2,400,000
Cloud Platform Engineer
Cloud Platform Engineer

Acunor • Hyderabad

On-site
INR 1,800,000 - 3,200,000
Platform Engineer
Platform Engineer

Outmarket AI • India

On-site
INR 2,500,000 - 4,000,000
Cloud Platform Engineer
Cloud Platform Engineer

Careernet • Chennai District

On-site
INR 1,800,000 - 3,000,000