Senior Site Reliability Engineer — CDN Infrastructure

Verge Cloud Pvt. Ltd.

India

On-site

INR 2,600,000 - 5,200,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Verge Cloud Pvt. Ltd. is seeking a Senior Site Reliability Engineer to operate and evolve the CDN infrastructure. You’ll work across baremetal edge servers, a Kubernetes-based control plane, and end-to-end observability pipelines to detect and resolve issues before customers are affected.

This hands-on, senior IC role demands moving between low-level networking debugging and platform automation, with the ability to drive decisions with minimal oversight.

Qualifications

  • 5+ years in an SRE/DevOps/infrastructure role.
  • Strong networking fundamentals — DNS, TCP/IP, and CDN concepts.
  • Solid experience operating Linux systems in production on self-managed baremetal infrastructure.
  • Hands-on with SaltStack or Ansible.
  • Experience with Kubernetes in production, including deploying via Helm.
  • Experience building CI/CD pipelines, ideally with GitHub Actions.
  • Working knowledge of Terraform or similar IaC tools.
  • Practical experience with Prometheus, Grafana, and Alertmanager; familiar with Grafana Tempo.
  • Ability to write Go or Python for automation.
  • Strong debugging skills from kernel/network to application levels.
  • Good communication in a small, high-ownership team.

Responsibilities

  • Operate and maintain our fleet of self-managed baremetal edge servers.
  • Use SaltStack (or Ansible) for configuration management and automation.
  • Manage and improve our managed Kubernetes cluster, deploying and maintaining services via Helm.
  • Build and maintain CI/CD pipelines (GitHub Actions) for infrastructure and service deployments.
  • Own and extend observability tooling — Prometheus, Grafana, Alertmanager, and Grafana Tempo.
  • Maintain and query ClickHouse at scale—millions of rows ingested daily, with focus on schema design and low-latency queries.
  • Manage cloud service configuration as code using Terraform.
  • Diagnose and resolve issues across the stack—from DNS resolution and BGP/routing to TCP/IP performance regressions and application problems.

Skills

Networking fundamentals
Linux in production
SaltStack
Ansible
Kubernetes
Helm
GitHub Actions
Terraform
Prometheus
Grafana
Alertmanager
Grafana Tempo
Go
Python
Debugging

Tools

ClickHouse

Job description

Senior Site Reliability Engineer — CDN Infrastructure

We're looking for a Senior SRE with 5+ years of experience to help operate and evolve

the infrastructure behind our CDN — from the baremetal edge nodes serving traffic, to

the Kubernetes-based control plane, to the observability and data pipelines that tell us

what's actually happening on the network. You'll work across the full stack: provisioning

and configuration management, CI/CD, Kubernetes, and the logging/monitoring systems

that let us catch problems before customers do.

This is a hands-on, senior individual-contributor role for someone who's comfortable

moving between low-level networking debugging and higher-level platform/infra

automation, and who can drive decisions with less oversight.

What You'll Do
  • Operate and maintain our fleet of self-managed baremetal edge servers, using
  • SaltStack (or Ansible) for configuration management and automation
  • Manage and improve our managed Kubernetes cluster, deploying and maintaining
  • services via Helm
  • Build and maintain CI/CD pipelines (GitHub Actions) for infrastructure and service
  • deployments
  • Own and extend observability tooling — Prometheus, Grafana, Alertmanager, and
  • distributed tracing with Grafana Tempo
  • Maintain and query ClickHouse at scale — millions of rows ingested daily, with a
  • focus on schema design and low-latency query tuning
  • Manage cloud service configuration as code using Terraform
  • Diagnose and resolve issues across the stack — from DNS resolution and
  • BGP/routing anomalies to TCP/IP-level performance regressions and application What
Requirements
What You'll Need
  • 5+ years of experience in an SRE, DevOps, or infrastructure/platform engineering role
  • Deep understanding of networking fundamentals — DNS, TCP/IP, and CDN concepts (caching, routing, anycast, edge delivery) — you should be comfortable reading a packet capture or debugging a DNS resolution chain
  • Solid experience operating Linux systems in production, including self-managed baremetal infrastructure
  • Hands-on experience with configuration management tools — SaltStack or Ansible
  • Experience with Kubernetes in production, including deploying and managing services via Helm
  • Experience building CI/CD pipelines, ideally with GitHub Actions
  • Working knowledge of Terraform or similar IaC tools
  • Practical experience with Prometheus, Grafana, and Alertmanager for monitoring and alerting; familiarity with distributed tracing (Grafana Tempo or similar)
  • Ability to write code in Go or Python for automation, tooling, or internal services
  • Strong debugging skills across layers — from kernel/network to application to
  • Good communication skills and comfort working in a small, high-ownership team
Nice to Have
  • Experience with BGP / anycast routing in a production CDN or network operator context
  • Experience with ClickHouse or another columnar/analytical database at scale
  • Experience with Kafka/Redpanda or similar streaming systems for log/data pipelines
  • Prior experience at a CDN, ISP, hosting provider, or similar network-heavy operator
  • Experience with GitOps workflows (ArgoCD or similar)
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer — CDN Infrastructure
Senior Site Reliability Engineer — CDN Infrastructure

VergeCloud • Bengaluru

On-site
INR 2,800,000 - 5,200,000
Site Reliability Engineer
Site Reliability Engineer

Innodata Inc. • India

On-site
INR 2,400,000 - 4,000,000
Senior Site Reliability Lead
Senior Site Reliability Lead

Generac • Pune District

On-site
INR 3,000,000 - 6,500,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

UST • Pune District

On-site
INR 1,800,000 - 3,000,000
Senior SRE Engineer
Senior SRE Engineer

Epam Systems • Bengaluru

On-site
INR 2,500,000 - 4,200,000
Site Reliability Engineer
Site Reliability Engineer

Indihire Consultants • Hyderabad

Hybrid
INR 1,500,000 - 2,800,000
Senior DevOps Engineer
Senior DevOps Engineer

Luxoft • Bengaluru

On-site
INR 3,000,000 - 6,000,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

NCR Voyix • Chennai District

On-site
INR 3,000,000 - 5,400,000
Lead DevOps Engineer
Lead DevOps Engineer

Lenskart • Gurugram District

On-site
INR 1,200,000 - 2,400,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Falabella India • Bengaluru

On-site
INR 4,000,000 - 7,000,000