Site Reliability Engineer

ValueFirst

Gurugram District

On-site

INR 800,000 - 1,200,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

ValueFirst is seeking a Site Reliability Engineer in Gurugram, India, to ensure the reliability and performance of their CPaaS systems. In this role, you'll manage the high availability of production systems, optimize messaging services, and automate operational tasks. Candidates should have a B.Tech/B.E in Computer Science, 2–3 years of relevant experience, and strong skills in Linux, MySQL, and cloud technologies. Responsibilities include monitoring system performance and handling production incidents.

Qualifications

  • 2–3 years of experience in SRE, DevOps, telecom, or CPaaS operations.
  • Solid understanding of networking fundamentals and production troubleshooting.
  • Proficiency in automation and reliability engineering.

Responsibilities

  • Ensure high availability and reliability of CPaaS production systems.
  • Own and improve SLIs, SLOs, and SLAs for messaging platforms.
  • Monitor system health, latency, and error rates using observability tools.

Skills

SMS gateways experience
Linux systems knowledge
MySQL administration
Shell scripting
tcpdump and Wireshark
Monitoring and alerting systems
Configuration management tools
Cloud platforms knowledge
Incident management skills

Education

B.Tech / B.E in Computer Science or related field

Tools

MySQL
MongoDB
Ansible
Datadog
ELK
Grafana
Apache
Nginx
Jira

Job description

The Site Reliability Engineering (SRE) team is responsible for ensuring the reliability, scalability, and performance of large-scale telecom and CPaaS platforms. This role combines software engineering and systems operations to build resilient, observable, and automated infrastructure that supports high-throughput messaging services. The team operates in a 24/7 environment and works closely with Engineering, CX and Products to maintain carrier-grade service reliability.

What you’ll be responsible for
  • Ensure high availability, performance, and reliability of CPaaS production systems speread across mutiple locations hosted over cloud and data centers
  • Own and improve SLIs, SLOs, and SLAs for messaging platforms and supporting services.
  • Monitor system health, latency, TPS, error rates, and delivery metrics using observability tools.
  • Participate in on‑call rotations and handle production incidents with a focus on fast recovery and root cause analysis.
  • Deploy, configure, and optimize for high-throughput messaging (multiple channels)
  • Troubleshoot telecom-specific issues including DLR failures, encoding problems, TPS dropsand routing issues.
  • Work directly with multiple teams for integrations, testing, and incident resolution.
  • Perform packet-level analysis using tcpdump and Wireshark to diagnose network and protocol‑level issues.
  • Write and maintain shell scripts and automation to eliminate repetitive operational tasks and reduce human intervention.
  • Contribute to infrastructure automation using tools like Ansible and CI/CD pipelines where applicable.
  • Improve deployment, configuration, and rollback processes for messaging services.
  • Design and enhance monitoring, alerting, and dashboards using tools such as Datadog, Site24x7, ELK and Grafana.
  • Administer and troubleshootLinux based servers in production environments.
  • Manage and optimize MySQL and MongoDB databases including performance tuning, backups, and recovery.
  • Works on API's and webhooks across the product & services. Its enhancements and troubleshooting.
  • Maintain web and application servers such as Apache, Nginx, and jboss (WildFly)
  • Support cloud-based and virtualized environments with exposure to auto-scaling and containerization concepts.
  • Collaborate with engineering teams on release planning, production deployments, and post‑release validation.
  • Lead or contribute to incident response & RCA focusing on long‑term reliability improvements.
  • Track issues, changes, and reliability work using Jira and related tools.
What you’d have
  • B.Tech / B.E in Computer Science or related field with 2–3 years of experience in SRE, DevOps, telecom, or CPaaS operations.
  • Hands‑on experience with SMS gateways and messaging workflows.
  • Solid understanding of Linux systems, networking fundamentals, and production troubleshooting.
  • Strong experience with MySQL & MongoDB administration, queries, and performance optimization.
  • Proficiency in shell scripting and a mindset toward automation and reliability engineering.
  • Hands‑on experience with tcpdump, Wireshark, and protocol‑level troubleshooting.
  • Experience with monitoring, logging, and alerting systems (Datadog, ELK, Grafana, Site24x7, etc.).
  • Familiarity with configuration management tools like Ansible and version control systems (Git).
  • Working knowledge of cloud platforms, virtualization, auto-scaling, and containerization.
  • Strong incident management, analytical thinking, and communication skills.
  • Certifications such as RHCE, AWS, or SRE-related credentials are a plus
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

Spot Your Leaders & Consulting • Pune District

On-site
INR 2,500,000 - 4,000,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Zorba AI • Chennai District

On-site
INR 1,200,000 - 2,400,000
SRE Lead
SRE Lead

Manatal • Bengaluru

On-site
INR 1,500,000 - 2,500,000
Lead Site Reliability Engineer
Lead Site Reliability Engineer

Sierra Ventures • Bengaluru

On-site
INR 3,500,000 - 5,500,000
SRE - AWS, GCP & Azure
SRE - AWS, GCP & Azure

PibyThree • Thane

On-site
INR 1,200,000 - 1,500,000
SRE - AWS, GCP & Azure
SRE - AWS, GCP & Azure

PibyThree • Navi Mumbai

On-site
INR 1,200,000 - 1,800,000
Resilience and Reliability Engineer
Resilience and Reliability Engineer

EY • Pune District, Gurugram District, Bengaluru

Hybrid
INR 1,800,000 - 2,800,000
Lead SRE
Lead SRE

Cvent, Inc. • India

On-site
INR 2,500,000 - 4,500,000
Site Reliability Engineer
Site Reliability Engineer

Metlife • Hyderabad

Hybrid
INR 1,500,000 - 2,600,000
Senior Associate Site Reliability Engineer
Senior Associate Site Reliability Engineer

NTT DATA BUSINESS SOLUTIONS • Hyderabad

On-site
INR 1,400,000 - 2,000,000