SRE Engineer II/III

Pace Industries, LLC

Hyderabad

On-site

INR 1,500,000 - 2,300,000

Full time

3 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Pace Industries, LLC is seeking a Site Reliability Engineer II/III to join our Hyderabad team. This role focuses on maintaining reliability, availability, and performance of production environments, with on-call responsibilities and cross-functional collaboration with engineering and operations.

You will drive root cause analysis, reduce toil via scripting and tooling, and own incident management across L2/L3. Prior experience in cloud, automation, and observability is essential for success.

Qualifications

  • Hands-on with Linux (RHEL/Ubuntu) and Windows Server basics.

Responsibilities

  • Own L2/L3 incident response, RCA, and post-mortems for production issues.
  • Monitor system health and ensure SLA/SLO adherence.
  • Automate operational tasks to reduce toil and improve reliability.
  • Collaborate with development teams on deployment reliability and capacity planning.
  • Participate in on-call rotations and maintain runbooks.

Skills

Linux systems
TCP/IP troubleshooting
HTTP/HTTPS debugging
Automation mindset
On-call experience

Tools

Terraform
Ansible
Kubernetes
Docker
Prometheus
Grafana
Datadog

Job description

## SRE Engineer II/IIIApply: India, Hyderabad: Full time: Posted Today: R-23604**Hiring:** Site Reliability Engineer**Location:** Hyderabad **Work Mode:** Work from Office **Work Shift:** 24/7 **Experience:** 3–8 Years **Level:** L2/L3 Engineer**About the Role**We are looking for a **Site Reliability Engineer – L2/L3** to join our team in Hyderabad. The ideal candidate will be responsible for maintaining the reliability, availability, and performance of production environments while working closely with engineering and operations teams. You serve as an escalation point, drive root cause analysis, and reduce toil through scripting and tooling.**Key Responsibilities*** Own L2/L3 incident response, RCA, and post-mortems for production issues.* Monitor system health and maintain SLA/SLO adherence.* Automate operational tasks to eliminate repetitive toil.* Collaborate with dev teams on deployment reliability and capacity planning.* Participate in on-call rotation and maintain runbooks.**Required Skills:****Operating Systems** Hands-on with Linux (RHEL/Ubuntu) — system, process management, file systems, performance tuning. Working knowledge of Windows Server and event log analysis.**Cloud** Practical experience on AWS / Azure / GCP — compute, storage, IAM, networking, and managed services. Familiarity with Terraform or equivalent IaC tools.**Scripting & Automation** Proficiency in Python and Bash for automation, API interaction, and operational tooling. Exposure to Ansible or similar config management is a plus.**Network Troubleshooting (In-Depth)** Strong command of TCP/IP internals — handshake lifecycle, connection states (TIME\\_WAIT, CLOSE\\_WAIT, SYN\\_FLOOD), packet flow, and socket behavior. Hands-on with tools like tcpdump, Wireshark, netstat/ss, traceroute, mtr, and dig. Solid understanding of DNS resolution, TLS/SSL negotiation, NAT, firewalls, and routing. Able to diagnose latency, packet loss, port exhaustion, and network-level bottlenecks at the OS and infrastructure layer.**Application & HTTP Troubleshooting** Deep understanding of HTTP/HTTPS methods, status codes, headers, and request lifecycle. Comfortable debugging through curl, Postman, access logs, and reverse proxy configs (Nginx / HAProxy).**Observability** Experience with Prometheus, Grafana, Datadog, or ELK. Ability to build dashboards, configure meaningful alerts, and trace issues end-to-end.Good to Have* Kubernetes / Docker experience.* Familiarity with message queues (Kafka, RabbitMQ).* Basic database troubleshooting (MySQL / PostgreSQL / Redis).* ITIL fundamentals and ITSM tools (Jira SM / ServiceNow).Soft SkillsStrong analytical thinking, clear communication under pressure, and a bias toward automation and continuous improvement.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

SRE Engineer II
SRE Engineer II

Webhosting • Hyderabad

On-site
INR 1,400,000 - 2,000,000
SRE
SRE

Metlife • Hyderabad

Hybrid
INR 1,500,000 - 2,100,000
SRE Engineer II/III
SRE Engineer II/III

Rackspace Technology • Hyderabad

On-site
INR 1,000,000 - 1,800,000
SRE Reliability Engineer
SRE Reliability Engineer

NTT DATA BUSINESS SOLUTIONS • Bengaluru

On-site
INR 2,500,000 - 4,000,000
Site Reliability Engineer
Site Reliability Engineer

Arch Systems • Hyderabad

On-site
INR 2,800,000 - 4,200,000
SRE - Site Reliability Engineering
SRE - Site Reliability Engineering

Build & Hire • Pune District

On-site
INR 1,500,000 - 2,300,000
SRE Engineer @ Investment Banking | Mumbai
SRE Engineer @ Investment Banking | Mumbai

Net Connect Global • Bengaluru, Mumbai

Hybrid
INR 1,800,000 - 2,400,000
SRE - 3i Infotech - BKC
SRE - 3i Infotech - BKC

3i Infotech • Mumbai Suburban, Vasai-Virar

On-site
INR 900,000 - 1,000,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

VMC Soft Technologies, Inc • Hyderabad

Hybrid
INR 1,500,000 - 2,000,000
SRE Lead
SRE Lead

3across • Bengaluru

Hybrid
INR 1,500,000 - 2,300,000