Site Reliability Specialist - High Availability

Freelanceshop

Gwalior District

Hybrid

INR 1,200,000 - 2,000,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Competitive salary
Health and life insurance
Flexible work arrangements
Professional development programs
Paid time off and wellness programs

Job summary

Freelanceshop is looking for a proficient Site Reliability Specialist – High Availability to enhance our technology operations. This critical role demands expertise in designing and maintaining resilient systems, working with cloud platforms, and automating tasks.

The candidate will ensure system reliability and efficiency while collaborating with cross-functional teams. A Bachelor's degree and 4–8 years of industry experience are required. Flexibility in working hours and hybrid work options are part of the offer.

Qualifications

  • Strong experience in Site Reliability Engineering (SRE), DevOps, or Production Engineering roles.
  • Hands-on expertise with cloud platforms such as AWS, Azure, or Google Cloud.
  • Solid understanding of high availability architectures, load balancing, clustering, and failover mechanisms.

Responsibilities

  • Design, implement, and manage highly available and fault-tolerant systems.
  • Monitor system performance, availability, latency, and capacity.
  • Lead root cause analysis (RCA) for major incidents.

Skills

Site Reliability Engineering
DevOps
Production Engineering
Cloud Platforms (AWS, Azure, Google Cloud)
Scripting (Python, Bash, Go, Java)
Containerization (Docker, Kubernetes)
Monitoring Tools (Prometheus, Grafana, ELK)
Networking Concepts

Education

Bachelor's degree in Computer Science, Engineering or related field

Tools

Jenkins
GitLab CI
Terraform
Ansible

Job description

Job Summary

Global MNC Tech is seeking a highly skilled and proactive Site Reliability Specialist – High Availability to join our global technology operations team. This role is critical in ensuring the reliability, scalability, and performance of our mission‑critical systems that support millions of users worldwide. You will work at the intersection of software engineering and IT operations, applying engineering principles to build resilient, self‑healing, and highly available platforms.

As a Site Reliability Specialist, you will be responsible for designing, implementing, and maintaining systems that meet strict uptime and performance targets. You will collaborate closely with development, infrastructure, security, and business teams to continuously improve system reliability while driving automation and operational excellence.

Key Responsibilities
  • Design, implement, and manage highly available and fault‑tolerant systems across cloud and hybrid environments.
  • Monitor system performance, availability, latency, and capacity using advanced observability tools.
  • Develop and maintain automated monitoring, alerting, and incident response frameworks.
  • Lead root cause analysis (RCA) for major incidents and implement long‑term corrective actions.
  • Drive automation of operational tasks using scripting and infrastructure‑as‑code (IaC) tools.
  • Participate in on‑call rotations and provide support for production systems to ensure 24/7 reliability.
  • Collaborate with engineering teams to improve system architecture and deploy best practices for high availability.
  • Conduct regular disaster recovery (DR) drills and ensure business continuity plans are up to date.
  • Optimize system performance, reduce downtime, and improve mean time to recovery (MTTR).
  • Define and track Service Level Objectives (SLOs), Service Level Indicators (SLIs), and error budgets.
Required Skills and Qualifications
  • Strong experience in Site Reliability Engineering (SRE), DevOps, or Production Engineering roles.
  • Hands‑on expertise with cloud platforms such as AWS, Azure, or Google Cloud.
  • Solid understanding of high availability architectures, load balancing, clustering, and failover mechanisms.
  • Proficiency in scripting and programming languages such as Python, Bash, Go, or Java.
  • Experience with containerization and orchestration tools (Docker, Kubernetes).
  • Knowledge of monitoring and observability tools like Prometheus, Grafana, ELK, Datadog, or New Relic.
  • Familiarity with CI/CD pipelines and automation tools (Jenkins, GitLab CI, Terraform, Ansible).
  • Strong understanding of networking concepts, security principles, and system performance tuning.
Experience
  • Bachelors degree in Computer Science, Information Technology, Engineering, or a related field.
  • 4–8 years of experience in SRE, DevOps, Systems Engineering, or similar roles.
  • Proven track record of managing high‑availability production systems in large‑scale environments.
  • Experience working in fast‑paced, high‑growth technology organizations is highly desirable.
Working Hours
  • Full‑time position with flexible working hours.
  • Participation in 24/7 on‑call rotation as part of a global support team.
  • Hybrid or remote work options depending on business requirements and location.
Knowledge, Skills and Abilities
  • Deep understanding of distributed systems and reliability engineering principles.
  • Strong analytical and problem‑solving skills with attention to detail.
  • Ability to work under pressure and manage critical incidents effectively.
  • Excellent communication skills and ability to collaborate with cross‑functional teams.
  • Strong documentation and knowledge‑sharing mindset.
  • Passion for automation, continuous improvement, and operational excellence.
Benefits
  • Competitive salary and performance‑based incentives.
  • Comprehensive health and life insurance coverage.
  • Flexible work arrangements and work‑life balance initiatives.
  • Professional development programs and technical training opportunities.
  • Access to cutting‑edge technologies and global projects.
  • Paid time off, holidays, and wellness programs.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

NTT DATA BUSINESS SOLUTIONS • Hyderabad

On-site
INR 1,800,000 - 3,200,000
Lead Site Reliability Engineer Expert Palo Alto And Versa SD‑WAN Experience
Lead Site Reliability Engineer Expert Palo Alto And Versa SD‑WAN Experience

SITA • India

Hybrid
INR 350,000 - 650,000
Flex Week
Flex-Location: up to 30 days remote
Site Reliability Engineer/ Expert/ Specialist (Must have strong experience in Windows Server, A[...]
Site Reliability Engineer/ Expert/ Specialist (Must have strong experience in Windows Server, A[...]

SITA • Delhi

Hybrid
INR 3,500,000 - 7,000,000
Flex Week: work from home up to 2 days
Flex Location: up to 30 days travel
Employee Wellbeing program
+2
Site Reliability Engineer
Site Reliability Engineer

Snapmint • Gurugram District

On-site
INR 800,000 - 1,200,000
Lead Site Reliability Engineer/ Expert (Palo Alto & Versa SD‑WAN Experience)
Lead Site Reliability Engineer/ Expert (Palo Alto & Versa SD‑WAN Experience)

SITA • Delhi

Hybrid
INR 350,000 - 700,000
Site Reliability Engineer
Site Reliability Engineer

SourcingXPress • Mumbai

On-site
INR 800,000 - 1,200,000
Application Support Engineer
Application Support Engineer

Accenture in India • Indore District

On-site
INR 1,200,000 - 2,000,000
Site Reliability Engineer
Site Reliability Engineer

Spot Your Leaders & Consulting • Pune District

On-site
INR 2,500,000 - 4,000,000
Site Reliability Engineer
Site Reliability Engineer

Arch Systems • Hyderabad

On-site
INR 2,800,000 - 4,200,000
Lead Site Reliability Engineer/ Expert
Lead Site Reliability Engineer/ Expert

SITA Group • Delhi

On-site
INR 1,200,000 - 2,400,000