Site Reliability Engineer, Cloud Platform

Qualys

Maharashtra

On-site

INR 1,200,000 - 1,800,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Qualys is seeking a skilled professional to develop and support cloud platform services in Maharashtra, India. The ideal candidate will have over 4 years of experience managing distributed systems, along with expertise in programming languages such as Java, Python, or Go. Responsibilities include automating deployment processes, enhancing system performance through monitoring, and leading incident responses. Candidates must hold a relevant degree and possess strong knowledge in systems programming and security practices.

Qualifications

  • 4+ years of experience in running distributed systems at scale in production.
  • Expertise in Java, Python, or Go.
  • Proficient in bash scripting.
  • Good understanding of SQL and NoSQL systems.
  • Knowledge of JVM concepts like garbage collection.
  • Demonstrated experience with production issue handling.

Responsibilities

  • Develop and maintain cloud platform services throughout their lifecycle.
  • Automate production system changes and evaluate results.
  • Support the cloud platform's production release activities.
  • Ensure cloud technology maintenance through performance monitoring.
  • Participate in incident responses, leading investigations and reports.

Skills

Distributed systems management
Java
Python
Bash scripting
SQL
NoSQL
Systems programming
Network elements
Performance tuning
Security best practices

Education

BS/MS degree in Computer Science or related field

Job description

Description

Co-develop and participate in the full lifecycle development of cloud platform services from inception and design, deployment, operation and improvement by applying scientific principles.

Increase the effectiveness, reliability and performance of cloud platform technologies by identifying and measuring key indicators, making changes to the production systems in an automated way and evaluating the results.

Support the cloud platform team before technologies are pushed for production release through activities such as system design, capacity planning, automation of key deployments, building a strategy for production monitoring and alerting, and participating in the testing/verification process.

Ensure that the cloud platform technologies are maintained properly by measuring and monitoring availability, latency, performance and system health.

Advise the cloud platform team to improve the reliability of systems in production and scale them based on need.

Participate in the development process by supporting new features, services, releases and hold an ownership mindset for the cloud platform technologies.

Develop tools and automate the process for achieving large scale provisioning and deployment of cloud platform technologies.

Participate in on-call rotation for cloud platform technologies. At times of incidents, lead incident response and be part of writing detailed postmortem analysis reports which are brutally honest with no-blame.

Propose improvements and drive efficiencies in systems and processes related to capacity planning, configuration management, scaling services, performance tuning, monitoring, alerting and root cause analysis.

Requirements
  • 4+ years of relevant experience in running distributed systems at scale in production.
  • Expertise in one of the programming languages: Java, Python, or Go.
  • Proficient in writing bash scripts.
  • Good understanding of SQL and NoSQL systems.
  • Good understanding of systems programming (network stack, file system, OS services).
  • Understanding of network elements such as firewalls, load balancers, DNS, NAT, TLS/SSL, VLANs, etc.
  • Skilled in identifying performance bottlenecks, identifying anomalous system behavior, and determining the root cause of incidents.
  • Knowledge of JVM concepts like garbage collection, heap, stack, profiling, class loading, etc.
  • Knowledge of best practices related to security, performance, high‑availability, and disaster recovery.
  • Demonstrate a proven record of handling production issues, planning escalation procedures, conducting post‑mortems, impact analysis, risk assessments and other related procedures.
  • Able to drive results and set priorities independently.
  • BS/MS degree in Computer Science, Applied Math or related field.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Platform Engineer
Platform Engineer

LambdaTest is now TestMu AI • Dadri

On-site
INR 1,200,000 - 1,800,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Keka Technologies Private Limited • Bengaluru

On-site
INR 4,000,000 - 7,000,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Akamai Technologies • Bengaluru

On-site
INR 2,400,000 - 3,400,000
Senior Cloud Site Reliability Engineer
Senior Cloud Site Reliability Engineer

Augusta Infotech • Bengaluru

Hybrid
INR 1,500,000 - 2,500,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Five9 • Bengaluru

On-site
INR 4,000,000 - 7,000,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

iLink Digital • Chennai

On-site
INR 1,200,000 - 1,800,000
Lead Site Reliability Engineer
Lead Site Reliability Engineer

Sierra Ventures • Bengaluru

On-site
INR 3,500,000 - 5,500,000
Site Reliability Engineer
Site Reliability Engineer

Indihire Consultants • Hyderabad

Hybrid
INR 1,500,000 - 2,800,000
Site Reliability Engineer(SRE)
Site Reliability Engineer(SRE)

MetaForgeIT • Hyderabad

On-site
INR 1,000,000 - 1,500,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Headout • Bengaluru

On-site
INR 1,200,000 - 1,800,000