A healthcare institution in Karachi seeks a Site Reliability Engineer to ensure system reliability while applying software engineering principles. You will automate workflows, manage incidents, and maintain uptime in production environments. The role requires expertise in cloud platforms and strong scripting abilities, along with collaboration across development teams. Suitable candidates should have at least 2 years of experience in similar roles, with a strong focus on scalability and performance.
Qualifications
2 years of experience in Site Reliability Engineering or related field.
Strong expertise in cloud platforms and automation tools.
Ability to work collaboratively across teams.
Responsibilities
Monitor and maintain reliability of critical production systems.
Automate infrastructure tasks to eliminate operational toil.
Lead incident response and conduct post-incident reviews.
Define and track SLIs, SLOs, and error budgets.
Build and maintain CI/CD pipelines and deployment strategies.
Implement observability using metrics, logs, and traces.
Collaborate with developers to embed reliability in design.
Conduct chaos engineering experiments to identify system weaknesses.
Skills
Proficiency in Python, Go, or Bash scripting languages
Expertise in AWS, GCP, or Azure cloud platforms
Strong knowledge of Kubernetes, Docker, and containerization
Experience with Prometheus, Grafana, and Datadog monitoring
Infrastructure as Code skills using Terraform or CloudFormation
Strong problem-solving, communication, and cross-team collaboration
Job description
# Site Reliability EngineerMarch 26, 20262515000 - 3655000 / year### Job DescriptionA Site Reliability Engineer applies software engineering principles to infrastructure and operations, ensuring system reliability, scalability, and performance. SREs bridge development and operations, automating workflows, managing incidents, and maintaining uptime across production environments at scale.### Key Responsibilities* Monitor and maintain reliability of critical production systems.* Automate infrastructure tasks to eliminate operational toil.* Lead incident response and conduct post-incident reviews.* Define and track SLIs, SLOs, and error budgets.* Build and maintain CI/CD pipelines and deployment strategies.* Implement observability using metrics, logs, and traces.* Collaborate with developers to embed reliability in design.* Conduct chaos engineering experiments to identify system weaknesses.### Skill & Experience* Proficiency in Python, Go, or Bash scripting languages.* Expertise in AWS, GCP, or Azure cloud platforms.* Strong knowledge of Kubernetes, Docker, and containerization.* Experience with Prometheus, Grafana, and Datadog monitoring.* Infrastructure as Code skills using Terraform or CloudFormation.* Strong problem-solving, communication, and cross-team collaboration.**Note**: Salary depends on experience and skills and is paid in local currency.LocationExperience2 Year## Job Location## Job Skills### Location: