Azure Platform Site Reliability Engineer

Foss United

India

On-site

INR 2,800,000 - 4,600,000

Full time

6 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Foss United seeks an experienced Reliability & Performance Engineer to ensure Azure platforms run with high reliability, scalability, and efficiency. You will design and maintain scalable Azure infrastructure and implement IaC with Terraform and Ansible, while building robust monitoring and incident response practices.

You will lead cost optimization, capacity planning, and security/compliance across Azure environments, collaborating with data scientists and developers to deliver resilient

Qualifications

  • Extensive experience with Azure services for reliability, availability and performance.
  • 8+ years in site reliability engineering, cloud operations or similar with a focus on Azure technologies.
  • Proficiency in scripting (PowerShell, Python) and CI/CD tooling.
  • Certifications in Azure (AZ-305 or equivalent) are preferred.

Responsibilities

  • Ensure reliability, availability, and performance of Azure-based platforms and services.
  • Design, deploy, and manage scalable Azure infrastructure using IaC (Terraform, Ansible).
  • Implement monitoring, alerting, and incident response strategies to mitigate issues.
  • Optimize costs and capacity, perform capacity planning and forecasting.
  • Lead incident response, document procedures, and mentor team members.
  • Collaborate with data scientists and developers to deliver reliable AI and analytics solutions.
  • Stay current with Azure, data analytics, ML, AI trends and introduce improvements.

Skills

Azure
AKS
Terraform
Ansible
Azure DevOps
Python
PowerShell
CI/CD
Monitoring
Security best practices

Education

Bachelor’s degree in CS/IT
8+ years in SRE/Cloud Ops

Tools

Terraform
Ansible
Azure DevOps
PowerShell
Python

Job description

Reliability & Performance: Ensure the reliability, availability, and performance of Azurebased platforms and services. Implement monitoring, alerting, and incident response strategies to address and mitigate issues proactively.

Infrastructure Management: Design, deploy, and manage scalable and fault-tolerant Azure infrastructure.Utilize Infrastructure as Code (IaC) tools such as Terraform and Ansible for automated provisioning and configuration.

Data Analytics: Ensure high performance and availability of data pipelines and analytics platforms.

Machine Learning & Generative AI: Using AIOps to ensure these systems in Azure Platform are scalable, secure, and optimized for performance.

AKS Management: Architect, deploy, and manage Azure Kubernetes Service (AKS). Optimize AKS clusters for performance, scalability, and cost-efficiency, and ensure best practices for container orchestration and management.

Automation & CI/CD: Develop and maintain automation workflows using Terraform and Ansible. Implement and manage CI/CD pipelines with Azure DevOps to streamline deployment processes and ensure continuous integration and delivery for day to day usage by project enablement team and making sure that the modules are kept up to date w.r.t. version and new policies introduced in the environment.

Monitoring & Observability: Implement and maintain comprehensive monitoring and observability solutions using Azure Monitor, Application Insights, and other tools. Analyze metrics, logs, and traces to identify and resolve performance bottlenecks and reliability issues.

Cost Optimization: Analyze and optimize Azure costs by implementing reservations, savings plans, and other cost-management strategies. Monitor usage patterns and provide recommendations for resource optimization and cost reduction.

Capacity Planning: Perform capacity planning and forecasting to ensure adequate resources are available to meet demand. Implement scaling strategies and optimize resource utilization to balance performance and cost.

Security & Compliance: Ensure that Azure environments adhere to security best practices and compliance requirements. Implement security measures, conduct regular audits, and address vulnerabilities proactively.

Documentation & Knowledge Sharing: Create and maintain detailed documentation for operational processes, incident response procedures, and infrastructure designs. Share knowledge and provide training to team members and stakeholders.

Collaboration & Stakeholder Engagement: Work closely with development teams, data scientists, and other stakeholders to understand requirements and deliver solutions that meet their needs. Communicate effectively on operational status, incidents, and improvements.

Continuous Improvement: Stay current with emerging Azure technologies, data analytics, machine learning, generative AI, and industry trends. Identify opportunities for innovation and contribute to the development of new tools, processes, and best practices to enhance platform reliability and performance.

Required Qualifications:

Technical Expertise: Extensive experience with Azure services, including Azure Data Analytics (Synapse Analytics, Data Lake, Data Factory, Power BI), Azure Machine Learning, Azure Cognitive Services, Generative AI technologies, and Azure Kubernetes Service (AKS). Proficiency in Terraform, Ansible, and Azure DevOps.

Experience: 8+ years of experience in site reliability engineering, cloud operations, or a similar role with a focus on Azure technologies. Proven track record of managing large-scale, high-availability systems and supporting data analytics and AI solutions.

Skills: Strong problem-solving and troubleshooting skills, with experience in incident management, performance optimization, and automation. Proficiency in scripting languages (e.g., PowerShell, Python) and CI/CD tools.

Certifications: Microsoft Certified: Azure Solutions Architect Expert (AZ-305) or equivalent advanced certification required. Additional certifications in Azure Data Engineering, Machine Learning, or DevOps are a plus.

Desired Attributes:

Operational Excellence: Demonstrated ability to maintain high standards of reliability and performance in complex cloud environments.

Cost Optimization: Experience with cost management strategies, including reservations and savings plans, to optimize Azure expenditures.

Leadership: Ability to lead incident response efforts, mentor team members, and drive continuous improvement initiatives.

Customer Focus: Strong commitment to delivering high-quality solutions that meet stakeholder needs and enhance user experience.

Innovative Mindset: Ability to drive innovation in data analytics, AI, and container management, applying creative solutions to complex challenges.

Team Collaboration: Excellent interpersonal skills with the ability to work effectively with cross-functional teams and foster a collaborative work environment.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer (Azure) - S
Site Reliability Engineer (Azure) - S

Tata Consultancy Services • Bengaluru

On-site
INR 2,500,000 - 4,500,000
Senior Site Reliability Engineer (Azure) - S
Senior Site Reliability Engineer (Azure) - S

Tata Consultancy Services • Kolkata District, Chennai District, Bengaluru

On-site
INR 2,400,000 - 4,200,000
Sr DevOps Engineer II
Sr DevOps Engineer II

MetLife • Maharashtra

On-site
INR 2,500,000 - 4,000,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Clarus Advisers • Hyderabad

On-site
INR 1,800,000 - 2,800,000
Lead Engineer - Cloud Engg & AI
Lead Engineer - Cloud Engg & AI

Anblicks Inc. • Hyderabad

On-site
INR 4,200,000 - 6,500,000
Platform Manager
Platform Manager

ACG • Mumbai

On-site
INR 3,500,000 - 7,000,000
Senior Platform Engineer
Senior Platform Engineer

Cube Asia • Bengaluru

On-site
INR 1,600,000 - 2,600,000
Site Reliability Engineer
Site Reliability Engineer

PwC India • Bengaluru

On-site
INR 1,500,000 - 2,500,000
Senior Azure Administrator / DevOps Engineer
Senior Azure Administrator / DevOps Engineer

Competent Groove Private Limited • Mohali

On-site
INR 3,500,000 - 6,000,000
Cloud Platform Engineer -(Terraform | Azure | AKS | Networking)
Cloud Platform Engineer -(Terraform | Azure | AKS | Networking)

Qnity • Hyderabad

On-site
INR 1,500,000 - 3,000,000