AI-Ops Specialist (SRE Infrastructure)

Softility Tech

Hyderabad

Hybrid

INR 3,500,000 - 5,200,000

Full time

6 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Softility Tech seeks experienced AI-OPS Engineers blending SRE, cloud operations, automation, observability, and GenAI. You will design intelligent ops solutions to improve incident detection, diagnosis, remediation, and efficiency.

Candidates should have 7+ years in Infra/SRE/DevOps, 3+ years on cloud platforms (AWS/GCP), and hands-on automation with self-healing capabilities. AWS/GCP, Kubernetes, and AI governance are preferred.

Qualifications

  • 7+ years in Infrastructure Operations, SRE, Platform Engineering, DevOps, or AIOps.
  • 3+ years working with Cloud Platforms (AWS/GCP).
  • Experience implementing automation and self-healing solutions.
  • Experience building AI-assisted operational workflows.
  • Experience with enterprise monitoring and observability platforms.

Responsibilities

  • Design and implement self-healing operational workflows.
  • Develop AI-assisted RCA and operational intelligence capabilities.
  • Build and maintain knowledge and runbook copilots.
  • Improve monitoring, observability, and incident response processes.
  • Automate operational tasks using Infrastructure as Code and orchestration tools.
  • Collaborate with SRE, platform, and application teams to improve reliability and operational efficiency.
  • Evaluate and implement Agentic AI solutions for autonomous operations.

Job description

Role Summary

We are seeking experienced AI-OPS Engineers with a strong blend of Site Reliability Engineering (SRE), Cloud Operations, Automation, Observability, and Generative AI expertise. Candidates should be capable of designing and implementing intelligent operational solutions that improve incident detection, diagnosis, remediation, and operational efficiency.


Required Technical Skills

AI/ML & Generative AI


  • OpenAI / GenAI solutions on GCP or AWS

  • Machine Learning fundamentals

  • MLOps

  • AI Agents / Agentic AI

  • Retrieval Augmented Generation (RAG)

  • Enterprise Search

  • AI Governance

  • Prompt Engineering


Programming & Automation


  • Python (strong requirement)

  • PowerShell

  • Ansible

  • Terraform

  • Infrastructure as Code (IaC)

  • DevOps practices


Cloud & Platform Engineering


  • GCP and/or AWS

  • Kubernetes

  • Container-based platforms

  • Cloud-native operational tooling


Observability & AIOps


  • Splunk

  • Observability platforms and monitoring tools

  • Incident Analytics

  • Event Correlation

  • RCA Analytics

  • Predictive Alerting

  • Self-Healing Automation


SRE & IT Operations


  • Reliability Engineering

  • Incident Management

  • Production Support

  • ITSM platforms (Remedy preferred)

  • Problem Management

  • Operational Excellence


Preferred Experience


  • 7+ years in Infrastructure Operations, SRE, Platform Engineering, DevOps, or AIOps

  • 3+ years working with Cloud Platforms (AWS/GCP)

  • Experience implementing automation and self-healing solutions

  • Experience building AI-assisted operational workflows

  • Experience with enterprise monitoring and observability platforms


Preferred Certifications


  • AWS Certified Solutions Architect / DevOps Engineer

  • Google Professional Cloud Architect

  • Certified Kubernetes Administrator (CKA)

  • ITIL Foundation

  • AI/ML or Generative AI certifications


Key Responsibilities


  • Design and implement self-healing operational workflows.

  • Develop AI-assisted RCA and operational intelligence capabilities.

  • Build and maintain knowledge and runbook copilots.

  • Improve monitoring, observability, and incident response processes.

  • Automate operational tasks using Infrastructure as Code and orchestration tools.

  • Collaborate with SRE, platform, and application teams to improve reliability and operational efficiency.

  • Evaluate and implement Agentic AI solutions for autonomous operations.


Ideal Candidate Profile

Candidates should possess a strong combination of:



  • SRE/Operations background

  • Cloud Engineering expertise

  • Automation and Infrastructure as Code experience

  • Observability and Incident Management knowledge

  • AI/ML and Generative AI capabilities

  • Excellent troubleshooting, RCA, and problem-solving skills


Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior AI Ops Architect
Senior AI Ops Architect

EPAM Systems • Hyderabad

On-site
INR 3,000,000 - 4,200,000
Senior AI Ops Architect
Senior AI Ops Architect

EPAM Systems • Maharashtra

On-site
INR 3,500,000 - 6,000,000
AIOps leadership experience
AWS ML certs
Senior AI Ops Architect
Senior AI Ops Architect

EPAM Systems • Bengaluru

On-site
INR 4,200,000 - 7,000,000
AI Solutions and Platforms Operations Engineer
AI Solutions and Platforms Operations Engineer

PepsiCo • Hyderabad

On-site
INR 1,000,000 - 2,000,000
AI Site Reliability Engineer (AI SRE)
AI Site Reliability Engineer (AI SRE)

Elfonze Technologies • India

On-site
INR 3,500,000 - 5,500,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Clarus Advisers • Hyderabad

On-site
INR 1,800,000 - 2,800,000
AI Engineer
AI Engineer

Allegis Group • Hyderabad, Bengaluru

Hybrid
INR 1,200,000 - 2,500,000
Systems Architect - Cloud AIOps
Systems Architect - Cloud AIOps

EPAM Systems • Gurugram District

On-site
INR 4,200,000 - 6,600,000
Systems Architect - Cloud AIOps
Systems Architect - Cloud AIOps

EPAM Systems • Maharashtra

On-site
INR 4,200,000 - 7,000,000
Systems Architect - Cloud AIOps
Systems Architect - Cloud AIOps

EPAM Systems • Coimbatore District

On-site
INR 5,500,000 - 7,500,000