AIOps Engineer
Experience: 3–8 Years
Location: Seattle, WA | San Francisco, CA | Austin, TX | New York, NY
Salary: $110,000–$190,000+ per year
Job Type: Full-Time
About The Role
We are seeking a skilled AIOps Engineer to build and optimize intelligent IT operations using AI, DevOps, automation, observability, and cloud technologies. You will help automate incident response, detect anomalies, improve system reliability, and reduce manual operational effort.
Key Responsibilities
- Design and implement AI-driven IT operations and automation solutions.
- Build automated workflows for monitoring, incident response, remediation, and deployment.
- Integrate AI/ML capabilities into DevOps and production operations.
- Analyze logs, metrics, traces, events, and telemetry to identify anomalies and operational issues.
- Develop automated root-cause analysis and intelligent incident-triage workflows.
- Build and maintain CI/CD pipelines and infrastructure automation.
- Create Python, Bash, or PowerShell scripts for operational automation.
- Work with cloud-native environments, containers, and Kubernetes.
- Integrate observability and monitoring platforms with AIOps workflows.
- Implement predictive monitoring, alert correlation, noise reduction, and automated remediation.
- Collaborate with DevOps, SRE, Cloud, Security, and Engineering teams.
- Continuously identify opportunities to improve reliability, performance, and operational efficiency.
Required Skills
- 3–8 years of experience in AIOps, DevOps, SRE, Cloud Engineering, or IT Operations.
- Strong understanding of AI/ML concepts and AI-powered automation.
- Strong DevOps and CI/CD experience.
- Proficiency in Python and/or Bash.
- Experience with Docker and Kubernetes.
- Hands‑on experience with AWS, Azure, or GCP.
- Experience with Terraform, Ansible, or similar automation tools.
- Knowledge of monitoring and observability tools such as Prometheus, Grafana, Datadog, Splunk, Elastic, or Dynatrace.
- Strong understanding of Linux, APIs, networking, and cloud infrastructure.
- Experience with incident management, troubleshooting, and root-cause analysis.
Preferred Qualifications
- Experience with LLMs, RAG, AI agents, or Generative AI.
- Experience integrating OpenAI or other LLM APIs into operational workflows.
- Knowledge of LangChain, LangGraph, or similar AI frameworks.
- Experience with ServiceNow, Jira, or ITSM platforms.
- Familiarity with MLOps and production AI systems.
- Experience building autonomous or AI-assisted remediation workflows.
What You'll Do
You will help transform traditional IT operations into intelligent, automated, and proactive operations, using AI and DevOps practices to improve reliability, accelerate incident resolution, and reduce operational overhead. Modern AIOps roles commonly combine telemetry integration, anomaly detection, automation, incident analysis, and operational optimization.
Skills: ai,devops,automation