SRE AIOps Engineer (cod: RN - DY)

Nybble Group

Northern (KY)

Hybrid

USD 120,000 - 180,000

Full time

10 days ago
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

2 weeks PTO
Birthday off
Sick leave
Study leave
Discounts platform
Connectivity expenses

Job summary

Nybble Group is seeking an SRE AIOps Engineer to join our team in Kentucky. You will build AI-powered automation to streamline Site Reliability workflows, blending software development with operations across CMS and cloud environments.

You will design AI/ML solutions for incident triage, develop Python services, manage IaC with Terraform/CloudFormation, and monitor health with Splunk, New Relic, and AppDynamics. Strong communication across time zones is essential.

Qualifications

  • 4–8 years in SRE/DevOps with a focus on automation.
  • 3–6 years supporting CMS platforms with hands-on AEM experience preferred.
  • Strong Python development: services, CLIs, integrations and tests.
  • Experience with AI/ML/LLM tools and integrating AI into operations.
  • Experience with AI/LLM APIs and AI agents.
  • Proficiency with AWS, Azure, or GCP.
  • Infrastructure as Code using Terraform, CloudFormation, or CDK.
  • Advanced monitoring/observability with AppDynamics, Splunk, New Relic.
  • Understanding of AEM Sites, DAM/Assets, authoring/publishing flows.

Responsibilities

  • Design AI/ML and LLM-based solutions to automate incident triage and remediation.
  • Develop Python-based services, CLIs, and scripts to automate SRE tasks.
  • Manage IaC for tooling using Terraform and CloudFormation.
  • Collaborate with SCOUT AI Tools and Atlassian ROVO for AI-driven ops.
  • Measure KPIs like MTTR/MTTD and reduce alert noise.
  • Serve as L2 contact for AEM Sites, AEM Assets, Adobe Target and integrations.
  • Validate tickets and perform incident analysis with log extracts.
  • Maintain runbooks, SOPs and knowledge articles; improve alerts/dashboards.

Skills

SRE
DevOps
Python
AI/ML
Cloud
IaC
Monitoring
AEM

Tools

Splunk
New Relic
AppDynamics
Terraform
CloudFormation
CDK
AEM Console

Job description

Job Overview:

We are looking for an SRE AIOps Engineer to join our growing team and help build AI-powered automation and tools that streamline Site Reliability Engineering workflows. The ideal candidate will combine strong software development and automation skills with hands‑on experience in SRE, DevOps, application support, and AI/LLM technologies.

Key Responsibilities:
  • Design and implement AI/ML and LLM-based solutions to automate incident triage and enrichment, generate incident summaries and stakeholder updates, and recommend remediation actions using runbooks and historical data.
  • Develop and maintain Python-based services, CLIs, and scripts to automate repetitive SRE tasks, execute runbooks, process logs, metrics, and events, and integrate with monitoring, ticketing, chat, and other external or internal tools.
  • Implement and manage Infrastructure as Code solutions using technologies such as Terraform and CloudFormation for SRE tooling and observability.
  • Work with SCOUT AI Tools and Atlassian ROVO to advance AI-driven operational capabilities.
  • Measure and improve operational performance through KPIs such as MTTR, MTTD, alert noise reduction, and on‑call load.
  • Serve as an L2 point of contact for AEM Sites, AEM Assets, Adobe Target, and related integrations.
  • Validate, categorize, and prioritize incoming tickets while assessing impact across pages, journeys, brands, and devices.
  • Perform application triage by reviewing AEM error logs and author/publish status, checking dispatcher/CDN behavior and cache‑related symptoms, and validating configuration and content changes.
  • Provide L3 and engineering teams with detailed incident descriptions, impact analysis, reproduction steps, and relevant log extracts.
  • Partner with content authors and marketing teams to diagnose content and publishing issues.
  • Monitor health dashboards and alerts using tools such as Splunk, New Relic, AppDynamics, and AEM consoles to proactively identify anomalies.
  • Maintain and improve runbooks, knowledge base articles, and standard operating procedures for common incident types.
  • Recommend improvements to alerts, dashboards, and thresholds based on recurring issues and operational trends.
Required Qualifications:
  • 4–8 years of experience in SRE, DevOps, or related roles, with a strong emphasis on automation and tooling.
  • 3–6 years of experience supporting CMS-based platforms, with hands‑on Adobe Experience Manager (AEM) experience strongly preferred.
  • Strong Python development skills, including experience building services, CLIs, integrations, and tests.
  • Experience working with AI/ML/LLM tools and integrating AI capabilities into operational workflows.
  • Experience with AI/LLM APIs and AI agents.
  • Proficiency with at least one major cloud platform such as AWS, Azure, or GCP.
  • Working knowledge of Infrastructure as Code technologies such as Terraform, CloudFormation, or CDK.
  • Advanced experience with monitoring and observability tools such as AppDynamics, Splunk, and New Relic.
  • Understanding of AEM Sites, including pages, components, templates, and content hierarchy.
  • Familiarity with AEM authoring and publishing flows, workflows, DAM/Assets, and related application functionality.
  • Experience diagnosing AEM application and content‑related issues.
  • Strong troubleshooting, analytical, and problem‑solving skills.
  • Excellent written and verbal English communication skills.
  • Ability to explain technical findings clearly to both technical and non‑technical stakeholders.
  • Comfortable working with distributed, cross‑functional teams across multiple time zones.
Preferred Qualifications:
  • Experience working with SCOUT AI Tools and Atlassian ROVO.
  • Experience using AI assistants for log summarization, pattern identification, incident analysis, and incident reporting.
  • Experience designing AI‑powered operational tools and automation.
  • Experience with Adobe Target and related AEM integrations.
  • Experience working with global, always‑on eCommerce platforms or other high‑availability environments.
  • Experience participating in on‑call rotations and managing high‑severity production incidents.
Work Environment:
  • Participation in a 24x7 on‑call rotation may be required, including evenings, weekends, and holidays.
  • This role supports a global, always‑on multi‑brand eCommerce platform with executive‑level visibility.
  • The successful candidate must be prepared to respond to high‑severity incidents outside standard working hours.
What We Offer:
  • Competitive salary and benefits package.
  • Professional growth opportunities and continuous learning.
  • A collaborative and innovative work environment.
  • 2 weeks PTO
  • Birthday off
  • Sick leave
  • Study leave
  • Discounts platform
  • Connectivity expenses
About Us:

Nybble Group is a leading digital transformation consulting company and technology solutions provider with over 20 years of experience, dedicated to driving innovation and growth for our clients. Our people‑first approach combines a deep passion for excellence with deep technology expertise to create transformative solutions that redefine customer experiences, systems, and business processes. Fueled by proven capabilities and an agile mindset, we challenge the status quo to solve real‑world problems and deliver impactful digital transformation.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Data Solutions (Artificial Intelligence) Specialist
Data Solutions (Artificial Intelligence) Specialist

ECLARO • Lawrence Township (NJ)

Hybrid
USD 90,000 - 120,000
401k Retirement Savings Plan
Medical, Dental & Vision Insurance
AI & Data Engineer
AI & Data Engineer

Socket.dev • Lawrence Township (NJ)

On-site
USD 140,000 - 190,000
Site Reliability Engineer (SRE) – AEM
Site Reliability Engineer (SRE) – AEM

Highbrow LLC • Atlanta (GA)

On-site
USD 100,000 - 130,000
Site Reliability Engineer
Site Reliability Engineer

Highbrow LLC • Atlanta (GA)

On-site
USD 100,000 - 130,000
Site Reliability Engineer [AQ-18653]
Site Reliability Engineer [AQ-18653]

Aquent • Austin (TX)

On-site
USD 140,000 - 190,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Spectraforce Technologies • Austin (TX)

Hybrid
USD 130,000 - 170,000
Sr. AEM developer
Sr. AEM developer

enableIT • Boston (MA)

On-site
USD 150,000 - 200,000
Health insurance
Health savings account
Vision insurance
+2
Site Reliability Engineer
Site Reliability Engineer

Skill • Austin (TX)

On-site
USD 140,000 - 190,000
Subsidized health plan
Retirement plan with match
Paid sick leave
Sr SRE Automation Engineer
Sr SRE Automation Engineer

Compunnel, Inc. • Austin (TX), Northern (KY)

Hybrid
USD 130,000 - 180,000
Site Reliability Engineer
Site Reliability Engineer

Aquent • Austin (TX)

On-site
USD 140,000 - 175,000
Subsidized health plan
Vision plan
Dental plan
+2