Job Overview:
We are looking for an SRE AIOps Engineer to join our growing team and help build AI-powered automation and tools that streamline Site Reliability Engineering workflows. The ideal candidate will combine strong software development and automation skills with hands‑on experience in SRE, DevOps, application support, and AI/LLM technologies.
Key Responsibilities:
- Design and implement AI/ML and LLM-based solutions to automate incident triage and enrichment, generate incident summaries and stakeholder updates, and recommend remediation actions using runbooks and historical data.
- Develop and maintain Python-based services, CLIs, and scripts to automate repetitive SRE tasks, execute runbooks, process logs, metrics, and events, and integrate with monitoring, ticketing, chat, and other external or internal tools.
- Implement and manage Infrastructure as Code solutions using technologies such as Terraform and CloudFormation for SRE tooling and observability.
- Work with SCOUT AI Tools and Atlassian ROVO to advance AI-driven operational capabilities.
- Measure and improve operational performance through KPIs such as MTTR, MTTD, alert noise reduction, and on‑call load.
- Serve as an L2 point of contact for AEM Sites, AEM Assets, Adobe Target, and related integrations.
- Validate, categorize, and prioritize incoming tickets while assessing impact across pages, journeys, brands, and devices.
- Perform application triage by reviewing AEM error logs and author/publish status, checking dispatcher/CDN behavior and cache‑related symptoms, and validating configuration and content changes.
- Provide L3 and engineering teams with detailed incident descriptions, impact analysis, reproduction steps, and relevant log extracts.
- Partner with content authors and marketing teams to diagnose content and publishing issues.
- Monitor health dashboards and alerts using tools such as Splunk, New Relic, AppDynamics, and AEM consoles to proactively identify anomalies.
- Maintain and improve runbooks, knowledge base articles, and standard operating procedures for common incident types.
- Recommend improvements to alerts, dashboards, and thresholds based on recurring issues and operational trends.
Required Qualifications:
- 4–8 years of experience in SRE, DevOps, or related roles, with a strong emphasis on automation and tooling.
- 3–6 years of experience supporting CMS-based platforms, with hands‑on Adobe Experience Manager (AEM) experience strongly preferred.
- Strong Python development skills, including experience building services, CLIs, integrations, and tests.
- Experience working with AI/ML/LLM tools and integrating AI capabilities into operational workflows.
- Experience with AI/LLM APIs and AI agents.
- Proficiency with at least one major cloud platform such as AWS, Azure, or GCP.
- Working knowledge of Infrastructure as Code technologies such as Terraform, CloudFormation, or CDK.
- Advanced experience with monitoring and observability tools such as AppDynamics, Splunk, and New Relic.
- Understanding of AEM Sites, including pages, components, templates, and content hierarchy.
- Familiarity with AEM authoring and publishing flows, workflows, DAM/Assets, and related application functionality.
- Experience diagnosing AEM application and content‑related issues.
- Strong troubleshooting, analytical, and problem‑solving skills.
- Excellent written and verbal English communication skills.
- Ability to explain technical findings clearly to both technical and non‑technical stakeholders.
- Comfortable working with distributed, cross‑functional teams across multiple time zones.
Preferred Qualifications:
- Experience working with SCOUT AI Tools and Atlassian ROVO.
- Experience using AI assistants for log summarization, pattern identification, incident analysis, and incident reporting.
- Experience designing AI‑powered operational tools and automation.
- Experience with Adobe Target and related AEM integrations.
- Experience working with global, always‑on eCommerce platforms or other high‑availability environments.
- Experience participating in on‑call rotations and managing high‑severity production incidents.
Work Environment:
- Participation in a 24x7 on‑call rotation may be required, including evenings, weekends, and holidays.
- This role supports a global, always‑on multi‑brand eCommerce platform with executive‑level visibility.
- The successful candidate must be prepared to respond to high‑severity incidents outside standard working hours.
What We Offer:
- Competitive salary and benefits package.
- Professional growth opportunities and continuous learning.
- A collaborative and innovative work environment.
- 2 weeks PTO
- Birthday off
- Sick leave
- Study leave
- Discounts platform
- Connectivity expenses
About Us:
Nybble Group is a leading digital transformation consulting company and technology solutions provider with over 20 years of experience, dedicated to driving innovation and growth for our clients. Our people‑first approach combines a deep passion for excellence with deep technology expertise to create transformative solutions that redefine customer experiences, systems, and business processes. Fueled by proven capabilities and an agile mindset, we challenge the status quo to solve real‑world problems and deliver impactful digital transformation.