Devops Engineer

Prodapt Solutions Private Limited

Irving (TX)

On-site

USD 120,000 - 170,000

Full time

3 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Prodapt Solutions Private Limited in Irving, TX seeks an experienced Site Reliability Engineer to strengthen production support for mission-critical applications. You will improve monitoring, incident management, and platform reliability across a complex enterprise ecosystem.

You will work with Dynatrace, ELK, Grafana and AI-enabled automation to reduce outages, optimize performance, and drive proactive remediation in a fast-paced telecom/tech environment. U.S. relocation not required.

Qualifications

  • Overall experience: 5+ years in Production Support for mission-critical, high-performance applications.
  • Bachelor’s degree in Computer Science, IT, Engineering or related field.
  • Experience with Docker, Kubernetes and Microsoft Azure Cloud, Unix, Networking and troubleshooting knowledge.
  • Experience with Dynatrace, Elastic, Kibana, Grafana and ELK stack for monitoring and observability.
  • Creation of dashboards on Dynatrace, ELK and Grafana; debugging Java/microservices logs.
  • Experience with relational and NoSQL databases (Oracle, Cassandra).
  • Strong incident, problem, root-cause analysis, and service restoration skills.
  • Experience working across multiple technical teams in large-scale enterprise environments.
  • Excellent written and verbal communication.

Responsibilities

  • Serve as a technical responder during production incidents and service disruptions.
  • Coordinate issue resolution across multiple technical teams and stakeholders.
  • Manage incident lifecycle: detection, triage, resolution, communication, and RCA.
  • Develop and maintain operational runbooks, knowledge articles, and support documentation.
  • Build and optimize monitoring dashboards and health checks across platforms.
  • Leverage AI/GenAI and automation to reduce manual operational effort.

Skills

Strong communication
Incident Management
Problem Management
Root Cause Analysis
Cross-team collaboration

Education

Bachelor’s degree in Computer Science, Information Technology, Engineering, or a related field

Tools

Docker
Kubernetes
Microsoft Azure
Unix
Networking
Dynatrace
ELK Stack
Kibana
Grafana
Python

Job description

Overview

Prodapt is the largest and fastest-growing specialized player in the Connectedness industry, recognized by Gartner as a Large, Telecom-Native, Regional IT Service Provider across North America, Europe and Latin America. With its singular focus on the domain, Prodapt has built deep expertise in the most transformative technologies that connect our world. Prodapt is a trusted partner for enterprises across all layers of the Connectedness vertical. Prodapt designs, configures, and operates solutions across their digital landscape, network infrastructure, and business operations – and craft experiences that delight their customers. Today, Prodapt’s clients connect 1.1 billion people and 5.4 billion devices, and are among the largest telecom, media, and internet firms in the world. Prodapt works with Google, Amazon, Verizon, Vodafone, Liberty Global, Liberty Latin America, Claro, Lumen, Windstream, Rogers, Telus, KPN, Virgin Media, British Telecom, Deutsche Telekom, Adtran, Samsung, and many more. A“Great Place To Work®Certified™” company, Prodapt employs over 6,000 technology and domain experts in 30+ countries across North America, Latin America, Europe, Africa, and Asia. Prodapt is part of the 130‑year-old business conglomerate The Jhaver Group, which employs over 30,000 people across 80+ locations globally.

Unlike a traditional production support role, this position requires strong engineering aptitude and a proactive operational mindset. The ideal candidate combines customer/agent facing application ecosystem knowledge, application monitoring expertise, incident management experience, automation skills, and a passion for leveraging AI and GenAI technologies to improve operational efficiency.

You will partner closely with engineering teams, release management, security, compliance, and business stakeholders to identify risks, reduce incidents, enhance monitoring capabilities, and drive platform reliability across a complex ecosystem supporting hundreds of enterprise applications.

Responsibilities
Site Reliability & Operations
  • Serve as a key member of the Engineering Operations organization supporting self-assist Web & Mobile App and related business applications.
  • Provide Tier 1/2 operational support for production systems in a 24x7 environment.
  • Monitor application health, performance, availability, and customer experience across the platform.
  • Drive proactive issue detection and prevention rather than relying solely on customer-reported incidents.
  • Participate in incident response, triage, war rooms, major incident management, and post-incident reviews.
  • Perform root cause analysis (RCA) and identify opportunities to improve platform stability and resiliency.
  • Partner with Tier 1, Tier 2, Tier 3, infrastructure, security, and application teams to rapidly resolve issues.
  • Create and maintain operational runbooks, knowledge articles, and support documentation.
Observability & Monitoring
  • Build, maintain, and optimize monitoring dashboards, alerts, and health checks.
  • Analyze application logs, API activity, transactions, and performance metrics.
  • Utilize observability and monitoring platforms including:
    • Dynatrace
    • ELK Stack (Elasticsearch, Logstash, Kibana)
    • Catchpoint or other synthetic monitoring solutions
    • Quantum Metrics or other user session replay solutions
  • Reduce alert fatigue through automation, threshold tuning, and intelligent event correlation.
  • Develop and enhance monitoring strategies to provide end-to-end visibility across Digital ecosystem and integrated systems.
Incident Management & Problem Management
  • Act as a technical responder during production incidents and service disruptions.
  • Coordinate issue resolution efforts across multiple technical teams and stakeholders.
  • Manage incident lifecycle activities including:
    • Detection
    • Triage
    • Resolution
    • Communication
    • Root cause analysis
  • Identify recurring issues and lead problem management initiatives to eliminate operational inefficiencies.
Automation & AI Enablement
    • Develop innovative approaches to reduce manual operational effort through automation.
    • Leverage AI, GenAI, agentic workflows, and intelligent operational tooling where appropriate.
    • Create automation solutions to improve:
      • Shift handoffs
      • Incident reporting
      • Alert management
      • Knowledge management
      • Operational reporting
    • Contribute to internal AI initiatives that improve engineering productivity and service reliability.
    • Evaluate and implement automation opportunities across monitoring, ticketing, collaboration, and support workflows.
Requirements
Required Qualifications
  • Overall experience: 5+ years experience performing Production Support for Mission Critical, high-performance applications (Customer Care, Retail and eCommerce customer/agent facing application experience preferred)
  • Bachelor’s degreein Computer Science, Information Technology, Engineering, or a related field
  • Experience using Docker, Kubernetes and Microsoft Azure Cloud, Unix, Networking and troubleshooting knowledge
  • Experience with enterprise monitoring and observability tools such as:
    • Application & Infrastructure Performance Monitoring tools like Dynatrace
    • Application Log Analytics tools like Elastic
    • Visualization tools like Kibana and Grafana. EFK stack experience preferred
  • Creation of Dashboards on Dynatrace, ELK and Grafana
  • Debugging java log, debugging microservices log
  • Experience in Relational & NoSQL databases like Oracle & Cassandra
  • Strong experience with:
    • Incident Management
    • Problem Management
    • Root Cause Analysis
    • Service restoration processes
  • Experience working across multiple technical teams in large-scale enterprise environments.
  • Strong written and verbal communication skills.
Preferred Qualifications
  • Site Reliability Engineering (SRE) experience.
  • Experience building operational dashboards and observability platforms.
  • Hands‑on experience with AI, GenAI, Copilot, & operational automation.
  • Experience with scripting or automation using Python, PowerShell, JavaScript, or similar technologies.
  • Exposure to release management, change management, and production deployments.
  • Understanding of security, SOX, compliance, and enterprise operational controls.
  • Experience supporting large-scale environments with hundreds of integrated applications.
  • Generative AI and Workflow Automation skills
    • GPT-4o LLM – Advanced
    • LangGraph and LangChain
    • Google DialogFlow(Api.ai)
    • Google Vertex
    • Databricks, Spark, Snowflake
    • Python, Java, SQL
    • Selenium, Playwright Robotic Process Automation (RPA)
    • Power Automate, Automation Anywhere
  • Demonstrated experience leveraging AI-driven tools for automating end-to-end operational workflows
  • Demonstrated experience using Text Generative & Code Generative AI Models
  • Automation, Gen AI & Agentic Workflow Technical Skills to include
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

AI Full Stack Engineer
AI Full Stack Engineer

Prodapt • Dallas (TX)

On-site
USD 120,000 - 180,000
.net Developer [AQ-20989]
.net Developer [AQ-20989]

Aquent • Austin (TX)

On-site
USD 120,000 - 180,000
Application Engineer View role →
Application Engineer View role →

Nrnptech • Northern (KY)

On-site
USD 90,000 - 140,000
Sr. Systems Engineer - AI
Sr. Systems Engineer - AI

Dairy Farmers of America • Kansas City (KS)

On-site
USD 130,000 - 170,000
Sr. Systems Engineer - AI
Sr. Systems Engineer - AI

Kansas Ag Connection • Kansas City (KS)

On-site
USD 140,000 - 190,000
Data Scientist
Data Scientist

Naveera IT Consulting Pvt Ltd • United States

Hybrid
USD 16,000 - 22,000
Staff Software Engineer, AI Systems
Staff Software Engineer, AI Systems

Dolby • Atlanta (GA)

On-site
USD 150,000 - 210,000
Senior IT Reliability & Automation Lead
Senior IT Reliability & Automation Lead

First Horizon Bank • Memphis (TN)

On-site
USD 120,000 - 180,000
Lead Software Engineer
Lead Software Engineer

The Sherwin-Williams Company • Cleveland (OH)

On-site
USD 120,000 - 150,000
Technology Support Lead - AI and Production Management Tools
Technology Support Lead - AI and Production Management Tools

JPMorgan Chase & Co. • Columbus (OH)

On-site
USD 180,000 - 240,000