Production Monitoring Engineer

Infilect Technologies Pvt. Ltd.

India

On-site

INR 600,000 - 900,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Infilect Technologies Pvt. Ltd. is seeking a Technical Support / Production Monitoring Engineer to monitor production systems, troubleshoot incidents, and ensure availability of applications and services.

This hands-on role requires diagnosing issues, resolving problems where possible, communicating with customers via email and calls, and coordinating with Engineering/DevOps for deeper fixes. The position involves rotating shifts to provide continuous monitoring and support for production

Qualifications

  • 1–4 years of experience in Technical Support, Application Support, Production Support, IT Operations, NOC or a similar technical role.
  • Good understanding of application and system troubleshooting.
  • Experience working with production environments.
  • Ability to read and interpret application/server logs.
  • Basic SQL knowledge.
  • Understanding of REST APIs and HTTP status codes.
  • Basic understanding of Linux/Unix systems and command-line troubleshooting.
  • Familiarity with monitoring and alerting tools.
  • Ability to troubleshoot issues systematically rather than relying only on predefined solutions.
  • Good understanding of incident management and escalation processes.
  • Strong documentation and communication skills.
  • Willingness to work in rotational shifts.

Responsibilities

  • Monitor production applications, services, APIs, servers, databases, integrations and scheduled jobs.
  • Monitor system health, availability, performance and error alerts.
  • Identify system failures, abnormal behaviour and potential production incidents.
  • Respond promptly to alerts and incidents based on defined severity and response timelines.
  • Ensure critical production issues are tracked until resolution.
  • Investigate application and system issues reported through monitoring systems, customers or internal teams.
  • Analyse application logs, server logs and error messages to identify the cause of incidents.
  • Troubleshoot API failures, integration issues, data issues, service failures and configuration problems.
  • Perform basic service restarts, job reruns, configuration changes and other approved operational fixes.
  • Resolve recurring operational issues and identify opportunities for automation.
  • Follow established runbooks and troubleshooting procedures.
  • Provide Level 1/Level 2 technical support for production systems.
  • Investigate application errors and determine whether an issue is related to application code, infrastructure, configuration, data or external dependencies.
  • Perform basic SQL/database checks and queries where required.
  • Troubleshoot API requests/responses and integration failures.
  • Validate fixes after deployment and confirm that systems are functioning normally.
  • Support production deployments and post‑deployment monitoring where required.
  • Identify genuine application defects during production monitoring and troubleshooting.
  • Create detailed technical tickets for Engineering with: Issue description, Impact, Time of occurrence, Error/log details, Steps to reproduce, Screenshots, Initial troubleshooting performed.
  • Coordinate with Engineering/DevOps teams for resolution of issues requiring code or infrastructure changes.
  • Track incidents through to closure and validate the final fix.
  • Classify incidents based on severity and business impact.
  • Follow defined escalation procedures for critical incidents.
  • Communicate clearly with relevant stakeholders during major incidents.
  • Maintain incident records and resolution timelines.
  • Participate in post‑incident reviews and contribute to Root Cause Analysis (RCA).
  • Identify recurring issues and recommend preventive measures.
  • Work in rotational shifts as required, including weekends and/or night shifts.
  • Maintain a clear shift log covering: Open incidents, Resolved incidents, Pending issues, Alerts requiring monitoring, Escalations, Planned maintenance or deployments.
  • Ensure no critical issue is missed during shift transitions.
  • The person in this role will be successful if they can: Respond to alerts within the required SLA; Resolve routine production issues independently; Reduce unnecessary escalations to Engineering; Provide Engineering with high‑quality technical information when escalation is required; Ensure critical incidents are properly tracked and communicated; Identify recurring problems and help prevent them from happening again; Keep production systems stable and operational.

Skills

Technical Support
Production Monitoring
Incident management
Logging & debugging
SQL basics
REST APIs
Linux/Unix
Monitoring tools
Documentation
Rotational shifts

Tools

Docker/Kubernetes
Grafana
Prometheus
Datadog
ELK/Kibana

Job description

Department: [Technology / Operations]

Reporting To: [Engineering Manager / Tech Lead / Operations Manager]

About the Role

We are looking for a Technical Support / Production Monitoring Engineer who will be responsible for monitoring our production systems, identifying and troubleshooting technical issues, resolving operational incidents, and ensuring the availability and smooth functioning of our applications and services.

This is a hands‑on technical role that requires the ability to monitor systems, investigate incidents, identify root causes, resolve issues wherever possible, communicate with customers over emails and calls, and coordinate with Engineering/DevOps teams for issues requiring deeper technical intervention.

The role will involve working in rotational shifts to provide continuous monitoring and support for production systems.

Key Responsibilities
1. Production Monitoring
  • Monitor production applications, services, APIs, servers, databases, integrations and scheduled jobs.
  • Monitor system health, availability, performance and error alerts.
  • Identify system failures, abnormal behaviour and potential production incidents.
  • Respond promptly to alerts and incidents based on defined severity and response timelines.
  • Ensure critical production issues are tracked until resolution.
2. Incident Troubleshooting & Resolution
  • Investigate application and system issues reported through monitoring systems, customers or internal teams.
  • Analyse application logs, server logs and error messages to identify the cause of incidents.
  • Troubleshoot API failures, integration issues, data issues, service failures and configuration problems.
  • Perform basic service restarts, job reruns, configuration changes and other approved operational fixes.
  • Resolve recurring operational issues and identify opportunities for automation.
  • Follow established runbooks and troubleshooting procedures.
3. Application & Technical Support
  • Provide Level 1/Level 2 technical support for production systems.
  • Investigate application errors and determine whether an issue is related to application code, infrastructure, configuration, data or external dependencies.
  • Perform basic SQL/database checks and queries where required.
  • Troubleshoot API requests/responses and integration failures.
  • Validate fixes after deployment and confirm that systems are functioning normally.
  • Support production deployments and post‑deployment monitoring where required.
4. Bug Identification & Engineering Coordination
  • Identify genuine application defects during production monitoring and troubleshooting.
  • Create detailed technical tickets for Engineering with:
    • Issue description
    • Impact
    • Time of occurrence
    • Error/log details
    • Steps to reproduce
    • Screenshots or relevant evidence
    • Initial troubleshooting performed
  • Coordinate with Engineering/DevOps teams for resolution of issues requiring code or infrastructure changes.
  • Track incidents through to closure and validate the final fix.
  • Classify incidents based on severity and business impact.
  • Follow defined escalation procedures for critical incidents.
  • Communicate clearly with relevant stakeholders during major incidents.
  • Maintain incident records and resolution timelines.
  • Participate in post‑incident reviews and contribute to Root Cause Analysis (RCA).
  • Identify recurring issues and recommend preventive measures.
6. Shift Operations
  • Work in rotational shifts as required, including weekends and/or night shifts.
  • Maintain a clear shift log covering:
    • Open incidents
    • Resolved incidents
    • Pending issues
    • Alerts requiring monitoring
    • Escalations
    • Planned maintenance or deployments
  • Ensure no critical issue is missed during shift transitions.
Required Skills
  • 1–4 years of experience in Technical Support, Application Support, Production Support, IT Operations, NOC or a similar technical role .
  • Good understanding of application and system troubleshooting.
  • Experience working with production environments.
  • Ability to read and interpret application/server logs.
  • Basic SQL knowledge.
  • Understanding of REST APIs and HTTP status codes.
  • Basic understanding of Linux/Unix systems and command‑line troubleshooting.
  • Familiarity with monitoring and alerting tools.
  • Ability to troubleshoot issues systematically rather than relying only on predefined solutions.
  • Good understanding of incident management and escalation processes.
  • Strong documentation and communication skills.
  • Willingness to work in rotational shifts.
Good to Have
  • Prior experience of working with customers on IT management
  • Experience with cloud platforms such as AWS, Azure or GCP.
  • Experience with tools such as Grafana, Prometheus, Datadog, New Relic, ELK/Kibana or similar monitoring platforms.
  • Basic knowledge of Docker/Kubernetes.
  • Experience with Git.
  • Basic scripting knowledge using Python, Shell or similar.
  • Experience with Jira or other incident/ticket management systems.
  • Experience supporting SaaS or B2B technology products.
  • Experience working with AI/ML platforms or data‑intensive applications.
Key Attributes
  • Calm and methodical during production incidents.
  • Comfortable working independently during shifts.
  • Takes ownership of incidents until closure.
  • Able to distinguish between an issue that can be fixed independently and one that requires escalation.
  • Strong communication and handover discipline.
  • Willingness to learn the product and underlying technology deeply.
What Success Looks Like

The person in this role will be successful if they can:

  • Respond to alerts within the required SLA.
  • Resolve routine production issues independently.
  • Reduce unnecessary escalations to Engineering.
  • Provide Engineering with high‑quality technical information when escalation is required.
  • Ensure critical incidents are properly tracked and communicated.Identify recurring problems and help prevent them from happening again.
  • Keep production systems stable and operational.
Suggested KPIs
  • Production monitoring coverage
  • Alert response time
  • Mean Time to Acknowledge (MTTA)
  • Mean Time to Resolve (MTTR)
  • Percentage of incidents resolved without Engineering escalation
  • Number of recurring incidents
  • SLA adherence
  • Quality of incident documentation
  • Shift handover accuracy
  • Number of preventive fixes/improvements implemented
Reporting Structure

Reports To: [Engineering Manager / Tech Lead / Head of Engineering]

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Production Monitoring Engineer
Production Monitoring Engineer

Infilect • Bengaluru

On-site
INR 300,000 - 800,000
Tech Support Engineer
Tech Support Engineer

Infilect Technologies Pvt. Ltd. • Bengaluru

On-site
INR 600,000 - 800,000
Product Support Engineer
Product Support Engineer

Datavail Corp. • Mumbai

On-site
INR 2,800,000 - 6,000,000
Production Support Engineer
Production Support Engineer

Moofwd • Pune District

On-site
INR 600,000 - 1,000,000
Production and Support Engineer
Production and Support Engineer

Fulcrum Worldwide Software • Pune District

Hybrid
INR 1,200,000 - 1,800,000
Production Support / Application Support
Production Support / Application Support

Cloudxtreme • Pune District

On-site
INR 1,200,000 - 1,800,000
Application /Production Support
Application /Production Support

DMart • Coimbatore District

On-site
INR 1,800,000 - 2,400,000
Support Analyst
Support Analyst

Movate Technologies • Chennai District, Bengaluru

On-site
INR 900,000 - 1,500,000
Production Support Analyst
Production Support Analyst

Movate Technologies • Chennai District, Bengaluru

Hybrid
INR 900,000 - 1,300,000
Application Support Engineer
Application Support Engineer

Isoft Services Noida • Hyderabad

On-site
INR 600,000 - 1,000,000