Production Support Specialist

Capco

Hyderabad

On-site

INR 900,000 - 1,200,000

Full time

14 hours ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Capco in Hyderabad is seeking an Operations / Production Engineer to keep business-critical services stable, secure and available day to day. You’ll combine strong Linux/Unix fundamentals, practical SQL skills and structured incident management to diagnose issues, restore services quickly and improve long-term reliability.

This role suits someone who enjoys solving real-world production problems, working calmly under pressure and turning recurring incidents into automation, monitoring and

Qualifications

  • Experience in Operations Engineering, Production Support, Site Reliability Engineering, Infrastructure Support or a similar role.
  • Understanding of incident, change and problem-management practices aligned to IT service-management principles.
  • Ability to assess impact, prioritise incidents and work effectively under pressure.
  • Strong written and verbal communication skills.
  • A disciplined approach to documentation, risk management and operational controls.
  • Commitment to security, resilience, service quality and continuous improvement.
  • Working experience on PostgreSQL and Linux platform is a must
  • Knowledge of ITIL practices and service-management tooling.
  • Familiarity with observability platforms, metrics, dashboards, alerting and distributed tracing.
  • Scripting or automation experience using languages such as Bash, Python or PowerShell.
  • Experience with deployment pipelines, version control and infrastructure-as-code.
  • Understanding of resilience testing, disaster recovery and capacity management.
  • Experience working in regulated, financial-services or other highly controlled environments.

Responsibilities

  • Monitor the health, performance and availability of production services.
  • Investigate and resolve incidents, outages, performance degradation and service alerts.
  • Perform structured triage, identify business impact and prioritise response accordingly.
  • Use Linux/Unix tools to diagnose system, process, network, memory and storage issues.
  • Analyse application and system logs to identify root causes and contributing factors.
  • Execute approved operational procedures, recovery activities and service restoration plans.
  • Participate in on-call or out-of-hours support arrangements, where required.
  • Develop and improve monitoring, alerting and service-health checks.
  • Reduce manual effort through scripting, automation and repeatable operational tooling.
  • Identify recurring incidents and deliver preventative improvements.
  • Contribute to capacity, resilience, disaster-recovery and operational-readiness activities.
  • Improve runbooks, standard operating procedures and knowledge articles.
  • Manage incidents from initial report through diagnosis, escalation, resolution and closure.
  • Communicate clearly with stakeholders throughout the incident lifecycle.
  • Escalate to specialist teams and suppliers when required, providing useful evidence and impact details.
  • Support root-cause analysis and post-incident reviews.
  • Track corrective and preventative actions through to completion.
  • Maintain accurate incident, change and problem records.
  • Excellent knowledge of SQL queries, joins, aggregation etc.
  • Able to identify non-performing SQL and optmise it in coordination with development team.

Skills

Operations Engineering
Production Support
Site Reliability Engineering
Infrastructure Support
IT service-management
Incident management
Communication skills
Documentation
Security
Resilience
PostgreSQL
Linux
ITIL practices
Observability
Bash scripting
Python scripting
PowerShell
Deployment pipelines
Version control
Infrastructure-as-code
Disaster recovery
Capacity management
Regulated environments

Tools

PostgreSQL
Linux

Job description

About the role

We’re looking for an Operations / Production Engineer to keep business-critical services stable, secure and available day to day. You’ll combine strong Linux/Unix fundamentals, practical SQL skills and structured incident management to diagnose issues, restore services quickly and improve long-term reliability.

This role suits someone who enjoys solving real-world production problems, working calmly under pressure and turning recurring incidents into automation, monitoring and preventative improvements.

Key responsibilities
Production support and reliability
  • Monitor the health, performance and availability of production services.
  • Investigate and resolve incidents, outages, performance degradation and service alerts.
  • Perform structured triage, identify business impact and prioritise response accordingly.
  • Use Linux/Unix tools to diagnose system, process, network, memory and storage issues.
  • Analyse application and system logs to identify root causes and contributing factors.
  • Execute approved operational procedures, recovery activities and service restoration plans.
  • Participate in on-call or out-of-hours support arrangements, where required.
Monitoring and automation
  • Develop and improve monitoring, alerting and service-health checks.
  • Reduce manual effort through scripting, automation and repeatable operational tooling.
  • Identify recurring incidents and deliver preventative improvements.
  • Contribute to capacity, resilience, disaster-recovery and operational-readiness activities.
  • Improve runbooks, standard operating procedures and knowledge articles.
Incident and problem management
  • Manage incidents from initial report through diagnosis, escalation, resolution and closure.
  • Communicate clearly with stakeholders throughout the incident lifecycle.
  • Escalate to specialist teams and suppliers when required, providing useful evidence and impact details.
  • Support root-cause analysis and post-incident reviews.
  • Track corrective and preventative actions through to completion.
  • Maintain accurate incident, change and problem records.
Database and data investigation
  • Excellent knowledge of SQL queries, joins, aggregation etc.
  • Able to identify non-performing SQL and optmise it in coordination with development team
Essential skills and experience
  • Experience in Operations Engineering, Production Support, Site Reliability Engineering, Infrastructure Support or a similar role.
  • Understanding of incident, change and problem-management practices aligned to IT service-management principles.
  • Ability to assess impact, prioritise incidents and work effectively under pressure.
  • Strong written and verbal communication skills.
  • A disciplined approach to documentation, risk management and operational controls.
  • Commitment to security, resilience, service quality and continuous improvement.
  • Working experience on PostgreSQL and Linux platform is a must
  • Knowledge of ITIL practices and service-management tooling.
  • Familiarity with observability platforms, metrics, dashboards, alerting and distributed tracing.
  • Scripting or automation experience using languages such as Bash, Python or PowerShell.
  • Experience with deployment pipelines, version control and infrastructure-as-code.
  • Understanding of resilience testing, disaster recovery and capacity management.
  • Experience working in regulated, financial-services or other highly controlled environments.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Production Support Engineer
Production Support Engineer

Moofwd • Pune District

On-site
INR 600,000 - 1,000,000
Production Support / Application Support
Production Support / Application Support

Cloudxtreme • Pune District

On-site
INR 1,200,000 - 1,800,000
Analyst II, Production Support
Analyst II, Production Support

fis • Pune District

On-site
INR 1,500,000 - 2,300,000
Production Monitoring Engineer
Production Monitoring Engineer

Infilect Technologies Pvt. Ltd. • India

On-site
INR 600,000 - 900,000
Product Support Engineer
Product Support Engineer

Datavail Corp. • Mumbai

On-site
INR 2,800,000 - 6,000,000
Production Support Engineer - Distributed Systems
Production Support Engineer - Distributed Systems

Hiringeye Solutions • Hyderabad

On-site
INR 900,000 - 1,300,000
VS01700 - SRE & Production Reliability Engineer
VS01700 - SRE & Production Reliability Engineer

E4 Software Services Pvt Ltd. • Bengaluru

On-site
INR 1,800,000 - 2,400,000
VS01700 - SRE & Production Reliability Engineer
VS01700 - SRE & Production Reliability Engineer

E4 Software Services Pvt Ltd. • India

On-site
INR 2,000,000 - 4,000,000
Software Integration Engineer II – Production Support Developer
Software Integration Engineer II – Production Support Developer

Arthrex India • Maharashtra

On-site
INR 1,200,000 - 1,800,000
Software Engineer
Software Engineer

NatWest Group • Gurugram District

On-site
INR 1,200,000 - 2,400,000