Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.
Neev is seeking a senior Production Engineer to lead operational improvements and incident recovery for payment applications. You will combine deep technical expertise with leadership to maintain highly available production systems and drive automation.
The role emphasizes incident management, cross-team coordination, and continuous improvement in a corporate banking environment. Strong Kubernetes, Linux, and scripting skills are essential.
Within the Production Engineering team, you'll play a senior role supporting business-critical payment applications while leading operational improvements and incident recovery activities.
You'll combine strong technical expertise with leadership, helping maintain highly available production systems while driving automation and operational excellence. Acting as a senior technical contributor, you'll coordinate incident resolution, mentor team members, and work closely with engineering and infrastructure teams to improve service reliability.
Job Description:
6+ years of experience in Production Support, Application Support or Software Engineering, preferably within Corporate Banking or Financial Services.
Strong Unix/Linux skills for production troubleshooting, including navigating the operating system, checking logs and investigating application behaviour.
Ability to write complex SQL queries.
Experience developing scripts in Bash, PowerShell, or Python.
Practical experience supporting and maintaining Kubernetes environments.
Experience supporting Java/J2EE applications.
Experience with Oracle, MQ and WebSphere.
Experience using monitoring and observability tools such as Splunk and AppDynamics.
Strong understanding of IT infrastructure and distributed systems.
Experience leading incident management and coordinating technical teams.
Demonstrated track record of implementing automation and operational improvements.
Practical experience using AI productivity tools (e.g. Copilot, Claude) to support automation, documentation or operational analysis.
Strong communication, stakeholder management, and collaboration skills.
Knowledge of payment systems or payment processing concepts.
Network troubleshooting knowledge.
ITIL certification.
Lead service recovery during production incidents and coordinate technical resolver teams.
Investigate application and infrastructure issues using logs, SQL queries, and Unix/Linux tools.
Respond to user queries regarding application behaviour and service usage.
Communicate incident impact, progress, and recovery plans to stakeholders.
Escalate critical incidents appropriately and drive timely resolution.
Problem management & continuous improvement
Lead or contribute to Root Cause Analysis (RCA) activities.
Identify recurring operational issues and implement permanent improvements.
Develop operational documentation, runbooks and knowledge sharing materials.
Improve monitoring, alerting and operational processes to increase service resilience.
Automation & Operational Excellence
Design and implement automation using Bash, PowerShell, or Python.
Develop internal operational tools that reduce manual effort.
Demonstrate practical use of AI tools (e.g. Copilot, Claude) to automate operational activities and improve team productivity.
Change & Release Support
Participate in change reviews ensuring operational readiness and compliance with CLIENT standards.
Support patching, platform maintenance, and disaster recovery activities.
Identify operational risks and recommend appropriate mitigations.