A technology solutions provider is seeking an experienced professional to serve as a Subject Matter Expert for hybrid cloud applications. Responsibilities include leading incident management, ensuring system performance, and driving operational improvements. Candidates need a Bachelor's degree and at least 5 years of experience in critical systems support, with strong skills in Linux administration and automation tools. This role is crucial for maintaining high-availability production systems in a collaborative DevOps environment.
Qualifications
Minimum 5 years of experience supporting critical, high-availability production systems.
Proven success in cross-functional collaboration within modern DevOps environments.
Experience supporting and maintaining PaaS environments and scalable, resilient architectures.
Responsibilities
Serve as Subject Matter Expert (SME) for distributed applications on hybrid cloud platforms.
Lead incident management and troubleshooting.
Proactively monitor, troubleshoot, and optimize application performance.
Skills
Linux Administration & Troubleshooting
Distributed Applications
Logging & Monitoring
Version Control
Automation using Bash/Python
Education
Bachelor’s degree in Information Technology, Engineering, or a related technical field
Tools
RHEL
CentOS
Ubuntu
Microservices
Splunk
Grafana
Prometheus
Git
GitHub
GitLab
Job description
Responsibilities
Serve as Subject Matter Expert (SME) for distributed applications on hybrid cloud platforms, documenting best practices and providing guidance to peers.
Champion continuous operational improvements informed by metrics analysis and customer feedback.
Lead incident management, troubleshooting, response coordination, and conduct comprehensive post‑incident reviews.
Clearly communicate complex technical issues to development teams, document root causes, and collaborate internally to create robust solutions.
Manage, deploy, and maintain enterprise applications and cloud‑based systems using secure, scalable, and reliable frameworks.
Proactively monitor, troubleshoot, and optimize the health, performance, and reliability of applications and platforms.
Perform detailed log analysis and utilize stack traces to debug and resolve issues reported by partners and end‑users.
Develop comprehensive documentation covering operational procedures, system configurations, and environment setups.
Continuously identify and implement automation opportunities to reduce manual tasks and operational overhead.
Train junior engineers in different subjects of expertise.
Participate in a 24x7 shifting rotation.
Your Qualifications
Bachelor’s degree in Information Technology, Engineering, or a related technical field.
Minimum 5 years of experience supporting critical, high‑availability production systems with a focus on automation, reliability, and operational excellence.
At least 5+ years of hands‑on experience in at least 1–2 tools per domain:
Linux Administration & Troubleshooting: RHEL, CentOS, Ubuntu, or similar Unix-based OS.
Distributed Applications: Microservices architecture and distributed application support.