Get more replies from employers
Send a job-specific resume in minutes.
TerraBarn Inc. is seeking a Production Operations Engineer to focus on operations, incident response, and deep troubleshooting of distributed systems. You will work with development and platform teams to maintain reliability and drive improvements.
The role requires strong experience with Linux, Kubernetes, and cloud platforms, plus monitoring, logging, and automation skills. You will participate in a 24x7 rotation and mentor junior engineers.
This role focuses on production operations, incident response, and deep troubleshooting of distributed systems. You will work closely with development and platform teams to maintain system reliability, investigate issues, and improve operational processes.
Support and maintain microservices-based applications across Linux, Kubernetes, and cloud platforms.
Monitor system health and respond to production incidents and service degradations.
Troubleshoot application and platform issues using logs, metrics, and system diagnostics.
Work with development and platform teams to resolve service and infrastructure issues.
Conduct root cause analysis and contribute to operational improvements.
Maintain runbooks, operational documentation, and troubleshooting guides.
Identify opportunities to automate repetitive operational tasks.
Mentor junior engineers and participate in a 24x7 rotational support schedule.
3+ years of experience supporting production applications or distributed systems environments.
Application Support: Strong understanding of HTTP/HTTPS requests, APIs, and TLS/SSL certificates, with experience troubleshooting application connectivity and service issues.
Microservices Systems: Experience supporting microservices architectures, including service deployments, scalability considerations, and reliability of distributed applications.
Platform Expertise: Hands-on experience working with Linux systems, Kubernetes/container platforms and AWS or cloud infrastructure.
Observability & Operations Tools: Experience using monitoring, logging, and metrics tools to investigate incidents and analyze service health and KPIs.
Operational Engineering Skills: Familiarity withGit workflows and basic scripting (Bash, Python, or similar)for operational tasks and automation.