Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.
LTM is seeking an experienced Site Reliability Engineer/DevOps professional to ensure reliability, scalability, and performance of our software systems through automation, monitoring, and incident management.
You will collaborate with development and operations to design scalable systems, develop automation tools for deployment and incident response, monitor performance, troubleshoot downtime, and participate in on-call rotations to drive root-cause analysis and proactive capacity planning.
Responsible for ensuring the reliability scalability and performance of software systems through automation monitoring and incident management
Collaborate with development and operations teams to design and implement scalable and reliable systems
Develop and maintain automation tools for deployment monitoring and incident response
Monitor system performance and troubleshoot issues to minimize downtime
Implement best practices for system security backup and disaster recovery
Participate in oncall rotations to provide timely incident resolution and root cause analysis
Continuously improve system reliability through proactive maintenance and capacity planning
Disclaimer: This job posting has been aggregated from external source. Role details, content, and availability are subject to change. Applicants are advised to confirm the latest information directly on the company website before applying.