Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.
Client of Salt is seeking a hands-on Senior DevOps Site Reliability Engineer to build and operate the delivery and runtime foundations for a portfolio of modern enterprise applications, workflow platforms and AI-enabled products.
You will work directly with software engineers to create reliable, secure and automated paths from code to production, owning deployment automation, runtime reliability, observability, infrastructure-as-code and operational readiness.
We are seeking a hands-on Senior DevOps Site Reliability Engineer to build and operate the delivery and runtime foundations for a portfolio of modern enterprise applications workflow platforms and AI-enabled products This is not primarily an infrastructure-administration role You will work directly with software engineers to create reliable secure and automated paths from code to production You will own deployment automation runtime reliability observability infrastructure-as-code and operational readiness across applications that integrate with critical enterprise systems
Build and maintain cloud infrastructure using Infrastructure as Code Design secure repeatable environments across development test staging and production Manage containerized workloads and Kubernetes-based deployments where appropriate Define standard application deployment patterns for backend frontend and AI services Implement secure secrets and configuration management Support network identity and connectivity requirements for enterprise integrations
Build automated CI CD pipelines Standardize build test security scanning and deployment processes Automate environment provisioning and configuration Reduce manual deployment steps and production configuration drift Work closely with engineering teams to improve release frequency and reliability
Establish logging metrics tracing and alerting Define service-level indicators and operational thresholds Build dashboards for system health and application performance Implement incident-response and production-support practices Design for graceful degradation retries failover and recovery Lead root-cause analysis of production incidents
Implement least-privilege access and secure deployment patterns Support auditability of infrastructure and production changes Integrate security checks into delivery pipelines Work with security and infrastructure teams to meet enterprise control requirements
Support business continuity and disaster-recovery design Define backup restore and recovery procedures Test operational recovery rather than relying solely on documented plans