Get more replies from employers
Send a job-specific resume in minutes.
Deploy reliable AI infrastructure for enterprise production
full-time
Highly competitive
Own the infrastructure that runs AI systems in production. You'll design deployment architectures, set up monitoring and alerting, ensure security compliance, and keep everything running smoothly at enterprise scale. This isn't ticket-driven ops work. You're building the platform that enables rapid, reliable delivery across multiple client environments. Expect to write code, design systems, and debug complex distributed failures.
Python AWS
Morning: Review overnight alerts (none, because you built good alerts). Mid-morning: Design review for multi-region deployment. Afternoon: Implement automated failover for critical service. Evening: Write runbook for new deployment pattern.
Build infrastructure for systems that millions depend on. Make architectural decisions that matter. Work with modern tools and patterns. No legacy baggage to maintain.
This role requires completing a technical challenge as part of the application process. Challenge: Medium: High-Availability Service
Rotating on-call (one week per month). Incidents are rare because we invest in reliability. Compensation for on-call time and incident response.