Company Overview
At Logixgrid, we build and support mission-critical logistics platforms used by global clients. We are looking for a Senior DevOps Engineer who can ensure seamless deployments, high system uptime, and strong SLA adherence across a complex, multi-cloud environment.
This role requires someone who can balance scalable infrastructure management, rigorous release governance, and client-facing reliability expectations while working closely with product, support, and implementation teams.
Key Responsibilities
1. Infrastructure & Multi-Cloud Management
- Manage, scale, and optimize infrastructure across a multi-cloud ecosystem, including Amazon Web Services (AWS) and Google Cloud Platform (GCP).
- Oversee and maintain legacy or external client infrastructures hosted on third-party providers like GoDaddy (managing DNS, VPS, SSL certificates, and domain migrations).
- Ensure high availability, auto-scaling, fault tolerance, and security of our multi-tenant SaaS logistics platforms.
- Handle environment provisioning and isolation for new enterprise client onboarding (Production, Staging, UAT).
2. Release & Deployment Management
- Own the end-to-end release lifecycle for SaaS product versions and client-specific deployments.
- Plan, schedule, and execute blue-green or canary deployments to achieve zero-downtime upgrades.
- Maintain rigorous release checklists, automated rollback strategies, and strict version control.
3. SLA & Production Support
- Ensure strict adherence to defined SaaS SLAs (uptime, response time, and resolution time).
- Monitor production systems globally and proactively address cross-cloud performance degradation.
- Support critical incidents, conduct Root Cause Analysis (RCA), and implement preventive engineering actions.
- Collaborate with the Client Success Team during enterprise escalations.
4. CI/CD & Automation (Infrastructure as Code)
- Design, maintain, and secure robust CI/CD pipelines for continuous, automated delivery.
- Enforce Infrastructure as Code (IaC) to ensure environment consistency across AWS, GCP, and other hosting environments.
- Automate routine operational tasks, patching, and configuration management to minimize manual intervention.
5. Monitoring & Site Reliability Engineering (SRE)
- Implement centralized multi-cloud monitoring, distributed tracing, and aggregated logging systems.
- Track system health, API latency, database performance, and cross-cloud networking metrics.
- Ensure zero/low downtime during peak global logistics operations and high-transaction windows.
Required Skills & Qualifications
Cloud & Infrastructure Expertise
- 4+ years of experience in DevOps, Cloud Architecture, or Site Reliability Engineering (SRE).
- Advanced AWS Architecture Ecosystem: Deep hands-on experience with services required to run a scalable SaaS application:
- Compute & Orchestration: EKS (Elastic Kubernetes Service), ECS (Elastic Container Service), AWS Fargate, and Lambda for serverless scaling.
- Data & Caching: Aurora MySQL/PostgreSQL (with read replicas), DynamoDB, ElastiCache (Redis/Memcached), and Amazon OpenSearch (Elasticsearch).
- Networking & Content Delivery: Route 53, CloudFront (CDN), ALB/NLB (Load Balancers), VPC Peering, and AWS Transit Gateway.
- Security & Governance: AWS IAM, Secrets Manager, KMS (Key Management Service), WAF (Web Application Firewall), and Shield.
- Google Cloud Platform (GCP): Solid experience managing GKE (Google Kubernetes Engine), Compute Engine, Cloud SQL, Cloud DNS, and GCP IAM networks.
- External & Hybrid Cloud Management: Proven experience handling DNS management, registrar configurations, web hosting, and resource migration out of platforms like GoDaddy.
DevOps Tooling & Engineering
- Containerization: Mastery of Docker and production-grade Kubernetes management (ingress controllers, service meshes, auto-scaling policies).
- Infrastructure as Code: Strong proficiency with Terraform for multi-cloud provisioning (AWS & GCP modules).
- CI/CD Tools: Advanced experience with GitHub Actions, Jenkins, or GitLab CI.
- Observability Stack: Hands-on experience with tools like Prometheus, Grafana, Datadog, ELK Stack (Elasticsearch, Logstash, Kibana), or AWS CloudWatch/GCP Operations Suite.
- Systems & Scripting: Solid understanding of Linux/Unix administration and proficiency in Python, Go, or Bash for automation.
Good to Have (Logistics Domain Advantage)
- Experience working with high-transaction logistics, supply chain, or fintech SaaS platforms.
- Exposure to high-throughput systems handling real-time IoT tracking APIs, webhooks, and heavy order volumes.
- Experience in 24x7 production support or follow-the-sun engineering models.
Key Success Metrics (KPIs)
- SLA Adherence: Maintain >99.9% platform uptime and minimize Mean Time to Resolution (MTTR).
- Deployment Velocity: Increase deployment frequency while maintaining a near-zero rollback rate.
- Automation Coverage: Percentage reduction in manual environment provisioning and configuration drift.
- Cloud Cost Optimization: Efficient resource utilization across AWS, GCP, and third-party hostings.
Soft Skills
- Ownership Mindset: High accountability for system uptime and critical infrastructure.
- Calm Under Pressure: Ability to methodically triage and resolve critical production incidents.
- Clear Communication: Ability to bridge technical infrastructure realities with business and client success teams.
Why Join Logixgrid
- Architect and manage a sophisticated, multi-cloud infrastructure powering real-world global supply chains.
- Work with a modern, enterprise-grade cloud-native tech stack (Kubernetes, Serverless, IaC).
- Collaborate in a growth-oriented, engineering-first culture where automation is valued over firefighting.