Lead DevOps Engineer

Skit

Hinoba-an

On-site

PHP 1,500,000 - 2,600,000

Full time

38 hours ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Skit.ai in the Philippines is seeking a Lead DevOps Engineer to own our multi-cloud, real-time media platform designed for regulated enterprises. You will shape architecture across AWS, GCP, and Azure, with self-hosted LiveKit and SIP, and GPU-based model serving to meet sub-500ms budgets.

You will drive reliability, observability, and security, mentor SREs, and push IaC-first standards across the team to sustain scale.

Qualifications

  • 6+ years hands-on cloud infrastructure experience.
  • 3+ years operating multiple clouds in production.
  • Deep expertise in at least two of AWS/GCP/Azure.
  • Real-time audio/video systems in production experience.
  • Strong networking, VPC/VNet design and private connectivity.
  • Kubernetes at scale with stateful workloads and service mesh.
  • Terraform-based IaC across multi-account estates.
  • AI/ML serving in production with GPU scheduling.
  • Secure CI/CD with integrated security scanning.
  • Production STT/TTS/LLM API operations knowledge

Responsibilities

  • Own multi-cloud substrate across AWS, GCP, and Azure.
  • Design real-time media plane with autoscaling and LiveKit/SIP.
  • Manage GPU fleets for self-hosted ASR and open-weight LLMs.
  • Establish reliability and observability with tracing and runbooks.
  • Lead cost engineering to optimize per-minute serving costs.
  • Enforce Zero Trust, IAM/RBAC, and auditable controls.
  • Provide technical leadership and enforce IaC standards.

Skills

Cloud infrastructure
Multi-cloud production
Real-time media systems
Networking depth
Kubernetes at scale
Infrastructure as Code
AI/ML model serving
Streaming protocols
Security fundamentals
CI/CD pipelines

Tools

Terraform
Kubernetes (EKS/GKE/AKS)
Istio/Linkerd
LiveKit
WebRTC
Triton/vLLM

Job description

Skit.ai runs autonomous voice agents for regulated enterprises — India's largest banks and telcos, and US collections operations. Every call is a live distributed system: PSTN/SIP → media server → ASR → LLM → TTS → back, spread across three clouds and multiple vendors, with a conversational response budget measured in hundreds of milliseconds.

The platform peaks at roughly **1 million calls per hour**. Billed minutes grew **5,000x+ in eight months**. At this scale, infrastructure is not a support function — latency, cost-per-minute, and auditability are product features. When infra degrades, a customer mid-sentence hears silence.

We're hiring a Lead DevOps Engineer to own this substrate and keep it ahead of the growth curve.

What you'll own:
  • Multi-cloud substrate : Production infrastructure across AWS, GCP, and Azure. Private interconnects (Direct Connect, Cloud Interconnect, ExpressRoute), transit/hub-spoke topologies, and cross-cloud latency managed as an explicit budget — p95 per hop in tens of milliseconds, not "best effort."
  • Real-time media plane: Self-hosted LiveKit and SIP infrastructure at scale. Media servers are stateful; you'll design session-affine, event-driven autoscaling (KEDA-class) that survives traffic tripling within an hour.
  • Model-serving infrastructure : GPU fleets (A100/H100/B200-class) for self-hosted ASR and open-weight LLMs — inference optimization, prefix caching, sticky-session routing, sub-500ms TTFT budgets — alongside managed APIs (Vertex AI/Gemini, Bedrock, Azure). Vendor failover is your design, not your incident.
  • Reliability & observability : OTel-native tracing (Grafana/Tempo stack), per-turn latency attribution across telephony/ASR/LLM/TTS, automated incident response and self-healing. You'll act as incident commander for infrastructure and write the runbooks you'd want at 3 a.m.
  • Cost engineering : Cost-per-minute is an SLO here. We cut per-minute serving cost ~18x in six months through caching, rightsizing, autoscaling, and workload re-architecture — you'll own the next 10x.
  • Security & compliance : Zero Trust across clouds: private endpoints/PrivateLink, IAM/RBAC, secrets management with rotation, WAF/DDoS protection. Operate controls for SOC 2 and ISO/IEC 27001; working command of ISO/IEC 42001:2023 (AI management systems) — hands-on preferred, rigorous theoretical grounding acceptable. You'll face bank and telecom auditors directly, including data-residency requirements.
  • Technical leadership : Terraform-first IaC standards, production-readiness reviews, mentoring SREs. "Lead" means you raise the floor of the whole team.
Problems on our plate right now
  • Migrating LLM inference from managed APIs to self-hosted open-weight models on GPUs without breaking TTFT budgets
  • ASR, LLM, and telephony living in different clouds: interconnect topology that keeps the packet path short and private
  • Multi-region DR that satisfies bank audits without doubling spend

If these read as interesting rather than terrifying, keep reading.

Must-have
  • 6+ years hands-on cloud infrastructure; 3+ years operating multiple clouds simultaneously in production; deep expertise in at least two of AWS/GCP/Azure
  • Real-time audio/video systems in production — WebRTC, SIP/PSTN, or streaming media; you've debugged jitter, not just read about it
  • Networking depth: VPC/VNet design, load balancing, DNS, NAT; private connectivity (Direct Connect / Cloud Interconnect / ExpressRoute, PrivateLink / Private Service Connect); transit gateways and cross-cloud mesh
  • Kubernetes at scale (EKS/GKE/AKS), Helm, and scaling *stateful* workloads; service mesh familiarity (Istio/Linkerd)
  • Infrastructure as Code: Terraform (non-negotiable) across multi-account/multi-project estates; drift is a bug
  • AI/ML serving in production: GPU allocation and scheduling, inference servers or serverless GPU platforms (vLLM / Triton / Modal / Baseten-class), streaming protocols (WebRTC, WebSocket, gRPC)
  • Production STT/TTS/LLM API operations: streaming integrations, quota management, multi-vendor failover (Deepgram / Google / Azure / Whisper-class ASR; ElevenLabs / Azure-class TTS)
  • Security fundamentals: IAM/RBAC, secrets management (Vault or cloud-native), encryption and key rotation
  • CI/CD: GitHub Actions or GitLab CI with security scanning integrated into the pipeline
Strong signal (nice-to-have)
  • LiveKit, pipecat, Twilio, or comparable real-time platforms; SIP trunking and PSTN integration
  • KEDA or other event-driven autoscaling used in anger
  • MLOps: model versioning, canary and A/B rollout
  • FinOps discipline: reserved/spot strategy, unit-economics reporting
  • Certifications: AWS SA Professional, GCP Professional Cloud Architect, Azure Solutions Architect Expert
  • ISO/IEC 42001:2023 exposure
What we're NOT looking for
  • Single-cloud depth with documentation-level knowledge of the other two
  • Tool-checklist DevOps without production AI/ML serving scars
  • "Can learn quickly" as the primary qualification — this role needs day-one production credibility
  • Anyone who has never traced a packet across a cloud boundary
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Devops Engineer
Devops Engineer

Acquire Intelligence • Metro Manila

On-site
PHP 1,200,000 - 1,800,000
Senior Ai Engineer Voice Ai Agentic Systems Danish Mullaji Gurugram
Senior Ai Engineer Voice Ai Agentic Systems Danish Mullaji Gurugram

Vibehackers • Hinoba-an

On-site
PHP 2,309,000 - 4,618,000
Forward Deployed Engineer Hybrid / Full time Ahmedabad, India
Forward Deployed Engineer Hybrid / Full time Ahmedabad, India

Attri Inc. • Hinoba-an

On-site
PHP 1,200,000 - 1,800,000
Software Engineer- AI-Driven SRE & Cloud SRE
Software Engineer- AI-Driven SRE & Cloud SRE

Keka Technologies Private Limited • Mexico

On-site
PHP 2,215,000 - 3,322,000
AI Solutions Architect
AI Solutions Architect

Willis Towers Watson • Philippines

On-site
PHP 1,800,000 - 3,200,000
Equal opportunity employer
AI/ DevOps Engineer
AI/ DevOps Engineer

Universal Access and Systems Solutions Inc. • Angeles

On-site
PHP 600,000 - 1,000,000
Member of Technical Staff, Distributed Systems
Member of Technical Staff, Distributed Systems

logcat.ai • Hinoba-an

On-site
PHP 1,200,000 - 2,400,000
Founding equity
Medical insurance for you and family
AI Developer – Backend & LLM Systems
AI Developer – Backend & LLM Systems

Salvo Software LLC • Mexico

On-site
PHP 1,228,000 - 2,150,000
Full-Stack Engineer — Systems Conductor
Full-Stack Engineer — Systems Conductor

Randstad (Schweiz) AG • Philippines

On-site
PHP 900,000 - 1,500,000
Senior Data Engineer
Senior Data Engineer

FinStrat Management • Manila

On-site
PHP 2,000,000 - 4,500,000
Unlimited vacation
Education stipends
Bonuses