Infrastructure Engineer (DevOps)

Pratilipi

Bengaluru

On-site

INR 2,500,000 - 4,200,000

Full time

6 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Pratilipi is building a next-generation storytelling platform and is hiring an Infrastructure Engineer to own AWS-based infrastructure, CI/CD pipelines, and AI infra in a high-growth team. The role involves multi-region hosting, cost optimisation, and working with a self-hosted data layer plus AI workloads.

You will collaborate with product and ML teams to ensure scalable, reliable operations, robust observability, and secure, compliant cloud usage across services.

Responsibilities

  • Own and operate AWS infrastructure (ECS on EC2, VPCs, IAM, autoscaling) with Terraform.
  • Build fast, safe Jenkins pipelines and improve deploy experience.
  • Drive cost optimisation across infrastructure.
  • Manage media delivery via Cloudflare CDN and transcoding pipelines.
  • Operate the self-managed data layer (RDS MySQL, MongoDB, Valkey) including migrations.
  • Set up AI production substrate: GPU capacity, inference gateways, observability for LLM agents.
  • Collaborate with AI/ML teams on self-hosted models, capacity planning, and canary rollouts.
  • Define reliability with SLOs, incident response, and blameless postmortems.
  • Create guardrails: IaC checks, deploy gates, and error-reducing paths.

Job description

About the role

Pratilipi is building a generational company in storytelling, and infrastructure is core to that bet. We're looking for an Infrastructure Engineer— a hands-on role embedded in a high-growth product team using AWS, CI/CD, our self-managed data layer, and the operational substrate for the AI workloads now landing in production. You'll help us self-host our own models and take the infrastructure multi-region as we expand globally.

Expect a lot of breadth, with depth in a few things that matter. You won't be a specialist hiding behind a narrow remit — you'll move across AWS, CI/CD, databases, CDN, and AI infra, and go deep where it counts.

What you'll do
  • Own our AWS infrastructure through Terraform — ECS on EC2, VPCs, IAM, autoscaling, rolling deploys. Clickops-free console as the target.
  • Build Jenkins pipelines that are fast, safe, and self-serve. Be in the deploy paths, not just the platform.
  • Drive cost optimisation.
  • Own media delivery at scale — Cloudflare CDN, image/video/audio transcoding, format/device delivery — tuned for latency, cache hit ratio, and egress cost.
  • Operate the self-managed data layer with service teams — RDS MySQL, MongoDB, Valkey, plus managed MSK and ScyllaDB Cloud. Lead migrations end-to-end (e.g. the ongoing Redis → Valkey).
  • Stand up the operational substrate for AI in production — GPU capacity, inference gateways, key rotation, rate limits, cost-per-request visibility. Build observability for LLM agents: traces, token/cost accounting, eval hooks, alerts on silent regressions.
  • Partner with the AI/ML team on self-hosting models — capacity planning, vLLM / TGI serving, canary rollouts, and cost/latency trade-offs vs managed providers.
  • Own reliability across services and infra — SLOs, alerting, incident response, blameless postmortems. Move teams from firefighting to proactive reliability.
  • Once you have the context, add the guardrails — IaC checks, deploy gates, paved paths — that make the right thing the easy thing and quietly remove whole classes of human error.
What we're looking for
  • 4–6 years in DevOps, SRE, or infrastructure — production systems at meaningful scale, with ownership beyond tickets.
  • A problem solver with high agency. You reason from first principles, don't wait to be told, and dig in rather than deflect — whether it's a developer stuck on Terraform or an ML engineer asking for GPUs.
  • Strong AWS hands-on — ECS on EC2, VPCs, IAM, multi-AZ design — and proficiency with Terraform (you write modules others reuse).
  • Jenkins in production plus a real sense of developer experience in CI/CD — you've looked at deploy-time metrics and changed them.
  • Python, Ansible, and shell — you automate work rather than repeat it.
  • Operated databases in production — at least some of RDS MySQL, MongoDB, Redis/Valkey. Done migrations, failovers, and perf tuning, not just provisioning. Working familiarity with Kafka and Cassandra/ScyllaDB.
  • Strong grasp of Linux internals and networking fundamentals (TCP/IP, DNS, TLS, load balancing) and common failure modes.
  • Some hands-on exposure to LLM-based systems — inference endpoints, agent tracing, token/cost, or evals in CI. Not a researcher; you should reason clearly about latency, cost, and failure modes of LLM workloads.
  • Comfortable with cloud security fundamentals: IAM least-privilege, secrets management, network segmentation. Prior exposure to ISO 27001 / DPDP-style controls is a plus.
Tech stack

AWS · ECS (on EC2) · Terraform · Jenkins · Ansible · Python · Cloudflare · Prometheus · Grafana · InfluxDB · RDS MySQL · MongoDB · Valkey · MSK · ScyllaDB Cloud · growing: GPU inference · vLLM / TGI · OpenTelemetry GenAI · Langfuse-style tracing · multi-region AWS

Security & Data Handling

All employees are expected to handle sensitive data responsibly in compliance with the DPDP Act, ISO-27001:2022, and Pratilipi's internal security policies — ensuring data privacy, confidentiality, and NDA obligations at all times.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Devops Engineer
Devops Engineer

Infer • Karnataka

On-site
INR 800,000 - 1,200,000
Devops Engineer
Devops Engineer

Synth (YC S21) • Bengaluru

On-site
INR 1,500,000 - 2,500,000
Platform Engineer - ML Data Pipelines
Platform Engineer - ML Data Pipelines

Humyn Labs • Bengaluru

On-site
INR 4,000,000 - 7,000,000
DevOps Engineer
DevOps Engineer

KnowledgeWorks Global Ltd. • Mumbai

On-site
INR 1,800,000 - 3,000,000
Member of Technical Staff (MTS) - DevOps & Infrastructure
Member of Technical Staff (MTS) - DevOps & Infrastructure

Fi Money • Bengaluru

On-site
INR 1,000,000 - 1,500,000
Site Reliability Engineer
Site Reliability Engineer

United States Digital Space LLC • Karnataka

On-site
INR 900,000 - 1,200,000
Significant equity in a venture-backed company
Opportunity to work with modern tech stack
Backend / DevOps Engineer · Mid
Backend / DevOps Engineer · Mid

Think Right Advisory Services Pvt. Ltd. • Bengaluru Urban

Hybrid
INR 1,800,000 - 2,800,000
Cloud & AI Infrastructure Engineer (AWS)
Cloud & AI Infrastructure Engineer (AWS)

Wisemonk • Hyderabad

On-site
INR 1,800,000 - 3,000,000
AI Infrastructure / DevOps Engineer (AI-Native, Agentic) 5–10 Years
AI Infrastructure / DevOps Engineer (AI-Native, Agentic) 5–10 Years

Sprouts.ai • Chandigarh

Hybrid
INR 1,200,000 - 1,800,000
Ownership of infra decisions
Fast track to leadership roles
Exposure to frontier AI systems
DevOps Engineer
DevOps Engineer

NAVVYASA CONSULTING PRIVATE LIMITED • Gurugram District

On-site
INR 800,000 - 1,200,000