Software Engineer, Infrastructure & Reliability

CrewAI, Inc.

Northern (KY)

Hybrid

USD 150,000 - 190,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

CrewAI, Inc. is seeking an infrastructure platform engineer to own and scale the cloud foundation behind our multi‑agent AI platform.

You will work across AWS, container tech, and CI/CD pipelines to improve reliability, security, and velocity for production deployments. You'll collaborate with runtime and product teams to optimize PostgreSQL/Redis workloads, implement robust on‑call practices, and build self‑hosted install tooling for customers.

Qualifications

  • Strong infra/platform engineering experience in production SaaS.
  • Deep experience with AWS, Docker, CI/CD, and containerized services.
  • Comfort with ECS/Kubernetes; Helm is a plus.
  • Proficient with PostgreSQL, Redis, and background job systems.
  • Ability to write reliable automation in Python, Ruby, Go, or Bash.
  • Security‑minded approach to IAM, secrets, and access control.
  • Calm, rigorous incident management and deployment practices.

Responsibilities

  • Own and improve infrastructure running CrewAI's platform across cloud providers.
  • Build and maintain CI/CD pipelines for builds, tests, image publishing, migrations, and deployments.
  • Improve reliability with health checks, alerting, and recovery planning; participate in on‑call rotation.
  • Collaborate with runtime and product engineers on production behavior and scaling.
  • Manage production observability from logs to dashboards and telemetry export.
  • Harden security posture across IAM, secrets management, and vulnerability handling.
  • Develop tooling and automation for self‑hosted installs and operator tooling.
  • Reduce operational toil through automation of recurring workflows.

Skills

AWS
Docker
Kubernetes
CI/CD
GitHub Actions
PostgreSQL
Redis
Automation scripting
Security IAM
Incident management
Celery
FastAPI
Python
Go
Ruby
Bash

Education

Tools

ECS/ECR
Kubernetes
Helm
PostgreSQL
Redis
Sentry
OpenTelemetry

Job description

About CrewAI

CrewAI is the leading framework and enterprise platform for building and orchestrating multi‑agent AI systems, powering 300M+ agent executions per month across thousands of companies. The Agent Management Platform is our control plane for deploying, monitoring, governing, and scaling agents in production. This role owns the infrastructure foundation that keeps it reliable, secure, and fast.

The Role

You'll build and operate the platform infrastructure behind CrewAI's cloud and enterprise deployments. You'll work across multiple hyperscalers - AWS, Azure, and GCP. You’ll work on containers, CI/CD, deployment automation, observability, secrets, networking, and runtime reliability. Your job is to make the product and runtime teams faster while making customer’s production environments safer.

This is not a pure DevOps support role. You'll write code, improve systems, design deployment paths, harden production, and build the internal platform that lets CrewAI scale and scale our customer deployments.

What You'll Do
  • Own and improve the infrastructure that runs CrewAI's platform: AWS, ECS/ECR, Docker, Kubernetes/Helm, networking, secrets, databases, Redis, and related services.
  • Build and maintain CI/CD pipelines for build, test, image publishing, migrations, environment promotion, rollbacks, and deploy safety.
  • Improve reliability across cloud and enterprise deployments: health checks, alerting, incident response, capacity planning, recovery paths, and operational runbooks - and own the front-line on‑call rotation and its SLAs.
  • Partner with runtime engineers on Celery/FastAPI/Redis workloads and with product engineers on Rails/Solid Queue/Postgres production behavior.
  • Manage production observability and telemetry infrastructure: logs, metrics, traces, dashboards, Sentry/OpenTelemetry plumbing, actionable alerts, and telemetry export to customers' own monitoring systems.
  • Harden security and compliance posture across IAM, workload identity, secrets management, vulnerability scanning, dependency/image hygiene, and least‑privilege access.
  • Build the tooling and automation that lets field engineers and customers run self‑hosted installs themselves - Helm charts, environment config, release artifacts, pre‑flight checks, and install runbooks - so engineering does fewer hands‑on installs over time.
  • Reduce operational toil by automating recurring workflows and making deployments boring.
What We're Looking For
  • Strong infrastructure/platform engineering experience in production SaaS environments.
  • Deep practical experience with AWS, Docker, CI/CD, GitHub Actions, and containerized services.
  • Experience with ECS and/or Kubernetes; Helm experience is a strong plus.
  • Comfort operating PostgreSQL, Redis, background job systems, queues, and web services in production.
  • Strong debugging instincts across app, infra, network, deploy, and dependency layers.
  • Security‑minded approach to IAM, secrets, workload identity, vulnerability management, and production access.
  • Ability to write reliable automation in Python, Ruby, Go, Bash, or similar.
  • Calm, rigorous approach to incidents, rollbacks, migrations, and production change management.
Bonus
  • Experience with AI/agent platforms, workflow runtimes, or high‑volume async execution systems.
  • Experience supporting enterprise/self‑hosted deployments.
  • Terraform or other IaC experience.
  • SRE background: SLOs, incident review, capacity planning, load testing.
  • Familiarity with Rails, FastAPI, Celery, OpenTelemetry, or multi‑service observability.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Software Engineer, Infrastructure & Reliability
Software Engineer, Infrastructure & Reliability

CrewAI • United States

On-site
USD 130,000 - 210,000
Full-Stack Engineer, Agent Management Platform
Full-Stack Engineer, Agent Management Platform

crewAI, Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
Member of Technical Staff – Infrastructure
Member of Technical Staff – Infrastructure

Jobtailor • New York (NY)

On-site
USD 170,000 - 230,000
Senior/Staff AI Infrastructure Engineer
Senior/Staff AI Infrastructure Engineer

Echelon • San Francisco (CA)

On-site
USD 180,000 - 280,000
Health insurance
Dental insurance
Vision insurance
+3
Senior DevOps Engineer – LATAM
Senior DevOps Engineer – LATAM

Luxury Presence • United States

On-site
USD 140,000 - 210,000
Member of Technical Staff, Infrastructure
Member of Technical Staff, Infrastructure

Sycamore • Palo Alto (CA)

On-site
USD 180,000 - 260,000
Infrastructure Engineer
Infrastructure Engineer

Optimized, Inc. • San Francisco (CA)

On-site
USD 170,000 - 230,000
Software Engineer, Open Source
Software Engineer, Open Source

crewAI, Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 140,000 - 210,000
Robotics Infra & AI Automation Engineer
Robotics Infra & AI Automation Engineer

Tutor Intelligence • City of Watertown (NY)

On-site
USD 120,000 - 160,000
Robotics Infrastructure Engineer
Robotics Infrastructure Engineer

Tutor Intelligence • City of Watertown (NY)

On-site
USD 120,000 - 160,000