AI Runtime & Platform Engineer for Production Systems

crewAI, Inc.

San Francisco (CA)

On-site

USD 180,000 - 240,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

CrewAI, Inc. in San Francisco is seeking a senior Python backend engineer to build and scale the enterprise runtime layer that turns open-source Crews and Flows into secure, observable production systems.

You will own APIs, workers, and deployment tooling, collaborate with open-source, product, and infra teams, and focus on reliability, security, and observability across distributed services. This role demands practical experience with FastAPI, Redis, Celery, and OpenTelemetry to ship robust,

Qualifications

  • Strong Python backend/platform engineering experience building production services.
  • Experience with FastAPI or similar API frameworks, Celery or other job systems, Redis, Pydantic.
  • Good instincts for distributed systems: retries, idempotency, async execution, status tracking, race conditions, and failure recovery.
  • Comfort with auth and security-sensitive systems: JWTs, webhooks, signatures, secrets, IAM/workload identity.
  • Practical observability experience: tracing, structured logging, metrics, Sentry/OpenTelemetry, and debugging multi-service failures.
  • Strong testing habits and comfort with CI, package/version management, and release discipline.

Responsibilities

  • Build and maintain the Python enterprise runtime around CrewAI: FastAPI services, Celery workers, Redis-backed state, execution APIs, and deployment-facing tools.
  • Extend open-source CrewAI behavior for enterprise environments while preserving compatibility with upstream framework changes.
  • Own production execution flows: crew and flow kickoff, status, retries, cancellation, checkpoint restore and fork, chat/session state, and human-in-the-loop resume paths.
  • Build secure integration surfaces: JWT auth, signed webhooks, token refresh, file handling, secret fetching, and workload identity across AWS, GCP, and Azure.
  • Improve observability across distributed execution: OpenTelemetry traces, structured logs, Sentry, event tracking, and debuggability across API, worker, and platform boundaries.
  • Maintain strong test coverage for async/runtime behavior using pytest, mypy, ruff, mocks/fakes, and e2e deployment harnesses.
  • Partner with the Agent Management Platform team on API contracts, versioning, enterprise client behavior, deployment status, and failure reporting.

Skills

Python
Backend development
Distributed systems
Security & auth
Observability
Testing & CI
CI/CD
Communication

Tools

FastAPI
Celery
Redis
Pydantic
OpenTelemetry
Sentry
pytest

Job description

CrewAI, Inc. in San Francisco is seeking a senior Python backend engineer to build and scale the enterprise runtime layer that turns open-source Crews and Flows into secure, observable production systems.

You will own APIs, workers, and deployment tooling, collaborate with open-source, product, and infra teams, and focus on reliability, security, and observability across distributed services. This role demands practical experience with FastAPI, Redis, Celery, and OpenTelemetry to ship robust,

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Software Engineer, AI Runtime & Platform Services
Software Engineer, AI Runtime & Platform Services

crewAI, Inc. • San Francisco (CA)

On-site
USD 180,000 - 240,000
Software Engineer, Infrastructure & Reliability
Software Engineer, Infrastructure & Reliability

crewAI, Inc. • San Francisco (CA)

On-site
USD 140,000 - 210,000
AI Platform Engineer — Production Agent Systems & Observability
AI Platform Engineer — Production Agent Systems & Observability

build • San Francisco (CA)

On-site
USD 120,000 - 150,000
Ownership of foundational systems
Work on hard engineering problems
Exceptional team collaboration
+1
Senior AI Infrastructure Engineer - Platform & Cloud
Senior AI Infrastructure Engineer - Platform & Cloud

6AM City, LLC • California (MO)

On-site
USD 140,000 - 200,000
Backend Platform Engineer — Secure & Scalable AI Infra
Backend Platform Engineer — Secure & Scalable AI Infra

Recruiting from Scratch • San Francisco (CA)

On-site
USD 140,000 - 210,000
Bi-annual bonuses
Relocation assistance
Housing stipend
+5
Platform Reliability Engineer - Cloud Infra & CI/CD
Platform Reliability Engineer - Cloud Infra & CI/CD

crewAI, Inc. • San Francisco (CA)

On-site
USD 140,000 - 210,000
AI Platform Engineer — Build & Run Production Systems
AI Platform Engineer — Build & Run Production Systems

Long Lake • San Francisco (CA), New York (NY)

On-site
USD 120,000 - 150,000
Senior Engine & Distributed Systems Engineer
Senior Engine & Distributed Systems Engineer

Stack AI, Inc. • New York (NY)

On-site
USD 120,000 - 160,000
Senior Backend Engineer, AI Agents Platform
Senior Backend Engineer, AI Agents Platform

close • United States

Remote
USD 140,000 - 190,000
Senior AI Platform Engineer - Production Automation
Senior AI Platform Engineer - Production Automation

IMO Health • Chicago (IL)

On-site
USD 140,000 - 200,000