Software Engineer, AI Runtime & Platform Services

crewAI, Inc.

San Francisco (CA)

On-site

USD 180,000 - 240,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

CrewAI, Inc. in San Francisco is seeking a senior Python backend engineer to build and scale the enterprise runtime layer that turns open-source Crews and Flows into secure, observable production systems.

You will own APIs, workers, and deployment tooling, collaborate with open-source, product, and infra teams, and focus on reliability, security, and observability across distributed services. This role demands practical experience with FastAPI, Redis, Celery, and OpenTelemetry to ship robust,

Qualifications

  • Strong Python backend/platform engineering experience building production services.
  • Experience with FastAPI or similar API frameworks, Celery or other job systems, Redis, Pydantic.
  • Good instincts for distributed systems: retries, idempotency, async execution, status tracking, race conditions, and failure recovery.
  • Comfort with auth and security-sensitive systems: JWTs, webhooks, signatures, secrets, IAM/workload identity.
  • Practical observability experience: tracing, structured logging, metrics, Sentry/OpenTelemetry, and debugging multi-service failures.
  • Strong testing habits and comfort with CI, package/version management, and release discipline.

Responsibilities

  • Build and maintain the Python enterprise runtime around CrewAI: FastAPI services, Celery workers, Redis-backed state, execution APIs, and deployment-facing tools.
  • Extend open-source CrewAI behavior for enterprise environments while preserving compatibility with upstream framework changes.
  • Own production execution flows: crew and flow kickoff, status, retries, cancellation, checkpoint restore and fork, chat/session state, and human-in-the-loop resume paths.
  • Build secure integration surfaces: JWT auth, signed webhooks, token refresh, file handling, secret fetching, and workload identity across AWS, GCP, and Azure.
  • Improve observability across distributed execution: OpenTelemetry traces, structured logs, Sentry, event tracking, and debuggability across API, worker, and platform boundaries.
  • Maintain strong test coverage for async/runtime behavior using pytest, mypy, ruff, mocks/fakes, and e2e deployment harnesses.
  • Partner with the Agent Management Platform team on API contracts, versioning, enterprise client behavior, deployment status, and failure reporting.

Skills

Python
Backend development
Distributed systems
Security & auth
Observability
Testing & CI
CI/CD
Communication

Tools

FastAPI
Celery
Redis
Pydantic
OpenTelemetry
Sentry
pytest

Job description

About CrewAI

CrewAI is the leading framework and enterprise platform for building and orchestrating multi-agent AI systems, powering 300M+ agent executions per month across thousands of companies. The Agent Management Platform is our control plane for deploying, monitoring, governing, and scaling agents in production.

The Role

You’ll work on the enterprise runtime layer that turns CrewAI’s open-source Crews and Flows into secure, observable, remotely executable production systems. This is the layer between the framework and the platform: APIs, workers, checkpoints, webhooks, auth, deployment behavior, telemetry, and enterprise extensions that make CrewAI run reliably in real customer environments.

You’ll partner closely with the open-source, product, and infrastructure teams, but your center of gravity is production execution: making agent workflows resumable, inspectable, authenticated, observable, and safe to operate at scale.

What You’ll Do
  • Build and maintain the Python enterprise runtime around CrewAI: FastAPI services, Celery workers, Redis-backed state, execution APIs, and deployment-facing tools.
  • Extend open-source CrewAI behavior for enterprise environments while preserving compatibility with upstream framework changes.
  • Own production execution flows: crew and flow kickoff, status, retries, cancellation, checkpoint restore and fork, chat/session state, and human-in-the-loop resume paths.
  • Build secure integration surfaces: JWT auth, signed webhooks, token refresh, file handling, secret fetching, and workload identity across AWS, GCP, and Azure.
  • Improve observability across distributed execution: OpenTelemetry traces, structured logs, Sentry, event tracking, and debuggability across API, worker, and platform boundaries.
  • Maintain strong test coverage for async/runtime behavior using pytest, mypy, ruff, mocks/fakes, and e2e deployment harnesses.
  • Partner with the Agent Management Platform team on API contracts, versioning, enterprise client behavior, deployment status, and failure reporting.
What We’re Looking For
  • Strong Python backend/platform engineering experience, especially building production services rather than only libraries.
  • Experience with FastAPI or similar API frameworks, Celery or other job systems, Redis, Pydantic, and typed Python.
  • Good instincts for distributed systems: retries, idempotency, async execution, status tracking, race conditions, and failure recovery.
  • Comfort with auth and security-sensitive systems: JWTs, webhooks, signatures, secrets, IAM/workload identity, and least-privilege thinking.
  • Practical observability experience: tracing, structured logging, metrics, Sentry/OpenTelemetry, and debugging multi-service failures.
  • Ability to work at the boundary between an open-source framework and a hosted enterprise platform without creating brittle coupling.
  • Strong testing habits and comfort with CI, package/version management, and release discipline.
Bonus
  • Experience operating AI/agent runtimes, workflow engines, or distributed task systems.
  • Cloud platform experience with AWS ECS/ECR, Kubernetes, Helm, GCP/Azure identity, or secret managers.
  • Experience with enterprise SaaS constraints: auditability, tenant isolation, customer environments, deployment rollbacks, and supportability.
  • Familiarity with Rails/SaaS platforms is useful, but not required.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Software Engineer, Infrastructure & Reliability
Software Engineer, Infrastructure & Reliability

crewAI, Inc. • San Francisco (CA)

On-site
USD 140,000 - 210,000
AI Runtime & Platform Engineer for Production Systems
AI Runtime & Platform Engineer for Production Systems

crewAI, Inc. • San Francisco (CA)

On-site
USD 180,000 - 240,000
Software Engineer - AI Agent Platforms
Software Engineer - AI Agent Platforms

Workato • San Francisco (CA)

On-site
USD 140,000 - 200,000
Senior Software Engineer, Engine & Distributed Systems
Senior Software Engineer, Engine & Distributed Systems

Stack Ai • United States

On-site
USD 170,000 - 260,000
Senior Software Engineer, Engine & Distributed Systems
Senior Software Engineer, Engine & Distributed Systems

Stack AI, Inc. • New York (NY)

On-site
USD 120,000 - 160,000
Senior Software Engineer, Engine & Distributed Systems
Senior Software Engineer, Engine & Distributed Systems

StackAI • San Francisco (CA)

On-site
USD 120,000 - 160,000
Senior Software Engineer, Engine & Distributed Systems
Senior Software Engineer, Engine & Distributed Systems

StackAI • New York (NY)

On-site
USD 202,000 - 230,000
Forward Deployment Engineer
Forward Deployment Engineer

TechDigital Group • Bloomfield (CT)

On-site
USD 120,000 - 190,000
Founding AI Agent Engineer
Founding AI Agent Engineer

Allus AI (YC F25) • Atlanta (GA)

On-site
USD 120,000 - 160,000
Founding AI Agent Engineer
Founding AI Agent Engineer

Allus AI (YC F25) • Cupertino (CA)

On-site
USD 120,000 - 150,000