Senior AI Platform Reliability Engineer

FloQast

San Jose (CA)

Hybrid

USD 186,000 - 282,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Medical insurance
Dental insurance
Vision insurance
Family forming benefits
Life & Disability Insurance
Unlimited vacation
Employee stock program

Job summary

FloQast is seeking an experienced DevOps/AI infrastructure engineer to own the AI runtime across US, EU, and AU, including Bedrock and sandboxed code execution. You’ll implement IaC, throughput control, and cost attribution for multi-region AI workloads.

The role emphasizes observability, security, and reliability, with on-call responsibilities and a focus on scalable, auditable infrastructure in a Production AI environment.

Qualifications

  • 5+ years in DevOps, SRE, platform, or infrastructure engineering with on-call experience.
  • Deep, hands-on AWS: ECS/Fargate, Lambda, SQS, S3, IAM, VPC and networking, ALB/NLB.
  • Terraform at production scale: modules, state management, multi-region, multi-account.
  • CI/CD and container ownership: GitHub Actions, Docker, image supply chain, scaling policies.
  • Production AI infrastructure: operated LLM-backed or ML-serving workloads in production.

Responsibilities

  • Own AI runtime infrastructure across US, EU, AU regions (Bedrock and AgentCore).
  • Operate sandboxed execution environments for AI-generated code.
  • Write and review Terraform across multi-account AWS estate.
  • Extend Grafana observability for AI workloads (signals, latency, token spend).
  • Manage AI spend as a cost line with tagging for FinOps.

Skills

DevOps/SRE experience
AWS expertise
Terraform at scale
CI/CD engineering
Python/TypeScript
Observability tooling

Tools

GitHub Actions
Docker
Terraform
NX monorepos
Grafana

Job description

FloQast is seeking an experienced DevOps/AI infrastructure engineer to own the AI runtime across US, EU, and AU, including Bedrock and sandboxed code execution. You’ll implement IaC, throughput control, and cost attribution for multi-region AI workloads.

The role emphasizes observability, security, and reliability, with on-call responsibilities and a focus on scalable, auditable infrastructure in a Production AI environment.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff AI Engineer: Lead Production ML & Platform
Staff AI Engineer: Lead Production ML & Platform

FloQast • San Jose (CA)

Hybrid
USD 164,000 - 246,000
Medical
Dental
Vision
+3
Senior Platform Architect – AI Workflows & SaaS
Senior Platform Architect – AI Workflows & SaaS

Sapphire Partners • San Jose (CA)

On-site
USD 180,000 - 240,000
Medical insurance
Dental insurance
Vision insurance
+3
Senior AI Platform Engineer (Remote)
Senior AI Platform Engineer (Remote)

Facility Grid • United States

Remote
USD 180,000 - 225,000
Health insurance
Paid time off
Vision insurance
+3
Senior AI Platform Engineer: Scalable AI Infra
Senior AI Platform Engineer: Scalable AI Infra

Socket.dev • San Francisco (CA)

On-site
USD 150,000 - 190,000
Senior Software Engineer & Tech Lead of Scalable AI Systems
Senior Software Engineer & Tech Lead of Scalable AI Systems

FloQast • Los Angeles (CA)

Hybrid
USD 144,000 - 216,000
Medical, Dental, Vision
Family Forming benefits
Life & Disability Insurance
+1
Senior Platform Reliability Engineer – AI-Driven FinServ
Senior Platform Reliability Engineer – AI-Driven FinServ

interface.ai • San Francisco (CA)

On-site
USD 150,000 - 210,000
100% paid health, dental & vision
401(k) & financial wellness
Daily meals on us
+3
Staff AI Platform Engineer: Architect the AI Factory
Staff AI Platform Engineer: Architect the AI Factory

Relevance AI • United States

Hybrid
USD 127,000 - 162,000
ESOP – Employee Stock Ownership Plan
AI Productivity Benefit – Get up to $1
Parental Leave – 12 weeks of paid time
+5
Senior AI Cloud Platform Engineer - Reliability & Scale
Senior AI Cloud Platform Engineer - Reliability & Scale

Crusoe • San Francisco (CA)

On-site
USD 170,000 - 205,000
Health insurance
RSUs
401(k) match
+2
Staff Platform Engineer - Remote AI Infra
Staff Platform Engineer - Remote AI Infra

Front Door Defense • United States

On-site
USD 147,000 - 276,000
Remote-eligible role
Competitive compensation
Equity participation
Senior Backend Reliability Engineer — AI‑Driven Platform (Remote)
Senior Backend Reliability Engineer — AI‑Driven Platform (Remote)

Affirm • Riverside (OH)

Remote
USD 173,000 - 233,000
Health coverage
FSA Wallets
Time off
+1