AI Infra Engineer: Multi-Cloud & Observability

Socket.dev

New York (NY)

On-site

USD 150,000 - 210,000

Full time

10 days ago
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Percepta is hiring an AI Infrastructure Engineer to own the infrastructure, deployment, and reliability powering our autonomous AI systems. You’ll tighten Terraform footprints, strengthen pipelines, and define new patterns for agentic workloads.

You’ll collaborate with teams across engineering to meet SOC 2, HIPAA and regulated environment requirements while advancing observability and reliability at scale. This is a high-autonomy role with real ownership.

Qualifications

  • 5+ years building and operating production infrastructure in DevOps or SRE roles
  • The kind of engineer who sees a manual process and can't rest until it's automated well, not just scripted
  • Strong hands-on Terraform experience
  • Deep experience with at least 1 major cloud provider (AWS, GCP, or Azure): networking, IAM, cost management, the operational realities of production workloads
  • Solid Docker and Kubernetes experience in production. We run managed clusters across all 3 major clouds; this is a core part of the role
  • Experience designing and maintaining CI/CD pipelines (GitHub Actions, GitLab CI, or similar)
  • Scripting proficiency in Python, Bash, or similar
  • High agency: you don't wait for a ticket to fix what's broken, but you communicate, collaborate, and bring the team along
  • Genuine curiosity about AI systems, not just the infrastructure running them. You want to understand what you're operating
  • You find it interesting (not alarming) that some systems you'll operate will be making decisions on their own

Responsibilities

  • Define infrastructure patterns for multi-agent systems that need to be observable, controllable, and recoverable in ways traditional apps don't require
  • Own and evolve our IaC stack: Terraform and Kubernetes across AWS, GCP, and Azure
  • Build observability primitives for agentic workflows, tracing agent decisions and execution paths, not just service latency and pod health
  • Design and maintain CI/CD pipelines that give teams fast, trustworthy feedback from commit to production
  • Build operational foundations: monitoring, alerting, incident response, and the new patterns that emerge when AI systems are participants in that response
  • Work across engineering teams to meet the reliability and compliance requirements of the institutions we serve (SOC 2, HIPAA, regulated environments in healthcare and energy)

Skills

Terraform
Kubernetes
Cloud platforms
CI/CD
Python
Bash
Docker
Observability
SRE

Tools

GitHub Actions
GitLab CI
Grafana
Prometheus
Loki

Job description

Percepta is hiring an AI Infrastructure Engineer to own the infrastructure, deployment, and reliability powering our autonomous AI systems. You’ll tighten Terraform footprints, strengthen pipelines, and define new patterns for agentic workloads.

You’ll collaborate with teams across engineering to meet SOC 2, HIPAA and regulated environment requirements while advancing observability and reliability at scale. This is a high-autonomy role with real ownership.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

AI Infrastructure Engineer for Autonomous Agent Systems
AI Infrastructure Engineer for Autonomous Agent Systems

Percepta • New York (NY)

On-site
USD 140,000 - 190,000
Senior AI Systems Engineer — Cloud & ML Infra
Senior AI Systems Engineer — Cloud & ML Infra

MSM Technology • United States

On-site
USD 150,000 - 200,000
Senior AI Infra Architect Terraform & Hybrid Cloud
Senior AI Infra Architect Terraform & Hybrid Cloud

Cognizant • Princeton (NJ)

Hybrid
USD 83,000 - 110,000
Medical/Dental/Vision/Life Insurance
Paid holidays plus PTO
401(k) plan and contributions
+3
Senior Cloud Infrastructure Engineer for Enterprise AI
Senior Cloud Infrastructure Engineer for Enterprise AI

Scale AI, Inc. • New York (NY)

On-site
USD 216,000 - 270,000
Health, dental & vision coverage
Retirement benefits
Learning & development stipend
+2
Platform Engineer – AI Infra, CI/CD & Observability
Platform Engineer – AI Infra, CI/CD & Observability

Outmarket AI • San Francisco (CA)

On-site
USD 120,000 - 150,000
AI-Driven Cloud Infrastructure Automation Lead
AI-Driven Cloud Infrastructure Automation Lead

Ex • New York (NY)

Hybrid
USD 150,000 - 170,000
Senior Cloud Engineer — AI-Driven Infra, Remote + Equity
Senior Cloud Engineer — AI-Driven Infra, Remote + Equity

Apply • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
Equity
401(k) match
Parental leave
+3
AI Infrastructure Engineer
AI Infrastructure Engineer

Socket.dev • New York (NY)

On-site
USD 150,000 - 210,000
Platform Infra Engineer for AI Agents & Cloud
Platform Infra Engineer for AI Agents & Cloud

Pace • New York (NY)

On-site
USD 140,000 - 200,000
Equity
Health insurance
Unlimited PTO
+1
AI Infra Platform Lead — Kubernetes, Terraform & Ansible
AI Infra Platform Lead — Kubernetes, Terraform & Ansible

Seekr • Reston (VA)

Hybrid
USD 180,000 - 240,000
Equity RSUs
Unlimited PTO
Hybrid work environment
+1