Infrastructure Engineer – Self-Hosted AI Platform

Serval, Inc.

San Francisco (CA)

On-site

USD 190,000 - 240,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Serval, Inc. is building an AI-native automation platform. We are seeking a Software Engineer, Infrastructure to design, implement, and operate large-scale distributed systems powering Serval's AI agents and workflow orchestration.

You will focus on self-hosted deployments for enterprise customers and maintain production-grade infrastructure. You will work with Terraform, Docker, and Kubernetes across AWS, GCP, or Azure, ensuring high availability, security, and scalable data pipelines.

Qualifications

  • 3+ years building and operating large-scale distributed systems in production environments.
  • Strong experience writing and maintaining Terraform for infrastructure provisioning and management.
  • Deep knowledge of at least one major cloud provider (AWS, GCP, or Azure), including compute, networking, storage, and managed services.
  • Experience building, packaging, and supporting self-hosted or on-premises software deployments for enterprise customers.
  • Proficiency in Python, Go, or similar languages for building automation, tooling, and infrastructure services.
  • Strong understanding of networking, databases, containerization (Docker, Kubernetes), and orchestration systems.
  • Experience with monitoring, logging, alerting, and incident management tools (e.g., Datadog, Prometheus, Grafana, PagerDuty).
  • Ability to communicate technical concepts clearly to customers and provide infrastructure support and guidance.
  • Ability to debug complex system issues, analyze performance bottlenecks, and implement effective solutions.

Responsibilities

  • Design, implement, and operate large-scale distributed systems that power Serval's AI agents, workflow orchestration, and data pipelines.
  • Write and maintain Terraform modules to provision and manage cloud infrastructure across AWS, GCP, or Azure environments.
  • Build and maintain deployment packages, installation scripts, and infrastructure templates that enable customers to self-host Serval in their own environments.
  • Provide technical guidance and troubleshooting support to enterprise customers deploying and operating self-hosted instances of Serval.
  • Ensure high availability, performance, and reliability of production systems through monitoring, alerting, incident response, and capacity planning.
  • Build internal tools and platforms that enable product engineers to deploy, test, and operate services efficiently.
  • Collaborate with engineering teams to design resilient, scalable architectures that support both cloud-hosted and self-hosted deployment models.
  • Profile and optimize system performance, including compute, storage, networking, and database layers.
  • Implement security best practices and ensure infrastructure meets enterprise compliance requirements for both managed and self-hosted deployments.

Skills

Distributed systems
Terraform
Cloud platforms
Python
Go
Networking
Containers
Monitoring
Debugging
Communication
Kubernetes in production

Tools

Terraform
Docker
Kubernetes
Datadog
Prometheus
Grafana

Job description

Serval, Inc. is building an AI-native automation platform. We are seeking a Software Engineer, Infrastructure to design, implement, and operate large-scale distributed systems powering Serval's AI agents and workflow orchestration.

You will focus on self-hosted deployments for enterprise customers and maintain production-grade infrastructure. You will work with Terraform, Docker, and Kubernetes across AWS, GCP, or Azure, ensuring high availability, security, and scalable data pipelines.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Infra Engineer, Self-Hosted AI Platform
Infra Engineer, Self-Hosted AI Platform

Serval • San Francisco (CA)

On-site
Software Engineer, Infrastructure
Software Engineer, Infrastructure

Serval • San Francisco (CA)

On-site
USD 120,000 - 160,000
Key player in product success
Rapid growth opportunities
Culture of innovation and accountability
Software Engineer, Infrastructure
Software Engineer, Infrastructure

Serval, Inc. • San Francisco (CA)

On-site
USD 190,000 - 240,000
AI-Driven Internal Ops Automation Engineer
AI-Driven Internal Ops Automation Engineer

SERVAL • New York (NY)

On-site
USD 110,000 - 170,000
Head of Engineering, AI Automation Platform
Head of Engineering, AI Automation Platform

Serval, Inc. • San Francisco (CA)

On-site
USD 210,000 - 320,000
Impact
Growth opportunities
Culture
AI Automation Engineer — Internal Ops Platform
AI Automation Engineer — Internal Ops Platform

Serval • San Francisco (CA)

On-site
USD 140,000 - 190,000
Impact on product and operations
Growth opportunities
Culture of ownership and innovation
Applied AI Engineer: Build Production AI Agents
Applied AI Engineer: Build Production AI Agents

Serval • San Francisco (CA)

On-site
USD 140,000 - 210,000
Software Engineer, Applied AI
Software Engineer, Applied AI

Serval • San Francisco (CA)

On-site
USD 140,000 - 210,000
Software Engineer, Applied AI
Software Engineer, Applied AI

Latitude • San Francisco (CA), Northern (KY)

Hybrid
USD 200,000 - 325,000
AI-Driven Automation Engineer for Internal Ops
AI-Driven Automation Engineer for Internal Ops

Serval Inc. • New York (NY)

On-site
USD 90,000 - 130,000
Impact
Growth
Culture