AI Infrastructure Engineer

Palona AI

New York (NY)

On-site

USD 110,000 - 170,000

Full time

10 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Competitive salary
Stock option plan
Medical, dental, vision benefits
Retirement plan
Family leave
Short/Long-term disability
Paid time off and holidays
Learning and development support

Job summary

Palona AI is seeking a cloud infrastructure engineer to design and evolve secure, scalable systems powering real-time AI services. You will work on architecture, performance, data flows, and operational readiness to support new capabilities and complex AI workloads.

You will drive reliability through SLOs, capacity planning, and incident prevention, while building deployment systems and IaC patterns across development, staging, and production.

Qualifications

  • 3+ years industrial experience in relevant technical domain.
  • Experience building/operating production distributed systems.
  • Hands-on with a major cloud platform (AWS preferred).
  • Experience with containers, IaC, CI/CD, monitoring, alerting, and production debugging.
  • Ability to write reliable automation and services in Python or another modern language.
  • Judgment on availability, latency, scalability, security, and cost tradeoffs.
  • History of diagnosing ambiguous operational problems to durable resolution.
  • Clear communication during architecture reviews, launches, and incidents.
  • AI-native operating habits and curiosity about LLM- and agent-powered systems.

Responsibilities

  • Design, build, and evolve secure, scalable cloud infrastructure for real-time AI services and customer-facing applications.
  • Improve service reliability through clear SLOs, observability, capacity planning, failure testing, and incident prevention.
  • Build deployment and release systems that are fast, repeatable, auditable, and safe.
  • Own infrastructure as code, environment consistency, and reusable platform patterns across environments.
  • Partner with product and AI engineers on architecture, performance, data flows, and operational readiness for new capabilities.
  • Diagnose complex distributed-system failures across application, network, database, model-provider, and third-party boundaries.
  • Reduce infrastructure and model-serving cost without harming user experience or velocity.
  • Strengthen secrets management, access controls, backups, vulnerability management, and security foundations.
  • Build internal tooling and paved paths to ship and operate services with less manual work.
  • Participate in incident response and turn incidents into better systems and processes.

Skills

Distributed systems
Python automation
System reliability
Incident response
Communication
Security awareness
Cost optimization
AI-native operations

Tools

AWS
Kubernetes
Docker
Terraform
CI/CD tooling
Monitoring

Job description

Palona's AI agents operate continuously in production, handle real-time guest interactions, integrate with restaurant systems, and face sharp traffic peaks. Infrastructure is therefore part of the product: latency, reliability, deployment safety, observability, security, and cost directly shape the guest and operator experience.

What you will own:
  • Design, build, and evolve secure, scalable cloud infrastructure for real-time AI services and customer-facing applications
  • Improve service reliability through clear SLOs, actionable observability, capacity planning, failure testing, and pragmatic incident prevention
  • Build deployment and release systems that make production changes fast, repeatable, auditable, and safe
  • Own infrastructure as code, environment consistency, and reusable platform patterns across development, staging, and production
  • Partner with product and AI engineers on architecture, performance, data flows, and operational readiness for new capabilities
  • Diagnose complex distributed-system failures across application, network, database, model-provider, and third-party integration boundaries
  • Reduce infrastructure and model-serving cost without compromising customer experience or engineering velocity
  • Strengthen secrets management, access controls, backup and recovery, vulnerability management, and other practical security foundations
  • Build internal tooling and paved paths that let engineers ship and operate services with less manual work
  • Participate in incident response and turn incidents into better systems, automation, documentation, and engineering judgment
Requirements
  • 3+ years industrial experience in relevant technical domain
  • Strong software engineering fundamentals and experience building or operating production distributed systems
  • Hands-on experience with a major cloud platform; AWS experience is especially relevant
  • Experience with containers, infrastructure as code, CI/CD, monitoring, alerting, and production debugging
  • Ability to write reliable automation and services in Python or another modern programming language
  • Sound judgment around availability, latency, scalability, security, and cost tradeoffs
  • A track record of taking ambiguous operational problems from diagnosis through durable resolution
  • Clear communication during architecture reviews, launches, and incidents
  • AI-native working habits and curiosity about the operational behavior of LLM- and agent-powered systems
Benefits
  • Competitive Salary and Stock Option Plan
  • Medical, dental, vision, retirement, leave, and disability benefits as applicable
  • Family Leave
  • Short Term & Long Term Disability
  • Paid time off and company holidays
  • Learning and development support
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Infrastructure Engineer
AI Infrastructure Engineer

Socket.dev • Los Altos (CA)

On-site
USD 110,000 - 170,000
Competitive Salary
Stock Option Plan
Medical, dental, vision
+3
AI Infrastructure Engineer
AI Infrastructure Engineer

Palona AI • Los Altos (CA)

On-site
USD 150,000 - 230,000
Stock options
Medical, dental, vision
Paid time off
+1
AI Engineer
AI Engineer

Teserac, Inc. • Sunnyvale (CA)

On-site
USD 100,000 - 130,000
Health Care Plan (Medical, Dental & Vision)
Paid Time Off (Vacation, Sick & Public Holidays)
Free Food & Snacks
+2
AI Infrastructure Engineer
AI Infrastructure Engineer

Capstone Investment Advisors • New York (NY)

On-site
USD 120,000 - 150,000
AI Infrastructure Engineer — Real-Time, Secure Cloud, Stock Options
AI Infrastructure Engineer — Real-Time, Secure Cloud, Stock Options

Palona AI • New York (NY)

On-site
USD 110,000 - 170,000
Competitive salary
Stock option plan
Medical, dental, vision benefits
+5
AI Engineer, AIOps & Infrastructure
AI Engineer, AIOps & Infrastructure

eloquentai • San Francisco (CA)

On-site
USD 130,000 - 160,000
AI Engineer
AI Engineer

AvantStay • United States

On-site
USD 100,000 - 140,000
Equity
Generous paid time off including holidays
100% remote – work from anywhere
+3
AI Infrastructure Engineer
AI Infrastructure Engineer

Hedge Fund • New York (NY)

On-site
USD 150,000 - 230,000
AI Engineer, Agent
AI Engineer, Agent

eloquentai • San Francisco (CA)

On-site
USD 120,000 - 160,000
AI Software Engineer, Product
AI Software Engineer, Product

Palona AI • New York (NY)

On-site
USD 140,000 - 190,000
Stock options
Health benefits
Family leave
+3