AI Infrastructure Engineer

Socket.dev

Los Altos (CA)

On-site

USD 110,000 - 170,000

Full time

13 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Competitive Salary
Stock Option Plan
Medical, dental, vision
Retirement plan
Paid time off
Learning and development

Job summary

Socket.dev is seeking an Infrastructure Engineer to design and operate our real-time AI platform. You will build scalable cloud infrastructure, improve reliability, and own deployment pipelines. The role emphasizes production-grade automation, observability, and cost-aware engineering decisions.

You will work with Python, Docker, AWS, and Terraform/OpenTofu, while shaping architecture with product and AI teams. Strong judgment on latency, security, and scalability is essential.

Qualifications

  • 3+ years industrial experience in a relevant technical domain.
  • Strong fundamentals in software engineering and production distributed systems.
  • Hands-on experience with a major cloud platform, AWS preferred.
  • Experience with containers, IaC, CI/CD, monitoring, and production debugging.
  • Ability to write reliable automation and services in Python or similar.
  • Judgment on availability, latency, scalability, security, and cost.

Responsibilities

  • Design, build, and evolve secure, scalable cloud infrastructure for real-time AI services.
  • Improve service reliability with clear SLOs, observability, and incident prevention.
  • Build deployment and release systems that are fast, repeatable, auditable, and safe.
  • Own infrastructure as code and reusable platform patterns across environments.
  • Collaborate with product and AI engineers on architecture and data flows.
  • Diagnose complex distributed-system failures across boundaries.
  • Reduce infra and model-serving costs without harming customer experience.
  • Strengthen secrets management, access controls, backup, and vulnerability management.
  • Build internal tooling to reduce manual work for engineers.
  • Participate in incident response and drive durable improvements.

Skills

Cloud infrastructure design
Distributed systems
Python
Docker
AWS
Terraform / IaC
CI/CD
Observability
Security considerations

Tools

AWS
Docker
ECS
Lambda
Terraform/OpenTofu
Datadog
CI/CD tooling
APIGateway

Job description

Palona’s AI agents operate continuously in production, handle real-time guest interactions, integrate with restaurant systems, and face sharp traffic peaks. Infrastructure is therefore part of the product: latency, reliability, deployment safety, observability, security, and cost directly shape the guest and operator experience.

We are looking for an Infrastructure Engineer who combines cloud and reliability depth with strong software engineering judgment. You will build and operate the platform beneath Palona’s AI products, improve how engineers ship, and turn production signals into durable system improvements. This is not a ticket-driven IT or operations role. You will write production code, design systems, automate repetitive work, and own outcomes across the full service lifecycle.

Our current environment includes Python services, Docker, AWS and selected Azure services, ECS and Lambda workloads, API Gateway, load balancers, relational data systems, OpenTofu/Terraform, Datadog, and CI/CD automation. We value the ability to learn and make sound tradeoffs more than exact tool-for-tool matching.

What you will own:
  • Design, build, and evolve secure, scalable cloud infrastructure for real-time AI services and customer-facing applications.
  • Improve service reliability through clear SLOs, actionable observability, capacity planning, failure testing, and pragmatic incident prevention.
  • Build deployment and release systems that make production changes fast, repeatable, auditable, and safe.
  • Own infrastructure as code, environment consistency, and reusable platform patterns across development, staging, and production.
  • Partner with product and AI engineers on architecture, performance, data flows, and operational readiness for new capabilities.
  • Diagnose complex distributed-system failures across application, network, database, model-provider, and third-party integration boundaries.
  • Reduce infrastructure and model-serving cost without compromising customer experience or engineering velocity.
  • Strengthen secrets management, access controls, backup and recovery, vulnerability management, and other practical security foundations.
  • Build internal tooling and paved paths that let engineers ship and operate services with less manual work.
  • Participate in incident response and turn incidents into better systems, automation, documentation, and engineering judgment.
Requirements
  • 3+ years industrial experience in relevant technical domain.
  • Strong software engineering fundamentals and experience building or operating production distributed systems.
  • Hands-on experience with a major cloud platform; AWS experience is especially relevant.
  • Experience with containers, infrastructure as code, CI/CD, monitoring, alerting, and production debugging.
  • Ability to write reliable automation and services in Python or another modern programming language.
  • Sound judgment around availability, latency, scalability, security, and cost tradeoffs.
  • A track record of taking ambiguous operational problems from diagnosis through durable resolution.
  • Clear communication during architecture reviews, launches, and incidents.
  • AI-native working habits and curiosity about the operational behavior of LLM- and agent-powered systems.
Benefits
  • Competitive Salary and Stock Option Plan.
  • Medical, dental, vision, retirement, leave, and disability benefits as applicable.
  • Family Leave
  • Short Term & Long Term Disability
  • Paid time off and company holidays.
  • Learning and development support.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Infrastructure Engineer
AI Infrastructure Engineer

Palona AI • Los Altos (CA)

On-site
USD 150,000 - 230,000
Stock options
Medical, dental, vision
Paid time off
+1
AI Infrastructure Engineer
AI Infrastructure Engineer

Palona AI • New York (NY)

On-site
USD 110,000 - 170,000
Competitive salary
Stock option plan
Medical, dental, vision benefits
+5
AI Software Engineer, Growth
AI Software Engineer, Growth

Palona AI • Los Altos (CA)

On-site
USD 140,000 - 190,000
Stock options
Health insurance
Retirement plan
+1
AI Software Engineer, Growth
AI Software Engineer, Growth

Palona AI • New York (NY)

On-site
USD 120,000 - 170,000
Competitive Salary
Stock Option Plan
Medical, dental, vision, retirement
+3
AI Infrastructure Engineer — Real-Time, Secure Cloud, Stock Options
AI Infrastructure Engineer — Real-Time, Secure Cloud, Stock Options

Palona AI • New York (NY)

On-site
USD 110,000 - 170,000
Competitive salary
Stock option plan
Medical, dental, vision benefits
+5
AI Software Engineer, Growth
AI Software Engineer, Growth

Worky • Los Altos (CA)

On-site
USD 120,000 - 190,000
Stock options
Benefits package
Paid time off
Real-Time AI Infra Engineer: Cloud, Reliability
Real-Time AI Infra Engineer: Cloud, Reliability

Palona AI • Los Altos (CA)

On-site
USD 150,000 - 230,000
Stock options
Medical, dental, vision
Paid time off
+1
AI Software Engineer, Growth
AI Software Engineer, Growth

Worky • New York (NY)

On-site
USD 120,000 - 160,000
Competitive Salary
Stock Option Plan
Medical, dental, vision
+5
AI Modeling Engineer
AI Modeling Engineer

Socket.dev • Los Altos (CA)

On-site
USD 140,000 - 200,000
Competitive Salary and Stock Options
Medical, dental, vision, retirement, &
Family Leave
+3
AI Modeling Engineer
AI Modeling Engineer

Palona AI • Los Altos (CA)

On-site
USD 140,000 - 210,000
Stock options
Benefits: medical/dental/vision/ret/le
Family leave
+3