Production Engineer

Socket.dev

San Mateo (CA)

On-site

USD 180,000 - 240,000

Full time

8 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Stock options
Health insurance
Relocation assistance
401K

Job summary

Skydio, a US-based drone company, is seeking a hands-on Site Reliability Engineer to own production infrastructure across Kubernetes, AWS, IaC, and CI/CD. You'll focus on reliability, scalability, and on-call readiness to ensure uptime for critical drone software.

Build and operate Kubernetes/EKS clusters, manage AWS networking and IAM, and advance Terraform-based infrastructure. You will partner with software, security, and SRE teams to reduce toil and accelerate deployments.

Qualifications

  • 3+ years in Production Engineer, SRE, DevOps or equivalent.
  • Strong hands-on Kubernetes experience beyond basic deployments.
  • Solid AWS fundamentals: VPCs, subnets, IAM, EKS, databases.
  • Terraform or similar IaC tooling experience.
  • Experience with CI/CD systems such as Argo CD, Spinnaker, GitHub Actions, GitLab CI/CD or Jenkins.
  • Ability to diagnose production infrastructure and networking issues.
  • Willingness to participate in on-call rotations.

Responsibilities

  • Build, operate, and troubleshoot production Kubernetes/EKS clusters.
  • Perform Kubernetes upgrades, node rollouts, and cluster maintenance.
  • Manage AWS infrastructure: VPCs, subnets, load balancers, IAM, EKS, databases.
  • Define and maintain infrastructure using Terraform.
  • Build and operate CI/CD and deployment infrastructure.
  • Troubleshoot production issues across Kubernetes, AWS, Linux, networking, and databases.
  • Build monitoring and observability for critical infrastructure.
  • Participate in on-call rotations and respond to incidents.
  • Identify and fix scaling and reliability problems.
  • Automate operational tasks with Python or Go.
  • Help expand infrastructure to new regions.

Skills

Kubernetes
AWS
Python
Go
CI/CD tooling
Infrastructure as Code

Tools

Terraform
Argo CD
Spinnaker
GitHub Actions
GitLab CI/CD
Jenkins
Datadog

Job description

Skydio is the leading US drone company and the world leader in autonomous flight, the key technology for the future of drones and aerial mobility. The Skydio team combines deep expertise in artificial intelligence, best-in-class hardware and software product development, operational excellence, and customer obsession to empower a broader, more diverse audience of drone users, from utility inspectors to first responders, soldiers in battlefield scenarios, and beyond.

About the role:

We are looking for a hands‑on Site Reliability Engineer to build, operate, and scale the cloud infrastructure that powers our products. This role is focused on owning production infrastructure, including Kubernetes, AWS, infrastructure as code, CI/CD, observability, networking, and reliability.

You don't need to be an expert in every area, but you should have strong Kubernetes and cloud fundamentals with meaningful depth in at least one infrastructure domain. Our technology helps save lives. You’ll play a critical role in keeping the infrastructure behind it reliable, scalable, and available when it matters most.

How you'll make an impact:
  • Build, operate, and troubleshoot production Kubernetes/EKS clusters.

  • Perform Kubernetes upgrades, node rollouts, and cluster maintenance.

  • Build and manage AWS infrastructure including VPCs, networking, subnets, load balancers, IAM, EKS, databases, and storage.

  • Define and maintain infrastructure using Terraform.

  • Build and operate CI/CD and deployment infrastructure.

  • Troubleshoot production issues across Kubernetes, AWS, Linux, networking, and databases.

  • Build monitoring, alerting, and observability for critical infrastructure.

  • Participate in on‑call rotations and respond to production incidents.

  • Identify and solve infrastructure scaling and reliability problems.

  • Automate operational work using Python, Go, or similar languages.

  • Help expand infrastructure across new regions and deployment environments.

What makes you a good fit:
  • 3+ years of experience as a Production Engineer, SRE, DevOps, SRE or equivalent infrastructure role.

  • Strong hands‑on experience operating Kubernetes, not simply deploying applications to existing clusters.

  • Experience managing Kubernetes/EKS upgrades and production clusters.

  • Strong AWS fundamentals, including VPCs, public/private subnets, networking, load balancers, EKS, IAM, and databases.

  • Production experience with Terraform or similar infrastructure-as-code tooling.

  • Experience owning or maintaining CI/CD and deployment systems such as Argo CD, Spinnaker, GitHub Actions, GitLab CI/CD, or Jenkins.

  • Experience diagnosing production infrastructure and networking problems.

  • Experience solving meaningful scaling or reliability challenges.

  • This position requires access to export‑controlled technical data, restricted government information, and/or information systems subject to U.S. government security and access‑control requirements. Employment in this role is contingent upon verification of U.S. person status and the ability to access controlled or restricted information as required for the position.

Bonus points:
  • Helm and GitOps experience.

  • Datadog or similar observability tooling.

  • PostgreSQL/database operations experience.

  • Multi‑region infrastructure experience.

  • On‑premises or disconnected deployment experience.

  • Streaming or high‑throughput distributed systems experience.

Compensation:

At Skydio, our compensation packages for regular, full‑time employees include competitive base salaries, equity in the form of stock options, and comprehensive benefits packages. Compensation will vary based on factors, including skill level, proficiencies, transferable knowledge, and experience. Relocation assistance may also be provided for eligible roles. The annual base salary range for this position is $180,000 - 240,000*. Fundamentally, we believe that equity is the key to long‑term financial growth, and we ensure all regular, full‑time employees have the opportunity to significantly benefit from the company’s success. Regular, full‑time employees are eligible to enroll in the Company’s group health insurance plans. Regular, full‑time employees are eligible to receive the following benefits: Paid vacation time, sick leave, holiday pay and 401K savings plan. This position and all associated benefits are subject to applicable federal, state, and local laws, as well as the Company’s policies and eligibility criteria.

*For some positions the pay may be dependent upon the individual’s regional location.

#LI-WA1

At Skydio we believe that diversity drives innovation. We have created a multidisciplinary environment that embraces the power of diverse perspectives to create elegant solutions for complex problems. We are committed to growing our network of people, programs, and resources to nurture an inclusive culture.

Qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, disability, protected veteran status, or other characteristics protected by federal, state or local anti‑discrimination laws.

For positions located in the United States of America, Skydio, Inc. uses E‑Verify to confirm employment eligibility. To learn more about E‑Verify, including your rights and responsibilities, please visit https://www.e-verify.gov/

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff Site Reliability Engineer
Staff Site Reliability Engineer

Socket.dev • San Mateo (CA)

On-site
USD 240,000 - 300,000
Stock options
Health insurance
401K savings plan
Site Reliability Engineer
Site Reliability Engineer

Skydio • United States

On-site
USD 180,000 - 240,000
Stock options
Health insurance
Paid vacation
+3
Staff Site Reliability Engineer
Staff Site Reliability Engineer

Skydio • San Mateo (CA)

On-site
USD 240,000 - 300,000
Equity
Health insurance
Paid time off
+1
Staff Site Reliability Engineer
Staff Site Reliability Engineer

Skydio • United States

On-site
USD 240,000 - 300,000
Stock options
Health insurance
Paid vacation
+4
Site Reliability Engineer
Site Reliability Engineer

Booster • United States

Hybrid
USD 140,000 - 210,000
Senior Software Engineer, Infrastructure
Senior Software Engineer, Infrastructure

Skydio • United States

Hybrid
USD 170,000 - 258,000
Paid vacation time
Comprehensive health insurance
401K savings plan
Production Supervisor - Dayshift
Production Supervisor - Dayshift

Socket.dev • Hayward (CA)

On-site
USD 114,000 - 146,000
Stock options
Relocation assistance
Health insurance
+1
Production Supervisor - Nightshift
Production Supervisor - Nightshift

Socket.dev • Hayward (CA)

On-site
USD 114,000 - 146,000
Staff Software Engineer – Infrastructure
Staff Software Engineer – Infrastructure

Skydio • San Mateo (CA)

Hybrid
USD 230,000 - 275,000
Competitive base salaries
Stock options
Comprehensive benefits packages
+2
Staff Software Engineer - Infrastructure
Staff Software Engineer - Infrastructure

Skydio • United States

On-site
USD 230,000 - 275,000
Equity in the form of stock options
Comprehensive health insurance
Paid vacation time
+1