Infrastructure Production Engineer

Socket.dev

United States

Remote

USD 60,000 - 80,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Health insurance
401k matching
Professional development
Paid time off
Sabbatical program
Remote office stipend
Internet reimbursement
Gym membership
Wellable subscription

Job summary

Vultr is seeking a highly skilled Associate Infrastructure Production Engineer to ensure GPU hardware validation and readiness across our global production environment. The role requires analytical problem-solving, strong Linux skills, and Python scripting, with collaboration across engineering to refine testing processes and automate validations.

You will design automated diagnostic and remediation frameworks, develop Python-based agents and Ansible automation, and improve deployment

Qualifications

  • Strong Linux proficiency with hands-on experience in production environments.
  • Proficient in Python scripting for automation and data collection.
  • Experience writing and maintaining shell scripts for diagnostics and automation.
  • Familiarity with Ansible for configuration management and automation tasks.
  • Ability to document work clearly, including commands, observations, and outcomes.
  • Excellent communication and collaboration with engineering teams.

Responsibilities

  • Design, develop, and maintain automated diagnostic and validation frameworks for production GPU hardware.
  • Build and support Python-based agents, services, APIs, and Ansible playbooks for provisioning, telemetry, and health monitoring.
  • Analyze workload performance and diagnostics to improve validation methodologies and readiness standards.
  • Evolve testing and verification processes for onboarding and RMAs, increasing reliability and coverage.
  • Create tooling to gate hardware deployment and enforce quality standards to reduce risk.
  • Document system designs and operational guidance; maintain Jira records of activities and outcomes.
  • Identify gaps in testing or automation and design solutions with engineering teams.

Skills

Linux proficiency
Python scripting
Shell scripting
Ansible
Documentation
Strong problem-solving
Networking basics

Tools

Ansible
Python

Job description

Who We Are

Vultr is on a mission to make high-performance cloud infrastructure easy to use, affordable, and locally accessible for enterprises and AI innovators around the world. With 33 global cloud data center locations, Vultr is trusted by hundreds of thousands of active customers across 185 countries for its flexible, scalable, global Cloud Compute, Cloud GPU, Bare Metal, and Cloud Storage solutions. In December 2024 Vultr announced an equity financing at a $3.5 billion valuation. Founded by David Aninowsky and self-funded for over a decade, Vultr has grown to become the world’s largest privately‑held cloud infrastructure company.

Vultr Cares
  • 100% company-paid insurance premiums for employee medical, dental and vision plans.
  • 401(k) plan that matches 100% up to 4%, with immediate vestingProfessional Development Reimbursement of $2,500 each year11 Holidays + Paid Time Off Accrual + Rollover PlanCommitment matters to Vultr!
  • Increased PTO at 3 year and 10 year anniversary + 1 month paid sabbatical every 5 years + Anniversary Bonus each year$500 stipend for remote office setup in first year + $400 each following yearInternet reimbursement up to $75 per monthGym membership reimbursement up to $50 per monthCompany paid Wellable subscription
Join Vultr

Vultr is seeking a highly skilled and experienced Associate Infrastructure Production Engineer to ensure GPU hardware validation and readiness across our global production environment. The ideal candidate is an analytical problem-solver with strong Linux skills, Python scripting ability, and a detail-oriented approach to infrastructure validation and automation. This is a highly visible role in a high-growth technology company, which will require collaboration with engineering teams to refine testing processes, close gaps in automation coverage, and build technical solutions that reduce deployment risk and improve production standards. This is your opportunity to join our fast growing team and leave your mark on Vultr and the future of Cloud Infrastructure.

Key Responsibilities
  • Design, develop, and maintain automated diagnostic, validation, and remediation frameworks for production GPU hardware (NVIDIA and AMD) using vendor tooling, Python, and infrastructure automation technologies.
  • Engineer and support Python-based agents, services, APIs, and Ansible automation (playbooks, roles, pipelines) that orchestrate hardware provisioning, telemetry collection, health monitoring, and production onboarding workflows.
  • Analyze workload performance, utilization, thermals, and diagnostic output to identify hardware and system issues, enhance validation methodologies, and improve infrastructure readiness standards.
  • Execute and evolve testing and verification processes for production onboarding and Return Material Authorization (RMA), developing automation enhancements to improve reliability, scalability, and coverage.
  • Contribute to production stability by building tooling that gates hardware deployment, enforces quality standards, and reduces systemic infrastructure risk.
  • Document system designs, automation logic, validation methodologies, and operational guidance; maintain accurate Jira records reflecting engineering activities, findings, and outcomes.
  • Identify gaps or inefficiencies within testing, validation, or automation processes and design technical solutions to address them in collaboration with engineering teams.
Qualifications
  • Strong analytical skills and attention to detail with the ability to evaluate, enhance, and optimize validation and testing methodologies.
  • Ability to clearly document work performed, including commands executed, observations, and outcomes.
  • Strong Linux proficiency, and experience working within Linux production environments in command-line system contexts.
  • Ability to design, write, and modify Python and shell scripts to support infrastructure diagnostics, automation, and validation workflows (required after initial training).
  • Familiarity with infrastructure automation or configuration management frameworks such as Ansible is strongly preferred.
  • Strong written and verbal communication skills and the ability to collaborate effectively with technical teams.
  • Understanding of data center networking concepts and protocols such as DHCP, IPv6, and ICMP.
  • Ability to collaborate with engineering teams to deliver automation solutions, reliability improvements, and production-ready tooling.
Compensation

$60,000 - $80,000

Final compensation will vary depending on years of experience, background/skill set, location, and applicable laws

Inclusion & Privacy

We are an equal opportunity employer and are committed to creating an inclusive environment for all employees. We welcome applications from individuals of all backgrounds and experiences, and we prohibit discrimination based on race, color, religion, sex, sexual orientation, gender identity, national origin, age, disability, veteran status, or any other protected status under applicable laws. Vultr will consider qualified applicants with arrest or conviction records in accordance with applicable laws and will not conduct a background check until after an offer of employment has been extended and accepted.

We also take your privacy seriously. We handle personal information responsibly and follow applicable laws, including U.S. privacy rules and India’s Digital Personal Data Protection Act, 2023. Your data is used only for legitimate business purposes and is protected with proper security measures. Where allowed by law, applicants may request details about the data we collect, access or delete their information, withdraw consent for its use, and opt out of nonessential communications. For more details, please see our Privacy Policy.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Infrastructure Production Engineer
Infrastructure Production Engineer

Vultr • United States

On-site
USD 60,000 - 80,000
Premium health insurance
401(k) with company match
Professional development reimbursement
+6
Infrastructure Production Engineer
Infrastructure Production Engineer

Webhosting • Northern (KY)

Hybrid
USD 60,000 - 80,000
Insurance
401(k) matching
Professional development reimbursement
+5
Staff AI/ML Infrastructure Engineer
Staff AI/ML Infrastructure Engineer

Vultr • United States

Remote
USD 145,000 - 160,000
100% company-paid insurance premiums
401(k) plan with matching
Professional Development Reimbursement
+4
Data Center Technician (Kansas City)
Data Center Technician (Kansas City)

Vultr • Kansas City (MO)

On-site
USD 60,000 - 75,000
100% company-paid insurance premiums
401(k) matching
Professional Development Reimbursement
+5
Software Engineer, Core Cloud Engineering
Software Engineer, Core Cloud Engineering

Webhosting • Northern (KY)

Hybrid
USD 80,000 - 95,000
Company-paid health plan
401(k) match
Professional development allowance
+3
Data Center Technican (Eagan, MN)
Data Center Technican (Eagan, MN)

Vultr • Eagan (MN)

On-site
USD 60,000 - 95,000
100% company-paid insurance premiums
401(k) plan with company match
Professional development reimbursement
+2
Senior Technical Project Manager, Data Center & Network Delivery
Senior Technical Project Manager, Data Center & Network Delivery

Vultr • United States

On-site
USD 110,000 - 140,000
100% company-paid insurance premiums for medical, dental, and vision plans
401(k) plan with 100% match up to 4%
Professional Development Reimbursement of $2,500 each year
+6
Data Center Technician (Springfield, OH) - Multiple Openings
Data Center Technician (Springfield, OH) - Multiple Openings

Vultr • Springfield (OH)

On-site
USD 65,000 - 75,000
100% company-paid insurance premiums
401(k) plan with matching
Professional Development Reimbursement
+2
Software Engineer, Core Cloud Engineering
Software Engineer, Core Cloud Engineering

Vultr • United States

On-site
USD 80,000 - 95,000
100% company-paid insurance premiums
401(k) employer match
Professional Development Reimbursement
+6
Senior Linux System Administrator
Senior Linux System Administrator

Webhosting • Northern (KY)

Hybrid
USD 80,000 - 100,000
Health insurance
401(k) with match
Professional development reimbursement
+6