GPU-Accelerated AI Cloud Engineer

Alarm

Tysons, Northern (VA, KY)

Hybrid

USD 100,000 - 135,000

Full time

7 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Medical plans with subsidies
Health Savings Account with company -?
401(k) with employer match
Paid vacation
Paid holidays and wellness time

Job summary

Alarm.com is seeking a Cloud Systems Engineer to deploy and operate GPU-accelerated AI and HPC compute infrastructure. You will manage Linux environments, hardware lifecycle, and incident response while collaborating with AI engineering, networking, and storage teams.

Ideal candidates have hands-on Linux expertise, hardware troubleshooting skills, and experience with GPU-centric data workloads. The role emphasizes reliability, scalability, and operational excellence in a fast-paced environment.

Qualifications

  • 3-5 years of Linux systems administration experience in production environments.
  • 3-5 years of experience supporting enterprise server infrastructure.
  • Experience supporting large-scale compute environments, HPC platforms, AI infrastructure, or GPU-enabled systems.
  • Experience performing hardware diagnostics, firmware management, and lifecycle maintenance.
  • Experience working within datacenter operations environments.
  • Bash, Python, PowerShell, or similar scripting languages.
  • Operating system performance tuning and monitoring.
  • Storage and networking fundamentals.
  • Experience with infrastructure monitoring and observability platforms, e.g. Grafana.
  • Hardware and firmware lifecycle management.

Responsibilities

  • Deploy, configure, and maintain GPU-accelerated compute infrastructure.
  • Manage OS, firmware, BIOS, BMC, driver, and software lifecycle updates.
  • Monitor system health, performance, utilization, and capacity.
  • Support AI model training, inference, and data processing workloads.
  • Develop and maintain runbooks and maintenance procedures.
  • Participate in on-call support and incident response.
  • Administer enterprise Linux environments (Ubuntu, Red Hat).
  • Perform system patching, hardening, and OS lifecycle management.
  • Troubleshoot OS, kernel, storage, networking, and app-level issues.
  • Develop automation to streamline deployment and monitoring.
  • Support security and compliance initiatives across AI infra.

Skills

Linux administration
Bash
Python
PowerShell
System performance tuning
On-call support
Networking fundamentals
Storage fundamentals
Observability (Grafana)
Hardware troubleshooting

Tools

Grafana
BMC

Job description

Alarm.com is seeking a Cloud Systems Engineer to deploy and operate GPU-accelerated AI and HPC compute infrastructure. You will manage Linux environments, hardware lifecycle, and incident response while collaborating with AI engineering, networking, and storage teams.

Ideal candidates have hands-on Linux expertise, hardware troubleshooting skills, and experience with GPU-centric data workloads. The role emphasizes reliability, scalability, and operational excellence in a fast-paced environment.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

AI & HPC Cloud Systems Engineer
AI & HPC Cloud Systems Engineer

Alarm Com • Tysons (VA)

On-site
USD 100,000 - 135,000
401(k) match
Health Insurance
Paid vacation
Cloud Systems Engineer
Cloud Systems Engineer

Alarm Com • Tysons (VA)

On-site
USD 100,000 - 135,000
401(k) match
Health Insurance
Paid vacation
AI/HPC Infrastructure Engineer: GPU Compute & Hybrid Cloud
AI/HPC Infrastructure Engineer: GPU Compute & Hybrid Cloud

Saigepartners • San Jose (CA)

Hybrid
USD 120,000 - 180,000
GPU-Accelerated AI Cloud Hardware Architect
GPU-Accelerated AI Cloud Hardware Architect

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 157,000 - 213,000
Health insurance
RSUs and competitive benefits
Paid time off
Senior GPU Data Center Engineer
Senior GPU Data Center Engineer

Prime Intellect AI • San Francisco (CA)

On-site
USD 150,000 - 300,000
Senior AI & GPU Cloud Solutions Architect
Senior AI & GPU Cloud Solutions Architect

NVIDIA • New Jersey

On-site
USD 152,000 - 287,500
Equity
Benefits
Senior GPU Compute Solutions Architect
Senior GPU Compute Solutions Architect

Computacenter AG & Co. oHG • Northern (KY)

Hybrid
USD 190,000 - 230,000
GPU Cloud Infrastructure Engineer
GPU Cloud Infrastructure Engineer

Prime Intellect • San Francisco (CA)

On-site
USD 150,000 - 300,000
AI & HPC Infra Engineer: GPU Compute & Cloud Automation
AI & HPC Infra Engineer: GPU Compute & Cloud Automation

Prodapt ASIC services (Formerly Innovative Logic) • San Jose (CA)

On-site
USD 150,000 - 210,000
Senior AI Infra Architect GPU & Cloud
Senior AI Infra Architect GPU & Cloud

Nvidia Corporation in • Washington

On-site
USD 184,000 - 357,000
Equity
Benefits