AI & HPC Cloud Systems Engineer

Alarm Com

Tysons (VA)

On-site

USD 100,000 - 135,000

Full time

7 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

401(k) match
Health Insurance
Paid vacation

Job summary

Alarm.com is seeking a Cloud Systems Engineer to support and operate large-scale AI and HPC environments in Tysons, VA. You will deploy and maintain GPU-accelerated compute infrastructure, manage OS and firmware lifecycles, and partner with AI engineering teams to ensure reliability and scalability of our AI workloads.

The role emphasizes hands-on Linux administration, hardware diagnostics, and automation to streamline deployment and operations across AI platforms.

Qualifications

  • 3-5 years of Linux systems administration experience in production environments.
  • 3-5 years of experience supporting enterprise server infrastructure.
  • Experience supporting large-scale compute environments, HPC platforms, AI infrastructure, or GPU-enabled systems.
  • Experience performing hardware diagnostics, firmware management, and lifecycle maintenance.
  • Experience working within datacenter operations environments.
  • Bash, Python, PowerShell, or similar scripting languages
  • Operating system performance tuning and monitoring
  • Storage and networking fundamentals
  • Experience with infrastructure monitoring and observability platforms, ex Grafana.

Responsibilities

  • Deploy, configure, and maintain GPU-accelerated compute infrastructure.
  • Manage OS, firmware, BIOS, BMC, driver, and software lifecycle updates.
  • Monitor system health, performance, utilization, and capacity across AI infrastructure environments.
  • Support infrastructure for AI model training, inference, and data processing workloads.
  • Develop and maintain operational standards, runbooks, and maintenance procedures.
  • Participate in on-call support and incident response activities.
  • Administer enterprise Linux environments, including Ubuntu and Red Hat-based distributions.
  • Perform system patching, hardening, and OS lifecycle management.
  • Troubleshoot OS, kernel, storage, networking, and application-level issues.
  • Develop automation to streamline deployment, monitoring, and operational processes.
  • Support security and compliance initiatives across AI infrastructure platforms.
  • Install, configure, maintain, and troubleshoot enterprise compute hardware.
  • Diagnose and resolve issues involving GPUs, CPUs, memory, storage, power, and networking components.
  • Perform firmware upgrades and hardware lifecycle management activities.
  • Coordinate hardware replacements, vendor support engagements, and warranty services.
  • Participate in rack-and-stack deployments, datacenter expansions, and technology refresh projects.
  • Maintain accurate asset inventories and operational documentation.
  • Support high-performance networking technologies, including Ethernet and InfiniBand environments.
  • Collaborate with networking, storage, cloud, and AI engineering teams on infrastructure design and operations.
  • Assist with scalability, resiliency, and performance optimization initiatives.
  • Perform root-cause analysis of infrastructure failures and develop preventative measures.
  • Other duties as assigned.

Skills

Linux administration
Bash
Python
PowerShell
Hardware troubleshooting
System patching
Monitoring
Networking basics
Grafana

Tools

BMC tools
Automation tooling

Job description

Alarm.com is seeking a Cloud Systems Engineer to support and operate large-scale AI and HPC environments in Tysons, VA. You will deploy and maintain GPU-accelerated compute infrastructure, manage OS and firmware lifecycles, and partner with AI engineering teams to ensure reliability and scalability of our AI workloads.

The role emphasizes hands-on Linux administration, hardware diagnostics, and automation to streamline deployment and operations across AI platforms.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

GPU-Accelerated AI Cloud Engineer
GPU-Accelerated AI Cloud Engineer

Alarm • Tysons (VA), Northern (KY)

Hybrid
USD 100,000 - 135,000
Medical plans with subsidies
Health Savings Account with company -?
401(k) with employer match
+2
Cloud Systems Engineer New Tysons, Virginia
Cloud Systems Engineer New Tysons, Virginia

Alarm • Tysons (VA), Northern (KY)

On-site
USD 100,000 - 135,000
Medical plans with subsidies
Health Savings Account with company -?
401(k) with employer match
+2
AI & HPC Systems Administrator - Hybrid (3 onsite/2 remote)
AI & HPC Systems Administrator - Hybrid (3 onsite/2 remote)

100 Eli Lilly and Company • South San Francisco (CA)

Hybrid
USD 141,000 - 231,000
401(k)
Pension
Vacation benefits
+1
AI Systems Engineer: HPC & GPU Clusters
AI Systems Engineer: HPC & GPU Clusters

Advanced Micro Devices, Inc. • San Jose (CA)

On-site
USD 180,000 - 260,000
AI Systems Engineer: HPC & GPU Clusters
AI Systems Engineer: HPC & GPU Clusters

AMD • San Jose (CA)

On-site
USD 180,000 - 260,000
AMD benefits
AI Systems Engineer: Cloud, GPUs & Automation
AI Systems Engineer: Cloud, GPUs & Automation

MCI • United States

On-site
USD 90,000 - 130,000
Cloud Systems Engineer
Cloud Systems Engineer

Alarm Com • Tysons (VA)

On-site
USD 100,000 - 135,000
401(k) match
Health Insurance
Paid vacation
Hybrid AI/HPC Systems Administrator - GPU & Cloud
Hybrid AI/HPC Systems Administrator - GPU & Cloud

Eli Lilly and Company • South San Francisco (CA)

Hybrid
USD 141,000 - 231,000
Hybrid work schedule
Comprehensive benefits
HPC AI Systems Administrator
HPC AI Systems Administrator

MRE Consulting • Houston (TX)

On-site
USD 120,000 - 180,000
Competitive salary
Comprehensive benefits
Professional development support
AI/HPC Infrastructure Engineer: GPU Compute & Hybrid Cloud
AI/HPC Infrastructure Engineer: GPU Compute & Hybrid Cloud

Saigepartners • San Jose (CA)

Hybrid
USD 120,000 - 180,000