Systems Engineer

Jobtailor

Nashville (TN)

Hybrid

USD 130,000 - 185,000

Full time

4 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Wicked Problems Lab in Nashville seeks a seasoned systems/infrastructure administrator to design and operate a hybrid AWS and on‑prem GPU/AI computing environment. You will scale GPU capacity, manage storage, and ensure disaster recovery with robust monitoring and security.

Responsibilities include workload placement, Linux administration, Terraform/Ansible for IaC, and database tuning. You will support researchers with AI tooling and secure data handling in a distributed fleet.

Qualifications

  • Bachelor's in Computer Science/Engineering or related field, or equivalent experience.
  • 4+ years hands-on systems/infrastructure administration is necessary.
  • Demonstrated command of AWS and Linux administration.
  • Experience operating GPU compute for AI/ML, including CUDA and model serving/fine-tuning.
  • Infrastructure-as-code using Terraform/Ansible.
  • Scripting using Bash/Python.
  • Database administration PostgreSQL, NoSQL, or comparable databases.
  • ETL/data-movement pipelines.
  • Working knowledge of infrastructure security.

Responsibilities

  • Design, architect, and operate the lab's hybrid AWS and on-prem GPU/AI computing environment.
  • Scale storage and GPU capacity to meet workload demand.
  • Engineer resilience through monitoring, backup, and disaster recovery.
  • Own workload placement across cloud and local systems based on cost, performance, and data sensitivity.
  • Administer Linux servers end to end.
  • Manage configuration as code using Terraform and Ansible.
  • Own SSH/key lifecycle and access across a distributed fleet.
  • Stand up, secure, tune, and back up research databases.
  • Build and maintain ETL and data-transfer pipelines.
  • Harden systems through patching, secrets management, and endpoint protection.
  • Enforce access control and sensitive-data handling requirements.
  • Serve as first point of contact for researchers' systems needs and AI/developer tooling.
  • Configure and troubleshoot routers, switches, VPNs, and network segmentation.
  • Capture and analyze network traffic; deploy IDS/IPS systems.
  • Report to the Director of the Wicked Problems Lab.

Skills

AWS Administration
Linux Administration
Infrastructure-as-Code
CUDA for AI/ML
Bash
Python
PostgreSQL
NoSQL
Scripting
Disaster Recovery

Education

Bachelor's in Computer Science/Engineering

Tools

Terraform
Ansible
CUDA
PostgreSQL
NoSQL
VPNs

Job description

  • Design, architect, and operate the Wicked Problems Lab's hybrid AWS and on-premises GPU/AI computing environment
  • Scale storage and GPU capacity to meet workload demand
  • Engineer resilience through monitoring, backup, and disaster recovery
  • Own workload placement across cloud and local systems based on cost, performance, and data sensitivity
  • Administer Linux servers end to end
  • Manage configuration as code using Terraform, Ansible, and scripting
  • Own SSH/key lifecycle and access across a distributed fleet
  • Stand up, secure, tune, and back up research databases
  • Build and maintain reliable ETL and data-transfer/movement pipelines
  • Harden systems through patch/vulnerability management, segmentation, secrets management, and endpoint detection and response
  • Enforce access control and sensitive-research data handling requirements
  • Serve as the first point of contact for researchers' systems needs and AI/developer tooling
  • Configure and troubleshoot routers, switches, VPNs, and network segmentation
  • Capture and analyze network traffic at the packet level
  • Deploy and tune intrusion detection/prevention systems to monitor and investigate anomalous activity
  • Report administratively and functionally to the Director of the Wicked Problems Lab
  • Perform other duties as needed
Requirements
  • Bachelor's in Computer Science/Engineering or related field is necessary; equivalent experience may substitute
  • 4+ years hands-on systems/infrastructure administration is necessary
  • Demonstrated command of AWS and strong Linux administration
  • Experience operating GPU compute for AI/ML, including CUDA stack and model serving/fine-tuning, is necessary
  • Infrastructure-as-code using Terraform/Ansible
  • Scripting using Bash/Python
  • Database administration with PostgreSQL, NoSQL, or comparable databases
  • ETL/data-movement pipelines
  • Working systems and network security knowledge
  • Relevant AWS/Linux/security certifications preferred
  • Experience with enterprise-class NVIDIA GPU systems preferred
  • U.S. citizenship and ability to obtain/maintain a U.S. security clearance preferred
Core Competencies

Demonstrates expertise in designing and operating hybrid AWS and on-premises GPU/AI computing environments, with strong skills in Linux administration, infrastructure-as-code, and database management. Proficient in implementing security measures and managing data pipelines to support research initiatives.

Highest-signal resume keywords
  • AWS Administration
  • Linux Administration
  • Infrastructure-As-Code
  • Database Administration
  • GPU Compute for AI/ML
ATS Optimization Keywords
Hard Skills
  • AWS
  • Linux
  • Terraform
  • Ansible
  • Bash
  • Python
  • PostgreSQL
  • NoSQL
  • ETL
  • CUDA
Certifications & Qualifications
  • AWS Certification
  • Linux Certification
  • Security Certification
Industry Keywords
  • AI
  • Machine Learning
  • Data Sensitivity
  • Disaster Recovery
  • Access Control
Tools & Technologies
  • GPU Systems
  • Research Databases
  • Network Security
  • Intrusion Detection/Prevention Systems
  • VPNs
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Member of Technical Staff – AI Cloud Infrastructure
Member of Technical Staff – AI Cloud Infrastructure

Jobtailor • California (MO)

On-site
USD 150,000 - 190,000
Staff Engineer, Senior Manager
Staff Engineer, Senior Manager

Jobtailor • Connecticut

On-site
USD 140,000 - 190,000
Software Engineer – AI & Cloud Engineering
Software Engineer – AI & Cloud Engineering

Jobtailor • Massachusetts

On-site
USD 110,000 - 170,000
Principal Cloud Engineer – AI
Principal Cloud Engineer – AI

Jobtailor • West Chester

On-site
USD 150,000 - 210,000
AI and ML Infra Software Engineer, GPU Clusters
AI and ML Infra Software Engineer, GPU Clusters

Jobtailor • California (MO)

On-site
USD 120,000 - 190,000
Forward Deployed Engineer
Forward Deployed Engineer

Jobtailor • Dallas (TX)

On-site
USD 150,000 - 190,000
Staff Engineer, CI/CD & Cloud Infrastructure
Staff Engineer, CI/CD & Cloud Infrastructure

San Diego Stealth Startup • San Diego (CA)

On-site
USD 175,000 - 185,000
Platform Automation Engineer
Platform Automation Engineer

Jobtailor • Maryland

On-site
USD 120,000 - 180,000
Solution Engineer, Data Engineering – Manager
Solution Engineer, Data Engineering – Manager

Jobtailor • California (MO)

On-site
USD 120,000 - 160,000
Member of Technical Staff, ML Engineer
Member of Technical Staff, ML Engineer

Jobtailor • Boston (MA)

On-site
USD 120,000 - 160,000