- Design, architect, and operate the Wicked Problems Lab's hybrid AWS and on-premises GPU/AI computing environment
- Scale storage and GPU capacity to meet workload demand
- Engineer resilience through monitoring, backup, and disaster recovery
- Own workload placement across cloud and local systems based on cost, performance, and data sensitivity
- Administer Linux servers end to end
- Manage configuration as code using Terraform, Ansible, and scripting
- Own SSH/key lifecycle and access across a distributed fleet
- Stand up, secure, tune, and back up research databases
- Build and maintain reliable ETL and data-transfer/movement pipelines
- Harden systems through patch/vulnerability management, segmentation, secrets management, and endpoint detection and response
- Enforce access control and sensitive-research data handling requirements
- Serve as the first point of contact for researchers' systems needs and AI/developer tooling
- Configure and troubleshoot routers, switches, VPNs, and network segmentation
- Capture and analyze network traffic at the packet level
- Deploy and tune intrusion detection/prevention systems to monitor and investigate anomalous activity
- Report administratively and functionally to the Director of the Wicked Problems Lab
- Perform other duties as needed
Requirements
- Bachelor's in Computer Science/Engineering or related field is necessary; equivalent experience may substitute
- 4+ years hands-on systems/infrastructure administration is necessary
- Demonstrated command of AWS and strong Linux administration
- Experience operating GPU compute for AI/ML, including CUDA stack and model serving/fine-tuning, is necessary
- Infrastructure-as-code using Terraform/Ansible
- Scripting using Bash/Python
- Database administration with PostgreSQL, NoSQL, or comparable databases
- ETL/data-movement pipelines
- Working systems and network security knowledge
- Relevant AWS/Linux/security certifications preferred
- Experience with enterprise-class NVIDIA GPU systems preferred
- U.S. citizenship and ability to obtain/maintain a U.S. security clearance preferred
Core Competencies
Demonstrates expertise in designing and operating hybrid AWS and on-premises GPU/AI computing environments, with strong skills in Linux administration, infrastructure-as-code, and database management. Proficient in implementing security measures and managing data pipelines to support research initiatives.
Highest-signal resume keywords
- AWS Administration
- Linux Administration
- Infrastructure-As-Code
- Database Administration
- GPU Compute for AI/ML
ATS Optimization Keywords
Hard Skills
- AWS
- Linux
- Terraform
- Ansible
- Bash
- Python
- PostgreSQL
- NoSQL
- ETL
- CUDA
Certifications & Qualifications
- AWS Certification
- Linux Certification
- Security Certification
Industry Keywords
- AI
- Machine Learning
- Data Sensitivity
- Disaster Recovery
- Access Control
Tools & Technologies
- GPU Systems
- Research Databases
- Network Security
- Intrusion Detection/Prevention Systems
- VPNs