Onemind Services, LLC, is a distinguished MSP, CSP, and ISP, creating its own proprietary products. Our flagship, Cloudmylab, is the world's leading hosted lab solution, built entirely within our data centers. Serving the DoD and Fortune 100 companies, we’re expanding our Public Cloud worldwide, powered by robust, self-owned infrastructure and unparalleled engineering expertise. Dedicated to innovation and security, our team is committed to excellence, providing scalable, efficient, and reliable solutions for digital transformation.
Job Description
This is a remote position.
Responsibilities
- OpenStack Architecture & Platform Engineering
- Design production‑grade OpenStack environments across controller, compute, and storage nodes.
- Architect HA control planes using HAProxy, Keepalived, Galera, and RabbitMQ clustering.
- Build scalable cell‑based Nova architectures.
- Implement multi‑region replication strategies.
- Perform platform capacity modeling and growth forecasting.
- Nova scheduler tuning and filters.
- CPU pinning and isolation.
- NUMA topology alignment.
- Live migrations and evacuations.
- GPU passthrough and SR‑IOV provisioning.
- Hypervisor stack includes KVM, QEMU, Libvirt, and Virt‑IO.
- VXLAN, Geneve, VLAN overlays.
- DVR and L3 routing.
- SR‑IOV and DPDK acceleration.
- Integration with BGP EVPN, MPLS, VRFs, and SD‑WAN.
- Ceph (Primary Requirement)
- RBD block storage.
- CephFS and RGW object storage.
- BlueStore performance tuning.
- NVMe and SSD tiering.
- Additional exposure to Linstor, DRBD, iSCSI, and NVMe‑oF preferred.
- Image & Lifecycle Services
- Glance image pipelines.
- QCOW2 optimization.
- Identity & Access (Keystone)
- RBAC modeling.
- LDAP/AD integration.
- Orchestration & Automation
- Heat orchestration templates.
- Ansible playbooks.
- CI/CD for infrastructure.
- Deployment frameworks include Kolla‑Ansible, OpenStack‑Ansible, TripleO, and MAAS/Juju.
- Helm/Operator‑based deployments.
- Pod and persistent volume troubleshooting.
- Hardware introspection.
- Integration with MAAS/Foreman.
- Observability & Reliability Engineering
- Prometheus and Grafana monitoring.
- Incident response and RCA.
- SLA tracking and alert tuning.
- Upgrade & Lifecycle Management
- Database migrations.
- Zero‑downtime patching.
Requirements
- Required Technical Experience
- 8–12+ years Linux systems engineering.
- 5+ years OpenStack production operations.
- Networking: BGP, VXLAN, EVPN.
- Storage: Ceph production operations.
- Automation: Ansible/Terraform.
- Scripting: Python/Bash.
- Preferred Skills
- Platform9 / Canonical / Red Hat OpenStack.