System Software EngineerNewHybrid

SpringTime Ventures

Albuquerque (NM)

On-site

USD 90,000 - 120,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

SpringTime Ventures is looking for a System Software Engineer based in Albuquerque, New Mexico, to help build and operate their multi-cloud computational platform. This role focuses on implementing and automating production systems, working closely with senior engineers.

The ideal candidate will have expertise in Kubernetes, cloud infrastructure, and observability practices, alongside experience in development and systems operation.

Join SpringTime and influence the growth trajectory of cutting-edge AI solutions.

Qualifications

  • Bachelor's degree and 3 years experience in relevant field.
  • Experience in cloud infrastructure and backend engineering roles.
  • Working knowledge of Kubernetes and production workloads.
  • Hands-on experience with cloud providers like AWS or GCP.

Responsibilities

  • Implement and maintain Kubernetes workloads and resources.
  • Deploy and operate model-serving workloads on GPU.
  • Support model training on distributed GPU systems.
  • Build and maintain instrumentation for observability.
  • Improve CI/CD pipelines and automation.

Skills

Kubernetes expertise
Cloud infrastructure experience
Observability data instrumentation
CI/CD pipeline familiarity
Scripting proficiency (Python, Go, Rust, or Bash)

Education

Bachelor's degree in Computer Science or related field

Tools

Prometheus
Grafana
OpenTelemetry
Ansible
Puppet
Helm

Job description

About Hoonify

Hoonify delivers secure, sovereign AI infrastructure designed for the next generation of inference workloads. Powered by TurbOS®, our platform enables organizations and NeoCloud/data center operators to transform CPU/GPU infrastructure into production-ready AI environments—supporting local LLMs, agentic copilots, RAG, and embeddings. We empower teams with robust model lifecycle management, multi-tenant controls, usage metering, and fully auditable operations.

The Role

We are seeking a System Software Engineer to help build, deploy, and operate our multi‑cloud computational platform and model‑serving infrastructure underpinning our AI/ML developer platform. This role focuses on implementation, automation, and day‑to‑day operation of production systems, working under the technical direction of senior engineers and the platform's established architectural patterns.

The successful candidate will deliver well‑engineered, well‑tested infrastructure changes, and grow their depth across Kubernetes, GPU‑backed workloads, observability, and continuous delivery in a production environment.

This role enables meaningful growth in cloud infrastructure, distributed systems, and ML serving. You will work directly with senior engineers on real production systems, receive code and design review on your work, and have a clear path to expand scope and ownership as your experience deepens.

Core Responsibilities
  • Implement and maintain Kubernetes workloads and supporting resources, including manifests, Helm charts, controllers, and configuration for networking, ingress, and storage, following established platform patterns.
  • Deploy and operate model‑serving workloads on GPU and accelerator node pools, including configuring autoscaling policies, resource requests and limits, and tenant‑specific deployment configurations.
  • Support model training and simulation workloads on distributed GPU systems.
  • Build and maintain instrumentation on Prometheus, Grafana, and OpenTelemetry, including authoring dashboards, alerting rules, and trace and metric instrumentation for new services.
  • Implement and improve CI/CD pipelines, including build, test, and deployment automation, and contribute to progressive delivery practices already in use on the platform.
  • Develop and maintain infrastructure‑as‑code modules and automation scripts in support of repeatable, auditable infrastructure changes across cloud environments.
  • Support response to production incidents, execute documented runbooks, and contribute to post‑mortems and follow‑up remediation work.
  • Investigate and resolve issues across the stack, including container, node, network, and accelerator‑level problems, escalating appropriately when scope exceeds the role.
  • Write clear documentation, including runbooks, internal references, and design notes for the changes you ship.
  • Participate in code and design reviews, both as author and reviewer, and incorporate feedback from senior engineers into your work.
Required Qualifications
  • Bachelor's degree in Computer Science, Computer Engineering, or Information Technology, plus three (3) years relevant work experience or equivalent combination of education and relevant experience.
  • Professional experience in cloud infrastructure, DevOps, site reliability, or backend engineering roles involving production system operation.
  • Working knowledge of Kubernetes in a production context, including writing and debugging manifests, understanding core resource types, and operating production workloads.
  • Hands‑on experience with at least one major cloud provider (e.g. AWS, GCP, or OpenStack), including its compute, networking, and identity primitives.
  • Experience instrumenting services and consuming observability data, including writing Prometheus queries, building Grafana dashboards, or working with distributed traces.
  • Familiarity with CI/CD systems and the basic mechanics of automated build, test, and deployment pipelines.
  • Experience in configuration management and infrastructure as code tools (e.g. Ansible, Puppet, and Helm).
  • Proficiency in at least one programming or scripting language used for infrastructure work (Python, Go, Rust, or Bash).
  • Comfort working in a Linux environment and with standard developer tooling, including Git‑based workflows.
Preferred Qualifications
  • Exposure to GPU or accelerator workloads in any production or research context.
  • Experience working in a multi‑cloud environment, or strong interest in developing it.
  • Exposure to model‑serving runtimes such as vLLM, or NVIDIA Triton.
  • Hands‑on experience compiling, packaging, and configuring Linux software for integration into a target system.
  • Experience with HPC batch schedulers and MPI based workloads.
Why Join Hoonify

You’ll have a direct line to leadership and genuine influence over the company’s growth trajectory. This is a rare opportunity to build a cutting‑edge multi‑cloud computational platform at a company doing meaningful work in AI — with the autonomy and resources to make it your own.

Hoonify is an equal opportunity employer. We welcome applicants from all backgrounds and are committed to building a diverse and inclusive team. Must be eligible to obtain and maintain a US government security clearance.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Software Engineer, Infrastructure Software for AI (Centralized AI Data Centers & Distrib[...]
Senior Software Engineer, Infrastructure Software for AI (Centralized AI Data Centers & Distrib[...]

Intelliswift - An LTTS Company • Sunnyvale (CA)

On-site
USD 120,000 - 150,000
Competitive salary
Health insurance
Flexible work hours
Staff Software Engineer (Platform/Infrastructure)
Staff Software Engineer (Platform/Infrastructure)

Hyperbolic • San Francisco (CA)

On-site
USD 130,000 - 170,000
GPU Cloud Platform Engineer
GPU Cloud Platform Engineer

Yotta Labs • United States

Remote
USD 120,000 - 160,000
Flexible remote work environment
Innovative team collaboration
Cutting-edge technology challenges
Hyperbolic Labs - Senior Site Reliability Engineer
Hyperbolic Labs - Senior Site Reliability Engineer

deCircle • San Francisco (CA)

On-site
USD 120,000 - 150,000
Senior GPU Infrastructure Engineer
Senior GPU Infrastructure Engineer

Hyperbolic • San Francisco (CA)

On-site
Staff HPC Engineer
Staff HPC Engineer

Biohub • San Francisco (CA)

Hybrid
USD 214,000 - 268,000
401(k) employer match
Paid volunteer time off
Relocation support
HPC AI Systems Administrator
HPC AI Systems Administrator

MRE Consulting • Houston (TX)

On-site
USD 95,000 - 140,000
Software Engineer, Compute Foundations Systems
Software Engineer, Compute Foundations Systems

OpenAI • San Francisco (CA)

On-site
USD 150,000 - 190,000
Member of Technical Staff - GPU Infrastructure
Member of Technical Staff - GPU Infrastructure

Prime Intellect • United States

On-site
USD 120,000 - 150,000
Platform Engineer - AI/ML Infrastructure (Kubernetes & Terraform)
Platform Engineer - AI/ML Infrastructure (Kubernetes & Terraform)

Madrona Venture Labs • United States

Hybrid
USD 180,000 - 260,000