GPU Cluster SRE: Secure Infra & Automation

Meitner Energy

Buenos Aires

Presencial

ARS 182.877.000 - 274.315.000

Jornada completa

Hace 5 días
Sé de los primeros/as/es en solicitar esta vacante
Generador de candidaturas

Una candidatura hecha para este puesto de trabajo — un currículum y una carta de presentación adaptados que responden directamente a la oferta.

Supera los filtros ATS

Ventajas ofrecidas por este puesto de trabajo

Health insurance
401(k) retirement plan
Employee stock option program
On-site fitness facility

Descripción de la vacante

Meitner Energy seeks a Site Reliability/Infrastructure Engineer to own the GPU cluster and security boundary for its regulated nuclear AI platform. You will operate bare-metal GPU infrastructure, enforce zero-trust access, and manage deployment automation with IaC, staging, and incident response.

You'll collaborate with the AI Platform Lead and corporate IT, build observable systems, and maintain backups, capacity, and cost telemetry. This role is based in Dallas and requires on-site presence.

Formación

  • Bachelor’s degree in computer science, engineering, or a related technical discipline, or equivalent professional experience.
  • 4+ years of professional experience in site reliability engineering, infrastructure engineering, or systems engineering.
  • Strong Linux systems administration, networking fundamentals, and firewall and access control list management, with the ability to diagnose a network problem and trace a storage failure independently.
  • Production container orchestration experience and demonstrated declarative infrastructure-as-code proficiency.
  • Experience operating GPU clusters or comparable high-performance compute infrastructure.
  • Hands‑on experience with production metrics, logging, and alerting stacks, including building alerts that reflect real failure modes.
  • Demonstrated on‑call ownership with incident response and post‑incident review practice, including verified restore testing rather than backup reports alone.
  • Experience implementing network segmentation, default‑deny policies, and privileged access controls in a production environment, and the ability to work on‑site in Dallas.

Responsabilidades

  • Operate and maintain the GPU cluster serving both the regulated and the non-regulated tiers, enforcing default-deny egress, network segmentation, and management-plane protection through zero-trust access patterns, bastion hosts, and privileged access controls.
  • You will own the network architecture that separates export-controlled and otherwise sensitive workloads from general-purpose compute, and ensure those controls satisfy the export-control and controlled-information requirements applicable to advanced nuclear technology, working alongside the AI Platform Lead and compliance stakeholders.
  • You will also partner with corporate IT on shared networking and identity infrastructure while maintaining clean boundaries between platform and corporate systems.
  • Own the infrastructure-as-code for the platform environment, using a declarative and version-controlled toolchain, so that every infrastructure change is version-controlled, reviewed, and reproducible from source.
  • You will build and maintain staging and test environments that mirror production closely enough to catch failures before they reach the live cluster, and implement deployment automation that supports reliable, auditable releases in coordination with the platform engineers who depend on it.
  • Define and enforce backup procedures, restore testing cadence, and recovery objectives for the platform, ensuring recovery procedures are documented and verified rather than assumed.
  • You will own the platform observability stack covering metrics, logs, and alerting, building dashboards and alerts that surface problems before they become incidents, and maintain the capacity and cost telemetry for GPU and storage resources that informs infrastructure investment decisions.
  • You will carry on-call responsibility for the platform infrastructure, respond to incidents with urgency, and conduct blameless post-incident reviews that close the loop on root cause rather than stopping at service restoration.

Conocimientos

Linux administration
Networking fundamentals
On-call incident response
Zero-trust concepts
English proficiency

Educación

Bachelor’s degree in computer science or related field

Herramientas

Kubernetes
Terraform / IaC
GPU clusters
Monitoring/Observability
Zero-trust networking tools

Descripción del empleo

Meitner Energy seeks a Site Reliability/Infrastructure Engineer to own the GPU cluster and security boundary for its regulated nuclear AI platform. You will operate bare-metal GPU infrastructure, enforce zero-trust access, and manage deployment automation with IaC, staging, and incident response.

You'll collaborate with the AI Platform Lead and corporate IT, build observable systems, and maintain backups, capacity, and cost telemetry. This role is based in Dallas and requires on-site presence.

Consigue la evaluación confidencial y gratuita de tu currículum.

o arrastra y suelta tu archivo aquí

Similar jobs

Puestos de trabajo similares que vale la pena comparar

Site Reliability / Infrastructure Engineer
Site Reliability / Infrastructure Engineer

Meitner Energy • Buenos Aires

Presencial
ARS 182.877.000 - 274.315.000
Health insurance
401(k) retirement plan
Employee stock option program
+1
Platform Security Engineer - AI & Export Controls
Platform Security Engineer - AI & Export Controls

Meitner Energy • Buenos Aires

Presencial
ARS 182.877.000 - 243.836.000
Health insurance
401(k) retirement plan
Stock option program
+1
Remote GPU & ML Infrastructure Engineer
Remote GPU & ML Infrastructure Engineer

Svitla Systems, Inc. • Argentina

Híbrido
ARS 182.877.000 - 274.315.000
Remote or office workspace
Technical webinars and meetups
Bonuses for talks and activities
+1
GPU & ML Infrastructure Engineer
GPU & ML Infrastructure Engineer

Svitla Systems, Inc. • Argentina

Híbrido
ARS 182.877.000 - 274.315.000
Remote or office workspace
Technical webinars and meetups
Bonuses for talks and activities
+1
Remote GPU Infra Engineer - Terraform, Go, Linux
Remote GPU Infra Engineer - Terraform, Go, Linux

OpenRelay, Inc. • Argentina

A distancia
ARS 91.051.000 - 136.577.000
Security Engineer, Security Operations and Governance
Security Engineer, Security Operations and Governance

Meitner Energy • Buenos Aires

Presencial
ARS 182.877.000 - 243.836.000
Health insurance
401(k) retirement plan
Stock option program
+1
Senior SRE - AI Reliability & Platform Engineer
Senior SRE - AI Reliability & Platform Engineer

t2s - Group International . your partner in executive search • Argentina

A distancia
ARS 2.400.000 - 4.200.000
Senior SRE: AI-Driven, Scalable Infra (Remote)
Senior SRE: AI-Driven, Scalable Infra (Remote)

Embedded Shishya • Argentina

A distancia
ARS 117.063.521 - 200.680.322
Remote work
Culture of learning
Senior AI Platform Engineer & Architecture Lead
Senior AI Platform Engineer & Architecture Lead

Meitner Energy • Buenos Aires

Presencial
ARS 274.315.000 - 365.753.000
Health insurance
401(k) plan
Stock options
SRE II: Cloud-Native Reliability & Automation Engineer
SRE II: Cloud-Native Reliability & Automation Engineer

Medallia • Buenos Aires

Híbrido
ARS 2.400.000 - 3.600.000