Site Reliability / Infrastructure Engineer

Meitner Energy

Buenos Aires

Presencial

ARS 182.877.000 - 274.315.000

Jornada completa

Hace 5 días
Sé de los primeros/as/es en solicitar esta vacante
Generador de candidaturas

Destaca en este puesto — crea un currículum adaptado y una carta de presentación en aproximadamente un minuto.

Supera los filtros ATS

Ventajas ofrecidas por este puesto de trabajo

Health insurance
401(k) retirement plan
Employee stock option program
On-site fitness facility

Descripción de la vacante

Meitner Energy seeks a Site Reliability/Infrastructure Engineer to own the GPU cluster and security boundary for its regulated nuclear AI platform. You will operate bare-metal GPU infrastructure, enforce zero-trust access, and manage deployment automation with IaC, staging, and incident response.

You'll collaborate with the AI Platform Lead and corporate IT, build observable systems, and maintain backups, capacity, and cost telemetry. This role is based in Dallas and requires on-site presence.

Formación

  • Bachelor’s degree in computer science, engineering, or a related technical discipline, or equivalent professional experience.
  • 4+ years of professional experience in site reliability engineering, infrastructure engineering, or systems engineering.
  • Strong Linux systems administration, networking fundamentals, and firewall and access control list management, with the ability to diagnose a network problem and trace a storage failure independently.
  • Production container orchestration experience and demonstrated declarative infrastructure-as-code proficiency.
  • Experience operating GPU clusters or comparable high-performance compute infrastructure.
  • Hands‑on experience with production metrics, logging, and alerting stacks, including building alerts that reflect real failure modes.
  • Demonstrated on‑call ownership with incident response and post‑incident review practice, including verified restore testing rather than backup reports alone.
  • Experience implementing network segmentation, default‑deny policies, and privileged access controls in a production environment, and the ability to work on‑site in Dallas.

Responsabilidades

  • Operate and maintain the GPU cluster serving both the regulated and the non-regulated tiers, enforcing default-deny egress, network segmentation, and management-plane protection through zero-trust access patterns, bastion hosts, and privileged access controls.
  • You will own the network architecture that separates export-controlled and otherwise sensitive workloads from general-purpose compute, and ensure those controls satisfy the export-control and controlled-information requirements applicable to advanced nuclear technology, working alongside the AI Platform Lead and compliance stakeholders.
  • You will also partner with corporate IT on shared networking and identity infrastructure while maintaining clean boundaries between platform and corporate systems.
  • Own the infrastructure-as-code for the platform environment, using a declarative and version-controlled toolchain, so that every infrastructure change is version-controlled, reviewed, and reproducible from source.
  • You will build and maintain staging and test environments that mirror production closely enough to catch failures before they reach the live cluster, and implement deployment automation that supports reliable, auditable releases in coordination with the platform engineers who depend on it.
  • Define and enforce backup procedures, restore testing cadence, and recovery objectives for the platform, ensuring recovery procedures are documented and verified rather than assumed.
  • You will own the platform observability stack covering metrics, logs, and alerting, building dashboards and alerts that surface problems before they become incidents, and maintain the capacity and cost telemetry for GPU and storage resources that informs infrastructure investment decisions.
  • You will carry on-call responsibility for the platform infrastructure, respond to incidents with urgency, and conduct blameless post-incident reviews that close the loop on root cause rather than stopping at service restoration.

Conocimientos

Linux administration
Networking fundamentals
On-call incident response
Zero-trust concepts
English proficiency

Educación

Bachelor’s degree in computer science or related field

Herramientas

Kubernetes
Terraform / IaC
GPU clusters
Monitoring/Observability
Zero-trust networking tools

Descripción del empleo

Meitner Energy is developing advanced nuclear energy solutions intended to deliver reliable, scalable, and carbon-free electricity for industrial and grid applications.

Our international team brings together nuclear, engineering, commercial, regulatory, and project-development experience. We are building an organization focused on disciplined engineering, responsible execution, and the deployment of nuclear energy at meaningful scale.

The Opportunity

Meitner’s AI platform runs on infrastructure that has to be right, and this role owns it. You will operate the bare-metal GPU cluster, the private networking and storage beneath it, and the compliance boundary that keeps regulated nuclear material separated from general-purpose workloads. You will also own the deployment automation and staging environments that other platform engineers rely on to ship reliably, which makes this the layer everything else is built on.

The work rewards someone who treats uptime and security posture as fixed requirements, can implement a default-deny network architecture and explain every exception in it, and has the discipline to automate and document rather than carry the environment in their head. You will work directly with the AI Platform Lead and partner with corporate IT on shared networking and identity where that makes sense.

This role is ideal for someone who:

  • Takes personal ownership of uptime and treats incidents as failures of process rather than bad luck.
  • Builds infrastructure through code and automation instead of one-off manual changes.
  • Is comfortable in regulated, security-conscious environments where the network boundary is a hard requirement.
  • Wants to be the person the rest of the engineering team relies on when the cluster is the constraint.

What You'll Do

Cluster Operations and Security Boundaries

Operate and maintain the GPU cluster serving both the regulated and the non-regulated tiers, enforcing default-deny egress, network segmentation, and management-plane protection through zero-trust access patterns, bastion hosts, and privileged access controls. You will own the network architecture that separates export-controlled and otherwise sensitive workloads from general-purpose compute, and ensure those controls satisfy the export-control and controlled-information requirements applicable to advanced nuclear technology, working alongside the AI Platform Lead and compliance stakeholders. You will also partner with corporate IT on shared networking and identity infrastructure while maintaining clean boundaries between platform and corporate systems.

Infrastructure as Code and Deployment Automation

Own the infrastructure-as-code for the platform environment, using a declarative and version-controlled toolchain, so that every infrastructure change is version-controlled, reviewed, and reproducible from source. You will build and maintain staging and test environments that mirror production closely enough to catch failures before they reach the live cluster, and implement deployment automation that supports reliable, auditable releases in coordination with the platform engineers who depend on it.

Reliability, Observability, and Capacity

Define and enforce backup procedures, restore testing cadence, and recovery objectives for the platform, ensuring recovery procedures are documented and verified rather than assumed. You will own the platform observability stack covering metrics, logs, and alerting, building dashboards and alerts that surface problems before they become incidents, and maintain the capacity and cost telemetry for GPU and storage resources that informs infrastructure investment decisions.

You will carry on-call responsibility for the platform infrastructure, respond to incidents with urgency, and conduct blameless post-incident reviews that close the loop on root cause rather than stopping at service restoration.

What We're Looking For

Required Qualifications

  • Bachelor’s degree in computer science, engineering, or a related technical discipline, or equivalent professional experience.
  • 4+ years of professional experience in site reliability engineering, infrastructure engineering, or systems engineering.
  • Strong Linux systems administration, networking fundamentals, and firewall and access control list management, with the ability to diagnose a network problem and trace a storage failure independently.
  • Production container orchestration experience and demonstrated declarative infrastructure-as-code proficiency.
  • Experience operating GPU clusters or comparable high-performance compute infrastructure.
  • Hands‑on experience with production metrics, logging, and alerting stacks, including building alerts that reflect real failure modes.
  • Demonstrated on‑call ownership with incident response and post‑incident review practice, including verified restore testing rather than backup reports alone.
  • Experience implementing network segmentation, default‑deny policies, and privileged access controls in a production environment, and the ability to work on‑site in Dallas.

Preferred Qualifications

  • Working proficiency in Spanish. Spanish is preferred because the position involves infrastructure collaboration with colleagues in Argentina.
  • Experience implementing zero‑trust network access, bastion, or privileged access management patterns in production.
  • Familiarity with regulated or export‑controlled data environments, including controlled unclassified information handling.
  • Experience with storage systems for high‑throughput machine learning workloads, including high‑speed local storage, distributed file systems, and object storage.
  • Background in nuclear energy, defense, aerospace, or another security‑sensitive industrial sector.

The Candidate We Are Seeking

The strongest candidate will be an engineer who has carried a pager for infrastructure they built themselves and can describe what they learned the hard way. This may be an excellent next step for a site reliability engineer, infrastructure engineer, or systems engineer who wants full ownership of a cluster and its security boundary rather than a share of a large fleet. Candidates should be prepared to discuss the worst outage they owned, not the cleanest one.

Why Join Meitner Energy?

Consequential Work. Build and operate the infrastructure foundation supporting the delivery of reliable, scalable, carbon‑free energy.

Direct Ownership. Own the cluster, the network boundary, and the recovery plan outright rather than filing requests against another team. Meitner is small enough that your decisions visibly shape the company.

International Scope. Work with professionals across the United States, Argentina, and the United Kingdom.

Strong Benefits. Meitner offers comprehensive health insurance, a 401(k) retirement plan, and participation in the company's employee stock option program, subject to plan terms and eligibility. The final offer will reflect the candidate's depth and relevant skills, including infrastructure depth, security design experience, Spanish proficiency, and regulated‑industry background.

High-Quality Workplace. Work from a modern Dallas office designed to support collaboration, productivity, and employee well‑being, including an on‑site fitness facility.

Consigue la evaluación confidencial y gratuita de tu currículum.

o arrastra y suelta tu archivo aquí

Similar jobs

Puestos de trabajo similares que vale la pena comparar

AI Platform Engineer / Solutions Architect
AI Platform Engineer / Solutions Architect

Meitner Energy • Buenos Aires

Presencial
ARS 274.315.000 - 365.753.000
Health insurance
401(k) plan
Stock options
Security Engineer, Security Operations and Governance
Security Engineer, Security Operations and Governance

Meitner Energy • Buenos Aires

Presencial
ARS 182.877.000 - 243.836.000
Health insurance
401(k) retirement plan
Stock option program
+1
GPU Cluster SRE: Secure Infra & Automation
GPU Cluster SRE: Secure Infra & Automation

Meitner Energy • Buenos Aires

Presencial
ARS 182.877.000 - 274.315.000
Health insurance
401(k) retirement plan
Employee stock option program
+1
SSr. Automation, Control Room & HSI Engineer
SSr. Automation, Control Room & HSI Engineer

Meitner Energy • Partido de Vicente López

Presencial
ARS 137.157.000 - 228.596.000
Contracts Specialist
Contracts Specialist

Meitner Energy • Partido de Vicente López

Presencial
ARS 1.800.000 - 2.400.000
AWS Cloud Security Engineer
AWS Cloud Security Engineer

ON.energy • Argentina

Presencial
ARS 136.147.000 - 196.657.000
Senior AWS Cloud Engineer
Senior AWS Cloud Engineer

On.Energy • Buenos Aires

Presencial
ARS 1.800.000 - 3.200.000
Medical, dental, and vision insurance
Professional development opportunities
SSr. Nuclear Instrumentation Engineer
SSr. Nuclear Instrumentation Engineer

Meitner Energy • Partido de Vicente López

Presencial
ARS 2.700.000 - 4.500.000
SSr. Automation and Control Engineer
SSr. Automation and Control Engineer

Meitner Energy • Partido de Vicente López

Presencial
ARS 1.800.000 - 2.400.000
Sr. Safety PSA Engineer
Sr. Safety PSA Engineer

Meitner Energy • Partido de Vicente López

Presencial
ARS 137.157.000 - 198.116.000