Cloud / DevOps Engineer (Infra & IaC)

Weekday 1

United States

On-site

USD 214,906,000 - 315,195,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Weekday 1 is seeking an experienced Cloud / DevOps Engineer (Infra & IaC) to contribute to a cutting-edge GenAI environment. The role focuses on building and improving large-scale AI training and inference infrastructure with Kubernetes, AWS, and IaC tooling.

You will drive realistic tasks, reference solutions, and evaluation frameworks for production-grade cloud infrastructure, CI/CD pipelines, and reliability practices, ensuring high-quality outputs for AI systems.

Qualifications

  • 4+ years of professional experience in Cloud Infrastructure, DevOps, SRE, Platform Engineering, or related field.
  • Hands-on experience managing Kubernetes in production environments and diagnosing outages.
  • Proven production experience with Infrastructure-as-Code (Terraform and/or AWS CDK).
  • Strong knowledge of AWS services and production integrations (Lambda, API Gateway, DynamoDB).
  • Experience designing and maintaining CI/CD pipelines for production workloads.

Responsibilities

  • Collaborate with research and engineering teams to identify knowledge gaps and improve AI model performance across cloud infrastructure, DevOps, Kubernetes, and IaC domains.
  • Design tasks covering Kubernetes troubleshooting, AWS service integration, infrastructure automation, and production operations.
  • Develop reference solutions for complex infrastructure engineering scenarios.
  • Review AI-generated technical solutions for correctness, reliability, and production readiness.
  • Provide structured feedback highlighting gaps and opportunities for improvement.
  • Create evaluation criteria, rubrics, and benchmarks for Kubernetes, IaC, and CI/CD reasoning.
  • Develop scenarios involving cluster failures and deployment workflows.
  • Maintain consistency and quality across evaluation datasets with SMEs.
  • Translate production experience into structured guidance to improve AI models.

Skills

Kubernetes in prod
Terraform
AWS CDK
AWS cloud services
CI/CD
Observability

Tools

Terraform
AWS CDK

Job description

This role is for one of our clients

Compensation: $75 - $110 per hour

We are seeking an experienced Cloud / DevOps Engineer (Infra & IaC) to contribute to a cutting-edge GenAI environment focused on building and improving large-scale AI training and inference infrastructure.

The ideal candidate will bring strong, hands‑on expertise in Kubernetes, AWS cloud services, Infrastructure-as-Code (IaC), and CI/CD. You will apply your real-world infrastructure engineering experience to evaluate technical workflows, create high-quality reference solutions, identify gaps in AI-generated outputs, and help establish rigorous standards for cloud and DevOps reasoning.

This is a full‑time engagement requiring 40 hours per week, Monday through Friday.

Requirements

Key Responsibilities
  • Collaborate with research and engineering teams to identify knowledge gaps and improve AI model performance across cloud infrastructure, DevOps, Kubernetes, and Infrastructure-as-Code domains.
  • Design realistic and technically challenging tasks covering Kubernetes troubleshooting, AWS service integration, infrastructure automation, and production operations.
  • Develop accurate, detailed reference solutions for complex infrastructure engineering scenarios.
  • Review and evaluate AI‑generated technical solutions for correctness, reliability, scalability, security, and adherence to production best practices.
  • Provide clear, structured written feedback highlighting technical gaps, incorrect assumptions, and opportunities for improvement.
  • Create detailed evaluation criteria, rubrics, and benchmarks for assessing Kubernetes troubleshooting, IaC architecture, AWS integrations, and CI/CD reasoning.
  • Develop scenarios involving cluster failures, infrastructure automation, deployment workflows, service integrations, and operational reliability.
  • Work closely with other technical subject matter experts to maintain consistency, accuracy, and quality across evaluation datasets.
  • Translate practical production experience into structured guidance that can be used to improve AI‑generated infrastructure solutions.
Core Qualifications
  • 4+ years of professional experience in Cloud Infrastructure, DevOps, Site Reliability Engineering, Platform Engineering, or a closely related field.
  • Strong hands‑on experience managing Kubernetes in production environments, including diagnosing, troubleshooting, and resolving cluster failures and operational issues.
  • Experience with Kubernetes beyond simply writing manifests or consuming managed Kubernetes control planes.
  • Proven production experience with Infrastructure-as-Code, particularly Terraform and/or AWS CDK.
  • Strong practical knowledge of AWS cloud services, including production integration with services such as:
    • AWS Lambda
    • API Gateway
    • DynamoDB
  • Experience designing, implementing, and maintaining CI/CD pipelines for production workloads.
  • Strong understanding of cloud architecture, infrastructure automation, deployment strategies, observability, reliability, and operational best practices.
  • Demonstrated career progression with increasing ownership and responsibility in infrastructure, DevOps, or platform engineering.
  • Ability to commit reliably to 40 hours per week during standard weekdays.
  • Excellent written and verbal communication skills, with the ability to explain complex technical concepts and engineering decisions clearly.
  • Strong analytical and troubleshooting abilities, particularly when diagnosing distributed systems and infrastructure failures.
Preferred Skills
  • Experience working with large‑scale cloud infrastructure or highly distributed systems.
  • Familiarity with Kubernetes networking, security, storage, scaling, and cluster lifecycle management.
  • Experience implementing infrastructure security and reliability best practices.
  • Knowledge of AWS architecture patterns and cloud‑native application design.
  • Experience with GitOps, containerization, monitoring, logging, and observability platforms.
  • Familiarity with modern DevOps and platform engineering methodologies.
  • Experience reviewing or evaluating technical documentation, engineering solutions, or AI‑generated outputs.
What You’ll Contribute

In this role, your production infrastructure expertise will help establish high‑quality standards for AI systems working with complex Cloud, DevOps, Kubernetes, AWS, and IaC problems.

You will play a key role in transforming practical engineering knowledge into structured tasks, reference solutions, evaluation frameworks, and high‑quality technical feedback that can improve the capabilities of next‑generation AI models.

Equal Opportunity

We are committed to providing equal employment opportunities to all qualified candidates. Employment decisions are made without regard to legally protected characteristics, and reasonable accommodations are available throughout the hiring and engagement process upon request.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Cloud / DevOps Engineer (Infra & IaC)
Cloud / DevOps Engineer (Infra & IaC)

Weekday AI (YC W21) • United States

On-site
USD 103,000 - 152,000
Cloud / DevOps Engineer (Infra & IaC)
Cloud / DevOps Engineer (Infra & IaC)

HumanitApp • Northern (KY)

On-site
USD 103,000 - 152,000
Senior Cloud & DevOps Engineer – Infra & IaC
Senior Cloud & DevOps Engineer – Infra & IaC

Weekday AI (YC W21) • United States

On-site
USD 103,000 - 152,000
Cloud DevOps Engineer - AI/ML
Cloud DevOps Engineer - AI/ML

Mercor • New York (NY)

On-site
USD 120,000 - 180,000
Cloud DevOps Engineer - AI/ML
Cloud DevOps Engineer - AI/ML

Obsidian • New York (NY)

On-site
USD 130,000 - 180,000
GenAI Cloud & DevOps Engineer
GenAI Cloud & DevOps Engineer

Weekday 1 • United States

On-site
USD 214,906,000 - 315,195,000
Infrastructure Specialist
Infrastructure Specialist

Obsidian • Miami (FL)

On-site
USD 140,000 - 180,000
Infrastructure Specialist
Infrastructure Specialist

Mercor • Miami (FL)

On-site
USD 120,000 - 170,000
GenAI Cloud DevOps Engineer | Kubernetes & AWS Expert
GenAI Cloud DevOps Engineer | Kubernetes & AWS Expert

Obsidian • New York (NY)

On-site
USD 130,000 - 180,000
Remote | DevOps Engineer — $50–$100/hour
Remote | DevOps Engineer — $50–$100/hour

engineeringjobs.net, Inc. • New York (NY)

Remote
USD 69,000 - 138,000