Infra Engineer, Cloud & AI Agent Platform

Agentrys

San Jose (CA)

Hybrid

USD 140,000 - 180,000

Full time

43 hours ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Agentrys is seeking an Infrastructure Engineer to design and operate the platform behind our agentic design‑automation systems, spanning multi‑cloud infrastructure and on‑prem, private‑cloud, and air‑gapped environments.

You will own the Kubernetes/container foundation, build enterprise security end‑to‑end, and scale AI/ML infrastructure, GPU clusters, and data pipelines to support diverse customer deployments.

Qualifications

  • Experience operating production infrastructure on a major cloud (GCP, AWS, or Azure).
  • Expert Kubernetes and container skills (Docker/OCI, Helm).
  • Proven track record delivering software into on‑prem, private‑cloud, or air‑gapped environments.
  • Strong enterprise security—RBAC/ReBAC, identity/SSO, secrets, and audit.
  • Proficiency in Go, Node, Python, or Rust and infrastructure‑as‑code (Terraform).
  • CI/CD and artifact/release management experience.

Responsibilities

  • Architect and operate multi‑cloud infra across GCP, AWS, and Azure.
  • Design on‑prem/private‑cloud deployments, including air‑gapped environments.
  • Own Kubernetes/container foundation: cluster, autoscaling, multi‑tenancy, upgrades.
  • Build enterprise access and security end‑to‑end: RBAC/ReBAC, SSO, secrets, audit.
  • Develop compute infra for EDA workloads within the Agentrys fabric.
  • Scale AI/ML infra: GPU clusters, scheduling, model serving & env management.
  • Create data/artifact pipelines and release processes for cloud/on‑prem/a/gapped customers.

Skills

Cloud infra ops
Kubernetes
Containerization
On‑prem/private‑cloud
RBAC/ReBAC security
Go/Node/Python/Rust
Terraform
CI/CD

Tools

Docker
Helm
Terraform

Job description

About the Role

Agentrys seeks an Infrastructure Engineer to design and operate the platform that runs our agentic design-automation systems — both on our own multi-cloud infrastructure and inside our customers' on-prem, private-cloud, and air-gapped environments. This position spans cloud, Kubernetes and containers, enterprise access and security, the compute fabric that runs EDA tools in our on-premises deployment, and the AI/ML and data infrastructure — GPU clusters, data pipelines, and artifact delivery — behind our agents. You'll work on how the platform is built, secured, packaged, and shipped so it runs reliably everywhere our customers do.

Key Responsibilities

The role spans multiple technical areas including:

  • Architecting and operating our internal multi-cloud infrastructure across GCP, AWS, and Azure — provisioning, networking, infrastructure, observability, reliability, and cost.
  • Designing and delivering on-prem and private-cloud deployments — including air-gapped environments — packaged to drop cleanly into each customer's existing infrastructure.
  • Owning the Kubernetes and container foundation: cluster architecture, Helm/packaging, autoscaling, multi-tenancy, and safe lifecycle and upgrades across every cloud and on-prem target.
  • Building enterprise access and security end to end — RBAC and ReBAC authorization, SSO/identity integration, secrets management, and audit.
  • Building the compute infrastructure that runs EDA tools under the Agentrys compute fabric — scheduling, isolation, and resource management for licensed EDA workloads that fit and federate into diverse customer environments.
  • Standing up and scaling AI/ML infrastructure — GPU clusters and scheduling, distributed training, model serving and inference, and model/environment management.
  • Building the data and artifact layer — data pipelines and storage, artifact and model repository management, and the release/“ship” pipeline that packages, signs, and distributes builds and models to cloud, on-prem, and air-gapped customers.
Required Qualifications
  • Deep experience operating production infrastructure on a major cloud (GCP, AWS, or Azure), with working knowledge of more than one.
  • Expert-level Kubernetes and container skills (Docker/OCI, Helm) — cluster operations, workload isolation, and multi-tenancy.
  • Proven track record delivering software into on-prem, private-cloud, or air-gapped customer environments.
  • Strong grasp of authorization and enterprise security — RBAC/ReBAC, identity/SSO, secrets, and audit.
  • Proficiency in a systems/automation language (Go, Node, Python, or Rust) and infrastructure-as-code (e.g., Terraform).
  • CI/CD and artifact/release management experience — build pipelines, registries, signing, and distribution.
Particularly Valuable Experience
  • GPU cluster operations and ML infrastructure (Kubernetes device plugins or Slurm, distributed training, high-throughput inference serving).
  • EDA / HPC / licensed-tool compute environments and schedulers (LSF, SGE, Slurm).
  • Fine-grained / ReBAC authorization systems (Zanzibar-style, e.g., OpenFGA or SpiceDB).
  • Data pipeline / data-platform work (orchestration, lineage) and distributing large model/artifact bundles to air-gapped sites.
  • Building portable, packaged deployments — Helm charts, operators, offline install bundles — across heterogeneous customer infrastructure.
  • Enterprise security & compliance (SOC 2, supply-chain/SBOM, artifact signing).
Why Agentrys

At Agentrys, you will have the opportunity to:

  • Help define a new category of semiconductor design technology.
  • Invent the agent-native algorithms and tools that will form the foundation of future automated design workflows.
  • Develop GPU-accelerated algorithms that make previously impractical design and optimization workflows possible.
  • Build AI systems that perform complex, consequential engineering work—not just generate recommendations.
  • Work with real semiconductor workflows, tools, and private engineering knowledge.
  • See your research deployed directly with leading chip-design organizations.
  • Work in a small, highly technical team where individual contributions can shape the product and company.
  • Collaborate with colleagues across San Jose, Austin, and Taiwan.
  • Change how chips are designed, rather than focus on only one design or one point tool.

Agentrys is an equal opportunity employer. We welcome candidates from diverse backgrounds who are excited to combine ambitious research with meaningful engineering impact.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Research Engineer, AI for Chip Design
Research Engineer, AI for Chip Design

Agentrys • San Jose (CA), Northern (KY)

Hybrid
USD 150,000 - 230,000
AI Engineer, Model Training, Inference & Infra
AI Engineer, Model Training, Inference & Infra

Agentrys • San Jose (CA), Northern (KY)

Hybrid
USD 170,000 - 210,000
Solutions Engineer, AI for Chip Design
Solutions Engineer, AI for Chip Design

Agentrys • San Jose (CA), Northern (KY)

Hybrid
USD 110,000 - 170,000
Applied AI Engineer, Silicon Engineering
Applied AI Engineer, Silicon Engineering

The Consensus • San Jose (CA)

On-site
USD 180,000 - 240,000
Full medical, dental, and vision
Housing subsidy $2,000/month
Daily lunch and dinner in office
+2
Senior Design Engineer
Senior Design Engineer

ChipAgents • San Jose (CA), Northern (KY)

Hybrid
USD 180,000 - 350,000
Salary + equity
Unlimited PTO
Full benefits
+1
Research Scientist - AI for Electronic Design Automation
Research Scientist - AI for Electronic Design Automation

ScOp Venture Capital • Santa Clara (CA)

On-site
USD 150,000 - 230,000
AI Research Engineer for Agentic Chip Design
AI Research Engineer for Agentic Chip Design

Agentrys • San Jose (CA), Northern (KY)

Hybrid
USD 150,000 - 230,000
Staff Chip Engineer- Agentic Workflow
Staff Chip Engineer- Agentic Workflow

Yugal Tech Academy • Redwood City (CA)

On-site
USD 150,000 - 200,000
Applied AI Engineer, Silicon Engineering
Applied AI Engineer, Silicon Engineering

Etched.ai, Inc. • San Jose (CA)

On-site
USD 1,000 - 2,000
Full medical, dental, and vision packages
Housing subsidy of $2,000/month
Daily lunch and dinner in the office
+2
Applied AI Engineer, Silicon Engineering
Applied AI Engineer, Silicon Engineering

Etched • San Jose (CA)

On-site
USD 130,000
Full medical, dental, and vision coverage
Housing subsidy of $2,000/month
Daily lunch and dinner in the office
+2