Software Engineer, Infrastructure, AI Labs

epiqsystems

New York (NY)

Hybrid

USD 145,000 - 195,000

Full time

6 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Epiq AI Labs seeks an experienced Infrastructure Engineer to build and operate the cloud backbone of its AI platform for legal teams. You will design scalable Terraform-based infrastructure, run Kubernetes in production, and own CI/CD release pipelines, security patterns, and observability.

The role emphasizes end-to-end ownership, collaboration with backend and AI teams, and a hybrid office schedule. You will contribute to reliability, security, and scalable deployment patterns while shaping

Qualifications

  • 3+ years of experience in infrastructure engineering, platform engineering, or site reliability engineering.
  • Production infrastructure experience with on-call responsibility for systems you designed.
  • Hands-on with at least one major cloud platform (AWS, GCP, or Azure).
  • Experience with infrastructure-as-code tools, particularly Terraform, with reusable modules.
  • Production Kubernetes experience including scaling, upgrades, resource limits, network policy, and troubleshooting.
  • Ownership of CI/CD pipelines using GitHub Actions, Azure DevOps, or a comparable platform.
  • Experience with observability tools such as Prometheus, Grafana, or OpenTelemetry and defining SLOs.
  • Incident-command experience in production environments.
  • Experience with secrets and certificate management at organizational scale.
  • Proficiency in Python or Go for tooling and automation.
  • Strong system-design and architectural documentation experience.

Responsibilities

  • Design and implement cloud infrastructure using Terraform, including reusable modules and drift remediation.
  • Operate Kubernetes in production with orchestration, autoscaling, and cluster lifecycle management.
  • Build and maintain CI/CD and release infrastructure with progressive delivery and rollback mechanisms.
  • Define platform SLIs/SLOs and implement metrics, tracing, alerting, and error budgeting practices.
  • Implement security and compliance infrastructure including data residency controls and audit logging.
  • Build and maintain network, secrets, key, credential, TLS, and certificate lifecycle infrastructure.
  • Contribute to incident response, post-incident reviews, and platform readiness.

Skills

Infrastructure engineering
Cloud experience
Terraform
Kubernetes
CI/CD
Observability
Incident command
Secrets management
Python/Go
System design

Tools

Terraform
Kubernetes
Docker
GitHub Actions
Azure DevOps
Prometheus
Grafana
OpenTelemetry
PostgreSQL

Job description

At Epiq , your work contributes to complex, global legal outcomes. You'll join a values‑driven community where integrity guides decisions, relentless service sets the bar, and we thrive on big challenges together. We invest in your growth with enterprise‑wide learning and mobility. We celebrate who you are, and we respect life beyond work with flexibility that's recognized externally. Enabled by modern platforms and AI, you'll do the most meaningful work of your career and see your impact at scale.

Job Description:

Epiq AI Labs is the innovation and engineering hub behind Epiq's next-generation AI platform for corporate legal departments and global law firms. Operating with the speed and autonomy of a startup and the resources of a global alternative legal services provider, the team builds intelligent agents, reasoning engines, knowledge systems, and structured workflows for litigation, investigations, compliance, and corporate knowledge work.

The team is highly collaborative, deeply technical, and focused on rapid iteration, thoughtful design, and end-to-end ownership.

The Opportunity

You will build the cloud infrastructure beneath Epiq AI Labs' AI platform, spanning infrastructure as code, Kubernetes, CI/CD and release engineering, networking, secrets management, observability, security, and compliance infrastructure. Ownership extends from initial design through production operation.

Infrastructure is treated as a product whose reliability, scalability, and security shape the capabilities of the entire platform. You will partner closely with backend engineering, AI engineering, product management, and security to establish deployment, observability, and security patterns that can scale with the organization . This role will be in the office 3~4 days a week.

Essential Job Responsibilities
  • Design and implement cloud infrastructure using Terraform, including reusable modules, environment topology, and drift detection and remediation.
  • Operate Kubernetes in production, including orchestration, autoscaling, resource governance, network policy, and cluster lifecycle management.
  • Build and maintain CI/CD and release infrastructure with progressive delivery, rollback mechanisms, and efficient paths from merge to production.
  • Define platform service-level objectives and build the metrics, tracing, alerting, and error-budget practices required to support them.
  • Implement security and compliance infrastructure, including hardening, audit logging, data-residency controls, retention, legal hold, and audit evidence collection.
  • Build and maintain network, secrets, key, credential, TLS, and certificate-lifecycle infrastructure.
  • Contribute to incident response, post-incident review, developer tooling, technical design documentation, architectural review, and platform operational readiness.
Required Qualifications
  • 3+ years of experience in infrastructure engineering, platform engineering, or site reliability engineering.
  • Demonstrated experience building and operating production infrastructure, including on-call responsibility for systems of your own design.
  • Hands-on experience with at least one major cloud platform such as AWS, GCP, or Azure.
  • Experience with infrastructure-as-code tools, particularly Terraform, including reusable module design.
  • Production Kubernetes experience, including scaling, upgrades, resource limits, network policy, and troubleshooting.
  • Ownership of CI/CD pipelines using GitHub Actions, Azure DevOps, or a comparable platform.
  • Experience with observability tooling such as Prometheus, Grafana, or OpenTelemetry, including defining and maintaining service-level objectives.
  • Demonstrated incident-command experience in production environments.
  • Experience with secrets and certificate management at organizational scale.
  • Proficiency in Python, Go, or a comparable language sufficient to build tooling and automation.
  • Strong system-design and architecture experience, including production of technical design documents.
Technology Stack
  • Azure
  • Terraform
  • Kubernetes
  • Docker
  • CI/CD platform to be confirmed
  • Prometheus
  • Grafana
  • OpenTelemetry
  • PostgreSQL
  • RabbitMQ
  • Python

The Compensation range for this role is $145,000 -$195,000 USD annually and may be eligible for an annual bonus.

In compliance with federal law, all persons hired will be required to verify identity and eligibility to work in the United States and to complete the required employment eligibility verification form upon hire.

Must be authorized to work in the United States for any employer.

Your specific salary will be determined based on several factors:

  • Location-based market rate for the role
  • Your abilities in relation to the job specification
  • Performance during screening and interview
  • Pay parity with the wider team in the considered location

Further details about the package will be provided during the initial screening call with the Talent Acquisition Team.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Software Engineer, Infrastructure, AI Labs
Software Engineer, Infrastructure, AI Labs

Epiq • New York (NY)

On-site
USD 145,000 - 195,000
Software Engineer, Infrastructure, AI Labs
Software Engineer, Infrastructure, AI Labs

Epiq Systems, Inc. • New York (NY)

Hybrid
USD 145,000 - 195,000
Senior Software Engineer, Data Platform, AI Labs
Senior Software Engineer, Data Platform, AI Labs

epiqsystems • New York (NY)

Hybrid
USD 180,000 - 240,000
Software Engineer, Infrastructure, AI Labs
Software Engineer, Infrastructure, AI Labs

Epiq Company • New York (NY)

Hybrid
USD 145,000 - 195,000
Senior Software Engineer, Data Platform, AI Labs
Senior Software Engineer, Data Platform, AI Labs

Epiq • New York (NY)

Hybrid
USD 180,000 - 240,000
Senior Software Engineer, Agentic Platform, AI Labs
Senior Software Engineer, Agentic Platform, AI Labs

Epiq Company • New York (NY)

Hybrid
USD 195,000 - 255,000
Senior Manager, Legal AI Engineer
Senior Manager, Legal AI Engineer

Epiq • New York (NY)

On-site
USD 250,000 - 290,000
Forward Deployed Legal Engineer
Forward Deployed Legal Engineer

Epiq • New York (NY)

On-site
USD 140,000 - 175,000
Senior Director, AI Programmatic Sales
Senior Director, AI Programmatic Sales

Epiq Company • New York (NY)

On-site
USD 180,000 - 240,000
Senior Director, AI Programmatic Sales
Senior Director, AI Programmatic Sales

Epiq Company • Chicago (IL), Northern (KY)

On-site
USD 180,000 - 260,000