At Epiq , your work contributes to complex, global legal outcomes. You'll join a values-driven community where integrity guides decisions, relentless service sets the bar, and we thrive on big challenges together. We invest in your growth with enterprise-wide learning and mobility. We celebrate who you are, and we respect life beyond work with flexibility that's recognized externally. Enabled by modern platforms and AI, you'll do the most meaningful work of your career and see your impact at scale.
Job Description
Epiq AI Labs is the innovation and engineering hub behind Epiq's next-generation AI platform for corporate legal departments and global law firms. Operating with the speed and autonomy of a startup and the resources of a global alternative legal services provider, the team builds intelligent agents, reasoning engines, knowledge systems, and structured workflows for litigation, investigations, compliance, and corporate knowledge work.
The team is highly collaborative, deeply technical, and focused on rapid iteration, thoughtful design, and end-to-end ownership.
The Opportunity
You will build the data infrastructure beneath Epiq AI Labs' AI platform, including high-volume document ingestion, distributed and asynchronous processing, and the storage and indexing architecture used by the broader platform. You will own systems from design through production operation.
This work directly supports retrieval, reasoning, and legal workflows by ensuring that data is complete, reliable, and indexed correctly. You will help establish an architecture designed for the next order of magnitude in scale. This role will be in the office 3~4 days a week.
Essential Job Responsibilities
- Design scalable data pipelines for multi-million-file workloads, including parsing, ingestion, storage, partitioning, and indexing strategies.
- Build queue-based asynchronous processing frameworks with fault tolerance, retry, idempotency, dead-letter handling, and elastic capacity.
- Design storage and indexing architecture across relational, object, search, and vector stores, including consistency, backfill, and reconciliation mechanisms.
- Design multi-tenant data architectures with tenant isolation, per-tenant resource controls, and workload management for AI workloads.
- Establish pipeline observability and service-level objectives covering throughput, latency, queue depth, error classification, and data quality.
- Plan and execute modernization and migration strategies that preserve production availability and data integrity.
- Produce technical design documentation, participate in architectural review, collaborate across engineering, product, and security, and strengthen standards for testing, CI, code quality, and operational readiness.
Required Qualifications
- 5+ years of experience designing, building, and operating data-intensive or distributed backend systems in production.
- Demonstrated experience with large-scale data-processing pipelines, including operational responsibility.
- Production experience with message brokers and queue-based architectures, including scaling, delivery and ordering guarantees, and failure handling.
- Strong relational database skills, including query optimization, indexing, partitioning, and zero-downtime schema migration.
- Experience with workflow-orchestration frameworks such as Temporal, Airflow, or Dagster.
- Experience with search or vector-index technologies and maintaining consistency between an index and a system of record.
- Proficiency in Python, or strong proficiency in a comparable language with the ability to work primarily in Python.
- Experience with observability tooling such as Prometheus, Grafana, or OpenTelemetry.
- Experience operating in containerized cloud environments using Azure or AWS, Kubernetes, and Docker.
- Strong system-design and architecture experience, including production of technical design documents.
- Demonstrated proficiency using AI tools in software development.
Preferred Qualifications
- Experience with document-processing technologies, including text extraction, OCR, format conversion, and archive or email extraction.
- Experience modernizing or replacing production systems while maintaining availability.
- Experience with infrastructure-as-code tools, particularly Terraform.
- Experience with multi-region deployments and data-residency or sovereignty requirements.
- Experience with AI/ML data infrastructure and retrieval technologies such as Solr or Qdrant.
- Experience in legal technology or another regulated industry.
Technology Stack
- Python - PostgreSQL - RabbitMQ - Solr - Qdrant - OAuth2 / OIDC / SAML - Azure - Kubernetes - Docker - Terraform - Prometheus - Grafana - OpenTelemetry.
#LI-KS1
The Compensation range for this role is $180,000 -$240,000 USD annually and may be eligible for an annual bonus.
In compliance with federal law, all persons hired will be required to verify identity and eligibility to work in the United States and to complete the required emp