Get more replies from employers
Send a job-specific resume in minutes.
CaseGuild in Seattle is seeking an engineer to own, operate, and improve a major vertical of our distributed infrastructure. You will design systems that handle millions of queries, process terabytes of data, and scale under real-world workloads on shared infrastructure.
This is an end-to-end ownership role: you will write production code, diagnose performance and reliability issues, manage deployments, observability, and incident response, and set patterns others can follow.
Own, operate, and improve a major vertical of CaseGuild’s distributed infrastructure.
CaseGuild runs large-scale services that execute millions of queries, process hundreds of millions of tokens every minute, and ingest, transform, store, and retrieve substantial volumes of structured and unstructured data.
The difficult part is not simply adding capacity. Workloads vary enormously: one customer’s matter may be 1,000 times larger or more demanding than another’s while both run on shared infrastructure. You will design systems that remain fair, isolated, observable, and predictable under contention.
This is an end-to-end ownership role. There is no separate platform, SRE, infrastructure, or database team responsible for finishing the work. When you design a system, you will also own its infrastructure definitions, deployment configuration, production promotion, observability, operational behavior, and incident response.
This is not primarily an architecture or advisory position. You will write production code, investigate performance and reliability problems, operate what you build, and establish technical patterns that other engineers can use.
One customer’s workload cannot degrade another’s. Large jobs are isolated, admission is fair, and tail latency remains predictable under contention.
Critical services have clear ownership, strong observability, understood failure modes, and reliable recovery paths.
The system handles extreme variance in matter size, query patterns, ingestion volume, and processing demand without requiring manual intervention.
Bottlenecks across ingestion, storage, retrieval, orchestration, and AI-processing pipelines are identified and removed.
Infrastructure, application code, deployment configuration, and production operation are treated as one engineering responsibility rather than separate functions.
The engineering team makes better architectural decisions because you contribute both technical leadership and working implementations.
You have built and operated high-volume, low-latency services on shared infrastructure.
Your experience may include:
Distributed systems
Workload isolation and multi-tenancy
Admission control, queuing, scheduling, and backpressure
Tail-latency management
Reliability and failure recovery
Large relational and NoSQL data stores
Production infrastructure and deployment automation
You've rebuilt a system most people are afraid to touch, and you're AI-native enough that if you can dream it, you can drive AI to build it.
That's the job. If it reads like you, apply.