We are seeking aSenior Software Engineer to help build the foundational platforms that power enterprise AI products used by legal, tax, and government professionals.
As a member of the AI Platform Engineering team, you will design, build, and operate cloud-native platforms, APIs, developer tools, and infrastructure that enable teams to deliver reliable, compliant, and scalable AI solutions.
This role focuses heavily on the reliability and resilience of mission-critical AI platforms. You will design systems that can withstand provider and dependency failures, establish observability and service-level objectives, and keep the platform secure and current as the AI ecosystem evolves.
The platform includes Kubernetes-based platform-as-a-service frameworks, model-agnostic AI/LLM gateways, hybrid networking, observability, and self-service developer tooling. We are looking for a collaborative and self-motivated engineer who is comfortable navigating ambiguity, takes ownership, and can independently lead complex technical initiatives.
Key Responsibilities
- Improve platform reliability and resilience by designing for failure, defining and meeting SLOs, leading incident response, and reducing operational toil.
- Design, build, and operate Kubernetes-based PaaS frameworks, AI/LLM gateways, APIs, and self-service tools for enterprise AI applications.
- Develop model-agnostic gateway capabilities for providers such as OpenAI, Anthropic, Gemini, and AWS Bedrock.
- Implement intelligent routing, fallback mechanisms, retries, rate limiting, caching, and cost controls across multiple AI providers.
- Build observability solutions covering metrics, logs, distributed traces, dashboards, alerting, and proactive issue detection.
- Develop secure networking solutions connecting applications across public-cloud and on-premises environments.
- Provision and manage AWS infrastructure using Terraform and modern infrastructure-as-code practices.
- Maintain platform security and stability through dependency patching, runtime upgrades, migrations, and provider-integration updates.
- Create reusable frameworks, templates, tools, and workflows that improve developer productivity and reduce operational overhead.
- Evaluate emerging AI technologies and adapt the platform to support new development patterns and business use cases.
- Independently lead large technical initiatives from initial concept through production implementation.
Required Qualifications
- At least six years of professional software engineering experience.
- Strong Python development skills, including experience building production-grade backend services and APIs.
- Experience designing and operating distributed systems in public-cloud environments.
- Strong understanding of system failure modes, fault tolerance, high availability, and resilient design patterns.
- Hands‑on AWS experience, including EC2, S3, IAM, and container‑based workloads.
- Experience with Infrastructure as Code, preferably Terraform.
- Experience supporting production systems using metrics, logging, tracing, dashboards, alerting, SLOs, and incident‑response practices.
- Strong knowledge of software architecture, databases, networking, cloud infrastructure, and modern application development.
- Demonstrated ability to take ownership and independently lead complex technical projects.
- Bachelor’s degree in computer science, engineering, or a related field, or equivalent practical experience.
Preferred Qualifications
- Experience building or operating API gateways, LLM gateways, proxy services, or similar platform layers.
- Experience implementing routing, provider fallback, retries, rate limiting, caching, and usage or cost tracking.
- Experience with Kubernetes, preferably Amazon EKS.
- Familiarity with Kubernetes autoscaling technologies such as Karpenter.
- Experience with observability tools such as OpenTelemetry, Prometheus, Grafana, or Datadog.
- Knowledge of AWS networking and security, including VPC, Direct Connect, IAM, and cloud‑security controls.
- Experience with chaos engineering, load testing, failure testing, capacity planning, or disaster recovery.
- Experience developing AI‑powered applications, agent‑based systems, model‑inference services, or AI‑serving platforms.
- Working knowledge of machine‑learning concepts and the ML development lifecycle.
- Experience with AWS SageMaker, AWS Bedrock, PyTorch, TensorFlow, or scikit‑learn.
- Familiarity with AI‑assisted development tools such as Claude Code, Cursor, or GitHub Copilot.
Ideal Candidate
- The ideal candidate is a senior, hands‑on engineer who combines strong Python development skills with deep experience in cloud infrastructure, distributed systems, Kubernetes, and production reliability. You should be comfortable solving complex platform challenges, working independently, and taking large‑scale technical initiatives from architecture and design through production deployment.