We are seeking a Senior Software Engineer, AI Platform Engineering for a 6–12 month contract position in Manhattan, NY (hybrid). You will join a team building AI-powered platforms that enable legal, tax, and government professionals to research, analyze, and make faster decisions. The AI Platform Engineering team develops cloud-native platforms, APIs, developer tools, and infrastructure that power reliable, compliant, and scalable AI solutions across the organization.
In this role, you will focus on the reliability and resilience of mission-critical AI platforms. You will design systems that withstand provider and dependency failures, establish observability and service-level objectives, and keep the platform current as the AI ecosystem evolves. You will work on Kubernetes-based platform-as-a-service frameworks, model-agnostic AI and LLM gateways, hybrid networking, observability, and self-service developer tooling. We are looking for a collaborative, self-motivated engineer who is comfortable with ambiguity, takes ownership, and enjoys solving complex infrastructure challenges.
Responsibilities
- Improve platform reliability and resilience by designing for failure, defining and meeting SLOs, leading incident response, and reducing operational toil.
- Design, build, and operate Kubernetes-based PaaS frameworks, AI and LLM gateways, APIs, and self-service tools for AI applications.
- Develop model-agnostic gateway capabilities for providers such as OpenAI, Anthropic, Gemini, and AWS Bedrock, including routing, fallback, retries, rate limiting, and cost controls.
- Build observability systems covering metrics, logs, traces, dashboards, and alerting to detect and resolve issues before they affect clients.
- Develop networking solutions that connect applications across public-cloud and on-premises environments.
- Provision and manage cloud infrastructure using Terraform and modern software engineering practices.
- Keep platforms secure and current through dependency patching, runtime upgrades, migrations, and provider-integration updates.
- Create frameworks, templates, and workflows that improve developer productivity and reduce operational overhead.
- Evaluate emerging AI technologies and adapt the platform to support new development patterns and use cases.
Qualifications
- Required
- 6+ years of professional software engineering experience.
- Strong Python skills and experience developing production-grade backend services and APIs.
- Experience designing and operating distributed systems in public-cloud environments, with a strong understanding of failure modes and resilient design patterns.
- Hands-on AWS experience, including services such as EC2, S3, IAM, and container-based workloads.
- Experience with Infrastructure as Code, preferably Terraform.
- Experience with production operations, including metrics, logging, tracing, alerting, SLOs, and incident response.
- Strong knowledge of software architecture, databases, networking, cloud infrastructure, and modern application development.
- A degree in computer science, engineering, or a related field, or equivalent practical experience.
- Preferred
- Java experience.
- Experience building or operating API gateways, LLM gateways, or similar proxy layers with routing, fallback, rate limiting, caching, and cost tracking.
- Experience with OpenTelemetry, Prometheus, Grafana, Datadog, or similar observability tools.}