Our client, a leading global law firm, is seeking a Platform Engineer to join its technology team in a highly hands‑on role focused on building the reusable platforms, tools, and automation that support modern software and AI workloads across the firm. The successful candidate must be proficient in both Go and Python and will work extensively across CI/CD and pipelines-as-code, infrastructure-as-code, cloud platforms, Kubernetes, observability, test automation, platform security, and SRE practices. This role will build and maintain self‑service developer tooling, automate engineering and operational workflows, and help establish scalable, secure, and reliable engineering standards across the firm. Agile methodology experience is essential, along with a strong background in developer enablement and supporting business‑critical systems. This is an excellent opportunity for an experienced Platform, DevOps, or Infrastructure Engineer to join an innovative and collaborative team working with both traditional and emerging AI infrastructure. New York and Washington, D.C. are the preferred locations, with potential flexibility for candidates in Dallas or Houston. This is a hybrid role and in‑office expectations vary by office.
Technology Environment
- CI/CD, Source Control & Test Automation: GitHub Actions, Azure DevOps, GitLab CI, Jenkins, Git, JFrog Artifactory, Playwright, pytest/JUnit
- Infrastructure & Configuration as Code: Terraform, Ansible, Bicep/ARM, Helm, Kustomize, GitOps using Argo CD and Flux
- Cloud & Orchestration: AWS, Azure, Docker, Kubernetes
- AI Inference & Application Infrastructure: Frontier and open‑weight models through Anthropic, Azure OpenAI, and Amazon Bedrock, model gateways and routing, retrieval and hybrid search, document ingestion, tool/function calling, Model Context Protocol (MCP), and agent orchestration
- AI Evaluation & Quality: Evaluation harnesses and golden datasets, LLM-as-judge and human‑in‑the‑loop review, regression suites, and red‑teaming
- Observability & Monitoring: Prometheus, Grafana, Datadog, Splunk, Elastic/ELK, OpenTelemetry, GenAI tracing, and token, latency, and cost telemetry
- Platform Security & Policy-as-Code: HashiCorp Vault, OPA/Conftest, SAST/DAST
- Developer Portal & Self‑Service: Internal developer portals, CLIs/SDKs, and APIs
Developer Tools & Support
- Develop and maintain self‑service tools, CLIs, libraries, and templates as version‑controlled and tested code to simplify how development teams build, test, and release software
- Contribute enhancements and fixes to the internal developer platform and portal through pull requests and peer code review
- Provide ongoing support to developers using platform services, including triaging and resolving technical requests and issues
- Enable governed self‑service access to AI platform capabilities, including model access, retrieval tools, and evaluation workflows, helping teams move AI solutions from prototype to production using established platform patterns
Build, Release & Testing
- Build and maintain pipelines-as-code using technologies such as GitHub Actions and Azure DevOps YAML while following established engineering patterns
- Support release activities in accordance with defined GitOps and change‑management processes
- Develop and maintain automated testing and quality gates within deployment pipelines
- Implement evaluation‑driven quality gates for AI systems, including evals-as‑code, regression testing against golden datasets, and human‑review thresholds where appropriate
Cloud & Infrastructure
- Develop and maintain modular, tested infrastructure-as-code using technologies such as Terraform and Helm to provision and configure cloud and on‑premises resources
- Deploy and support workloads across AWS, Azure, and Kubernetes environments using GitOps tools such as Argo CD and Flux
- Follow established tagging, configuration, and cost‑management standards when provisioning resources
- Assist with deploying and operating AI infrastructure, including model gateways, retrieval services, orchestration components, and related cloud or Kubernetes resources
Monitoring & Reliability
- Instrument services and establish monitoring, logging, and alerting as code using tools such as Prometheus, Grafana, and OpenTelemetry
- Participate in the team's on‑call rotation, respond to incidents, and assist with restoring services
- Contribute to blameless post‑incident reviews and implement identified remediation actions through code
- Instrument AI services to provide visibility across prompts, tools, and agent workflows while monitoring latency, token consumption, cost, and quality regressions
Security & Automation
- Implement platform security controls and remediate vulnerabilities according to established standards, including policy‑as‑code checks using tools such as OPA and Conftest
- Securely manage secrets, access, and configurations using approved technologies such as Vault
- Apply security and confidentiality safeguards to AI workloads, including protections against prompt injection and data exfiltration, output filtering, and access controls across prompts, retrieval systems, and agent tools
- Automate repetitive operational activities through scripts and lightweight services
Engineering Practices
- Follow established engineering standards for source control, testing, code review, and development workflows
- Work collaboratively with platform and development teams while appropriately escalating complex technical issues
- Develop and maintain runbooks, SOPs, and documentation as code alongside the systems and tools they support
- Work effectively within an Agile engineering environment
Skills & Qualifications
- Proficiency writing production‑quality code in both Go and Python, along with PowerShell and/or Bash scripting experience
- Strong experience using Git, pull requests, peer code review, and automated testing practices
- Working proficiency developing pipelines-as-code with GitHub Actions, Azure DevOps, GitLab CI, or Jenkins
- Experience with infrastructure-as-code technologies such as Terraform, Ansible, and Helm, along with GitOps concepts
- Hands‑on experience with at least one major cloud platform, preferably AWS or Azure
- Experience working with Kubernetes and Docker
- Familiarity with observability technologies such as Grafana, Datadog, Splunk, ELK, and OpenTelemetry
- Understanding of foundational SRE practices
- Exposure to test automation, policy-as-code, and platform security practices
- Familiarity with ITIL practices covering incident, change, and problem management is preferred
- Experience working within Lean or Agile methodologies is preferred
- Relevant certifications such as AWS/Azure Associate or Certified Kubernetes Administrator are preferred
Background
- Several years of experience in platform engineering, DevOps, infrastructure engineering, software engineering, or a related discipline
- Bachelor's degree in Computer Science, Information Technology, Engineering, or a related field, or equivalent professional experience
- Experience supporting business‑critical technology environments, preferably within legal, professional services, or another regulated industry
Expected salary for this exempt role is $140,000 - $180,000, commensurate with experience, training, skills, qualifications, and other market factors.