About the Role
Our Platform Engineering team builds and operates the core backend systems that power CoCounsel's AI agent platform - the infrastructure that lets legal AI agents run reliably at scale. We own the services that sit between product-facing chat/agent experiences and the AI runtime layer, the CI/CD and deployment infrastructure that ships them safely, and the internal developer tooling that other engineering teams build on. We work closely with product, applied-AI, and platform teams to ship AI-driven capabilities faster and more reliably across the business.
As an Engineering Manager , you will lead a team of platform engineers, owning both their technical output and their growth - staying close enough to the architecture and code to make sound calls, while driving the team's roadmap and priorities in partnership with engineering leadership.
What you will do
- Own people-management responsibilities for a team of platform engineers - hiring, performance management, career development, and workload/priority balancing - while remaining technically engaged in architecture and code-level decisions.
- Set and communicate team-level technical strategy and roadmap in partnership with engineering leadership, translating business priorities into an execution plan.
- Guide the design and delivery of backend and platform services that power AI agent workflows - request/event ingestion, agent orchestration, document and file handling - ensuring reliability and delivery speed for product teams building on the platform.
- Steer the evolution of the CI/CD and progressive-delivery infrastructure that lets both customer-facing product systems and internal developer tooling ship safely and continuously - including release-decoupling work that separates code deployment from customer-facing exposure.
- Ensure the team owns and evolves the cloud infrastructure the platform runs on - provisioning, capacity planning, and environment configuration - using Infrastructure-as-Code practices.
- Champion sound backend architecture with a focus on scalability, reliability, and maintainability across a microservices/event-driven system, influencing long-term platform direction.
- Drive cross-functional technical initiatives with product, applied-AI, and infrastructure teams, aligning priorities and de-risking complex, multi-system projects from design through production.
- Raise the bar on engineering practices - code quality, automated testing, observability, CI/CD, and operational excellence - across the platform stack.
- Champion observability across the platform - metrics, logging, and tracing - ensuring your team and downstream teams have the visibility needed to operate complex systems with confidence.
- Ensure strong operational posture for the team's services - monitoring, incident response, and root-cause analysis for production systems running on managed cloud AI runtimes ( e.g. AWS Bedrock AgentCore ) and Kubernetes.
- Ensure the team maintains a healthy on-call rotation supporting our customer-facing systems and critical internal tooling, and participate in it yourself - treating every incident as an opportunity to protect the customer experience.
- Grow technical leadership within the team, fostering a culture of learning, ownership, and continuous improvement.
About you
- 6+ years of professional software engineering experience, including prior technical leadership or people-management experience, with a track record designing, building, and operating large-scale backend systems in production.
- Demonstrated experience leading a team directly - 1:1s, performance reviews, hiring, career development - while staying credible at the code/architecture level.
- Strong proficiency in Python ( FastAPI or similar) or another backend language, with deep experience in distributed systems, microservices, and cloud-native development.
- Hands-on expertise with relational databases (PostgreSQL), caching/messaging systems (Redis), API design, and AWS, including container orchestration with Kubernetes (EKS).
- Experience with CI/CD pipeline engineering and progressive/safe delivery patterns, and release engineering (decoupling deployment from release exposure).
- Infrastructure-as-Code practices for managing cloud infrastructure.
- Experience with observability tooling ( e.g. OpenTelemetry , Datadog) for building and operating production systems.
- Experience operating systems built on or integrating with LLM/Generative AI infrastructure is a strong plus.
- Demonstrated strength in system design, debugging, and performance optimization, with the ability to make thoughtful trade-offs between speed, quality, and long-term scalability.
Nice to Have
- Proven ability to influence technical direction without direct authority, and to communicate clearly with technical and non-technical audiences and leadership.
- Hands-on experience with AWS CDK specifically is a nice -to-have.