Senior Cloud Platform Engineer (AWS), AI Infrastructure - Evinova

AstraZeneca

Gaithersburg (MD)

Hybrid

USD 145,000 - 185,000

Full time

3 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

401(k) plan
Paid vacation
Health benefits

Job summary

AstraZeneca seeks a Senior Cloud Platform Engineer to design, build and operate the AWS-based platform that runs GenAI workloads. You will write infrastructure as code in AWS CDK (TypeScript/Python), create reusable constructs, and enable AI engineers to deploy at scale.

You will collaborate with ML/Data Science teams, implement observability, security, cost controls, and ensure compliance in a regulated environment. Hybrid work from Gaithersburg area is expected.

Qualifications

  • Minimum of 4 years in a platform engineering, SRE or DevOps role where infrastructure code was the main output.
  • Deep hands-on AWS cloud engineering experience including production infrastructure.
  • Strong experience with AWS services (Bedrock, ECS, IAM, VPC, S3, CloudWatch).
  • Experience with IaC: AWS CDK (Python/TypeScript) and Terraform.

Responsibilities

  • Design, build and operate scalable AWS cloud platform capabilities for production ML/AI workloads.
  • Write and maintain AWS IaC, including reusable CDK constructs.
  • Build and operate containerized workloads using ECS and AgentCore.
  • Develop reusable platform capabilities across compute, networking, IAM, secrets, storage and model access.
  • Build CI/CD and GitOps workflows for safe automated deployments.
  • Partner with ML teams to transition prototypes into production services.

Skills

AWS CDK
Python
TypeScript
Docker
ECS
SageMaker
Bedrock
Terraform
IaC

Tools

Terraform
AWS CDK
Docker
Kubernetes (EKS)

Job description

WHY JOIN US?

Evinova is a health-tech business focused on accelerating better health outcomes by advancing digital transformation across the life sciences sector. By combining science-based expertise, evidence-led rigor, and deep human insight, we design digital solutions that enable healthcare to work better for everyone.

Operating at the intersection of healthcare, technology, data, and analytics, we are helping unlock the full potential of digital health, transforming how clinical research is conducted, how care is delivered, and how patients experience healthcare. Our solutions are built to scale, driving efficiency, improving decision-making, and ultimately delivering better outcomes for patients worldwide.

At Evinova, we are driven by a shared purpose to transform health through data and digital innovation. Our teams collaborate across disciplines to solve complex challenges, continuously learning and evolving in a fast-paced, high-impact environment.

We also recognize the importance of flexibility and balance. Our ways of working support both individual needs and team collaboration. To foster connection and collaboration, employees are expected to work from the office three days per week, creating opportunities for in-person teamwork, innovation, and meaningful connection.

This role is located in Gaithersburg, Maryland and follows a hybrid work model. Candidates must reside within commuting distance of Gaithersburg or be willing to relocate for this opportunity.

Introduction to Role

The Machine Learning and Artificial Intelligence Operations team (ML/AI Ops) is the cloud platform engineering team responsible for building and operating the infrastructure that enables our AI Engineers and Data Scientists to deploy Generative AI applications reliably, securely, efficiently and at scale.

As a Senior Cloud Platform Engineer on the ML/AI Ops team, you will design, build and operate the AWS platform that runs our production Generative AI, agentic AI and conversational AI workloads. Most of the code you write will be AWS CDK (TypeScript and/or Python) that creates the infrastructure for other teams to build on.

This is a cloud and platform engineering role rather than an AI application development role. You will partner closely with the engineers and Data Scientists who build agents and models, and provide the infrastructure, deployment patterns, model access, observability and operational capabilities they need to move solutions from experimentation into reliable production environments.

You will work across AWS infrastructure, infrastructure as code (IaC), Amazon Bedrock AgentCore, Amazon ECS, CI/CD, SageMaker Unified Studio, AI gateways, observability, scalability, reliability, security, governance and cost optimization. Your work will establish reusable platform capabilities that allow teams across Evinova to deploy and operate solutions faster and more reliably while meeting the requirements of a highly regulated pharmaceutical environment.

Accountabilities
Cloud Platform Engineering
  • Design, build and operate scalable AWS cloud platform capabilities for production ML/AI and Generative AI workloads.
  • Create reusable infrastructure, tooling and deployment patterns that enable AI Engineers and Data Scientists to independently deploy and operate their applications.
  • Write and maintain AWS IaC primarily AWS CDK in TypeScript and/or Python, including reusable CDK constructs that other teams consume.
  • Build and operate containerized workloads using Amazon ECS and AgentCore.
  • Develop reusable platform capabilities across compute, networking, IAM, secrets management, storage, model access and workload isolation.
  • Build and maintain CI/CD and GitOps workflows that enable safe, automated and repeatable deployments across environments.
  • Partner with engineers and Data Scientists to transition prototypes and research workloads into resilient, production-grade services.
  • Build self-service capabilities and automation that improve developer experience and reduce operational toil.
Reliability, Scalability & Operational Excellence
  • Engineer platform capabilities that improve the availability, scalability, resiliency and performance of production GenAI workloads.
  • Design and implement autoscaling, load balancing, retries, timeouts, fallback, rate limiting and failure-recovery strategies.
  • Establish monitoring, alerting, SLOs, runbooks and production-readiness standards for ML/AI workloads.
  • Troubleshoot complex production issues across AWS infrastructure, container, application and model-provider layers.
  • Automate operational processes and proactively identify opportunities to improve platform reliability and performance.
  • Drive cloud and model cost optimization through capacity management, workload optimization and data-driven analysis.
AI Gateway & Model Access
  • Build and operate a centralized AI gateway that routes requests across model providers.
  • Provide secure, reliable and governed access to foundation models through Amazon Bedrock, OpenAI, Anthropic, Microsoft Foundry, Gemini Enterprise Agent Platform (formerly Vertex AI) and other model platforms.
  • Implement model/provider routing, fallback, authentication, rate limiting, quotas and cost controls.
  • Enable AI teams to evaluate and change model providers without tightly coupling their applications to individual model endpoints.
  • Provide the compute, networking, storage and runtime infrastructure to reliably operate agentic AI, RAG and conversational AI workloads in production.
Observability, Governance & Cost Optimization
  • Build platform-level observability for GenAI workloads, including token consumption, latency, throughput, errors, model/provider performance and cost attribution by team and application.
  • Implement standardized tracing, logging, metrics and alerting capabilities that can be adopted across AI applications.
  • Integrate observability technologies such as Amazon CloudWatch, OpenTelemetry, Datadog and Splunk.
  • Build the telemetry and data pipelines that AI teams use to evaluate and monitor production LLM behavior.
  • Build appropriate security, auditability and governance controls for AI workloads operating within a regulated environment.
  • Support compliance with applicable industry standards and practices, including Good Clinical Practice and Good Machine Learning Practice and other GxP related standards and practices.
Representative Projects
  • Build a library of AWS CDK constructs that gives an AI team a production-ready Amazon ECS service, with IAM, secrets, networking and observability, from a single import.
  • Deploy and operate a centralized AI gateway on Amazon ECS, with multi-provider routing, fallback, quotas and rate limiting.
  • Build multi-account CI/CD with CDK Pipelines or GitOps, with safe, staged deployments across environments.
  • Implement token cost attribution by team and application with SLO and burn-rate alerts.
  • Migrate container workloads from Amazon EKS to Amazon ECS without loss of reliability.
Essential Skills/Experience
  • Minimum of 4 years of hands-on experience in a platform engineering, infrastructure, site reliability engineering (SRE) or DevOps role where infrastructure code was your main output.
  • Deep hands-on AWS cloud engineering experience, including designing, deploying and operating production cloud-native infrastructure.
  • Strong experience with AWS services such as Bedrock, ECS, IAM, VPC, load balancing, S3, CloudWatch, Secrets Manager and related services.
  • Strong hands-on experience with IaC. AWS CDK using Python and/or TypeScript is preferred. Strong Terraform engineers who are willing to work in CDK are welcome.
  • Strong experience with Docker and Amazon ECS, including workload deployment, autoscaling, resource management, health checks, networking, security and production troubleshooting.
  • Experience designing and maintaining CI/CD and/or GitOps pipelines for production cloud workloads.
  • Strong understanding of cloud and platform engineering principles, including high availability, scalability, resiliency, observability, performance, security and cost optimization.
  • Experience with on-call rotations, incident response and root-cause analysis for production systems.
  • Strong software engineering skills in Python and/or TypeScript, with experience building production-quality platform tooling and automation.
  • Experience implementing production observability using technologies such as CloudWatch, OpenTelemetry, Prometheus, Splunk, Datadog, Pydantic Logfire, Langfuse or comparable tools.
  • Demonstrated experience partnering with software, data or ML/AI teams to transition workloads from experimentation into reliable production environments.
  • Proven ability to collaborate effectively across ML/AI engineering, software engineering, security, infrastructure and product teams.
  • Strong written and verbal communication skills, including the ability to clearly document infrastructure architecture, operational processes and platform standards.
Preferred Skills/Experience
  • Experience building internal developer platforms or self-service cloud capabilities used by multiple engineering teams.
  • Experience supporting Generative AI/LLM workloads in production, with an understanding of their unique operational challenges including model availability, latency, token consumption, evaluation and cost.
  • Hands-on experience with Amazon Bedrock or another enterprise foundation-model platform.
  • Experience implementing and/or configuring AI/model gateways, proxies or routing layers, including model routing, fallback, rate limiting and provider abstraction.
  • Experience operating infrastructure supporting agentic AI, multi-agent systems, RAG pipelines or conversational AI.
  • Experience with AgentCore or comparable managed agent-runtime capabilities.
  • Experience supporting Model Context Protocol (MCP) tools, servers or services.
  • Experience with LLM evaluation and observability platforms such as Arize Phoenix, Langfuse, Braintrust, Freeplay or Logfire.
  • Experience supporting production ML/AI infrastructure within pharmaceutical, healthcare, financial services or another highly regulated industry.
  • Understanding of GxP and/or other controls applicable to regulated ML/AI systems.
  • Experience with Kubernetes (Amazon EKS), including migrating container workloads to Amazon ECS.
This Role Is Not a Good Fit If
  • Your main experience is building agents, RAG pipelines, prompts or LLM applications, and you have not owned production infrastructure.
  • You have not written or maintained infrastructure as code for production systems.
  • You prefer research or model development to operating production platforms.
Personal Attributes
  • Platform-minded: You think beyond individual applications and build reusable capabilities that enable multiple engineering teams.
  • Operationally focused: You care about reliability, scalability, observability, security, performance and what happens after a workload reaches production.
  • Hands-on: You are comfortable writing code, building infrastructure, operating container platforms and troubleshooting production systems.
  • Customer-obsessed: You view the AI Engineers and Data Scientists using the platform as your customers and continually look for ways to improve their developer experience.
  • Organized and attentive to detail: You effectively manage multiple initiatives and competing priorities.
  • Collaborative and inclusive: You foster a positive engineering culture where teams share knowledge and solve problems together.
  • Comfortable working autonomously: You like working in an evolving environment where you can help establish new systems, patterns, and best practices.
  • Curious and committed to staying current: You love learning and applying that knowledge to the systems you build. You stay current with the latest AI trends and technologies and apply that knowledge to your work pragmatically.
Where can I find out more?
  • Explore what we’re building: www.evinova.com
  • Stay connected and see our impact in action: https://www.linkedin.com/company/evinova/

Evinova is an equal opportunity employer that is committed to diversity and inclusion and providing a workplace that is free from discrimination. Evinova is committed to accommodating persons with disabilities. Such accommodation is available on request in respect of all aspects of the recruitment, assessment and selection process and may be requested by emailing AZCHumanResources@astrazeneca.com.

The annual base pay for this position ranges from $145,000 - $185,000 USD. Base pay offered may vary depending on multiple individualized factors, including market location, job-related knowledge, skills, and experience.

In addition, our positions offer a short-term incentive bonus opportunity; eligibility to participate in our equity-based long-term incentive program. Benefits offered included a qualified retirement program [401(k) plan]; paid vacation and holidays; paid leaves; and, health benefits including medical, prescription drug, dental, and vision coverage in accordance with the terms and conditions of the applicable plans.

Additional details of participation in these benefit plans will be provided if an employee receives an offer of employment. If hired, employee will be in an "at-will position" and the Company reserves the right to modify base pay (as well as any other discretionary payment or compensation program) at any time, including for reasons related to individual performance, Company or individual department/team performance, and market factors.

Date Posted 01-Oct-2026

Closing Date 22-Oct-2026

Our mission is to build an inclusive environment where equal employment opportunities are available to all applicants and employees. In furtherance of that mission, we welcome and consider applications from all qualified candidates, regardless of their protected characteristics. If you have a disability or special need that requires accommodation, please complete the corresponding section in the application form.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Generative AI Cloud Operations Engineer - Evinova
Generative AI Cloud Operations Engineer - Evinova

AstraZeneca GmbH • Gaithersburg (MD)

On-site
USD 134,866 - 202,299
Comprehensive relocation packages
Health benefits
401(k) retirement plan
Digital Patient Scientist - Evinova
Digital Patient Scientist - Evinova

Evinova • Gaithersburg (MD)

On-site
USD 116,000 - 152,000
401(k) plan
Paid vacation
Paid leaves
+1
Principal Software Engineering Lead (eCOA) - Evinova
Principal Software Engineering Lead (eCOA) - Evinova

AstraZeneca plc • Gaithersburg (MD)

On-site
USD 165,656 - 217,424
401(k) retirement program
Paid vacation and holidays
Health benefits including medical and dental coverage
Technical Product Manager - Reporting, Dashboards, and Analytics
Technical Product Manager - Reporting, Dashboards, and Analytics

AstraZeneca plc • Gaithersburg (MD)

Hybrid
USD 186,000 - 279,000
Principal Product Engineer - Evinova
Principal Product Engineer - Evinova

AstraZeneca GmbH • Gaithersburg (MD)

On-site
USD 172,000 - 259,000
401(k) plan
Health benefits
Paid vacation and holidays
+1
Associate Strategy Director (Associate Consultant) - Clinical Development - Evinova
Associate Strategy Director (Associate Consultant) - Clinical Development - Evinova

Evinova • Boston (MA)

On-site
USD 142,000 - 213,000
Short-term incentive bonuses
Equity-based awards
Health, dental and vision coverage
+2
Digital Patient Scientist - Evinova
Digital Patient Scientist - Evinova

AstraZeneca plc • Gaithersburg (MD)

On-site
USD 116,000 - 152,000
Service Improvement Manager - Evinova
Service Improvement Manager - Evinova

AstraZeneca GmbH • Durham (NC)

On-site
USD 130,000 - 190,000
Office-based 3 days/week
Flexible working
Service Improvement Manager - Evinova
Service Improvement Manager - Evinova

Evinova • Durham (NC)

Hybrid
USD 120,000 - 160,000
Office in-office 3 days/week
Hybrid work model
Competitive compensation
AI Security Engineer I
AI Security Engineer I

Sky Mavis • Overland Park (KS)

On-site
USD 90,000 - 120,000