Staff Platform Engineer, AI/ML Infrastructure

Pfizer

Bethesda, Northern (MD, KY)

Hybrid

USD 180,000 - 240,000

Full time

9 days ago
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Pfizer is seeking a Staff Platform Engineer for AI/ML infrastructure to lead cloud platform strategy across AWS, Kubernetes, and serverless environments. You will shape architecture, tooling, and reliability practices, partnering with software, security, and operations teams.

The role requires hands-on cloud engineering with staff-level influence, mentoring engineers and guiding multi-team initiatives to raise operational maturity of AI platforms.

Qualifications

  • 7+ years of experience in DevOps, platform engineering, cloud infrastructure, SRE, or software engineering.
  • Strong hands-on experience with AWS/Azure/GCP infrastructure and services.
  • Experience designing and operating production systems on Kubernetes, ECS/Fargate, or comparable container orchestration platforms.
  • Proficiency with infrastructure-as-code, especially CloudFormation, Terraform, Helm, or similar tooling.
  • Strong CI/CD experience with GitHub Actions or similar platforms, including reusable workflows and automated testing.
  • Experience building observability solutions using CloudWatch, Prometheus, Grafana, OpenSearch.
  • Strong understanding of cloud security practices, IAM, secrets management, audit logging, and compliance.
  • Experience leading technical design, mentoring engineers, and influencing engineering practices across teams.

Responsibilities

  • Define and drive the technical strategy for AI/ML platform infrastructure supporting generative AI applications, LLM integrations, model routing, and enterprise AI services.
  • Architect, build, and operate scalable cloud platforms using AWS services (EKS, ECS/Fargate, Lambda, DynamoDB, S3, OpenSearch).
  • Establish reusable infrastructure patterns using CloudFormation, Helm, and Terraform for multi-environment deployments.
  • Lead CI/CD architecture with GitHub Actions, reusable workflows, OIDC-based AWS authentication, quality gates, and approvals.
  • Improve observability across AI platforms with CloudWatch, Prometheus/Grafana, OpenSearch, and Langfuse.
  • Build platform capabilities for GenAI workloads including model availability monitoring.
  • Partner with software engineering to improve deployment reliability, health checks, autoscaling, and performance.
  • Define and enforce security and compliance practices for infrastructure (IAM, Secrets Manager, tagging, audit logs).
  • Provide leadership for cost optimization, capacity planning, and operational resilience across environments.
  • Mentor engineers and influence platform engineering practices across teams.

Skills

DevOps
Platform engineering
AWS
Kubernetes
CI/CD
Observability
Security
Python
Terraform

Education

Bachelor's degree in Computer Science or related field

Tools

CloudFormation
Terraform
Helm
GitHub Actions
Prometheus
Grafana
OpenSearch

Job description

**Staf****f Platform Engineer, AI/ML Infrastructure**Department:AI Software & Operations**Role Summary** The Staff Platform Engineer, AI/ML Infrastructure will provide technical leadership for thecloud platforms, deployment systems, and operational foundations that power enterprise-scalegenerative AI applications. This role will define and evolve the infrastructure architecture for AI/ML platforms running across AWS,Kubernetes, serverless, and containerized environments. The engineer will lead platform standards forreliability, scalability, observability, CI/CD, security, and developer enablement, while partnering closelywith software engineering, AI engineering, security, and operations teams. The ideal candidate combines deep hands-on cloud engineering experience with staff-level technicalinfluence. They are comfortable designing infrastructure patterns, writing infrastructure-as-code,improving delivery pipelines, mentoring engineers, and making architectural decisions that raise theoperational maturity of AI platforms across multiple teams. **Key Responsibilities** Define and drive the technical strategy for AI/ML platform infrastructure supporting generative AIapplications, LLM integrations, model routing, and enterprise AI services. Architect, build, and operate scalable cloud platforms using AWS services such as EKS, ECSFargate, Lambda, DynamoDB, S3, OpenSearch, Secrets Manager, CloudWatch, ALB, and MWAA. Establish reusable infrastructure patterns using CloudFormation, Helm, and Terraform to supportreliable multi-environment and multi-region deployments. Lead CI/CD architecture using GitHub Actions, reusable workflows, OIDC-based AWSauthentication, automated quality gates, deployment promotion, and environment approvals. Design and improve observability across AI platforms, including CloudWatch dashboards, logs,alarms, Prometheus/Grafana, OpenSearch, Langfuse, and LLM-specific operational metrics. Build platform capabilities for GenAI workloads, including model availability monitoring. Partner with software engineering teams to improve deployment reliability, rollback strategies,health checks, autoscaling, load testing, and runtime performance. Define and enforce security and compliance practices for infrastructure, including IAM permissionboundaries, Secrets Manager usage, secret scanning, audit logging, tagging standards, andchange-management controls. Provide technical leadership for cost optimization, capacity planning, environment standardization,and operational resilience across development, test, production, and sandbox environments. Mentor engineers, review architecture and infrastructure designs, and influence platformengineering practices across teams.**Basic Qualifications** Bachelor’s degree in Computer Science, Engineering, Information Technology, or a relatedtechnical field, or equivalent practical experience. 7+ years of experience in DevOps, platform engineering, cloud infrastructure, site reliabilityengineering, or software engineering roles. Strong hands-on experience with AWS/Azure/GCP infrastructure and services, including container,serverless, networking, storage, observability, and security services. Experience designing and operating production systems on Kubernetes, ECS/Fargate, orcomparable container orchestration platforms. Proficiency with infrastructure-as-code, especially CloudFormation, Terraform, Helm, or similartooling. Strong CI/CD experience with GitHub Actions or similar platforms, including reusable workflows,automated testing, deployment gates, and cloud authentication. Experience building and operating observability solutions using CloudWatch, Prometheus/Grafana,OpenSearch, or similar tools. Strong understanding of cloud security practices, IAM, secrets management, least-privilegeaccess, audit logging, and compliance requirements. Experience supporting distributed systems, microservices, APIs, asynchronous workloads, andmulti-environment deployments. Demonstrated ability to lead technical design, mentor engineers, and influence engineeringpractices across teams.**Preferred Qualifications** Experience supporting AI/ML or generative AI platforms, including LLM gateways, model routing,prompt observability, token metering, or model failover. Experience operating platforms in regulated enterprise environments, ideally healthcare,pharmaceutical, finance, or life sciences. Experience with multi-account, multi-region AWS architectures and enterprise governancepatterns. Experience with cost optimization, autoscaling strategies, capacity planning, and cloud budgetmonitoring. Experience with load testing and performance validation using tools such as Locust or comparableframeworks. Strong Python or scripting skills for platform automation, operational tooling, and CI/CD extensions. Ability to communicate complex technical decisions clearly to engineering, security, operations,and leadership audiences. Technical Environment This role works across a modern AI platform ecosystem including: Cloud: AWS EKS, ECS Fargate, Lambda, DynamoDB, S3, OpenSearch, CloudWatch, SecretsManager, ALB, VPC, IAM Infrastructure-as-Code: CloudFormation, Helm, Terraform CI/CD: GitHub Actions, reusable workflows, OIDC federation, environment approvals, automatedrelease promotion AI/ML Platform: AWS Bedrock, Azure OpenAI, LiteLLM, Langfuse Observability: CloudWatch dashboards and alarms, Prometheus, Grafana, OpenSearch, Langfuse,custom metrics Security & Governance: IAM permission boundaries, secret scanning, audit logging, taggingcompliance, change-management automation Engineering Practices: Docker, Python, pre-commit, automated testing, load testing, code qualitygates, monorepo service standards**Leadership Expectations** As a J090 Staff-level engineer, this role is expected to operate beyond individual delivery. The engineerwill identify systemic platform gaps, define technical direction, create reusable standards, and raiseengineering maturity across multiple teams. Success in this role requires strong judgment, ownership, and communication. The engineer should beable to balance hands-on implementation with architectural leadership, guide teams through ambiguoustechnical decisions, and build platform capabilities that make AI product teams faster, safer, and morereliable.Work location assignment : Remote
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Platform Engineer (Cloud & AI Platform)
Senior Platform Engineer (Cloud & AI Platform)

OEC • Alpharetta (GA)

On-site
USD 140,000 - 190,000
Senior Software Engineer – AI Platform
Senior Software Engineer – AI Platform

PRI Technology • New York (NY)

On-site
USD 160,000 - 240,000
AI Platform Engineer
AI Platform Engineer

Corebridge Financial, Inc. • Houston (TX)

Hybrid
USD 120,000 - 180,000
Hybrid work policy
Competitive benefits
AWS AgentCore Platform Engineer - 67417
AWS AgentCore Platform Engineer - 67417

Hitachi Automotive Systems Americas, Inc. • Reading

On-site
USD 120,000 - 180,000
Lead Platform Engineer
Lead Platform Engineer

Heitmeyer Consulting • Plano (TX)

On-site
USD 150,000 - 190,000
Cloud Solutions Architect
Cloud Solutions Architect

Anblicks Inc. • Dallas (TX), Northern (KY)

Hybrid
USD 150,000 - 210,000
AI Infrastructure Engineer
AI Infrastructure Engineer

dicedemo • Boston (CT)

On-site
USD 130,000 - 170,000
Senior / Lead / Principal Platform Engineer (DevOps / Cloud Infrastructure)
Senior / Lead / Principal Platform Engineer (DevOps / Cloud Infrastructure)

CB Smart Recruit • Los Angeles (CA)

On-site
USD 200,000 - 300,000
Competitive sign-on bonus
Comprehensive benefits package
Long-term career growth opportunities
AWS AI Platform Engineer
AWS AI Platform Engineer

Agility Partners • Cincinnati (OH)

On-site
USD 120,000 - 180,000
Senior Platform Engineer: AI-Driven Infra & Security
Senior Platform Engineer: AI-Driven Infra & Security

7AI • Boston (MA)

On-site
USD 120,000 - 160,000