Principal AI/ML Engineer

Soteria Reinsurance Ltd.

Durham, Northern (NC, KY)

Hybrid

USD 180,000 - 250,000

Full time

7 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Soteria Reinsurance Ltd. is seeking a Principal AI/ML Engineer to design and deploy advanced ML systems across enterprise platforms. You will lead secure model development, governance, and scalable MLOps using SageMaker, Kubernetes, and cloud-native tools.

You will mentor teams on robust AI methodologies, multi-agent safety, and compliant deployment in a fast-paced, security-focused environment in the US. Prior leadership and cloud experience required.

Qualifications

  • DE delivering scalable, secure, distributed apps with IAM and KMS.
  • Experience building AI/ML systems on cloud platforms (AWS/Azure/GCP).
  • Expertise in MLOps, CI/CD, and reproducible model deployment.
  • Strong knowledge of federated learning, adversarial robustness, governance.
  • Proficiency in Python/Java and multi-agent AI workflows.

Responsibilities

  • Architect AI/ML systems for training, inference, observability, and monitoring across cloud environments.
  • Develop trustworthy AI frameworks with governance and risk mitigation.
  • Build agentic data pipelines and lifecycle automation for models.
  • Collaborate with security, cloud, and data science teams to ensure safety and compliance.
  • Mentor engineering teams on ML algorithms and secure development practices.

Skills

Secure AI
ML platform development
MLOps
Federated learning
Adversarial robustness
Kubernetes
SageMaker
CI/CD
Python
Java
IAM & encryption
Retrieval Augmented Generation
Vector databases
Distributed systems
Cloud services (AWS/Azure/GCP)

Education

Bachelor’s degree in CS/Engineering/IT or related
Five years as Principal AI/ML Engineer
Master’s degree in CS/Engineering/IT or related
Foreign education equivalent acceptable

Tools

Terraform
Kubeflow
OpenAI Swarm
Strands
CrewAI
LangGraph
AGP protocol
SageMaker
Airflow
AWS Step Functions
Jenkins
JFrog Artifactory
MLflow
Redis
OpenSearch
Hadoop

Job description

## Job Description:**Note: Fidelity will not provide immigration sponsorship for this position.**Position Description:Designs and develops advanced Machine Learning (ML) and trustworthy AI systems that support large-scale, enterprise-wide platforms. Focuses on secure model development, AI safety, multi-agent orchestration, and cloud-scale ML infrastructure, enabling reliable, transparent, and high-assurance deployment of AI across critical business functions. Builds robust pipelines for model training, inference, and continuous monitoring, ensuring compliance and resilience under real-world conditions. Designs model lifecycle operations (MLOps) infrastructure including SageMaker and Kubernetes for reproducible, scalable, and secure ML deployment. Responsible for research and prototyping of cutting-edge AI methodologies, including federated learning, adversarial robustness, and multi-agent safety mechanisms.Primary Responsibilities**:*** Architects and implements AI/ML systems that support model training, inference, observability, and continuous monitoring across distributed cloud environments.* Develops secure and trustworthy AI frameworks, including adversarial robustness pipelines, anomaly detection models, and model governance mechanisms that ensure compliance, transparency, and risk mitigation.* Builds and optimizes agentic AI workflows that automate data pipelines, model lifecycle operations, and system self-diagnostics.* Supports research and prototyping of advanced AI/ML methodologies, including hybrid neural architectures, federated learning, adversarial learning, and multi-agent AI safety mechanisms.* Conducts performance, reliability, and robustness evaluations of AI systems under real-world constraints, high-throughput workloads, and adversarial conditions.* Collaborate with cybersecurity, cloud engineering, and data science teams to integrate AI safety and model assurance into enterprise architectures.* Collaborates with cross-functional teams to integrate AI safety and governance into enterprise architectures while mentoring engineering teams on advanced ML algorithms and secure development practices.* Delivers technical guidance and mentorship to cross-functional engineering teams on advanced ML algorithms, infrastructure patterns, and secure development practices.Education and Experience:Bachelor’s degree in Computer Science, Engineering, Information Technology, Information Systems, or a closely related field (or foreign education equivalent) and five (5) years of experience as a Principal AI/ML Engineer (or closely related occupation) developing ML platform applications for Cloud infrastructures (Amazon Web Services (AWS), Azure, Google, and IBM) using agile methodologies.Or, alternatively, Master’s degree in Computer Science, Engineering, Information Technology, Information Systems, or a closely related field (or foreign education equivalent) and three (3) years of experience as a Principal AI/ML Engineer (or closely related occupation) developing ML platform applications for Cloud infrastructures (Amazon Web Services (AWS), Azure, Google, and IBM) using agile methodologies.Skills and Knowledge:Candidate must also possess:* Demonstrated Expertise (“DE”) delivering scalable, secure, and distributed applications with robust Identity and Access Management (IAM) and encryption (Knowledge Management System (KMS)), by architecting and deploying enterprise-scale AI/ML systems and auto-ML infrastructure through Infrastructure as code (IAC) and using Cloud-native platforms (Terraform, AWS, Azure, or GCP) and container orchestration frameworks (Kubeflow).* DE architecting and engineering high-performance big data applications (AWS Glue, EMR, Kinesis, Athena, and Dynamo DB) and autonomous multi-agent systems; designing batch processing jobs and Extract, Transform, Load (ETL) pipelines to support predictive analytics using Hadoop, MongoDB, AWS, and PostgreSQL; developing agentic workflows using frameworks including Strands, CrewAI, LangGraph, and OpenAI Swarm with protocols -- Model Context Protocol (MCP) Server and Accelerated Graphics Port (AGP); and driving context aware predictive analytics and optimization models by implementing multi threaded, asynchronous solutions in Python and Java, supported by short term and long term memory management architectures for Retrieval Augmented Generation (RAG) pipelines using vector databases, OpenSearch, and high performance caching solutions (Redis or Memcached).* DE automating end-to-end application deployment and MLOps via workflow orchestration tools (Airflow and AWS Step Functions) and CI/CD pipelines (Jenkins and Git); performing model metadata processing to deliver lineage tracking, version control, auditability, reproducibility, and real time analysis of model performance, parameters, and deployment history across distributed ML workflows, using AWS DynamoDB, AWS RDS, MLflow, and AWS Athena; enabling reproducible builds, integrity checks, and scalable CI/CD by securely handling, versioning, and distributing container images, ML models, and software dependencies using JFrog Artifactory; and accelerating model development, monitoring, drift detection, and interpretability within Agile environments by engineering bridge solutions, using Go and FastAPI.* DE implementing secure AI through adversarial robustness, federated learning, and governance across distributed ML systems using AI Generative Adversarial Network (AIGAN), Google Federated Learning Framework, SageMaker Clarify, and MLflow; enhancing low latency, high throughput model serving and maximizing central processing unit (CPU) or graphics processing unit (GPU) utilization through deployment on accelerated inference servers, using Deep Java Library (DJL), Triton, and Flask; and evaluating ML model inference performance using statistical analysis, monitoring tools (CloudWatch, Datadog, and Splunk), and dashboards including Streamlit and Gradio.[Experience and/or expertise may be gained during doctoral program.]
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Principal AI/ML Engineer
Principal AI/ML Engineer

Fidelity Investments Inc. • North Carolina

On-site
USD 180,000 - 240,000
Senior Consultant, AI/ML Engineer
Senior Consultant, AI/ML Engineer

Hollstadt Consulting • Minnesota

On-site
USD 150,000 - 210,000
AI/ML Engineer
AI/ML Engineer

Veritis Group Inc • Chicago (IL)

On-site
USD 140,000 - 210,000
Principal AI/ML Engineer
Principal AI/ML Engineer

Habitat For Humanity Of Durham • Durham (NC)

On-site
USD 140,000 - 180,000
Lead AI Engineer
Lead AI Engineer

Sherwin-Williams • Cleveland (OH)

Remote
USD 140,000 - 190,000
? AI/ML Engineer
? AI/ML Engineer

Xperteez Technology • Washington

Hybrid
USD 130,000 - 190,000
Senior Software Engineer/Developer - AI
Senior Software Engineer/Developer - AI

Pyramid Systems, Inc. • Washington, Northern (KY)

On-site
USD 180,000 - 280,000
Lead AI Engineer
Lead AI Engineer

RedStream Technology • Lewisville (TX)

On-site
USD 180,000 - 240,000
AI/ML Engineer
AI/ML Engineer

Winaxis LLC • Dallas (TX)

On-site
USD 120,000 - 160,000
Gen AI Architect
Gen AI Architect

GlobalPoint • Charlotte (NC)

On-site
USD 180,000 - 240,000