Turn this role into an interview — a resume and cover letter built around what this employer wants.
AppliedAI in Abu Dhabi is seeking an Associate ML Ops Engineer to join our growing team. You will work at the intersection of operations and development to keep our platform reliable, scalable, and secure as teams deliver AI solutions for regulated industries.
The ideal candidate is early in their SRE/MLOps career with 1–2 years of relevant experience, familiar with AWS services, automation, and modern CI/CD practices.
AppliedAI is a pioneering AI technology company headquartered in Abu Dhabi, UAE. We are committed to innovation and excellence in artificial intelligence solutions across regulated industries such as healthcare, insurance, government, and financial services.
AppliedAI is at the forefront of redefining the future of work through cutting-edge AI solutions. We empower organizations by automating complex, document-heavy processes, ensuring unparalleled efficiency and accuracy. Our commitment to synergizing human intelligence with artificial intelligence helps our clients excel in highly regulated industries, including healthcare, finance, and insurance.
We're seeking an Associate ML Ops Engineer to join our growing team. In this role, you'll work at the intersection of operations and development, supporting our platform's reliability, performance, and security while helping our development teams build and maintain scalable solutions. This is a strong opportunity for someone early in their SRE/MLOps career to build technical depth alongside a senior team.
Key Responsibilities
Help monitor and maintain production, development and staging environments, supporting high availability and performance of our architecture
Collaborate with DevOps, MLOps and Development teams to help troubleshoot and resolve issues
Support our observability stack, with guidance from senior engineers
Help maintain and improve our incident response processes
Assist with capacity planning and performance optimization for systems handling up to 10,000 requests per minute
Support compliance with security standards and regulatory requirements across our infrastructure
Participate in on-call rotation during regular business hours, with senior engineer backup for escalations
Help implement and maintain SLOs, SLIs, and SLAs
Support cloud cost optimization efforts alongside system performance and reliability
Support our CI/CD pipelines and deployment processes
Required Skills & Experience
Infrastructure Best Practices
Familiarity with infrastructure best practices such as the AWS and Azure Well-Architected Framework
Basic understanding of infrastructure patterns for high availability, fault tolerance, and disaster recovery
Understanding of infrastructure security fundamentals (principle of least privilege, network segmentation, encryption at rest/in transit)
Awareness of infrastructure compliance and governance frameworks
Interest in cost optimization strategies and FinOps practices
Exposure to Infrastructure as Code concepts (modularity, reusability, versioning)
Understanding of observability patterns (logging, metrics, tracing)
Technical Experience
1-2 years of experience in SRE, DevOps, or a similar role (internships and hands-on project experience will be considered)
Some experience with AWS services, including:
Exposure to monitoring and observability tools
Basic knowledge of infrastructure as code (CDK and/or Terraform)
Understanding of event-driven architecture concepts
Some experience with containerization and microservices
Basic scripting and automation skills
Good problem-solving abilities and a systematic approach to debugging
Experience working in Agile environments is a plus
Preferred Qualifications
Experience with Next.js, Node.js and Python
Familiarity with authentication systems (Auth0, SSO)
Awareness of regulatory compliance requirements (SOC 2, HIPAA, GDPR, PCI DSS)
Interest in ML/LLM operations
Exposure to multi-region AWS deployments
Interest in working with high-traffic systems
Basic experience with database management and optimization
Awareness of caching strategies and CDN implementations
Understanding of data lifecycle management and ETL processes
Exposure to vector and graph databases is a plus
What We Offer
Opportunity to work with cutting-edge technologies
Mentorship and guidance from senior SRE, Architecture, DevOps and MLOps engineers as you build your skills
Collaborative environment with dedicated Architecture, Development, DevOps & MLOps teams
Growth potential in a rapidly scaling startup
Work with a globally distributed team
Regular working hours with flexibility for occasional emergency support
A clear path to take on more ownership of SRE practices as you grow in the role
Required Qualities
Problem-solving mindset
Team player attitude
Self-motivated and proactive approach
Ability to work in a fast-paced startup environment
Strong interest in continuous learning and skill development
The ideal candidate is early in their reliability engineering career, eager to build technical depth in a collaborative environment, and comfortable working in a dynamic startup setting while developing toward higher standards of system reliability and performance over time.
Opportunity to work with a leading AI technology company.
Collaborative and innovative work environment.
Growing, entrepreneurial and forward-thinking culture.
Career growth and professional development opportunities.
Exposure to a thriving ecosystem working from our Abu Dhabi HQ.