- As a Staff MLOps Engineer, you will build and own the infrastructure, tooling, and scalable systems that make high-impact AI possible
- You’ll architect and maintain the platforms that power data ingestion, feature computation, model training, automated evaluation, deployment, and ongoing monitoring for the ML teams building recommendations, LLM-based experiences, ads, visual search, growth, and trust & safety
- You will design foundational systems that allow our ML engineers to experiment faster, ship models more reliably, and operate them with confidence in production
- We’re looking for an exceptional MLOps engineer who’s passionate about enabling ML at scale, (>6M daily active users and 100’s of millions of daily user interactions)
- Someone who loves building robust, automated pipelines; creating reliable production training and inference systems; and establishing the infrastructure and processes that accelerate ML product development across the organization
- In this role, you will shape and execute the strategy for Grindr’s ML platform and end-to-end model lifecycle
- Build and maintain end-to-end ML pipelines for data ingestion, feature computation, model training, validation, deployment, and inference, all at substantial scale of data
- Stand up and manage a feature store, ensuring feature consistency, lineage, and reuse across teams
- Expertise with best in class tools for managing deployment, scheduling, and environments and how to use them in the specialized regime of ML Infrastructure
- Develop automated model deployment workflows with CI/CD, safe rollout strategies, and reproducibility guarantees
- Implement monitoring and observability for ML systems, including data quality checks, drift detection, performance metrics, and alerting
- Build and support training environments with experiment tracking, distributed training, hyperparameter tuning, and artifact and environment management
- Collaborate with ML engineers and data engineers to streamline workflows, improve model iteration speed, and enforce MLOps best practices
- Ensure reliability, scalability, and maintainability of ML systems through strong engineering and operational rigor
Benefits
- 100% employer-paid medical, dental, and vision premiums for employees
- 401k Matching (up to 6%)
- Flexible PTO, sick-time, holidays, and two company wide shut-down weeks + an annual travel and leisure stipend
- Equity for all employees
- Monthly lifestyle stipend for food and wellness as well as fully catered onsite breakfast, lunch, and snacks
- Gender Affiliation Fund for transgender and gender non-conforming employees
- Queer-inclusive health benefits including LGBTQ consultative healthcare and expanded HRT shipped directly to employees in eligible states
- Monthly commuter stipend to get you to and from the office on in-person collaboration days
- Community projects and charitable giving stipend to support LGBTQ+ non-profit organizations alongside Grindr for Equality
Excellent engineering fundamentals: Python, SQL, bash, GitStrong understanding of cloud platforms (AWS, GCP, or Azure)Strong experience building production ML pipelines and supporting end-to-end ML workflowsBachelor’s degree in CS, Engineering, Mathematics, or related fieldExperience with big data and distributed compute: Snowflake, Spark/pySpark, Airflow, Kubernetes, Docker, Helm5+ years experience in MLOps, ML platform engineering, ML infrastructure, or similar rolesAbility to produce well-engineered, maintainable software with tests, documentation, and operational rigorExperience with data quality frameworks, observability tooling, or experiment tracking systemsExperience with ML frameworks (PyTorch, TensorFlow) sufficient to support training pipelines and deployment workflowsExperience implementing full model lifecycle management (from data → training → deployment → monitoring)Experience with vector databases, embeddings pipelines, or retrieval systemsFamiliarity with NLP/LLM-based data pipelines or image/vision data workflowsExperience with recommendation system infrastructureStrong grasp of classical ML concepts as they relate to platform designKnowledge of data governance, compliance, retention, and classificationTrack record of partnering with research/ML teams to operationalize models at scale