About the Team
Our Data & ML Platform team powers the core intelligence of our lending operations. We own critical real-time decisioning systems that process 100,000+ requests daily with strict sub-second latency SLAs. These systems drive underwriting, credit risk, and fraud detection across 30+ Loan Product Categories. We operate a highly scalable, event-driven cloud architecture on AWS and are looking for engineers who excel at bridging the gap between data science and robust software engineering.
The Role
We are looking for a high-impact Machine Learning Engineer with 2-3 years of experience to help us deploy, scale, and maintain our production ML systems. In this role, your primary focus will be on the engineering and infrastructure lifecycle of machine learning. You will take predictive models built by our Data Science team, containerize them, optimize them for high concurrency, and deploy them as highly available REST APIs on our AWS cloud infrastructure.
Key Responsibilities
- Containerization & Deployment: Package machine learning models (XGBoost, LightGBM, Scikit-learn) into efficient Docker containers. Deploy and orchestrate these containers using AWS ECS, EKS (Kubernetes), or similar cloud-native platforms.
- Building Scalable APIs: Design, build, and maintain high-throughput, low-latency RESTful APIs (using FastAPI/Flask) to serve real-time model inference to downstream decisioning engines.
- Cloud Engineering & Infrastructure: Architect and manage resilient cloud infrastructure on AWS. Implement Infrastructure as Code (IaC) using Terraform or AWS CloudFormation to automate deployment environments.
- MLOps & CI/CD: Build robust CI/CD pipelines (via GitHub Actions, Jenkins, or GitLab CI) for automated model testing, integration, and continuous deployment. Manage model versioning and registry.
- System Reliability at Scale: Monitor model API performance, optimize memory usage, and configure auto-scaling rules to handle traffic spikes. Integrate asynchronous logging (Kafka/Kinesis) to ensure audit compliance without impacting API response times.
- Code Refactoring: Partner closely with Data Scientists to refactor experimental notebook code into modular, production-ready, and fully tested Python code.
Basic Qualifications
- Experience: 2 to 3 years of hands‑on experience in Machine Learning Engineering, Backend Engineering, or Cloud Infrastructure with a strong focus on serving ML models.
- Programming: Advanced proficiency in Python (writing production‑grade, object‑oriented code) and API frameworks (FastAPI, Flask, etc.). Proficient in SQL.
- Containerization & Orchestration: Deep hands‑on experience with Docker. Experience managing containerized workloads using AWS ECS, Kubernetes (EKS), or Docker Swarm.
- Cloud Platforms: Solid practical experience with AWS core services (EC2, S3, IAM, CloudWatch, Lambda, Application Load Balancers).
- Software Engineering Practices: Strong understanding of version control (Git), unit testing, and CI/CD automation.
Preferred Qualifications (Stand‑Out Skills)
- Big Data & Databricks: Experience with Apache Spark (PySpark) and the Databricks ecosystem (Delta Lake, Databricks Model Serving, MLflow) is a strong plus.
- High‑Performance Databases: Familiarity with real‑time NoSQL databases (e.g., DynamoDB, Redis) for sub‑millisecond feature lookups during model inference.
- Event‑Driven Architectures: Experience working with high‑throughput messaging queues (Apache Kafka, AWS Kinesis, or SQS).
- Domain Knowledge: Previous experience in FinTech, lending, or building high‑availability financial systems.
What We Offer
- Opportunity to work on high‑stakes, real‑time systems where your cloud engineering directly impacts the company's daily revenue and risk exposure.
- A culture that values automation, strict engineering standards, and scalable architectural design.
- Collaboration with cross‑functional teams of highly skilled Data Engineers and Data Scientists.