Experience: 5 years +Work Mode: Hybrid, 3 days/week from office (mandatory)
Base Location: Islamabad, Lahore
Data Scale: Terabyte-scale datasets, across text and image domains
Core Domains: Text, Image, and Agentic AI
Core Stack: Python, PyTorch, Agentic AI/LLMs, Docker, Kubernetes
Arbisoft is hiring a Machine Learning Engineer – AI/Agentic Systems for one of our client projects.
Our client is a globally recognized travel search platform, used by millions of travelers worldwide to search, compare, and book flights, hotels, and rental cars. The client is investing heavily in AI-powered and agentic travel planning experiences, alongside its core large-scale search and personalization systems.
What You'll Be Doing
- Working with Large-Scale Data: We work with very large datasets, often at terabyte scale, across both text and image domains.
- Building Labeled Data: Many of our problems text, image, and agentic AI start with little or no labeled data. You'll get time and ownership to build/curate labeled datasets using manual labeling, statistical/ML techniques, and LLM-based approaches.
- Owning the Full Project Lifecycle: Every team member is responsible for the complete project lifecycle, from R&D and experimentation to production deployment.
- Solving Open-Ended Problems: We work on image, text, and agentic AI problems, so experience solving open-ended ML/AI problems is important.
- Engineering Data at Scale: You should be comfortable handling, loading, processing, and parallelizing large-scale datasets efficiently.
What We’re Looking For- Strong hands-on experience with Python and PyTorch
- Demonstrated experience owning ML projects end-to-end from problem definition to production deployment
- Experience working with large-scale datasets (text and/or image) with strong data engineering fundamentals
- Exposure to building labeled datasets from limited/no ground truth via manual labeling, statistical/ML techniques, or LLM-based approaches
- Familiarity with LLMs and agentic AI concepts production experience is a strong plus, but hands-on experimentation, side projects, or research exposure will also be considered
- Working knowledge of Docker and Kubernetes for deployment
- Ability to independently explore a problem, experiment with different approaches, build the required data pipeline, and take the solution to production
Nice to Have- Experience with distributed data processing tools (Spark, Ray, Dask)
- Open-source contributions or public technical writing related to ML/LLMs
- Prior experience in a product company or applied AI/data-focused software house serving global clients
What We Value: Someone who can independently explore a problem, experiment with different approaches, build the required data pipeline, and ultimately take the solution to production.