A complete application in a minute — tailored resume and cover letter, ready to send.
Bybit in Hong Kong is seeking an eager intern for data governance and AI data-platform work. You will explore governance approaches, build AI-assisted validation, and automate data quality and metadata workflows on real datasets.
Ideal candidates have SQL, Python, and data warehouse knowledge, plus exposure to Spark/Flink/Hive and LLM tooling. This is a 3-month, on-site internship with Mandarin language requirements.
Established in 2018, Bybit is one of the world’s leading cryptocurrency exchanges and digital financial platforms, serving over 80 million users across more than 200 countries and regions. Powered by world-class technology and a user-first mindset, Bybit delivers a seamless ecosystem across trading, payments, wealth management, custody, institutional services, and Web3 — connecting users to the future of digital finance.
Our core values define how we build. We listen, care and improve to create products and experiences that put users first. Backed by a global team of ambitious builders, problem-solvers, and innovators, we foster a high-performance and fast-moving environment where talent is empowered to drive real impact at the global scale. Supported by 24/7 multilingual customer service and a strong commitment to innovation, we are shaping the future of finance through technology, collaboration, and bold execution.
Today, Bybit is recognized as one of the most trusted and transparent platforms in the digital asset industry, continuing to expand its global presence while building the infrastructure for the next generation of financial services.
Research & Benchmarking — Survey industry and open-source approaches in AI + data governance (metadata & lineage platforms, semantic layers, data quality frameworks, NL2SQL / NL2Metric); produce a best-practice proposal adapted to our tech stack. Data Model Governance — Build AI-assisted model review capabilities: naming convention and layering (ODS/DWD/DWS/ADS) validation, duplicate and redundant model detection, lineage-based identification of unused / low-value / high-cost assets, and automated refactoring recommendations. Metric Semantic Automation — Automatically extract and standardize metric definitions from SQL, lineage, and documentation; detect definition conflicts and redundant builds of the same metric across different reports; maintain a machine-readable semantic layer to support natural language metric queries. Data Quality Automation — Auto-generate quality rules based on data profiling and lineage (replacing hand-written rules); perform anomaly detection on data volume, distribution, timeliness, and schema drift; conduct root cause analysis along lineage; implement automated alert grading and remediation recommendations (or auto-remediation). MVP Delivery — Run at least one end-to-end implementation across the three areas: problem definition → solution design → prototype → deployment on a real data domain → quantified results (coverage, precision/recall of issue detection, manual effort saved) → iteration. Documentation & Communication — Produce solution designs, evaluation methodologies, and results; present findings to platform and data stakeholders; deliver a reusable framework rather than one-off scripts.
Undergraduate or graduate student in Computer Science, Data Science, Statistics, or a related field. Solid proficiency in SQL and Python. Understanding of data warehouse fundamentals — dimensional modeling, layered architecture, metadata, and lineage. Experience with Spark / Flink / Hive / StarRocks is a plus. Hands-on experience with LLM application development: prompt engineering, RAG, Agent / tool-calling frameworks (e.g., LangChain, LlamaIndex, MCP), with the ability to evaluate whether an LLM system is actually effective. Structured thinking: able to distill a vague governance pain point into a well-defined problem with quantifiable success criteria, and honestly articulate what the MVP validated and what it did not. Self-driven and comfortable with ambiguity — this is an exploratory project with no predetermined answers. Bonus: experience with DataHub / OpenMetadata / Atlas, dbt, Great Expectations / Deequ, or any metrics / semantic layer tooling. Able to read technical materials in English; clear written communication skills. Minimum 3-month internship commitment, 5 days per week on-site. Fluent in Mandarin is required.
At Bybit, we are committed to fostering a supportive and enriching work environment.
Our benefits include: