Senior Data Engineer - Vice President

Citigroup Inc.

Irving (TX)

On-site

USD 126,000 - 189,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Citi in Irving, Texas seeks a Senior Data Developer / Lead to design, implement, and optimize real-time data pipelines within our Hadoop ecosystem and streaming platforms. This hands-on role emphasizes end-to-end delivery, strong Python/PySpark skills, and Kafka-based streaming in a banking environment.

The successful candidate will integrate AI/ML capabilities, mentor teammates, and ensure performance, security, and governance of data infrastructure in a small, autonomous squad.

Qualifications

  • 7+ years in Hadoop-based and real-time data platforms
  • Proven AI/ML exposure and coding tools
  • Experience with Kafka and real-time streaming
  • Strong Python, PySpark, SQL, Unix
  • Banking domain knowledge and regulatory awareness
  • End-to-end project delivery track record

Responsibilities

  • Develop scalable data solutions in Hadoop ecosystem
  • Lead real-time streaming using Kafka
  • Mentor squad members and share best practices
  • Optimize Hadoop clusters for performance and cost
  • Ensure data security and governance in banking context
  • Collaborate with engineers and scientists to translate requirements

Skills

Hadoop
Kafka
Python
PySpark
AI/ML
Spark
SQL
Unix

Education

Bachelor’s degree
Master’s degree preferred

Tools

HDFS
YARN
Hive
FastAPI
Spark
Git

Job description

We are strategically building anA-teamof highly skilled and autonomous individuals. We are seeking aSenior Data Developer /Lead to join one of our small, co-located, high-performing squads. This role demands a hands-on contributor who can develop, implement, and optimize robust, scalable, and real-time data solutions within our Hadoop ecosystem and streaming platforms. The ideal candidate will possess deep technical expertise in big data technologies, including real-time data streaming with Kafka, and a proven track record of end-to-end project delivery within the banking domain.

We are looking for anAI-first thinkerwho is adept at leveraging cutting-edge AI coding tools to maximize efficiency and innovation. This role requires strong technical contribution in solution development, proficiency in developing data ingestion, advanced analytics, and comprehensive reporting solutions, and active collaboration within agile, focused teams. The successful candidate will ensure optimal performance, security, and governance of our data infrastructure, with a keen focus on integrating and advancing AI/ML capabilities, while also mentoring and elevating the skills of others. High autonomy and agency are baseline expectations.

Responsibilities:

  • Develop and implement scalable and efficient data solutions within the Hadoop ecosystem, ensuring alignment with enterprise data strategy and best practices for small, co-located teams.
  • Develop, implement, and manage complex Hadoop-based data pipelines and real-time streaming solutions, demonstrating high autonomy and end-to-end ownership.
  • Implement robust real-time data streaming solutions using technologies like Apache Kafka, proactively identifying and integrating performance enhancements.
  • Collaborate effectively within a small, agile squad, translating diverse data requirements from engineers and scientists into detailed technical specifications and practical implementations.
  • Optimize Hadoop clusters and real-time streaming applications for performance, scalability, and efficient resource utilization, ensuring high throughput and low latency.
  • Maintain and monitor the Hadoop and streaming infrastructure to ensure high availability, reliability, and data integrity, proactively identifying and resolving issues with minimal oversight.
  • Implement comprehensive data security and governance policies, adhering to strict regulatory and compliance standards prevalent in the banking sector, with a deep understanding of their functional impact.
  • Stay abreast of the latest advancements and emerging trends in big data technologies, real-time analytics,AI/ML, and AI coding tools, recommending and integrating innovative, AI-first solutions to enhance team capabilities.
  • Troubleshoot and resolve complex issues within the Hadoop ecosystem and real-time streaming platforms efficiently, acting as a mentor for less experienced team members.
  • Develop Spark-based solutions to support near real-time data ingestion, advanced analytics, and comprehensive reporting, demonstrating hands-on expertise.
  • Contribute to the implementation of AI/ML models within data pipelines, focusing on data preparation, feature engineering, and seamless model deployment in production environments, and mentor others on best practices.
  • Participate in MLOps initiatives to streamline the machine learning lifecycle, including continuous integration, continuous delivery, and continuous training, emphasizing efficiency gains through AI-driven development.
  • Take full ownership of end-to-end project delivery for data initiatives, from conceptualization through implementation, testing, and deployment, ensuring projects are delivered on time, within scope, and with exceptional quality.

Recommended Qualifications:

  • 7+ years of proven experience in developing, implementing, and optimizing complex Hadoop-based and real-time data platforms, with a strong emphasis on senior, hands-on contributions.
  • Solid hands-on experience in AI/ML and modern AI coding tools.
  • Practical exposure to building ML models andCoding tools like Devin and Copilot for development is preferable.
  • Strong understanding of Hadoop ecosystem components, including HDFS, YARN, MapReduce, Hive, HBase, FastAPI and Spark.
  • Extensive hands-on experience with real-time data streaming technologies, particularly Apache Kafka, and a track record of optimizing their performance.
  • Demonstrable proficiency and experience using AI coding tools for accelerated and efficient development.
  • Strong hands-on knowledge of Python, PySpark, FastAPI, Unix, and SQL.
  • Familiarity with major cloud platforms such as AWS, Azure, and Google Cloud.
  • Experience with advanced data modeling principles and practices, including dimensional modeling
  • Proficiency in developing and implementing complex ETL/ELT processes for both batch and real-time data.
  • In-depth knowledge of data warehousing concepts and best practices.
  • Strong understanding and practical experience with AI/ML lifecycle management, MLOps practices, and the integration of machine learning models into data solutions, with an "AI-first" mindset.
  • Experience with Generative AI (GenAI) concepts and tools, and the ability to leverage them in practical applications, is a significant advantage.
  • Proven track record of successful end-to-end project delivery for complex data initiatives.
  • Deep and strong domain understanding within the banking or financial services sector, with a clear ability to articulate the functional impact and importance of technical work, and adherence to relevant data standards and regulations.
  • Excellent problem-solving skills and the ability to operate with high autonomy as part of a small, high-performing team.
  • Strong communication skills, both written and verbal, with the ability to articulate complex technical concepts, mentor others, and contribute to team growth.

Education:

  • Bachelor’s degree/University degree or equivalent experience
  • Master’s degree preferred

This job description provides a high-level review of the types of work performed. Other job-related duties may be assigned as required.

Job Family Group:

Technology

Job Family:

Applications Development

Time Type:

Full time

Primary Location:

Irving Texas United States

Primary Location Full Time Salary Range:

$125,760.00 - $188,640.00

In addition to salary, Citi’s offerings may also include, for eligible employees, discretionary and formulaic incentive and retention awards. Citi offers competitive employee benefits, including: medical, dental & vision coverage; 401(k); life, accident, and disability insurance; and wellness programs. Citi also offers paid time off packages, including planned time off (vacation), unplanned time off (sick leave), and paid holidays. For additional information regarding Citi employee benefits, please visit citibenefits.com. Available offerings may vary by jurisdiction, job level, and date of hire.

Most Relevant Skills

Please see the requirements listed above.

Other Relevant Skills

For complementary skills, please see above and/or contact the recruiter.

Anticipated Posting Close Date:

Sep 22, 2026

Citi is an equal opportunity employer, and qualified candidates will receive consideration without regard to their race, color, religion, sex, sexual orientation, gender identity, national origin, disability, status as a protected veteran, or any other characteristic protected by law.

If you are a person with a disability and need a reasonable accommodation to use our search tools and/or apply for a career opportunity review Accessibility at Citi. View Citi’s EEO Policy Statement and the Know Your Rights poster.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Data Engineer - Vice President
Senior Data Engineer - Vice President

Citi • Irving (TX)

On-site
USD 125,760 - 188,640
Senior Data Developer Tech Lead - Vice President
Senior Data Developer Tech Lead - Vice President

Citi • New York (NY)

Hybrid
USD 126,000 - 189,000
AI-First Data Tech Lead: Real-Time Hadoop & Kafka
AI-First Data Tech Lead: Real-Time Hadoop & Kafka

Citi • New York (NY)

On-site
USD 126,000 - 189,000
Senior Data Engineer - Vice President
Senior Data Engineer - Vice President

Citi • New York (NY)

Hybrid
USD 126,000 - 189,000
Medical, dental & vision coverage
401(k)
Paid time off
+2
SR. Data Engineer - Assistant Vice President
SR. Data Engineer - Assistant Vice President

Citi • New York (NY)

On-site
USD 107,000 - 161,000
Medical, dental & vision coverage
401(k)
Paid time off packages
Senior Data Engineer & AVP — Hybrid Data Platform Leader
Senior Data Engineer & AVP — Hybrid Data Platform Leader

Citi • New York (NY)

On-site
Lead Data Engineer – Vice President
Lead Data Engineer – Vice President

Citigroup Inc. • Jersey City (NJ)

Hybrid
USD 142,320 - 213,480
Hybrid working model
Continuous learning and professional发展
Lead Data Engineer – Vice President
Lead Data Engineer – Vice President

Citi • Jersey City (NJ)

Hybrid
USD 142,000 - 214,000
Hybrid work model
Lead enterprise data initiatives
Continuous learning
+4
Big Data/PySpark Engineering Lead - Vice President
Big Data/PySpark Engineering Lead - Vice President

Citi • Tampa (FL)

On-site
USD 113,000 - 171,000
Senior Big Data Developer
Senior Big Data Developer

Citi • New York (NY)

Hybrid
USD 114,000 - 171,000