ForgeApply · Job listing
Lead Data Engineer – Vice President
Citi
See all 29 open roles at Citi →
Tailor your resume for this Citi job in about a minute.
ForgeApply tailors your resume and cover letter to this exact posting, then hands you a ready-to-submit application for Citi's site. Free trial, no card required.
About this role
Citi is looking for a visionary and highly technical Lead Data Engineer – AI & Distributed Analytics within the Metrics AI & Analytics Group to architect, scale, and optimize the enterprise data platforms that power our advanced analytics, machine learning, and Generative AI capabilities. In this high-impact role, you will lead a team of data engineers to design and deliver production-grade, high-performance data pipelines that bridge raw data and intelligent applications across the enterprise. You will work at the forefront of data innovation, collaborating directly with Data Scientists, AI Researchers, Product Managers, Enterprise Architects, and business stakeholders to shape the technical direction of Citi's data ecosystem. The ideal candidate is a hands-on engineering leader with deep expertise in big data technologies and hands-on experience building and scaling high-volume distributed data lakes, including modern data orchestration, ingestion, processing, and distribution platforms, with a proven track record of delivering production-grade data solutions at scale. Responsibilities • Define and drive the technical vision, architecture, and roadmap for Citi's enterprise Data Lake and Analytics platform, working closely with Citi Enterprise Architects to ensure alignment with current and future cloud-native and hybrid-cloud strategies.
• Lead, mentor, and develop a team of data engineers, setting technical standards, conducting architecture and code reviews, fostering technical excellence, and promoting a culture of continuous learning and agile delivery.
• Collaborate with business leaders, product owners, data scientists, and enterprise architects to translate complex business requirements into scalable technical solutions and platform capabilities.
• Provide hands-on technical leadership in the design, development, deployment, and support of enterprise-scale data engineering solutions. Hands-on development is required.
• Design, build, and maintain robust batch and real-time streaming data pipelines using Apache Iceberg, Starburst, Startree/Apache Pinot, Apache Kafka, Apache Flink, and related technologies. Demonstrate expertise designing large-scale data transmission, ingestion, transformation, and processing pipelines using a Java-based technology stack.
• Architect and implement Feature Stores to standardize feature engineering and ensure consistent, reliable data delivery for model training and real-time ML inference.
• Optimize data storage and query performance across Enterprise Data Lakes and Lakehouses, leveraging platforms such as Delta Lake, Apache Iceberg, Databricks, and other modern data warehousing technologies.
• Implement end-to-end metadata management and data lineage tracking to provide full visibility into how data flows from source systems through transformations to analytics consumers and AI models.
• Establish automated data quality frameworks, including validation rules, reconciliation checks, monitoring, and controls to ensure high-fidelity data and compliance with enterprise governance standards.
• Ensure all data platforms and pipelines meet enterprise security requirements, encryption standards, and data privacy regulations such as GDPR and CCPA.
• Evaluate and prototype emerging data technologies, frameworks, and tools, translating findings into actionable recommendations that keep Citi's data platform at the cutting edge.
• Drive continuous improvement initiatives focused on platform scalability, reliability, operational excellence, and performance optimization across distributed data systems.
Required qualifications & skills • Bachelor's degree, university degree, or equivalent professional experience in Computer Science, Data Engineering, Information Systems, or a quantitative field.
• 6+ years of professional experience in data engineering, software engineering, or data platform development, including 3+ years in a technical lead or engineering leadership role.
• Expert-level proficiency in Java (required) and SQL, with additional expertise in Python highly desirable.
• Demonstrated expertise designing and building large-scale data transmission, ingestion, transformation, and processing pipelines using Java-based technologies.
• Deep expertise in big data frameworks including Apache Iceberg, Startree/Apache Pinot, Hive/HDFS, and distributed computing principles.
• Hands-on experience building and operating real-time streaming pipelines using Apache Kafka and Apache Flink at high volume and scale.
• Demonstrated ability to design and build data infrastructure specifically supporting machine learning, advanced analytics, or AI applications in production environments.
• Experience working with modern table formats and transactional storage layers such as Apache Iceberg, Delta Lake, or Apache Hudi, with strong understanding of lakehouse architectures.
• Solid experience with workflow orchestration tools such as Apache Airflow or equivalent modern orchestrators to manage complex pipeline dependencies.
• Strong understanding of data governance, metadata management, data lineage, data quality frameworks, and enterprise security practices.
• Experience building and scaling high-volume distributed data lakes, including modern data orchestration, ingestion, processing, and distribution capabilities.
• Exceptional communication skills with the ability to articulate complex technical concepts and data architectures to both technical and non-technical stakeholders.
• Strong problem-solving capabilities with demonstrated success troubleshooting and optimizing complex distributed systems.
Beneficial skills & qualifications • Master's degree in Computer Science, Data Engineering, Information Systems, or a related quantitative discipline.
• Additional programming proficiency in Python and modern data engineering frameworks.
• Experience with cloud data platforms such as Databricks, including cluster tuning, platform administration, a
Salary insight
The midpoint of this range ($178k) is about 10% above the median disclosed salary for New York roles listed on ForgeApply ($162k across 9,593 jobs).
See full Data Engineer salary data for New York →
Based on live postings with disclosed pay on ForgeApply; refreshed daily. Not an estimate of this employer's offer.
Tailor your resume for this Citi role before you apply.
Tailor my resume for this jobSimilar jobs
- Lead Data Engineer — Raymond James · Saint Petersburg, Florida - United States
- Lead Data Engineer — Inspire · Atlanta Support Center
- Lead Data Engineer — Capital One · Chicago, IL | McLean, VA
- Lead Data Engineer — Thomson Reuters · Texas, Frisco, United States | Ontario, Toronto, Canada | Minnesota, Eagan, United States
- Lead Data Engineer — Disney · Orlando, FL
- Lead Data Engineer — Aloyoga · Beverly Hills, California, United States
- Lead Data Engineer — 3M · Remote - Minnesota
- Lead Data Engineer — Atticus · Remote
More like this: Data Engineer Jobs · Data Engineer Jobs in New York · Browse all jobs
Free ATS checker · How to Tailor Your Resume to a Job Description (Step by Step)