Jobs and Careers
CI

Senior AI & Data Engineering Lead - Senior Vice President

Citi
Jersey City, United Statesfull_timeVerifiedPosted 16 Jun 2026
💰 $265,080/yr($176,720/yr$265,080/yr)

About the role

This job description outlines a senior-level role for a data architect or lead data engineer within a Data Services team. The position is centered on building and managing the data infrastructure required to support large-scale Generative AI and Machine Learning initiatives. Below is a detailed breakdown of the responsibilities and the skills required for such a role

Expanded Responsibilities

This role combines deep technical expertise in data engineering with strategic thinking and leadership. The core responsibilities can be broken down into three main pillars:

1. Strategic AI Enablement

This goes beyond just building databases; it's about designing the entire data foundation for the company's AI strategy.

  • Data Ecosystem Architecture: You will be responsible for the high-level design of the data platform. This includes:

    • Data Lake/Lakehouse Design: Implementing a central repository to store vast amounts of structured, semi-structured, and unstructured data from various sources. This could involve technologies like AWS S3, Azure Data Lake Storage, or Google Cloud Storage.
    • Federated Querying: Leveraging technologies like Starburst (commercial Trino) to create a virtual data warehouse. This allows data consumers (analysts, data scientists, AI models) to query data across different sources (e.g., data lakes, relational databases, NoSQL databases) with a single SQL query, without needing to move or copy the data.
    • Scalability and Performance: Ensuring the architecture can scale horizontally to handle petabytes of data and a high volume of concurrent queries, which is critical for pre-training large language models (LLMs).

2. Advanced AI Ops & Data Pipelines

This is the hands-on engineering aspect of the role, focused on the movement and processing of data.

  • High-Throughput Data Pipelines: You will lead the development of the data "plumbing" that powers the AI systems. This includes:
    • Batch Processing: Using Apache Spark for large-scale data transformation, cleaning, and feature engineering on historical data.
    • Real-time Stream Processing: Using Apache Kafka as a messaging bus to ingest real-time data from sources like application logs, IoT devices, or clickstreams. Apache Flink would be used for complex event processing on these streams (e.g., fraud detection, real-time recommendations).
  • Optimization and Reliability: Your pipelines must be not only fast but also resilient. This involves:
    • Low Latency: Tuning jobs and infrastructure to minimize the time it takes for data to travel from source to destination.
    • High Availability: Implementing failover mechanisms, monitoring, and alerting to ensure the data pipelines are always running and the AI models have uninterrupted access to fresh data.
    • CI/CD for Data: Implementing DevOps and AI Ops best practices for data pipelines, including automated testing, deployment, and data quality checks.

3. AI Governance & Leadership

This pillar focuses on the "people" and "process" aspects of the role, ensuring data is used responsibly and effectively.

  • Data Governance for AI: As AI systems become more critical, the data they use must be trustworthy. You will establish frameworks for:
    • Data Quality: Implementing automated checks and monitoring to ensure data is accurate, complete, and consistent.
    • Data Provenance & Lineage: Creating systems to track where data comes from, how it has been transformed, and how it is used. This is crucial for debugging models and for regulatory compliance.
    • Data Security: Working with security teams to implement access controls, data masking, and encryption to protect sensitive information, especially in the context of training AI models.
  • Team Leadership and Mentorship: This is a leadership role where you will be expected to:
    • Mentor Data Engineers: Guide junior and mid-level engineers, conduct code reviews, and establish best practices for the team.
    • Foster Innovation: Stay up-to-date with the latest technologies and methodologies in the data and AI space and encourage a culture of experimentation and continuous improvement.
    • Cross-functional Collaboration: Work closely with data scientists, ML enginee

Apply for this role

Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.

Apply Now →Generate Application Kit

Free account required — sign up in 30s

Company

Citi

View company profile →