Jobs and Careers
AL

Lead Data Engineer (Remote)

Allocate
Menlo Park, United StatesRemotefull_timeVerifiedPosted 18 Jul 2025

About the role

Summary

Allocate is looking for a Lead Data Engineer to design and implement the foundational data infrastructure and strategy that will power our analytics, reporting, and data-driven product features. As a fintech startup on a mission to make investing in top-tier private markets more accessible, we have a wealth of financial and investment data to harness. In this role, you will take ownership of how data flows through our systems, such as modeling core financial entities, integrating internal and external sources, and enabling our engineering and product teams to make informed decisions and build compelling features. This is a fully remote position where you’ll work closely with our backend team (C#/.NET) and frontend team (Node/Vue.js) to integrate data pipelines into our platform. If you’re a hands-on engineer who wants to shape and scale a modern data strategy from the ground up in a collaborative startup environment, we want to hear from you.

Responsibilities

  • Design and Build Data Architecture: Architect and implement Allocate's data lakehouse on AWS – combining data lake storage and warehouse technologies to store diverse financial datasets. Develop a knowledge graph to model key relationships (investors, funds, companies, etc.) and integrate a vector database for storing embeddings to enable semantic search and retrieval for our AI agents across models and providers
  • Develop Data Pipelines: Create robust ETL/ELT pipelines to ingest, clean, and transform data from various sources (internal application data and third-party APIs). Ensure both batch processing and real-time data streaming are handled to support up-to-date analytics and recommendations. Build pipelines with an eye on scalability (able to handle increasing data volume and complexity) and reliability (proper error handling and monitoring).
  • Enable AI/ML Capabilities: Work closely with our data science and engineering team to provision the data and infrastructure needed for machine learning models and AI features. This includes preparing training datasets, setting up feature stores, and orchestrating workflows that feed LLM-based agents with the context they need (e.g. retrieving relevant data via vector similarity search). You will also implement systems to serve AI model outputs (such as recommendations) back into the product in real time.
  • Technical Leadership & Collaboration: Serve as the subject matter expert for data engineering and AI infrastructure within Allocate. Provide architectural guidance and best practices to engineers who consume data in their services. Work in cross-functional squads to incorporate data-driven features into the product roadmap. As we grow, help mentor junior engineers and potentially lead a small "AI Data" team, setting coding standards and fostering a culture of data excellence.
  • Infrastructure & DevOps: Collaborate with our DevOps engineers to deploy and maintain data services. Containerize and orchestrate data tools (using Docker/Kubernetes on AWS EKS) for production use. Implement CI/CD pipelines for data workflows, so that changes to data processing or models are tested and deployed automatically. Monitor the health and performance of our data platforms (setting up alerts, dashboards) and be ready to troubleshoot and resolve issues in a production environment to ensure uptime of critical data and AI services
  • Continuous Improvement: Stay up-to-date with the latest in data engineering and AI (from new AWS offerings to open-source ML tools). Evaluate and recommend new technologies – for example, assessing if a stream processing platform like Kafka/Kinesis or an orchestration tool like Airflow could improve pipeline reliability. Challenge conventions and innovate: we encourage rethinking how things are done as we push to build a world-class, intelligent platform.

What You'll Need to Succeed
  • - Extensive Data Engineering Experience: 5+ years of hands-on experience in data engineering (or related fields), including designing and building large-scale data pipelines and storage solutions. You should have taken projects through the full lifecycle from architecture design to production deployment.
  • Cloud Proficiency (AWS): Strong experience working with AWS cloud services for data. You should be comfortable with tools like S3, EC2, ECS, EKS, Athena, Redshift, Glue, and Step Functions. Experience setting up infrastructure-as-code (Terraform/CloudFormation) for these servic

Apply for this role

Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.

Apply Now →Generate Application Kit

Free account required — sign up in 30s

Company

Allocate

View company profile →