Jobs and Careers
EL

Data Engineer

Elicit
United StatesRemotefull_timeVerifiedPosted 16 Apr 2025
💰 $230,000/yr($195,000/yr$230,000/yr)

About the role

About Elicit

Elicit is an AI research assistant that uses language models to help researchers figure out what’s true and make better decisions, starting with common research tasks like literature review.

What we're aiming for:

  1. Elicit radically increases the amount of good reasoning in the world.

    • For experts, Elicit pushes the frontier forward.

    • For non-experts, Elicit makes good reasoning more affordable. People who don't have the tools, expertise, time, or mental energy to make well-reasoned decisions on their own can do so with Elicit.

  2. Elicit is a scalable ML system based on human-understandable task decompositions, with supervision of process, not outcomes. This expands our collective understanding of safe AGI architectures.

Visit our Twitter to learn more about how Elicit is helping researchers and making progress on our mission.

Why we're hiring for this role

Our users responded enthusiastically to the latest evolution of Elicit which we launched in late 2023. We introduced Elicit Plus and Pro—our monthly subscription plans—and added thousands of paying users in a matter of months as well as hundreds of thousands of new sign-ups. This has been energizing for our team, but we want to ship more useful functionality to our users even faster.

Our academic research paper search pipeline is at the heart of Elicit's capabilities, and we need a skilled Data Engineer to take it to the next level.

Our mission is to make Elicit the most complete and up-to-date database of scholarly sources. As we continue to add new data sources and expand our coverage of academic literature, we face challenges in efficiently processing, deduplicating, and indexing hundreds of millions of research papers. In addition, enterprise customers will need to be able to upload, process, and manage their own documents in private corpora for their teams to use.

We're looking for someone who can architect and implement robust, scalable solutions to handle our growing data needs while maintaining high performance and data quality.

Our tech stack

  • Data pipeline: Python, Flyte, Spark

  • Frontend: Next.js, TypeScript, and Tailwind

  • Backend: Node and Python

  • We like static type checking in Python and TypeScript

  • All infrastructure runs in Kubernetes across a couple of clouds

  • We use GitHub for code reviews and CI

Am I a good fit?

Consider the questions:

  • How would you optimize a Spark job that's processing a large amount of data but running slowly?

  • What are the differences between RDD, DataFrame, and Dataset in Spark? When would you use each?

  • How does data partitioning work in distributed systems, and why is it important?

  • How would you implement a data pipeline to handle regular updates from multiple academic paper sources, ensuring efficient deduplication?

If you have a solid answer for these—without reference to documentation—then we should chat!

Location and travel

We have a lovely office in Oakland, CA, but we don't all work from there all the time. It's important to us to spend time with our teammates, however, so we ask that all Elicians spend 1 week out of every 6 with teammates.

  • We have a quarterly team retreat, normally in and around the SF bay area.

  • We have quarterly co-working weeks (offset from the team retreats) in our Oakland office.

  • If you come to the retreats and co-working weeks, you'll meet our expectations for in-person time!

  • There is flexibility around the specifics here: if you're not sure you can make this work, get in touch.

What you'll bring to the role

  • 5+ years of experience as a data engineer building core datasets and supporting business verticals with high data volumes

  • Strong proficiency in Python (5+ years experience)

  • Experience with architecting and optimizing large data pipelines with Spark

  • Strong SQL skills, including understanding of aggregation functions, window functions, UDFs, self-joins, partitioning, and clustering approaches

  • Experience with Parquet file formats and other columnar data storage formats

  • Strong data quality management skills

  • Ability to balance technical expertise with creative problem-solving

  • Excited to interact with product/web app and help ship new features (e.g., real-time updates, advanced filtering options)

  • <

Apply for this role

Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.

Apply Now →Generate Application Kit

Free account required — sign up in 30s

Company

Elicit

View company profile →