Senior Data Engineer
AffinityAbout the role
Affinity stitches together billions of data points from massive datasets to create a powerful, accurate representation of the world's professional relationship graph. Based on this data, we offer our users the insights and visibility they need to nurture and tap into their team's network of opportunities.
This role is part of the AI Insights team, which owns the services that power Affinity's industry-leading relationship intelligence platform. Our team extracts and retrieves information from billions of structured and unstructured data points to deliver insights to our customers.
As a Senior Data Engineer, you will collaborate with machine learning engineers, software engineers, and product managers to shape the future of private capital's leading CRM platform. This involves designing and building scalable, efficient data extraction, load, and transform (ELT) solutions, monitoring and managing data quality, and ensuring data security and best practices.
What you’ll be doing:
- Design scalable and reliable data pipelines to consume, integrate and analyze large volumes of complex data from different sources, supporting the evolving needs of our business.
- Help define our data roadmap. You'll collaborate with our team of machine learning engineers, software engineers, product, and business leaders to use data to shape product development.
- Build and maintain frameworks for measuring and monitoring data quality and integrity.
- Establish and optimize CI/CD processes, test frameworks, and infrastructure-as-code tooling.
- Build and implement robust data solutions using Spark, Python, Databricks, Kafka, and the AWS ecosystem (including S3, Redshift, EMR, Athena, Glue).
- Identify skill and process gaps within the team, and develop processes to drive team effectiveness and success.
- Articulate the trade-offs of different approaches to building ETL pipelines and storage solutions, providing clear recommendations aligned with product and business requirements.
- To confirm you have read this entire description, please include the word '#AI-Insights' in your answer to the first application question.
Qualifications:
Don’t meet every single requirement? Studies have shown that women and people of color are less likely to apply for jobs unless they meet every qualification. At Affinity, we are dedicated to building a diverse, inclusive, and authentic workplace, so if you’re excited about this role but your past experience doesn’t perfectly align with the qualifications above, we encourage you to apply anyways. You may be just the right candidate for this or other roles.
Required:
- 5+ years of experience as a Data Engineer or Data Platform Engineer, working on complex, sometimes ambiguous engineering projects across team boundaries.
- Proficiency in data modeling, data warehousing, and ETL pipeline development is essential.
- Proven hands-on experience building scalable data platforms and reliable data pipelines using Spark and Databricks, and familiarity with Hadoop, AWS SQS, AWS Kinesis, Kafka, or similar technologies.
- Comfortable working with large datasets and high-scale data ingestion, transformation, and distributed processing tools such as Apache Spark (Scala or Python).
- Strong proficiency in SQL.
- Familiar with industry-standard databases and analytics technologies, including Data Warehousing and Data Lakes.
- Experience with cloud platforms such as AWS, Databricks, GCP, Azure or related technologies.
- Familiar with CI/CD processes and test frameworks.
- Comfortable partnering with product and machine learning teams on large, strategic data projects.
Nice to have:
- Hands-on experience with both relational and non-relational database/data stores, including vector databases (e.g. Weaviate, Milvus), graph databases, and text search engines (e.g. OpenSearch or Vespa clusters), with a focus on indexing and query optimization.
- Experience with Infrastructure as Code (IaC) tools, such as Terraform.
- Experience implementing data consistency measures using validation and monitoring tools.
- Please include your favorite programming language at the very end of your resume, outside of your skills section, with the word '#filter' next to it.
Tech Stack: Our Data stack includes tools to build data pipelines between AWS RDS and DBX via scheduled batch jobs and streaming syncing. Spark SQL and MLlib for large-scale data processing in DBX. We also build data pipelines between RDS and other search-optimized engines, such as openSearch. In-house data quality tools and governance tools to ensure data quality, security and compliance.
How we wor
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s