Senior Big Data Engineer
Definitive HealthcareAbout the role
At Definitive Healthcare, our passion is to transform data, analytics and expertise into healthcare commercial intelligence. We help clients uncover the right markets, opportunities and people, so they can shape tomorrow’s healthcare industry. Our SaaS platform creates new paths to commercial success in the healthcare market, so companies can identify where to go next.
Our employees are kind, collaborative, energetic, approachable and driven. On top of that, we value the unique perspectives, backgrounds and voices of our employees. Why? Because their diverse experiences drive new ideas and help us build a better community.
For over 10 years, we’ve built a collaborative culture driven by employees who share a passion for improving the healthcare ecosystem, enjoy giving back to the local community and value diversity and inclusion.
One of the hallmarks of our culture is our commitment to community service. Through the DefinitiveCares program, employees can work with their choice of more than 40 charitable organizations, supporting causes from hunger and homelessness to healthcare, LGBTQ+ issues, racial justice, women’s initiatives and more. 2021 marked the sixth year that we had 100% employee participation in DefinitiveCares.
We also provide a range of opportunities for employees to connect with each other. Employees can join any of our employee run affinity groups supporting causes such as women’s empowerment, LGBTQ+, Black, indigenous and people of color (BIPOC), disabilities and working parents and potential for many more. Affinity groups often enable greater education companywide through training, events and speaker series.
We’re also a great place to work. For five years in a row, we’ve been recognized by the Boston Business Journal and the Boston Globe as a best place to work in Massachusetts. In 2022, Energage recognized us for Culture Excellence in Compensation & Benefits, Innovation, Great Leadership, Purpose & Value and Work-Life Flexibility!
Think you’d be a good addition to our team? Explore our available positions here. We’d love the chance to get to know you.
Responsibilities:
- Design and Develop Data Pipelines:
- Build and maintain scalable data pipelines using Python, Spark, and Databricks.
- Implement data workflows and ETL processes using Apache Airflow.
- Data Integration and Management:
- Integrate data from various sources (AWS, GCP, on-premises) into a unified data warehouse.
- Handle variety of data formats such as csv, text, xml, parquet, delta etc.,
- Ensure data quality and integrity through effective data cleansing and curation practices.
- Manage and optimize data storage solutions, ensuring high availability and performance.
- Automate observability of data and workloads
- Metadata Management and Governance:
- Implement and manage Unity Catalog for metadata management.
- Ensure data governance policies are followed, including data security, privacy, and compliance.
- Develop and maintain data documentation and data dictionaries.
- Automate data observability across pipelines
- Performance Tuning and Troubleshooting:
- Optimize Spark jobs for performance and efficiency.
- Investigate and resolve performance bottlenecks in Spark applications.
- Utilize JVM tuning techniques to improve application performance.
- Data Maturity Lifecycle:
- Implement and manage the Medallion architecture for data maturity lifecycle.
- Ensure data is appropriately processed and categorized at different stages (bronze, silver, gold) to maximize its usability and value.
- Collaboration and Continuous Improvement:
- Work closely with data scientists, analysts, and other stakeholders to understand data needs and deliver solutions.
- Implement CI/CD pipelines to automate deployment and testing of data infrastructure.
- Stay up to date with the latest industry trends and technologies to continuously improve data engineering practices.
Required Skills and Qualifications:
- Technical Skills:
- Hands-on Python or Scala programming.
- Strong experience with Apache Spark and Databricks.
- Hands-on experience with Apache Airflow or similar workflow orchestration tools.
- Data modeling and processing fundamentals with large-scale volume of d
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s