Data Engineer
Convey Health SolutionsAbout the role
Job Title: Data Engineer
Position Summary:
The Data Engineer must have expertise in SQL, Python and PySpark, with extensive knowledge and practical experience in utilizing AWS services. The Senior Data Engineer has a strong background in data engineering, with a focus on building scalable and efficient data pipelines. They will work with a wide array of healthcare data for ingestion, processing and consumption including but not limited to eligibility, claims, payments and risk adjustment. The Senior Data Engineer supports Pareto Operations.
Key Duties and Responsibilities:
- Design, develop, and maintain robust data pipelines using Python and PySpark to process large volumes of healthcare data efficiently in a multitenant analytics platform.
- Collaborate with cross-functional teams to understand data requirements, implement data models, and ensure data integrity throughout the pipeline.
- Optimize data workflows for performance and scalability, considering factors such as data volume, velocity, and variety.
- Implement best practices for data ingestion, transformation, and storage in AWS services such as S3, Glue, EMR, Athena, and Redshift.
- Model data in relational databases (e.g., PostgreSQL, MySQL) and file-based databases to support data processing requirements.
- Design and implement ETL processes using Python and PySpark to extract, transform, and load data from various sources into target databases.
- Troubleshoot and enhance existing ETLs and processing scripts to improve efficiency and reliability of data pipelines.
- Develop monitoring and alerting mechanisms to proactively identify and address data quality issues and performance bottlenecks.
Education and Experience:
- Bachelor's or Master's degree in Computer Science, Data Engineering, or a related field.
- Minimum of 5 years of experience in data engineering, with a focus on building and optimizing data pipelines.
- Expertise in Python programming and hands-on experience with SQL and PySpark for data processing and analysis.
- Proficiency in Python frameworks and libraries for scientific computing (e.g. Numpy, Pandas, SciPy, Pytorch, Pyarrow).
- Strong understanding of AWS services and experience in deploying data solutions on cloud platforms.
- Experience working with healthcare data, including but not limited to eligibility, claims, payments, and risk adjustment datasets.
- Expertise in modeling data in relational databases (e.g., PostgreSQL, MySQL) and file-based databases, ETL processes and data warehousing concepts.
- Proven track record of designing, implementing, and troubleshooting ETL processes and processing scripts using Python and PySpark.
- Excellent problem-solving skills and the ability to work independently as well as part of a team.
- Relevant certifications in AWS or data engineering would be a plus.
Knowledge, Skills, and Abilities:
- Expertise in Python programming language for data processing and analysis.
- Expertise in PySpark for building scalable data pipelines.
- Hands-on experience with SQL query authoring for data analyses and validation
- In-depth knowledge of AWS services such as S3, Glue, EMR, Athena, and Redshift for data storage and processing.
- Familiarity with relational databases (e.g., PostgreSQL, MySQL) and file-based databases for data modeling and storage.
- Understanding of data modeling, ETL processes, and data warehousing concepts.
- Knowledge of best practices in data engineering and experience in optimizing data workflows for performance and scalability.
- Experience in healthcare data domains, including eligibility, claims, payments, and risk adjustment datasets.
- Up-to-date knowledge of emerging technologies and trends in data engineering.
- Strong problem-solving skills and the ability to troubleshoot and optimize data pipelines and ETL processes.
- Excellent communication and collaboration skills to work effectively with cross-functional teams.
- Proficient in designing, implementing, and maintaining data pipelines for processing large volumes of data.
- Ability to model
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s