Data Scientist
University of ChicagoAbout the role
Department
BSD OBG - Griffith Lab
About the Department
Job Summary
The Data Scientist/Statistician I will implement the research analyses, executing large-scale data harmonization, statistical analysis, and modeling – including predictive models for Alzheimer’s disease with a particular focus on sex-specific (female) risk. The harmonization pipeline will integrate female-specific and cognitive variables/items from large population-based cohort studies and longitudinal secondary health datasets drawn from multiple U.S. and international sources. Responsibilities include acquiring, cleaning, and organizing datasets; mapping and assessing available cohorts and sources by profiling variable coverage, coding systems, and cognitive instruments; prototyping an end-to-end harmonization on subsets of several cohorts with documented variable/value maps and quality-control checks; validating measurement invariance and IRT linking on one to two cognitive scales to produce crosswalks and uncertainty summaries; and delivering an analysis-ready, versioned dataset accompanied by a data dictionary and a concise user guide. Success will be demonstrated by a documented harmonization pipeline with repeatable builds and validation, clear crosswalks and comparability statements for key cognitive measures, and a dataset stakeholders can use with confidence – complete with known limitations, provenance, and QC metrics. The role also includes preparing high-quality reports, visualizations, and peer-reviewed publications as needed. The Data Scientist will serve as a key analytical expert and integral team member, contributing specialized skills in data integration, statistical modeling, and interpretation to advance the project’s goals.
This at-will position is wholly or partially funded by contractual grant funding which is renewed under provisions set by the grantor of the contract. Employment will be contingent upon the continued receipt of these grant funds and satisfactory job performance.
Responsibilities
Leads the acquisition, cleaning, and harmonization of secondary datasets from the multiple sources, including international cohort studies, with support from Dr. Farina and the project team.
Conducts data exploration and statistical analyses to extract meaningful insights from large, complex datasets, with support from Dr. Farina, Dr. Capuano and the project team.
Unify different types of data including cognitive instruments (e.g., MMSE, MoCA, Trails, Digit Symbol, HVLT, etc.): perform measurement invariance testing; build IRT/linking models and score crosswalks; document comparability limits.
Correct site/batch effects and temporal drift using mixed-effects models, empirical Bayes approaches, and sensitivity analyses.
Handle missing with principled methods (e.g., MICE, IPW); quantify robustness.
Maintain privacy-conscious data handling (HIPAA/GDPR concepts).
Maintains and analyzes statistical models using best practices in machine learning, statistical inference, and reproducible research workflows.
Prepares publication-ready tables, figures, and statistical summaries for interim and final reports.
Develops tailored statistical procedures and visualizations for specific research questions.
Analyzes moderately complex data sets for the purpose of extract
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s