2026 Future Talent Program – Precision Genetics Computational – Co-op
MSDAbout the role
Job Description
Empowering Precision Medicine via Large-Scale Bulk RNA-Seq Foundation Models
The Future Talent Program features Cooperative (Co-op) education that lasts up to 6 months and will include one or more projects. These opportunities in our Research and Development Division can provide you with great development and a chance to see if we are the right company for your long-term goals.
Companion diagnostic (CDx) assays are essential tools in precision medicine, enabling clinicians to identify patients most likely to benefit from specific therapeutic interventions. Large foundation models trained on collections of disease-relevant datasets offer the potential to stratify patients and predict treatment responses across multiple related conditions, such as various cancer types, by leveraging patterns learned from diverse biological contexts rather than being constrained to single-disease paradigms. The extensive availability of transcriptomics datasets through public repositories and biobanks makes RNA-seq data particularly well-suited for training these foundation models.
Transcriptomics-based foundation models can be trained on either single-cell or bulk RNA-seq datasets, each offering distinct advantages. While single-cell models offer high-resolution insights into tissue heterogeneity and rare cell types, bulk transcriptome models excel in sample-level tasks like patient stratification and biomarker identification due to greater coverage and more stable representation of sample biology1,2. A recent foundation model trained on ~10,000 bulk transcriptomics samples from The Cancer Genome Atlas (TCGA) database showed promising results in cancer subtyping and patient stratification3. Such models could significantly benefit CDx discovery and patient stratification in autoimmune diseases with high heterogeneity in symptoms and prognosis, such as inflammatory bowel disease (IBD) and rheumatoid arthritis (RA).
Unlike cancer research with the well-established TCGA database, autoimmune disease research lacks standardized databases containing sufficient samples (~10,000) for large-scale model development. To address this limitation, our group is systematically cataloging and collecting disease-relevant bulk RNA-seq datasets from public repositories (GEO, ArrayExpress) to create the database for robust autoimmune disease models. Our approach consists of three key phases:
1. Data Collection and Curation: Systematically gather and curate autoimmune disease bulk RNA-seq datasets from public repositories.
2. Data Processing and Model Development: Process raw FASTQ files using a standardized pipeline to minimize technical variability, followed by foundation model construction.
3. Model Validation: Fine-tune models using bulk RNA-seq data from clinical trials post-treatment and assess predictive accuracy for treatment response.
The recruited Co-op student will engage in all project phases, gaining hands-on experience in large-scale omics analysis, AI/ML model development, companion diagnostic discovery, and precision medicine for autoimmune diseases. This role offers a unique opportunity to work at the intersection of computational biology, machine learning, and translational medicine, advancing precision medicine in autoimmune diseases.
Required Education and Skills:
• Candidate must be currently enrolled in a graduate program (MSc or PhD) in Biomedical Engineering, Computer Science, Biological Sciences, or a related field. PhD candidates are especially encouraged to apply.
• Candidate must have availability to work full-time on-site for a 6-month period in 2026.
Preferred Experience and Skills:
• Candidate should have proficiency in R or Python programming, with a solid foundation in biostatistics and experience analyzing bulk or single-cell RNA-seq datasets.
• Candidate should have experience or strong familiarity with developing and applying machine learning models to biological data.
• Candidate should have background or keen interest in immunology, with a focus on bioinformatics applications.
• Candidate should have excellent academic record and strong analytical skills.
• Candidate should have outstanding communication and interpersonal abilities, with a proven capacity to thrive in a collaborative team environment.
Please note that this position may be closed before the posted end date or may remain open longer, at the discretion of the company.
Under New York City, Colorado State, Washington State, and California State law, the Company is required to provide a reasonable estimate of the salary range for this job. Final determinations with respect to salary will take into account a number of factors,
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s