Summer Intern/Data Engineering
GSKAbout the role
Why GSK?
Uniting science, technology and talent to get ahead of disease together.
GSK is a global biopharma company with a special purpose – to unite science, technology and talent to get ahead of disease together – so we can positively impact the health of billions of people and deliver stronger, more sustainable shareholder returns – as an organisation where people can thrive. We prevent and treat disease with vaccines, specialty and general medicines. We focus on the science of the immune system and the use of new platform and data technologies, investing in four core therapeutic areas (infectious diseases, HIV, respiratory/ immunology and oncology).
Our success absolutely depends on our people. While getting ahead of disease together is about our ambition for patients and shareholders, it’s also about making GSK a place where people can thrive. We want GSK to be a place where people feel inspired, encouraged and challenged to be the best they can be. A place where they can be themselves – feeling welcome, valued, and included. Where they can keep growing and look after their wellbeing. So, if you share our ambition, join us at this exciting moment in our journey to get Ahead Together.
Department Description
Onyx Data Engineering Team harnesses the power of data to drive innovation and support GSK’s strategic goals. By building robust, scalable data solutions, we empower all teams to make informed decisions, ultimately enhancing patient outcomes and advancing our mission.
Our Data Engineering interns will serve as technical contributors, helping teams translate well-defined specifications into functioning components—such as pipelines, services, APIs, or functions. They will follow best practices for software development and data engineering, including code quality, documentation, DevOps, and testing.
At GSK, we have a leading portfolio of vaccines, respiratory and specialty medicines as well as R&D based on immune system and genetics science. GSK’s ambition and purpose are to unite science, talent and technology to get ahead of disease together – all with the clear ambition of delivering human health impact; stronger and more sustainable shareholder returns; and as a new GSK where outstanding people thrive.This internship will support the Onyx Data Engineering Team to deliver tech innovation that will support our scientists to drive scientific breakthroughs that will change lives all over the world.
Something to get you excited about Onyx:
What is Onyx? Hear directly from Shobie, GSK’s Chief Digital and Technology Officer: Welcome to GSK Onyx - YouTube
Why Onyx? Hear from Nick, VP of Onyx Research Data Platform and Kim,SVP and Head of AI/ML: Why Onyx? - YouTube
Watch this before you apply! GSK Early Careers Top Tips:
Learn practical advice from GSK on how to prepare a standout application and succeed in your early careers journey: GSK Early Careers Top Tips
Job Description
- Create modular code and services using modern data engineering tools (Python, Spark, Kafka) and orchestration platforms (Google Workflow, Airflow).
- Develop well-engineered solutions with automated test suites and comprehensive documentation.
- Maintain consistent logging and data lineage by enforcing platform abstractions.
- Adhere to QMS (Quality Management System) frameworks and CI/CD best practices for reliable deployments.
- Troubleshoot and resolve issues in existing tools, services, and pipelines.
Minimum Qualifications
- Pursuing a Bachelor’s or Master’s in Computer Science or related disciplines.
- Proficiency in at least one programming language (Python, Java, or Scala).
- Basic knowledge of SQL and relational databases.
- Familiarity with data structures, algorithms, and simple ETL concepts.
- Must be able to work full-time (35-40 hours/week) throughout the 12-week Internship (May/June-August 2026).
- Must have an active student status and/or within 12 months post-graduation from a BS or MS degree program. Post-doctoral candidates are not eligible.
Preferred Qualifications
- Prior hands-on experience in data engineering or software development through internships, hackathons, or robust academic projects.
- Exposure to big data and streaming technologies, such as Apache Spark and Kafka.
- Familiarity with modern software development workflows, including Git/GitHub and basic DevOps principles.
- Experience using, or a strong eagerness to learn, AI assistants (e.g., GitHub Copilot, Claude Code, Gemini) to accelerate development, troubleshoot code, and understand modern workflows.
- Proficiency in Microsoft Word,
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s