Jobs and Careers
GS

Staff Scientific Knowledge Engineer

GSK
United Statesfull_timeVerifiedPosted 13 Mar 2024
💰 $240,638/yr($177,863/yr$240,638/yr)

About the role

The Onyx Research Data Platform organization represents a major investment by GSK R&D and Digital & Tech, designed to deliver a step-change in our ability to leverage data, knowledge, and prediction to find new medicines.  We are a full-stack shop consisting of product and portfolio leadership, data engineering, infrastructure and DevOps, data / metadata / knowledge platforms, and AI/ML and analysis platforms, all geared toward:

  • Building a next-generation data experience for GSK’s scientists, engineers, and decision-makers, increasing productivity and reducing time spent on “data mechanics”
  • Providing best-in-class AI/ML and data analysis environments to accelerate our predictive capabilities and attract top-tier talent
  • Aggressively engineering our data at scale to unlock the value of our combined data assets and predictions in real-time

The Knowledge Graph Platform Engineering team is responsible for the design, delivery, and maintenance of a world-class, scalable, and industrialized Knowledge Graph platform. They deliver a petabyte scale Knowledge Graph into production that is resilient, available, and most importantly scalable. They support and maintain the operations of the Knowledge Graph using a site reliability approach through monitoring, auditing, and alerting to intercept potential issues before they reach the analysis and end users. They deliver the infrastructure, IAC, and microservices used by application teams to create subgraphs that power artificial intelligence and analysis with the goal of accelerating drug discovery. They deliver the event driven microservices to bridge the gap between end user subgraph queries, data management, ontology management, and data governance systems.

This role is responsible for architecting, building, and maintaining a world-beating Knowledge Graph Platform.  The Senior Knowledge Graph Engineer is a leading technical contributor who can consistently design, scope, and deliver data projects. They should be deeply familiar with the languages and tools of modern data engineering (e.g., Scala, Spark, Kafka, ...), and engaged with the open-source community surrounding them. They support the Director of Knowledge Graph Platform Engineering in building a strong culture of accountability and ownership, as well as model best-in-class engineering practices (e.g., testing, code reviews, documentation, and DevOps-forward ways of working). They work in harmony with teammates and in close partnership with Product, Platform, and user groups such as AI/ML engineers to ensure the right data orchestration and robustness of our services. 

This role will provide YOU the opportunity to lead key activities to progress YOUR career.  These responsibilities include some of the following:

  • Designs, builds, and operates data tools, services, workflows, etc on petabytes of data on Cloud by leveraging modern data engineering tools and orchestration tools. 
  • Measure, optimize, and architect high performance systems, especially, evaluate and optimize Knowledge Graph data storage and query performance.
  • Resolve customer-facing issues and fix bugs. Debug and resolve complex issues related to knowledge graph construction and management in a timely manner.
  • Stay up-to-date with emerging trends and technologies in knowledge graph and streaming data processing.
  • Collaborate with cross-functional teams (product, platform, Quality, and DevOps) to translate business problems into technical solutions that leverage the knowledge graph. 
  • Fully versed in coding best practices and ways of working, participates in code reviews and provide constructive feedback to improve code quality and team’s standards.
  • Design, debug, and scale core query language engine.
  • Deploy to GCP using CI/CD best practices, monitor and manage GCP resources.
  • Develop secure, auditable, and performant graph query services for consumers such as AI/ML and other research teams, and integrate the query services into data catalogue, governance, and security services.

Why you?

Basic Qualifications:

We are looking for professionals with these required skills to achieve our goals:

  • Bachelor’s degree ​in computer science, Software Engineering, or related discipline.

  • Experience with industry standard big data technologies e.g., Spark, BigQuery, Kafka, HDFS, Delta Lake.
  • Experience using Scala, including toolchain, documentation, testing, and operations / observability.
  • Programming background. Experience with parser combinators, relational algebra.
  • Experience with linked data, especially RDF.
  • Experience with various data storage solutions (SQL, key-value, column, document, graph stores). 
  • Experience with data modelling, particularly involving the use of semantic data and ontologies/taxo

Apply for this role

Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.

Apply Now →Generate Application Kit

Free account required — sign up in 30s

Company

GSK

View company profile →