Staff Data Scientist
GE VernovaAbout the role
Job Description Summary
At GE Vernova Power, our Data Science and AI (DSAI) team spearheads innovation by integrating advanced data science and AI solutions to transform business operations. We are looking for a creative and detail-oriented Staff Data Scientist to be instrumental in operationalizing our Generative AI strategy, with a critical focus on data.This role serves as a crucial bridge between our complex business data and advanced AI models. You will lead teams developing statistical, machine learning, and AI solutions for Gas Power stakeholders. Your core mission involves deep exploratory analysis, strategic curation, and rigorous management of business-specific data to generate high-quality, reliable input for Large Language Model (LLM) applications.
You will contribute to deploying modern machine learning, operational research, and semantic analysis methods to derive insights and achieve Gas Power's strategic objectives. Importantly, this position focuses not on building models from scratch, but on ensuring that models developed by our central AI Foundry effectively understand and address specific business challenges. This role is ideal for a data expert passionate about uncovering hidden context and enabling transformative business solutions.
Job Description
As a Staff Data Scientist, you will be part of a data science or cross-disciplinary team developing innovative solutions, typically involving large, complex data sets to achieve business outcomes. These teams will include statisticians, computer scientists, software developers, engineers, product managers, and functional stakeholders. In addition to hands on development, the Staff Data Scientist will lead extended team members from the Emerging Technology Guild and functional DT teams to develop and operationalize data science solutions are ready for scale-up.
- Perform comprehensive exploratory data analysis (EDA) on diverse and complex business datasets, using statistical analysis, Natural Language Processing (NLP), and unsupervised clustering techniques to uncover patterns, identify quality issues, and extract meaningful insights.
- Collaborate closely with business Subject Matter Experts (SMEs) to translate their deep domain knowledge into structured, AI-ready datasets for use in prompt engineering, Retrieval-Augmented Generation (RAG), and model fine-tuning.
- Develop and prepare "golden datasets" that serve as pristine examples of our business processes, significantly reducing the iteration time for prompt engineering and AI development teams.
- Design, create, and maintain a suite of data benchmarks that represent our core business use cases. These benchmarks will be the definitive standard for evaluating the real-world performance of AI methods within our BU.
- Establish and enforce rigorous data quality standards and validation protocols, ensuring the accuracy, relevance, and integrity of all data used in our GenAI applications.
- Proactively identify and document potential data biases, working with stakeholders to develop mitigation strategies that promote responsible and fair AI outcomes.
- Serve as the primary steward for the BU’s curated AI datasets, defining and implementing a clear data management strategy that includes versioning, access controls, and a lifecycle management plan.
- Create and maintain comprehensive documentation for all curated datasets (e.g., "datasheets for datasets"), detailing their origin, schema, limitations, and intended use to ensure transparency and reusability.
- Continuously survey the BU's data landscape to identify new high-value data sources and champion their integration into our Generative AI ecosystem.
- Act as the primary data liaison between the Business Unit, prompt engineers, and the central AI Foundry.
- Rigorously test and validate the effectiveness of generalized tools and methods provided by the AI Foundry against your BU-specific data benchmarks.
- Provide precise, data-driven feedback and recommendations to the Foundry, collaborating to refine and enhance central AI capabilities to ensure they meet our specific business needs.
- Communicate methods, findings, and hypotheses with stakeholders
Qualifications
Required Qualifications:
- Bachelor’s or Master’s degree in a quantitative field such as Data Science, Computer Science, Statistics, Economics, or a related discipline.
- 3-5+ years of professional experience as a Data Scientist, Data Analyst, or in a similar role with a heavy emphasis on data exploration, manipulation, and preparation.
- Strong proficiency in Python and core data science libraries (e.g., pandas, NumPy, scikit-learn, spaCy, NLTK).
- Demonstrated experience with a wide range of exploratory data analysis and unsupervised machine learn
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s