Jobs and Careers
G2
Data Scientist/Software Engineer
G2 Ops, Inc.United StatesRemotefull_timeVerifiedPosted 17 Apr 2025
💰 $150,000/yr($130,000/yr – $150,000/yr)
About the role
Quick Position Facts!
Location: Virginia Beach, VA at our wonderful G2 Ops office.
Work Setting: In person, some remote opportunity and/or flexible working hours, not a fully remote position.
Looking to Start: May 2025
Salary Range: $130-$150K plus benefits
Openings: 1 Full-Time Role
Years of Industry Experience: 5+ years of relevant experience
Security Clearance Requirement: Must be able to obtain and maintain Active DoD Secret Clearance
Knowledge Requirements: Qualifications:
- Data Science Degree preferred
- Technical Skills:
- Proficiency in Python and R for machine learning and data analysis.
- Proficiency in using libraries and frameworks such as Hugging Face Transformers, TensorFlow, and PyTorch for model development and fine-tuning
- Experience with data manipulation and analysis utilizing Numpy and Pandas and utilizing libraries such as Matplotlib and Plotly for data visualizations
- Expertise in fine tuning large language models (LLMs) such as GPT, BERT, Llama, and T5, with a strong understanding of their architectures and training methodologies
- Experience with APIs for large language models (LLMs), such as OpenAI's API, to integrate advanced text generation capabilities into applications
- Ability to apply LLMs to diverse NLP tasks, optimizing model performance and utilizing prompt engineering to enhance model outputs for specific applications and use cases
- Strong foundation in statistical analysis, including proficiency in descriptive and inferential statistics, probability theory, and multivariate analysis. Ability to apply statistical techniques to draw meaningful insights from data, perform hypothesis testing, and build predictive models. Experience with time series analysis and familiarity with statistical software and tools for data exploration and visualization
- Ability to design and implement efficient training pipelines that automate the process of data ingestion, preprocessing, model training, testing, evaluation, and validation
- Experience in implementing Retrieval Augmented Generation (RAG), embedding models, and vector databases to combine retrieval mechanisms with generative models to enhance the accuracy and relevance of generated content
- Proficiency in Extract, Transform, Load (ETL) processes for converting raw data into structured formats suitable for analysis and model training. Experience in annotating and labeling datasets to enhance model accuracy for supervised learning
- Experience with hyperparameter fine tuning and optimization techniques to improve model performance
- Experience in data preprocessing, data cleaning, normalization, augmentation, and annotation
- Experience with deep learning models, including CNNs for image processing and RNNs for sequence data
- Understanding of transformer architectures and attention mechanisms for advanced NLP tasks, including language translation and text generation
- Database Expertise:
- Proficient in vector, relational, and NoSQL databases (e.g., Oracle, MySQL, MongoDB, Apache Cassandra)
- SDLC Involvement:
- Active in all software development stages, including system requirement analysis and deployment of machine learning models
- Software Design & Documentation:
- Stro
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s