Jobs and Careers
PR

Junior Mid-Level Machine Learning Data Engineer

PredictX
Gdańsk, Polandfull_timeVerifiedPosted 19 Nov 2025

About the role

About PredictX
Make a real difference at one of London’s foremost SaaS scale-ups: Be ready to pioneer the future of AI, data analytics, and technology. Step into PredictX, where we don't just see AI as a fashionable bandwagon but have lived and breathed AI & ML in every aspect of our product for the past decade. As an Enterprise SaaS provider, we're revolutionising critical decision-making for many of the world’s largest businesses, including 3 FAANGs, seeking empowerment through our integrative AI technology and Predictive Analytics.  We pride ourselves on our commitment to staying at the forefront of technological advancements. You'll be joining a team that actively explores and integrates the latest innovations to maintain our competitive edge. The RoleAs a Junior/Mid-Level ML Data Engineer, you’ll help build and maintain the data pipelines and infrastructure that power our AI solutions. You’ll work closely with senior engineers and data scientists, gaining exposure to LLMs, big data tools, and real-world ML deployment. This is a growth role for someone eager to learn, contribute, and evolve in the AI space

Key Responsibilities

  • Assist to develop and maintain scalable and efficient data pipelines using technologies such as Spark, Python, and relevant ETL tools to support our machine learning models, including those leveraging LLMs, and analytical needs.
  • Support with the architect and implement robust data warehousing solutions and data models that ensure data quality, integrity, and performance, catering to the specific data requirements of advanced AI models.
  • Assist in the development, testing, and deployment of machine learning models, including exploration and integration of Large Language Models (LLMs) and other novel AI architectures, collaborating closely with Data Scientists to productionize innovative solutions.
  • Help transform engineer approaches for storing, transforming, transporting, synchronising, archiving, and securing large and complex datasets, including unstructured and semi-structured data crucial for training and deploying advanced AI models.
  • Participate in the evaluation and testing of new machine learning models and frameworks, including LLMs, to assess their potential and applicability to our products.
  • Assist to identify and resolve performance bottlenecks, data quality issues, and other pain points within our data and ML infrastructure. Proactively recommend and implement solutions for optimisation and improvement, especially in the context of deploying large-scale AI models.
  • Define and govern data modelling and design standards, best practices, and development methodologies within the team, considering the unique challenges and opportunities presented by LLMs and other advanced AI.
  • Create and maintain comprehensive technical documentation for data pipelines, data models, and machine learning workflows, including details specific to LLM integration and testing.
  • Collaborate effectively with Business Analysts, Data Scientists, and other engineering teams to understand data requirements and deliver impactful data and AI solutions.
  • Stay abreast of the latest advancements in data engineering, machine learning (including LLMs and generative AI), and big data technologies, and actively participate in the evaluation and integration of promising new technologies.

Experience/Skills

  • 1 - 3 years of experience in data engineering or machine learning roles.
  • Proficiency in data engineering tools and technologies, including Spark (PySpark and/or Scala), Python, SQL, and various ETL/ELT tools.
  • An understanding of data modelling techniques (e.g., star schema, dimensional modelling) and data warehousing concepts.
  • Knowledge of data governance, data quality principles, and data security best practices, with an awareness of the specific security and ethical considerations related to AI models.
  • Exposure with data integration, data cleansing, and data transformation processes on large datasets, including data preparation for machine learning and LLMS.
  • Familiarity with data profiling and data lineage tools.
  • An ability to identify, diagnose, and resolve data issues, performance bottlenecks, and data quality problems effectively, including those encountered when working with large AI models.
  • Analytical and problem-solving skills to analyse data sets and translate them into actionable technical solutions, with an aptitude for understanding the nuances of AI model performance.
  • Excellent written and verbal communication skills to effectively convey technical concepts to both technical and non-technic

Apply for this role

Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.

Apply Now →Generate Application Kit

Free account required — sign up in 30s

Company

PredictX

View company profile →