Senior Data Engineer (Machine Learning) - PA or NJ
GuardianAbout the role
Guardian is seeking an innovative and dedicated Data Engineer to join our Enterprise Data & Analytics engineering team. The ideal candidate will work closely with our Data Science team, helping to enable cutting edge AI & machine learning solutions which will contribute towards enhancing wellbeing of our customers, foster growth, maintain competitive advantage, and customer satisfaction.
We strive to solve meaningful problems, create superior value, and pioneer new paths in our industry. Through your passion for Data Engineering, machine learning and data-driven decision making, you will be instrumental in shaping a future where data not only informs but also drives us forward. You will work at the forefront of technology, disrupting the status quo and enabling our business to navigate the unknown.
As a Data Engineer, you will play a key role in this exciting journey. Your contributions will go beyond coding, as you'll help bring life to ideas, transforming innovative ideas into tangible solutions that directly impact our business and customers.
You'll work in an innovative, fast-paced environment, collaborating with bright minds while enjoying a balance between strategic and hands-on work. We value continuous learning, and you will have the chance to expand your skillset, mastering new tools and technologies that advance our company's goals.
We look forward to welcoming a committed team player who thrives on creating value through innovative solutions and is eager to make a significant impact.
You will
- Collaborate with data scientists and analysts to understand data requirements and translate them into scalable, high performant data pipeline solutions.
- Support data discovery & data preparation for model development. Perform detailed analysis of raw data sources by applying business context and collaborate with cross-functional teams to transform raw data into curated & certified data assets to be used for ML and BI use cases.
- Extract text data from variety of sources like documents (Word, PDFs, Text Files, JSON etc.), logs, text notes stored in databases, using Web scrapping method from web pages to support development of NLP / LLM solutions.
- Monitor and troubleshoot data pipeline performance, identifying and resolving bottlenecks and issues.
- Collaborate with data science and data engineering team to build scalable and reproducible machine learning pipelines for training and inference.
- Implement machine learning models into operations and processes via batch, streaming and API methods.
- Develop, test, and maintain robust tools, frameworks, and libraries that standardize and streamline the data & machine learning lifecycle.
- Contribute to developing and maintaining end-to-end MLOps lifecycle to automate machine learning solutions development and delivery.
- Implement robust monitoring framework for model performance.
- Collaborate with cross-functional teams of Data Science, Data Engineering, business units and various IT teams.
- Create and maintain effective documentation for project and practices ensuring transparency and effective team communication.
- Stay up-to-date with the latest trends in modern data engineering, machine learning & AI, ensuring that our company remains at the cutting edge of industry advancements.
You Have
- Bachelor’s or Master’s degree with 5+ years of experience in Computer Science, Data Science, Engineering, or a related field.
- 4+ years of experience in working with Python, SQL, PySpark and bash scripts. Proficient in software development lifecycle and software engineering practices.
- 3+ years of experience in developing and maintaining robust data pipelines for both structured and unstructured data to be used by Data Scientists to build ML Models.
- 3+ years of hands-on experience in operationalizing Machine Learning solutions which are used in live production processes.
- 3+ years of experience working with Cloud Data Warehousing (Redshift, Snowflake, Databricks SQL or equivalent) platforms and experience in working with distributed framework like Spark.
- 2+ years of hands-on experience in using Databricks platform for MLOps using MLFlow, Model Registry and Databricks Workflow.
- Solid understanding of machine learning life cycle, data mining, and ETL techniques.
- Experience with machine learning frameworks (like Keras or PyTorch) and libraries (like scikit-learn, xgboost).
- Proficiency in API development using Flask / FastAPI frameworks and familiarity with containerization technologies like docker or Kubernetes.
- Hands-on experience in building and maintaining tools and libraries which have been used by multiple teams across organization.
- Proficient in understanding and incorporating software e
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s