Jobs and Careers
AN

Data Engineer (f/m/x)

ANNEA
GermanyRemotefull_timeVerifiedPosted 15 Sept 2023

About the role

<h3>Your mission</h3> <p>We are looking for a passionate Data Engineer who can help us automate and scale our data pipelines. You are a data geek with experience working as an ML/data engineer or Data Scientist, ready to level up and build production-grade data pipelines focused on time series data. You are probably subscribing to +10 ML newsletters, have arXiv sanity as your homepage and a picture of Andrew Ng on your bedside table, and you read PEP8 as a bedtime story.</p><p><br/> </p><p>RESPONSIBILITIES </p><ul><li>Clean and reshape raw data from varying sources and in different formats </li><li>Conduct data analyses </li><li>Ensure data quality</li><li>Implement, run, maintain, and evaluate data pipelines </li><li>Interpret, reflect, and report on the results with strong focus on feasibility and Client value </li><li>Develop and improve new tools using the latest technologies focusing on efficiency and automation </li><li>Use machine learning and statistical techniques to create scalable solutions for timeseries and computer vision problems</li><li>Detect and eliminate bugs in the data pipelines</li><li>Deploy the developed data pipelines in the cloud</li><li>Keep up to date with, implement, and create your own state of the artdata best practices</li></ul><p> </p><p>Ultimately, we expect deep industry knowledge as well as technical and excellent industry-related engineering expertise.</p> <h3>Your profile</h3> <p>QUALIFICATIONS </p><ul><li>Degree (BSc or MSc) in data science, computer science, engineering, mathematics (or comparable field) </li><li>+5 years of practical work experience in Data Engineering, Machine Learning, or Data Science </li><li>In-depth experience with<ul><li>Python data science stack (NumPy, SciPy, Pandas, Scikit-Learn, Jupyter and IPython.)</li><li>PIP packages</li><li>Testing frameworks (e.g. tox, pytest, pylint, flake8, black, isort, pydocstyle)</li><li>Running data pipelines in a production setting (deployed)</li><li>Working with relational and time-series database (e.g. Postgres and Timescaledb)</li><li>UNIX and git</li><li>Agile software development</li><li>Theoretical and practical knowledge of</li><li>Reliability theory and predictive maintenance</li><li>Python frameworks such as MLFlow, Prophet, Sklearn, xgboost, etc</li><li>Data orchestration tools (e.g. Airflow, Luigi, Prefect, etc)</li><li>Common machine learning algorithms (e.g. random forests, boosting algorithms, etc.) </li><li>Statistics, hypothesis testing, model evaluation, etc. </li></ul></li><li>Bonus if you have experience with<ul><li>Neural networks, deep learning, and convolutional neural networks and their implementation in Python (e.g. Keras, etc.) </li><li>Working with very large data sets and parallel processing (e.g. Dask, Spark, Hadoop, …)</li><li>Implementing CI/CD pipelines (e.g., Gitlab CI, Circle CI).</li><li>Cloud architectures in AWS, Azure, GCP, or similar</li><li>Docker, Kubernetes, and Helm</li><li>PEP8, PEP20, TDD, Clean Code</li></ul></li></ul><p> </p> <h3>Why us?</h3> <ul><li>Strong focus on R&amp;D </li><li>Freedom to develop &amp; implement your own ideas and learn new things </li><li>A fast-paced, dynamic environment with flat hierarchies in which you can put your skills to the test </li><li>The opportunity to become part of a young team that is redefining the Renewable Energy market </li><li>Office in the heart of Lisbon </li><li>Hybrid and remote options  </li></ul><p> </p><p> </p><p>Are you up for the challenge? Please apply through our Career page. Please provide CV and cover letter in English. We are team of many nations and languages, but commonly speak English for work. We look forward to your application.<br/>  </p>

Apply for this role

Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.

Apply Now →Generate Application Kit

Free account required — sign up in 30s

Company

ANNEA

View company profile →