Jobs and Careers
CY
Data Engineer (m/f/d)
Cyber InsightGermanyfull_timeVerifiedPosted 27 Oct 2025
About the role
<p>At Cyber Insight, we are building the next generation of AI-driven platforms for IT security and risk management. Our mission is to empower companies to gain deep insights into their IT landscapes and proactively mitigate risks in an increasingly complex digital world.</p>
<p>As a fast-growing startup, we combine expertise in cybersecurity, data engineering, and artificial intelligence to deliver solutions that automate risk assessments, predict potential threats, and help organizations stay ahead of evolving cyber risks. Our team thrives on innovation, collaboration, and a shared passion for making a real impact in the cybersecurity space.</p>
<p>We are looking for a <strong>hands-on Data Engineer</strong> who is passionate about building reliable, scalable, and secure data systems. You’ll help shape our data architecture and pipelines that feed our AI models and risk assessment engines — including the crucial task of mapping vulnerabilities (CVEs) to specific software and system components.</p>
<h2 id="tasks">Tasks</h2>
<ul>
<li>Design, build, and maintain <strong>data pipelines</strong> and <strong>ETL/ELT workflows</strong> across GCP and on-prem environments.</li>
<li>Ingest and process <strong>cybersecurity-relevant data sources</strong> such as CVE feeds, software inventories, vulnerability databases, and event logs.</li>
<li>Develop and maintain transformation logic and data models linking vulnerabilities (CVEs) to affected software and assets.</li>
<li>Implement and automate <strong>data validation</strong>, <strong>consistency checks</strong>, and <strong>quality assurance</strong> using tools like <strong>Great Expectations</strong> or <strong>Deequ</strong>.</li>
<li>Collaborate with AI and graph modeling teams to structure and prepare data for <strong>threat intelligence</strong> and <strong>risk quantification models</strong>.</li>
<li>Manage and optimize data storage using <strong>BigQuery</strong>, <strong>PostgreSQL</strong>, and <strong>Cloud Storage</strong>, ensuring scalability and performance.</li>
<li>Automate data workflows and testing through <strong>CI/CD pipelines</strong> (GitHub Actions, GCP Cloud Build, Jenkins).</li>
<li>Implement monitoring and observability for pipelines using <strong>Prometheus</strong>, <strong>Grafana</strong>, and <strong>OpenTelemetry</strong>.</li>
<li>Apply a <strong>security-focused mindset</strong> in data handling, ensuring safe ingestion, processing, and access control of sensitive datasets.</li>
</ul>
<h2 id="requirements">Requirements</h2>
<p>3+ years of experience in <strong>data engineering</strong>, <strong>backend data systems</strong>, or <strong>cybersecurity data processing</strong>.</p>
<ul>
<li>Strong Python skills and experience with <strong>pandas</strong>, <strong>PySpark</strong>, or <strong>Dask</strong> for large-scale data manipulation.</li>
<li>Proven experience with <strong>data orchestration and transformation frameworks</strong> (Airflow, dbt, or Dagster).</li>
<li>Solid understanding of <strong>data modeling</strong>, <strong>data warehousing</strong>, and <strong>SQL optimization and ETL pipelines (Kafka)</strong>.</li>
<li>Familiarity with <strong>CVE data structures</strong>, vulnerability databases (e.g. NVD, CPE, CWE), or security telemetry.</li>
<li>Experience integrating heterogeneous data sources (APIs, CSV, JSON, XML, or event streams).</li>
<li>Knowledge of <strong>GCP data tools</strong> (BigQuery, Pub/Sub, Dataflow, Cloud Functions) or equivalent in Azure/AWS.</li>
<li>Experience with <strong>containerized environments</strong> (Docker, Kubernetes) and infrastructure automation (Terraform or Pulumi).</li>
<li>Understanding of <strong>data testing</strong>, <strong>validation</strong>, and <strong>observability practices</strong> in production pipelines.</li>
<li>A structured and security-aware approach to building data products that support <strong>AI-driven risk analysis</strong>.</li>
</ul>
<p>Nice to Have</p>
<ul>
<li>Experience working with <strong>graph databases</strong> (Neo4j, ArangoDB) or <strong>ontology-based data modeling</strong>.</li>
<li>Familiarity with <strong>ML pipelines</strong> (Vertex AI Pipelines, MLflow, or Kubeflow).</li>
<li>Understanding of <strong>software composition analysis</strong> (SCA) or vulnerability scanning outputs (e.g. Trivy, Syft).</li>
<li>Background in <strong>threat intelligence</strong>, <strong>risk scoring</strong>, or <strong>cyber risk quantification</strong>.</li>
<li>Experience in <strong>multi-cloud or hybrid setups</strong> (GCP, Azure, on-prem).</li>
</ul>
<h2 id="benefits">Benefits</h2>
<ul>
<li>Freedom to design and shape a modern, secure data platform from the ground up.</li>
<li>A collaborative startup environment where your work directly supports AI and cybersecurity products.</li>
<li>Flexible working hours and remote-friendly setup.</li>
<li>Exposure to cutting-edge technologies in <strong>AI</strong>, <strong>data engineering</strong>, and <strong>c
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s