Principal / Senior Data Engineer
OnapsisAbout the role
About the job
The world’s most critical--and at risk--business applications have been neglected for far too long. Onapsis eliminates this blind spot by providing cybersecurity solutions dedicated to business-critical applications. Whether running on premises, in the cloud, or in a hybrid environment, Onapsis helps nearly 30% of the Forbes Global 100 understand the threats and risks across their SAP and Oracle landscapes.
We are seeking a Senior Data Engineer to join our mission-driven team. This role is ideal for experienced data engineers with a proven track record in architecting scalable data pipelines, leveraging cloud technologies, and contributing to high-impact cybersecurity solutions. You will be responsible for building high-performance ETL frameworks, optimizing data platforms, and contributing directly to the enhancement of our customers' threat detection, response, and remediation capabilities.
What you will be doing, your legacy:
You will be working directly with company Principal Engineers evaluating, scoping, proposing, and building features to fulfill business solution requirements to protect our customers. You will be working directly setting the foundation of a new product. Additionally, you will be working with Engineering and DevOps to deliver high-quality products and services while also working closely with security and IT professionals to ensure safe and secure best practices are followed.
Responsibilities:
- Architect and Design Scalable Data Solutions: Design, develop, and maintain highly-scalable ETL/ELT pipelines across diverse data domains using cloud technologies like AWS (Glue, Redshift, Lambda, EMR, S3) and Azure (Data Factory, Synapse, Databricks).
- Data Pipeline Development: Implement data models and data processing frameworks (Spark, Kafka, Snowflake) to ingest, transform, and load large datasets (100+ TB), ensuring high availability and reliability of data.
- Advanced Data Integration: Develop solutions that integrate multiple data sources into Snowflake or similar data warehouses to enable real-time analytics and reporting across dashboards.
- AI/ML Integration: Collaborate with cross-functional teams to co-develop AI-driven features like text summarization and chatbot functionalities using AWS Bedrock, SageMaker, or similar AI/ML technologies, reducing response times and enhancing decision-making capabilities.
- Compliance and Security: Ensure compliance with industry standards and secure best practices (SOX, SOC 1/2), by implementing data governance frameworks, monitoring data pipelines, and optimizing cloud database architectures to protect sensitive information.
- Stakeholder Collaboration: Work closely with stakeholders, including analysts, engineers, and product managers, to understand their data needs, propose solutions, and drive data-driven decision-making by delivering actionable insights.
- Data Infrastructure Monitoring: Continuously monitor, troubleshoot, and enhance data pipelines, leveraging CI/CD tools (Docker, Jenkins, GitHub Actions) and orchestrating workflows using Apache Airflow to maintain operational efficiency.
- Leadership and Mentorship: Provide technical leadership within the data platform organization, leading the implementation of cutting-edge cloud technologies and mentoring junior data engineers in best practices and advanced data management techniques.
- Cloud Migration: Lead large-scale database migrations from on-premises environments (Oracle, SQL Server) to cloud-based solutions like Snowflake and AWS, improving query performance and reducing technical debt.
- Documentation and Governance: Establish comprehensive documentation for data architecture, governance, and processes to ensure scalability, compliance, and security.
Qualifications:
- 5+ years of proven experience as a Data Engineer or in a similar role with a deep understanding of data architecture and cloud-based ETL/ELT frameworks.
- Strong experience with AWS and/or Azure cloud services, particularly with Glue, Redshift, Lambda, Step Functions, Databricks, Synapse, and Snowflake.
- Proficiency in big data technologies such as Apache Spark, Kafka, Hadoop, and Databricks for distributed data processing.
- Strong programming skills in Python and SQL, with experience in advanced data modeling (star, snowflake schemas) and partitioning techniques.
- Hands-on experience in building real-time data processing and AI/ML-driven analytics solutions (SageMaker, Bedrock, NLP, Power BI).
- Proven ability to architect and manage data warehouse solutions (e.g., Snowflake, Redshift) for enterp
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s