Skywise Engineer
AirbusAbout the role
Job Description:
Education: Engineer (or equivalent) degree in computer engineering or computer science. Experience: 6-8 Years of Experience
Job Description
Develop and maintain data pipelines for efficient data extraction, transformation, and loading (ETL/ELT) processes, utilizing PySpark for distributed big data processing and Polars/Rust for maximal performance and memory safety in single-node bottlenecks.
Design, implement, and own internal process improvements: automating manual processes, optimizing data delivery latency using high-speed language components, and re-designing infrastructure for greater scalability and cost efficiency.
Work on the data pipeline operations to operate, maintain, and evolve our decoding pipelines, proposing improvements to automatize and industrialize all processes and ways of working with a focus on data quality and platform stability.
Support the ramp-up and installations of data pipeline for future airline/aircraft deployment.
Integrate high-performance Rust-compiled routines (e.g., UDFs) into Python and PySpark workflows to resolve critical performance issues.
Technical Skills
Core High-Performance & Distributed Computing
Rust: Strong proficiency in Rust for developing memory-safe, highly concurrent, and low-latency data processing microservices or core pipeline components. Familiarity with the Cargo package manager.
Polars (Expert Level): Mastery of Polars for high-speed, multi-threaded data manipulation on single machines. Deep understanding of the Lazy API, Apache Arrow columnar format, and query optimization techniques (e.g., predicate and projection pushdown).
Python: Deep expertise in writing production-grade, modular, and reusable code (including Python packaging/wheels). Proven ability to orchestrate complex workflows and tooling.
PySpark (Expert Level): Mastery of PySpark internals, including:
Advanced Performance Tuning: Expertise in diagnosing and resolving bottlenecks using the Spark UI. Deep understanding of Adaptive Query Execution (AQE), data skew mitigation, and optimizing shuffles.
Transactional Data: Experience with Delta Lake, Apache Hudi, or Apache Iceberg for building reliable, ACID-compliant Data Lakehouse architectures (handling UPSERTs, Time Travel).
SQL (Expert Level): Expert proficiency in analytical SQL (window functions, CTEs) and database optimization (indexing, partitioning, query plan analysis).
Data Orchestration & Platform Ownership
Workflow Orchestration (e.g., Apache Airflow, Dagster): Proven experience in designing, deploying, and maintaining complex, dependency-driven DAGs in production environments, including failure recovery and alerting.
Infrastructure-as-Code (IaC): Hands-on experience with Terraform or CloudFormation to provision, manage, and secure cloud data resources and compute clusters.
Cloud Platforms (AWS/Azure/GCP): Hands-on experience with core cloud data services and cost-management practices (e.g., EMR/Dataproc, S3/ADLS/GCS, and serverless compute).
Data Governance & Software Engineering Practices
Data Quality & Lineage: Experience implementing robust data quality checks (e.g., Great Expectations) and integrating with Data Catalog/Lineage tools.
Security & Access Control: Practical experience implementing least-privilege access (IAM/RBAC), data encryption, and data masking/tokenization for sensitive data (PII/PHI).
CI/CD & Testing: Strong knowledge in building automated test, lint, and deployment pipelines (Jenkins, GitLab CI) for data services. Ability to write comprehensive
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s