Jobs and Careers
RE
Senior Data Engineer
Red HatUnited StatesRemotefull_timeVerifiedPosted 3 Aug 2026
💰 $180,000/yr($158,309/yr – $180,000/yr)
About the role
*Telecommuting role to be performed anywhere in the U.S.
Architect and implement complex, high-volume data pipelines between Snowflake and Databricks utilizing PySpark for distributed data processing and dbt for SQL-based data transformation, including developing macro-driven data quality tests and validation frameworks.
What You Will Do:
- Orchestrate pipeline scheduling, dependency management, and automated failure recovery using Apache Airflow to deliver B2B marketing attribution and multi-touch targeting analytics.
- Administer the enterprise Databricks platform by configuring IAM roles for secure Amazon S3 bucket access, managing application credentials and secrets through Databricks' built-in vault system and OpenShift secrets, and establishing workspace governance policies and cluster configurations for cross-functional data science and engineering teams.
- Design and deploy intelligent retrieval architecture and AI-driven workflows using vector-based search methods and enterprise data platforms, building marketing retrieval and decision-automation applications that integrate multiple data sources and APIs.
- Operationalize MLOps methodologies using MLflow for experiment tracking and model registry management, and Lakehouse monitoring for automated post-production model performance tracking to optimize predictive accuracy and increase marketing return on investment.
- Implement end-to-end machine learning models and deliver stakeholder-facing analytical outputs by building Named Entity Recognition (NER) systems using TextBlob, gensim, and fastText for enterprise systems analysis, developing predictive models using XGBoost and Scikit-learn, constructing deep learning architectures using Keras, and designing time-series forecasting models for event-based user adoption prediction.
- Manage CI/CD pipelines using Git and Tekton to ensure reliable, repeatable code delivery for production applications.
- Build and manage container images using buildah and skopeo, pushing to internal container registries for deployment.
- Lead the deployment and maintenance of containerized data science models and enterprise applications on Red Hat OpenShift (Kubernetes), managing network routes, TLS termination, and container orchestration for highavailability services.
- Lead application security initiatives by completing comprehensive enterprise security compliance assessments encompassing 20+ security controls across the full technology stack, aligned with industry frameworks such as NIST and CIS Controls.
- Perform static application security testing (SAST) using SonarQube, execute vulnerability scanning using Qualys and pip-audit, complete Privacy Impact Assessments (PIA), and conduct STRIDE-based threat modeling.
- Collaborate with enterprise information security teams to remediate identified vulnerabilities, navigate compliance audits, and maintain centralized logging and monitoring through Splunk.
What You Will Bring:
- Master's degree (U.S. or foreign equivalent) in Computer Science or related field and three (3) years of experience in the job offered or related role OR Bachelor's degree (U.S. or foreign equivalent) in Computer Science or related field and five (5) years of experience in the job offered or related role.
- Must have three (3) years of experience with: architecting and implementing high-volume data pipelines between cloud data warehouse (Snowflake) and lakehouse (Databricks) platforms using PySpark for distributed data processing and dbt for SQL-based data transformation, including developing macro-driven data quality test frameworks and validation logic; orchestrating and scheduling data pipeline workflows using Apache Airflow, including configuring DAG-based dependency management, automated failure recovery, and pipeline monitoring for enterprise analytics workloads; administering enterprise Databricks environments, including configuring IAM roles for secure cloud object storage (Amazon S3) access, managing application secrets through platform vault systems and OpenShift secrets, and establishing workspace governance and cluster policies for cross-functional teams; implementing end-to-end machine learning models by: 1) building Named Entity Recognition (NER) systems using TextBlob, gensim, and fastText for enterprise text analysis; 2) developing predictive models using gradient boosting frameworks (XGBoost) and Scikit-learn; 3) constructing deep learning architectures using Keras; and 4) designing time-series forecasting models for event-based prediction; delivering full-scale information retrieval systems for enterprise data by researching, evaluating, and implementing Transformer architectures and Transfer Learning methodologies using deep learning frameworks for semantic search, text classification, and vector-based clustering
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s