ForgeApply · Job listing
Senior Data Engineer
Red Hat
See all 38 open roles at Red Hat →
Tailor your resume for this Red Hat job in about a minute.
ForgeApply tailors your resume and cover letter to this exact posting, then hands you a ready-to-submit application for Red Hat's site. Free trial, no card required.
About this role
*Telecommuting role to be performed anywhere in the U.S.
Architect and implement complex, high-volume data pipelines between Snowflake and Databricks utilizing PySpark for distributed data processing and dbt for SQL-based data transformation, including developing macro-driven data quality tests and validation frameworks.
What You Will Do: • Orchestrate pipeline scheduling, dependency management, and automated failure recovery using Apache Airflow to deliver B2B marketing attribution and multi-touch targeting analytics. • Administer the enterprise Databricks platform by configuring IAM roles for secure Amazon S3 bucket access, managing application credentials and secrets through Databricks' built-in vault system and OpenShift secrets, and establishing workspace governance policies and cluster configurations for cross-functional data science and engineering teams. • Design and deploy intelligent retrieval architecture and AI-driven workflows using vector-based search methods and enterprise data platforms, building marketing retrieval and decision-automation applications that integrate multiple data sources and APIs. • Operationalize MLOps methodologies using MLflow for experiment tracking and model registry management, and Lakehouse monitoring for automated post-production model performance tracking to optimize predictive accuracy and increase marketing return on investment. • Implement end-to-end machine learning models and deliver stakeholder-facing analytical outputs by building Named Entity Recognition (NER) systems using TextBlob, gensim, and fastText for enterprise systems analysis, developing predictive models using XGBoost and Scikit-learn, constructing deep learning architectures using Keras, and designing time-series forecasting models for event-based user adoption prediction. • Manage CI/CD pipelines using Git and Tekton to ensure reliable, repeatable code delivery for production applications. • Build and manage container images using buildah and skopeo, pushing to internal container registries for deployment. • Lead the deployment and maintenance of containerized data science models and enterprise applications on Red Hat OpenShift (Kubernetes), managing network routes, TLS termination, and container orchestration for highavailability services. • Lead application security initiatives by completing comprehensive enterprise security compliance assessments encompassing 20+ security controls across the full technology stack, aligned with industry frameworks such as NIST and CIS Controls. • Perform static application security testing (SAST) using SonarQube, execute vulnerability scanning using Qualys and pip-audit, complete Privacy Impact Assessments (PIA), and conduct STRIDE-based threat modeling. • Collaborate with enterprise information security teams to remediate identified vulnerabilities, navigate compliance audits, and maintain centralized logging and monitoring through Splunk.
What You Will Bring: • Master's degree (U.S. or foreign equivalent) in Computer Science or related field and three (3) years of experience in the job offered or related role OR Bachelor's degree (U.S. or foreign equivalent) in Computer Science or related field and five (5) years of experience in the job offered or related role. • Must have three (3) years of experience with: architecting and implementing high-volume data pipelines between cloud data warehouse (Snowflake) and lakehouse (Databricks) platforms using PySpark for distributed data processing and dbt for SQL-based data transformation, including developing macro-driven data quality test frameworks and validation logic; orchestrating and scheduling data pipeline workflows using Apache Airflow, including configuring DAG-based dependency management, automated failure recovery, and pipeline monitoring for enterprise analytics workloads; administering enterprise Databricks environments, including configuring IAM roles for secure cloud object storage (Amazon S3) access, managing application secrets through platform vault systems and OpenShift secrets, and establishing workspace governance and cluster policies for cross-functional teams; implementing end-to-end machine learning models by: 1) building Named Entity Recognition (NER) systems using TextBlob, gensim, and fastText for enterprise text analysis; 2) developing predictive models using gradient boosting frameworks (XGBoost) and Scikit-learn; 3) constructing deep learning architectures using Keras; and 4) designing time-series forecasting models for event-based prediction; delivering full-scale information retrieval systems for enterprise data by researching, evaluating, and implementing Transformer architectures and Transfer Learning methodologies using deep learning frameworks for semantic search, text classification, and vector-based clustering; operationalizing MLOps methodologies using MLflow for experiment tracking and model registry management, and implementing automated post-production model monitoring to track performance degradation and optimize predictive accuracy; managing CI/CD pipelines using Git and Tekton, building and publishing container images using buildah and skopeo to internal container registries, and deploying containerized applications on Red Hat OpenShift (Kubernetes) with network route management, TLS termination, and high-availability configurations; and leading enterprise security compliance assessments, including performing static application security testing (SAST) using SonarQube, executing vulnerability scans using Qualys, completing Privacy Impact Assessments (PIA), and conducting STRIDE-based threat modeling.
#LI-DNI
The salary range for this position is $ 158,309 - $180,000 /year. Actual offer will be based on your qualifications.
Pay Transparency Red Hat determines compensation based on several factors including but not limited to job location, experience, applicable skills and training, external market value, and internal pay equity. Annual salary is
Salary insight
The midpoint of this range ($169k) is about 21% above the median disclosed salary for Raleigh-Durham roles listed on ForgeApply ($140k across 309 jobs).
See full Data Engineer salary data for Raleigh-Durham →
Based on live postings with disclosed pay on ForgeApply; refreshed daily. Not an estimate of this employer's offer.
Tailor your resume for this Red Hat role before you apply.
Tailor my resume for this jobSimilar jobs
- Senior Data Engineer — Gordon Food Service · Wyoming, Michigan | Atlanta, Georgia
- Senior Data Engineer — GDIT · USA VA Sterling
- Senior Data Engineer — Akumincorp · Remote
- Senior Data Engineer — Abbott · Des Plaines, Illinois, United States
- Senior Data Engineer — Adobe · San Jose
- Senior Data Engineer — Guidehouse · US - VA, McLean
- Senior Data Engineer — Travelers · CT - Hartford | GA - Atlanta | MN - St. Paul
- Senior Data Engineer — Rocket · Seattle, WA
More like this: Data Engineer Jobs · Remote Data Engineer Jobs · Data Engineer Jobs in Raleigh-Durham · Browse all jobs
Free ATS checker · How to Tailor Your Resume to a Job Description (Step by Step)