ForgeApply · Job listing
AWS Lakehouse Data Engineer
Guidehouse
See all 462 open roles at Guidehouse →
Tailor your resume for this Guidehouse job in about a minute.
ForgeApply tailors your resume and cover letter to this exact posting, then hands you a ready-to-submit application for Guidehouse's site. Free trial, no card required.
About this role
Job Family : Software Development & Support Travel Required : None Clearance Required : Ability to Obtain Public Trust
AWS Lakehouse Data Engineer We are seeking an AWS Lakehouse Data Engineer to design, implement, and operate the cloud-native data platform that powers AI/ML, analytics, reporting, and data visualization. You will build a modern lakehouse on Amazon S3 using AWS-native services and open table formats, providing Databricks-like capabilities while maintaining portability, strong governance, cost efficiency, and operational control. You will also develop scalable batch and streaming ingestion, Python and PySpark ETL/ELT pipelines, metadata and governance services, and automated cloud provisioning and CI/CD across environments. This role is ideal for an engineer who enjoys platform building, automation, performance optimization, and enabling advanced analytics through trusted, secure, and well-governed data.
What You Will Do Build and Operate Data Pipelines (Batch and Streaming) • Design and implement batch and streaming ingestion from APIs, relational databases, file drops, event streams, and external partners.
• Implement, test, and optimize ETL/ELT pipelines using Python and PySpark to produce curated, analytics-ready datasets for reporting, visualization, and machine learning.
• Implement incremental processing, change data capture (CDC), data contracts, schema validation, and reusable transformation frameworks.
• Improve pipeline reliability through automated testing, orchestration, monitoring, retry handling, and operational runbooks.
Deliver an AWS-Native Lakehouse Data Platform • Design and implement a Delta Lakehouse-style data platform using AWS-native services to provide Databricks-like capabilities for data engineering, analysis, and data visualization.
• Build and manage a scalable lakehouse on Amazon S3 using Apache Iceberg and open columnar formats such as Apache Parquet.
• Implement SQL-like table reliability for data stored in Amazon S3, including ACID transactions, schema evolution, partition evolution, snapshot isolation, time travel, and rollback capabilities using Apache Iceberg.
• Enable fast, interactive querying of lakehouse data using AWS-native query and compute services such as Amazon Athena, Amazon EMR, AWS Glue, and Amazon Redshift where appropriate.
• Optimize performance and cost through partitioning, compaction, file sizing, statistics, caching, lifecycle policies, and efficient separation of compute and storage.
• Establish standardized development, test, and production environments with consistent configuration and controlled promotion across stages.
Metadata, Governance, Access Control, Lineage, and Quality • Implement data governance and fine-grained access control using AWS-native services, including AWS Lake Formation, AWS Glue Data Catalog, AWS Identity and Access Management (IAM), AWS Key Management Service (KMS), and related security services.
• Implement a managed metadata repository for dataset cataloging, ownership, business definitions, tagging, classification, and discoverability.
• Enable end-to-end lineage from source through transformation and consumption to support auditability, impact analysis, and regulatory requirements.
• Apply policy-based access, least-privilege permissions, row-, column-, and cell-level controls where required, data classification, retention, encryption, and secure data handling.
• Build operational data quality checks for freshness, completeness, uniqueness, validity, consistency, and anomaly detection, and publish measurable SLAs/SLOs.
AWS Automation, CI/CD, and Operations • Implement automated AWS provisioning using Infrastructure as Code (IaC) to create consistent environments and secure-by-default baselines.
• Build and enhance CI/CD for data pipelines and lakehouse components, including automated tests, security checks, validation gates, packaging, deployment, promotion, and rollback strategies.
• Implement observability with centralized metrics, logs, traces, alerts, dashboards, runbooks, and incident-response procedures.
• Continuously evaluate platform performance, scalability, reliability, security, and cost, and implement measurable improvements.
Cross-Team Collaboration and Documentation • Work closely with data, application, analytics, AI/ML, security, networking, and cloud platform teams to support mission needs and delivery timelines.
• Maintain high-quality engineering documentation, including architecture diagrams, data models, SOPs, interface specifications, operational runbooks, and secure configuration baselines.
• Present technical findings, trade-offs, risks, and recommendations clearly to technical and non-technical stakeholders.
What You Will Need • Bachelor's degree in Engineering, Information Technology, Computer Science, Data Engineering, or a related field, or FOUR (4) years equivalent practical experience in leu of degree.
• SIX (6) years of relevant experience.
• Hands-on experience implementing AWS-native data lake or lakehouse architectures using Amazon S3 and services such as AWS Glue, Amazon Athena, Amazon EMR, AWS Lake Formation, and Amazon Redshift.
• Strong experience developing production ETL/ELT pipelines using Python and PySpark, including data modeling, transformation, testing, performance tuning, and error handling.
• Hands-on experience with Apache Iceberg, including ACID transactions, snapshots, schema and partition evolution, time travel, table maintenance, and query optimization.
• Advanced SQL skills and experience supporting analytical queries, semantic layers, reporting tools, and data visualization workloads.
• Experience implementing metadata management and governance capabilities, including cataloging, lineage, ownership, classification, policy enforcement, and fine-grained access controls.
• Experience with AWS security fundamentals, including I
Tailor your resume for this Guidehouse role before you apply.
Tailor my resume for this jobSimilar jobs
- AWS Data Engineer — Clera · Remote
- Senior Software Engineer, Data Lake — Robinhood · Bellevue, WA
- AWS Database Engineer — Accenturefederalservices · Arlington, VA
- AWS Database Engineer — Guidehouse · US - VA, McLean
- AWS Cloud Engineer — Accenturefederalservices · Tampa, FL
- Azure Data Engineer — Booz Allen Hamilton · Remote
- Cloud Data Warehouse Engineer III — Modivcare · Remote
- Cloud Data Engineer — Norwegian Cruise Line Holdings (NCLH) (NCL Shoreside Careers) · Miami, Florida
More like this: Data Engineer Jobs · Remote Data Engineer Jobs · Browse all jobs
Free ATS checker · How to Tailor Your Resume to a Job Description (Step by Step)