ForgeApply · Job listing
Junior Data Engineer (New York, NY)
Blab
Apply in about a minute — without sacrificing quality.
ForgeApply autofills this application and tailors your resume to this exact posting. You review everything before it's sent. Free trial, no card required.
About this role
This is a Full-Time Role (40 hours per week, 5 days per week) with no option for part-time work. While this is a remote-first opportunity, the candidate filling this role must be a resident of Pennsylvania, New York, or Brazil at the start of employment. Additionally, they must be within commuting distance of our office in Philadelphia, New York City, or São Paulo. Please visit our Careers page to review all opportunities and submit your application for the role(s) that best fit your location and work authorization.
About the Team
The Data & AI team is responsible for building B Lab's data platform, delivering on business priorities, and scaling AI adoption across the organization. The team reports directly to the CTDO. The Data & ML Platforms team owns the data platform, ML infrastructure, and the foundational data models the rest of the Data & AI team depends on — the supply-side capability layer that Business & Data Priorities and AI Enablement build on.
About the Opportunity
As a Junior Data Engineer within the Data & ML Platforms pillar, you build and maintain the data pipelines and infrastructure the rest of the Data & AI team depends on. Your first priority is unblocking the onboarding of new data sources — currently one of the team's top hiring priorities, paused pending this hire. You will absorb data engineering work currently split between the Senior Machine Learning Engineer and the Pillar Lead, freeing them to focus on ML infrastructure and platform strategy respectively. You will work closely with the Senior Analytics Engineer on core data modeling and with the Data Governance Lead on data standards, definitions, and compliance.
Core Responsibilities
Data Pipeline Development & New Source Onboarding (55%):
• Own Pipeline Development End-to-End: Design, build, and maintain robust, scalable ETL/ELT pipelines that reliably deliver clean data to the platform.
• Add New Data Sources: Evaluate, scope, and integrate new data sources as they're identified.
• Partner Across Pillars: Work continuously with Network Priorities, Regional Enablement and AI Enablement to understand incoming data needs and translate them into pipeline work.
• Monitor & Troubleshoot: Proactively identify and resolve pipeline failures and data quality issues before they affect downstream users.
Platform Support & Cross-Team Collaboration (35%):
• Support Foundational Data Models: Work with the Senior Analytics Engineer to maintain the core data models the rest of the team depends on.
• Ensure Data Availability for Consumers: Make sure the data needed by Data Analysts and the Senior Machine Learning Engineer is reliably available.
• Follow Data Governance Standards: Apply the data standards, definitions, and sensitivity classifications set by the Data Governance Lead.
Strategic Innovation & Business Impact (10%):
• Evaluate Pipeline Tooling: Explore and pilot new ETL/ELT tools or approaches that could improve onboarding speed or pipeline reliability.
• Quantify Impact: Track and articulate how pipeline reliability and onboarding speed affect downstream analytics and ML work.
Major Objectives/Project for the role in the first 6-12 months
• Add new data sources to the Data Platform
• Improve data infrastructure allowing for less downtime and more proactive monitoring
• Enable transformation of data and data models
• Improve dependency handling between tables
• Improve data labelling and documentation for AI usage
About You
• A BA/BS in Computer Science, Information Technology, or a related field strongly preferred
• Minimum of 2+ years of experience in data engineering
• Experience working in a DevOps-oriented culture that prioritizes continuous integration and continuous deployment
• Proficiency with Git and collaborative version control workflows (e.g., branching, pull requests, code review)
• Experience with Infrastructure as Code (e.g., Terraform, CloudFormation, or CDK) for provisioning and managing cloud infrastructure
• Proven experience in designing and deploying data solutions
• Experience designing, building, and onboarding new data sources into ETL/ELT pipelines
• Ability and desire to take product/project ownership
• Proficiency in SQL and experience with scripting languages such as Python, Java, or Scala
• Experience with data pipeline and workflow management tools
• Strong knowledge of big data tools and frameworks such as Hadoop, Spark, or Hive is a plus
• Experience using AI coding assistants and other AI tools to improve development speed and productivity
• Excellent communication skills
Compensation Details
B Lab has a compensation plan that includes:
• An annual salary in the range of $62,900 - $83,000 based on experience and skills
• Excellent health benefits package including access to medical, vision and dental coverage
• Paid time off for vacation - in your first year, you’ll start with 15 days (prorated in a to your start date)
• Additional paid time off for organizational closures
• 403(b) with a match of up to 3%
• Unlimited sick and personal time - if you need it, use it
• After your first year of employment, 40 hours paid time off for community service; paid parental leave; and time and budget for your professional development (we assess this PD budget annually)
• A remote-first workplace
• A flexible work environment with the ability to plan your work week around your personal commitments
This is a Full-Time Role (40 hours per week, 5DWW) with no option for part-time work.
This job ad is for New York, NY. While this is a remote-first opportunity, the candidate filling this role must hold U.S. work authorization without any time limitations or any other restrictions, and they must be a resident of New York State at the start of employment. Additionally, they must be within commuting distance of our office in New York City. If you wish to be based in one of our other locations listed for this role, please visit our Careers page and sub
Salary insight
The midpoint of this range ($73k) is about 59% below the median disclosed salary for New York roles listed on ForgeApply ($177k across 4,579 jobs).
See full Data Engineer salary data for New York →
Based on live postings with disclosed pay on ForgeApply; refreshed daily. Not an estimate of this employer's offer.
Ready to apply to Blab?
Apply in about a minuteSimilar jobs
- Junior Data Engineer (Philadelphia, PA) — Blab · Philadelphia, PA metro area
- Junior Data Engineer — Capitaltg · Remote
- Senior Data Engineer — Employerdirecthealthcare · Dallas, TX - Hybrid (3x in office/week)
- Senior Data Engineer — Checkr · Denver, Colorado, United States; San Francisco, California, United States
- Senior Data Engineer — Interrahealth · Remote
- Senior Data Engineer — Coreweave · New York, NY / Sunnyvale, CA / Bellevue, WA
- Senior Data Engineer — Talkspace · New York, NY (Hybrid)
- Senior Data Engineer — Babylist · Remote
More like this: Data Engineer Jobs · Data Engineer Jobs in New York · More jobs at Blab · Browse all jobs