ForgeApply · Job listing
Data Engineer
Matterworks
Apply in about a minute — without sacrificing quality.
ForgeApply autofills this application and tailors your resume to this exact posting. You review everything before it's sent. Free trial, no card required.
About this role
ABOUT US
Most of the molecules driving human biology are invisible to us. Mass spectrometers already detect metabolites, lipids, and peptides, but the vast majority of those signals never get identified. A typical experiment names a small fraction of its features and discards the rest. We call this biology's dark matter. It’s signal-rich and mechanism-defining, yet almost entirely opaque.
Matterworks is building the foundation models that make that dark matter legible. Our Large Spectral Models do for biochemical biology what AlphaFold and ESM did for proteins: turning a library-bound discipline into something predictable and generative, and embedding it at every stage of R&D.
Come build the future of biological discovery with us.
Position Overview
Matterworks is seeking a Data Engineer to build and run the pipelines behind our models. Our platform acquires mass spectrometry and molecular data at scale and turns it into the datasets our AI team trains on and our product reasons over. You will own real pieces of that path end to end.
Our data serves two different customers. The AI team needs training corpora that are complete, correctly split, and reproducible. The product and agentic layer needs values a scientist can explain and we can safely show a customer. You will build inside the contracts and checks that keep both honest, and you will help extend them.
This is a hands-on role with a clear growth path. You will start by owning well-scoped pipelines and datasets and grow toward owning larger parts of the platform. You will report to the Head of Engineering and work daily with our machine learning researchers, scientists, and product team.
Key Responsibilities
- Build and Operate Pipelines: Own well-scoped pipelines end to end: designed, tested, instrumented, documented, and running on a schedule.
- Labels and Enrichment: Turn raw data into datasets people can actually use, with consistent schemas, trustworthy metadata, and documented definitions.
- Data Quality Checks: Design, build and extend our quality checks so that each build gets compared against the last one before it publishes, and a failure stops the pipeline instead of shipping.
- Ingest and Acquisition: Bring new public and partner datasets into the platform: fetching, converting, validating, and reconciling them against what we already hold. Expect messy scientific and vendor formats and file that require continuous improvements to our systems to handle at scale.
About You
- 2+ years of professional experience building data pipelines in production.
- Proficient in Python and SQL.
- Working knowledge of cloud data infrastructure. We run Argo Workflows and Metaflow on EKS, Glue and Athena over Apache Iceberg and Parquet, DuckDB, and Terraform. Depth in any comparable stack transfers fine.
- Demonstrated experience owning a pipeline or dataset end to end, including the tests, the monitoring, and the failures.
- Comfort with messy data and messy formats, and the patience to track down why two sources disagree.
- Daily use of AI coding tools, paired with healthy skepticism about their output on questions of production data correctness.
- Clear written communication, particularly when explaining what broke and what you changed.
- Curiosity about the science. Experience in life sciences, biotechnology, or biochemistry is a plus but not a requirement, and you will work alongside strong in-house chemistry every day.
- A passion for contributing to an early-stage startup where autonomy, eagerness to learn, and enthusiasm for solving novel scientific challenges prevail over rigid processes and egos.
WORKING AT MATTERWORKS
Given the cross-disciplinary and innovative nature of our work, effective collaboration and communication are critical to our progress. We operate in a flexible hybrid model that accommodates both fully remote team members and those who work full-time from our Somerville, MA office. While some positions may require regular in-person presence for hands-on work or local collaboration, many roles can be performed remotely with team members distributed across various locations.
COMPENSATION AND BENEFITS
Matterworks offers full-time employees a competitive base salary, stock options, and benefits (health & dental, vision, long- and short-term disability, life insurance, 401k with company match). Employees enjoy a flexible work & unlimited time away policy, commuter benefits and parking, regular team meals and outings, and company support for continued education/coursework and conference participation.
Matterworks, Inc. is an equal opportunity employer. All candidates for employment at Matterworks are considered without regard to race, color, religion, national origin, age, sex, marital status, ancestry, physical or mental disability, veteran status, gender identity, sexual orientation, or any other category protected by law.
Ready to apply to Matterworks?
Apply in about a minuteSimilar jobs
- Data Engineer — M9solutions · Springfield, VA - TS/SCI clearance required
- Data Engineer — Lightningai · New York, New York, United States
- Data Engineer — Base-power · Austin, TX
- Data Engineer — Capitaltg · Remote
- Data Engineer — Bruntworkwear · North Reading, MA
- Data Engineer — Vardaspace · El Segundo, California, United States
- Data Engineer — M9solutions · Bethesda, MD - TS/SCI clearance required
- Data Engineer — Candidhealth · San Francisco
More like this: Data Engineer Jobs · Remote Data Engineer Jobs · More jobs at Matterworks · Browse all jobs