ForgeApply · Job listing
Principal Data Scientist
Wiley
See all 17 open roles at Wiley →
Tailor your resume for this Wiley job in about a minute.
ForgeApply tailors your resume and cover letter to this exact posting, then hands you a ready-to-submit application for Wiley's site. Free trial, no card required.
About this role
Job Description:
We believe in bold ideas, diverse perspectives, and the drive to transform knowledge into impact. Here, your curiosity fuels progress, your voice shapes innovation, and your ambition helps redefine what’s possible within science and learning. We are a culture that obsesses over impact, challenges, and drives what’s next to power infinite possibilities for our customers, colleagues and society at large.
About the Role:
We're building the systems that turn one of the world's largest scientific corporation into research intelligence. That means production NLP pipelines running over millions of journal articles, extracting entities, classifications, claim tuples, and summaries optimized for use by downstream agentic applications . We're looking for a principal data scientist to own domain-specific content modeling work end to end, from the eval set through the pipeline stage that ships it.
You'll join a small, senior team where data scientists own their models in production. You'll write the code, own the evaluations, ship the changes, and stay accountable for the outcomes. This is a hands-on role for someone who wants to see their models through to real users in a rapidly evolving market .
Job Responsibilities:
• Design and build NLP enrichment pipelines that extract entities, classifications, claims, and summaries from scientific full-text at scale.
• Compare NLP approaches to extraction and enrichment against LLM-based approaches, and pick the right tool for each task. That means putting traditional NLP (NER, sequence labeling, classification), embedding-based retrieval, LLM prompting, and fine-tuned smaller models on the same table, and defending each choice with evaluation, cost, and operational tradeoffs. This is a core part of the job, not an occasional exercise.
• Own evaluation. Build the golden sets in consultation with SMEs and vendors, choose the metrics, and make productive tradeoffs between speed, quality, and cost.
• Write production-quality Python. Manage concurrency and cost for high-volume LLM workloads. Structure code that engineers can ship and other data scientists can extend.
• Collaborate with a team of data engineers to o rchestrate work in data pipeline and data build tools like Airflow and Dagster . Design idempotent, retryable , evaluable pipeline stages that stay reliable when a run fails at scale.
• Contribute to agentic AI application work: tool-using systems that reason over the enriched corpus, where your NLP and evaluation background will shape how the agent grounds and defends its answers.
• Work directly with editors, product managers, and engineers. Bring the modeling perspective into product decisions, and translate stakeholder pushback into concrete modeling work.
Job Requirements:
• Deep Python. You've written it in production, at scale, for years. You know when to reach for asyncio versus threads versus a queue, and you can explain the tradeoff clearly.
• Strong NLP background across modern (LLMs, transformers, embeddings, retrieval) and classical (NER, classification, sequence labeling) approaches. You've built evaluations and learned from the results .
• A habit of comparing approaches and choosing the right one for the task. You can defend "prompt a large LLM" and "train a small classifier on 2,000 labels" with equal seriousness, back the choice with an eval and a cost estimate, and know what to do when performance drifts.
• A track record of shippin g – not just prototypes and papers, but systems that deliver value to real users.
Preferred:
• Experience working with scientific or scholarly text.
• Familiarity with AWS (S3, Batch, Lambda, SageMaker) and Parquet or Iceberg data lake patterns.
• Experience running LLMs under real cost and latency budgets in production.
• Some exposure to agentic AI applications: tool use, multi-step reasoning, guardrails, and evaluation of trajectories rather than single-turn outputs.
We power infinite possibilities.
For more than 200 years, we've transformed knowledge into discoveries that shape the world. Today, our global team of innovators, creators, and experts is driving what's next in science, education, and publishing—creating impact that reaches everywhere. We're not just observers of progress. We're the ones accelerating scientific breakthroughs, advancing learning, and sparking innovation that redefines entire fields and improves lives.
Here, your talent matters. Your ideas have room to grow. And your work creates breakthroughs that can change everything.
Wiley is an equal opportunity/affirmative action employer. We evaluate all qualified applicants and treat all qualified applicants and employees without regard to race, color, religion, sex, sexual orientation, gender identity or expression, national origin, disability, protected veteran status, genetic information, or based on any individual's status in any group or class protected by applicable federal, state or local laws. Wiley is also committed to providing reasonable accommodation to applicants and employees with disabilities. Applicants who require accommodation to participate in the job application process may contact tasupport@wiley.com for assistance.
We are proud that our workplace promotes continual learning and internal mobility. Our values support courageous teammates, needle movers, and learning champions all while striving to support the health and well-being of all employees. We offer meeting-free Friday afternoons allowing more time for heads down work and professional development, and through a robust body of employee programing we facilitate a wide range of opportunities to foster community, learn, and grow.
We are committed to fair, transparent pay, and we strive to provide competitive compensation in addition to a comprehensive benefits package. The range below re
Salary insight
This posting doesn't disclose pay. Across 10,065 New York jobs with disclosed salaries on ForgeApply, the median is $160k.
See full Data Scientist salary data for New York →
Based on live postings with disclosed pay on ForgeApply; refreshed daily. Not an estimate of this employer's offer.
Tailor your resume for this Wiley role before you apply.
Tailor my resume for this jobSimilar jobs
- Principal Data Scientist — Databricks · Mountain View, California; San Francisco, California
- Principal Data Scientist — Veracyte · Remote
- Principal Data Scientist — Crate & Barrel · Remote
- Principal Data Scientist — Warner Bros. Discovery · CA San Francisco 153 Kearny Street | NY New York 30 Hudson Yards | WA Bellevue 205 108th Avenue NE, Suite S200
- Principal Data Scientist — Columbia · Portland, Oregon, United States
- Principal Data Scientist — Red Hat · Remote
- Principal Data Scientist — Blinkhealth · New York, NY
- Principal Data Scientist — Autodesk · San Francisco, CA
More like this: Data Scientist Jobs · Remote Data Scientist Jobs · Data Scientist Jobs in New York · Browse all jobs
Free ATS checker · No Salary on the Job Posting? How to Find the Number Before You Interview