Remote
Full time
Remote
Engineering
About NewtonX
NewtonX is a B2B insights company trusted by the world's most innovative companies to make high-stakes decisions with confidence. We combine a verified network of business professionals with AI-powered research tools to deliver research intelligence faster, more precise, and more defensible than traditional methods.
Our clients include Google, Microsoft, TikTok, DoorDash, Stripe, and Coinbase. Our research has been cited by Fortune, Forbes, TechCrunch, Adweek, and the Wall Street Journal.
NewtonX has raised $47M from investors including Two Sigma Ventures, Third Prime, XFund, and Citi Ventures.
About the Role
NewtonX is rapidly expanding into the AI data annotation and RLHF (Reinforcement Learning from Human Feedback) space. We are leveraging our core superpower—recruiting the world’s leading domain experts—to provide high-quality, expert-led data labeling for AI labs and enterprises. Beyond our unmatched B2B recruiting, we utilize powerful, automated project management processes that allow us to scale projects rapidly, adapt to shifting requirements, and manage our subject matter experts with professional, industry-best practices and fair compensation.
This role is the technical owner of data quality. You will partner directly with the Program Lead and act as technical lead in communicating with clients, interpreting their core AI model testing goals and assisting the Program Lead in creating concrete technical specs that will accomplish these goals.
The core question you own: will the data we produce do its job once the customer trains on it or evaluates with it? Clean data that passes every operational gate can still fail — a training set that yields a weak or misleading signal, or an eval that technically runs but doesn't surface the weaknesses that matter. You are the person who can look at an approved submission and say "this is technically correct and still won't do the job, and here's why."
You are hands-on and close to the work. This is a foundational role — what it covers will grow as the business does.
In this role you will:
Training Signal Integrity (core)
Own the judgment of whether task designs and rubrics produce a useful training signal for the consuming method (SFT, RLHF, RLVR, agentic RL, CoT, evals). Catch mismatches between what the data rewards and what the customer is actually training for.
Design the task, environment, and rubric structure for agentic workflows.
Own the ML-level validation of the dataset before delivery — not just whether individual submissions meet spec.
Technical Feedback Loop with Operations
Partner with the Program Lead to convert customer requirements into concrete technical specs: expert profiles, screener trees, task interfaces, task templates, QC rubrics, statistical thresholds.
Partner with the Program Lead to define and defend quality metrics: inter-annotator agreement targets, gold-standard injection rates, statistical power thresholds — providing the statistical and methodological grounding.
Calibrate the ops team on what "good" looks like per engagement; run alignment sessions when standards shift.
Customer Technical Credibility
Serve as the technical counterpart to the customer's ML, applied science, and product teams. Hold your ground on the technical questions that decide data quality across training and evaluation — task and reward design, data quality for SFT and preference methods, RL and RLVR, contamination, statistical rigor, and agentic workflows.
Help diagnose data-quality questions when a customer reports the data underperformed — reason through whether the issue is the data, the quantity, or something in their training setup, and make the case defensibly.
Who you are
Deep applied ML experience centered on post-training human data — you've owned a human-data workstream as an applied scientist or ML engineer.
At least 2 years working hands-on with RL, including how tool-use trajectories are rewarded and evaluated. If you're not fluent in RL, this isn't the role — it's the foundation of the core judgment you'll be making.
Working fluency across modern LLM post-training and evaluation: SFT, RLHF/preference data quality, RLVR, chain-of-thought, eval harness construction, contamination handling, statistical significance, and agentic/tool-use evaluation.
Genuine understanding of how training data becomes model behavior — you can reason about what a model will learn from a given dataset, not just whether the data meets spec.
Strong programming foundation: read and reason about an eval harness, write Python comfortably, work with model APIs, prototype scoring pipelines. Not a production engineer, but not hands-off.
Statistical fluency: you know when an effect is real vs. noise and can defend a sample size or significance threshold.
Client-facing presence: you've defended technical design choices in real time to skeptical audiences and adjusted scope without losing rigor. Range matters — you can talk to a Series B CTO and a Fortune 100 AI lead in the same week.
Strong written communication: methodology sections, technical reports, and specs that hold up to expert review.
If the profile above describes you and your passions, we'd love to hear from you!
What we offer
Massive Impact: Opportunity to have an astounding impact, build a brand new business unit from the ground up, and have direct C-level influence at an extremely fast-growing late-stage startup.
Fast-track career growth: This foundational role will enable you to progress quickly within NewtonX towards commercial and operational leadership.
Comprehensive Benefits: Excellent medical, dental, and vision insurance.
Retirement: 401k match with immediate vesting.
Perks: Health savings/flexible savings account, and pre-tax commuter benefits.
Work-Life Balance: Paid time off: vacation, holidays, sick, and parental leave.
Great Culture: A diverse, collaborative, and positive culture where we invest in and celebrate each other's success (happy hours, team projects, and retreats).
Visa sponsorship is not available for this role.
NewtonX is proud to be an equal opportunity workplace. We do not discriminate based upon race, religion, color, national origin, sex, sexual orientation, gender identity/expression, age, status as a protected veteran, status as an individual with a disability, or any other applicable legally protected characteristics.
Compensation Range: $180K - $260KSign in to browse authentic reviews, anonymous ratings and salary data before you apply.