Blueprint 01
RLHF preference data project
A preference dataset for reward modeling, built by experts in the target domain.
Workflow
- 01Scoping call and NDA, then written task definition agreed with the lab
- 02Guideline drafting, with edge cases and refusal criteria documented
- 03Calibration round on a small task set, reviewed jointly with the lab
- 04Expert staffing from the vetted network, matched by domain
- 05Live ranking production in batches, with rationales captured per comparison
- 06Delivery in the lab's schema, with a guideline changelog
Quality gates
- —Calibration sign off before live production begins
- —Independent QC review on a second pass of every batch
- —Random audits across contributors throughout the project
- —Inter annotator agreement checks on overlapping items