hirq
← All jobs

Whisk

AI Quality & Evaluation Lead

Poland · Remote · Temporary · Product

Apply well, not just fast

Create a free account and upload your resume to get a match score, keyword gaps, a tailored resume, a cover letter and interview prep for this job.

About the role

LLMsExcelProduct ManagementUser ResearchUX Design
About Samsung Food (Whisk) At Samsung Food (you might know us as Whisk), we are pioneering the future of everyday care and health — helping individuals and families around the world live better through connected, proactive, and personalized experiences. Our mission is to connect food, health, and home across the Samsung ecosystem, turning data from your devices into personalized guidance that supports you and your loved ones. We work at the intersection of AI, nutrition, behavior change, and digital health. The Samsung Food app was included in Google Play’s Best of 2020 Everyday Essentials list, has been regularly featured on the Apple App Store, and was nominated for a 2021 Webby Award. We’ve also launched Vision AI, which recognises food and meals using our global Food Genome, and developed deep integrations with Samsung Health and SmartThings. We are a remote-first team of more than 100 people in over 30 countries, united by our commitment to innovation and impact. As part of Samsung Electronics, we combine the agility of a startup with the scale of the world’s largest consumer electronics company. Join us and help build technology that improves daily life for tens of millions of people across phones, wearables, TVs, and kitchen appliances in the years to come. Learn https://samsungfood.com/working-at-samsung-food/ more about how we’ve shaped a high-performing global team over the past 14 years. Summary Overview We're building an AI coaching experience inside Samsung Health. Every piece of coaching a user sees is generated, not written. Our platform partner owns the evaluation runtime and functional evals (did the pipeline execute, did the tool call succeed). Nobody owns the layer above that: is the output actually good — right framing, right tone, right structure, does it deliver the coaching logic, does it feel like it's talking to me. This engagement exists to build that layer: author the failure taxonomy, codify it into a binary rubric, validate an LLM judge against your own human grades, and hand the whole thing to an internal owner. We have a v0.1 rubric and golden scenarios from our product lead, and a functional rubric owned by engineering. We don't have the experience rubric — that's what this engagement produces. Deliverables by Phase (Overall project ~13 weeks) 1. Failure taxonomy and golden set (within 4 weeks from the contract start date) At least 150 traces hand-graded with open-coded notes. At least 40 of them come from sparse-data synthetic profiles. The taxonomy has 5–10 categories, each with a frequency count, and the top three failures are shown with their share of all failures. Saturation is evidenced: the last 20 traces produced no new category. At least 50 traces are frozen as a holdout. Head of Product signs off the taxonomy. 2. Rubric v1 and calibration (within 4 weeks from phase 1 completion) One binary pass/fail criterion per taxonomy category, each with a definition, a pass example and a fail example. Written annotation guidelines. At least 30 traces are double-coded with a second grader, and every disagreement is logged and resolved in writing. A guideline revision log records each change and its reason. Head of Product signs off the rubric. 3. Validated judge and weekly readout (within 4 weeks from phase 2 completion) Judge prompts for every criterion. A validation report giving true positive rate and true negative rate per criterion on both the dev set and the holdout. A weekly readout on helpfulness, relevance and tone, run at least twice: once by the consultant, and once by the internal owner with the consultant shadowing. A documented revalidation routine, triggered by every model or prompt change and run quarterly regardless. 4. Playbook and handover test (within 1 week from phase 3 completion) A written playbook covering the whole method: grading, updating the taxonomy, revising the rubric, revalidating the judge and running the readout. The handover test passes when the internal owner reruns the judge validation unaided, gets within ±5 points, and runs one weekly readout with no consultant involvement. What This Role Is Not - Not building eval infrastructure, harnesses, or dashboards (owned elsewhere) - Not functional/module-level evals or accuracy metrics - Not clinical safety validation - Not high-volume annotation — grading a few hundred traces to build the taxonomy; volume work goes to vendors/the judge once the rubric exists What We're Looking For You've run this loop end-to-end at least once on conversational or generated-text output: error analysis on real traces → a failure taxonomy you built yourself → binary criteria and annotation guidelines → an LLM judge validated against your own labels. Background: conversation design, AI product quality, model behaviour/policy, content design, human data operations, UX research with strong qualitative coding, or RLHF. Applied linguistics and product management backgrounds also fit. You'll be strong in this role if you: - Can describe a failure mode you personally discovered by reading output — one nobody had named before you - Are comfortable being the arbiter and documenting what you overruled and why - Understand the rubric is discovered through grading, not written upfront - Can explain why a judge that agrees with you 94% of the time might still be useless - Write annotation guidelines that don't need interpreting - Are comfortable working in notebooks and spreadsheets (no production code required) Not required: domain expertise in nutrition, weight management, or behaviour change — we have that in-house.