Evaly v praxi: od narativu ke ground truth
A 45-minute hands-on workshop with Veronika Kráľ. What LLM evals look like and what they are actually for, on one small real project: a bot that scores rental listings against a written brief. We walk through the three ways we got the grading wrong — one holistic narrative, then weighted criteria with rubrics, then ground truth for what the model can't read — and what each mistake taught us.
The worksheet we handed out — Flat Scout: pracovní list, the exercise that turns the two-page narrative brief into weighted criteria with rubrics — is a printable PDF ↗. The deck is downloadable too ↗.
Want this talk for your audience?Invite me to speak ↗
