DevFlow Conf 2026

Evaluating Model Output When There Is No Right Answer

15 May 2026, 10:15 – 10:45 · Room 2B · AI Engineering · Talk (30 min)

In clinical summarisation there is no gold answer to score against, and accuracy is the wrong question anyway. How Kestrel Health built an evaluation set clinicians agreed with, what inter-rater disagreement told us about the task, and the failure modes that only ever showed up in review.

Speakers

Materials

← Back to the schedule