Human Usefulness
An analyst, doctor, and loan applicant can look at the same explanation and need three different things. Technical quality does not settle whether the explanation helps a person act.
Human usefulness must be evaluated with the intended user and task.
Task-based evaluation
Useful measures include decision accuracy, error detection, calibration, time to decision, appropriate reliance, ability to contest an outcome, and cognitive load. Preference ratings alone are weak because people often prefer confident, colorful explanations.
Match the Explanation to the Job
Choose a user. The highlighted explanation format is the one with the highest task-success estimate, because usefulness depends on what the person must do next.
Task: Identify a feasible next step.
heatmap
31%
task success
feature list
55%
task success
recourse
93%
task success
A doctor may need evidence linked to the image and patient record. An applicant may need actionable recourse. An auditor may need reproducibility and logs.
Test whether the explanation improves the real task compared with a strong baseline such as the prediction and confidence alone.
What is a better usefulness test than asking users which explanation they like?