Human Usefulness

An analyst, doctor, and loan applicant can look at the same explanation and need three different things. Technical quality does not settle whether the explanation helps a person act.

Human usefulness must be evaluated with the intended user and task.

Task-based evaluation

Useful measures include decision accuracy, error detection, calibration, time to decision, appropriate reliance, ability to contest an outcome, and cognitive load. Preference ratings alone are weak because people often prefer confident, colorful explanations.

Match the Explanation to the Job

Choose a user. The highlighted explanation format is the one with the highest task-success estimate, because usefulness depends on what the person must do next.

Task: Identify a feasible next step.

heatmap

31%

task success

feature list

55%

task success

recourse

93%

task success

Preference is not enough. Evaluate whether the explanation improves a real decision, error check, or recourse task.

A doctor may need evidence linked to the image and patient record. An applicant may need actionable recourse. An auditor may need reproducibility and logs.

Test whether the explanation improves the real task compared with a strong baseline such as the prediction and confidence alone.

Checkpoint

What is a better usefulness test than asking users which explanation they like?