Data

Transparent AI starts with transparent data. A model inherits the collection choices, label definitions, missing values, measurement errors, sampling gaps, and privacy constraints of the data used to build or operate it.

Data transparency asks practical questions: Who is represented? Who is missing? What does each label mean? How often are labels wrong? Which features are proxies for sensitive traits? Which fields should never be used for this purpose?

Fill a Datasheet for Hospital Readmission Data

Select a datasheet section to read the question it must answer, then mark it filled once your team has a real answer — not a placeholder.

Motivation
Composition
Coverage
Missingness
Privacy constraints
Maintenance plan

Datasheet coverage

33%

Coverage

Which patients, clinics, or regions are underrepresented?

If left unanswered:

This dataset underrepresents rural clinics and out-of-network care, so rural predictions are unreliable.

A datasheet is not a formality. 'This dataset underrepresents rural clinics' is a specific, testable warning. 'May not generalize' is not.

Datasheets for datasets and data cards are common ways to record this information. They should describe motivation, collection process, composition, recommended use, known limitations, maintenance, and ethical constraints.

Data transparency is not a promise that the dataset is perfect. It is a map of what is known, what is uncertain, and what users should not assume.

Checkpoint

Why should a data card describe missingness instead of only reporting dataset size?