Evidence

Naming a feature is easy. Showing that the name is correct takes converging evidence. Two views of a candidate feature are cheap to gather and often disagree with each other.

Feature visualization

Feature visualization optimizes an input to increase a component's activation. Dataset examples provide a second view by showing real images that activate it. Agreement between the two is stronger than either alone.

Evidence for features: feature visualization and dataset examples
Evidence for features: feature visualization and dataset examples [Source]

High-low frequency detectors look for low-frequency patterns on one side of their receptive field, and high-frequency patterns on the other side. These are typically found in families of features that look for the same thing in different orientations.

High-low frequency detectors responding to low-frequency and high-frequency patterns on opposite sides of the receptive field
High-low frequency detectors look for low-frequency patterns on one side of their receptive field, and high-frequency patterns on the other side. [Source]

Another evidence for features is the pose invariant dog head detector. Feature visualization allows us to establish a causal link (shown on the left), while dataset examples test the neuron's use in practice and whether there are a second type of stimuli that it reacts to (shown on the right).

Pose invariant dog head detector: feature visualization on the left, dataset examples on the right
The pose invariant dog head detector. Feature visualization establishes a causal link (left), while dataset examples test the neuron's use in practice (right). [Source]

Some features cross surface form. A multimodal model may respond to photographs of Spider-Man, drawings of Spider-Man, and the word “spider.” The shared representation is more abstract than any one pixel pattern.

High-level features can also represent tone, relationships, political activity, time of day, or an organization. Their breadth makes them powerful and harder to name cleanly.

Multimodal features across images, drawings, and text
Large neural networks trained on multimodal data (text and images) exhibit multimodality, showing one concept across several surface forms, similar to human neurons. [Source]

Abstraction

An abstract feature responds to a shared concept across varied low-level forms. Its invariances are part of its meaning: which changes leave the feature active?