Neurons and Polysemanticity

So far, we have presented the ideal situation, where a concept is mapped cleanly to a neuron. Curve detectors and pose-invariant dog detectors make it appear like the story is simple. However, this is not always the case, as we see in the below figure. Our car detector spreads into many seemingly unrelated neurons.

Cars in superposition
The car detector neuron exhibits superposition, where it spreads into seeminly unrelated neurons. [Source].

Monosemantic neuron

A monosemantic neuron responds to one coherent feature across the relevant input distribution. Its activation has a stable interpretation.

One Neuron, One Feature? Toggle the Scenario

The clean story maps one neuron to one feature. Toggle to the real scenario to see why that mapping breaks down: a single neuron can drive several unrelated features, while another neuron's role stays unclear.

Neuron 1Neuron 2Neuron 3Feature 1Feature 2Feature 3
In the ideal scenario, each neuron has a single, stable, nameable meaning.
Checkpoint

What evidence would justify calling a neuron monosemantic?

A single neuron's top activating images showing cat faces, fronts of cars, and cat legs
A polysemantic neuron responding to cat faces, car fronts, and cat legs. [Source].

In the figure above, the same neuron responded to cat faces, fronts of cars, and cat legs. Those categories do not provide a satisfying single label. The neuron is polysemantic: several features share the same coordinate.

Polysemanticity

A neuron is polysemantic when its activation reflects multiple distinct features. The features may be represented by different directions that all use that neuron.

Checkpoint

Why can a polysemantic neuron still participate in reliable computation?