LGB 2
I’m reading The Laws Governing Belief. This note is on What is Evidence?.
This article is saying three things: evidence needs to be causally related (“entangled”) to the target of inquiry; evidence is valid insofar as it can accurately distinguish between different states that the target of inquiry can be in; and beliefs well-processed are contagious, such that the processor is also usefully causally entangled with the inquiry.
The following paragraph confused me:
To say it abstractly: For an event to be evidence about a target of inquiry, it has to happen differently in a way that’s entangled with the different possible states of the target. (To say it technically: There has to be Shannon mutual information between the evidential event and the target of inquiry, relative to your current state of uncertainty about both of them.)
I spent way too much time trying to understand the relevant information/probability theory to know what this means. I still don’t understand it exactly. My current understanding is this.
Think of my current beliefs about a target of inquiry as a probability distribution over t and not-t. The entropy of the distribution \(H(T)\) represents my uncertainty about the inquiry; the conditional entropy of T given the evidence E \(H(T|E)\) represents my expected uncertainty after I see the presence or absence of the evidence. The mutual information \(I(T; E) = H(T)-H(T|E)\) tells me how much, on average, seeing the evidence will reduce my uncertainty about the target.
If \(H(T|E)=0\), that means that the evidence perfectly distinguishes between different states of the target, such that seeing the presence of evidence would make me perfectly certain of one state of the target and absence of evidence would make me perfectly certain of another. If \(H(T|E)=H(T)\), then the evidence has no distinguishing value: I’m left with the probability distribution that I started with. (Equivalently, the distributions for \(E\) when \(T=t\) and \(T=\neg t\) are the same.)
Mutual information \(I(T; E)\) ranges between \([0, H(T)]\): you can remove anywhere from 0 uncertainty to all the uncertainty that you possessed. E is evidence about T when the mutual information is greater than 0.
Something that initially confused me was why we could subtract conditional entropy from entropy. The entropy of a probability distribution and the expected entropy of conditional distributions felt like very different things. But the following analogy helped. Think of \(H(T)\) as my “expected uncertainty” for T, conditional on no evidence, and \(H(T|E)\) as my expected uncertainty for T, conditional on E. (In the former case, imagine conditioning on null evidence, evidence that has no information for the target of inquiry.) Then, both of them are in the form of conditional entropy, and I feel happy. (Presumably, I can also think of these simply as “my expected uncertainty about T before learning E” and “my expected uncertainty about T after learning E”, but eh.)
“Beliefs are contagious when processed” is also an interesting phrase. The idea is that, given a “correct processor” (what does this mean? I’m assuming for now that it just implies that the belief and the processor have positive mutual information), observing something modifies the processor; “observing” the processor by some other processor (say, by communicating your beliefs via language) modifies the other processor; and so on. Shoelaces untied -> light -> eyes -> brain -> words -> other brain.