Uncertainty heads for NNUE

So the normal NNUE loss is this:

The idea is to instead have the network produce two outputs:

Then make the loss encompass both minimising MSE and predicting it: