The Math Behind Neural Networks, Part 1.4 — Which Loss to Choose, and Why
The output distribution was a choice. Other choices give other costs, each predicting a different statistic — and why cross-entropy beats squared error when the gradient would otherwise fall silent.


