The double category of mixed probabilities [lcc-000T]
The double category of mixed probabilities [lcc-000T]
In probability theory, a random variable is a (measurable) map x: \Omega \to S from some implicit "background space" \Omega , which is taken to be equipped with a probability measure \mathbb {P} --- the image measure x(\mathbb {P}) describes the distribution of x.
Since the background space itself is supposed not to matter too much, one might thing we could just as well replace x with its image distribution, but in fact the explicit presentation plays a very important role, most notably because, given two variables x,y: \Omega \to S, we get to talk about their joint distribution on S^2
(Of course, we could simply study this distribution in itself, but first, this is not always possible for infinite-dimensional cases such as the trajectories of stochastic dynamical systems, and second, it is often more convenient to work with variables than joint distributions)
From a categorical point of view, we can think of "parametrized stochastic maps" as parametrized deterministic maps equipped with a distribution on the parameterizing object. Much of this drops out of well-known constructions: Given any monoidal category \mathcal {C}, the slice category \mathcal {C}_{/I} is again monoidal, and has a forgetful functor to \mathcal {C}, and therefore acts on it. We can therefore form the double category \mathsf {\mathbb Para}_{\mathcal {C}_{/I}}(\mathcal {C}) of parametrized morphisms (see para construction under "Generalizations, and this talk by David Jaz Myers). We will be restrict.ting to the sub-double category of this containing only those horizontal morphisms where the map \Omega \otimes X \to Y is deterministic.
Since we are generally only interested in morphisms up to almost-certain equality, we may want to expand the set of 2-cells, requiring only that the square commutes up to \mathbb {P}-almost certain equality. This has the result of rendering two horizontal maps isomorphic if they are equal almost certainly, but note that for this to have well-defined composition requires some assumptions on \mathcal {C} (see Fritz 2019, Prop. 13.9)
Explicitly, we have the following definition:
Given a parametrized map f: \Omega \otimes X \to Y, \mathbb {P}: I \to \Omega , there is an obvious stochastic map \hat {f}: X \to Y given by composing. This is a sort of half-companion to f: we have a canonical 2-cell filling this square:
This 2-cell has a sort of universal property---if given another 2-cell from the identity horizontal cell 1_A: A = A to f, it must factor over the above one (this amounts to the claim that, given such as 2-cell, the kernel A \to Y is the composition of the kernel A \to X and \hat {f}, which is not hard to see)
In the other direction, we have no hope of building a 2-cell, since that would require two maps X \times \Omega \to Y to be (almost certainly) equal, one of which is given by first applying the deletion \Omega \to I. We shouldn't't expect these two maps to be proper companions, since this would entail that if \hat {f} = \hat {g} then f = g, defeating the whole point of the exercise