Continuous binomial distribution

From HandWiki
Short description: Continuous probability distribution on the unit interval

Continuous binomial (cobin)
Parameters

θ∈ℝ (natural parameter)

λ∈{1,2,3,…} (inverse dispersion)
Support x∈[0,1] if λ=1, x∈(0,1) if λ≥2
PDF

f(x;θ,λ)=h(x;λ)exp⁡(λθx−λB(θ)),0≤x≤1,
with B(θ)={log⁡(eθ−1θ),θ≠0,0,θ=0, and

h(x;λ)=λ(λ−1)!∑k=0λ(−1)k(λk)max⁡(0,λx−k)λ−1.
Mean B′(θ)={eθeθ−1−1θ,θ≠0,12,θ=0,
Variance 1λB″(θ)={1λ(1θ2−eθ(eθ−1)2),θ≠0,112λ,θ=0.

In probability theory and statistics, the continuous binomial distribution (also called the cobin distribution) is a family of continuous probability distributions on the unit interval that belongs to an exponential dispersion family. It was introduced as a response distribution for generalized linear models for continuous proportional data, proposed as an alternative to beta regression.[1] The special case λ=1 coincides with the continuous Bernoulli distribution[2].

Definition

A random variable X is said to follow a continuous binomial (cobin) distribution with natural parameter θ and inverse dispersion λ∈{1,2,…}, written X∼cobin(θ,λ−1), if it has density on [0,1] given by

f(x;θ,λ)=h(x;λ)exp⁡(λθx−λB(θ)),0≤x≤1,

where the log-partition function is

B(θ)={log⁡(eθ−1θ),θ≠0,0,θ=0,

and the base measure h(x;λ) is

h(x;λ)=λ(λ−1)!∑k=0λ(−1)k(λk)max⁡(0,λx−k)λ−1,

with h(x;1)=1. The function h(x;λ)/λ coincides with the probability density function of the Irwin–Hall distribution with parameter n=λ, evaluated at λx.

When λ is fixed, the cobin distribution belongs to a one-parameter natural exponential family in θ.

  • Bates distribution: when θ=0, the density reduces to h(x;λ), corresponding to the distribution of the mean of λ independent Uniform(0,1) random variables (equivalently, a scaled Irwin–Hall distribution or Bates distribution).
  • If X1,…,Xλ are independent and identically distributed continuous Bernoulli random variables with common natural parameter θ, then
X¯=1λ∑i=1λXi∼cobin(θ,λ−1).

Properties

Mean and variance

The mean and variance of X∼cobin(θ,λ−1) can be expressed in terms of derivatives of B(θ):

  • E⁡(X)=B′(θ)=eθeθ−1−1θ, for θ≠0.
  • Var⁡(X)=1λB″(θ)=1λ(1θ2−eθ(eθ−1)2), for θ≠0.

If θ=0, then E⁡(X)=12 and Var⁡(X)=112λ.

Sufficient statistic for the mean

If X1,…,Xn are independent and identically distributed continuous binomial random variables with common natural parameter θ and fixed inverse dispersion parameter λ, then the sample mean

X¯=1n∑i=1nXi

is a sufficient statistic for θ.

This is in contrast with the beta distribution: under a mean–precision parameterisation Xi∼Beta(μϕ,(1−μ)ϕ) with fixed ϕ, a sufficient statistic for the mean μ is

∑i=1nlog⁡(Xi1−Xi),

not the sample mean X¯.

Applications

The cobin distribution has been proposed as a response distribution for generalized linear models of continuous proportional data, as an alternative to beta regression, including extensions with random effects.

  1. ↑ Lee, Changwoo J.; Dahl, Benjamin K.; Ovaskainen, Otso; Dunson, David B. (2026-05-18). "Scalable and robust regression models for continuous proportional data". Journal of the American Statistical Association. doi:10.1080/01621459.2026.2626081. ISSN 0162-1459. PMID 42169758. PMC 13188389. https://pmc.ncbi.nlm.nih.gov/articles/PMC13188389/. 
  2. ↑ Loaiza-Ganem, Gabriel; Cunningham, John (2019). "The continuous Bernoulli: fixing a pervasive error in variational autoencoders". Advances in Neural Information Processing Systems (Curran Associates, Inc.) 32. https://proceedings.neurips.cc/paper/2019/hash/f82798ec8909d23e55679ee26bb26437-Abstract.html.