Differences between revisions 1 and 11 (spanning 10 versions)
Revision 1 as of 2024-06-14 15:43:35
Size: 981
Comment: Initial commit
Revision 11 as of 2026-08-12 22:16:57
Size: 2681
Comment: Notes
Deletions are marked like this. Additions are marked like this.
Line 3: Line 3:
The '''Bernoulli distribution''' is a discrete propability distribution that gives 1 with probability ''p'' and 0 with probability ''q = 1 - p''. The '''Bernoulli distribution''' is a discrete probability density function, specifically giving outcomes 0 or 1.
Line 11: Line 11:
== Statistics == == Description ==
Line 13: Line 13:
The expected value of a Bernoulli-ditributed variable is ''E[X] = p''. The distribution gives outcome 1 with probability ''p'', and 0 with probability ''q = 1 - p''. It is appropriate for modeling any binary event.
Line 15: Line 15:
The variance of a Bernoulli-distributed variable is ''Var[X] = p(1 - p) = pq''. A variable distributed this way is notated like ''X ~ Bernoulli(p)''. (Sometimes shortened to 'Bern'.)
Line 17: Line 17:
The sum of repeated and independent Bernoulli-distributed events are described by the [[Statistics/BinomialDistribution|binomial distribution]]. The sum of repeated and independent Bernoulli-distributed events are described by the [[Analysis/BinomialDistribution|binomial distribution]].
Line 23: Line 23:
== Sampling == == Moments ==
Line 25: Line 25:
If all frame listings have an equal probability of selection, sampling can be implemented like: The [[Analysis/ExpectedValue|expected value]] is given as ''E[X] = p''.

[[Analysis/Variance|Variance]] is given as ''Var[X] = p(1 - p) = pq''. Observe that the standard calculation for variance of a discrete quantitative random variable simplifies quickly, since there are only two possible values:

{{attachment:var1.svg}}

Furthermore note that ''E[X] = p'', that ''Prob(X=1) = p'', and that ''Prob(X=0) = 1 - p'':

{{attachment:var2.svg}}

It should be clear that this simplifies to ''p(1-p)'' immediately.

----



== Usage ==



=== Bernoulli Trials ===

Consider a Bernoulli random variable with probability ''p''. A single trial to observe this variable is called a '''Bernoulli trial'''. The sum of repeated (independent) Bernoulli trials is known to follow a [[Analysis/BinomialDistribution|binomial distribution]]. In other words, if ''X ~ Bernoulli(p)'' then ''nX̅ ~ Binomial(n,p)''.



=== Sampling ===

Selection for an experiment can be conceptualized as a Bernoulli event.

The natural implementation of a sample is:
Line 34: Line 64:
The expected number of cases sampled is ''np''; the sample size is described by the [[Statistics/BinomialDistribution|binomial distribution]]. The sample size ''np'' is known to follow a [[Analysis/BinomialDistribution|binomial distribution]].



=== Conservative Confidence Interval ===

The variance of a Bernoulli random variable, i.e. ''p(1 - p)'', is maximized by ''p = 1/2''. (The same is true for a [[Analysis/BinomialDistribution|binomial]] random variable; ''Var[X] = np(1 - p)''.) Therefore ''Var[X] ≤ 1/4''. This leads to the concept of a '''conservative confidence interval''' that simply assumes the 'worst case' and sets variance to ''1/4''.

Following from the '''De Moivre-Laplace theorem''', with a large enough ''n'', a binomial random variable is approximately [[Analysis/NormalDistribution|normal]] as ''nX̅ ~ N(X̅, X̅(1-X̅)/n)''.

This leads to a simple formula for the '''Wald interval''':

{{attachment:wald.svg}}

Bernoulli Distribution

The Bernoulli distribution is a discrete probability density function, specifically giving outcomes 0 or 1.


Description

The distribution gives outcome 1 with probability p, and 0 with probability q = 1 - p. It is appropriate for modeling any binary event.

A variable distributed this way is notated like X ~ Bernoulli(p). (Sometimes shortened to 'Bern'.)

The sum of repeated and independent Bernoulli-distributed events are described by the binomial distribution.


Moments

The expected value is given as E[X] = p.

Variance is given as Var[X] = p(1 - p) = pq. Observe that the standard calculation for variance of a discrete quantitative random variable simplifies quickly, since there are only two possible values:

var1.svg

Furthermore note that E[X] = p, that Prob(X=1) = p, and that Prob(X=0) = 1 - p:

var2.svg

It should be clear that this simplifies to p(1-p) immediately.


Usage

Bernoulli Trials

Consider a Bernoulli random variable with probability p. A single trial to observe this variable is called a Bernoulli trial. The sum of repeated (independent) Bernoulli trials is known to follow a binomial distribution. In other words, if X ~ Bernoulli(p) then nX̅ ~ Binomial(n,p).

Sampling

Selection for an experiment can be conceptualized as a Bernoulli event.

The natural implementation of a sample is:

scalar p = .2 /* Probability of selection */
set seed 123456789
generate double r = runiform()
generate sampled = (r < p)

The sample size np is known to follow a binomial distribution.

Conservative Confidence Interval

The variance of a Bernoulli random variable, i.e. p(1 - p), is maximized by p = 1/2. (The same is true for a binomial random variable; Var[X] = np(1 - p).) Therefore Var[X] ≤ 1/4. This leads to the concept of a conservative confidence interval that simply assumes the 'worst case' and sets variance to 1/4.

Following from the De Moivre-Laplace theorem, with a large enough n, a binomial random variable is approximately normal as nX̅ ~ N(X̅, X̅(1-X̅)/n).

This leads to a simple formula for the Wald interval:

wald.svg


CategoryRicottone

Analysis/BernoulliDistribution (last edited 2026-08-12 22:16:57 by DominicRicottone)