Share to: share facebook share twitter share wa share telegram print page

Hypergeometric distribution

Hypergeometric
Probability mass function
Hypergeometric PDF plot
Cumulative distribution function
Hypergeometric CDF plot
Parameters
Support
PMF
CDF where is the generalized hypergeometric function
Mean
Mode
Variance
Skewness
Excess kurtosis

MGF
CF

In probability theory and statistics, the hypergeometric distribution is a discrete probability distribution that describes the probability of successes (random draws for which the object drawn has a specified feature) in draws, without replacement, from a finite population of size that contains exactly objects with that feature, wherein each draw is either a success or a failure. In contrast, the binomial distribution describes the probability of successes in draws with replacement.

Definitions

Probability mass function

The following conditions characterize the hypergeometric distribution:

  • The result of each draw (the elements of the population being sampled) can be classified into one of two mutually exclusive categories (e.g. Pass/Fail or Employed/Unemployed).
  • The probability of a success changes on each draw, as each draw decreases the population (sampling without replacement from a finite population).

A random variable follows the hypergeometric distribution if its probability mass function (pmf) is given by[1]

where

  • is the population size,
  • is the number of success states in the population,
  • is the number of draws (i.e. quantity drawn in each trial),
  • is the number of observed successes,
  • is a binomial coefficient.

The pmf is positive when .

A random variable distributed hypergeometrically with parameters , and is written and has probability mass function above.

Combinatorial identities

As required, we have

which essentially follows from Vandermonde's identity from combinatorics.

Also note that

This identity can be shown by expressing the binomial coefficients in terms of factorials and rearranging the latter. Additionally, it follows from the symmetry of the problem, described in two different but interchangeable ways.

For example, consider two rounds of drawing without replacement. In the first round, out of neutral marbles are drawn from an urn without replacement and coloured green. Then the colored marbles are put back. In the second round, marbles are drawn without replacement and colored red. Then, the number of marbles with both colors on them (that is, the number of marbles that have been drawn twice) has the hypergeometric distribution. The symmetry in and stems from the fact that the two rounds are independent, and one could have started by drawing balls and colouring them red first.

Note that we are interested in the probability of successes in draws without replacement, since the probability of success on each trial is not the same, as the size of the remaining population changes as we remove each marble. Keep in mind not to confuse with the binomial distribution, which describes the probability of successes in draws with replacement.

Properties

Working example

The classical application of the hypergeometric distribution is sampling without replacement. Think of an urn with two colors of marbles, red and green. Define drawing a green marble as a success and drawing a red marble as a failure. Let N describe the number of all marbles in the urn (see contingency table below) and K describe the number of green marbles, then N − K corresponds to the number of red marbles. Now, standing next to the urn, you close your eyes and draw n marbles without replacement. Define X as a random variable whose outcome is k, the number of green marbles drawn in the experiment. This situation is illustrated by the following contingency table:

drawn not drawn total
green marbles k Kk K
red marbles nk N + k − n − K N − K
total n N − n N

Indeed, we are interested in calculating the probability of drawing k green marbles in n draws, given that there are K green marbles out of a total of N marbles. For this example, assume that there are 5 green and 45 red marbles in the urn. Standing next to the urn, you close your eyes and draw 10 marbles without replacement. What is the probability that exactly 4 of the 10 are green?

This problem is summarized by the following contingency table:

drawn not drawn total
green marbles k = 4 Kk = 1 K = 5
red marbles nk = 6 N + k − n − K = 39 N − K = 45
total n = 10 N − n = 40

To find the probability of drawing k green marbles in exactly n draws out of N total draws, we identify X as a hyper-geometric random variable to use the formula

To intuitively explain the given formula, consider the two symmetric problems represented by the identity

  1. left-hand side - drawing a total of only n marbles out of the urn. We want to find the probability of the outcome of drawing k green marbles out of K total green marbles, and drawing n-k red marbles out of N-K red marbles, in these n rounds.
  2. right hand side - alternatively, drawing all N marbles out of the urn. We want to find the probability of the outcome of drawing k green marbles in n draws out of the total N draws, and K-k green marbles in the rest N-n draws.

Back to the calculations, we use the formula above to calculate the probability of drawing exactly k green marbles

Intuitively we would expect it to be even more unlikely that all 5 green marbles will be among the 10 drawn.

As expected, the probability of drawing 5 green marbles is roughly 35 times less likely than that of drawing 4.

Symmetries

Swapping the roles of green and red marbles:

Swapping the roles of drawn and not drawn marbles:

Swapping the roles of green and drawn marbles:

These symmetries generate the dihedral group .

Order of draws

The probability of drawing any set of green and red marbles (the hypergeometric distribution) depends only on the numbers of green and red marbles, not on the order in which they appear; i.e., it is an exchangeable distribution. As a result, the probability of drawing a green marble in the draw is[2]

This is an ex ante probability—that is, it is based on not knowing the results of the previous draws.

Tail bounds

Let and . Then for we can derive the following bounds:[3]

where

is the Kullback-Leibler divergence and it is used that .[4]

Note: In order to derive the previous bounds, one has to start by observing that where are dependent random variables with a specific distribution . Because most of the theorems about bounds in sum of random variables are concerned with independent sequences of them, one has to first create a sequence of independent random variables with the same distribution and apply the theorems on . Then, it is proved from Hoeffding [3] that the results and bounds obtained via this process hold for as well.

If n is larger than N/2, it can be useful to apply symmetry to "invert" the bounds, which give you the following: [4] [5]

Statistical Inference

Hypergeometric test

The hypergeometric test uses the hypergeometric distribution to measure the statistical significance of having drawn a sample consisting of a specific number of successes (out of total draws) from a population of size containing successes. In a test for over-representation of successes in the sample, the hypergeometric p-value is calculated as the probability of randomly drawing or more successes from the population in total draws. In a test for under-representation, the p-value is the probability of randomly drawing or fewer successes.

Biologist and statistician Ronald Fisher

The test based on the hypergeometric distribution (hypergeometric test) is identical to the corresponding one-tailed version of Fisher's exact test.[6] Reciprocally, the p-value of a two-sided Fisher's exact test can be calculated as the sum of two appropriate hypergeometric tests (for more information see[7]).

The test is often used to identify which sub-populations are over- or under-represented in a sample. This test has a wide range of applications. For example, a marketing group could use the test to understand their customer base by testing a set of known customers for over-representation of various demographic subgroups (e.g., women, people under 30).

Let and .

  • If then has a Bernoulli distribution with parameter .
  • Let have a binomial distribution with parameters and ; this models the number of successes in the analogous sampling problem with replacement. If and are large compared to , and is not close to 0 or 1, then and have similar distributions, i.e., .
  • If is large, and are large compared to , and is not close to 0 or 1, then

where is the standard normal distribution function

The following table describes four distributions related to the number of successes in a sequence of draws:

With replacements No replacements
Given number of draws binomial distribution hypergeometric distribution
Given number of failures negative binomial distribution negative hypergeometric distribution

Multivariate hypergeometric distribution

Multivariate hypergeometric distribution
Parameters




Support
PMF
Mean
Variance



The model of an urn with green and red marbles can be extended to the case where there are more than two colors of marbles. If there are Ki marbles of color i in the urn and you take n marbles at random without replacement, then the number of marbles of each color in the sample (k1, k2,..., kc) has the multivariate hypergeometric distribution:

This has the same relationship to the multinomial distribution that the hypergeometric distribution has to the binomial distribution—the multinomial distribution is the "with-replacement" distribution and the multivariate hypergeometric is the "without-replacement" distribution.

The properties of this distribution are given in the adjacent table,[8] where c is the number of different colors and is the total number of marbles in the urn.

Example

Suppose there are 5 black, 10 white, and 15 red marbles in an urn. If six marbles are chosen without replacement, the probability that exactly two of each color are chosen is

Occurrence and applications

Application to auditing elections

Samples used for election audits and resulting chance of missing a problem

Election audits typically test a sample of machine-counted precincts to see if recounts by hand or machine match the original counts. Mismatches result in either a report or a larger recount. The sampling rates are usually defined by law, not statistical design, so for a legally defined sample size n, what is the probability of missing a problem which is present in K precincts, such as a hack or bug? This is the probability that k = 0 . Bugs are often obscure, and a hacker can minimize detection by affecting only a few precincts, which will still affect close elections, so a plausible scenario is for K to be on the order of 5% of N. Audits typically cover 1% to 10% of precincts (often 3%),[9][10][11] so they have a high chance of missing a problem. For example, if a problem is present in 5 of 100 precincts, a 3% sample has 86% probability that k = 0 so the problem would not be noticed, and only 14% probability of the problem appearing in the sample (positive k ):

The sample would need 45 precincts in order to have probability under 5% that k = 0 in the sample, and thus have probability over 95% of finding the problem:

Application to Texas hold'em poker

In hold'em poker players make the best hand they can combining the two cards in their hand with the 5 cards (community cards) eventually turned up on the table. The deck has 52 and there are 13 of each suit. For this example assume a player has 2 clubs in the hand and there are 3 cards showing on the table, 2 of which are also clubs. The player would like to know the probability of one of the next 2 cards to be shown being a club to complete the flush.
(Note that the probability calculated in this example assumes no information is known about the cards in the other players' hands; however, experienced poker players may consider how the other players place their bets (check, call, raise, or fold) in considering the probability for each scenario. Strictly speaking, the approach to calculating success probabilities outlined here is accurate in a scenario where there is just one player at the table; in a multiplayer game this probability might be adjusted somewhat based on the betting play of the opponents.)

There are 4 clubs showing so there are 9 clubs still unseen. There are 5 cards showing (2 in the hand and 3 on the table) so there are still unseen.

The probability that one of the next two cards turned is a club can be calculated using hypergeometric with and . (about 31.64%)

The probability that both of the next two cards turned are clubs can be calculated using hypergeometric with and . (about 3.33%)

The probability that neither of the next two cards turned are clubs can be calculated using hypergeometric with and . (about 65.03%)

Application to Keno

The hypergeometric distribution is indispensable for calculating Keno odds. In Keno, 20 balls are randomly drawn from a collection of 80 numbered balls in a container, rather like American Bingo. Prior to each draw, a player selects a certain number of spots by marking a paper form supplied for this purpose. For example, a player might play a 6-spot by marking 6 numbers, each from a range of 1 through 80 inclusive. Then (after all players have taken their forms to a cashier and been given a duplicate of their marked form, and paid their wager) 20 balls are drawn. Some of the balls drawn may match some or all of the balls selected by the player. Generally speaking, the more hits (balls drawn that match player numbers selected) the greater the payoff.

For example, if a customer bets ("plays") $1 for a 6-spot (not an uncommon example) and hits 4 out of the 6, the casino would pay out $4. Payouts can vary from one casino to the next, but $4 is a typical value here. The probability of this event is:

Similarly, the chance for hitting 5 spots out of 6 selected is while a typical payout might be $88. The payout for hitting all 6 would be around $1500 (probability ≈ 0.000128985 or 7752-to-1). The only other nonzero payout might be $1 for hitting 3 numbers (i.e., you get your bet back), which has a probability near 0.129819548.

Taking the sum of products of payouts times corresponding probabilities we get an expected return of 0.70986492 or roughly 71% for a 6-spot, for a house advantage of 29%. Other spots-played have a similar expected return. This very poor return (for the player) is usually explained by the large overhead (floor space, equipment, personnel) required for the game.

See also

References

Citations

  1. ^ Rice, John A. (2007). Mathematical Statistics and Data Analysis (Third ed.). Duxbury Press. p. 42.
  2. ^ http://www.stat.yale.edu/~pollard/Courses/600.spring2010/Handouts/Symmetry%5BPolyaUrn%5D.pdf [bare URL PDF]
  3. ^ a b Hoeffding, Wassily (1963), "Probability inequalities for sums of bounded random variables" (PDF), Journal of the American Statistical Association, 58 (301): 13–30, doi:10.2307/2282952, JSTOR 2282952.
  4. ^ a b "Another Tail of the Hypergeometric Distribution". wordpress.com. 8 December 2015. Retrieved 19 March 2018.
  5. ^ Serfling, Robert (1974), "Probability inequalities for the sum in sampling without replacement", The Annals of Statistics, 2 (1): 39–48, doi:10.1214/aos/1176342611.
  6. ^ Rivals, I.; Personnaz, L.; Taing, L.; Potier, M.-C (2007). "Enrichment or depletion of a GO category within a class of genes: which test?". Bioinformatics. 23 (4): 401–407. doi:10.1093/bioinformatics/btl633. PMID 17182697.
  7. ^ K. Preacher and N. Briggs. "Calculation for Fisher's Exact Test: An interactive calculation tool for Fisher's exact probability test for 2 x 2 tables (interactive page)".
  8. ^ Duan, X. G. "Better understanding of the multivariate hypergeometric distribution with implications in design-based survey sampling." arXiv preprint arXiv:2101.00548 (2021). (pdf)
  9. ^ Glazer, Amanda; Spertus, Jacob (10 February 2020) [8 March 2020]. Start spreading the news: New York's post-election audit has major flaws (white paper). Elsevier. doi:10.2139/ssrn.3536011. SSRN 3536011. SSRN 3536011. Retrieved 4 December 2023 – via SSRN.com.
  10. ^ "State audit laws". Verified Voting. 10 February 2017. Retrieved 2 April 2018.
  11. ^ "Post-election audits". ncsl.org. National Conference of State Legislatures. Retrieved 2 April 2018.

Sources

Read other articles:

2003 video gameJames Bond 007: Everything or NothingDeveloper(s)Griptonite GamesPublisher(s)Electronic ArtsProducer(s)Steve EttingerMichelle GingrichProgrammer(s)Stephen NguyenWriter(s)Michael HumesComposer(s)Ian StockerSeriesJames BondPlatform(s)Game Boy AdvanceReleaseNA: November 17, 2003Genre(s)Third-person shooterMode(s)Single-player, multiplayer James Bond 007: Everything or Nothing is a third-person shooter video game, developed by Griptonite Games and published by Electronic Arts for the …

Renu Setna:capellán ,Josephine Welcome:Kattrin, Margaret Robertson: Madre Coraje, Madre Coraje y sus hijos, Bertolt Brecht, Teatro Internacionalista Angelique Rockas ,Carmen e Okon Jones, El balcón Teatro Internacionalista (Internationalist Theatre) es la compañía de teatro de Londres fundada por la actriz sudafricana de origen griega Angelique Rockas en abril de 1981 para ser la pionera en la actuación de dramaturgia clásica y piezas contemporáneas con elencos multiraciales y multi-nacio…

Johan Maurits MohrLahir(1716-08-18)18 Agustus 1716Eppingen, Elektorat Pfalz, Kekaisaran Romawi SuciMeninggal27 Oktober 1775(1775-10-27) (umur 59)Batavia, Hindia BelandaAlmamaterUniversitas GroningenTahun aktif1737–1775Dikenal atasLaporan transit Venus pada tahun 1761 dan 1769Suami/istriJohanna Cornelia van der Sluys ​ ​(m. 1739; meninggal 1750)​ Anna Elisabeth van ’t Hoff ​ ​(m. 1752)​ Johan Maurits …

Maiensäss Matschwitz im Montafon (ca. 1905) Häuser in Matschwitz; Richtung Golmerbahn (2005) Das Maiensäss (bzw. Maiensäß), auch Maisäss (Maisäß), Maien, Vorsäss (Vorsäß), Hochsäß, Niederleger, Unterstafel, in Graubünden auch rätoromanisch acla,[1] im Tessin Monti,[2] im Unterwallis frankoprovenzalisch Mayens, ist eine Sonderform der Alm/Alp: eine gerodete Fläche mit Hütten und Ställen. Auf jedem Maiensäss steht mindestens ein kleines Haus und ein Stall; als En…

يفتقر محتوى هذه المقالة إلى الاستشهاد بمصادر. فضلاً، ساهم في تطوير هذه المقالة من خلال إضافة مصادر موثوق بها. أي معلومات غير موثقة يمكن التشكيك بها وإزالتها. (مارس 2016) كأس آسيا لكرة القدم للسيدات 20012001年亞足聯女子亞洲杯تفاصيل المسابقةالبلد المضيف تايبيه الصينيةالتواريخ4–16 ديسم

العلاقات الأندورية الجنوب أفريقية أندورا جنوب أفريقيا   أندورا   جنوب أفريقيا تعديل مصدري - تعديل   العلاقات الأندورية الجنوب أفريقية هي العلاقات الثنائية التي تجمع بين أندورا وجنوب أفريقيا.[1][2][3][4][5] مقارنة بين البلدين هذه مقارنة عامة ومرج

Song written by Bob Crewe and Bob Gaudio The Sun Ain't Gonna Shine (Anymore)Single by Frankie Vallifrom the album Solo B-sideThis Is GoodbyeReleasedAugust 1965RecordedJuly 1965GenrePop rockLength3:26LabelSmashSongwriter(s)Bob CreweBob GaudioProducer(s)Bob CreweFrankie Valli singles chronology Please Take a Chance (1959) The Sun Ain't Gonna Shine (Anymore) (1965) (You're Gonna) Hurt Yourself (1966) The Sun Ain't Gonna Shine (Anymore) is a song written by Bob Crewe and Bob Gaudio. It was originall…

Tómbola Programa de televisiónGénero Talk showPresentado por Ximo RoviraTema principal Tómbola(compuesto por Augusto Algueró)País de origen España EspañaIdioma(s) original(es) Valenciano, CastellanoN.º de temporadas 7N.º de episodios 383ProducciónDuración 285 min.Empresa(s) productora(s) RTVVProducciones 52LanzamientoMedio de difusión Cadena Original Canal Nou (1997-2004) Miembros de la FORTA Telemadrid(1997-2001) Canal Sur(1997)Otras cadenas Canal 4 Castilla y León(1997-2004)…

American TV series or program True Confessions of a Hollywood StarletDVD coverBased onTrue Confessions of a Hollywood Starletby Lola DouglasScreenplay byElisa BellDirected byTim MathesonStarringJoanna JoJo LevesqueValerie BertinelliLynda BoydShenae GrimesLeah CudmoreIan NelsonCountry of originUnited StatesOriginal languageEnglishProductionProducerMark WinemakerCinematographyDavid HerringtonEditorCharles BornsteinRunning time87 minutesProduction companyStarlet ProductionsOriginal releaseNetw…

Jembatan VauxhallKoordinat51°29′15″N 0°07′37″W / 51.48750°N 0.12694°W / 51.48750; -0.12694Koordinat: 51°29′15″N 0°07′37″W / 51.48750°N 0.12694°W / 51.48750; -0.12694Moda transportasiJalan A202MelintasiSungai ThamesLokalLondon, InggrisStatus cagar budayaBangunan terdaftarSebelumnyaJembatan Regent (Jembatan Vauxhall lama) 1816–1898KarakteristikDesainJembatan lengkungBahan bakuBaja dan granitPanjang total809 kaki (247 m)…

Wali Kota Surakarta Joko Widodo (kanan) bersama dengan wakilnya, FX Hadi Rudyatmo, pada tahun 2011. Joko Widodo adalah Presiden Indonesia ke-7 (2014-2024), mantan Gubernur DKI Jakarta ke-14 (2012–2014), serta Wali Kota Surakarta ke-16 (2005–2012). Dalam rekam jejak politiknya, Joko Widodo tidak pernah dikalahkan dalam pemilihan umum.[1][2] Pemilihan Wali Kota Surakarta Foto resmi Joko Widodo sebagai Wali Kota Surakarta pada tahun 2005. 2005 Artikel utama: Pemilihan umum Wali …

Hifdzi KhoirLahirMuhammad Hifdzi Khoir4 Oktober 1991 (umur 32)Lampung, Bandar Lampung, IndonesiaAlmamaterUniversitas Gadjah MadaPekerjaanPelawak tunggalaktorpenyanyiTahun aktif2012—sekarangSuami/istriRinita Dini Eka Sari ​ ​(m. 2019)​Anak1 Muhammad Hifdzi Khoir, S.S. (lahir 4 Oktober 1991) adalah pelawak tunggal, aktor, dan penyanyi berkebangsaan Indonesia. Hifdzi adalah salah satu kontestan Stand Up Comedy Indonesia Kompas TV musim keempat (SUCI 4)…

Yellow brick For other uses, see Dutch brick (disambiguation). Close-up of Dutch bricks with inscription Spiral staircase to the carillon of the Dutch Reformed Church of IJsselstein Dutch brick (Dutch: IJsselsteen) is a small type of red brick made in the Netherlands, or similar brick, and an architectural style of building with brick developed by the Dutch. The brick, made from clay dug from river banks or dredged from river beds of the river IJssel[1] and fired over a long period of ti…

1995 single by Tlot TlotThe Girlfriend SongSingle by Tlot TlotB-sideIn the SummertimeStinkSunny Delirious (live)Released27 February 1995[1]Recorded1993GenreAlternative rockLength3:40LabelEMISongwriter(s)Owen Bolwell, Stanley PaulzenProducer(s)Siew, Tlot TlotTlot Tlot singles chronology Old Mac (1992) The Girlfriend Song (1995) The Girlfriend Song is the second and final single by the Australian rock band Tlot Tlot. The single was released in 1995 and was nominated for the ARIA award for …

Kasuari gelambir-ganda Status konservasi Risiko Rendah (IUCN 3.1) Klasifikasi ilmiah Kerajaan: Animalia Filum: Chordata Kelas: Aves Ordo: Struthioniformes Famili: Casuariidae Genus: Casuarius Spesies: Casuarius casuarius Nama binomial Casuarius casuarius(Linnaeus, 1758) Sebaran geografis Kasuari gelambir-ganda (Casuarius casuarius) adalah salah satu burung dari tiga spesies kasuari, yaitu kasuari gelambir ganda (Casuarius casuarius), kasuari gelambir tunggal (Casuarius unappendiculatus), da…

Ung thư học (từ tiếng Hy Lạp Cổ Đại ὄγκος onkos, to, lớn, khối, và tiền tố -λογία -logia, nghiên cứu) là một nhánh y học nghiên cứu về ung thư. Các bác sĩ chuyên về ung thư được gọi là nhà ung thư học.[1] Ung thư học có liên quan đến: Triệu chứng của bất kỳ loại ung thư nào ở con người Điều trị (chẳng hạn ngoại khoa, hóa trị liệu, điều trị phóng xạ và các hình thức k…

2005 Ohio's 2nd congressional district special election ← 2004 August 2, 2005 2006 →   Majority party Minority party   Candidate Jean Schmidt Paul Hackett Party Republican Democratic Popular vote 59,671 55,886 Percentage 51.63% 48.35% County results Schmidt:      50–60% Hackett:      50–60%      60–70% U.S. Representative before election Rob Portman Republican Elected U.S. Repres…

Universität Félix Houphouët-Boigny (Université de Cocody-Abidjan, Sigilium Universitatis Abidjansis)bis August 1996 University of Abidjan-Cocody Motto Scientia et Sapientia Via Mea Gründung 9. Januar 1964 Trägerschaft öffentlich Ort Abidjan Land Elfenbeinküste Elfenbeinküste Professeur Bakayoko Ly Ramata Studierende ca.60.000 Netzwerke FUIW[1] Website univ-fhb.edu.ci Campuseingang Die Universität Félix Houphouët-Boigny (französisch Université Félix Houphouët-Boigny, …

Al este del Edén de John Steinbeck Género Novela Subgénero Saga familiar Ambientada en Gilded Age Salinas Idioma Inglés e inglés Título original East of Eden Editorial Viking Press País Estados Unidos Fecha de publicación Septiembre de 1952 [editar datos en Wikidata] Al este del Edén o El este del Edén, a veces también traducida como Al este del Paraíso (East of Eden) es una novela de 1952 del novelista y premio Nobel estadounidense John Steinbeck. Narra la historia de…

يفتقر محتوى هذه المقالة إلى الاستشهاد بمصادر. فضلاً، ساهم في تطوير هذه المقالة من خلال إضافة مصادر موثوق بها. أي معلومات غير موثقة يمكن التشكيك بها وإزالتها. (ديسمبر 2018) حارة اللقية  - حارة -  تقسيم إداري البلد  اليمن المحافظة محافظة عدن المديرية مديرية دار سعد الح…

Kembali kehalaman sebelumnya

Lokasi Pengunjung: 18.118.227.16