Розподіл вибірки з двох незалежних популяцій Бернуллі


17

Припустимо, що у нас є вибірки двох незалежних випадкових величин Бернуллі, Ber(θ1) і Ber(θ2) .

Як ми доводимо, що

(X¯1−X¯2)−(θ1−θ2)θ1(1−θ1)n1+θ2(1−θ2)n2→dN(0,1)
?

Припустимо, що n1≠n2 .


Z_i = X_1i - X_2i - це послідовність iid rv кінцевих середніх та дисперсійних. Отже, вона задовольняє центральну граничну теорему Леві-Ліндерберга, з якої випливають ваші результати. Або ти просиш доказувати сам кельт?
— Три Діаг

@ThreeDiag Як ви застосовуєте LL-версію CLT? Я не думаю, що це правильно. Напишіть мені відповідь, щоб перевірити деталі.
— Старий чоловік у морі.

Усі деталі вже є. Для застосування LL вам потрібна послідовність iid rv з кінцевою середньою та дисперсією. Змінна Z_i = X_i1 і X_i2 задовольняє всі три вимоги. Незалежність випливає з незалежності двох оригінальних версій Бернуллі, і ви можете бачити, що E (Z_i) і V (Z_i) є кінцевими, застосовуючи стандартні властивості E і V
— Три діаг.

1
"вибірки двох незалежних випадкових змінних Бернуллі" - неправильне вираження. Повинно бути: "два незалежні вибірки з розподілу Бернуллі".
— Віктор

1
Будь ласка, додайте "як ". n1,n2→∞
— Віктор

Відповіді:


10

Покладіть ,b=√a=θ1(1−θ1)n1 , A=(ˉX1-θ1)/a, B=(ˉX2-θ2)/b. Маємо A→dN(0,1),B→dN(0,1). З точки зору характерних функцій це означає ϕA(t)≡Eeb=θ2(1−θ2)n2A=(X¯1−θ1)/aB=(X¯2−θ2)/bA→dN(0,1), B→dN(0,1) Ми хочемо довести, що D:= a

ϕA(t)≡EeitA→e−t2/2, ϕB(t)→e−t2/2.
D:=aa2+b2A−ba2+b2B→dN(0,1)

Оскільки і B незалежні, ϕ D ( t ) = ϕ A ( aAB як ми бажаємо, щоб це було.

ϕD(t)=ϕA(aa2+b2t)ϕB(−ba2+b2t)→e−t2/2,

Цей доказ є неповним. Тут нам потрібні деякі оцінки для рівномірного зближення характерних функцій. Однак у розглянутому випадку ми можемо робити чіткі розрахунки. Покладіть . ϕ X 1 , 1 ( t )p=θ1, m=n1

ϕX1,1(t)=1+p(eit−1),ϕX¯1(t)=(1+p(eit/m−1))m,ϕX¯1−θ1(t)=(1+p(eit/m−1))me−ipt,ϕA(t)=(1+p(eit/mp(1−p)−1))me−iptm/p(1−p)=((1+p(eit/mp(1−p)−1))e−ipt/mp(1−p))m=(1−t22m+O(t3m−3/2))m
as t3m−3/2→0. Thus, for a fixed t,
ϕD(t)=(1−a2t22(a2+b2)n1+O(n1−3/2))n1(1−b2t22(a2+b2)n2+O(n2−3/2))n2→e−t2/2
(even if a→0 or b→0), since |e−y−(1−y/m)m|≤y2/2m  when  y/m<1/2 (see /math/2566469/uniform-bounds-for-1-y-nn-exp-y/ ).

Note that similar calculations may be done for arbitrary (not necessarily Bernoulli) distributions with finite second moments, using the expansion of characteristic function in terms of the first two moments.


This seems correct. I'll get back to you later on, when I have time to check everything. ;)
— An old man in the sea.

-1

Proving your statement is equivalent to proving the (Levy-Lindenberg) Central Limit Theorem which states

If {Zi}i=1n is a sequence of i.i.d random variable with finite mean E(Zi)=μ and finite variance V(Zi)=σ2 then

n(Z¯−μ)→dN(0,σ2)

Here Z¯=∑iZi/n that is the sample variance.

Then it is easy to see that if we put

Zi=X1i−X2i
with X1i,X2i following a Ber(θ1) and Ber(θ2) respectively the conditions for the theorem are satisfied, in particular

E(Zi)=θ1−θ2=μ

and

V(Zi)=θ1(1−θ1)+θ2(1−θ2)=σ2

(There's a last passage, and you have to adjust this a bit for the general case where n1≠n2 but I have to go now, will finish tomorrow or you can edit the question with the final passage as an exercise )


I could not obtain what I wanted exactly because of the possibility of n1≠n2
— An old man in the sea.

I will show later if you can't get it. Hint: compute the variance of the sample mean of Z and use that as the variable in the theorem
— Three Diag

Three, could you please add the details for when n1≠n2? Thanks
— An old man in the sea.

Will do as soon as find a little timr. There was in fact a subtlety that prevents from using LL clt without adjustment. There are three ways to go, the simplest of which is invoking the fact that for large n1 and n2, X1 and X2 go in distribution to normals, then a linear combination of normal is also normal. This is a property of normals that you can take as given, otherwise you can prove it by characteristic functions.
— Three Diag

The other two require either a different clt (Lyapunov possibly) or alternatively treat n1 = i and n2= i +k. Then for large i you can essentially disregard k and you can go back to apply LL (but still it will require some care to nail the right variance)
— Three Diag
Використовуючи наш веб-сайт, ви визнаєте, що прочитали та зрозуміли наші Політику щодо файлів cookie та Політику конфіденційності.
Licensed under cc by-sa 3.0 with attribution required.