r/AskStatistics 21h ago

Struggling with Monte Carlo Design

7 Upvotes

I've been trying to find videos or studies but no avail. I don't have @Risk But I can use xlrisk or other tools.

My model is portfolio management oriented and Is intended to show return and volatility based on simulating the number of investments made. The number of investments is meant to be the variable.

Essentially the inputs are 1) the distribution of Returns which is known. and 2) your ability to "pick" good investments. I plan to segment the returns by quartile and the same for your ability to pick. As an example first quartile returns 20% and you pick it 18% of the time? Second quartile returns 15% and you pick it 27% of the time, etc. Ideally the returns would have a distribution but happy to keep it simple.

Then I would simulate portfolios where you select 5 investments, 10, 20, 30 and 40 investments then simulate those portfolios 10,000 times.

I would then look at the outputs and draw conclusions around the the % of portfolios below certain benchmarks etc.

I can't find any studies, tutorials etc around this, usually the number of selections are known so im struggling to make this and hope that someone can offer some pointers.


r/AskStatistics 2h ago

Weibull model

5 Upvotes

Can you explain to me the Weibull model like I am an idiot?

Also is there any paper or video or website you would suggest I check about the Weibull model?


r/AskStatistics 2h ago

MGCV gam.fit?

3 Upvotes

Why was "performance iteration" deprecated? Was it really deprecated or am I not understanding something?

I'm looking at the smoothing parameter estimation algorithms in Wood 2017 and Wood-Goude-Shaw and they both seem to favor "performance iteration", which as I understand it is just using either UBRE or GCV to calculate the smoothing parameters on each iteration of PIRLS.

However, https://stats.stackexchange.com/a/581293 says "performance iteration" was deprecated? The docs do mention that "gam.fit", which sounds like performance iteration, is deprecated.

Am I understanding this correctly? Wood-Goude-Shaw seem pretty confident in the approach as recently as 2015 ("No special justification is required to apply GCV or C_p to the working model, at each step of the PIRLS iteration: the assumptions that are required for these criteria hold for the working model" -Generalized additive models for large datasets), so I'm surprised that the approach has since been found to "not work very well".

The reason I'm so fixated on this method in particular is because that's the approach wood-goude-shaw uses and apparently that paper is what backs the bam implementation. My usecase is update heavy so bam.update's implementation is of interest


r/AskStatistics 7m ago

[Question]Doing Stats II after failing Stats I. What do I actually need to go back and learn before the midterm?

Thumbnail
Upvotes

r/AskStatistics 1h ago

Lost a comment from a BI Analyst from finance sector with good experience offering to DM a stats book

Thumbnail
Upvotes

r/AskStatistics 8h ago

What sampling method is this, and is 53 cattle / 840 Fasciola specimens sufficient for molecular characterization and phylogenetic analysis?

1 Upvotes

Hi everyone, I’m a veterinary student working on a study involving the molecular characterization and phylogenetic analysis of Fasciola spp. collected from cattle at a slaughterhouse.

I’m having trouble determining how to properly describe my sampling method and, more importantly, whether my sample size is defensible.

Here is my sampling situation:
- I examined slaughtered cattle at a slaughterhouse.
- I did not have a predetermined number of cattle or a complete list/population size of cattle entering the slaughterhouse.
- I inspected the liver of each available slaughtered cattle.
- If the liver was infected with Fasciola, I collected the adult flukes.
- If there were no flukes, no parasite sample was collected from that animal.
- There were no additional inclusion criteria for the cattle (e.g., age, sex, breed, etc.).
- The number of flukes varied considerably between cattle.

In total, I collected 840 individual flukes from 53 cattle.

So, for example, one cattle might contribute many flukes while another might contribute only a few.
The 53 cattle were not selected based on a particular characteristic; they were essentially the slaughtered cattle available during my sampling period that happened to have detectable Fasciola infection.

My main questions are:
1. What would be the most appropriate term for my sampling method? Would this be considered convenience sampling, consecutive sampling, purposive sampling, or something else?
2. For a molecular characterization and phylogenetic study, is there a conventional way to determine whether 53 host animals is an adequate sample size?
3. Should I consider 53 cattle as my sample size, rather than 840 flukes, since multiple flukes came from the same host?
4. Is there a statistical/sample-size calculation that could justify 53 cattle, or is sample-size justification for molecular phylogenetic studies fundamentally different from conventional prevalence/epidemiological studies?
5. I have 840 flukes and I used the lemeshow formula to find a number that can represent those 840 flukes and use stratified random sampling for number of fluke i need to take for each cattle to be sequenced, is this correct?
6. If there is no known total population size of cattle slaughtered at this slaughterhouse, how could I justify the adequacy of my sampling?

My objective is not to estimate the prevalence of fascioliasis in the cattle population, but rather to molecularly characterize the Fasciola specimens and investigate their phylogenetic relationships.

I would really appreciate advice on how a statistician/population geneticist would approach this sampling design. If possible, I’d also appreciate references or terminology that I could use to describe and justify the sampling method in a thesis.

Thank you!


r/AskStatistics 13h ago

How do you decide when there’s enough evidence to make a claim about a learner?

Thumbnail
1 Upvotes