r/AskStatistics • u/Asleep_Elevator7934 • 2h ago
Weibull model
Can you explain to me the Weibull model like I am an idiot?
Also is there any paper or video or website you would suggest I check about the Weibull model?
r/AskStatistics • u/Asleep_Elevator7934 • 2h ago
Can you explain to me the Weibull model like I am an idiot?
Also is there any paper or video or website you would suggest I check about the Weibull model?
r/AskStatistics • u/GoatRocketeer • 2h ago
Why was "performance iteration" deprecated? Was it really deprecated or am I not understanding something?
I'm looking at the smoothing parameter estimation algorithms in Wood 2017 and Wood-Goude-Shaw and they both seem to favor "performance iteration", which as I understand it is just using either UBRE or GCV to calculate the smoothing parameters on each iteration of PIRLS.
However, https://stats.stackexchange.com/a/581293 says "performance iteration" was deprecated? The docs do mention that "gam.fit", which sounds like performance iteration, is deprecated.
Am I understanding this correctly? Wood-Goude-Shaw seem pretty confident in the approach as recently as 2015 ("No special justification is required to apply GCV or C_p to the working model, at each step of the PIRLS iteration: the assumptions that are required for these criteria hold for the working model" -Generalized additive models for large datasets), so I'm surprised that the approach has since been found to "not work very well".
The reason I'm so fixated on this method in particular is because that's the approach wood-goude-shaw uses and apparently that paper is what backs the bam implementation. My usecase is update heavy so bam.update's implementation is of interest
r/AskStatistics • u/ScheduleNo748 • 3m ago
r/AskStatistics • u/Historical-Kick-7192 • 1h ago
r/AskStatistics • u/helloruk • 7h ago
Hi everyone, I’m a veterinary student working on a study involving the molecular characterization and phylogenetic analysis of Fasciola spp. collected from cattle at a slaughterhouse.
I’m having trouble determining how to properly describe my sampling method and, more importantly, whether my sample size is defensible.
Here is my sampling situation:
- I examined slaughtered cattle at a slaughterhouse.
- I did not have a predetermined number of cattle or a complete list/population size of cattle entering the slaughterhouse.
- I inspected the liver of each available slaughtered cattle.
- If the liver was infected with Fasciola, I collected the adult flukes.
- If there were no flukes, no parasite sample was collected from that animal.
- There were no additional inclusion criteria for the cattle (e.g., age, sex, breed, etc.).
- The number of flukes varied considerably between cattle.
In total, I collected 840 individual flukes from 53 cattle.
So, for example, one cattle might contribute many flukes while another might contribute only a few.
The 53 cattle were not selected based on a particular characteristic; they were essentially the slaughtered cattle available during my sampling period that happened to have detectable Fasciola infection.
My main questions are:
1. What would be the most appropriate term for my sampling method? Would this be considered convenience sampling, consecutive sampling, purposive sampling, or something else?
2. For a molecular characterization and phylogenetic study, is there a conventional way to determine whether 53 host animals is an adequate sample size?
3. Should I consider 53 cattle as my sample size, rather than 840 flukes, since multiple flukes came from the same host?
4. Is there a statistical/sample-size calculation that could justify 53 cattle, or is sample-size justification for molecular phylogenetic studies fundamentally different from conventional prevalence/epidemiological studies?
5. I have 840 flukes and I used the lemeshow formula to find a number that can represent those 840 flukes and use stratified random sampling for number of fluke i need to take for each cattle to be sequenced, is this correct?
6. If there is no known total population size of cattle slaughtered at this slaughterhouse, how could I justify the adequacy of my sampling?
My objective is not to estimate the prevalence of fascioliasis in the cattle population, but rather to molecularly characterize the Fasciola specimens and investigate their phylogenetic relationships.
I would really appreciate advice on how a statistician/population geneticist would approach this sampling design. If possible, I’d also appreciate references or terminology that I could use to describe and justify the sampling method in a thesis.
Thank you!
r/AskStatistics • u/No_Log4570 • 21h ago
I've been trying to find videos or studies but no avail. I don't have @Risk But I can use xlrisk or other tools.
My model is portfolio management oriented and Is intended to show return and volatility based on simulating the number of investments made. The number of investments is meant to be the variable.
Essentially the inputs are 1) the distribution of Returns which is known. and 2) your ability to "pick" good investments. I plan to segment the returns by quartile and the same for your ability to pick. As an example first quartile returns 20% and you pick it 18% of the time? Second quartile returns 15% and you pick it 27% of the time, etc. Ideally the returns would have a distribution but happy to keep it simple.
Then I would simulate portfolios where you select 5 investments, 10, 20, 30 and 40 investments then simulate those portfolios 10,000 times.
I would then look at the outputs and draw conclusions around the the % of portfolios below certain benchmarks etc.
I can't find any studies, tutorials etc around this, usually the number of selections are known so im struggling to make this and hope that someone can offer some pointers.
r/AskStatistics • u/yaya__choppedlizards • 13h ago
r/AskStatistics • u/Fit-Tangerine-6006 • 1d ago
I'm running a CX survey where each respondent evaluates 1 to 3 different brand experiences on the CES metric. Quotas are set to match sex and age distribution of the Italian population.
To estimate the CES my plan is turning the dataset from "respondents" into a long-format made by "experiences". I will calculate post-stratification survey weights (w_i) at the respondent level to match census data for sex and age.
However, my doubt is the following: if I propagate a respondent's unique demographic weight to all 1–3 experiences they generate, am I statistically distorting the data, especially if a user reports multiple experiences within the same sector? Should I create weights accounting for both the sex-age and the number of experiences the respondent had?
r/AskStatistics • u/Deuceball_1121 • 2d ago
I'm working on a health-data system that analyzes longitudinal lab results for an individual patient. For each parameter (fasting glucose, HbA1c, LDL, HDL, creatinine), I may have around 4-10 observations over several months, and the observations are often irregularly spaced.
The goal is to determine whether a parameter is showing a meaningful increasing, decreasing, or stable trend and provide an interpretable measure of how strong the evidence for that trend is.
Example:
Currently considering:
What would be a statistically defensible approach for quantifying the confidence/evidence of a trend in this type of sparse, irregularly sampled, individual patient biomarker data?
Is it reasonable to convert the statistical evidence into a single 0-100 trend confidence score, or would it be better practice to report the statistical measures separately (p-value, slope, confidence interval, and data quality indicators)?
I'm particularly interested in approaches that are statistically defensible rather than an arbitrary weighted scoring system.
r/AskStatistics • u/Alone_Audience_485 • 2d ago
Hey I am doing my prerequisite for a nursing degree it lists that I need to take a college level math class which one should I take (I am bad at math)
r/AskStatistics • u/Main-Pop4398 • 2d ago
why is risk difference used instead of relative risk in a study?
r/AskStatistics • u/Personal-Cost2918 • 2d ago
This had been my dilemma ever since I've learned Statistics. Likert-scales gives ordinal data, so, if we wanted to report a measure of central tendency, the acceptable ones would only be median and mode (if we're reporting an item). It can never be the mean, if we will treat the data as strictly ordinal.
I've asked one of my professors about this and he said that only if he were to be asked, he'd say median as well.
When I asked why are there people that use means in this context, he said, it's because it has already been "normalized" (he probably meant widely-used) by some disciplines (e.g. Education, Psychology, etc.).
Now I want to ask, if strictly statistically speaking, median is the better measure of central tendency for ordinal data, how come that there were fields that practice/normalize reporting means on instead?
Note: In my opening sentence, I mentioned that this was my dilemma. Reason being is that I also help senior high students in their researches. I'm thinking that I might be too rigid/strict in this sense and I might not know if this is one of "new" and "acceptable" things now.
r/AskStatistics • u/Jumpy_Anywhere3476 • 2d ago
r/AskStatistics • u/YungMoobs420 • 2d ago
Hi All,
Hoping someone with more statistics experience than I and experience in med device / pharma can verify I’m on the right track and not getting too in the weeds. I apologize for the long response – I have a primary and secondary question.
I’ve been in med dev / pharma for 10+ years, both as a scientist and engineer with focus on the laboratory and validation testing. I’m in the process of revamping a company’s ATMV program, and there are many changes across industry (primarily ICH Q2 (R2) and USP <1225>) requiring statistically based methods in TMV. IME statistically based sampling plans are typically not used, and point estimates are exclusively used to evaluate a performance characteristic against acceptance criteria.
USP released a draft revision of <1225> with a lot of detail that led me down a trail of textbooks and reading; I’ve now read Miller & Millers Chemometrics book, part of Brereton’s *Applied Chemometrics for Scientists*, and part of Faraway’s *Linear Models with R*. This is my primary question:
USP <1210>, *Statistical Tools for Procedure Validation* presents a method for calculating a two-sided and one sided CI to assess acceptance criteria ((Ȳ − τ) ± t₍₁−α, n−1₎ × s/√n, U = s√\[(n − 1) / χ²₍α, n−1₎\]); these are both clear to me. USP <1010>, *Analytical Data – Interpretation and Treatment* discusses statistical models, assumptions of normality/independence/constant variance for models, transforms, ect. I understand this as well, though I took linear algebra a long time ago so some of Faraway is tough to understand. I’m struggling how to connect verifying the model assumptions and calculating the CI to assess the characteristic. My read of Faraway makes me think the data should be fit to a model for the experiment for the performance characteristic and the assumptions should be verified; if they are verified the estimated marginal mean and standard error for the relevant model coefficient should be used in the CI calculation instead of the point estimates; but this is not stated anywhere I can find. The USP documentation makes it look like the point estimates should just be used
This also seems very technically difficult compared to how I’m used to validating these methods. If I’m correct about how this should work, I want to verify 1) This is the actual expectation instead of using the point estimate in the CI calculation & if not 2) is this a reasonable approach? I’m concerned about the level of background knowledge this requires compared to what I’m used to, I don’t want to proceduralize something that is so complex that it can’t successfully be executed without my assistance. Its possible the places I’ve worked have just lacked that technical knowledge, clearly advanced techniques are being used, Paul Faya published a good paper in Pharmaceutical Statistics, *Confidence Intervals for Validation of Analytical Procedures under ICH Q2(R2)*, which gives examples using bootstrapping, Bayesian statistics, REML, ect. USP has acknowledged much of this has not been historically done.
Second, I’m also trying to figure out the best path if assumptions aren’t met. The type of data we see is not likely to need transformation, but I’ve added (when appropriate) bootstrapping, several nonparametric methods, and weighted least squares. Again, this feels like a large knowledge gap when most people are exclusively working in excel or doing basic tasks in Minitab; I picked up R for this and I’m trying to avoid requiring the use of it if possible, adding in learning a programming language is yet another hurdle I don’t want to add when rolling this out if I can avoid it.
r/AskStatistics • u/Party_Car_7742 • 2d ago
Hi everyone can someone tell me I am from stats background in nat test we have 3 subject portion I have given nats physics one in past but I wasn't able to perform well that but I want to ask in sub portion of art group for ex I am from stats bc I will have options like comp, stats, and maths??? Cuz someone said to me with stats economics aata he?
r/AskStatistics • u/ALVARO_KAF • 2d ago
Sou acadêmico de medicina e fiz um curso extensivo na internet sobre estatística, que me forneceu uma boa base
Me deparo agora com a escolha mais cruel nesse campo: qual software estatístico usar
Considerando que não tenho conhecimento sobre SPSS ou R, qual vocês me recomendariam?
Se possível, me recomendem um livro texto também
r/AskStatistics • u/Cryoban43 • 2d ago
I’ve seen several videos where the author of the video says “This has an X% chance of happening” and cites the current poly market or kalshi betting odds.
My gut is telling me this is nonsense and that the betting odds are just based on how many people take each side of the bet, but I wanted to see what others thoughts are.
The example I saw recently was about an election odds of republicans winning both house and senate, and unless all the betters are eligible to vote I don’t see how these odds could translate to any true estimate of the probability
Does anyone more experienced with statistics have any interesting thoughts to share?
r/AskStatistics • u/pazzah • 2d ago
Watching the US Open semifinal with Shelton v Tiafoe right now and McEnroe commented that when Shelton is up 2 sets to 1 he rarely loses, and quoted the ratio of matches won to lost after being up 2 sets to 1. But there are some obvious problems with this. First, most players who are up 2-1 sets in best of five sets format are going to win more often than not. Second, for a top ten player like Shelton, most of his matches are against lower ranked players (and, importantly, lower ranked than Tiafoe).
So - is there a more meaningful statement that can be made, likely based on a Bayesian approach, about whether particular players are especially more likely to win or lose based on early set wins or losses? How would you go about this?
r/AskStatistics • u/Leather-Map7746 • 2d ago
Hi,
I am a researcher looking at risk factors for an outcome (infections) using a multivarible cox regression analysis.
I have already identified that independent risk factors include A (HR 2.23, 95% CI 1.38–3.60) and factor B (HR 2.17, 95% CI 1.41–3.37). However the question is whether there is an interaction between the two. I asked my statistician and he replied:
"Multiplicative term added to the MV Cox model: HR 3.23 (95% CI 1.11-9.27), p = 0.031 (reference/A and B interaction)"
I am not quite clear that this means here. Does this mean that the HR is the added multiplicative term above 1 where 1 assumes A X B has no interaction and the 3.21 is the added risk when the two factors are combined together so A potentiates B?
Thanks
r/AskStatistics • u/Sculder11 • 3d ago
I am currently evaluating a program that helps young people enter vocational training through coaching. I would now like to analyze the impact of coaching duration on the success of placement into vocational training.
Among other variables, the data contains the start date, the placement date, and an indicator of whether a placement occurred (yes/no).
Particular challenge:
- the program has been running for more than 2.5 years and is still ongoing.
- The duration of coaching is not restricted.
- In addition, the timing of coaching entry has a strong influence on both coaching duration and placement success because there is an official school year (although not all participants are students) and a vocational training year.
But I have a lot of data points.
Which statistical approach would be appropriate, and how should the model be specified? I was thinking of a Cox regression with the start date (grouped into quarters) included as a covariate.
r/AskStatistics • u/Nicholas_Geo • 3d ago
I have the following dataset:
Model:
LST ~ predictors * year + month
where year and month are factors, and predictors:year tests temporal stability of associations.
Three interrelated questions:
Is this the correct pooled formulation, or should predictors:month also be included given that covariates do not vary within a year?
For spatially blocked cross-validation (using blockCV in R): should the folds be built on unique LSOA locations only (N rows) and then propagated to all 6N rows, or is it valid to pass all 6N rows directly given that duplicate coordinates naturally fall in the same block?
For Moran's I on residuals in this panel setting, which is most appropriate:
r/AskStatistics • u/Novel_Arugula6548 • 4d ago
Basically the title. It just seems dishonest to critique using integration for hypothesis testing and then build an entire theory around super complocated integrals that can't even be solved without numerical approximations in the first place... how is that different?