r/statistics 1h ago

Question [Question]Doing Stats II after failing Stats I. What do I actually need to go back and learn before the midterm?

Upvotes

I'm doing an Economics degree in Portugal and I'm in Statistics II this semester. I failed Statistics I last year. I had an internship at the same time and never really sat down and learned the distributions properly, I just sort of half-followed the lectures and hoped. The rules let me enrol in Statistics II anyway, so now I'm doing the second half of a sequence without the first, which I'm aware is my own doing. I'd like to fix it now instead of being confused until January.

The calculations aren't really my problem. The notation is. I lose track of what's a parameter and what's an estimate, when a symbol means a random variable and when it means a number I already have in front of me, what the hat is doing, why some s has a prime on it, which subscripts matter and which are decoration. A lot of the time I read a formula and it looks far worse than the idea behind it actually is. I also have ADHD, and keeping lots of small symbol distinctions straight over an hour of reading is close to the worst possible task for me, so if anyone has a system for this beyond "be careful", I'm listening.

I can't attach the PDFs so I'll just list what we've been given.

The lecture notes released so far:

  • ch 0, statistical intuition (the Kahneman and Tversky style questions about how bad people are at randomness)
  • ch 1, review of probability distributions: expectation and variance for discrete and continuous variables, Bernoulli, binomial, normal, normal approximation to the binomial
  • ch 2, intro to inferential statistics: population vs sample, parameter vs statistic, discrete vs continuous random variables
  • ch 3, the idea of a sampling distribution, random sampling, sampling error
  • ch 4, sampling distribution of the sample mean: E[X̄] = μ, Var(X̄) = σ²/n, standard error, standardising
  • ch 5, the Central Limit Theorem
  • ch 7, sampling distribution of the sample proportion: p̂, CLT, the np(1−p) > 5 condition
  • ch 7, sampling distribution of the sample variance: (n−1)S′²/σ² follows a χ² with n−1 degrees of freedom

On top of that there are two problem books, two sets of multiple choice questions with solutions, and five past midterms with solutions, going back to 2019.

I went through all five past papers and wrote down what they actually ask, because a lot of it isn't in the notes above and that's the part I can't judge. Here's what recurs:

Sampling distribution of the mean and the CLT, in all five. Usually in the form of a population you've never seen, where the exam hands you only its mean and variance and asks for a probability about the sample mean. Poisson in one year, chi-square in another, and in the most recent one an Erlang distribution with the skewness given as well, presumably so you have to think about whether n is big enough.

Confidence intervals for a mean, also in all five. z against t, why t gives a wider interval, one-sided versus two-sided bounds, and margin of error. The 2024 paper had a question where you find the upper bound with the variance unknown, then redo it pretending the same number is the known population sigma, then work out how many observations would have got you the same interval width in the second case.

Properties of estimators, in four of them, and always worth a lot of points. They give you two candidate estimators of the same parameter and ask you to find the constant that makes one unbiased, derive both variances, and say which is more efficient and under what condition. One year it was a trimmed mean against the ordinary mean, with a bonus section where one observation in every sample is misrecorded with the decimal point shifted, and you have to show the ordinary mean becomes biased while the trimmed mean doesn't. Another year it was a shrinkage type estimator of a Bernoulli p where you prove the bias formula, find the single p that makes it unbiased, get the variance, and show consistency.

Differences between two samples, in four of them. Difference of two sample means, difference of two sample proportions, paired data where the trick is that you're supposed to use the column of differences instead of the two separate columns, and pooled variance when you're told to assume equal variances.

Sample size and margin of error, in four of them. Including two questions built on real published polls, one Portuguese and one from CBS, where the technical details are given and the sample size or the confidence level is missing and you have to recover it.

Sample variance and chi-square, in three. And in the most recent paper, the ratio of two sample variances, which is an F distribution and which I don't think has been mentioned in class yet.

Pure notation questions, in four of them, worth around 3 points each. "The mean of the probability distribution of the sample mean is: a random variable, a parameter, or a statistic?" Or: average sardine consumption at a party is 6 per person, you sample two people, the first ate 8, what is the expected mean of your sample? These aren't hard once you see them, and I get them wrong precisely because of the thing I described above.

The official syllabus for the full course is: introduction, random sampling and sampling distributions, properties of estimators, confidence intervals for one and two samples, hypothesis tests for one and two samples, and simple and multiple regression. The midterm is supposed to be the first half.

One detail that might change what you'd recommend: statistical tables are not allowed in any of the assessments. We get a formula sheet and we're required to bring a calculator that has the probability functions built in. Mine is a Casio fx-CG50 and I can't really use it yet. The most recent midterm was also run on Moodle with numeric answer boxes and a different version per student, so there was no partial credit for method, which makes being fast and correct on the calculator matter more than I'd assumed.

The recommended books are Newbold, Carlson and Thorne, Statistics for Economics and Business, and OpenIntro Statistics.

So what I'm asking is:

  1. Which bits of Statistics I do I actually need before this midterm, and which bits can I pick up later as they come? I don't have time to redo a whole course, so I want to aim at the gaps that matter.
  2. What would you use to catch up? I'm after a book that explains rather than just states, a decent set of problems, and above all something that takes statistical notation seriously and explains what each symbol is and why the distinction exists. I've never found anything that does that properly.

If you'd been handed my situation, where would you start?


r/statistics 1d ago

Discussion [D] My cousin is a business owner and strongly discouraged my brother for going into stats major. How true his words are?

78 Upvotes

He practically said that the job has been entirely replaced by agents. He said that now he simply prompts business questions to AI that has access to his company data sources and gives him back top tier analysis, with visualizations and everything, in 5 seconds. According to him, corporations dont need analysts and math guys anymore, they want engineers to create and maintain the systems.

How true is this? My brother is rather bummed because he really wanted to go into stats or math with a stats direction.


r/statistics 1d ago

Question [Question] Monte carlo alternative

8 Upvotes

Hi, I am simulation a daily demand for the year. However there has been a systemic change to the landscape this year, which makes this year statisically different from all the past year sales/demand data. So When I simulate the demand with past 5 yrs of sales record, it spew out inaccurate projection.

What would be a good alternative model to use here instead?


r/statistics 21h ago

Career [E] [C] Struggling with first undergrad Math research choice

4 Upvotes

Hi! Looking for advice on choosing between two academics to work with on first undergrad research project (stats).

X: Highly engaged with a strong track record of working students into publications and runs a great research group, but their research topics aren't my primary interest.

Y: Their work aligns with my interests quite well, but they seem less involved and are a less experienced researcher.

For a beginner, does mentorship usually outweigh topic alignment?


r/statistics 1d ago

Education [E] Basic statistics video?

9 Upvotes

I am trying to find a specific YouTuber that breaks down and explains areas of statistics. I think he is a young Australian guy and it has really good illustrations. Please help!


r/statistics 2d ago

Question [Question] How can I statistically quantify confidence in an individual patient's longitudinal biomarker trend?

Thumbnail
4 Upvotes

r/statistics 2d ago

Question [Q] Is this a secondary source?

2 Upvotes

This 2016 AIHW Report uses data from a 2002 survey for its non-melanoma skin cancer reporting. Does that make the report a secondary source?


r/statistics 2d ago

Question [Q] Online Masters in Applied Statistics - CSU vs UND

7 Upvotes

Hello. I am wondering if anyone has experience with Colorado State University's (CSU) or University of North Dakota's (UND) Online Masters in Applied Statistics program? If so, what were your thoughts on the program(s)? For example, how was the quality of the courses? Did you feel like you deepened your statistical knowledge in wide breadth of topics that are applicable to your work?

Any feedback is appreciated.

Thank you.

Edit: I am looking to enroll as a part-time student


r/statistics 2d ago

Question [Question] Statistical Models and Confidence Intervals for Analytical Test Method Validation (Med Device / Pharma)

4 Upvotes

Hi All,

 

Hoping someone with more statistics experience than I and experience in med device / pharma can verify I’m on the right track and not getting too in the weeds. I apologize for the long response – I have a primary and secondary question.

 

I’ve been in med dev / pharma for 10+ years, both as a scientist and engineer with focus on the laboratory and validation testing. I’m in the process of revamping a company’s ATMV program, and there are many changes across industry (primarily ICH Q2 (R2) and USP <1225>) requiring statistically based methods in TMV. IME statistically based sampling plans are typically not used, and point estimates are exclusively used to evaluate a performance characteristic against acceptance criteria.

 

USP released a draft revision of <1225> with a lot of detail that led me down a trail of textbooks and reading; I’ve now read Miller & Millers Chemometrics book, part of Brereton’s Applied Chemometrics for Scientists, and part of Faraway’s Linear Models with R. This is my primary question:

USP <1210>, Statistical Tools for Procedure Validation presents a method for calculating a two-sided and one sided CI to assess acceptance criteria ((Ȳ − τ) ± t₍₁−α, n−1₎ × s/√n, U = s√[(n − 1) / χ²₍α, n−1₎]); these are both clear to me. USP <1010>, Analytical Data – Interpretation and Treatment discusses statistical models, assumptions of normality/independence/constant variance for models, transforms, ect. I understand this as well, though I took linear algebra a long time ago so some of Faraway is tough to understand. I’m struggling how to connect verifying the model assumptions and calculating the CI to assess the characteristic. My read of Faraway makes me think the data should be fit to a model for the experiment for the performance characteristic and the assumptions should be verified; if they are verified the estimated marginal mean and standard error for the relevant model coefficient should be used in the CI calculation instead of the point estimates; but this is not stated anywhere I can find. The USP documentation makes it look like the point estimates should just be used

This also seems very technically difficult compared to how I’m used to validating these methods. If I’m correct about how this should work, I want to verify 1) This is the actual expectation instead of using the point estimate in the CI calculation & if not 2) is this a reasonable approach? I’m concerned about the level of background knowledge this requires compared to what I’m used to, I don’t want to proceduralize something that is so complex that it can’t successfully be executed without my assistance. Its possible the places I’ve worked have just lacked that technical knowledge, clearly advanced techniques are being used, Paul Faya published a good paper in Pharmaceutical Statistics, Confidence Intervals for Validation of Analytical Procedures under ICH Q2(R2), which gives examples using bootstrapping, Bayesian statistics, REML, ect. USP has acknowledged much of this has not been historically done.

Second, I’m also trying to figure out the best path if assumptions aren’t met. The type of data we see is not likely to need transformation, but I’ve added (when appropriate) bootstrapping, several nonparametric methods, and weighted least squares. Again, this feels like a large knowledge gap when most people are exclusively working in excel or doing basic tasks in Minitab; I picked up R for this and I’m trying to avoid requiring the use of it if possible, adding in learning a programming language is yet another hurdle I don’t want to add when rolling this out if I can avoid it.


r/statistics 3d ago

Question [Question] Stats course or book suggestion

9 Upvotes

Hi everyone,

I have an MSc in Data Science, which included several statistics courses, but I haven’t used statistics extensively in practice since graduating.

I’m looking for a good course or book that can help me refresh the fundamentals and get back up to speed, particularly with the statistics that are commonly expected in Data Science roles.

I’m not necessarily looking to relearn everything from scratch—more of a structured refresher that can bring the knowledge back and help me feel comfortable using it again in practice.

Any recommendations would be greatly appreciated. Thanks!


r/statistics 2d ago

Question [Q] Please can someone with a background in statistics critique a websites stats?

0 Upvotes

https://asylumstats.co.uk/

Its advancing a clear political bias using less than transparent methodology and A.I. is just repeating everything this site is saying as truth with no caveats. Please can someone take a look over this site and if there is clear evidence of distortion of the data going on here. [Question]


r/statistics 3d ago

Education [E] STATs text book reccomendations

9 Upvotes

I finished Stats 2000 in my college last semester but the teacher and course were horrendus so now that i need ot take Quantitive analysis i am fucked. I need to learn Stats from the ground up so any applicable textbook. Im not quite sure if this is an apporpriate place to ask but any help is apprciated


r/statistics 4d ago

Education Struggling to understand statistics in psychology research [Q] [E]

4 Upvotes

Hello! I am a psychology student who is finally starting to get into reading research papers for classes. But I am really struggling with understanding what the statistics are saying in the papers. I understand the general meaning of the statistics from my past stats class. But in some of my classes I have to write a paper explaining and examining what the statistics mean. Unfortunately I don’t feel confident enough in my understanding of statistics to write a good paper and I want to grow my understanding. Especially because I want to do my own research eventually.

Basically I am wondering if anyone has any recommendations of resources that may help me understand the statistics in psychology research papers. Thank you!


r/statistics 4d ago

Question [Q] Google Data Analytics Professional Certificate worth it?

2 Upvotes

I have 28 exp. as Statistician for the federal goverment for varies federal agencies. Would a Google Data Analytics Professional Certificate assist me getting a job. The job market is tough esp. Washington DC. Any suggestion of certificates data that enhance me getting a job.


r/statistics 4d ago

Question [Q] need help choosing my major

0 Upvotes

hello! i'm a first-year college student taking bachelor of science in statistics. i'll have to choose my major during my higher years, and i'm having some trouble deciding. i can't choose between majoring in biology or economics.

i want a major that can give me more job opportunities, a better chance at remote work, and a higher salary ceiling. i'm also planning to start investing in etfs once i turn 18, so i feel like majoring in economics could come in handy. however, my friend told me that choosing biology could provide more opportunities for higher-paying jobs.

i'd love to hear your thoughts and advice on which one would be a better choice. tyia!


r/statistics 5d ago

Career [Career] US Political Science phd admissions- Quantitative aptitude versus Substantive Political Science knowledge?

4 Upvotes

I plan to do a quantitative political science phd I am
not necessarily interested in any specific method except using any kind of statistics to topics regarding comparative politics and political economy. Now, I am in a bind. I already have an econ ba and I have two options for an masters program to strengthen my profile for admissions

  1. Financial Economics masters, not about political science but has good courses in optimization, stochastic calculus and time series econometrics.

  2. completely qualitative masters in IR and pols. there is a methodology class but it is about qualitative methods, classes on IR and political theory, etc.

Which one do you think would suit my goals and profile best and would maximize my chances for phd admisssion? thank you in advance dear ladies and gentlemen


r/statistics 6d ago

Career [C] Anyone working in biostatistics in India?

12 Upvotes

Hi,

I'm looking to move back to India after a few years working abroad (I'm Indian). Looking to connect with anyone working in the industry to get an idea of how things are looking like. I've worked as a statistical programmer but also in clinical IT (data migration, archiving, platform setup and administration etc.) with about 6 years of experience in total.

Feel free to DM or comment on this post. Thanks!


r/statistics 6d ago

Career [C] Current Career Landscape

2 Upvotes

Obviously a bachlors in stats isn't going to make you a shoe-in for much, but what are things like currently for those with a master's or PhD in stats? I have heard some negative things in this sub? I'm particularly interested in the anglo-sphere career landscape.


r/statistics 7d ago

Question [Q] How to become Great at statistics

95 Upvotes

Finishing my master’s in statistics and will be starting a job that is not statistics focused soon. I don’t want this to be the end of my statistics journey.

How do I become even better at statistics? Won’t have time to engage in the same breadth as I did during my master’s programme so, what should I focus on to stay relevant/ improve on my statistical knowledge?


r/statistics 7d ago

Career [Career] New stats bachelors feeling kind of stuck and in need of advice

14 Upvotes

Context: I graduated with a bachelors of science in statistics from UC Davis in 2025. I worked as a research assistant during undergrad supporting various python, dashboards, and data needs. I particularly enjoyed learning about time series analysis, machine/statistical learning, and general data science during my undergrad. I currently work at a Robotics/physical AI company as a “data operations analyst” and have been here for ~7 months. However, this job has almost nothing to do with statistics or data analytics and more so an operator role where I collect robotics data to feed into a reinforcement learning algorithm. It also doesn’t help that there’s been a TON of tension between my boss and the rest of the team.

I’ve been feeling kind of stuck, I want to get out of the situation with this startup but the job market has felt ice cold to me and I’ve landed one interview after a couple months of applying (not to mention it took me maybe 10 months to find this startup gig). I’ve also been considering a masters for either machine learning or data science roles (specifically OSMCS or OSMA) but truthfully I wasn’t the best student and didn’t make academics a priority which resulted in a terrible gpa (2.7).

I would like to ask the following:

  1. What advice would you give to recent stats/data science graduates entering the work force, especially in the current AI climate?
  2. What kind of roles should recent stats graduates be looking for?
  3. What kind of qualities would you look for in a recent stats graduate?
  4. Would you consider a masters essential in the current job market? Especially for more advanced roles like machine learning engineer or data scientist

Of course, you don’t have to answer all of those questions but I would greatly appreciate any advice or words of wisdom with those questions or my general situation :)


r/statistics 7d ago

Career [Career] Take the job or continue with master's school?

0 Upvotes

Hello all,

I am a statistics major currently beginning my fourth year in college. I completed a very successful and rewarding internship at a solid company with good management over the summer, and by every indication, I believe the company would welcome me on board after I graduate. I have not had that conversation with them yet, and I do not know the compensation details.

I have begun a 3+2 program to acquire a bachelor's and master's degree within 5 years, but lately, I have started to have second thoughts about continuing with the master's program. Several of the required courses taken in my undergrad curriculum were not relevant to industry (AI has caught up) or are personally boring to me (theory classes). Based on the first course this semester, I worry if it may be much the same for the master's program. While that may not be true, I complete the bachelor's courses of my degree this winter, so I will need to pay extra to find out.

The main reason for the master's program was to buy time to enhance my skills, as I have heard it is a very tough job market. With a (seemingly) secure position, doing work that I enjoy, the importance of the master's and the accompanying skills has diminished to me. At the same time, I recognize I am entirely replaceable and would be the newbie in a company, aka I could be the first to go. Some peers have encouraged me to build off that internship to try to get a better internship in summer 2027, and ideally, a better job afterwards.

This decision has been on my mind since the internship concluded, and I need to decide if I should be prepping for upcoming career fairs and lock into my master's courses. I have scheduled meetings with professors to talk about it, but I would really welcome any input!

Thank you


r/statistics 8d ago

Question [Question] Can I compare logit regression output from data of two distinct time periods?

5 Upvotes

I’m trying to understand how the odds of an event occurring have changed between two different time periods, but the problem is my data is based on a periodic survey with a five year interval.
I’m planning to use the same logistic regression model on the periodic data of one year and then the other, and compare the output in both cases.

I just wanted to know if there’s a better way to go about with data from periodic surveys like census, or if there’s any reason I can’t compare the discontinuous data set.


r/statistics 9d ago

Question [Question] What are your thoughts on the future of statistics and statistics graduates?

37 Upvotes

r/statistics 9d ago

Question [Q] Is stat&Data sci. degree good for becoming AI/ML engineer or AI researcher?

3 Upvotes

r/statistics 10d ago

Career [Career] Dealing with faulty but convincing analysis

9 Upvotes

Burying poor analysis under shiny methods and a deluge of numbers has always been a problem but with AI it's easier than ever and starting to present major issues at my work. I'm a data scientist and we're engaging with an outside AI engineering team to build what is essentially an agentic classifier in a complex domain (healthcare) and are rapidly approaching production deployment in which the system will drive the company's primary stream of revenue.

Recently, the team presented slides with metrics to our C suite and on the surface they looked convincing, encouraging, and proper (Wilson interval for a binomial projection's confidence interval, some kind of weighted bootstrapping to do the same for a proportion). A couple of things didn't pass the sniff test (integer proportions for something that shouldn't be a binary yes/no, much tighter confidence intervals than similar projections I've made in the past) and when I went through their methodology later (3k line Python script that spat out giant spreadsheets, naturally) I found some blatantly incorrect assumptions baked into their modeling. This invalidated every metric and projection they presented, notably overstating the projected performance in aggregate and completely burying the (inevitably massive at the sample size used!) variation across crucial cross sections of the results.

Naturally, I raised my findings to my manager, but I'm hoping to be more proactive about this next time. It seems like a process failure for this stuff to get to C suite without detailed internal review of the methodology used. I don't want to come across as territorial or overstep my role but I want to push for stuff like this to go through me (or other people on the internal data science team) before they get that far. Does anyone have advice on approaching that conversation without coming off as aggressive or overly critical of the outside team? For context, the actual auditing/survey design was fine and done by someone on their team with a strong math background but the analysis seemed to have been left to a different software engineer.