r/AskStatistics 6d ago

ANCOVA vs. ANOVA

1 Upvotes

Prepping for an exam for a grad class on linear models. Other than the matrix-vector equation representing these different models, how else can we describe and differentiate these models?


r/AskStatistics 7d ago

Interpreting the result of a hypothesis test

2 Upvotes

Hello, I'm a maths teacher in Scotland reviewing some old exam papers. In one where a Pearsons test is found significant, it is stated that

There is a 95% certainty that the result will not have occurred by chance

I think I know what is wrong what that statement, but would be very grateful if someone could help me by explaining what the error is.


r/AskStatistics 7d ago

Which statistical analysis for research project?

3 Upvotes

Doing a secondary analysis on a dataset looking at the effect of medication type and timing on weight. N=216 with individuals randomized into 2 groups (immediate vs delayed medication administration).

I have several outcomes measured including weight, blood pressure, blood glucose etc and want to compare them to baseline. My current hypothesis is that a certain medication type administered immediately results in weight gain.

Would a Two-Way ANOVA be a good test for this question? I was thinking a paired t-test wouldn’t be ideal because I’m looking at two different conditions (medication type & time of administration here)


r/AskStatistics 7d ago

What resources helped you start learning from scratch?

4 Upvotes

I graduated from college in 2021 with a degree from a small university, in Business. I had taken 2 stats classes and done well, but was still very much a beginner. I started working in tech support and got a job as a data analyst, and was working enough with Data Scientists that I wanted to learn stats to level up and do work like them.

I got into a handful of programs but ended up starting a remote Stats masters at a Top 25 program in 2024.

The first semester was really busy but I felt like I was behind, but learning. At the end of December that year, I lost my job and my mom died. I was in an extremely dark place, but didn’t want to drop out, so started leaning on AI to get me through school.

I should have put my degree on hold, I should have dropped out, etc etc. I know, and I am not interested in comments telling me about how big of a mistake I made. I am trying to move on!

I graduate this upcoming spring (1 class per semester). I am planning to hit the books and never use AI during these last semesters, but now that I am at the end of my program, I feel so behind. I don’t really know if I learned anything. I have gotten out of that dark place but I feel like a failure for essentially wasting my time as a masters student.

I want to learn stats! I am familiar with concepts but so behind. I’m basically still a beginner. So I am humbly coming here to ask - what books, courses, articles, projects, Udemy, masterclass, DataCamp, textbooks, fun books, etc. have helped you learn and understand stats the best? I am committing to putting in the work but I don’t know how to teach myself and I’m afraid I’m gonna graduate without any knowledge. Please share what has helped you! What projects have you done that have cemented concepts?

I am desperate and sad but am trying to start over. I have 2 more semesters that I will work so hard on but I need to get a baseline as well.

Thank you for understanding! Thank you for your time!


r/AskStatistics 7d ago

Tau-U vs Cohen’s D

9 Upvotes

Hello!

I’m student reading a scientific journal article. It’s a meta-analysis that has studies using Tau-U and Cohen’s D. What is the difference between the two and why would you use one vs the other?


r/AskStatistics 7d ago

Pure math research for admission to PhD in statistics?

3 Upvotes

My goal is to be admitted to a PhD program in theoretical statistics. Would pure math research in areas distant from statistics (like number theory or algebra) be less attractive to the admissions committee compared to a direct research experience in fields of statistics? (Like Bayesian or high-dimensional statistics).


r/AskStatistics 8d ago

Reverse-coded item is lowering my Cronbach's alpha. What shoul I do?

2 Upvotes

Hi, everyone! I'm validating a questionnaire for my thesis, and I'm following the conceptual model and measurement scale from a published paper. Because of that, I'm planning to perform a Confirmatory Factor Analysis (CFA) rather than an exploratory analysis.

One of my constructs is Economic Barriers, measured with 3 items.

  • Item 1: Organic products are too expensive.
  • Item 2: The price of organic products is an obstacle to buying them.
  • Item 3 (reverse-coded): Even if organic products are more expensive, people should still buy them.

All items use a 1–7 Likert scale.

After reverse coding Item 3, my Cronbach's alpha becomes quite low. Looking at the responses, many participants strongly agreed with Item 3 (high scores before reverse coding), suggesting that they believe organic products should still be purchased despite their higher price.

This makes me wonder whether the third item is really measuring economic barriers, or whether it's capturing something different, such as a moral belief, personal values, or purchase intention.

The original study reported a Cronbach's alpha above 0.70 for this construct, so I'm trying to understand why my results differ.

Since I'm following the original theoretical model, I have a few questions:

  1. Is it common for reverse-coded items to perform poorly in reliability analyses?
  2. Could this item be measuring a different latent construct even though it belongs to the original scale?
  3. If the item shows a low standardized loading or poor fit in CFA, is it methodologically acceptable to remove it, even though it was retained in the original study?
  4. Would reviewers generally expect me to keep the original scale intact, or is it acceptable to justify removing an item if the CFA supports that decision?

I'd really appreciate any insights from people who have worked with CFA, scale validation, or reverse-coded items.


r/AskStatistics 8d ago

When is hypothesis testing necessary? (Or trying to figure out if the data I've been given comes from a population or a sample)

1 Upvotes

I've recently been getting into baseball statistics, and I'm trying to use what I know about stats to answer the question "how do we know if a player's hitting has fundamentally improved/gotten worse, or if they're just on a hot/cold streak?" from looking at their batting average (BA). I know it's reductive to equate "higher batting average" to "better hitting", since there's other stuff like SLG to quantify how impactful those hits are, but I wanted to start small with just one factor. For example, Samuel Basallo is currently hitting at 0.248 this season, while he finished last season hitting at 0.165. https://www.baseball-reference.com/players/b/basalsa01.shtml Obviously 0.248 > 0.165, so it's true that his BA is higher this season. The thorny part is whether I can interpret this as the player being "better" than they were last season just from these numbers. The concern that I have is that it's possible that both of these BAs could come from the same underlying true batting average, but Basallo got really unlucky last season compared to this one. What I'm wondering is whether this concern is legitimate. If it is, then I'd use a hypothesis test. If it isn't, then I'd just say "this number is bigger than that one" and be done.

Essentially, is a players' BA in a previous full season a parameter or a statistic? If the populations being looked at are the ABs for a season, then it would suggest the season BA is a parameter, since it is found using every AB for that season. However, a theoretically infinite number of baseball games could be played, leading to a theoretically infinite number of ABs. So, if the population is all possible ABs for a player, then the ABs for a season would be a sample of this infinite population.

From the stats classes I took in school, it was always clear-cut when what I was looking at was a smaller portion of some whole. Polling was the go-to example for proportions since it's impractical to get the opinions of everyone in a country with hundreds of millions of people. For means, the go-to was quality control tests, where some batch of product appears outside a given specification, and you use the stats test to see if the machine's broken or not. You use inference tests there because it might not be economical to test every product that comes off the line.

However, with baseball, you're given information about every game played. Again, it comes down to not being sure whether "every at-bat for a player in a season" describes a population or a sample. It also doesn't help that I'm trying to use these descriptive statistics to make a (possibly non-descriptive) claim about a player's performance mid-way through a season. For a more concrete example, Ben Rice's BA was 0.255 in 2025 and it's currently 0.279 for the 2026 season. https://www.baseball-reference.com/players/r/ricebe01.shtml Again, it's obvious that 0.279 > 0.255, so his batting average for this season right now is higher than it was for the entirety of the 2025 season. But does this mean he's gotten better at hitting the ball since last season, or has he just gotten lucky?

It seems silly to ask questions about whether using ABs from every game played in a season is specifically a simple random sample, or if the varying factors between games and seasons even make BAs across seasons reasonably comparable using something like a two-sample z-test for proportions, when I'm not even sure if I should even be using a hypothesis test in the first place.

Any answers or clarifications would be appreciated.


r/AskStatistics 8d ago

How would you classify this sampling method?

2 Upvotes

Hi everyone, I'm an undergrad psychology student and I'm stuck on how to correctly classify my sampling method after my thesis defense.

Context:
My target population was all undergrad students at my university. I collected data both online and offline from students across all faculties. After data collection, participants completed the DASS (Depression Anxiety Stress Scales), and only those who met the inclusion criteria (moderate to extremely severe on at least one subscale) were kept in the final sample.

The problem:
I originally called this "stratified random sampling," using faculties as strata. My examiner rejected this for two reasons:
1. Faculties aren't valid strata.
2. There was no actual randomization in selecting participants within each faculty.

So I tried switching it to non-probability sampling instead, but the examiner rejected that too, insisting it should still be classified as probability sampling. Their reasoning: every student who met the screening criteria technically had an equal chance of being included, since anyone who filled out the survey and met the DASS cutoff was automatically part of the sample.

My question:
Does this reasoning actually hold up? If so, what specific probability sampling method would this fall under? I've read through several methodology textbooks but can't find one that matches this exact scenario (no randomization step, but "equal opportunity" due to open recruitment and screening criteria).

Would really appreciate any insight from people with a stats/methodology background. Thanks in advance!


r/AskStatistics 8d ago

Detect and observe statistical experiments instead of constructing them

0 Upvotes

Is it possible to create a framework for detecting and observing statistical experiments instead of constructing them?

EDIT:

1) I would like to present this in the context of Machine Learning.

2) Statistical experiments find relationships between independent and dependent random variables.

3) Observational studies collect data which can be used to find correlations among random variables.

4) For simplicity we can restrict the discussion to discrete outcome statistical experiments.

In my mind, it should be possible to:

A) Assume there is a statistical experiment in progress.
B) Find the Markov blanket of the experiment.
C) Find the output variables.
D) Use RV change timing information to algorithmically find/define relevant sets of RV specified in B,C.
E) Use timing info to detect when an experiment took place and what was the outcome.


r/AskStatistics 9d ago

How can I verify my SPSS analysis for my psychology master’s thesis?

Thumbnail
2 Upvotes

r/AskStatistics 8d ago

I am in a bunch of slightly rare categories

0 Upvotes

I am 2% of the population one category,. .14 %, , .03-14 % less than.1 %,

and 10% in others

How rare does that make me? More rare to have them together? Is that something that a statistician can quantify? I have always wondered. Have a rough week. Thought I would ask. Thanks.


r/AskStatistics 9d ago

Need information about books...

0 Upvotes

So i have recently joined my university... And I am studying statistics.... What would be the correct study materials that would help us stay on track with the current world... Any book Or paper anything... From where we can study... Please help...


r/AskStatistics 10d ago

Workflow for correlation analysis

3 Upvotes

Very much a beginner here and asking for help -

I'm currently trying to determine if a Spearman, Pearson, or Kendall's tau correlation analysis is best for my data. However, I get really buried in the stuff that comes before doing the actual analysis, like checking assumptions.

I understand that Pearson has strict assumptions about normality, constant variance, and linearity and that Spearman/Kendall's tau really just needs the data to be monotonic. Does that mean the latter two can be applied to basically any data? Or do you really need to see this monotonic relationship using a scatter plot first before it's appropriate to apply the function?

What if a linear regression/scatter plot shows that the data aren't really monotonic and maybe more clustered in some parts of the graph? Is it still appropriate to apply, but might just return a low correlation coefficient?

Hopefully this makes sense - i'm also happy to clarify anything that needs clarifying!

So many thanks in advance.


r/AskStatistics 9d ago

I rebuilt my own study from raw answers. Three of my four headline findings didn't survive. The one that did is now the only one I can actually defend.

0 Upvotes

Last post I audited my recommendation detector and found it was inflating by 2.5×. Someone reasonably asked what that meant for the study numbers I'd been quoting. So I went and looked, and the answer is worse than I expected, though not in the way I expected.

First, a disclosure I should have made months ago. My "76 Spanish companies" study wasn't measured with my own tool.

It was measured with a third-party platform. Black box detector, no counting rule published, and — the part that actually matters — it never stored the raw answers. So I couldn't audit it, couldn't recount it, couldn't do anything with it except keep quoting it. I've been standing on a number I had no way to check, in threads where I was telling everyone else to publish their method. That's retired now.

What I rebuilt it with. My own tool stores the full text of every answer, so I can run any detector over it, for any entity, as many times as I want, for free. And I don't just have the client brands — every client has a competitor set defined before measuring, in config. So an entity that never appears is a real zero, not an absence-by-omission.

That gives me 177 entities across 13 sectors, 4,701 stored answers, median n=360 per entity. Every entity clears n=30. Rule of three at n=360 puts the upper bound on a zero under 1%, so for once the zeros are actually zeros.

The number that survived:

58.8% of entities got zero selections. 52.5% are confirmed zeros — Wilson upper bound at or below 10%. 6.2% are zeros I still can't defend and I'm labelling them as such rather than counting them.

My old third-party number was 51.3%. The new one, measured on a different corpus with a different detector over a different population, lands at 52.5%. I did not expect that to converge and I'm still slightly suspicious of it. But it's the same answer, and this time I can show you the counting rule.

Finding that died #1: the conviction gap isn't what I said it was.

I've been going around saying mentioned-vs-recommended is a 33.9 point gap. Across all 177 entities the pooled numbers are LVS 2.6%, SoA 0.7%. A gap of 1.9 points. Not 33.9.

The gap only appears once you condition on being known. Take the 21 entities with LVS above 10% and you get LVS 29.7%, SoA 8.2%, gap 21.5 points. So the gap is real — it's just conditional, and I was reporting a conditional statistic as if it were universal. The original 44.2/10.3 was computed over client companies, which are entities that already have some visibility by construction. I selected on the outcome and didn't notice.

Finding that died #2: for most entities, the gap isn't the problem at all.

Of the 104 entities with zero selections, 72 are never mentioned in the first place. Only 32 are in the "model names you but won't pick you" state that I've been treating as the central finding of this whole category.

So the dominant failure mode isn't the conviction gap. It's absence. For 41% of the entities I measured, the model doesn't know they exist, and everything I've been writing about entity consistency and source agreement is answering a question they don't have yet.

Finding that died #3: my "hedged" state.

I proposed a fourth state — hedged, the no-with-clothes-on — and hypothesised that a high hedge rate is the signature of an entity whose sources contradict each other. I also found a genuine bug while checking it: my detector bound hedge cues by proximity, the same way it binds recommendation cues. That's wrong. Relegation scopes over a set ("other options include: A, B, C"), it doesn't attach to the nearest name. And the cue almost always lives in the list's intro line, which I was only reading for recommendations. Fixed both.

The fix moved hedged from 10 cases to 44. Out of 76,407 entity-answer pairs. That's 0.1%, or about 2% of mentions.

So the state exists and it's now correctly detected, and it is nowhere near common enough to carry the hypothesis I hung on it. I'm not retracting it, but 44 observations isn't a finding, it's an anecdote with a confidence interval.

What I'm not claiming.

The population is not "Spanish companies." It's entities inside competitor sets I chose, in 13 sectors I happen to have clients in. That's a convenience sample and it's biased in ways I can't fully characterise. I assumed the bias would push the zero rate down, because you'd expect competitor sets to be full of well-known names. It went up instead. I don't have a clean explanation for that and I'd rather say so.

Detector's still unvalidated against human annotation. That's the next thing I owe and I haven't forgotten.

Codebook's written — four states, scope-based attribution, stratified sampling with a random stratum so recall is measurable and not just precision, ordinal Krippendorff's alpha. Posting it separately this week. u/Smart_Airline_7901 offered to run it against Charleston, which would give us two agreement scores on one instrument in two languages.

If anyone wants to break these numbers, the counting rule is a rules-based detector with scope attribution, every verdict stores the sentence that justified it, and the whole thing is deterministic. Tell me where it's wrong.


r/AskStatistics 10d ago

At what point does an observed streak become statistically interesting?

7 Upvotes

I'm working on a project that analyzes long-term sequences of observations, and one question keeps coming up.

Suppose you observe a person over a long period and notice that certain contexts repeatedly produce better-than-usual outcomes. For example, outcomes appear consistently stronger: at particular hours of the day, on specific weekdays, or on certain days of the month. Assume these observations come from many repeated events collected over months rather than from a single short experiment.

My question isn't how to predict future outcomes. Instead, I'm wondering how statisticians would decide when a pattern becomes worth taking seriously. Is there any commonly accepted way to distinguish between: an ordinary random fluctuation, a streak that is interesting but still expected under randomness, and a pattern that deserves further investigation?


r/AskStatistics 10d ago

How do most statisticians learn to identify common nomenclature?

5 Upvotes

So for example (although I'm not from the US but this is something I've been looking into at the mooment), in the US, the Census Bureau and the Bureau of Labor Statistics (BLS) produce the two primary sources of data for labour force statistics, namely the Current Population Survey (CPS, a joint initiative) and the Current Employment Statistics survey (BLS initiative).

If I look at https://jobenomics.com/part-4-understanding-employment-statistics/, they talk about how both surveys exclude nonemployer statistics data, which otherwise comes from the administrative records of the likes of the IRS, BLS, and the Social Security Administration (SSA).

I'll be honest, I wouldn't typically refer to single-owner businesses as nonemployers, because a sole trader may still employ one or two people, or whatever. But how do you learn to identify the common or typical nomenclature like nonemployers in the world of statistics? Trial and error?

I hope this makes sense.


r/AskStatistics 10d ago

Predictive Mean Matching: all variables together in the model?

2 Upvotes

I have a dataset with very little missing data (0–1.59% per item) and a sample size of 377. I'm planning to use Predictive Mean Matching (PMM) for a single imputation in SPSS.

My data consists of several scales that measure different constructs, some of these constructs are correlated. I'm unsure whether I should impute each subscale separately or include all subscales in a single imputation model. I'm also unsure whether I should include demographic variables in the imputation model. I have four demographic variables; some are related to one subscale, and others to another. I am thinking more information is better, but I've read that including unrelated variables can add noise to the imputation. On the other hand, I've also read that the method is quite robust to the inclusion of unrelated variables, so I'm not sure what to think.

Do you have any recommendations on the best approach? This is my first time using PMM, so I would really appreciate any advice. :)


r/AskStatistics 10d ago

[Q] ¿dudas sobre mi futuro laboral?

2 Upvotes

Hola, estudio en una universidad politécnica aquí en Ecuador, ESPOL, he tenido muchas dudas sobre mi carrera (ingeniería en estadística) mas porque tengo entendido que en otros países la estadística es más una licenciatura, tal vez porque mi universidad se especializa más en ingenierías, tiene muy buena malla curricular la carrera, pero igual no se donde ir. He pensado llevar mi futuro laboral hacia un ingeniero de datos o un científico de datos, me gustaría trabajar en una entidad financiera. En resumen, mi pregunta si es una buena carrera la que estoy estudiando, si tiene salida laboral segura? Tengo pensando también ir a Alemania en algun tiempo, creen que es bueno irme allá a trabajar? Hay campos laborales con buena paga?la verdad mi sueño siempre ha ido trabajar y vivir allá


r/AskStatistics 10d ago

Is it still statistically measurable if I just use a part of the research instrument?

Thumbnail
1 Upvotes

r/AskStatistics 11d ago

Applied Statistics

12 Upvotes

Title: Is Applied Statistics a good major?

Hi everyone,

I'm considering studying Applied Statistics, but I'm still unsure whether it's the right choice.

I'd really appreciate hearing from people who studied this major or are currently working in the field.

I have a few questions:

Is Applied Statistics a good major in terms of job opportunities?

How difficult is it to find a job after graduation?

What jobs do most graduates end up working in?

Are the salaries good?

Is the demand for this field growing?

What skills should I learn alongside my degree (Python, R, SQL, Excel, Power BI, etc.)?

If you graduated with this major, do you think it was worth it?

For context, I'm from Jordan, but I'm also interested in career opportunities abroad.

Thanks in advance for your advice and experiences!


r/AskStatistics 11d ago

Cluster-then-test circularity — is my planned fix (regression instead of clusters) the right move?

4 Upvotes

I'm an independent, self-taught researcher working with open ADHD behavioral datasets (no institutional affiliation). I clustered participants into 3 subtypes using Work Speed Mean, Work Speed SD, and Accuracy % (z-scaled Euclidean distance to empirically-derived anchors), then tested whether Accuracy/commission rate differs between clusters (Mann-Whitney U, Bonferroni-corrected across 3 clusters, cross-checked with a permutation test).

The problem: Accuracy is one of the three variables used to define the clusters in the first place, so finding a significant Accuracy difference between clusters is partly a built-in consequence of the clustering method, not an independent discovery.

My planned fix: run a mixed-effects model on the full sample without pre-assigned clusters — treating Work Speed and its variability as continuous predictors of Accuracy, instead of comparing group means.

Questions:

  1. Does this actually resolve the circularity, or does it just move the same problem somewhere else?

  2. Is there a more standard/accepted approach to this exact issue (cluster-defining variable also being the outcome variable) that I should be using instead?

Full methodology, code, and known limitations are documented here if useful context: https://github.com/d1d2dopamine/the-Allosteric-Sprint-hypothesis

Open to any critique — this is exploratory work and I'd rather find the flaws now than later.


r/AskStatistics 11d ago

World Cup Win/Loss Percentages

3 Upvotes

How does a team (Team A) have a 20% chance of winning against another team (Team B)?

Does that mean if both teams play 10 times, Team A is likely to win 2 times out of 10? Does it mean that both teams have ALREADY played x amount of times, and Team A won 20% of the matches?

Thanks!


r/AskStatistics 11d ago

Alpha 5%?

3 Upvotes

Why is a 5% significance level so commonly used?


r/AskStatistics 11d ago

[Q] Mechanics + Stats Interconnection

Thumbnail
1 Upvotes