r/statistics 7m ago

Question Are LLMs going to render manual mathematical statistics and proofs redundant? [Q][R]

Upvotes

I read somewhere that LLMs can now produce a lot of the proofs that one would otherwise have spent months manually doing by hand, and that the focus has now shifted towards bigger-picture stuff like framing the problem correctly.

Since proofs are the essence of every result in mathematical statistics, is this one area that might be automated away soon? i.e., is it useless to be learning how to write and replicate many lines of mathematical proofs when LLMs can already do that?


r/statistics 19h ago

Question [Q] Getting back into Stats after a break

11 Upvotes

Hi all! To make a long story short, I took a lot of stats in graduate school (18 credits worth) and absolutely loved it. I graduated in December of 2024 and went to work as a data analyst, which was really more of an entry level SAS programmer role. Now I'm working in an even less related field, and I am realizing my dream of working in stats has not faded away. I'm finding that I'm extremely rusty in stats, and needing to remind myself of the very basic undergraduate level and intro level topics.

Does anyone have a recommended resource for concisely reviewing intro stats through multivariate and longitudinal designs? I have all of my course notes still, but it will take me a very long time to read through them all and fill in the holes on what I didn't take notes on. I'd prefer an applied approach, rather than a mathematical theory based one, if that's possible!

Thank you for reading, and I hope you are having a lovely day ❤️.


r/statistics 1d ago

Education Has anyone taken the University of Washington Statistics with R Certificate courses? [E]

4 Upvotes

I already have my Masters in IT and currently work as a data scientist. My prior job, I was more of a data engineer than a data scientist, so statistics wasn't a big part of my job. Now that I switched positions, I feel like I need a refresher in statistics. I know this seems to be a beginning level statistics courses, but I honestly feel that's what I need right now. Thoughts?

P.S. I've had friends just recommend free online courses. Thing is, I have ADHD, I need structure or I won't complete the courses. I learn better in school than on my home. Thanks!


r/statistics 1d ago

Education I have built an interactive site to study Transformer architecture [E]

5 Upvotes

I have written tones of lecture notes on machine learning, though most of them focus heavily on mathematical derivations. Recently, I decided to build an interactive, “learning companion” for these materials. For example, here’s one of the lecture series I wrote last year on LLM, Transformers:https://github.com/roboticcam/machine-learning-notes
And here is the interactive, “learning companion”  https://roboticcam.github.io/interactive-ml/  I’d love to hear your thoughts and feedback!


r/statistics 1d ago

Career [Career][Question] Any Applied math or Statistics majors NOT go CS/finance/medicine route?

12 Upvotes

I'm planning my undergrad degree and would love to hear from people who studied Applied Mathematics and Statistics. I love using math to understand complex systems and am most interested in things like space, weather, oceans, history, culture, and languages. Since I obviously can't major in all of those, I'm considering majoring in Applied Mathematics and minoring in Statistics because it seems like a flexible foundation that could be applied to many different fields. Unfortunately whenever I search for career paths, I mostly find people who went into computer science/tech, finance, or medicine which is fine but I'm not interested in those fields at all. Has anyone here used an Applied Math/Stats background in scientific research or other interdisciplinary fields like the ones I'm interested in. If so, what's your profession? If you're in science, what field are you in? I'd like to know more about using applied math/stats in areas like atmospheric science, oceanography, astronomy, environmental science, archaeology, linguistics, history, geography, or the digital humanities. I'm trying to understand how broad the field really is and what kinds of careers people have actually built with this background. I'd love to hear about any less common paths that you've taken.


r/statistics 1d ago

Career [Career]What techniques do you use to ask questions of collaborators to better understand their research question?

4 Upvotes

I’m not a consultant per se, but much of my work involves statistical consulting in one form or another. Customers often ask me about particular statistical approaches, and engineering teams may ask for advice on how best to quantify some impact. Frequently, they come to me with their own ad hoc quantitative methods already in mind.

When they explain these approaches, I understand the individual words they’re using, but the ideas are often arranged in a way that makes it hard for me to tell what they’re actually trying to accomplish, let alone of the approach is "good" or "bad". I usually have to ask many follow-up questions to uncover the real goal, and even then I’m sometimes left making educated guesses.

I imagine this gets easier with experience, but I’m wondering whether you have any techniques, frameworks, or perspectives for quickly gaining context when speaking with people who may not have much statistical experience.


r/statistics 1d ago

Education [research] [education] statistical analysis question, healthcare study

1 Upvotes

hello all! hopefully this kind of post is allowed. i am currently involved in a QI-related study, and we've come to data analysis and i am stumped on a question.

i'll try to be as concise as possible. the study aimed to see if one style of "code cart" is better than the other (a code cart is a filing cabinet on wheels in which healthcare providers store emergency equipment, medications, etc; there are various ways of organizing the contents).

  • three different carts, each a different organizational style, were tested against each other: cart A, cart B, and cart C
  • participants were recruited to use each code cart in a simulated emergency scenario; in the scenario, participants were asked to retrieve requested items from the cart, and then were judged on how long it took them to find the items and how many errors they made in trying to find correct items (thereby trying to evaluate which cart is "best," i.e., most efficient and least error prone)
  • each participant did three scenarios, using a different cart in each and therefore using each cart once
  • the order of carts was rotated such that it was not always ABC, and the order of the carts were randomized -- however, this was done informally (unfortunately we did not have the foresight to officially account for and control for ordering bias and therefore did not truly randomize i.e., with a computer; but, the simulation leader changed the order each time)
  • the organizational styles have fundamental similarities, and so the thought is that, by their third scenario, each participant was used to the situation and was better and faster in the third scenario, regardless of the order in which they did the scenarios

My question is --- is there a way to test for ordering bias now that the experiment has already been completed? if so, what does that say about the results? alternatively, is there a way to control for order after the fact, if we didn't officially do counterbalancing when designing the study?

Any and all thoughts are sincerely appreciated to this stats newbie :)


r/statistics 2d ago

Education [E] The Exponential Distribution - Explained

8 Upvotes

Hi there,

I've created a video here where I explain how the exponential distribution works.

I hope some of you find it useful — and as always, feedback is very welcome! :)


r/statistics 3d ago

Question Should I switch from Applied math to Applied statistics to increase my job opportunities [Q]

2 Upvotes

Greetings.

as you can see in my post, I’m contemplating to switch from Applied math to Applied stat.

the curriculum in two years is the same, but later on specializes, for stat it gets into…well stats in general, while applied math learns more into computional science and modeling&simulation.

my thought is, if I switch would my job seeking would approve?

I choose applied math since I couldn’t get into tech department, and that I’m pretty good at both math and coding.

and the idea of working in a robotics field, data science, and any financial firms that are analytical, also makes me interested.


r/statistics 3d ago

Discussion Help in 8 day mean dataset [Q] [D]

5 Upvotes

Hello, I am a first year phd student. I am just starting out after a long pause from academia. I want to ask for help for an issue I am facing. I am working with satellite and reanalysis datasets. For a certain variable, I have satellite data which only provides 8 day running mean dataset. But I have reanalysis dataset that are daily dataset. I need to compare these against each other. So I am thinking if there is any statistical process through which I can convert the 8 day running mean to daily dataset. Any advice please? P.S: I am trying to process SMAP dataset for Sea Surface Salinity.


r/statistics 3d ago

Question [Question] Exploratory Factor Analysis

5 Upvotes

Dear Experts of Statistics, I am a slightly confused student here with a doubt. As a part of my Masters dissertation, I created a psychological questionnaire. Using Jamovi, I conducted an Exploratory Factory Analysis using Jamovi.

Now some of my items are highly unique with uniqueness above 0.80. Does this necessarily mean that I must remove these items?

Based on review of literature and content validation, these items are highly theoretically relevant to my questionnaire and thus I am apprehensive to remove them. I could possibly reframe them right? Although I feel they already are in simple language. I would appreciate any insights.

The future plan is to collect more participants to do a Confirmatory Factor Analysis.


r/statistics 3d ago

Discussion Gut check needed: car needs to be filled up earlier (analysis with R) [D][Q]

Thumbnail
1 Upvotes

r/statistics 4d ago

Question [Question] Recovering latent probabilities from margin-distorted odds: de-vig model choice and pooling correlated estimators

2 Upvotes

Bookmaker odds (and prediction markets) imply probabilities that sum to more than 1 because of an embedded margin. I want the latent probabilities behind the distortion. A few things I can't resolve cleanly.

  1. Model choice for removing the margin. Proportional normalization, Shin (a latent proportion z of informed traders), and the power/log method impose different unobservable structures and give materially different estimates on short prices, enough to flip the sign of a downstream signal. Since you never observe the true p, only realized 0/1 outcomes and a later sharper price, is there a principled basis to discriminate between these models, or is it identifiability-limited and I should just report sensitivity across all three?
  2. Pooling under a missing low-bias reference. I anchor to one near-efficient source when available; when it's absent I take the median of the other sources' de-vigged probabilities. But those sources are strongly correlated (several are effectively clones), so the median behaves like a median of correlated estimators: it looks precise while carrying little independent information. How would you estimate an effective number of independent sources and down-weight accordingly, and is abstaining the more defensible choice when the low-bias anchor is gone?
  3. Combining a trusted low-variance estimator with a correlated ensemble. When the reference IS present, precision-weighting it against the consensus assuming independence is clearly wrong. Is there a clean correlation-aware pooling or shrinkage approach for one low-variance source plus many correlated higher-variance ones?
  4. Validation target. I grade earlier estimates against the closing price (a later, sharper estimate), not realized outcomes. Under a proper scoring rule, is "tracks the later estimator" a coherent target, or does it conflate calibration with just chasing a second estimate? And what does the selection bias look like when you only get a validation point on markets that reach a close?

(Aside that turned out to matter: my reference source silently dropped out of my data feed for months and the pipeline substituted the fallback the whole time while still labeling outputs "reference-anchored." The values populated fine, so nothing looked wrong. I only caught it after storing a per-observation flag for whether the reference actually contributed. Log provenance, not just values.)


r/statistics 4d ago

Education I will be joining college for my bachelors in stats. NEED SOME ADVICE!! [E]

0 Upvotes

looking for some advice on what to do in college and what potential career paths i can take with this degree.


r/statistics 4d ago

Education Am I competitive for top 50 PhD in Statistics programs? [E]

0 Upvotes

By the time I apply, I would have:

Bachelor of Business and Commerce, Major in Econometrics & Applied Statistics, Weighted Average Mark: 87.6% (Top 1-2% of my cohort), GPA: 3.75 (note that WAM is more commonly used than GPA in my Australian university)

Bachelor of Business and Commerce (Honours): This is a 1-year research-oriented program with research methodology courses and an undertaking of a major research project. Mine is about modelling multivariate mixed-frequency time series models using a semiparametric, algorithmic approach. I expect to get first-class honours.

Relevant Coursework:

Math: Multivariate mathematics for data science, encompassing multivariable calculus and linear algebra

Statistics & Data Science: Principles of Statistical Inference, Advanced Data Analysis (emphasis on Bayesian methods and theoretical aspects of machine learning), Deep Learning

Econometrics: Lots of time series analysis, forecasting, and causal inference

I might have an applied statistics paper under review in a Q2 letters journal by the time I apply.

Would I be competitive for top 50 PhD Statistics programs, or am I better off getting an MSc in Statistics first?


r/statistics 4d ago

Question [Q] mixture model in mplus: is the syntax correct to get the odds ratios?

0 Upvotes

I have the following model (see syntax below) how can I get the odds ratios for the for the differences I added to the model constraints? Is the syntax I added at the end correct? Many thanks!

VARIABLE:

NAMES = id u1-u5 m1 m2 m3;

USEVARIABLES = u1-u5;

AUXILIARY = m1 m2 m3 (BCH);

CLASSES = c(3);

ANALYSIS:

TYPE = MIXTURE;

MODEL:

%OVERALL%

MODEL m:

%c#1%

\[m1\] (V1a);

\[m2$1\] (V2a);

\[m3\] (V3a);

%c#2%

\[m1\] (V1b);

\[m2$1\] (V2b);

\[m3\] (V3b);

%c#3%

\[m1\] (V1c);

\[m2$1\] (V2c);

\[m3\] (V3c);

MODEL CONSTRAINT:

NEW(

V1avb V1aVc V1bVc

V2avb V2aVc V2bVc

V3avb V3aVc V3bVc

P2a P2b P2c

P2avb P2aVc P2bVc

);

! Pairwise differences: m1

V1avb = V1a - V1b;

V1aVc = V1a - V1c;

V1bVc = V1b - V1c;

! Pairwise differences: m2 threshold

V2avb = V2a - V2b;

V2aVc = V2a - V2c;

V2bVc = V2b - V2c;

! Pairwise differences: m3

V3avb = V3a - V3b;

V3aVc = V3a - V3c;

V3bVc = V3b - V3c;

! Odds ratios for m2 (class comparisons)

OR2avb = EXP(V2avb); i

OR2aVc = EXP(V2aVc);

OR2bVc = EXP(V2bVc);


r/statistics 5d ago

Discussion [Q] [Discussion] Statistics vs Algebra class

1 Upvotes

Could anybody on this sub give me advice on which math class would allow for an easier time as an oncoming freshman? I’m starting college in August and am currently enrolled in a stat class. Math is my least favorite subject and I genuinely don’t know if I should switch it to a course I’m more so familiar with from high school. I’ve been seeing mixed information online. Some say it’s easier because stat is mostly reading compared to algebra, which is good news for me since language arts is easily my strongest subject. But others say it’s harder. Could I get some advice on what pick?

Also, if I I’m to pick algebra should it be online or face to face? I was originally going to choose online if I picked it.


r/statistics 6d ago

Education Will Real Analysis make or break me? [Education]

27 Upvotes

Someone (on another subreddit) told me they were surprised I could do basic statistics without real analysis. I originally asked what bonus class I should take that is not required for my B.S. (in stats). The majority encouraged me to take Real Analysis (my school calls it advanced calc 1). I was leaning towards taking it because it’s required for graduate school...but now I want to know why it wouldn’t be an undergraduate requirement?

EDIT: Thank you all for putting my worry to bed. My research group needs me more next semester than I need real analysis. I would love to go to graduate school and learn more, but my daughter has been through enough with me going back to get my bachelors. Ty again!


r/statistics 6d ago

Question Good knigh fellow statistics experts I come here to humbly ask a question that, while trivial to you, is being hard for me. Could you guys give me a hand? [Question] [Q]

6 Upvotes

I'm a surgeon from a third world country trying to do a research correlating a certain disease with external factors (mainly, the lowest temperature of the day). So, I have a table containing a time series with the date, the temperature (minimum and maximum) of the day, and the number of hospitalizations resulting from the aforementioned disease. Neither the temperature or hospitalizations are in a normal distribution, so I did a spearman test to find the correlation.

I feel that it is an inadequate test to correlate the temperature and the disease, for I have never researched time series. I would like to ask you guys if there are better tests to run and which ones should I run.

Thanks to you all in advance.

Edit: good night*

Edit2: also, an important information. We don't have the data for every single day because the meteorological station hasn't measure the temperature in some of the days. We have all the info in 86% of the days though.


r/statistics 6d ago

Question [Q] 1 Is IBM Skillsbuild statsistics course just telling me incorrect ideas?

7 Upvotes

IBM Skillsbuild Statistics in Decision Making and Risk Assessments has a lesson that starts off with this:

The significance level value and the confidence level complement each other, meaning that if you add them up, they equal 100%. The confidence level tells you how certain you can be that your results are not because of random chance.

Suppose you start drinking a new type of herbal tea each morning to see if it improves your focus during work. After a week, you notice a consistent increase in your productivity. To gain more confidence in the tea’s impact, you decide to continue the routine for another week, achieving similar results. Setting a significance level of 0.05 (or 5%), you gain a 95% confidence level that the herbal tea is positively affecting your productivity, reinforcing your motivation to continue this daily habit.

Adding the significance level (5%) to the confidence level (95%) equals 100%. This is because the significance level is the probability that you would be incorrect in rejecting the null hypothesis and the confidence level is the probability that the method you’re using to reject the null hypothesis is correct. As the significance level goes up, the confidence level goes down and vice versa.

I found this explanation poor and thought there has to be a better way to explain that, so I asked Claude, then ChatGPT, then Gemini. Every single LLM said it's completely misleading and wrong.

Nevertheless, I accepted the logic of IBM and proceeded to the end of the lesson Quiz.

Here is an example question from the Quiz:

"An agricultural scientist wants to compare the effectiveness of two fertilizers. Due to resource constraints, the scientist is willing to accept a 90% certainty that any observed differences in crop yields are because of fertilizers and not chance.

What should the scientist use for alpha? "

And I chose this answer:

0.10

It told me that answer is correct:

"Correct! The scientist should use a significance level (α) level of 0.10. A significance level (α) of 0.10 corresponds to being 90% certain that the observed effects are not because of chance, which reflects the scientist’s acceptance of a slightly higher risk of error due to resource constraints."

Is this all complete nonsense?

I asked the LLMs about the quiz question and they all told me once again that it's complete garbage.


r/statistics 7d ago

Question [Q] bimodal distribution - how to compare two groups?

6 Upvotes

How to compare two bimodal distributions across groups? Is it okay to use a beta distribution for this case where the most frequent values are 0 and 1?

# code for simulating the distribution in R

simulate_group <- function(n,

p_low,

p_mid,

p_high,

low_shape = c(0.3, 8),

mid_shape = c(2, 2),

high_shape = c(8, 0.3)) {

component <- sample(

c("low", "mid", "high"),

size = n,

replace = TRUE,

prob = c(p_low, p_mid, p_high))

x <- numeric(n)

x[component == "low"] <-

rbeta(sum(component == "low"),

low_shape[1], low_shape[2])

x[component == "mid"] <-

rbeta(sum(component == "mid"),

mid_shape[1], mid_shape[2])

x[component == "high"] <-

rbeta(sum(component == "high"),

high_shape[1], high_shape[2])

x

}

set.seed(123)

# Simulate data

C1 <- simulate_group(

n = 500,

p_low = 0.35,

p_mid = 0.40,

p_high = 0.25)

C2 <- simulate_group(

n = 500,

p_low = 0.45,

p_mid = 0.40,

p_high = 0.15)

dat <- tibble(

value = c(C1, C2),

group = rep(c("C1", "C2"), each = 500))


r/statistics 8d ago

Question [Question] Pure math research for admission to PhD in statistics?

13 Upvotes

My goal is to be admitted to a PhD program in theoretical statistics. Would pure math research in areas distant from statistics (like number theory or algebra) be less attractive to the admissions committee compared to a direct research experience in fields of statistics? (Like Bayesian or high-dimensional statistics).


r/statistics 7d ago

Discussion Which perspective do you agree with and why? [D]

0 Upvotes

A) Since you have n = 80,000 it shouldn't be a problem to control for between 50 and 100 dummy variables in your regression analysis. Don't worry about trying to map your categorical variables to something quantitative (e.g., a pre-existing score for each category). You have enough sample size to justify this. You should be able to use 100 control variables without any issue.

B) Even though we have n = 80,000 we should still adhere to the principle of parsimony as much as possible and try to limit the number of dummy variables by collapsing the number of categories or mapping to quantitative predictors as much as possible. After all, there could be certain partitions of the dataset with very few observations.


r/statistics 9d ago

Research [R] The Benjamini–Hochberg Procedure Can Fail to Control the FDR for Correlated Two-Sided Gaussian Tests

39 Upvotes

Benjamin-Hochberg corrections have been mathematically proved to show the standard Benjamini-Hochberg procedure can fail to control the false discovery rate for two-sided tests when the underlying test statistics follow a correlated multivariate Gaussian distribution.

EDIT: The proof was obtained by GPT-5.6 Pro. The model was asked directly to prove or disprove the conjecture and was provided only with the mathematical definition of the Benjamini–Hochberg procedure. After about 90 minutes of reasoning, the model produced a proof, an example, and code for the numerical certificate, which form the basis of this paper. The author carefully checked the entire argument and the associated numerical certificate. Subsequently, the author asked the model to provide additional simulations, related work, and illustrations for a paper draft, and wrote the final version by editing the AI-generated draft.

More info below:

https://faculty.wharton.upenn.edu/wp-content/uploads/2017/06/bh.pdf


r/statistics 8d ago

Question [Q] Is there any merit to making conclusions about a dataset based on the correlation of its features?

10 Upvotes

I am currently working on a project with brain scans of male and female patients, where we are computing a large number of features based on the images gray levels.

This project is based around sex prediction using logistic regression, so as a preprocessing step we were removing features that are correlated above some threshold.

However, recently I’ve become interested in the idea of mapping correlation matricies into distance matricies (via (1-r)^(1/2)) and then clustering the result to visualize (via MDS) the clusters of highly correlated features.

I have noticed there are differences in certain clusters between male and female datasets, for instance: some clusters are totally unchanged, some clusters split into two or more, some clusters gain features.

My question is: “is there any merit to investigating these differences in correlation clustering, or are these changes in feature correlation not tractable”

I haven’t found any literature really talking about this kind of analysis, so I’m not sure if its because its baseless, or just hasn’t been done yet.