Browse guidance on statistical fundamentals and variation, practical advice on study design and preparing data and tools for visualisation and exploration.
Fundamentals, context and study design
-
Thinking about variation
Variation is at the heart of statistics. It’s everywhere, and without it, data are not very interesting. Understanding this is central to the tasks we carry out in visualising, summarising and analysing data.
-
Thinking about distributions
What makes the variable random is that — unlike the kind of variable we see in a quadratic equation, for example — we cannot say what the observed value of the random variable is until we actually carry out the procedure.
-
The three Rs of study design
“You can’t fix by analysis what you bungled by design.” Light, Singer and Willett (1990, page v). Good study design is paramount. Here are the three Rs of study design.
-
Take care with context
A basic lesson in data recording and management is to carefully and accurately describe the variables in a data file. This includes describing the measurement scale used and the units for numerical variables.
Preparing your data
-
Starting with Excel
If you are bringing your data along to a consultant, there are several things you might consider doing before you meet.
-
Errors in the data
Before a consultant can start on serious analysis of your data, it is vital to be confident that the data are "clean".
-
Outliers
Some textbooks recommend removing outliers or even "adjusting" outliers. We strongly recommend that you do not remove or adjust data.
-
Statistical software
Statistical software for researchers at the University of Melbourne.
-
Missing data
'Missing data’ refers to data that was intended to have been collected but was not.
Data visualisation and exploration
-
Five principles of good graphs
You may have met five principles for creating good graphs in a course or seminar given by the Statistical Consulting Centre.
-
Why you shouldn’t use pie charts
Pie charts and their close cousins, doughnut charts, are part of our staple diet of graphs of data – we find them on blogs, in newspapers, in textbooks and in presentations at work.
-
Graphing complexity
Putting together a graph to summarise a data set where multiple explanatory variables are involved.
-
Understanding boxplots
A boxplot is a visualisation of a numerical variable based on summary statistics.
-
Error bars on graphs
In many publications, you will see error bars around an estimate, such as a mean or a mean difference.
-
Tricks for plotting confidence intervals in Minitab
Minitab can provide confidence intervals for means, for example, via the Graphs > Interval Plot menu.
Effect sizes, P-values and confidence intervals
-
What is a P-value?
P-values come from testing theories using sample data.
-
How to report a P-value
You will routinely find P-values in the output of statistical software.
-
There's more to life than statistical significance
Finding a small P-value in the results of an analysis.
-
What is an effect size?
There is a renewed emphasis on reporting effect sizes, with the hope that this will encourage researchers to rely less on statistical significance.
-
What is a confidence interval?
Confidence intervals appear in reports of statistical analysis, academic papers and sometimes in mainstream media.
-
Interactive app: Confidence intervals and P-values
A confidence interval contains a range of values for the parameter of interest that are plausible, given the data.
Statistical practice
-
Reproducible analysis
Terms like "reproducible research” and “open research” get used a lot and may mean documenting methods, publishing data and code so that others can check the conclusions, or being able to redo analysis at some later point.
-
Worried about Normality?
A concern about "normality" often arises when using a statistical analysis for a numerical outcome such as an independent samples t-test, analysis of variance, regression or linear model.
-
Understanding levels of variation and mixed models
Using a statistical model that doesn’t account for the hierarchical nature of this data will give incorrect results.
Reporting statistical inference
-
Inference for two independent samples (t-test)
An example of inference for two independent samples, a test of the theory of Lamarckian inheritance.
-
Linear regression
An example of linear regression. The data for this example comes from measurements made by the US Federal Trade Commission on 25 different varieties of cigarettes.
-
Rescaling explanatory variables in linear regression
Have you ever seen a very small estimated regression coefficient (e.g. 0.0023) for an explanatory variable that is reported with a very small P-value (e.g. P < 0.001)?
-
One way ANOVA
This example uses a dataset called ‘chickwts’. It is one of the inbuilt datasets provided in R software, and is popular for illustrating methods of one-way analysis of variance.
-
Understanding two way interactions
A concern during the Second World War was the provision of vitamin C to soldiers, and in this broad context, the effects of ascorbic acid and orange juice were studied in animals.
-
Reporting two way interactions
Analysis involving a two-way interaction, using the example of Crampton’s (1947) study of the effects of dose and method of delivery of vitamin C on the development of teeth in guinea pigs.
-
Logistic regression
Logistic regression is used to model a binary response variable in terms of explanatory variables.
-
Transforming explanatory variables in logistic regression
Have you ever seen an estimated odds ratio that is very close to 1 for a numerical explanatory variable which is reported with a small P-value?
Tips for using R
-
Plotting multiple variables
In exploratory data analysis, it’s common to want to make similar plots of a number of variables at once.
-
Basic inference in R
Navigating R can be daunting with the vast number of packages and options for obtaining analyses.
-
Fitting many statistical models at once using dplyr
One common task in applied statistics is to fit and interpret a number of statistical models at once.
Teaching and learning resources
-
RealStat
Real case studies based on genuine consulting projects. Learn about the planning and design of each study, the methods of data collection and the data that are available for analysis.
-
What is statistics?
Statistics is a discipline concerned with designing experiments and other data collection, summarising information, drawing conclusions from data, and estimating the present or predicting the future.
-
OzDASL: Australasian Data and Story Library
A library of data sets and associated stories with the greatest emphasis on the Australasian context.
-
Data and Story Library
An online library of data files and stories that illustrate the use of basic statistical methods.