• BB852
  • 1 Preface
    • 1.1 How to use this book
    • 1.2 Data wrangling
    • 1.3 Data visualisation
    • 1.4 Statistics
    • 1.5 Data sources
    • 1.6 Your instructor(s)
    • 1.7 Expectations
    • 1.8 Your feedback
    • 1.9 Assessment
    • 1.10 Acknowledgements
  • 2 Schedule
  • 3 Getting set up
    • 3.1 R and RStudio are two different things
    • 3.2 Installing R and RStudio
    • 3.3 A quick tour of RStudio
    • 3.4 Get the course data
    • 3.5 Key takeaways
    • 3.6 Common pitfalls recap
    • 3.7 Mini-project
  • 4 Paths and projects
    • 4.1 File paths in plain language
      • 4.1.1 Absolute paths (full address)
      • 4.1.2 Relative paths (recommended)
    • 4.2 Use an RStudio Project
      • 4.2.1 How to create a project
    • 4.3 Suggested folder structure
    • 4.4 Key takeaways
    • 4.5 Common pitfalls recap
    • 4.6 Mini‑project
  • 5 An R refresher
    • 5.1 Getting started with RStudio
    • 5.2 Getting help
    • 5.3 R as a calculator
    • 5.4 Objects and vectors
    • 5.5 Manipulating vectors
    • 5.6 Missing values, infinity, and NaN
    • 5.7 Data frames
    • 5.8 Classes and factors
    • 5.9 Organising your work and importing data
    • 5.10 Inspecting the data
    • 5.11 Tables and summary statistics
    • 5.12 Basic plotting
    • 5.13 R packages
    • 5.14 Exercise: Californian bird diversity
      • 5.14.1 The data
      • 5.14.2 Your tasks
    • 5.15 Key takeaways
    • 5.16 Common pitfalls recap
    • 5.17 Mini-project
  • 6 Tips and tricks
    • 6.1 Appearance
    • 6.2 Shortcuts (top six)
    • 6.3 Code style (keep it readable)
    • 6.4 The plots pane
    • 6.5 Tables
    • 6.6 Importing data: common problems
    • 6.7 Key takeaways
    • 6.8 Common pitfalls recap
    • 6.9 Mini‑project
  • 7 Additional recommended reading
    • 7.1 Papers and chapters
    • 7.2 Useful websites
  • I Data Wrangling
  • 8 Data wrangling with dplyr
    • 8.1 select
    • 8.2 filter
    • 8.3 arrange
    • 8.4 summarise and group_by
    • 8.5 Using pipes, saving data.
    • 8.6 Exercise: Wrangling the Amniote Life History Database
  • 9 Combining data sets
    • 9.1 Using join
    • 9.2 Using pivot_longer
    • 9.3 Exercise: Temperature effects on egg laying dates
  • II Data visualisation
  • 10 Visualising data with ggplot2
    • 10.1 The data we will use
    • 10.2 Histograms
    • 10.3 Comparing groups in a histogram
    • 10.4 Facets (split across panels)
    • 10.5 Box plots
    • 10.6 Lines and points (trends over time)
    • 10.7 Scatter plots
    • 10.8 Bar plots (use with care)
    • 10.9 Key takeaways
    • 10.10 Common pitfalls recap
    • 10.11 Mini-project
  • 11 Distributions and summarising data
    • 11.1 Relationships in Data: Response and Explanatory Variables
    • 11.2 Populations, Samples, and Bias
    • 11.3 Probability, odds, and uncertainty
      • 11.3.1 Probability
      • 11.3.2 Odds
      • 11.3.3 Uncertainty and certainty
      • 11.3.4 Probability density
      • 11.3.5 Distributions
    • 11.4 Normal distribution
    • 11.5 Comparing normal distributions
    • 11.6 Poisson distribution
    • 11.7 Comparing normal and Poisson distributions
    • 11.8 The law of large numbers
      • 11.8.1 Coin flipping
    • 11.9 Exercise: Virtual dice
  • 12 Pimping your plots
    • 12.1 A basic plot
    • 12.2 Axis limits
    • 12.3 Transforming the axis (log scale)
    • 12.4 Changing the axis tick marks
    • 12.5 Axis labels
    • 12.6 Colours
    • 12.7 Themes
    • 12.8 Moving the legend
    • 12.9 Combining multiple plots
    • 12.10 Saving your plot
    • 12.11 Final word on plots
  • III Statistics
  • 13 Randomisation Tests
    • 13.1 Randomisation test in R
      • 13.1.1 Calculate the observed difference
      • 13.1.2 Null distribution
      • 13.1.3 Testing significance
      • 13.1.4 Testing the hypothesis
      • 13.1.5 Writing it up
    • 13.2 Paired Randomisation Tests
      • 13.2.1 The randomisation test
      • 13.2.2 Null distribution
      • 13.2.3 The formal hypothesis test
    • 13.3 Exercise: Sexual selection in Hercules beetles
  • 14 t-test: Comparing two means
    • 14.1 Some theory
    • 14.2 One sample t-test
    • 14.3 Doing it “by hand” - where does the t-statistic come from?
    • 14.4 Paired t-test
    • 14.5 A paired t-test is a one-sample test.
    • 14.6 Two sample t-test
    • 14.7 t-tests are linear models
    • 14.8 Exercise: Sex differences in fine motor skills
    • 14.9 Exercise: Therapy for anorexia
    • 14.10 Exercise: Compare t-tests with randomisation tests (optional)
  • 15 Assumptions in linear models
    • 15.1 The assumptions
  • 16 ANOVA: Linear models with a single categorical explanatory variable
    • 16.1 One-way ANOVA
    • 16.2 Fitting an ANOVA in R
      • 16.2.1 Where are the differences?
      • 16.2.2 Tukey’s Honestly Significant Difference (HSD)
    • 16.3 ANOVA calculation “by hand”.
    • 16.4 Exercise: Apple tree crop yield
  • 17 Linear regression: models with a single continuous explanatory variable
    • 17.1 Some theory
    • 17.2 Evaluating a hypothesis with a linear regression model
    • 17.3 Assumptions
    • 17.4 Worked example: height-hand length relationship
    • 17.5 Exercise: Chirping crickets
  • 18 ANCOVA: Linear models with categorical and continuous explanatory variables
    • 18.1 The height ~ hand length example.
    • 18.2 Summarising with anova
    • 18.3 The summary of coefficients (summary)
  • 19 n-way ANOVA: Linear models with >1 categorical explanatory variables
    • 19.1 Fitting a two-way ANOVA model
    • 19.2 Summarising the model (anova)
    • 19.3 Summarising the model (summary)
    • 19.4 Exercise: Fish behaviour
  • 20 Evaluating linear models
    • 20.1 R-squared value
    • 20.2 Akaike Information Criterion (AIC)
    • 20.3 Variance partitioning
    • 20.4 Conclusion
  • 21 Generalised linear models
    • 21.1 Count data with Poisson errors.
      • 21.1.1 Example: Number of offspring in foxes.
      • 21.1.2 Example: Cancer clusters
      • 21.1.3 Overdispersion: What It Is and Why It Matters
      • 21.1.4 Overdispersion in practice: elephant poaching
    • 21.2 Exercise: Maze runner
  • 22 Extending use cases of GLM
    • 22.1 Binomial response data
    • 22.2 Example: NFL field goals
      • 22.2.1 DHARMa
      • 22.2.2 Continuing the analysis
    • 22.3 Example: Sex ratio in turtles
    • 22.4 Example: Smoking
    • 22.5 Overdispersion in binomial models
    • 22.6 Exercise: Forensic footprints
  • 23 GLM families and use cases
    • 23.1 Summary Table of GLM Families
    • 23.2 Common GLM Families
    • 23.3 Quasi-family models
  • 24 Power analysis by simulation
    • 24.1 Type I and II errors and statistical power
    • 24.2 What determines statistical power?
    • 24.3 An example of calculating statistical power.
      • 24.3.1 The simulation
      • 24.3.2 Some questions for you to address:
    • 24.4 Summary
    • 24.5 Power for other models
      • 24.5.1 A binomial GLM
      • 24.5.2 A linear regression
    • 24.6 Extending the simulation (optional, advanced)
    • 24.7 Exercise 1: Snails on the move
    • 24.8 Exercise 2: Mouse lemur strength
  • IV Appendix
  • 25 Examples of statistics reporting
    • 25.1 t-test
    • 25.2 Simple linear regression model
    • 25.3 A Generalised linear model (GLM)
  • 26 An example of a past written assignment (2020)
  • 27 Leveraging ChatGPT for R Programming Assistance
    • 27.1 Introduction
    • 27.2 Overview
    • 27.3 What is a Large Language Model (LLM) and how do they work?
    • 27.4 Limitations
      • 27.4.1 Limited knowledge
      • 27.4.2 Hallucination
      • 27.4.3 Numerical ability
    • 27.5 Ethics of using LLMs in education
    • 27.6 Use cases in R
      • 27.6.1 Finding errors.
      • 27.6.2 Explaining code
      • 27.6.3 Interpreting output
      • 27.6.4 Translating code (e.g. from Python/Matlab to R)
      • 27.6.5 Solving modelling problems
      • 27.6.6 Helping with documentation/comments
      • 27.6.7 Finding alternative/better ways
    • 27.7 Some tips
  • V Solutions
  • 28 Exercise Solutions
    • 28.1 Californian bird diversity
    • 28.2 Wrangling the Amniote Life History Database
    • 28.3 Temperature effects on egg laying dates
    • 28.4 Virtual dice
    • 28.5 Sexual selection in Hercules beetles
    • 28.6 Sex differences in fine motor skills
    • 28.7 Therapy for anorexia
    • 28.8 Compare t-tests with randomisation tests
    • 28.9 Apple tree crop yield
    • 28.10 Chirping crickets
    • 28.11 Fish behaviour
    • 28.12 Maze runner
    • 28.13 Forensic footprints
    • 28.14 Snails on the move
    • 28.15 Mouse lemur strength
  • Published with bookdown

BB852 - Data handling, visualisation and statistics

Chapter 7 Additional recommended reading

Why this section exists

These readings and websites give helpful background, examples, and visual inspiration. None are required, but all are useful if you want to go deeper.

7.1 Papers and chapters

  • Broman, K. W., & Woo, K. H. (2018). Data Organization in Spreadsheets. The American Statistician, 72(1), 2–10.
  • Gotelli, N. J., & Ellison, A. M. (2013). Chapter 4, Framing and Testing Hypotheses, in A Primer of Ecological Statistics. Sinauer.
  • Petchey, O., Beckerman, A., & Childs, D. (2009). Shock and Awe by Statistical Software - Why R? Bulletin of the British Ecological Society, 40(4), 55–58.
  • Weissgerber, T. L., Milic, N. M., Winham, S. J., & Garovic, V. D. (2015). Beyond bar and line graphs: time for a new data presentation paradigm. PLoS Biology, 13(4), e1002128. doi: 10.1371/journal.pbio.1002128
  • Wickham, H. (2014). Tidy Data. Journal of Statistical Software, 59(10), 1–23.

7.2 Useful websites

  • The R Graph Gallery: https://www.r-graph-gallery.com/
  • STHDA ggplot2 essentials: http://www.sthda.com/english/wiki/ggplot2-essentials
  • STHDA R basics: http://www.sthda.com/english/wiki/r-basics-quick-and-easy

Tip: Pick one reading and one website and spend 30 minutes exploring. A little browsing goes a long way.