R语言数据科学工具链:基础统计分析与高效可视化实践 (版本 2024.08.15)
TL;DR: 本文档旨在为初学者提供一个R语言进行基础统计分析与数据可视化的实践指南。内容覆盖R环境配置、关键包安装、数据导入、描述性统计计算及使用ggplot2进行常见图表绘制。所有操作均基于命令行或代码块,确保可复现性。
R环境准备与核心包安装
版本信息: R 4.3.2, RStudio 2023.12.1
在进行任何数据分析前,需确保R环境已正确配置。此步骤主要包含R及RStudio的安装,并安装必要的数据处理与可视化包。本教程不会详细讲解R语言下载和RStudio下载的步骤,官网均有详细说明。
1. 安装R及RStudio
请从官方网站获取最新稳定版。
2. 安装并加载核心R包
tidyverse包集成了数据处理与可视化的常用工具,包括dplyr和ggplot2。
install.packages("tidyverse")
Expected Output:
Installing package into ‘/path/to/R/library’
(as ‘lib’ is unspecified)
...
* DONE (tidyverse)
library(tidyverse)
Expected Output:
── Attaching packages ─────────────────────────────────────── tidyverse 2.0.0 ──
✔ ggplot2 3.5.1 ✔ purrr 1.0.2
✔ tibble 3.2.1 ✔ dplyr 1.1.4
✔ tidyr 1.3.1 ✔ stringr 1.5.1
✔ readr 2.1.5 ✔ forcats 1.0.0
── Conflicts ────────────────────────────────────────── tidyverse_conflicts() ──
✖ dplyr::filter() masks stats::filter()
✖ dplyr::lag() masks stats::lag()
ℹ Use the conflicted package (<http://conflicted.r-lib.org/>) to force all conflicts to become errors
Note: Conflicts提示是正常现象,表示某些函数名在不同包中重复,tidyverse会优先使用自身版本。
数据导入与基础统计分析
本节将演示如何导入CSV数据并进行基本的描述性统计分析。
1. 数据准备
创建一个名为sample_data.csv的文件,内容如下:
ID,Group,Value
1,A,10
2,B,12
3,A,15
4,A,11
5,B,13
6,B,16
7,A,14
8,B,18
2. 导入数据
使用read_csv()函数导入。
data <- read_csv("sample_data.csv")
Expected Output:
Rows: 8 Columns: 3
── Column specification ────────────────────────────────────────────────────────
Delimiter: ","
chr (1): Group
dbl (2): ID, Value
ℹ Use `spec()` to retrieve the full column specification for this data.
ℹ Specify the column types or set `show_col_types = FALSE` to quiet this message.
3. 描述性统计
计算Value列的均值、中位数、标准差。并按Group分组计算。
summary(data$Value)
Expected Output:
Min. 1st Qu. Median Mean 3rd Qu. Max.
10.00 11.75 13.50 13.62 15.25 18.00
data %>%
group_by(Group) %>%
summarise(
Mean_Value = mean(Value),
Median_Value = median(Value),
SD_Value = sd(Value)
)
Expected Output:
# A tibble: 2 × 4
Group Mean_Value Median_Value SD_Value
<chr> <dbl> <dbl> <dbl>
1 A 12.5 12.5 2.38
2 B 14.7 14.5 2.60
数据可视化实战:ggplot2入门
ggplot2是R中最强大的可视化工具之一。本节展示如何绘制箱线图和直方图。
1. 箱线图 (Boxplot)
展示不同Group下Value的分布。
ggplot(data, aes(x = Group, y = Value, fill = Group)) +
geom_boxplot() +
labs(title = "Value Distribution by Group",
x = "Group",
y = "Value") +
theme_minimal()
Expected Output: (Plot will be rendered in RStudio Plots pane or saved to file)
此命令将生成一个箱线图,可视化不同分组间数值的分布情况。
2. 直方图 (Histogram)
展示Value的频率分布。
ggplot(data, aes(x = Value)) +
geom_histogram(binwidth = 2, fill = "steelblue", color = "black") +
labs(title = "Histogram of Value",
x = "Value",
y = "Frequency") +
theme_minimal()
Expected Output: (Plot will be rendered in RStudio Plots pane or saved to file)
此命令将生成一个直方图,展示Value数据的分布模式。
Warning: binwidth参数对直方图的视觉呈现影响显著,应根据数据特性进行调整。
3. 图表保存
将生成的图表保存为高分辨率文件,例如PDF。这也是R语言统计分析教程中常见的高级操作。
ggsave("boxplot_by_group.pdf", width = 6, height = 4, units = "in")
Expected Output: (File boxplot_by_group.pdf will be created in the working directory)
Saving 6 x 4 in image
ggsave("histogram_value.png", width = 800, height = 600, units = "px")
Expected Output: (File histogram_value.png will be created in the working directory)
Saving 800 x 600 px image
Note: ggsave()函数默认保存最后生成的ggplot2图形,文件格式由扩展名决定。
References
- Wickham, H., Çetinkaya-Rundel, M., & Grolemund, G. (2023). R for Data Science (2nd ed.). O'Reilly Media.
- ggplot2 Documentation: roxi.cc/ggplot2-docs