首页文献管理数据分析开源社区写作排版
首页 › 数据分析 › R语言数据科学工具链:基础统计分析与

R语言数据科学工具链:基础统计分析与高效可视化实践 (版本 2024.08.15)

Roxi
Roxi 加速器 — 稳定·快速·安全
全球节点覆盖,支持所有主流平台,一键连接无需配置。新用户免费试用。
立即体验 →

TL;DR: 本文档旨在为初学者提供一个R语言进行基础统计分析与数据可视化的实践指南。内容覆盖R环境配置、关键包安装、数据导入、描述性统计计算及使用ggplot2进行常见图表绘制。所有操作均基于命令行或代码块,确保可复现性。

R环境准备与核心包安装

中国45美国30日本12韩国8其他5

版本信息: R 4.3.2, RStudio 2023.12.1

在进行任何数据分析前,需确保R环境已正确配置。此步骤主要包含R及RStudio的安装,并安装必要的数据处理与可视化包。本教程不会详细讲解R语言下载和RStudio下载的步骤,官网均有详细说明。

1. 安装R及RStudio

请从官方网站获取最新稳定版。

2. 安装并加载核心R包

tidyverse包集成了数据处理与可视化的常用工具,包括dplyr和ggplot2。

install.packages("tidyverse")

Expected Output:

Installing package into ‘/path/to/R/library’
(as ‘lib’ is unspecified)
...
* DONE (tidyverse)
library(tidyverse)

Expected Output:

── Attaching packages ─────────────────────────────────────── tidyverse 2.0.0 ──
✔ ggplot2 3.5.1      ✔ purrr   1.0.2 
✔ tibble  3.2.1      ✔ dplyr   1.1.4 
✔ tidyr   1.3.1      ✔ stringr 1.5.1 
✔ readr   2.1.5      ✔ forcats 1.0.0 
── Conflicts ────────────────────────────────────────── tidyverse_conflicts() ──
✖ dplyr::filter() masks stats::filter()
✖ dplyr::lag()    masks stats::lag()
ℹ Use the conflicted package (<http://conflicted.r-lib.org/>) to force all conflicts to become errors

Note: Conflicts提示是正常现象,表示某些函数名在不同包中重复,tidyverse会优先使用自身版本。

数据导入与基础统计分析

品牌定位清晰视觉体系统一内容矩阵搭建社媒运营规划效果追踪复盘

本节将演示如何导入CSV数据并进行基本的描述性统计分析。

1. 数据准备

创建一个名为sample_data.csv的文件,内容如下:

ID,Group,Value
1,A,10
2,B,12
3,A,15
4,A,11
5,B,13
6,B,16
7,A,14
8,B,18

2. 导入数据

使用read_csv()函数导入。

data <- read_csv("sample_data.csv")

Expected Output:

Rows: 8 Columns: 3
── Column specification ────────────────────────────────────────────────────────
Delimiter: ","
chr (1): Group
dbl (2): ID, Value

ℹ Use `spec()` to retrieve the full column specification for this data.
ℹ Specify the column types or set `show_col_types = FALSE` to quiet this message.

3. 描述性统计

计算Value列的均值、中位数、标准差。并按Group分组计算。

summary(data$Value)

Expected Output:

Min. 1st Qu.  Median    Mean 3rd Qu.    Max. 
10.00   11.75   13.50   13.62   15.25   18.00 
data %>%
  group_by(Group) %>%
  summarise(
    Mean_Value = mean(Value),
    Median_Value = median(Value),
    SD_Value = sd(Value)
  )

Expected Output:

# A tibble: 2 × 4
  Group Mean_Value Median_Value SD_Value
  <chr>      <dbl>        <dbl>    <dbl>
1 A           12.5         12.5     2.38
2 B           14.7         14.5     2.60

数据可视化实战:ggplot2入门

第1周环境搭建第2周核心开发第3周测试优化第4周正式发布

ggplot2是R中最强大的可视化工具之一。本节展示如何绘制箱线图和直方图。

1. 箱线图 (Boxplot)

展示不同Group下Value的分布。

ggplot(data, aes(x = Group, y = Value, fill = Group)) +
  geom_boxplot() +
  labs(title = "Value Distribution by Group",
       x = "Group",
       y = "Value") +
  theme_minimal()

Expected Output: (Plot will be rendered in RStudio Plots pane or saved to file)

此命令将生成一个箱线图,可视化不同分组间数值的分布情况。

2. 直方图 (Histogram)

展示Value的频率分布。

ggplot(data, aes(x = Value)) +
  geom_histogram(binwidth = 2, fill = "steelblue", color = "black") +
  labs(title = "Histogram of Value",
       x = "Value",
       y = "Frequency") +
  theme_minimal()

Expected Output: (Plot will be rendered in RStudio Plots pane or saved to file)

此命令将生成一个直方图,展示Value数据的分布模式。

Warning: binwidth参数对直方图的视觉呈现影响显著,应根据数据特性进行调整。

3. 图表保存

将生成的图表保存为高分辨率文件,例如PDF。这也是R语言统计分析教程中常见的高级操作。

ggsave("boxplot_by_group.pdf", width = 6, height = 4, units = "in")

Expected Output: (File boxplot_by_group.pdf will be created in the working directory)

Saving 6 x 4 in image
ggsave("histogram_value.png", width = 800, height = 600, units = "px")

Expected Output: (File histogram_value.png will be created in the working directory)

Saving 800 x 600 px image

Note: ggsave()函数默认保存最后生成的ggplot2图形,文件格式由扩展名决定。

References

上一篇学术会议投稿与Rebuttal策略:SRE视角的系统化流程与风险规避 (Vers 下一篇提升协作效率:GitHub PR流程的标准化实践 (版本 2024.07.31)

猜你喜欢

延伸阅读