Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

R is the language; packages are add-ons that provide functions, documentation, sample data and sometimes compiled code. This practical list covers 11 widely used, beginner-relevant packages for importing, cleaning, visualizing, modeling and sharing analyses. “Popular” here means broadly useful and well-supported, not a verified download ranking.

Learn packages by task rather than installing everything. Install a package once, load it in each session, and use explicit namespaces when you want dependencies to be obvious:

install.packages("dplyr")
library(dplyr)
dplyr::filter(data, score > 80)

Package and documentation links below were checked against the cited official pages; package versions change, so check CRAN before installing.

Choose packages by the work you need to do

Package Main job Suggested timing Good first function
dplyr Transform tables Start here filter()
ggplot2 Charts Start here ggplot()
tidyr Reshape data Learn soon pivot_longer()
readr CSV and delimited files Start here read_csv()
readxl Excel workbooks Start here when needed read_excel()
lubridate Dates and times Learn soon ymd()
stringr Text processing Learn soon str_detect()
janitor Names and simple checks Learn soon clean_names()
data.table Fast tables and files When scale matters fread()
tidymodels Modeling workflows After wrangling basics initial_split()
shiny Interactive web apps When you need an app shinyApp()

The tidyverse is a coordinated collection, not one all-purpose package. Its core includes ggplot2, dplyr, tidyr, readr and purrr; see the tidyverse overview and package directory.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Installation, loading and package management

install.packages() downloads from repositories such as CRAN and normally runs once per computer or project environment. library() attaches an installed package for the current R session.

install.packages("ggplot2")  # usually once
library(ggplot2)             # each session

In reusable scripts, ggplot2::ggplot(data) calls a function without attaching the package. Posit explains package installation and CRAN management in its RStudio package guide.

Install only what your task requires

install.packages(c("dplyr", "ggplot2", "tidyr", "readr"))
install.packages(c("readxl", "janitor"))
install.packages(c("lubridate", "stringr"))
install.packages("data.table")
install.packages("tidymodels")
install.packages("shiny")

Installing tidyverse or tidymodels also installs dependencies. That is normal, although larger dependency trees can take longer to install and troubleshoot.

1. dplyr: readable data transformation

dplyr supplies verbs for selecting rows and columns, creating variables, grouping, summarizing, sorting and joining tables.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
library(dplyr)
summary <- starwars |>
  filter(!is.na(height)) |>
  group_by(gender) |>
  summarise(average_height = mean(height), people = n(), .groups = "drop")

mutate() preserves the row structure; summarise() generally reduces it. Grouping persists until removed or replaced. Use is.na(x), not x == NA. Joins can duplicate rows when the right-hand key is not unique, so check key relationships first. Base R indexing and merge() and the data.table syntax are alternatives. Documentation: CRAN dplyr page.

2. ggplot2: build charts in layers

ggplot2 maps variables to visual aesthetics and adds layers called geoms.

library(ggplot2)
ggplot(mtcars, aes(wt, mpg, color = factor(cyl))) +
  geom_point() +
  labs(x = "Weight", y = "Miles per gallon", color = "Cylinders") +
  theme_minimal()

ggplot() starts the plot, aes() maps data, and geoms such as geom_point(), geom_col() and geom_line() draw marks. facet_wrap() makes small multiples; theme() changes presentation. Put a constant colour outside aes() when every mark should have that colour. Label units and do not infer causation from a visual association. Base graphics, lattice and plotly are valid alternatives. See the CRAN page.

3. tidyr: put tables into a useful shape

tidyr handles wide-to-long and long-to-wide transformations and straightforward missing-value operations.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
library(tidyr)
long_data <- pivot_longer(data, starts_with("sales_"), names_to = "month", values_to = "sales")

Also learn pivot_wider(), separate_wider_delim(), separate_longer_delim(), drop_na(), replace_na() and fill(). Duplicate key combinations can make pivot_wider() produce list columns or warnings. Tidy data is a useful convention, not a rule that every reporting table must be long.

4. readr: import delimited text

readr provides consistent parsers and diagnostics for rectangular text files.

library(readr)
sales <- read_csv("sales.csv")
# Other common choices: read_csv2(), read_tsv(), read_delim()

Check the delimiter, decimal mark, encoding and inferred column types. Date columns may arrive as character values, and a relative path fails if R is running from another directory. Prefer project-relative paths. data.table::fread() and base R are alternatives; CRAN lists current requirements.

5. readxl: read Excel workbooks

readxl reads both .xls and .xlsx files without requiring Excel.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
library(readxl)
excel_sheets("report.xlsx")
data <- read_excel("report.xlsx", sheet = "January")

Spreadsheets often contain title rows, merged cells, subtotals or several tables on one sheet. Mixed numbers and text can confuse type inference. readxl reads workbook data; it is not a replacement for Excel’s formatting or formula workflow. openxlsx is an alternative when you also need broader workbook writing and manipulation.

6. lubridate: parse and calculate dates

lubridate makes common date and date-time parsing and arithmetic more readable.

library(lubridate)
dates <- ymd(c("2026-01-15", "2026-02-20"))
dates + months(1)

Use ymd(), mdy() or dmy(); extract components with year(), month() and day(); and round with floor_date(). Ambiguous dates, daylight-saving changes, locales and time zones require explicit decisions. A month is not a fixed number of days.

7. stringr: consistent text operations

stringr uses discoverable str_ functions for searching, extracting, replacing, splitting and trimming text.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
library(stringr)
emails <- c("alice@example.com", "not-an-email")
str_detect(emails, fixed("@"))

Useful functions include str_length(), str_sub(), str_replace_all(), str_extract(), str_split() and str_trim(). Regular expressions are powerful but easy to misread; use fixed() for literal matching. Account for case, missing values and Unicode text. Base R and stringi are alternatives.

8. janitor: make imported tables easier to work with

janitor is a small utility package for cleaning names and producing simple frequency tables.

library(janitor)
library(readr)
data <- read_csv("messy_export.csv") |> clean_names()
tabyl(data, region)

clean_names() turns spaces, punctuation and inconsistent capitalization into code-friendly names. adorn_totals() can add totals to suitable tables. Name cleaning does not validate units, duplicates, provenance or whether a variable means what its label suggests.

9. data.table: an alternative for speed and large tables

data.table combines compact syntax for subsetting, grouping, joins and updates with performance-oriented implementations.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
library(data.table)
sales <- fread("sales.csv")
sales[, .(average_sales = mean(amount, na.rm = TRUE), records = .N), by = region]

The DT[i, j, by] model takes practice and may be less immediately familiar than dplyr. Choose dplyr for readable verbs and a broad beginner ecosystem; choose data.table when large files, performance or team conventions matter. Learn both later if your work warrants it.

10. tidymodels: structured modeling and machine learning

tidymodels is a collection rather than a single modeling algorithm. Its components include recipes for preprocessing, parsnip for model specifications, rsample for resampling, yardstick for metrics, tune for tuning and workflows for combining steps.

install.packages("tidymodels")
library(tidymodels)
split <- initial_split(data, prop = 0.8)
train <- training(split)
test <- testing(split)

Learn it after data frames, missing values, predictors, outcomes and train/test evaluation. Avoid leakage by fitting preprocessing only on training data, and choose metrics that match the outcome. caret, mlr3 and direct model packages remain legitimate alternatives.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

11. shiny: turn analysis into an interactive app

shiny lets R users build interactive web applications. The user interface defines controls, the server function defines reactive behavior, and outputs render results.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
library(shiny)
ui <- fluidPage(
  sliderInput("n", "Number of points", 10, 100, 50),
  plotOutput("plot")
)
server <- function(input, output, session) {
  output$plot <- renderPlot(plot(runif(input$n)))
}
shinyApp(ui, server)

Local success does not remove deployment, authentication, privacy, maintenance or performance concerns. Large or expensive computations may need caching, preprocessing or a database. Quarto, flexdashboard and JavaScript frameworks suit other publishing needs.

A complete beginner workflow

This example connects import, cleaning, transformation and visualization:

library(readr)
library(janitor)
library(dplyr)
library(tidyr)
library(ggplot2)

summary <- read_csv("sales.csv") |>
  clean_names() |>
  drop_na(region, amount) |>
  group_by(region) |>
  summarise(total_sales = sum(amount), .groups = "drop")

ggplot(summary, aes(region, total_sales)) + geom_col()

Inspect missingness, duplicates, units, key uniqueness, impossible values and provenance before trusting the result. No package can decide whether a dataset is conceptually correct.

Which learning path fits your goal?

  • Data analysis: dplyr, ggplot2, tidyr, readr, then readxl.
  • Statistics and prediction: add tidymodels after the wrangling and visualization fundamentals.
  • Automation and apps: add stringr, purrr and shiny as tasks demand.
  • Large files: learn data.table; consider arrow later for columnar data.

Troubleshooting common package problems

  • “There is no package called …”: run install.packages("package"), then restart R and call library(package).
  • Function not found: load the package or use package::function(); check spelling and capitalization.
  • Name conflicts: qualify the intended function with its package namespace.
  • File not found: inspect the project directory and use a correct relative path.
  • Parsing warnings: review the importer’s problems and specify column types or locale deliberately.
  • Version or system errors: check R.version.string, sessionInfo() and .libPaths(), then consult the package’s CRAN installation notes. Updating may help: update.packages(ask = FALSE, checkBuilt = TRUE).

Keep projects reproducible

Use one R project per analysis and load only the packages the script needs. For a recorded project library, renv provides:

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
install.packages("renv")
renv::init()
renv::snapshot()
renv::restore()

Learn base R fundamentals—vectors, data frames, indexing, functions, conditions, loops and NA handling—alongside packages. Packages extend R; they do not replace the language.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.