Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Johns Hopkins’ COVID-19 data archive is historical: its repositories cover data collected from January 22, 2020, through March 10, 2023. To analyze those CSVs in R, import the appropriate global or U.S. time-series file, reshape its date columns into rows, parse the dates, and aggregate at a clearly defined geographic level. Daily-report files need an extra step because their columns changed over time.

Which Johns Hopkins COVID-19 files can you download?

The Johns Hopkins Coronavirus Resource Center (CRC) began its dashboard on January 22, 2020, expanded into the CRC on March 3, 2020, and stopped collecting data after three years as reporting cadences changed. Its archived repositories cover information collected through March 10, 2023. See the CRC archive information.

For chronological time-series data, open the csse_covid_19_data repository, then csse_covid_19_time_series. Choose the file for the geography and measure you need, open its raw view, and save the CSV. The available file families include:

  • confirmed_global and deaths_global for global time series.
  • confirmed_US and deaths_US for U.S. time series.

The CRC describes this access process in its data download guide. Confirmed cases, deaths, and recovered cases are separate measures; do not assume every file has identical dimensions or fields.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

How do you import and inspect a time-series CSV in R?

Read each downloaded CSV into its own data frame, then inspect its dimensions and structure before transforming it. For example:

confirmed <- read.csv("time_series_covid19_confirmed_global.csv")
deaths <- read.csv("time_series_covid19_deaths_global.csv")
recovered <- read.csv("time_series_covid19_recovered_global.csv")

str(confirmed)
dim(confirmed)
names(confirmed)

Use the actual local filenames if they differ. Repeat the inspection for each imported file: the University of Toronto’s R tutorial notes that these datasets can have different row and column counts. Its examples use read.csv() and str() to inspect confirmed, deaths, and recovered data (tutorial).

How do you reshape the wide table into tidy rows?

In the original time-series layout, geography is stored in identifier columns and each reporting date has its own column. For grouped analysis, reshape the date columns into a date field and a value field, while retaining the geographic identifiers: Country.Region, Province.State, Lat, and Long. The tutorial demonstrates gathering those date columns, then grouping by country and date and summing the measure.

Conceptually, each resulting row represents one geographic unit on one date. Apply the same transformation separately to confirmed, deaths, and recovered values. Once each measure has been made long, full-join the three tables using their shared geographic and date keys, so rows found in only one measure are retained. Check the join keys and output dimensions rather than assuming the three source files line up exactly.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

How do you parse dates and aggregate by country or worldwide?

R may prepend X to date-column names when it imports a CSV. Remove that prefix before parsing the labels. The date format used in the JHU time-series files described by the tutorial is month.day.two-digit-year:

date_labels <- sub("^X", "", date_labels)
dates <- as.Date(date_labels, format = "%m.%d.%y")

Here, date_labels represents the imported date-column names. Ensure the resulting dates are valid before relying on them for grouping. A practical validation is to inspect the parsed range with min(dates) and max(dates).

For country-level cumulative totals, group the long data by country and date and sum the geographic rows that belong to each country. The tutorial also describes creating a cumulative confirmed count and an elapsed days variable within each country, then aggregating by date to make a world table. Keep the aggregation grain explicit: province/state rows should be summed once into country totals, and country totals can then be summed by date for a world total. Avoid mixing a country total with its constituent province/state rows in the same sum.

Why do daily JHU files have different columns?

Daily-report files can have schema drift: their fields were added or changed as governments changed reporting and as mapping requirements introduced latitude and longitude columns. The R README documenting this issue recommends adding missing columns as NA before binding the daily files into one table, then saving the cleaned result as an RDS file for later work (JHU COVID-19 data README).

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

In practice, compare column names across files, make each file conform to a common set of columns, and only then row-bind them. This preserves a consistent table shape without treating an absent field in an older file as if it were a reported zero. Keep the original CSVs as well as the cleaned RDS so the transformation can be reproduced.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

How can you check that the cleaned table is sound?

  • Check schema and size: compare dimensions and column names after each import and after each reshape or join.
  • Check dates: inspect the parsed date range and look for missing or malformed dates.
  • Check aggregation: verify that country and world totals sum the intended geographic rows exactly once.
  • Check missingness: distinguish an absent schema field, represented as NA, from an observed zero.
  • Keep raw inputs: retain downloaded files alongside the cleaned table and transformation code.

These checks matter for interpretation as well as code correctness. The workflow documentation warns that country-specific data are not accurate enough for reliable cross-country comparisons because source coverage and reporting practices differ; it also notes that confirmed case counts do not correlate with country population size. Treat the archive as a record of reported data, not as a directly comparable country ranking. See the JHU data README.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.