Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Yes. GDELT data has been made available through Google BigQuery’s Public Dataset Program, which covers dataset storage costs. Query processing is a separate matter: Google’s public-dataset documentation says the first 1 TB of query data processed per month is free, subject to its query pricing details. Processing beyond applicable free use may cost money.

What “free access” to GDELT in BigQuery means

GDELT is an open global-news and events dataset. In 2015, the GDELT Project announced BigQuery access to its Event, Mentions, and Global Knowledge Graph (GKG) tables. Google’s Public Dataset Program makes participating datasets publicly accessible and covers their storage costs; users pay for the queries they run. Google documents the first 1 TB of query data processed per month as free, subject to query pricing details. See Google Cloud’s public datasets documentation for the current terms.

The distinction is between data access and computation. You do not pay to store the public dataset, but a query can process a large amount of data, and usage beyond the applicable free allowance can lead to charges. Google says a project must be created or selected to use BigQuery; billing must be enabled if you intend to go beyond free usage.

How to find and query GDELT

Use the BigQuery console

  1. Create or select a Google Cloud project. In the Google Cloud console, open BigQuery and choose the project you want to use. Review its billing status before running queries that could exceed free usage.
  2. Find the GDELT dataset and table. Browse the Explorer pane or search for GDELT. Table names and availability can change, so use the current Explorer listing rather than relying on a name from an old tutorial.
  3. Inspect the table before writing SQL. Open its schema and preview sample rows. Note the fields, location, partitioning, and last-modified information shown for the table.
  4. Write a bounded query. Select only the columns you need and add a date range or other selective conditions where appropriate. Avoid an unrestricted scan unless you have a reason to process the whole table.
  5. Check the estimate before running. BigQuery displays an estimate of bytes to be processed. Review it before execution; Google’s cost guidance recommends estimating query data and controlling usage with quotas where suitable.

Google also supports BigQuery through the bq command-line tool, the REST API, and client libraries. Whichever route you use, query processing must use a location compatible with the dataset’s location; check the dataset details and the location selected for the job. Google explains the location requirement in its public-dataset documentation.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

How to keep GDELT queries within free use

  • Read the estimate first. On-demand compute is based on the amount of data processed, so a large scan can consume the monthly free allowance quickly.
  • Request fewer columns. Name the fields you need instead of using SELECT *; this can reduce the amount read, depending on the table and query.
  • Limit the date range. Add a selective date condition when it fits your analysis.
  • Use a partition filter only when supported. If the specific table is partitioned, filtering on its partition column can reduce the data scanned. Confirm the current table’s partitioning and column in BigQuery before relying on that optimization.
  • Set a quota if you need a guardrail. Google recommends custom daily query quotas to limit usage. A quota can help control query volume, but it does not replace checking estimates or understanding the project’s billing setup.

Partitioning can make a substantial difference, but historical figures are not current performance promises. In an August 2016 post, the GDELT Project said its GKG table then contained 353 million records and totaled 3.6 TB. For one example query in that post, an unpartitioned scan processed 423 GB, compared with 15 GB for the date-partitioned version using a partition filter. Those values describe that table and those queries at that time, not today’s table sizes or expected savings. See GDELT’s 2016 partitioning announcement.

Try the BigQuery sandbox without a billing account

Google’s BigQuery sandbox lets you explore public datasets without attaching a billing account. Google documents a monthly free compute limit of 1 TiB of processed query data, a lifetime storage quota of 10 GiB, and a 60-day default expiration for sandbox datasets, tables, views, and partitions. The sandbox has limited features, so it may not suit a longer-lived or larger workflow. Check Google’s sandbox documentation for its current limits.

Google’s public-dataset page states its free query amount as 1 TB per month, while the sandbox page expresses its limit as 1 TiB per month. Keep those page-specific units distinct when assessing your usage.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Verify the table before relying on an old tutorial

GDELT’s 2015 launch announcement said its Event, Mentions, and GKG tables were updated every 15 minutes at that time. That launch-era description does not establish the current refresh rate for every GDELT table. Before building an analysis, confirm the selected table’s current availability, schema, location, partitioning, and last-modified information in BigQuery. The announcement remains useful for historical context: GDELT 2.0: Our Global World in Realtime.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.