Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

Google File System (GFS) was a distributed file system designed for Google’s large, data-intensive applications—not a system that literally stored “a planet” of data. Its published design used a master to coordinate metadata, chunkservers to store replicated file chunks, and clients that moved data directly to those servers. GFS is historically important, but it should not be mistaken for Google’s current storage system: Google identifies Colossus as its successor.

What was Google File System?

GFS was an internal distributed file system whose design Google researchers described in a 2003 paper. The authors, Sanjay Ghemawat, Howard Gobioff, and Shun-Tak Leung, characterized it as a scalable file system for large, distributed, data-intensive applications. It was built to tolerate failures on inexpensive commodity hardware while delivering high aggregate performance to many clients. Google Research’s paper page provides the original publication.

The “planet” in the headline is a scale metaphor, not a claim about how much data GFS stored. The paper’s largest reported cluster had hundreds of terabytes spread across thousands of disks and more than a thousand machines, with hundreds of clients accessing it concurrently. Those figures describe the deployment reported in 2003, not Google’s present-day infrastructure.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

How did GFS work?

GFS separated the work of coordinating files from the work of transferring their contents. The master managed metadata, while clients exchanged file data directly with chunkservers. That division kept the master from handling every byte in common read and write operations.

#1 Best Overall
Seagate 8TB IronWolf Internal NAS Hard Drive | SATA 6 Gb/s (ST8000VNZ04)
  • IronWolf internal hard drives are the ideal solution for up to 8-bay, multi-user NAS environments craving powerhouse performance.date transfer rate:6.0 gigabits_per_second
  • Store more and work faster with a NAS-optimized hard drive providing 8TB and cache of up to 256MB
  • Purpose built for NAS enclosures, IronWolf delivers less wear and tear, little to no noise/vibration, no lags or down time, increased file-sharing performance, and much more
  • Easily monitor the health of drives using the integrated IronWolf Health Management system and enjoy long-term reliability with 1M hours MTBF
  • Three-year limited product warranty protection plan and three year Rescue Data Recovery Services included

The master: metadata and coordination

The master maintained the file namespace, the mapping from files to chunks, and information about chunk replicas. Clients contacted it for metadata and chunk locations. Chunkservers also reported their state to the master, allowing it to coordinate the system and respond to failures.

Chunkservers: local storage and replicas

Files were divided into large chunks stored by chunkservers on local disks. The original paper specified a 64 MB chunk size—a detail of the published design, not a current Google-wide storage standard. Large chunks reduced metadata overhead and the frequency of client interactions with the master. GFS kept replicas and used background repair to address failures in its commodity-machine environment. The original GFS paper (PDF) describes the chunk, replication, and repair design.

Clients: locate first, transfer directly

  1. Ask the master: A client requests the metadata and locations needed to access a file’s chunk.
  2. Contact a chunkserver: Once it has a location, the client transfers file data directly with a chunkserver rather than routing the data through the master.
  3. Rely on coordination and repair: The master tracks system metadata and replica state; replication and background repair help the system recover when machines or disks fail.

This arrangement balanced centralized coordination with distributed data transfer. It was a design for the workloads and failure conditions described in the paper, not a universal recipe for every file system.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Rank #2
Seagate 8TB BarraCuda Internal Hard Drive | SATA 6 Gb/s (ST8000DM004)
  • Store more, compute faster, and do it confidently with the proven reliability of BarraCuda internal hard drives
  • Build a power house gaming computer or desktop setup with a variety of capacities and form factors
  • The go to SATA hard drive solution for nearly every PC application from music to video to photo editing to PC gaming. Ax. Sustained transfer rate OD: 190MB/s
  • Confidently rely on internal hard drive technology backed by 20 years of innovation
  • Frustration Free Packaging - This is just an anti-static bag. No cables, no box.

What workloads and consistency model did GFS target?

The design reflected Google’s data-intensive processing needs, especially large reads and record append. Its assumptions matter: GFS should not be treated as if it were a local file system in which arbitrary concurrent writes automatically behave as users might expect.

The paper describes specialized mutation and consistency behavior suited to its target workloads. Large chunks and direct client-to-chunkserver data transfer helped support high-throughput work, while replication and repair addressed routine hardware failure. These choices involved trade-offs; a system tuned for large-scale processing is not necessarily the right model for interactive, small-file, general-purpose use.

Why did GFS reach limits?

Google’s SRE case study describes operational constraints in the historical GFS architecture, including its single master and in-memory chunk map. As production systems grew, the master’s metadata responsibilities became a scaling concern. In the case study, restarting a GFS cell required the master to retrieve the full chunk inventory, a process reported to take 10–30 minutes. That is a historical observation from the case study, not a current restart-time estimate. Google’s SRE case study on managing critical state explains the issue.

Rank #3
Seagate BarraCuda 2TB Internal Hard Drive HDD – 3.5 Inch SATA 6Gb/s 7200 RPM 256MB Cache – Frustration Free Packaging (ST2000DM008/ST2000DMZ08)
  • Migrate and clone data from old drives with ease using our free Seagate DiscWizard software tool
  • Store more, compute faster, and do it confidently with the proven reliability of BarraCuda internal hard drives
  • Build a powerhouse gaming computer or desktop setup with a variety of capacities and form factors
  • The go to SATA hard drive solution for nearly every PC application—from music to video to photo editing to PC gaming
  • Confidently rely on internal hard drive technology backed by 20 years of innovation

Is Google still using GFS, and what replaced it?

GFS is best understood as a historically important published design, not as Google’s current storage system. Google describes Colossus as GFS’s successor, developed to address scaling limits that included metadata management. Its public account describes a scalable metadata service, with file metadata stored in Bigtable, and clients exchanging data directly with file servers. Google’s Colossus overview outlines that successor architecture.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Bigtable and Colossus are not interchangeable names for the same system. Colossus is the storage system in this account; Bigtable provides a metadata role for Colossus. Separately, Google Cloud’s current Bigtable overview says Bigtable tables are stored on Colossus. That relationship does not mean Bigtable replaced GFS. Google Cloud’s Bigtable overview describes Bigtable’s storage context.

Google’s public descriptions explain the broad architecture and successor relationship, but they are not a complete current implementation specification. The historical GFS paper and later Colossus accounts should therefore be read as descriptions of their respective designs, not as a full map of Google’s present infrastructure.

Rank #4
Sale
Seagate IronWolf 4TB NAS Internal Hard Drive CMR 3.5 Inch SATA 6Gb/s 5400 RPM 64MB Cache for RAID Network Attached Storage Rescue Services (ST4000VNZ06/006)
  • IronWolf internal hard drives are the ideal solution for up to 8-bay, multi-user NAS environments craving powerhouse performance
  • Store more and work faster with a NAS-optimized hard drive providing ultra-high capacity up to 16TB and cache of up to 256MB
  • Purpose built for NAS enclosures, IronWolf delivers less wear and tear, little to no noise/vibration, no lags or down time, increased file-sharing performance, and much more
  • Easily monitor the health of drives using the integrated IronWolf Health Management system and enjoy long-term reliability with 1M hours MTBF
  • Three-year limited warranty protection plan included and three year Rescue Data Recovery Services included
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

How to place GFS in context

GFS is useful to study as an example of designing storage around a particular workload: large-scale data processing on machines expected to fail. When comparing it with another distributed storage system, the meaningful questions are how it handles metadata and control-plane scaling, the sizes and access patterns it targets, how it replicates and repairs data, what write and consistency semantics it offers, and whether it is an internal platform or a system users can deploy.

The published GFS design does not establish a quantitative, apples-to-apples comparison with HDFS or modern cloud file and object stores. Nor does it provide a current figure for data stored by GFS today. For a broader introduction to storage and distributed data systems, O’Reilly’s Designing Data-Intensive Applications, 2nd Edition covers the wider subject; it is not a GFS-specific manual.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Quick Recap

Bestseller No. 2
Seagate 8TB BarraCuda Internal Hard Drive | SATA 6 Gb/s (ST8000DM004)
Seagate 8TB BarraCuda Internal Hard Drive | SATA 6 Gb/s (ST8000DM004)
Confidently rely on internal hard drive technology backed by 20 years of innovation; Frustration Free Packaging - This is just an anti-static bag. No cables, no box.
$279.99
Bestseller No. 3

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.