Massive Data Storage
Core principles and architectures for handling petabyte-scale datasets.
Key Architectures
Google File System (GFS)
Foundational design for distributed storage, emphasizing fault tolerance and scalability over consistency.
- Origin: Designed for Google’s internal services, notably YouTube and search indexing.
- Core Philosophy: Optimized for large files and sequential writes; assumes hardware failures are common.
- Key Mechanism: Master server coordinates metadata; chunk servers handle actual data storage.
- Relevance: Precursor to Hadoop Distributed File System (HDFS) and modern cloud storage solutions.
- Detailed Analysis: See Google File System: Scalable, Fault-Tolerant Distributed Storage for Massive Data for a deep dive into its replication and fault-handling strategies.
Related Concepts
- distributed-file-system
- Data Replication
- Sharding
- MapReduce