Recommended Free Tools
Hulu chose Cassandra for real-time watch history and cross-device session continuity after evaluating HBase and Riak. In the company’s 2014 account, Cassandra best matched its combination of heavy write traffic, range queries, replication, reliability and a small team’s need for lower operational overhead. Hadoop remained in use for long-term storage, so the decision was about workload fit—not replacing every data system with Cassandra.
This article describes Hulu’s reported evaluation in 2014, not Hulu’s current architecture or a modern benchmark of the three databases.
The workload Hulu needed to support
Hulu was rewriting part of its service because the previous system could not scale writes as its user base grew. A viewer might start a program on one device, stop, and resume on another. The service therefore needed to save session state and make it available in real time when someone watched a video or received a recommendation.
According to Jason Verge’s July 31, 2014 report for Data Center Knowledge, the Cassandra-backed system stored subscriber watch history and later supported social data, messaging and using a phone as a remote for a connected device.
Free tools Windows power users keep installed
One-click scans. No signup required.
#1 Best Overall
How the three options compared for Hulu
| Criterion | Cassandra | HBase | Riak |
|---|---|---|---|
| Write-heavy, real-time workload | Reportedly handled Hulu’s load | Could fit teams with an existing Hadoop environment, but required substantial operational attention in Hulu’s experience | Could scale, but its performance was reported as less suitable for Hulu’s needs |
| Range queries | Reportedly available without the limitations Hulu encountered elsewhere | Not stated in the report | Hulu said range queries were not supported at the time |
| Operations | Considered easier for Hulu’s small team to maintain | More complex to set up and maintain; Hadoop and HDFS added work | Team lacked Erlang experience, making adoption less straightforward |
| Failure and replication concerns | Hulu rated replication and reliability favorably | Team worried about an HDFS NameNode single point of failure and had observed cascading failures affecting region servers | Not stated beyond the reported fit and performance concerns |
| Best contextual fit | Hulu’s real-time watch-history and session workload | An organization that already operated a Hadoop cluster | Not established for Hulu’s specific requirements |
These are descriptions of Hulu’s evaluation at that time. The report provides no standardized, independently measured head-to-head test.
Why Riak lost the comparison
Hulu’s team did not reject Riak because it was incapable of scaling. Andres Rangel, then a senior software development lead, said Riak could scale but was less performant for Hulu’s workload. Two practical issues mattered:
- Range queries: Hulu needed them, and Rangel said Riak did not support them at the time.
- Team expertise: The team did not have Erlang experience, increasing the learning and support burden.
Rangel characterized Riak as a poor fit for the service’s real-time requirements and for what his team could operate comfortably. That is a 2014, team-specific assessment—not a claim about current Riak releases.
Why HBase did not win
HBase was initially the front-runner, but its surrounding Hadoop infrastructure created concerns for Hulu.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Setup and maintenance effort
Rangel said setting up Hadoop instances took considerable work and that HBase was more complex to set up and maintain than Cassandra in Hulu’s experience.
HDFS and failure concerns
HBase’s reliance on HDFS led the team to worry about a NameNode single point of failure. Rangel also said Hulu had seen cascading failures take down region servers. The report notes that Hulu tested a newer HBase version with high-availability improvements, so the concern should be read as an operational experience during that evaluation, not a universal limitation of HBase.
Rank #3
When HBase could still make sense
Rangel’s qualification was that an organization with an existing Hadoop cluster might reasonably choose HBase. Hulu did not have a need to accept the additional operational complexity merely to obtain real-time access to this workload.
Why Cassandra won
Rangel summarized Cassandra’s appeal as its ability to handle the load, reliability, range-query support and ease of maintenance. He also said Cassandra performed better for replication in Hulu’s evaluation. The team changed hardware because Cassandra had different specifications, and the report describes the deployment as optimized for SSDs.
The Tool Desk
Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →The choice was therefore a fit across several dimensions rather than a single throughput number:
- High-volume writes for watch history and session updates
- Low-latency access while a viewer watched or received recommendations
- Range queries needed by the application
- Replication across geographically separate data centers
- A maintenance model suitable for a relatively small engineering team
Rangel said, “With Cassandra, it managed to handle the load, it’s very reliable, it allows range queries without limitations, and it’s easy to maintain.” He also described the production experience as better than expected and free of bad experiences at the time of the interview.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.What Hulu’s Cassandra deployment looked like
The report describes a primary Cassandra cluster of 32 nodes across two data centers, one on the US East Coast and one on the West Coast. The watch-history keyspace contained several billion CQL3 rows and approximately 1 TB of unreplicated data per data center. Those figures are measurements reported for the 2014 system, not current Hulu capacity.
Hulu continued using Hadoop for long-term storage while Cassandra served real-time access. This split illustrates the architectural decision: Cassandra handled interactive application reads and writes, while Hadoop retained a role suited to longer-term data storage and processing.
Best Value
- Used Book in Good Condition
What the 2014 scale figures mean—and do not mean
Data Center Knowledge reported more than 6 million paid subscribers by April 2014 and access on about 400 million internet-connected devices. Both numbers describe Hulu’s reach at the time of publication in July 2014. They should not be used as current subscriber, device or infrastructure figures.
The practical lesson for database selection
Hulu’s example supports evaluating a database against the application’s access patterns and the team’s operating reality:
- Define whether the workload is dominated by writes, reads, range scans or another access pattern.
- Separate real-time serving needs from long-term retention and analytics.
- Include replication and failure behavior in the design review, not only nominal scalability.
- Account for the skills required by the implementation language and surrounding platform.
- Measure setup and maintenance effort for the team that will actually run the system.
- Check whether an existing platform, such as Hadoop, changes the operational trade-off.
On those criteria, Hulu reported Cassandra as the least burdensome option for its real-time watch-history service in 2014. The conclusion is useful as a case study in workload-specific engineering, but it is not evidence that Cassandra is universally better than HBase or Riak.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.
Do these 3 things before closing this tab:
1Fix the driver behind crashes, sound loss and screen glitches2Repair Windows errors before they cause bigger problems3Scan for outdated or missing drivers - takes under a minute

