What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Fencing tokens stop an expired lock holder from overwriting newer work only when the protected resource checks the token on every write. The coordinator issues a strictly increasing token for each successful acquisition; the resource remembers the greatest token it has accepted and rejects any lower one. Twitter’s documented architecture used ZooKeeper for coordination, but its Snowflake design also shows why Twitter did not put ZooKeeper in every ID-generation operation.

How fencing tokens prevent stale writes

A distributed lock or lease grants a client temporary authority. It cannot reach into a paused process, withdraw a request already in flight, or stop a partitioned client from resuming and sending delayed work. If the protected storage accepts writes based only on the client’s belief that its lease is still valid, an old holder can overwrite a newer holder’s data.

Fencing adds an ordered value to each acquisition and requires the protected resource to enforce that order. For example, the coordinator issues token 33 to client A and, after A’s lease expires, token 34 to client B. The storage records the highest accepted token. Once it has accepted a write with 34, it rejects a later write carrying 33. Martin Kleppmann describes this approach in “How to do distributed locking” (2016), emphasizing that the token must accompany every write to the storage service.

  1. Acquire: The coordination service grants a lock and a monotonically increasing token, such as 33.
  2. Attach: The client includes that token with every operation that could change the protected state.
  3. Enforce atomically: The resource compares the token with its stored high-water mark as part of applying the operation. It rejects a lower token; for a higher token, it records the new maximum and applies the write as one atomic operation.

The atomicity matters: checking a token separately from applying the write can leave a race in which a stale request passes the check and then changes data after a newer request. Fencing is therefore a rule at the resource boundary, not merely a property of the lock service or client library.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

What happens during a pause or partition

Suppose A receives 33, then pauses in a long garbage-collection cycle or becomes isolated long enough for its lease to expire. B acquires the lock with 34 and writes first. The resource records 34. When A resumes and its delayed write arrives with 33, the resource rejects it. A’s local timer, clock, or belief about its lease cannot override the resource’s recorded ordering.

The protection is based on the order in which writes reach and are accepted by the resource. If A’s token-33 write arrives and is accepted before the resource has seen token 34, fencing alone cannot know that B has already acquired a newer lease elsewhere. A later token-34 write can supersede it, but preventing every old write immediately upon lease expiry requires additional coordination with the resource. Fencing’s key guarantee is that, after the resource accepts a newer token, an older token cannot subsequently roll the protected state back.

What Twitter used ZooKeeper for

Twitter Engineering’s 2018 account, “ZooKeeper at Twitter,” calls Apache ZooKeeper “a system for distributed coordination.” It documents ZooKeeper as a coordination kernel for distributed locks, leader or master election, service discovery, and critical metadata. It also cautions against treating ZooKeeper as a generic strongly consistent in-memory key-value store: the guidance is to keep its data small and mostly out of the performance-critical path.

That distinction matters to fencing. ZooKeeper can help coordinate who is allowed to act and can provide ordered metadata, but the protected storage still has to reject stale operations. The lock decision and the storage write are separate parts of the design.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Snowflake: coordinate worker assignment, not every ID

Twitter’s Snowflake announcement describes a deliberate boundary: worker numbers were selected at startup through ZooKeeper, while each generated ID combined a timestamp, worker number, and sequence number. Twitter considered ZooKeeper sequential nodes for ID generation but rejected that approach because it could not meet the desired performance characteristics and the team feared that tighter coordination would reduce availability without enough benefit.

Snowflake is an example of choosing where coordination belongs, not an example of fencing a storage write. ZooKeeper helped assign worker identity at startup; the design avoided making every generated ID depend on a coordinated ZooKeeper operation.

Manhattan: elected log writers and failover

Twitter’s Manhattan storage design used per-shard logs. Coordinators mapped keys to shards and submitted operations to those logs; storage nodes applied each shard’s operations sequentially as replicated state machines. Each log had an elected writer, and ZooKeeper supplied failover when a writer failed during network partitions, hardware failures, or planned maintenance.

This illustrates coordination around a shared resource and an ordered operation stream. It should not be read as proof that ZooKeeper alone fences every Manhattan write: the published design description establishes the writer-election and failover roles, while the stale-write safety of any particular implementation depends on how the log or storage layer validates authority.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Choosing and enforcing a fencing token

A useful token is not just unique. It must be strictly increasing in the scope of the resource it protects, and that ordering must remain valid across acquisitions and failover. If two independent resources have independent token sequences, their numbers cannot automatically be compared as one global order. Define the scope explicitly: for example, one token sequence per shard or per protected record, depending on the coordination design.

  • Use an ordered source: The coordinator must produce a newer token for each successive successful acquisition in the relevant scope. Kleppmann notes that a ZooKeeper zxid or znode version can serve when the implementation provides the necessary monotonicity and scope.
  • Validate at the resource: Every mutation must carry the token, and the resource must reject tokens below its recorded maximum. A token checked only by the client or lock service does not fence storage.
  • Make comparison and mutation atomic: Store the high-water mark alongside the protected state, or otherwise ensure they are updated in one atomic transaction. Define how retries and duplicate requests behave.
  • Plan for multiple writes per lease: A holder may make several writes with the same token. The resource can allow equal tokens for that holder while rejecting lower ones; if operations under one lease also need ordering or deduplication, add a separate per-operation sequence or idempotency mechanism.
  • Observe rejection and recovery: Record the resource, token, and rejection reason in logs or metrics. A rejected stale write is a safety signal, not a reason to blindly retry using the old token. The client should reacquire authority and reconcile whether its operation remains valid.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

How fencing differs from related approaches

Approach What it establishes What it does not establish
Lease without fencing A coordinator’s time-bounded grant to a client. It cannot retract delayed work or stop an expired client from contacting a resource that does not check authority.
Lease with fencing Ordered authority plus resource-side rejection of lower tokens after a newer token has been accepted. It does not make a resource aware of a newer lease before that resource receives a write carrying the newer token.
Unique random lock value Can identify a particular lock attempt. Uniqueness is not monotonic ordering. Kleppmann’s analysis says Redlock’s random value is not a fencing token.
ZooKeeper coordination Can coordinate locks, elections, discovery, metadata, and ordered values when the chosen primitive has the required guarantees. Using ZooKeeper does not by itself make an unrelated storage system reject stale writes; the resource must enforce the token.

Failure cases to test

  • Long process pause: Pause a lock holder beyond lease expiry, let another client acquire a newer token and write, then resume the first client. Confirm the resource rejects its old-token write.
  • Delayed network request: Hold an old-token request in transit until after a newer-token write has been accepted. Verify the same rejection path applies to delayed packets, not just active clients.
  • Coordinator failover: Restart or fail over the token-issuing component and verify that the next token remains greater in the protected resource’s scope. Do not assume identifiers are monotonic merely because they come from ZooKeeper; verify the specific primitive’s guarantees.
  • Resource failover or restore: Confirm that replicas and restored backups retain a safe high-water mark. If the maximum token is lost or rolled back, an old client may appear current again.
  • Concurrent writes and retries: Exercise equal-token writes, duplicate submissions, and a higher token arriving while a lower-token transaction is in progress. The comparison and state mutation must have defined atomic behavior.

What this says about Twitter’s design

Twitter’s documented examples point to a selective role for coordination: ZooKeeper supported control-plane decisions such as worker assignment, leadership, and failover, while performance-sensitive data paths had their own designs and trade-offs. Snowflake’s startup worker assignment and Manhattan’s elected log writers illustrate different boundaries. Neither example justifies the broader claim that Twitter’s present-day X architecture is unchanged, or that ZooKeeper automatically provides storage-level fencing wherever it is used.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.