Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Prevent stale player presence by separating a player’s stable identity from each WebSocket connection, refreshing connection liveness with heartbeats, and expiring connections that stop refreshing. A close event is useful for prompt cleanup, but it is not reliable enough to be the only way a player goes offline: a network failure can leave the server unaware that the socket has disappeared.

Why a player can remain online after a WebSocket disconnect

A WebSocket close handler runs when the server learns that a connection has closed. With abrupt network loss, a sleeping device, or a broken link, the server may not receive a clean close notification. If your application removes a player only in that handler, the player can remain marked online indefinitely. Microsoft’s ASP.NET Core WebSockets guidance explains that the server is not automatically informed when a client disappears because of lost connectivity, and recommends timeout-based detection.

The underlying issue is state ownership: a socket is a temporary transport, while a player’s identity and presence are application state. A reconnect creates a new connection; it does not necessarily create a new player or session.

Separate player identity from connection identity

Assign each authenticated player or logical session a stable ID, and give every WebSocket connection its own ID. Track active connection IDs under the stable identity rather than treating one socket ID as the player’s entire presence record.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  • Player or session ID: identifies the logical participant across reconnects.
  • Connection ID: identifies one specific socket and its liveness state.
  • Presence state: represents whether the player has one or more valid active connections.

This distinction handles reconnect overlap safely. A replacement socket may join before the old socket’s delayed close is processed. Removing the old connection must not mark the player offline while the replacement remains valid.

Use heartbeats and expiry as the recovery path

Refresh a per-connection last-seen timestamp or lease when a heartbeat arrives. If a connection stops refreshing, expire it after a configured stale threshold and remove its association with the player. The expiry mechanism handles abrupt loss, server restarts, and other cases where a clean close callback never arrives.

You can use WebSocket protocol Ping/Pong or application-level heartbeat messages. Choose based on which endpoint needs to detect a failure and whether the application must observe the heartbeat itself. The Python websockets 13.0 keepalive documentation describes a configurable Ping/Pong loop: its example waits 20 seconds before sending a Ping and expects a Pong within 20 seconds. Those are library example values, not universal settings. Shorter intervals detect failures sooner but increase traffic and sensitivity to latency.

Set the heartbeat cadence and stale threshold against your offline-detection target, expected network jitter, and the shortest relevant idle timeout along the connection path. Allow slack beyond the expected heartbeat interval so a delayed packet does not immediately evict a live player. Microsoft’s guidance likewise recommends timeouts and periodic client pings for connections expected to be idle; verify applicable framework defaults for the version you deploy.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Make connect, reconnect, and cleanup race-safe

  1. Authenticate and identify: on connect, authenticate the player and resolve the stable player or session ID.
  2. Register the new socket: assign a fresh connection ID, associate it with the logical identity, and initialize its liveness timestamp.
  3. Resume application state: restore subscriptions or other state needed by the new connection, subject to your session policy.
  4. Refresh on heartbeat: update only the liveness record for the connection that sent the heartbeat.
  5. Remove only the closing socket: on a clean close, delete that connection ID, not the player identity wholesale.
  6. Expire stale connections: on lease or session expiry, remove the expired association and emit a player-leave event only if no other valid connection remains.

Make cleanup idempotent: processing the same close or expiry more than once should have the same result as processing it once. Also check ownership or connection identity before applying a delayed cleanup. In particular, a close from an older socket must not erase a newer socket’s registration or publish an offline event for a player who is still connected.

Choose storage that matches your deployment

For a single-process service, in-memory connection tracking may be sufficient, provided expiry handles missed closes and process restarts are acceptable for your presence model. In a multi-instance deployment, presence must be shared between instances or reconnects must be routed to the instance that owns the resumable state. Otherwise, one server may not see a connection registered by another.

Redis documents expiring session keys for automatic cleanup and sets for tracking multiple sessions associated with a user. See its session-store guidance. A per-connection lease or timestamped presence set can also support stale filtering; choose the structure according to how you query, refresh, and clean up presence. Redis does not by itself resolve update-ordering races: connection IDs, ownership checks, and idempotent leave-event handling remain application responsibilities.

For an implementation example, Vercel’s presence guide uses heartbeat-refreshed timestamps and stale cleanup; its local reference-count approach avoids removing an identity while another local socket remains. These are design examples rather than WebSocket protocol requirements. A practical reconnection guide also discusses separating connection identity from resumable session identity and bounding resumable state with a TTL.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Best Value
Sale
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Reconnect with randomized, bounded backoff

When many clients reconnect immediately after an outage, they can create a synchronized retry storm. RFC 6455 Section 7.2.3 recommends randomizing the first retry after abnormal closure and increasing delays after repeated failures, such as with truncated binary exponential backoff. Its example of a randomized initial delay from 0 to 5 seconds is an example, not a mandatory universal setting. The RFC does not prescribe a maximum retry count; choose a retry limit or elapsed retry window based on your product’s recovery behavior.

Apply jitter to retry delays, cap the delay, and bound how long the client keeps trying. On the server, bound resumable session state with a TTL so abandoned sessions do not persist indefinitely. The reconnect policy and session TTL should work together: a client should not keep attempting to resume state that the server has already expired.

Set thresholds by balancing detection and false departures

There is no single heartbeat interval or stale TTL that fits every game, network, and deployment. Use these trade-offs to choose and validate values:

  • Detection delay: decide how quickly a room must show a silently disconnected player as offline.
  • False-offline tolerance: account for jitter, device sleep and wake, and packet delays before expiring an otherwise live connection.
  • Connection policy: decide whether players may have multiple tabs or devices active at once, and how overlapping old and new sockets affect presence.
  • State ownership: establish whether presence is process-local, shared across instances, or attached to a resumable session store.
  • Operational cost: account for heartbeat traffic, storage updates, expiry checks, and the fan-out caused by presence changes.

Test the chosen policy under packet loss, device sleep and wake, abrupt process termination, reconnect overlap, and simultaneous sessions. Measure both how quickly stale players disappear and how often live players are incorrectly marked offline before settling production thresholds.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.