PC Slower Than It Used to Be?
A free scan shows the junk files, broken settings and background clutter dragging Windows down - then fixes them in one click.Free scan · Windows 10 & 11Crashes, No Sound, or Screen Glitches?
Random freezes, missing sound and display glitches usually trace back to one bad driver. Find and replace yours safely.Free scan · under a minuteiTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more
DeepSeek V4 combines two attention paths: Compressed Sparse Attention (CSA) and Heavily Compressed Attention (HCA). CSA compresses the key-value (KV) cache and uses DeepSeek Sparse Attention (DSA) to select relevant entries; HCA compresses more heavily and applies dense attention to the compressed representation. Manifold-Constrained Hyper-Connections (mHC) is separate: it governs how residual information flows between layers, not which tokens are attended to.
How the four mechanisms fit together
The terms describe different parts of the architecture rather than four equivalent attention methods:
- CSA and HCA are complementary attention paths in V4. Both operate on compressed representations of the KV cache, but use different compression and selection strategies.
- DSA is the sparse-selection mechanism used within CSA. It ranks entries and selects a subset for the core attention operation.
- mHC constrains mixing among residual streams across layers. It is not an attention path and does not select historical tokens.
In shorthand: DSA selects within CSA; CSA and HCA are the two attention routes; mHC shapes inter-layer information flow.
How DSA selects entries
DeepSeek’s V3.2 report describes a learned “lightning indexer” that scores preceding KV entries for each query. A top-k selector then retains a subset for the core attention operation. Here, k is the number of selected entries, while L represents sequence length.
#1 Best Overall
DeepSeek reports that this changes the core attention operation’s complexity from O(L²) to O(Lk). That reduction does not eliminate all quadratic work: the report says the lightning indexer itself still has O(L²) complexity. The distinction matters when interpreting DSA’s complexity claim.
CSA versus HCA
Both paths reduce the sequence dimension of the KV cache, but they make different tradeoffs. CSA keeps a less heavily compressed pool and uses DSA to select entries; HCA compresses more heavily and uses dense attention across its resulting representation.
Rank #2
| Path | KV-cache compression | Selection behavior | Architectural emphasis |
|---|---|---|---|
| CSA | Lower compression, with overlapping windows described in Hugging Face’s implementation documentation | A Lightning Indexer gathers top-k entries before core attention | Selective access to a less-compressed pool |
| HCA | Heavier compression | Dense attention over the compressed pool; no indexer is used in the documented implementation | Broad attention across a more-compressed representation |
The model card presents CSA and HCA together as V4’s hybrid attention design. It would be incomplete to describe either path alone as the full V4 attention architecture.
Quick wins for a faster PC:
Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Repair Windows errors before they cause bigger problemsFix Now →What mHC changes
mHC, or Manifold-Constrained Hyper-Connections, addresses residual connections rather than attention over the sequence. DeepSeek describes constraining residual mapping to the manifold of doubly stochastic matrices, called the Birkhoff polytope, with the stated aim of stabilizing signal propagation while retaining expressivity. Hugging Face’s documentation describes parallel residual streams mixed through a doubly stochastic projection.
Rank #3
That makes mHC orthogonal to DSA, CSA, and HCA: it concerns how information is mixed across layers, not how the model compresses or selects KV entries.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.What the V4 context specification does—and does not—say
DeepSeek AI’s V4 model card, published April 27, 2026, specifies a 1M context length. This is a model-card specification, not an independent benchmark or a guarantee that every deployment configuration exposes that full context.
Rank #4
The architecture descriptions and complexity figures should likewise be kept distinct from end-to-end performance results. The O(Lk) figure applies to DSA’s core attention operation, with the indexer retaining O(L²) complexity; it is not a blanket performance measurement for the whole model. A later CSA indexer implementation study reports synthetic V4-shaped indexer-step experiments and explicitly does not claim real-checkpoint, end-to-end performance.
Quick Recap
Best Value
Sources
- DeepSeek AI, V4 model card (published April 27, 2026).
- DeepSeek, V3.2 technical report.
- Hugging Face Transformers documentation for DeepSeek V4.
- StreamIndex study of memory-bounded CSA indexer processing.
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

