Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Intel’s Skylake-SP generation of Xeon Scalable processors replaced the earlier on-die ring interconnect with a two-dimensional mesh to better accommodate more cores and greater memory and I/O bandwidth. The mesh connects resources inside a processor die; Intel UPI is a separate coherent link between processor sockets. Intel describes the mesh as a scalability design, not as a guarantee that every access or workload is faster.

What Intel means by mesh architecture

Skylake-SP was the codename for the Intel Xeon Scalable family discussed in Intel’s technical overview. Its mesh is an on-die network that connects processor cores and other uncore resources, including last-level cache (LLC) slices, memory controllers and I/O. Instead of sending traffic around a ring, the design provides communication paths in vertical and horizontal directions.

Intel describes a route as moving vertically to the appropriate row and then horizontally to the destination column. That gives traffic a path through the mesh toward its destination; it does not mean every transaction follows an identical route or takes a fixed number of hops. The overview’s description of a shortest path is an explanation of the topology, not a workload-level latency guarantee. Intel Xeon Processor Scalable Family Technical Overview (updated December 1, 2022).

Why Intel moved from rings to a mesh

Intel says earlier Xeon generations, including Haswell- and Broadwell-era processors, used ring architecture to connect cores, LLC, memory controllers, I/O and QPI ports. As core counts grew, Intel identified increasing access latency and declining bandwidth per core as scaling concerns. Splitting the design into two rings partly addressed the problem, but the later Xeon Scalable family added cores as well as memory and I/O bandwidth, increasing pressure on the on-die interconnect.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
#1 Best Overall
Intel XEON 22 CORE Processor E5-2699V4 2.2GHZ 55MB Smart Cache 9.6 GT/S QPI TDP 145W
  • Intel Xeon E5-2699 V4 Docosa-core (22 Core) 2.20 Ghz Processor - Socket Lga 2011-v3 - 5.50 Mb - 55 Mb Cache - 64-bit Processing - 14 Nm - 145 W

The mesh was Intel’s architectural response to that scaling problem: organize communication paths and supporting functions across the die rather than rely on a ring whose contention and bandwidth demands could become more difficult to manage as resources increased. Intel’s 2022 overview describes the Purley-platform Xeon Scalable family as supporting up to 28 cores; that is a family/platform maximum, not a claim that every Xeon Scalable processor has 28 cores.

What the CHA does in the mesh

Each core and LLC slice is associated with a combined Caching and Home Agent (CHA). The CHA helps determine where an address should be handled: an LLC bank, a memory controller or an I/O subsystem. It also provides routing information for traffic through the mesh. Distributing these caching, home-agent and I/O-related functions is intended to avoid concentrating work at one central point.

This role matters because a mesh is more than a grid of wires. Requests still need to find the right cache or memory destination, and coherency needs to be maintained when data may be cached by more than one core. The CHA participates in that organization; it does not make every access local or eliminate contention in every workload.

Ring and mesh: what changes

Aspect Earlier Xeon ring design Skylake-SP Xeon Scalable mesh
Topology Resources communicate over ring paths; Intel says some prior designs used two rings to mitigate scaling pressure. Vertical and horizontal paths provide routes across rows and columns; routing depends on source and destination.
Scaling concern Intel says growing core counts increased access latency and reduced bandwidth per core. Intel designed the distributed mesh to scale with more cores and higher memory and I/O bandwidth; this is a design rationale, not a universal speedup result.
Cache and routing functions Intel’s overview describes the earlier ring connecting cores, LLC, memory controllers, I/O and QPI ports. A CHA at each core/LLC slice supplies address-mapping and routing functions across the mesh.
Cache hierarchy in Intel’s comparison 256 KB mid-level cache (MLC) and 2.5 MB LLC per core, for the previous generation described by Intel. 1 MB MLC and 1.375 MB LLC per core, for Xeon Scalable as described by Intel.
Scope These are on-die interconnect and cache-design comparisons. Neither topology should be confused with the socket-to-socket UPI link.

The cache figures are Intel’s family-level comparison in its 2022 overview, not values to apply indiscriminately to every Xeon SKU or configuration. Intel says the larger MLC is intended to improve hit rate and reduce demand on the LLC and mesh. Because the LLC is non-inclusive, data missing from the LLC may still be present in a private cache; a snoop filter tracks such cache lines.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Rank #3
Intel Xeon E5-2690 V4 SR2N2 14-Core 2.6GHz 35MB LGA 2011-3 Processor (Renewed)
  • Total Cores 14
  • Total Threads 28
  • Processor Base Frequency 2.60 GHz
  • Max Turbo Frequency 3.50 GHz
  • Sockets Supported LGA2011-3

Mesh versus UPI: on-die and between sockets

The mesh carries traffic among resources within a processor die. Intel Ultra Path Interconnect (UPI), which replaced QPI for this Xeon Scalable family, is the coherent connection between processor sockets. A multisocket system can therefore involve both: local on-die traffic uses the mesh, while coherent traffic between processors uses UPI.

Intel’s 2022 overview says Xeon Scalable processors support two or three UPI links, depending on the processor, and gives a maximum link operating speed of 10.4 GT/s. The number of links is processor-dependent, and 10.4 GT/s is the stated maximum rather than a speed to assume for every link or system. For platform context, Intel’s approximately 2017 Xeon Scalable brief lists maxima of six memory channels, 48 PCIe 3.0 lanes and up to 28 cores; these are platform-level figures, not per-processor guarantees. Intel Xeon Scalable Platform Product Brief.

Rank #4
Sale
Intel Xeon E5-2699v4 2.2/55/2400 22C 145 (E5-2699v4) (Renewed)
  • Manufacturer: Intel CPU Frequency: 2.20 GHz CPU Max Turbo Frequency: 3.60 GHz Number of Cores: 22 Threads: 44 Cache: 55 MB Intel Smart Cache Number of UPI Links: 0 Lithography: 14 nm Thermal Design Power: 145 W Memory Types: DDR4 1600/1866/2133/2400 Max Memory Size: 1.5 TB Max # Memory Channels: 4 Sockets Supported: FCLGA2011-3 E5-2699v4
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Does mesh make Xeon faster than ring?

Intel’s architecture overview explains why the company adopted a mesh and what scaling benefits it expected. It does not provide an isolated, universal benchmark showing a fixed mesh-versus-ring speedup. It would therefore be too broad to claim that every mesh route is faster than every ring route, or that every application improves by a particular percentage.

Observed cache and uncore behavior also depends on operating conditions. In a 2019 study of Skylake-SP energy-efficiency features, Schöne, Ilsche, Bielert, Gocht and Hackenberg measured LLC access times of 119 cycles at a 1.4 GHz uncore frequency and 83 cycles at 2.4 GHz in their test setup. Those are configuration-specific measurements, not universal processor specifications and not a ring-versus-mesh comparison. The authors also measured about 9.8 ms of additional delay from the default uncore-frequency control loop before it adapted to a changed workload pattern; this concerns uncore frequency control, not mesh traversal latency.

Free tools Windows power users keep installed

One-click scans. No signup required.