Yes—you can build an Arm system-on-chip without designing a CPU core. License Cortex or Neoverse processor IP, then create differentiation around the memory system, coherent interconnect, accelerators, security, I/O, packaging and software. Cortex covers application, real-time and microcontroller roles; Neoverse targets infrastructure platforms.
What Cortex and Neoverse mean in an SoC
Cortex and Neoverse are processor-IP families, not complete chips. The licensed CPU is one component alongside interconnect, memory controllers, physical interfaces, security blocks, debug and trace, system controllers and software.
| Family | Best fit | Design emphasis |
|---|---|---|
| Cortex-A | General-purpose application processors | Application performance at an appropriate power level |
| Cortex-R | Deterministic or safety-sensitive embedded control | Real-time response and safety behavior |
| Cortex-M | Energy-efficient microcontrollers and control processors | Low-power embedded operation, system management and security control |
| Neoverse | Infrastructure, cloud, edge, networking, storage, HPC and automotive central compute | Throughput, scalable performance, coherency, virtualization, RAS and platform integration |
Choose Cortex when the SoC is primarily an embedded, real-time or application device. Choose Neoverse when the CPU subsystem must scale across infrastructure workloads or provide server-class coherency, virtualization and reliability features. A single chip can also combine families—for example, high-performance application processors with a Cortex-M security or management controller.
Which Neoverse core should you license?
| CPU | Workload and platform fit | Relevant capabilities described by Arm | What to validate in your design |
|---|---|---|---|
| Neoverse N1 | General infrastructure, cloud and edge systems where balanced efficiency and server features matter | Armv8.2-A; server-class RAS, virtualization, power management, cache stashing, profiling and coherency | Memory bandwidth, core count, cache capacity, virtual-machine density and total platform cost |
| Neoverse E1 | Throughput-oriented edge-to-core data transport and networking | SMT, AArch64 and Armv8.2-A compatibility, with scaling aimed at throughput compute | Packet or stream concurrency, latency targets, I/O balance and whether SMT benefits the workload |
| Neoverse V1 | HPC, cloud HPC and AI/ML workloads needing higher per-core performance | Arm reports a 50% IPC uplift over N1; two 256-bit SVE vector units; support for systems using DDR5 and HBM2e/3 | Vector utilization, compiler and library support, sustained memory bandwidth, thermal limits and accelerator coupling |
| Neoverse V3 | Current high-performance cloud, HPC and machine-learning platforms | Arm describes double-digit improvements over V2 and the first Neoverse support for Arm Confidential Computing Architecture | Confidential-computing requirements, software enablement, fabric topology and production availability for your program |
| Neoverse V3AE | Automotive central compute, autonomous driving, ADAS and cockpit workloads | Automotive-focused platform paired with CMN S3AE and related safety-island technology | Safety case, isolation, diagnostics, real-time domains, certification schedule and vehicle I/O |
Arm’s published N1 page also claims up to 40% better price performance for AWS Graviton2 than comparable x86 instances. That is Arm’s product-page claim, not an independent benchmark; your result will depend on software, instance configuration, utilization and commercial terms.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
#1 Best Overall
How to decide between per-core performance and throughput
Start with the bottleneck
- Latency-bound or lightly threaded code: prioritize per-core performance, cache behavior and frequency, making V-series options more natural candidates.
- Highly parallel services or data movement: prioritize core count, SMT behavior, memory bandwidth and I/O throughput; E1 or N1 may fit better.
- Vector and AI workloads: quantify the fraction of code that can use SVE or an attached accelerator rather than assuming a wider vector engine automatically improves the whole application.
- Safety or confidentiality requirements: treat RAS, safety islands, isolation and confidential-computing support as architectural requirements, not optional features added late.
Check the complete platform, not only the CPU
A fast core can be starved by an undersized cache, memory controller or coherent fabric. Before selecting an IP configuration, model working-set size, read/write ratios, accelerator traffic, DMA, PCIe or Ethernet bursts, and the number of coherent agents. Confirm that the selected interconnect and memory subsystem support the required bandwidth and topology.
What a no-custom-core Arm SoC contains
Arm’s IP catalog places CPU cores beside CoreLink interconnect, memory and system controllers, CoreSight debug and trace, physical IP, security IP and Corstone subsystems. The resulting chip may contain several processor classes:
- Application or infrastructure CPUs for operating systems and user workloads.
- Cortex-M or similar control processors for boot, power, safety monitoring and runtime security.
- Dedicated accelerators for networking, storage, graphics, signal processing or machine learning.
- Coherent interconnect and memory controllers connecting CPUs, accelerators and I/O.
- Security roots, key storage, firewalls, secure boot and debug controls.
- Physical interfaces such as PCIe, Ethernet, chiplet die-to-die links and memory PHYs.
Arm’s RD-V3-R1 reference design illustrates this composition: Neoverse Poseidon-V3 application processors connect through CMN S3, with AXI expansion for coherent PCIe, Ethernet and offload traffic, while Cortex-M55 processors handle runtime security processing. A reference design is not a turnkey production chip, but it demonstrates the architectural relationship between CPU IP and the surrounding platform.
How to differentiate without designing a CPU core
Memory hierarchy and coherency
Cache sizes, directory organization, bandwidth, latency, NUMA behavior, memory-channel count and accelerator coherency can materially change performance. These choices often matter more than a small difference in nominal CPU frequency.
Rank #3
Interconnect and chiplets
For a multi-die design, define which agents must be coherent, how traffic crosses die-to-die links, where memory is attached and how failures are contained. A scalable coherent mesh can let CPU, accelerator and I/O chiplets share data, but it also adds verification, packaging, signal-integrity and thermal constraints.
Accelerators and workload-specific engines
Video, packet processing, cryptography, storage compression and AI engines can deliver larger gains per watt than asking general-purpose cores to perform every operation. Specify queueing, DMA, memory consistency and software APIs alongside the accelerator hardware.
Security and safety architecture
Use secure boot, hardware roots of trust, isolation domains, memory protection, key management and controlled debug access as system-level features. Automotive and other safety-critical products also need independent monitoring, fault containment and a safety process compatible with the target certification.
I/O and software
Differentiate through Ethernet or PCIe topology, storage interfaces, memory technology, firmware, compilers, libraries, hypervisors, drivers and observability. A theoretically faster CPU is not useful if the operating system, toolchain or accelerator stack cannot exploit it.
Recommended Free Tools
Best Value
Licensing and integration path
Arm describes Total Access as an annual subscription that can provide access to IP products, tools and models, support, training, software and manufacture rights, including Cortex and Neoverse CPUs. The exact scope, permitted implementations and commercial terms depend on the agreement, so confirm current Arm partner terms before committing a program.
- Define the workload: record latency, throughput, concurrency, vector use, memory bandwidth, virtualization, safety and security requirements.
- Select the processor family: decide whether the system needs Cortex-A, Cortex-R, Cortex-M or a Neoverse N-, E- or V-series platform.
- Build a platform model: size caches, memory controllers, coherent agents, interconnect links, accelerators and I/O together.
- Check software readiness: verify operating-system support, compiler and library maturity, firmware, hypervisor behavior and accelerator drivers for the chosen architecture.
- Review physical and manufacturing constraints: evaluate power, thermal limits, die area, package or chiplet strategy, foundry process and manufacturing rights.
- Prototype representative traffic: use architectural models, emulation or FPGA platforms to test contention, coherency, boot, security and failure handling before RTL completion.
- Freeze the licensing scope: ensure the selected cores, interconnect, models, tools, support and production permissions cover the intended product variants.
Common selection mistakes
- Choosing a core from a benchmark headline without matching the benchmark’s software, memory and pricing assumptions.
- Comparing IPC, clock speed or vector width while ignoring cache misses, fabric contention and accelerator traffic.
- Leaving security, safety islands or debug policy until after the CPU subsystem is fixed.
- Assuming a reference design can be manufactured unchanged instead of adapting it to the product’s I/O, package and verification requirements.
- Treating “no custom core” as “no architectural work.” The surrounding platform and software still determine much of the product’s differentiation and schedule.
A practical selection rule
Use Cortex-A, Cortex-R or Cortex-M when the product’s defining requirement is application processing, deterministic control or ultra-efficient embedded operation. Use Neoverse E1 or N1 for scalable infrastructure throughput and balanced server features; use V1 or V3 when per-core, vector and HPC/ML performance dominates; and use V3AE when automotive central compute and safety integration are the governing constraints. In every case, license the CPU as part of a complete coherent SoC architecture rather than as an isolated block.

