Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

A scalable search system is designed around the workload it must serve: how quickly documents arrive, how much data it stores, which queries users run, and how much delay or downtime the service can tolerate. Nodes supply capacity, shards divide an index into parallel work, and replicas provide redundancy and additional read capacity. The difficult part is choosing their number and placement without making every query needlessly expensive.

Start with representative data and query benchmarks, then set clear limits for indexing, query latency, recovery, and growth. The same principles apply whether you use Elasticsearch, SolrCloud, OpenSearch, or a managed service, but their coordination and scaling models differ.

How do nodes, shards, and replicas scale search?

These three components solve different problems. Elastic’s documentation describes adding nodes to increase cluster capacity, with Elasticsearch distributing data and query load across available nodes. Within an index, shards divide the data into partitions; replicas are copies of those partitions.

  • Nodes are the servers or instances that provide compute, memory, and storage. Adding nodes can increase the resources available to a cluster, though actual gains depend on placement, workload, and whether the cluster can rebalance effectively.
  • Primary shards partition an index’s data. A search over that index may need to reach multiple shards, so adding partitions can enable parallel work but also increases coordination overhead.
  • Replica shards copy primary shards. They can preserve availability when a node fails and provide additional copies on which searches may run.

In Elasticsearch, the primary-shard count is set when an index is created; changing that count later requires an index-level operation such as reindexing into a newly configured index. The replica count, by contrast, can be changed without interrupting indexing or search operations. This makes replica adjustment a more flexible way to respond to read demand, but it does not replace planning the original partitioning.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
#1 Best Overall
Sale
UGREEN NAS DH2300 2-Bay for Beginners & Personal Users, Phone Backup
  • Entry-level NAS Personal Storage:UGREEN NAS DH2300 is your first and best NAS made easy. It is designed for beginners who want a simple, private way to store videos, photos and personal files, which is intuitive for users moving from cloud storage or external drives and move away from scattered date across devices. This entry-level NAS 2-bay perfect for personal entertainment, photo storage, and easy data backup (doesn't support Docker or virtual machines).
  • Set Your Devices Free, Expand Your Digital World: This unified storage hub supports massive capacity up to 64TB.*Storage drives not included. Stop Deleting, Start Storing. You can store 22 million 3MB images, or 2 million 30MB songs, or 43K 1.5GB movies or 67 million 1MB documents! UGREEN NAS is a better way to free up storage across all your devices such as phones, computers, tablets and also does automatic backups across devices regardless of the operating system—Window, iOS, Android or macOS.
  • The Smarter Long-term Way to Store: Unlike cloud storage with recurring monthly fees, a UGREEN NAS enclosure requires only a one-time purchase for long-term use. For example, you only need to pay $459.98 for a NAS, while for cloud storage, you need to pay $719.88 per year, $2,159.64 for 3 years, $3,599.40 for 5 years. You will save $6,738.82 over 10 years with UGREEN NAS! *NAS cost based on DH2300 + 12TB HDD; cloud cost based on 12TB plan (e.g. $59.99/month).
  • Blazing Speed, Minimal Power: Equipped with a high-performance processor, 1GbE port, and 4GB RAM on Board, this NAS handles multiple tasks with ease. File transfers reach up to 125MB/s—a 1GB file takes only 8 seconds. Don't let slow clouds hold you back; they often need over 100 seconds for the same task. The difference is clear.
  • Let AI Better Organize Your Memories: UGREEN NAS uses AI to tag faces, locations, texts, and objects—so you can effortlessly find any photo by searching for who or what's in it in seconds. It also automatically finds and deletes similar or duplicate photo, backs up live photos and allows you to share them with your friends or family with just one tap. Everything stays effortlessly organized, powered by intelligent tagging and recognition.

How many shards should you use?

There is no reliable universal shard count or shard size. It depends on document size and distribution, indexing rate, query mix, hardware, retention period, and the amount of parallelism the workload actually needs. Elastic recommends benchmarking production data on production hardware with production-like indexing loads and queries.

A shard is not free. Elastic documents that each shard runs a search on a single CPU thread. A request that fans out across many shards therefore consumes more search work and coordination, and too many concurrent shard searches can exhaust search thread pools and reduce throughput. More shards are not automatically faster.

Benchmark the workload before fixing the design

  1. Build representative test data. Match the expected document shapes, size distribution, mappings, update patterns, and total data volume as closely as possible.
  2. Replay realistic traffic. Include the common queries, filters, aggregations, tenant or time constraints, and concurrent indexing that production will generate.
  3. Compare candidate shard layouts. Measure query latency percentiles, throughput, indexing lag, resource use, and failure behavior—not just the response time of one query at low load.
  4. Test growth and recovery. Check how rebalancing, node loss, replica recovery, and retention cleanup affect service while the system is under load.
  5. Choose the least complex layout that meets the service objectives. Leave room to scale, but do not create shards merely to use every available node.

For Elasticsearch specifically, treat initial primary-shard count as an important index-design decision because it is fixed at index creation. Replica count is adjustable as needs change, but replicas also consume storage and resources.

Rank #2
Sale
UGREEN NAS DXP2800 2-Bay for Advanced Home Users, Remote Workers & Creators
  • 【Advanced Home Data & Media Hub】For advanced home users who need phone backup, file storage, and centralized data management. Centralize family photos, 4K videos, movies, computer backups, and personal files in one place while running multiple apps for home entertainment and everyday data management. Suitable for households with growing digital libraries and multiple NAS use cases.
  • 【Built for Creators, Media Servers & Advanced Apps】Powered by the Intel N100 Quad-Core CPU, 8GB DDR5 RAM, 2.5GbE networking, and dual M.2 NVMe slots, DXP2800 handles large files and heavier workloads with ease. Run Docker, virtual machines, and media server applications compatible with Plex—ideal for content creators, tech enthusiasts, and advanced home users managing 4K videos, RAW photos, personal media libraries, and multiple NAS apps.
  • 【Up to 80TB for Growing Digital Libraries】 Supports up to 80TB of storage using two HDD bays and two M.2 NVMe SSD slots for family photos, movies, RAW photos, 4K videos, work files, and device backups. AI photo management supports recognition of people, objects, scenes, and locations, album organization, and duplicate photo detection. HDDs and SSDs are not included.
  • 【AI-powered Home Surveillance】Turn DXP2800 into a centralized home surveillance hub by connecting compatible network cameras and storing recordings locally on your NAS. AI-powered features include Face Recognition, People Detection, and Pet Detection, helping advanced home users review important events more efficiently while managing home surveillance and personal data in one place.
  • 【One data Center Across Your Devices】Keep files from desktops, laptops, phones, tablets, and other devices together instead of scattered across cloud accounts and external drives. Access, back up, organize, and share data across Windows, macOS, Android, iOS, web browsers, and compatible smart TVs—ideal for creators and advanced home users working across multiple devices.

How should data ingestion and index lifecycle be designed?

Indexing choices affect both search quality and operational stability. Normalize documents before writing them, use explicit mappings or schemas where field types and query behavior need to remain stable, and batch writes when the application and latency requirements allow it. Monitor indexing throughput and the delay between an update and its visibility to searchers.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Keep ingestion and query-serving workloads isolated only when measured contention or service objectives justify the added complexity. Separate clusters or dedicated resources can protect search latency from heavy writes, but they also create more infrastructure to operate and keep consistent.

Match index boundaries to retention

For data with a predictable retention window—such as dated events or logs—time-based indices or collections can make lifecycle management easier. Removing an entire expired index can free resources more efficiently than deleting many individual documents: deleted documents may remain in underlying segments until merges reclaim their space. Choose retention boundaries and rollover behavior with expected data volume and operational recovery in mind.

Rank #3
TP-Link 24 Port Gigabit Ethernet Switch Desktop/ Rackmount Plug & Play Shielded Ports Sturdy Metal Fanless Quiet Traffic Optimization Unmanaged (TL-SG1024S)
  • 𝙊𝙣𝙚 𝙎𝙬𝙞𝙩𝙘𝙝 𝙈𝙖𝙙𝙚 𝙩𝙤 𝙀𝙭𝙥𝙖𝙣𝙙 𝙉𝙚𝙩𝙬𝙤𝙧𝙠: 24 port of 10/100/1000Mbps RJ45 Ports supporting Auto Negotiation and Auto MDI/MDIX
  • 𝙂𝙞𝙜𝙖𝙗𝙞𝙩 𝙩𝙝𝙖𝙩 𝙎𝙖𝙫𝙚𝙨 𝙀𝙣𝙚𝙧𝙜𝙮: Latest innovative energy-efficient technology greatly expands your network capacity with much less power consumption and helps save money
  • 𝙍𝙚𝙡𝙞𝙖𝙗𝙡𝙚 𝙖𝙣𝙙 𝙌𝙪𝙞𝙚𝙩: IEEE 802. 3X flow control provides reliable data transfer and Fanless design ensures whisper quiet operation
  • 𝙋𝙡𝙪𝙜 𝙖𝙣𝙙 𝙋𝙡𝙖𝙮: Easy setup with no software installation or configuration needed, just plug it in and start
  • 𝙈𝙚𝙩𝙖𝙡 𝘾𝙖𝙨𝙞𝙣𝙜: Metal-cased switches provide superior durability, heat dissipation, and EMI protection, making them the clear choice for reliable performance over cheaper plastic switches.

How do replicas, routing, and failure domains affect availability?

Replicas improve resilience only if their placement avoids the same failure that would take out their primaries. Where the platform supports it, distribute copies across separate nodes and availability zones. Set recovery expectations, understand how the cluster rebalances after loss, and regularly test snapshots by restoring them; a snapshot that has never been restored is not a proven recovery path.

Routing also shapes query cost. Elasticsearch’s adaptive replica selection considers prior response time, prior search duration, and queue size when selecting a copy to handle a request. Explicit preference values can make repeat requests more likely to reach the same shard copies, helping cache locality. A routing key can constrain work when queries naturally target a tenant, region, or other partition.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Routing is not a substitute for balanced data. A poorly chosen key can concentrate documents and traffic on a small subset of shards, creating hot spots. Use it only when the query pattern has a genuine locality boundary and test the distribution under realistic traffic. Limit concurrent shard requests when needed to contain fan-out pressure; in Elasticsearch, the documented default maximum for max_concurrent_shard_requests is five per node, a version-sensitive product default rather than a universal target.

Rank #4
2 Bay DIY NAS Kit, x86 Home Server, Intel Quad-Core, 16GB RAM,
  • 【Build Your Own NAS & Homelab — Not Just Storage】 More than a traditional NAS, ZimaBlade 7700 is a flexible x86 mini server for building your own homelab, personal cloud, or Docker host. Perfect for DIY NAS, self-hosting, container apps, and even retro systems — not limited like typical ARM-based NAS devices.
  • 【x86 Platform — Broad Compatibility, Real Freedom】 Powered by an Intel quad-core x86 processor, it runs a wide range of operating systems and software with native compatibility. Ideal for Linux, Docker, CasaOS, and more — designed for flexibility and experimentation rather than locked-down appliance use.
  • 【16GB RAM for Smooth Multi-Service Workloads】 Handle file sharing, media streaming, backups, and multiple lightweight services at once. Optimized for low-power, always-on operation — a great fit for home labs and personal servers running 24/7.
  • 【Smooth 4K Media Streaming — Plex Direct Play Ready】 Stream your personal media library smoothly with Plex and similar media servers. Supports 4K playback on compatible devices via direct play, delivering a reliable home media experience without the need for heavy transcoding.
  • 【Complete 2-Bay NAS Kit — Ready to Build】 Includes power supply, 16GB RAM, metal drive cage for 2 HDD/SSD, and dual SATA cables — everything you need to start building your own NAS right out of the box.

How do you avoid slow distributed queries?

Distributed search latency is affected by the slowest work in the request as well as the coordination needed to gather results. A query that touches many shards can multiply CPU demand and queueing even if each individual shard is fast. Use query scope, routing, and concurrency controls to keep unnecessary fan-out down.

  • Scope queries deliberately. Filter by time, tenant, region, or another meaningful partition where the application can do so without changing the result semantics.
  • Use stable preference where appropriate. Repeatable shard-copy selection may improve cache reuse, but avoid a single fixed destination becoming a hot spot.
  • Measure concurrency under load. Increasing concurrent shard work may reduce the duration of an individual request while overloading shared thread pools and harming overall throughput.
  • Track the full latency distribution. Watch p50, p95, and p99 latency, errors, partial or failed shards, and queue pressure. A good average can conceal an unacceptable tail.
  • Review index and query design together. Large fan-outs may reflect a mismatch between partitioning and the application’s common search boundaries.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

What should you monitor and when should you scale?

Define service-level objectives before setting autoscaling or capacity triggers. Track the metrics that show both demand and saturation, so a team can distinguish a genuine resource shortfall from an inefficient query, skewed routing, or a recovery event.

  • Query latency percentiles, request rate, concurrency, and error rate.
  • Shard failures, queue depth, and search thread-pool pressure.
  • Indexing throughput, indexing lag, and refresh or search-visibility delay.
  • Heap and other memory pressure, disk use and watermarks, and segment merge pressure.
  • Cache hit rates, shard movement, rebalancing activity, and replica recovery progress.

Set capacity triggers against document count, stored bytes, QPS, concurrent requests, indexing rate, and latency objectives rather than relying on one metric. A rise in latency may call for more capacity, but it may also indicate excessive fan-out, a hot shard, merge pressure, or a query pattern that needs correction.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Best Value
Synology 2-Bay DiskStation DS223j (Diskless)
  • Secure private cloud - Enjoy 100% data ownership and multi-platform access from anywhere
  • Easy sharing and syncing - Safely access and share files and media from anywhere, and keep clients, colleagues and collaborators on the same page
  • Automated Backup Protection - Set-and-forget backups for Macs, PCs and mobile devices to multiple destinations including cloud and external drives
  • Home Security System - Record and monitor your property 24/7 with support for multiple IP cameras and remote viewing
  • 2-Year Warranty - Reliable hardware backed by Synology's expert customer support team and ongoing software updates

Should you use managed search or run it yourself?

A managed service can reduce the work of provisioning and operating infrastructure, but it does not remove the need to understand capacity, workload, or recovery. Self-managed clusters offer more direct control over topology and operating procedures, in exchange for responsibility for upgrades, monitoring, failure handling, and capacity planning.

Amazon CloudSearch is one documented managed model: AWS says it adjusts instance size and count for data and traffic, partitions indexes after the largest instance type is insufficient, and adds duplicate instances when request load rises. Automatic scaling can still have setup delay during a sudden traffic increase, with transient errors possible. Confirm the service’s current limits, regional availability, and behavior against your workload before relying on it for an SLO.

For other managed offerings, verify the exact scaling controls and responsibilities of the specific product and plan. The available platform facts here establish Elasticsearch’s cluster model and CloudSearch’s scaling behavior, but do not establish comparable managed-service limits, pricing, or scaling guarantees for every provider.

How do Elasticsearch, SolrCloud, OpenSearch, and CloudSearch differ?

These products do not expose identical operational models. Compare how they coordinate cluster state, place replicas, route queries, scale capacity, recover from failure, and fit your security and monitoring environment. The table separates documented distinctions from details that are not established here.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Platform Coordination and partition model Replica or routing characteristics Scaling behavior established here
Elasticsearch Nodes, shards, and replicas are integrated into the cluster model. Primary shard count is fixed at index creation; replica count can be changed during operation. (Elastic documentation) Adaptive replica selection and request controls support load-aware routing; preference and routing values can steer requests. (Elastic documentation) Adding nodes increases cluster capacity, and Elasticsearch distributes data and query load across available nodes. (Elastic documentation)
SolrCloud Uses ZooKeeper for orchestration, shard routing, and leader election. (Solr documentation) NRT, TLOG, and PULL replica types trade freshness, write cost, and query availability differently. Further routing controls are not stated here (Solr documentation). Not stated in the official-source facts summarized here (Solr documentation).
OpenSearch AWS describes integrated cluster management with manager-eligible nodes and primary and replica shards, without a separate ZooKeeper service. (AWS documentation) Not stated in the official-source facts summarized here (AWS documentation). Not stated in the official-source facts summarized here (AWS documentation).
Amazon CloudSearch Managed service; detailed shard and coordination behavior is not stated here (AWS documentation). Replica freshness and request-routing controls are not stated here (AWS documentation). AWS says it scales instance size and count for data and traffic, partitions indexes when the largest instance type is insufficient, and adds duplicate instances as request load rises (AWS documentation).

SolrCloud’s replica types make freshness and write/query trade-offs explicit; Elasticsearch’s documented model emphasizes integrated nodes, shards, replicas, and adaptive replica selection. AWS’s description of OpenSearch highlights cluster management without a separate ZooKeeper service. CloudSearch’s documented scaling approach is more managed, but its specific behavior should be evaluated against the required query patterns and failure objectives.

A practical architecture decision sequence

  1. Write down the workload. Estimate document growth, data size, retention, indexing rate, query mix, peak request concurrency, and freshness requirements.
  2. Set operational objectives. Define acceptable latency percentiles, error rates, recovery time, and how much indexing delay is tolerable.
  3. Choose the operating boundary. Decide which components your team can operate reliably and whether the managed service’s scaling and recovery controls satisfy the objectives.
  4. Design schemas and lifecycle. Normalize documents, stabilize mappings or schemas, and choose retention boundaries that support efficient cleanup.
  5. Benchmark partitioning and routing. Test candidate shard layouts and keys with realistic data and query mixes on representative hardware.
  6. Place replicas across failure domains. Establish recovery and snapshot-restore procedures, then test node-loss and recovery scenarios.
  7. Set monitoring and scaling triggers. Base alerts and capacity actions on both demand and symptoms such as queueing, latency, disk pressure, or indexing lag.
  8. Revisit the design with evidence. Re-benchmark when data shape, query behavior, retention, software version, hardware, or traffic changes materially.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.