iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more
To speed up ClickHouse ingestion, stop sending frequent tiny synchronous inserts. Batch rows at the producer when practical; when producers cannot buffer effectively, test asynchronous inserts with wait_for_async_insert=1 so clients receive flush failures. ClickHouse’s current guidance recommends at least 1,000 rows per synchronous insert and calls 10,000–100,000 rows an ideal range, but your workload—not a universal number—should determine the final batch size.
Why row-by-row inserts slow ClickHouse down
Each insert can create data parts that ClickHouse must later merge. Frequent small inserts can create parts faster than background merges consolidate them, consuming CPU and I/O and potentially affecting query performance. Batching reduces that part-creation pressure; it does not eliminate merge work or make partition design irrelevant.
In a ClickHouse 2023 illustrative workload, 200 synchronous inserts every 10 seconds produced around 200 new parts per second. The authors report that the workload reached the active-parts safeguard after five minutes and was aborted. That example shows how quickly tiny inserts can accumulate parts in a particular setup; it is not a general throughput limit. ClickHouse’s async-insert article describes the example and the underlying issue.
How large should a ClickHouse insert batch be?
ClickHouse’s resource guidance recommends at least 1,000 rows per synchronous insert and identifies 10,000–100,000 rows as an ideal range. Treat the larger range as a starting point for testing, not a fixed prescription: row width, insert latency targets, memory limits, schema, partitioning, and concurrency all affect a useful batch size. ClickHouse’s bulk-insert guidance provides the current recommendation.
#1 Best Overall
If the producer can buffer rows without unacceptable latency or memory use, aggregate them and send fewer, larger inserts. Measure end-to-end visibility latency as well as throughput: waiting to accumulate a batch means newly produced rows may take longer to reach queries.
Producer batching or asynchronous inserts?
| Consideration | Producer-side batching | Asynchronous inserts |
|---|---|---|
| Where batching happens | The application or client buffers rows before sending a synchronous insert. | ClickHouse buffers compatible incoming inserts on the server and flushes them later. |
| Best fit | Producers can conveniently aggregate data and tolerate the buffer’s latency and memory cost. | Many independent producers cannot conveniently coordinate or buffer their own batches. |
| Acknowledgment and errors | The synchronous insert response reports the result of that insert. | With wait_for_async_insert=1, acknowledgment follows the flush and flush errors are returned to the client. |
| Visibility timing | Rows wait for the producer’s batch to fill or flush before the insert is sent. | Rows wait in a server-side buffer until its applicable flush trigger. |
| Parts and merges | Larger inserts generally reduce part creation compared with sending the same rows in many tiny inserts. | Buffering can combine inserts, but each flush still consumes server resources and can create multiple parts. |
Asynchronous inserts move batching to ClickHouse; they do not make ingestion free of part creation or merge work. A flush may create multiple parts because of different partitions, oversized data, or separate buffers or nodes. Insert shape and settings also matter: only compatible incoming queries can collect together in the relevant buffers. Test with the actual distribution of producers and inserts.
Rank #2
Configure asynchronous inserts for observable failures
For production workloads where the client needs to know whether buffered data was flushed successfully, use wait_for_async_insert=1. ClickHouse documents this as the recommended production mode: the client is acknowledged after the flush, and flush errors are returned. Fire-and-forget mode acknowledges memory buffering instead. That can hide errors from the client and carries the risk of buffered data loss, so use it only if the application accepts those trade-offs. ClickHouse documents these behaviors in its asynchronous-insert guidance.
Windows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallOutdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchMake the setting explicit in a query rather than relying on an assumed default:
Rank #3
INSERT INTO events SETTINGS async_insert = 1, wait_for_async_insert = 1 FORMAT JSONEachRow
{"event":"page_view","user_id":42}
For a client or driver that sets query parameters separately, configure the equivalent settings there. Confirm the syntax supported by your client and the effective server settings for your deployed version.
Check version behavior before changing settings
ClickHouse’s 26.3 LTS release announcement says asynchronous inserts are enabled by default starting with version 26.3. It describes buffer flushes triggered by a timeout, accumulated size, or number of inserts. The announcement also notes adaptive async-insert timeouts in 24.2 and a consistent deduplication mechanism for asynchronous inserts with materialized views in 26.1. Defaults and behavior can depend on the deployed version and configuration, so verify the effective settings rather than assuming a release-wide default applies to your installation. See the ClickHouse 26.3 LTS release announcement.
Choose an input format and compression by testing
Batch size is only one part of ingestion performance. ClickHouse’s FastFormats benchmark considered more than 70 input formats and found that Native led in essentially all tested scenarios. The same vendor benchmark identifies LZ4 as a strong compression choice; ZSTD can be relevant when network bandwidth is the constraint. These are benchmark findings, not guarantees for every schema, client, or machine. Read ClickHouse’s input-format benchmark and reproduce comparisons on your workload, including implementation complexity and client support.
The benchmark article also reports that Netflix processes about 5 PB of logs per day after adopting native-protocol encoding with LZ4. This is a vendor-reported case study, not an independently verified benchmark or a throughput target for other deployments.
Best Value
Benchmark the full ingestion path
Change one factor at a time and use representative traffic. A faster insert benchmark alone may not reveal higher query latency, excessive memory use, or a growing part backlog.
- Batch size and producer parallelism: Test sensible batch sizes, including ClickHouse’s recommended range, alongside the number of concurrent producers.
- Insert mode and acknowledgment: Compare synchronous producer batching with asynchronous inserts, retaining
wait_for_async_insert=1when clients need flush errors. - Format and compression: Compare realistic client-supported formats and compression choices using the same data and workload.
- Partitioning and data shape: Use your actual schema, row widths, and partition distribution; one flush can still create several parts.
- Operational effects: Track insert throughput and latency, query effects, CPU, memory, and part counts while running the production workload mix. ClickHouse’s resource-usage guidance emphasizes evaluating the workload as a whole rather than sizing query classes in isolation.
Keep the setting that improves ingestion without creating unacceptable visibility delays, resource pressure, or query impact. If neither batching strategy performs well, investigate the full ingestion path—especially partitioning, input format, producer concurrency, and available server resources—instead of treating batch size as the only control.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.

