Benchmark S3 with repeatable uploads and downloads, not a single timed transfer. Start with a serial baseline, then change one variable at a time—such as concurrency, multipart threshold or part size—while recording object size, client and bucket Regions, elapsed time, throughput, latency, retries and errors. Repeat each run and compare medians and tail results; a fast run that produces 503 responses or consumes excessive client resources is not a useful performance win.
What a useful S3 benchmark measures
S3 performance depends on more than the bucket. Record the conditions around every result so that configurations can be compared fairly:
- Transfer: PUT or GET, object size, elapsed time and calculated throughput.
- Configuration: client-to-bucket network path, client and bucket Regions, concurrency, multipart threshold, multipart part size and retry policy.
- System behavior: per-request latency, retry counts, HTTP 5xx responses, CPU, memory and network utilization.
- Cost context: the number and type of requests and the amount of data transferred. A faster configuration may use more requests or more client resources.
AWS recommends measuring network throughput, CPU, DRAM, DNS lookup time, latency, transfer speed and 503 responses when optimizing S3. Keep test objects and the client environment consistent; otherwise, a change in the network or machine can look like a change in S3 performance.
Run a repeatable Python benchmark
The example below uses Boto3’s managed transfer methods and TransferConfig. It generates fixed-size local files, warms each size with an unmeasured upload and download, then times repeated uploads and downloads for each concurrency setting. It reports each sample and the median and 95th-percentile duration and throughput. It removes the test objects and local files when it finishes.
#1 Best Overall
Install Boto3, configure AWS credentials with permission to upload, download and delete objects in the test bucket, and make sure the client’s configured Region is appropriate for the bucket. Save this as s3_bench.py:
import argparse
import math
import os
import statistics
import tempfile
import time
import boto3
from boto3.s3.transfer import TransferConfig
def write_test_file(path, size_bytes):
# Write actual bytes rather than creating a sparse file.
block = b"x" * (1024 * 1024)
remaining = size_bytes
with open(path, "wb") as f:
while remaining:
chunk = block[: min(len(block), remaining)]
f.write(chunk)
remaining -= len(chunk)
def percentile(values, fraction):
ordered = sorted(values)
index = max(0, math.ceil(fraction * len(ordered)) - 1)
return ordered[index]
def measure_transfer(operation, source_or_destination, bucket, key, config):
start = time.perf_counter()
if operation == "PUT":
bucket.upload_file(source_or_destination, key, Config=config)
else:
bucket.download_file(key, source_or_destination, Config=config)
return time.perf_counter() - start
def main():
parser = argparse.ArgumentParser(description="Repeatable Boto3 S3 transfer benchmark")
parser.add_argument("bucket", help="S3 bucket name")
parser.add_argument("--prefix", default="benchmark-test", help="Test-object key prefix")
parser.add_argument("--sizes-mib", nargs="+", type=int, default=[16, 128])
parser.add_argument("--concurrency", nargs="+", type=int, default=[1, 4, 8])
parser.add_argument("--runs", type=int, default=5)
parser.add_argument("--part-mib", type=int, default=16)
parser.add_argument("--threshold-mib", type=int, default=16)
args = parser.parse_args()
if args.runs < 1 or args.part_mib < 1 or args.threshold_mib < 1:
parser.error("runs, part-mib and threshold-mib must be positive")
if any(size < 1 for size in args.sizes_mib):
parser.error("all sizes-mib values must be positive")
if any(value < 1 for value in args.concurrency):
parser.error("all concurrency values must be positive")
session = boto3.Session()
s3 = session.resource("s3")
bucket = s3.Bucket(args.bucket)
print(f"Client Region: {session.region_name or 'not set in Boto3 session'}")
print(f"Bucket: {args.bucket}; runs per operation/configuration: {args.runs}")
for size_mib in args.sizes_mib:
size_bytes = size_mib * 1024 * 1024
key = f"{args.prefix.rstrip('/')}/{size_mib}MiB.bin"
local_objects = []
try:
with tempfile.TemporaryDirectory() as temp_dir:
source = os.path.join(temp_dir, "source.bin")
destination = os.path.join(temp_dir, "download.bin")
write_test_file(source, size_bytes)
# Warm credentials, DNS and the transfer path; do not include these timings.
warm_config = TransferConfig(
multipart_threshold=args.threshold_mib * 1024 * 1024,
multipart_chunksize=args.part_mib * 1024 * 1024,
max_concurrency=1,
use_threads=False,
)
bucket.upload_file(source, key, Config=warm_config)
bucket.download_file(key, destination, Config=warm_config)
for concurrency in args.concurrency:
config = TransferConfig(
multipart_threshold=args.threshold_mib * 1024 * 1024,
multipart_chunksize=args.part_mib * 1024 * 1024,
max_concurrency=concurrency,
use_threads=(concurrency > 1),
)
for operation in ("PUT", "GET"):
durations = []
for run in range(1, args.runs + 1):
seconds = measure_transfer(
operation,
source if operation == "PUT" else destination,
bucket,
key,
config,
)
durations.append(seconds)
mib_per_second = size_bytes / seconds / (1024 * 1024)
print(
f"size={size_mib}MiB concurrency={concurrency} "
f"operation={operation} run={run} seconds={seconds:.3f} "
f"throughput={mib_per_second:.2f}MiB/s"
)
median_seconds = statistics.median(durations)
p95_seconds = percentile(durations, 0.95)
median_rate = size_bytes / median_seconds / (1024 * 1024)
p95_rate = size_bytes / p95_seconds / (1024 * 1024)
print(
f"SUMMARY size={size_mib}MiB concurrency={concurrency} "
f"operation={operation} median={median_seconds:.3f}s "
f"p95={p95_seconds:.3f}s median_rate={median_rate:.2f}MiB/s "
f"p95_duration_rate={p95_rate:.2f}MiB/s"
)
finally:
bucket.Object(key).delete()
if __name__ == "__main__":
main()
Run it with, for example:
python s3_bench.py my-test-bucket --sizes-mib 16 128 --concurrency 1 4 8 --runs 5
The script treats concurrency 1 as a serial control by disabling transfer threads. For values above 1, Boto3 can use threads up to the configured concurrency. The script holds the multipart threshold and part size constant during its concurrency sweep; change those values in a separate sweep rather than changing several settings at once. Its throughput is calculated from the known object size divided by transfer-call wall time. It does not capture request-level latency, retries, HTTP status codes, CPU, memory or network utilization, so collect those separately when those measurements matter. Run order is fixed in this example; for careful comparisons, randomize configuration order and repeat the full sweep.
Understand the Boto3 transfer settings
| Setting | What it controls | Benchmark implication |
|---|---|---|
multipart_threshold |
Size at which managed transfers use multipart behavior. | Keep fixed while testing concurrency; then compare thresholds with other settings held steady. |
multipart_chunksize |
Size of each multipart part. | Compare part sizes for large objects and record the value with each result. |
max_concurrency |
Maximum concurrent transfer work for managed multipart transfers. Boto3 documents a default of 10. | Lower values can reduce bandwidth use; higher values can use more bandwidth. Sweep values rather than assuming the maximum is fastest. |
use_threads |
Enables or disables threaded transfer work. | Set to False for a serial control run. When threads are disabled, max_concurrency has no effect. |
num_download_attempts |
Controls download attempts in managed transfers. | Record the setting and retry behavior; otherwise, a result with retries may not be comparable to one without them. |
io_chunksize |
Controls the I/O chunk size for downloads. | Change it independently and observe both transfer results and client resource use. |
The Boto3 high-level upload_file and download_file methods manage multipart and non-multipart transfers and include retry handling. For a lower-level implementation, use exponential backoff for retryable failures and, where appropriate, retry over a fresh connection. Record the retry policy because retries affect elapsed time and successful-transfer throughput.
Rank #2
How to interpret throughput and latency
Calculate throughput using the actual object size divided by elapsed transfer time, and state whether the reported value is bytes per second, MiB/s or another unit. Report the median across repeated runs and a tail measure such as the 95th-percentile duration. Keep the individual samples: a median alone can hide occasional slow transfers or retries.
Compare more than aggregate throughput. Include median and tail latency, CPU and memory consumption, error and retry rates, object size, part size, concurrency, client-to-bucket distance and transfer cost. A configuration that improves throughput while sharply increasing retries, 5xx responses or client load may not be suitable for production.
Start with a single-request baseline, then increase concurrency gradually while watching the client and S3 error metrics. AWS describes multiple concurrent requests over separate connections as a way to use available bandwidth. Its performance guidance also advises tracking throughput and retrying the slowest 5 percent of large, variably sized requests—for example, requests over 128 MB. That is guidance to evaluate for a workload, not a universal threshold or guaranteed improvement.
Rank #3
When multipart transfers and byte ranges help
Large uploads
Compare a single-stream transfer with multipart parallel transfer for large objects. Multipart lets a managed transfer work on parts concurrently, but the useful part size and concurrency depend on the client, network and workload. Keep the object fixed while testing these settings, and watch memory, CPU and retry behavior as well as total time.
Large downloads
For large downloads, compare managed transfers with concurrent byte-range GETs or parallel retrieval of multipart parts. AWS recommends aligning GET ranges with the original multipart boundaries where possible. The benchmark should use the same object and range layout across repetitions so the comparison measures the retrieval strategy rather than different data.
Do these 3 things before closing this tab:
1Scan for outdated or missing drivers - takes under a minute2Clear out junk files and repair common Windows errors3Fix the driver behind crashes, sound loss and screen glitchesSmall objects and request-heavy workloads
For small objects, request latency and request rate can matter more than bulk throughput. AWS states that applications can achieve thousands of S3 transactions per second. Its current guidance gives reference figures of at least 3,500 PUT/COPY/POST/DELETE requests per second or 5,500 GET/HEAD requests per second per partitioned S3 prefix. These are service guidance figures, not a promise for a particular bucket or benchmark; actual results vary by workload, client configuration, object size, network and Region, and scaling is gradual.
Rank #4
Diagnose 503 Slow Down responses
A 503 spike does not automatically mean that a bucket is permanently slow. S3 may return temporary 503 responses while adapting to a new request rate. A sudden increase in request volume or a concentrated prefix can also contribute to a spike. Ramp request rates progressively rather than jumping straight to a high concurrency setting.
- Record request counts, 5xx responses, retries and latency for each test configuration.
- Check CloudWatch S3 request metrics, S3 Storage Lens or server access logs for 5xx responses when those monitoring sources are available and enabled.
- If errors appear after an increase, reduce the rate, allow the workload to stabilize, then raise it in measured steps.
- Keep SDK retries enabled and record their behavior; retries can mask transient errors in a success-only throughput figure.
Do not treat a request-rate reference as an individual-bucket target. The measured outcome depends on the workload and setup, so use the rate figures as context while watching the actual error and latency data.
Control Region, network path and transfer acceleration
Run the client near the bucket’s AWS Region when practical: reducing distance can lower latency and transfer cost. Record both Regions and the network path in benchmark notes. Results from a client on a different network or in a different Region are not directly comparable.
Quick wins for a faster PC:
Scan for outdated or missing drivers - takes under a minuteDriver Scan →Clear out junk files and repair common Windows errorsFree Scan →Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Best Value
Amazon S3 Transfer Acceleration is a separate option for long-distance transfers. Measure it against the same objects and conditions rather than assuming it will be faster; include its costs in any comparison. Also record DNS lookup time, network throughput and client resource use, because bottlenecks outside S3 can limit observed transfer speed.
Clean up and make results reproducible
The example deletes its named test object and temporary local files when it exits. Use a dedicated test bucket or a unique prefix to keep test data separate from production objects. For workloads using multipart uploads, account for incomplete multipart uploads as well; a bucket lifecycle rule can be used to abort incomplete uploads, or clean them up through your own maintenance process.
- Record the test date, client Region, bucket Region, object sizes, operation, concurrency, multipart settings, retry policy and network path alongside results.
- Repeat each configuration enough times to compare medians and tail behavior, and avoid concurrent unrelated transfers during a controlled run.
- Randomize test order when practical, so a time-dependent change in the network or service does not consistently favor one configuration.
- For high-rate tests, ramp up progressively and monitor 5xx responses instead of treating an isolated peak as a sustainable result.
AWS Labs’ aws-crt-s3-benchmarks repository includes comparisons of S3 libraries and Python runners such as boto3-classic. It can serve as a reference for test orchestration; your own results still need the Region, network, object-size and concurrency assumptions recorded.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.
The Tool Desk
Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →

