Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

A FastAPI and Celery deployment on a 4 GB EC2 instance reportedly reduced memory use from about 3 GB to 1.5 GB after its author changed API concurrency and recycled both API and task-worker processes. Those are observations from one deployment, not a controlled or independently reproduced benchmark. The useful lesson is the combination of fewer processes and bounded process lifetimes—not a guarantee that the same settings will halve memory elsewhere.

What changed in the reported deployment?

Rohan Sen Sharma described a Python backend with a FastAPI API served by Gunicorn and background jobs handled by Celery. His account compares the deployment before and after changing API process/thread configuration, lowering Celery concurrency, and recycling processes. The article does not publish a workload profile or a controlled comparison that isolates the effect of each change.

Measure Before After
Reported process or worker count 12 total: four API workers and eight Celery workers distributed across three queues (Sharma, DEV Community) Six total, described as API threads plus reduced Celery concurrency (Sharma, DEV Community)
Reported memory use About 3 GB, or 75% of the 4 GB instance (Sharma, DEV Community) About 1.5 GB, or 40% of the 4 GB instance (Sharma, DEV Community)
Per-concurrency comparison About 500 MB for equivalent concurrency using processes (Sharma, DEV Community) About 250 MB using threads (Sharma, DEV Community)

Sharma also estimated that each Python worker process used about 250 MB on average. The article does not specify the measurement method, observation period, traffic mix, or exact thread and concurrency settings. Treat every figure as the author’s reported deployment value, not a general memory baseline.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The reported reduction has two distinct parts: fewer concurrent processes and process recycling. Because the article changed multiple settings together, it does not establish how much of the reduction came from each part, nor whether the service sustained the same throughput and latency.

Which settings did the author change?

API concurrency: threads rather than as many processes

The account describes changing the API’s process/thread configuration and reports lower memory for equivalent concurrency when using threads. Gunicorn’s threads setting applies to its gthread worker type; it is not a generic replacement for choosing a worker class. Consult the Gunicorn 20.1.0 settings documentation for the relationship between worker type and thread settings.

Thread-based concurrency can be worth evaluating for an I/O-bound API that spends time waiting on network or database operations. It is not an automatic fit for CPU-bound work, and it does not make each request’s memory requirements disappear. Compare representative throughput and latency as well as memory before changing the deployment.

Celery: replace children after a task limit

The author configured Celery with --max-tasks-per-child=100, limiting how many tasks a pool child handles before it is replaced. This can contain high-water memory behavior or growth that accumulates in a child over time; it does not fix an underlying task that allocates too much memory or a genuine leak.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Celery also documents worker_max_memory_per_child, which replaces a prefork child after it exceeds a configured resident-memory threshold. Task-count and memory-based limits address different triggers. Celery cautions that excessively frequent replacement wastes processing capacity, so the author’s value of 100 is a workload-specific example, not a recommended default. See Celery’s optimizing guide.

Gunicorn: recycle requests and stagger restarts

The reported API configuration used --max-requests=1000 and --max-requests-jitter=50. Gunicorn defines max_requests as the number of requests a worker handles before restarting; jitter adds a randomized offset so workers do not all restart together. These controls can limit long-lived worker growth and stagger the resulting restarts, but they do not establish that 1,000 requests is suitable for another application. The mechanism and settings are documented in Gunicorn’s version 20.1.0 settings reference.

How to decide whether these changes fit your service

  1. Measure the right processes. Track resident memory (RSS) for the Gunicorn master and workers, and for Celery’s master and child processes. Observe both ordinary operation and workload peaks; a single service-wide snapshot can hide which process is growing.
  2. Separate API and task-worker decisions. Gunicorn’s worker and thread configuration governs API request handling. Celery’s pool and concurrency govern background tasks. Changing one does not imply changing the other.
  3. Test realistic work. Compare memory, throughput, and latency under representative I/O-bound and CPU-bound traffic. Include task duration and peak allocations for Celery jobs rather than judging only steady-state averages.
  4. Account for recycling costs. Record restart frequency, task or request interruption behavior, and the overhead of starting replacement processes. Tune task-count, memory, and request thresholds against observed growth and acceptable restart overhead.
  5. Check pool feature compatibility. Celery identifies prefork as its default and generally recommended starting point. Its concurrency documentation warns that alternate pool models may silently disable features, including max_tasks_per_child. Verify the behavior required by your deployment before selecting a different pool. See Celery’s concurrency guide.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

What the reported result does—and does not—show

The author’s figures show that his deployment reported lower memory use after reducing process count and introducing recycling. They do not demonstrate that threads alone caused the reduction, that the before-and-after workloads were equivalent, or that the revised configuration preserved performance. The source does not provide an independently reproducible benchmark or enough workload detail to transfer its thresholds directly.

Celery’s guidance explains why child replacement can help when a prefork child retains a high-water memory footprint: the memory may not be released until that child exits. Recycling can contain the symptom while you investigate. If task peaks are the cause, reducing a task’s memory demand or processing work in smaller chunks may address the underlying pressure more directly; if memory keeps growing, distinguish child growth from growth in the master process. See Celery’s optimizing guide.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.