Free tools Windows power users keep installed
One-click scans. No signup required.
A load spike can overwhelm a Liberty-based service even when its gateway stays healthy. In a staging incident described by Josephine Eskaline Joyce and Ajay Chebbi, the response was not a single Liberty setting: the team addressed application threads, database pooling, inconsistent timeouts, and an unregulated private ingress path. Their account is a useful case study, not a set of universal configuration defaults.
What failed during the load test
In their March 3, 2025 DZone case study, Joyce and Chebbi describe a tenant sending rapidly growing traffic through a private endpoint. During staging tests, Java microservices saw rising CPU and memory use, threads hung, and JMeter requests timed out. The Go-based gateway remained stable while the Liberty-based Java app server hung; the authors also observed that database connections did not increase as they expected.
The system ran on Kubernetes distributed across three zones and included Istio, a Go gateway, a Liberty Java app server, PostgreSQL, Redis, and pgBouncer. Public traffic passed through IBM Cloud Internet Services with rate limiting, but the private Istio ingress gateway had no rate limit. The incident therefore involved more than Liberty: the team identified an unregulated traffic route, inconsistent gateway timeouts, and application-thread and database-connection behavior as parts of the problem.
How the team changed the system
The authors applied controls at different layers, aiming to limit work entering the service and constrain the resources consumed downstream. These were decisions for their described environment; another deployment needs workload and capacity analysis before adopting similar values.
Do these 3 things before closing this tab:
1Repair Windows errors before they cause bigger problems2Scan for outdated or missing drivers - takes under a minute3Clear out junk files and repair common Windows errors#1 Best Overall
| Layer | Team’s change | Failure mode addressed |
|---|---|---|
| Liberty request threads | The authors set the maximum thread count (maxTotal) to 200, which they said matched the maximum HTTP request threads available in their setup. They also report changing related pool parameters. |
Uncontrolled thread growth and hung threads in the application server. |
| pgBouncer and PostgreSQL | Initially, pgBouncer used session pooling with max_client_conn set to 200 per instance. With three instances, the authors concluded this could permit more connections than PostgreSQL’s configured maximum of 400. They switched to transaction pooling and set max_client_conn to 100 per instance. |
A mismatch between possible client connections and the database’s configured connection limit. |
| Timeouts | The authors aligned the previously inconsistent Nginx and Istio timeout settings at 60 seconds. | Different timeout behavior across the request path. |
| Private ingress | They added rate limiting at the Istio private gateway and a Retry-After response header. |
Excess requests arriving through the previously unregulated private endpoint. |
| pgBouncer version | They upgraded pgBouncer; the authors said this did not directly affect resilience. | A version update was included in the response, but was not credited with the resilience improvement. |
Why the controls work as a set
Each change constrains a different part of the request path. Rate limiting curbs the rate of incoming private traffic; the Liberty thread limit constrains concurrent application work; and transaction pooling plus a lower per-instance client limit changes how database connections are managed. Aligning gateway timeouts makes the request path less inconsistent, though a timeout by itself does not add capacity. The authors present these as complementary controls rather than a single cure.
What results the authors reported
For a GET request retrieving 122 KB and involving approximately 7–9 database calls, under a load of 400 concurrent API requests, Joyce and Chebbi reported latency falling from 9 seconds to 2 seconds. They also reported a fivefold increase in concurrently handled requests. These are the authors’ staging case-study results, not independently audited benchmark measurements or a guarantee for other workloads.
Rank #2
The authors said errors fell substantially and customers then mainly saw HTTP 429 responses when they sent too many requests within a period. They did not provide a quantified error-rate measurement. A 429 is an explicit indication that rate limiting is applying; the Retry-After header can tell clients when to try again, so clients should respect it rather than immediately retrying and adding more load.
How to apply the lessons to another Liberty deployment
Use the case study as a diagnostic sequence, not a configuration recipe. Establish what your system can safely process, then test the boundaries at each layer.
- Trace every ingress path. Inventory public and private routes through gateways and service meshes. Confirm that rate limits cover the paths tenants and internal callers can actually use.
- Observe the failure as it develops. Log request failures and inspect hung threads, private-endpoint traffic, and active database connections alongside CPU and memory. The authors stress logging and monitoring because the gateway’s apparent health did not reveal the Liberty server’s problem.
- Set limits from capacity, not from this case’s numbers. Relate application thread limits and database connection ceilings to the workload, downstream capacity, and the total number of service or pool instances. A per-instance setting multiplies across replicas.
- Check pooling semantics and connection arithmetic. Compare the possible client connections across all pgBouncer instances with PostgreSQL’s configured capacity. If changing pooling mode, verify that the application’s transaction behavior is compatible with it.
- Make timeout behavior deliberate across the route. Review proxy and mesh timeouts together, and ensure their relationship fits the expected request duration. A longer timeout may keep work occupied longer; a shorter one can reject work sooner but does not fix a saturated dependency.
- Load-test and tune incrementally. Exercise both public and private paths, track latency, failures, thread behavior, and database connections, and change one relevant control at a time where possible. Keep the findings specific to the test conditions.
What this case study does—and does not—establish
The account supports a practical lesson: when a Liberty service hangs under load, investigate the entire path into and through it, including private ingress, gateway timeouts, application concurrency, and database pooling. The published account describes the architecture, selected settings, and outcomes, but does not provide an independently verified benchmark protocol, code repository, or controlled comparison of alternative settings. Treat its figures and configuration values as one team’s report from staging, not as Liberty or pgBouncer defaults.
As the authors put it in their conclusion, “Resilience isn’t a one-time fix — it’s a mindset.”
Quick Recap
Best Value
Rank #4
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

