iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more
A single GraphQL request can reduce the number of calls a client makes to a server, but it does not guarantee one backend call, parallel execution, or a faster complete response. Nested resolvers can still trigger repeated data loads, and federated services can still need to fetch data in sequence. The waterfall may disappear from the browser’s network panel while continuing inside the server.
Which GraphQL waterfall are you trying to remove?
“Waterfall” can describe several different costs. Distinguishing them matters because a fix for one may leave the others unchanged.
- Client-to-server round trips: how many requests the application sends before it has the data it needs. GraphQL can combine related selections into one operation, potentially reducing these trips and avoiding over-fetching. The GraphQL FAQ describes these as potential benefits, not a speed guarantee: GraphQL FAQ.
- Backend calls: how many database, service, or subgraph requests the server makes to resolve that operation. One incoming request can fan out into many calls.
- Time to useful UI: when the client can render information that is meaningful to the user. A response that waits for every field may arrive later than a partial response, even if both involve the same underlying work.
Measure these separately. A lower browser request count does not prove fewer backend calls, and an earlier first payload does not prove that the whole operation finished sooner.
Why one GraphQL operation can still create an N+1 problem
GraphQL lets a client request nested data in one operation, but the server’s resolvers determine how that selection is loaded. If a resolver fetches a list of records and then another resolver loads related data once for each record, the service can issue repeated backend requests. This is the N+1 problem: one load for the list, followed by one load per item.
#1 Best Overall
For example, a query for a set of events and each event’s venue may look like one coherent client request. If the venue resolver independently queries the database for every event, the server can still make many database calls. Apollo’s request-waterfall article uses an events application to illustrate this behavior and approaches to batching across different implementations; it is an implementation example, not a universal benchmark: Optimizing Your GraphQL Request Waterfalls.
Batch repeated loads at the data-access boundary
A common remedy is to collect related identifiers and load them together, rather than issuing one backend request per identifier. Request-scoped loading tools such as DataLoader can batch loads over a short interval and often cache duplicate keys within the request. Other server implementations may translate a GraphQL selection set into a more efficient source query. The appropriate method depends on the backend and resolver architecture.
Batching changes how repeated loads are performed; it does not make expensive fields free or guarantee that every resolver can be combined. Check that a batch preserves the requested ordering and handles missing records correctly, and avoid sharing request-specific cache state across users or requests. The GraphQL performance guide discusses N+1 behavior, batching, and related techniques: GraphQL performance guidance.
Recommended Free Tools
Why federation can preserve a serial waterfall
In a federated graph, a router may call multiple subgraphs to assemble one result. Some calls can happen in parallel, but a later fetch must wait when it needs a value produced by an earlier one.
Rank #3
Apollo’s documented Products/Reviews query plan shows this dependency: the router first fetches products, then uses product identifiers to request their reviews. Because the second sub-query needs data from the first, those sub-queries run serially. Combining the work into one client operation reduces client coordination, but it cannot remove the data dependency. See Apollo Router’s @defer documentation for that example and its support details.
Inspect the plan, not just the operation text
For a federated query, examine the router’s query plan and trace spans to identify which subgraph fetches are independent and which depend on prior results. A long operation can contain parallel branches; a short one can contain a strictly sequential chain. The dependency order—not the number of fields in the query—explains whether the router must wait.
When @defer can help—and what it cannot do
The @defer directive can let a compatible server send ready, non-deferred data before slower deferred fields are available. The client can render that initial portion and incorporate the later payload when it arrives. This may improve perceived responsiveness when the early data is useful on its own.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Incremental delivery changes when results are delivered; it does not eliminate a dependency such as “fetch products, then use their IDs to fetch reviews.” The later work still has to happen, and the total completion time may be unchanged. A product should use this approach only when partial results improve the experience enough to justify the additional client and server complexity.
Best Value
Check compatibility before adopting it
Apollo documents Router support for @defer in Router v1.8.0 and newer, with clients required to handle multipart HTTP responses. Those are the documented requirements in that source; verify compatibility against the versions and configuration actually deployed.
The GraphQL Working Group’s defer/stream RFC is a working draft, identified in its introduction as September 2024. It says servers are not required to implement the directives and describes cases where clients must tolerate a server not deferring or streaming as requested. Its draft also discusses possible extra latency, client resource contention, higher server or data-layer costs, and repeated client rendering. Read the defer/stream RFC as a protocol proposal, not proof that every GraphQL stack supports incremental delivery.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Choose the fix for the cost you measured
| Technique | What it targets | What it does not guarantee |
|---|---|---|
| Batching and request-scoped loading | Repeated backend loads, including resolver-level N+1 patterns | Fewer client round trips, faster completion for every query, or lower total work in every case |
| Client-side caching | Repeated requests for data the client can reuse | Efficient first-time loading or efficient resolver behavior |
| Persisted query hashes | Request handling and the overhead of sending full query text, where supported | Fewer dependent backend fetches |
| GET requests for queries | Cacheability, where the server and client support query requests over GET | Resolution of N+1 loads or serial subgraph dependencies |
| Compression | Transferred payload size | Less backend work or earlier completion of a dependency chain |
| Pagination | Oversized result sets and the work needed for a page of data | Faster resolution of every individual field |
| Incremental delivery with @defer | Time until useful initial data can be delivered and rendered | Removal of underlying work or a shorter time to the final payload |
| Depth, breadth, batch, or query-cost controls | Protection against expensive or excessive operations | Optimization of a legitimate query’s resolver or subgraph plan |
These tools address different dimensions. GraphQL’s performance guidance also covers monitoring and persisted queries, while its security guidance explains why batching alone does not neutralize excessive nesting or costly combinations of fields. See performance guidance and GraphQL security guidance.
Do these 3 things before closing this tab:
1Repair Windows errors before they cause bigger problems2Fix the driver behind crashes, sound loss and screen glitches3Clear out junk files and repair common Windows errorsHow to diagnose the waterfall in your application
- Record the client experience. Capture request timing, time to first useful content, and time until the complete result is available. If using incremental delivery, distinguish the initial payload from later payloads.
- Trace server resolution. Add or inspect resolver spans so you can see which fields consume time and whether nested resolvers repeat loads.
- Count backend work. Measure database calls, service requests, and subgraph fetches for the operation. Compare the count and timing with the request trace rather than assuming one GraphQL request means one source call.
- Inspect the federation query plan. Look for serial fetches and identify the values that create each dependency. Determine whether a later fetch can be redesigned or whether the dependency is inherent in the data relationship.
- Apply one targeted change and remeasure. Batch repeated loads to address N+1 behavior; use caching for reusable results; consider incremental delivery when early data is useful and the stack supports it. Compare first-useful-content time, total completion time, backend work, payload size, and client rendering behavior.
- Keep demand controls in place. Set suitable limits for depth, breadth, batches, or query cost so that a single operation cannot request unbounded work. GraphQL’s security guidance covers these risk controls.
There is no universal speedup to promise: performance depends on the query, resolver design, backend, router plan, and client. The useful question is not simply whether GraphQL reduced request count, but which part of the measured path improved—and whether total work and completion time improved too.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

