Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

An abstraction can make an operation easy to call while hiding the work it triggers. That work may be measurable in a benchmark, yet impossible to infer from the source line alone: the same-looking call might use a fast user-space path, cross into the kernel, perform many smaller operations, or batch work efficiently. The answer is not to avoid abstraction, but to understand what it hides and measure it under the workload you actually run.

What “invisible at the call site” means

A call site is the line of code where a function, method, or interface is invoked. Abstraction intentionally lets that line describe what the program wants rather than every step of how the operation is performed. The implementation details may be in a library, runtime, kernel, or remote service.

That separation is useful, but it makes costs harder to spot by reading the source. A cost can be measurable without being recognizable in the syntax: a brief-looking call may trigger allocation, copying, synchronization, a system call, or network activity. Chris’s September 8, 2026 article distinguishes these measurable-but-hidden costs from costs that are obvious in code and costs that are difficult to measure at all. The hidden category is easy to overlook because the call itself gives little reason to budget for or investigate that work. Chris’s article.

As Chris puts it, “You can’t tell which one you’re looking at from the call site, because the call site is doing its job.” The call site is not necessarily misleading; it is simply not a full account of the implementation.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

What work can a simple call conceal?

Different layers can sit behind familiar syntax. These examples illustrate the kinds of hidden work to look for, not universal performance rankings.

A user-space function or a kernel transition

Some operations that look like ordinary library calls can be handled without entering the kernel on supported systems. On Linux, the kernel maps a virtual dynamic shared object (vDSO) into user-space processes. A C library can use it for certain supported operations, avoiding a system call on that path. Which operations are available, and how they are implemented, depends on the architecture and kernel. The Linux vDSO manual.

That makes two calls that appear similarly small in source potentially different at the boundary: one may stay in user space, while another requires a transition into the kernel. Chris reports that, on the author’s laptop, clock_gettime(CLOCK_MONOTONIC, ...) took about 17 nanoseconds and syscall(SYS_getpid) about 107 nanoseconds—roughly six times as long for the latter in that comparison. These are observations reported for one laptop, not general constants for other CPUs, kernels, C libraries, or clock sources. Chris’s article.

Extra copies or work repeated per call

A high-level operation may move data through user space even when the program’s goal is simply to transfer it. Linux sendfile() transfers data between file descriptors within the kernel. Its manual describes this as more efficient than a read()/write() sequence in the sense that the latter requires transferring data to and from user space. That is a design advantage, not a promise of a fixed speedup for every workload. The Linux sendfile(2) manual.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Other familiar syntax can hide repeated work too. Chris’s article cites ORM N+1 queries, remote method calls, and string concatenation in a hot loop as examples. Their actual cost depends on the implementation and workload; the examples are prompts to inspect what happens beneath the call, not evidence that every use is slow. Chris’s article.

Coordination and batching choices

Linux io_uring provides shared submission and completion queues and supports asynchronous requests and batching. This can change how often an application must submit work, but it also adds configuration choices. For example, submission-queue polling (SQPOLL) can avoid some submission calls while its polling thread consumes CPU when active. Whether that trade is worthwhile depends on the workload; polling is not an automatic optimization. The io_uring(7) manual and io_uring_sqpoll(7) manual.

Why frequency and workload shape change the cost

The cost of an operation is not just the time for one call. It also depends on how often the program makes that call and how much useful work each call accomplishes. A fixed overhead that is negligible beside a large transfer can dominate when repeated for tiny pieces of work.

Chris illustrates this with a hundred-nanosecond operation repeated across a million tiny reads: multiplying those figures gives a tenth of a second of transition time. That is explanatory arithmetic, not a separate measured result or a claim that each real read incurs exactly that cost. Chris’s article.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Buffering and batching can amortize per-operation overhead by grouping useful work into fewer calls. The trade is that buffering can affect when data is delivered, memory use, and latency, while batching can change responsiveness and resource consumption. The right choice depends on the application’s requirements, not just the lowest time per operation.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

How to assess a hidden cost

When two abstractions express similar intent, examine the implementation path and the workload together. These questions help turn an invisible cost into a testable one:

  • What work is behind the call? Check whether it allocates, copies data, synchronizes, enters the kernel, performs I/O, or invokes a remote service.
  • How often is it called, and how much useful work does each call do? A small fixed cost may matter in a tight loop or a stream of tiny operations, but not in the same way for large, infrequent work.
  • Can the work be amortized? Compare per-item calls with buffering or batching, while accounting for latency, memory, and CPU use.
  • Does the compiler or runtime change the path? Inlining and other optimizations can make source-level boundaries differ from executed work.
  • Which platform and configuration are involved? Hardware, architecture, kernel, library, compiler, and options such as polling can alter the implementation or trade-offs.
  • Does the measurement represent the real workload? Measure the latency and resource use that matter for your application, rather than treating one isolated call time as a complete answer.

A benchmark result describes the tested machine and conditions, not an abstraction’s universal cost. Chris’s nanosecond figures are useful for showing that different paths can have different measurable costs, but they do not establish what another system will observe. Likewise, a mechanism such as sendfile() or io_uring suggests a path worth evaluating; it does not substitute for workload-specific measurement.

When to keep the abstraction—and when to investigate

Abstraction earns its place when it makes code easier to reason about, change, and reuse. Its cost is not, by itself, a reason to discard it. Investigate when profiling or a concrete workload points to a hot path, excessive calls, avoidable copying, or a resource trade-off that matters. Then compare implementations under the same realistic conditions and preserve the clearer abstraction unless the measured benefit justifies a more specialized path.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.