Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

In the author’s account published July 15, 2026, the benchmark stopped at N=22 because Python tried to convert a very large integer to a string and hit CPython’s configured integer-string digit limit. Removing a conversion the tool did not need restored the N=24 result. The apparent cutoff was only the first of nine problems: issues in parsing, precision, run identity, and chart interpretation also obscured or distorted results. After the reported fixes, the author says the sweep produced 96 of 96 datapoints. These are the author’s findings, not independently reproduced measurements. Read the author’s account.

Why the benchmark stopped at N=22

The benchmark compared four agents implemented in Python, Go, Node.js, and Rust while computing Mersenne primes with the Lucas–Lehmer test. Its harness swept N=1–24, but the earlier script and chart ended at N=22. According to the author, the Python agent failed at the larger input because it attempted an integer-to-string conversion that exceeded CPython’s configured limit.

At N=24, the reported error said that integer-to-string conversion exceeded the 4,300-digit limit. The code converted each prime to a string even though the tool returned only elapsed time; removing that unnecessary conversion restored the datapoint. The author reports a Python N=24 result of 2,425.9 ms and says the 24th Mersenne prime, 219937−1, has 6,002 digits. Both figures are from the author’s 2026 account, not an independent reproduction. The account describes the failure and fix.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The larger lesson is that a benchmark’s stopping point is not necessarily a meaningful algorithmic boundary. A failed operation can be hidden when someone narrows the test range instead of following the error to its source. As the author puts it: “The workaround you commit is the bug you keep.”

Eight more bugs distorted the measurements

Fixing the integer conversion explained the cutoff, but it did not make the benchmark trustworthy by itself. The author’s account describes additional failures in the harness and in how results were compared.

1. A workaround concealed a runtime error

Shortening the sweep to N=22 made the script appear to have a deliberate limit, while leaving the failed conversion unexplained. The author also notes that Go performed an unnecessary string conversion inside its timed region. Work that does not belong to the operation under measurement can both create failures and inflate timing.

2. Timing depended on model-generated prose

The harness’s main parser looked for a particular phrase in Gemini’s prose. That is fragile: a wording change can break extraction even when a valid measurement exists. The author says Python results were available only because a fallback read the structured tool artifact. Machine-consumed timing should come from structured output, not from phrases an agent may paraphrase.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

3. Reported prime counts were not validated

Node and Rust could report finding 100 primes even though their exponent tables contained 26. A plausible-looking count is not evidence that the output is complete or correct. The harness should check results against the expected input set and independently validate the count.

4. Rounding erased the fastest timings

Rust measurements were formatted to two decimal places in milliseconds. Very fast results consequently became 0.00 ms, which could not be plotted on a logarithmic chart. Keep raw timing values at adequate precision and round only for display; do not convert a small nonzero result into zero.

5. The duration parser missed nanoseconds

After formatting work was removed, Go emitted nanosecond durations for small inputs. The parser recognized microseconds, milliseconds, and seconds, but not nanoseconds. A parser must handle every unit the runtime can produce, or normalize durations to a single unit before parsing.

6. Reused session IDs carried history into reruns

Deterministic context IDs let ADK retain prior conversation state. During a rerun, Gemini responded: “I already did that. Do you want to do it again?” The author says unique IDs for each run fixed the missing datapoints. When an agent framework preserves session context, each independent benchmark run needs an independent context.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

7. The chart implied a language-only comparison

Python and Go used Gemini tool calling through ADK, while Node.js and Rust used direct HTTP handlers. The chart therefore mixed distinct execution paths while presenting results as a language comparison. The author reports median round-trip times of 2.6 ms and 4.6 ms for the direct agents, versus approximately 1.6 s and 1.8 s for the Gemini-routed agents. These numbers reflect different architectures as well as different implementations; they do not establish that one language is inherently faster.

8. End-to-end latency was easy to confuse with computation time

A tool-routed request includes work beyond the Lucas–Lehmer calculation, such as the model-mediated call path. A direct handler and a Gemini-routed agent are not measuring the same thing merely because both eventually compute the same prime. Name the measured quantity—calculation time, handler time, or end-to-end round-trip latency—and keep comparisons within the same execution path.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

How to debug a benchmark that stops unexpectedly

  1. Find the first failing operation. Reproduce the last successful input and the first failing one, then inspect the error or traceback. Do not assume the final successful N is an intentional benchmark boundary.
  2. Remove work outside the measurement. If a tool returns elapsed time, do not stringify a huge result unless the string itself is part of the task. Keep setup, formatting, and logging out of the timed region unless the benchmark explicitly measures them.
  3. Make results structured and checkable. Return numeric timing fields and explicit status or count fields. Validate expected result counts and reject missing, malformed, or inconsistent records instead of silently plotting them.
  4. Normalize duration units without losing precision. Convert all supported units—including nanoseconds—to a common numeric unit. Preserve raw values; round only when rendering labels or tables.
  5. Isolate each run’s state. Use a fresh context identifier when the framework can retain conversation history. Confirm that a rerun actually executes rather than reusing a previous answer.
  6. Audit the chart against the data. Check that every expected N appears, no zero or missing values are silently substituted, and logarithmic axes receive positive values. Label the architecture and the measured interval so readers can tell what the plotted numbers mean.

What the final 96/96 result does—and does not—show

The author says the corrected sweep returned 96/96 datapoints. That indicates the reported run completed across the intended set after the fixes; it does not, by itself, prove that the implementations are equivalent or that the measurements isolate programming-language performance. For a fair comparison, hold the execution path and measured interval constant, preserve input and output validation, and show missing or failed runs rather than hiding them.

The account is a first-person description published July 15, 2026. The stated digit limit, timings, and completion count are attributed to that account and should be read as reported results, not independently verified measurements. See the full debugging account.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.