Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

In one reported test, adding a two-line cache change reduced a Python script’s runtime from 3.71 seconds to 1.20 seconds. The gain came from parsing repeated date strings—not from a faster interpreter or CSV reader—and it is not a general promise that caching makes Python scripts three times faster.

The useful lesson is to find the work consuming time, check whether its inputs repeat, and measure a focused change against the same workload.

What made this Python script slow?

The example program generated and processed one million sales rows. It parsed each row’s date, aggregated revenue by month and region, then wrote a text report. In the author’s profile, strptime was the leading self-time function: one million calls accounted for 3.045 seconds of self time in an 8.440-second profiled run.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

That profile helped identify date parsing as a hot spot, but its runtime is not comparable to the unprofiled 3.71-second baseline. Profilers add overhead; the Python documentation says they are designed to provide an execution profile, not for benchmarking. It recommends timeit for reasonably accurate timing and says cProfile is suitable for most users. See The Python Profilers.

Why did caching speed up the example?

Although the program processed one million rows, the dates were not all different. The author reports just 365 distinct date strings in the input. That meant the parser could perform the conversion once per distinct string and reuse the result for later occurrences.

The reported cache statistics were 365 misses and 999,635 hits. A hit returns the saved result for an argument already seen; a miss requires the function to calculate and store a result. The high hit count explains why this workload benefited.

What changed in the code?

The change imports lru_cache and decorates the existing parser:

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
from functools import lru_cache

@lru_cache(maxsize=None)
def parse_date(s):
    return datetime.strptime(s, "%Y-%m-%d %H:%M:%S")

This is memoization: when the function receives the same argument again, it can return the cached result instead of parsing the string again. The example uses maxsize=None, which allows the cache to grow without a size limit. Python’s documentation also notes that cache arguments must be hashable and that cache_info() reports hits, misses, maximum size, and current size. See functools.lru_cache.

Unbounded caching can use increasing amounts of memory as new inputs arrive. Check the cache’s size and the lifetime of its entries in your application; an unbounded cache is not automatically appropriate for a long-running process.

How much faster was it, and when did caching help?

The article reports two timing sets, taken with different scopes. Its headline comparison is a single in-script run; the table below gives medians from five whole-process runs per version. Treat the table’s plain and cached values as comparisons within each row, not as replacements for or direct comparisons with the headline timings.

Distinct date strings Plain median Cached median Reported speedup Output
365 3.96 s 1.43 s 2.77× Same
20,000 3.70 s 1.49 s 2.48× Same
1,000,000 3.76 s 3.98 s 0.95× Same

These are the author’s results for a synthetic workload on one Mac mini M4 Pro with 48 GB of memory, using Python 3.14.6. The table’s figures are five-run medians and whole-process timings. In the all-unique case, caching was about 6% slower, because there were no repeated inputs to reuse and the cache still incurred overhead. The author also reports that the headline single-run output files were byte-for-byte equal.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

How to test whether caching will help your script

  1. Profile to locate a hot spot. Use a profiler such as cProfile to see which functions consume time and how often they run. Use the profile to guide investigation, not to compare benchmark runtimes.
  2. Check whether the inputs repeat. Count distinct arguments as well as total calls. A costly function called frequently may still be a poor cache candidate if nearly every argument is unique.
  3. Confirm the function is safe to cache. Caching is appropriate only when returning a prior result for the same input preserves the program’s intended behavior. The function should not rely on side effects or need to produce a fresh mutable object on every call. Arguments must be hashable.
  4. Make one focused change. Add the cache to the candidate function, then inspect cache_info() and, where relevant, current cache size.
  5. Benchmark under matching conditions. Compare unprofiled runs of the original and changed versions with the same input and timing scope. Use repeated measurements rather than drawing a conclusion from one run; Python’s documentation points to timeit for reasonably accurate timing.
  6. Verify correctness and resource use. Compare outputs and check whether the cache’s memory growth is acceptable for the workload and process lifetime.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

When should you avoid an unbounded cache?

A cache is less likely to help when inputs rarely repeat, the function is inexpensive, or lookup and memory costs outweigh saved computation. The million-distinct-date case demonstrates that trade-off: the cached version was slower in the reported five-run medians.

Prefer a bounded cache when retaining every distinct input could consume too much memory, or skip caching when reuse is too low to justify it. Choose based on measurements from your own input pattern rather than assuming that an expensive-looking function will benefit.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.