Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

A row-ID retention cap can make a scheduled data pipeline forget rows that are still present in its source. When that happens, a later poll may treat previously delivered rows as new, send them again, and charge for them again. FetchSmith reported finding this failure in its watch-mode Actors; the safeguards it described made truncation visible, but did not prevent duplicate delivery.

How a bounded baseline can lead to duplicate delivery

In the account published by FetchSmith on September 23, 2026, watch-mode Actors poll a source and deliver rows considered new since an earlier run. Each Actor keeps IDs for previously delivered rows in a key-value store. A configured limit—examples in the post include 5,000, 20,000, or 60,000 IDs—keeps that store from growing indefinitely.

The failure comes from what the cap means for correctness. FetchSmith says its eviction logic removed the oldest IDs when the baseline exceeded its limit, without signaling that IDs had been discarded. If a run returned more IDs than the baseline could retain, an older row could remain available from the upstream source after its ID had been evicted. On a later poll, the Actor no longer recognized that row and could deliver it again. For a service charging per delivered row, that can mean charging for a row a buyer already received.

This is not simply a storage housekeeping issue: the retained baseline is part of the pipeline’s memory of what it has already delivered. Its capacity has to be considered against poll volume and how long old rows remain visible upstream.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
#1 Best Overall
Sale
Data Recovery Stick for Windows Data Recovery Software – Photos, Files
  • The Data Recovery Stick requires no technical skills — simply plug it into your Windows computer, click Start, and the software automatically begins scanning and recovering lost files within minutes. Compatible with Windows Vista, 7, 8, 10, & 11, it's designed to be a reliable first step when accidental deletion occurs.
  • Recover photos (JPG, BMP, PNG, TIFF), Microsoft Office documents (Word, Excel, PowerPoint, Publisher, Access), Open Office files, MP3 music files, PDFs, RTF documents, AutoCAD files, and HTML web pages. Whether it's personal memories or critical business files, the Data Recovery Stick covers the file types that matter most.
  • Works with hard drives, USB drives, SD cards, memory sticks, and other common storage formats that use FAT or NTFS file systems — making it a single solution for hard drive recovery, USB drive recovery, SD card recovery, and more. Note: a media reader is required for micro SD cards and some mass storage devices.
  • No Installation Required - The Data Recovery Stick runs entirely from the USB drive with no software installation on your computer — helping prevent new data from overwriting the files you're trying to recover. This also makes it ideal for use across multiple computers or in emergency situations where installation isn't practical.
  • Use the Data Recovery Stick on as many computers as often as needed — simply clear the recovered data between uses to free up storage space. Software updates keep the tool compatible with newer systems and devices, backed by 25+ years of data software expertise from Paraben Consumer Software.

What FetchSmith reported in its reproduction

FetchSmith described a reproduction in its google-play-reviews-scraper Actor with WATCH_KEEP = 20000. In the reported seed run, the Actor recorded 1,000 IDs and dropped 997 by the cap. In the following incremental run, it delivered 40 rows and skipped zero as already seen. FetchSmith described the zero-skipped result as the signature of a broken watch in this setup, and said those 40 rows had already been paid for in the seed run. These are the publisher’s reported figures, not an independently audited test.

FetchSmith said it reproduced the same failure pattern on 22 Actors, including ones for Google Play reviews, court records, clinical trials, public tenders, FDA recalls, grants, and job listings—sources where a poll could plausibly return more IDs than the configured baseline retained. The post does not quantify actual duplicate charges, affected buyers, or refunds, so the reproduction should not be read as a financial-impact total.

Why successful runs did not reveal the problem

The reported runs could still finish with a SUCCEEDED status and a plausible row count. From the Actor’s perspective, a row whose ID is missing from its stored baseline looks new; the run can therefore complete normally even though the baseline lost information that mattered to billing correctness. FetchSmith says the issue emerged in later billing rather than in the run that evicted the IDs.

An unexpectedly high delivered-row count combined with zero rows skipped as previously seen is a useful incident clue when an incremental poll should overlap substantially with the prior result. It is not a universal test: a source may genuinely have many new rows, or its behavior may produce little overlap. Treat the pattern as a prompt to inspect truncation and baseline history, not proof by itself.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

What the safeguards changed—and what they did not

FetchSmith says it added signals intended to expose baseline truncation to both operators and downstream systems:

  • A warning when saving after IDs have been evicted.
  • A note in the run status.
  • Last-run and cumulative truncation counts in the watch record.
  • baselineTruncated and baselineTruncatedTotal fields in RUN_SUMMARY and in any configured webhook payload.
  • README documentation explaining the cap.

These changes improve visibility and measurement; they do not stop an evicted ID from being forgotten or prevent a later re-delivery. The practical claim is that truncation became observable, not that duplicate delivery or billing was eliminated.

A related Apify UK tender Actor listing gives a concrete implementation example: it documents a 60,000-ID watch-record cap, warns that an ID that falls off can be returned and charged again, and describes warnings, status messages, stored counters, and run-summary and webhook fields. The listing recommends narrowing watches when truncation occurs. It is a related product listing, not independent verification of FetchSmith’s fleet-wide report, and its live limits and product details may change.

How to think about prevention and capacity

Detection and prevention are separate design goals. A warning lets an operator investigate; it does not restore discarded IDs. FetchSmith’s post mentions unbounded retention and capacity tied more intelligently to poll volume as possible design directions, but does not benchmark them or identify a universally best choice.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

When evaluating a baseline design, consider the source’s largest plausible result in one poll, the configured retention capacity, and how long previously delivered rows can remain visible upstream. Also check whether a truncation signal reaches only a log or run status, or is carried into counters and webhooks that downstream monitoring can act on. Any cap should be treated as a correctness boundary as well as a storage setting.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Why API retry idempotency is related but different

Stripe’s official API documentation says, “The API supports idempotency for safely retrying requests without accidentally performing the same operation twice.” An idempotency key can make eligible repeated API requests return the original result, subject to Stripe’s key-retention rules and matching request parameters. Stripe’s idempotent requests documentation says keys may be pruned after at least 24 hours and that parameters are checked for a match; confirm the current reference before relying on those details in an implementation.

That mechanism addresses retries of a particular API operation. The FetchSmith incident concerns a different layer: an application-maintained set of row IDs that no longer remembers an item after eviction. Idempotency at a payment or delivery API can help guard repeated requests at that API, but it does not by itself repair a pipeline’s incomplete deduplication baseline. Stripe is not reported as involved in FetchSmith’s incident.

Sources

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.