Free tools Windows power users keep installed
One-click scans. No signup required.
iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more
A nightly job that rebuilds a GitHub-backed alternatives directory can delete valid rows without failing once. The cleanup step treats “not in today’s result” as “no longer valid,” so any gap in the result becomes data loss. Three failure modes create that gap. Each has a fix that makes the cleanup defer instead of guess, and the rule behind all three is simple: an absence can authorize deletion only when the input that produced it was complete and plausible.
The rule behind all three fixes
Deleting a row because a fetch did not return it is safe only when the fetch was complete and the resulting set makes sense. Each failure below breaks one of those two conditions. An empty successful response and a failed request look similar in code, but they mean different things. The first is an answer. The second is the absence of an answer. A catch block that returns an empty list turns the second into the first, and the cleanup then acts on it.
1. A failed fetch builds an incomplete keep-list
A typical refresh loop does three things for each SaaS entry. It fetches the GitHub alternatives listed for that entry, collects the repository names that came back into a keep list, and then deletes stored rows whose names are missing from that list. The delete is usually a single statement shaped like this (illustrative, not copied from any specific codebase):
The Tool Desk
Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →DELETE FROM alternatives
WHERE saas_slug = $1
AND repo_full_name NOT IN ($2, $3, ...);
If GitHub returns a 403 or 429 for one alternative, that repository is missing from the keep list even though the seed file still lists it. The statement then removes a valid row. The next run may succeed, and the row comes back. That is why the symptom resembles flickering: present one night, gone the next, back after a clean retry. It is easy to misdiagnose as a problem with the data source.
#1 Best Overall
The reported fix
- Count failed fetches for each SaaS slug during the loop.
- If the count is greater than zero, skip the stale-row delete for that slug and leave its rows untouched.
- If every fetch for the slug succeeded, build the keep list from the successful results and run the delete.
- Log the slug, the failure count, and the fact that the prune was skipped.
This trades a little staleness for safety. A row that should have been removed survives one more cycle, which is a cheap error. A valid row removed on incomplete evidence is an expensive one, because nothing fails and nothing alerts. The engineer who documented this fix put it plainly: “A stale row persisting an extra night is much cheaper than a valid row vanishing without an error message.”
Making the skip visible
A deferred prune is safe only if someone can see it. Otherwise, deferred rows quietly become old data. The following additions are suggestions from this guide rather than part of the reported fix:
- Print the skip count for each run in the job summary, not only as a debug line.
- Alert when the same slug is skipped on several consecutive runs. The threshold is a judgment call; pick one that matches how quickly you expect GitHub to recover.
- Store a last-confirmed timestamp on each row, so staleness can be queried directly instead of inferred.
2. A seed spelling is not the repository’s identity
Seed files are written by people. A repository may be listed with different capitalization, an outdated owner name, or a slug that has since changed. GitHub returns its own identity in the repository response. If the keep list is built from the seed spelling, the fetched repository may not match the seed entry. The valid row is then left out of the keep set and pruned.
Do these 3 things before closing this tab:
1Clear out junk files and repair common Windows errors2Fix the driver behind crashes, sound loss and screen glitches3Repair Windows errors before they cause bigger problemsRank #2
The reported fix compares against the canonical full_name field from the repository response. That field is already in the response alongside the other repository details, so it adds no request per repository.
| Comparison key | Where it comes from | Risk in the keep-list and delete comparison | Extra requests |
|---|---|---|---|
| Seed spelling | Seed file, typed by a person | Mismatch with GitHub’s identity can cause a valid row to be pruned | None |
Canonical full_name |
Repository response from GitHub | Matches the object GitHub actually returned | None; the field is already in the response |
Use that one key everywhere: when you upsert, when you build the keep set, and when you compare for deletion. Normalizing in one place and using the raw seed string in another brings the bug back.
3. A truncated seed makes valid rows look stale
A second cleanup pass compares the stored slugs for each SaaS with the current seed file and removes rows that no longer appear in it. If a merge conflict or an editing mistake truncates the seed, a large share of stored rows suddenly looks stale. The cleanup cannot distinguish a deliberate removal from a damaged file.
Rank #3
The ratio guard
The reported safeguard is a circuit breaker. It counts the apparent stale rows and compares that count with a limit. If the count is above the limit, the prune is skipped. The limit is the larger of two values: 10% of the stored rows being compared, and a floor of three rows.
| Stored rows for the SaaS | 10% of stored rows | Limit applied (larger of 10% or 3) | Prune is skipped when apparent stale count is |
|---|---|---|---|
| 12 | 1.2 | 3 | more than 3 |
| 200 | 20 | 20 | more than 20 |
| 500 | 50 | 50 | more than 50 |
The 10% ratio and three-row floor are the example configuration reported by the engineer. They are implementation choices, not an established standard. Set your own values based on how often your seed legitimately shrinks.
When the breaker trips
A tripped breaker is a question, not a verdict. The guard catches implausible input, but it cannot prove the seed is correct. A reasonable response is to work through these checks in order:
- Compare the seed’s line count and the stored slug count with the previous run and with the last commit that touched the seed.
- If the file is truncated or the merge is damaged, restore the file and rerun. The breaker should then pass.
- If the removals are real, such as a deliberate purge of retired tools, run a manual prune against a reviewed list of slugs, or raise the threshold for that single run.
A legitimate large cleanup needs an explicit route. Without one, the breaker will keep blocking valid work, and people will be tempted to disable it.
Why event feeds and incremental syncs need the same caution
Teams often try to avoid full refreshes by reading change events. That path has its own limits, and it is where the same mistake returns.
An event feed is not a complete ledger
GitHub’s Events API documentation states that public events are limited to the most recent 30 days and at most 300 events. It also describes event latency of 30 seconds to six hours depending on time of day, and says the API is not intended for real-time use. A bounded, delayed feed can tell you that something changed. It cannot prove that a record is absent. Use events to decide what to re-fetch, and use a complete fetch before deleting anything.
Best Value
Event type matters too. GitHub’s webhook documentation pairs label changes with labeled and unlabeled actions, and milestone changes with milestoned and demilestoned. Its REST issue-events documentation lists event types such as unlabeled and head_ref_deleted. Choosing the right event lets you react to a specific change; it does not turn the feed into a snapshot.
Cursor boundaries can drop or duplicate a record
For an incremental sync that uses a timestamp cursor, Airbyte’s GitHub source documentation describes a boundary problem. The GitHub since filter is inclusive, so a record whose timestamp equals the saved cursor is returned again. If a connector then keeps only records strictly newer than the cursor, that boundary record can be dropped. According to Airbyte’s documentation, version 2.4.0 keeps the boundary record on the affected streams instead. The trade-off is one extra row per repository in append-only destinations. Destinations that append and then deduplicate collapse that extra row on the primary key.
| Boundary approach | Record at the cursor timestamp | Trade-off |
|---|---|---|
| Strictly newer than the saved cursor | Can be dropped | An update that shares the cursor timestamp may never reach the destination |
| Inclusive boundary, deduplicated on the primary key | Re-emitted and kept | One extra row per repository in append-only destinations, unless the destination deduplicates |
Set the high-water mark from the cursor values the source returns, not from the worker’s wall clock. A general incremental-sync design recommendation holds that clock skew between the worker and the source can silently skip rows. That advice is not specific to GitHub’s API, but it applies to any timestamp cursor you store.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Choosing a cleanup policy
There are two policies to compare: prune on every run, or prune only after a complete and plausible observation. The table below uses the same criteria for both.
| Criterion | Prune on every run | Prune only after a complete, plausible observation |
|---|---|---|
| Data-loss risk from partial input | High: an incomplete fetch deletes valid rows | Low: an incomplete scope is skipped |
| Stale-row duration | Shortest when the input is correct | Longer for a skipped scope, usually by one cycle |
| Operational visibility | Failures can stay hidden behind a successful job | Requires logging and alerting on every skip |
| Recovery cost after a bad run | Often needs restoring deleted rows from backup or re-seeding | Usually means clearing a backlog of stale rows after the scope recovers |
For most directories that rebuild from a seed file, the second policy is the safer default. Its cost is a few extra stale rows during a bad night. The first policy’s cost is a valid row that disappears with no error to investigate.
Evidence limits
The three fixes, the breaker values, and the detection-lag anecdote come from one engineer’s account of an implementation they ran. The author reports their effects; no independent test has run the code, queried the live GitHub API, or measured how often these failures occur. The 10% ratio and three-row floor are configuration choices, not benchmarks. The GitHub and Airbyte statements above come from their respective documentation, and the incremental-sync design advice is a general recommendation rather than GitHub-specific behavior.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

