The Tool Desk
Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more
Geoff Cox says he used a required CI gate, human review and production monitoring while AI coding agents wrote much of TopSet, his live-trading system. His experience shows how to make failures harder to merge—and why a green build still cannot prove that trading code is correct or behaving as intended.
Cox’s account is a 2025 retrospective about a one-person project that he says trades with his own capital. It is not an independent audit, a controlled study or proof that this particular test suite makes a trading system safe. Its most useful lesson is narrower: automated checks can catch many regressions, but they can also pass when tests share a false assumption, invalid inputs produce plausible-looking output, or the changed code never runs.
What the merge gate checked
Cox describes a required pull-request gate: a failed check blocked a merge. The per-PR suite took about ten minutes, according to his account. He reports roughly 7,500 tests across 331 modules, as well as about 1,600 pull requests over 21 months and 161 report scripts. These are figures he reported, not independently verified measurements.
| When it ran | Checks Cox says were included | What the check was meant to catch |
|---|---|---|
| Every pull request | Ruff linting and mypy type checks; about 7,500 tests across 331 modules; a check that the test database was empty afterward; migration rollback and reapplication from the base; Terraform validation against two AWS accounts. | Code and type errors, behavioral regressions, test data left behind, broken migration paths and infrastructure configuration errors. |
| Weekly scheduled suite | Two full rebalancing end-to-end tests against a mock broker; one integration test repeated 100 times; nine deterministic model-training snapshots pinned in Docker; two training-pipeline leakage checks. | Workflow failures, intermittent concurrency bugs, changes to model outputs and data leakage. |
In the repeated integration test, the mock broker varied fill timing, partial completion and prices. That variation mattered because two runs of the same workflow could encounter different event orderings. The aim was to expose failures that a single deterministic pass might miss, not to claim that 100 repetitions could cover every possible interleaving.
#1 Best Overall
Test the risky workflows, not only the individual functions
The money-related end-to-end scenarios included resubmitted orders; restarting a rebalance with partial fills or with buys and sells still outstanding; cancelling smart orders mid-flight; and deposits or repeated withdrawals arriving during a rebalance. These cases exercise interactions among persisted state, retries, timing and order execution—the places where individually reasonable code paths can produce a broken workflow when combined.
How green tests gave the wrong answer
Cox’s examples are a reminder that tests establish agreement with their assertions, not truth about the production system. As he puts it, “Tests prove that the code does what the tests say.”
A fixture repeated an incorrect data convention
The code treated a broker’s stock-dividend rate as a fraction such as 0.05. Cox says the records he inspected represented new shares divided by old shares instead. In the data he examined, all 101 records across 43 symbols had rates of at least 1.0016. A test fixture repeated the mistaken convention, so the test and implementation agreed with each other while disagreeing with the data.
Quick wins for a faster PC:
Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Clear out junk files and repair common Windows errorsFree Scan →Rank #2
Cox reports that this error deflated a nine-event price history by about 490 times and created an apparent rise of more than 1,000-fold. Those figures describe his examined data and calculation; they are not general statistics about stock dividends or market data.
NaN values made two model selectors look identical
An alternate model selector appeared to agree 100% with the existing selector. The apparent match was misleading: missing input columns made all the scores NaN, and both selectors fell through to the same fixed tie-break. Cox says the fix made missing or invalid scoring inputs a hard error and pinned the test to a version that included the required columns. A perfect match is not meaningful if the comparison never received valid scores.
The changed training target was not being used
A new training target produced exactly the same picks as the baseline. Cox found that the runner routed to another function that ignored the new field. Before the path was fixed, he reports 238 identical picks out of 238; afterward, he reports 0 identical picks out of 238. The regression test then required the modes to differ.
Rank #3
Different results alone do not establish that a new model is better. In this case, the test addressed a more basic question: did the changed setting reach code that used it? Cox’s warning is apt: “Identical results are not a pass. They’re a smell.”
Do these 3 things before closing this tab:
1Fix the driver behind crashes, sound loss and screen glitches2Repair Windows errors before they cause bigger problems3Scan for outdated or missing drivers - takes under a minuteA zero-value order triggered a retry loop
A pending buy, adjusted for withdrawals, could reach zero. A guard rejected negative values but allowed zero through, and an affordability check accepted 0 as affordable because 0 is greater than or equal to 0. The database then rejected the order under an amount-greater-than-zero constraint.
The rollback caused a second problem: it discarded the completing order’s status update along with the rejected transaction. The scheduler therefore retried with the same inputs. Cox says the first proposed correction guarded the promotion step but missed a mutation to a live ORM object that could be persisted during autoflush. He reproduced the failure and reran the proposed fix against that reproduction. The episode shows why a database constraint, transaction boundaries and retry behavior need to be tested together.
Rank #4
How to tell whether a change really worked
A build can be green while the result is wrong, meaningless or unchanged for the wrong reason. Cox’s account supports treating test output as evidence to inspect, not a substitute for checking the behavior the change was supposed to produce.
- Verify the inputs. Treat fixtures as claims about production data. Compare sample values with source records, especially when a field’s units or convention determine a calculation.
- Reject invalid states loudly. Missing columns, non-finite scores, impossible order amounts and other unusable values should raise clear errors rather than flow into a default or tie-break that looks valid.
- Ask what should differ. For an experiment or new configuration, define an expected observable effect. If outputs are identical, trace whether the changed value reaches the intended function and whether valid inputs were present.
- Check results independently. Export calculations and inspect the arithmetic; spot-check the inputs against raw records rather than relying only on a test that may share the implementation’s assumptions. Cox contrasts a test assertion with a spreadsheet’s ability to ask whether a number is true.
- Reproduce the failure. When a bug involves retries, persisted state or timing, first make the failure repeatable, then run the proposed fix against that case. Review whether transactions and in-memory objects can still cause side effects.
What remains a human responsibility
Automation can enforce checks, but people still decide what the checks mean and whether a change is acceptable. Cox says he reviewed model-training and execution approaches deeply while relying on tests for line-level behavior. That division depends on tests that actually represent the domain, which is exactly what the fixture and NaN failures put in doubt.
- Choose the architecture, define expected behavior and decide which workflows deserve end-to-end coverage.
- Review outputs and data assumptions as well as the code diff; a plausible-looking change can be a no-op or a calculation based on the wrong convention.
- Keep ownership with the person approving the change. Cox’s formulation is: “The author owns the change, whoever typed it.”
- Strengthen the gate before increasing the volume of agent-written changes. Track escaped defects and hotfix rates rather than treating lines of code or pull-request counts as quality measures.
Cox also recalls that automated pull-request testing at a healthcare company reduced active medium- and high-severity bugs by 72% and weekly hotfixes from seven to 1.5. He provides no underlying study, measurement method or company name for that recollection, so it should not be read as a general estimate of what CI will achieve elsewhere.
Best Value
Why monitoring still matters after a merge
Pre-merge tests cannot establish that a live system is behaving correctly under every production condition. Cox says TopSet used CloudWatch production ERROR alarms sent to Discord. He investigated incidents using logs and a local copy of the database. He explicitly notes that tests did not catch the first failures; they helped prevent those failures from recurring.
That distinction matters: CI is a barrier against known classes of mistakes, while production monitoring helps reveal behavior the test suite did not anticipate. A useful gate therefore needs an incident feedback loop—investigate what happened, reproduce it where possible, and add a check that would make the same failure harder to reintroduce.
What this account can—and cannot—show
Cox describes a single-person system operating with his own capital and cautions that it is not the same problem as a team of fifteen. His experience illustrates concrete engineering practices and failure modes; it does not establish that the gate guarantees safety, that his reported figures generalize, or that agent-written trading software is safe simply because tests pass. The practical standard is more demanding: require broad checks before merge, verify that those checks encode reality, review the behavior a change produces, and keep watching after deployment.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

