Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

In Yurii Tor’s reported AI coding experiment, every displayed configuration passed the original acceptance suite, yet a later audit found an expired-lease bug in both Astra + Luna runs. The key review question is whether a time-sensitive decision happens before or after a transaction wait: a timestamp captured before SQLite grants a write lock can be stale by the time code uses it.

What the benchmark tested—and what it found

Tor’s task was to build a durable TypeScript/SQLite reminder queue that survives restarts, retries failed deliveries, and handles competing workers. The author compared four configurations, but the reported table highlights Astra running solo and Astra + Luna. Each displayed configuration had two runs.

The original acceptance result and the later diagnostic audit measure different things. The acceptance suite passed in all four displayed runs; the subsequent audit exposed a defect in the paired configuration. The figures below are reported by Tor in 2026 and are not independently verified here.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Measure Astra solo Astra + Luna
Original acceptance 2/2 runs passed 2/2 runs passed
Later diagnostic checks 7/7 in each of 2 runs 4/7 in each of 2 runs
Mean fixed-rate estimate 38.681850 units 20.384872 units
Mean elapsed time 543.302 seconds 795.081 seconds

In these runs, Tor calculated that Astra + Luna had a 47.3% lower fixed-rate estimate and took 46.3% longer than Astra solo. Astra’s planning and review accounted for 95.6% of the paired workflow’s fixed-rate estimate.

Those estimated cost units were derived from token counts multiplied by fixed historical rates. They are not a bill, a measured subscription deduction, or a demonstrated subscription-quota saving. The experiment covers one task and two runs per setup, has no Sol-only control, and changed CLI version and executor-selection protocol before the Astra + Luna runs. The diagnostic audit was retrospective, and three of its seven checks probe the same clock-after-lock defect. These results therefore do not establish a general model ranking or general economics of orchestration.

How a SQLite lock wait can expire a lease

A lease typically records an owner token and an expiration time. A worker uses those values to decide whether it may claim work or whether it still owns work it is completing or failing. SQLite permits a writer to wait for another transaction to release its write lock. If application code reads the current time before requesting that lock, the captured value can age while the transaction waits.

Rank #2

For example, code might sample time, attempt to begin a write transaction, wait behind another writer, and then use the old timestamp to make a lease decision. The comparison is performed later, but it reflects the earlier moment. A claim may therefore be given an expiration based on stale time, or an ownership check may accept a lease that has already expired.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

This is the specific mechanism the retrospective audit identified; it does not show that every lease or timing bug has the same cause.

Order the transaction and clock read deliberately

For this failure mode, acquire the write transaction first and read the clock only after SQLite has granted the write lock. Keep the time comparison and the related state update in that same transaction. Preserve owner-token checks, and make sure the implementation follows the queue’s intended clock semantics.

  1. Begin the write transaction and wait until the write lock is held.
  2. Read the current time after acquiring the lock.
  3. Evaluate expiry and ownership using that post-wait time and the expected owner token.
  4. Apply the claim, completion, or failure state change in the same transaction.

Moving the clock read after lock acquisition prevents this particular pre-wait timestamp from aging unnoticed. It is not, by itself, proof that all concurrency, clock, or lease-lifecycle problems are solved.

Make the lock-wait regression test deterministic

A useful regression test must force the ordering that causes the defect instead of hoping normal test timing happens to reproduce it. Tor’s outline uses two independent SQLite connections, a barrier, and an injected clock:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  1. On connection A, hold an immediate write transaction behind a test-controlled barrier.
  2. Start a claim operation on connection B and confirm that it has reached the write-lock boundary and is waiting.
  3. While B is blocked, advance the injected clock beyond the relevant deadline.
  4. Release A so B can acquire the lock and continue.
  5. Check claim, completion, and failure behavior against the post-wait time and the expected owner token.

In the audit’s example, the clock advanced from 0 to 10 while the claim waited, and the lease duration was 5. A fresh claim should therefore expire at 15. The audit reports that both Astra + Luna runs instead returned a claim ending at 5. It also reports that both accepted expired ownership at time 5. These are distinct checks: reclaiming a claim and rejecting an expired owner during completion or failure are not interchangeable behaviors.

Real sleeps alone are a weak substitute for the barrier and controllable clock in this scenario. A sleep does not prove that the second connection reached the lock boundary before the clock changed, so the test can pass or fail according to scheduling luck.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Read the benchmark results with the right boundaries

The result is a caution about what an acceptance suite can miss, not proof that a particular model or workflow is broadly better. The original passing results remain the original results; the later audit adds diagnostic evidence and should not be presented as if those checks were part of the initial acceptance suite.

  • Acceptance is not exhaustive: passing the reported suite did not expose the lock-wait timing defect.
  • The audit is narrow: it tested one reminder-queue task, with two runs per displayed setup, and multiple checks overlap on one defect mechanism.
  • The comparison protocol was not fully consistent: the CLI version and executor-selection protocol changed before the paired runs.
  • The cost and time values describe these runs: they should not be generalized to other tasks, clients, or subscription usage.

Tor’s article links to an ORCH-1 report, methodology, and run-level results, but the measurements above are attributed to the author’s article rather than independently verified underlying data. For a stronger comparison, the author suggests adding a Sol-only control under the same client and protocol, testing more tasks, and freezing expanded diagnostic checks before candidate runs.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.