iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more
FORGE looked like a working AI-agent product on day one because its dashboard was already full: a workflow canvas, event stream, replay, counters and tools. But those signals came from simulated agents and sample data, not a system doing the work the interface implied. The build account’s central lesson is practical: every visible status, metric and answer needs a real event behind it—or an unmistakable simulation label.
What made the first version look complete?
Ted describes FORGE as a home-hosted interface and workflow for agents that plan, research, write and review answers. The first version, however, was a dashboard for a system that did not yet exist. A simulated clock and fake agents supplied plausible activity, making the canvas, event stream, replay view, counters, builder and workflow designer appear populated.
One design decision did survive the transition to real agents: treating each run as an event log. Both the live display and replay view came from that log, and replay could stop at a selected point. That gave the interface a coherent source of truth for showing what happened during a run; it did not, by itself, make simulated events genuine.
Recommended Free Tools
Why was a convincing answer the most serious failure?
The trust problem was not merely that a demo might fail or crash. In one reported case, a simulator ignored the user’s question but returned a polished answer anyway, with no visible indication that the run was simulated. The output looked like the result of research and reasoning even though the underlying workflow had not done that work.
#1 Best Overall
An interface therefore cannot establish that an answer is trustworthy just by displaying progress, counters or a fluent response. Those claims need provenance: a real search event, pages actually read, work completed by the relevant agent, and a clear record of what a reviewer checked. When the system is only demonstrating a workflow, the user should be able to tell before relying on its answer.
What changed when real agents replaced the simulation?
Connecting actual agents exposed operational failures that the simulated activity had hidden. On Ted’s server, requests timed out in an environment without IPv6; some model responses were empty when reasoning used up the output budget; researchers reached their step limit without recording notes; and page fetches could be slow. These are problems from this particular setup, not universal prescriptions for other deployments.
- Connection behavior: Ted reports preferring IPv4 and increasing the connection-attempt time after timeouts on a server without IPv6.
- Empty responses: The implementation retried empty model outputs with more room for output after reasoning consumed the budget.
- Research progress: Agents were told how many rounds remained, addressing cases where researchers hit the step limit without writing notes.
- Slow fetches: Page fetches were capped at 20 seconds, with a host skipped after a timeout.
The broader point is that a simulated path exercises the display of activity, not necessarily the failure modes of the real network, model calls or agent loop. A product can look smooth under simulation and still have basic operational behavior left to discover.
Do these 3 things before closing this tab:
1Scan for outdated or missing drivers - takes under a minute2Clear out junk files and repair common Windows errors3Fix the driver behind crashes, sound loss and screen glitchesHow did FORGE make simulation and stored state more honest?
Label simulation and limit when it runs
During what Ted calls an “honesty pass,” the build turned up invented provider usage, tool-success rates for tools that had never run, and sample run history. The simulator had also answered a real question confidently despite ignoring it. Ted says simulation was then labeled throughout the interface and restricted to an explicit dry-run action. That separation lets a user explore the workflow without mistaking a demonstration for completed work.
Make the server the owner of run history
Initially, FORGE stored data in each browser’s local storage. Ted reports that desktop and laptop state diverged, while independently assigned run IDs allowed one browser to overwrite a run. The reported fix was a server-owned SQLite database, live updates for open tabs, a one-time merge of existing browser data, and run IDs issued by the server.
This change addressed two related trust issues: users could see different histories on different devices, and two clients could claim the same run identifier. In the revised arrangement, the server—not each browser separately—owned persistent run records and their IDs.
Rank #3
How did the workflow distinguish quick answers from checked answers?
FORGE offered three modes with different review and research paths. The author’s account says Quick is marked not fact-checked and Verified is the default. Verified adds a reviewer that checks cited pages; Parallel assigns three researchers through a lead before writing and review. The author also describes agents asking teammates follow-up questions when research notes leave a gap.
| Mode | Reviewer and verification | Research arrangement | Ted’s reported typical cost and duration |
|---|---|---|---|
| Quick | No reviewer is described; explicitly marked not fact-checked | Planner, researcher and writer | About $0.005 and 1–2 minutes |
| Verified | Reviewer checks cited pages; described as the default | Planner, researcher, writer and reviewer | $0.02–$0.04 and 1–4 minutes |
| Parallel | Includes review | A lead assigns three researchers before writing and review | About $0.04 and about five minutes |
These amounts and durations are Ted’s reported typical figures in his September 27, 2026 account, not independent measurements or performance promises. They are not interchangeable with examples he separately gives: one finished Quick run is captioned as three agents taking 1 minute 40 seconds at about a tenth of a cent, while a separate six-agent run is reported at $0.468.
What did source verification and failure monitoring add?
The reviewer’s role was to open cited pages and check claims, rather than treat the presence of citations as proof. Ted says a low source count produces an unverified warning, and the reviewer checks two or three cited pages, reusing pages already fetched. The interface also warns when researchers have read fewer than two pages.
Search reliability needed its own handling. Ted reports a run with seven failed searches; he interpreted the logs as pointing to a short local network outage rather than provider-specific throttling. The implementation used a 12-second search limit, one retry, and a 30-second wait after three consecutive failures. These settings describe his response to that observed run, not a guarantee that the same diagnosis or limits fit another system.
For operations, Ted says FORGE appears in Operator Pulse, which tracks server state, recent runs, success rate and remaining OpenRouter credit; scheduled questions are tracked as jobs. That describes how this build was monitored, not a general assessment of the dashboard’s availability or suitability.
What do the reported model costs and timings mean?
The figures below are the author’s reported values in the September 27, 2026 post. They describe his implementation and should not be read as current API pricing, reproducible benchmarks or estimates for a different workload.
Best Value
- Ted reports about $0.15 per million input tokens for GLM-5.3 Flash through OpenRouter.
- For a review, he reports about $0.03 using Claude Sonnet at high effort, about $0.014 at lower effort, and $0.009 using Claude Haiku.
- In a three-sentence prompt comparison, he reports 93 seconds and 3,200 reasoning tokens without a reasoning-effort setting, versus 9 seconds with effort set to low. He also describes a later parallel question finishing in 5 minutes 10 seconds for four cents after setting effort for each call.
The account connects those results to practical tuning: reasoning settings can affect both elapsed time and how much output remains for the task. Ted’s figures are observations from his own calls, not a controlled comparison of models or a prediction of what another run will cost.
What is the useful lesson for evaluating an AI app demo?
Look beyond whether the screen is populated. Ask what event produced each status, whether the displayed history survives across clients, and whether a simulated run can be mistaken for a real one. For research answers, check whether the system distinguishes “sources attached” from “claims checked,” and whether it makes a low-evidence result visible rather than silently presenting it as verified.
Ted summarizes the distinction: “A demo shows that something can work. Making it trustworthy meant finding every place it only looked like it worked.” His account is a first-person build report: its architecture, fixes, timings and costs are self-reported, not independently tested.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

