What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more
Bloom gives two language models the same creative-coding prompt, runs their p5.js sketches side by side, and compares what the rendered canvases and provider usage data show. In Harish Kotra’s example, both models are asked to animate an ocean current. Bloom measures observable differences—including whether sampled pixels actually change—rather than declaring a winner with a supposedly objective aesthetic score.
What Bloom compares
Bloom is a practical instrument for a paired model comparison: each model receives the same request, and its generated p5.js sketch appears in a sandboxed browser iframe beside the other result. The example prompt asks for an animated ocean-current artwork and requires code only.
This setup helps make a creative task easier to inspect. You can check whether each sketch parses and runs, whether its image changes over time, and what descriptive visual measurements result. Provider-reported usage data can add another comparison axis when the provider returns it. The measurements describe the outputs; they do not establish which artwork is more beautiful or which model is better overall.
How Bloom turns sketches into measurements
Motion from sampled pixels
A sketch can call a draw loop without visibly changing its image. Bloom therefore does not rely on a model claiming that its output is animated. In the reported implementation, the iframe samples the canvas on a 48 × 48 RGB grid every fifth frame and sends those pixels to the parent page. The motion reading averages absolute pixel differences between consecutive samples; zero means the sampled pixels did not change.
#1 Best Overall
Kotra describes this as keeping the measurement honest: “That keeps the measurement honest — the sketch can’t self-report "I animate, trust me".” The quote is from the article’s author, not an independent evaluation.
Four descriptive visual measures
Bloom reports four summaries of the rendered image. Each captures a limited property of the pixels, not artistic merit:
- Distinct colors: the number of colors after quantization.
- Mean luminance: average brightness using Rec. 709 luminance.
- Edge density: a measure based on luminance changes between neighboring pixels.
- Composition symmetry: a correlation-based estimate of left-right symmetry.
These descriptors can help explain how two images differ, but none is a substitute for a viewer’s judgment. Bloom does not combine them into a single objective aesthetic score.
How the sketches are run and constrained
The backend makes model calls, which allows the project to work with local providers such as Ollama and LM Studio without configuring browser CORS for those calls; it also keeps API keys out of the browser bundle. The frontend sends generated code to a reusable iframe through postMessage. Kotra says p5.js is bundled locally rather than loaded from a CDN.
The project uses several defenses around generated JavaScript. Before execution, an Acorn abstract syntax tree scan flags operations such as network calls, module loading, workers, storage access, parent-window access, and imports; code that triggers a violation is refused. The execution iframe uses sandbox="allow-scripts", and its bootstrap disables several network and storage interfaces before the generated code runs.
Kotra reports headless-browser probes of selected restrictions. That is evidence about the checks he describes, not proof that every possible hostile script is contained. A code scan, disabled interfaces, and iframe sandboxing are defense layers; they should not be read as an independent security assessment or a guarantee against arbitrary generated code.
Rank #4
What the reported run found
In one run with two local models, Kotra reports motion readings of 1.05 and 0.82/255. These are readings from that project run, not general performance statistics or evidence that one model is better across prompts. The author also reports that browser-level verification ran both canvases at approximately 60 fps.
Recommended Free Tools
For verification, Kotra says a server self-test covered 39 checks. He also reports browser testing with two local LM Studio models, including sandbox probes and the pause, reseed, and poster-rendering controls. These are the author’s reported checks, not an independently reviewed test suite or a replicated benchmark.
Best Value
What a fair comparison can—and cannot—show
Bloom’s strongest contribution is methodological: compare what two outputs do under a shared task, and make the observations visible. For a useful paired run, record the prompt, models, provider, settings, and run conditions, then examine several distinct questions:
- Did each sketch parse and run?
- Did the sampled pixels change, and by how much?
- How do the four visual descriptors differ?
- What usage information did each provider return?
- Which safety probes were attempted, and what happened?
The project appends a fresh random nonce to prompts and records hashes of prompt-plus-nonce, an approach intended to reduce accidental cache reuse. It also uses seeded rendering. Those choices support repeatability, but the article does not establish a full controlled benchmark protocol or broad statistical conclusions. A result from one paired run should stay a result from that run.
Implementation details for people reproducing the approach
Bloom handles provider differences selectively: it shows reasoning-token data only when returned, conditionally sends a thinking-related parameter for a specified provider/model case, reads available model lists live without making that step a prerequisite, and retries selected budget errors. These details help keep the interface usable across providers with different capabilities; they do not imply that every provider returns the same data.
Windows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallCrashes, No Sound, or Screen Glitches?
Random freezes, missing sound and display glitches usually trace back to one bad driver. Find and replace yours safely.Free scan · under a minuteThe author’s proposed next steps are bracket mode for more than two models, a judge slot, replay files containing prompt, nonce, seeds, code, and metrics, and time-lapse export. These are ideas for future work, not capabilities confirmed as available in the described project.
Source and scope
The project and its reported results are described by Harish Kotra in “Bloom: I Made Two LLMs Paint the Same Sentence and Measured What Happened”, published September 29, 2026. The account explains the implementation and selected checks, but does not provide an independent reproduction or a general ranking of language models.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

