Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

BeyondBug is an MIT-licensed hackathon submission and judging platform built for DOGFOOD 2026, and it runs locally with Docker Compose. Its central feature is a judge-severity correction. The platform estimates how strict or generous each judge is, adjusts their reviews, and keeps the original scorecards for comparison. In the project’s official fixture, that correction changed the rank of 33 of 40 ranked projects. The second feature is the access boundary: checks run on the server before any protected record is read or changed, so hiding a button in the interface is never treated as protection. A separate anomaly model flags reviews for organizers to inspect, but it cannot change a score or a result.

The details below come from the project author’s own article, kadhiravan’s DEV Community post of September 29, 2026, BeyondBug: The Score That Moved, the Boundary That Held. Every figure in it is either a fixture result or a synthetic test reported by the author, and each is labelled that way where it appears. The author states the aim this way: “The objective was software another organizer could evaluate, operate and extend, not a checklist with hidden gaps.”

What BeyondBug covers and who can do what

The platform is designed to run the full life of a hackathon: event setup, registration, teams, submissions, judging, community voting, results publication, feedback, awards and certificates. The article distinguishes five roles: visitor, participant, judge, organizer and administrator. Roles are scoped to each event, so one person can hold different permissions in different events.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  • Participants submit projects. Their requests to the peer-score route are forbidden.
  • Judges can see only the projects assigned to them.
  • Organizers pass the authorization checks for ranking and export functions, and they are the only role that can open the anomaly inspection queue.
  • Visitors and administrators are named as roles, but the article does not list their permissions in detail.

Why the score moved

The primary ranking starts with criterion scores from 0 to 5, combined using organizer-defined positive weights. The author’s premise is that a simple average can be pushed up or down by which judges happened to review a project, so the platform adds a severity adjustment on top of that weighted score.

The two-way additive severity model

The adjustment uses a regularized two-way additive model that estimates two quantities at once: the quality of each project and the severity of each judge. Reviews are then adjusted for the estimated severity. The stored original scorecard is retained alongside the adjusted value, so an organizer can always see both. The author presents the model as a way to make a panel’s scoring tendencies inspectable. It is not presented as proof that the statistics reveal objective truth.

Fixture results

The official fixture contains 41 project records from 40 teams, including one deliberate duplicate. Excluding that duplicate leaves 40 ranked projects, built from 126 historical scorecards and 122 completed reviews. The fixture has 30 judges in one connected overlap component, plus a judge who gives constant scores. The results below are the author’s reported figures for that fixture.

Project Raw rank Adjusted rank Adjusted score
Iron Switch 2 1 4.316
Salt Ledger 1 2 4.295
Dry Relay 4 3 4.176
Salt Loom 5 4 4.069
Salt Kiln 6 5 4.043
Open Beacon 26 19 not stated
Paper Anchor 21 28 not stated

The author reports that 33 of the 40 ranked projects change position after adjustment. The table shows the five adjusted positions the article lists, plus two further examples. The article does not give adjusted scores for the last two projects.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

What the reshuffle shows and what it does not

The author uses the swap between Iron Switch and Salt Ledger to show two things: judge severity can change a simple average, and the correction can be reproduced on this data. The article is explicit that the reshuffle does not prove the adjusted order is objectively correct. A reproducible correction shows consistency in the method, not that the resulting order is the right one. Because the raw scorecards remain on file, an organizer can trace any adjusted position back to the individual reviews behind it.

Why the boundary held: server-side authorization

The article’s governing rule is that access is decided in the backend before any protected record is read or changed. Hiding a control in the interface is explicitly not treated as the security boundary.

A judge requests another judge’s scores

The article’s main example is a judge asking for another judge’s scores. The server derives the requester’s identity from the session and then checks whether that judge owns the assignment. It does not trust a user ID supplied by the browser, so altering an identifier in a request does not change who the server believes is asking. Unauthorized peer-score requests receive a 403 response. Rankings and exports require organizer authorization on the server.

Session and credential handling

The author describes the following protections. They are implementation details reported in the article, not findings from an independent security review.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  • Session tokens are opaque, and SQLite stores only their SHA-256 digests.
  • Passwords are stored as salted PBKDF2-HMAC-SHA256 hashes.
  • Cookies are HttpOnly and SameSite=Strict. Secure cookies are optional and intended for HTTPS deployments.
  • Logout and password change revoke existing sessions.
  • Writes carrying a foreign Origin header are rejected.
  • Login attempts are throttled.

Deadlines, publication locks and scorecard history

Deadlines are enforced inside database transactions, so the check does not depend on the interface. The article also describes publication locks. Each scorecard keeps the rubric version it was written against, together with the original scores, which is what makes the adjusted view traceable.

Community voting and the identity limit

Voting has its own controls: ballot limits per event and per account, rejection of self-votes and duplicate project votes, tallies hidden until publication, and configuration locks once voting begins. The article’s threat discussion is direct about what remains open. An account does not prove that one person controls it, email matching does not prove that someone owns an inbox, and shared networks complicate IP-based limits. For high-stakes community prizes, the author recommends curated invitations.

The anomaly model: advisory, and evaluated on synthetic data

The anomaly component is an organizer-only inspection queue. The article presents it as advisory software, and its headline evaluation comes from simulated events rather than real ones.

Why the first Isolation Forest was rejected

The author’s first Isolation Forest attempt was dropped for five stated reasons. The list is useful for anyone building similar review-checking tools.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  • Its training contract used a different score scale from the platform’s.
  • It depended on fields that were not available at scoring time.
  • Its peer and history features were prone to leakage.
  • Its evaluation split was unsuitable.
  • Its dependencies were incompatible with the offline Docker image.

The integrated version exports the trees to JSON and runs inference with the Python standard library, so inference needs no extra packages.

Synthetic evaluation

The model card describes 120 simulated events with 30 projects per event and four reviews per project, giving 14,400 simulated reviews, with about 4.6% of them given injected anomalies. The Isolation Forest is configured with 300 trees and a contamination setting of 0.05. The held-out test covers simulated events 108 to 119. The author reports these results on that synthetic test:

Metric Reported value (synthetic held-out test)
Precision 0.52
Recall 0.56
F1 0.54
Overall accuracy 0.95
Decision-score gap 0.137

The article warns that when anomalies are rare, accuracy can look strong while the rare class remains hard to classify. Read plainly, about half of the items the model flagged in the synthetic test were real injected anomalies (precision 0.52), and it caught a little over half of the real ones (recall 0.56). These figures describe simulated data and do not show how real judges would be scored.

Fixture signals and false-alarm rates

The official fixture has no anomaly labels, so its 15 advisory signals cannot be scored for accuracy. The author did report how often the model flagged simulated judges of each type:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Simulated judge type Reported false-alarm rate
Normal 0.8%
Inconsistent 2.9%
Strict 5.2%
Generous 7.5%

Strict and generous scoring look unusual to the model and are flagged several times more often than normal judges. A panel with a genuinely harsh or lenient style will therefore fill the queue faster than a typical panel, which is one reason the queue has to be read as advisory.

What the queue can and cannot do

  • It can surface reviews for an organizer to inspect.
  • It cannot write or change scores, alter normalization or ranking, assign judges, disqualify participants, choose winners, issue certificates or expose peer scores to judges.

It should be described as advisory software evaluated on synthetic data, not as an automated fraud detector or a verdict on any judge.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Running it locally

The repository bundles its dependencies for offline operation. The stack uses FastAPI and SQLite, with local fonts, templates and scripts, the exported model files, fixture data and pinned Python wheels.

Install with Docker Compose

  1. Clone the repository: git clone https://github.com/BeyondBug/DogFood.git
  2. Change into the cloned directory.
  3. Start the stack from that directory with docker compose up.
  4. Find the local address and port in the container output. The article does not state them.

Supported topology and performance evidence

The stated supported deployment is one Uvicorn worker with one SQLite database. The warm local read probes the author ran are short tests. They are not a production service-level objective, not a simultaneous-user rating, and they do not measure write contention. The article does not test multiple workers or multiple application instances, so it offers no basis for scaling beyond the stated topology.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Backups

Backups are local SQLite snapshots, with integrity-check and restore procedures described in the article. Because they are local, the platform’s protection against losing the machine itself depends on what you add, as the decision table below shows.

Decision guide: where the platform fits and what to add

Situation What the article supports What you must add or verify
Single-site event run by one organizer on a local machine Local Docker Compose install, offline operation, one worker and one SQLite database Off-host backups. The article lists off-host disaster recovery as a missing feature.
Event where participants may lose access or need email Login, sessions, logout and password-change revocation Account recovery and email delivery. The article lists both as missing.
Event that publishes results and may correct scores later Original scorecards and adjusted rankings are stored side by side A versioned republication workflow. The article says score correction after publication still needs this, and treats it as future work.
Event that issues certificates for outside verification Certificates can be publicly verified against the local database Certificates are not cryptographically signed, so verification depends on access to that database.
Event that accepts repository submissions from many teams Duplicate detection on identical nonempty repository URLs Manual review for the same project submitted under different URLs, which the rule does not catch.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.