Recommended Free Tools
iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more
The business logic in a codebase is the code that runs when the system does its real work, and ordinary code search cannot show you which code that is. Mikhail’s bootstrap pipeline tries to answer the question that motivates his post, “where is the real business logic here, and what can I safely throw away?” It does so by recording which source functions each test actually executes, then feeding those links into search. In his experiments, the simpler shortcuts did not identify the functions a test runs, so he used a runtime trace. The trace narrows where to look. It does not certify that anything is safe to delete, and the deletion decision still depends on history, ownership and production evidence.
What the source is, and how far its numbers reach
The account is a first-person engineering case study by Mikhail, posted to DEV Community on September 22, 2026, and labeled as draft material. Every figure in this article is the author’s own measurement from his experiments. He did not provide independent replication, and none of the numbers is a general performance guarantee or an industry statistic. Each one is given below with the sample it came from.
Why ordinary code search misses the answer
Code search indexes text and symbols. It can tell you that a function exists, what it is called and where it is referenced, but not whether a test or a request path actually runs it. A file that imports a module and a function that executes are different facts, and a plain search only sees the first. That gap is what makes triage hard: a large, well-named helper can look central while rarely running, and a small, obscure function can sit on every request path.
The Tool Desk
Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →The four stages of the bootstrap
The pipeline builds its map in four passes, and each pass adds a different kind of evidence.
1. Entities
Types and data classes describe what the system stores and passes around. They supply the nouns of the map, so later stages can connect behavior to the data it handles.
2. Entry points
Entry points are the places where outside input arrives, such as functions decorated with @mcp_app.tool. The pipeline uses them as the starting points for reading the system.
3. Tests as execution evidence
Each test is linked to the source functions it executes. This is the stage that carries most of the technical difficulty, and the next two sections explain why it needed a runtime trace rather than a simpler method.
Windows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallCrashes, No Sound, or Screen Glitches?
Random freezes, missing sound and display glitches usually trace back to one bad driver. Find and replace yours safely.Free scan · under a minute4. Git history
Commit history is mined for architectural decision records, which explain why a module is shaped the way it is. That context turns a list of functions into a list of decisions someone can review.
Linking a test to the code it runs
Connecting a test to the function it actually executes is the central technical problem. The author tried four approaches before settling on a trace.
| Approach | What the author measured | Outcome in the post |
|---|---|---|
| Name matching | None of 109 sampled tests named the function they executed | Rejected |
| File-level import matching | Reached 77.9% of tests | Judged too coarse, because a file can contain many functions |
| Tarantula-style ranking heuristic (a spectrum-based fault-localization technique) | Supplied rank-at-most-three candidates for 22.6% of tests, and rank one for 7.5% | Rejected as a universal primary-target selector |
Full execution trace from a custom sys.settrace plugin |
1,551 of 1,727 tests (89.8%) executed at least one source function; 1,212 unique source functions were reached | Kept as the basis for TESTS edges |
The trace matters because one test usually touches many functions. Across the linked tests, the average was 10.1 source functions per test, but the median was 6 and the range ran from 1 to 118. A model that assumes one test maps to one function is wrong for most tests, and a tool that shows only a single “primary” function per test hides most of what the suite exercises.
What tracing costs
A trace slows the suite down. The author measured two setups, and they were separate comparisons rather than one controlled run.
Do these 3 things before closing this tab:
1Clear out junk files and repair common Windows errors2Scan for outdated or missing drivers - takes under a minute3Repair Windows errors before they cause bigger problems| Setup | Baseline | Traced | Overhead reported |
|---|---|---|---|
Custom sys.settrace plugin, same-session comparison over the 1,727-test suite |
174.8 s | 198.6 s | 13.6% |
coverage run on Python 3.14 using sys.monitoring |
184.88 s | 221.78 s | 19.96% |
The author describes the second setup as about 1.5 times slower than his plugin. Whether either overhead matters depends on how often the trace runs. If it runs on every push, the cost is paid on every push; if it runs on a schedule, the cost is limited to that schedule.
Static analysis as a companion, not the edge driver
The author compared three static signals against the dynamic trace, which served as the reference.
| Signal | Hit | Recall | Precision | Mean candidates |
|---|---|---|---|---|
| L1: AST direct calls | 88.4% | 30.3% | 68.0% | 2.9 |
| L2: name tokens | 17.7% | 3.8% | 12.1% | Not stated |
| L3: file imports | 91.6% | 72.0% | 21.8% | 41.4 |
| Union of L1, L2 and L3 | 90.4% | 70.0% | 20.6% | Not stated |
The signals trade precision against coverage. Direct calls are the most precise but recover only 30.3% of what the trace finds. File imports recover most of it, but they bring about 41 candidates per test and only 21.8% precision. The author’s conclusion is that static analysis is a useful companion for tests that rely on mocks, where it can supply candidates, while the dynamic trace remains the driver of TESTS edges.
The mock case is the reason for the companion. The author found that 176 tests (10.2% of the suite) executed no source functions at all, and he attributes these gaps to mocks, which can stand in for the code a test would otherwise run. Static companions covered 88 of those 176 tests.
From trace to search signal
In the build the author labels E17, the trace produced 16,172 TESTS edges across 1,595 Test nodes, linked to 1,132 covered functions, from the 1,727-test suite. The search layer then uses those edges as follows:
Rank #4
- The search finds function results as it normally would.
- For each returned function, it calls
SymbolIndexAdapter.get_tests_for_symbol()to fetch the tests linked to it. - It appends up to three tests per function through
Searcher._append_tests_signal(). - The appended tests are capped per query with
min(len, 6), a ceiling of six. - Tests receive a graph_score of 0.4, against 1.0 for definitions, so they supplement the results rather than outranking them.
The behavior sits behind a toggle named MSCODEBASE_TESTS_SIGNAL, which was off by default in the implementation described.
What the search experiments showed
A seven-function A/B panel
With the signal on and off, function ranking stayed the same, with MRR of 1.000 in both arms. Six of the seven queries received relevant covering tests.
The 35-query panel
| Metric | Signal off | Signal on |
|---|---|---|
| hit@1 | 33/35 (94.3%) | 33/35 (94.3%) |
| hit@3 | 34/35 (97.1%) | 34/35 (97.1%) |
| MRR | 0.957 | 0.957 |
The signal added covering tests to 34 of 35 responses (97.1%), and it left hit@1, hit@3 and MRR unchanged. Read the table that way: the signal adds context to results that were already found, and it does not improve top-one retrieval. The author puts it directly: “TESTS-signal does not improve hit@1 (off=on). It only adds context (tests) to already-found results.”
Latency in the wider panel
Across the 35-query panel, average graph-stage time rose from 6.52 ms to 7.53 ms, roughly 15%. That is a measurement from one panel, not a forecast for production latency. The author lists a real LLM-pipeline check and caching as unfinished work.
Best Value
External projects: small checks, not a portability benchmark
The author also ran the pipeline against three other projects. Those checks show that the pipeline ran on other repositories, not how well it performs across codebases in general, and the samples are small.
| Project | Language | Tests | Reported result |
|---|---|---|---|
gemma_agent |
Python | 2,882 tests, 2,874 passing | 97.3% linked (2,805 tests); overhead of 17.4% (71.4 s against 60.8 s) |
commit- |
Python | 27 tests | 100% linked |
codebase-memory-mcp |
Go | 27 test functions | Coverage of 51.0% at package level and 22.2% per test |
Language coverage is the biggest limit
The dynamic edge builder is Python-first. In the implementation described, TESTS edges exist for 1,108 of 3,256 Python functions (34.0%). They exist for none of the Go and Rust functions (716 in the listed group) or the TypeScript functions (11). Connectors for Go and TypeScript are proposed but not built. If your codebase is mostly outside Python, the method as described does not yet apply to it.
Where the approach fails
- Failing tests drop their edges. A failing test can remove its links from the map, so read the graph from a run in which the suite passes.
- Utility functions create noise. Widely used helpers run under many tests, so they link to a large share of the suite and crowd out the functions you are looking for.
- Some test nodes have line 0. Their location data is unreliable, so a jump to the test’s line can land in the wrong place.
- The graph_score of 0.4 is untested against BM25 and reranker interactions. Its effect in a full search stack that uses those components is unknown.
- Reindexing can shift node order. Tie-breaking between builds may therefore differ even when the code is unchanged.
- Very large suites may not fit a CI window. A traced run of a very large suite may not finish within a CI time limit.
- Verification has been local so far. At publication, clean CI confirmation was pending a pull request merge.
- The search evaluation is narrow. The 35-query panel used one primary codebase, so its results may not carry to other projects.
Deciding what you can safely throw away
Neither the trace nor the search signal certifies code as dead. The trace shows what the suite reaches, not what production traffic, scripts or outside callers reach. Treat the output as a ranked worklist. The steps below are a suggested triage built from the trade-offs above; the source does not test this sequence.
Free tools Windows power users keep installed
One-click scans. No signup required.
- Start from the entry points and follow the TESTS edges to see which functions the suite covers beneath each one.
- Flag functions with no TESTS edge and no visible caller from an entry point. These are candidates to investigate, not deletions.
- Read the Git history for each flagged module and look for the architectural decision record that explains why it exists.
- Confirm with the module’s owners and with production or log evidence, which the test suite cannot provide.
Questions to ask before adopting a similar approach
Use these questions to compare a test-to-code linker with the approach described here:
Quick Recap
- Does it report precision and candidate volume alongside recall?
- What does execution tracing add to the suite’s runtime, and how often will that cost be paid?
- Which languages and test frameworks does it cover today?
- Does its output improve ranking, or only add context to results?
- How does it behave with failing tests, mocks and very large suites?
- How broad is its validation, and has it been exercised by a downstream LLM pipeline?
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

