Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

A deletion test can change a model’s confidence without changing its predicted label. That is why flip rate—the share of inputs whose predicted class changes after an explainer’s top-ranked feature is removed—cannot, by itself, tell you whether the explanation was informative. A useful evaluation records the confidence trajectory as well as flips, and compares results across baseline-confidence and input-type groups.

What a flip rate does—and does not—measure

In a deletion test, an evaluator ranks input features using an explanation method, removes or alters one or more features, and observes the model’s output. A label flip occurs when the predicted class after the intervention differs from the original predicted class. Flip rate is the proportion of evaluated inputs for which that happens.

That binary outcome discards how the model arrived there. If the original class remains the prediction while its probability falls sharply, the test records no flip. If a small change nudges two nearly tied classes across the decision boundary, it records a flip. Those outcomes are not equivalent evidence about the importance of the removed feature.

Record the original-class probability or logit at each deletion step, not only the final label. The resulting curve preserves the direction and size of output changes that a single flip/no-flip field hides. Insertion and deletion evaluations are widely used, but their behavior depends on metric settings and can be affected by out-of-distribution inputs; Wang and Wang’s TRACE paper examines these issues and offers guidance for using the metrics (ICML 2024 paper).

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Why baseline confidence changes how to read a flip

A model’s starting confidence affects how much a label-based statistic can reveal. With a highly confident prediction, deleting a feature may reduce the original class’s probability substantially yet leave it above every alternative. The label stays fixed even though the output changed. Near a decision boundary, a smaller probability shift may be enough to change the winner.

These cases make flip rate conditional on the starting point: it is not a confidence-independent measure of explanation quality. Stratifying by baseline confidence makes that dependence visible. It does not make one confidence band inherently better, nor does it prove that a low flip rate means an explanation failed. Interpret the label outcome alongside the probability or logit trajectory and the intervention used.

What the titled post reports

Parshvi Jain’s DEV Community post, indexed as published September 28, 2026, reports a LIME deletion evaluation using distilbert-base-uncased-finetuned-sst-2-english, model revision 714eb0fa, num_samples=300, num_features=10, and five random seeds. The author says the evaluation began with 30 pre-registered sentiment inputs across six categories, then excluded two inputs deemed structurally invalid, leaving 28.

Reported result Author-reported value What it describes
Aggregate flip rate 39.3% (11/28) Share of the 28 tested inputs with a changed predicted label
Directional correctness 89.3% (25/28) Reported directional measure; the indexed extract does not provide enough detail to define its scoring rule
Mean top-5 Jaccard stability 0.81 Reported stability summary; the indexed extract does not state further calculation details
High-confidence group 39.1% (23 inputs; p ≥ 0.99) Reported flip rate for that baseline-confidence group
“Strong baselines” category 0% flips Reported result for the named input category; its count is not stated in the extract
“Lexical shortcuts” category 100% flips Reported result for the named input category; its count is not stated in the extract

These are the post author’s figures, not independently verified benchmark results. The post’s direct page was not available for inspection, and the indexed extract does not provide an independently inspectable run artifact. The reported pattern illustrates why group-level context matters, but it cannot establish a general flip rate for LIME, DistilBERT, or saturated models.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

How to design a confidence-stratified deletion evaluation

  1. Define the outcome before running the test. Specify what counts as a flip, which class is the target, and whether the recorded model output is a label, probability, or logit. For a diagnostic evaluation, retain the original target-class probability or logit at every deletion step as well as the label.
  2. Specify the intervention precisely. Record the explainer’s feature ordering, the number of features removed per step, and whether removal means deletion, masking, blurring, or replacement. State how structurally invalid inputs are handled and report exclusions with their counts and reasons.
  3. Set strata in advance. Choose baseline-confidence bands and meaningful input categories before examining outcomes. Report each group’s sample size alongside its flip rate and confidence or logit changes. Small or uneven groups can make percentages unstable, so keep the counts visible and avoid over-interpreting apparent differences.
  4. Plot the trajectory. For each deletion step, show the target output (such as the original class probability) or summarize trajectories in a way that preserves their change over steps. Include flip timing where useful. A final binary flip rate cannot substitute for the path that produced it.
  5. Apply the same budget across methods. When comparing explainers, use the same inputs, target output, deletion steps, and perturbation budget. Report run-to-run stability and computational cost alongside the deletion results so readers can distinguish explanation ordering from implementation or sampling differences.
  6. Pair deletion with other evidence. Treat deletion as one diagnostic of faithfulness rather than a universal score. Depending on the use case, add suitable complementary metrics or human evaluation; make clear what each measure tests.

Why the removal rule matters

Deletion is a simulated intervention, not a direct window into a feature’s causal role. Covert, Lundberg, and Lee describe removal-based explanations as methods that simulate feature removal to quantify influence: “We describe a new unified class of methods, removal-based explanations, that are based on the principle of simulating feature removal to quantify each feature’s influence.” Their framework emphasizes that results depend on how features are removed, which model behavior is explained, and how influence is summarized (JMLR 2021 paper).

For text, deleting a token, replacing it with a mask token, and substituting another word are different interventions. Each can produce a different input and a different model response. Describe the operation rather than treating all forms of “removal” as interchangeable. Also report whether the resulting inputs remain meaningful for the task; the deletion score alone does not establish that they do.

Perturbation realism is a documented concern in image saliency evaluation: Gomez, Fréour, and Mouchère note that progressive masking or blurring can create inputs unlike those seen during training. Their analysis concerns image evaluation, so it should not be presented as direct proof of the same effect in text; it does reinforce the need to examine what an intervention feeds to the model (2022 analysis of DAUC and IAUC).

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

What to report when comparing explainers

A fair comparison makes the choices behind each score inspectable. Use a compact report that covers:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Best Value
  • Feature ranking and the exact deletion, masking, or replacement operator.
  • Target output and whether the analysis uses labels, probabilities, or logits.
  • Per-step trajectories as well as any aggregate statistic.
  • Baseline-confidence strata and input categories, with sample counts.
  • Perturbation realism concerns and treatment of invalid inputs.
  • Run-to-run stability, shared inputs and perturbation budget, and computational cost.

Metric-aware methods are also being studied: Yoshikawa and Iwata’s ID-ExpO work proposes differentiable regularizers for insertion/deletion metrics and reports experiments on image and tabular datasets. It is a research approach, not evidence that one universally settled evaluation metric exists (AISTATS 2024 paper).

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.