A package a detector flagged and a human then cleared is not a wasted scan. It is a labelled example of exactly where a heuristic over-reached — which is the single most useful thing you can feed back into it. We treat every confirmed false positive as an asset and keep it forever.
The loop
A finding does not go straight from a detector to a verdict. The path is deliberately boring:
- A detector flags a package on a concrete signal — a lifecycle hook reaching the network, a high-entropy string, a native binary in the tree.
- A human reviews it against the artifact and decides: real, or over-reach.
- If it is over-reach, the reason is recorded — not just "benign", but which heuristic fired and why it was wrong here.
- That reason is written back as a negative fixture and an allowlist entry, so the same package cannot be re-flagged on the same mistake.
- The detector is tuned, and the fixture stays in the suite as a regression guard against the next tuning pass undoing today's correction.
What a false positive actually teaches
Each cleared case marks a boundary the heuristic did not know about. A few of the boundaries we have had to learn, which keep recurring:
- Minified first-party bundles. A vendor's own release, run through a minifier or bundler, trips the same dynamic-execution and obfuscation signals a dropper does. The build step is the tell, not malice.
- Bundled media binaries. A creative or video tool that ships
ffmpeglooks like a "native binary dropper" until you notice the binary is the declared purpose of the package. - Declared-purpose network calls. A CLI that fetches from its own service, or a loopback /
localhostcallout, is not exfiltration — the destination matters as much as the act. - Public keys that look secret. Some analytics and client keys are the same length and shape as a real secret, but are meant to ship in the open. Flagging one as a leak is its own kind of error.
Tune, don't suppress
There is a lazy version of this where you simply mute the noisy rule. It makes the dashboard quieter and the detector blinder — a muted class of verdict is indistinguishable from a detector that never fires, and that is how real findings get lost. An allowlist entry with a recorded reason is the opposite: it is scoped to the case it explains, it is reviewable, and it leaves the underlying detector free to fire on everything the reason does not cover.
We correct in public, too
The same discipline applies when a mistake has already left the building. If we ever flag a public key as a leaked secret and notify a maintainer, the fix is not a quiet database edit — it is a correction sent to the same person, saying plainly what we got wrong. A research program that cannot admit a false positive has no business publishing true ones.
Cleared packages are never pruned from the corpus. The benign set is the negative class every detector is measured against — without it, you can make a scanner flag everything and call it sensitive. The false positives are half of the ground truth.