← All posts
Method feedsmethodology

Why we attribute every indicator to a package

Most threat feeds hand you a list of domains and no provenance. Ours name the package, version and file each indicator came from — so you can check whether it even concerns you, and verify it before you block.

28 Jul 2026 5 min read By codelake Research

An indicator you cannot trace back is a request to trust us. A domain on a blocklist, a hash in a feed — on their own they tell you to act without telling you why. We would rather hand you something you can check: every indicator we publish names the package, the version and the file it was extracted from.

The problem with a flat list

The common shape of a threat feed is a column of domains, IPs or hashes with maybe a date attached. It looks comprehensive, and it is almost impossible to act on with confidence. You cannot tell whether an entry touches anything you actually run, you cannot verify the claim without re-doing the research yourself, and a single over-broad entry — a shared CDN host, a pastebin, a popular analytics domain — turns into a self-inflicted outage the moment you push it to a blocklist.

The missing piece is always the same: provenance. Where did this indicator come from, and what was it doing there?

What attribution looks like

Every indicator in our feeds carries the chain of custody that produced it:

  • Registry, package and version. Not "this domain is bad" but "this domain is contacted by [email protected]."
  • The file inside the artifact. The archive-relative path the indicator was lifted from, so you can open the same file.
  • A content hash. A SHA-256 of that file, so the thing you inspect is provably the thing we inspected.
  • The role it plays. Loader, beacon, exfil endpoint — what the indicator does in the chain, not just that it appeared.

The same indicator can be benign in general and malicious in context. A hosting domain is not an IOC; a hosting domain that a lifecycle hook posts environment variables to, in a specific package version, is. Attribution is what lets the feed say the second thing instead of the first.

A domain is not an indicator. A domain, tied to the package and version that abused it, is.

Why it changes what a defender can do

Once each row is anchored to a concrete artifact, triage stops being guesswork:

  • Check relevance first. If you do not depend on the package, the indicator is context, not an alert — you can deprioritise it instead of blocking a domain blind.
  • Verify before you block. Pull the named file at the named hash, read it, and confirm the extraction yourself. The feed is a claim you can falsify, not an article of faith.
  • Scope to your estate. Cross the package list against your lockfiles and you know exactly which repositories and build pipelines are in scope.
  • Keep the behaviour, not just the string. Knowing an endpoint was reached from a postinstall hook tells you where to look in your own logs.

Reproducible by design

Attribution is only credible if the source survives. Malicious releases are frequently unpublished within hours of discovery, so we archive the original tarball and its metadata at capture time. The sample outlives the registry entry, which means the file an indicator points to is still there to be re-examined long after the package has vanished from the registry.

That is the whole point: the indicator, the file, the hash and the archived artifact form a line you can walk end to end. You never have to take our word for it — and neither do we.

Keep reading
All posts →
Stay current

New editions, in your inbox.

Get notified when codelake Research publishes a new report, threat brief or quarterly advisory roundup. No marketing — just the research.

We use your address only to send research updates. Unsubscribe anytime.