Deduping millions of findings with one content hash
Part of my series on RiskRancher, the open-source vulnerability manager I build. Source: github.com/Kuebiko-LLC/risk-rancher-core.
Intro
Scanners are repetitive by nature. You run Trivy on the same image every build, Nessus on the same hosts every week, and the overwhelming majority of what comes back is findings you already have. If you store every finding from every scan as a new row, you drown in duplicates fast, and the “system of record” becomes useless. So deduplication isn’t a nice-to-have in a vulnerability manager, it’s the thing that makes the data worth keeping. Here’s how RiskRancher does it with a single hash, and the interesting part is what I chose to leave out of that hash.
I. A finding’s stable identity
The question is: what makes two findings “the same”? My answer is three things: which scanner reported it, what asset it’s about, and what the finding is. So the dedupe key is the scanner source, the asset identifier, and the title, hashed:
hashInput := ticket.Source + "|" + ticket.AssetIdentifier + "|" + ticket.Title
hash := sha256.Sum256([]byte(hashInput))
ticket.DedupeHash = hex.EncodeToString(hash[:])
That DedupeHash is a finding’s fingerprint. The same vulnerability, from the same scanner, on the same asset, always produces the same hash. So when next week’s scan arrives, RiskRancher recognizes the finding it already has instead of adding a near-identical twin. One deterministic hash turns “store everything and sort it out later” into “recognize what we’ve seen.”
II. What I deliberately left out
Here’s the design decision I think is worth sharing: the hash is built from source, asset, and title, and nothing else. Not the severity. Not the description. Not the scan timestamp.
That’s on purpose, and it’s the opposite of what feels natural. Your instinct is to throw more fields into the key to be “more precise.” But think about what changes between scans of the same real vulnerability. The scanner bumps the severity after a CVSS update. The description gets reworded. The remediation text changes. If any of those were in the hash, the same underlying finding would get a new fingerprint and show up as a brand-new duplicate, which is exactly the thing dedup is supposed to prevent.
So the rule is: hash the stable identity, not the volatile details. Source, asset, and title are what make a finding that finding. Everything else is an attribute of it that’s allowed to change without making it a different finding.
III. Grouping, for free
Once findings are normalized and fingerprinted, grouping falls out naturally. RiskRancher buckets findings by source and asset:
key := groupKey{Source: ticket.Source, Asset: ticket.AssetIdentifier}
groupedTickets[key] = append(groupedTickets[key], ticket)
So the dashboard can show you “everything wrong with this host” or “everything this scanner found,” because the data was organized by asset at ingest time. Dedup and grouping are the same idea applied at two scopes: one hash identifies a single finding, and the (source, asset) pair gathers them into something a human can actually work through.
Conclusion
Deduplication is what keeps a vulnerability manager from becoming a landfill, and it comes down to one question: what makes two findings the same? Answer it with the stable identity (source, asset, title), hash that, and pointedly keep the fields that drift between scans out of the key. The result is that re-scanning the same target forever doesn’t grow your database, it just refreshes what you already know. Thanks for reading, and may your duplicates collapse quietly.