Every scanner speaks a different language: normalizing findings into one model
Part of my series on RiskRancher, the open-source vulnerability manager I build. Source: github.com/Kuebiko-LLC/risk-rancher-core.
Intro
The first hard problem in vulnerability management isn’t storing findings, it’s that every tool describes a finding differently. Trivy buries vulnerabilities under nested result arrays, Qualys hands you one giant CSV, Nessus has its own format, and Dependabot speaks yet another JSON dialect. One calls it title, the next name, the third vuln_id. One says severity is HIGH, another says 7.5, another says Critical. If you want all of that in one place, you have to translate. This post is about how RiskRancher does that translation once, into a single model, so nothing downstream has to care where a finding came from.
I. The usual pain: one parser per tool
The naive approach, and the one I kept doing before this project, is to write a little parser for each scanner. A Trivy script, a Qualys script, a Nessus script. Each one knows that tool’s exact shape, and each one breaks the day the tool changes its export format. You end up maintaining a pile of brittle one-offs, and adding a new scanner means writing another one. That doesn’t scale, and it’s exactly the busywork a vulnerability manager is supposed to remove.
II. One shape to rule them all
RiskRancher’s answer is a single internal model that every finding becomes, no matter its source. A finding, once ingested, is a Ticket:
type Ticket struct {
Source string
AssetIdentifier string
Title string
Severity string
Description string
RecommendedRemediation string
// ...status, dedupe hash, timestamps
}
That’s the whole contract. Source is which scanner it came from, AssetIdentifier is what it’s about (a host, an image, a repo), Title is the finding, and the rest is detail. Every scanner’s output, however baroque, gets flattened into this. Trivy’s nested vulnerabilities and Qualys’s flat CSV rows arrive at the same place, looking identical.
The payoff is that everything after ingestion is written against this one shape. Dedup, asset grouping, SLAs, the dashboard, the reports: none of them have a special case for Trivy or Qualys, because by the time they see a finding, it’s just a Ticket. Normalize once at the edge, and the whole rest of the system gets simpler.
III. JSON and CSV, because that’s what scanners emit
Scanner exports come in two flavors in practice: structured JSON and flat CSV. RiskRancher ingests both. A JSON export gets decoded and the findings array walked; a CSV export gets read row by row:
reader := csv.NewReader(file)
// ...each row becomes a finding, each column a field to map
Either way, the output of this stage is the same list of normalized Tickets. The input format is an implementation detail that stops mattering immediately.
IV. The part I deliberately didn’t hardcode
Here’s the key decision: I did not bake the mapping for each scanner into the code. There’s no if source == "trivy" anywhere. Instead, the knowledge of “where does the title live in this tool’s output” is configuration, not code, defined per scanner in what I call an Adapter. That’s its own post, the no-code Adapter Builder, but the principle is what matters here: the translation layer is data you can add to, not code you have to ship. Supporting a new scanner shouldn’t require a release.
Conclusion
The messiest part of vulnerability management is that every tool disagrees about how to describe a vulnerability. The fix is to translate everything into one model at the moment it enters the system, so every feature downstream can be written once against one shape. Normalize at the edge, keep the mapping as configuration rather than code, and adding the tenth scanner costs the same as adding the second. Thanks for reading, and may your findings all speak the same language.