Data minimisation is the rare compliance principle that everyone agrees with and almost no one can evidence. “Collect only what you need” is one line in a policy and a genuinely hard thing to prove against a live codebase. The schema that a service ships with is rarely the schema someone designed on purpose. Fields arrive for a feature, the feature ships and later gets cut, and the field stays in the model, quietly holding data that nothing reads.
A spreadsheet inventory will not catch that. The spreadsheet records what someone believed the system collected on the day they filled it in. The code records what it collects today. So we pointed Scrutora at a real one and let it derive the map from source.
Metriport is a good subject for this because it is open, it is a real healthcare API, and it handles data that regulation actually cares about. Everything below is checkable against the repository, field by field, down to the file and line.
What the Scan Settles First
Before a single flow is drawn, the question is which fields are personal data at all.
Scrutora begins by reading the codebase for every field name that looks like personal data and proposing a classification for each. That first pass is deliberately wide. Name-based detection catches real identifiers and also catches false positives: a config parameter called max_age, an entity property called name, a format string. So the map does not treat the first pass as truth. It asks for confirmation, and it shows its own uncertainty on screen.

On the tags themselves
Scrutora scans against more than one framework, so the same field can carry a different tag depending on which regime applies to the codebase. A date of birth is plain personal data in one context and a protected health identifier in another. The tag follows the framework, not the field name, which is why confirmation matters and why the map is built from confirmed fields rather than raw matches.
Then It Follows Each Field
Collection, processing, and every sink the field reaches.
For each confirmed field, the scan traces where the value goes: into a database store, into application logs, out through an external transmit. The map is a live picture of that, the auto-detected flows merged with any destinations added by hand. In this run it resolves to 372 sinks across 211 entry points and 237 processing points.

This is already more than a spreadsheet can hold, because it is derived rather than recalled. But the finding that stays with you is not on any of the loaded paths. It is the field that has no path at all.
Sixty-Four Fields That Go Nowhere
Declared in the code, never traced to a store, a log, or a recipient.
The map carries a sink of its own for these: Declared, no flow traced. Sixty-four fields are defined in the codebase and never reach a store, a log, or an external recipient that the code can demonstrate. They sit in the schema. Something, at some point, decided to collect them.

That list is the most actionable minimisation artifact we know of. A privacy engineer cannot review a whole schema on instinct. They can review 64 named fields, and for each one answer a question that a RoPA spreadsheet never poses: do we still need this, and if not, why are we holding it?
Why Derive It From the Code
The usual way to build a Records of Processing inventory is to ask teams what their systems collect and write it down. That produces a document that is accurate on day one and drifts every day after, because the code keeps changing and the document does not. Deriving the map from source inverts the default. Instead of asking people to recall what they hold, you start from what the code demonstrably touches, and the fields that touch nothing fall out on their own.
- The inventory is a product of the current code, not a memory of it, so it does not drift between audits.
- Unused fields surface as a side effect of mapping, rather than needing a separate hunt.
- Every entry is evidence: a field name, a file, a line, checkable by anyone with the repo.
- The uncertain cases are labelled as uncertain instead of being presented as fact.
The Honest Limits
Two, stated plainly. The classification pass is wide by design and needs a human to confirm it; the 33 fields ruled out in this run are that step, not an afterthought. And the flow map is static, so it sees imports and assignments rather than runtime behaviour. Neither limit weakens the minimisation finding, because a field with no traced flow is exactly the field a human should look at next. It does mean the number is a starting point for review, not a verdict.
If you want to check any of this, the repository is public and the field names above are real. That is the point. A data map is only worth as much as its evidence, and evidence is checkable or it is nothing.
Vimalendukumar Dwivedi is Co-founder of Scrutora, a Code Compliance Platform.
Find your own 64
Scrutora traces every personal-data field in your codebase to wherever it actually lands, at rest or in motion, and hands you the evidence instead of another meeting.
Share this article