Data Inventory: Metric Reference
This page is the source of truth for the in-app Explain this panels on the
Data Inventory page (/data-inventory). Each section is
written once as a content partial under _explain/data-inventory/ and rendered both here and
inside the app's info panel (scripts/build-explain.mjs compiles the registry).
All KPIs and the table respect the topbar instance filter and only cover PostgreSQL and SQL Server, the two classified engines. The single-table view is documented separately in the Detail reference.
Where sensitive data lives right now across monitored databases: KPIs, a risk-scored table of sensitive tables, and filters to narrow by engine, schema, or search.
How it's calculated
- Scope: PostgreSQL and SQL Server are the only classified engines today. Oracle, MongoDB, Redis, and Valkey are selectable in the Engine filter but labelled "Not classified yet" and return no rows.
- Classification runs from schema metadata only — schema, table, column name, and data type. No row values are read. A model assigns up to 4 categories per column, each kept at confidence ≥ 50%.
- A daily sweep and schema-change events keep results current; there is no manual trigger. KPIs and the table respect the topbar instance filter and are cached for about 5 minutes.
- The surface is marked Beta.
Reading it
Start at the KPI strip for exposure at a glance, then work the Sensitive Tables list — sorted by risk by default — to triage. The same data is available as a point-in-time PDF under Reports.
Tables With PII
How many classified tables have at least one column flagged as PII.
How it's calculated
- Counts distinct tables with ≥1 column classified into the PII category specifically — the other 7 categories (Financial, Credentials, Payment, Health, Location, Legal, HR) don't count toward this KPI.
- Only PostgreSQL and SQL Server tables are classified, so this is a lower bound on the real estate, not organization-wide coverage.
- Respects the topbar instance filter. Cached for about 5 minutes and refreshed after each successful classification run.
Reading it
A rising count without a matching schema change usually means the daily sweep reached a database for the first time. Cross-check against Sensitive Column Density to see whether PII is concentrated or spread thin.
High-Risk Tables
How many sensitive tables carry a High or Critical computed risk score.
How it's calculated
- Risk = 0.5 × sensitivity + 0.3 × encryption + 0.2 × access, bucketed as Critical (≥0.75), High (≥0.55), Medium (≥0.35), or Low. This KPI counts the Critical and High buckets.
- Sensitivity comes from the table's highest-sensitivity category — Credentials and Payment score highest, then Health, then PII and Financial.
- Encryption and access-level inputs are described under Unencrypted Tables and the Sensitive Tables table. Respects the topbar instance filter and the same 5-minute cache as the other KPIs.
Reading it
A table can be High risk from access alone, even when encrypted — check the row's Access level column before assuming encryption is the gap. Open a row for the full breakdown.
Unencrypted Tables
How many sensitive tables are confirmed unencrypted.
How it's calculated
- Counts sensitive tables (≥1 classified column) whose encryption status is Unencrypted. Tables with status Unknown are excluded from this count, not treated as unencrypted.
- Only SQL Server carries a real encryption value, from the database's transparent data encryption (TDE) flag. Every other engine reports Unknown, so in practice this KPI counts SQL Server tables without TDE.
- Respects the topbar instance filter and the shared 5-minute KPI cache.
Reading it
Zero does not mean everything is encrypted — it can mean nothing sensitive is on SQL Server yet, or that TDE status hasn't been read. Check the list's Encryption column, where PostgreSQL tables always show "—".
Sensitive Column Density
What share of the classified catalog's columns are sensitive.
How it's calculated
- Sensitive columns divided by all live catalog columns, as a percentage to 1 decimal, on PostgreSQL and SQL Server only. The hint below the number shows the raw sensitive-column count.
- Shows "—" when the total can't be read, is 0, or comes back smaller than the sensitive count. An organization with no classifications at all shows 0.0%, not "—".
- This is density across the whole catalog, not coverage. Classification coverage — the share of columns in databases with at least one completed run — is a related but different number shown only in the Data Inventory PDF report.
Reading it
Rising density can mean either newly classified sensitive columns or a shrinking denominator (fewer live catalog columns). Pair with Tables With PII to see whether the change is concentrated in a few tables.
Filters
Narrows the Sensitive Tables list by search text, engine, and schema.
How it's calculated
- Search matches database, table, or instance name, debounced as you type.
- Engine is a multi-select chip; unclassified engines are shown but labelled "Not classified yet".
- Schema is a multi-select chip populated only with schemas already present in the inventory — it can't be used to browse schemas that have no classified table.
- Clear filters resets all of the above. The topbar instance filter and the table's own header-menu filters (Database, Encryption, Risk) apply on top of this row. Filter state is kept in the URL, so a filtered view is shareable.
Reading it
Combine Engine and Schema to scope a review to one part of the estate before sorting the table by Risk. A cleared filter row still respects the topbar instance selection.
Sensitive Tables
Every table with at least one sensitive column, ranked by computed risk.
How it's calculated
- Only tables with ≥1 classified column appear. Default sort is Risk descending; 20 rows per page.
- Columns: Database, Instance, Table, Data types (first 3 categories, then "+N" for the rest), Encryption, Access level, Risk.
- Sortable columns: Database, Instance, Table, Encryption, Risk. Header-menu filters are available on Database, Encryption, and Risk, on top of the Filters row above.
- Access level is per schema — every table in a schema shows the same value, labelled like "High (12 roles)" from distinct roles in the latest permissions snapshot.
Reading it
Click a row to open the table's detail view for its full column-level classification. Sort by Risk to triage, or filter by Encryption to find sensitive tables still missing TDE.