Skip to main content

Insights: Metric Reference

This page is the source of truth for the in-app Explain this panels on the Insights page (/insights). Each section is written once as a content partial under _explain/insights/ and rendered both here and inside the app's info panel (scripts/build-explain.mjs compiles the registry).

The KPIs follow only the topbar instance and time filters; the status tabs and chip filters narrow the list alone. See the Insights guide for detector families and engine coverage.

Findings that background detectors raise when a database drifts from its own recent baseline: a KPI strip, a Confidence mix bar, and one filterable list.

How it's calculated​

  • Detectors run about every 15 minutes per database, using metrics Logstag has already collected. They never open new connections to the monitored database.
  • Three detector families:
    • N+1 pattern: a query called many times that returns about one row per call. PostgreSQL, SQL Server, and Oracle only.
    • Regression: slower queries on PostgreSQL, SQL Server, and Oracle; slower collection operations on MongoDB; slower or failing commands on Redis and Valkey.
    • Workload anomaly: the database's overall workload shape has changed. Runs on all six engines.
  • Every finding is compared against the database's own history, not a global threshold. Low-severity findings are not kept, so the list shows Medium, High, and Critical findings.
  • The KPIs and the list follow the topbar instance and time filters. A finding is in range if it was active at any point in the selected window, not only if it was first detected there. Results are cached for about 5 minutes; dismissing a finding refreshes them.
  • The surface is marked Beta.

Reading it​

Use the KPI strip to see how much is happening and how much of it you can trust. Then work the list from the top: open a finding to see what changed, the recommended fix, and the evidence behind it.

Databases at risk​

How many distinct databases have at least one finding in the current scope.

How it's calculated​

  • Counts distinct databases across all findings that match the topbar instance and time filters.
  • Counts every state (Active, Dismissed, and No longer detected), and ignores the status tabs and the chip filters below. Switching tabs does not change this number.
  • A finding counts toward the selected window if it was active at any point in it.

Reading it​

This is the blast radius, not a severity score: one database with ten findings counts once. Filter the list to Active to see which of these databases still have open work.

New​

Findings first detected inside the selected time range.

How it's calculated​

  • Counts findings whose first detection falls inside the topbar time window. A finding that started earlier and is still firing is in scope for the page, but it is not New.
  • If no window is applied, this falls back to findings first detected in the last 7 days. The card's hint says which definition is in use.
  • Recurrence is ignored: a finding first detected in the window that has already fired several times is still New.
  • Uses the same definition as the New badge on list rows, and ignores the status tabs and chip filters.

Reading it​

A spike in New after a deploy or a configuration change is the signal to look at first. Filter by Category to see whether the new findings are query-level regressions or database-wide workload changes.

Recurring​

Findings that have been detected more than once.

How it's calculated​

  • Counts findings in the current scope with more than one occurrence. Each detector pass that sees the same problem again on the same database adds one occurrence and updates Last seen.
  • A finding that stops firing moves to No longer detected after about 45 minutes. If it comes back within 7 days, the same finding is reactivated and its count keeps growing. After 7 days it returns as a new finding with a count of 1.
  • Ignores the status tabs and chip filters.

Reading it​

Recurring findings are persistent problems, not one-off spikes. Sort the list by Most recurring to see them first. The row chip N× · time ago shows how often a finding fired and when it last did.

High-confidence​

Findings rated High or Very high confidence.

How it's calculated​

  • Sums the High and Very high findings in the current scope. The Confidence bar under the KPI strip shows the full mix: Very high, High, Medium, and Low.
  • Confidence measures how much data backs a finding, not how bad it is (that is severity):
    • N+1: how consistently the pattern appeared across the last 15 minutes. At least 90% of minutes gives Very high; at least 70% gives High.
    • Regression: how many hours of baseline history the comparison had. At least 100 hours gives Very high; at least 40 gives High. A perfectly flat baseline caps it at Medium.
    • Workload anomaly: at most High, which needs at least 4 samples from the same weekday and hour. It is never Very high.
  • Confidence is recalculated on every detection pass, so it can go up or down as history builds.

Reading it​

Start with high-confidence findings: they are the least likely to be noise. Low-confidence findings on a new or recently reset database usually firm up once more baseline history exists.

Category​

Narrows the list to one or more finding categories.

How it's calculated​

  • Multi-select chips. Selecting several shows findings in any of them, and no selection shows every category. Changing the selection returns to page 1.
  • Two categories have detectors today:
    • Query lifecycle: N+1 patterns and regressions (queries, MongoDB collection operations, Redis and Valkey commands).
    • Workload anomaly: changes in a database's overall workload shape, such as read/write mix, connections, or cache behavior.
  • Query cost, Schema health, Security posture, and Fleet pattern are reserved for upcoming detectors and currently return no findings.
  • Applies to the list only. The KPI strip ignores it.

Reading it​

Use Query lifecycle to find specific queries or commands to fix, and Workload anomaly for database-wide shifts that usually trace back to a deploy, a traffic change, or a new job. Combine it with the Severity, Confidence, and Database filters in the same row.

Insights list​

Every finding in scope, one row each, filtered by status tab and chips and sorted by most recent activity by default.

How it's calculated​

  • Status tabs:
    • Active: still firing.
    • Dismissed: closed by a user.
    • No longer detected: the detector has not seen the problem for about 45 minutes. This does not mean anyone confirmed a fix.
    • Expired: reserved; no finding moves to this state today.
    • All: the default.
  • Each row shows a title, the database and instance, severity (Critical, High, or Medium), confidence, a New badge if the finding was first detected in the selected window, and N× · time ago once it has recurred.
  • Expanding a row shows:
    • What changed: baseline versus current values.
    • Why: likely causes, such as data growth, a plan change, or lock contention.
    • Recommended fix: numbered steps.
    • Evidence: the metrics behind the finding.
    • First seen, Last seen, and Occurrences.
  • Dismiss (Active rows only) asks for an optional reason, up to 500 characters. The reason and time appear on the row afterwards. A dismissed finding stays silent for about 30 days even if the detector sees it again. There is no undo. Dismissing changes nothing in the monitored database.
  • Dismissed and No longer detected findings are removed after 30 days.

Reading it​

Expand a finding before you act on it. The Evidence and What changed sections show whether the baseline was solid. Dismiss with a reason when a finding is expected, such as a planned migration, so the next reviewer knows why it was closed.