Skip to main content

Dashboard: Metric Reference

This page is the source of truth for the in-app Explain this panels on the organization Dashboard (/). Each section is written once as a content partial under _explain/ and rendered both here and inside the app's info panel (scripts/build-explain.mjs compiles the registry).

The topbar instance filter and time window apply to the widgets below as noted in each section. Platform super-admin home is a different surface and is not covered here.

Instances​

How many database servers (instances) are connected in this organization right now.

How it's calculated​

  • Served by GET /api/v1/health-checks/overview and overlaid live onto the dashboard response — not read from a cached snapshot of the KPI strip.
  • Counts every server belonging to the organization. When the topbar instance filter is set, only the selected server IDs are included (InstanceFilter).
  • The trend label ("vs. last week") is the difference between the current count and how many of those same servers already existed seven days ago (CreatedAt ≤ now − 7 days). A new server registered this week raises the trend; removing one that was present last week lowers it.
  • Independent of the topbar time-window chip — this is a live estate count, not a windowed series.

Reading it​

Use this as the inventory baseline for the rest of the page. When the instance filter is on, every other KPI and chart on the dashboard narrows to the same selection, so a lower Instances number should match a quieter Connections and Alerts strip.

Databases​

How many monitored databases belong to the selected instances right now.

How it's calculated​

  • Served by GET /api/v1/health-checks/overview and overlaid live onto the dashboard response.
  • Counted in SQL as the current number of databases for the organization, optionally scoped to the topbar's instance (server) IDs. An empty filter match returns zero rather than falling back to org-wide.
  • The trend ("vs. last week") compares that current count to the count as of seven days ago (created_at ≤ now − 7 days for the previous baseline).
  • Independent of the topbar time-window chip.

Reading it​

Databases grow with new registrations and shrink when databases are removed from monitoring. A jump without a matching Instances change usually means more databases were attached to existing servers.

Connections​

How many client connections are open across the selected instances at this moment.

How it's calculated​

  • Served by GET /api/v1/metrics/connections — a point-in-time snapshot, not a series over the topbar time window.
  • Each engine (PostgreSQL, SQL Server, Oracle) contributes its latest per-session sample; Logstag merges the engines by summing the state buckets (total, active, idle, waiting, and related splits).
  • When the topbar instance filter is set, only those servers' databases are included. The card's hint line breaks the same snapshot into active · idle · waiting.
  • The snapshot is independent of the time-window chip: Refresh re-fetches "now", but changing Last 1 hour → Last 24 hours does not redefine this number.

Reading it​

Treat this as current load on the estate. Correlate spikes with the Utilization and Wait Events trend tiles below — a high total that is mostly idle is a pool-sizing story; a high total that is mostly active or waiting is a contention story.

Alerts​

How many alerts are currently open (actionable) across the selected instances.

How it's calculated​

  • Served by GET /api/v1/health-checks/overview and overlaid live from the alerts table — not derived from health-check findings.
  • Current count: unresolved, unmuted alerts right now (state not Resolved / AutoResolved, is_muted = false). Optionally scoped to the topbar instance filter via each alert's server (through its database, or the alert's own server_id).
  • Trend ("vs. last week"): the same statistic reconstructed as of seven days ago from created_at / resolved_at (past alert state is not versioned). Positive trend means more open alerts than a week ago.
  • Independent of the topbar time-window chip — this is "open now", not "fired in the window". For windowed activity, use Alert History and Latest 10 Alerts.

Reading it​

Zero is the healthy baseline. When the count is up, open Latest 10 Alerts or the Alerts Explorer to see what is still firing; the Alert History chart below shows how that load arrived over the selected window.

Alert History​

Alerts that were firing in each interval versus alerts closed in it, across the selected databases.

How it's calculated​

  • Served by GET /api/v1/alerts/raised-resolved-history with the same from / to / bucketSize the dashboard derives from the topbar time window (presets re-anchor on Refresh; custom ranges keep fixed bounds).
  • Active (raised): an overlap gauge — every unmuted alert whose firing window [first_triggered_at, last_triggered_at] overlaps the bucket (edges clamped to the request window). A chronic alert that keeps re-firing appears in every bucket it was active, not only the day its row was created.
  • Resolved: unmuted alerts whose resolved_at falls inside the bucket.
  • Muted alerts are excluded. Instance scope uses resolved instance names (same filter family as the alerts list).
  • Buckets that end at or before monitoringStartedAt (when the earliest server in scope was registered) are drawn as gaps — they predate monitoring, so a flat zero would be misleading.

Reading it​

Read Active as "how much was on fire" and Resolved as "how much was closed" in each interval. A quiet estate shows a flat Active line near zero; a rising Active line with a flat Resolved line means the backlog is building. Hover stays in sync with the other three trend tiles so you can line a spike up with utilization or wait events at the same bucket.

Utilization​

CPU, memory and disk usage across every monitored host in the selected window, one line per resource.

How it's calculated​

  • Served by GET /api/v1/metrics/utilization with the same from / to / bucketSize (1 minute / 1 hour / 1 day / 1 week) the dashboard derives from the topbar time window, plus the optional instance filter (empty = whole org). Results are cached for 30 s.
  • Sourced from VictoriaMetrics host stats reported by the agent, covering every monitored host regardless of database engine.
  • CPU (%): each host's CPU usage averaged over the bucket, then averaged across hosts — every host weighted equally.
  • Memory (%): same averaging method; used vs. total RAM only — swap is not included.
  • Disk (%): the fullest single volume across all hosts at its peak in the bucket (max of max), not an average. /boot, /boot/efi, and /snap/* are ignored. This is the alerting-relevant figure.
  • Buckets with no data come back null and are drawn as gaps; a real zero stays 0. The in-progress last bucket is a partial average.

Reading it​

CPU and Memory are fleet averages, so one saturated host can hide behind several quiet ones; drill into the instance when a line creeps up. Disk is the single fullest volume at its peak, so a rising Disk line means at least one volume is filling up, even if the rest of the estate has room. Hover stays in sync with the other three trend tiles.

Query Performance​

Average and worst per-query latency across the selected window.

How it's calculated​

  • Served by GET /api/v1/metrics/query-performance with the dashboard's from / to / bucketSize and optional instance filter (empty = whole org). Results are cached for 30 s.
  • Covers PostgreSQL, SQL Server and Oracle only — MongoDB and MySQL are not included.
  • Avg Latency (ms): total query execution time ÷ total executions in the bucket, across all engines (execution-weighted, so busier engines pull the average toward themselves).
  • Max Latency (ms): the slowest query's average latency in the bucket — the highest per-query mean, not the slowest single execution.
  • Per engine: PostgreSQL uses pg_stat_statements deltas, collected every ~60 s. SQL Server uses the Query Store interval max where available, otherwise the per-execution average. Oracle uses per-plan averages from SQL statistics deltas; a statement seen only once in a bucket contributes nothing.
  • Oracle has no historical rollup, so it only contributes to windows within the last ~7 days.
  • Buckets with no data come back null and are drawn as gaps; a real zero stays 0. A failed read on one engine silently drops that engine's contribution.

Reading it​

Avg Latency tracks overall load; Max Latency tracks the worst offender. A rising Max with a flat Avg usually means one query regressed, not the whole workload. On windows reaching back more than ~7 days, older buckets reflect PostgreSQL and SQL Server only, since Oracle data is no longer available there. Hover stays in sync with the other three trend tiles.

Wait Events​

Typical number of sessions waiting at once, by wait category, across the selected databases.

How it's calculated​

  • Served by GET /api/v1/metrics/wait-events with the dashboard's from / to / bucketSize and optional instance filter (empty = whole org). Results are cached for 30 s.
  • Covers PostgreSQL, SQL Server and Oracle only.
  • Unit: average number of sessions waiting at once in the bucket, per database, then summed across databases and engines — averaged only over the samples that had at least one waiting session, so rare, bursty waits read higher than a strict time-average would. Sessions running on CPU (no wait) are not counted.
  • Sampling cadence: PostgreSQL pg_stat_activity ~every 10 s; SQL Server sessions ~every 30 s; Oracle sessions in WAITING state.
  • Category: Locks — PostgreSQL Lock, BufferPin; SQL Server LCK_*; Oracle Concurrency, Application.
  • Category: Lightweight Locks — PostgreSQL LWLock; SQL Server LATCH_*; Oracle Configuration.
  • Category: I/O — PostgreSQL IO; SQL Server PAGEIOLATCH_*, WRITELOG, IO_*, PAGELATCH_*, HADR_*, LOGMGR*; Oracle User I/O, System I/O.
  • Category: Client / Network — PostgreSQL Client, IPC; SQL Server ASYNC_NETWORK_IO, *NETWORK*, SNI_*; Oracle Network.
  • Category: Other — everything else (e.g. PostgreSQL Activity, Timeout; SQL Server CXPACKET, SOS_SCHEDULER_YIELD).
  • Buckets with no data come back null and are drawn as gaps; a real zero stays 0. A failed read on one engine silently drops that engine's contribution.

Reading it​

Idle PostgreSQL connections wait on Client, so a steady Client / Network band usually means idle pooled connections, not a network problem; idle Oracle sessions ("SQL*Net message from client") land in Other. Look at Locks, Lightweight Locks and I/O for real contention. Hover stays in sync with the other three trend tiles.

Latest 10 Alerts​

A compact digest of the ten most recently triggered alerts in the current time window and instance scope.

How it's calculated​

  • Uses the same alerts list API as the Alerts Explorer (queryAlerts), limited to page size 10, sorted by last_triggered_at descending, with status set to all (including resolved).
  • Time window: for a preset, only from = now − preset duration is sent (through now); for a custom range, both from and to are sent. Filtering is on last_triggered_at, matching the shared topbar window the trend tiles use.
  • Instance scope: when the topbar filter is set, alerts are narrowed by instance name after ID→name resolution. The digest waits for that resolution so it never briefly shows an unfiltered org-wide list.
  • Compact mode hides search, filters, and pagination — View opens the alert's detail page carrying the same window and instance scope.

Reading it​

This is a glance, not the full workspace. Use View all for filters, bulk actions, and history charts. If the KPI Alerts count is high but this list looks quiet, the open alerts may last have triggered outside the selected window — widen the chip or open the Alerts Explorer.