Dashboard: Metric Reference
This page is the source of truth for the in-app Explain this panels on the
organization Dashboard (/). Each section is written once as a
content partial under _explain/ and rendered both here and inside the app's info panel
(scripts/build-explain.mjs compiles the registry).
The topbar instance filter and time window apply to the widgets below as noted in each section. Platform super-admin home is a different surface and is not covered here.
Instances
How many database servers (instances) are connected in this organization right now.
How it's calculated
- Served by
GET /api/v1/health-checks/overviewand overlaid live onto the dashboard response — not read from a cached snapshot of the KPI strip. - Counts every server belonging to the organization. When the topbar instance filter is set, only the selected server IDs are included (
InstanceFilter). - The trend label ("vs. last week") is the difference between the current count and how many of those same servers already existed seven days ago (
CreatedAt ≤ now − 7 days). A new server registered this week raises the trend; removing one that was present last week lowers it. - Independent of the topbar time-window chip — this is a live estate count, not a windowed series.
Reading it
Use this as the inventory baseline for the rest of the page. When the instance filter is on, every other KPI and chart on the dashboard narrows to the same selection, so a lower Instances number should match a quieter Connections and Alerts strip.
Databases
How many monitored databases belong to the selected instances right now.
How it's calculated
- Served by
GET /api/v1/health-checks/overviewand overlaid live onto the dashboard response. - Counted in SQL as the current number of databases for the organization, optionally scoped to the topbar's instance (server) IDs. An empty filter match returns zero rather than falling back to org-wide.
- The trend ("vs. last week") compares that current count to the count as of seven days ago (
created_at ≤ now − 7 daysfor the previous baseline). - Independent of the topbar time-window chip.
Reading it
Databases grow with new registrations and shrink when databases are removed from monitoring. A jump without a matching Instances change usually means more databases were attached to existing servers.
Connections
How many client connections are open across the selected instances at this moment.
How it's calculated
- Served by
GET /api/v1/metrics/connections— a point-in-time snapshot, not a series over the topbar time window. - Each engine (PostgreSQL, SQL Server, Oracle) contributes its latest per-session sample; Logstag merges the engines by summing the state buckets (total, active, idle, waiting, and related splits).
- When the topbar instance filter is set, only those servers' databases are included. The card's hint line breaks the same snapshot into active · idle · waiting.
- The snapshot is independent of the time-window chip: Refresh re-fetches "now", but changing Last 1 hour → Last 24 hours does not redefine this number.
Reading it
Treat this as current load on the estate. Correlate spikes with the Utilization and Wait Events trend tiles below — a high total that is mostly idle is a pool-sizing story; a high total that is mostly active or waiting is a contention story.
Alerts
How many alerts are currently open (actionable) across the selected instances.
How it's calculated
- Served by
GET /api/v1/health-checks/overviewand overlaid live from the alerts table — not derived from health-check findings. - Current count: unresolved, unmuted alerts right now (
statenot Resolved / AutoResolved,is_muted = false). Optionally scoped to the topbar instance filter via each alert's server (through its database, or the alert's ownserver_id). - Trend ("vs. last week"): the same statistic reconstructed as of seven days ago from
created_at/resolved_at(past alert state is not versioned). Positive trend means more open alerts than a week ago. - Independent of the topbar time-window chip — this is "open now", not "fired in the window". For windowed activity, use Alert History and Latest 10 Alerts.
Reading it
Zero is the healthy baseline. When the count is up, open Latest 10 Alerts or the Alerts Explorer to see what is still firing; the Alert History chart below shows how that load arrived over the selected window.
Alert History
Alerts that were firing in each interval versus alerts closed in it, across the selected databases.
How it's calculated
- Served by
GET /api/v1/alerts/raised-resolved-historywith the samefrom/to/bucketSizethe dashboard derives from the topbar time window (presets re-anchor on Refresh; custom ranges keep fixed bounds). - Active (raised): an overlap gauge — every unmuted alert whose firing window
[first_triggered_at, last_triggered_at]overlaps the bucket (edges clamped to the request window). A chronic alert that keeps re-firing appears in every bucket it was active, not only the day its row was created. - Resolved: unmuted alerts whose
resolved_atfalls inside the bucket. - Muted alerts are excluded. Instance scope uses resolved instance names (same filter family as the alerts list).
- Buckets that end at or before
monitoringStartedAt(when the earliest server in scope was registered) are drawn as gaps — they predate monitoring, so a flat zero would be misleading.
Reading it
Read Active as "how much was on fire" and Resolved as "how much was closed" in each interval. A quiet estate shows a flat Active line near zero; a rising Active line with a flat Resolved line means the backlog is building. Hover stays in sync with the other three trend tiles so you can line a spike up with utilization or wait events at the same bucket.
Utilization
CPU, memory and disk usage across every monitored host in the selected window, one line per resource.
How it's calculated
- Served by
GET /api/v1/metrics/utilizationwith the samefrom/to/bucketSize(1 minute / 1 hour / 1 day / 1 week) the dashboard derives from the topbar time window, plus the optional instance filter (empty = whole org). Results are cached for 30 s. - Sourced from VictoriaMetrics host stats reported by the agent, covering every monitored host regardless of database engine.
- CPU (%): each host's CPU usage averaged over the bucket, then averaged across hosts — every host weighted equally.
- Memory (%): same averaging method; used vs. total RAM only — swap is not included.
- Disk (%): the fullest single volume across all hosts at its peak in the bucket (max of max), not an average.
/boot,/boot/efi, and/snap/*are ignored. This is the alerting-relevant figure. - Buckets with no data come back
nulland are drawn as gaps; a real zero stays 0. The in-progress last bucket is a partial average.
Reading it
CPU and Memory are fleet averages, so one saturated host can hide behind several quiet ones; drill into the instance when a line creeps up. Disk is the single fullest volume at its peak, so a rising Disk line means at least one volume is filling up, even if the rest of the estate has room. Hover stays in sync with the other three trend tiles.
Query Performance
Average and worst per-query latency across the selected window.
How it's calculated
- Served by
GET /api/v1/metrics/query-performancewith the dashboard'sfrom/to/bucketSizeand optional instance filter (empty = whole org). Results are cached for 30 s. - Covers PostgreSQL, SQL Server and Oracle only — MongoDB and MySQL are not included.
- Avg Latency (ms): total query execution time ÷ total executions in the bucket, across all engines (execution-weighted, so busier engines pull the average toward themselves).
- Max Latency (ms): the slowest query's average latency in the bucket — the highest per-query mean, not the slowest single execution.
- Per engine: PostgreSQL uses
pg_stat_statementsdeltas, collected every ~60 s. SQL Server uses the Query Store interval max where available, otherwise the per-execution average. Oracle uses per-plan averages from SQL statistics deltas; a statement seen only once in a bucket contributes nothing. - Oracle has no historical rollup, so it only contributes to windows within the last ~7 days.
- Buckets with no data come back
nulland are drawn as gaps; a real zero stays 0. A failed read on one engine silently drops that engine's contribution.
Reading it
Avg Latency tracks overall load; Max Latency tracks the worst offender. A rising Max with a flat Avg usually means one query regressed, not the whole workload. On windows reaching back more than ~7 days, older buckets reflect PostgreSQL and SQL Server only, since Oracle data is no longer available there. Hover stays in sync with the other three trend tiles.
Wait Events
Typical number of sessions waiting at once, by wait category, across the selected databases.
How it's calculated
- Served by
GET /api/v1/metrics/wait-eventswith the dashboard'sfrom/to/bucketSizeand optional instance filter (empty = whole org). Results are cached for 30 s. - Covers PostgreSQL, SQL Server and Oracle only.
- Unit: average number of sessions waiting at once in the bucket, per database, then summed across databases and engines — averaged only over the samples that had at least one waiting session, so rare, bursty waits read higher than a strict time-average would. Sessions running on CPU (no wait) are not counted.
- Sampling cadence: PostgreSQL
pg_stat_activity~every 10 s; SQL Server sessions ~every 30 s; Oracle sessions inWAITINGstate. - Category: Locks — PostgreSQL
Lock,BufferPin; SQL ServerLCK_*; Oracle Concurrency, Application. - Category: Lightweight Locks — PostgreSQL
LWLock; SQL ServerLATCH_*; Oracle Configuration. - Category: I/O — PostgreSQL
IO; SQL ServerPAGEIOLATCH_*,WRITELOG,IO_*,PAGELATCH_*,HADR_*,LOGMGR*; Oracle User I/O, System I/O. - Category: Client / Network — PostgreSQL
Client,IPC; SQL ServerASYNC_NETWORK_IO,*NETWORK*,SNI_*; Oracle Network. - Category: Other — everything else (e.g. PostgreSQL
Activity,Timeout; SQL ServerCXPACKET,SOS_SCHEDULER_YIELD). - Buckets with no data come back
nulland are drawn as gaps; a real zero stays 0. A failed read on one engine silently drops that engine's contribution.
Reading it
Idle PostgreSQL connections wait on Client, so a steady Client / Network band usually means idle pooled connections, not a network problem; idle Oracle sessions ("SQL*Net message from client") land in Other. Look at Locks, Lightweight Locks and I/O for real contention. Hover stays in sync with the other three trend tiles.
Latest 10 Alerts
A compact digest of the ten most recently triggered alerts in the current time window and instance scope.
How it's calculated
- Uses the same alerts list API as the Alerts Explorer (
queryAlerts), limited to page size 10, sorted bylast_triggered_atdescending, with status set to all (including resolved). - Time window: for a preset, only
from = now − preset durationis sent (through now); for a custom range, bothfromandtoare sent. Filtering is onlast_triggered_at, matching the shared topbar window the trend tiles use. - Instance scope: when the topbar filter is set, alerts are narrowed by instance name after ID→name resolution. The digest waits for that resolution so it never briefly shows an unfiltered org-wide list.
- Compact mode hides search, filters, and pagination — View opens the alert's detail page carrying the same window and instance scope.
Reading it
This is a glance, not the full workspace. Use View all for filters, bulk actions, and history charts. If the KPI Alerts count is high but this list looks quiet, the open alerts may last have triggered outside the selected window — widen the chip or open the Alerts Explorer.