Server Metrics Pipeline
This is the implementation reference for historical server metrics: CPU, RAM, disk, files, and process counts. All historical timestamps and aggregation windows are UTC.
Short version
configured node API
-> servers_metrics_collect.php (every 5 min)
-> server_metrics_raw
-> servers_metrics_aggregate.php (every 15 min)
-> server_metrics_hourly and server_metrics_daily
-> /api/metrics/{server}/{metric}
-> shared server-card modal
The current server card and the historical modal are different views:
- The server card shows the latest data fetched/cached for the dashboard.
- The modal shows rows persisted by the historical pipeline below.
- A current card value does not prove that historical rows exist. A reachable node and a successful collection run are required for history.
Schedules and programs
The installed manifest is
docker/cron/crontab.
Cron runs as www-data in Docker. Shared-hosting installations can use the
shared-hosting crontab template.
| When | Program | Purpose | Writes |
|---|---|---|---|
| Every 5 min | cron/servers_data_fetch.php |
Dashboard/node collector setup | not the historical tables |
| Every 5 min | cron/servers_metrics_collect.php |
Fetch one snapshot per server not marked disabled |
server_metrics_raw |
| Every 15 min | cron/servers_metrics_aggregate.php |
Repair bounded completed hourly and daily windows | server_metrics_hourly, server_metrics_daily |
| Daily, 02:25 | cron/retention_cleanup.php |
Delete old rows according to retention settings | deletes from all metric tiers |
servers_metrics_collect.php skips configuration keys beginning with _ and
servers with disabled set to true. Each remaining server needs api_url
and api_key in its server configuration.
Cron has a restricted environment. Container startup writes only the required
DB_* values to /etc/mata/cron.env; each cron command sources it before PHP
bootstrap. Check that exact context with:
docker exec <web-container> mata-cron-smoke-check
1. Collection: node response to raw snapshot
cron/servers_metrics_collect.php creates an authenticated NodeApiClient, then
uses ServerDataCollector and ServerMonitor to fetch server health data from
the configured node API.
buildSnapshotFromPayload() maps the node response into
ServerMetricRawSnapshot:
| Raw snapshot value | Node response source | Notes |
|---|---|---|
cpu_usage |
cpu_usage |
Percent |
ram_used, ram_total |
ram_used / ram_capacity, or RAM/memory alternatives |
KiB |
ram_usage_percent |
calculated as ram_used / ram_total * 100 |
Null if total is absent or zero |
disk_used, disk_total |
disk/storage alternatives | KiB |
disk_usage_percent |
calculated as disk_used / disk_total * 100 |
Null if total is absent or zero |
total_file_count |
total_system_filecount or file-system alternatives |
Count |
process_count, php_process_count |
processes.* or alternatives |
Count |
php_version, os_version, server_uptime |
corresponding node data | Metadata |
collected_at |
node timestamp; otherwise cache timestamp adjusted by cache age, then collection time | UTC |
ingested_at |
dashboard insertion time | UTC |
The collector identity is read from data.collector when present; otherwise it
uses configured fallback values. ServerMetricsRepository resolves/creates
servers.id and metric_collectors.id, normalizes non-negative counts and
0–100 percentages, then uses ServerMetricRawUpserter for the database upsert.
Duplicate protection is database-enforced by unique (server_id, collector_id,
collected_at) and (server_id, payload_hash, collected_at) keys. A duplicate
updates the existing row; it does not add another point.
Collection health is recorded in:
logs/cron/app-servers-metrics-collect.loglogs/health/collect_server_metrics_health.loglogs/cron/heartbeat-servers-metrics-collect.jsonl
The health report contains per-server inserted count, latest sample timestamp, and ingestion lag.
2. Raw table
server_metrics_raw is the detailed tier. Expected density is about twelve
points per hour when collection succeeds every five minutes.
| Columns |
|---|
id, server_id, collector_id |
collected_at, ingested_at, payload_hash |
cpu_usage |
ram_used, ram_total, ram_usage_percent |
disk_used, disk_total, disk_usage_percent |
total_file_count, process_count, php_process_count |
php_version, os_version, server_uptime |
Most metric columns are nullable. NULL means the node did not provide that
measurement. It is not equivalent to zero.
3. Aggregation
cron/servers_metrics_aggregate.php is orchestration only. It creates:
AggregationWindowPlanner: completed UTC windows.ServerMetricsAggregator: SQL statistics for one server/window.ServerMetricsRepository: id resolution and upserts.
Hourly planning
With the default configuration (lag_minutes: 5, catch_up_hours: 24), every
15-minute run plans the last 24 completed UTC hours. It intentionally
upserts them again. This bounded overlap repairs late-arriving snapshots and
missed aggregation runs.
For an hour starting at 10:00 UTC, raw rows are selected with:
collected_at >= 10:00:00 UTC
collected_at < 11:00:00 UTC
The start is inclusive and end is exclusive. No raw rows means no hourly row; the job does not write a fake zero aggregate.
Hourly aggregates compute AVG, MIN, and/or MAX for the available metric,
plus sample_count, first_sample, and last_sample.
Daily planning
Daily planning does not depend on an hourly row being inserted in the same run.
With lag_minutes: 30 and configured catch_up_days: 7, it upserts
recent completed UTC days after the 30-minute post-midnight delay.
A daily aggregate normally reads raw rows for:
collected_at >= <day-start> 00:00:00 UTC
collected_at < <next-day> 00:00:00 UTC
If raw retention has already removed the entire day, it falls back to hourly
rows. Average values use SUM(hourly_average * sample_count) / SUM(sample_count);
it never averages averages without weights. Daily critical-event and warning-event counts are loaded separately from
events; they are not counts of logical alarms. If that query fails, metric
values are still saved and the failure is logged.
Aggregate tables
Both tables use a unique server/window key and INSERT ... ON DUPLICATE KEY
UPDATE, so a rerun updates the same row.
| Table | Window key | Columns |
|---|---|---|
server_metrics_hourly |
(server_id, hour_timestamp) |
cpu_avg/min/max; ram_used_avg/min/max; ram_percent_avg/max; disk_used_avg/min/max; disk_percent_avg/max; file_count_avg/max; process_count_avg/max; sample_count; first_sample; last_sample |
server_metrics_daily |
(server_id, date) |
The same CPU/RAM/disk/file/process aggregate fields; sample_count; alarm_count; warning_count |
Aggregation logs to logs/cron/app-servers-metrics-aggregate.log and writes
heartbeat-servers-metrics-aggregate.jsonl.
4. Retention
Aggregation does not delete data. cron/retention_cleanup.php owns deletion.
Settings are in the active data_retention.json; the example is
config/example/data_retention.json.example.
Default retention:
| Tier | Default |
|---|---|
| Raw | 30 days |
| Hourly | 90 days |
| Daily | 365 days |
Set a retention value to 0 to disable cleanup for that tier. Do not set raw
retention shorter than the daily aggregation/recovery window unless relying on
the hourly weighted fallback is acceptable.
5. Reading and displaying history
API read path
MetricsController::getServerMetricSeries() handles:
GET /api/metrics/{server}/{metric}?range=24h&granularity=hourly
ServerMetrics\MetricsRequestValidatorvalidates HTTP query inputs usingMetricCatalogand sharedMetricQueryPolicynormalization/range rules.MetricsCacheServicechecks the response cache unlessforce=1.MetricsViewBuilderpicksServerMetricsQuery::getRawMetrics,getHourlyMetrics, orgetDailyMetrics.ServerMetricsQuerymaps database rows to DTOs.MetricsViewBuilderextracts the requested field, labels it in UTC, and calculates min, max, and average from non-null values.MetricsRendererreturns JSON.
Queries are ordered ascending and use inclusive start/exclusive end boundaries.
A valid request with no rows returns HTTP 200 with empty labels/data. A null
measurement remains JSON null.
Supported API metric keys:
| Key | Source by tier | Unit |
|---|---|---|
cpu |
cpu_usage / cpu_avg |
% |
ram_percent |
ram_usage_percent / ram_percent_avg |
% |
ram_used |
ram_used / ram_used_avg |
kilobytes |
disk_percent |
disk_usage_percent / disk_percent_avg |
% |
disk_used |
disk_used / disk_used_avg |
kilobytes |
file_count |
total_file_count / file_count_avg |
count |
process_count |
process_count / process_count_avg |
count |
ram_used, disk_used |
raw value / matching *_avg |
KiB |
Shared server-card modal
templates/partials/cards/server-stats.html.twig turns CPU, RAM, disk, and
file-count card stats into buttons. Each opens:
GET /watch/modal/{server}/metrics?metric=<key>&range=24h
MetricsController::renderServerMetricsModal() only accepts cpu,
ram_percent, disk_percent, and file_count. The route renders
templates/partials/modals/server-metrics-modal.html.twig once.
public/assets/js/server-metrics-modal.js handles selectors, refresh, request
cancellation, loading/error/empty/partial states, Chart.js lifecycle, and
unit formatting. It fetches the canonical API endpoint when a selector changes.
The modal uses a fixed tier mapping:
| UI range | Table read |
|---|---|
| 6 hours | server_metrics_raw |
| 24 hours | server_metrics_hourly |
| 7 days | server_metrics_hourly |
| 30 days | server_metrics_daily |
Process count deliberately remains a non-interactive card statistic. Cronjobs have a separate existing modal.
Operations and diagnosis
Run these inside the web container:
# Row count and latest timestamp for each tier
php /var/www/html/scripts/server_metrics_diagnostic.php
# Verify cron-user config and database bootstrap
mata-cron-smoke-check
For per-server detail, use SQL:
SELECT s.name, COUNT(*) AS raw_rows, MAX(r.collected_at) AS latest_raw
FROM servers s
LEFT JOIN server_metrics_raw r ON r.server_id = s.id
GROUP BY s.id, s.name
ORDER BY s.name;
SELECT s.name,
MAX(h.hour_timestamp) AS latest_hourly,
MAX(d.date) AS latest_daily
FROM servers s
LEFT JOIN server_metrics_hourly h ON h.server_id = s.id
LEFT JOIN server_metrics_daily d ON d.server_id = s.id
GROUP BY s.id, s.name
ORDER BY s.name;
If raw rows stay at zero, inspect collection logs first. Common causes are a
disabled server configuration, missing api_url/api_key, node DNS/network
failure, node authentication failure, or cron database bootstrap failure.