1
0
mirror of git://f0xx.org/ac/ac-docs synced 2026-07-29 05:59:06 +03:00

new DR and spec for cluster monitoring, docs updated

This commit is contained in:
Anton Afanasyeu
2026-06-26 19:47:17 +02:00
parent df726f5993
commit aab2b8512d
53 changed files with 5505 additions and 737 deletions

View File

@@ -0,0 +1,830 @@
# Infrastructure health monitoring — specification
<!-- doc-meta:start -->
| Field | Value |
|---|---|
| Author | Anton Afanasyeu |
| Revision | R0 |
| Creation date | 2026-06-26 |
| Last modification date | 2026-06-26 |
| Co-authored | Cursor Agent (project assistant) |
| Severity | high |
| State | accepted |
| Document type | spec |
| Supersedes | — |
| Design review | [docs/DRs/20260626_cluster_monitoring.md](../DRs/20260626_cluster_monitoring.md) (frozen R0) |
<!-- doc-meta:end -->
---
**Document type:** SPEC
**Source draft:** [docs/drafts/20260626_cluster_monitoring.txt](../drafts/20260626_cluster_monitoring.txt)
**Design review:** [docs/DRs/20260626_cluster_monitoring.md](../DRs/20260626_cluster_monitoring.md) (frozen)
**PDF:** [20260626_cluster_monitoring.pdf](20260626_cluster_monitoring.pdf) · Regenerate: `bash ac-docs/build-pdf.sh`
**Status:** **Accepted** — implement per §7 deployment plan
**Scope:** Passive infrastructure monitoring and alerting for XEN cluster cast01cast04; Track A only
**Related:** [INFRA.md](../INFRA.md) · [DRs/20260626_cluster_monitoring.md](../DRs/20260626_cluster_monitoring.md) · [specs/20100612_1_scaling.md](20100612_1_scaling.md) · [20260620_postgres_timescale_db_platform.md](../20260620_postgres_timescale_db_platform.md)
**Track B (future):** [drafts/20260626_cluster_orchestration.txt](../drafts/20260626_cluster_orchestration.txt)
**Documentation index:** [README.md](../README.md)
---
## Table of contents
<!-- toc -->
- [1. Purpose](#1-purpose)
- [2. Scope](#2-scope)
- [3. Requirements](#3-requirements)
- [3.1 Must](#31-must)
- [3.2 Should](#32-should)
- [3.3 Non-goals](#33-non-goals)
- [4. Monitoring domains](#4-monitoring-domains)
- [4.1 Hardware metrics](#41-hardware-metrics)
- [4.2 Software and service metrics](#42-software-and-service-metrics)
- [4.3 External availability](#43-external-availability)
- [4.4 Provider-layer fault escalation](#44-provider-layer-fault-escalation)
- [5. Topology — cast04 as dedicated monitoring node](#5-topology--cast04-as-dedicated-monitoring-node)
- [6. Component catalog](#6-component-catalog)
- [6.1 Alpine apk packages (cast04)](#61-alpine-apk-packages-cast04)
- [6.2 Alpine apk packages (cast01cast03)](#62-alpine-apk-packages-cast0103)
- [6.3 Docker-based (cast04)](#63-docker-based-cast04)
- [6.4 External SaaS (no install)](#64-external-saas-no-install)
- [7. Deployment plan](#7-deployment-plan)
- [7.1 Phase 0 — cast04 provisioning](#71-phase-0--cast04-provisioning)
- [7.2 Phase 1 — Prometheus + node_exporter](#72-phase-1--prometheus--node_exporter)
- [7.3 Phase 2 — Service probes and MariaDB exporter](#73-phase-2--service-probes-and-mariadb-exporter)
- [7.4 Phase 3 — Alertmanager](#74-phase-3--alertmanager)
- [7.5 Phase 4 — Grafana dashboards](#75-phase-4--grafana-dashboards)
- [7.6 Phase 5 — Uptime Kuma + UptimeRobot](#76-phase-5--uptime-kuma--uptimerobot)
- [7.7 ac-deploy integration](#77-ac-deploy-integration)
- [8. Alert routing](#8-alert-routing)
- [8.1 Alert severity levels](#81-alert-severity-levels)
- [8.2 Routing rules](#82-routing-rules)
- [8.3 Alertmanager SMTP config (skeleton)](#83-alertmanager-smtp-config-skeleton)
- [8.4 Telegram receiver config (skeleton)](#84-telegram-receiver-config-skeleton)
- [8.5 Alert rule catalog](#85-alert-rule-catalog)
- [9. Grafana dashboard catalog](#9-grafana-dashboard-catalog)
- [10. Pre-requisite accounts](#10-pre-requisite-accounts)
- [11. TimescaleDB relationship](#11-timescaledb-relationship)
- [12. Verification and acceptance](#12-verification-and-acceptance)
- [13. Risks and mitigations](#13-risks-and-mitigations)
- [14. Related documents](#14-related-documents)
- [15. Changelog](#15-changelog)
<!-- /toc -->
---
## 1. Purpose
Define the **normative** implementation of passive infrastructure monitoring and alerting for the AndroidCast XEN-based VM cluster. The system must:
- Record hardware and software health metrics from all cluster nodes continuously
- Alert operators promptly when measured values exceed defined thresholds
- Provide visual dashboards for trend analysis and incident investigation
- Detect external (provider-layer) unavailability via an independent probe
This SPEC covers **Track A only** (observe, record, alert). Automated cluster remediation is Track B and is explicitly out of scope — see [drafts/20260626_cluster_orchestration.txt](../drafts/20260626_cluster_orchestration.txt).
---
## 2. Scope
| In scope | Out of scope |
|----------|--------------|
| Hardware metrics: CPU, RAM, disk, NIC per node | Automated node provisioning / deprovisioning |
| Software metrics: process health, HTTP probes, MariaDB | Log aggregation (Grafana Loki — planned R1) |
| Alert routing: email + Telegram | Auto-scaling or auto-failover reactions |
| External uptime probe (UptimeRobot) | Changes to FE nginx or router software |
| Grafana dashboards (infra layer) | Business / application metrics (→ TimescaleDB) |
| cast04 as dedicated monitoring VM | Kubernetes / service mesh |
| ac-deploy integration (monitoring role) | SaaS external aggregator for metric data |
| Pre-requisite account setup guide | iOS / mobile performance profiling |
**Alpha constraint:** Monitoring stack MAY be deployed incrementally per §7 phases. Each phase MUST leave prior phases operational.
---
## 3. Requirements
### 3.1 Must
| ID | Requirement |
|----|-------------|
| R-MON-1 | cast04 dedicated monitoring VM provisioned as Alpine LTS paravirt XEN guest |
| R-MON-2 | node_exporter running on cast01, cast02, cast03, cast04 |
| R-MON-3 | Prometheus scraping all node_exporters on 15s interval |
| R-MON-4 | blackbox_exporter probing all critical HTTP endpoints (§4.2) from cast04 |
| R-MON-5 | mysqld_exporter running on MariaDB primary node |
| R-MON-6 | Alertmanager routing CRITICAL alerts to email within 2 min of threshold breach |
| R-MON-7 | Grafana serving dashboards on cast04 (internal access via SSH tunnel at R0) |
| R-MON-8 | UptimeRobot external HTTP probe active for `apps.f0xx.org` |
| R-MON-9 | All software installed via Alpine apk or Docker (no source compilation) |
| R-MON-10 | No new software installed on FE VM or router |
| R-MON-11 | Metric data retained locally on cast04 for minimum 90 days |
| R-MON-12 | Alert for CPU steal time > 10% sustained 5 min (Xen overcommit indicator → notify Anton) |
| R-MON-13 | Alert for disk write latency > 100ms sustained 5 min on MariaDB node |
| R-MON-14 | Alert for MariaDB replication lag > 30s |
| R-MON-15 | Hardware guarantee baseline alerts (§4.1 table) — notify when measured < guaranteed |
### 3.2 Should
| ID | Requirement |
|----|-------------|
| R-MON-S1 | Alertmanager Telegram receiver configured alongside email |
| R-MON-S2 | Uptime Kuma status page deployed on cast04 (Docker) |
| R-MON-S3 | Grafana provisioned with community dashboard for node_exporter (ID 1860) |
| R-MON-S4 | Grafana provisioned with community dashboard for MariaDB |
| R-MON-S5 | Prometheus metric retention configurable via `--storage.tsdb.retention.time` |
| R-MON-S6 | Alertmanager grouping configured to suppress alert storms (group_wait 30s, group_interval 5m) |
| R-MON-S7 | ac-deploy `monitoring/` role added to cluster0 scripts for reproducible deployment |
### 3.3 Non-goals
- Automated node scale-out or scale-in (Track B)
- DDoS automatic mitigation (Track B)
- Storage automatic failover (Track B)
- Centralized log search (Grafana Loki — R1)
- Business metrics dashboards (TimescaleDB + Grafana — separate track)
- Monitoring of router or FE internal metrics (config-only access; UptimeRobot covers external availability)
---
## 4. Monitoring domains
### 4.1 Hardware metrics
Collected by **node_exporter** on every VM (cast01cast04). Alert thresholds are guidelines — tune after collecting baseline.
| Metric | Prometheus metric name | WARNING threshold | CRITICAL threshold | Xen note |
|--------|----------------------|-------------------|--------------------|-----------|
| CPU utilization | `rate(node_cpu_seconds_total{mode!="idle"}[5m])` | > 80% (5 min) | > 95% (5 min) | |
| **CPU steal time** | `rate(node_cpu_seconds_total{mode="steal"}[5m])` | > 5% (5 min) | > 10% (5 min) | Xen overcommit → notify Anton |
| Load average 1m | `node_load1 / node_cpu_count` | > 1.5 | > 3.0 | |
| RAM used % | `1 - node_memory_MemAvailable_bytes / node_memory_MemTotal_bytes` | > 80% | > 90% | |
| Swap used % | `node_memory_SwapFree_bytes / node_memory_SwapTotal_bytes` | > 50% | > 80% | |
| Disk used % | `1 - node_filesystem_avail_bytes / node_filesystem_size_bytes` | > 75% | > 90% | |
| Disk write latency | `rate(node_disk_write_time_seconds_total[5m]) / rate(node_disk_writes_completed_total[5m])` | > 50ms | > 100ms | Critical on MariaDB node |
| Disk IOPS | `rate(node_disk_io_time_seconds_total[5m])` | > 80% | > 95% | |
| NIC ingress bytes/s | `rate(node_network_receive_bytes_total[5m])` | > 150 Mbit/s | > 190 Mbit/s | Guaranteed: 200 Mbit/s |
| NIC egress bytes/s | `rate(node_network_transmit_bytes_total[5m])` | > 150 Mbit/s | > 190 Mbit/s | Guaranteed: 200 Mbit/s |
| NIC packet drops | `rate(node_network_receive_drop_total[5m])` | > 0.1% | > 1% | |
| NIC errors | `rate(node_network_receive_errs_total[5m])` | > 0 | > 10/min | |
**Hardware guarantee baseline check:** If `node_memory_MemTotal_bytes` < 7.5 GB, or CPU core count < 8, or disk total < 30 GB → CRITICAL alert (guaranteed resources not provided → notify Anton).
### 4.2 Software and service metrics
#### Process / systemd health (node_exporter `systemd` collector or blackbox TCP probe)
| Service | Check | Alert trigger |
|---------|-------|---------------|
| nginx | TCP :80 probe | Connection refused → CRITICAL |
| php-fpm | TCP socket probe | No response → CRITICAL |
| MariaDB | TCP :3306 probe | Connection refused → CRITICAL |
| sshd | TCP :22 probe | Connection refused → WARNING |
| Per-microservice HTTP endpoint | HTTP GET /health or / | Non-2xx or timeout > 5s → WARNING; > 30s → CRITICAL |
#### MariaDB metrics (mysqld_exporter)
| Metric | Prometheus metric | WARNING | CRITICAL |
|--------|------------------|---------|---------|
| Replication lag | `mysql_slave_status_seconds_behind_master` | > 10s | > 30s |
| Open connections | `mysql_global_status_threads_connected` | > 80% of max | > 95% of max |
| Slow queries/s | `rate(mysql_global_status_slow_queries[5m])` | > 1/s | > 5/s |
| InnoDB buffer pool hit | `mysql_global_status_innodb_buffer_pool_reads / mysql_global_status_innodb_buffer_pool_read_requests` | < 95% | < 90% |
| Aborted connections | `rate(mysql_global_status_aborted_connects[5m])` | > 1/min | > 10/min |
#### nginx metrics (via node_exporter nginx stub or blackbox HTTP probe)
| Metric | Source | WARNING | CRITICAL |
|--------|--------|---------|---------|
| 5xx response rate | blackbox HTTP probe returning 5xx | Any 5xx returned | > 10% of requests |
| Response time | blackbox `probe_duration_seconds` | > 2s | > 10s |
#### php-fpm metrics (via php-fpm status endpoint + blackbox or custom exporter)
| Metric | WARNING | CRITICAL |
|--------|---------|---------|
| Active workers % of max | > 70% | > 90% |
| Queue length | > 5 | > 20 |
### 4.3 External availability
**Tool:** UptimeRobot free tier — no installation; checks from UptimeRobot's global nodes.
| Monitor | URL | Check interval | Alert on |
|---------|-----|---------------|---------|
| Frontend TLS | `https://apps.f0xx.org` | 5 min | Non-2xx or timeout |
| API health | `https://apps.f0xx.org/app/androidcast_project/` | 5 min | Non-2xx or timeout |
**Internal cross-check:** blackbox_exporter on cast04 also probes the FE endpoint from inside the cluster. If UptimeRobot fires but blackbox does NOT → provider/uplink fault (Anton's layer). If both fire → likely cluster-side issue.
### 4.4 Provider-layer fault escalation
When monitoring detects that the root cause is in Anton's infrastructure layer:
| Symptom | Likely cause | Action |
|---------|-------------|--------|
| CPU steal > 10% sustained | Xen host overcommit | Email Anton; await response |
| Disk write latency persistent despite low I/O | Xen host disk contention | Email Anton |
| NIC saturation below 200 Mbit guarantee | Uplink / Xen host NIC | Email Anton |
| UptimeRobot alert + cast04 unreachable | Xen host or uplink down | Email Anton; check intra-network |
| Entire cluster unreachable | Router / uplink / payment | Email Anton; check calendar |
These events are **alert-only** at R0. No automated remediation.
---
## 5. Topology — cast04 as dedicated monitoring node
```text
10.7.0.0/8 internal network
┌────────────────────────────────────────┐
│ cast04 — monitoring VM (Alpine LTS) │
│ │
cast01 ──[ :9100 node_exporter ]──┤ │
cast02 ──[ :9100 node_exporter ]──┤──► Prometheus :9090 │
cast03 ──[ :9100 node_exporter ]──┤ ├── scrape all :9100 │
cast04 ──[ :9100 node_exporter ]──┘ ├── scrape :9104 (mysqld) │
MariaDB node ─[ :9104 mysqld_exp ] └── evaluate alert rules │
│ │
▼ │
Alertmanager :9093 │
├── email (SMTP relay) │
└── Telegram bot API │
blackbox_exporter :9115 │
└── HTTP/TCP probes → cast0103, │
FE endpoint, per-service /health │
Grafana :3000 │
├── datasource: Prometheus │
└── datasource: TimescaleDB (future) │
Uptime Kuma :3001 (Docker) │
└── per-service status page │
└───────────────────────────────────────┘
External (no install):
UptimeRobot ──► HTTP probe ──► apps.f0xx.org (from internet)
```
**cast04 minimum provisioning request (to Anton):**
```
VM name: cast04
VM type: paravirt XEN guest
OS: Alpine Linux latest LTS (64-bit)
CPU: 2 vCPU
RAM: 2 GB
Disk: 32 GB (non-shared)
Network: Internal 10.7.0.0/8 + internet egress (for Alertmanager SMTP / Telegram API calls)
```
---
## 6. Component catalog
### 6.1 Alpine apk packages (cast04)
```bash
apk update && apk upgrade
apk add prometheus alertmanager grafana \
prometheus-node-exporter \
prometheus-blackbox-exporter \
docker docker-cli docker-compose
```
| Package | Alpine repo | Version policy |
|---------|------------|---------------|
| `prometheus` | community/edge | latest stable in repo |
| `alertmanager` | community/edge | latest stable |
| `grafana` | community/edge | latest stable |
| `prometheus-node-exporter` | community | latest stable |
| `prometheus-blackbox-exporter` | community | latest stable |
| `docker` | community | for Uptime Kuma container |
If a package is only in Alpine **edge** repo, add the edge community URL to `/etc/apk/repositories` and pin via `/etc/apk/world`.
### 6.2 Alpine apk packages (cast01cast03)
Install only the node exporter on each workload node:
```bash
apk add prometheus-node-exporter
rc-update add node-exporter default
rc-service node-exporter start
```
On the MariaDB primary node, additionally:
```bash
apk add prometheus-mysqld-exporter
# Create limited MariaDB user for exporter:
mysql -e "CREATE USER 'prom_exporter'@'localhost' IDENTIFIED BY '<password>';"
mysql -e "GRANT PROCESS, REPLICATION CLIENT, SELECT ON *.* TO 'prom_exporter'@'localhost';"
# Configure /etc/prometheus-mysqld-exporter/mysqld_exporter.env
# DATA_SOURCE_NAME='prom_exporter:<password>@(localhost:3306)/'
rc-update add mysqld-exporter default
rc-service mysqld-exporter start
```
### 6.3 Docker-based (cast04)
```bash
rc-update add docker default
rc-service docker start
# Uptime Kuma
docker run -d --restart=always \
--name uptime-kuma \
-p 3001:3001 \
-v uptime-kuma:/app/data \
louislam/uptime-kuma:latest
```
### 6.4 External SaaS (no install)
| Service | URL | Action |
|---------|-----|--------|
| UptimeRobot | https://uptimerobot.com | Create free account; add HTTP monitor for `https://apps.f0xx.org`; configure email alert |
---
## 7. Deployment plan
Deploy phases in order. Each phase MUST be verified before starting the next.
### 7.1 Phase 0 — cast04 provisioning
**Goal:** cast04 VM exists and is reachable from cast01cast03 on internal network.
| Action | Detail |
|--------|--------|
| Request VM | Ask Anton for cast04 — spec per §5 table |
| Base install | Alpine latest LTS; configure apk repos (community + edge) |
| SSH access | Add developer SSH key; test login |
| Network | Confirm cast04 can reach cast0103 IPs on internal network |
| NTP | `apk add chrony && rc-service chronyd start` |
**Verify:** `ssh cast04 ping cast01` succeeds; `apk update` succeeds.
### 7.2 Phase 1 — Prometheus + node_exporter
**Goal:** Prometheus scraping hardware metrics from all nodes.
```bash
# On cast04:
apk add prometheus prometheus-node-exporter
# /etc/prometheus/prometheus.yml
global:
scrape_interval: 15s
evaluation_interval: 15s
scrape_configs:
- job_name: 'node'
static_configs:
- targets:
- 'cast01:9100'
- 'cast02:9100'
- 'cast03:9100'
- 'cast04:9100'
rc-update add prometheus default
rc-update add node-exporter default
rc-service prometheus start
rc-service node-exporter start
```
On cast01cast03: install and start `prometheus-node-exporter` (§6.2).
**Verify:** `curl http://cast04:9090/api/v1/targets` shows all 4 nodes in state `up`.
### 7.3 Phase 2 — Service probes and MariaDB exporter
**Goal:** HTTP/TCP probes and MariaDB metrics flowing.
Add to `prometheus.yml`:
```yaml
- job_name: 'blackbox_http'
metrics_path: /probe
params:
module: [http_2xx]
static_configs:
- targets:
- https://apps.f0xx.org
- http://cast01:80 # nginx on workload nodes
# add per-microservice /health endpoints as they exist
relabel_configs:
- source_labels: [__address__]
target_label: __param_target
- source_labels: [__param_target]
target_label: instance
- target_label: __address__
replacement: cast04:9115 # blackbox_exporter
- job_name: 'blackbox_tcp'
metrics_path: /probe
params:
module: [tcp_connect]
static_configs:
- targets:
- cast01:3306 # MariaDB
- cast02:3306 # replica
- cast01:22
- cast02:22
- cast03:22
relabel_configs:
- source_labels: [__address__]
target_label: __param_target
- source_labels: [__param_target]
target_label: instance
- target_label: __address__
replacement: cast04:9115
- job_name: 'mariadb'
static_configs:
- targets: ['<mariadb-primary-node>:9104']
```
Start services:
```bash
rc-service blackbox-exporter start
rc-service mysqld-exporter start # on MariaDB node
```
**Verify:** `curl "http://cast04:9115/probe?target=https://apps.f0xx.org&module=http_2xx"` returns `probe_success 1`.
### 7.4 Phase 3 — Alertmanager
**Goal:** Alerts route to email and (optionally) Telegram on threshold breach.
See §8 for full Alertmanager config skeletons. Deploy skeleton, then fill in SMTP credentials and Telegram token.
```bash
apk add alertmanager
# Edit /etc/alertmanager/alertmanager.yml (see §8.3 + §8.4)
rc-update add alertmanager default
rc-service alertmanager start
```
Add `alerting:` and `rule_files:` blocks to `prometheus.yml` (see §8.5 for rule file examples).
**Verify:** `amtool --alertmanager.url=http://cast04:9093 alert` returns empty (no active alerts); send a test alert and confirm email receipt.
### 7.5 Phase 4 — Grafana dashboards
**Goal:** Visual dashboards accessible for all metrics collected in phases 13.
```bash
apk add grafana
rc-update add grafana default
rc-service grafana start
# Access: http://cast04:3000 (via SSH tunnel: ssh -L 3000:cast04:3000 cast04)
```
In Grafana UI:
1. Add Prometheus datasource: `http://localhost:9090`
2. Import community dashboard **1860** ("Node Exporter Full") — covers §4.1 metrics
3. Import community dashboard **7362** or **13106** (MariaDB / mysqld_exporter)
4. Create custom panel for blackbox probe results (status per service endpoint)
5. Create overview panel showing hardware guarantee baseline vs measured values
**Verify:** All four dashboards render data; node_exporter CPU steal panel shows non-zero (or zero — both valid).
### 7.6 Phase 5 — Uptime Kuma + UptimeRobot
**Goal:** Service status page + external probe active.
```bash
# Uptime Kuma on cast04 (Docker):
rc-service docker start
docker run -d --restart=always --name uptime-kuma \
-p 3001:3001 -v uptime-kuma:/app/data louislam/uptime-kuma:latest
# Access: http://cast04:3001 (via SSH tunnel)
```
Configure monitors in Uptime Kuma UI:
- Add HTTP monitors for each microservice `/health` endpoint
- Add TCP monitors for MariaDB :3306, nginx :80
- Configure notification to same email as Alertmanager
**UptimeRobot:**
- Register free account at https://uptimerobot.com
- Add HTTP(s) monitor: `https://apps.f0xx.org`, 5-minute interval
- Add email alert notification (same ops email as Alertmanager)
**Verify:** Uptime Kuma dashboard shows all green; UptimeRobot sends test notification successfully.
### 7.7 ac-deploy integration
Add a `monitoring/` role to cluster0 deploy scripts in `ac-deploy/sim/cluster0/`:
```
ac-deploy/sim/cluster0/
monitoring/
deploy-monitoring.sh # installs cast04 components
prometheus.yml.j2 # Jinja2/envsubst template with node IPs
alertmanager.yml.j2 # template with SMTP/Telegram creds from env
alert-rules/
hardware.yml # §4.1 rules
services.yml # §4.2 rules
grafana/
datasources.yml # auto-provision Prometheus datasource
dashboards/ # provisioned dashboard JSONs
```
Credentials (SMTP password, Telegram token) MUST be passed via environment variables or a secrets file outside git — never committed.
**Verify:** `bash deploy-monitoring.sh` on a fresh cast04 produces a running stack within 10 minutes.
---
## 8. Alert routing
### 8.1 Alert severity levels
| Level | Meaning | Response |
|-------|---------|---------|
| `critical` | Service down or hardware limit exceeded | Alert immediately; wake operator if outside hours |
| `warning` | Approaching threshold; degraded | Alert during business hours; batch in digest |
| `info` | Notable event; no action required | Log only; no alert |
### 8.2 Routing rules
```
critical → email (immediately) + Telegram (immediately)
warning → email (grouped, 5 min window)
info → no alert (recorded in Prometheus only)
```
### 8.3 Alertmanager SMTP config (skeleton)
```yaml
# /etc/alertmanager/alertmanager.yml
global:
smtp_smarthost: 'smtp.gmail.com:587' # or Mailgun: smtp.mailgun.org:587
smtp_from: 'alerts@your-domain.com'
smtp_auth_username: 'your-gmail@gmail.com'
smtp_auth_password: '<APP_PASSWORD>' # from env / secrets file
smtp_require_tls: true
route:
group_by: ['alertname', 'instance']
group_wait: 30s
group_interval: 5m
repeat_interval: 4h
receiver: 'ops-email'
routes:
- match:
severity: critical
receiver: 'ops-critical'
group_wait: 10s
repeat_interval: 1h
receivers:
- name: 'ops-email'
email_configs:
- to: 'ops@your-domain.com'
send_resolved: true
- name: 'ops-critical'
email_configs:
- to: 'ops@your-domain.com'
send_resolved: true
telegram_configs:
- bot_token: '<TELEGRAM_BOT_TOKEN>' # from env / secrets file
chat_id: <TELEGRAM_CHAT_ID>
send_resolved: true
inhibit_rules:
- source_match:
severity: 'critical'
target_match:
severity: 'warning'
equal: ['instance']
```
**Note:** Replace `<APP_PASSWORD>` and `<TELEGRAM_BOT_TOKEN>` with values from environment — never commit credentials to git.
### 8.4 Telegram receiver config (skeleton)
Obtain Telegram credentials:
1. Open Telegram, message `@BotFather`
2. `/newbot` → follow prompts → save `TOKEN`
3. Add the bot to your ops group or DM it, then: `curl https://api.telegram.org/bot<TOKEN>/getUpdates` → find `chat.id`
4. Set `TELEGRAM_BOT_TOKEN=<TOKEN>` and `TELEGRAM_CHAT_ID=<CHAT_ID>` in the secrets file used by deploy script
### 8.5 Alert rule catalog
Create `/etc/prometheus/rules/hardware.yml`:
```yaml
groups:
- name: hardware
rules:
- alert: CpuHighUtilization
expr: 100 - (avg by(instance) (rate(node_cpu_seconds_total{mode="idle"}[5m])) * 100) > 95
for: 5m
labels:
severity: critical
annotations:
summary: "CPU critical on {{ $labels.instance }}"
description: "CPU utilization {{ $value | humanize }}% > 95%"
- alert: XenCpuStealHigh
expr: avg by(instance) (rate(node_cpu_seconds_total{mode="steal"}[5m])) * 100 > 10
for: 5m
labels:
severity: critical
annotations:
summary: "Xen CPU steal critical on {{ $labels.instance }} — notify Anton"
description: "CPU steal {{ $value | humanize }}% > 10% — hypervisor overcommit"
- alert: RamCritical
expr: (1 - node_memory_MemAvailable_bytes / node_memory_MemTotal_bytes) * 100 > 90
for: 5m
labels:
severity: critical
annotations:
summary: "RAM critical on {{ $labels.instance }}"
- alert: DiskCritical
expr: (1 - node_filesystem_avail_bytes{fstype!~"tmpfs|overlay"} / node_filesystem_size_bytes) * 100 > 90
for: 5m
labels:
severity: critical
annotations:
summary: "Disk critical on {{ $labels.instance }} mount {{ $labels.mountpoint }}"
- alert: DiskWriteLatencyHigh
expr: >
rate(node_disk_write_time_seconds_total[5m])
/ rate(node_disk_writes_completed_total[5m]) * 1000 > 100
for: 5m
labels:
severity: critical
annotations:
summary: "Disk write latency critical on {{ $labels.instance }} — possible HW fault; notify Anton"
- alert: HardwareGuaranteeRamBreach
expr: node_memory_MemTotal_bytes < 7.5 * 1024 * 1024 * 1024
for: 1m
labels:
severity: critical
annotations:
summary: "RAM below guaranteed 8 GB on {{ $labels.instance }} — notify Anton"
```
Create `/etc/prometheus/rules/services.yml`:
```yaml
groups:
- name: services
rules:
- alert: ServiceDown
expr: probe_success == 0
for: 1m
labels:
severity: critical
annotations:
summary: "Service unreachable: {{ $labels.instance }}"
- alert: ServiceSlowResponse
expr: probe_duration_seconds > 10
for: 2m
labels:
severity: warning
annotations:
summary: "Slow response from {{ $labels.instance }}: {{ $value | humanize }}s"
- alert: MariaDbReplicationLag
expr: mysql_slave_status_seconds_behind_master > 30
for: 2m
labels:
severity: critical
annotations:
summary: "MariaDB replication lag {{ $value | humanize }}s on {{ $labels.instance }}"
- alert: MariaDbDown
expr: mysql_up == 0
for: 1m
labels:
severity: critical
annotations:
summary: "MariaDB exporter cannot reach database on {{ $labels.instance }}"
```
Add to `prometheus.yml`:
```yaml
rule_files:
- "/etc/prometheus/rules/*.yml"
alerting:
alertmanagers:
- static_configs:
- targets: ['cast04:9093']
```
---
## 9. Grafana dashboard catalog
| Dashboard | Grafana ID | Source | Covers |
|-----------|-----------|--------|--------|
| Node Exporter Full | 1860 | community | CPU, RAM, disk, NIC, load — all §4.1 metrics |
| MariaDB Overview | 7362 | community | Connections, queries, replication lag, InnoDB |
| Blackbox Exporter | 7587 | community | HTTP probe success/latency per target |
| Custom: Hardware Guarantee | — | create | Measured vs Anton's guaranteed values |
| Custom: Service Health Overview | — | create | All service probes in one panel grid |
**Provisioning via ac-deploy:** export dashboard JSON from Grafana UI after customization and store in `ac-deploy/sim/cluster0/monitoring/grafana/dashboards/`. Add datasource YAML to `grafana/datasources/prometheus.yml`.
---
## 10. Pre-requisite accounts
The following must be created by the project owner before Phase 3 deploy:
| Account | Service | Purpose | Notes |
|---------|---------|---------|-------|
| Gmail app password | Gmail Settings → Security → App passwords | Alertmanager SMTP relay | App password, not account password. Alternatively: Mailgun free (10k/month) |
| Telegram bot | @BotFather in Telegram | Alertmanager Telegram receiver | `/newbot` → TOKEN; `getUpdates` → CHAT_ID |
| UptimeRobot free | https://uptimerobot.com | External HTTP probe | Add monitor for `https://apps.f0xx.org` |
**Important:** ImprovMX and addy.io provide inbound email routing (MX records) — they are not outbound SMTP relay services and cannot be used with Alertmanager. Use Gmail app password or a dedicated transactional email service (Mailgun, SendGrid) for outbound alerts.
---
## 11. TimescaleDB relationship
This SPEC implements infra-layer monitoring. Business and application metrics are a separate concern handled by the TimescaleDB platform plan ([20260620_postgres_timescale_db_platform.md](../20260620_postgres_timescale_db_platform.md)).
| Layer | Store | Examples | Retention |
|-------|-------|----------|-----------|
| Infrastructure | Prometheus TSDB on cast04 | CPU steal, disk latency, service probe | 90 days |
| Business / application | TimescaleDB (PostgreSQL) | Crash rate, session count, graph analytics | Long-term (configurable) |
Grafana on cast04 can query both stores simultaneously. No ETL or data duplication is required between them. The TimescaleDB datasource is added to Grafana as a second datasource in a future step — it does not affect R0 monitoring deployment.
---
## 12. Verification and acceptance
All phases must pass their gate before the next phase begins.
| Phase | Acceptance criteria |
|-------|---------------------|
| 0 | cast04 reachable from cast0103; `apk update` succeeds |
| 1 | `curl http://cast04:9090/api/v1/targets` shows all 4 nodes state=up |
| 2 | Blackbox probe for `https://apps.f0xx.org` returns `probe_success 1`; mysqld_exporter metrics visible in Prometheus |
| 3 | Test alert fires and email is received within 2 minutes; Telegram notification received (if configured) |
| 4 | Grafana dashboard 1860 renders data for all nodes; MariaDB dashboard shows replication status |
| 5 | Uptime Kuma shows all services green; UptimeRobot sends test notification |
| ac-deploy | Fresh cast04 reproduced by running `deploy-monitoring.sh` within 10 minutes |
**Ongoing regression:** After any cluster change (new microservice, new VM, nginx config change), add corresponding blackbox probe target and Grafana panel within 48 hours.
---
## 13. Risks and mitigations
| Risk | Mitigation |
|------|------------|
| cast04 itself fails — monitoring blind | UptimeRobot external probe continues; Alertmanager writes state to disk and resumes after VM restart |
| XEN host failure takes cast04 offline | UptimeRobot fires from external — operator is alerted even without cast04 |
| Alpine edge package breaks Prometheus | Pin version in `/etc/apk/world`; test upgrades on cast04 before rolling to workload nodes |
| SMTP rate limit during incident storm | Alertmanager `group_interval: 5m` + `repeat_interval: 4h` limit volume; Telegram is parallel channel |
| Prometheus disk overflow | 32 GB cast04 disk; estimated < 5 GB for 3 nodes at 15s scrape over 90 days; monitor cast04 disk in Grafana |
| Credentials leaked in git | Secrets must be in environment variables or a secrets file excluded from git (`.gitignore`); never in `prometheus.yml` or `alertmanager.yml` committed to repo |
| FE/router metrics unavailable | FE nginx stub_status available via config if Anton allows; not mandatory — blackbox HTTP probe sufficient at R0 |
---
## 14. Related documents
| Document | Role |
|----------|------|
| [DRs/20260626_cluster_monitoring.md](../DRs/20260626_cluster_monitoring.md) | Frozen design review (source of this SPEC) |
| [drafts/20260626_cluster_orchestration.txt](../drafts/20260626_cluster_orchestration.txt) | Track B — auto-remediation placeholder |
| [INFRA.md](../INFRA.md) | FE/BE topology, hosts, ports |
| [specs/20100612_1_scaling.md](20100612_1_scaling.md) | VM roles and cluster layout |
| [specs/20260618_repos_reorganizing.md](20260618_repos_reorganizing.md) | Microservice split (each new MS needs a blackbox probe) |
| [20260620_postgres_timescale_db_platform.md](../20260620_postgres_timescale_db_platform.md) | Business metrics store (complements Prometheus) |
| [GRAFANA_vs_others_graphvis_pivot.md](../GRAFANA_vs_others_graphvis_pivot.md) | Earlier Grafana evaluation for application analytics |
| [orchestration/sim/cluster0/ARCHITECTURE.md](../../orchestration/sim/cluster0/ARCHITECTURE.md) | Lab cluster node layout |
---
## 15. Changelog
| Rev | Date | Change |
|-----|------|--------|
| R0 | 2026-06-26 | SPEC converted from draft + DR R0; all decisions locked; deployment plan §7; alert rules §8 |