Cluster self-healing and automated orchestration draft (Track B) milestone: R1+ severity: high state: placeholder — not for immediate implementation; referenced by monitoring SPEC (Track A) author: Anton Afanasyeu date: 2026-06-26 related to: 20260626_cluster_monitoring*, 20260618_repos_reorganizing*, microservice split Track A (prerequisite, must be complete first): - Full passive monitoring + alerting operational (see specs/20260626_cluster_monitoring.md) - Alert routing proven stable (email + Telegram) - At least 4 weeks of baseline metrics collected (needed to calibrate thresholds) The idea: Once the monitoring layer (Track A) is stable and producing reliable alert signals, the next step is to make the cluster respond to those signals automatically — rather than requiring a human to act on every alert. This is commonly called cluster self-healing, auto-remediation, or a "smart hypervisor" layer. It is NOT part of the monitoring SPEC; it is an orchestration concern. Target scenarios (desired future automation): Scenario 1 — CPU overload → horizontal scale-out Trigger: microservice X CPU utilization > threshold for N minutes (Alertmanager signal) Action: provision new node Y (via Xen/cloud API), deploy microservice X replica on it, update load balancer / nginx upstream, re-route portion of traffic Complexity: HIGH — requires VM provisioning API + dynamic nginx config reload Scenario 2 — Idle scale-in (cost saving) Trigger: microservice X CPU < threshold for Y hours AND overall cluster load < 75% Action: gracefully drain microservice X instance, deprovision node, update upstream config Complexity: HIGH — requires graceful drain protocol per microservice type Scenario 3 — DDoS / network spike response Trigger: NIC ingress/egress > threshold OR nginx 5xx rate spikes (Alertmanager signal) Sub-action A (DDoS suspected): block suspect IPs, rate-limit at iptables, notify Anton Sub-action B (organic spike): scale out cluster capacity, increase bandwidth (notify Anton for bandwidth tier change — cannot be auto-done from cluster) Complexity: VERY HIGH — requires traffic analysis, IP reputation, coordinated with upstream Scenario 4 — Storage fault recovery Trigger: disk SMART pre-failure / disk write latency > threshold on node Y Action: elect replica to master (if MariaDB GTID replication is active), notify Anton to provision spare node, re-sync replica on new node Complexity: HIGH — MariaDB GTID failover + new node provisioning Common prerequisites for all scenarios: 1. Track A monitoring + alerting fully operational (Alertmanager webhook available) 2. VM provisioning API (Xen xl toolstack or cloud provider API with automation credentials) 3. Configuration management (Ansible / ac-deploy role) for dynamic service placement 4. Graceful drain protocol defined per microservice (health endpoint + connection draining) 5. Load balancer / nginx upstream reconfiguration automation (e.g. consul-template or custom script) 6. Decision logic (thresholds, hysteresis, human-in-the-loop gates vs fully-automatic) Approach options (TBD — to be designed in R1 DR): a) Kubernetes operators (standard industry path; requires migrating to K8s) b) Ansible event-driven (Ansible EDA) — react to Alertmanager webhooks with playbooks c) Home-grown orchestration scripts called by Alertmanager webhook receiver on cast04 d) HashiCorp Nomad + Consul (lighter than K8s; service placement + dynamic upstream config) Open questions (for R1 DR): - Does Anton's infrastructure support Xen VM provisioning via API / xl scripting? - What is the minimum cluster node count for safe scale-in (must keep >= N nodes)? - Is Kubernetes acceptable given paravirt constraint (k3s on Alpine paravirt is feasible)? - What is the human-in-the-loop policy: always require approval for scale-out/in, or fully automatic within defined bounds? - Cost model for scale-out: how are additional nodes billed / what is the budget ceiling? Related documents: - Track A: docs/DRs/20260626_cluster_monitoring.md + docs/specs/20260626_cluster_monitoring.md - Scaling SPEC: docs/specs/20100612_1_scaling.md - Repo reorganization SPEC: docs/specs/20260618_repos_reorganizing.md - ac-deploy cluster scripts: orchestration/sim/cluster0/