Monitoring & Alerts
What metrics do we monitor?
We provide comprehensive monitoring across servers and Kubernetes clusters using our open-source operations frameworks - LinuxAid and KubeAid - powered by Prometheus.
1. Server Monitoring (LinuxAid)
For servers, all metrics and alerting rules are defined in our LinuxAid stack here:
These include:
- Host and hardware metrics
- CPU, RAM, I/O, disk health
- Filesystem and storage monitoring
- Network performance
These rules are maintained as open source, updated regularly and we have also written a test suite for them.
2. Kubernetes Monitoring (KubeAid)
For Kubernetes environments, we use Kube Prometheus as our foundation and extend it with additional rules and mixins.
Baseline alerting rules:
Additional alerts and mixins:
Our Kubernetes alerting includes:
- Node, pod, and container metrics
- Control-plane and API server health
- etcd performance and availability
- Scheduler and controller-manager signals
- Ingress, service and network health
- Application-level alerts for selected open-source components
- Optional mixin-based alert packs that can be enabled or disabled per cluster
This layered approach ensures platform-wide visibility with the flexibility to adapt to each customer’s stack.
Security monitoring and mitigation
Fixes roll out in the service windows you choose. When one needs a change in the config repository, such as a new KubeAid version or snapshot date, it comes first as a branch to merge.
Servers (LinuxAid)
- Known vulnerabilities. The portal lists each server's open CVEs, with severity and the next update window.
- End of life. An alert from 75 days before the OS release stops getting free security updates. Where the release has paid extended support, it keeps firing until that ends too.
- Firewall. Each server reports its firewall's packet counts per rule as metrics, but no alert fires on dropped traffic.
- Hardening. SSH only allows strong key exchange, ciphers and MACs, except on RHEL 6. SELinux and auditd can be switched on per server.
- Repo snapshots. Where your config pins a snapshot date, packages come from that snapshot. The day before each update cycle we open a pull request that moves the date forward, so every server in the cycle gets the same updates.
Kubernetes (KubeAid)
- Supply chain security.
- Image scanning. Trivy scans the container images you run and keeps the critical and high findings that have a fix. On clusters that report to Obmondo, the portal shows them.
- Registry mirror. Harbor can mirror those images and scan them too.
- Firewall. Cilium network policies filter ingress and egress traffic. Cilium exports every dropped packet as a Hubble metric, but no alert fires when denied traffic stays high.
- Runtime security. Tetragon records process execution, kernel module loads and privilege changes, and follows network and file access only once a policy for it is added. KubeArmor can block what your policies forbid.
- Admission policies. Kyverno or Gatekeeper can stop workloads that break your rules before they start.
Image scanning, runtime detection and admission policies are options you switch on per cluster.
How we handle alerts and incident escalation
Prometheus routes alerts to our central alert ingestion endpoint (with the option to exclude hosts or components upon request).
Alerts are then processed by our internal operations platform, which:
- Automatically creates or updates alert tickets
- Notifies our 24/7 operations team
- Applies your SLA to correctly prioritize and set deadlines
- Tracks acknowledgment and resolution progress in real time
- Performs automated escalations - first to the on-call engineer, then to management if needed
- Some of our alerts are preemptive in nature i.e. fire before things can be broken/down.
We always prioritise restoring service as quickly as possible and then focus on addressing the root cause to prevent future recurrence.