Capacity and risk review¶
Per workload: what it actually consumes, what it reserved, whether it is
bounded at all, and whether anything is failing. That takes usage, spec, the
ownership chain and events together, which is what no single kubectl
command can do, and Kubernetes forgets: events expire after about an hour and
usage is only ever now.
The output on this page is from a CI run against the shop namespace the
Examples use, on a one-node kind cluster.
The tables involved¶
Usage comes from metrics-server as metrics_k8s_io_pods and
metrics_k8s_io_nodes. Events arrive as core_events and
events_k8s_io_events. Every table is named for its API group, which is why
core pods are core_pods beside metrics_k8s_io_pods.
Quantities are strings until you convert them¶
Kubernetes reports measurements as strings with unit suffixes, 49903n of CPU
and 14488Ki of memory. Postgres cannot sum those, and > on them is a string
comparison that answers wrongly without an error. axiom_quantity() converts
exactly:
SELECT axiom_quantity('100m') AS "100m",
axiom_quantity('128Mi') AS "128Mi",
axiom_quantity('1M') = axiom_quantity('1Mi') AS "1M = 1Mi";
100m | 128Mi | 1M = 1Mi
-------+-----------+----------
0.100 | 134217728 | f
(1 row)
M is a thousand thousand and Mi is 1024 × 1024, so they are not equal. It
returns NULL rather than raising for anything that is not a quantity, so one
odd field cannot fail a query that spans a cluster.
The whole fleet in one statement¶
WITH usage AS (
SELECT m.namespace, m.name AS pod,
sum(axiom_quantity(c->'usage'->>'memory')) AS mem_used
FROM k8s.metrics_k8s_io_pods m, jsonb_array_elements(m.containers) c
GROUP BY 1, 2),
spec AS (
SELECT p.namespace, p.name AS pod,
p.metadata->'ownerReferences'->0->>'name' AS rs,
sum(axiom_quantity(c->'resources'->'requests'->>'memory')) AS mem_req,
bool_or(c->'resources'->'limits' IS NULL) AS no_limits
FROM k8s.core_pods p, jsonb_array_elements(p.spec->'containers') c
GROUP BY 1, 2, 3),
owner AS (
SELECT r.namespace, r.name AS rs,
coalesce(r.metadata->'ownerReferences'->0->>'name', r.name) AS workload
FROM k8s.apps_replicasets r),
warn AS (
SELECT e.namespace, e.involved_object->>'name' AS pod,
sum(coalesce(e.count, 1)) AS warnings,
max(coalesce(e.last_timestamp, e.event_time, e.creation_timestamp)) AS at,
(array_agg(e.reason ORDER BY coalesce(e.last_timestamp, e.event_time,
e.creation_timestamp) DESC))[1] AS why
FROM k8s.core_events e
WHERE e.type = 'Warning' AND e.involved_object->>'kind' = 'Pod'
GROUP BY 1, 2)
SELECT s.namespace,
coalesce(o.workload, s.pod) AS workload,
count(*) AS pods,
round(sum(u.mem_used) / 1024 / 1024, 1) AS mem_used_mib,
round(sum(s.mem_req) / 1024 / 1024, 1) AS mem_requested_mib,
CASE WHEN sum(s.mem_req) > 0
THEN round(100 * sum(u.mem_used) / sum(s.mem_req)) END AS pct_of_request,
bool_or(s.no_limits) AS unbounded,
coalesce(sum(w.warnings), 0) AS warnings,
(array_agg(w.why ORDER BY w.at DESC NULLS LAST))[1] AS latest_warning
FROM spec s
LEFT JOIN usage u ON u.namespace = s.namespace AND u.pod = s.pod
LEFT JOIN owner o ON o.namespace = s.namespace AND o.rs = s.rs
LEFT JOIN warn w ON w.namespace = s.namespace AND w.pod = s.pod
GROUP BY s.namespace, coalesce(o.workload, s.pod)
ORDER BY mem_used_mib DESC NULLS LAST;
namespace | workload | pods | mem_used_mib | mem_requested_mib | pct_of_request | unbounded | warnings | latest_warning
--------------------+-------------------------------------------------+------+--------------+-------------------+----------------+-----------+----------+------------------
kube-system | kube-apiserver-axiom-e2e-control-plane | 1 | 356.4 | | | t | 1 | NodeNotReady
kube-system | etcd-axiom-e2e-control-plane | 1 | 77.6 | 100.0 | 78 | t | 1 | NodeNotReady
kube-system | kube-controller-manager-axiom-e2e-control-plane | 1 | 74.4 | | | t | 0 |
kube-system | coredns | 2 | 30.0 | 140.0 | 21 | f | 2 | FailedScheduling
kube-system | kube-scheduler-axiom-e2e-control-plane | 1 | 23.4 | | | t | 3 | NodeNotReady
kube-system | metrics-server | 1 | 21.9 | 200.0 | 11 | t | 0 |
axiom-system | axiom-gateway | 1 | 17.1 | 64.0 | 27 | f | 0 |
kube-system | kube-proxy-cdpt6 | 1 | 16.0 | | | t | 0 |
kube-system | kindnet-fm56z | 1 | 13.4 | 50.0 | 27 | f | 0 |
local-path-storage | local-path-provisioner | 1 | 8.8 | | | t | 1 | FailedScheduling
shop | web | 2 | 0.4 | 32.0 | 1 | t | 0 |
axiom-e2e | web-0 | 1 | 0.2 | | | t | 0 |
shop | catalog | 1 | 0.2 | 256.0 | 0 | f | 0 |
axiom-e2e | db-0 | 1 | 0.2 | | | t | 0 |
axiom-e2e | web-1 | 1 | 0.2 | | | t | 0 |
shop | checkout | 1 | | | | t | 1 | FailedScheduling
shop | checkout-worker | 1 | | | | t | 7 | BackOff
shop | report | 1 | | | | t | 8 | Failed
(18 rows)
In shop, catalog reserves 256 MiB and uses 0.2, which the scheduler
cannot know. checkout, checkout-worker and report use nothing because
they never started, and why is on the same row. The largest consumer is the
API server, at 356 MiB with no limit set; on a cluster of your own, the top
rows are your workloads.
Keep usage on a LEFT JOIN. An inner join drops exactly the workloads
that have no metrics because they never started, which are the ones you most
want to see.
Act on it¶
The review produces a list; a write turns it into an action. Find it and fix it in one statement annotates each Deployment in a namespace that uses under a fifth of the memory it requests.
Requirements¶
Usage tables need metrics-server installed in the cluster; without it
metrics.k8s.io does not exist and the tables are simply absent. The gateway
also needs list on that group. The shipped RBAC grants it, so metrics-server
installed later is picked up on the next discovery refresh with no RBAC edit.
Custom and external metrics APIs, from an adapter such as KEDA or
prometheus-adapter, are granted the same way.