Skip to content

Capacity and risk review

Per workload: what it actually consumes, what it reserved, whether it is bounded at all, and whether anything is failing. That takes usage, spec, the ownership chain and events together, which is what no single kubectl command can do, and Kubernetes forgets: events expire after about an hour and usage is only ever now.

The output on this page is from a CI run against the shop namespace the Examples use, on a one-node kind cluster.

The tables involved

Usage comes from metrics-server as metrics_k8s_io_pods and metrics_k8s_io_nodes. Events arrive as core_events and events_k8s_io_events. Every table is named for its API group, which is why core pods are core_pods beside metrics_k8s_io_pods.

Quantities are strings until you convert them

Kubernetes reports measurements as strings with unit suffixes, 49903n of CPU and 14488Ki of memory. Postgres cannot sum those, and > on them is a string comparison that answers wrongly without an error. axiom_quantity() converts exactly:

SELECT axiom_quantity('100m')  AS "100m",
       axiom_quantity('128Mi') AS "128Mi",
       axiom_quantity('1M') = axiom_quantity('1Mi') AS "1M = 1Mi";
 100m  |   128Mi   | 1M = 1Mi 
-------+-----------+----------
 0.100 | 134217728 | f
(1 row)

M is a thousand thousand and Mi is 1024 × 1024, so they are not equal. It returns NULL rather than raising for anything that is not a quantity, so one odd field cannot fail a query that spans a cluster.

The whole fleet in one statement

WITH usage AS (
  SELECT m.namespace, m.name AS pod,
         sum(axiom_quantity(c->'usage'->>'memory')) AS mem_used
    FROM k8s.metrics_k8s_io_pods m, jsonb_array_elements(m.containers) c
   GROUP BY 1, 2),
spec AS (
  SELECT p.namespace, p.name AS pod,
         p.metadata->'ownerReferences'->0->>'name' AS rs,
         sum(axiom_quantity(c->'resources'->'requests'->>'memory')) AS mem_req,
         bool_or(c->'resources'->'limits' IS NULL) AS no_limits
    FROM k8s.core_pods p, jsonb_array_elements(p.spec->'containers') c
   GROUP BY 1, 2, 3),
owner AS (
  SELECT r.namespace, r.name AS rs,
         coalesce(r.metadata->'ownerReferences'->0->>'name', r.name) AS workload
    FROM k8s.apps_replicasets r),
warn AS (
  SELECT e.namespace, e.involved_object->>'name' AS pod,
         sum(coalesce(e.count, 1)) AS warnings,
         max(coalesce(e.last_timestamp, e.event_time, e.creation_timestamp)) AS at,
         (array_agg(e.reason ORDER BY coalesce(e.last_timestamp, e.event_time,
                                               e.creation_timestamp) DESC))[1] AS why
    FROM k8s.core_events e
   WHERE e.type = 'Warning' AND e.involved_object->>'kind' = 'Pod'
   GROUP BY 1, 2)
SELECT s.namespace,
       coalesce(o.workload, s.pod) AS workload,
       count(*) AS pods,
       round(sum(u.mem_used) / 1024 / 1024, 1) AS mem_used_mib,
       round(sum(s.mem_req) / 1024 / 1024, 1) AS mem_requested_mib,
       CASE WHEN sum(s.mem_req) > 0
            THEN round(100 * sum(u.mem_used) / sum(s.mem_req)) END AS pct_of_request,
       bool_or(s.no_limits) AS unbounded,
       coalesce(sum(w.warnings), 0) AS warnings,
       (array_agg(w.why ORDER BY w.at DESC NULLS LAST))[1] AS latest_warning
  FROM spec s
  LEFT JOIN usage u ON u.namespace = s.namespace AND u.pod = s.pod
  LEFT JOIN owner o ON o.namespace = s.namespace AND o.rs = s.rs
  LEFT JOIN warn  w ON w.namespace = s.namespace AND w.pod = s.pod
 GROUP BY s.namespace, coalesce(o.workload, s.pod)
 ORDER BY mem_used_mib DESC NULLS LAST;
     namespace      |                    workload                     | pods | mem_used_mib | mem_requested_mib | pct_of_request | unbounded | warnings |  latest_warning  
--------------------+-------------------------------------------------+------+--------------+-------------------+----------------+-----------+----------+------------------
 kube-system        | kube-apiserver-axiom-e2e-control-plane          |    1 |        356.4 |                   |                | t         |        1 | NodeNotReady
 kube-system        | etcd-axiom-e2e-control-plane                    |    1 |         77.6 |             100.0 |             78 | t         |        1 | NodeNotReady
 kube-system        | kube-controller-manager-axiom-e2e-control-plane |    1 |         74.4 |                   |                | t         |        0 | 
 kube-system        | coredns                                         |    2 |         30.0 |             140.0 |             21 | f         |        2 | FailedScheduling
 kube-system        | kube-scheduler-axiom-e2e-control-plane          |    1 |         23.4 |                   |                | t         |        3 | NodeNotReady
 kube-system        | metrics-server                                  |    1 |         21.9 |             200.0 |             11 | t         |        0 | 
 axiom-system       | axiom-gateway                                   |    1 |         17.1 |              64.0 |             27 | f         |        0 | 
 kube-system        | kube-proxy-cdpt6                                |    1 |         16.0 |                   |                | t         |        0 | 
 kube-system        | kindnet-fm56z                                   |    1 |         13.4 |              50.0 |             27 | f         |        0 | 
 local-path-storage | local-path-provisioner                          |    1 |          8.8 |                   |                | t         |        1 | FailedScheduling
 shop               | web                                             |    2 |          0.4 |              32.0 |              1 | t         |        0 | 
 axiom-e2e          | web-0                                           |    1 |          0.2 |                   |                | t         |        0 | 
 shop               | catalog                                         |    1 |          0.2 |             256.0 |              0 | f         |        0 | 
 axiom-e2e          | db-0                                            |    1 |          0.2 |                   |                | t         |        0 | 
 axiom-e2e          | web-1                                           |    1 |          0.2 |                   |                | t         |        0 | 
 shop               | checkout                                        |    1 |              |                   |                | t         |        1 | FailedScheduling
 shop               | checkout-worker                                 |    1 |              |                   |                | t         |        7 | BackOff
 shop               | report                                          |    1 |              |                   |                | t         |        8 | Failed
(18 rows)

In shop, catalog reserves 256 MiB and uses 0.2, which the scheduler cannot know. checkout, checkout-worker and report use nothing because they never started, and why is on the same row. The largest consumer is the API server, at 356 MiB with no limit set; on a cluster of your own, the top rows are your workloads.

Keep usage on a LEFT JOIN. An inner join drops exactly the workloads that have no metrics because they never started, which are the ones you most want to see.

Act on it

The review produces a list; a write turns it into an action. Find it and fix it in one statement annotates each Deployment in a namespace that uses under a fifth of the memory it requests.

Requirements

Usage tables need metrics-server installed in the cluster; without it metrics.k8s.io does not exist and the tables are simply absent. The gateway also needs list on that group. The shipped RBAC grants it, so metrics-server installed later is picked up on the next discovery refresh with no RBAC edit. Custom and external metrics APIs, from an adapter such as KEDA or prometheus-adapter, are granted the same way.