Skip to content

Setting up Axiom (for coding agents)

A procedure for an autonomous agent to bring up a working Axiom environment and prove it works. Same result as the quick start, written step by step, with a check after each step, for a reader that cannot see the person's screen.

Every command is non-interactive and safe to re-run. Each step states how to verify it before moving on. Do not proceed past a failed verification — Failures lists the ones with unhelpful messages and what they actually mean.

Ask before you act

Settle these with the user first. Each changes what you do, and guessing any of them wrong costs more than the question:

  1. A local demo, or a real installation? This procedure builds a throwaway local environment on kind. For a real cluster and an existing Postgres, follow Install and Initialize instead, and ask the questions below.
  2. Which Kubernetes cluster? The kubectl context, and whether you may create a namespace, a Deployment and cluster-wide RBAC in it.
  3. Which Postgres? Its major version (16, 17 or 18), how it was installed, and whether you may restart it: Axiom must be added to shared_preload_libraries, which takes a restart.
  4. How does Postgres reach the cluster? The address of the gateway as seen from the database host. Without one, nothing works; do not invent it.
  5. Anything already there? An existing kind cluster named axiom, or a Postgres on port 5432 or 55432. Never delete or replace someone's environment to make room.

If the answer to 1 is "a local demo" and the user is happy to run a script, quickstart.sh does everything below in one command. Use the steps here when you need to check each one, or cannot run that script.

What you are building

Three pieces:

  1. a Kubernetes cluster (kind, created here if absent);
  2. the gateway, a Deployment in that cluster holding the cluster credentials and exposing gRPC over TLS on NodePort 30443;
  3. Postgres, a container with the Axiom extension preinstalled, on the same Docker network, querying the gateway.

Postgres joins the kind Docker network and dials the node by container name, so no port mapping is needed on the cluster.

Before you start

docker version --format '{{.Server.Version}}'
kubectl version --client -o json | head -5
kind version
uname -m

All four must succeed. The first three are required tools; uname -m is recorded only so a bug report can name the machine — it changes nothing below. arm64/aarch64 and x86_64 are both supported natively, every image this guide uses is published for both, and adding a --platform flag would pin you to a slice your machine then has to emulate.

Step 0 — somewhere to work

This procedure writes a TLS private key to disk. It must not land in a git checkout, where it would be untracked and one git add -A away from being committed.

mkdir -p ~/axiom-quickstart

Every path below is absolute, deliberately. Do not rewrite them as relative paths after a cd: if you run each command in its own shell — which is what most tool-using agents do — a cd in one step is gone by the next, and ./certs would then resolve to wherever that shell happened to start. That is exactly the mistake this step exists to prevent.

Two forms appear below and they are not interchangeable. ~ expands only at the start of a word, so it works for docker -v ~/axiom-quickstart/... but not after an = sign: --from-file=tls.crt=~/... is passed through literally and the file is not found. Those use "$HOME/..." instead.

Step 1 — cluster

This procedure requires kind. Not because Axiom needs it, but because Postgres reaches the gateway over kind's Docker network in step 4. On any other cluster the gateway has to be exposed some other way first, and this procedure does not cover that.

It creates a cluster named axiom and uses that name throughout. Do not substitute an existing cluster with a different name unless you also change every later use of axiom-control-plane.

Stop if a cluster of that name already exists. Step 3 replaces axiom-gateway-tls and restarts axiom-gateway in whatever cluster this resolves to, so adopting someone's existing axiom cluster would rewrite the TLS material of a running deployment — and steps 3 to 6 would then pass while having broken it. Deleting it is the caller's decision, not yours:

if kind get clusters | grep -qx axiom; then
  echo "a kind cluster named 'axiom' already exists; stopping." >&2
  echo "This procedure would replace its gateway TLS secret and restart its" >&2
  echo "gateway, which may not be disposable. Ask which cluster to use." >&2
  exit 1
fi
kind create cluster --name axiom

Do not resolve this by deleting the cluster. It may be someone's working environment — the report that prompted this check was run on exactly that. Ask. And do not continue to step 2: the later steps would succeed against that cluster while having broken it.

Verify. The node is NotReady for a while after creation, so wait for it rather than reading get nodes once:

kubectl --context kind-axiom wait --for=condition=Ready node --all --timeout=180s
kubectl --context kind-axiom get nodes

Expect one node named axiom-control-plane, Ready. That name is used verbatim later; do not substitute a different one.

Step 2 — TLS keypair

The gateway has no plaintext mode. Generate the keypair inside this container — macOS ships LibreSSL, whose output the gateway rejects for three different reasons.

mkdir -p ~/axiom-quickstart/certs
docker run --rm -v ~/axiom-quickstart/certs:/certs -w /certs \
  --entrypoint /bin/sh alpine/openssl:3.3.3 -c "
    openssl req -x509 -newkey ec -pkeyopt ec_paramgen_curve:prime256v1 -nodes \
      -days 365 -subj '/CN=axiom-dev-ca' -keyout ca.key -out ca.crt
    openssl req -newkey ec -pkeyopt ec_paramgen_curve:prime256v1 -nodes \
      -subj '/CN=gateway' -keyout gateway.key -out gateway.csr
    printf 'subjectAltName=DNS:axiom-control-plane,DNS:axiom-gateway.axiom-system.svc,DNS:localhost,IP:127.0.0.1\nextendedKeyUsage=serverAuth\n' > san.cnf
    openssl x509 -req -in gateway.csr -CA ca.crt -CAkey ca.key -CAcreateserial \
      -days 365 -extfile san.cnf -out gateway.crt
    rm -f gateway.csr san.cnf ca.srl ca.key
    chown $(id -u):$(id -g) ca.crt gateway.crt gateway.key
  "

Verify:

ls ~/axiom-quickstart/certs/
docker run --rm -v ~/axiom-quickstart/certs:/certs:ro alpine/openssl:3.3.3 \
  x509 -in /certs/gateway.crt -noout -ext subjectAltName

Check it in the container too. openssl x509 -ext is OpenSSL 1.1.1 and later; the openssl on macOS is LibreSSL, which does not have the flag and fails with unknown option -ext.

Expect exactly ca.crt, gateway.crt, gateway.key, and a SAN list containing DNS:axiom-control-plane. If that name is missing, the TLS handshake in step 5 fails; regenerate rather than continuing.

Step 3 — gateway

RAW=https://raw.githubusercontent.com/dhilipkumars/axiom/main/deploy/k8s

kubectl --context kind-axiom apply -f "$RAW/gateway-rbac.yaml"

kubectl --context kind-axiom -n axiom-system delete secret axiom-gateway-tls --ignore-not-found
kubectl --context kind-axiom -n axiom-system create secret generic axiom-gateway-tls \
  --from-file=tls.crt="$HOME/axiom-quickstart/certs/gateway.crt" \
  --from-file=tls.key="$HOME/axiom-quickstart/certs/gateway.key"

kubectl --context kind-axiom apply -f "$RAW/gateway-deployment.yaml"
kubectl --context kind-axiom -n axiom-system rollout restart deploy/axiom-gateway
kubectl --context kind-axiom -n axiom-system rollout status deploy/axiom-gateway --timeout=180s

The delete secret --ignore-not-found before create is what makes this step re-runnable; create secret alone fails on a second attempt.

Verify the rollout, and the port step 5 depends on:

kubectl --context kind-axiom -n axiom-system get svc axiom-gateway \
  -o jsonpath='{.spec.ports[0].port}:{.spec.ports[0].nodePort}{"\n"}'

Expect 8443:30443. Step 5 dials axiom-control-plane:30443, so if this prints anything else, fix the endpoint there rather than meeting it as an opaque TLS error two steps later.

rollout status exits 0 and prints deployment "axiom-gateway" successfully rolled out. If it times out:

kubectl --context kind-axiom -n axiom-system get pods
kubectl --context kind-axiom -n axiom-system logs deploy/axiom-gateway --tail=30

Step 4 — Postgres with Axiom preinstalled

docker rm -f axiom-postgres 2>/dev/null || true
docker run -d --name axiom-postgres \
  --network kind \
  -e POSTGRES_PASSWORD=axiom \
  -v ~/axiom-quickstart/certs:/certs:ro \
  ghcr.io/dhilipkumars/axiom-postgres:latest-pg17

Wait for readiness — poll, do not sleep a fixed amount:

for i in $(seq 1 90); do
  docker exec axiom-postgres pg_isready -U postgres >/dev/null 2>&1 && break
  sleep 2
done
docker exec axiom-postgres pg_isready -U postgres

Expect /var/run/postgresql:5432 - accepting connections.

Verify the extension is loadable, which is what the preload buys:

docker exec axiom-postgres psql -U postgres -tAc "SHOW shared_preload_libraries"

Expect axiom. Anything else means step 5 will fail fatally, not degrade.

Step 5 — connect Postgres to the gateway

Use docker exec -i with a heredoc. Never -it: with no TTY attached it either errors or hangs waiting for one.

docker exec -i axiom-postgres psql -U postgres -v ON_ERROR_STOP=1 <<'SQL'
CREATE EXTENSION IF NOT EXISTS axiom;
DROP SERVER IF EXISTS prod CASCADE;
CREATE SERVER prod
  FOREIGN DATA WRAPPER axiom_fdw
  OPTIONS (endpoint 'https://axiom-control-plane:30443', ca_cert '/certs/ca.crt');
CREATE USER MAPPING FOR CURRENT_USER SERVER prod;
CREATE SCHEMA IF NOT EXISTS k8s;
IMPORT FOREIGN SCHEMA k8s FROM SERVER prod INTO k8s;
SQL

DROP SERVER IF EXISTS ... CASCADE and the IF NOT EXISTS clauses make this re-runnable. ON_ERROR_STOP=1 matters: without it psql continues past a failed statement and exits 0, and you would conclude the step succeeded.

Verify:

docker exec axiom-postgres psql -U postgres -tAc \
  "SELECT string_agg(table_name, ',' ORDER BY table_name) FROM information_schema.tables
    WHERE table_schema='k8s' AND table_name IN ('core_pods','core_configmaps','apps_deployments','networking_k8s_io_networkpolicies','core_secrets')"

Expect apps_deployments,core_configmaps,core_pods,networking_k8s_io_networkpolicies, with no core_secrets.

The bundled RBAC reads broadly: everything in Kubernetes' view role plus nodes, storage, CRDs, RBAC objects, events and metrics. The whole schema holds several dozen tables, and the exact count depends on the Kubernetes version, so check for these four rather than for a number. Secrets are never granted, and their absence is part of the check.

Custom resources appear when their operator ships an aggregate-to-view role, or once a read-only ClusterRole for them is labelled axiom.dhilipkumars.github.io/aggregate-to-gateway: "true". A new grant is seen on the next import; after revoking one, restart the gateway (it caches what it is allowed). Either way, import again.

Step 6 — success criterion

docker exec axiom-postgres psql -U postgres -c \
  "SELECT name, namespace, phase, node FROM k8s.core_pods WHERE namespace='kube-system' ORDER BY name"

The environment is working when this returns the cluster's kube-system pods — on a fresh kind cluster, rows including etcd-axiom-control-plane and kube-apiserver-axiom-control-plane, all Running, all on node axiom-control-plane. Compare against the cluster directly if you want a second source:

kubectl --context kind-axiom -n kube-system get pods

The two must agree. They are the same data by different routes, which is the point of the project.

Failures

Symptom Cause Action
no matching manifest for linux/arm64/v8 the tag has no slice for this machine. Versions before 0.1.1 are amd64-only and give this legitimately; every tag from 0.1.1 on, including the latest-pgNN this guide uses, should be multi-architecture if you pinned an older version, use latest-pgNN or 0.1.1+. Otherwise the publish is broken — confirm with docker manifest inspect ghcr.io/dhilipkumars/axiom-postgres:latest-pg17 and report it. Retrying will not help: a manifest list is tagged atomically
denied on docker pull image is private or the tag does not exist check the tag; do not retry with credentials
FATAL: cannot create PGC_POSTMASTER variables after startup Axiom loaded without shared_preload_libraries use the published image, or preload it
CREATE EXTENSION closes the connection same as above as above
TLS handshake failure on the first IMPORT certificate SAN does not cover axiom-control-plane regenerate the keypair, step 2
Cancelled: Timeout expired on IMPORT whole-cluster import exceeded rpc_timeout_secs ALTER SERVER prod OPTIONS (ADD rpc_timeout_secs '120'); SET only if it is already set
rollout status times out image pull or crash loop kubectl -n axiom-system logs deploy/axiom-gateway
psql exits 0 but nothing was created missing ON_ERROR_STOP=1 re-run with it

Teardown

Leaving a kind cluster running consumes memory indefinitely. Tear down unless asked to keep it:

docker rm -f axiom-postgres
kind delete cluster --name axiom
rm -rf ~/axiom-quickstart

Constraints worth knowing before you suggest things

  • Axiom cannot be installed on managed Postgres. It is not a trusted extension and needs shared_preload_libraries, so RDS, Cloud SQL and Aurora cannot run it. Do not suggest them.
  • Installing into an existing Postgres uses a release package, published per major and architecture from v0.1.1 onward — postgresql-<major>-axiom as a .deb, axiom_<major> as an .rpm, or a tarball where neither applies. v0.1.0 has none, so check the releases page for the version you name before telling someone to download one. Recommend the package over the tarball: it refuses to install on a glibc below 2.34 rather than failing at the next postmaster start. The images remain the path for trying Axiom without touching an existing Postgres.
  • Every image is published for amd64 and arm64 from v0.1.1 on, so nothing runs under emulation and no --platform flag is needed. Versions before that are amd64-only.
  • What a query can reach is bounded by the gateway's RBAC, not the SQL user's. Reads are broad by default and never include Secrets. To expose a custom resource, label a read-only ClusterRole axiom.dhilipkumars.github.io/aggregate-to-gateway: "true"; to make one writable, grant its verbs in the axiom-gateway ClusterRole. Then re-import — foreign tables are catalog objects and do not follow the change. Only revoking a grant needs a gateway restart.