Setting up Axiom (for coding agents)¶
A procedure for an autonomous agent to bring up a working Axiom environment and prove it works. Same result as the quick start, written step by step, with a check after each step, for a reader that cannot see the person's screen.
Every command is non-interactive and safe to re-run. Each step states how to verify it before moving on. Do not proceed past a failed verification — Failures lists the ones with unhelpful messages and what they actually mean.
Ask before you act¶
Settle these with the user first. Each changes what you do, and guessing any of them wrong costs more than the question:
- A local demo, or a real installation? This procedure builds a throwaway local environment on kind. For a real cluster and an existing Postgres, follow Install and Initialize instead, and ask the questions below.
- Which Kubernetes cluster? The
kubectlcontext, and whether you may create a namespace, a Deployment and cluster-wide RBAC in it. - Which Postgres? Its major version (16, 17 or 18), how it was installed,
and whether you may restart it: Axiom must be added to
shared_preload_libraries, which takes a restart. - How does Postgres reach the cluster? The address of the gateway as seen from the database host. Without one, nothing works; do not invent it.
- Anything already there? An existing kind cluster named
axiom, or a Postgres on port 5432 or 55432. Never delete or replace someone's environment to make room.
If the answer to 1 is "a local demo" and the user is happy to run a script,
quickstart.sh does everything below in one command. Use the
steps here when you need to check each one, or cannot run that script.
What you are building¶
Three pieces:
- a Kubernetes cluster (kind, created here if absent);
- the gateway, a Deployment in that cluster holding the cluster credentials and exposing gRPC over TLS on NodePort 30443;
- Postgres, a container with the Axiom extension preinstalled, on the same Docker network, querying the gateway.
Postgres joins the kind Docker network and dials the node by container name,
so no port mapping is needed on the cluster.
Before you start¶
docker version --format '{{.Server.Version}}'
kubectl version --client -o json | head -5
kind version
uname -m
All four must succeed. The first three are required tools; uname -m is
recorded only so a bug report can name the machine — it changes nothing
below. arm64/aarch64 and x86_64 are both supported natively, every
image this guide uses is published for both, and adding a --platform flag
would pin you to a slice your machine then has to emulate.
Step 0 — somewhere to work¶
This procedure writes a TLS private key to disk. It must not land in a git
checkout, where it would be untracked and one git add -A away from being
committed.
mkdir -p ~/axiom-quickstart
Every path below is absolute, deliberately. Do not rewrite them as relative
paths after a cd: if you run each command in its own shell — which is what
most tool-using agents do — a cd in one step is gone by the next, and
./certs would then resolve to wherever that shell happened to start. That is
exactly the mistake this step exists to prevent.
Two forms appear below and they are not interchangeable. ~ expands only at
the start of a word, so it works for docker -v ~/axiom-quickstart/... but
not after an = sign: --from-file=tls.crt=~/... is passed through
literally and the file is not found. Those use "$HOME/..." instead.
Step 1 — cluster¶
This procedure requires kind. Not because Axiom needs it, but because Postgres reaches the gateway over kind's Docker network in step 4. On any other cluster the gateway has to be exposed some other way first, and this procedure does not cover that.
It creates a cluster named axiom and uses that name throughout. Do not
substitute an existing cluster with a different name unless you also change
every later use of axiom-control-plane.
Stop if a cluster of that name already exists. Step 3 replaces
axiom-gateway-tls and restarts axiom-gateway in whatever cluster this
resolves to, so adopting someone's existing axiom cluster would rewrite the
TLS material of a running deployment — and steps 3 to 6 would then pass while
having broken it. Deleting it is the caller's decision, not yours:
if kind get clusters | grep -qx axiom; then
echo "a kind cluster named 'axiom' already exists; stopping." >&2
echo "This procedure would replace its gateway TLS secret and restart its" >&2
echo "gateway, which may not be disposable. Ask which cluster to use." >&2
exit 1
fi
kind create cluster --name axiom
Do not resolve this by deleting the cluster. It may be someone's working environment — the report that prompted this check was run on exactly that. Ask. And do not continue to step 2: the later steps would succeed against that cluster while having broken it.
Verify. The node is NotReady for a while after creation, so wait for it
rather than reading get nodes once:
kubectl --context kind-axiom wait --for=condition=Ready node --all --timeout=180s
kubectl --context kind-axiom get nodes
Expect one node named axiom-control-plane, Ready. That name is used
verbatim later; do not substitute a different one.
Step 2 — TLS keypair¶
The gateway has no plaintext mode. Generate the keypair inside this container — macOS ships LibreSSL, whose output the gateway rejects for three different reasons.
mkdir -p ~/axiom-quickstart/certs
docker run --rm -v ~/axiom-quickstart/certs:/certs -w /certs \
--entrypoint /bin/sh alpine/openssl:3.3.3 -c "
openssl req -x509 -newkey ec -pkeyopt ec_paramgen_curve:prime256v1 -nodes \
-days 365 -subj '/CN=axiom-dev-ca' -keyout ca.key -out ca.crt
openssl req -newkey ec -pkeyopt ec_paramgen_curve:prime256v1 -nodes \
-subj '/CN=gateway' -keyout gateway.key -out gateway.csr
printf 'subjectAltName=DNS:axiom-control-plane,DNS:axiom-gateway.axiom-system.svc,DNS:localhost,IP:127.0.0.1\nextendedKeyUsage=serverAuth\n' > san.cnf
openssl x509 -req -in gateway.csr -CA ca.crt -CAkey ca.key -CAcreateserial \
-days 365 -extfile san.cnf -out gateway.crt
rm -f gateway.csr san.cnf ca.srl ca.key
chown $(id -u):$(id -g) ca.crt gateway.crt gateway.key
"
Verify:
ls ~/axiom-quickstart/certs/
docker run --rm -v ~/axiom-quickstart/certs:/certs:ro alpine/openssl:3.3.3 \
x509 -in /certs/gateway.crt -noout -ext subjectAltName
Check it in the container too. openssl x509 -ext is OpenSSL 1.1.1 and later;
the openssl on macOS is LibreSSL, which does not have the flag and fails with
unknown option -ext.
Expect exactly ca.crt, gateway.crt, gateway.key, and a SAN list
containing DNS:axiom-control-plane. If that name is missing, the TLS
handshake in step 5 fails; regenerate rather than continuing.
Step 3 — gateway¶
RAW=https://raw.githubusercontent.com/dhilipkumars/axiom/main/deploy/k8s
kubectl --context kind-axiom apply -f "$RAW/gateway-rbac.yaml"
kubectl --context kind-axiom -n axiom-system delete secret axiom-gateway-tls --ignore-not-found
kubectl --context kind-axiom -n axiom-system create secret generic axiom-gateway-tls \
--from-file=tls.crt="$HOME/axiom-quickstart/certs/gateway.crt" \
--from-file=tls.key="$HOME/axiom-quickstart/certs/gateway.key"
kubectl --context kind-axiom apply -f "$RAW/gateway-deployment.yaml"
kubectl --context kind-axiom -n axiom-system rollout restart deploy/axiom-gateway
kubectl --context kind-axiom -n axiom-system rollout status deploy/axiom-gateway --timeout=180s
The delete secret --ignore-not-found before create is what makes this step
re-runnable; create secret alone fails on a second attempt.
Verify the rollout, and the port step 5 depends on:
kubectl --context kind-axiom -n axiom-system get svc axiom-gateway \
-o jsonpath='{.spec.ports[0].port}:{.spec.ports[0].nodePort}{"\n"}'
Expect 8443:30443. Step 5 dials axiom-control-plane:30443, so if this
prints anything else, fix the endpoint there rather than meeting it as an
opaque TLS error two steps later.
rollout status exits 0 and prints
deployment "axiom-gateway" successfully rolled out. If it times out:
kubectl --context kind-axiom -n axiom-system get pods
kubectl --context kind-axiom -n axiom-system logs deploy/axiom-gateway --tail=30
Step 4 — Postgres with Axiom preinstalled¶
docker rm -f axiom-postgres 2>/dev/null || true
docker run -d --name axiom-postgres \
--network kind \
-e POSTGRES_PASSWORD=axiom \
-v ~/axiom-quickstart/certs:/certs:ro \
ghcr.io/dhilipkumars/axiom-postgres:latest-pg17
Wait for readiness — poll, do not sleep a fixed amount:
for i in $(seq 1 90); do
docker exec axiom-postgres pg_isready -U postgres >/dev/null 2>&1 && break
sleep 2
done
docker exec axiom-postgres pg_isready -U postgres
Expect /var/run/postgresql:5432 - accepting connections.
Verify the extension is loadable, which is what the preload buys:
docker exec axiom-postgres psql -U postgres -tAc "SHOW shared_preload_libraries"
Expect axiom. Anything else means step 5 will fail fatally, not degrade.
Step 5 — connect Postgres to the gateway¶
Use docker exec -i with a heredoc. Never -it: with no TTY attached it
either errors or hangs waiting for one.
docker exec -i axiom-postgres psql -U postgres -v ON_ERROR_STOP=1 <<'SQL'
CREATE EXTENSION IF NOT EXISTS axiom;
DROP SERVER IF EXISTS prod CASCADE;
CREATE SERVER prod
FOREIGN DATA WRAPPER axiom_fdw
OPTIONS (endpoint 'https://axiom-control-plane:30443', ca_cert '/certs/ca.crt');
CREATE USER MAPPING FOR CURRENT_USER SERVER prod;
CREATE SCHEMA IF NOT EXISTS k8s;
IMPORT FOREIGN SCHEMA k8s FROM SERVER prod INTO k8s;
SQL
DROP SERVER IF EXISTS ... CASCADE and the IF NOT EXISTS clauses make this
re-runnable. ON_ERROR_STOP=1 matters: without it psql continues past a
failed statement and exits 0, and you would conclude the step succeeded.
Verify:
docker exec axiom-postgres psql -U postgres -tAc \
"SELECT string_agg(table_name, ',' ORDER BY table_name) FROM information_schema.tables
WHERE table_schema='k8s' AND table_name IN ('core_pods','core_configmaps','apps_deployments','networking_k8s_io_networkpolicies','core_secrets')"
Expect apps_deployments,core_configmaps,core_pods,networking_k8s_io_networkpolicies,
with no core_secrets.
The bundled RBAC reads broadly: everything in Kubernetes' view role plus
nodes, storage, CRDs, RBAC objects, events and metrics. The whole schema holds
several dozen tables, and the exact count depends on the Kubernetes version,
so check for these four rather than for a number. Secrets are never granted,
and their absence is part of the check.
Custom resources appear when their operator ships an aggregate-to-view
role, or once a read-only ClusterRole for them is labelled
axiom.dhilipkumars.github.io/aggregate-to-gateway: "true". A new grant is
seen on the next import; after revoking one, restart the gateway (it caches
what it is allowed). Either way, import again.
Step 6 — success criterion¶
docker exec axiom-postgres psql -U postgres -c \
"SELECT name, namespace, phase, node FROM k8s.core_pods WHERE namespace='kube-system' ORDER BY name"
The environment is working when this returns the cluster's kube-system
pods — on a fresh kind cluster, rows including etcd-axiom-control-plane and
kube-apiserver-axiom-control-plane, all Running, all on node
axiom-control-plane. Compare against the cluster directly if you want a
second source:
kubectl --context kind-axiom -n kube-system get pods
The two must agree. They are the same data by different routes, which is the point of the project.
Failures¶
| Symptom | Cause | Action |
|---|---|---|
no matching manifest for linux/arm64/v8 |
the tag has no slice for this machine. Versions before 0.1.1 are amd64-only and give this legitimately; every tag from 0.1.1 on, including the latest-pgNN this guide uses, should be multi-architecture |
if you pinned an older version, use latest-pgNN or 0.1.1+. Otherwise the publish is broken — confirm with docker manifest inspect ghcr.io/dhilipkumars/axiom-postgres:latest-pg17 and report it. Retrying will not help: a manifest list is tagged atomically |
denied on docker pull |
image is private or the tag does not exist | check the tag; do not retry with credentials |
FATAL: cannot create PGC_POSTMASTER variables after startup |
Axiom loaded without shared_preload_libraries |
use the published image, or preload it |
CREATE EXTENSION closes the connection |
same as above | as above |
TLS handshake failure on the first IMPORT |
certificate SAN does not cover axiom-control-plane |
regenerate the keypair, step 2 |
Cancelled: Timeout expired on IMPORT |
whole-cluster import exceeded rpc_timeout_secs |
ALTER SERVER prod OPTIONS (ADD rpc_timeout_secs '120'); SET only if it is already set |
rollout status times out |
image pull or crash loop | kubectl -n axiom-system logs deploy/axiom-gateway |
psql exits 0 but nothing was created |
missing ON_ERROR_STOP=1 |
re-run with it |
Teardown¶
Leaving a kind cluster running consumes memory indefinitely. Tear down unless asked to keep it:
docker rm -f axiom-postgres
kind delete cluster --name axiom
rm -rf ~/axiom-quickstart
Constraints worth knowing before you suggest things¶
- Axiom cannot be installed on managed Postgres. It is not a trusted
extension and needs
shared_preload_libraries, so RDS, Cloud SQL and Aurora cannot run it. Do not suggest them. - Installing into an existing Postgres uses a release package, published
per major and architecture from v0.1.1 onward —
postgresql-<major>-axiomas a.deb,axiom_<major>as an.rpm, or a tarball where neither applies. v0.1.0 has none, so check the releases page for the version you name before telling someone to download one. Recommend the package over the tarball: it refuses to install on a glibc below 2.34 rather than failing at the next postmaster start. The images remain the path for trying Axiom without touching an existing Postgres. - Every image is published for amd64 and arm64 from v0.1.1 on, so nothing runs under emulation and no
--platformflag is needed. Versions before that are amd64-only. - What a query can reach is bounded by the gateway's RBAC, not the SQL
user's. Reads are broad by default and never include Secrets. To expose a
custom resource, label a read-only ClusterRole
axiom.dhilipkumars.github.io/aggregate-to-gateway: "true"; to make one writable, grant its verbs in theaxiom-gatewayClusterRole. Then re-import — foreign tables are catalog objects and do not follow the change. Only revoking a grant needs a gateway restart.