Fri 18 Sep · 17:00 CEST · onlineCommunity Town Hall: the roadmap, the three colours, and a live provisioning demoAgenda and registration →

neon-multi-node

Context Skill

Use when self-hosted Neon says "server does not support SSL, but SSL was required", compute_ctl logs "unknown TLS key type", a safekeeper outage rehearsal fails with "FAIL AssertionError", psql returns "INSERT 0 1" instead of the expected value, a timed-out quorum write appears after recovery, cleanup reports Ansible ENOENT, repeat delete reports migration after removing its backend, or three safekeepers and S3 are being mistaken for complete HA. Also use when changing Red/Blue runtimes on one profile, negative subprocess exits bypass failure gates, or Bun cold dependency resolution fails. Verified AWS multi-node TLS, SQL witness, WAL quorum, S3 recovery and lifecycle contracts.

Installation
npx skills add https://github.com/getcolors/skills --skill neon-multi-node

Neon Multi-Node

Route by symptom

Symptom or situation Read
SSL required but server has SSL off TLS feature gate
unknown TLS key type, connection closes after enabling experimental TLS Certificate parser
Outage write returns FAIL AssertionError despite an inserted row psql command tags
Timed-out no-quorum write exists after quorum returns Uncertain transaction outcome
Three safekeepers or S3 presented as complete HA or zero RPO Acceptance boundaries
Installed Ansible reports ENOENT during deletion Scaffold event
Repeat delete reports migration after backend retirement Missing bucket vs missing key
Switching native runtimes on an existing deployment Runtime handoff proof
Timeout or signal is reported as success SDK exit-status boundary
Copied Red launcher fails to resolve pinned Git dependencies Cold Bun resolution
Resource template cannot be found, or ingress validation fails offline Build traps

Provenance and ownership

This knowledge comes from the September 10, 2026 live AWS deployment of getcolors/neon-multi-node, with evidence in neon-multi-node-aws. Claims below were verified against that deployment unless labeled source review, offline validation, or unverified. The exact pins delimit what was tested. This skill carries reasoning; the companion package owns all working files. Do not reconstruct its scripts from this prose.

The verified shape is five Ubuntu machines in one AWS availability zone: compute-0, pageserver-0 with the broker, and safekeeper-0/1/2. Six containers run on those machines. Three independent safekeeper disks provide a WAL quorum; the one compute and one pageserver remain service failure points. Native PostgreSQL TLS is reached through a DNS-only Cloudflare record and a restricted client CIDR. Managed S3 backend and application buckets are separate resources.

The original Green recovery and deletion proof below is retained separately from the later native Red/Blue verification. Runtime handoff was verified in both directions across two fresh lifecycles, with recovery and complete cleanup separately checked for each native runtime. Switching still requires matching profile, dependency pins and state contracts; invoke only one lifecycle operation for a profile at a time.

Diagnose the boundary that actually failed

A valid TLS field in a compute specification did not enable TLS at the pinned release. Reading compute.rs::tls_config showed the experimental feature gate. Enabling it exposed a second, independent issue: tls.rs::verify_key_cert rejected the actual ACME certificate's ECDSA/SHA384 signature algorithm. The accepted deployment uses native PostgreSQL SSL settings and its mounted certificate/key, bypassing that helper. See the source contract before treating the feature flag as a fix.

A successful SQL operation and a successful harness assertion are different facts. The first safekeeper rehearsal acknowledged its INSERT, then asserted on psql's trailing command tag. The row existed, but its witness ledger did not. Quiet psql output fixed the harness; a new complete rehearsal supplied the recovery proof. That failed attempt must not be relabeled a passed gate.

A client timeout likewise does not establish rollback. With two safekeepers stopped, the write was not acknowledged within the bounded client wait. After quorum returned, the uncertain row was present. Applications need a way to reconcile that outcome; an automatic retry must not assume the first attempt never happened.

Hold success to these gates

Use the companion's implementation and acceptance doctrine: external hostname verification, authentication and plaintext negatives, exact persistent witnesses, per-member outage witnesses, observable S3 uploads, scoped-storage denial checks, and recovery readback. Container liveness and an arbitrary S3 object are supporting observations, not substitutes.

The verified recovery preserved the original random external witness and six unique acknowledged outage witnesses through compute recreation, individual safekeeper outages, quorum interruption, broker restart, and pageserver tenant cache removal followed by reattachment, and empty local-state reconstruction of one safekeeper through the other two. Safekeeper WAL survived the pageserver recovery. This does not prove recovery after simultaneous loss of every storage disk, automatic compute/pageserver failover, multi-zone availability, or a numerical full-storage-loss RPO.

Deletion is a separate acceptance boundary. The package's guarded workflow requires current-run proof that every writer stopped before application S3 purge, then completes compute and backend retirement. Treat final cloud absence as an independent check; complete cleanup and repeated deletion passed, with zero remaining deployment resources independently observed. Details are in acceptance.