Writing The Day CREATE EXTENSION Became a Platform Decision

Explore PostgreSQL extension security, compatibility, package provenance and lifecycle governance with clear ownership and reliable recovery

profile_img

DATABASE ARCHITECTURE

The Day CREATE EXTENSION Became a Platform Decision

Why useful PostgreSQL capabilities need a catalogue, compatibility evidence and an owner.

A developer asked for one extension to solve one problem. The SQL took seconds. The review uncovered an operating-system package, a shared library, a preload change, background workers and an upgrade dependency that would live for years.

Start with one workload | Scale into a complete platform | Open Source Freedom | Enterprise-Grade Operations.

Published | January 2026

The question behind the obvious answer

A developer asked for one extension to solve one problem. The SQL took seconds. The review uncovered an operating-system package, a shared library, a preload change, background workers and an upgrade dependency that would live for years.

Extensions are part of PostgreSQL’s strength, but unmanaged extensions turn database freedom into hidden platform coupling.

An Extension Is Code Inside The Database Trust Boundary.

The turning point

The platform introduced an approved catalogue that connected business use, PostgreSQL version, package provenance, runtime requirements, licensing and rollback constraints.

The team used OSS Manager to make approved PostgreSQL extension lifecycle reviewable and repeatable. The engine still owned its native behaviour; the control plane made intent, prerequisites, execution and evidence visible to the people responsible for the service.

What had to become explicit

  • The service boundary: An approved extension catalog maps PostgreSQL and operating-system versions to signed packages, preload requirements, database enablement, rollback constraints, and test evidence.
  • The operational controls: allow-list, package signature and SBOM, binary compatibility, preload change review, dependency graph, database scope, backup, rollback, and license review.
  • The capacity conversation: extension workload, memory and background workers, index growth, maintenance cost, query patterns, package footprint, and upgrade duration.
  • The human boundary: who may observe, who may approve, who may execute, and who decides whether the application is ready.

The architecture that changed the conversation

The diagram is deliberately centred on the decision the team had to make. It is not a product inventory. It shows where authority sits, what crosses the boundary, and where a failed assumption must stop the workflow.


The practical design choices

GateQuestionEvidence
CompatibilityWill it survive the next major version?Version matrix
RuntimeDoes it need preload or workers?Restart and resource plan
RecoveryCan backups restore the objects?Restore test
SecurityWhat code enters the trust boundary?Package provenance and privilege review
LicensingCan it be redistributed and supported?Approved licence and ownership record

These choices are intentionally small enough to review and test. They keep the architecture tied to operating behaviour instead of allowing a visually impressive diagram to hide unclear ownership.

What OSS Manager changed - and what it did not

OSS Manager brought discovery, planning, guarded execution and normalised status into one path for PostgreSQL. It did not replace the engine's correctness model or the application's responsibility for data semantics.

  • Before change, validate versions, hosts, identities, artifacts, topology and the recovery boundary without mutating the service.
  • During change, persist progress, expose stop conditions and prevent a partial result from being mistaken for success.
  • After change, observe extension version; load success; background worker health; error log; query latency; resource use; package drift; unsupported dependency and run representative application journeys, not only process checks.
  • For recovery, protect the last known-good authority and require an explicit decision before promotion, rollback or destructive cleanup.

The failure modes worth rehearsing

The test plan should make room for missing shared library, incompatible ABI, preload error, dependency conflict, extension update failure, catalog drift, unsupported downgrade, and license exception. The purpose is not to produce a longer checklist. It is to learn whether operators can recognise the failure and choose the safe next action while the evidence is incomplete.

A human operating model

L1 operators need a plain-language answer to what is healthy, what is delayed and whether a change is in progress. Platform engineers need topology and evidence. Application owners need to know what users will experience. A useful platform connects those views without giving every person the same privileges.

The catalogue did not reduce innovation. It shortened the path from a good idea to a supportable production capability.

Where this pattern earns its place

Media and OTT

This pattern is relevant where teams need regional continuity, burst traffic, low-latency serving and predictable recovery. The architecture must still be adjusted for local data classification, recovery objectives, workload shape and application behaviour.

Retail

This pattern is relevant where teams need campaign peaks, catalogue freshness, transaction continuity and rapid rollback. The architecture must still be adjusted for local data classification, recovery objectives, workload shape and application behaviour.

Telecom

This pattern is relevant where teams need high event volume, distributed operations, identity boundaries and service assurance. The architecture must still be adjusted for local data classification, recovery objectives, workload shape and application behaviour.

The lesson we would carry into the next project

The catalogue did not reduce innovation. It shortened the path from a good idea to a supportable production capability.

The strongest open-source platforms are not the ones with the most automation. They are the ones where automation makes ownership, risk and recovery easier for people to understand.

Start with one workload. Prove the operating model. Then scale it into a complete platform.