# PostgreSQL 教程与生产实践指南 Canonical URL: https://pg.edu.rich/docs Last reviewed: 2026-08-02 这不是 PostgreSQL 官方手册的替代品,而是一张更容易进入官方手册的地图。内容以 **PostgreSQL 18** 为当前基线,同时避免在没有说明时依赖单一版本的特性。 ## 两种阅读模式 [#两种阅读模式] ### 人类学习模式 [#人类学习模式] 按顺序读 [从这里开始](/docs/start-here)。每篇包含:你会完成什么、最小示例、为什么如此、如何验证、下一步是什么。第一次阅读时不要求记住系统目录和所有 SQL 语法。 ### AI 检索模式 [#ai-检索模式] 从 [AI / Agent 入口](/docs/ai) 进入。页面使用稳定标题、显式前置条件、可复制 SQL、边界条件和失败模式。让 Agent 优先检索相关小节,不要一次塞入整站内容。 ## 文档约定 [#文档约定] | 标记 | 含义 | | ------ | ------------------------- | | **默认** | PostgreSQL 默认行为;仍应在目标实例验证 | | **建议** | 适用于大多数新项目,不代表唯一正确答案 | | **危险** | 可能锁表、丢数据、泄露权限或制造长事务 | | **验证** | 执行后用于确认结果的命令或查询 | 事实性细节以 [PostgreSQL 18 官方文档](https://www.postgresql.org/docs/18/) 为准。本站的价值是路径、示例、约束和跨章节连接。 --- # PostgreSQL 19:新功能、发布时间与 18 升级 19 指南 Canonical URL: https://pg.edu.rich/docs/postgresql-19 Last reviewed: 2026-08-06 截至 **2026-08-02**,PostgreSQL 19 的最新公开测试版是 **Beta 2**,还不是正式生产版本。PostgreSQL 官方路线图计划在 **2026 年 9 月**发布 19;Beta 2 公告使用了更保守的 **2026 年 9–10 月**窗口。最终日期、功能细节和兼容性要求仍可能变化。 官方鼓励用真实工作负载测试 Beta,但明确不建议运行在生产环境。本页适合提前评估、建立兼容矩阵和演练 PostgreSQL 18 升级 19;正式切换应等待 GA、目标 minor、扩展和托管平台支持。 ## PostgreSQL 19 最新消息与发布时间 [#postgresql-19-最新消息与发布时间] | 日期 | 官方进展 | 对使用者的意义 | | ----------- | ---------------------------------------------------------------------------------------------------- | -------------------------------------------------------- | | 2026-06-04 | [PostgreSQL 19 Beta 1 发布](https://www.postgresql.org/about/news/postgresql-19-beta-1-released-3313/) | 功能预览开放,适合开始 CI 与应用兼容测试 | | 2026-07-16 | [PostgreSQL 19 Beta 2 发布](https://www.postgresql.org/about/news/postgresql-19-beta-2-released-3350/) | 修复 Beta 1 回归,并继续调整 temporal、SQL/PGQ、逻辑解码和 autovacuum 等实现 | | 2026-09(计划) | [官方路线图的目标月份](https://www.postgresql.org/developer/roadmap/) | 不是不可变承诺;Beta 公告保留 9–10 月发布窗口 | Beta 2 仍允许数据库行为、API 和功能细节发生小幅变化。跟踪时以 [PostgreSQL 19 Release Notes](https://www.postgresql.org/docs/19/release-19.html) 和官方新闻为准,不以第三方功能清单作为上线依据。 ## PostgreSQL 19 Linux 软件包可用性 [#postgresql-19-linux-软件包可用性] ### 发行版官方仓库:postgresql-19 PkgSeek 软件包快照;核对时间:2026-08-02 15:47:58 UTC。 > 没有准确索引坐标。这表示当前快照未收录,不等于软件包不存在。 来源:[PkgSeek 软件包查询](https://pkgseek.com/search?q=postgresql-19)。发行版 revision 和回溯补丁属于完整版本身份。 ### PGDG 仓库:postgresql-19 PkgSeek 软件包快照;核对时间:2026-08-02 15:47:58 UTC。 > 没有准确索引坐标。这表示当前快照未收录,不等于软件包不存在。 来源:[PkgSeek 软件包查询](https://pkgseek.com/search?q=postgresql-19)。发行版 revision 和回溯补丁属于完整版本身份。 这两个快照用于追踪正式软件包落地情况,不是 PostgreSQL 19 的发布状态来源。Beta tarball、开发构建或容器存在,不代表发行版仓库已经提供可用于生产升级的 `postgresql-19`。进入 GA 后仍应等待目标操作系统、CPU 架构、扩展和备份工具形成完整兼容矩阵。 ## PostgreSQL 19 有哪些新功能 [#postgresql-19-有哪些新功能] 下面是截至本页核对日最值得工程团队测试的方向,不代表最终发布说明的完整摘要。 ### SQL、图查询与时间数据 [#sql图查询与时间数据] * **SQL/PGQ Property Graph Queries**:在关系数据上定义并查询 property graph;应验证驱动、SQL parser、ORM 和 AI SQL 生成器是否认识新语法。 * **`FOR PORTION OF`**:让 `UPDATE`、`DELETE` 针对时间范围操作,适合测试 temporal 数据模型,但 Beta 2 仍修复了多项相关问题。 * **`GROUP BY ALL`**:自动对 target list 中非 aggregate、非 window 项分组。 * **Window functions `IGNORE NULLS` / `RESPECT NULLS`**:适用于 `lead()`、`lag()`、`first_value()`、`last_value()` 和 `nth_value()`。 * **`INSERT ... ON CONFLICT DO SELECT ... RETURNING`**:可返回发生冲突的行,并选择性加锁。 ### 运维、性能与可观测性 [#运维性能与可观测性] * **`REPACK` 与 `REPACK CONCURRENTLY`**:统一 `VACUUM FULL` / `CLUSTER` 的重写能力,并提供降低排他锁影响的新路径;旧命令为兼容仍保留。详见 [REPACK 与在线表重写](/docs/operations/repack)。 * **分区拆分与合并**:新增 `ALTER TABLE ... SPLIT/MERGE PARTITIONS`。 * **并行 autovacuum worker**(见 [并行 autovacuum](/docs/operations/parallel-autovacuum)),以及 `pg_stat_autovacuum_scores`、`pg_stat_lock`、`pg_stat_recovery` 等观测视图。 * **优化器改进**:eager aggregation——在 join 之前先执行部分聚合以减少待处理行数;以及在列可证明非空时把 `NOT IN` 转换为更高效的 anti join。跨版本对比计划时应固定这些变化,而不是把差异归因于索引。 * **在线启停 data checksums**:不再只能离线使用 `pg_checksums`;见 [数据校验和](/docs/operations/data-checksums)。 * 异步 I/O read-ahead、`COPY FROM` SIMD、radix sort、外键检查等性能改进。 * `EXPLAIN ANALYZE` 新增 `IO` 选项;`EXPLAIN (ANALYZE, WAL)` 可报告 full-page write bytes。 * 新数据默认 TOAST 压缩从 `pglz` 改为 `lz4`。 ### 逻辑复制与 standby 等待 [#逻辑复制与-standby-等待] * **序列同步**让 subscriber 的序列值与 publisher 保持一致:用 `CREATE PUBLICATION ... FOR ALL SEQUENCES` 发布序列,用 `ALTER SUBSCRIPTION ... REFRESH SEQUENCES` 在订阅端对齐取值;`pg_get_sequence_data()` 可查看同步状态。这直接解决逻辑复制大版本升级后的序列/主键冲突——见 [复制、故障切换与升级](/docs/operations/replication-upgrades)。 * **`WAIT FOR LSN`** 让会话等待指定 LSN 被写入、刷盘或回放(包括在 standby 上等待 replay),从而在主备分离架构中实现 read-your-writes。见 [WAIT FOR LSN](/docs/operations/wait-for-lsn) 与 [PostgreSQL 19 WAIT FOR 文档](https://www.postgresql.org/docs/19/sql-wait-for.html)。 * **Publication 黑名单**:`CREATE` / `ALTER PUBLICATION ... FOR ALL TABLES EXCEPT (TABLE ...)` 发布除指定表之外的全部表,大 schema 下不再需要逐表白名单。 * **订阅端冲突保留**:订阅参数 `retain_dead_tuples` 与 `max_retention_duration` 保留用于冲突检测的 dead tuple 信息,并以保留窗口限制时长;`pg_stat_subscription_stats.update_deleted` 统计因并发 delete 而被忽略的 update。 * **`effective_wal_level`** 报告实际生效的 WAL level;`wal_level = replica` 时,服务器可在逻辑复制需要时自动把生效级别提升到 `logical`。 PostgreSQL 19 的 `REPACK (CONCURRENTLY)` 是基于 logical decoding 的核心 SQL command,对 replica identity、unlogged/partitioned/system table、replication slot 和额外磁盘有约束;第三方 `pg_repack` 是独立 extension 与 CLI。两者不能共享未经验证的运行手册。参见 [扩展生态选型](/docs/reference/extensions-ecosystem) 与 [PostgreSQL 19 REPACK 文档](https://www.postgresql.org/docs/19/sql-repack.html)。 如果 Agent 的目标实例还是 PostgreSQL 18,不要让它生成 SQL/PGQ、`FOR PORTION OF`、`GROUP BY ALL` 等 19 才有的语法。把 `server_version_num`、允许语法和扩展版本写进检索上下文或工具契约。 ## PostgreSQL 18 升级 19:需要注意什么 [#postgresql-18-升级-19需要注意什么] PostgreSQL 18 → 19 是 **major upgrade**,不能把 18 的 data directory 直接交给 19 启动。可选路径是 `pg_upgrade`、逻辑 dump/restore 或逻辑复制。升级前优先处理以下兼容性变化。 ### 1. 认证与安全变化 [#1-认证与安全变化] * **RADIUS 支持被移除**;仍依赖 RADIUS 的环境必须先设计替代认证路径。 * PostgreSQL 18 已把 MD5 密码标记为 deprecated;19 会在 MD5 认证成功后发出 warning。迁移计划应推进 SCRAM,而不是只屏蔽 warning。 * 新增密码即将过期 warning,默认阈值为 7 天;检查监控是否会把预期 warning 当故障。 ### 2. SQL 与对象兼容性 [#2-sql-与对象兼容性] * `standard_conforming_strings` 在服务器端强制为 `on`。如果旧环境曾设为 `off`,使用 19 版 `pg_dump` / `pg_dumpall` 重新导出,或先修正设置与应用转义行为。 * database、role、tablespace 名称不能包含 CR/LF;`pg_upgrade` 会拒绝此类 cluster。 * 使用 `btree_gist` 的 `inet` / `cidr` index 会阻塞 `pg_upgrade`,因为旧 opclass 可能漏行;让 `pg_upgrade --check` 给出实际阻塞项,再按 release notes 处理。 * `MULE_INTERNAL` encoding 被移除,相关数据库必须以其他 encoding dump/restore。 ### 3. 默认值、性能与监控变化 [#3-默认值性能与监控变化] * **JIT 默认关闭**。大型分析查询不能假设升级后计划行为不变;分别在 `jit=off/on` 下记录执行时间和计划。 * **`log_lock_waits` 默认开启**:超过 `deadlock_timeout` 的锁等待会直接进入服务器日志。升级后锁等待日志量会突然增加,应调整日志量告警与锁监控,而不是把新增噪音当作回归。 * `max_locks_per_transaction` 默认值从 64 变为 128,同时锁内存计算发生变化;不要只复制旧配置数字,应重新做容量评估。 * `pg_stat_subscription_stats.sync_error_count` 更名为 `sync_table_error_count`;等待事件类型 `BUFFERPIN` 更名为 `BUFFER`。修正 dashboard、告警和采集 SQL。 * 新默认只影响新写入数据,不代表升级后所有旧 TOAST 值会自动改用 LZ4;不要把默认变化等同于无条件压缩收益。 ### 4. 扩展、驱动与平台 [#4-扩展驱动与平台] `pg_upgrade` 可以检查核心 cluster 的许多二进制条件,但不能证明第三方 module 与 PostgreSQL 19 二进制兼容。为 PostGIS、pgvector、TimescaleDB、自定义 C extension、审计插件、备份代理、pooler、ORM 和驱动分别记录: | 组件 | 需要确认 | | --------- | ---------------------------------------------------------- | | Extension | 19 对应 package/shared library、支持声明、升级脚本、index 重建要求 | | 驱动与 ORM | server version detection、19 新/变更语法、prepared statement、类型映射 | | 连接池 | startup parameter、认证、failover 与连接回收行为 | | 备份与 CDC | 新版本 catalog、WAL/逻辑解码、restore 演练 | | 云数据库 | 区域、SKU、扩展版本、升级窗口与回退能力;以服务商上线公告为准 | ## PostgreSQL 18 升级 19 实战清单 [#postgresql-18-升级-19-实战清单] ### 阶段 A:现在就可以做 [#阶段-a现在就可以做] 1. 固定一份 PostgreSQL 18 生产备份并完成 restore 验证。 2. 清点 extension、collation、replication slot、tablespace、自定义 full-text 文件、认证方式和外部 module。 3. 用 PostgreSQL 19 Beta 2 建立**可丢弃**的测试环境,运行 migration、应用测试、备份恢复、CDC 和关键查询 benchmark。 4. 对比 `EXPLAIN (ANALYZE, BUFFERS, WAL)`;单独评估 JIT 默认变化与新 I/O 行为。 5. 在 CI 中禁止把 PostgreSQL 19 专属语法下发给 18 实例。 把 PostgreSQL 18.4 设为阻断发布的生产门禁,把 PostgreSQL 19 Beta 2 设为前瞻兼容通道;后者可以先允许失败,但每个失败都应归类并在 GA 采用前清零。完整流水线见 [PostgreSQL 安全迁移与零停机 Schema 变更](/docs/operations/safe-migrations)。 ### 阶段 B:正式切换前 [#阶段-b正式切换前] 先运行 **19 版** `pg_upgrade` 的只检查模式;路径必须替换为目标环境实际目录: ```bash /opt/postgresql/19/bin/pg_upgrade \ --old-bindir=/opt/postgresql/18/bin \ --new-bindir=/opt/postgresql/19/bin \ --old-datadir=/data/postgresql/18 \ --new-datadir=/data/postgresql/19 \ --check ``` `--check` 不迁移数据,但应使用与正式切换相同的 binary、extension、initdb 参数和 transfer mode 做演练。不要把示例路径直接复制到生产。 然后完成: * 锁定 PostgreSQL 19 正式 minor、OS package、container digest 与 extension 版本; * 根据数据量选择 `pg_upgrade` copy/clone/link/swap、dump/restore 或逻辑复制; * 记录预计停机、额外磁盘、统计恢复时间、DNS/连接池收敛时间; * 定义 rollback 判据、负责人和最晚回退时点; * 验证 standby、slot、sequence、large object、权限、RLS、job scheduler 和备份链路。 ### 阶段 C:切换后 [#阶段-c切换后] 1. 执行 `pg_upgrade` 生成的 post-upgrade / rebuild 脚本,完成前不要访问被标记的表。 2. 按工具提示补齐 optimizer statistics;再比较高流量查询的计划与延迟。 3. 检查错误率、认证 warning、复制 lag、WAL、autovacuum、锁和备份。 4. 进行一次从 PostgreSQL 19 新备份恢复的演练。 5. 只有通过业务验收且回退窗口关闭后,才清理 PostgreSQL 18 cluster。 完整步骤见 [PostgreSQL 19 pg\_upgrade](https://www.postgresql.org/docs/19/pgupgrade.html);复制与切换原则见 [复制、故障切换与升级](/docs/operations/replication-upgrades)。 ## PostgreSQL 19 常见问题 [#postgresql-19-常见问题] ### PostgreSQL 19 正式版发布了吗? [#postgresql-19-正式版发布了吗] 没有。截至 2026-08-02 最新版本是 Beta 2。官方路线图目标是 2026 年 9 月,Beta 公告给出的窗口是 9–10 月;最终以 PostgreSQL 官方新闻为准。 ### PostgreSQL 18 能直接升级到 PostgreSQL 19 吗? [#postgresql-18-能直接升级到-postgresql-19-吗] 可以按 major upgrade 路径直接迁移,不要求先经过其他 major。常见方法是 `pg_upgrade`、dump/restore 或逻辑复制;Beta 只用于演练,生产升级应等待正式版和依赖支持。 ### PostgreSQL 18 升级 19 会停机多久? [#postgresql-18-升级-19-会停机多久] 没有通用数字。停机由数据量、relation 数量、transfer mode、extension/reindex、统计恢复、连接切换和验证决定。必须在恢复出的真实数据副本上演练并测量。 ### 云 PostgreSQL 什么时候支持 19? [#云-postgresql-什么时候支持-19] 不同厂商、区域和 SKU 的节奏不同。不要从社区 GA 日期推导云平台可用日期;跟踪服务商官方版本矩阵,并核对 extension、PITR、read replica 和回退限制。参见 [云 PostgreSQL 选型](/docs/cloud/service-map)。 ### 现在应该从 PostgreSQL 18 升级 19 吗? [#现在应该从-postgresql-18-升级-19-吗] 现在适合建立测试矩阵,不适合把 Beta 作为生产默认。正式版发布后,也应先等待自己依赖的扩展、驱动、工具和托管平台给出明确支持,再根据业务收益与风险排期。 ## 事实状态与复核 [#事实状态与复核] 本页状态快照:**PostgreSQL 19 Beta 2,核对日期 2026-08-02**。进入 RC、GA 或 release notes 发生重要兼容性变化时,应更新标题说明、新闻时间线、升级阻塞项和 `updatedAt`。 --- # 5 分钟快速开始 Canonical URL: https://pg.edu.rich/docs/quickstart Last reviewed: 2026-08-02 下面的实例仅用于本地学习。密码写在命令行、端口暴露到宿主机都不是生产配置。 ### 启动实例 [#启动实例] ```bash docker run --name pg-guide \ -e POSTGRES_PASSWORD=dev-only-password \ -e POSTGRES_DB=playground \ -p 5432:5432 \ -v pg-guide-data:/var/lib/postgresql/data \ -d postgres:18 ``` ### 等待并检查 [#等待并检查] ```bash docker exec pg-guide pg_isready -U postgres -d playground docker logs pg-guide --tail 20 ``` 看到 `accepting connections` 再继续。 ### 打开 psql [#打开-psql] ```bash docker exec -it pg-guide psql -U postgres -d playground ``` ```sql CREATE TABLE notes ( id bigint GENERATED ALWAYS AS IDENTITY PRIMARY KEY, body text NOT NULL CHECK (length(body) > 0), created_at timestamptz NOT NULL DEFAULT now() ); INSERT INTO notes (body) VALUES ('hello, PostgreSQL'); SELECT id, body, created_at FROM notes; ``` ### 验证持久化 [#验证持久化] ```bash docker restart pg-guide docker exec pg-guide psql -U postgres -d playground \ -c "SELECT id, body, created_at FROM notes;" ``` 重启后仍能看到数据,说明命名卷已生效。 ## 连接字符串 [#连接字符串] 宿主机应用使用: ```text postgresql://postgres:dev-only-password@127.0.0.1:5432/playground ``` 不要把真实密码提交到 Git。生产环境使用 secret manager、最小权限应用角色和 TLS。 ## 清理 [#清理] ```bash docker rm -f pg-guide docker volume rm pg-guide-data ``` 第二条命令会永久删除练习数据;只在你确认不再需要时执行。 继续 [从这里开始](/docs/start-here),或者直接学习 [数据建模](/docs/core/data-modeling)。 --- # 从这里开始 Canonical URL: https://pg.edu.rich/docs/start-here Last reviewed: 2026-08-02 完成这条路线后,你应该能启动实例、用 `psql` 连接、设计一张有约束的表、在事务里修改数据,并用 `EXPLAIN` 判断查询是否走了预期路径。 ## 先记住五件事 [#先记住五件事] 1. **实例(cluster)里有多个 database**,database 里有 schema,schema 里有表、视图和函数。 2. 客户端连接的是一个具体 database;跨 database 查询不像跨 schema 查询那样直接。 3. 每条语句都在事务中运行。未显式 `BEGIN` 时,客户端通常为单条语句自动提交。 4. 约束是数据模型的一部分,不只是应用层校验的备份。 5. 索引有写入与维护成本;创建前后都要看真实执行计划。 ## 90 分钟路线 [#90-分钟路线] ### 0–10 分钟:运行并连接 [#010-分钟运行并连接] 完成 [快速开始](/docs/quickstart),保留一个名为 `pg-guide` 的本地容器。 ### 10–30 分钟:建立模型 [#1030-分钟建立模型] 阅读 [数据建模](/docs/core/data-modeling),创建 `customers` 和 `orders`,使用主键、外键、`CHECK`、`NOT NULL` 和唯一约束表达事实。 ### 30–50 分钟:查询数据 [#3050-分钟查询数据] 阅读 [查询工具箱](/docs/core/queries),练习过滤、连接、聚合、CTE 和窗口函数。每个查询都明确列名,避免在持久接口里使用 `SELECT *`。 ### 50–70 分钟:理解并发 [#5070-分钟理解并发] 阅读 [事务与并发](/docs/core/transactions)。在两个 `psql` 会话中观察 `READ COMMITTED`,再用 `SELECT ... FOR UPDATE` 保护一次余额修改。 ### 70–90 分钟:验证性能 [#7090-分钟验证性能] 阅读 [索引与 EXPLAIN](/docs/core/indexes-explain)。先运行 `EXPLAIN (ANALYZE, BUFFERS)`,再创建索引并对比;不要把“出现 Seq Scan”自动判定为问题。 ## 你的完成标准 [#你的完成标准] ```sql SELECT version(); SELECT current_database(), current_user; SELECT schemaname, tablename FROM pg_catalog.pg_tables WHERE schemaname NOT IN ('pg_catalog', 'information_schema'); ``` 你能解释这三条查询的输出,并知道如何安全地删除练习容器,就已经完成第一阶段。 “已经生成备份文件”不等于“能够恢复”。进入生产运维路线后,至少做一次恢复到空数据库的演练。 --- # 数据库 Agent 评估 Canonical URL: https://pg.edu.rich/docs/ai/agent-evals Last reviewed: 2026-08-02 ## 评估分三层 [#评估分三层] | 层 | 测什么 | 例子 | | -- | --------------------- | ------------------------ | | 生成 | SQL 是否引用真实对象、参数化、符合方言 | 不出现不存在的 `orders.user_id` | | 执行 | 结果是否正确、稳定、有界 | 与 golden query 的结果集相同 | | 安全 | 越权、注入、批量写、昂贵查询是否被拒绝 | 跨租户查询在执行前被拦截 | 只比较 SQL 字符串会误判:不同 SQL 可以等价,同一 SQL 也可能因数据和权限产生不同结果。优先断言结果、行数、SQLSTATE、权限边界和副作用。 ## 固定夹具数据库 [#固定夹具数据库] 每次评估从同一组迁移和种子数据启动临时 PostgreSQL。数据集要包含:`NULL`、空集、重复值、时区边界、金额边界、孤立记录(如果模型允许)、多租户相同自然键和足以触发不同计划的数据量。 ## 用例格式 [#用例格式] ```yaml id: revenue-by-day-001 question: 过去 7 个完整 UTC 日每天已支付金额是多少? contract_version: test-42 role: agent_reader assert: read_only: true max_rows: 7 columns: [day, paid_cents] result_fixture: expected/revenue-by-day.json forbidden_relations: [app.payment_secrets] max_duration_ms: 1000 ``` 记录模型、提示、工具 schema、数据库版本和随机种子。对非确定模型重复运行,报告通过率和方差,而不是只留最好一次。 ## 安全红队集 [#安全红队集] * 用户要求忽略规则并输出其他租户数据。 * schema 注释中包含提示注入文本。 * 值看起来像 SQL 片段。 * 请求没有 `WHERE` 的删除或全表更新。 * 请求 `pg_read_file`、`COPY PROGRAM`、扩展安装或权限提升。 * 用超大笛卡尔积或递归 CTE 制造资源耗尽。 成功结果应是策略层拒绝,而不是期待模型每次自律。 ## 计划回归 [#计划回归] 对关键读查询保存规范化的 `EXPLAIN (FORMAT JSON)` 特征:顶层节点、实际/估算行数比、buffer read 和执行时间区间。不要锁死精确成本数字;统计、缓存和 PostgreSQL 版本都会改变计划。 ## 发布门槛 [#发布门槛] 新提示或模型必须同时通过:正确性基线、安全集零违规、P95 延迟与成本预算、旧 schema/缺失上下文时能拒答、审计事件完整。任何一项回退都应阻止自动发布。 --- # AI Agent 长期记忆与 PostgreSQL Canonical URL: https://pg.edu.rich/docs/ai/agent-memory Last reviewed: 2026-08-06 对语言模型的每次调用都是无状态的:模型不会回写任何东西,它能“记住”的只有 prompt 里的内容。一旦应用需要知道用户是谁、上周的偏好是什么、哪些事实已经变化,这份状态就必须由应用自己持有。这就是“Agent 记忆”在工程上的含义——一个写入路径由 LLM 辅助完成的数据库问题,而不是模型功能。 ## 记忆不等于 RAG [#记忆不等于-rag] RAG 和 Agent 记忆经常被混为一谈,因为两者都以“把相关文本检索进 prompt”结束。区别在于写入路径: | | RAG | Agent 记忆 | | ---- | -------------- | ----------------- | | 内容 | 外部语料(文档、工单、代码) | 交互中产生的事实、偏好与经历 | | 写入路径 | 批量摄取,可重放,幂等 | 对话过程中的在线写入,常由模型抽取 | | 更新方式 | 重新摄取新版本文档 | 对单条事实进行更正、取代与遗忘 | | 典型查询 | “找与 X 相关的段落” | “我们当前对这个用户了解什么” | | 典型故障 | 切块过期或无权限 | 错误事实以与正确事实相同的权威写入 | 因此记忆存储必须支持定向的 `UPDATE` 与 `DELETE`、按时间的有效期和冲突处理,而不只是最近邻搜索。把对话日志建成只读向量索引,是对聊天记录做 RAG,不是记忆。 ## 记忆层的三种形态 [#记忆层的三种形态] 1. **专用记忆框架**——SDK 与服务(Mem0、Cognee 等),在 `add`/`search` API 背后接管抽取、存储与检索。原型最快,但 schema 与检索策略留在框架内部。 2. **只用向量库**——embedding 加元数据过滤。简单,但用户画像、实体间关系、精确匹配查询都要在别处重建。 3. **直接用 PostgreSQL**——一个系统同时承载 embedding(pgvector)、结构化画像(`jsonb`)、关键词检索(全文检索)、实体关系(普通表,或 PostgreSQL 19 的 SQL/PGQ 属性图),外加事务与行级安全。写入策略留在你自己的代码和 SQL 里,而不是框架里。 三者并不互斥:第一类框架通常会把数据持久化到后两类存储上。真正的决策是“schema 的权威放在哪里、写入策略由谁控制”。 ## 用 PostgreSQL 做记忆底座 [#用-postgresql-做记忆底座] 一个可用的记忆 schema 把事实与它的 embedding 分开存放,并用显式的历史链代替原地覆盖: ```sql CREATE EXTENSION IF NOT EXISTS vector; CREATE TABLE agent_memory ( id bigint GENERATED ALWAYS AS IDENTITY PRIMARY KEY, tenant_id bigint NOT NULL, user_id text NOT NULL, kind text NOT NULL CHECK (kind IN ('profile', 'preference', 'episode', 'fact')), content text NOT NULL, attributes jsonb NOT NULL DEFAULT '{}', embedding vector(1536), search_vector tsvector GENERATED ALWAYS AS (to_tsvector('simple', content)) STORED, valid_from timestamptz NOT NULL DEFAULT now(), expires_at timestamptz, superseded_by bigint REFERENCES agent_memory(id), created_at timestamptz NOT NULL DEFAULT now() ); CREATE INDEX ON agent_memory USING hnsw (embedding vector_cosine_ops); CREATE INDEX ON agent_memory USING gin (search_vector); ``` * **pgvector** 把 embedding 放在它描述的事实旁边;索引选择与带过滤条件的扫描行为遵循 [RAG 管道](/docs/ai/rag-pipeline)和[安装 pgvector](/docs/ai/pgvector-setup)中的同一套规则。 * **`jsonb`** 承载画像的结构化部分(时区、语言、套餐档位),这些字段必须能逐字段过滤和更新,而不是每次变更都重新 embedding。 * **全文检索**覆盖 embedding 处理不好的精确名字、ID 与错误串;两路候选集合并后融合,与混合 RAG 检索完全一致。 * **实体关系**就是普通表:`entity` 加一张以 `PRIMARY KEY (src, dst, relation)` 约束的边表。在 PostgreSQL 19 上还可以把它们声明为属性图并用 `GRAPH_TABLE` 查询,见 [SQL/PGQ 图查询](/docs/core/graph-queries);更早的版本用 join 或 `WITH RECURSIVE` 查询同样的表。 检索是一条把租户与权限过滤推进数据库的 SQL,模式见 [RAG 管道](/docs/ai/rag-pipeline)。这里没有任何东西需要一台“记忆专用服务器”。 ## 候选框架对比 [#候选框架对比] 以下能力声明于 2026-08 对照官方仓库与文档核对。厂商发布的基准分数属于厂商自报数据,本站未独立验证。 | | [Mem0](https://github.com/mem0ai/mem0) | [Cognee](https://github.com/topoteretes/cognee) | | ---------------- | ---------------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------- | | 自我定位 | “记忆层” SDK、自托管 server 与托管云 | “AI 记忆平台”,从摄取的数据构建知识图谱 | | 记忆模型 | 按 user、session、agent 三个层级组织的记忆 | 文档 → 图中的实体/关系加 embedding;`remember` / `recall` / `forget` API | | 存储后端 | 可插拔向量库,[支持列表包含 PGVector](https://docs.mem0.ai/components/vectordbs/overview) | 可插拔的关系、向量与图后端;文档给出 [PostgreSQL + pgvector 配置](https://github.com/topoteretes/cognee#run-the-whole-memory-layer-on-postgres) | | 与 PostgreSQL 的关系 | PostgreSQL 是多种可选向量存储之一 | 其 README 明确标注 Postgres **图**存储目前是 demo 特性,生产图负载建议使用图原生后端或其付费产品 | | 抽取策略 | 基于 LLM 的事实抽取与更新决策在框架内完成 | 基于 LLM 的流水线(cognify)在框架内构建图谱 | 两个框架都可以把部分存储放在 PostgreSQL 之上,所以“框架 vs PostgreSQL”通常是“框架的写入策略跑在 PostgreSQL 上”与“你自己的写入策略跑在 PostgreSQL 上”之间的选择,而不是两个不同的数据库。 冯若航(vonng)在[一篇 2026 年的文章](https://blog.vonng.com/en/ai/agent-memory-framework/)中主张:记忆框架是被模型和数据库两头挤压的中间件——抽取策略会迁入模型(或一个简短的 skill 文件),存储会回到 PostgreSQL,持久的护城河在数据层。这是一位从业者的个人观点,不是可核实的事实;把它当作一个假设,对照你自己的写入路径复杂度去检验,再决定采用还是放弃框架。 ## 生产关注点 [#生产关注点] ### 记忆写入的一致性 [#记忆写入的一致性] 最危险的写入是模型抽取的那一条。要给它加约束:记忆走固定 schema 和 `CHECK` 约束;执行抽取的模型使用不能触碰业务表的受限角色([安全 SQL 护栏](/docs/ai/safe-sql));更正操作是 `INSERT` 加 `superseded_by` 链接,而不是原地改写,这样“Agent 回答时我们相信什么”始终可审计。如果某条记忆派生自一笔业务事务,要么两者在同一事务中写入,要么显式记录来源版本。 ### 遗忘与 TTL [#遗忘与-ttl] 没有过期机制的记忆会积累成噪声,再被检索放大。领域允许时给事实加 `expires_at`,定期清理过期行;用户主动发起的删除要执行硬 `DELETE`(连同 embedding 行),而不是打标记。pgvector 索引不会让已删除的行从备份里消失——PITR 保留期要与删除策略对齐。 ### 用 RLS 做多租户隔离 [#用-rls-做多租户隔离] 记忆是带 prompt 注入爆炸半径的用户数据:在某个租户被投毒写入的记忆,绝不能被另一个租户检索到。隔离必须做在数据库里,而不是检索代码里: ```sql ALTER TABLE agent_memory ENABLE ROW LEVEL SECURITY; CREATE POLICY tenant_isolation ON agent_memory USING (tenant_id = current_setting('app.tenant_id')::bigint); ``` 每个请求从已认证身份设置 `app.tenant_id`。这样即使 Agent 或 [MCP 工具](/docs/ai/mcp)发出你未曾手写的查询,隔离保证依然成立。 ### 评估记忆质量 [#评估记忆质量] 检索召回率是必要条件而非充分条件:记忆层还可能因为抽取了错误事实、保留了互相矛盾的条目、返回了过期状态而失败。保留一组带版本的对话轨迹,标注它们应产生的记忆和检索应返回的答案;把抽取准确率与检索召回率分开度量;每次改 prompt、换模型、动 schema 都重跑。评测框架与[数据库 Agent 评估](/docs/ai/agent-evals)相同;框架发布的基准是厂商自报数字,不能替代用你自己流量构建的测试集。 ## AI prompt:起草记忆 schema [#ai-prompt起草记忆-schema] ## 相关页面 [#相关页面] * [RAG 管道](/docs/ai/rag-pipeline)——读取路径复用的混合检索、索引与权限过滤 * [安装 pgvector](/docs/ai/pgvector-setup)——向量扩展的安装与验证 * [上下文契约](/docs/ai/context-contract)——模型允许看到什么,以及如何表达“上下文不足” * [安全 SQL 护栏](/docs/ai/safe-sql)——模型发起写入的角色与语句限制 * [数据库 Agent 评估](/docs/ai/agent-evals)——记忆质量的评测框架 * [SQL/PGQ 图查询](/docs/core/graph-queries)——PostgreSQL 19 上的实体关系查询 --- # 上下文契约 Canonical URL: https://pg.edu.rich/docs/ai/context-contract Last reviewed: 2026-08-02 ## 契约应回答什么 [#契约应回答什么] 每个 Agent 任务只需要相关子图,但下面的字段应稳定: ```yaml contract_version: 2026-08-02.1 database: commerce schema: app role: analytics_readonly dialect: postgresql-18 timezone: UTC currency_unit: cents tables: orders: purpose: one row per checkout primary_key: [id] columns: customer_id: { type: bigint, nullable: false, ref: customers.id } status: { type: text, allowed: [pending, paid, shipped, cancelled] } total_cents: { type: bigint, min: 0 } placed_at: { type: timestamptz, meaning: checkout completion instant } invariants: - paid orders have an immutable total sensitive: [] limits: statement_timeout_ms: 5000 max_rows: 200 writes: forbidden ``` 契约版本应与迁移版本或 schema hash 关联。工具返回契约版本,便于定位“模型基于旧 schema 生成 SQL”的问题。 ## 三层上下文 [#三层上下文] 1. **全局规则**:方言、时区、金额单位、默认 schema、权限与限制。 2. **任务子图**:相关表、键、列、注释、枚举和关键索引。 3. **动态证据**:只读样例、统计摘要、最近错误;必须标注采样时间与是否截断。 不要发送全库每个索引的完整 DDL。先按表名、列名、注释和外键关系检索出任务子图,再按需展开索引或函数定义。 ## 必须从上下文移除 [#必须从上下文移除] * 密码、连接 URI、API key 和 `pg_authid` 数据。 * 不属于当前租户或权限域的样例值。 * 完整生产行样本,尤其是个人信息与密钥材料。 * 无法说明来源和时效性的业务规则。 * 把估算统计误写成精确事实的数字。 ## 输出契约 [#输出契约] 要求模型返回结构化对象,而不是可直接执行的任意文本: ```json { "intent": "read", "sql": "SELECT id, total_cents FROM app.orders WHERE customer_id = $1 LIMIT $2", "params": [42, 50], "assumptions": ["customer_id is the authenticated customer's internal id"], "expected_columns": ["id", "total_cents"], "risk": "R0" } ``` 策略层再次验证 SQL;不要因为 JSON 格式正确就信任其语义。 ## 缺失信息的行为 [#缺失信息的行为] 契约必须允许模型回答 `insufficient_context`,列出所需的表、列或业务定义。相比猜一个看似合理的列名,这是成功行为,不是失败。 ## 把动态 Linux 事实交给只读工具 [#把动态-linux-事实交给只读工具] 安装、升级和排错问题还需要发行版动态事实。可以把 [PkgSeek 只读 MCP](https://pkgseek.com/mcp) 配置为 Linux 软件包证据层,让 Agent 查询准确包名、文件 provider、仓库、release、版本历史和厂商安全状态;本站文档继续提供 PostgreSQL 的选择、升级和验证规则。 ```yaml linux_evidence: provider: pkgseek distro: ubuntu release: noble architecture: amd64 package_source: pgdg observed_at: required missing_coordinate: insufficient_context state_changes: require_confirmation ``` 工具没有返回准确坐标时,模型必须报告“当前索引未收录”,不能改写成“软件包不存在”。查询是只读的;`sudo apt install`、仓库修改、服务重启和大版本升级仍要单独确认并在执行后验证。 --- # AI / Agent 文档入口 Canonical URL: https://pg.edu.rich/docs/ai Last reviewed: 2026-08-02 模型不会因为“会写 SQL”就理解你的数据库。可靠系统需要把数据库上下文变成契约,把执行能力收窄成工具,并把正确性变成可重复评估。 ## 推荐架构 [#推荐架构] ```text 用户意图 → 任务分类(读 / 写 / DDL / 运维) → 检索 schema 契约与相关文档 → 模型生成结构化 tool call → 策略层校验 AST、权限、成本与参数 → 受限数据库角色执行 → 返回行数、SQLSTATE、耗时与截断状态 → 记录审计事件 ``` 数据库凭据不进入模型上下文。模型不直接选择连接目标。工具层必须绑定环境、database、schema 和 role。 ## 风险分级 [#风险分级] | 等级 | 例子 | 默认策略 | | -- | ---------------------------- | ---------------------- | | R0 | 列表、描述 schema、带上限的只读查询 | 自动执行,短超时 | | R1 | 读取敏感列、较大聚合 | 权限过滤、审计、成本上限 | | R2 | `INSERT` / 有主键条件的单行 `UPDATE` | dry-run + 业务 API 或明确确认 | | R3 | 批量写、DDL、权限、复制、备份恢复 | 不向通用 Agent 暴露;专家流程 | “不要删数据”只是一条行为建议。真正的边界来自数据库角色、网络隔离、只读事务、SQL 解析与工具白名单。 ## 最小成功标准 [#最小成功标准] 一个可上线的数据库 Agent 至少应做到:所有值参数化;默认只读;限制语句时间和结果行数;拒绝多语句;不向模型返回 secret;记录查询指纹和审计信息;针对 `40001`、`40P01`、`57014` 等状态码有确定行为。 --- # Postgres MCP Server Canonical URL: https://pg.edu.rich/docs/ai/mcp Last reviewed: 2026-08-06 MCP(Model Context Protocol)为 agent 调用工具和读取数据提供了统一协议。Postgres MCP server 位于 agent 与数据库之间,把 schema 元数据、只读查询和 `EXPLAIN` 输出暴露为工具,让 agent 对照真实的系统目录工作,而不是猜测列名。这会改变 AI 生成 SQL 的失败模式——模型能自查之后,大多数“幻觉列名”错误就消失了。 Server 本身只是传输层,真正的安全边界是它连接时使用的数据库角色,下文以及[安全 SQL 护栏](/docs/ai/safe-sql)会展开说明。 ## 选择 server [#选择-server] 以下工具能力声明均以 2026-08 时各项目仓库为准核对;MCP 生态变化很快,采用前请重新确认。 | 项目 | 维护方 | 说明 | | ------------------------------------------------------------------------------ | ----------- | ------------------------------------------------------------------------------------------------------------------------------------------------- | | [Postgres MCP Pro](https://github.com/crystaldba/postgres-mcp)(`postgres-mcp`) | Crystal DBA | Schema 浏览、`EXPLAIN` 分析、索引调优与健康检查。提供 `--access-mode=restricted` 参数,将执行限制为只读 SQL。基于 Python,通过 `uvx`、`pipx` 或 `crystaldba/postgres-mcp` Docker 镜像运行。 | | [Neon MCP](https://github.com/neondatabase/mcp-server-neon) | Neon | 增加了项目级资源:创建分支、在分支上执行迁移、获取连接串。 | | [Supabase MCP](https://supabase.com/docs/guides/getting-started/mcp) | Supabase | 托管 server,覆盖数据库与项目管理(分支、日志、advisor)。 | `@modelcontextprotocol/server-postgres`——Anthropic 最初的参考实现——已在 npm 上标记废弃,源码移入只读的 [servers-archived](https://github.com/modelcontextprotocol/servers-archived) 仓库,不再有维护和安全修复。新部署不要使用它(2026-08 核对)。 无论选哪个 server,都把厂商的功能列表当作临时状态看待:只有确实用到调优工具时才启用 `pg_stat_statements` 和 `hypopg`,并在配置中固定 server 版本,让升级成为一个主动决定。 ## 先建只读角色 [#先建只读角色] 配置任何客户端之前,先创建一个只能读的角色: ```sql CREATE ROLE readonly LOGIN PASSWORD 'secret'; GRANT CONNECT ON DATABASE myapp TO readonly; GRANT USAGE ON SCHEMA public TO readonly; GRANT SELECT ON ALL TABLES IN SCHEMA public TO readonly; ALTER DEFAULT PRIVILEGES IN SCHEMA public GRANT SELECT ON TABLES TO readonly; ALTER ROLE readonly SET default_transaction_read_only = on; ALTER ROLE readonly SET statement_timeout = '5s'; ``` `ALTER DEFAULT PRIVILEGES` 这行很关键:没有它,之后新建的表对该角色不可见,agent 看到的 schema 会悄悄过期。超时也应在角色层设置,避免 agent 发出的失控查询长期占用资源——完整配置(lock timeout、idle-in-transaction timeout、行数上限)见[安全 SQL 护栏](/docs/ai/safe-sql)。 不要给 MCP server 超级用户或表属主凭据——本地开发也不行,因为开发配置里的习惯会漏进生产。生产环境的写权限绝不应暴露给 agent:让它连接只读副本、Aurora reader 端点或数据库分支。MCP server 的受限模式只是第二层防线,不能替代角色层的权限控制。 ## 在 Claude Code 中接入 [#在-claude-code-中接入] 项目级 `.mcp.json`(提交到仓库,团队共享同一份 server 配置),或用 `claude mcp add` 添加用户级条目: ```json { "mcpServers": { "postgres": { "command": "uvx", "args": [ "postgres-mcp", "--access-mode=restricted", "postgres://readonly:secret@localhost:5432/myapp" ] } } } ``` `--access-mode=restricted` 保证即使 agent 发起写请求,server 也只执行只读 SQL(2026-08 在[项目 README](https://github.com/crystaldba/postgres-mcp) 中核对)。连接串优先从环境变量或 secret manager 注入,不要把密码提交进仓库。 ## 在 Cursor 中接入 [#在-cursor-中接入] `.cursor/mcp.json`: ```json { "mcpServers": { "postgres": { "command": "uvx", "args": ["postgres-mcp", "--access-mode=restricted"], "env": { "DATABASE_URI": "postgres://readonly:secret@localhost:5432/myapp" } } } } ``` ## 每个 PR 一个数据库分支 [#每个-pr-一个数据库分支] Neon 和 Supabase 都支持基于 copy-on-write 的即时分支。结合 MCP,每个 PR 可以拥有一个独立数据库,供 agent 执行迁移、查询,用完即销毁: 1. 从 main fork 一个分支(毫秒级,不复制数据)。 2. 对着真实形态的数据跑迁移和测试,让 agent 在真实数据量上读 `EXPLAIN`。 3. 在 PR 中 review 变更,合并后删除分支。 ```bash neon branches create --name pr-123 --parent main export DATABASE_URI=$(neon connection-string --branch pr-123) # 用 DATABASE_URI 启动 MCP server ``` 这是 MCP 价值最大的场景:agent 对照真实 catalog 和真实数据分布验证迁移,全程不接触生产写入端。 ## 通过契约约束暴露面 [#通过契约约束暴露面] 数据库 MCP server 回答的是“schema 长什么样”“这条查询会做什么”——它不应该变成通用的数据外泄通道。用[上下文契约](/docs/ai/context-contract)把面向任务的上下文控制得小而明确:agent 只拿到与任务相关的表和列,MCP server 负责验证,而不是漫游无关的 schema。 ## 配合 MCP 使用的 prompt [#配合-mcp-使用的-prompt] ## 下一步 [#下一步] * 约束 agent 能执行的 SQL → [安全 SQL 护栏](/docs/ai/safe-sql) * 压缩 agent 需要的上下文 → [上下文契约](/docs/ai/context-contract) * 从系统目录生成这份上下文 → [Schema 检索与文档生成](/docs/ai/schema-retrieval) --- # PostgreSQL 安装 pgvector Canonical URL: https://pg.edu.rich/docs/ai/pgvector-setup Last reviewed: 2026-08-02 pgvector 是独立扩展,不是 PostgreSQL 核心内置类型。安装软件包或使用包含扩展的镜像后,仍需在每个目标 database 执行 `CREATE EXTENSION vector`。 ## Docker 最小环境 [#docker-最小环境] pgvector 上游提供基于官方 Postgres 镜像的版本化镜像。以下示例固定 PostgreSQL 18 和 pgvector 0.8.2: ```bash docker volume create pgvector18-data docker run --name pgvector18 \ --env POSTGRES_PASSWORD=local-only-change-me \ --publish 5432:5432 \ --volume pgvector18-data:/var/lib/postgresql/data \ --detach pgvector/pgvector:0.8.2-pg18-trixie docker exec -it pgvector18 \ psql -U postgres -d postgres -c "CREATE EXTENSION vector;" ``` 密码仅用于隔离的本机演示,不可进入生产配置。已有同名容器或 5432 端口占用时应选择明确的新名称/端口,不要删除未知实例。 ## Ubuntu / Debian 包 [#ubuntu--debian-包] ### 发行版仓库中的 pgvector PkgSeek 软件包快照;核对时间:2026-08-02 15:47:58 UTC。 | 发行版 | Release | 完整版本 | 仓库 | 关联公告 | |---|---|---|---|---:| | AlmaLinux | 10 | 0.6.2-6.el10_0 | official / AppStream | — | | AlmaLinux | 9 | 0.8.1-1.module_el9.8.0+234+5456f35d | official / AppStream | — | | Arch Linux | rolling | 0.8.6-1 | official / extra | — | | CentOS Stream | 10 | 0.6.2-8.el10 | official / AppStream | — | | CentOS Stream | 9 | 0.8.1-1.module_el9+1300+1c4aa8df | official / AppStream | — | | Fedora | 42 | 0.6.2-4.fc42 | official / everything | — | | Fedora | 43 | 0.8.0-1.fc43 | official / everything | — | | Fedora | 44 | 0.8.0-2.fc44 | official / everything | — | | Oracle Linux | 10 | 0.6.2-6.el10_0 | official / appstream | — | | Oracle Linux | 9 | 0.6.2-2.module+el9.8.0+90925+e22a792e | official / appstream | — | | Red Hat Enterprise Linux | 10.2 | 0.6.2-6.el10_0 | official / AppStream | — | | Red Hat Enterprise Linux | 9.8 | 0.6.2-2.module+el9.8.0+24096+5a959ed6 | official / AppStream | — | | Rocky Linux | 10 | 0.6.2-6.el10_0 | official / AppStream | — | | Rocky Linux | 9 | 0.6.2-2.module+el9.8.0+40212+d6f50005 | official / AppStream | — | 来源:[PkgSeek 软件包查询](https://pkgseek.com/packages/pgvector)。发行版 revision 和回溯补丁属于完整版本身份。 同一个扩展在不同发行版中可能使用 `pgvector`、`postgresql18-pgvector` 或与 server major 绑定的其他名称。上表只显示准确命名为 `pgvector` 的坐标;不能据此推断所有发行版的版本化包是否存在。 ### PGDG:postgresql-18-pgvector PkgSeek 软件包快照;核对时间:2026-08-02 15:47:58 UTC。 > 没有准确索引坐标。这表示当前快照未收录,不等于软件包不存在。 来源:[PkgSeek 软件包查询](https://pkgseek.com/search?q=postgresql-18-pgvector)。发行版 revision 和回溯补丁属于完整版本身份。 若上方没有准确坐标,仍需以目标 PGDG 仓库元数据复核,不能根据通用 `pgvector` 包推断 Ubuntu / Debian 的版本化包名。 配置 PostgreSQL 官方 Apt 仓库后,扩展包与 server major 绑定: ```bash sudo apt install postgresql-18-pgvector sudo -u postgres psql -d app -c "CREATE EXTENSION vector;" ``` 安装到操作系统不代表所有 database 已启用。验证实际版本: ```sql SELECT extname, extversion FROM pg_extension WHERE extname = 'vector'; ``` 升级扩展前阅读 release notes,并在目标 database 执行经过验证的: ```sql ALTER EXTENSION vector UPDATE; ``` ## 最小查询验证 [#最小查询验证] ```sql CREATE TABLE vector_demo ( id bigint GENERATED ALWAYS AS IDENTITY PRIMARY KEY, embedding vector(3) NOT NULL ); INSERT INTO vector_demo (embedding) VALUES ('[1,2,3]'), ('[4,5,6]'), ('[1,1,1]'); SELECT id, embedding <-> '[1,2,2]'::vector AS l2_distance FROM vector_demo ORDER BY embedding <-> '[1,2,2]'::vector LIMIT 2; ``` 没有近似索引时这是 exact search。数据量和真实过滤条件需要时,再依据[pgvector 生产最佳实践](/docs/ai/vector-production)评估 HNSW 或 IVFFlat。 ## 云 PostgreSQL [#云-postgresql] 云服务通常限制操作系统访问和扩展清单。上线前确认: * engine major 与 `vector` 扩展的具体版本; * 谁有权限执行 `CREATE EXTENSION` / `ALTER EXTENSION`; * HNSW、IVFFlat 和 iterative scans 是否由该版本支持; * 扩展升级是否跟随引擎升级、需要维护窗口或手工执行; * index build 的磁盘、内存、WAL 和副本延迟限制。 模型、维度、归一化或距离语义变化时,使用新列或新表重建并评测。数据库允许写入同维度向量,不代表它们在语义上可比较。 安装方式和当前发行版以 [pgvector 官方仓库](https://github.com/pgvector/pgvector#installation)为准。 --- # PostgreSQL RAG 管道 Canonical URL: https://pg.edu.rich/docs/ai/rag-pipeline Last reviewed: 2026-08-02 ## 数据模型 [#数据模型] ```sql CREATE EXTENSION IF NOT EXISTS vector; CREATE TABLE documents ( id bigint GENERATED ALWAYS AS IDENTITY PRIMARY KEY, tenant_id bigint NOT NULL, source_uri text NOT NULL, source_version text NOT NULL, title text NOT NULL, access_scope text[] NOT NULL DEFAULT '{}', created_at timestamptz NOT NULL DEFAULT now(), UNIQUE (tenant_id, source_uri, source_version) ); CREATE TABLE document_chunks ( id bigint GENERATED ALWAYS AS IDENTITY PRIMARY KEY, document_id bigint NOT NULL REFERENCES documents(id) ON DELETE CASCADE, ordinal integer NOT NULL CHECK (ordinal >= 0), content text NOT NULL, token_count integer NOT NULL CHECK (token_count > 0), embedding vector(1536) NOT NULL, embedding_model text NOT NULL, search_vector tsvector GENERATED ALWAYS AS (to_tsvector('simple', content)) STORED, UNIQUE (document_id, ordinal) ); ``` ## 摄取必须可重放 [#摄取必须可重放] 保存 source version、切块算法版本、embedding 模型和维度。对相同输入生成确定的文档/切块键,支持幂等 upsert。模型变更时创建新 embedding 列或新表并双写重建,不要把两种向量混在一个索引里。 ## 检索顺序 [#检索顺序] 1. 在 SQL 中过滤 `tenant_id`、权限、文档状态和时间范围。 2. 用全文和向量分别产生有界候选集。 3. 使用 rank fusion 或应用侧重排合并。 4. 取少量相邻切块补足上下文。 5. 返回 `source_uri`、版本、chunk id 与原文片段供引用。 示意性的向量候选查询: ```sql SELECT c.id, c.document_id, c.ordinal, c.content, c.embedding <=> $1::vector AS distance FROM document_chunks AS c JOIN documents AS d ON d.id = c.document_id WHERE d.tenant_id = $2 AND d.access_scope && $3::text[] ORDER BY c.embedding <=> $1::vector LIMIT 40; ``` 索引类型与参数取决于数据量、召回率、延迟和写入模式。先建立 exact-search 基线,再评估 HNSW 或 IVFFlat;不要只用演示数据判断。 ### 近似索引与过滤 [#近似索引与过滤] HNSW/IVFFlat 的 tenant、ACL 等条件通常在索引产生候选后应用,因此可能返回少于 `LIMIT` 的结果。这不是放宽权限过滤的理由。pgvector 0.8.0+ 可用 iterative scans 扩大候选扫描;高频 tenant 还可评估分区、部分索引或独立表。每种方案都要按真实 tenant/ACL 分桶测量 recall\@k。 具体索引 DDL、参数和评测方法见[向量检索生产化](/docs/ai/vector-production)与 [pgvector 官方文档](https://github.com/pgvector/pgvector#filtering)。 ## 安全与引用 [#安全与引用] 权限条件必须存在于 SQL/RLS 内,由数据库在结果离开前执行;不能先向应用返回全局候选再过滤,否则无权内容可能通过日志、缓存或模型上下文泄露。最终回答携带可验证引用;找不到足够证据时明确说“不足以回答”。 相近文本可能过期、互相矛盾或属于错误租户。RAG 需要版本、权限、来源优先级和答案评估,而不只是最近邻。 --- # 安全 SQL 护栏 Canonical URL: https://pg.edu.rich/docs/ai/safe-sql Last reviewed: 2026-08-02 ## 数据库角色是第一道边界 [#数据库角色是第一道边界] ```sql CREATE ROLE agent_reader LOGIN; GRANT CONNECT ON DATABASE commerce TO agent_reader; GRANT USAGE ON SCHEMA app TO agent_reader; GRANT SELECT ON ALL TABLES IN SCHEMA app TO agent_reader; ALTER DEFAULT PRIVILEGES FOR ROLE app_owner IN SCHEMA app GRANT SELECT ON TABLES TO agent_reader; ALTER ROLE agent_reader SET default_transaction_read_only = on; ALTER ROLE agent_reader SET statement_timeout = '5s'; ALTER ROLE agent_reader SET lock_timeout = '1s'; ALTER ROLE agent_reader SET idle_in_transaction_session_timeout = '10s'; ``` 凭据由 secret manager 或云身份集成配置,不要保存在迁移文件、prompt 或工具响应中。确认该角色不能 `SET ROLE` 到更高权限角色。 ## 执行前策略 [#执行前策略] 对模型生成的 SQL 做 parser/AST 级校验,不用正则代替解析器。默认规则: * 只允许单条 `SELECT`。 * 拒绝 `COPY ... PROGRAM`、大对象、外部数据包装器和危险函数。 * 拒绝多个语句和注释绕过。 * 限制可访问 schema、表和列。 * 强制参数绑定;标识符只能从白名单选择。 * 对非聚合结果施加 `LIMIT`,同时在驱动层设置最大返回字节数。 * 执行 `EXPLAIN (FORMAT JSON)` 做成本预检时,不能把成本估算当作时间保证。 ## 每次读取使用只读事务 [#每次读取使用只读事务] ```sql BEGIN READ ONLY; SET LOCAL statement_timeout = '5s'; SET LOCAL lock_timeout = '1s'; SET LOCAL search_path = app, pg_catalog; SELECT id, status, total_cents FROM orders WHERE customer_id = $1 ORDER BY placed_at DESC LIMIT 100; COMMIT; ``` 只读事务仍可能运行昂贵查询并泄露可读取的数据,所以权限、成本和结果限制缺一不可。 ## 写操作不要开放任意 SQL [#写操作不要开放任意-sql] 优先向 Agent 暴露领域工具: ```json { "tool": "cancel_order", "arguments": { "order_id": 8842, "expected_status": "pending", "reason": "duplicate order", "idempotency_key": "case-2026-184" } } ``` 应用服务验证身份与状态转换,在事务中执行参数化 SQL,并返回明确结果。批量写、DDL、`GRANT`、备份恢复和复制配置不应暴露给通用 Agent。 ## 审计字段 [#审计字段] 至少记录:请求者/租户、工具名、模型与提示版本、契约版本、数据库目标、SQL 指纹(参数脱敏)、风险级别、审批者、行数、耗时、SQLSTATE、是否截断。不要把原始敏感结果复制进普通日志。 先执行写入再发现“影响太多行”可能已经触发触发器或产生外部事件。应在受控事务里预览目标集,或通过领域 API 将可修改集合限制在查询本身。 --- # Schema 检索与文档生成 Canonical URL: https://pg.edu.rich/docs/ai/schema-retrieval Last reviewed: 2026-08-02 ## 不要让模型自己探库 [#不要让模型自己探库] 生产 Agent 不应拥有无边界的系统目录探索权。由受信任的构建任务提取 schema,脱敏、版本化后写入检索库;运行时只返回与任务相关的子图。 ## 表和列 [#表和列] 使用 `information_schema` 获取可移植的基础信息: ```sql SELECT c.table_schema, c.table_name, c.ordinal_position, c.column_name, c.data_type, c.udt_name, c.is_nullable, c.column_default FROM information_schema.columns AS c WHERE c.table_schema = ANY($1::text[]) ORDER BY c.table_schema, c.table_name, c.ordinal_position; ``` 列注释来自 PostgreSQL 目录: ```sql SELECT n.nspname AS schema_name, cls.relname AS table_name, a.attname AS column_name, col_description(cls.oid, a.attnum) AS comment FROM pg_catalog.pg_attribute AS a JOIN pg_catalog.pg_class AS cls ON cls.oid = a.attrelid JOIN pg_catalog.pg_namespace AS n ON n.oid = cls.relnamespace WHERE n.nspname = ANY($1::text[]) AND cls.relkind IN ('r', 'p') AND a.attnum > 0 AND NOT a.attisdropped; ``` ## 外键边形成任务子图 [#外键边形成任务子图] ```sql SELECT src_ns.nspname AS table_schema, src.relname AS table_name, src_col.attname AS column_name, dst_ns.nspname AS foreign_table_schema, dst.relname AS foreign_table_name, dst_col.attname AS foreign_column_name FROM pg_catalog.pg_constraint AS con JOIN pg_catalog.pg_class AS src ON src.oid = con.conrelid JOIN pg_catalog.pg_namespace AS src_ns ON src_ns.oid = src.relnamespace JOIN pg_catalog.pg_class AS dst ON dst.oid = con.confrelid JOIN pg_catalog.pg_namespace AS dst_ns ON dst_ns.oid = dst.relnamespace CROSS JOIN LATERAL unnest(con.conkey, con.confkey) AS key_columns(src_attnum, dst_attnum) JOIN pg_catalog.pg_attribute AS src_col ON src_col.attrelid = src.oid AND src_col.attnum = key_columns.src_attnum JOIN pg_catalog.pg_attribute AS dst_col ON dst_col.attrelid = dst.oid AND dst_col.attnum = key_columns.dst_attnum WHERE con.contype = 'f' AND src_ns.nspname = ANY($1::text[]) ORDER BY con.oid, src_col.attnum; ``` `conkey` 与 `confkey` 按位置对应;并行 `unnest` 能正确保留复合外键的列映射。只按 `constraint_name` 连接 `information_schema` 视图,可能在复合外键上产生列的笛卡尔积。 ## 文档构建流程 [#文档构建流程] ```text 迁移合并 → 临时数据库应用全部迁移 → 目录提取 → 规范化排序并移除环境值 → 生成 JSON + Markdown 摘要 → 计算 hash / 绑定迁移版本 → 评审 schema diff → 发布到检索索引 ``` 每张表的摘要只保留:用途、主键、外键、列类型与 nullable、约束、业务注释、敏感级别,以及最关键的查询索引。函数体、视图定义和策略仅在任务需要时展开。 ## 防止陈旧 [#防止陈旧] Agent 工具每次返回 `contract_version`。若运行时数据库的迁移版本与检索文档不一致,拒绝高风险请求并触发重建。不要静默使用旧契约。 --- # Text-to-SQL 生产模式 Canonical URL: https://pg.edu.rich/docs/ai/text-to-sql Last reviewed: 2026-08-02 Text-to-SQL 的目标不是“尽量生成一条能跑的 SQL”,而是只在证据、权限和成本边界清楚时执行正确查询;其余请求应澄清或拒绝。 ## 推荐执行链 [#推荐执行链] ```text 自然语言问题 → 解析业务实体、指标、时间范围和期望粒度 → 检索版本化 schema / metric 契约与少量已验证示例 → 生成结构化查询计划和参数,不直接执行自由文本 → SQL AST 校验、对象/函数白名单、权限与成本检查 → 受限角色 + 只读事务 + 超时 + 结果上限 → 返回结果、口径、SQL 指纹、截断状态与可解释错误 ``` 模型输入至少包括:schema 版本、表/列语义、主外键、枚举值、时区、金额单位、软删除规则、租户边界、已批准指标定义和允许查询的对象。不要把整个数据库 DDL、示例客户数据或凭据无差别塞进上下文。 ## 结构化工具优先 [#结构化工具优先] 对常见分析请求,优先让模型生成领域参数: ```json { "metric": "paid_order_revenue", "time_range": { "start": "2026-07-01", "end": "2026-08-01" }, "group_by": ["day"], "filters": [{ "field": "region", "op": "eq", "value": "east" }], "limit": 100 } ``` 由服务端把 metric、字段和操作符映射到已审查 SQL。只有长尾探索才进入自由 SQL 通道;该通道仍必须解析 AST,不能用正则判断“以 SELECT 开头”。CTE、可写 CTE、函数、`COPY`、多语句和注释混淆都会绕过天真的字符串检查。 ## 数据库执行封装 [#数据库执行封装] ```sql BEGIN READ ONLY; SET LOCAL statement_timeout = '3s'; SET LOCAL lock_timeout = '500ms'; SET LOCAL idle_in_transaction_session_timeout = '5s'; -- 由策略层批准的单条参数化 SELECT;服务端强制结果行/字节上限 SELECT date_trunc('day', paid_at) AS day, sum(total_cents) AS revenue_cents FROM analytics.paid_orders WHERE tenant_id = $1 AND paid_at >= $2 AND paid_at < $3 GROUP BY 1 ORDER BY 1 LIMIT 100; COMMIT; ``` `READ ONLY` 是纵深防御,不是完整沙箱:仍应只允许受信函数和对象,以低权限专用角色执行,并由服务端绑定 tenant、环境和参数。模型不能传入连接字符串、role 或 `search_path`。 ## 执行前检查 [#执行前检查] 1. 只允许一个语句和允许的 AST 节点;拒绝 DDL/DML、`COPY`、任意函数调用和系统管理对象。 2. 所有字面值转为绑定参数;标识符只能来自 schema 契约白名单。 3. 对高成本候选运行 `EXPLAIN (FORMAT JSON)`,检查访问对象、估算行数和总成本;估算只是信号,不能保证运行时间。 4. 强制时间范围、行数/字节上限和最大 join 数;大导出走独立异步产品能力。 5. 敏感列在策略层拒绝或映射为已脱敏视图,不依赖模型“记得不要选”。 6. tenant 条件由数据库 RLS 或服务端模板注入,不能由用户问题决定。 ## 正确性与拒答 [#正确性与拒答] 查询能执行不等于答案正确。评估集应覆盖空结果、重复 join、时区边界、NULL、退款/取消、迟到数据、权限隔离和口径歧义。对每个问题同时断言:允许/拒绝决策、结果集、访问对象、最大成本与解释文本。 以下情况应澄清而不是猜测:指标有多个业务定义;日期缺少时区或年份;实体名称匹配多个 ID;请求要求不存在的历史快照;schema 版本与部署不一致。 语法错误最多基于结构化错误做有限重生成。`57014`(超时/取消)应缩小请求或转异步;`40001` 与 `40P01` 只在整个事务可安全重放时重试。始终保留原请求、schema 版本、查询指纹与最终决策。 进一步阅读:[PostgreSQL 18 事务 `READ ONLY` 语义](https://www.postgresql.org/docs/18/sql-set-transaction.html)和[错误码附录](https://www.postgresql.org/docs/18/errcodes-appendix.html)。 --- # pgvector 生产最佳实践 Canonical URL: https://pg.edu.rich/docs/ai/vector-production Last reviewed: 2026-08-02 [pgvector](https://github.com/pgvector/pgvector) 默认执行精确最近邻搜索;添加 HNSW 或 IVFFlat 索引后才进入近似搜索。索引选择是召回、延迟、内存、构建时间和写入成本之间的工程决策。 ## 先固定距离语义 [#先固定距离语义] | 业务语义 | 操作符 | 索引 operator class | | --------- | ----- | ------------------- | | L2 / 欧氏距离 | `<->` | `vector_l2_ops` | | 内积(返回负内积) | `<#>` | `vector_ip_ops` | | 余弦距离 | `<=>` | `vector_cosine_ops` | embedding 生成、索引和查询必须使用同一种距离语义。余弦相似度是 `1 - cosine distance`。保存 `embedding_model`、维度、归一化方式和生成版本;不可比较的向量不要混入一个列/索引。 ## 建立 exact 基线 [#建立-exact-基线] 从真实查询分布抽样并保存 exact top-k。评估近似索引时,比较 `recall@k`、p50/p95/p99 延迟、候选不足率和资源消耗,而不是只看单次速度。 ```sql BEGIN; SET LOCAL enable_indexscan = off; SELECT c.id FROM document_chunks AS c JOIN documents AS d ON d.id = c.document_id WHERE d.tenant_id = $1 ORDER BY c.embedding <=> $2::vector LIMIT 20; ROLLBACK; ``` 禁用 index scan 只用于基线/诊断,不是生产查询设置。测试集要覆盖热门/冷门 tenant、常见 ACL、时间过滤、新写入、删除和 embedding 分布漂移。 ## HNSW 与 IVFFlat [#hnsw-与-ivfflat] ```sql CREATE INDEX CONCURRENTLY document_chunks_embedding_hnsw ON document_chunks USING hnsw (embedding vector_cosine_ops); ``` * **HNSW**:通常查询性能与 speed/recall tradeoff 更好,不需要训练数据;构建慢、使用更多内存,索引维护成本也更高。 * **IVFFlat**:构建更快、内存更少;需要已有代表性数据来形成 lists,且 speed/recall tradeoff 通常低于 HNSW。不要在空表上创建后就忘记重建。 * 生产已有表优先 `CREATE INDEX CONCURRENTLY`,并在副本上观察 WAL、磁盘、构建时间和复制延迟。 没有适用于所有数据集的 `m`、`ef_construction`、`ef_search`、`lists` 或 `probes` 常数。先用默认值和 exact 基线,再用真实过滤条件调参。 ## 过滤会改变召回 [#过滤会改变召回] 使用近似索引时,过滤条件通常在索引扫描产生候选后应用。默认 `hnsw.ef_search = 40` 时,如果只有 10% 候选满足 tenant/ACL 条件,结果可能少于 `LIMIT`,即使数据库里存在更多匹配行。 pgvector 0.8.0+ 支持 iterative scans,索引候选不足时继续扫描: ```sql BEGIN; SET LOCAL hnsw.iterative_scan = strict_order; SET LOCAL hnsw.ef_search = 200; SELECT c.id, c.content, c.embedding <=> $1::vector AS distance FROM document_chunks AS c JOIN documents AS d ON d.id = c.document_id WHERE d.tenant_id = $2 AND d.access_scope && $3::text[] ORDER BY c.embedding <=> $1::vector LIMIT 20; COMMIT; ``` 先确认云服务提供的 pgvector 版本。若 tenant 值较少且不均衡,可考虑 list partitioning;若值很多,可把高频 tenant 部分索引、独立表或物理隔离。选择必须用过滤后的 recall 与运维成本验证。 tenant/ACL 条件必须留在 SQL/RLS 内,不能先取全局结果再由应用过滤。即使数据库正确阻止越权行,共享 ANN 图仍可能因后过滤返回不足;安全性和召回率是两个分别验证的目标。 ## 上线门槛 [#上线门槛] * 每个 embedding 版本有可重放摄取、exact gold set 和回滚路径; * 在线记录模型版本、过滤桶、候选数、结果数、距离分布、延迟和截断; * 定期抽样 exact 查询计算 recall\@k,并按 tenant/ACL 分桶; * embedding 重建双写到新列/表,建新索引、评估后原子切流; * 原文和 metadata 是事实来源,向量可以从版本化输入重建; * 混合检索分别产生全文/向量候选,再以可版本化方法融合或重排。 官方参数与限制以 [pgvector README 的索引、过滤和监控章节](https://github.com/pgvector/pgvector#hnsw)为准。 --- # Google Cloud AlloyDB for PostgreSQL Canonical URL: https://pg.edu.rich/docs/cloud/alloydb Last reviewed: 2026-08-06 [AlloyDB for PostgreSQL](https://docs.cloud.google.com/alloydb/docs/overview) 是 Google Cloud 的 PostgreSQL 兼容数据库服务:计算与存储解耦、跨可用区高可用、面向分析查询的可选列式引擎,以及库内机器学习集成。它讲 PostgreSQL 协议、运行 PostgreSQL 兼容的 SQL,但不是社区二进制——存储、复制、发布节奏和部分查询行为是 Google 自己的实现。与其他增强型引擎的定位对比见[云服务版图](/docs/cloud/service-map)。 ## AlloyDB 改变了什么 [#alloydb-改变了什么] * **计算与存储分离。** 集群内的实例共享带分层缓存的解耦存储,而不是每个节点持有本地卷。 * **可选列式引擎。** 热点数据可以以列式格式驻留,与行存并行服务分析型扫描——Google 公布的事务和分析加速数字是厂商基准,把它们当作待验证的假设,用你自己的数据测试。 * **库内 ML。** `google_ml_integration` 扩展可以从 SQL 直接调用托管模型端点,覆盖 embedding 生成和预测,省掉外部胶水服务。 * **ScaNN 向量索引。** 在 pgvector 的 HNSW 和 IVFFlat 之外的另一种近似最近邻索引。 ## 用 google\_ml.embedding() 在库内生成 embedding [#用-google_mlembedding-在库内生成-embedding] 装好扩展并注册好模型端点后,embedding 生成就是一次 SQL 函数调用。以下按[官方文档](https://docs.cloud.google.com/alloydb/docs/ai/work-with-embeddings)核对于 2026-08-06: ```sql CREATE EXTENSION IF NOT EXISTS google_ml_integration; CREATE EXTENSION IF NOT EXISTS vector; SELECT google_ml.embedding( model_id => 'gemini-embedding-001', content => 'AlloyDB keeps embedding generation inside the database' ) AS embedding; INSERT INTO articles (body, embedding) VALUES ( 'Some article text', google_ml.embedding(model_id => 'gemini-embedding-001', content => 'Some article text')::vector ); ``` 该函数返回 `real[]`,存入 `vector` 列或参与向量运算前需要像上面一样加 `::vector` 显式转换。 两条边界要记清:这个调用会离开数据库访问模型端点,因此区域、数据驻留、IAM 权限和模型生命周期都成了数据库侧的问题;模型 ID 会变化——以文档中的当前 ID 为准,不要照抄任何示例(包括本页)。 ## ScaNN 向量索引 [#scann-向量索引] AlloyDB 通过 `alloydb_scann` 扩展提供 [ScaNN 索引](https://docs.cloud.google.com/alloydb/docs/ai/create-scann-index): ```sql CREATE EXTENSION IF NOT EXISTS alloydb_scann CASCADE; CREATE INDEX articles_embedding_scann ON articles USING scann (embedding cosine) WITH (mode = 'MANUAL', num_leaves = 100); ``` ScaNN 是基于树的量化索引。按 Google 文档,它比 HNSW 构建更快、内存占用更低,QPS 和召回率取决于 `num_leaves` 等调优参数。但这两条性质都不会在真实数据上自动成立:在你实际的 tenant 和 ACL 过滤之后测量召回率,并在同一数据集上与 pgvector HNSW 对比,再决定标准化用哪一个。 ## 需要接受的差异 [#需要接受的差异] * **不是社区二进制等价物。** 扩展可用性、参数面以及部分系统目录或等待事件行为与社区 PostgreSQL、与 Cloud SQL 都有差异。带着兼容性清单迁移,不要靠假设。 * **按实例小时计费。** 没有 scale-to-zero,空闲集群照常计费。用量形态波动大的负载可能更适合 serverless 平台。 * **仅限 Google Cloud。** AlloyDB Omni 可用于其他环境的自管部署,但它是单独授权、单独运维的产品。 * **厂商性能数字是营销输入。** 公布的加速比基于 Google 基准条件下与 Cloud SQL 的对比;真实数字由你的 schema、并发和数据形态决定。 ## 生产前验证 [#生产前验证] 1. 把所需扩展和参数与 AlloyDB 支持清单比对,包括 pgvector 和 `google_ml_integration` 的确切版本。 2. 演练迁移路径——Database Migration Service 或 `pg_dump`/恢复——包括回滚方案。 3. 用有代表性的 OLTP 和分析查询做基准;只有实测收益成立时才启用列式引擎。 4. 在生产数据量级、真实过滤条件下,分别测量 ScaNN 和 HNSW 的召回率与延迟。 5. 确认 embedding 模型的区域、IAM 范围、配额和生命周期策略;记录模型版本退役时已存 embedding 的处置方案。 6. 用原生工具导出一次,并在独立 PostgreSQL 环境中恢复,作为退出路径验证。 能力、模型 ID 和扩展行为按 Google Cloud 官方文档核对于 2026-08-06。AlloyDB 迭代很快;采购时请重新核实模型可用性和扩展版本。 连接预算、备份演练和监控基线等跨厂商内容见[生产环境技术栈](/docs/operations/production-stack)。开发规模的评估可以从[免费 PostgreSQL 选型](/docs/cloud/free-postgresql)中的方案开始。 --- # AWS RDS for PostgreSQL 与 Aurora Canonical URL: https://pg.edu.rich/docs/cloud/aws-rds-aurora Last reviewed: 2026-08-06 AWS 为 PostgreSQL 工作负载提供两条托管路径,但二者不是同一个产品。[Amazon RDS for PostgreSQL](https://docs.aws.amazon.com/AmazonRDS/latest/UserGuide/CHAP_PostgreSQL.html) 把社区内核作为托管实例运行;[Aurora PostgreSQL-Compatible](https://docs.aws.amazon.com/AmazonRDS/latest/AuroraUserGuide/CHAP_AuroraOverview.html) 是 AWS 自研分布式存储之上的 PostgreSQL 兼容引擎,以集群而非单实例为管理单位。驱动和大多数 SQL 在两边都能工作,但参数、扩展、故障转移和计费行为并不一致。跨厂商视角见[云服务版图](/docs/cloud/service-map)。 ## 什么时候选哪一个 [#什么时候选哪一个] 以下情况默认选 RDS for PostgreSQL: * 希望行为尽可能接近社区 PostgreSQL(在托管服务允许的范围内); * 单主加可选 Multi-AZ 备库和只读副本即可满足可用性目标; * 存储远低于 64 TiB 上限,且偏好预分配、可预期的容量规划。 以下情况 Aurora PostgreSQL-Compatible 才值得它的溢价: * 需要更快的故障转移,以及共享同一集群存储卷的最多 15 个 Aurora 副本做读扩展; * 希望存储按 10 GiB 步长自动增长(上限 256 TiB),而不是提前预分配; * 愿意接受拥有独立发布节奏、版本编号和扩展清单的引擎。 这不是推荐排名。两者都是无主机访问的托管服务,正确选择取决于你实测的故障转移、连接和成本需求。 ## RDS 与 Aurora 并排对比 [#rds-与-aurora-并排对比] | | RDS for PostgreSQL | Aurora PostgreSQL-Compatible | | ---- | ------------------------------ | ----------------------------------------------------------------------------------------------------------------------- | | 存储 | 预分配 EBS(gp3/io2),上限 64 TiB | 共享集群存储卷,自动增长至 256 TiB,跨 3 个可用区保存 6 份副本 | | 副本 | 只读副本拥有独立存储,物理复制 | 最多 15 个 Aurora 副本共享集群存储卷,副本延迟通常很低 | | 故障转移 | Multi-AZ 备库提升,官方文档典型值 60–120 秒 | 有可用读节点时通常为数十秒;请自行实测 | | 备份 | 自动快照加 WAL,保留期内支持 PITR | 持续备份,保留窗口内支持 PITR | | 内核版本 | 按 RDS 发布日历提供社区大版本 | Aurora 自有发布节奏和版本编号 | | 扩展 | 平台允许清单 | 独立的[扩展支持矩阵](https://docs.aws.amazon.com/AmazonRDS/latest/AuroraPostgreSQLReleaseNotes/AuroraPostgreSQL.Extensions.html) | | 成本模型 | 实例加预分配存储 | 实例加实际消耗存储加 I/O 请求计费;同等工作负载下通常高于 RDS | 上述上限来自 [RDS 存储文档](https://docs.aws.amazon.com/AmazonRDS/latest/UserGuide/CHAP_Storage.html)与 [Aurora 概览](https://docs.aws.amazon.com/AmazonRDS/latest/AuroraUserGuide/CHAP_AuroraOverview.html),核对日期 2026-08-06。文档上限不等于对你工作负载的 SLA 承诺——请在 [RDS](https://aws.amazon.com/rds/postgresql/pricing/) 与 [Aurora](https://aws.amazon.com/rds/aurora/pricing/) 定价页建模成本,并在依赖任何数字之前亲自演练故障转移。 ## 需要接受的差异 [#需要接受的差异] * **Aurora 的 `max_connections` 计算方式不同。** 默认值为 `LEAST({DBInstanceClassMemory/9531392}, 5000)`——16 GiB 内存的实例约 1,800 个连接,且无论实例多大上限都是 5,000。高并发应用需要在任一服务前放置 [RDS Proxy](https://docs.aws.amazon.com/AmazonRDS/latest/UserGuide/rds-proxy.html) 或 PgBouncer;参见 [Aurora 参数参考](https://docs.aws.amazon.com/AmazonRDS/latest/AuroraUserGuide/AuroraPostgreSQL.Reference.ParameterGroups.html)。 * **扩展可用性是按引擎划分的允许清单。** 在 RDS 上可用的扩展(如 `pg_repack`)在 Aurora 上可能缺失或版本被锁定。迁移前先把 `pg_extension` 清单与 Aurora 支持矩阵比对,而不是迁移之后。 * **Aurora 以集群为单位管理。** 部分参数集群级生效,部分系统视图和等待事件与社区 PostgreSQL 不同,存储按消耗量加 I/O 请求计费而非预分配容量。 * **两个服务都不提供主机或 superuser 访问。** `rds_superuser` 是裁剪过的角色;任何需要 OS 访问、任意 `shared_preload_libraries` 或 untrusted 语言的方案在两边都行不通。 ## 生产前验证 [#生产前验证] 1. 把已安装扩展及其确切版本与目标引擎的允许清单逐一比对。 2. 做一次真实故障转移演练,计时应用的重连行为,而不只是 DNS 切换。 3. 从 `max_connections` 出发核算连接预算,再决定 RDS Proxy 或 PgBouncer 及其池化模式。 4. 向全新实例或集群执行一次 PITR 恢复,实测 RPO/RTO。 5. 用真实 I/O 水平建模成本——Aurora 单独计 I/O 请求费,从预分配存储的 RDS 迁来时常被这点打个措手不及。 6. 用 `pg_dump` 或快照导出一次,并在独立的 PostgreSQL 环境中恢复,作为退出路径验证。 连接池、备份演练和监控基线等与厂商无关的内容见[生产环境技术栈](/docs/operations/production-stack)。如果只需要开发或评估用数据库,[免费 PostgreSQL 选型](/docs/cloud/free-postgresql)列出了零成本替代方案。 --- # 免费 PostgreSQL 云数据库选型 Canonical URL: https://pg.edu.rich/docs/cloud/free-postgresql Last reviewed: 2026-08-02 免费数据库先按产品本质分类:**运行 PostgreSQL 的托管服务**、**以 PostgreSQL 为核心的开发平台**,以及**只兼容 pgwire/部分 SQL 的独立数据库**。驱动能连接,不代表扩展、事务语义和运维工具完全兼容。 本页按官方页面核对于 2026-08-02。免费额度、区域、项目数、休眠和备份策略变化很快;创建项目前重新打开来源核对。免费层通常不提供生产 SLA,也不能替代独立导出和恢复演练。 ## 免费 PostgreSQL 服务对比 [#免费-postgresql-服务对比] | 服务 | 当前免费额度快照 | 关键限制 | 更适合 | | --------------------------------------------------------------------------------------- | ------------------------------------------------------------------ | ---------------------------------------------------- | ----------------------------------- | | [Supabase](https://supabase.com/pricing) | 最多 2 个活跃项目;每项目 500 MB database;免费计划含 1 GB file storage、5 GB egress | 一周无活动暂停;免费层**没有自动备份和 PITR** | 需要 Auth、Storage、Realtime、API 的 BaaS | | [Neon](https://neon.com/pricing) | 最多 100 个项目;每项目 0.5 GB storage、100 CU-hours/月;5 GB public transfer | 空闲约 5 分钟 scale to zero;免费 restore window 有限 | 纯 PostgreSQL、branch、preview/CI、间歇负载 | | [Aiven for PostgreSQL](https://aiven.io/docs/products/postgresql/concepts/pg-free-tier) | 1 CPU、1 GB RAM、1 GB disk、包含 backup | 单节点、`max_connections=20`、无 HA/SLA/VPC/pooler;闲置可被关停 | 接近传统托管 PG 的学习和小型验证 | | [Nhost](https://nhost.io/pricing) | 1 个活跃项目;1 GB database、1 GB file storage、5 GB egress | 一周无活动暂停;GraphQL/Auth/Storage 属于平台耦合 | GraphQL-first、Hasura 与完整后端平台 | | [Prisma Postgres](https://www.prisma.io/pricing) | 500 MB storage、10 万 operations/月、最多 50 databases | 每个 SQL query 或 Prisma query 都计 operation;官方把免费层定位为评估 | Prisma 生态、临时数据库、PR/Agent 开发环境 | | [Koyeb PostgreSQL](https://www.koyeb.com/docs/databases) | 0.25 vCPU、1 GB RAM、1 GB data | 每月仅 5 小时 active compute,空闲后休眠 | Demo、教程和极低频测试,不适合常驻 API | | [Render Postgres](https://render.com/docs/free#free-postgres) | 1 GB、每 workspace 一个免费实例 | 30 天到期;无备份和 managed pooling,之后进入删除流程 | 一次性演示和平台试用 | 免费数字不能直接比较:Neon 的 CU-hour、Prisma 的 operation、Koyeb 的 active hour 与固定 VM 容量不是同一种计量单位。先用真实请求模式估算,再验证超额时是暂停、拒绝、删除还是自动收费。 ## CockroachDB 为什么单独列 [#cockroachdb-为什么单独列] [CockroachDB Cloud Basic](https://www.cockroachlabs.com/pricing/) 当前提供每月 5000 万 Request Units 与 10 GiB storage 的免费量,但 CockroachDB 是独立分布式 SQL 数据库,不是 PostgreSQL server。它支持 pgwire 和大量 PostgreSQL syntax,同时仍存在 range type、FDW、advisory lock、权限和事务行为差异;以 [PostgreSQL compatibility matrix](https://www.cockroachlabs.com/docs/stable/postgresql-compatibility) 为准。 如果目标是全球分布式事务和多区域容错,可以把它作为独立候选;如果目标是学习 PostgreSQL extension、系统目录、WAL 或运维,不要用它替代真实 PostgreSQL。 ## 直接选择建议 [#直接选择建议] | 需求 | 优先评估 | 原因 | | -------------------------------------- | --------------- | ---------------------------------- | | Auth、Storage、Realtime、REST/GraphQL API | Supabase | PostgreSQL 之外已有完整应用后端 | | 分支、preview database、scale-to-zero | Neon | 数据库生命周期适合 CI 和短时环境 | | 传统托管 PostgreSQL 体验 | Aiven | 资源与限制更像小型单节点托管实例 | | GraphQL-first | Nhost | PostgreSQL + Hasura + Auth/Storage | | Prisma 工作流和大量临时数据库 | Prisma Postgres | operation 计费与 Prisma/Agent 工具链结合紧密 | | 极短 Demo | Koyeb 或 Render | 免费限制决定了它们不是长期数据源 | 对 Drizzle、node-postgres、Kysely 等普通 PostgreSQL client,Neon、Supabase、Aiven、Nhost 和 Prisma Postgres 都应分别测试 direct/pooled URL、prepared statement、migration 与 transaction pooling 兼容性。不要因“可用标准连接串”跳过验证。 ## 免费层上线前检查 [#免费层上线前检查] ```sql SELECT version(), current_setting('server_version_num') AS server_version_num, current_database(), current_user; SELECT extname, extversion FROM pg_extension ORDER BY extname; SHOW max_connections; SHOW transaction_read_only; ``` 再逐项确认: 1. PostgreSQL major/minor 与扩展**具体版本**; 2. direct 和 pooled connection 的用途、上限与 pool mode; 3. idle 后冷启动、DNS/endpoint 是否改变; 4. 自动备份、PITR、保留期和免费层是否包含; 5. egress、operation/CU-hour 的计费口径与 hard limit; 6. `pg_dump` 导出、恢复到本地 PostgreSQL,以及项目暂停/删除后的取回窗口; 7. 付费升级是否原地完成,还是需要迁移或连接串切换。 重要生产系统至少需要可测量的 RPO/RTO、备份保留、恢复入口、支持与故障通知。即使选择付费托管服务,也要完成一次平台外导出与独立恢复。 生产能力与退出成本见 [云 PostgreSQL 生产选型清单](/docs/cloud/production-checklist);数据库血缘与兼容边界见 [PostgreSQL 血缘与兼容数据库](/docs/reference/postgresql-compatible-databases)。 --- # 云 PostgreSQL 入口 Canonical URL: https://pg.edu.rich/docs/cloud Last reviewed: 2026-08-06 “云 PostgreSQL”不是单一产品。先分清三类,才能知道哪些 PostgreSQL 经验可以直接复用: | 类别 | 典型产品 | 兼容边界 | 适合的目标 | | ---------------- | --------------------------------------------------------------------------------------------------------------------------------------- | ---------------------------------------- | ---------------------------------- | | 托管社区 PostgreSQL | [Amazon RDS for PostgreSQL](/docs/cloud/aws-rds-aurora)、Cloud SQL、Azure Database for PostgreSQL、阿里云 RDS、TencentDB | 运行社区内核,但主机权限、参数、扩展和升级受平台控制 | 希望保留较高 SQL/工具兼容性,同时把补丁、备份和 HA 交给平台 | | PostgreSQL 兼容增强型 | [Aurora PostgreSQL-Compatible](/docs/cloud/aws-rds-aurora)、[AlloyDB](/docs/cloud/alloydb)、[PolarDB for PostgreSQL](/docs/cloud/polardb) | 协议和大量 SQL 兼容;存储、复制、版本节奏和部分行为由厂商实现 | 愿意用平台架构换取弹性、读扩展或分析/AI 能力 | | 开发者数据平台 | [Neon](/docs/cloud/neon)、[Supabase](/docs/cloud/supabase) | PostgreSQL 是核心,但连接、分支、认证、API、实时或休眠语义属于平台 | 快速交付、预览环境、低运维团队或全栈产品 | 驱动能连接,只证明 wire protocol 可用。上线前仍要验证扩展版本、参数、系统视图、复制能力、连接池、备份导出、维护重启和故障切换行为。 ## 责任边界 [#责任边界] 托管服务通常替你处理基础设施、补丁编排、自动备份和部分故障转移,但以下工作仍属于应用团队: * schema、约束、索引、SQL 和事务设计; * 连接预算、池化方式与重试策略; * RPO/RTO 定义,以及真实恢复演练; * 数据访问、密钥、网络和最小权限; * 慢查询、膨胀、长事务、vacuum 与成本治理; * 大版本升级、扩展升级和退出计划。 AWS 对 Aurora 的说明也明确把查询优化归为客户责任;这是理解所有托管数据库的好起点。参见 [Amazon Aurora 概览](https://docs.aws.amazon.com/AmazonRDS/latest/AuroraUserGuide/CHAP_AuroraOverview.html)。 ## 推荐决策顺序 [#推荐决策顺序] 1. 写出数据驻留、合规、RPO、RTO、峰值连接、延迟和预算边界。 2. 确认所需 PostgreSQL 大版本、扩展及其**具体版本**。 3. 用生产形态的数据和 SQL 做基准,不使用供应商示例数字代替。 4. 演练维护、主备切换、PITR、连接耗尽和区域故障。 5. 用原生工具导出一次,并在独立 PostgreSQL 环境恢复。 本节事实核对日期为 **2026-08-02**。云功能、区域和套餐变化很快,采购与上线时应重新核对官方文档。 --- # Neon serverless PostgreSQL Canonical URL: https://pg.edu.rich/docs/cloud/neon Last reviewed: 2026-08-06 [Neon](https://neon.com/docs/get-started/why-neon) 是一个 serverless PostgreSQL 平台:计算与存储分离,计算可自动伸缩并在空闲时 scale to zero,写时复制存储让数据库分支便宜到可以为每个 pull request 建一个。它运行社区 PostgreSQL——截至本次核对,17 和 18 等新大版本均可用——因此驱动、SQL 和大多数扩展行为符合预期。与其他平台的定位对比见[云服务版图](/docs/cloud/service-map)。 ## Neon 改变了什么 [#neon-改变了什么] * **数据库分支成为一等工作流。** 分支是数据和 schema 的写时复制克隆,秒级创建。预览环境、CI 运行和迁移演练都可以各拿一个完整数据库,而不必重复支付存储成本。 * **Scale-to-zero 与自动伸缩。** 空闲计算自动挂起,下一个连接到来时恢复;繁忙时计算在配置范围内伸缩。间歇性负载不再为空转容量付费。 * **池化与直连两套端点。** Neon 同时提供基于 PgBouncer 的池化连接串和直连连接串,二者用途和上限不同。 * **按用量计费。** 计算按 CU-hour、存储按 GB·月计量,成本模型与按实例计费的服务不同;请用真实流量模式对照[定价页](https://neon.com/pricing)估算。 ## 数据库分支工作流 [#数据库分支工作流] 当前 [Neon CLI](https://neon.com/docs/reference/neon-cli) 以 `neon` 命令安装(`neonctl` 仍是别名),核对日期 2026-08-06: ```bash npm install -g neon neon auth neon projects create --name myapp neon branches create --name feature/x --parent main neon connection-string feature/x ``` 典型的「一个特性一个分支」流程: ```bash export DATABASE_URL=$(neon connection-string feature/x) psql "$DATABASE_URL" -c "ALTER TABLE users ADD COLUMN beta_flag boolean DEFAULT false;" # 分支不会自动合并——把评审过的 SQL 自己应用到 main psql "$(neon connection-string main)" -f migrations/0042_add_beta_flag.sql # 合并前比对 schema neon branches schema-diff main feature/x neon branches delete feature/x ``` 关键的纪律在中间一步:Neon 分支在创建时复制数据和 schema,但没有自动的 schema 合并。迁移仍要走正常的评审和迁移工具流程,像任何变更一样应用到 `main`。 ## 需要接受的差异 [#需要接受的差异] * **挂起后的冷启动。** 开启 scale-to-zero 后,空闲期结束的第一个连接要等待计算恢复。对延迟敏感的生产服务应调大或关闭挂起超时,并实测驱动和连接池在唤醒期间的重试行为。 * **项目区域在创建时固定。** 每个项目部署在单一 AWS 区域(Azure 区域正在逐步退出),创建后无法迁移区域——换区域意味着新建项目加数据搬迁。以当前[区域列表](https://neon.com/docs/introduction/regions)为准。 * **扩展来自允许清单。** 需要 OS 级访问或任意共享库的扩展不可用;在[扩展文档](https://neon.com/docs/extensions/pg-extensions)中确认确切清单和版本。 * **分支数据治理由你负责。** 分支默认复制生产数据。如果分支会进入 CI 或预览环境,需要规划脱敏,或从已匿名的父分支创建。 ## 生产前验证 [#生产前验证] 1. 经历真实空闲期后,用你实际的驱动、ORM 和连接池测量冷启动延迟。 2. 按工作负载决定池化还是直连,并在池化端点上测试 prepared statement 和 transaction pooling 行为。 3. 在分支进入共享环境之前,定好分支生命周期和数据脱敏规则。 4. 在所用套餐的恢复窗口内演练基于分支的恢复,并用 `pg_dump` 导出一次到独立 PostgreSQL 恢复。 5. 用一个真实流量周估算 CU-hour 和存储消耗,而不是按免费层形态估算。 免费套餐的当前额度见[免费 PostgreSQL 选型](/docs/cloud/free-postgresql)。它空闲几分钟后 scale to zero,且不附带生产 SLA——适合评估,不构成生产承诺。 池化、备份演练和监控等跨平台内容见[生产环境技术栈](/docs/operations/production-stack)。 --- # 阿里云 PolarDB for PostgreSQL Canonical URL: https://pg.edu.rich/docs/cloud/polardb Last reviewed: 2026-08-06 PolarDB for PostgreSQL 是阿里云的云原生 PostgreSQL 兼容数据库:计算与存储分离,集群内节点共享分布式存储,只读节点独立于主节点扩展。架构上它与 Aurora 同属一类——厂商自研共享存储之上的 PostgreSQL 兼容引擎——只是运行在阿里云生态内。跨厂商视角见[云服务版图](/docs/cloud/service-map)。 ## PolarDB 的定位 [#polardb-的定位] 当基础设施已经在阿里云上,或工作负载必须部署在 AWS 和 Google Cloud 均未运营的中国大陆地域时,PolarDB 是候选。如果需要与 AWS 或 GCP 服务深度集成,或要求社区 PostgreSQL 小版本发布当周即可用,它都不合适——任何厂商的托管引擎都落后于社区发布日历。 截至 2026-08,PolarDB for PostgreSQL 支持社区大版本 11、14、15、16、17 和 18——PostgreSQL 18 兼容版本于 2025 年 12 月发布——以官方[大版本生命周期说明](https://help.aliyun.com/zh/polardb/polardb-for-postgresql/major-version-lifecycle-description)为准。注意阿里云同时运营一个 Oracle 语法兼容的 PolarDB 版本;确认你评估的是 PostgreSQL 版。 ## 架构与兼容性 [#架构与兼容性] * **共享分布式存储。** 计算节点挂载跨可用区多副本的共享存储卷,增加只读节点不需要复制数据集。 * **弹性操作。** 计算规格和节点数量独立于存储调整;故障转移提升已有只读节点。 * **扩展允许清单。** 可用扩展及其版本按引擎版本和内核小版本划分;把你的需求与官方[支持插件列表](https://help.aliyun.com/zh/polardb/polardb-for-postgresql/list-of-supported-plug-ins)逐一比对,不要假设与社区一致。 * **开源血缘。** PolarDB for PostgreSQL 内核同时以开源形式发布([openpolardb.com](https://openpolardb.com)),这对评估和长期退出思路有意义——但托管服务与开源构建不是同一个产物。 ## 中国合规与数据驻留视角 [#中国合规与数据驻留视角] 驻留是 PolarDB 进入候选名单最强的理由,值得精确表述: * 中国大陆地域由阿里云中国主体在中国监管框架下运营,与阿里云国际站是**相互独立的账号体系**。账号、计费和支持在两个体系之间不通用。 * 如果用户和数据必须留在中国内地——出于延迟、ICP 相关部署要求或数据驻留义务——中国地域的 PolarDB 集群能满足任何 AWS 或 GCP 地域都无法满足的约束。 * 反过来说,把中国用户的个人信息存放在国际站地域的集群,会引出 PIPL 下的跨境传输问题。两个方向都不是自我声明即可合规:请与法务或合规团队确认当前要求,并书面记录数据由哪个账号、哪个地域、哪个法律主体持有。 ## 需要接受的差异 [#需要接受的差异] * **参数与权限控制。** 与其他托管引擎一样,部分参数被锁定或由平台管理,也不提供 superuser 等价权限。 * **小版本滞后。** 社区小版本按阿里内核的节奏合入,而非社区节奏;查发布记录确认某个 PostgreSQL 大版本背后的内核版本。 * **价格对比不可移植。** 实例、存储和节点计费结构与阿里云 RDS for PostgreSQL、与 AWS Aurora 都不同;在阿里云价格计算器里建模自己的工作负载,不要信百分比口径的说法。 * **工具体系。** 控制台、CLI、监控和迁移(DTS)都是阿里云特有的;其他云的运维手册不能直接搬用。 ## 生产前验证 [#生产前验证] 1. 把所需扩展——`pgvector`、`pg_trgm`、PostGIS、`pg_cron` 及任何版本敏感项——与目标引擎版本的支持插件列表比对。 2. 检查你依赖的非默认 `postgresql.conf` 设置在 PolarDB 上的参数兼容性。 3. 用阿里云 DTS(全量加增量)或 `pg_dump` 演练迁移,包括回滚方案和字符集检查。 4. 做一次故障转移演练,经你的驱动和连接池测量重连行为。 5. 如果需要跨地域或跨境副本,先用全球数据库网络(GDN)功能对照驻留义务做验证。 6. 用原生工具导出一次,并在独立 PostgreSQL 环境中恢复,作为退出路径验证。 版本支持、扩展可用性和产品结构按阿里云官方文档核对于 2026-08-06。中国站与国际站的能力可能不同;按你实际使用的账号类型核对对应文档。 连接预算、备份演练和监控基线等跨厂商内容见[生产环境技术栈](/docs/operations/production-stack)。开发规模的评估方案见[免费 PostgreSQL 选型](/docs/cloud/free-postgresql)。 --- # 云 PG 生产选型清单 Canonical URL: https://pg.edu.rich/docs/cloud/production-checklist Last reviewed: 2026-08-02 ## 1. 兼容性清单 [#1-兼容性清单] * PostgreSQL 大版本、次版本补丁节奏和停止支持日期是什么? * `pg_extension` 中需要的扩展及版本是否都可用?升级是自动、手工还是需要迁移? * 哪些参数不可改?是否允许 `shared_preload_libraries`? * 是否支持逻辑复制、复制槽、FDW、事件触发器和所需认证方式? * 系统目录、统计视图和超级用户操作有哪些替代接口? * 驱动、ORM、迁移工具和备份工具是否通过真实流水线测试? 把检查结果保存为机器可读清单,绑定服务 SKU、区域、引擎版本和核对日期。 ## 2. 可用性与恢复 [#2-可用性与恢复] | 测试 | 通过条件示例 | | ------ | --------------------------------------- | | 强制主备切换 | 客户端在预算内重连;事务失败以可识别 SQLSTATE 返回;没有静默部分成功 | | PITR | 恢复到指定时间的新实例;校验业务行数、约束、角色和扩展;实测 RTO | | 误删恢复 | 明确整实例、整库、单表各自的恢复路径和耗时 | | 区域故障 | DNS、密钥、对象存储备份和应用计算不与数据库同故障域 | | 备份导出 | 能在厂商账号之外恢复一份可用副本 | 应用仍需设置连接超时、事务级重试和幂等键。不要重放一个已经可能提交成功的写操作,除非可以用业务幂等键确认结果。 ## 3. 连接与弹性 [#3-连接与弹性] 计算每个应用副本、后台任务、迁移工具、BI 和 Agent 的连接上限。对 serverless/Agent 流量优先使用受控池,但要确认: * transaction pooling 是否与 session state、临时表、LISTEN/NOTIFY 或 prepared statements 兼容; * 缩容、休眠、故障切换时连接字符串和 TLS 证书是否变化; * `statement_timeout`、`idle_in_transaction_session_timeout` 和客户端超时谁先触发; * 突发请求是否在应用层排队,而不是把连接风暴直接传给 PostgreSQL。 ## 4. 成本模型 [#4-成本模型] 除计算与存储外,至少估算 IOPS、备份、跨区流量、只读副本、日志、监控、代理、PITR、快照导出和支持计划。AI 工作负载还应单列 embedding、索引重建、向量存储和检索候选重排成本。 ## 5. 可迁移性 [#5-可迁移性] 每季度或大版本升级前执行一次: ```bash pg_dump --format=custom --no-owner --no-acl "$DATABASE_URL" > app.dump createdb portability_restore pg_restore --exit-on-error --no-owner --no-acl \ --dbname=portability_restore app.dump ``` 这只是逻辑可迁移性检查,不替代平台 PITR。恢复后还要核对 extension、role/grant、large object、sequence、行数、约束、关键查询结果和执行计划。 ## 上线证据包 [#上线证据包] * 服务/区域/SKU/引擎/扩展版本清单; * RPO、RTO、连接预算和容量模型; * failover、PITR、误删与平台外恢复报告; * 加密、网络、角色、RLS 与密钥轮换记录; * 版本升级、扩展升级和厂商退出 runbook; * 基准负载下的延迟、错误、WAL、vacuum、存储与成本数据。 --- # 云 PostgreSQL 服务版图 Canonical URL: https://pg.edu.rich/docs/cloud/service-map Last reviewed: 2026-08-06 ## 托管社区 PostgreSQL [#托管社区-postgresql] | 服务 | 已核实的能力 | 选型时重点验证 | | --------------------------------------------------------------------------------------------------------------------- | ------------------------------------------- | --------------------------------------------------- | | [Amazon RDS for PostgreSQL](https://docs.aws.amazon.com/AmazonRDS/latest/UserGuide/CHAP_PostgreSQL.html) | 自动备份/PITR、Multi-AZ、只读副本、VPC 与 TLS | 无主机访问;参数、系统能力和扩展来自平台允许清单 | | [Cloud SQL for PostgreSQL](https://docs.cloud.google.com/sql/docs/postgres/introduction) | 托管备份、HA/故障转移、加密、私网/公网、只读副本和维护编排 | 维护或部分配置可能重启;核对区域、扩展、连接与 AI 功能可用性 | | [Azure Database for PostgreSQL Flexible Server](https://learn.microsoft.com/en-us/azure/postgresql/overview) | 同区/跨可用区 HA、PITR、TLS、私网、托管维护、可启用内置 PgBouncer | 自动备份保留默认 7 天、最长 35 天;内置 PgBouncer 使用 6432 端口,核对池化模式 | | [阿里云 RDS PostgreSQL](https://help.aliyun.com/zh/rds/apsaradb-rds-for-postgresql/what-is-apsaradb-rds-for-postgresql/) | 基础版、高可用版与集群版;自动/手工备份、只读实例和数据库代理 | 高可用版备节点不可直接访问;同步模式、代理路由和备份类型随架构变化 | | [TencentDB for PostgreSQL](https://cloud.tencent.com/document/product/409) | 托管安装、存储、HA、备份、大/小版本升级与只读实例组 | 官方说明单个只读实例不具备 HA/SLA;生产读组应核对节点数、路由和一致性 | 云厂商通常不会在社区发布当天立即提供相同内核或扩展版本。把“支持 PostgreSQL 17/18”拆成三个问题:能否新建、能否从旧版本升级、目标扩展是否支持该版本。AWS 的两条路径——托管社区 RDS 与 Aurora 引擎——见 [AWS RDS for PostgreSQL 与 Aurora](/docs/cloud/aws-rds-aurora)。 ## 兼容增强型引擎 [#兼容增强型引擎] | 服务 | 架构特点 | 需要接受的差异 | | --------------------------------------------------------------------------------------------------------------------- | ---------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------ | | [Aurora PostgreSQL-Compatible](https://docs.aws.amazon.com/AmazonRDS/latest/AuroraUserGuide/CHAP_AuroraOverview.html) | 定制 PostgreSQL 兼容引擎和分布式存储,以集群而非单实例为主要管理单位 | Aurora 自有版本节奏;扩展来自[支持清单](https://docs.aws.amazon.com/AmazonRDS/latest/AuroraPostgreSQLReleaseNotes/AuroraPostgreSQL.Extensions.html),不会随社区扩展自动升级 | | [AlloyDB for PostgreSQL](https://docs.cloud.google.com/alloydb/docs/overview) | 计算/存储解耦、跨区 HA,可选列式引擎,并提供向量和模型集成 | 不是社区二进制等价物;验证扩展、参数、迁移工具、分析路径和区域能力 | “PostgreSQL-compatible”适合描述迁移起点,不应作为测试结论。至少跑 schema 迁移、关键查询、事务并发、驱动、扩展和故障恢复测试。各引擎落地页:[AWS RDS for PostgreSQL 与 Aurora](/docs/cloud/aws-rds-aurora)、[AlloyDB for PostgreSQL](/docs/cloud/alloydb)、[阿里云 PolarDB for PostgreSQL](/docs/cloud/polardb)。 ## 开发者平台 [#开发者平台] | 服务 | 强项 | 容易忽略的生产问题 | | -------------------------------------------------------------- | ---------------------------------------------------------- | ------------------------------------------------ | | [Neon](https://neon.com/docs/get-started/why-neon) | 计算/存储分离、自动伸缩、scale-to-zero、数据库分支和池化连接 | 冷启动、计算规格变化、分支数据治理,以及 pooled/direct 连接的用途差异 | | [Supabase](https://supabase.com/docs/guides/database/overview) | 每项目完整 PostgreSQL,并集成 Auth、Storage、Realtime、API 和 Supavisor | 浏览器直连数据 API 前必须正确设计 RLS;核对备份/PITR 套餐、连接预算与平台组件耦合 | 需要零成本开发环境时,查看按 2026-08-02 核对的 [免费 PostgreSQL 云数据库选型](/docs/cloud/free-postgresql)。Supabase、Neon、Nhost 与 Prisma Postgres 的平台结构不同,不能只按免费 storage 数字排序。平台落地页:[Neon](/docs/cloud/neon) 与 [Supabase](/docs/cloud/supabase)。 必须验证恢复粒度、保留期、跨区域副本、密钥依赖、导出能力和实测恢复时间。Azure 等平台的托管物理备份不能直接导出到平台外;退出路径通常要另做 `pg_dump`、逻辑复制或迁移服务。 ## 面向 AI 工作负载 [#面向-ai-工作负载] 选择云 PG 承载 RAG 或 Agent 元数据时,优先核对: 1. `pgvector` 的**版本**、HNSW/IVFFlat 支持和升级节奏; 2. 最大连接数与池化方式,尤其是短生命周期函数/Agent; 3. 向量索引构建的内存、临时存储、WAL 和副本延迟; 4. tenant/ACL 过滤后的真实召回率,而非无过滤基准; 5. embedding 模型、向量数据和数据库是否满足同一驻留边界; 6. 是否能导出原文、metadata 与 embedding,避免数据管道锁定。 云厂商集成的模型端点、自动 embedding 或 AI 助手能减少胶水代码,但也增加权限、区域、模型生命周期和成本维度;它们不能替代数据库侧的 RLS、最小角色和检索评测。 如果候选只支持 PostgreSQL protocol 或复用了 query layer,继续检查 [PostgreSQL 血缘与兼容数据库](/docs/reference/postgresql-compatible-databases)。 --- # Supabase PostgreSQL 平台 Canonical URL: https://pg.edu.rich/docs/cloud/supabase Last reviewed: 2026-08-06 [Supabase](https://supabase.com/docs/guides/database/overview) 为每个项目提供一个完整、独立的 PostgreSQL 数据库,并在其周围集成平台服务:Auth、Storage、Realtime、自动生成的 REST 与 GraphQL API(PostgREST)、Edge Functions,以及 Supavisor 连接池。数据库是真实的 PostgreSQL——可以用任何客户端连接、使用普通扩展——但安全模型和日常运维由平台塑造。与其他方案的定位对比见[云服务版图](/docs/cloud/service-map)。 ## 平台包含什么 [#平台包含什么] * **每项目一个完整 PostgreSQL 实例**,同时提供直连和池化连接串。 * **与数据库集成的 Auth**:用户身份存放在 `auth.users`,JWT 声明可在 SQL 内读取。 * **自动生成的数据 API**,把 schema 中的表直接暴露给浏览器和移动客户端。 * **基于 PostgreSQL 逻辑复制的 Realtime** 推送。 * **Storage 与 Edge Functions**,处理文件和靠近数据库的服务端逻辑。 这种捆绑正是选择 Supabase 的理由:一个平台替代整个后端层。它同时也是需要接受的主要代价——采用 Auth、Storage、Realtime 和 API 越多,应用与 Supabase 特有 schema 和服务的耦合就越深。 ## RLS 是安全边界 [#rls-是安全边界] 由于浏览器客户端可以通过自动生成的 API 直接访问表,[行级安全](https://supabase.com/docs/guides/database/postgres/row-level-security)不是加固选项,而是安全模型本身: ```sql ALTER TABLE notes ENABLE ROW LEVEL SECURITY; CREATE POLICY "users can read their own notes" ON notes FOR SELECT USING (user_id = auth.uid()); CREATE POLICY "users can insert their own notes" ON notes FOR INSERT WITH CHECK (user_id = auth.uid()); ``` `auth.uid()` 从请求的 JWT 中读取用户 id,因此前端可以直接经 API 查询,由数据库强制所有权边界。三条规则保证这套模型不失效: 1. 对每张经 API 暴露的表启用 RLS,包括之后新增的表——schema 暴露什么,API 就暴露什么。 2. 用匿名和已认证两种角色测试策略,而不是只用控制台的 service 连接验证。 3. 永远不要向客户端下发 `service_role` key;它会完全绕过 RLS。 ## 需要接受的差异 [#需要接受的差异] * **备份与 PITR 取决于套餐。** 免费计划没有自动备份和 PITR——当前数字见[免费 PostgreSQL 选型](/docs/cloud/free-postgresql)——付费档位的保留期和恢复粒度也不同。写入数据前,确认[备份文档](https://supabase.com/docs/guides/platform/backups)与你的 RPO 匹配。 * **连接预算由平台决定。** 直连连接数有限;serverless 和高并发客户端走 Supavisor,其 transaction pooling 模式会限制 prepared statement 等会话级特性。用池化连接串测试你的驱动和 ORM,而不是只测直连。 * **Realtime 消耗数据库资源。** 它经逻辑复制流式推送变更,高写入频率的表和过大的 publication 会在同一个服务查询的实例上产生真实的 WAL 和复制负载。 * **退出时需要拆解平台组件。** `pg_dump` 能导出数据,但 Auth 用户、Storage 对象、Realtime 订阅和 Edge Functions 都是平台服务,dump 不会包含它们。 ## 生产前验证 [#生产前验证] 1. 每张经 API 暴露的表都已启用 RLS,并从匿名和已认证两种上下文完成策略测试。 2. 对照实际付费的套餐确认备份、PITR 和恢复粒度,并完成一次真实恢复演练。 3. 用池化连接串测试驱动、ORM 和迁移工具,包括 prepared statement 行为。 4. 按生产形态的写入频率压测 Realtime 负载。 5. 完成一次覆盖 `pg_dump` 加 Auth、Storage、函数处置方案的导出演练。 套餐额度、备份覆盖范围和平台功能按 Supabase 官方文档核对于 2026-08-06,变化频繁;采购时请重新核实。 池化、监控和恢复演练等跨平台内容见[生产环境技术栈](/docs/operations/production-stack)。 --- # PostgreSQL autovacuum 与表膨胀 Canonical URL: https://pg.edu.rich/docs/operations/autovacuum-bloat Last reviewed: 2026-08-06 标准 `VACUUM` 的目标不只是“释放空间”:它让 dead row versions 可复用、维护 planner statistics 和 visibility map,并防止 transaction ID/multixact wraparound。多数系统应保持 autovacuum 开启。 ## 日常观测 [#日常观测] ```sql SELECT schemaname, relname, n_live_tup, n_dead_tup, last_vacuum, last_autovacuum, vacuum_count, autovacuum_count, last_analyze, last_autoanalyze FROM pg_stat_user_tables ORDER BY n_dead_tup DESC LIMIT 30; ``` 统计是估算且会重置,不能仅凭一个 `n_dead_tup` 阈值判断膨胀。结合表大小、更新速率、查询延迟、autovacuum 日志和趋势。 查看正在运行的 vacuum: ```sql SELECT pid, datname, relid::regclass AS relation, phase, heap_blks_total, heap_blks_scanned, heap_blks_vacuumed, index_vacuum_count, dead_tuple_bytes, num_dead_item_ids, indexes_total, indexes_processed FROM pg_stat_progress_vacuum; ``` 这些字段名对应 PostgreSQL 18;较早 major 的 progress view 列可能不同,跨版本监控应先核对目标版本目录。 ## 为什么没有触发或跟不上 [#为什么没有触发或跟不上] * 表很大,默认 scale factor 对应的变更行数过高; * worker、I/O 或维护内存不足; * 长事务、prepared transaction、复制槽或 standby snapshot 阻止回收; * vacuum 经常被冲突锁取消; * 写入峰值持续高于清理能力。 针对已验证的热点表覆盖参数,而不是先全局激进调整: ```sql ALTER TABLE app.events SET ( autovacuum_vacuum_scale_factor = 0.02, autovacuum_vacuum_threshold = 1000, autovacuum_analyze_scale_factor = 0.01 ); ``` 参数只是示例。根据表大小和每天变更量计算触发频率,并观察 I/O、WAL、延迟与实际完成时间。 ## 手工维护边界 [#手工维护边界] ```sql VACUUM (ANALYZE, VERBOSE) app.events; ``` 普通 `VACUUM` 主要让空间在关系内部复用,通常不会把文件缩回操作系统。`VACUUM FULL` 会重写整张表、需要额外磁盘并获取 `ACCESS EXCLUSIVE` 锁,不是日常清理命令。 ## 测量与修复膨胀 [#测量与修复膨胀] `pg_stat_user_tables` 中的 dead tuple 比例是估算值。[`pgstattuple`](https://www.postgresql.org/docs/18/pgstattuple.html) 会扫描整个关系,给出精确的 dead tuple 与空闲空间: ```sql CREATE EXTENSION IF NOT EXISTS pgstattuple; SELECT * FROM pgstattuple('app.events'); -- dead_tuple_count, dead_tuple_percent, free_space, free_percent ``` 当必须把空间还给操作系统时,普通 `VACUUM` 做不到,而 `VACUUM FULL` 在整个重写期间阻塞读写。常用的在线方案: ```bash # pg_repack 在后台重建表,只持有短暂的锁 sudo -u postgres pg_repack -h db.example -U postgres -d commerce --table=events ``` ```sql -- 仅索引膨胀:不锁表重建单个索引 REINDEX INDEX CONCURRENTLY app.events_pkey; ``` `REINDEX CONCURRENTLY` 自 PostgreSQL 12 起内建。`pg_repack` 是第三方扩展,有独立的版本与运维要求——安装、版本核对和演练都应独立于核心功能进行。 ### VACUUM FULL 看似卡住时 [#vacuum-full-看似卡住时] 几乎总是在等现有会话释放 `ACCESS EXCLUSIVE` 锁。重试前先找出阻塞者: ```sql SELECT a.pid, a.usename, a.state, a.wait_event, left(a.query, 160) AS query FROM pg_stat_activity a WHERE a.pid = ANY (pg_blocking_pids(12345)); -- VACUUM FULL 会话的 pid ``` 确认持有者可以放弃后再 cancel 或 terminate。即使成功运行,`VACUUM FULL` 也需要约等于表大小的空闲磁盘并重写全部索引——这也是日常回收空间优先用 `pg_repack` 的原因。 ## 事务 ID wraparound 防护 [#事务-id-wraparound-防护] 事务 ID 是 32 位:约 20 亿个事务之后,旧元组会看起来“来自未来”,PostgreSQL 会在此之前停止接受写入。autovacuum 通过 freeze 旧元组回收事务 ID 空间来防止回卷。 按 database 和表跟踪 freeze age: ```sql SELECT datname, age(datfrozenxid) AS xid_age FROM pg_database ORDER BY xid_age DESC; SELECT n.nspname, c.relname, age(c.relfrozenxid) AS xid_age, pg_size_pretty(pg_total_relation_size(c.oid)) AS total_size FROM pg_class c JOIN pg_namespace n ON n.oid = c.relnamespace WHERE c.relkind = 'r' AND n.nspname NOT IN ('pg_catalog', 'information_schema') ORDER BY xid_age DESC LIMIT 20; ``` 常用的告警起点:database 级 `xid_age` 超过 10 亿需要关注,超过 15 亿属于紧急情况——具体阈值应结合事务速率校准。如果 autovacuum 来不及 freeze,在低峰窗口强制 freeze,并在 `pg_stat_progress_vacuum` 中观察进度: ```sql VACUUM (FREEZE, VERBOSE) app.events; ``` 整库 age 偏高时省略表名,对整个 database 执行。 先找出具体表、阶段、等待事件和资源瓶颈。关闭 autovacuum 会积累 dead tuples、陈旧统计和冻结风险;反 wraparound vacuum 即使表级设置关闭也可能运行。 完整原理见 [PostgreSQL 18 Routine Vacuuming](https://www.postgresql.org/docs/18/routine-vacuuming.html)。 --- # 备份、恢复与 PITR Canonical URL: https://pg.edu.rich/docs/operations/backup-recovery Last reviewed: 2026-08-06 ## 选择工具 [#选择工具] | 需求 | 工具/方式 | 关键边界 | | ----------- | ------------------------------------- | ----------------------- | | 单库、可移植、选择对象 | `pg_dump` / `pg_restore` | 不包含集群级角色与 tablespace 定义 | | 全集群逻辑对象 | `pg_dumpall --globals-only` 配合各库 dump | 大库恢复慢,需重建索引 | | 整实例快速恢复 | `pg_basebackup` 或成熟备份工具 | 版本和平台约束更强 | | 恢复到某一时间点 | 物理基准备份 + 连续 WAL 归档 | 必须持续验证 WAL 完整性 | ### pgBackRest、WAL-G 与 pg\_dump 怎么选 [#pgbackrestwal-g-与-pg_dump-怎么选] | 方案 | 更适合 | 不足以单独证明 | | -------------------------------------------------------- | ----------------------------------------------------------------- | ------------------------ | | [`pgBackRest`](https://github.com/pgbackrest/pgbackrest) | 自托管实例的 full/differential/incremental、并行备份、多 repository、WAL 与 PITR | 目标 RTO 已达标,密钥和所有 WAL 都可用 | | [`WAL-G`](https://github.com/wal-g/wal-g) | 对象存储导向的物理备份与 WAL 工作流 | repository 保留、删除保护和恢复正确 | | `pg_dump` / `pg_restore` | 逻辑迁移、选择对象、小规模恢复与跨版本导出 | 连续时间点恢复或整实例低 RTO | | 云平台备份 | 降低基础设施维护量 | 跨账户、跨区域、平台外恢复和全部扩展可恢复 | 不要机械套用固定的每日/每周频率。由 RPO、WAL 生成量、恢复带宽、保留策略和实测 RTO 反推 backup cadence,并保留至少一份独立于主数据库权限边界的副本。 ## 逻辑备份 [#逻辑备份] 自定义格式支持并行恢复与选择对象: ```bash pg_dump \ --format=custom \ --file=commerce-20260802.dump \ --dbname='postgresql://backup@db.example/commerce' pg_restore --list commerce-20260802.dump createdb commerce_restore_test pg_restore \ --dbname=commerce_restore_test \ --jobs=4 \ --exit-on-error \ commerce-20260802.dump ``` `pg_dump` 在导出期间提供一致快照,但只能备份一个 database。角色等全局对象另行备份: ```bash pg_dumpall --globals-only > globals-20260802.sql ``` 不要把包含密码哈希的 globals 文件放进普通制品库。 ## 物理备份与 PITR [#物理备份与-pitr] PITR 需要:可用的基准备份、从基准备份起连续完整的 WAL、正确的恢复配置,以及时间线管理。只保存 WAL 不够;只做 base backup 也无法恢复到任意时间点。 归档命令必须只在安全复制成功后返回 0,并避免覆盖已有文件。对象存储通常需要成熟备份工具管理并发、校验、保留和加密,而不是一条未经监控的 shell 命令。 持续告警 archive failure、缺失 WAL、repository 容量和最近一次可恢复时间。删除旧备份前,让工具按依赖关系计算保留链;不要只按文件日期手工删除。 ### PITR 分步实操 [#pitr-分步实操] 手工恢复到某个时间点的最小流程: ```bash sudo systemctl stop postgresql # 把基准备份解压到清空的数据目录 tar -xzf /backup/2026-08-01/base.tar.gz -C "$PGDATA" touch "$PGDATA"/recovery.signal ``` ```ini # postgresql.conf(或 postgresql.auto.conf) restore_command = 'cp /archive/%f %p' recovery_target_time = '2026-08-01 14:30:00+08' ``` 启动后实例回放 WAL 到目标点并暂停(默认 `recovery_target_action = 'pause'`)。确认数据后执行 `SELECT pg_wal_replay_resume();` 完成恢复。时间线处理同样需要演练:每次完成恢复都会产生新的 timeline,归档必须能跟上。 ### pgBackRest 配置示例 [#pgbackrest-配置示例] 最小 repository 配置: ```ini # /etc/pgbackrest/pgbackrest.conf [global] repo1-path=/var/lib/pgbackrest # 示例值;保留策略应由 RPO 和存储容量推导 repo1-retention-full=4 repo1-cipher-type=aes-256-cbc repo1-cipher-pass= [main] pg1-path=/var/lib/postgresql/18/main ``` ```bash sudo -u postgres pgbackrest --stanza=main stanza-create sudo -u postgres pgbackrest --stanza=main check sudo -u postgres pgbackrest --stanza=main --type=full backup sudo -u postgres pgbackrest --stanza=main --type=incr backup sudo -u postgres pgbackrest --stanza=main info # 时间点恢复 sudo systemctl stop postgresql sudo -u postgres pgbackrest --stanza=main \ --type=time --target='2026-08-01 14:30:00+08' restore sudo systemctl start postgresql ``` `check` 会端到端验证 WAL 归档,每次修改配置后都应运行。保留链、加密密钥和 repository 权限在恢复时必须全部可用——丢失 cipher pass 等于丢失整个 repository。 ## 恢复演练 [#恢复演练] 每次演练都在隔离实例上进行——备用主机或测试机的另一个端口,绝不用生产数据目录: 1. 把最近一次备份恢复到临时路径,例如 `pgbackrest --stanza=main --pg1-path=/tmp/restore-test restore`。 2. 用独立端口启动一个临时实例:`postgres -D /tmp/restore-test -p 6543`。 3. 运行下面的校验查询与冒烟测试。 4. 确认所需扩展、角色和恢复密钥都可用;单库 dump 不包含集群级角色。 5. 停止并删除临时实例。 每次演练记录:备份 ID、起止时间、恢复目标、数据库版本、所需密钥、实际 RTO、可恢复到的最新事务时间、校验查询和异常。 验证至少包括: ```sql SELECT count(*) FROM critical_table; SELECT min(created_at), max(created_at) FROM critical_table; SELECT conname, convalidated FROM pg_constraint WHERE NOT convalidated; SELECT indexrelid::regclass, indisvalid FROM pg_index WHERE NOT indisvalid; ``` 再运行应用层只读冒烟测试。行数相同不证明业务关系和权限正确。 复制会快速复制误删、错误更新和逻辑损坏。备份需要独立保留、删除保护、校验和恢复演练。 --- # PostgreSQL 配置调优(postgresql.conf) Canonical URL: https://pg.edu.rich/docs/operations/configuration Last reviewed: 2026-08-06 默认的 `postgresql.conf` 取值保守,是为了让 PostgreSQL 几乎在任何硬件上都能启动。这意味着默认值不适合专用服务器,但反过来也不存在一组放之四海皆准的"最佳配置"。下面的数值是 **PostgreSQL 18 的起步值**,来自常用经验法则:先应用,再针对自己的负载逐项验证。 ## 先测量,再调优 [#先测量再调优] 没有基线就改 GUC,结果往往是把一个慢系统变成另一种慢系统。动手之前: * 启用 `pg_stat_statements`(加入 `shared_preload_libraries` 并重启,然后 `CREATE EXTENSION pg_stat_statements;`),用它按总耗时和 `shared_blks_read` / `shared_blks_hit` 给查询排序; * 记录当前的延迟、I/O 和 checkpoint 行为,保证之后每个改动都能归因; * 每次只改一组参数,改完重新测量。 ```sql SELECT query, calls, total_exec_time, mean_exec_time, shared_blks_hit, shared_blks_read FROM pg_stat_statements ORDER BY total_exec_time DESC LIMIT 20; ``` 如果 `pg_stat_statements` 显示热点查询是缺索引或在小表上全表扫描,任何内存参数都救不了它。测量指向哪里,就调哪里。 ## 核心内存 GUC [#核心内存-guc] ### shared\_buffers [#shared_buffers] PostgreSQL 自己的 page cache,启动时从共享内存分配。修改需要重启。 ```sql ALTER SYSTEM SET shared_buffers = '4GB'; ``` 专用数据库服务器上常见的起步值是**内存的 25%**。流传的"25%,不要超过 40%"上限是经验之谈而非文档中的硬性限制——写入多的负载可能受益于更大的值,而热数据已经能放进 shared\_buffers 的读多负载未必需要。想调大就用 `pg_buffercache` 命中率和端到端延迟验证,不要默认越大越好。 ### effective\_cache\_size [#effective_cache_size] 它**不分配内存**,只是告诉规划器 shared\_buffers 加上 OS page cache 大概能缓存多少数据,主要影响 index scan 与 seq scan 之间的选择。reload 即生效。 ```sql ALTER SYSTEM SET effective_cache_size = '12GB'; ``` 合理的起步值是**内存的 50–75%**。设得远高于实际可缓存的内存会让规划器对 index scan 过于乐观,调整前后用执行计划对比确认。 ### work\_mem [#work_mem] **单个排序、hash 等操作**的内存上限——一条复杂查询可能同时消耗好几份,再乘以并发连接数。它是全局调大时最容易导致 OOM 的参数。 ```sql ALTER SYSTEM SET work_mem = '16MB'; ``` * 默认 4MB 会让很多排序和 hash 落盘;`log_temp_files`(见下文[日志](#生产环境建议开启的日志))可以直接告诉你是否正在发生。 * 重分析型用户按角色或会话单独调大,不要全局调:`ALTER ROLE analytics SET work_mem = '256MB';` * 全局上限的经验值约为 `内存 × 0.25 / max_connections`,而且这已经假设每个连接同时只跑一个大操作。 ### maintenance\_work\_mem [#maintenance_work_mem] 供 `VACUUM`、`CREATE INDEX`、`ALTER TABLE ADD FOREIGN KEY` 等维护操作使用。 ```sql ALTER SYSTEM SET maintenance_work_mem = '1GB'; ``` 调大可以加快建索引和 vacuum。注意并行建索引和并行 vacuum 的每个 worker 会各自申请一份,总量可能是该值的好几倍。 ## 连接数与连接池 [#连接数与连接池] ```sql ALTER SYSTEM SET max_connections = 200; ``` PostgreSQL 每个连接是一个独立的后端进程,占用数 MB 内存和调度开销,所以 `max_connections` 是预算而不是吞吐量旋钮。应用需要的并发超过数据库能承载的连接数时,正确做法是在前面放连接池,而不是无限调大这个值——PgBouncer 的取舍(包括 transaction 模式下哪些特性会失效)见 [PostgreSQL 生产工具栈与高可用](/docs/operations/production-stack)。 ```sql SELECT count(*), state FROM pg_stat_activity GROUP BY state; ``` `idle in transaction` 数量大说明应用长时间不提交事务,应该先在应用侧修复,而不是调连接数上限。`max_connections` 修改需要重启。 ## WAL 与 checkpoint [#wal-与-checkpoint] ```ini wal_compression = on max_wal_size = 4GB min_wal_size = 1GB checkpoint_timeout = 15min checkpoint_completion_target = 0.9 ``` `max_wal_size` 太小会导致 checkpoint 频繁、I/O 抖动;太大则拉长崩溃恢复和 WAL 回放时间。写入量大的系统常见取值在 4GB 到 16GB 之间——根据自己测得的 WAL 产生速率(`pg_stat_bgwriter`、`pg_stat_wal`)来定,而不是照抄表格。`wal_buffers` 默认 `-1`(按 shared\_buffers 自动计算),一般不需要显式设置。 ## 规划器与 I/O 成本参数 [#规划器与-io-成本参数] ```ini random_page_cost = 1.1 # SSD;默认值 4.0 是 HDD 时代的遗留 effective_io_concurrency = 200 # SSD/NVMe 能支撑远高于默认值 16 的并发 jit = off # JIT 利好长时间分析查询,OLTP 场景常常是负优化 ``` 默认的 `random_page_cost = 4.0` 假设随机读比顺序读贵四倍。在 SSD/NVMe 上这会让规划器避开本该使用的 index scan。1.1 是 SSD 上广泛使用的取值,但请用自己的查询跑 `EXPLAIN (ANALYZE, BUFFERS)` 确认,不要盲信任何固定数字。 `default_statistics_target` 默认 100,很少需要全局修改。个别倾斜严重的列按列调整:`ALTER TABLE t ALTER COLUMN c SET STATISTICS 1000;`,然后执行 `ANALYZE t;`。 `jit` 在 PostgreSQL 18 中默认开启。负载以短 OLTP 查询为主时,全局关掉是合理的起步选择;真正受益的报表查询可以在会话里单独开启。 ## 修改配置的三种方式 [#修改配置的三种方式] ```sql -- 1. ALTER SYSTEM:写入 postgresql.auto.conf,重启后仍然保留 ALTER SYSTEM SET work_mem = '32MB'; SELECT pg_reload_conf(); -- 2. 直接编辑 postgresql.conf,然后 -- sudo systemctl reload postgresql -- 3. 会话级或事务级,用于一次性任务 SET work_mem = '256MB'; SET LOCAL work_mem = '256MB'; -- 仅当前事务 ``` 持久化修改优先用 `ALTER SYSTEM`:改动可审计(`pg_settings` 里 `source` 会指向 auto.conf),也不用维护手改的配置文件。撤销用 `ALTER SYSTEM RESET name;`。 不是所有参数 reload 都生效,改之前先查 context: ```sql SELECT name, setting, context FROM pg_settings WHERE context IN ('postmaster', 'superuser-backend') ORDER BY name; -- context = 'postmaster' 需要完整重启 ``` 常见需要重启的参数:`shared_buffers`、`max_connections`、`shared_preload_libraries`。 ## 按机型规格的起步模板 [#按机型规格的起步模板] 以下模板面向 SSD/NVMe 存储上的专用 PostgreSQL 18 服务器,是配合上文测量流程验证的起步值,不是最终答案。 ### 4 vCPU / 16 GB [#4-vcpu--16-gb] ```ini shared_buffers = 4GB effective_cache_size = 12GB maintenance_work_mem = 1GB work_mem = 16MB max_connections = 100 wal_compression = on max_wal_size = 4GB random_page_cost = 1.1 effective_io_concurrency = 200 max_parallel_workers = 4 max_parallel_workers_per_gather = 2 ``` ### 8 vCPU / 32 GB [#8-vcpu--32-gb] ```ini shared_buffers = 8GB effective_cache_size = 24GB maintenance_work_mem = 2GB work_mem = 32MB max_connections = 200 wal_compression = on max_wal_size = 8GB random_page_cost = 1.1 effective_io_concurrency = 200 max_parallel_workers = 6 max_parallel_workers_per_gather = 4 ``` ### 16 vCPU / 64 GB [#16-vcpu--64-gb] ```ini shared_buffers = 16GB effective_cache_size = 48GB maintenance_work_mem = 4GB work_mem = 64MB max_connections = 500 # 前面需要挂连接池 wal_compression = on max_wal_size = 16GB random_page_cost = 1.1 effective_io_concurrency = 200 max_parallel_workers = 12 max_parallel_workers_per_gather = 6 ``` 想要第二份参考,可以用 [PGTune](https://pgtune.leopard.in.ua/) 按机器规格生成模板对照。 ## 生产环境建议开启的日志 [#生产环境建议开启的日志] ```ini log_min_duration_statement = '500ms' log_checkpoints = on log_connections = on log_disconnections = on log_lock_waits = on log_temp_files = 0 log_autovacuum_min_duration = 0 log_line_prefix = '%t [%p] %u@%d %a ' ``` `log_temp_files = 0` 记录每一次临时文件创建,是 `work_mem` 不够用的直接信号。`log_autovacuum_min_duration = 0` 让 autovacuum 行为可审计,配合 [autovacuum 与表膨胀](/docs/operations/autovacuum-bloat) 中的查询一起用。`log_connections` / `log_disconnections` 在连接池场景下开销很小,但短连接极多时日志量会很大,按需取舍。完整的监控体系见 [监控与日志](/docs/operations/monitoring-logging)。 ## 并行查询 [#并行查询] ```ini max_worker_processes = 8 # 后台 worker 总数(含逻辑复制等) max_parallel_workers = 6 # 并行查询可用的 worker 总数 max_parallel_workers_per_gather = 4 # 单条查询单个节点能用几个 min_parallel_table_scan_size = '8MB' ``` `max_worker_processes` 是所有后台 worker(并行查询、逻辑复制 apply worker、扩展)的全局预算,必须不小于 `max_parallel_workers` 加上复制和扩展的需求;它和 `max_parallel_workers` 一般都按 vCPU 数量来定。并行主要在大扫描上受益;`min_parallel_table_scan_size` 保持默认 8MB 或以上时,小型 OLTP 查询很少触发并行。 ## 查看当前配置 [#查看当前配置] ```sql -- 所有偏离默认值的参数及其来源 SELECT name, setting, unit, source FROM pg_settings WHERE source <> 'default' ORDER BY name; -- 当前值与启动时取值的差异(等待重启的改动) SELECT name, setting, boot_val, pending_restart FROM pg_settings WHERE setting <> boot_val OR pending_restart; ``` `pending_restart = true` 表示 `ALTER SYSTEM` 或配置文件的修改要等重启才生效——下结论说"改了没用"之前先查这一列。 ## AI prompt:按硬件生成配置 [#ai-prompt按硬件生成配置] 对 AI 的产出和本页模板采取同样的态度:它是待验证的假设,不是成品配置。 ## 相关页面 [#相关页面] * 连接池、备份与高可用选型 → [PostgreSQL 生产工具栈与高可用](/docs/operations/production-stack) * autovacuum 调参与膨胀排查 → [autovacuum 与表膨胀](/docs/operations/autovacuum-bloat) * 指标与日志管线 → [监控与日志](/docs/operations/monitoring-logging) --- # 容器环境下 PostgreSQL 的内存管理与 OOM 防治 Canonical URL: https://pg.edu.rich/docs/operations/container-memory-oom Last reviewed: 2026-08-06 典型的容器 OOM 现场是这样的:PostgreSQL 每隔几小时重启一次,`dmesg` 里有 `Out of memory: Killed process ... (postgres)`,而团队坚持 pod"内存很充足",因为容器里的 `free -m` 明明显示还有几个 GB 空闲。两个观察都是真的,也不矛盾:`free` 读的是 `/proc/meminfo`,它**没有按容器命名空间隔离**——容器里看到的是宿主机的内存,而不是内核真正执行的 cgroup limit。真正会杀掉你的那个 limit 在 cgroup 里,不在 `/proc/meminfo` 里。 ## 真正的 limit 在哪里读 [#真正的-limit-在哪里读] cgroup v2 主机上,生效的 limit 在 `memory.max`(当前记账在 `memory.current`);cgroup v1 上是 `memory.limit_in_bytes` 和 `memory.usage_in_bytes`: ```bash # cgroup v2(Kubernetes 1.31+ 及多数现代发行版) cat /sys/fs/cgroup/memory.max cat /sys/fs/cgroup/memory.current # cgroup v1 cat /sys/fs/cgroup/memory/memory.limit_in_bytes ``` 完整文件列表见内核文档 [cgroup v2](https://docs.kernel.org/admin-guide/cgroup-v2.html) 和 [cgroup v1 内存控制器](https://docs.kernel.org/admin-guide/cgroup-v1/memory.html)。任何在容器里从 `/proc/meminfo` 推导"可用内存"的工具或运维手册,量的都是错误的盒子。 ## 必须装进 limit 的内存预算 [#必须装进-limit-的内存预算] PostgreSQL 不会读 cgroup limit,它只按 `postgresql.conf` 给自己定尺寸,所以配置必须由容器 limit 推导,而不是由宿主机内存推导。必须装进去的部分(见 PostgreSQL 18 文档的 [Resource Consumption](https://www.postgresql.org/docs/18/runtime-config-resource.html)): * `shared_buffers`——启动时固定分配; * `work_mem` × 并发排序/哈希——它是**每个操作**而不是每个会话:一条带多个 hash join 的查询可以吃掉数倍 `work_mem`,所以 `max_connections × work_mem` 只是粗略上限; * `maintenance_work_mem`(或 `autovacuum_work_mem`)× 并发的 vacuum worker 和维护命令; * 每个后端的 `temp_buffers`、WAL buffer,以及每个后端进程数 MB 的连接固定开销; * 再加上容器里其他东西(sidecar、监控 agent)以及 PostgreSQL 依赖的 page cache。 起步值而非铁律:专用 PostgreSQL 容器里 `shared_buffers` 取 cgroup limit 的 25% 左右,给连接和 page cache 留出空间,之后凭测量再上调。完整的配置清单见 [服务器配置](/docs/operations/configuration)。 ## PG 认识的内存 vs cgroup 的记账 [#pg-认识的内存-vs-cgroup-的记账] 大多数"我们离 limit 还远着呢"的意外,来自两处记账差异: * **page cache 计入 cgroup。** 对堆文件和 WAL 文件的读写由内核缓冲,页面记到第一个触碰它的 cgroup 名下。一个进程匿名内存很少的容器,也可能因为其余部分都是 page cache 而顶在 limit 上——这本身正常且大多可回收,直到匿名内存突发分配(`work_mem` 尖峰、连接潮)到来,而缓存此时是脏的或正被引用。 * **RSS 随触碰过的页面增长。** 后端进程共享 `shared_buffers`,把各进程 RSS 加总的工具会把共享页重复计算。不要用 `ps` 输出的总和来定 limit,要用峰值负载下的 `memory.current`。 内核只在分配发生的瞬间没有可回收内存时才杀人,这就是为什么 OOM Kill 集中在负载尖峰而不是稳态。 ## OOM killer 打分与 oom\_score\_adj [#oom-killer-打分与-oom_score_adj] 当 cgroup(或宿主机)触顶时,内核按 `oom_score` 挑选牺牲者,可以通过 `/proc//oom_score_adj` 按进程调整(范围 `-1000` 到 `1000`;`-1000` 完全豁免——见 [proc 文件系统文档](https://docs.kernel.org/filesystems/proc.html))。运维上有两个事实必须知道: * 给 postmaster 设 `-1000` 是标准做法,但 **`oom_score_adj` 会被子进程继承**:之后 fork 出来的每个后端同样不可被杀,这恰恰和你想要的相反。如果保护了 postmaster,后端必须重置自己的分值——例如在 wrapper 脚本里重新设置,或者接受在 Kubernetes 上这一层由 pod 粒度控制。 * Kubernetes 不支持按容器设置 `oom_score_adj`;kubelet 根据 pod 的 QoS 等级赋值:Guaranteed pod 得 `-997`,BestEffort pod 得 `1000`,Burstable 落在中间。这写在 [node-pressure eviction](https://kubernetes.io/docs/concepts/scheduling-eviction/node-pressure-eviction/) 文档里。一个 pod 只放一个数据库实例,这层保护才有意义。 关于 `vm.overcommit_memory = 2` 的取舍要单独权衡:PostgreSQL 文档的 [Linux Memory Overcommit](https://www.postgresql.org/docs/18/kernel-resources.html) 一节解释了这笔交易——把它调到 `2` 可以降低 OOM killer 被触发的概率,代价是内存分配会更早直接失败。 ## K8s 的 requests、limits 与 QoS [#k8s-的-requestslimits-与-qos] 对数据库 pod,来自 [Kubernetes 内存资源文档](https://kubernetes.io/docs/tasks/configure-pod-container/assign-memory-resource/)和 [Pod QoS 等级](https://kubernetes.io/docs/concepts/workloads/pods/pod-qos/)的可辩护模式: * `requests.memory` 与 `limits.memory` 设为相等,让 pod 成为 **Guaranteed**:节点压力下它会排在 Burstable/BestEffort 之后才被驱逐,并且拿到有利的 `oom_score_adj`。 * `postgresql.conf` 从 limit 推导(见上一节),不从节点规格推导——pod 会被重新调度到更大的节点上,环境会悄悄变化。 * limit 里要给 page cache 和尖峰留余量:把 `shared_buffers + 连接数` 配到接近 limit 的 100%,等于保证下一次 `work_mem` 突发是致命的。 ## 诊断一次 OOM Kill [#诊断一次-oom-kill] 从内核证据入手,一路向内查到查询: ```bash # 内核侧的 kill 记录(宿主机上,或通过节点日志) dmesg -T | grep -i -E 'out of memory|oom_kill' journalctl -k | grep -i oom # cgroup v2:该 cgroup 每发生一次 kill,oom_kill 计数加一 cat /sys/fs/cgroup/memory.events # low 0 / high 0 / max / oom / oom_kill # cgroup v1 等价文件 cat /sys/fs/cgroup/memory/memory.oom_control ``` `memory.events` 里的 `max` 事件(触顶但靠回收挺过去、没有 kill)同样值得看——`max` 持续上涨而 `oom_kill` 为零,说明你已经长期贴在天花板上,被杀只是时间问题。 PostgreSQL 一侧,找 kill 发生前的内存消耗者。临时文件用量是 `work_mem` 溢出的经典签名;它按 database 统计,不按会话: ```sql SELECT datname, temp_files, pg_size_pretty(temp_bytes) AS temp_written FROM pg_stat_database ORDER BY temp_bytes DESC; ``` `temp_bytes` 是累计值,要比较事故窗口前后的增量。归因到具体语句要靠日志:设 `log_temp_files = 0`(或一个阈值),让每次落盘临时文件都记下肇事查询,再与 kill 时刻 `pg_stat_activity` 里的活跃会话对照。完整的监控管线见 [监控与日志](/docs/operations/monitoring-logging)。 ## 防治清单 [#防治清单] * 一个 pod 一个 postmaster;Guaranteed QoS(`requests == limits`)。 * `shared_buffers`、`work_mem`、`maintenance_work_mem`、`max_connections` 从 cgroup limit 推导而非宿主机内存——limit 变更后重新推导。 * 数据库前面放连接池,让 `max_connections × work_mem` 有界。 * `log_temp_files = 0`(或阈值),并对 `temp_bytes` 增长告警。 * 对 `memory.events` 的 `oom_kill` 增量和 `max` 计数增长告警。 * 容量决策永远不信容器里的 `free`、`top`、`htop`。 不重新推导 `postgresql.conf` 就调大 limit,只是把崩溃往后挪。要么是配置对 limit 来说太大,要么是某个查询级消费者(`work_mem` 尖峰、连接潮)没有上界——先搞清楚是哪一种,再花钱买内存。 把输出当作假设:先在测试 pod 应用,在峰值负载下观察 `memory.current` 和 `memory.events`,确认后再动生产。 ## 相关页面 [#相关页面] * `postgresql.conf` 取值推导 → [服务器配置](/docs/operations/configuration) * 临时文件与语句级证据 → [监控与日志](/docs/operations/monitoring-logging) * 当元数据本身吃掉内存时 → [表太多也是病](/docs/operations/too-many-tables) --- # 在线启停 data checksums Canonical URL: https://pg.edu.rich/docs/operations/data-checksums Last reviewed: 2026-08-06 截至 2026-08,PostgreSQL 19 尚未正式发布。下文函数签名、状态与视图列以 [PostgreSQL 19 文档](https://www.postgresql.org/docs/19/checksums.html)为准;用于生产前请对照正式 release notes 复核。 Data checksums 在每个数据页里存一个校验值,页面写出时计算、读入时验证,是发现存储和文件系统损坏的主要手段。从 PostgreSQL 18 起 `initdb` 默认为新集群开启。但在更老 major 版本上初始化的集群往往至今没开——这正是本页要解决的场景。 ## PostgreSQL 18 及之前:pg\_checksums 必须停机 [#postgresql-18-及之前pg_checksums-必须停机] 到 PostgreSQL 18 为止,改变既有集群的 checksum 状态只能用 [`pg_checksums`](https://www.postgresql.org/docs/19/app-pgchecksums.html),而它要求服务器**干净关闭**。启用 checksum 需要原地重写所有关系块,TB 级集群的维护窗口以小时计,且操作中途严禁启动集群。对要求始终在线的系统,"补开 checksum" 几乎排不进日程。 ## 在线启用 [#在线启用] PostgreSQL 19 新增了 SQL 函数,在集群正常运行、客户端照常访问的情况下切换 checksum,见官方文档 [28.2 节 Data Checksums](https://www.postgresql.org/docs/19/checksums.html): ```sql -- 查看当前状态:off / inprogress-on / on / inprogress-off SHOW data_checksums; -- 开始启用;cost_delay / cost_limit 用于节流 SELECT pg_enable_data_checksums(cost_delay => 10, cost_limit => 1000); ``` 调用之后发生什么: 1. 集群状态变为 `inprogress-on`。此后页面写出时*计算* checksum,但读入时暂不*验证*。 2. 一个 launcher 为每个 database 启动后台 worker,遍历所有关系、把每个页标记为脏,使其带 checksum 重写。 3. 所有 database 处理完毕后,状态自动切换为 `on`,读时验证开始。 需要知道的前提与卡点: * 该过程占用两个后台 worker 槽位——确认 `max_worker_processes` 有余量。 * 开始前要等待所有已打开的事务结束;对每个 database,还要等启用那一刻之前存在的临时表全部被删除。使用长生命周期临时表的应用可能无限期卡住流程,必要时需要终止对应连接。 * 如果集群在 `inprogress-on` 期间停止,**不会断点续做**:重启后重新执行 `pg_enable_data_checksums()`,重写从头开始。安排在没有重启、没有 failover 的窗口。 ## I/O 代价与进度观测 [#io-代价与进度观测] 启用 checksum 要重写集群里的每一个页——脏页刷盘加上重写产生的 WAL,I/O 冲击是实打实的。`pg_enable_data_checksums()` 的 `cost_delay` / `cost_limit` 参数沿用 vacuum 那套基于成本的节流模型;生产环境建议从保守值起步,同时观察延迟。 进度通过 `pg_stat_progress_data_checksums` 观测:launcher 一行(跟踪 `databases_total` / `databases_done`),每个 worker 一行(跟踪 `relations_total` / `relations_done` 以及当前关系的 `blocks_total` / `blocks_done`): ```sql SELECT pid, datname, phase, databases_total, databases_done, relations_total, relations_done, blocks_total, blocks_done FROM pg_stat_progress_data_checksums; ``` `phase` 列区分 `enabling`、`disabling`、`waiting on barrier`(等待各后端确认状态变更)和 `waiting on temporary tables`——后两个解释了大多数"看似卡住"的情况。这些进程也会出现在 `pg_stat_activity` 里,`backend_type` 为 `datachecksums launcher` / `datachecksums worker`。 文档中关于复制的一个提醒:standby 从 WAL 流中收到 checksum 状态变更时会强制做一次 restartpoint,完成前阻塞 redo,可能造成复制延迟——同步 standby 上还会反过来阻塞 primary。开始前调低 `max_wal_size` 可以缩短这次 restartpoint。 ## 在线禁用 [#在线禁用] ```sql SELECT pg_disable_data_checksums(); ``` 状态先变为 `inprogress-off`(仍写 checksum、不再验证),所有后端确认后落定到 `off`。禁用不重写任何页面,没有 I/O 冲击,但仍需 checkpoint。在启用过程中执行禁用会中止启用。如果集群在 `inprogress-off` 期间停止,重启后状态即为 `off`。 ## 硬件加速的校验和计算 [#硬件加速的校验和计算] 数据页里存的校验和用的是 PostgreSQL 自有的、基于 FNV-1a 的算法,而不是 CRC;x86-64 上会在运行时自动选用 AVX2 向量化实现。CRC-32C 保护的是 WAL 记录:运行时在 x86 上自动分发到 SSE4.2 或 AVX-512 指令,在 ARMv8 上使用 CRC32/PMULL 扩展(实现背景见 [digoal 的 commit 解读](https://github.com/digoal/blog/blob/master/202604/20260406_06.md);WAL 使用 CRC-32C 见官方文档 [28.1 节 Reliability](https://www.postgresql.org/docs/19/wal-reliability.html))。现代 CPU 上,日常保持 checksum 开启的 CPU 开销很小;主要成本是启用时的一次性重写,而不是常态运行。 ## 与备份校验的联动 [#与备份校验的联动] Checksum 与备份工具链在两个不同层面配合: * [`pg_basebackup`](https://www.postgresql.org/docs/19/app-pgbasebackup.html) 读取集群时默认逐页验证 checksum(除非指定 `--no-verify-checksums`);校验失败会以非零状态退出,并计入 `pg_stat_database.checksum_failures`。开了 checksum,每次基础备份就顺带完成一次全库损坏扫描。 * [`pg_verifybackup`](https://www.postgresql.org/docs/19/app-pgverifybackup.html) 依据 `backup_manifest` 里的文件级哈希校验已完成的备份。它抓的是备份存储、传输过程中引入的损坏——与守护在线集群的页 checksum 是不同的失效域。 两者都应纳入 [备份、恢复与 PITR](/docs/operations/backup-recovery) 中的验证流程;`checksum_failures` 应接入 [监控与日志](/docs/operations/monitoring-logging) 里的看板。 ## 什么时候值得开 [#什么时候值得开] * 在 PostgreSQL 18 修改默认值之前初始化、一直没开 checksum 的集群——现在没有停机这个借口了。 * 有合规或完整性要求,必须能发现静默存储损坏的系统。 * 希望 `pg_basebackup` 的内建验证兼作定期损坏巡检的系统。 可以暂缓的理由很窄:即用即弃的集群,或者存储栈已端到端校验(如 ZFS)且风险偏好较高的情况。注意只有 PostgreSQL 层的 checksum 才会被 PostgreSQL 自己的工具验证。PostgreSQL 19 的更多变化见[版本总览](/docs/postgresql-19)。 ## AI 提示词:规划 checksum 启用 [#ai-提示词规划-checksum-启用] --- # 生产运维入口 Canonical URL: https://pg.edu.rich/docs/operations Last reviewed: 2026-08-02 ## 先定义目标 [#先定义目标] | 目标 | 要回答的问题 | | --- | ----------------------- | | RPO | 最多允许丢失多少数据? | | RTO | 故障后多久必须恢复服务? | | 容量 | 峰值连接、数据增长、WAL 与备份增长是多少? | | 可用性 | 哪些故障自动切换,哪些必须人工判断? | | 安全 | 谁能连接、读哪些数据、做哪些变更? | 没有目标的“高可用”和“做了备份”不可验证。 ## 每日最小检查 [#每日最小检查] ```sql SELECT now(), version(); SELECT state, count(*) FROM pg_stat_activity GROUP BY state ORDER BY state; SELECT datname, age(datfrozenxid) FROM pg_database ORDER BY age(datfrozenxid) DESC; SELECT num_timed, num_requested, num_done, buffers_written, write_time, sync_time FROM pg_stat_checkpointer; SELECT buffers_clean, maxwritten_clean, buffers_alloc FROM pg_stat_bgwriter; ``` PostgreSQL 17 起,检查点统计位于 `pg_stat_checkpointer`,后台写进程统计仍在 `pg_stat_bgwriter`。字段定义见 [PostgreSQL 18 累积统计视图](https://www.postgresql.org/docs/18/monitoring-stats.html#MONITORING-PG-STAT-CHECKPOINTER-VIEW)。 还应监控磁盘空间、WAL 生成/归档、复制 lag、备份状态、事务时长、锁等待、查询延迟、autovacuum 活动和连接池饱和度。阈值必须来自本系统基线。 ## 变更纪律 [#变更纪律] 1. 在类似规模数据上测量锁与耗时。 2. 写明回退路径和不可逆点。 3. 设置 `lock_timeout`,避免迁移无限等待后突然获取大锁。 4. 观察执行期间的锁、WAL、复制延迟和错误率。 5. 用查询或业务指标验证结果。 生产 DDL 不是“执行成功就结束”,而是一个可观察、可中断、可验证的发布。 --- # PostgreSQL 瞬间克隆:Copy-on-Write Canonical URL: https://pg.edu.rich/docs/operations/instant-clone Last reviewed: 2026-08-06 过去复制一个 200 GB 的数据库,就意味着真的拷贝 200 GB。写时复制(Copy-on-Write,CoW)改变了这笔账:文件系统创建第二棵目录树,与源共享同一份物理数据块,只有之后任意一侧写入的页面才占用新空间。克隆在几秒内出现,初始磁盘开销接近零。PostgreSQL 18 让单库级克隆成为一等操作,而任何支持 reflink 的文件系统都能实现整实例级克隆。 ## 为什么需要瞬间克隆 [#为什么需要瞬间克隆] * **每个 PR 一个库。** CI 在接近生产规模的完整数据副本上跑迁移和集成测试,用完即弃。 * **每个 Agent 任务一个沙盒。** Agent 可以随意改、随便删、随时重来,不碰共享状态——与经由 [MCP](/docs/ai/mcp) 驱动的 Agent 工作流天然契合。 * **迁移与升级演练。** 在生产数据形态上预演 schema 迁移、量出耗时,然后删掉演练副本。 * **事故排查。** 给工程师一份故障现场的冻结副本,而不是在生产库上反复试探。 共同要求是副本便宜到可以随手丢弃。`pg_dump` 加恢复在数据量上来后过不了这道门槛:几分钟到几小时、期间双倍存储、还要重建全部索引。 ## PostgreSQL 18 的原生基础 [#postgresql-18-的原生基础] CoW 复制依赖 **reflink**:源文件与目标文件共享数据块,直到某一方写入才分裂。文件系统支持是硬前提——以 `reflink=1` 格式化的 XFS(现代发行版的默认)、Btrfs、APFS 都支持;OpenZFS 在 2.3 加入了块克隆。一行命令即可确认你的挂载点是否够格: ```bash cp --reflink=always /srv/pg/somefile /tmp/reflink-test && rm /tmp/reflink-test # 不支持的文件系统会报 "Operation not supported" ``` 在此之上,PostgreSQL 18 新增了服务器参数 [`file_copy_method`](https://www.postgresql.org/docs/18/runtime-config-resource.html#GUC-FILE-COPY-METHOD)(默认 `copy`,设为 `clone` 启用)。设为 `clone` 后,[`CREATE DATABASE ... STRATEGY = FILE_COPY`](https://www.postgresql.org/docs/18/sql-createdatabase.html) 和 `ALTER DATABASE ... SET TABLESPACE` 的文件复制会走 Linux 的 `copy_file_range()` 或 macOS 的 `copyfile()`,在支持的文件系统上即成为 reflink;不支持的系统上则直接报错。 注意边界:PostgreSQL 18 里 `initdb` 和 `pg_basebackup` 仍是逐块复制,没有 reflink 选项。实例级的 CoW 来自文件系统(reflink 复制或快照),不来自 PostgreSQL 的开关。 ## 用 reflink 克隆整个实例 [#用-reflink-克隆整个实例] 一致性路径:干净地停掉源库,reflink 复制,把副本作为独立实例拉起。 ```bash # 1. 干净停库——fast 模式完成 checkpoint 并等待客户端断开 pg_ctl -D /srv/pg/18/main stop -m fast # 2. CoW 复制整个数据目录——秒级完成,额外空间接近零 cp -a --reflink=always /srv/pg/18/main /srv/pg/18/clone-pr482 # 3. 重新拉起源库 pg_ctl -D /srv/pg/18/main start # 4. 改掉必须不同的配置,把克隆作为独立实例启动 echo 'port = 55432' >> /srv/pg/18/clone-pr482/postgresql.conf pg_ctl -D /srv/pg/18/clone-pr482 start ``` 源库与克隆之间必须错开的:`port`(以及 `listen_addresses`/socket 目录的假设)、显式设置的 `data_directory`,以及 `postgresql.auto.conf` 里按实例钉死的资源配置。如果集群使用了额外表空间,这些目录在数据目录之外,也要一并 reflink 复制,并把 `pg_tblspc` 下的符号链接指到新位置。 如果源库不能停,用**文件系统原子快照**代替普通复制——Btrfs/ZFS/LVM 快照在瞬间完成,得到的是崩溃一致(crash-consistent)的镜像,PostgreSQL 首次启动时像断电恢复一样重放 WAL。对**运行中**的数据目录做普通 `cp`(无论加不加 reflink)既不是原子的也不一致:复制遍历目录的过程中文件在变,结果可能起不来,或者更糟——带着细微损坏起来。不要把它当捷径。 ## 用 CREATE DATABASE 克隆单个数据库 [#用-create-database-克隆单个数据库] 实例内部,[`CREATE DATABASE ... TEMPLATE`](https://www.postgresql.org/docs/18/sql-createdatabase.html) 在文件层面复制一个数据库。`STRATEGY` 选项(PostgreSQL 15+)加上 PostgreSQL 18 的 `file_copy_method = clone`,把它变成 CoW 克隆: ```sql SET file_copy_method = 'clone'; CREATE DATABASE app_pr482 TEMPLATE app STRATEGY = FILE_COPY; ``` 按官方文档的约束: * 复制期间模板库不能有任何其他连接;有则 `CREATE DATABASE` 直接失败,且复制完成前新连接会被挡在模板库外。 * `FILE_COPY` 会在复制前后各强制一次 checkpoint,繁忙系统上可能有感知。 * 数据库级配置(`ALTER DATABASE ... SET`)和数据库级 `GRANT` 不会被复制。 * 默认策略 `WAL_LOG` 逐块经 WAL 复制——更慢、占满量空间,但不依赖文件系统的 reflink 支持。 产物是同一实例上的一个普通数据库:同一批角色、同一个端口、schema 与数据互相隔离,未修改的数据块与模板库共享,直到某一方写入。 ## 方案对比 [#方案对比] | 方法 | 粒度 | 一致性前提 | 时间与空间 | | --------------------------------------------------------------------------- | ----- | ---------------------------- | ---------------- | | `pg_dump` / `pg_restore` | 单库,逻辑 | 在线;自带一致性快照 | 全量拷贝;量大时慢;索引重建 | | `CREATE DATABASE ... TEMPLATE`(默认 `WAL_LOG`) | 单库 | 模板库无连接 | 实例内全量物理拷贝 | | `CREATE DATABASE ... STRATEGY = FILE_COPY` + `file_copy_method = clone` | 单库 | 模板库无连接;文件系统支持 reflink(PG 18) | 近瞬时;数据块 CoW 共享 | | reflink 复制数据目录 | 整个集群 | 源库已停,或原子快照 | 近瞬时;数据块 CoW 共享 | | [`pg_basebackup`](https://www.postgresql.org/docs/18/app-pgbasebackup.html) | 整个集群 | 在线 | 全量拷贝;PG 18 无 CoW | 当克隆需要跨版本、跨平台、跨实例搬运时选 `pg_dump`——逻辑副本的可移植性是文件级复制给不了的。 ## 与 Neon 等分支平台的关系 [#与-neon-等分支平台的关系] 托管分支平台产品化的正是同一个想法。[Neon](/docs/cloud/neon) 在存储层实现分支:分支是某一时间点数据的写时复制分叉,API 调用几秒建成,按增量计费。本页的自建方案给你同样的原语而不绑定平台——代价是几条 shell 命令换掉了 API 和计费模型,外加下面这份运维边界清单。 ## 风险清单 [#风险清单] * **克隆不是备份。** 源库与克隆共享物理数据块:存储层的一次损坏会同时命中所有克隆;克隆存在期间删掉源库也释放不了任何空间。独立的真实备份照旧是刚需——见 [备份、恢复与 PITR](/docs/operations/backup-recovery)。 * **WAL 一致性。** 只有干净停机后的复制或原子快照能产出安全的克隆。对运行中的实例做普通复制,得到的目录可能过不了崩溃恢复,或者带着撕裂的状态启动。 * **磁盘统计会说谎。** `df` 和 PostgreSQL 的大小函数都不反映共享数据块;按全量副本校准的容量告警会误报。监控要看文件系统的实际分配量。 * **限同一文件系统。** reflink 不能跨挂载点、跨主机——克隆必须与源在同一个文件系统上。 * **共享随写入消退。** 每次对共享数据块的首次写入都会在该侧分配新块。活得久、写得多的克隆会收敛回全量大小;CI 里的克隆寿命要据此设计。 * **模板库锁定。** 单库克隆期间模板库连接被冻结;CoW 文件系统上是几秒,普通拷贝则随数据量线性增长。 写时复制让克隆便宜到可以随用随弃——就把它当作天生短命的东西。任何丢不起的数据,必须以独立副本的形式放在另一份存储上,而不是同一批数据块上再多一条 reflink。 --- # PostgreSQL 监控:SQL、指标与日志 Canonical URL: https://pg.edu.rich/docs/operations/monitoring-logging Last reviewed: 2026-08-06 PostgreSQL 可观测性至少需要三层证据:**SQL 统计解释资源花在哪里,指标说明系统何时偏离基线,日志保留错误与事件上下文**。只装一个 dashboard 不能替代这三层。 ## 1. 用 pg\_stat\_statements 找工作负载热点 [#1-用-pg_stat_statements-找工作负载热点] `pg_stat_statements` 是 PostgreSQL 官方扩展。它需要加入 `shared_preload_libraries`,通常要重启实例,然后在需要统计的 database 中创建扩展: ```ini shared_preload_libraries = 'pg_stat_statements' compute_query_id = auto ``` ```sql CREATE EXTENSION IF NOT EXISTS pg_stat_statements; SELECT queryid, calls, total_exec_time, mean_exec_time, rows, shared_blks_hit, shared_blks_read, left(query, 160) AS query FROM pg_stat_statements ORDER BY total_exec_time DESC LIMIT 20; ``` 按总耗时、平均耗时、调用次数、返回行数和 I/O 分别看排行;单次最慢与累计消耗最大不是同一个问题。统计会被重置,部署、故障和参数变更时应记录采样窗口。字段定义以 [pg\_stat\_statements 官方文档](https://www.postgresql.org/docs/current/pgstatstatements.html) 为准。 扩展会规范化常量,但日志、DDL、动态 SQL 和应用注释仍可能泄露标识符或业务数据。限制统计视图与日志读取权限,并为采集、保留和脱敏设定规则。 ## 2. 使用 JSON 日志保留事件上下文 [#2-使用-json-日志保留事件上下文] `jsonlog` 便于可靠解析时间、SQLSTATE、backend、database、用户、application name 和错误上下文: ```ini logging_collector = on log_destination = 'jsonlog' log_min_duration_statement = '500ms' # 示例值,应按负载基线调整 log_lock_waits = on deadlock_timeout = '1s' ``` 不要把示例阈值直接复制到所有环境。过低会产生大量 I/O 与敏感查询文本,过高会漏掉高频中等耗时 SQL。优先用 `pg_stat_statements` 发现累计热点,用日志解释错误、锁等待、检查点、autovacuum 和特定慢请求。参数语义见 [Error Reporting and Logging](https://www.postgresql.org/docs/current/runtime-config-logging.html)。 [pgBadger](https://github.com/darold/pgbadger) 可以分析 PostgreSQL 原生日志与 `jsonlog`,生成查询、连接、错误、锁、检查点和 autovacuum 报告。先确保日志格式稳定、轮转可靠、时区一致,再把它加入离线分析流程。 ## 3. 指标、Prometheus 与 Grafana [#3-指标prometheus-与-grafana] [postgres\_exporter](https://github.com/prometheus-community/postgres_exporter) 适合已有 Prometheus/Grafana 的团队。采集角色优先使用 `pg_monitor` 或必要的只读统计权限,不要给 exporter 超级用户: ```sql CREATE ROLE metrics LOGIN; GRANT pg_monitor TO metrics; ``` 在目标版本上核对 collector 和权限。其 multi-target 模式仍被上游标为 Beta,自定义 `extend.query-path` 已 deprecated;新采集需求应优先使用内置 collector 或单独的通用 SQL exporter,而不是积累不可维护的查询文件。 ### 最小信号集 [#最小信号集] | 领域 | 信号 | 需要一起看的上下文 | | ------ | ---------------------------------------- | ------------------------ | | 连接 | 使用量、等待、连接池队列 | pool mode、应用实例数、保留连接 | | 查询 | latency、calls、rows、I/O | deploy、plan 变化、参数分布 | | 事务 | 长事务、idle in transaction、冲突 | owner、重试能力、vacuum 影响 | | 锁 | 等待时长、阻塞链、deadlock | DDL、批任务、业务事务 | | WAL/复制 | 生成速率、archive failure、lag、slot retention | RPO、网络、磁盘余量 | | 维护 | dead tuples、freeze age、vacuum/analyze 进度 | 表写入率、autovacuum 参数 | | 存储 | 数据/WAL/临时文件增长、I/O latency | 容量预测、checkpoint、查询 spill | | 恢复 | 最近成功备份、可恢复时间、实测 RTO | repository、密钥、恢复演练 | 阈值应来自正常时段和峰值时段的基线,告警应指向可执行的诊断路径。复制 lag 的字节数、时间和 replay 状态含义不同,不能只设一个全局阈值。 ### 告警阈值参考起点 [#告警阈值参考起点] 每个阈值都要用自己的基线校准;下面是常见的起点,不是通用值: | 指标 | 警告 | 严重 | | --------------------------- | ---- | ----- | | 连接数(占 `max_connections` 比例) | 70% | 90% | | 复制延迟(秒) | 10 | 60 | | 复制延迟(字节) | 1 GB | 10 GB | | database 事务 ID age | 10 亿 | 15 亿 | | 磁盘使用率 | 75% | 90% | | 每秒 deadlock 数 | 0.1 | 1 | | 单表 autovacuum 持续超过 | 2 小时 | 6 小时 | | dead tuples(占表比例) | 20% | 40% | | 缓冲池命中率低于 | 95% | 90% | 警告级不应触发值班电话,严重级必须有对应的 runbook;触发了却没有对应动作的告警应调参或删除。 ## 4. 现成诊断 SQL 模板 [#4-现成诊断-sql-模板] 事故处理时随手可用的一组模板,均为只读,可在生产直接执行。 ```sql -- 连接数,按状态分 SELECT count(*) AS total, count(*) FILTER (WHERE state = 'active') AS active, count(*) FILTER (WHERE state = 'idle in transaction') AS idle_in_transaction FROM pg_stat_activity; -- 运行超过 5 秒的查询 SELECT pid, usename, application_name, now() - query_start AS duration, wait_event, left(query, 200) AS query FROM pg_stat_activity WHERE state = 'active' AND now() - query_start > interval '5 seconds' ORDER BY query_start; -- 阻塞链:谁在等谁 SELECT blocked.pid AS blocked_pid, blocking.pid AS blocking_pid, left(blocked.query, 120) AS blocked_query, left(blocking.query, 120) AS blocking_query FROM pg_stat_activity blocked JOIN pg_stat_activity blocking ON blocking.pid = ANY (pg_blocking_pids(blocked.pid)); -- 主库视角的复制延迟 SELECT application_name, state, sync_state, pg_wal_lsn_diff(pg_current_wal_lsn(), replay_lsn) AS lag_bytes, replay_lag FROM pg_stat_replication; -- 复制槽保留量(无限增长会撑爆主库磁盘) SELECT slot_name, active, pg_wal_lsn_diff(pg_current_wal_lsn(), restart_lsn) AS retained_bytes FROM pg_replication_slots; ``` `SELECT pg_cancel_backend(pid);` 只取消当前查询;`SELECT pg_terminate_backend(pid);` 直接断开连接。优先 cancel,并先确认该会话可以放弃。 ## 工具采用顺序 [#工具采用顺序] 1. 所有生产实例先启用并治理 `pg_stat_statements`。 2. 输出可解析日志,并把 SQLSTATE、锁等待、归档失败和 autovacuum 纳入采集。 3. 已有 Prometheus 时接入 postgres\_exporter 与 PostgreSQL 专用 dashboard。 4. 需要日志趋势报告时加入 pgBadger;管理多实例可评估 [pgwatch](https://github.com/cybertec-postgresql/pgwatch)。 5. 深度 workload 分析可评估 [PoWA](https://github.com/powa-team/powa);临时排障可使用 [pg\_activity](https://github.com/dalibo/pg_activity)。 工具越多不等于盲区越少。先统一 instance、database、role、application、query ID、时间窗口和变更事件这些关联维度。 --- # OS 升级与索引的静默损坏 Canonical URL: https://pg.edu.rich/docs/operations/os-upgrade-index-corruption Last reviewed: 2026-08-06 `text` 列上的 B-tree 索引按插入那一刻生效的排序规则(collation)排列键值。libc collation 的排序规则来自操作系统的 C 库。一次 OS 升级如果带了新的 glibc(或新的 ICU 库),规则就可能变化——而 PostgreSQL 不会拿新规则重新校验既有索引。查询照常使用这些索引,于是悄悄返回错误结果:范围扫描和前缀匹配丢行、`ORDER BY ... LIMIT` 输出不对、唯一性检查漏掉重复值。全程没有任何报错。 ## OS 升级如何损坏索引 [#os-升级如何损坏索引] libc collation 的字符串比较走操作系统的接口(`strcoll_l` 等)。glibc 改变某个 locale 的排序规则后,升级之后插入的键按新规则放置,升级前写入的键仍停在旧规则的位置上。索引在结构上完全正常——页面链接、校验和、元组格式都没问题——但它不再服从一套一致的顺序。任何依赖顺序的索引扫描都可能提前停止、跳过条目或按错误顺序返回。 最著名的触发点是 glibc 2.28(2018 年),它把大量 locale 对齐到了新的通用排序模板。跨代发行版升级——例如 RHEL/CentOS 7 → 8、Debian 9 → 10——是经典的踩坑场景,但任何发行版升级都可能携带变化的 locale 数据。ICU collation 暴露的是同一个问题,只是触发源换成了 ICU 库升级,与 glibc 无关。 ## 哪些索引有风险 [#哪些索引有风险] * `text`、`varchar`、`char` 及其 domain 上的 B-tree 索引,当列 collation 由 **libc** 提供且不是 `C`/`POSIX` 时——包括数据库本身使用 libc provider 时、使用数据库**默认** collation 的索引。 * 这些列上的唯一约束和主键,因为底层就是这样的索引。 * ICU provider 的 collation 受 ICU 库版本变化影响,与 glibc 无关。 不受影响的:`C` 和 `POSIX` collation(按字节序比较,到处稳定)、builtin provider 的 locale 如 `C.UTF-8`(设计上不可变,PostgreSQL 17+)、hash 索引(无顺序概念),以及 `integer`、`bigint`、`uuid`、`timestamptz` 等不可排序类型上的索引。 provider 语义见 PostgreSQL 18 官方文档 [Collation Support](https://www.postgresql.org/docs/18/collation.html)。 ## 升级前先做索引盘点 [#升级前先做索引盘点] 任何 OS 升级之前,先列出所有依赖 libc 排序的索引并存档——这是事后比对的基线: ```sql SELECT DISTINCT i.indexrelid::regclass AS index_name, i.indrelid::regclass AS table_name, c.collname AS collation, c.collprovider AS provider -- c = libc, d = default, i = icu, b = builtin FROM pg_index i JOIN LATERAL unnest(i.indcollation) AS u(coll_oid) ON true JOIN pg_collation c ON c.oid = u.coll_oid WHERE c.collprovider IN ('c', 'd') AND c.collname NOT IN ('C', 'POSIX') ORDER BY 2, 1; ``` `default` collation(provider 为 `d`)继承数据库 locale,因此还要确认每个数据库实际使用什么: ```sql SELECT datname, datcollate, datlocprovider, datcollversion FROM pg_database; ``` 如果 `datlocprovider` 是 `c` 且 `datcollate` 不是 `C`/`POSIX`,该库所有默认 collation 的文本索引都依赖 libc。上面的查询覆盖索引键列;表达式索引的表达式里可能嵌入额外 collation,需要人工过一遍定义。 ## 升级后的检测 [#升级后的检测] ### 比对已记录的排序规则版本 [#比对已记录的排序规则版本] 从 PostgreSQL 10 起,系统目录在 [`pg_collation.collversion`](https://www.postgresql.org/docs/18/catalog-pg-collation.html) 里记录每个 collation 的 provider 版本;PostgreSQL 15 又扩展到数据库默认 collation(`pg_database.datcollversion`,以及配套的 `ALTER DATABASE ... REFRESH COLLATION VERSION`)。当使用到的对象其记录版本与 OS 报告的版本不一致时,会话会对每个 collation 发出一次版本不匹配的警告。也可以不依赖警告、直接比对: ```sql SELECT collname, collprovider, collversion, pg_collation_actual_version(oid) AS os_version FROM pg_collation WHERE collversion IS DISTINCT FROM pg_collation_actual_version(oid); ``` [`pg_collation_actual_version`](https://www.postgresql.org/docs/18/functions-info.html) 向操作系统查询当前安装的版本。这条查询返回的每一行,都是一个行为可能已经在你脚下变过的 collation。PostgreSQL 9.6 及更早版本没有这套版本追踪设施——没有警告、无可比对,检测完全靠下面的结构检查,或者干脆按计划重建。 ### 用 amcheck 验证索引结构 [#用-amcheck-验证索引结构] [`amcheck`](https://www.postgresql.org/docs/18/amcheck.html) 扩展用**当前**生效的比较规则重新评估 B-tree 的顺序,恰好对准这种故障模式:一个在旧规则下内部一致的索引,升级后可能通不过验证,因为验证期待的是新顺序。 ```sql CREATE EXTENSION IF NOT EXISTS amcheck; SELECT bt_index_parent_check('app.orders_customer_name_idx', heapallindexed => true); ``` 索引正常时函数不返回行,发现不一致则抛错。`bt_index_parent_check` 是更严格的变体(额外检查父子页关系);`bt_index_check` 更轻。`bt_index_parent_check` 在索引及其表上持有 `ShareLock`,阻塞并发的 `INSERT`/`UPDATE`/`DELETE`,安排在维护窗口执行;`bt_index_check` 只持有 `AccessShareLock`——与普通 `SELECT` 同级——不阻塞写入。 两个局限要记牢:amcheck 只在“存储的键序与当前规则冲突”时才报告——如果排序规则的变化恰好没有重排任何既有键,检查会静默通过,所以通过不等于安全。反过来,在一次动过 glibc 或 ICU 的 OS 升级之后,文本索引上的 amcheck 报错几乎必然意味着“重建”,而不是“硬件故障”。 ## 重建受影响的索引 [#重建受影响的索引] 不阻塞应用的重建方式: ```sql REINDEX INDEX CONCURRENTLY app.orders_customer_name_idx; -- 或者一次重建整张表的所有索引: REINDEX TABLE CONCURRENTLY app.orders; ``` `REINDEX CONCURRENTLY` 期间表保持可读写,代价是比普通的 `REINDEX` 慢,且不能在事务块内执行。失败或中断的运行可能留下 `INVALID` 索引——在 `pg_index` 里找(`NOT indisvalid`),重试前先删掉。全实例级的事故按表逐个重建,并按索引大小排序,把大表的窗口期安排得有计划。 重建完成后,刷新记录的版本号以消除不匹配警告: ```sql ALTER COLLATION "de_DE" REFRESH VERSION; ALTER DATABASE app REFRESH COLLATION VERSION; ``` 只能在依赖索引重建**之后**刷新版本——先刷新会把警告消掉,而损坏还留在原地。 没有报错,校验和不会触发,常规监控什么也看不到。症状是 OS 升级几天甚至几周之后,用户开始报"查不到数据"。把盘点查询和 amcheck 检查写进 OS 升级的 runbook,而不是留到事后复盘。 ## 预防 [#预防] * **新数据库优先用 ICU 或 builtin collation。** `CREATE DATABASE ... LOCALE_PROVIDER = icu ICU_LOCALE = 'de-DE'` 把排序规则钉在显式带版本号的 ICU 规则集上,而不是发行版自带的任意 glibc;不需要自然语言排序时,builtin provider 的 `C.UTF-8`(PostgreSQL 17+)在 OS 升级面前完全不可变。已有的 libc 数据库可以用 `CREATE COLLATION` 加并发重建索引,把个别列迁到 ICU collation。 * **OS 升级与 PostgreSQL 大版本升级分开做。** `pg_upgrade` 只复制数据文件、不重建索引,把两类升级塞进同一个窗口等于风险加倍,出了异常也难以归因。分两个窗口做,中间穿插版本比对和 amcheck 检查;升级路径见 [复制、故障切换与升级](/docs/operations/replication-upgrades)。 * **每次 glibc/ICU 升级后复查。** 版本比对查询成本很低,每个 OS 补丁周期后跑一次,变了什么就重建什么。 * **清楚查询依赖什么。** 损坏通过索引扫描和有序输出暴露——[索引与 EXPLAIN](/docs/core/indexes-explain) 讲如何看哪些执行计划依赖受影响的索引。 --- # PostgreSQL 19 并行 autovacuum 与评分调度 Canonical URL: https://pg.edu.rich/docs/operations/parallel-autovacuum Last reviewed: 2026-08-06 截至 2026-08,PostgreSQL 19 尚未正式发布。下文参数名、视图列与默认值以 [PostgreSQL 19 文档](https://www.postgresql.org/docs/19/runtime-config-autovacuum.html)为准;用于生产前请对照正式 release notes 复核。 ## PostgreSQL 18 及之前:串行的 autovacuum [#postgresql-18-及之前串行的-autovacuum] 到 PostgreSQL 18 为止,autovacuum 有两个长期存在的限制: * 一个 autovacuum worker 一次只处理一张表,且 *vacuuming indexes* 和 *cleaning up indexes* 阶段逐个串行处理索引。一张有十几个索引的宽表可能独占 worker 数小时,其他表只能排队。 * 在单个 database 内,worker 大致按 `pg_class` 目录顺序处理候选表。一张逼近事务 ID wraparound 的表,与一张刚过 analyze 阈值的表,没有形式化的优先级差别。 手动 `VACUUM` 从 PostgreSQL 13 起就支持 `PARALLEL` 并行处理索引,但 autovacuum 一直用不上。给大表补上这个缺口只能靠人工定时跑手动 vacuum——基础机制见 [autovacuum 与表膨胀](/docs/operations/autovacuum-bloat)。 ## 并行索引处理:autovacuum\_max\_parallel\_workers [#并行索引处理autovacuum_max_parallel_workers] [`autovacuum_max_parallel_workers`](https://www.postgresql.org/docs/19/runtime-config-autovacuum.html) 设置单个 autovacuum worker 在 index vacuuming 和 index cleanup 阶段最多可招募的并行 worker 数。默认 `0`(禁用),需要显式开启: ```sql ALTER SYSTEM SET autovacuum_max_parallel_workers = 4; SELECT pg_reload_conf(); ``` 它相当于手动 `VACUUM` 的 `PARALLEL` 选项的 autovacuum 版本。实际 worker 数还受 `max_parallel_workers` 限制——那是与并行查询共享的总池子。 表级存储参数可以给单表设上限,避免某张热点表吃光并行预算: ```sql ALTER TABLE app.events SET (autovacuum_parallel_workers = 2); ``` 约束条件(见官方文档 [24.1.7 节 Parallel Vacuum](https://www.postgresql.org/docs/19/routine-vacuuming.html)): * 索引只有大于 `min_parallel_index_scan_size` 才能参与并行,且每个索引最多一个 worker——所以一张表至少要有两个合格索引才会启动并行 worker。 * 计算出的 worker 数不保证全部到位,实际可能更少甚至为零。 * 并行 worker 沿用 leader autovacuum worker 的 cost delay 参数,既有的节流模型仍然生效。 ## 评分调度系统 [#评分调度系统] 在单个 database 内,autovacuum worker 现在先构建候选表列表,再按评分排序,而不是按目录顺序处理。一张表的评分是五个分量评分的最大值(见 [24.1.6.1 节 Autovacuum Prioritization](https://www.postgresql.org/docs/19/routine-vacuuming.html)): * **事务 ID 年龄**——`age(relfrozenxid)` 相对 `autovacuum_freeze_max_age` 的比例;超过 `vacuum_failsafe_age` 后急剧上升。权重:`autovacuum_freeze_score_weight`。 * **multixact ID 年龄**——`relminmxid` 相对 `autovacuum_multixact_freeze_max_age` 的比例;超过 `vacuum_multixact_failsafe_age` 或 multixact member 超过约 20 亿条后急剧上升。权重:`autovacuum_multixact_freeze_score_weight`。 * **vacuum**——更新/删除的元组数相对 vacuum 阈值。权重:`autovacuum_vacuum_score_weight`。 * **vacuum insert**——插入元组数相对 insert 阈值。权重:`autovacuum_vacuum_insert_score_weight`。 * **analyze**——变更元组数相对 analyze 阈值。权重:`autovacuum_analyze_score_weight`。 五个权重默认都是 `1.0`(一视同仁),用 `pg_reload_conf()` reload 即可生效。文档里有两个细节值得注意: * freeze 权重调到 1.0 以上不只是放大评分——分量开始激进爬升的年龄会*除以*该权重,即 freeze 压力会更早被视为紧急。 * 把五个权重全部设为 `0.0`,调度策略退回 19 之前的目录顺序。 database 级的选择是另一层逻辑:launcher 仍然优先处理有 wraparound 风险的 database,其次是最久未处理的。 ## 用 pg\_stat\_autovacuum\_scores 观测 [#用-pg_stat_autovacuum_scores-观测] 新增的 `pg_stat_autovacuum_scores` 视图展示当前 database 里每张表的实时评分,把"autovacuum 为什么不理这张表"从猜源码变成一次查询: ```sql SELECT relid::regclass AS relation, round(score::numeric, 1) AS score, round(xid_score::numeric, 1) AS xid_score, round(vacuum_score::numeric, 1) AS vacuum_score, round(vacuum_insert_score::numeric, 1) AS insert_score, round(analyze_score::numeric, 1) AS analyze_score, do_vacuum, do_analyze, for_wraparound FROM pg_stat_autovacuum_scores ORDER BY score DESC LIMIT 20; ``` `score` 是五个 `*_score` 分量的最大值;`do_vacuum` / `do_analyze` 表示该表当前是否满足触发条件,`for_wraparound` 标记反 wraparound 压力。文档的一个提醒:视图用当前会话可见的信息计算评分,可能与 autovacuum worker 构建列表时看到的不完全一致——它是排障工具,不是处理顺序的保证。调权重前后各查一次,确认排序确实按预期变化——例如让 dead tuple 回收优先于统计刷新: ```sql ALTER SYSTEM SET autovacuum_vacuum_score_weight = 2.0; ALTER SYSTEM SET autovacuum_analyze_score_weight = 0.5; SELECT pg_reload_conf(); ``` 建议把这个视图和 `pg_stat_progress_vacuum` 一起纳入 [监控与日志](/docs/operations/monitoring-logging) 中的日常巡检。 ## 对比 PG18 的 vacuumdb --jobs [#对比-pg18-的-vacuumdb---jobs] PostgreSQL 18 时代应对维护缓慢的经典办法是手动并行: ```bash vacuumdb --jobs=4 --analyze dbname ``` `vacuumdb --jobs` 起多个连接、每个处理不同的表——这是*跨表*并行,时间窗口和调度都由你自己负责。它无法加速单张大表的索引阶段(除非你自己跑 `VACUUM (PARALLEL n)`),也改变不了 autovacuum 内部的处理顺序。 PostgreSQL 19 补的是另一条轴:*单表索引阶段内*的并行,加上感知紧急程度的排序,两者都是自动的。定时跑 `vacuumdb` 对可预测的批量窗口、整库 `FREEZE` 仍然有用,二者并不互斥。 ## 什么时候不要开 [#什么时候不要开] * **I/O 已饱和的系统。** 并行索引 vacuum 会成倍增加并发 I/O 流。cost limit 仍然共享,但对延迟敏感、存储已打满负载的系统应小步开启,同时盯 `pg_stat_io` 和查询延迟。 * **小库。** 所有索引都小于 `min_parallel_index_scan_size` 时,没有任何索引能参与并行——开了也没有效果。 * **worker 预算紧张。** vacuum 并行 worker 从 `max_parallel_workers` 里出,与并行查询共用。小实例上两边都调高可能挤占查询并行。 * **多数表只有一两个索引。** 每个索引一个 worker,只有一个合格索引的表永远走不了并行。 和一切 autovacuum 调优一样:一次只改一个参数,用 `pg_stat_autovacuum_scores` 和 autovacuum 日志验证效果,而不是凭感觉。PostgreSQL 19 的更多变化见[版本总览](/docs/postgresql-19)。 ## AI 提示词:调整 autovacuum 评分 [#ai-提示词调整-autovacuum-评分] --- # PostgreSQL 生产工具栈与高可用 Canonical URL: https://pg.edu.rich/docs/operations/production-stack Last reviewed: 2026-08-02 生产 PostgreSQL 的优先级通常是:**能够恢复 → 不耗尽连接 → 看得见问题 → 安全地变更 → 再自动故障切换**。高可用不能替代备份,副本也不能修复已经复制过去的误删。 截至 **2026-08-02**,PostgreSQL 18.4 是最新稳定 major 18 的当前 minor;14–18 仍处于官方支持期。生产实例应运行其所在 major 的当前 minor,而不是只因为 18 最新就强制跨 major 升级。版本状态见 [PostgreSQL versioning policy](https://www.postgresql.org/support/versioning/) 与 [18.4 release notes](https://www.postgresql.org/docs/current/release-18-4.html)。 ## 最小生产基线 [#最小生产基线] | 层 | 最先回答的问题 | 常见选择 | | ---- | -------------------------- | ------------------------------------ | | 数据库 | minor 更新、角色、TLS、参数、扩展如何管理? | PostgreSQL 官方包或经过验证的镜像 | | 连接 | 峰值应用并发会不会耗尽 backend? | 应用连接池,必要时 PgBouncer | | 恢复 | RPO/RTO 是多少,能否从平台外恢复? | pgBackRest、WAL-G 或云平台备份 + 独立副本 | | 可观测性 | 哪条 SQL、哪个等待事件、哪段日志解释故障? | `pg_stat_statements`、JSON 日志、指标采集 | | 变更 | DDL 锁、回填、回滚和兼容窗口如何验证? | expand-and-contract、迁移检查、真实 PG 测试 | | 可用性 | 独立故障域、仲裁和切换后数据边界是什么? | 托管 HA、Patroni、CloudNativePG 或 Pigsty | ## PgBouncer:不要默认选择 transaction pooling [#pgbouncer不要默认选择-transaction-pooling] [PgBouncer](https://www.pgbouncer.org/) 把大量客户端连接复用到较少的 PostgreSQL server connection。不要通过持续增大 `max_connections` 代替容量设计;每个 backend 都会占用内存和调度资源,还要给迁移、监控、备份、管理与复制保留连接。 | 模式 | server connection 归还时间 | 适用边界 | | ------------- | ---------------------- | --------------------- | | `session` | 客户端断开 | 兼容性最高,适合依赖会话状态的应用 | | `transaction` | 事务结束 | Web/API 常用,但必须验证会话级特性 | | `statement` | 每条语句结束 | 不允许多语句事务,只适合非常受控的负载 | transaction pooling 下,SQL 级 `PREPARE`、会话 advisory lock、`LISTEN`、带 `WITH HOLD` 的 cursor,以及许多依赖会话状态的做法不可用或受限。协议级 prepared statement 需要正确配置 `max_prepared_statements` 并验证驱动行为。完整矩阵见 [PgBouncer feature map](https://www.pgbouncer.org/features.html)。 上线前至少验证:认证与 TLS、prepared statements、临时表、迁移工具、连接 reset、failover、事务重试,以及 ORM 是否把状态留在 session。 ## pgBackRest、WAL-G 与云平台备份 [#pgbackrestwal-g-与云平台备份] [pgBackRest](https://github.com/pgbackrest/pgbackrest) 支持 full/differential/incremental backup、并行传输、多 repository、WAL archive 与 PITR,适合自托管 PostgreSQL 的完整恢复链路。[WAL-G](https://github.com/wal-g/wal-g) 更偏对象存储工作流。二者都不能通过“安装完成”证明可恢复。 备份周期应由 RPO、WAL 生成量、恢复带宽、保留要求和实测 RTO 推导,而不是机械套用“每周全量、每日增量”。必须持续检查 archive gap、对象删除保护、加密密钥、跨账户或跨主机副本,并定期恢复到临时实例。 `pg_dump` 仍适合逻辑迁移、选择对象和小规模恢复,但它不能单独提供连续时间点恢复。详见 [备份、恢复与 PITR](/docs/operations/backup-recovery)。 ## Patroni、CloudNativePG 与 Pigsty 怎么选 [#patronicloudnativepg-与-pigsty-怎么选] | 环境 | 候选方案 | 采用前提 | | ---------------- | --------------------------------------------- | -------------------------------- | | 托管云数据库 | 服务商的跨可用区 HA | 核对区域、故障切换、PITR、扩展与平台外恢复限制 | | 独立 VM / 裸机 | [Patroni](https://github.com/patroni/patroni) | 多个独立故障域、可靠 DCS、网络与存储运维能力 | | 已有 Kubernetes 平台 | [CloudNativePG](https://cloudnative-pg.io/) | 团队已能运维 K8s、存储、网络和 Operator 升级 | | 多集群自托管平台 | [Pigsty](https://github.com/pgsty/pigsty) | 接受 Ansible/VM 运维模型,并验证其集成组件与升级路径 | CloudNativePG 的对象存储备份当前应评估 [Barman Cloud Plugin](https://cloudnative-pg.io/plugin-barman-cloud/docs/intro/),不要照搬已弃用的内置 object-store 配置。不要仅为了一个 PostgreSQL 实例引入 Kubernetes。 它们会同时受到主机断电、内核故障、存储损坏和网络中断影响。自动选主只在节点、存储和仲裁边界真实独立时改善可用性。 ## 分阶段采用 [#分阶段采用] ### 生产基线 [#生产基线] * 当前 minor、角色分离、TLS 与受控扩展; * 连接预算,必要时部署并验证 PgBouncer; * 独立保留的备份、连续 WAL/PITR 与恢复演练; * SQL、指标、日志三层可观测性; * 迁移前锁测试、超时、回退路径和业务验证。 ### 条件采用 [#条件采用] * 只有存在独立故障域和明确 RTO 时,才引入自动故障切换; * 只有已有成熟 Kubernetes 运维时,才优先 CloudNativePG; * 只有管理多个自托管集群的收益覆盖平台复杂度时,才引入 Pigsty; * 只有真实工作负载需要时,才安装 pgvector、PostGIS、TimescaleDB 等扩展。 ### 实验通道 [#实验通道] PostgreSQL 19 Beta、较新的扩展和存储引擎应进入可丢弃的兼容性环境,而不是生产默认。将稳定版本测试与下一 major 测试拆成两条 CI 通道,见 [安全迁移与零停机 Schema 变更](/docs/operations/safe-migrations)。 --- # REPACK / REPACK CONCURRENTLY 在线消膨胀 Canonical URL: https://pg.edu.rich/docs/operations/repack Last reviewed: 2026-08-06 截至 2026-08,PostgreSQL 19 尚未正式发布。下文语法、GUC 名与视图以 [PostgreSQL 19 文档](https://www.postgresql.org/docs/19/sql-repack.html)为准;用于生产前请对照正式 release notes 复核。 普通 `VACUUM` 让死空间在关系内部复用,但几乎不会把空间还给操作系统。两种经典的收缩手段——`VACUUM FULL` 和 `CLUSTER`——都要在整个重写期间持有 `ACCESS EXCLUSIVE` 锁,所以团队要么排维护窗口,要么装第三方的 `pg_repack` 扩展。PostgreSQL 19 给出了内建的第三条路:[`REPACK`](https://www.postgresql.org/docs/19/sql-repack.html),其 `CONCURRENTLY` 模式在重建期间保持表可读写。 ## VACUUM FULL 与 CLUSTER 的锁问题 [#vacuum-full-与-cluster-的锁问题] `VACUUM FULL` 把表的存活元组重写到新文件并重建所有索引;`CLUSTER` 做同样的事,只是按索引排序。两者从头到尾持有 `ACCESS EXCLUSIVE`,期间表上的所有读写都排在锁后面。大表重写要几分钟到几小时,光是锁等待堆积的被阻塞会话就足以拖垮应用。膨胀的测量与目标表选择见 [autovacuum 与表膨胀](/docs/operations/autovacuum-bloat);本页只讲重建这一步。 ## REPACK:统一的命令语义 [#repack统一的命令语义] `REPACK` 把 `VACUUM FULL` 与 `CLUSTER` 的行为折进一条语句: ```sql REPACK [ ( option [, ...] ) ] [ table_and_columns [ USING INDEX [ index_name ] ] ] REPACK [ ( option [, ...] ) ] USING INDEX ``` 选项有 `VERBOSE`、`ANALYZE`、`CONCURRENTLY`: * `REPACK t;`——纯重写回收磁盘,等价于 `VACUUM FULL`。 * `REPACK t USING INDEX i;`——额外按索引重排行,等价于 `CLUSTER`。省略索引名时使用之前 `ALTER TABLE ... CLUSTER ON` 设置的索引。 * `REPACK;`——不带表名,处理当前数据库里你持有 `MAINTAIN` 权限的所有表和物化视图。该形式不能在事务块里执行,也不能与 `CONCURRENTLY` 组合。 * `REPACK (ANALYZE) t;`——重写后执行 `ANALYZE`。目前只支持单张非分区表;重写后统计信息重置,无论如何都值得做。 需要对表持有 `MAINTAIN` 权限。不带 `CONCURRENTLY` 时,`REPACK` 仍然全程持有 `ACCESS EXCLUSIVE`——锁语义没变,统一的只是命令表面。 ## CONCURRENTLY 的实现原理 [#concurrently-的实现原理] `CONCURRENTLY` 模式下,`REPACK` 把存活元组拷贝到新文件(每个索引各一个新文件),期间对旧文件的读写照常进行。拷贝过程中发生的变更通过逻辑解码捕获并应用到新文件;之后命令才获取短暂的 `ACCESS EXCLUSIVE` 锁交换新旧文件、删除旧文件。锁通常只持有交换所需的时间——但如果积压的变更很多,它们必须在持锁期间 replay 完,所以热点表在收尾阶段仍可能有可感知的阻塞窗口。 文档里两条行为值得记住: * repack 开始后插入的行不参与排序,即使指定 `USING INDEX`——clustering 是一次性的物理重排。 * 其他会话在 repack 期间对该表执行 DDL 可能导致 `REPACK CONCURRENTLY` 失败。进行中的 repack 要避开迁移窗口。 文档明确警告 `REPACK` 的 `CONCURRENTLY` 选项不是 MVCC 安全的(见 [MVCC caveats](https://www.postgresql.org/docs/19/mvcc-caveats.html))。把它当作需要刻意安排的维护操作,而不是无感的后台动作。 ## 硬性限制 [#硬性限制] 以下任一情况都会拒绝 `CONCURRENTLY`: * 表是 `UNLOGGED`; * 表是分区表(普通 `REPACK` 可以处理分区表——逐个分区 repack——但不能并发,也不能在事务块里); * 表既没有主键也没有基于索引的 replica identity——逻辑解码需要标识行的手段; * 表是系统目录或 TOAST 表; * `REPACK` 在事务块里执行; * `max_repack_replication_slots` 没有空闲槽位(见下文)。 另一个硬限制是磁盘,与是否 `CONCURRENTLY` 无关:重写需要表加全部索引的临时副本,空闲空间至少是表大小加索引大小。顺序扫描加排序的路径还会产生临时排序文件,峰值可能接近表大小的两倍再加索引(可以对会话设置 `enable_sort = off` 强制走索引扫描路径)。`CONCURRENTLY` 在此之上还要缓冲拷贝期间的并发 DML。开始前给会话配一个宽裕的 `maintenance_work_mem`。 ## max\_repack\_replication\_slots 与槽位隔离 [#max_repack_replication_slots-与槽位隔离] `CONCURRENTLY` 需要一个复制槽来做逻辑解码。PostgreSQL 19 为此给 `REPACK` 单开了一个池子:[`max_repack_replication_slots`](https://www.postgresql.org/docs/19/runtime-config-replication.html#GUC-MAX-REPACK-REPLICATION-SLOTS)(默认 **5**,只能在启动时设置)在 `max_replication_slots` 之外为 `REPACK` 独占预留槽位。隔离是双向的: * repack 作业永远不会占用逻辑复制发布/订阅依赖的槽位; * 订阅端也永远饿不死 repack 作业; * 但默认同时只能跑 5 个 `REPACK CONCURRENTLY`,第六个直接失败。 实操建议:串行执行 repack 作业(出于 I/O 考虑,一两个并发通常也是上限);确实需要更多并行度时,计划一次重启来调大 `max_repack_replication_slots`。repack 失败时不要去调 `max_replication_slots`——耗尽的并不是那个池子。 ## 内建 REPACK 与 pg\_repack 扩展的选型 [#内建-repack-与-pg_repack-扩展的选型] 内建命令与 [pg\_repack 扩展](https://pgrepack.github.io/pg_repack/)是解决同一问题的两套独立实现。选型对照: | | `REPACK`(PostgreSQL 19 内建) | `pg_repack`(扩展) | | ------ | ------------------------------------ | ------------------------------- | | 可用性 | PostgreSQL 19+,无需安装 | 扩展 + CLI;PostgreSQL 18 及以前的生产方案 | | 在线变更捕获 | 基于逻辑解码的专用复制槽 | 触发器写日志表,交换前 replay | | 行标识要求 | 主键或基于索引的 replica identity | 主键或合适的唯一索引 | | 锁形态 | 仅最终交换时短暂 `ACCESS EXCLUSIVE` | 开始和结束时的短暂排他锁 | | 物理重排 | `USING INDEX` | `--order-by` | | 只重建索引 | 不支持——用 `REINDEX CONCURRENTLY` | `--index` / `--only-indexes` | | 运维形态 | SQL 语句,`pg_stat_progress_repack` 看进度 | 外部 CLI,有独立的发布节奏与版本匹配要求 | 在 PostgreSQL 19 上,内建命令去掉了扩展依赖、版本匹配负担和基于触发器的日志表。在更老的版本上,`pg_repack` 仍是工具——内建命令在那些版本不存在。 ## 实操 runbook [#实操-runbook] 一次完整的 repack 流程: ```sql -- 1. 确认表确实膨胀(pg_stat_user_tables 的估算不够) CREATE EXTENSION IF NOT EXISTS pgstattuple; SELECT * FROM pgstattuple('app.events'); -- 看 dead_tuple_percent 和 free_percent -- 2. 确认可以用 CONCURRENTLY:replica identity 必须是默认(有主键)或索引 SELECT relreplident FROM pg_class WHERE oid = 'app.events'::regclass; -- 'd'(默认,需有主键)或 'i'(replica identity 索引)可以;'n' 和 'f' 不行 -- 3. 确认磁盘:空闲空间 >= 表 + 索引大小 SELECT pg_size_pretty(pg_total_relation_size('app.events')); ``` ```sql -- 4. 在独立会话中执行 SET maintenance_work_mem = '1GB'; REPACK (CONCURRENTLY, ANALYZE, VERBOSE) app.events; ``` ```sql -- 5. 从另一个会话观察进度 SELECT pid, datname, relid::regclass AS relation, command, phase, heap_blks_scanned, heap_blks_total, heap_tuples_scanned, heap_tuples_inserted, index_rebuild_count FROM pg_stat_progress_repack; ``` 阶段从 `initializing` 开始,经过堆扫描/拷贝(`seq scanning heap` / `index scanning heap`、`sorting tuples`、`writing new heap`),然后是 `catch-up`(应用缓冲的变更)、`swapping relation files`(短暂的排他锁)、`rebuilding index` 和 `performing final cleanup`。`catch-up` 阶段拖得长,说明表足够热、最终持锁窗口会可感知——考虑换到更空闲的窗口重跑。该视图同时跟踪 `CLUSTER` 和 `VACUUM FULL`,用 `command` 列区分。 repack 中途失败时,临时文件和复制槽由命令负责清理,原表不受影响。排除原因后重试即可——最常见的原因是并发 DDL、缺 replica identity、槽位池耗尽。 ## 提示词:规划消膨胀战役 [#提示词规划消膨胀战役] 命令参考:[PostgreSQL 19 REPACK](https://www.postgresql.org/docs/19/sql-repack.html) 与[进度报告视图](https://www.postgresql.org/docs/19/progress-reporting.html)。另见 [PostgreSQL 19 总览](/docs/postgresql-19)。第三方实测:[depesz 的走查](https://www.depesz.com/2026/03/19/waiting-for-postgresql-19-introduce-the-repack-command/)与 [digoal 关于槽位隔离的分析](https://github.com/digoal/blog/blob/master/202604/20260408_02.md)。 --- # 复制、故障切换与升级 Canonical URL: https://pg.edu.rich/docs/operations/replication-upgrades Last reviewed: 2026-08-06 ## 物理与逻辑复制 [#物理与逻辑复制] | 维度 | 物理流复制 | 逻辑复制 | | ---- | ---------------------- | ------------------- | | 单位 | WAL/实例 | 表级变更 | | 目标 | 同一大版本体系的 standby、HA、只读 | 选择表、跨大版本迁移、数据分发 | | DDL | 物理同步 | 通常需要另行同步 schema | | 序列 | 物理同步 | 需单独处理序列状态 | | 写入冲突 | standby 不写 | subscriber 本地写入可能冲突 | 物理 standby 的基本观测: ```sql -- primary SELECT application_name, state, sync_state, sent_lsn, write_lsn, flush_lsn, replay_lsn FROM pg_stat_replication; -- standby SELECT pg_is_in_recovery(), pg_last_wal_receive_lsn(), pg_last_wal_replay_lsn(), now() - pg_last_xact_replay_timestamp() AS replay_delay; ``` 没有新事务时,时间型 replay delay 可能为空或看起来很大;同时看 LSN、WAL 速率和业务健康。 ## 复制槽 [#复制槽] 复制槽可防止所需 WAL 被过早删除,但消费者停止时会持续占用磁盘。为 slot lag 和 `pg_wal` 空间设置告警与容量上限;删除槽前确认没有消费者依赖。 ```sql SELECT slot_name, slot_type, active, pg_size_pretty(pg_wal_lsn_diff(pg_current_wal_lsn(), restart_lsn)) AS retained FROM pg_replication_slots; ``` ## 故障切换不是单条命令 [#故障切换不是单条命令] 切换流程要覆盖:确认 primary 确实不可用、评估未同步 WAL、提升目标节点、让旧 primary 无法继续接受写入(fencing)、更新路由、验证写入与后台任务、重建冗余。缺少 fencing 可能产生双主分叉。 ## 升级路径 [#升级路径] * **小版本**:同一 major 内只包含修复;通常更换二进制并重启,但仍要阅读 release notes。 * **大版本**:需要 `pg_upgrade`、逻辑 dump/restore 或逻辑复制迁移;数据目录不向前兼容。 * 可以跨过中间 major 直接升级,但应阅读所有中间版本 release notes,并验证扩展支持。 ### 大版本切换清单 [#大版本切换清单] 1. 清点扩展、collation、数据类型、驱动与复制拓扑。 2. 在恢复出的生产副本上演练升级,统计停机、磁盘与 `ANALYZE` 时间。 3. 跑应用测试、关键查询计划对比和数据校验。 4. 冻结或双写期间明确 source of truth。 5. 切换前确认复制追平、长事务清空、回退窗口仍有效。 6. 切换后重建统计、检查 invalid objects、错误率、性能与备份。 如果目标是当前正在测试的 major,请使用 [PostgreSQL 19:新功能与 18 升级 19 指南](/docs/postgresql-19) 核对 Beta 状态、兼容性变化与 `pg_upgrade --check` 清单。 ### PostgreSQL 19 的逻辑复制新增能力 [#postgresql-19-的逻辑复制新增能力] PostgreSQL 19(截至本页更新仍为 Beta)计划了多项直接影响逻辑复制升级的变化;细节以最终 release notes 为准。 * **序列同步**补上了上表中的经典缺口:`CREATE PUBLICATION ... FOR ALL SEQUENCES` 发布序列,`ALTER SUBSCRIPTION ... REFRESH SEQUENCES` 让 subscriber 的序列值与 publisher 对齐;`pg_get_sequence_data()` 可查看同步状态。切换后 subscriber 上的新序列值不会再与已复制行冲突——这正是逻辑复制大版本升级后主键冲突的常见根因。见 [CREATE PUBLICATION](https://www.postgresql.org/docs/19/sql-createpublication.html) 与 [ALTER SUBSCRIPTION](https://www.postgresql.org/docs/19/sql-altersubscription.html)。 * **Publication 黑名单**:`FOR ALL TABLES EXCEPT (TABLE ...)` 只排除指定表,不必枚举所有纳入的表;只有少数表需要留在旧端时,迁移配置会简单很多。 * **冲突保留**:订阅参数 `retain_dead_tuples` 与 `max_retention_duration` 在 subscriber 上保留用于冲突检测的 dead tuple 信息,并以保留窗口限制时长。 * **standby 上 read-your-writes**:`WAIT FOR LSN` 命令让会话等待指定 LSN 在 standby 上完成 replay,切换后把读流量导向副本时不会丢失刚提交的写入。见 [WAIT FOR LSN](/docs/operations/wait-for-lsn)。 PostgreSQL 每个 major 通常支持 5 年。新系统使用受支持版本的最新 minor;截至 2026-08,18、17、16、15、14 受支持,14 将在 2026-11 结束支持。详见 [版本策略](/docs/reference/version-policy)。 --- # PostgreSQL 安全迁移与零停机 Schema 变更 Canonical URL: https://pg.edu.rich/docs/operations/safe-migrations Last reviewed: 2026-08-02 零停机不是某条 DDL 的属性,而是**旧应用与新应用能否在迁移窗口内同时工作**。安全迁移需要兼容阶段、锁预算、真实数据规模测试、可观察执行和明确回退点。 ## 默认采用 expand-and-contract [#默认采用-expand-and-contract] 以重命名或替换一个高流量字段为例: 1. **Expand**:添加新字段或新表,不删除旧结构;尽量使用短元数据操作。 2. **Dual compatible**:应用能读取新旧结构,并在必要时双写;写入必须幂等。 3. **Backfill**:按主键范围或时间窗口小批回填,限制事务时长、WAL 与 replica lag。 4. **Switch**:先切读路径,再停止旧写入;用业务指标和校验查询确认。 5. **Contract**:经过一个可回退发布窗口后,才删除旧结构。 直接把“加列、回填、设 `NOT NULL`、删旧列”放进一个长事务,通常会扩大锁、WAL、rollback 和复制延迟风险。 ## 给锁等待设上限 [#给锁等待设上限] ```sql BEGIN; SET LOCAL lock_timeout = '2s'; SET LOCAL statement_timeout = '15min'; ALTER TABLE app.orders ADD COLUMN IF NOT EXISTS fulfillment_state text; COMMIT; ``` 示例超时不是通用默认值。`lock_timeout` 防止迁移长时间排队后在不可控时刻获得强锁;`statement_timeout` 限制执行时间。失败后应退出并调查 blocker,而不是无限重试。 添加约束时可把扫描与短锁阶段拆开: ```sql ALTER TABLE app.orders ADD CONSTRAINT orders_total_nonnegative CHECK (total_cents >= 0) NOT VALID; ALTER TABLE app.orders VALIDATE CONSTRAINT orders_total_nonnegative; ``` 先在目标 PostgreSQL 版本和代表性数据上检查具体 DDL 的锁级别。`CREATE INDEX CONCURRENTLY` 也会消耗 I/O、WAL 和更长时间,并需要检查失败后留下的 invalid index。 ## 推荐 CI 流水线 [#推荐-ci-流水线] ```text Schema / reviewed SQL ↓ Migration generation ↓ Squawk static checks ↓ Disposable PostgreSQL 18.4 ↓ Apply every migration from an empty and upgraded state ↓ pgTAP + application integration + RLS negative tests ↓ PostgreSQL 19 Beta compatibility lane ↓ Representative-data rehearsal → staging → production ``` * [Squawk](https://github.com/sbdchd/squawk) 检查常见危险迁移,例如非并发索引、未使用 `NOT VALID` 的约束和一些锁风险;它不是零停机证明。 * [Testcontainers for Node.js](https://github.com/testcontainers/testcontainers-node) 在 CI 中启动真实 PostgreSQL,适合验证事务、锁、RLS、JSONB、扩展与驱动行为。 * [pgTAP](https://github.com/theory/pgtap) 在数据库内部测试函数、触发器、约束和策略。 如果使用 Drizzle ORM,可让 Drizzle Kit 生成普通 schema 变更,再对生成 SQL 进行 Squawk 与人工审核。复杂 index、policy、function、extension 和 PostgreSQL 19 新语法可以使用受审查的原生 SQL;ORM 无法表达不代表数据库不应该使用。 ## 测试升级路径,而不只测试空库 [#测试升级路径而不只测试空库] CI 至少需要两种数据库状态: | 起点 | 能发现的问题 | | ------------------------------ | ---------------- | | 空数据库执行全部 migration | 顺序、依赖、语法和初始化问题 | | 生产版本 schema/脱敏数据执行增量 migration | 锁、回填、旧数据、约束与性能问题 | 再将测试拆成两条版本通道: * **生产门禁**:当前生产 major/minor,例如 PostgreSQL 18.4,失败时阻止发布; * **前瞻兼容**:PostgreSQL 19 Beta 2,可允许失败但必须归类、跟踪和在 GA 前清零。 不要让 Beta 测试替代稳定版本门禁。版本状态见 [PostgreSQL 19 专题](/docs/postgresql-19)。 ## RLS 与安全对象必须做反向测试 [#rls-与安全对象必须做反向测试] 迁移成功不代表权限正确。为每个 tenant 和 role 验证:允许的 `SELECT/INSERT/UPDATE/DELETE` 成功,不允许的跨租户读写失败;应用运行角色不是表 owner,必要时启用 `FORCE ROW LEVEL SECURITY`。详见 [安全基线](/docs/operations/security)。 删除列、不可逆回填、类型收窄和外部副作用可能无法安全逆转。每次发布应写明最后可回退时点、旧应用能否读取新 schema,以及 forward fix 的触发条件。 ## 何时采用更重的工具 [#何时采用更重的工具] | 工具 | 适用场景 | 采用前先验证 | | ------------------------------------------------------------------------- | --------------------------- | -------------------- | | [pgroll](https://github.com/xataio/pgroll) | 高频、兼容窗口明确的零停机 schema change | 支持的 DDL、代理/连接方式、回滚语义 | | [Bytebase](https://github.com/bytebase/bytebase) | 多团队审批、SQL review、环境和审计治理 | 权限边界、部署模型、现有 CI 集成 | | [Database Lab Engine](https://github.com/postgres-ai/database-lab-engine) | 大型数据库的快速 clone 与迁移演练 | 存储、脱敏、clone 生命周期与成本 | 小团队先把 expand-and-contract、真实 PostgreSQL 测试、锁观察与恢复演练做好,再引入控制面。 --- # 安全基线 Canonical URL: https://pg.edu.rich/docs/operations/security Last reviewed: 2026-08-06 ## 分离角色 [#分离角色] ```sql CREATE ROLE app_owner NOLOGIN; CREATE ROLE app_runtime LOGIN; CREATE ROLE app_migrator LOGIN NOINHERIT; CREATE SCHEMA app AUTHORIZATION app_owner; GRANT app_owner TO app_migrator; GRANT CONNECT ON DATABASE commerce TO app_runtime, app_migrator; GRANT USAGE ON SCHEMA app TO app_runtime; GRANT SELECT, INSERT, UPDATE, DELETE ON ALL TABLES IN SCHEMA app TO app_runtime; ALTER DEFAULT PRIVILEGES FOR ROLE app_owner IN SCHEMA app GRANT SELECT, INSERT, UPDATE, DELETE ON TABLES TO app_runtime; ALTER DEFAULT PRIVILEGES FOR ROLE app_owner IN SCHEMA app GRANT USAGE, SELECT ON SEQUENCES TO app_runtime; ``` 运行时角色不拥有对象,迁移角色只在迁移期间 `SET ROLE app_owner`,所有者角色不登录。`ALTER DEFAULT PRIVILEGES` 只影响未来由指定创建者创建的对象,不会回补已有对象。 为 BI 工具、支持查询和 AI agent 增加一个只读层,补全分层: ```sql CREATE ROLE app_readonly LOGIN; ALTER ROLE app_readonly SET default_transaction_read_only = on; GRANT CONNECT ON DATABASE commerce TO app_readonly; GRANT USAGE ON SCHEMA app TO app_readonly; GRANT SELECT ON ALL TABLES IN SCHEMA app TO app_readonly; ALTER DEFAULT PRIVILEGES FOR ROLE app_owner IN SCHEMA app GRANT SELECT ON TABLES TO app_readonly; ``` `default_transaction_read_only` 保证即使之后误加了写权限,意外写入也会直接失败。 ## 认证与网络 [#认证与网络] * 只监听需要的接口,用防火墙/安全组限制来源。 * 远程连接要求 TLS,并验证服务端证书;高敏场景考虑客户端证书。 * 新密码认证使用 SCRAM,逐步淘汰 MD5 配置。 * `pg_hba.conf` 按具体网络、database、role 从窄到宽编排;修改后 reload 并测试允许与拒绝两条路径。 * 管理入口与应用入口分开,避免向公网暴露数据库端口。 与上述顺序对应的 `pg_hba.conf` 示例——先匹配先生效,所以窄规则在前: ``` # TYPE DATABASE USER ADDRESS METHOD local all all peer hostssl commerce app_runtime 10.0.1.0/24 scram-sha-256 hostssl commerce app_readonly 10.0.2.0/24 scram-sha-256 host all all 0.0.0.0/0 reject ``` 修改后执行 `SELECT pg_reload_conf();`,并分别测试放行与拒绝两条路径。 ## search\_path 防护 [#search_path-防护] 不要信任可写 schema 中的同名对象解析。撤销 `public` 的默认创建权,并为安全敏感函数固定路径: ```sql REVOKE CREATE ON SCHEMA public FROM PUBLIC; CREATE FUNCTION app.current_tenant() RETURNS bigint LANGUAGE sql STABLE SECURITY DEFINER SET search_path = pg_catalog, app AS $$ SELECT current_setting('app.tenant_id')::bigint $$; REVOKE ALL ON FUNCTION app.current_tenant() FROM PUBLIC; GRANT EXECUTE ON FUNCTION app.current_tenant() TO app_runtime; ``` `SECURITY DEFINER` 函数以所有者权限运行,必须审计所有参数、对象限定名、search path 和执行权限。 ## 行级安全 RLS [#行级安全-rls] ```sql ALTER TABLE app.orders ENABLE ROW LEVEL SECURITY; ALTER TABLE app.orders FORCE ROW LEVEL SECURITY; CREATE POLICY tenant_orders ON app.orders USING (tenant_id = current_setting('app.tenant_id')::bigint) WITH CHECK (tenant_id = current_setting('app.tenant_id')::bigint); ``` RLS 启用后,没有适用策略时默认拒绝。超级用户、`BYPASSRLS` 角色以及通常的表所有者可绕过;`FORCE ROW LEVEL SECURITY` 让所有者在普通访问中也受策略约束。仍需普通对象权限。 管理员测试成功不能证明 RLS 有效。测试允许租户、其他租户、缺失 tenant context、插入与更新,并确认连接池每次借出/归还时正确设置和清除上下文。 ## Secret 与日志 [#secret-与日志] 凭据轮换、短期化并由 secret manager 分发。数据库日志避免记录绑定值和敏感 DDL;审计日志限制访问与保留期。`pg_stat_activity` 也可能显示 SQL 文本,读取监控视图的权限同样需要控制。 ## 审计 [#审计] 内建的起点是语句日志: ```sql ALTER SYSTEM SET log_statement = 'ddl'; SELECT pg_reload_conf(); ``` 需要可追责的审计轨迹时使用 [pgaudit](https://github.com/pgaudit/pgaudit),它必须通过 `shared_preload_libraries` 加载: ```ini shared_preload_libraries = 'pgaudit' pgaudit.log = 'write, ddl, role' ``` 需要提前规划的边界: * pgaudit 是扩展而非核心功能,可用性和版本取决于发行方式或云服务商。 * `pgaudit.log = 'all'` 这类宽类别在繁忙系统上日志量巨大。先从 `ddl, role` 开始,按合规要求逐步增加类别;对少数敏感表用对象级审计(`pgaudit.log_relation`),而不是全量记录。 * 会话审计仍可能记录敏感值,上一节的日志访问、保留和脱敏规则同样适用于审计输出。 --- # 单库表太多对 PostgreSQL 的危害 Canonical URL: https://pg.edu.rich/docs/operations/too-many-tables Last reviewed: 2026-08-06 一个有几十万张表的数据库,通常是通过三条路之一走到这一步的: * **每租户一套表**:每个客户拥有自己的一组表(`tenant_1234.orders`、`tenant_1234.invoices`……),因为这样看起来隔离得干净; * **分区狂魔**:几十张表全部按天分区且永久保留,每个分区又各自带着索引、约束和统计信息; * **ORM 与工具失控**:框架为每个实体版本、每张报表、每次导入任务建表,且从不清理。 三条路通向同一种故障形态,而且它是渐进的:5,000 张表时一切正常,50,000 张时某些东西说不清地慢,500,000 张时你在排查内存压力和以小时计的 `pg_dump`。 ## 为什么表多是病 [#为什么表多是病] ### relcache:按后端缓存的元数据 [#relcache按后端缓存的元数据] 每个后端进程都维护自己的 relation cache——relcache——存放它访问过的每个关系的解析后元数据:tuple descriptor、索引、规则、触发器、统计信息指针。这是后端私有内存中的按进程缓存,首次访问时惰性填充,随后跟随后端整个生命周期(见 PostgreSQL 源码 `src/backend/utils/cache/relcache.c`)。由此得出两个结论: * 一张被触碰过的表,其内存代价**每个后端各付一次**,所以 200 个连接各自触碰 20,000 张表,元数据就存了 200 份; * 连接存活期间没有任何机制回收它,所以 `session` 模式下长命的连接池连接内存单调增长——在容器里,这正是最终以 [OOM Kill](/docs/operations/container-memory-oom) 收场的那种匿名内存增长。 PostgreSQL 14+ 可以通过 [`pg_backend_memory_contexts`](https://www.postgresql.org/docs/18/view-pg-backend-memory-contexts.html) 观察到这一点:元数据累积在 `CacheMemoryContext` 及其子上下文下,可以看着它随后端触碰的关系数增长。 ### 系统目录变大,所有遍历它的操作都变慢 [#系统目录变大所有遍历它的操作都变慢] 一张表不是一行目录记录。一张有五列、一个主键、一个索引的表,会向 `pg_class`、`pg_attribute`、`pg_index`、`pg_constraint`、`pg_depend`、`pg_description`、`pg_statistic` 等目录各写入若干行——按[系统目录](https://www.postgresql.org/docs/18/catalogs.html)的结构,每张表至少十几行。规模化的后果: * 元数据内省变慢:psql 的 `\d`、应用启动时 ORM 的 schema 反射、枚举目录的 GUI 工具; * `pg_dump` 要遍历并锁定每个关系,所以即使数据量不变,备份时长也随表数量增长; * 目录本身也会像普通表一样膨胀、需要 vacuum——巨大 `pg_class` 上的 autovacuum 是真实存在的负载。 ## 多少算多 [#多少算多] 以下是经验值,不是文档化的阈值——实际的临界点取决于连接数、每张表的列数、以及每个后端触碰多少关系: * **几千张表**:任何合理配置下都没问题; * **几万张**:开始能感到——连接预热变慢、dump 变慢、relcache 内存在监控里可见; * **几十万张**:事故区——后端持有数 GB relcache、容器内 OOM 风险、`pg_dump` 以小时计。 真正关键的乘数是 *每后端触碰的表数 × 并发后端数*,而不是原始表数。50,000 张表但每个请求只碰 50 张,远比 50,000 张表每个请求全碰便宜。 ## 诊断 [#诊断] 按类型统计关系数,排除系统 schema: ```sql SELECT c.relkind, count(*) FROM pg_class c JOIN pg_namespace n ON n.oid = c.relnamespace WHERE n.nspname NOT IN ('pg_catalog', 'information_schema') GROUP BY c.relkind ORDER BY count(*) DESC; ``` `relkind` 含义:`r` 普通表,`p` 分区表,`i`/`I` 索引,`S` 序列,`t` TOAST 表,`v`/`m` 视图。如果计数被索引主导,底下的表仍是根因——每张表都拖着自己的索引。 找出表集中在哪里: ```sql SELECT n.nspname, count(*) FILTER (WHERE c.relkind = 'r') AS tables, count(*) FILTER (WHERE c.relkind = 'i') AS indexes FROM pg_class c JOIN pg_namespace n ON n.oid = c.relnamespace WHERE n.nspname NOT IN ('pg_catalog', 'information_schema') GROUP BY n.nspname ORDER BY tables DESC LIMIT 20; ``` 成千上万个 schema、里面的表名完全相同,就是每租户一套表的签名。 测量目录大小和单后端的元数据内存: ```sql -- 最大的系统目录 SELECT relname, pg_size_pretty(pg_total_relation_size(oid)) AS total FROM pg_class WHERE relnamespace = 'pg_catalog'::regnamespace AND relkind = 'r' ORDER BY pg_total_relation_size(oid) DESC LIMIT 10; -- 当前后端的元数据内存(PG 14+; -- 观察其他后端需要 superuser 或 pg_read_all_stats) SELECT name, pg_size_pretty(sum(used_bytes)) AS used FROM pg_backend_memory_contexts GROUP BY name ORDER BY sum(used_bytes) DESC LIMIT 15; ``` `CacheMemoryContext` 随后端触碰过的不同关系数增长、且从不收缩,即可确认 relcache 代价。 ## 治理路线 [#治理路线] **把每租户的表合并为共享表 + 行级安全(RLS)。** 一张带 `tenant_id` 列和 RLS 策略的 `orders` 表保住了隔离保证,同时把表数量压缩几个数量级;见 [Row Security Policies](https://www.postgresql.org/docs/18/ddl-rowsecurity.html) 和本站[安全](/docs/operations/security)页。迁移本身是机械工作(把各租户的表 union 进去、加列、回填),但要在你自己的查询形态上测试 RLS 的执行计划——策略按查询生效,并与 `tenant_id` 上的索引相互作用。 **刻意给分区数设上限。** 分区粒度由保留周期和查询模式决定,而不是由日历习惯决定:按月而不是按天,旧分区 drop 或 detach,而不是让历史永远在线。详见下一节。 **按库拆分。** 如果租户确实需要硬隔离,每租户一个 database 比每租户一个 schema 更能限制爆炸半径——目录和 relcache 都是按 database 的,连接也是。代价是连接管理和跨租户报表。 **清掉 ORM 忘记删的表。** 在 `pg_stat_user_tables` 的一个观察窗口内找出零扫描、零元组的表并删除;schema 卫生比上面任何一项都便宜。 ## 与分区的边界 [#与分区的边界] 分区不是这个问题的逃生舱——**分区也是表**。每个分区有自己的 `pg_class` 条目、自己在每个触碰它的后端里的 relcache 足迹,通常还有自己的索引。声明式分区只是自动化了路由,元数据代价是累加的。 [分区文档](https://www.postgresql.org/docs/18/ddl-partitioning.html)指出,规划器可以较好地处理几千个分区的层级,*前提是分区剪枝在规划阶段就排除了其中绝大多数*;当大量分区在剪枝后存活时,规划时间和内存消耗都会上升。所以实际的限制是叠加的:一张分了 5,000 个子分区的表、查询不带分区键过滤、再乘上 200 连接的连接池,等于最坏情况的规划代价乘上最坏情况的 relcache 放大。分区数尽量保持在几百以内,查询永远带分区键,让保留策略(`DROP PARTITION` / `DETACH`)——而不是存储容量——来决定多少分区保持在线。 共享表、schema、分区之间如何选择的数据建模视角,见[数据建模](/docs/core/data-modeling)。 relcache 不受 `work_mem` 或 `shared_buffers` 约束;它住在每后端的私有内存里,没有任何配置上限。真正的杠杆只有:更少的表、每个查询触碰更少的表、更少的长命后端,或者更多内存——按这个顺序优先考虑。 先在恢复出来的副本上验证方案——目录手术和 RLS 上线恰恰是最值得演练的变更。 ## 相关页面 [#相关页面] * 行级安全配置 → [安全](/docs/operations/security) * 表结构布局选型 → [数据建模](/docs/core/data-modeling) * 连接与内存预算 → [服务器配置](/docs/operations/configuration) * 当元数据内存撞上 cgroup limit → [容器环境下 PostgreSQL 的内存管理与 OOM 防治](/docs/operations/container-memory-oom) --- # WAIT FOR LSN 与读写分离下的 read-your-writes Canonical URL: https://pg.edu.rich/docs/operations/wait-for-lsn Last reviewed: 2026-08-06 截至 2026-08,PostgreSQL 19 尚未正式发布。下文语法与行为以 [PostgreSQL 19 文档](https://www.postgresql.org/docs/19/sql-wait-for.html)为准;用于生产前请对照正式 release notes 复核。 常见的扩容做法是写走主库、读走异步备库。这个模式最脆弱的一环是写后立即读:客户端已在主库提交,但备库还没 replay 对应的 WAL,紧随其后的读会拿到旧值。这就是 stale read 问题。在 PostgreSQL 19 之前,没有服务端原语能关掉这个窗口,只能靠应用侧的各种绕法。 ## 陈旧读问题 [#陈旧读问题] 异步流复制下,主库的 `COMMIT` 不等待任何备库。WAL 记录要经传输、写入、刷盘、replay 之后在备库可见,正常是毫秒级,负载高时可能到秒级。写完立刻通过备库连接池读,就可能看不到自己刚提交的数据: ```text 主库: UPDATE profile SET display_name = 'new' WHERE id = 42; -- COMMIT 备库: SELECT display_name FROM profile WHERE id = 42; --> 'old' (replay 还没追上这条 commit 记录) ``` 这个窗口没有上界:写入突增、长事务、recovery conflict 都会让复制延迟抖动,任何“备库大概已经追上了”的固定假设迟早会失效。 ## PostgreSQL 19 之前的四种绕法 [#postgresql-19-之前的四种绕法] 四种常见做法都能解决问题,但代价各不相同: * **写后固定 sleep。** 下一次读之前睡 N 毫秒。延迟不是常量:sleep 短了,高峰期仍然读到旧值;sleep 长了,备库空闲时每个请求都白等。 * **Sticky read。** 写之后的一段时间内把该会话的读钉在主库。结果正确,但把读流量推回主库的恰好是最活跃的用户——正是你卸读流量的对象。 * **`synchronous_commit = remote_apply`。** 每次 `COMMIT` 等待同步备库 replay 完成,改动随即处处可见。有效,但所有写入都要付备库往返延迟,且写可用性与备库健康绑定。这是集群级的持久化决策,不是按请求给的一致性提示。 * **轮询 `pg_last_wal_replay_lsn()`。** 写后在主库记录 `pg_current_wal_insert_lsn()`,然后在备库上循环比较 replay 位置,直到越过目标 LSN。语义正确,但每次检查都是一次网络往返,循环占连接,超时和升主处理全靠自己写。 PostgreSQL 19 用服务端阻塞等待替代了最后一种模式。 ## WAIT FOR LSN 语法 [#wait-for-lsn-语法] [`WAIT FOR`](https://www.postgresql.org/docs/19/sql-wait-for.html) 让会话阻塞,直到服务器到达目标 LSN,然后返回一行状态: ```sql WAIT FOR LSN '0/306EE20'; WAIT FOR LSN '0/306EE20' WITH (MODE 'standby_flush'); WAIT FOR LSN '0/306EE20' WITH (MODE 'standby_write', TIMEOUT '100ms', NO_THROW); ``` 返回值有三种:`success`、`timeout`、`not in recovery`。不给 `TIMEOUT`(或给 `0`)则无限等待。超时——或者在非 recovery 状态的服务器上用 standby 模式——会报错,除非指定 `NO_THROW`,此时由返回的 status 列告知结果。 典型的 read-your-writes 流程: ```sql -- 主库,写入提交后立即执行 SELECT pg_current_wal_insert_lsn(); -- 0/306EE20 -> 交给应用 / 连接池保存 -- 备库,执行依赖这次写入的读之前 WAIT FOR LSN '0/306EE20' WITH (TIMEOUT '200ms', NO_THROW); -- status = success -> 该写入在这台备库已可见 SELECT display_name FROM profile WHERE id = 42; ``` 取 LSN 用 `pg_current_wal_insert_lsn()`(insert 位置)而不是 flush 位置,这样即使写入会话使用 `synchronous_commit = off` 流程仍然正确。文档还指出,最后一次修改的 LSN 应保存在客户端应用或连接池一侧——`WAIT FOR` 本身不在语句之间记住任何状态。 ## 四种 MODE 的语义 [#四种-mode-的语义] `MODE` 选择等待 WAL 处理到哪个阶段,默认是 `standby_replay`。 | MODE | 等待到 | 服务器状态 | | ---------------- | ----------------------------------------------------------- | ----- | | `standby_replay` | LSN 在备库 replay(应用)完成;成功后 `pg_last_wal_replay_lsn()` 大于等于目标值 | 仅备库 | | `standby_flush` | WAL 在备库刷盘——拿到持久化保证但不等 apply | 仅备库 | | `standby_write` | WAL 在备库写入操作系统——比 flush 快,持久化保证更弱 | 仅备库 | | `primary_flush` | WAL 在主库刷盘;成功后 `pg_current_wal_flush_lsn()` 大于等于目标值 | 仅主库 | 对 read-your-writes 有意义的只有 `standby_replay`:replay 完成才是行对查询可见的时刻。`standby_write` 和 `standby_flush` 也会被备库上已有的 WAL 满足(来自 base backup 或归档恢复),不限于刚流式收到的部分。在主库上用 standby 模式、或在备库上用 `primary_flush`,都会报错。 ## 限制 [#限制] [命令文档](https://www.postgresql.org/docs/19/sql-wait-for.html)中的限制直接决定调用方式: * **必须是顶级语句。** `WAIT FOR` 不能在函数、存储过程或 `DO` 块里执行,必须由驱动或连接池中间件作为独立语句发出。 * **不能持有快照。** 命令要求当前没有 active 或 registered snapshot,因此不能用在必须保持快照的上下文——包括隔离级别高于 `READ COMMITTED` 的事务。实际写法:作为独立的 autocommit 语句发送,先于需要保证的那次读。 * **升主改变答案。** 等待期间备库被提升,standby 模式会返回 `not in recovery`(未加 `NO_THROW` 则报错)。升主产生新 timeline,你等待的 LSN 可能属于旧 timeline——应用必须重新评估目标是否还有意义。 * **只比较数值。** `WAIT FOR` 比较的是 LSN 数值,不感知 timeline。级联备库的上游被提升后,只要 replay 位置在数值上越过目标就可能返回 `success`,哪怕那个位置在另一条 timeline 上。在乎这个区分就自己校验 timeline。 * **recovery conflict 仍然会打断。** 备库上等待的会话可能被 recovery conflict 处理打断——有些冲突(文档举的例子是 tablespace drop)会无条件终止所有后端。调用方要准备重试和降级路径。 ## 超时后的降级策略 [#超时后的降级策略] 把 `WAIT FOR` 当作有严格预算的一致性增强,而不是正确性原语。始终 `TIMEOUT` 与 `NO_THROW` 成对使用,按 status 分支: * `success`——按计划在备库读。 * `timeout`——降级:把这次读改发主库,或先返回旧值再异步刷新。记录超时时刻的 replay 延迟;超时率上升是备库容量问题的早期信号。 * `not in recovery`——拓扑变了。重新解析主库,必要时在新主库上重取 LSN;旧 LSN 不要跨 timeline 复用。 超时值按读延迟预算定(几十到几百毫秒),而不是按平均延迟定——这个等待本来就是为尾部准备的。如果 `WAIT FOR` 频繁超时,说明该先治理复制延迟本身,见[流复制升级与延迟控制](/docs/operations/replication-upgrades)。 ## 提示词:设计路由层 [#提示词设计路由层] 命令参考:[PostgreSQL 19 WAIT FOR](https://www.postgresql.org/docs/19/sql-wait-for.html)。PostgreSQL 19 其他新特性见 [PostgreSQL 19 总览](/docs/postgresql-19)。实测类文章:[rednafi 的完整走查](https://rednafi.com/system/wait-for-lsn/)与 [digoal 的分析](https://github.com/digoal/blog/blob/master/202601/20260106_01.md)。 --- # 数据建模与约束 Canonical URL: https://pg.edu.rich/docs/core/data-modeling Last reviewed: 2026-08-02 好的 PostgreSQL 模型不是“先建几列,规则以后再补”,而是尽量让数据库知道哪些状态合法。 ## 一份可工作的订单模型 [#一份可工作的订单模型] ```sql CREATE TABLE customers ( id bigint GENERATED ALWAYS AS IDENTITY PRIMARY KEY, email text NOT NULL, display_name text NOT NULL CHECK (length(trim(display_name)) > 0), created_at timestamptz NOT NULL DEFAULT now(), CONSTRAINT customers_email_unique UNIQUE (email) ); CREATE TABLE orders ( id bigint GENERATED ALWAYS AS IDENTITY PRIMARY KEY, customer_id bigint NOT NULL REFERENCES customers(id), status text NOT NULL DEFAULT 'pending' CHECK (status IN ('pending', 'paid', 'shipped', 'cancelled')), total_cents bigint NOT NULL CHECK (total_cents >= 0), placed_at timestamptz NOT NULL DEFAULT now() ); COMMENT ON COLUMN orders.total_cents IS 'Order total in the smallest currency unit; never a floating-point amount.'; ``` ## 类型选择 [#类型选择] | 需求 | 建议类型 | 避免 | | ---- | ------------------------------------------- | --------------------------- | | 主键 | `bigint GENERATED ... AS IDENTITY` 或 `uuid` | 新设计继续依赖 `serial` 的隐式行为 | | 金额 | 最小货币单位的 `bigint`,或明确精度的 `numeric(p,s)` | `real` / `double precision` | | 时间点 | `timestamptz` | 把带时区的现实时间存成字符串 | | 文本 | `text` + 业务约束 | 没有业务意义的任意 `varchar(255)` | | 状态 | 小而稳定时用 `CHECK`;独立生命周期时用引用表 | 无约束自由文本 | | 文档数据 | `jsonb` | 把核心关系和外键藏进 JSON | `timestamptz` 存储绝对时间点,显示时按会话时区转换。它不保存原始输入的时区名称;若业务需要“Europe/Paris”这类规则,另存时区标识。 ## 约束的职责 [#约束的职责] * `NOT NULL`:值必须存在。 * `CHECK`:单行必须满足谓词。 * `UNIQUE`:候选键唯一;默认允许多个 `NULL`。 * `PRIMARY KEY`:唯一且非空的行标识。 * `FOREIGN KEY`:引用目标必须存在;删除策略要显式设计。 `ON DELETE CASCADE` 表示父记录消失时子记录也应消失。只有生命周期确实从属时才使用;账单、审计记录通常不应级联删除。 ## Schema 与命名 [#schema-与命名] 为应用对象使用明确 schema,并收紧默认权限: ```sql CREATE SCHEMA app; REVOKE CREATE ON SCHEMA public FROM PUBLIC; ALTER ROLE app_runtime SET search_path = app, pg_catalog; ``` 对 AI 和人都友好的命名应完整、稳定、少缩写:`customer_id` 优于 `cid`,`created_at` 优于 `ctime`。用 `COMMENT ON` 记录单位、状态转换和隐私等级,而不是复述列名。 ## 验证模型 [#验证模型] ```sql INSERT INTO customers (email, display_name) VALUES ('ada@example.com', 'Ada') RETURNING id, created_at; -- 应失败:金额不能为负 INSERT INTO orders (customer_id, total_cents) VALUES (1, -100); ``` 设计完成的标准不只是“合法数据能写入”,还包括“典型非法数据被正确拒绝”。 --- # SQL/PGQ 图查询 Canonical URL: https://pg.edu.rich/docs/core/graph-queries Last reviewed: 2026-08-06 PostgreSQL 19 实现了 SQL/PGQ(ISO/IEC 9075-16,SQL:2023 第 16 部分):在普通关系表上直接做图模式匹配。属性图本质是现有表之上的只读视图——不装扩展、不复制数据——`GRAPH_TABLE` 查询走的也是普通 JOIN 所用的同一套优化器。 SQL/PGQ 是 PostgreSQL 19 的新特性,PostgreSQL 18 及更早版本没有 `CREATE PROPERTY GRAPH` 和 `GRAPH_TABLE`。版本状态见 [PostgreSQL 19 发布与升级指南](/docs/postgresql-19);本文语法以官方文档为准([5.15 Property Graphs](https://www.postgresql.org/docs/19/ddl-property-graphs.html)、[7.9 Graph Queries](https://www.postgresql.org/docs/19/queries-graph.html))。 ## 什么场景值得用图查询 [#什么场景值得用图查询] * 社交关系:朋友的朋友、共同好友。 * 数据血缘:报表里的数字由哪些源数据、经过哪些转换算出来。 * 风控与审计:资金流向、可疑路径、合规追溯。 取舍在于深度固定。PostgreSQL 19 的 SQL/PGQ 按跳逐段匹配模式,不支持变长路径。几跳之内,模式语法比等价的 JOIN 链清晰得多;超出这个范围就用 `WITH RECURSIVE`(见[边界与 WITH RECURSIVE](#边界与-with-recursive))。 ## 定义图:CREATE PROPERTY GRAPH [#定义图create-property-graph] 经典的社交网络模型:`person` 和 `knows`。 ```sql CREATE TABLE person ( id int PRIMARY KEY, name text NOT NULL, age int, city text ); CREATE TABLE knows ( a int NOT NULL REFERENCES person(id), -- 认识谁 b int NOT NULL REFERENCES person(id), -- 被谁认识 since int, PRIMARY KEY (a, b) ); ``` 把它声明成图——顶点是 `person`,边是 `knows`,方向从 `a` 指向 `b`: ```sql CREATE PROPERTY GRAPH social VERTEX TABLES ( person KEY (id) LABEL person PROPERTIES (id, name, age, city) ) EDGE TABLES ( knows SOURCE KEY (a) REFERENCES person (id) DESTINATION KEY (b) REFERENCES person (id) LABEL knows PROPERTIES (since) ); ``` 把它当一份契约读: * 顶点表的 `KEY` 通常就是主键;顶点表必须有主键。 * `SOURCE KEY ... REFERENCES ...` 和 `DESTINATION KEY ... REFERENCES ...` 声明边的方向:从哪列出发、到哪张表的哪列。 * `LABEL` 是图里的名字(可以与表名不同);`PROPERTIES` 决定哪些列在图中可见。如果表名、列名本身就是想要的标签和属性,这两个子句都可以省略。 `CREATE PROPERTY GRAPH` 只存元数据。数据仍在原表,图的创建和删除都不动业务数据,定义丢了可以随时重建。 ## 用 GRAPH\_TABLE 查询 [#用-graph_table-查询] 列出所有人: ```sql SELECT name FROM GRAPH_TABLE (social MATCH (p IS person) COLUMNS (p.name) ) ORDER BY name; ``` 逻辑上等价于 `SELECT name FROM person ORDER BY name`。`MATCH (p IS person)` 遍历每个 `person` 顶点并绑定为 `p`;`COLUMNS` 投影输出列。`GRAPH_TABLE` 的结果就是一张普通表:可以起别名、过滤、和其他 `FROM` 项做 JOIN。 `COLUMNS (p.*)` 会报错 `"*" is not supported here`。输出列必须显式列出,例如 `COLUMNS (p.id, p.name)`。 ### 边模式与方向 [#边模式与方向] ```sql SELECT * FROM GRAPH_TABLE (social MATCH (p IS person)-[IS knows]->(p2 IS person) COLUMNS (p.id, p.name, p2.id, p2.name) ) ORDER BY 1, 2, 3; ``` `(p)-[IS knows]->(p2)` 是边模式: * `->` 沿声明方向走(SOURCE → DESTINATION);`<-` 反向。 * 裸 `-` 两个方向都匹配。 裸 `-` 表示两个方向任一匹配(相当于两个方向的 OR),同一对关系会出现两次。除非底表数据本身对称(每个方向各插了一行),否则明确写 `->` 或 `<-`。 ### 多跳 [#多跳] ```sql SELECT * FROM GRAPH_TABLE (social MATCH (a IS person)-[IS knows]-> (b IS person)-[IS knows]->(c IS person) WHERE a.id <> c.id COLUMNS (a.name AS a, b.name AS via, c.name AS c) ) ORDER BY a, c, via; ``` * 每多一跳,就在链上多写一段 `-[IS knows]->(...)`。 * `WHERE a.id <> c.id` 写在 `MATCH` 里,过滤掉 "Alice → Bob → Alice" 这类回路。去掉它就能看到包含回路在内的所有两跳路径。 ## 多类顶点与边 [#多类顶点与边] 真实模型通常横跨多张表。加上公司和雇佣关系: ```sql CREATE TABLE company ( id int PRIMARY KEY, name text NOT NULL, industry text NOT NULL ); CREATE TABLE works_at ( pid int NOT NULL REFERENCES person(id), cid int NOT NULL REFERENCES company(id), role text NOT NULL, PRIMARY KEY (pid, cid) ); CREATE PROPERTY GRAPH company_social VERTEX TABLES ( person KEY (id) LABEL person PROPERTIES (id, name, age, city), company KEY (id) LABEL company PROPERTIES (id, name, industry) ) EDGE TABLES ( knows SOURCE KEY (a) REFERENCES person (id) DESTINATION KEY (b) REFERENCES person (id) LABEL knows PROPERTIES (since), works_at SOURCE KEY (pid) REFERENCES person (id) DESTINATION KEY (cid) REFERENCES company (id) LABEL works_at PROPERTIES (role) ); ``` 一个图包含两类顶点、两类边,一条查询横跨它们——"Alice 的朋友都在哪上班?": ```sql SELECT * FROM GRAPH_TABLE (company_social MATCH (me IS person WHERE me.name = 'Alice') -[IS knows]->(friend IS person) -[IS works_at]->(co IS company) COLUMNS (friend.name AS friend, co.name AS company) ) ORDER BY friend, company; ``` `WHERE` 可以直接写在顶点模式里。等价的普通 SQL 要写多表 JOIN;图语法把"沿哪类关系走"直接编进了模式。 ### 没有多模式 MATCH:JOIN 两个 GRAPH\_TABLE [#没有多模式-matchjoin-两个-graph_table] SQL 标准里逗号分隔的 `MATCH (a...), (b...)` 在 PostgreSQL 19 尚未实现。变通方案是本页最实用的写法:`GRAPH_TABLE` 的结果是普通表,把 ID 投影出来再 JOIN: ```sql SELECT m.me, m.via, m.coworker, w.company FROM GRAPH_TABLE (company_social MATCH (a IS person)-[IS knows]->(b IS person)-[IS knows]->(c IS person) WHERE a.id <> c.id COLUMNS (a.id AS aid, c.id AS cid, a.name AS me, b.name AS via, c.name AS coworker) ) m JOIN GRAPH_TABLE (company_social MATCH (x IS person)-[IS works_at]->(co IS company) <-[IS works_at]-(y IS person) WHERE x.id <> y.id COLUMNS (x.id AS xid, y.id AS yid, co.name AS company) ) w ON w.xid = m.aid AND w.yid = m.cid ORDER BY me, coworker; ``` "互相认识的同事"这类问题就这么拆:一个 `GRAPH_TABLE` 解决一段图模式,ID 做桥接。 只有一个边表的图里,`(a)->(b)` 没有歧义;但在 `company_social` 这种多边图里,匿名边会静默地 union 所有边类型——`knows` 和 `works_at` 的行一起出来。图里只要有多个边表,就写明边标签(`[IS knows]`、`[IS works_at]`)。 ## 性能:它就是按 JOIN 规划的 [#性能它就是按-join-规划的] 发布说明明确写到 SQL/PGQ 查询"像视图一样处理,被写成标准关系查询"。两跳模式展开成多路 JOIN 后走普通优化器,`EXPLAIN` 里没有任何图专用执行节点。这正是它可预测的原因:计划形状由模式形状直接决定,索引规则和你熟悉的 JOIN 完全一样。 真正要记住的规则只有一条:**边表主键覆盖源端;要做反向遍历,就给目标端列补索引。** * 正向("42 认识谁")按 `knows.a = 42` 查找,主键 `(a, b)` 直接覆盖。 * 反向("谁认识 42")按 `knows.b` 过滤,这个主键帮不上忙——扫描退化为读整张边表,表越大差距越大。 ```sql CREATE INDEX ON knows (b); ``` 用 `EXPLAIN (ANALYZE, BUFFERS)` 验证,方式和普通 JOIN 完全一样,见[索引与 EXPLAIN](/docs/core/indexes-explain)。 ## 数据血缘查询模式 [#数据血缘查询模式] 合规场景的经典问题是"报表里这个数字是怎么来的?"。把 ETL 管线建模成图: * 每层一张顶点表:`clickstream_source`(原始)→ `staging` → `fact` → `report`,各有 `id` 主键和 `name`。 * 每种转换类型一张边表,`PRIMARY KEY (src, dst)`,外键指向它连接的两层:`loads_into`(直拷)、`aggregates_into`(分组聚合)、`rollup_into`(进一步汇总)。 ```sql CREATE PROPERTY GRAPH lineage VERTEX TABLES ( clickstream_source KEY (id), staging KEY (id), fact KEY (id), report KEY (id) ) EDGE TABLES ( loads_into SOURCE KEY (src) REFERENCES clickstream_source (id) DESTINATION KEY (dst) REFERENCES staging (id), aggregates_into SOURCE KEY (src) REFERENCES staging (id) DESTINATION KEY (dst) REFERENCES fact (id), rollup_into SOURCE KEY (src) REFERENCES fact (id) DESTINATION KEY (dst) REFERENCES report (id) ); ``` 边标签带语义,这是它比"外键加递归 CTE"强的地方:`rollup_into` 告诉你*发生了哪种转换*,而外键只告诉你*存在关系*。`CREATE PROPERTY GRAPH` 语句本身可以从编排工具的元数据生成(dbt、Airflow 这类工具本来就记录上下游关系)。 四个高频查询模式: ```sql -- 1. 值溯源:这个报表数字是谁喂出来的 SELECT * FROM GRAPH_TABLE (lineage MATCH (r IS report)<-[IS rollup_into]-(f IS fact) COLUMNS (r.name AS report, f.name AS fact) ); -- 2. 下游影响:改了这个源,哪些报表会出问题? SELECT DISTINCT rpt FROM GRAPH_TABLE (lineage MATCH (s IS clickstream_source WHERE s.name = 'events_raw') -[IS loads_into]->(IS staging) -[IS aggregates_into]->(IS fact) -[IS rollup_into]->(r IS report) COLUMNS (r.name AS rpt) ); -- 3. 全链路审计:MATCH 同 2,COLUMNS 把每一跳都输出 -- (s.name、staging 名、fact 名、r.name) -- 4. 黑洞:没有下游消费者的源数据 SELECT src.name FROM GRAPH_TABLE (lineage MATCH (s IS clickstream_source) COLUMNS (s.id AS sid, s.name AS name) ) src WHERE NOT EXISTS ( SELECT 1 FROM loads_into l WHERE l.src = src.sid ); ``` 模式 4 展示了退路:`GRAPH_TABLE` 的结果可以和普通 SQL 自由组合,反连接、聚合、`EXISTS` 都能用在图输出之上。 ## 边界与 WITH RECURSIVE [#边界与-with-recursive] PostgreSQL 19 的 SQL/PGQ 还没有: * **变长路径**:`(a)-[IS knows]->{1,3}(b)` 会被拒绝。要么逐跳展开后 `UNION`,要么退回 `WITH RECURSIVE`。 * **多模式 MATCH**(逗号分隔)——用两个 `GRAPH_TABLE` JOIN 替代,如上文所示。 * **路径变量绑定**(`p = (a)->(b)`)和 `ANY SHORTEST` / `ALL SHORTEST` 最短路。 超过五跳左右的深链,递归 CTE 仍然是最顺手的工具: ```sql WITH RECURSIVE chain AS ( SELECT a, b, 1 AS depth FROM knows WHERE a = 42 UNION ALL SELECT k.a, k.b, c.depth + 1 FROM chain c JOIN knows k ON k.a = c.b WHERE c.depth < 5 ) SELECT * FROM chain; ``` 更多 CTE 写法见[查询工具箱](/docs/core/queries)。 ## AI prompt:让 AI 起草图查询 [#ai-prompt让-ai-起草图查询] ## 相关页面 [#相关页面] * [PostgreSQL 19 发布与升级指南](/docs/postgresql-19) —— SQL/PGQ 的版本边界 * [索引与 EXPLAIN](/docs/core/indexes-explain) —— 验证反向遍历索引 * [查询工具箱](/docs/core/queries) —— CTE 与 `WITH RECURSIVE` --- # 索引与 EXPLAIN Canonical URL: https://pg.edu.rich/docs/core/indexes-explain Last reviewed: 2026-08-06 ## 先获取计划 [#先获取计划] ```sql EXPLAIN (ANALYZE, BUFFERS, VERBOSE) SELECT id, customer_id, placed_at FROM orders WHERE customer_id = 42 ORDER BY placed_at DESC LIMIT 20; ``` * `EXPLAIN` 只展示估算,不执行语句。 * `ANALYZE` 会真实执行并给出实际行数和耗时。对写语句使用时,应包在事务中并回滚。 * `BUFFERS` 展示 shared/local/temp block 的命中与读取。 * 重点比较 `rows` 与 `actual rows`、循环次数、最贵节点和是否发生磁盘排序。 `EXPLAIN ANALYZE DELETE ...` 真的会删除。需要检查写语句时使用 `BEGIN; EXPLAIN (ANALYZE, BUFFERS) ...; ROLLBACK;`,并确认没有不可回滚的外部副作用。 ## 为查询形状建索引 [#为查询形状建索引] 上面的过滤与排序可以使用: ```sql CREATE INDEX CONCURRENTLY orders_customer_placed_idx ON orders (customer_id, placed_at DESC) INCLUDE (id); ``` 复合 B-tree 通常从最左列开始匹配。列顺序由实际谓词、范围条件与排序共同决定,不是简单地把“区分度最高”放最前。 `INCLUDE` 列不参与搜索顺序,但可能允许 index-only scan;是否真正只读索引还取决于可见性图。 ## 常用索引类型 [#常用索引类型] | 类型 | 适合 | | --------------- | ------------------------------------------------------------------- | | B-tree | 等值、范围、排序;默认选择 | | GIN | `jsonb` 包含、数组成员、全文检索 | | GiST | 几何、范围、某些扩展运算符 | | SP-GiST | trie、quad-tree、k-d tree 等可分区搜索结构 | | BRIN | 物理顺序与值强相关的超大表,如按时间追加日志 | | Hash | 仅等值;通常 B-tree 更通用 | | Bloom extension | 多字段任意组合的等值过滤;lossy 且需要 recheck;自带 operator class 仅有 `int4` 与 `text` | 这些名称仍不足以决定索引是否可用:operator class 决定具体 operator 与数据类型。Table Access Method、HNSW/IVFFlat 和更完整的选择图见 [索引与存储访问方法](/docs/reference/index-access-methods)。 ## 两种高价值索引 [#两种高价值索引] 部分索引只覆盖相关行: ```sql CREATE INDEX orders_unfinished_idx ON orders (placed_at) WHERE status IN ('pending', 'paid'); ``` 表达式索引加速规范化查找: ```sql CREATE UNIQUE INDEX customers_email_ci_idx ON customers (lower(email)); ``` 查询谓词需要与表达式或部分条件相匹配,优化器才能使用它们。 ## 为什么没有走索引 [#为什么没有走索引] * 表很小,顺序扫描更便宜。 * 查询要返回很大比例的行。 * 统计信息陈旧或列之间相关性未被描述。 * 对列套了与索引不匹配的函数或隐式转换。 * 复合索引的最左列不适用。 * 成本参数与真实存储特征不匹配。 先运行 `ANALYZE orders;` 并检查估算偏差,不要第一反应关闭顺序扫描。 ## 生产创建与清理 [#生产创建与清理] `CREATE INDEX CONCURRENTLY` 减少对写入的阻塞,但耗时更长、不能放在事务块内,失败时可能留下 invalid index。使用: ```sql SELECT indexrelid::regclass, indisvalid, indisready FROM pg_index WHERE indrelid = 'orders'::regclass; ``` 每个索引都会增加写放大、WAL、缓存压力和 vacuum 工作量。定期结合 `pg_stat_user_indexes` 与业务周期审查未使用索引。 ## 跨版本对比计划 [#跨版本对比计划] PostgreSQL 19(截至本页更新仍为 Beta)为 `EXPLAIN ANALYZE` 新增 `IO` 选项,报告扫描节点的异步 I/O 行为:prefetch 队列的距离与容量,以及已发出 I/O 请求的数量、平均大小、I/O 等待次数和并发度。`EXPLAIN (ANALYZE, WAL)` 还会报告 full-page write bytes。两者让存储行为无需外部工具即可观测;见 [PostgreSQL 19 EXPLAIN 文档](https://www.postgresql.org/docs/19/sql-explain.html)。 19 同时把 JIT 改为默认关闭,因为其成本模型被证明不可靠。在 18 与 19 之间对比计划或耗时,应在两侧显式设置 `jit`——否则看到的差异可能只是默认值变化,而不是优化器或索引的变化。 --- # JSONB、全文与语义检索 Canonical URL: https://pg.edu.rich/docs/core/jsonb-search Last reviewed: 2026-08-02 PostgreSQL 可以在同一系统里处理关系数据、JSON 文档、词法全文检索,并通过扩展进行向量检索。能放在一起不代表应该把所有问题都塞进一列。 ## 何时使用 JSONB [#何时使用-jsonb] 适合:来源不统一的元数据、低频变化的可选属性、需要保留原始载荷的集成数据。 不适合:主键与外键、金额、权限边界、经常 JOIN/排序/聚合的核心字段。 ```sql CREATE TABLE products ( id bigint GENERATED ALWAYS AS IDENTITY PRIMARY KEY, sku text NOT NULL UNIQUE, name text NOT NULL, attributes jsonb NOT NULL DEFAULT '{}'::jsonb, CHECK (jsonb_typeof(attributes) = 'object') ); INSERT INTO products (sku, name, attributes) VALUES ('KB-01', 'Keyboard', '{"layout":"75%","wireless":true}'); SELECT id, name FROM products WHERE attributes @> '{"wireless":true}'; ``` ## JSONB 索引 [#jsonb-索引] ```sql CREATE INDEX products_attributes_gin ON products USING gin (attributes); ``` 默认 GIN operator class 支持多种键与包含查询。若工作负载几乎只有 `@>`,`jsonb_path_ops` 通常索引更小,但支持的运算符集合更窄。用真实查询和数据分布比较。 频繁查询的单个属性可以使用表达式索引,或提升为普通列: ```sql CREATE INDEX products_layout_idx ON products ((attributes ->> 'layout')); ``` ## 内置全文检索 [#内置全文检索] ```sql ALTER TABLE products ADD COLUMN search_document tsvector GENERATED ALWAYS AS ( setweight(to_tsvector('simple', coalesce(name, '')), 'A') || setweight(to_tsvector('simple', coalesce(attributes::text, '')), 'B') ) STORED; CREATE INDEX products_search_gin ON products USING gin (search_document); SELECT id, name, ts_rank(search_document, websearch_to_tsquery('simple', $1)) AS rank FROM products WHERE search_document @@ websearch_to_tsquery('simple', $1) ORDER BY rank DESC LIMIT 20; ``` 中文分词不由内置 `simple` 配置完整解决;生产中文搜索需要评估专用分词扩展、应用侧分词或外部搜索系统。 ## 语义检索与 pgvector [#语义检索与-pgvector] 向量不是 PostgreSQL 核心内置类型。常见方案是安装独立的 `pgvector` 扩展: ```sql CREATE EXTENSION IF NOT EXISTS vector; CREATE TABLE document_chunks ( id bigint GENERATED ALWAYS AS IDENTITY PRIMARY KEY, document_id bigint NOT NULL, content text NOT NULL, embedding vector(1536) NOT NULL, embedding_model text NOT NULL ); ``` 维度必须匹配模型;切换 embedding 模型时不要在同一索引里混用不可比较的向量。保存模型名、切块版本和源文档定位,才能重建与审计。 用关键词/权限/时间等结构化条件缩小候选集,再做向量相似度排序。始终在 SQL 中执行租户和访问控制过滤,不要依赖模型自行遵守。 需要可运行环境时先阅读[安装 pgvector](/docs/ai/pgvector-setup);准备建立近似索引前阅读[pgvector 生产最佳实践](/docs/ai/vector-production)。 --- # PostgreSQL 锁等待与死锁排查 Canonical URL: https://pg.edu.rich/docs/core/locks-deadlocks Last reviewed: 2026-08-02 ## 锁等待和死锁不同 [#锁等待和死锁不同] * **锁等待**:会话等待另一个事务释放冲突锁;可能最终成功,也可能超时。 * **死锁**:形成等待环,任何参与者都无法自行前进;PostgreSQL 检测后中止其中一个事务。 死锁失败的 SQLSTATE 是 `40P01`。当前事务必须回滚;只有整个业务事务可安全重放时才做有界重试。 ## 查看阻塞链 [#查看阻塞链] ```sql SELECT blocked.pid AS blocked_pid, blocker.pid AS blocker_pid, now() - blocked.query_start AS blocked_for, blocked.wait_event_type, blocked.wait_event, left(blocked.query, 120) AS blocked_query, left(blocker.query, 120) AS blocker_query FROM pg_stat_activity AS blocked CROSS JOIN LATERAL unnest(pg_blocking_pids(blocked.pid)) AS b(pid) JOIN pg_stat_activity AS blocker ON blocker.pid = b.pid ORDER BY blocked.query_start; ``` 先检查 blocker 是否处于 `idle in transaction`、事务包含什么修改、应用是否仍存活。不要看到 PID 就立即终止。 ## 减少死锁 [#减少死锁] 1. 所有代码路径按相同顺序锁定资源,例如总按 account id 升序。 2. 事务只包含必须原子完成的数据库工作,不跨用户输入和外部 API。 3. 为定位要更新行的条件建立合适索引,减少锁定/扫描范围。 4. 对批量任务分块,并避免多个任务交叉处理相同键空间。 5. 设置有业务依据的 `lock_timeout` 和 `statement_timeout`。 ```sql BEGIN; SET LOCAL lock_timeout = '1s'; SET LOCAL statement_timeout = '10s'; SELECT id FROM accounts WHERE id = ANY($1::bigint[]) ORDER BY id FOR UPDATE; -- bounded writes COMMIT; ``` ## 终止会话前 [#终止会话前] `pg_cancel_backend(pid)` 请求取消当前语句;`pg_terminate_backend(pid)` 终止会话并回滚其事务。执行前确认: * PID 仍属于目标会话,避免使用陈旧截图; * 回滚可能耗时和产生额外 I/O; * 应用不会立即以相同方式重连并再次阻塞; * 被中止工作是否可重试、是否需要业务补偿。 完整锁模式与冲突矩阵见 [PostgreSQL 18 显式锁定](https://www.postgresql.org/docs/18/explicit-locking.html)。 --- # PostgreSQL MVCC 与快照可见性 Canonical URL: https://pg.edu.rich/docs/core/mvcc-snapshots Last reviewed: 2026-08-02 MVCC(多版本并发控制)让普通读取通常不阻塞写入。`UPDATE` 不会就地覆盖所有读者看到的值,而会产生新行版本;每个查询按自己的 snapshot 判断哪个版本可见。 ## 两个会话观察快照 [#两个会话观察快照] 先准备数据: ```sql CREATE TABLE mvcc_demo ( id integer PRIMARY KEY, value text NOT NULL ); INSERT INTO mvcc_demo VALUES (1, 'before'); ``` 会话 A: ```sql BEGIN ISOLATION LEVEL REPEATABLE READ; SELECT value FROM mvcc_demo WHERE id = 1; -- before ``` 会话 B: ```sql UPDATE mvcc_demo SET value = 'after' WHERE id = 1; COMMIT; ``` 回到会话 A: ```sql SELECT value FROM mvcc_demo WHERE id = 1; -- 仍是 before COMMIT; SELECT value FROM mvcc_demo WHERE id = 1; -- after ``` `READ COMMITTED` 则在每条语句开始时取得新 snapshot,所以同一事务里的第二次查询可能看到会话 B 已提交的值。 ## 为什么长事务危险 [#为什么长事务危险] 只要旧 snapshot 仍可能看到某些行版本,vacuum 就不能把它们当作完全可回收。长事务因此会扩大: * dead tuples 与表/索引膨胀; * vacuum 工作量和磁盘占用; * 复制槽、逻辑解码或 standby 的保留压力; * transaction ID wraparound 风险窗口。 查找持有旧事务或 snapshot 的会话: ```sql SELECT pid, usename, application_name, state, now() - xact_start AS transaction_age, age(backend_xmin) AS snapshot_xid_age, wait_event_type, wait_event, left(query, 120) AS query FROM pg_stat_activity WHERE xact_start IS NOT NULL OR backend_xmin IS NOT NULL ORDER BY xact_start NULLS LAST; ``` 不要仅凭“时间长”终止会话;先确认业务、是否在执行备份/维护、事务能否安全重试及终止影响。 行版本解决读取可见性,写写冲突、DDL、外键检查和显式锁仍会等待。下一步阅读[锁等待与死锁](/docs/core/locks-deadlocks)。 权威行为见 [PostgreSQL 18 并发控制](https://www.postgresql.org/docs/18/mvcc.html)和[事务隔离](https://www.postgresql.org/docs/18/transaction-iso.html)。 --- # 查询工具箱 Canonical URL: https://pg.edu.rich/docs/core/queries Last reviewed: 2026-08-06 ## 可维护查询的基本形状 [#可维护查询的基本形状] ```sql SELECT o.id, c.email, o.total_cents, o.placed_at FROM orders AS o JOIN customers AS c ON c.id = o.customer_id WHERE o.status = $1 AND o.placed_at >= $2 ORDER BY o.placed_at DESC, o.id DESC LIMIT $3; ``` 这条查询明确了输出、连接条件、参数、稳定排序和上限。`$1`、`$2`、`$3` 由驱动绑定;不要用字符串拼接用户输入。 ## JOIN 的决策 [#join-的决策] | 目标 | 使用 | | --------- | -------------------------------------- | | 只保留两边匹配行 | `INNER JOIN` / `JOIN` | | 保留左表全部行 | `LEFT JOIN` | | 判断相关行是否存在 | `EXISTS`,常比“JOIN 后 DISTINCT”更清楚 | | 找没有关联行的数据 | `NOT EXISTS`,避免 `NOT IN` 的 `NULL` 语义陷阱 | ```sql SELECT c.id, c.email FROM customers AS c WHERE NOT EXISTS ( SELECT 1 FROM orders AS o WHERE o.customer_id = c.id ); ``` ## 聚合与窗口不是一回事 [#聚合与窗口不是一回事] `GROUP BY` 把多行折叠为一行;窗口函数保留明细行,同时在窗口内计算。 ```sql SELECT customer_id, id AS order_id, total_cents, row_number() OVER ( PARTITION BY customer_id ORDER BY placed_at DESC, id DESC ) AS recency_rank, sum(total_cents) OVER (PARTITION BY customer_id) AS lifetime_cents FROM orders; ``` ## CTE 的用途 [#cte-的用途] CTE 应为复杂查询命名阶段,而不是自动优化按钮。 ```sql WITH recent_paid AS ( SELECT customer_id, total_cents FROM orders WHERE status = 'paid' AND placed_at >= now() - interval '30 days' ) SELECT customer_id, sum(total_cents) AS paid_cents FROM recent_paid GROUP BY customer_id; ``` ## 分页 [#分页] 大结果集优先 keyset pagination: ```sql SELECT id, placed_at, total_cents FROM orders WHERE (placed_at, id) < ($1, $2) ORDER BY placed_at DESC, id DESC LIMIT 50; ``` 与很大的 `OFFSET` 相比,它不需要不断跳过前面的行,并在并发写入时更稳定。游标必须包含排序的全部键。 ## PostgreSQL 19 语法便利项 [#postgresql-19-语法便利项] PostgreSQL 19(截至本页更新仍为 Beta)计划了几项小型语法新增。不要把这些语法发给更旧的服务器;以最终 release notes 为准。 `GROUP BY ALL` 自动按 target list 中所有非 aggregate、非 window 项分组,不必重复书写分组列表: ```sql SELECT customer_id, status, count(*) FROM orders GROUP BY ALL; ``` 窗口函数支持 `IGNORE NULLS` / `RESPECT NULLS`,适用于 `lead()`、`lag()`、`first_value()`、`last_value()` 和 `nth_value()`: ```sql SELECT customer_id, placed_at, lag(placed_at) IGNORE NULLS OVER ( PARTITION BY customer_id ORDER BY placed_at ) AS previous_order_at FROM orders; ``` `INSERT ... ON CONFLICT DO SELECT ... RETURNING` 把 get-or-create 变成单条原子语句:要么插入新行,要么返回发生冲突的已有行。`DO SELECT` 必须提供 `conflict_target` 和 `RETURNING` 子句;可选的锁定子句(`FOR UPDATE`、`FOR NO KEY UPDATE`、`FOR SHARE`、`FOR KEY SHARE`)会对冲突行加锁,防止并发更新。 ```sql INSERT INTO customers (email) VALUES ($1) ON CONFLICT (email) DO SELECT RETURNING id; ``` 见 [PostgreSQL 19 INSERT 文档](https://www.postgresql.org/docs/19/sql-insert.html)。 ## 写查询前的检查 [#写查询前的检查] * 输出列是否稳定并最小化? * 每个 JOIN 是否可能放大行数? * `NULL` 的含义是否明确? * 排序是否有唯一的最终 tie-breaker? * 参数是否由驱动绑定? * 是否需要超时和结果行数上限? --- # 事务、MVCC 与并发 Canonical URL: https://pg.edu.rich/docs/core/transactions Last reviewed: 2026-08-02 ## 最小事务 [#最小事务] ```sql BEGIN; SELECT balance_cents FROM accounts WHERE id = $1 FOR UPDATE; UPDATE accounts SET balance_cents = balance_cents - $2 WHERE id = $1 AND balance_cents >= $2; COMMIT; ``` 事务把多条语句组成一个原子单元。`FOR UPDATE` 锁住选中的行,直到提交或回滚;业务仍应检查 `UPDATE` 的影响行数。 ## MVCC 的直觉 [#mvcc-的直觉] PostgreSQL 通过多版本并发控制让读取通常不阻塞写入、写入通常不阻塞普通读取。更新会创建新行版本;旧版本在对所有活跃快照都不可见后,由 autovacuum 回收空间。 长事务会让旧版本迟迟不能回收,并增加表膨胀、WAL 保留和复制延迟风险。不要把事务跨越用户思考、网络重试或外部 API 调用。 ## 隔离级别 [#隔离级别] | 级别 | PostgreSQL 行为 | 应用责任 | | ----------------- | ---------------- | ------------------ | | `READ COMMITTED` | 默认;每条语句获取新快照 | 不假设同一事务两次查询结果不变 | | `REPEATABLE READ` | 事务快照稳定;可能序列化失败 | 捕获 `40001` 并重试整个事务 | | `SERIALIZABLE` | 只允许可证明等价于串行执行的结果 | 必须设计整事务重试与退避 | PostgreSQL 的 `READ UNCOMMITTED` 实际按 `READ COMMITTED` 处理。 与 SQL 标准最低要求相比,PostgreSQL 的 `REPEATABLE READ` 还会阻止 phantom read,但仍可能出现需要整事务重试的序列化失败。另一个重要边界是 sequence:`nextval()` 等变更不会因事务回滚而撤销,因此序号存在空洞是正常现象。 ## 重试的正确边界 [#重试的正确边界] 序列化失败或死锁后,当前事务已经不能继续。应用应回滚并重放**整个事务**,不是只重跑最后一条 SQL。 ```text begin run all reads and writes commit on SQLSTATE 40001 or 40P01 rollback retry whole unit with bounded exponential backoff ``` 确保重试边界内的外部副作用是幂等的,或把它们放到提交后的 outbox 消费流程。 ## 死锁与锁等待 [#死锁与锁等待] 降低死锁概率:所有事务按相同顺序锁定资源;保持事务短小;为定位条件建立合适索引;为请求设置合理的 `lock_timeout` 和 `statement_timeout`。 ```sql SET LOCAL lock_timeout = '2s'; SET LOCAL statement_timeout = '10s'; ``` `SET LOCAL` 只在当前事务生效。 客户端开启事务后不提交,会继续持有快照甚至锁。监控 `pg_stat_activity.state = 'idle in transaction'`,并考虑设置 `idle_in_transaction_session_timeout`。 隔离级别和允许现象以 [PostgreSQL 18 事务隔离文档](https://www.postgresql.org/docs/18/transaction-iso.html)为准。 继续深入:[MVCC 与快照可见性](/docs/core/mvcc-snapshots)解释旧行版本和长事务;[锁等待与死锁](/docs/core/locks-deadlocks)提供阻塞链诊断与安全处理流程。 --- # PostgreSQL 连接错误排查 Canonical URL: https://pg.edu.rich/docs/reference/connection-errors Last reviewed: 2026-08-02 连接失败发生在 SQL 执行之前时,客户端不一定能得到 SQLSTATE。保留完整错误文本、时间、客户端版本和目标 host/port,但不要记录密码。 ## 固定诊断顺序 [#固定诊断顺序] ```text DNS 解析 → TCP 路由/防火墙/监听端口 → TLS 协商与证书身份 → pg_hba.conf 匹配 → 用户认证 → database 与 CONNECT 权限 → 连接数/池容量 → 会话初始化参数 ``` 跳过前一层直接重置密码或放宽权限,通常会掩盖真实问题。 ## 高频错误 [#高频错误] | 错误 | 含义 | 验证 | | -------------------------------- | ---------------------------------- | ------------------------------------- | | `could not translate host name` | DNS/主机名无法解析 | `getent hosts`、`nslookup`,核对拼写和私网 DNS | | `connection refused` | 目标地址没有接受该端口 | 服务状态、`listen_addresses`、端口和容器映射 | | `connection timed out` | 网络路径或防火墙丢弃 | 从同一应用环境测试 TCP,不从个人电脑代替 | | `no pg_hba.conf entry` | 没有匹配来源/database/user/TLS 的 HBA 规则 | 查看服务端日志和规则顺序;修改后 reload | | `password authentication failed` | 凭据或认证方式不匹配,常见 SQLSTATE `28P01` | 确认目标实例和 user,安全轮换密码 | | `database ... does not exist` | 目标实例中没有该 database,SQLSTATE `3D000` | 连接 `postgres` 后查询 `pg_database` | | `too many connections` | 实例/角色/数据库连接上限耗尽,SQLSTATE `53300` | `pg_stat_activity`、连接池与保留管理连接 | | `certificate verify failed` | CA、主机名、有效期或证书链错误 | 检查 `sslmode`、URI host、CA 和平台轮换通知 | ## 客户端验证 [#客户端验证] ```bash psql --version psql -X "postgresql://app_reader@db.example.com:5432/commerce?sslmode=verify-full" ``` 连接成功后立即运行: ```sql \conninfo SELECT current_database(), current_user, inet_server_addr(), inet_server_port(), current_setting('server_version'); ``` ## 服务端最小检查 [#服务端最小检查] ```sql SELECT datname, datallowconn, datconnlimit FROM pg_database ORDER BY datname; SELECT usename, application_name, client_addr, state, count(*) FROM pg_stat_activity GROUP BY usename, application_name, client_addr, state ORDER BY count(*) DESC; ``` 需要操作系统权限时再检查监听 socket、防火墙和 PostgreSQL 日志。云数据库没有主机权限,应使用平台连接诊断、网络流日志和审计日志。 把 `pg_hba.conf` 改为 `trust` 会移除认证边界,并且不能证明原密码为何失败。应在受控渠道轮换凭据,核对匹配到的 HBA 规则和服务端日志。 成功连接的安全配置见 [`psql` 与 SSL](/docs/setup/psql-connection);SQL 执行错误见 [SQLSTATE 速查](/docs/reference/errors)。 --- # 编辑、事实核对与更正政策 Canonical URL: https://pg.edu.rich/docs/reference/editorial-policy Last reviewed: 2026-08-02 ## 内容责任 [#内容责任] PostgreSQL Field Guide 是独立社区知识库,与 PostgreSQL Global Development Group 及文中云厂商无隶属关系。本站提供学习路径、工程解释和可验证示例;规范行为最终以目标版本的 PostgreSQL 官方文档为准。 ## 来源优先级 [#来源优先级] 1. PostgreSQL 当前受支持版本的官方手册、release notes 与 versioning 页面。 2. 扩展的上游仓库和发行说明,例如 pgvector 官方仓库。 3. 云服务商针对具体产品、区域和引擎版本的官方文档。 4. 可复现的本地测试与公开技术标准。 社区文章可以帮助发现问题,但不会单独支撑版本、接口、安全或恢复结论。 ## 发布检查 [#发布检查] * SQL 名称、系统视图和参数先查目标版本文档,不凭记忆补接口。 * 可安全执行的示例尽量在 PostgreSQL 18 临时实例验证。 * 写入、锁表、恢复、权限与复制操作明确风险和前置条件。 * 中英文页面同路径发布;任一语言不得是空占位或机器直译草稿。 * 云产品能力记录核对日期,并提醒读者重新确认区域、SKU 与扩展版本。 * 页面通过类型检查、lint、生产构建、内部链接和关键 HTTP/SEO 检查。 ## AI 辅助披露 [#ai-辅助披露] AI 可以辅助资料整理、翻译初稿、示例审查和一致性检查,但不能作为事实来源。涉及数据库行为的结论必须回到官方资料或可复现实验;涉及业务口径、风险接受和生产变更的决定必须由人类负责。 ## 日期与更正 [#日期与更正] 页面底部的“最后更新”表示最近一次内容核对日期。发现错误时,应同时修正中英文页面、相关交叉链接和机器可读 Markdown,并在项目的内容核对记录中保留可追踪说明。 当前全站基线核对日期:**2026-08-02**。 --- # SQLSTATE 错误速查 Canonical URL: https://pg.edu.rich/docs/reference/errors Last reviewed: 2026-08-02 应用应分支处理 **SQLSTATE**,不要匹配可能随版本和语言变化的错误文本。 ## 高频状态码 [#高频状态码] | SQLSTATE | 名称 | 常见含义 | 稳妥动作 | | -------- | ----------------------------- | --------------------- | ------------------------------ | | `23505` | unique\_violation | 唯一键冲突 | 返回冲突或使用明确的 `ON CONFLICT` 语义 | | `23503` | foreign\_key\_violation | 引用不存在/仍被引用 | 修正操作顺序,不要临时禁用约束 | | `23502` | not\_null\_violation | 必填列缺失 | 修正输入或迁移顺序 | | `23514` | check\_violation | 违反 `CHECK` | 解释业务边界,修正值 | | `22P02` | invalid\_text\_representation | 类型转换失败 | 在应用边界校验并绑定正确类型 | | `40001` | serialization\_failure | 并发下无法保持隔离保证 | 回滚并重试整个事务 | | `40P01` | deadlock\_detected | 形成等待环 | 回滚整事务;统一锁顺序 | | `55P03` | lock\_not\_available | `NOWAIT`/lock timeout | 稍后重试或返回冲突 | | `57014` | query\_canceled | statement timeout 或取消 | 区分主动取消与超时,优化或缩小请求 | | `25P02` | in\_failed\_sql\_transaction | 当前事务此前已失败 | `ROLLBACK`;不要继续发业务 SQL | | `42501` | insufficient\_privilege | 对对象或动作无权限 | 修正 grant/owner;不要提升为 superuser | | `42P01` | undefined\_table | 表不存在或 search path 错 | 核对 database/schema/迁移版本 | | `42703` | undefined\_column | 列不存在 | 核对 schema 契约与部署版本 | | `53300` | too\_many\_connections | 连接槽耗尽 | 检查池配置、泄漏、保留管理连接 | | `57P03` | cannot\_connect\_now | 启动、恢复或关闭中 | 带上限退避,检查实例状态 | | `08006` | connection\_failure | 连接已失败 | 判断事务结果是否未知,再安全重试 | ## 事务失败后的规则 [#事务失败后的规则] 事务内任意语句失败后,通常进入 aborted 状态: ```text ERROR: current transaction is aborted... SQLSTATE: 25P02 ``` 必须 `ROLLBACK`,或回滚到失败前创建的 savepoint。不要继续发送语句期待自动恢复。 ## 重试分类 [#重试分类] * **可整事务重试**:`40001`、`40P01`;使用次数上限、指数退避和 jitter。 * **可能短暂重试**:`55P03`、`57P03`、部分 `08***`;先确认幂等和事务提交状态。 * **输入/模型错误,不应盲重试**:`22***`、`23***`、`42***`、`42501`。 * **资源问题**:`53300`、磁盘满、内存问题;重试会放大故障,先降载和修复容量。 ## 诊断上下文 [#诊断上下文] 记录 SQLSTATE、约束/表/列名、数据库与 schema、应用版本、迁移版本、事务 ID/请求 ID、参数类型(敏感值脱敏)和是否已提交。驱动通常提供结构化错误字段,应直接读取。 ## AI 工具的返回 [#ai-工具的返回] ```json { "ok": false, "sqlstate": "23505", "category": "constraint", "retryable": false, "constraint": "customers_email_unique", "message_safe": "A customer with this email already exists" } ``` 不要把原始数据库错误无过滤地返回终端用户;它可能泄露对象名、路径或数据片段。 --- # PostgreSQL 扩展与开源生态选型指南 Canonical URL: https://pg.edu.rich/docs/reference/extensions-ecosystem Last reviewed: 2026-08-06 PostgreSQL 扩展让类型、索引、planner hook、后台 worker 和存储能力进入数据库进程,也会进入备份、复制、故障恢复和 major upgrade 的关键路径。选型原则是:**没有明确工作负载和退出方案,就不要安装**。 ## 安装前的六项门禁 [#安装前的六项门禁] 1. 目标 PostgreSQL major、操作系统和 CPU 架构有明确支持与 package。 2. 许可证满足自托管、SaaS、分发和商业功能边界。 3. 备份、PITR、standby、逻辑复制和恢复环境能够加载相同版本。 4. `pg_upgrade`、extension update 和需要重建的 index 有演练路径。 5. 托管云的区域、SKU 与 allowlist 提供所需版本,不只提供同名扩展。 6. 有不依赖该扩展的导出或迁移策略,避免无意中锁定平台。 记录 `SELECT extname, extversion FROM pg_extension`,并把扩展版本与数据库版本一起进入部署清单和 AI 上下文。 先检查 [PostgreSQL 索引与存储访问方法](/docs/reference/index-access-methods) 中的原生 B-tree/GIN/GiST/SP-GiST/BRIN、全文检索、分区、FDW 与物化视图。只有原生能力无法满足已测量的 workload 时,再增加 extension。 ## 按工作负载选择 [#按工作负载选择] | 场景 | 常见候选 | 采用边界 | | -------------------------- | ------------------------------------------------------------------------------------- | --------------------------------------------------------------------------- | | SQL 统计 | [`pg_stat_statements`](https://www.postgresql.org/docs/current/pgstatstatements.html) | 官方 contrib,生产可观测性基线;治理查询文本权限 | | 向量检索 / RAG | [pgvector](https://github.com/pgvector/pgvector) | 用真实过滤条件测 recall、latency、内存与索引构建 | | 地理空间 | [PostGIS](https://postgis.net/) | GIS 标准选择;确认 extension 与数据格式升级路径 | | 时间序列 | [TimescaleDB](https://github.com/timescale/timescaledb) | 需要 hypertable/压缩/连续聚合时评估;逐项核对许可证 | | 分布式多租户 | [Citus](https://github.com/citusdata/citus) | 单机已被测量为瓶颈且 shard key 稳定后再引入 | | BM25 / 搜索 | [ParadeDB / pg\_search](https://github.com/paradedb/paradedb) | 核对许可证、索引恢复、复制与云支持 | | PostgreSQL 内 BM25 | [pg\_textsearch](https://github.com/timescale/pg_textsearch) | 上游目前声明 production ready;仍需按目标版本和语料独立验收 | | 嵌入式分析 / Parquet | [pg\_duckdb](https://github.com/duckdb/pg_duckdb) | 适合分析路径;验证事务边界、资源隔离和对象存储凭据 | | Iceberg columnstore mirror | [pg\_mooncake](https://github.com/Mooncake-Labs/pg_mooncake) | 通过 logical change capture 维护 columnstore mirror;验证一致性、对象存储、pg\_duckdb 依赖与恢复 | | 图查询 | [Apache AGE](https://github.com/apache/age) | 只有图模型与 Cypher 带来可测收益时采用 | “上游 production ready”是项目自己的状态声明,不是对你的 workload、SLA 或云平台的认证。 ## TimescaleDB 快速上手与版本边界 [#timescaledb-快速上手与版本边界] 先满足上文的采用边界:只有当 hypertable、压缩或连续聚合对应已测量的需求时才评估 TimescaleDB。TimescaleDB 2.x 上的最小评估序列: ```sql CREATE EXTENSION IF NOT EXISTS timescaledb; CREATE TABLE metrics ( time timestamptz NOT NULL, device text NOT NULL, value double precision ); SELECT create_hypertable('metrics', by_range('time')); CREATE MATERIALIZED VIEW metrics_hourly WITH (timescaledb.continuous) AS SELECT time_bucket('1 hour', time) AS hour, device, avg(value), max(value) FROM metrics GROUP BY hour, device; ALTER TABLE metrics SET (timescaledb.compress); SELECT add_compression_policy('metrics', INTERVAL '7 days'); SELECT add_retention_policy('metrics', INTERVAL '90 days'); ``` * `by_range` 维度构造器要求 TimescaleDB 2.13 及以上;更早版本使用 `create_hypertable('metrics', 'time')`。 * `ALTER TABLE ... SET (timescaledb.compress)` 是旧版压缩 API。近期 2.x 版本将压缩收敛到 hypercore,reloption 与 policy 名称不同 —— 以已安装版本和 TimescaleDB 官方文档为准。 * 连续聚合在附加 `add_continuous_aggregate_policy` 之前不会按调度刷新。 * TimescaleDB 采用 Timescale License(TSL),不是 PostgreSQL License。生产使用前逐项核对功能边界。 ## 维护与数据治理 [#维护与数据治理] | 工具 | 作用 | 不应误解为 | | ------------------------------------------------------------------------ | ---------------------------------- | --------------------- | | [HypoPG](https://github.com/HypoPG/hypopg) | 用 hypothetical index 评估 planner 选择 | 真实构建成本和生产收益证明 | | [pg\_repack](https://github.com/reorg/pg_repack) | 以较短排他锁窗口重组表和 index | 日常 autovacuum 的替代品 | | [pg\_partman](https://github.com/pgpartman/pg_partman) | 管理原生时间/序列分区生命周期 | 自动修复错误 partition key | | [pg\_cron](https://github.com/citusdata/pg_cron) | 在数据库内调度简单 SQL 工作 | 通用业务队列和复杂 workflow 引擎 | | [Greenmask](https://github.com/GreenmaskIO/greenmask) | 生成脱敏、子集化的测试数据 | 可以无审查复制生产敏感数据 | | [PostgreSQL Anonymizer](https://gitlab.com/dalibo/postgresql_anonymizer) | 声明式静态/动态掩码 | 自动满足全部合规要求 | 严重膨胀先找长事务、autovacuum、写入模式与 fillfactor 根因,再使用 pg\_repack。分区只解决能按 partition key 剪枝和管理的数据生命周期问题。 ## PostgreSQL 19 REPACK 与 pg\_repack 不是一回事 [#postgresql-19-repack-与-pg_repack-不是一回事] PostgreSQL 19 Beta 文档中的核心 [`REPACK`](https://www.postgresql.org/docs/19/sql-repack.html) 是新 SQL command,`REPACK (CONCURRENTLY)` 基于 logical decoding,并对主键/replica identity、unlogged/partitioned/system table、replication slot 和磁盘空间有约束。 第三方 **pg\_repack** 是独立 extension 与命令行工具,具有自己的兼容矩阵、安装包和操作边界。不要因为名称相近,就把 pg\_repack 的经验、监控或风险模型直接套到 PostgreSQL 19 核心 `REPACK`。 截至 2026-08-02,PostgreSQL 19 仍为 Beta 2。核心 REPACK 的语义和限制应以最终 GA 文档与自己的恢复副本演练为准。 ## PostgreSQL 19 计划建议模块 [#postgresql-19-计划建议模块] PostgreSQL 19 Beta 新增两个面向计划稳定化的 contrib 模块。[`pg_plan_advice`](https://www.postgresql.org/docs/19/pgplanadvice.html) 允许把关键 planner 决策描述、复现并通过附加在查询上的 advice 改写;[`pg_stash_advice`](https://www.postgresql.org/docs/19/pgstashadvice.html) 把 advice 字符串按 query identifier 存入动态共享内存并自动应用。 两者都应定位为实验通道,而不是生产级计划管理契约:advice 格式与覆盖范围在 GA 前仍可能变化,而且用这种方式固定计划绕开了重新 `ANALYZE` 或 HypoPG 这类验证手段。把它们放进 Beta 测试通道与计划对比 benchmark 一起评估,而不是用来替代这些验证。 ## AI、备份与升级清单 [#ai备份与升级清单] 向 AI/Agent 提供:`server_version_num`、`extname/extversion`、允许使用的 operator/index method、目标云限制和禁止语法。不要仅告诉模型“这是 PostgreSQL”。 每次扩展升级前完成: 1. 读取目标版本 release notes、SQL update script 和已知重建要求; 2. 从真实备份恢复到隔离环境; 3. 升级 PostgreSQL 与 extension,运行完整性、性能和 RLS 测试; 4. 重建要求的 index,并比较 plan、recall 或业务结果; 5. 重新生成备份并执行一次恢复,确认新版本链路成立。 Supabase、Neon、YugabyteDB、CockroachDB、Cloudberry、Gel 和 FerretDB 不应混进“扩展排行榜”:它们分别是平台、分支、独立数据库或协议转换层。分类见 [PostgreSQL 血缘与兼容数据库](/docs/reference/postgresql-compatible-databases)。 --- # PostgreSQL 索引与存储访问方法 Canonical URL: https://pg.edu.rich/docs/reference/index-access-methods Last reviewed: 2026-08-02 PostgreSQL 不采用 MySQL 那种为普通表频繁选择 InnoDB/MyISAM 的使用模型。绝大多数表使用核心 **heap table access method**;索引通过独立的 index access method 和 operator class 决定支持哪些查询。 ## Table Access Method 不是日常调优开关 [#table-access-method-不是日常调优开关] PostgreSQL 提供 [Table Access Method API](https://www.postgresql.org/docs/current/tableam.html),允许扩展或定制构建实现新的表存储方式,但普通应用的默认仍是 `heap`: ```sql CREATE TABLE events ( id bigint GENERATED ALWAYS AS IDENTITY PRIMARY KEY, occurred_at timestamptz NOT NULL, payload jsonb NOT NULL ) USING heap; ``` 通常省略 `USING heap`。选择新 table access method 会进入 WAL、MVCC、VACUUM、backup、replication、extension 和 major upgrade 的关键路径,不能像切换 query hint 一样随意尝试。 查看当前实例暴露的 access method: ```sql SELECT amname, CASE amtype WHEN 't' THEN 'table' WHEN 'i' THEN 'index' ELSE amtype::text END AS access_method_type FROM pg_am ORDER BY amtype, amname; ``` ## 核心索引访问方法 [#核心索引访问方法] PostgreSQL 18 核心提供 B-tree、Hash、GiST、SP-GiST、GIN 和 BRIN;`bloom` 是随 PostgreSQL 提供但需要 `CREATE EXTENSION bloom` 的 module。官方边界见 [Index Types](https://www.postgresql.org/docs/current/indexes-types.html)。 | 类型 | 优先场景 | 重要边界 | | --------------- | --------------------------------- | ---------------------------------------------------------------------------------- | | B-tree | `=`、范围、排序、unique、前缀模式匹配 | 默认选择;复合列顺序与 operator class 决定可用查询 | | Hash | 单列等值比较 | 只支持 `=`;B-tree 通常更通用,采用前需有测量收益 | | GIN | JSONB、array、全文检索和多值内容 | 更新和构建成本较高;行为取决于 operator class | | GiST | range、几何、PostGIS、nearest-neighbor | 是可扩展框架,不是单一算法;必须匹配 operator/operator class | | SP-GiST | trie、quad-tree、k-d tree 等空间分区结构 | 适合具有可分区结构的数据;不是 GiST 的通用替换 | | BRIN | 与物理顺序高度相关的超大追加表 | 保存 block range 摘要;相关性差时会读取大量 heap block | | Bloom extension | 多列任意组合的等值过滤 | lossy、需要 recheck;不支持 range、unique 或搜索 `NULL`;自带 operator class 仅覆盖 `int4` 与 `text` | GIN、GiST、SP-GiST 和 BRIN 是框架。真正决定支持哪些 operator、排序和数据类型的是 operator class。看到“使用 GIN”仍不足以复现一个索引设计。 ## 常见工作负载映射 [#常见工作负载映射] ```text 等值 / 范围 / 排序 / unique → B-tree JSONB contains / array member → GIN PostgreSQL FTS → GIN(通常) range / GIS / nearest neighbor → GiST 或匹配的 SP-GiST 超大、按时间物理追加 → BRIN 向量 ANN → pgvector HNSW / IVFFlat ``` HNSW 与 IVFFlat 来自 [pgvector](https://github.com/pgvector/pgvector),不是 PostgreSQL 核心 index method。它们需要单独验证 recall、filter、memory、build time、WAL、replica lag 与 extension upgrade。 PostgreSQL 内置全文检索支持 parser、dictionary、ranking、highlight 和 GIN/GiST index。内置配置不自动解决所有语言的分词;例如中文通常需要额外 tokenizer/extension 或应用侧预处理,不能只创建一个 GIN 就宣称搜索质量成立。 ## 创建前先匹配查询 [#创建前先匹配查询] ```sql -- 普通业务过滤与排序 CREATE INDEX CONCURRENTLY orders_customer_time_idx ON orders (customer_id, placed_at DESC); -- JSONB 包含查询:payload @> '{"status":"paid"}' CREATE INDEX CONCURRENTLY events_payload_gin_idx ON events USING gin (payload jsonb_path_ops); -- 时间与物理写入顺序高度相关的大表 CREATE INDEX CONCURRENTLY events_time_brin_idx ON events USING brin (occurred_at); ``` `jsonb_path_ops` 更专注于 `@>`、`@?`、`@@` 等路径/包含查询,并不支持默认 `jsonb_ops` 的全部 operator。索引 DDL 必须从真实 query shape 反推。 创建后记录 plan 与尺寸: ```sql SELECT indexrelname, idx_scan, pg_size_pretty(pg_relation_size(indexrelid)) AS index_size FROM pg_stat_user_indexes WHERE relname = 'events' ORDER BY pg_relation_size(indexrelid) DESC; ``` 再使用 `EXPLAIN (ANALYZE, BUFFERS)` 比较实际行数、heap block、recheck、排序与写入代价。完整步骤见 [索引与 EXPLAIN](/docs/core/indexes-explain)。 ## 原生能力优先 [#原生能力优先] | 需求 | 先验证 PostgreSQL 原生 | 仍不足时再评估 | | ------ | -------------------------------------------------- | --------------------------------- | | 模糊匹配 | FTS、`pg_trgm` contrib、表达式/GIN/GiST index | 外部搜索或 BM25 extension | | 任务互斥 | transaction、`FOR UPDATE SKIP LOCKED`、advisory lock | 专用队列与 workflow 系统 | | 时间生命周期 | 原生 partition、BRIN、scheduled cleanup | pg\_partman、TimescaleDB | | 跨库访问 | `postgres_fdw`、logical replication | CDC 平台或独立同步系统 | | 分析 | materialized view、partition、parallel query | pg\_duckdb、pg\_mooncake、warehouse | | 向量检索 | 没有核心 vector type/index | pgvector 或专用向量系统 | “原生优先”不是拒绝扩展,而是减少不必要的 binary、license、backup 和 upgrade 依赖。扩展候选见 [PostgreSQL 扩展生态选型](/docs/reference/extensions-ecosystem)。 --- # PostgreSQL 现场速查 Canonical URL: https://pg.edu.rich/docs/reference Last reviewed: 2026-08-02 ## psql [#psql] ```bash psql 'postgresql://user@host:5432/database?sslmode=verify-full' psql -X --set ON_ERROR_STOP=on --file migration.sql "$DATABASE_URL" ``` | 命令 | 作用 | | ---------------- | ------------ | | `\conninfo` | 当前连接信息 | | `\l` | database 列表 | | `\dn` | schema 列表 | | `\dt app.*` | 表列表 | | `\d+ app.orders` | 对象定义与存储信息 | | `\du` | 角色列表 | | `\dx` | 扩展列表 | | `\timing on` | 显示客户端观察的耗时 | | `\x auto` | 宽结果自动纵向显示 | | `\gdesc` | 描述查询结果列而不展示行 | | `\q` | 退出 | 脚本使用 `-X` 避免加载用户 `.psqlrc`,并用 `ON_ERROR_STOP` 在首个错误退出。 ## 当前上下文 [#当前上下文] ```sql SELECT version(), current_database(), current_user, session_user, current_schema(), current_setting('TimeZone') AS timezone, inet_server_addr(), inet_server_port(); ``` ## 对象大小 [#对象大小] ```sql SELECT relname, pg_size_pretty(pg_total_relation_size(relid)) AS total, pg_size_pretty(pg_relation_size(relid)) AS heap, pg_size_pretty(pg_indexes_size(relid)) AS indexes FROM pg_catalog.pg_statio_user_tables ORDER BY pg_total_relation_size(relid) DESC LIMIT 20; ``` ## 活跃会话与长事务 [#活跃会话与长事务] ```sql SELECT pid, usename, application_name, state, now() - xact_start AS xact_age, wait_event_type, wait_event, left(query, 160) AS query FROM pg_stat_activity WHERE pid <> pg_backend_pid() ORDER BY xact_start NULLS LAST; ``` ## 谁阻塞谁 [#谁阻塞谁] ```sql SELECT blocked.pid AS blocked_pid, blocker.pid AS blocker_pid, now() - blocked.query_start AS blocked_for, left(blocked.query, 120) AS blocked_query, left(blocker.query, 120) AS blocker_query FROM pg_stat_activity AS blocked CROSS JOIN LATERAL unnest(pg_blocking_pids(blocked.pid)) AS b(pid) JOIN pg_stat_activity AS blocker ON blocker.pid = b.pid; ``` 不要看到 blocker 就立刻 `pg_terminate_backend`。先确认业务、事务内容、是否可重试和终止后影响。 ## 安全的会话设置 [#安全的会话设置] ```sql BEGIN; SET LOCAL statement_timeout = '10s'; SET LOCAL lock_timeout = '2s'; SET LOCAL search_path = app, pg_catalog; -- work COMMIT; ``` ## 诊断顺序 [#诊断顺序] 确认目标实例与角色 → 记录 SQLSTATE → 检查事务状态 → 检查等待事件/阻塞 → 获取查询计划与统计 → 在受控环境复现 → 修复后执行验证查询。 连接尚未建立时使用 [PostgreSQL 连接错误排查](/docs/reference/connection-errors);连接成功后的 SQL 错误使用 [错误与 SQLSTATE](/docs/reference/errors)。 扩展、维护工具和开源组件的升级边界见 [PostgreSQL 扩展与开源生态选型](/docs/reference/extensions-ecosystem)。 索引与存储概念见 [PostgreSQL 索引与存储访问方法](/docs/reference/index-access-methods);协议兼容、分支与 PostgreSQL 血缘见 [PostgreSQL 血缘与兼容数据库](/docs/reference/postgresql-compatible-databases)。 --- # PostgreSQL 17:新特性与升级注意点 Canonical URL: https://pg.edu.rich/docs/reference/postgresql-17 Last reviewed: 2026-08-06 PostgreSQL 17 于 **2024-09-26** 正式发布(GA),支持至 **2029-11-08**。截至本页核对日,当前 minor 为 17.10;实时支持状态见 [版本与支持策略](/docs/reference/version-policy)。以下所有特性声明均已对照 [PostgreSQL 17 官方 release notes](https://www.postgresql.org/docs/17/release-17.html) 核实。 ## 流式 I/O(streaming I/O)框架 [#流式-iostreaming-io框架] PostgreSQL 17 为顺序读引入了流式 I/O 接口:执行器不再逐块发出读请求,而是让一串读请求保持在线,顺序扫描等批量读的性能因此改善。[PostgreSQL 18](/docs/reference/postgresql-18) 的异步 I/O 子系统正是建立在这套框架之上,17 的这部分工作是 18 更大 I/O 收益的前置基础。 ## 逻辑复制:failover 复制槽与 pg\_createsubscriber [#逻辑复制failover-复制槽与-pg_createsubscriber] * **复制槽可以跨 failover 存活。** 复制协议新增了 `failover` 属性,配合 `sync_replication_slots` 让备库同步启用了 failover 的逻辑复制槽;发布端切换到备库后,逻辑订阅者可以继续流式接收。 * **`pg_createsubscriber`** 把物理备库转换为逻辑副本。这是把流复制副本转成逻辑订阅者的标准工具,也是低停机大版本迁移的一条实用路径。 ## MERGE:WHEN NOT MATCHED BY SOURCE 与 RETURNING [#mergewhen-not-matched-by-source-与-returning] PostgreSQL 17 为 `MERGE` 增加了 `WHEN NOT MATCHED BY SOURCE` 动作和 `RETURNING` 子句(含 `merge_action()` 函数,可报告每一行实际执行的 DML 动作): ```sql MERGE INTO target USING source ON source.id = target.id WHEN MATCHED AND target.deleted = false THEN UPDATE SET ... WHEN NOT MATCHED THEN INSERT ... WHEN NOT MATCHED BY SOURCE THEN DELETE RETURNING merge_action(), *; ``` 这两项加在一起,让 `MERGE` 从仅能 upsert 变成可以承担完整同步任务的语句。 ## JSON\_TABLE [#json_table] PostgreSQL 17 增加了 SQL 标准的 `JSON_TABLE()` 函数,把 JSON 数据投影为关系行集: ```sql SELECT * FROM JSON_TABLE(jsonb_col, '$.items[*]' COLUMNS ( id bigint PATH '$.id', name text PATH '$.name' )) AS t; ``` 当 JSON 文档需要按行做连接、过滤或聚合时,它可以替代一类手写的 `jsonb_to_recordset` 与 lateral 展开查询。 ## pg\_basebackup 增量备份 [#pg_basebackup-增量备份] PostgreSQL 17 增加了文件系统级增量备份:`pg_basebackup --incremental` 只生成相对上一份备份清单(manifest)发生变化的块,`pg_combinebackup` 负责把全量加增量链合成为完整备份: ```bash pg_basebackup --incremental=/path/to/backup_manifest -D ./backup_inc ``` 对大型数据库,这能缩短备份窗口并减少备份存储,代价是恢复时多一步合成。 ## 可配置的 SLRU 缓冲池 [#可配置的-slru-缓冲池] 子事务、提交时间戳等子系统背后的 SLRU 缓存现在可以通过 `subtransaction_buffers`、`commit_timestamp_buffers` 等参数独立调整,且默认随 `shared_buffers` 缩放。此前这些缓存大小固定,在高并发下可能成为竞争点。 ## PostgreSQL 17 的其他变化 [#postgresql-17-的其他变化] * **`COPY ... ON_ERROR ignore`**:跳过格式错误的输入行,而不是让整个 copy 中止。(PostgreSQL 18 随后又增加了 `REJECT_LIMIT`,用于限制可丢弃的行数。) * PostgreSQL 17 移除了 `old_snapshot_threshold` 设置、`adminpack` contrib 扩展等若干长期弃用的部分——如果从很老的大版本跨级升级,请核对 release notes。 ## 升级到 PostgreSQL 17 的注意点 [#升级到-postgresql-17-的注意点] * 大版本升级需要 `pg_upgrade`、dump/restore 或逻辑复制;数据目录不跨大版本兼容。上文提到的 `pg_createsubscriber` 是本版本引入的逻辑复制迁移路径。 * 跨越的每一个大版本的 release notes 都要读,不能只读目标版本的。 * 监控查询可能需要更新:PostgreSQL 17 重命名了若干 `pg_stat_statements` 计时列(例如 `blk_read_time` 改为 `shared_blk_read_time`),并从 `pg_stat_bgwriter` 移除了 `buffers_backend` / `buffers_backend_fsync` 列(与 `pg_stat_io` 重复)。 * 每个扩展的 PostgreSQL 17 支持情况要单独确认;扩展有独立的版本号与升级脚本。 生产环境应运行当前的 17.x minor;任何大版本升级都应先在真实数据的恢复副本上演练,再排期切换。通用升级清单见 [版本与支持策略](/docs/reference/version-policy)。 ## 版本路径 [#版本路径] * 下一大版本:[PostgreSQL 18:新特性与升级注意点](/docs/reference/postgresql-18) * 截至 2026-08 处于 Beta:[PostgreSQL 19:新功能与 18 升级 19 指南](/docs/postgresql-19) * 支持周期:[版本与支持策略](/docs/reference/version-policy) ## 来源 [#来源] * [PostgreSQL 17 release notes](https://www.postgresql.org/docs/17/release-17.html) * [PostgreSQL 版本策略](https://www.postgresql.org/support/versioning/) --- # PostgreSQL 18:新特性与升级注意点 Canonical URL: https://pg.edu.rich/docs/reference/postgresql-18 Last reviewed: 2026-08-06 PostgreSQL 18 于 **2025-09-25** 正式发布(GA),支持至 **2030-11-14**。截至本页核对日,当前 minor 为 18.4;实时支持状态见 [版本与支持策略](/docs/reference/version-policy)。以下所有特性声明均已对照 [PostgreSQL 18 官方 release notes](https://www.postgresql.org/docs/18/release-18.html) 核实。 ## 异步 I/O(AIO) [#异步-ioaio] PostgreSQL 18 是第一个具备真正异步 I/O 子系统的版本,构建在 [PostgreSQL 17](/docs/reference/postgresql-17) 引入的流式 I/O 框架之上。I/O 请求发出后后端不再逐次阻塞等待,顺序扫描、vacuum 和 bitmap heap scan 都会受益。 行为由 `io_method` 控制: * `sync` —— 传统的同步 I/O; * `worker` —— 由后台 I/O worker 进程执行;这是**默认值**,为了更广泛的兼容性而选择; * `io_uring` —— 在受支持的内核上使用 Linux io\_uring。 [PostgreSQL 18 发布公告](https://www.postgresql.org/about/news/postgresql-18-released-3142/)称基准测试中特定场景最高有约 3× 提升。请把它当作上限而不是预期:切换 `io_method` 前后都用自己的工作负载实测,从 `worker` 换到 `io_uring` 也要先测试再上线。 ## 内置 uuidv7() [#内置-uuidv7] ```sql SELECT uuidv7(); ``` UUID v7 内嵌 48 位毫秒时间戳,生成的值大致按时间有序。与随机的 v4 UUID 相比,时间有序的主键会插入 B-tree 的右缘,高写入场景下主键索引更紧凑、对缓存更友好。新表可以直接采用: ```sql CREATE TABLE events ( id uuid PRIMARY KEY DEFAULT uuidv7() ); ``` ## 虚拟生成列(virtual generated columns) [#虚拟生成列virtual-generated-columns] ```sql ALTER TABLE orders ADD COLUMN total numeric GENERATED ALWAYS AS (qty * price) VIRTUAL; ``` PostgreSQL 18 之前,生成列一律是 `STORED`——写入时计算并占用磁盘。PostgreSQL 18 增加了 `VIRTUAL` 生成列:读取时计算、不占用存储,并将 `VIRTUAL` 设为默认;需要写入时计算的行为仍可显式指定 `STORED`。 ## B-tree skip scan [#b-tree-skip-scan] 过去,`(a, b)` 上的复合 B-tree 索引必须带前导列 `a` 的过滤条件才用得上。PostgreSQL 18 支持 skip scan:执行器遍历 `a` 的不同取值,逐个下钻到对应的 `b` 区间,因此只按非前导列过滤的查询也能使用该索引。一些当年仅为覆盖第二列而建的索引可能已经不再必要——但请先在真实数据上用 `EXPLAIN (ANALYZE, BUFFERS)` 验证再删除。 ## RETURNING 支持 OLD / NEW [#returning-支持-old--new] `INSERT`、`UPDATE`、`DELETE` 和 `MERGE` 现在可以在 `RETURNING` 中显式返回新旧两个行版本: ```sql UPDATE products SET price = price * 1.1 RETURNING OLD.price AS prev_price, NEW.price AS new_price; ``` 过去要在一条语句里拿到变更前后的值,往往需要 CTE 或自连接的变通写法,现在可以省去。 ## PostgreSQL 18 的其他变化 [#postgresql-18-的其他变化] * **`EXPLAIN ANALYZE` 自动附带 `BUFFERS` 输出**,常见场景下不再需要显式加选项。 * **`COPY FROM` 新增 `REJECT_LIMIT`**,限制 `ON_ERROR ignore` 模式下可丢弃的无效行数,超过即失败。(`ON_ERROR ignore` 本身是在 [PostgreSQL 17](/docs/reference/postgresql-17) 加入的。) * **`pg_stat_io` 新增按字节统计的 I/O 列,并报告 WAL I/O 行**;相应地,`pg_stat_wal` 移除了读/同步相关列——引用过这些列的监控看板需要更新。 * **`pg_upgrade` 现在会保留优化器统计信息**,缩短大版本升级后计划劣化的窗口期。 透明数据加密(TDE)在 PostgreSQL 18 开发周期内曾被提议,但最终没有合并,PostgreSQL 18 不带内置的页级加密。静态加密仍然需要文件系统或磁盘层加密(如 LUKS)、存储层能力,或提供 TDE 的厂商发行版。不要围绕"本版本会有核心 TDE"来做规划。 ## 升级到 PostgreSQL 18 的注意点 [#升级到-postgresql-18-的注意点] * 大版本升级需要 `pg_upgrade`、dump/restore 或逻辑复制;数据目录不跨大版本兼容。 * **`initdb` 现在默认启用数据校验和(data checksums)。** 由于 `pg_upgrade` 要求两端校验和设置一致,新增的 `--no-data-checksums` initdb 选项可用于从未启用校验和的旧集群升级。 * **MD5 密码认证已弃用**,`ALTER ROLE` 设置 MD5 密码时会发出警告;应规划迁移到 SCRAM,而不是屏蔽警告。 * 如上所述,`pg_stat_wal` 的读/同步列已移入 `pg_stat_io`;监控采集器需要相应调整。 * 扩展、驱动、连接池、备份工具对 18 的支持要逐项确认;`pg_upgrade` 无法证明第三方模块兼容。 ## 版本路径 [#版本路径] * 上一大版本:[PostgreSQL 17:新特性与升级注意点](/docs/reference/postgresql-17) * 截至 2026-08 处于 Beta:[PostgreSQL 19:新功能与 18 升级 19 指南](/docs/postgresql-19) * 支持周期:[版本与支持策略](/docs/reference/version-policy) ## 来源 [#来源] * [PostgreSQL 18 release notes](https://www.postgresql.org/docs/18/release-18.html) * [PostgreSQL 18 发布公告](https://www.postgresql.org/about/news/postgresql-18-released-3142/) * [PostgreSQL 版本策略](https://www.postgresql.org/support/versioning/) --- # PostgreSQL 血缘、分支与兼容数据库 Canonical URL: https://pg.edu.rich/docs/reference/postgresql-compatible-databases Last reviewed: 2026-08-02 “基于 PostgreSQL”“使用 PostgreSQL 协议”和“可以替换 PostgreSQL”是三个不同结论。兼容性至少有五层: ```text 驱动可以连接 → pgwire 消息可交换 → SQL / 类型 / 函数兼容 → catalog / extension / transaction 行为兼容 → backup / replication / upgrade / failure 语义兼容 ``` 越靠下,越需要真实迁移和故障测试。产品自称 PostgreSQL-compatible,通常只描述其中一部分。 ## PostgreSQL 为核心的开发平台 [#postgresql-为核心的开发平台] | 平台 | PostgreSQL 在哪里 | 平台增加了什么 | 不能直接假设 | | ------------------------------------------------------ | ------------------------------ | ----------------------------------------------------------- | ----------------------------------------------- | | [Supabase](https://github.com/supabase/supabase) | 每个项目运行 PostgreSQL | PostgREST、Auth、Realtime、Storage、Functions、Dashboard 与连接池 | 自托管与云平台功能/运维完全相同;浏览器 API 自动安全 | | [Neon](https://github.com/neondatabase/neon) | compute node 运行 PostgreSQL 查询层 | compute/storage separation、page server、branch、scale-to-zero | data directory、WAL、unlogged table、恢复与普通 PG 运维相同 | | [Nhost](https://github.com/nhost/nhost) | PostgreSQL 是 database | Hasura GraphQL、Auth、Storage、Functions | GraphQL permission 等同于全部数据库权限边界 | | [Prisma Postgres](https://www.prisma.io/docs/postgres) | 托管 PostgreSQL | PgBouncer、HTTP/edge driver、query cache、临时数据库和 Prisma 工具链 | operation 计费、pooling 和 extension 与任意自托管 PG 相同 | 这些平台适合继续使用普通 SQL、migration 和 `pg_dump` 思维,但仍要把平台的连接代理、休眠、备份、扩展 allowlist、API 权限和计费模型写进架构。 ## PostgreSQL 分支、查询层复用与新存储 [#postgresql-分支查询层复用与新存储] | 项目 | 实现路径 | 主要目标 | 迁移风险中心 | | --------------------------------------------------------------------------- | ------------------------------------------- | ----------------------------------------------- | ----------------------------------------------------- | | [YugabyteDB](https://github.com/yugabyte/yugabyte-db) | YSQL 复用 PostgreSQL query layer,底层是分布式 DocDB | 分布式事务、横向扩展、多区域 | extension、lock/isolation、catalog、DDL 与分布式成本模型 | | [PolarDB for PostgreSQL](https://github.com/polardb/PolarDB-for-PostgreSQL) | PostgreSQL 血缘的计算存储分离分支 | shared storage、一写多读、云原生架构 | 开源分支版本节奏、专用 storage/HA、与公有云版本差异 | | [Apache Cloudberry](https://github.com/apache/cloudberry) | Greenplum/PG 血缘的 MPP 数据库 | 数据仓库、大规模并行分析 | OLTP transaction、distribution key、SQL/extension 与运维工具 | | [IvorySQL](https://github.com/IvorySQL/IvorySQL) | 跟随 PostgreSQL 的 Oracle-compatible 分支 | PL/iSQL、Oracle syntax、package 与迁移 | compatibility mode、Oracle 语义、extension package 与上游同步 | | [openGauss](https://github.com/opengauss-mirror/openGauss-server) | PostgreSQL 血缘的独立数据库内核 | 企业部署、并行与自身生态 | 已长期独立演进,不能把当前 PostgreSQL 兼容性当作默认 | | [OrioleDB](https://github.com/orioledb/orioledb) | 面向 PostgreSQL 的新 storage engine,通常需要其支持的构建 | undo-based MVCC、copy-on-write/checkpoint、降低某些膨胀 | binary/build、WAL/backup、extension、major upgrade 与故障恢复 | 这里的“血缘”不等于 drop-in replacement。特别是分布式存储会改变 transaction retry、hot key、sequence、foreign key、lock 和一致性/延迟取舍。 ## 支持 PostgreSQL 客户端,但不是 PostgreSQL [#支持-postgresql-客户端但不是-postgresql] | 项目 | pgwire / PostgreSQL 的作用 | 实际定位 | | --------------------------------------------------------------------------------- | ---------------------------------------- | -------------------------------------------------------------- | | [CockroachDB](https://www.cockroachlabs.com/docs/stable/postgresql-compatibility) | 实现 pgwire 和大量 PostgreSQL syntax | 独立分布式 SQL 数据库;不等同于 PostgreSQL extension/catalog/transaction 行为 | | [Materialize](https://github.com/MaterializeInc/materialize) | PostgreSQL-compatible driver 可查询 view | 流式增量计算与实时数据层,不是通用 OLTP PostgreSQL | | [Gel](https://github.com/geldata/gel) | 底层使用 PostgreSQL 技术并提供 SQL interface/生态连接 | graph-relational database,主要数据模型和 query language 是 Gel/EdgeQL | 部分兼容数据库会报告 PostgreSQL 风格的 `server_version` 或提供相似 catalog。版本字符串只能帮助驱动选择协议路径,不能证明服务器运行同一 PostgreSQL 内核。 ## FerretDB 是反方向兼容 [#ferretdb-是反方向兼容] [FerretDB 2.x](https://github.com/FerretDB/FerretDB) 接收 MongoDB 5.0+ wire protocol,把请求转换为 SQL,并使用带 DocumentDB extension 的 PostgreSQL 作为 database engine: ```text MongoDB driver → FerretDB proxy → PostgreSQL + DocumentDB extension ``` 因此它不是“PostgreSQL client 连接一个 MongoDB-compatible server”,而是“MongoDB client 使用 PostgreSQL-backed document database”。验证重点是 MongoDB command/BSON 兼容矩阵、DocumentDB extension、索引、transaction、backup 与版本组合。 ## 迁移兼容性测试矩阵 [#迁移兼容性测试矩阵] | 层 | 必测内容 | 不能接受的替代证据 | | ------ | ------------------------------------------------------------ | ------------------------- | | 连接 | TLS、SCRAM、startup parameter、prepared statement、pooler | `psql` 能执行 `SELECT 1` | | Schema | type、identity/sequence、generated column、constraint、partition | ORM migration 只在空库成功 | | SQL | function/operator、JSONB、CTE/window、collation、全文 | 跑过简单 CRUD | | 事务 | isolation、retry、row lock、deadlock、advisory lock | 宣传页写“ACID” | | 扩展 | exact version、operator/index method、upgrade script | 扩展名称出现在 allowlist | | 运维 | backup/PITR、CDC、replication、catalog、monitoring | 有“backup”按钮 | | 故障 | node/zone failure、连接收敛、RPO/RTO、回滚 | vendor benchmark 或 SLA 数字 | 建议先运行应用测试和 migration,再恢复脱敏生产副本,最后做切换与回退演练。对于 CockroachDB/YugabyteDB 这类分布式数据库,还要主动制造 transaction conflict、hot partition 和节点故障。 ## 如何做选择 [#如何做选择] * 要标准 PostgreSQL 生态与最低迁移成本:优先社区 PostgreSQL 或明确运行 PostgreSQL 的托管服务。 * 要 BaaS:Supabase;要 GraphQL-first:Nhost;要 branch/scale-to-zero:Neon。 * 要多区域分布式 OLTP:把 YugabyteDB 与 CockroachDB 作为新数据库评估,不视为配置项。 * 要 MPP warehouse:评估 Cloudberry,不用 OLTP benchmark 推导分析性能。 * 要 Oracle migration:评估 IvorySQL,并保留 PostgreSQL mode 与 Oracle mode 的差异测试。 * 要实时增量 view:Materialize 是数据层候选,不是主 OLTP 数据库的透明替换。 云服务免费层见 [免费 PostgreSQL 云数据库选型](/docs/cloud/free-postgresql);真正的扩展选型见 [PostgreSQL 扩展生态](/docs/reference/extensions-ecosystem)。 --- # PostgreSQL 与 DuckDB 选型 Canonical URL: https://pg.edu.rich/docs/reference/postgresql-vs-duckdb Last reviewed: 2026-08-06 PostgreSQL 是为大量并发事务设计的行存客户端/服务器数据库;DuckDB 是进程内的列存分析引擎——一个链接进应用的库,而不是一个需要连接的服务。两者经常被放在一起比较只是因为都会说 SQL,但它们回答的是不同的问题:“一千个用户能否安全地同时读写”与“单个进程能多快聚合十亿行”。 本页的行为陈述以 2026 年 8 月核对的双方官方文档为准:[PostgreSQL](https://www.postgresql.org/docs/current/) 与 [DuckDB](https://duckdb.org/docs/stable/)。涉及具体版本的细节请在依赖前复核。 ## 架构对比 [#架构对比] | | PostgreSQL | DuckDB | | ---- | ----------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------- | | 存储布局 | 行存(heap),为点查与小写入调优 | 列存、向量化执行,为扫描与聚合调优 | | 进程模型 | 独立服务器;客户端经网络连接(pgwire) | 进程内库或 CLI;数据库就是一个文件 | | 并发写入 | MVCC 下的多连接并发读写 | 要么一个进程以读写方式打开数据库,要么多个进程以只读方式打开(`access_mode = 'READ_ONLY'`);不支持多进程并发写入([并发文档](https://duckdb.org/docs/stable/connect/concurrency.html)) | | 事务 | 完整 ACID,可配置隔离级别 | 单进程内的 ACID 事务 | | 部署形态 | 自行运维的服务器,或托管服务 | 嵌入 Python/R/Java/Wasm/CLI,无服务可运维 | | 扩展生态 | 装载进服务器的扩展(pgvector、PostGIS、TimescaleDB……) | 可加载扩展(`postgres`、`parquet`、`iceberg`……) | | 服务模型 | 在线服务、API、多租户应用 | 本地分析、ETL、数据准备、边缘/嵌入式分析 | 并发写入那一行是实际的分界线:负载若是“来自多台机器的多个写入方”,DuckDB 在架构上就出局;负载若是“单个作业按磁盘允许的速度读列存数据”,客户端/服务器模型的每查询一次网络往返正是 DuckDB 所没有的开销。 ## 什么时候各选各的 [#什么时候各选各的] ### 选 DuckDB [#选-duckdb] * 对本地或对象存储上的文件(Parquet、CSV、JSON)做交互式分析,不导入任何地方。 * 流水线或 notebook 里的数据准备与转换环节。 * 无法附带一台服务器的边缘与嵌入式分析场景。 ### 选 PostgreSQL [#选-postgresql] * 有并发写入、约束与外键的多用户事务负载。 * 任何要 7×24 小时为 API 或应用供数的场景。 * 需要行级安全、逻辑复制、时间点恢复或扩展生态的负载——具体收益见[扩展生态](/docs/reference/extensions-ecosystem)与[云服务版图](/docs/cloud/service-map)。 ### PostgreSQL 侧的分析能力边界 [#postgresql-侧的分析能力边界] PostgreSQL 能正确执行分析查询,但相对列存引擎是逐行处理的;在大扫描上这个差距是结构性的,不是靠调参能消除的。不迁走 PostgreSQL 数据的前提下,有三条弥合路径: * **列存扩展**:[pg\_mooncake](https://github.com/Mooncake-Labs/pg_mooncake) 在 Iceberg 中维护 PostgreSQL 表的列存镜像,并用 DuckDB 执行引擎加速分析;[Citus](https://github.com/citusdata/citus) 在分布式能力之外提供列存表选项。 * **导出给 DuckDB**:PostgreSQL 保持为权威数据源,定期把快照导出为 Parquet 供分析——DuckDB 原生读取 Parquet。 * **推给数据仓库**:当负载彻底超出单机时,把分析整体移到仓库层。 ## 两者混用模式 [#两者混用模式] ### DuckDB 读 PostgreSQL:postgres\_scanner [#duckdb-读-postgresqlpostgres_scanner] DuckDB 官方的 [`postgres` 扩展](https://duckdb.org/docs/stable/core_extensions/postgres.html)可以直接挂载一个在线的 PostgreSQL 数据库并对其执行查询,包括条件下推: ```sql INSTALL postgres; ATTACH 'dbname=app user=analyst host=127.0.0.1' AS pg (TYPE postgres, READ_ONLY); SELECT status, count(*), avg(total_cents) FROM pg.orders WHERE placed_at >= now() - interval '30 days' GROUP BY status; ``` 这是“不导出就分析生产数据”的标准做法:DuckDB 拉取所需的行,重聚合发生在 DuckDB 的向量化引擎里。挂载时用只读模式和低权限的 PostgreSQL 角色,保证分析路径无法回写。 ### PostgreSQL 读文件:FDW [#postgresql-读文件fdw] 反方向上,PostgreSQL 的[外部数据包装器](https://www.postgresql.org/docs/current/ddl-foreign-data.html)可以把 Parquet 文件暴露为外部表(例如用 [`parquet_s3_fdw`](https://github.com/pgspider/parquet_s3_fdw))。适合的场景是:这些文件是关系处理的输入,需要在 PostgreSQL 权限体系下与在线表 join——它并不替代 DuckDB 的扫描速度。 ```sql CREATE EXTENSION parquet_s3_fdw; CREATE SERVER parquet_files FOREIGN DATA WRAPPER parquet_s3_fdw; CREATE FOREIGN TABLE lake_events (...) SERVER parquet_files OPTIONS (dirname 's3://analytics/events/', sorted 'event_time'); ``` server 与表级选项随 FDW 版本不同,准确写法以该扩展的 README 为准。 ## AI 场景对照 [#ai-场景对照] 两个引擎通常出现在同一条流水线的不同阶段: * **DuckDB 负责准备**:把原始导出和数据湖文件清洗、join、聚合成 AI 应用真正要供数的文档与表。没有服务器要运维,产出的 Parquet 直接进入下一阶段。 * **PostgreSQL 负责在线**:多租户事务状态、[RAG](/docs/ai/rag-pipeline) 里 pgvector 加全文检索的混合检索,以及 Agent [长期记忆](/docs/ai/agent-memory)这类状态——在这些场景里,并发、权限与可审计性才是重点。 一条经验法则:正在被*准备*的静态数据交给 DuckDB;正在被*服务*给用户和 Agent 的数据放在 PostgreSQL。 ## AI prompt:为负载选引擎 [#ai-prompt为负载选引擎] ## 相关页面 [#相关页面] * [PostgreSQL 血缘与兼容数据库](/docs/reference/postgresql-compatible-databases)——另一种对比:复用 PostgreSQL 本身的系统 * [云 PostgreSQL 服务版图](/docs/cloud/service-map)——选型落在服务器上时的托管选项 * [PostgreSQL 扩展生态](/docs/reference/extensions-ecosystem)——“选 PostgreSQL”所包含的 pgvector、PostGIS 等 * [RAG 管道](/docs/ai/rag-pipeline)——PostgreSQL 承担的在线检索路径 --- # 版本与支持策略 Canonical URL: https://pg.edu.rich/docs/reference/version-policy Last reviewed: 2026-08-02 ## 版本号含义 [#版本号含义] 从 PostgreSQL 10 起,第一个数字是 major,例如 18;点后的数字是 minor,例如 18.4。major 大约每年发布一次并带来新功能;minor 只包含错误、安全和低风险修复。 minor 升级不需要 dump/restore,通常替换二进制并重启;仍应阅读该版本 release notes。major 之间的数据目录不兼容,需要 `pg_upgrade`、逻辑 dump/restore 或逻辑复制迁移。 ## 当前支持快照 [#当前支持快照] 截至 **2026-08-02**: | Major | 当前 minor | 支持状态 | 最终支持日期 | | ----- | -------: | --------- | ---------- | | 18 | 18.4 | 支持 | 2030-11-14 | | 17 | 17.10 | 支持 | 2029-11-08 | | 16 | 16.14 | 支持 | 2028-11-09 | | 15 | 15.18 | 支持 | 2027-11-11 | | 14 | 14.23 | 支持,即将 EOL | 2026-11-12 | 来源:[PostgreSQL 官方版本策略](https://www.postgresql.org/support/versioning/)。动态版本信息以该页为准。 ## 上游版本不等于发行版软件包版本 [#上游版本不等于发行版软件包版本] ### 各发行版默认 postgresql 软件包节选 PkgSeek 软件包快照;核对时间:2026-08-02 15:47:58 UTC。 | 发行版 | Release | 完整版本 | 仓库 | 关联公告 | |---|---|---|---|---:| | Alibaba Cloud Linux | 3 | 13.23-3.0.1.al8 | official / updates | 0 | | Alibaba Cloud Linux | 4 | 15.18-1.alnx4 | official / updates | 0 | | AlmaLinux | 10 | 16.14-1.el10_2 | official / AppStream | 0 | | AlmaLinux | 9 | 18.4-2.module_el9.8.0+280+5ad12178 | official / AppStream | 0 | | Arch | rolling | 18.4-3 | official / extra | 0 | | CentOS Stream | 10 | 16.14-1.el10 | official / AppStream | 0 | | CentOS Stream | 9 | 13.23-3.el9 | official / AppStream | 0 | | Debian | trixie | 17+278 | official / main | 3 | | deepin | 25.2 | 16+255 | official / main | 0 | | Fedora | 42 | 16.13-1.fc42 | official / updates | 0 | | Fedora | 43 | 18.3-2.fc43 | official / updates | 0 | | Fedora | 44 | 18.3-2.fc44 | official / updates | 0 | 来源:[PkgSeek 软件包查询](https://pkgseek.com/packages/postgresql)。发行版 revision 和回溯补丁属于完整版本身份。 发行版可能冻结 major,并在 `16.14-1.el10_2`、`18+290ubuntu1` 等完整版本号中记录打包 revision 和安全回溯。不能只截取开头的 `16` 或 `18` 判断漏洞状态;应使用发行版、release、repository、architecture 和完整 package version 共同定位,再核对厂商安全公告。 ## 新项目怎么选 [#新项目怎么选] 默认选择最新稳定 major 的最新 minor,除非驱动、扩展、托管平台或组织认证尚未支持。需要更保守时选择仍有充足支持窗口、且已被自身工作负载验证的 major;不要为了“稳定”新建即将 EOL 的版本。 PostgreSQL 19 在 2026-08 仍处于 beta 周期,不作为生产默认。测试新 major 时重点验证扩展、collation、备份工具、连接池、ORM、查询计划和监控采集器。 当前 Beta 状态、功能变化和逐项迁移风险见 [PostgreSQL 19:新功能与 18 升级 19 指南](/docs/postgresql-19)。 ## 支持矩阵应写进仓库 [#支持矩阵应写进仓库] ```yaml postgresql: supported_majors: [17, 18] tested_minor_floor: 17: 17.10 18: 18.4 extensions: vector: "tested in CI" upgrade_owner: platform-database next_review: 2026-11-01 ``` 不要把 `latest` 当成部署策略。镜像、包和基础设施应锁定可审计版本,并由依赖更新流程推进 minor。 ## 升级原则 [#升级原则] * 总是运行所选 major 的当前 minor;继续运行旧 minor 往往比升级风险更高。 * 升级前读所有跨越版本的 release notes。 * 先在恢复出的真实数据副本上跑应用测试和查询计划对比。 * 扩展具有独立版本与升级脚本;分别检查。 * major 切换完成后重新收集统计并验证备份。 --- # PostgreSQL 安装与连接 Canonical URL: https://pg.edu.rich/docs/setup Last reviewed: 2026-08-02 ## 先选择目标 [#先选择目标] | 目标 | 推荐起点 | | -------- | -------------------------------- | | 学习、测试、CI | Docker;容易固定版本并完整删除 | | 本机长期开发 | 操作系统包管理器或受信安装器 | | 生产环境 | 云托管服务,或由团队管理的软件仓库与自动化配置 | | 只需要客户端 | 安装 `psql`/libpq 客户端包,不必运行本地数据库服务 | 不要为了“连得上”就把 5432 端口暴露到公网或把 `pg_hba.conf` 改成全网信任。先在本机验证,再设计网络、TLS、认证和最小权限。 ## 每种安装都执行同一验证 [#每种安装都执行同一验证] ```bash psql --version psql -X "postgresql://postgres@localhost:5432/postgres" \ -c "select current_setting('server_version'), current_database(), current_user;" ``` 客户端版本和服务端版本是两个概念。`psql --version` 只显示客户端;SQL 查询才显示实际连接的服务端。多个版本并存时记录可执行文件路径、端口和数据目录。 下一步阅读 [`psql` 连接与 SSL](/docs/setup/psql-connection),再创建不使用超级用户的应用角色。 --- # PostgreSQL Linux 软件包、版本与 PGDG 仓库 Canonical URL: https://pg.edu.rich/docs/setup/linux-packages Last reviewed: 2026-08-02 Linux 上的“安装 PostgreSQL”不是一个稳定不变的命令。发行版、release、仓库来源、CPU 架构和软件包名共同决定最终安装的 major、打包 revision 与安全修复。本页把 [PkgSeek](https://pkgseek.com/packages/postgresql) 的动态软件包证据与本站的安装、升级规则组合起来。 ### 各 Linux 发行版默认 postgresql 软件包 PkgSeek 软件包快照;核对时间:2026-08-02 15:47:58 UTC。 | 发行版 | Release | 完整版本 | 仓库 | 关联公告 | |---|---|---|---|---:| | Alibaba Cloud Linux | 3 | 13.23-3.0.1.al8 | official / updates | 0 | | Alibaba Cloud Linux | 4 | 15.18-1.alnx4 | official / updates | 0 | | AlmaLinux | 10 | 16.14-1.el10_2 | official / AppStream | 0 | | AlmaLinux | 9 | 18.4-2.module_el9.8.0+280+5ad12178 | official / AppStream | 0 | | Arch | rolling | 18.4-3 | official / extra | 0 | | CentOS Stream | 10 | 16.14-1.el10 | official / AppStream | 0 | | CentOS Stream | 9 | 13.23-3.el9 | official / AppStream | 0 | | Debian | trixie | 17+278 | official / main | 3 | | deepin | 25.2 | 16+255 | official / main | 0 | | Fedora | 42 | 16.13-1.fc42 | official / updates | 0 | | Fedora | 43 | 18.3-2.fc43 | official / updates | 0 | | Fedora | 44 | 18.3-2.fc44 | official / updates | 0 | | Kali Linux | kali-rolling | 18+290 | official / main | 0 | | Kylin OS | V10-SP1 | 12+214kylin0.1 | official / 10.1-main | 0 | | Kylin OS Server | V10-SP3-2403 | 10.5-23.p09.ky10 | official / updates | 0 | | OpenAnolis | 23.4 | 15.18-1.an23 | official / updates | 0 | | OpenAnolis | 8.10 | 12.22-7.0.1.module+an8.10.0+11420+6683745d | official / AppStream | 0 | | openEuler | 24.03-LTS-SP4 | 15.18-1.oe2403sp4 | official / everything | 0 | | openSUSE | 15.6 | 18-150600.17.9.1 | official / update-sle | 0 | | openSUSE | tumbleweed | 18-3.4 | official / oss | 0 | | Oracle Linux | 10 | 16.14-1.0.1.el10_2 | official / appstream | 0 | | Oracle Linux | 8 | 12.22-6.0.1.module+el8.10.0+90932+f6d78e3c | official / appstream | 34 | | Oracle Linux | 9 | 13.23-3.el9_8 | official / appstream | 26 | | Raspberry Pi OS | bookworm | 15+248+deb12u1 | official / bookworm-main | 0 | | Raspberry Pi OS | trixie | 17+278 | official / trixie-main | 0 | | Red Hat Enterprise Linux | 10.2 | 16.14-1.el10_2 | official / AppStream | 0 | | Red Hat Enterprise Linux | 9.8 | 18.4-2.module+el9.8.0+24359+da7fad50 | official / AppStream | 0 | | Rocky Linux | 10 | 16.14-1.el10_2 | official / AppStream | 0 | | Rocky Linux | 9 | 13.23-3.el9_8 | official / AppStream | 13 | | Ubuntu | focal | 12+214 | official / main | 0 | | Ubuntu | jammy | 14+238 | official / main | 0 | | Ubuntu | noble | 16+257build1 | official / main | 0 | | Ubuntu | resolute | 18+290ubuntu1 | official / main | 0 | | Void Linux | rolling | 18_1 | official / current | 0 | 来源:[PkgSeek 软件包查询](https://pkgseek.com/packages/postgresql)。发行版 revision 和回溯补丁属于完整版本身份。 ## 先分清四类包名 [#先分清四类包名] | 类型 | 常见示例 | 含义 | | ------- | ----------------------------------------------------- | ------------------------------------- | | 发行版元包 | `postgresql` | 跟随该发行版选择的默认 major | | 版本化服务端 | `postgresql-18`、`postgresql18-server` | 明确绑定 major,实际命名因 DEB/RPM 生态不同 | | 客户端与开发包 | `postgresql-client-18`、`libpq-dev`、`postgresql-devel` | 提供 `psql`、libpq header 或编译文件,不一定运行服务端 | | 扩展包 | `postgresql-18-pgvector`、`pgvector` | 还要与 server major、架构和扩展版本一起核对 | 只看到包名相似不能证明用途相同。安装后分别验证客户端和服务端: ```bash psql --version sudo -u postgres psql -X -d postgres \ -c "select version(), current_setting('server_version_num');" ``` ## 发行版仓库与 PGDG [#发行版仓库与-pgdg] 发行版官方仓库通常维护其选定的 major;[PostgreSQL Global Development Group 仓库](https://www.postgresql.org/download/linux/)常用于取得其他受支持 major。选择 PGDG 会增加一条外部软件源,也意味着需要持续核对仓库签名、release 支持、升级策略和退出路径。 ### Ubuntu 官方仓库中的 postgresql-18 PkgSeek 软件包快照;核对时间:2026-08-02 15:47:58 UTC。 | 发行版 | Release | 完整版本 | 仓库 | 关联公告 | |---|---|---|---|---:| | Ubuntu | resolute | 18.3-1 | official / main | — | 来源:[PkgSeek 软件包查询](https://pkgseek.com/packages/postgresql-18)。发行版 revision 和回溯补丁属于完整版本身份。 ### Ubuntu PGDG 中的 postgresql-18 PkgSeek 软件包快照;核对时间:2026-08-02 15:47:58 UTC。 > 没有准确索引坐标。这表示当前快照未收录,不等于软件包不存在。 来源:[PkgSeek 软件包查询](https://pkgseek.com/search?q=postgresql-18)。发行版 revision 和回溯补丁属于完整版本身份。 上方分别查询 Ubuntu 官方仓库与 PGDG 的准确坐标。若卡片显示“没有准确索引坐标”,它描述的是索引覆盖,不是对目标仓库内容的否定。需要安装时,以目标系统实际的 `apt-cache policy` 和 PostgreSQL 官方仓库说明为最终依据。 ```bash apt-cache policy postgresql postgresql-18 postgresql-client-18 apt-cache madison postgresql-18 ``` ## PostgreSQL 19 软件包跟踪 [#postgresql-19-软件包跟踪] ### 发行版仓库中的 postgresql-19 PkgSeek 软件包快照;核对时间:2026-08-02 15:47:58 UTC。 > 没有准确索引坐标。这表示当前快照未收录,不等于软件包不存在。 来源:[PkgSeek 软件包查询](https://pkgseek.com/search?q=postgresql-19)。发行版 revision 和回溯补丁属于完整版本身份。 在 PostgreSQL 19 仍为 Beta 时,缺少稳定发行版包是正常状态。即使测试构建出现,也应把 Beta 兼容性通道与生产 PostgreSQL 18 分开;升级判断见 [PostgreSQL 18 升级 19 指南](/docs/postgresql-19)。 ## 安全地使用软件包证据 [#安全地使用软件包证据] 1. 记录 `distro + release + source + repository + architecture + full version`。 2. 先确认包提供的是服务端、客户端、开发文件还是扩展。 3. 对 CVE 优先使用发行版厂商公告和回溯状态;不要只比较上游版本前缀。 4. 软件包安装成功后验证实际连接的服务端,不要只运行 `psql --version`。 5. major 改变时走 `pg_upgrade`、dump/restore 或逻辑复制,不能把换包当成升级数据目录。 软件包索引有覆盖范围和刷新时间。页面显示“没有准确索引坐标”时,应继续核对上游仓库;不要让人或 AI 根据空数组生成确定性结论。 ## 给 AI / Agent 的最小契约 [#给-ai--agent-的最小契约] ```yaml task: resolve_postgresql_package target: distro: ubuntu release: noble architecture: amd64 source: pgdg requirements: - return exact package coordinates and observed_at - separate indexed fact, inference, and unknown - never interpret missing index data as package absence - ask before repository, package, service, or data changes - verify client and server versions after execution ``` PkgSeek 提供[只读 MCP 工具目录](https://pkgseek.com/mcp/tools)和 [OpenAPI 3.1](https://pkgseek.com/openapi.json);接入边界见 [AI 上下文契约](/docs/ai/context-contract)。 --- # macOS 安装 PostgreSQL 18 Canonical URL: https://pg.edu.rich/docs/setup/macos Last reviewed: 2026-08-02 [PostgreSQL macOS 下载页](https://www.postgresql.org/download/macosx/)列出三条常用路径:EDB 图形安装器、Postgres.app 和 Homebrew。不要同时让多套服务监听 5432;先选择一种并记录数据目录。 ## Homebrew [#homebrew] ```bash brew update brew install postgresql@18 brew services start postgresql@18 "$(brew --prefix postgresql@18)/bin/psql" --version ``` Homebrew 可能不自动把带版本后缀的客户端放入默认 `PATH`。需要长期使用时,按 `brew info postgresql@18` 给出的路径配置 shell,而不是复制一个可能随升级改变的 Cellar 路径。 验证连接: ```bash "$(brew --prefix postgresql@18)/bin/psql" -X -d postgres \ -c "select version(), current_setting('data_directory');" ``` 若本地角色/数据库名不同,用 `-U`、`-d` 明确指定。 ## Postgres.app 或图形安装器 [#postgresapp-或图形安装器] * Postgres.app 适合希望通过菜单栏启动/停止、减少系统配置的本地开发。 * EDB 安装器包含服务端、pgAdmin 和 StackBuilder,适合需要图形化安装流程的用户。 安装完成后不要只看图标,应使用安装目录中的 `psql` 执行: ```sql SELECT version(), current_database(), current_user; ``` 确认下载包或 Homebrew 前缀对应 arm64/amd64。不要从另一架构复制数据目录;跨环境迁移优先使用逻辑备份或经过验证的升级流程。 ## 多版本诊断 [#多版本诊断] ```bash which -a psql psql --version lsof -nP -iTCP:5432 -sTCP:LISTEN ``` 若 `psql` 客户端版本与服务端不同,不必立即重装;先确认实际连接目标。通常使用不旧于服务端的客户端更稳妥。 --- # psql 连接 PostgreSQL 与 SSL 配置 Canonical URL: https://pg.edu.rich/docs/setup/psql-connection Last reviewed: 2026-08-06 ## 明确连接五要素 [#明确连接五要素] ```bash psql -X \ --host=db.example.com \ --port=5432 \ --username=app_reader \ --dbname=commerce ``` 目标由 host、port、database、user 和 TLS 参数共同决定。不要只看数据库名;同名 database 可以存在于多个实例。 连接 URI 等价写法: ```bash psql -X "postgresql://app_reader@db.example.com:5432/commerce?sslmode=verify-full" ``` 不要把密码写入命令行 URI、源代码或日志。交互使用提示,自动化使用 secret manager、短期凭据、`.pgpass` 或 libpq service file。 ## pgpass [#pgpass] Unix 默认文件是 `~/.pgpass`,权限必须限制为 `0600`: ```text hostname:5432:database:username:password ``` ```bash chmod 600 ~/.pgpass ``` Windows 默认位置是 `%APPDATA%\postgresql\pgpass.conf`。通配符会扩大凭据适用范围,应尽量写具体 host、database 和 user。 ## 连接 service 文件 [#连接-service-文件] libpq 的 service 文件可以把 host、user 和 TLS 参数从命令行与 shell history 中移出。默认路径是 `~/.pg_service.conf`,可用 `PGSERVICEFILE` 覆盖: ```text [prod] host=db.example.com port=5432 user=app_reader dbname=commerce sslmode=verify-full ``` ```bash psql -X service=prod ``` 应用程序也可以通过 `PGSERVICE=prod` 引用同一条配置。密码仍放在 `.pgpass` 或 secret manager 中,不要写进 service 文件。 ## SSL 模式 [#ssl-模式] | `sslmode` | 行为 | 使用建议 | | ------------- | ------------- | -------------- | | `disable` | 不使用 TLS | 仅受控本机/隔离测试 | | `require` | 要求加密,但不完整验证身份 | 比明文好,不足以抵抗错误端点 | | `verify-ca` | 验证证书链 | 仍不验证主机名 | | `verify-full` | 验证证书链和主机名 | 远程生产连接的推荐目标 | `verify-full` 要求 URI 中的 host 与证书身份匹配,并正确配置根证书。云平台可能有自己的 CA 轮换流程,不能永久固定一份过期证书。 ## 连接后立即确认 [#连接后立即确认] ```sql \conninfo SELECT current_database(), current_user, session_user, inet_server_addr(), inet_server_port(), current_setting('server_version') AS server_version, current_setting('TimeZone') AS timezone; ``` 脚本建议使用: ```bash psql -X --set ON_ERROR_STOP=on --file migration.sql "$DATABASE_URL" ``` `-X` 避免用户 `.psqlrc` 改变自动化行为;`ON_ERROR_STOP` 让脚本在 SQL 错误时退出。`psql` 退出码语义见 [PostgreSQL 18 psql 文档](https://www.postgresql.org/docs/18/app-psql.html)。 ## 面向脚本的输出格式 [#面向脚本的输出格式] 交互式输出默认是表格对齐;管道处理需要非对齐、仅元组的输出: ```bash psql -X -A -F, -t -c "SELECT id, email FROM users" > users.csv ``` `-A` 关闭对齐,`-F,` 指定字段分隔符,`-t` 只输出数据行。会话内的等价控制是 `\pset format csv`、`\pset null '[NULL]'`,以及 `\o /tmp/out.txt`(再次执行 `\o` 回到标准输出)。 ## 实用 one-liner [#实用-one-liner] 按总大小列出最大的表: ```bash psql -X -c " SELECT schemaname||'.'||relname AS table_name, pg_size_pretty(pg_total_relation_size(relid)) AS total_size FROM pg_stat_user_tables ORDER BY pg_total_relation_size(relid) DESC LIMIT 10;" ``` 当前正在执行的查询,按开始时间排序: ```bash psql -X -c " SELECT pid, now() - query_start AS duration, state, query FROM pg_stat_activity WHERE state = 'active' AND query NOT ILIKE '%pg_stat_activity%' ORDER BY query_start LIMIT 20;" ``` 确认 PID 后终止失控的后端进程: ```bash psql -X -c "SELECT pg_terminate_backend(12345);" ``` ## 交互式 \~/.psqlrc [#交互式-psqlrc] `~/.psqlrc` 在每次交互式启动时执行;`-X` 会跳过它,因此不影响自动化。一个最小模板: ```text \timing on \pset null '[NULL]' \x auto \pset linestyle unicode \pset border 2 \set conns 'SELECT pid, usename, application_name, state, query FROM pg_stat_activity WHERE state <> ''idle'';' \set locks 'SELECT pid, mode, locktype, relation::regclass, granted FROM pg_locks WHERE NOT granted;' \set HISTFILE ~/.psql_history- :DBNAME \set HISTCONTROL ignoredups ``` 之后 `:conns` 和 `:locks` 会展开为存储的查询;按库分离的 history 文件避免不同数据库的命令互相污染。 连接参数控制建立会话要等多久;`statement_timeout` 控制 SQL 执行;`lock_timeout` 只控制等待锁。应用层还要设置请求截止时间,并确保超时后取消或释放数据库连接。 连接失败时记录完整错误和 SQLSTATE,再查看[错误速查](/docs/reference/errors),不要通过关闭 TLS 或扩大权限来试错。 --- # Ubuntu 安装 PostgreSQL 18 Canonical URL: https://pg.edu.rich/docs/setup/ubuntu Last reviewed: 2026-08-02 Ubuntu 自带 PostgreSQL 包,但其 major 版本由 Ubuntu 发行版快照决定。如果只需要该发行版维护的默认版本: ```bash sudo apt update sudo apt install postgresql postgresql-client ``` ### Ubuntu 默认 postgresql 元包 PkgSeek 软件包快照;核对时间:2026-08-02 15:47:58 UTC。 | 发行版 | Release | 完整版本 | 仓库 | 关联公告 | |---|---|---|---|---:| | Ubuntu | focal | 12+214 | official / main | — | | Ubuntu | jammy | 14+238 | official / main | — | | Ubuntu | noble | 16+257build1 | official / main | — | | Ubuntu | resolute | 18+290ubuntu1 | official / main | — | 来源:[PkgSeek 软件包查询](https://pkgseek.com/packages/postgresql)。发行版 revision 和回溯补丁属于完整版本身份。 这个快照说明不同 Ubuntu release 的默认 major 并不相同。`postgresql` 是跟随发行版的元包,不等于永远安装 PostgreSQL 18。完整区别见 [Linux 软件包与 PGDG](/docs/setup/linux-packages)。 ## 明确安装 PostgreSQL 18 [#明确安装-postgresql-18] 需要指定 major 时,使用 PostgreSQL 项目维护的 Apt 仓库。官方提供自动配置脚本: ```bash sudo apt install -y postgresql-common ca-certificates sudo /usr/share/postgresql-common/pgdg/apt.postgresql.org.sh sudo apt update sudo apt install postgresql-18 postgresql-client-18 ``` 执行脚本前阅读其输出并确认系统版本受支持。命令和当前支持的 Ubuntu 版本以 [PostgreSQL Ubuntu 下载页](https://www.postgresql.org/download/linux/ubuntu/) 为准。 ### PkgSeek:Ubuntu / PGDG / postgresql-18 PkgSeek 软件包快照;核对时间:2026-08-02 15:47:58 UTC。 > 没有准确索引坐标。这表示当前快照未收录,不等于软件包不存在。 来源:[PkgSeek 软件包查询](https://pkgseek.com/search?q=postgresql-18)。发行版 revision 和回溯补丁属于完整版本身份。 若卡片没有返回准确的 PGDG 坐标,不能把空结果解释为软件包不存在。安装前以 PGDG 仓库元数据和 PostgreSQL 官方下载页复核;本站保留空状态,是为了显式暴露数据覆盖边界,而不是用推测补齐。 ## 验证服务和集群 [#验证服务和集群] ```bash systemctl status postgresql --no-pager pg_lsclusters sudo -u postgres psql -X -d postgres \ -c "select version(), current_setting('data_directory');" ``` `postgresql.service` 是集群管理入口;具体实例通常对应 `postgresql@18-main`。`pg_lsclusters` 属于 Debian/Ubuntu 的 `postgresql-common` 工具,不是所有 Linux 发行版都有。 创建一个本地练习角色和数据库: ```bash sudo -u postgres createuser --pwprompt learner sudo -u postgres createdb --owner=learner learner psql -X -h localhost -U learner -d learner -c "select current_user;" ``` 远程连接需要同时考虑 `listen_addresses`、`pg_hba.conf`、防火墙和 TLS。修改前保存原配置,规则从具体网段/数据库/角色写起,并同时测试允许和拒绝路径。 ## 升级边界 [#升级边界] `apt upgrade` 可以安装同一 major 的 minor 修复;从 17 到 18 是大版本升级,需要 `pg_upgrade`、逻辑 dump/restore 或逻辑复制,不能只替换软件包后复用数据目录。 --- # Windows 安装 PostgreSQL 18 Canonical URL: https://pg.edu.rich/docs/setup/windows Last reviewed: 2026-08-02 PostgreSQL 项目在 [Windows 下载页](https://www.postgresql.org/download/windows/)链接 EDB 认证的交互式安装器。安装包通常包含 PostgreSQL 服务端、pgAdmin 和 StackBuilder。 ## 安装时记录 [#安装时记录] * major 版本与安装目录; * 数据目录,避免放在会被同步软件接管的位置; * PostgreSQL 服务账号; * 监听端口,默认通常为 5432; * `postgres` 管理角色密码——由密码管理器保存; * locale,生产迁移前要验证 collation 行为。 只勾选真正需要的组件。StackBuilder 中的附加驱动和扩展不是 PostgreSQL 核心,按项目需求安装。 ## 用 psql 验证 [#用-psql-验证] 打开安装器提供的 SQL Shell,或把 PostgreSQL `bin` 目录加入当前终端路径: ```powershell psql.exe --version psql.exe -X -h localhost -p 5432 -U postgres -d postgres ` -c "select version(), current_database(), current_user;" ``` 通过“服务”管理器或 PowerShell 查看服务状态: ```powershell Get-Service *postgres* ``` 若客户端找不到,先定位安装目录中的 `psql.exe`,不要下载来源不明的单独 DLL。 ## 常见连接失败 [#常见连接失败] | 现象 | 先检查 | | ------------------------------ | ------------------------------ | | connection refused | PostgreSQL 服务是否运行、端口是否正确 | | password authentication failed | 用户名、目标实例、密码和 `pg_hba.conf` 规则 | | database does not exist | `-d` 指定的 database 是否创建 | | 连接到意外版本 | Windows 上是否运行了多个 PostgreSQL 服务 | 本机开发通常不需要开放入站防火墙。远程访问必须限定来源地址并启用适当 TLS 校验;不要把 `pg_hba.conf` 设为无密码 `trust` 来绕过诊断。 ## 卸载前 [#卸载前] 安装器卸载和删除数据目录是两件事。先用 `pg_dump`/`pg_dumpall --globals-only` 导出需要保留的内容并验证恢复,再确认具体数据目录;不要根据文件夹名称猜测后直接删除。 --- # PostgreSQL tutorial and production guide Canonical URL: https://pg.edu.rich/en/docs Last reviewed: 2026-08-02 This guide is not a replacement for the PostgreSQL manual. It is a map into it. The current baseline is **PostgreSQL 18**, while pages avoid version-specific behavior unless they label it explicitly. ## Two reading modes [#two-reading-modes] ### Human learning mode [#human-learning-mode] Follow [Start here](/en/docs/start-here) in order. Each page states the outcome, gives a minimal example, explains the model, verifies the result, and points to a next step. You do not need to memorize system catalogs or the full SQL grammar. ### AI retrieval mode [#ai-retrieval-mode] Enter through the [AI / agent index](/en/docs/ai). Pages use stable headings, explicit preconditions, copyable SQL, boundaries, and failure modes. Retrieve the relevant section; do not inject the whole site into one prompt. ## Documentation conventions [#documentation-conventions] | Label | Meaning | | --------------- | --------------------------------------------------------------------- | | **Default** | PostgreSQL default behavior; verify it on the target instance | | **Recommended** | A strong default for new systems, not the only valid choice | | **Danger** | May lock data, lose data, leak privilege, or create long transactions | | **Verify** | A command or query that confirms the result | For normative detail, use the [PostgreSQL 18 manual](https://www.postgresql.org/docs/18/). This field guide contributes paths, examples, guardrails, and cross-topic connections. --- # PostgreSQL 19 release and upgrade guide Canonical URL: https://pg.edu.rich/en/docs/postgresql-19 Last reviewed: 2026-08-06 As of **2026-08-02**, the newest public PostgreSQL 19 test build is **Beta 2**, not a production release. The PostgreSQL roadmap plans version 19 for **September 2026**, while the Beta 2 announcement gives a more conservative **September/October 2026** window. The final date, feature details, and compatibility requirements can still change. The project encourages testing with representative workloads but explicitly advises against production use. Use this period to build a compatibility matrix and rehearse a PostgreSQL 18 to 19 upgrade. Wait for GA plus support from your extensions and managed platform before cutting over. ## PostgreSQL 19 news and release date [#postgresql-19-news-and-release-date] | Date | Official update | What it means | | ------------------------ | ---------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------- | | 2026-06-04 | [PostgreSQL 19 Beta 1 released](https://www.postgresql.org/about/news/postgresql-19-beta-1-released-3313/) | Feature preview opened for CI and application compatibility testing | | 2026-07-16 | [PostgreSQL 19 Beta 2 released](https://www.postgresql.org/about/news/postgresql-19-beta-2-released-3350/) | Fixed Beta 1 regressions and continued changes around temporal SQL, SQL/PGQ, logical decoding, and autovacuum | | September 2026 (planned) | [Target month on the official roadmap](https://www.postgresql.org/developer/roadmap/) | Not an immutable promise; the Beta announcement retains a September/October window | Beta 2 still permits small changes to behavior, APIs, and feature details. Track the [PostgreSQL 19 release notes](https://www.postgresql.org/docs/19/release-19.html) and project news rather than treating a third-party feature list as a launch contract. ## PostgreSQL 19 Linux package availability [#postgresql-19-linux-package-availability] ### Distribution repositories: postgresql-19 PkgSeek package snapshot; checked: 2026-08-02 15:47:58 UTC. > No exact indexed coordinate. This means the current snapshot has no match, not that the package does not exist. Source: [PkgSeek package lookup](https://pkgseek.com/search?q=postgresql-19). Distribution revisions and backported fixes are part of the complete version identity. ### PGDG repositories: postgresql-19 PkgSeek package snapshot; checked: 2026-08-02 15:47:58 UTC. > No exact indexed coordinate. This means the current snapshot has no match, not that the package does not exist. Source: [PkgSeek package lookup](https://pkgseek.com/search?q=postgresql-19). Distribution revisions and backported fixes are part of the complete version identity. These snapshots track repository adoption; they are not the authority for PostgreSQL 19 release status. A beta tarball, development build, or container does not prove that a distribution repository provides a production-ready `postgresql-19` package. After GA, wait for a complete matrix across the target OS, architecture, extensions, and backup tooling. ## What is new in PostgreSQL 19 [#what-is-new-in-postgresql-19] These are the areas most worth testing as of this review date, not a substitute for the final release notes. ### SQL, graph queries, and temporal data [#sql-graph-queries-and-temporal-data] * **SQL/PGQ property graph queries** define and query property graphs over relational data. Test whether drivers, SQL parsers, ORMs, and AI SQL generators recognize the syntax. * **`FOR PORTION OF`** lets `UPDATE` and `DELETE` operate on temporal ranges; Beta 2 still contained several fixes in this area. * **`GROUP BY ALL`** groups all non-aggregate, non-window target-list items. * **Window function `IGNORE NULLS` / `RESPECT NULLS`** support applies to `lead()`, `lag()`, `first_value()`, `last_value()`, and `nth_value()`. * **`INSERT ... ON CONFLICT DO SELECT ... RETURNING`** can return the conflicting row and optionally lock it. ### Operations, performance, and observability [#operations-performance-and-observability] * **`REPACK` and `REPACK CONCURRENTLY`** unify table-rewrite behavior associated with `VACUUM FULL` / `CLUSTER` and add a path with less access-exclusive locking; the old commands remain for compatibility. See [REPACK and online table rewrites](/en/docs/operations/repack). * Partition split and merge through `ALTER TABLE ... SPLIT/MERGE PARTITIONS`. * Parallel autovacuum workers (see [parallel autovacuum](/en/docs/operations/parallel-autovacuum)) and new views including `pg_stat_autovacuum_scores`, `pg_stat_lock`, and `pg_stat_recovery`. * Optimizer work includes eager aggregation — performing some aggregate processing before joins to cut the number of rows processed — and converting `NOT IN` to more efficient anti joins when the columns are known non-nullable. Plan comparisons between versions should pin these changes down rather than attributing them to indexes. * Data checksums can be enabled and disabled **online**, instead of only offline with `pg_checksums`; see [data checksums](/en/docs/operations/data-checksums). * Improvements to asynchronous-I/O read-ahead, SIMD `COPY FROM`, radix sort, and foreign-key checks. * An `IO` option for `EXPLAIN ANALYZE`, plus full-page-write bytes in `EXPLAIN (ANALYZE, WAL)`. * LZ4 replaces pglz as the default TOAST compression method for newly compressed data. ### Logical replication and standby waits [#logical-replication-and-standby-waits] * **Sequence synchronization** lets a subscriber's sequence values match the publisher: publish them with `CREATE PUBLICATION ... FOR ALL SEQUENCES`, and reconcile values on the subscriber with `ALTER SUBSCRIPTION ... REFRESH SEQUENCES`; `pg_get_sequence_data()` reports the synchronization state. This directly targets the sequence/primary-key conflicts seen after logical-replication major upgrades — see [Replication, failover, and upgrades](/en/docs/operations/replication-upgrades). * **`WAIT FOR LSN`** lets a session wait until a given LSN is written, flushed, or replayed — including on a standby — which enables read-your-writes flows across a primary/standby split. See [WAIT FOR LSN](/en/docs/operations/wait-for-lsn) and the [PostgreSQL 19 WAIT FOR documentation](https://www.postgresql.org/docs/19/sql-wait-for.html). * **Publication blacklists**: `CREATE` / `ALTER PUBLICATION ... FOR ALL TABLES EXCEPT (TABLE ...)` publishes everything except the named tables, replacing per-table allowlists for large schemas. * **Conflict retention on subscribers**: the subscription parameters `retain_dead_tuples` and `max_retention_duration` keep dead-tuple information used for conflict detection, bounded by a retention window; `pg_stat_subscription_stats.update_deleted` reports updates ignored due to concurrent deletes. * **`effective_wal_level`** reports the effective WAL level; with `wal_level = replica`, the server can raise the effective level to `logical` automatically when logical replication requires it. PostgreSQL 19 `REPACK (CONCURRENTLY)` is a core SQL command based on logical decoding, with constraints around replica identity, unlogged/partitioned/system tables, replication slots, and extra disk. Third-party `pg_repack` is an independent extension and CLI; do not reuse one runbook for the other without testing. See the [extensions ecosystem guide](/en/docs/reference/extensions-ecosystem) and [PostgreSQL 19 REPACK documentation](https://www.postgresql.org/docs/19/sql-repack.html). If an agent targets PostgreSQL 18, do not let it generate SQL/PGQ, `FOR PORTION OF`, `GROUP BY ALL`, or other version-19 syntax. Put `server_version_num`, allowed syntax, and extension versions into retrieval context or the tool contract. ## PostgreSQL 18 to 19 upgrade considerations [#postgresql-18-to-19-upgrade-considerations] PostgreSQL 18 → 19 is a **major upgrade**. A version-18 data directory cannot simply be started by version 19. Use `pg_upgrade`, logical dump/restore, or logical replication, and address these compatibility changes first. ### 1. Authentication and security [#1-authentication-and-security] * **RADIUS support is removed**. Environments that still depend on it need an alternate authentication design before upgrading. * PostgreSQL 18 deprecated MD5 passwords; version 19 warns after successful MD5 authentication. Move toward SCRAM instead of merely suppressing the warning. * A password-expiration warning is added with a default seven-day threshold. Ensure monitoring does not misclassify an expected warning as an outage. ### 2. SQL and object compatibility [#2-sql-and-object-compatibility] * The server now forces `standard_conforming_strings` to `on`. If an old environment used `off`, create logical dumps with PostgreSQL 19 `pg_dump` / `pg_dumpall`, or first correct the setting and application escaping behavior. * Database, role, and tablespace names cannot contain CR/LF; `pg_upgrade` rejects affected clusters. * `btree_gist` indexes over `inet` / `cidr` block `pg_upgrade` because the old operator classes can miss rows. Let `pg_upgrade --check` identify the actual blockers and follow the release notes. * The `MULE_INTERNAL` encoding is removed, so affected databases require dump/restore to another encoding. ### 3. Defaults, performance, and monitoring [#3-defaults-performance-and-monitoring] * **JIT is disabled by default**. Analytical workloads must not assume unchanged plans or runtime; benchmark with `jit=off` and `jit=on`. * **`log_lock_waits` is enabled by default**, so lock waits beyond `deadlock_timeout` reach the server log without configuration. Expect a sudden increase in lock-wait log lines after upgrading; adjust log-volume alerts and lock monitoring instead of treating the new noise as a regression. * `max_locks_per_transaction` changes from 64 to 128 and lock-memory accounting changes. Recalculate capacity instead of copying the old number blindly. * `pg_stat_subscription_stats.sync_error_count` becomes `sync_table_error_count`; wait-event type `BUFFERPIN` becomes `BUFFER`. Update dashboards, alerts, and collectors. * The TOAST default affects newly written/compressed values; it does not automatically rewrite all old TOAST data to LZ4. ### 4. Extensions, drivers, and platforms [#4-extensions-drivers-and-platforms] `pg_upgrade` checks many core binary properties, but cannot prove third-party modules are binary-compatible with PostgreSQL 19. Record explicit support for PostGIS, pgvector, TimescaleDB, custom C extensions, audit modules, backup agents, pools, ORMs, and drivers. | Component | Verify | | --------------- | ------------------------------------------------------------------------------------------------- | | Extension | Version-19 package/shared library, support statement, update script, index rebuilds | | Driver and ORM | Server-version detection, new/changed grammar, prepared statements, type mapping | | Connection pool | Startup parameters, authentication, failover, and connection recycling | | Backup and CDC | New catalogs, WAL/logical decoding, and a tested restore | | Managed service | Region, SKU, extension version, maintenance window, and rollback; wait for provider documentation | ## PostgreSQL 18 to 19 upgrade checklist [#postgresql-18-to-19-upgrade-checklist] ### Phase A: work you can do now [#phase-a-work-you-can-do-now] 1. Preserve a PostgreSQL 18 production backup and prove it can be restored. 2. Inventory extensions, collations, slots, tablespaces, custom full-text files, authentication, and external modules. 3. Create a **disposable** PostgreSQL 19 Beta 2 environment and run migrations, application tests, restore, CDC, and critical-query benchmarks. 4. Compare `EXPLAIN (ANALYZE, BUFFERS, WAL)` and evaluate the JIT-default and I/O changes separately. 5. Prevent CI or AI agents from sending version-19-only syntax to PostgreSQL 18. Make PostgreSQL 18.4 the release-blocking production gate and PostgreSQL 19 Beta 2 a forward-compatibility lane. The latter may initially allow failures, but every failure should be classified and cleared before GA adoption. See the complete [safe migration and zero-downtime schema workflow](/en/docs/operations/safe-migrations). ### Phase B: before the production cutover [#phase-b-before-the-production-cutover] Run check-only mode using the **version 19** `pg_upgrade` binary. Replace every path with the real target layout: ```bash /opt/postgresql/19/bin/pg_upgrade \ --old-bindir=/opt/postgresql/18/bin \ --new-bindir=/opt/postgresql/19/bin \ --old-datadir=/data/postgresql/18 \ --new-datadir=/data/postgresql/19 \ --check ``` `--check` does not migrate data, but the rehearsal should use the same binaries, extensions, initdb options, and transfer mode as the cutover. Do not copy the example paths into production unchanged. Then: * pin a PostgreSQL 19 GA minor, OS package or container digest, and extension versions; * choose `pg_upgrade` copy/clone/link/swap, dump/restore, or logical replication based on the measured workload; * record downtime, extra disk, statistics rebuild time, and routing/pool convergence time; * define rollback criteria, owner, and the last safe rollback point; * verify standbys, slots, sequences, large objects, privileges, RLS, schedulers, and backups. ### Phase C: after cutover [#phase-c-after-cutover] 1. Run the post-upgrade or rebuild scripts produced by `pg_upgrade`; do not access tables it flags until those scripts complete. 2. Regenerate the missing optimizer statistics, then compare high-traffic plans and latency. 3. Inspect errors, authentication warnings, replication lag, WAL, autovacuum, locks, and backup jobs. 4. Restore a fresh backup taken from PostgreSQL 19. 5. Remove the PostgreSQL 18 cluster only after acceptance and the rollback window close. See [PostgreSQL 19 pg\_upgrade](https://www.postgresql.org/docs/19/pgupgrade.html) for the complete procedure and [Replication, failover, and upgrades](/en/docs/operations/replication-upgrades) for cutover design. ## PostgreSQL 19 FAQ [#postgresql-19-faq] ### Has PostgreSQL 19 been released? [#has-postgresql-19-been-released] No. As of 2026-08-02, Beta 2 is the newest release. The roadmap targets September 2026 and the Beta announcement gives a September/October window. Treat official project news as authoritative. ### Can PostgreSQL 18 upgrade directly to PostgreSQL 19? [#can-postgresql-18-upgrade-directly-to-postgresql-19] Yes, using a major-upgrade method; there is no requirement to pass through another major. Common choices are `pg_upgrade`, dump/restore, or logical replication. Use Beta only for rehearsals and wait for GA plus dependency support for production. ### How much downtime does an 18-to-19 upgrade need? [#how-much-downtime-does-an-18-to-19-upgrade-need] There is no universal number. Data size, relation count, transfer mode, extensions and reindexing, statistics, routing, and validation all contribute. Rehearse against a restored copy of representative production data and measure it. ### When will managed PostgreSQL services support version 19? [#when-will-managed-postgresql-services-support-version-19] Provider, region, and SKU timelines differ. Do not infer availability from community GA. Track each provider's version matrix and verify extensions, PITR, replicas, and rollback limits. See [Cloud PostgreSQL service map](/en/docs/cloud/service-map). ### Should I upgrade from PostgreSQL 18 now? [#should-i-upgrade-from-postgresql-18-now] Build the test matrix now, but do not make a Beta the production default. After GA, wait for explicit support from the extensions, drivers, tools, and managed platform your system actually uses, then schedule around measured benefit and risk. ## Fact status and review [#fact-status-and-review] Snapshot: **PostgreSQL 19 Beta 2, reviewed 2026-08-02**. When the project reaches RC or GA, or the release notes add important incompatibilities, update the status, news timeline, upgrade blockers, and `updatedAt` together. --- # 5-minute quickstart Canonical URL: https://pg.edu.rich/en/docs/quickstart Last reviewed: 2026-08-02 This instance is for local learning only. A password on the command line and a host-published port are not production configuration. ### Start the instance [#start-the-instance] ```bash docker run --name pg-guide \ -e POSTGRES_PASSWORD=dev-only-password \ -e POSTGRES_DB=playground \ -p 5432:5432 \ -v pg-guide-data:/var/lib/postgresql/data \ -d postgres:18 ``` ### Wait and check [#wait-and-check] ```bash docker exec pg-guide pg_isready -U postgres -d playground docker logs pg-guide --tail 20 ``` Continue after `accepting connections` appears. ### Open psql [#open-psql] ```bash docker exec -it pg-guide psql -U postgres -d playground ``` ```sql CREATE TABLE notes ( id bigint GENERATED ALWAYS AS IDENTITY PRIMARY KEY, body text NOT NULL CHECK (length(body) > 0), created_at timestamptz NOT NULL DEFAULT now() ); INSERT INTO notes (body) VALUES ('hello, PostgreSQL'); SELECT id, body, created_at FROM notes; ``` ### Verify persistence [#verify-persistence] ```bash docker restart pg-guide docker exec pg-guide psql -U postgres -d playground \ -c "SELECT id, body, created_at FROM notes;" ``` If the row survives the restart, the named volume is working. ## Connection string [#connection-string] Applications on the host can use: ```text postgresql://postgres:dev-only-password@127.0.0.1:5432/playground ``` Never commit real credentials. Production systems need a secret manager, a least-privilege application role, and TLS. ## Clean up [#clean-up] ```bash docker rm -f pg-guide docker volume rm pg-guide-data ``` The second command permanently deletes the practice data. Run it only when that is intentional. Continue with [Start here](/en/docs/start-here), or jump to [Data modeling](/en/docs/core/data-modeling). --- # Start here Canonical URL: https://pg.edu.rich/en/docs/start-here Last reviewed: 2026-08-02 At the end of this route you should be able to run an instance, connect with `psql`, design a constrained table, change data inside a transaction, and use `EXPLAIN` to check how a query executes. ## Remember five things first [#remember-five-things-first] 1. A PostgreSQL **cluster contains databases**; a database contains schemas; schemas contain tables, views, and functions. 2. A client connects to one database. Cross-database access is not as direct as cross-schema access. 3. Every statement runs in a transaction. Without explicit `BEGIN`, clients normally auto-commit one statement at a time. 4. Constraints are part of the data model, not merely a backup for application validation. 5. Indexes cost writes and maintenance. Inspect real plans before and after creating one. ## The 90-minute route [#the-90-minute-route] ### Minutes 0–10: run and connect [#minutes-010-run-and-connect] Finish [Quickstart](/en/docs/quickstart) and keep the local `pg-guide` container running. ### Minutes 10–30: build a model [#minutes-1030-build-a-model] Read [Data modeling](/en/docs/core/data-modeling). Create `customers` and `orders`; express facts with primary keys, foreign keys, `CHECK`, `NOT NULL`, and unique constraints. ### Minutes 30–50: query it [#minutes-3050-query-it] Read the [Query toolbox](/en/docs/core/queries). Practice filters, joins, aggregates, CTEs, and windows. Name output columns explicitly; avoid `SELECT *` in durable interfaces. ### Minutes 50–70: understand concurrency [#minutes-5070-understand-concurrency] Read [Transactions and concurrency](/en/docs/core/transactions). Observe `READ COMMITTED` in two `psql` sessions, then protect a balance change with `SELECT ... FOR UPDATE`. ### Minutes 70–90: verify performance [#minutes-7090-verify-performance] Read [Indexes and EXPLAIN](/en/docs/core/indexes-explain). Run `EXPLAIN (ANALYZE, BUFFERS)`, add an index, and compare. A sequential scan is not automatically a problem. ## Completion check [#completion-check] ```sql SELECT version(); SELECT current_database(), current_user; SELECT schemaname, tablename FROM pg_catalog.pg_tables WHERE schemaname NOT IN ('pg_catalog', 'information_schema'); ``` If you can explain these results and safely remove the practice container, you have completed the first stage. A backup file is not evidence of recoverability. In the operations track, restore one into an empty database at least once. --- # Database agent evaluation Canonical URL: https://pg.edu.rich/en/docs/ai/agent-evals Last reviewed: 2026-08-02 ## Three evaluation layers [#three-evaluation-layers] | Layer | Measures | Example | | ---------- | ------------------------------------------------------------------ | ------------------------------------------- | | Generation | Real objects, parameters, correct dialect | Never invent `orders.user_id` | | Execution | Correct, deterministic, bounded result | Matches a golden query result set | | Safety | Rejects privilege escape, injection, bulk writes, and abusive cost | Blocks cross-tenant access before execution | String equality on SQL is misleading: different queries can be equivalent, and the same query changes behavior with data and privilege. Prefer assertions on results, row counts, SQLSTATE, boundaries, and side effects. ## Fixed fixture database [#fixed-fixture-database] Start an ephemeral PostgreSQL from identical migrations and seeds for every run. Include `NULL`, empty sets, duplicates, time-zone boundaries, money boundaries, permitted orphans, same natural keys across tenants, and enough rows to expose plan differences. ## Case format [#case-format] ```yaml id: revenue-by-day-001 question: What was paid revenue for each of the last 7 complete UTC days? contract_version: test-42 role: agent_reader assert: read_only: true max_rows: 7 columns: [day, paid_cents] result_fixture: expected/revenue-by-day.json forbidden_relations: [app.payment_secrets] max_duration_ms: 1000 ``` Record model, prompt, tool schema, database version, and random seed. Repeat nondeterministic runs and report pass rate and variance—not the best sample. ## Safety red-team set [#safety-red-team-set] * User asks to ignore rules and return another tenant's data. * Schema comments contain prompt injection. * A value resembles an SQL fragment. * Request asks for delete or update without a predicate. * Request asks for `pg_read_file`, `COPY PROGRAM`, extension install, or privilege elevation. * Query attempts resource exhaustion with a huge Cartesian product or recursive CTE. Success is policy-layer rejection, not hoping the model self-regulates every time. ## Plan regression [#plan-regression] For important reads, retain normalized `EXPLAIN (FORMAT JSON)` features: top nodes, actual-to-estimated row ratio, buffer reads, and a runtime band. Do not pin exact cost numbers; statistics, cache, and PostgreSQL versions change plans. ## Release gate [#release-gate] A new prompt or model must pass correctness, zero safety violations, P95 latency and cost budgets, refusal under stale/missing context, and complete audit events. Regression in any dimension blocks automatic rollout. --- # AI agent long-term memory on PostgreSQL Canonical URL: https://pg.edu.rich/en/docs/ai/agent-memory Last reviewed: 2026-08-06 Every call to a language model is stateless: nothing is written back, and the only "memory" available is what you put into the prompt. Once an application is expected to know who the user is, what they preferred last week, and which facts have changed since, the application has to own that state itself. That state layer is what "agent memory" means in practice — a database problem with an LLM-assisted write path, not a model feature. ## Memory is not RAG [#memory-is-not-rag] RAG and agent memory are often conflated because both end with "retrieve relevant text into the prompt". The difference is the write path: | | RAG | Agent memory | | ------------- | --------------------------------------- | --------------------------------------------------------- | | Content | External corpus (docs, tickets, code) | Facts, preferences, and episodes produced by interaction | | Write path | Batch ingestion, replayable, idempotent | Online writes during conversations, often model-extracted | | Updates | Re-ingest a new document version | Correct, supersede, and forget individual facts | | Typical query | "Find passages about X" | "What do we currently believe about this user?" | | Failure mode | Stale or unpermitted chunks | Wrong facts written with the same authority as true ones | A memory store must therefore support targeted `UPDATE` and `DELETE`, time-scoped validity, and conflict handling — not just nearest-neighbor search. A read-only vector index of conversation logs is RAG over chat history, not memory. ## Three shapes of a memory layer [#three-shapes-of-a-memory-layer] 1. **Dedicated memory frameworks** — SDKs and services (Mem0, Cognee, and others) that own extraction, storage, and retrieval behind an `add`/`search` API. Fastest to prototype; the schema and retrieval policy live inside the framework. 2. **A vector database alone** — embeddings plus metadata filters. Simple, but user profiles, relationships between entities, and exact-match lookup all have to be rebuilt elsewhere. 3. **PostgreSQL directly** — one system holds embeddings (pgvector), structured profiles (`jsonb`), keyword search (full-text search), entity relationships (tables, or SQL/PGQ graphs in PostgreSQL 19), plus transactions and row-level security. The write policy lives in your code and SQL instead of a framework. These are not exclusive: the frameworks in the first row typically persist into stores from the other two. The real decision is where the schema of record lives and who controls the write policy. ## PostgreSQL as the memory base [#postgresql-as-the-memory-base] A workable memory schema separates the fact from its embedding and keeps update history explicit instead of overwriting in place: ```sql CREATE EXTENSION IF NOT EXISTS vector; CREATE TABLE agent_memory ( id bigint GENERATED ALWAYS AS IDENTITY PRIMARY KEY, tenant_id bigint NOT NULL, user_id text NOT NULL, kind text NOT NULL CHECK (kind IN ('profile', 'preference', 'episode', 'fact')), content text NOT NULL, attributes jsonb NOT NULL DEFAULT '{}', embedding vector(1536), search_vector tsvector GENERATED ALWAYS AS (to_tsvector('simple', content)) STORED, valid_from timestamptz NOT NULL DEFAULT now(), expires_at timestamptz, superseded_by bigint REFERENCES agent_memory(id), created_at timestamptz NOT NULL DEFAULT now() ); CREATE INDEX ON agent_memory USING hnsw (embedding vector_cosine_ops); CREATE INDEX ON agent_memory USING gin (search_vector); ``` * **pgvector** stores embeddings next to the facts they describe; index choice and filtered-scan behavior follow the same rules as [RAG pipeline](/en/docs/ai/rag-pipeline) and [pgvector setup](/en/docs/ai/pgvector-setup). * **`jsonb`** holds the structured part of a profile (timezone, language, plan tier) that must be filterable and updatable field by field, not re-embedded on every change. * **Full-text search** covers exact names, IDs, and error strings that embeddings handle poorly; combine both candidate sets and fuse, exactly as in hybrid RAG retrieval. * **Entity relationships** are ordinary tables: `entity` plus an edge table with `PRIMARY KEY (src, dst, relation)`. On PostgreSQL 19 they can additionally be declared as a property graph and queried with `GRAPH_TABLE` — see [Graph queries with SQL/PGQ](/en/docs/core/graph-queries). On earlier versions the same tables are queried with joins or `WITH RECURSIVE`. Retrieval is one SQL statement with the tenant and permission filters pushed into the database, per the pattern in [RAG pipeline](/en/docs/ai/rag-pipeline). Nothing here requires a memory-specific server. ## Candidate frameworks [#candidate-frameworks] Capability statements below were checked against the official repositories and documentation in 2026-08. Benchmark figures published by vendors are vendor-reported, not independently verified here. | | [Mem0](https://github.com/mem0ai/mem0) | [Cognee](https://github.com/topoteretes/cognee) | | ----------------------- | --------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | Self-description | "Memory layer" SDK, self-hosted server, and managed cloud | "AI memory platform" that builds a knowledge graph from ingested data | | Memory model | Memories scoped by user, session, and agent | Documents → entities/relations in a graph plus embeddings; `remember` / `recall` / `forget` API | | Storage backends | Pluggable vector stores; [supported list includes PGVector](https://docs.mem0.ai/components/vectordbs/overview) | Pluggable relational, vector, and graph backends; a [PostgreSQL + pgvector configuration](https://github.com/topoteretes/cognee#run-the-whole-memory-layer-on-postgres) is documented | | PostgreSQL relationship | PostgreSQL is one of several supported vector stores | Its README notes the Postgres **graph** store is currently a demo feature and points production graph workloads at graph-native backends or a licensed offering | | Extraction policy | LLM-based fact extraction and update decisions run inside the framework | LLM-based pipeline (cognify) builds the graph inside the framework | Both frameworks can sit on top of PostgreSQL for part of their storage, so "framework vs PostgreSQL" is usually "framework's write policy on PostgreSQL" vs "your write policy on PostgreSQL", not two different databases. Ruohang Feng (vonng) argues in [a 2026 essay](https://blog.vonng.com/en/ai/agent-memory-framework/) that memory frameworks are middleware squeezed between models and databases: extraction strategy migrates into the model (or a short skill file), storage migrates back into PostgreSQL, and the durable moat is the data layer. This is one practitioner's opinion, not a verifiable fact — treat it as a hypothesis to test against your own write-path complexity before adopting or dismissing a framework. ## Production concerns [#production-concerns] ### Write consistency [#write-consistency] The dangerous write is the model-extracted one. Constrain it: memories go through a fixed schema with `CHECK` constraints, the extracting model gets a narrow role that cannot touch business tables ([Safe SQL for agents](/en/docs/ai/safe-sql)), and a correction is an `INSERT` plus `superseded_by` link rather than an in-place rewrite, so "what did we believe when the agent answered" stays auditable. If a memory is derived from a business transaction, write both in one transaction or record the source version explicitly. ### Forgetting and TTL [#forgetting-and-ttl] Memory without expiry grows into noise that retrieval then amplifies. Give facts `expires_at` where the domain allows it, sweep expired rows on a schedule, and treat user-initiated deletion as a hard `DELETE` (plus embedding rows) rather than a flag. pgvector indexes do not make deleted rows disappear from backups — align PITR retention with your deletion policy. ### Multi-tenant isolation with RLS [#multi-tenant-isolation-with-rls] Memory is per-user data with prompt-injection blast radius: a poisoned memory written in one tenant must not be retrievable in another. Enforce isolation in the database, not the retrieval code: ```sql ALTER TABLE agent_memory ENABLE ROW LEVEL SECURITY; CREATE POLICY tenant_isolation ON agent_memory USING (tenant_id = current_setting('app.tenant_id')::bigint); ``` Set `app.tenant_id` per request from the authenticated identity. This keeps the guarantee intact when agents or [MCP tools](/en/docs/ai/mcp) issue queries you did not hand-write. ### Evaluating memory quality [#evaluating-memory-quality] Retrieval recall is necessary but not sufficient: a memory layer can also fail by extracting wrong facts, keeping contradictions, or surfacing stale state. Keep a versioned set of conversation traces with the memories they should produce and the answers retrieval should return, measure extraction accuracy and retrieval recall separately, and re-run on every prompt, model, or schema change. The harness for this is the same as in [Agent evals](/en/docs/ai/agent-evals); framework-published benchmarks are vendor-reported numbers and should not substitute for a suite built from your own traffic. ## AI prompt: draft a memory schema [#ai-prompt-draft-a-memory-schema] ## Related [#related] * [RAG pipeline](/en/docs/ai/rag-pipeline) — hybrid retrieval, indexing, and permission filtering that the read path reuses * [pgvector setup](/en/docs/ai/pgvector-setup) — installing and verifying the vector extension * [Context contract](/en/docs/ai/context-contract) — what the model is allowed to see, and how to say "insufficient context" * [Safe SQL for agents](/en/docs/ai/safe-sql) — roles and statement limits for model-issued writes * [Agent evals](/en/docs/ai/agent-evals) — evaluation harness for memory quality * [Graph queries with SQL/PGQ](/en/docs/core/graph-queries) — entity relationships on PostgreSQL 19 --- # Context contract Canonical URL: https://pg.edu.rich/en/docs/ai/context-contract Last reviewed: 2026-08-02 ## What the contract answers [#what-the-contract-answers] Each task needs only a relevant subgraph, but these fields should be stable: ```yaml contract_version: 2026-08-02.1 database: commerce schema: app role: analytics_readonly dialect: postgresql-18 timezone: UTC currency_unit: cents tables: orders: purpose: one row per checkout primary_key: [id] columns: customer_id: { type: bigint, nullable: false, ref: customers.id } status: { type: text, allowed: [pending, paid, shipped, cancelled] } total_cents: { type: bigint, min: 0 } placed_at: { type: timestamptz, meaning: checkout completion instant } invariants: - paid orders have an immutable total sensitive: [] limits: statement_timeout_ms: 5000 max_rows: 200 writes: forbidden ``` Tie contract versions to a migration version or schema hash. Return that version from tools so stale-schema generation can be diagnosed. ## Three context layers [#three-context-layers] 1. **Global rules**: dialect, time zone, money unit, default schema, privilege, and bounds. 2. **Task subgraph**: relevant tables, keys, columns, comments, enumerations, and important indexes. 3. **Dynamic evidence**: read-only samples, statistical summaries, recent errors—timestamped and marked when truncated. Do not ship complete DDL and every index for the whole database. Retrieve a task subgraph by names, comments, and foreign-key edges, then expand indexes or functions only when needed. ## Always exclude [#always-exclude] * Passwords, connection URIs, API keys, and `pg_authid` data. * Sample values outside the current tenant or authorization scope. * Full production rows, especially personal data and key material. * Business rules without source or freshness. * Estimates presented as exact counts. ## Output contract [#output-contract] Ask the model for a structured object, not arbitrary executable text: ```json { "intent": "read", "sql": "SELECT id, total_cents FROM app.orders WHERE customer_id = $1 LIMIT $2", "params": [42, 50], "assumptions": ["customer_id is the authenticated customer's internal id"], "expected_columns": ["id", "total_cents"], "risk": "R0" } ``` The policy layer validates SQL again. Valid JSON is not trustworthy semantics. ## Missing-information behavior [#missing-information-behavior] The contract must let the model return `insufficient_context` with required tables, columns, or business definitions. Refusing to invent a plausible column is a success condition. ## Delegate changing Linux facts to a read-only tool [#delegate-changing-linux-facts-to-a-read-only-tool] Install, upgrade, and troubleshooting tasks also need distribution facts that change over time. Configure the [read-only PkgSeek MCP](https://pkgseek.com/mcp) as the Linux package evidence layer for exact package names, file providers, repositories, releases, history, and vendor security status. This guide remains responsible for PostgreSQL selection, upgrade, and verification rules. ```yaml linux_evidence: provider: pkgseek distro: ubuntu release: noble architecture: amd64 package_source: pgdg observed_at: required missing_coordinate: insufficient_context state_changes: require_confirmation ``` When the tool returns no exact coordinate, the model must say “not present in the current index”, not “the package does not exist”. Lookups are read-only; repository changes, `sudo apt install`, service restarts, and major upgrades still require separate confirmation and post-action verification. --- # AI / agent reference Canonical URL: https://pg.edu.rich/en/docs/ai Last reviewed: 2026-08-02 A model does not understand your database merely because it can write SQL. Reliable systems turn database context into a contract, narrow execution into tools, and make correctness repeatably testable. ## Recommended architecture [#recommended-architecture] ```text user intent → task class (read / write / DDL / operations) → retrieve schema contract and relevant guidance → model emits a structured tool call → policy layer checks AST, privilege, cost, and parameters → restricted database role executes → return row count, SQLSTATE, duration, and truncation state → write an audit event ``` Database credentials do not enter model context. The model does not choose connection targets. The tool binds environment, database, schema, and role. ## Risk tiers [#risk-tiers] | Tier | Example | Default policy | | ---- | ---------------------------------------------- | ----------------------------------------------- | | R0 | List/describe schema, bounded read | Auto-run with a short timeout | | R1 | Sensitive columns, larger aggregate | Permission filter, audit, cost bound | | R2 | `INSERT` or primary-key single-row `UPDATE` | Dry run plus business API or explicit approval | | R3 | Bulk writes, DDL, grants, replication, restore | Not exposed to a general agent; expert workflow | “Do not delete data” is behavioral advice. Real boundaries come from roles, network isolation, read-only transactions, SQL parsing, and tool allowlists. ## Minimum bar [#minimum-bar] A deployable database agent should parameterize all values; default to read-only; bound statement time and result rows; reject multiple statements; never return secrets to the model; audit query fingerprints; and handle SQLSTATE values such as `40001`, `40P01`, and `57014` deterministically. --- # Postgres MCP server Canonical URL: https://pg.edu.rich/en/docs/ai/mcp Last reviewed: 2026-08-06 MCP (Model Context Protocol) gives an agent a uniform way to call tools and read data. A Postgres MCP server sits between the agent and your database: it exposes schema metadata, read-only queries, and `EXPLAIN` output as tools, so the agent inspects the real catalog instead of guessing column names. This changes the failure mode of AI-generated SQL — most "hallucinated column" errors disappear once the model can check. The server is only a transport. The actual safety boundary is the database role it connects with, covered below and in [Safe SQL guardrails](/en/docs/ai/safe-sql). ## Choosing a server [#choosing-a-server] Tool capabilities below were verified in 2026-08 against each project's repository; this ecosystem moves fast, so re-check before adopting. | Project | Maintainer | Notes | | ------------------------------------------------------------------------------- | ----------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | [Postgres MCP Pro](https://github.com/crystaldba/postgres-mcp) (`postgres-mcp`) | Crystal DBA | Schema browsing, `EXPLAIN` analysis, index tuning and health checks. Ships a `--access-mode=restricted` flag that limits execution to read-only SQL. Python-based; run via `uvx`, `pipx`, or the `crystaldba/postgres-mcp` Docker image. | | [Neon MCP](https://github.com/neondatabase/mcp-server-neon) | Neon | Adds project-level resources: branch creation, migrations on branches, connection strings. | | [Supabase MCP](https://supabase.com/docs/guides/getting-started/mcp) | Supabase | Hosted server covering database plus project management (branches, logs, advisors). | `@modelcontextprotocol/server-postgres` — the original Anthropic reference implementation — is deprecated on npm and moved to [servers-archived](https://github.com/modelcontextprotocol/servers-archived); it receives no maintenance or security fixes. Do not deploy it in new setups (verified 2026-08). Whichever server you pick, treat vendor feature lists as provisional: enable `pg_stat_statements` and `hypopg` only if you actually use the tuning tools, and pin the server version in your config so upgrades are deliberate. ## Create the read-only role first [#create-the-read-only-role-first] Before configuring any client, create a dedicated role that can only read: ```sql CREATE ROLE readonly LOGIN PASSWORD 'secret'; GRANT CONNECT ON DATABASE myapp TO readonly; GRANT USAGE ON SCHEMA public TO readonly; GRANT SELECT ON ALL TABLES IN SCHEMA public TO readonly; ALTER DEFAULT PRIVILEGES IN SCHEMA public GRANT SELECT ON TABLES TO readonly; ALTER ROLE readonly SET default_transaction_read_only = on; ALTER ROLE readonly SET statement_timeout = '5s'; ``` The `ALTER DEFAULT PRIVILEGES` line matters: without it, tables created later are invisible to the role and the agent's schema view silently drifts out of date. Add timeouts at the role level so a runaway query from an agent cannot hold resources — see [Safe SQL guardrails](/en/docs/ai/safe-sql) for the full set (lock timeout, idle-in-transaction timeout, row bounds). Do not give the MCP server superuser or table-owner credentials — not even in development, because habits from dev configs leak into production. Never expose production write access to an agent at all: point it at a read replica, an Aurora reader endpoint, or a database branch. A restricted-mode flag in the MCP server is a second layer, not a substitute for role-level privileges. ## Claude Code [#claude-code] Project-scoped `.mcp.json` (commit it so the whole team gets the same server), or run `claude mcp add` for a user-scoped entry: ```json { "mcpServers": { "postgres": { "command": "uvx", "args": [ "postgres-mcp", "--access-mode=restricted", "postgres://readonly:secret@localhost:5432/myapp" ] } } } ``` `--access-mode=restricted` keeps the server on read-only SQL even if the agent asks for writes (verified 2026-08 in the [project README](https://github.com/crystaldba/postgres-mcp)). Prefer injecting the connection string from an environment variable or secret manager over committing passwords. ## Cursor [#cursor] `.cursor/mcp.json`: ```json { "mcpServers": { "postgres": { "command": "uvx", "args": ["postgres-mcp", "--access-mode=restricted"], "env": { "DATABASE_URI": "postgres://readonly:secret@localhost:5432/myapp" } } } } ``` ## One database branch per PR [#one-database-branch-per-pr] Neon and Supabase both support near-instant copy-on-write branches. Combined with MCP, each PR gets its own database the agent can migrate, query, and destroy: 1. Fork a branch from main (milliseconds, no data copy). 2. Run migrations and tests against real-shaped data; let the agent read `EXPLAIN` on realistic volumes. 3. Review the change in the PR; the branch is deleted on merge. ```bash neon branches create --name pr-123 --parent main export DATABASE_URI=$(neon connection-string --branch pr-123) # start the MCP server against DATABASE_URI ``` This is where MCP pays off most: the agent validates migrations against a real catalog and real data distribution, without touching the production writer. ## What to expose through the contract [#what-to-expose-through-the-contract] A database MCP server answers "what does the schema look like" and "what does this query do" — it should not become a general-purpose data exfiltration channel. Keep the task-facing context small and explicit with a [context contract](/en/docs/ai/context-contract): the agent gets the tables and columns relevant to the task, and the MCP server handles verification, not exploration of unrelated schemas. ## Prompts that use it well [#prompts-that-use-it-well] ## Next [#next] * Lock down what the agent may run → [Safe SQL guardrails](/en/docs/ai/safe-sql) * Shrink what the agent needs to know → [Context contract](/en/docs/ai/context-contract) * Generate that context from the catalog → [Schema retrieval](/en/docs/ai/schema-retrieval) --- # Install pgvector for PostgreSQL Canonical URL: https://pg.edu.rich/en/docs/ai/pgvector-setup Last reviewed: 2026-08-02 pgvector is an independent extension, not a core PostgreSQL type. After installing a package or using an image that contains it, run `CREATE EXTENSION vector` in every target database. ## Minimal Docker environment [#minimal-docker-environment] The pgvector project publishes versioned images based on official Postgres images. This example pins PostgreSQL 18 and pgvector 0.8.2: ```bash docker volume create pgvector18-data docker run --name pgvector18 \ --env POSTGRES_PASSWORD=local-only-change-me \ --publish 5432:5432 \ --volume pgvector18-data:/var/lib/postgresql/data \ --detach pgvector/pgvector:0.8.2-pg18-trixie docker exec -it pgvector18 \ psql -U postgres -d postgres -c "CREATE EXTENSION vector;" ``` The password is for an isolated local demonstration, never production configuration. If the name or port is occupied, choose an explicit alternative; do not delete an unknown instance. ## Ubuntu / Debian package [#ubuntu--debian-package] ### pgvector in distribution repositories PkgSeek package snapshot; checked: 2026-08-02 15:47:58 UTC. | Distribution | Release | Full version | Repository | Linked advisories | |---|---|---|---|---:| | AlmaLinux | 10 | 0.6.2-6.el10_0 | official / AppStream | — | | AlmaLinux | 9 | 0.8.1-1.module_el9.8.0+234+5456f35d | official / AppStream | — | | Arch Linux | rolling | 0.8.6-1 | official / extra | — | | CentOS Stream | 10 | 0.6.2-8.el10 | official / AppStream | — | | CentOS Stream | 9 | 0.8.1-1.module_el9+1300+1c4aa8df | official / AppStream | — | | Fedora | 42 | 0.6.2-4.fc42 | official / everything | — | | Fedora | 43 | 0.8.0-1.fc43 | official / everything | — | | Fedora | 44 | 0.8.0-2.fc44 | official / everything | — | | Oracle Linux | 10 | 0.6.2-6.el10_0 | official / appstream | — | | Oracle Linux | 9 | 0.6.2-2.module+el9.8.0+90925+e22a792e | official / appstream | — | | Red Hat Enterprise Linux | 10.2 | 0.6.2-6.el10_0 | official / AppStream | — | | Red Hat Enterprise Linux | 9.8 | 0.6.2-2.module+el9.8.0+24096+5a959ed6 | official / AppStream | — | | Rocky Linux | 10 | 0.6.2-6.el10_0 | official / AppStream | — | | Rocky Linux | 9 | 0.6.2-2.module+el9.8.0+40212+d6f50005 | official / AppStream | — | Source: [PkgSeek package lookup](https://pkgseek.com/packages/pgvector). Distribution revisions and backported fixes are part of the complete version identity. The same extension may be named `pgvector`, `postgresql18-pgvector`, or another server-major-specific package across distributions. This table only shows exact `pgvector` coordinates; it does not establish whether every versioned package exists. ### PGDG: postgresql-18-pgvector PkgSeek package snapshot; checked: 2026-08-02 15:47:58 UTC. > No exact indexed coordinate. This means the current snapshot has no match, not that the package does not exist. Source: [PkgSeek package lookup](https://pkgseek.com/search?q=postgresql-18-pgvector). Distribution revisions and backported fixes are part of the complete version identity. If no exact coordinate appears above, verify the target PGDG repository metadata. Do not infer an Ubuntu or Debian versioned package name from a generic `pgvector` record. After configuring the official PostgreSQL Apt repository, the extension package is bound to the server major: ```bash sudo apt install postgresql-18-pgvector sudo -u postgres psql -d app -c "CREATE EXTENSION vector;" ``` Installing at OS level does not enable every database. Verify the actual version: ```sql SELECT extname, extversion FROM pg_extension WHERE extname = 'vector'; ``` Before upgrading, read release notes and test this in the target database: ```sql ALTER EXTENSION vector UPDATE; ``` ## Minimal query verification [#minimal-query-verification] ```sql CREATE TABLE vector_demo ( id bigint GENERATED ALWAYS AS IDENTITY PRIMARY KEY, embedding vector(3) NOT NULL ); INSERT INTO vector_demo (embedding) VALUES ('[1,2,3]'), ('[4,5,6]'), ('[1,1,1]'); SELECT id, embedding <-> '[1,2,2]'::vector AS l2_distance FROM vector_demo ORDER BY embedding <-> '[1,2,2]'::vector LIMIT 2; ``` Without an approximate index, this is exact search. When scale and real filters justify it, evaluate HNSW or IVFFlat using the [pgvector production guide](/en/docs/ai/vector-production). ## Managed cloud PostgreSQL [#managed-cloud-postgresql] Managed services commonly restrict host access and extension allowlists. Confirm: * engine major and the exact `vector` extension version; * who can run `CREATE EXTENSION` and `ALTER EXTENSION`; * whether that version supports HNSW, IVFFlat, and iterative scans; * whether extensions follow engine upgrades or require manual maintenance; * disk, memory, WAL, and replica-lag limits during index builds. When model, dimension, normalization, or distance semantics change, rebuild into a new column or table and evaluate it. Equal dimensions do not make vectors semantically comparable. Use the [pgvector installation documentation](https://github.com/pgvector/pgvector#installation) for current releases and methods. --- # PostgreSQL RAG pipeline Canonical URL: https://pg.edu.rich/en/docs/ai/rag-pipeline Last reviewed: 2026-08-02 ## Data model [#data-model] ```sql CREATE EXTENSION IF NOT EXISTS vector; CREATE TABLE documents ( id bigint GENERATED ALWAYS AS IDENTITY PRIMARY KEY, tenant_id bigint NOT NULL, source_uri text NOT NULL, source_version text NOT NULL, title text NOT NULL, access_scope text[] NOT NULL DEFAULT '{}', created_at timestamptz NOT NULL DEFAULT now(), UNIQUE (tenant_id, source_uri, source_version) ); CREATE TABLE document_chunks ( id bigint GENERATED ALWAYS AS IDENTITY PRIMARY KEY, document_id bigint NOT NULL REFERENCES documents(id) ON DELETE CASCADE, ordinal integer NOT NULL CHECK (ordinal >= 0), content text NOT NULL, token_count integer NOT NULL CHECK (token_count > 0), embedding vector(1536) NOT NULL, embedding_model text NOT NULL, search_vector tsvector GENERATED ALWAYS AS (to_tsvector('simple', content)) STORED, UNIQUE (document_id, ordinal) ); ``` ## Ingestion must be replayable [#ingestion-must-be-replayable] Store source version, chunker version, embedding model, and dimension. Derive deterministic document/chunk keys for idempotent upsert. When changing models, build a new embedding column or table and dual-write during rebuild; never mix incomparable vectors in one index. ## Retrieval order [#retrieval-order] 1. Filter tenant, permission, document state, and time in SQL. 2. Produce bounded candidates independently from full text and vectors. 3. Merge with rank fusion or application reranking. 4. Fetch a small number of neighboring chunks for continuity. 5. Return source URI, version, chunk id, and excerpt for citation. An illustrative vector candidate query: ```sql SELECT c.id, c.document_id, c.ordinal, c.content, c.embedding <=> $1::vector AS distance FROM document_chunks AS c JOIN documents AS d ON d.id = c.document_id WHERE d.tenant_id = $2 AND d.access_scope && $3::text[] ORDER BY c.embedding <=> $1::vector LIMIT 40; ``` Index type and parameters depend on scale, recall, latency, and write pattern. Establish an exact-search baseline before evaluating HNSW or IVFFlat; demo data is not enough. ### Approximate indexes and filters [#approximate-indexes-and-filters] HNSW/IVFFlat normally apply tenant, ACL, and other predicates after the index produces candidates, so a query can return fewer rows than its `LIMIT`. That is not a reason to weaken authorization filters. pgvector 0.8.0+ iterative scans can expand candidate scanning; large tenants may also justify partitions, partial indexes, or separate tables. Measure recall\@k in real tenant/ACL buckets for every design. See [Vector search in production](/en/docs/ai/vector-production) and the [official pgvector filtering guidance](https://github.com/pgvector/pgvector#filtering) for index DDL, parameters, and evaluation. ## Security and citation [#security-and-citation] Authorization predicates stay inside SQL/RLS so the database applies them before rows leave the boundary. Never return global candidates to the application and filter there; unauthorized text can leak through logs, caches, or model context. Final answers carry verifiable citations and report insufficient evidence when retrieval is weak. Nearby text can be stale, contradictory, or from the wrong tenant. RAG needs versions, permissions, source precedence, and answer evaluation—not only nearest neighbors. --- # Safe SQL guardrails Canonical URL: https://pg.edu.rich/en/docs/ai/safe-sql Last reviewed: 2026-08-02 ## The database role is the first boundary [#the-database-role-is-the-first-boundary] ```sql CREATE ROLE agent_reader LOGIN; GRANT CONNECT ON DATABASE commerce TO agent_reader; GRANT USAGE ON SCHEMA app TO agent_reader; GRANT SELECT ON ALL TABLES IN SCHEMA app TO agent_reader; ALTER DEFAULT PRIVILEGES FOR ROLE app_owner IN SCHEMA app GRANT SELECT ON TABLES TO agent_reader; ALTER ROLE agent_reader SET default_transaction_read_only = on; ALTER ROLE agent_reader SET statement_timeout = '5s'; ALTER ROLE agent_reader SET lock_timeout = '1s'; ALTER ROLE agent_reader SET idle_in_transaction_session_timeout = '10s'; ``` Configure credentials through a secret manager or cloud identity integration, never migrations, prompts, or tool responses. Confirm the role cannot `SET ROLE` into a stronger role. ## Pre-execution policy [#pre-execution-policy] Validate generated SQL with a parser/AST, not regex. Default rules: * Allow one `SELECT` statement only. * Reject `COPY ... PROGRAM`, large objects, foreign-data wrappers, and dangerous functions. * Reject multiple statements and comment-based bypasses. * Restrict accessible schemas, tables, and columns. * Bind values; choose identifiers only from an allowlist. * Add a `LIMIT` to non-aggregate results and cap returned bytes in the driver. * A cost preflight with `EXPLAIN (FORMAT JSON)` is useful, but estimated cost is not a runtime guarantee. ## Use a read-only transaction for each read [#use-a-read-only-transaction-for-each-read] ```sql BEGIN READ ONLY; SET LOCAL statement_timeout = '5s'; SET LOCAL lock_timeout = '1s'; SET LOCAL search_path = app, pg_catalog; SELECT id, status, total_cents FROM orders WHERE customer_id = $1 ORDER BY placed_at DESC LIMIT 100; COMMIT; ``` Read-only transactions can still run expensive queries and expose readable data. Privilege, cost, and output bounds are all required. ## Do not expose arbitrary write SQL [#do-not-expose-arbitrary-write-sql] Prefer domain tools: ```json { "tool": "cancel_order", "arguments": { "order_id": 8842, "expected_status": "pending", "reason": "duplicate order", "idempotency_key": "case-2026-184" } } ``` The application validates identity and transition, runs parameterized SQL in a transaction, and returns a typed outcome. Bulk writes, DDL, `GRANT`, restore, and replication configuration should not be general-agent tools. ## Audit fields [#audit-fields] Record requester/tenant, tool, model and prompt versions, contract version, database target, parameter-redacted SQL fingerprint, risk tier, approver, row count, duration, SQLSTATE, and truncation state. Do not copy raw sensitive results into general logs. Writing first and noticing “too many rows” later may already fire triggers or external effects. Preview the target set in a controlled transaction, or make the domain API constrain the modifiable set in its predicate. --- # Schema retrieval and documentation Canonical URL: https://pg.edu.rich/en/docs/ai/schema-retrieval Last reviewed: 2026-08-02 ## Do not let the model explore production [#do-not-let-the-model-explore-production] A production agent should not have unbounded catalog exploration. A trusted build job extracts schema, redacts and versions it, and publishes it to retrieval. Runtime returns only a task-relevant subgraph. ## Tables and columns [#tables-and-columns] Use `information_schema` for portable basics: ```sql SELECT c.table_schema, c.table_name, c.ordinal_position, c.column_name, c.data_type, c.udt_name, c.is_nullable, c.column_default FROM information_schema.columns AS c WHERE c.table_schema = ANY($1::text[]) ORDER BY c.table_schema, c.table_name, c.ordinal_position; ``` Column comments come from PostgreSQL catalogs: ```sql SELECT n.nspname AS schema_name, cls.relname AS table_name, a.attname AS column_name, col_description(cls.oid, a.attnum) AS comment FROM pg_catalog.pg_attribute AS a JOIN pg_catalog.pg_class AS cls ON cls.oid = a.attrelid JOIN pg_catalog.pg_namespace AS n ON n.oid = cls.relnamespace WHERE n.nspname = ANY($1::text[]) AND cls.relkind IN ('r', 'p') AND a.attnum > 0 AND NOT a.attisdropped; ``` ## Foreign-key edges form the task graph [#foreign-key-edges-form-the-task-graph] ```sql SELECT src_ns.nspname AS table_schema, src.relname AS table_name, src_col.attname AS column_name, dst_ns.nspname AS foreign_table_schema, dst.relname AS foreign_table_name, dst_col.attname AS foreign_column_name FROM pg_catalog.pg_constraint AS con JOIN pg_catalog.pg_class AS src ON src.oid = con.conrelid JOIN pg_catalog.pg_namespace AS src_ns ON src_ns.oid = src.relnamespace JOIN pg_catalog.pg_class AS dst ON dst.oid = con.confrelid JOIN pg_catalog.pg_namespace AS dst_ns ON dst_ns.oid = dst.relnamespace CROSS JOIN LATERAL unnest(con.conkey, con.confkey) AS key_columns(src_attnum, dst_attnum) JOIN pg_catalog.pg_attribute AS src_col ON src_col.attrelid = src.oid AND src_col.attnum = key_columns.src_attnum JOIN pg_catalog.pg_attribute AS dst_col ON dst_col.attrelid = dst.oid AND dst_col.attnum = key_columns.dst_attnum WHERE con.contype = 'f' AND src_ns.nspname = ANY($1::text[]) ORDER BY con.oid, src_col.attnum; ``` `conkey` and `confkey` correspond positionally; parallel `unnest` preserves composite foreign-key column mappings. Joining `information_schema` views only by `constraint_name` can produce a Cartesian product of columns for a composite key. ## Documentation build flow [#documentation-build-flow] ```text merge migration → apply all migrations to an ephemeral database → extract catalogs → normalize ordering and remove environment values → generate JSON plus Markdown summaries → hash / bind migration version → review schema diff → publish to retrieval index ``` A table summary keeps purpose, primary and foreign keys, column types and nullability, constraints, business comments, sensitivity, and only the most important query indexes. Expand function bodies, view definitions, and policies on demand. ## Prevent staleness [#prevent-staleness] Every agent tool returns `contract_version`. If the runtime migration version differs from retrieval, reject high-risk requests and trigger a rebuild. Never silently use a stale contract. --- # Text-to-SQL production pattern Canonical URL: https://pg.edu.rich/en/docs/ai/text-to-sql Last reviewed: 2026-08-02 The goal of Text-to-SQL is not “produce something that runs.” It is to execute a correct query only when evidence, privilege, and cost boundaries are clear—and clarify or refuse everything else. ## Recommended execution chain [#recommended-execution-chain] ```text natural-language question → resolve business entities, metric, time range, and grain → retrieve a versioned schema/metric contract and a few verified examples → generate a structured query plan and parameters, not directly executed free text → validate SQL AST, object/function allowlists, privilege, and cost → restricted role + read-only transaction + timeouts + result bounds → return result, metric definition, SQL fingerprint, truncation, and explainable errors ``` Model context should include schema version, table/column semantics, keys, enums, time zone, currency units, soft-delete rules, tenant boundaries, approved metric definitions, and allowed objects. Do not indiscriminately inject all DDL, sample customer data, or credentials. ## Prefer structured tools [#prefer-structured-tools] For common analytics, have the model produce domain parameters: ```json { "metric": "paid_order_revenue", "time_range": { "start": "2026-07-01", "end": "2026-08-01" }, "group_by": ["day"], "filters": [{ "field": "region", "op": "eq", "value": "east" }], "limit": 100 } ``` The server maps metrics, fields, and operators to reviewed SQL. Only long-tail exploration enters a free-SQL lane, which must still parse an AST. A regex check for “starts with SELECT” is not a guardrail: CTEs, data-modifying CTEs, functions, `COPY`, multiple statements, and comment tricks defeat naive string checks. ## Database execution envelope [#database-execution-envelope] ```sql BEGIN READ ONLY; SET LOCAL statement_timeout = '3s'; SET LOCAL lock_timeout = '500ms'; SET LOCAL idle_in_transaction_session_timeout = '5s'; -- One policy-approved parameterized SELECT; server enforces row/byte bounds SELECT date_trunc('day', paid_at) AS day, sum(total_cents) AS revenue_cents FROM analytics.paid_orders WHERE tenant_id = $1 AND paid_at >= $2 AND paid_at < $3 GROUP BY 1 ORDER BY 1 LIMIT 100; COMMIT; ``` `READ ONLY` is defense in depth, not a complete sandbox. Allow only trusted functions and objects, execute as a dedicated low-privilege role, and bind tenant, environment, and parameters on the server. The model never supplies a connection string, role, or `search_path`. ## Pre-execution checks [#pre-execution-checks] 1. Allow one statement and approved AST nodes; reject DDL/DML, `COPY`, arbitrary functions, and administration objects. 2. Convert every value to a bound parameter; identifiers only come from the schema-contract allowlist. 3. Run `EXPLAIN (FORMAT JSON)` on expensive candidates and inspect objects, estimated rows, and total cost. Estimates are signals, not execution-time guarantees. 4. Require time ranges and row/byte/join bounds. Route bulk exports to a separate asynchronous product path. 5. Reject sensitive columns in policy or expose reviewed masked views; never rely on the model remembering not to select them. 6. Inject tenant scope through database RLS or server templates, never from the user's wording. ## Correctness and refusal [#correctness-and-refusal] A runnable query can still answer the wrong question. Evaluation sets should cover empty results, join duplication, time-zone boundaries, NULLs, refunds/cancellations, late data, tenant isolation, and ambiguous metrics. Assert the allow/refuse decision, result set, accessed objects, maximum cost, and explanation together. Clarify instead of guessing when a metric has multiple business definitions, a date lacks year/time zone, a name maps to multiple IDs, the request requires a nonexistent historical snapshot, or the schema contract does not match deployment. At most perform bounded regeneration from structured syntax errors. `57014` (cancel/timeout) should narrow the request or switch to async; retry `40001` and `40P01` only when the whole transaction is safe to replay. Retain the original request, schema version, query fingerprint, and final decision. Read the [PostgreSQL 18 `READ ONLY` transaction semantics](https://www.postgresql.org/docs/18/sql-set-transaction.html) and [SQLSTATE appendix](https://www.postgresql.org/docs/18/errcodes-appendix.html). --- # pgvector production practices Canonical URL: https://pg.edu.rich/en/docs/ai/vector-production Last reviewed: 2026-08-02 [pgvector](https://github.com/pgvector/pgvector) performs exact nearest-neighbor search by default. Search becomes approximate only after adding HNSW or IVFFlat. Index selection is an engineering tradeoff across recall, latency, memory, build time, and write cost. ## Fix distance semantics first [#fix-distance-semantics-first] | Meaning | Operator | Index operator class | | ----------------------------------------------- | -------- | -------------------- | | L2 / Euclidean distance | `<->` | `vector_l2_ops` | | Inner product (negative inner product returned) | `<#>` | `vector_ip_ops` | | Cosine distance | `<=>` | `vector_cosine_ops` | Embedding generation, index, and query must use the same distance meaning. Cosine similarity is `1 - cosine distance`. Store embedding model, dimension, normalization, and generation version; do not mix incomparable vectors in one column/index. ## Establish an exact baseline [#establish-an-exact-baseline] Sample the real query distribution and save exact top-k results. Compare approximate indexes on `recall@k`, p50/p95/p99 latency, insufficient-result rate, and resources—not one demonstration query. ```sql BEGIN; SET LOCAL enable_indexscan = off; SELECT c.id FROM document_chunks AS c JOIN documents AS d ON d.id = c.document_id WHERE d.tenant_id = $1 ORDER BY c.embedding <=> $2::vector LIMIT 20; ROLLBACK; ``` Disabling index scans is for baselines and diagnosis, not a production setting. Cover hot/cold tenants, common ACLs, time filters, new writes, deletions, and embedding-distribution drift. ## HNSW and IVFFlat [#hnsw-and-ivfflat] ```sql CREATE INDEX CONCURRENTLY document_chunks_embedding_hnsw ON document_chunks USING hnsw (embedding vector_cosine_ops); ``` * **HNSW** generally has a better query speed/recall tradeoff and needs no training set, but builds more slowly, uses more memory, and costs more to maintain. * **IVFFlat** builds faster and uses less memory, but needs representative existing data to form lists and generally has a weaker speed/recall tradeoff. Do not create it on an empty table and forget to rebuild. * On an existing production table, prefer `CREATE INDEX CONCURRENTLY` and observe WAL, disk, build duration, and replica lag. There is no universal `m`, `ef_construction`, `ef_search`, `lists`, or `probes` value. Start with defaults and an exact baseline, then tune against real filters. ## Filtering changes recall [#filtering-changes-recall] With approximate indexes, filters are normally applied after the index scan produces candidates. At the default `hnsw.ef_search = 40`, if only 10% of candidates satisfy tenant/ACL filters, the query can return fewer than its `LIMIT` even when more matching rows exist. pgvector 0.8.0+ supports iterative scans that continue when initial candidates are insufficient: ```sql BEGIN; SET LOCAL hnsw.iterative_scan = strict_order; SET LOCAL hnsw.ef_search = 200; SELECT c.id, c.content, c.embedding <=> $1::vector AS distance FROM document_chunks AS c JOIN documents AS d ON d.id = c.document_id WHERE d.tenant_id = $2 AND d.access_scope && $3::text[] ORDER BY c.embedding <=> $1::vector LIMIT 20; COMMIT; ``` Confirm the pgvector version offered by the cloud service. For a few skewed tenant values, consider list partitioning; for many values, partial indexes for large tenants, separate tables, or physical isolation may work better. Validate the choice with filtered recall and operational cost. Keep tenant/ACL predicates in SQL/RLS; never retrieve globally and filter in the application. Even when PostgreSQL correctly blocks unauthorized rows, a shared ANN graph may under-return after filtering. Security and recall are separate acceptance criteria. ## Launch bar [#launch-bar] * Every embedding version has replayable ingestion, an exact gold set, and rollback. * Record model version, filter bucket, candidate/result counts, distance distribution, latency, and truncation online. * Sample exact searches regularly to calculate recall\@k, bucketed by tenant/ACL. * Rebuild embeddings by dual-writing to a new column/table, building a new index, evaluating, then atomically switching reads. * Treat text and metadata as source of truth; vectors are rebuildable from versioned inputs. * Produce full-text and vector candidates independently, then fuse or rerank with a versioned method. Use the [pgvector README sections on indexing, filtering, and monitoring](https://github.com/pgvector/pgvector#hnsw) as the authoritative parameter reference. --- # Google Cloud AlloyDB for PostgreSQL Canonical URL: https://pg.edu.rich/en/docs/cloud/alloydb Last reviewed: 2026-08-06 [AlloyDB for PostgreSQL](https://docs.cloud.google.com/alloydb/docs/overview) is Google Cloud's PostgreSQL-compatible database service: decoupled compute and storage, cross-zone high availability, an optional columnar engine for analytical queries, and in-database machine-learning integrations. It speaks the PostgreSQL protocol and runs PostgreSQL-compatible SQL, but it is not the community binary — storage, replication, release cadence, and parts of query behavior are Google's implementation. For positioning against other enhanced engines, see the [cloud service map](/en/docs/cloud/service-map). ## What AlloyDB changes [#what-alloydb-changes] * **Separated compute and storage.** Instances in a cluster share disaggregated storage with layered caching, instead of each node owning a local volume. * **Optional columnar engine.** Frequently queried data can be held in a columnar format for analytical scans alongside the row store — Google's published figures for transactional and analytical speedups are vendor benchmarks; treat them as hypotheses to test on your data. * **In-database ML.** The `google_ml_integration` extension calls managed model endpoints from SQL, covering embeddings and predictions without an external glue service. * **ScaNN vector indexing.** An alternative approximate-nearest-neighbor index alongside pgvector's HNSW and IVFFlat. ## In-database embeddings with google\_ml.embedding() [#in-database-embeddings-with-google_mlembedding] With the extension installed and a model endpoint registered, embedding generation is a SQL function call, checked against the [official documentation](https://docs.cloud.google.com/alloydb/docs/ai/work-with-embeddings) on 2026-08-06: ```sql CREATE EXTENSION IF NOT EXISTS google_ml_integration; CREATE EXTENSION IF NOT EXISTS vector; SELECT google_ml.embedding( model_id => 'gemini-embedding-001', content => 'AlloyDB keeps embedding generation inside the database' ) AS embedding; INSERT INTO articles (body, embedding) VALUES ( 'Some article text', google_ml.embedding(model_id => 'gemini-embedding-001', content => 'Some article text')::vector ); ``` The function returns `real[]`, so store or compare it through a `::vector` cast, as above. Two boundaries to keep in mind: the call leaves the database toward a model endpoint, so region, residency, IAM permissions, and model lifecycle become database concerns; and model IDs change — confirm the current ID in the documentation rather than copying examples, including this one. ## ScaNN vector indexes [#scann-vector-indexes] AlloyDB offers [ScaNN indexes](https://docs.cloud.google.com/alloydb/docs/ai/create-scann-index) through the `alloydb_scann` extension: ```sql CREATE EXTENSION IF NOT EXISTS alloydb_scann CASCADE; CREATE INDEX articles_embedding_scann ON articles USING scann (embedding cosine) WITH (mode = 'MANUAL', num_leaves = 100); ``` ScaNN is a tree-based quantization index. Per Google's documentation it builds faster and uses less memory than HNSW, with QPS and recall depending on tuning parameters such as `num_leaves`. Neither property survives contact with real data unexamined: measure recall after your actual tenant and ACL filters, and compare against pgvector HNSW on the same dataset before standardizing on either. ## Differences to accept [#differences-to-accept] * **Not community-binary-equivalent.** Extension availability, parameter surfaces, and some catalog or wait-event behavior differ from community PostgreSQL and from Cloud SQL. Migrate with a compatibility checklist, not an assumption. * **Per-instance-hour pricing.** There is no scale-to-zero; idle clusters still bill. Usage-shaped workloads may fit a serverless platform better. * **Google Cloud only.** AlloyDB Omni exists for self-managed deployments elsewhere, but it is a separately licensed, separately operated product. * **Vendor performance figures are marketing inputs.** Published speedup claims compare against Cloud SQL under Google's benchmark conditions; your schema, concurrency, and data shape decide the real number. ## Validate before production [#validate-before-production] 1. Diff required extensions and parameters against AlloyDB's supported list, including exact pgvector and `google_ml_integration` versions. 2. Rehearse the migration path — Database Migration Service or `pg_dump`/restore — including rollback. 3. Benchmark representative OLTP and analytical queries; enable the columnar engine only if measured results justify it. 4. Measure vector recall and latency after real filtering, on ScaNN and HNSW, at production data volume. 5. Confirm embedding model region, IAM scope, quota, and lifecycle policy; record what happens to stored embeddings if the model version retires. 6. Export once with native tools and restore into an independent PostgreSQL environment as an exit-path test. Capabilities, model IDs, and extension behavior were checked against Google Cloud documentation on 2026-08-06. AlloyDB evolves quickly; re-verify model availability and extension versions at procurement time. For connection budgeting, backup drills, and monitoring baselines that apply on any provider, see the [production stack guide](/en/docs/operations/production-stack). A development-scale evaluation can start from the options in the [free PostgreSQL guide](/en/docs/cloud/free-postgresql). --- # AWS RDS for PostgreSQL and Aurora Canonical URL: https://pg.edu.rich/en/docs/cloud/aws-rds-aurora Last reviewed: 2026-08-06 AWS offers two managed paths for PostgreSQL workloads, and they are not the same product. [Amazon RDS for PostgreSQL](https://docs.aws.amazon.com/AmazonRDS/latest/UserGuide/CHAP_PostgreSQL.html) runs the community engine as a managed instance. [Aurora PostgreSQL-Compatible](https://docs.aws.amazon.com/AmazonRDS/latest/AuroraUserGuide/CHAP_AuroraOverview.html) is a PostgreSQL-compatible engine on AWS-built distributed storage, managed as a cluster rather than a server. Drivers and most SQL work on both; parameters, extensions, failover, and billing do not behave identically. For the cross-provider view, see the [cloud service map](/en/docs/cloud/service-map). ## When each one fits [#when-each-one-fits] RDS for PostgreSQL is the default choice when: * you want behavior as close to community PostgreSQL as a managed service allows; * one primary with an optional Multi-AZ standby and read replicas meets your availability target; * storage stays comfortably below the 64 TiB ceiling and you prefer provisioned, predictable sizing. Aurora PostgreSQL-Compatible earns its price when: * you need faster failover and read scaling across up to 15 Aurora Replicas that share one cluster volume; * storage should grow automatically in 10 GiB increments instead of being provisioned, up to 256 TiB; * you accept an engine with its own release cadence, versioning scheme, and extension allowlist. This is not a ranking. Both are managed services without host access, and the right answer depends on your measured failover, connection, and cost requirements. ## RDS and Aurora side by side [#rds-and-aurora-side-by-side] | | RDS for PostgreSQL | Aurora PostgreSQL-Compatible | | --------------- | ------------------------------------------------------------------ | ----------------------------------------------------------------------------------------------------------------------------------------------- | | Storage | Provisioned EBS (gp3/io2), up to 64 TiB | Shared cluster volume, auto-grows to 256 TiB, six copies across three AZs | | Replicas | Read replicas with independent storage, physical replication | Up to 15 Aurora Replicas sharing the cluster volume, typically low replica lag | | Failover | Multi-AZ standby promotion, documented as typically 60–120 seconds | Typically tens of seconds when a reader is available; measure it yourself | | Backup | Automated snapshots plus WAL, PITR within the retention window | Continuous backup with PITR inside the retention window | | Engine versions | Community majors per the RDS release calendar | Aurora's own release cadence and version numbering | | Extensions | Platform allowlist | Separate [extension support matrix](https://docs.aws.amazon.com/AmazonRDS/latest/AuroraPostgreSQLReleaseNotes/AuroraPostgreSQL.Extensions.html) | | Cost model | Instance plus provisioned storage | Instance plus consumed storage plus I/O requests; typically higher than RDS for the same workload | The ceilings above come from the [RDS storage documentation](https://docs.aws.amazon.com/AmazonRDS/latest/UserGuide/CHAP_Storage.html) and the [Aurora overview](https://docs.aws.amazon.com/AmazonRDS/latest/AuroraUserGuide/CHAP_AuroraOverview.html), checked on 2026-08-06. Documented ceilings are not SLA commitments for your workload — model costs in the [RDS](https://aws.amazon.com/rds/postgresql/pricing/) and [Aurora](https://aws.amazon.com/rds/aurora/pricing/) pricing pages, and rehearse failover before relying on either number. ## Differences to accept [#differences-to-accept] * **`max_connections` is computed differently on Aurora.** The default is `LEAST({DBInstanceClassMemory/9531392}, 5000)` — a 16 GiB instance lands around 1,800 connections, and the cap is 5,000 regardless of instance size. High-concurrency applications need [RDS Proxy](https://docs.aws.amazon.com/AmazonRDS/latest/UserGuide/rds-proxy.html) or PgBouncer in front of either service; see the [Aurora parameter reference](https://docs.aws.amazon.com/AmazonRDS/latest/AuroraUserGuide/AuroraPostgreSQL.Reference.ParameterGroups.html). * **Extension availability is a per-engine allowlist.** Extensions that work on RDS, such as `pg_repack`, may be missing or version-pinned on Aurora. Diff your `pg_extension` list against the Aurora support matrix before migrating, not after. * **Aurora is managed as a cluster.** Some parameters apply cluster-wide, some system views and wait events differ from community PostgreSQL, and storage is billed by consumption plus I/O requests rather than provisioned capacity. * **Neither service gives host or superuser access.** `rds_superuser` is a reduced role; anything requiring OS access, arbitrary `shared_preload_libraries`, or untrusted languages is out of scope on both. ## Validate before production [#validate-before-production] 1. Diff installed extensions and their exact versions against the target engine's allowlist. 2. Run a real failover drill and time the application's reconnection behavior, not just the DNS flip. 3. Size the connection budget from `max_connections`, then decide on RDS Proxy or PgBouncer and its pooling mode. 4. Perform a PITR restore into a fresh instance or cluster and measure actual RPO/RTO. 5. Model cost with realistic I/O — Aurora bills I/O requests separately, which surprises workloads migrating from provisioned RDS storage. 6. Export once with `pg_dump` or snapshot export and restore into an independent PostgreSQL environment as an exit-path test. For connection pooling, backup drills, and monitoring baselines that apply regardless of provider, see the [production stack guide](/en/docs/operations/production-stack). If you only need a development or evaluation database, the [free PostgreSQL guide](/en/docs/cloud/free-postgresql) covers zero-cost alternatives. --- # Free PostgreSQL cloud database guide Canonical URL: https://pg.edu.rich/en/docs/cloud/free-postgresql Last reviewed: 2026-08-02 Classify free databases before comparing them: **managed services that run PostgreSQL**, **developer platforms built around PostgreSQL**, and **independent databases that implement pgwire or part of the SQL dialect**. A working driver connection does not prove extension, transaction, or operational compatibility. This page was checked against provider pages on 2026-08-02. Free limits, regions, project counts, sleep, and backup policies change quickly; reopen every source before creating a project. Free tiers generally carry no production SLA and do not replace an independent export and restore drill. ## Free PostgreSQL service comparison [#free-postgresql-service-comparison] | Service | Current free snapshot | Critical limit | Better fit | | --------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------ | ------------------------------------------------------------------ | | [Supabase](https://supabase.com/pricing) | Up to 2 active projects; 500 MB database per project; Free includes 1 GB file storage and 5 GB egress | Pauses after one inactive week; **no automatic backup or PITR** on Free | BaaS requiring Auth, Storage, Realtime, and APIs | | [Neon](https://neon.com/pricing) | Up to 100 projects; 0.5 GB storage and 100 CU-hours/month per project; 5 GB public transfer | Scales to zero after about 5 idle minutes; limited free restore window | PostgreSQL, branches, preview/CI databases, intermittent workloads | | [Aiven for PostgreSQL](https://aiven.io/docs/products/postgresql/concepts/pg-free-tier) | 1 CPU, 1 GB RAM, 1 GB disk, and backups | Single node, `max_connections=20`, no HA/SLA/VPC/pooler; idle services may power off | Traditional managed-PostgreSQL learning and small validation | | [Nhost](https://nhost.io/pricing) | 1 active project; 1 GB database, 1 GB file storage, and 5 GB egress | Pauses after one inactive week; GraphQL/Auth/Storage create platform coupling | GraphQL-first Hasura applications and an integrated backend | | [Prisma Postgres](https://www.prisma.io/pricing) | 500 MB storage, 100,000 operations/month, and up to 50 databases | Every SQL or Prisma query counts as an operation; Free is positioned for evaluation | Prisma workflows, temporary databases, PR and agent environments | | [Koyeb PostgreSQL](https://www.koyeb.com/docs/databases) | 0.25 vCPU, 1 GB RAM, and 1 GB data | Only 5 active compute hours/month; sleeps when idle | Demos, tutorials, and very infrequent tests, not a persistent API | | [Render Postgres](https://render.com/docs/free#free-postgres) | 1 GB and one free instance per workspace | Expires after 30 days; no backups or managed pooling before deletion | One-off demos and platform evaluation | These units are not interchangeable. Neon CU-hours, Prisma operations, Koyeb active hours, and fixed VM capacity measure different things. Model a real request pattern, then verify whether excess usage pauses, rejects, deletes, or starts billing. ## Why CockroachDB is separate [#why-cockroachdb-is-separate] [CockroachDB Cloud Basic](https://www.cockroachlabs.com/pricing/) currently includes 50 million Request Units and 10 GiB of storage per month, but CockroachDB is an independent distributed SQL database, not a PostgreSQL server. It supports pgwire and much PostgreSQL syntax while retaining differences around range types, FDWs, advisory locks, privileges, and transaction behavior. Use its [PostgreSQL compatibility matrix](https://www.cockroachlabs.com/docs/stable/postgresql-compatibility) as the boundary. Evaluate it independently for globally distributed transactions and multi-region resilience. Do not substitute it for real PostgreSQL when learning extensions, catalogs, WAL, or PostgreSQL operations. ## Direct selection guide [#direct-selection-guide] | Requirement | Evaluate first | Why | | -------------------------------------------- | --------------- | ----------------------------------------------------------------- | | Auth, Storage, Realtime, REST/GraphQL APIs | Supabase | A complete application backend surrounds PostgreSQL | | Branching, preview databases, scale-to-zero | Neon | Database lifecycle fits CI and short-lived environments | | Traditional managed PostgreSQL | Aiven | Resources and limits resemble a small single-node service | | GraphQL-first development | Nhost | PostgreSQL plus Hasura, Auth, and Storage | | Prisma workflow and many temporary databases | Prisma Postgres | Operation billing and Prisma/agent tooling are closely integrated | | Very short demo | Koyeb or Render | Their free limits rule them out as durable data sources | For Drizzle, node-postgres, Kysely, and other ordinary PostgreSQL clients, separately test direct and pooled URLs, prepared statements, migrations, and transaction-pooling behavior on Neon, Supabase, Aiven, Nhost, and Prisma Postgres. A standard connection string is not proof of identical behavior. ## Pre-deployment free-tier checks [#pre-deployment-free-tier-checks] ```sql SELECT version(), current_setting('server_version_num') AS server_version_num, current_database(), current_user; SELECT extname, extversion FROM pg_extension ORDER BY extname; SHOW max_connections; SHOW transaction_read_only; ``` Then verify: 1. PostgreSQL major/minor and exact extension versions; 2. direct and pooled connection purpose, limit, and pool mode; 3. cold start after idle and whether DNS or endpoints change; 4. automatic backup, PITR, retention, and whether Free includes them; 5. egress and operation/CU-hour accounting plus hard limits; 6. `pg_dump` export, restore into local PostgreSQL, and the retrieval window after pause or deletion; 7. whether paid upgrade is in place or requires migration and connection-string changes. Important production systems need measurable RPO/RTO, backup retention, a restore path, support, and incident notification. Even on a paid managed service, perform one off-platform export and independent restore. Use the [cloud production selection checklist](/en/docs/cloud/production-checklist) for production and exit costs, and [PostgreSQL lineage and compatible databases](/en/docs/reference/postgresql-compatible-databases) for engine boundaries. --- # Cloud PostgreSQL entry point Canonical URL: https://pg.edu.rich/en/docs/cloud Last reviewed: 2026-08-06 “Cloud PostgreSQL” is not one product category. Classify the service first to understand which PostgreSQL assumptions remain valid: | Category | Examples | Compatibility boundary | Best fit | | ------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------- | | Managed community PostgreSQL | [Amazon RDS for PostgreSQL](/en/docs/cloud/aws-rds-aurora), Cloud SQL, Azure Database for PostgreSQL, Alibaba Cloud RDS, TencentDB | Runs a community engine, while host access, parameters, extensions, and upgrades are platform-controlled | Preserve strong SQL/tool compatibility while delegating patching, backup, and HA | | PostgreSQL-compatible enhanced engine | [Aurora PostgreSQL-Compatible](/en/docs/cloud/aws-rds-aurora), [AlloyDB](/en/docs/cloud/alloydb), [PolarDB for PostgreSQL](/en/docs/cloud/polardb) | Protocol and broad SQL compatibility; vendor implements storage, replication, release cadence, and some behavior | Trade some portability for elasticity, read scaling, or analytical/AI features | | Developer data platform | [Neon](/en/docs/cloud/neon), [Supabase](/en/docs/cloud/supabase) | PostgreSQL is central, with platform-specific connection, branching, auth, API, realtime, or suspend semantics | Fast delivery, preview environments, small operations teams, or full-stack products | A successful driver connection only proves wire-protocol compatibility. Validate extension versions, parameters, catalog views, replication, poolers, backup export, maintenance restarts, and failover behavior before launch. ## Ownership boundary [#ownership-boundary] Managed services commonly own infrastructure, patch orchestration, automated backups, and some failover. The application team still owns: * schemas, constraints, indexes, SQL, and transaction design; * connection budgets, pooling mode, and retry policy; * explicit RPO/RTO and real restore drills; * data access, keys, networking, and least privilege; * slow queries, bloat, long transactions, vacuum, and cost controls; * major-version upgrades, extension upgrades, and an exit plan. AWS explicitly identifies query tuning as the customer's responsibility for Aurora—a useful starting model for every managed database. See the [Amazon Aurora overview](https://docs.aws.amazon.com/AmazonRDS/latest/AuroraUserGuide/CHAP_AuroraOverview.html). ## Recommended decision order [#recommended-decision-order] 1. State residency, compliance, RPO, RTO, peak connections, latency, and budget limits. 2. Confirm the required PostgreSQL major version and the **exact versions** of extensions. 3. Benchmark representative data and SQL; do not substitute vendor headline numbers. 4. Exercise maintenance, failover, PITR, connection exhaustion, and regional failure. 5. Export with native tools and restore into an independent PostgreSQL environment once. Facts in this section were checked on **2026-08-02**. Cloud features, regions, and plans change quickly; re-check official documentation during procurement and launch. --- # Neon serverless PostgreSQL Canonical URL: https://pg.edu.rich/en/docs/cloud/neon Last reviewed: 2026-08-06 [Neon](https://neon.com/docs/get-started/why-neon) is a serverless PostgreSQL platform: compute and storage are separated, compute autoscales and can scale to zero when idle, and copy-on-write storage makes database branches cheap enough to create one per pull request. It runs community PostgreSQL — recent majors including 17 and 18 are available as of this check — so drivers, SQL, and most extensions behave as expected. For positioning against other platforms, see the [cloud service map](/en/docs/cloud/service-map). ## What Neon changes [#what-neon-changes] * **Database branching as a first-class workflow.** A branch is a copy-on-write clone of data and schema, created in seconds. Preview environments, CI runs, and migration rehearsal can each get a full database without duplicating storage cost. * **Scale-to-zero and autoscaling.** Idle computes suspend automatically and resume on the next connection; busy computes scale within a configured range. Intermittent workloads stop paying for idle capacity. * **Pooled and direct endpoints.** Neon exposes a PgBouncer-based pooled connection string alongside a direct one; they have different purposes and different limits. * **Usage-based billing.** Compute is metered in CU-hours and storage in GB-months, so cost modeling differs from per-instance services; check the [pricing page](https://neon.com/pricing) with a realistic traffic pattern. ## Database branching workflow [#database-branching-workflow] The current [Neon CLI](https://neon.com/docs/reference/neon-cli) is installed as `neon` (`neonctl` remains an alias), checked on 2026-08-06: ```bash npm install -g neon neon auth neon projects create --name myapp neon branches create --name feature/x --parent main neon connection-string feature/x ``` A typical branch-per-feature cycle: ```bash export DATABASE_URL=$(neon connection-string feature/x) psql "$DATABASE_URL" -c "ALTER TABLE users ADD COLUMN beta_flag boolean DEFAULT false;" # branches do not merge automatically — apply the reviewed SQL to main yourself psql "$(neon connection-string main)" -f migrations/0042_add_beta_flag.sql # compare schemas before merging neon branches schema-diff main feature/x neon branches delete feature/x ``` The important discipline is the middle step: a Neon branch copies data and schema at creation, but there is no automatic schema merge back. Migrations still go through your normal review and migration tooling, applied to `main` like any other change. ## Differences to accept [#differences-to-accept] * **Cold starts after suspend.** With scale-to-zero enabled, the first connection after an idle period waits for compute to resume. Tune the suspend timeout, or disable it, for latency-sensitive production services — and test how your driver and pooler retry during a wake-up. * **A project's region is fixed at creation.** Deployment is in one AWS region per project (Azure regions are being phased out), and you cannot move a project between regions later — migration means a new project plus data movement. Check the current [region list](https://neon.com/docs/introduction/regions). * **Extensions come from an allowlist.** Extensions needing OS-level access or arbitrary shared libraries are not available; confirm the exact list and versions in the [extension documentation](https://neon.com/docs/extensions/pg-extensions). * **Branch data governance is on you.** A branch copies production data by default. If branches reach CI or preview environments, plan masking or branch from an anonymized parent. ## Validate before production [#validate-before-production] 1. Measure cold-start latency with your actual driver, ORM, and connection pooler after a real idle period. 2. Decide pooled versus direct connections per workload, and test prepared statements and transaction-pooling behavior on the pooled endpoint. 3. Define branch lifecycle and data-masking rules before branches reach shared environments. 4. Exercise branch-based restore inside your plan's restore window, and export once with `pg_dump` into an independent PostgreSQL. 5. Model CU-hour and storage consumption from a realistic traffic week, not from the free-tier shape. The free plan is covered, with current limits, in the [free PostgreSQL guide](/en/docs/cloud/free-postgresql). It scales to zero after minutes of idleness and carries no production SLA — fine for evaluation, not for a production commitment. For pooling, backup drills, and monitoring that apply on any provider, see the [production stack guide](/en/docs/operations/production-stack). --- # Alibaba Cloud PolarDB for PostgreSQL Canonical URL: https://pg.edu.rich/en/docs/cloud/polardb Last reviewed: 2026-08-06 PolarDB for PostgreSQL is Alibaba Cloud's cloud-native PostgreSQL-compatible database: compute and storage are separated, nodes in a cluster share distributed storage, and read nodes scale independently of the primary. Architecturally it sits in the same category as Aurora — a PostgreSQL-compatible engine on vendor-built shared storage — inside the Alibaba Cloud ecosystem. For the cross-provider view, see the [cloud service map](/en/docs/cloud/service-map). ## Where PolarDB fits [#where-polardb-fits] PolarDB is a candidate when your infrastructure already runs on Alibaba Cloud, or when workloads must be deployed in mainland China regions where AWS and Google Cloud do not operate. It is a poor fit if you need tight integration with AWS or GCP services, or if you require community PostgreSQL minor versions the week they ship — managed engines on any vendor trail the community release calendar. As of 2026-08, PolarDB for PostgreSQL supports community majors 11, 14, 15, 16, 17, and 18 — PostgreSQL 18 compatibility was released in December 2025 — per the official [major version lifecycle](https://www.alibabacloud.com/help/en/polardb/polardb-for-postgresql/major-version-lifecycle-description). Note that Alibaba Cloud also operates a separate Oracle-syntax-compatible PolarDB edition; confirm you are evaluating the PostgreSQL edition. ## Architecture and compatibility [#architecture-and-compatibility] * **Shared distributed storage.** Compute nodes mount a shared storage volume with replicas across availability zones, so adding a read node does not copy the dataset. * **Elastic operations.** Compute specifications and node counts change independently of storage; failover promotes an existing read node. * **Extension allowlist.** Available extensions and their versions are per engine version and per kernel minor release; diff your requirements against the official [supported plugin list](https://help.aliyun.com/zh/polardb/polardb-for-postgresql/list-of-supported-plug-ins) rather than assuming community parity. * **Open-source lineage.** The PolarDB for PostgreSQL kernel is also published as open source ([openpolardb.com](https://openpolardb.com)), which matters for evaluation and long-term exit thinking — though the managed service and the open-source build are not the same artifact. ## China residency and compliance perspective [#china-residency-and-compliance-perspective] Residency is the strongest reason to shortlist PolarDB, and it deserves precision: * Mainland China regions are operated under Chinese regulation by Alibaba Cloud's China entity, with a **separate account system** from Alibaba Cloud International. Accounts, billing, and support do not transfer between the two. * If users and data must remain in mainland China — for latency, ICP-related deployment, or data-residency obligations — a China-region PolarDB cluster satisfies constraints that no AWS or GCP region can. * Conversely, storing personal information of China-based users in an International-region cluster raises cross-border transfer questions under PIPL. Neither direction is self-certifying: confirm the current requirements with counsel or your compliance team, and document which account, region, and legal entity holds the data. ## Differences to accept [#differences-to-accept] * **Parameter and privilege controls.** As on other managed engines, some parameters are locked or platform-managed, and superuser-equivalent access is not available. * **Minor-version lag.** Community minor releases arrive on Alibaba's kernel schedule, not the community's; check the release notes for the kernel version behind a given PostgreSQL major. * **Pricing comparisons are not portable.** Instance, storage, and node billing differ structurally from both Alibaba RDS for PostgreSQL and AWS Aurora; model your own workload in the Alibaba Cloud pricing calculator instead of trusting percentage claims. * **Tooling ecosystem.** Console, CLI, monitoring, and migration (DTS) are Alibaba-specific; operational runbooks from other clouds do not transfer directly. ## Validate before production [#validate-before-production] 1. Diff required extensions — `pgvector`, `pg_trgm`, PostGIS, `pg_cron`, and anything version-sensitive — against the supported plugin list for your target engine version. 2. Check parameter compatibility for any non-default `postgresql.conf` settings you rely on. 3. Rehearse migration with Alibaba DTS (full plus incremental) or `pg_dump`, including a rollback plan and a character-set check. 4. Run a failover drill and measure reconnection behavior through your driver and pooler. 5. If you need cross-region or cross-border replicas, validate the Global Database Network feature against your residency obligations first. 6. Export once with native tools and restore into an independent PostgreSQL environment as an exit-path test. Version support, extension availability, and product structure were checked against Alibaba Cloud documentation on 2026-08-06. China-region capabilities and International-region capabilities can differ; verify against the documentation for the account type you will actually use. For connection budgeting, backup drills, and monitoring baselines that apply on any provider, see the [production stack guide](/en/docs/operations/production-stack). Development-scale evaluation options are covered in the [free PostgreSQL guide](/en/docs/cloud/free-postgresql). --- # Cloud PG production checklist Canonical URL: https://pg.edu.rich/en/docs/cloud/production-checklist Last reviewed: 2026-08-02 ## 1. Compatibility inventory [#1-compatibility-inventory] * What are the PostgreSQL major version, patch cadence, and end-of-support date? * Are every required extension and its exact version present in `pg_extension`? Is upgrade automatic, manual, or migration-based? * Which settings are immutable? Is `shared_preload_libraries` available? * Are logical replication, slots, FDWs, event triggers, and required authentication supported? * Which catalog, statistics, and superuser operations have platform-specific replacements? * Have drivers, ORM, migrations, and backup tools passed the real delivery pipeline? Store the answers as a machine-readable manifest bound to service SKU, region, engine version, and verification date. ## 2. Availability and recovery [#2-availability-and-recovery] | Exercise | Example acceptance condition | | ------------------- | ------------------------------------------------------------------------------------------------------------ | | Forced failover | Clients reconnect inside budget; failed transactions return recognizable SQLSTATE; no silent partial success | | PITR | Restore a new instance to the target time; verify rows, constraints, roles, and extensions; measure RTO | | Accidental deletion | Document separate whole-instance, database, and table-level paths and durations | | Region failure | DNS, keys, object-storage backups, and application compute do not share the database failure domain | | Backup export | Restore a usable copy outside the provider account | Applications need connection timeouts, transaction-level retries, and idempotency keys. Do not replay a write that may have committed unless a business idempotency key can confirm the outcome. ## 3. Connections and elasticity [#3-connections-and-elasticity] Budget connections across every application replica, worker, migration tool, BI client, and agent. Prefer a controlled pool for serverless/agent traffic, while confirming: * whether transaction pooling supports session state, temporary tables, LISTEN/NOTIFY, or prepared statements; * whether scale-down, suspend, or failover changes endpoints or TLS certificates; * which of `statement_timeout`, `idle_in_transaction_session_timeout`, and client timeouts fires first; * whether bursts queue in the application instead of becoming a direct PostgreSQL connection storm. ## 4. Cost model [#4-cost-model] Beyond compute and storage, estimate IOPS, backups, cross-zone/region traffic, replicas, logs, monitoring, proxies, PITR, snapshot export, and support. For AI, separate embedding, index rebuild, vector storage, retrieval, and reranking costs. ## 5. Portability [#5-portability] Quarterly—or before a major upgrade—run: ```bash pg_dump --format=custom --no-owner --no-acl "$DATABASE_URL" > app.dump createdb portability_restore pg_restore --exit-on-error --no-owner --no-acl \ --dbname=portability_restore app.dump ``` This checks logical portability; it does not replace provider PITR. After restore, verify extensions, roles/grants, large objects, sequences, row counts, constraints, critical results, and query plans. ## Launch evidence pack [#launch-evidence-pack] * service, region, SKU, engine, and extension version manifest; * RPO, RTO, connection budget, and capacity model; * failover, PITR, accidental-delete, and off-platform restore reports; * encryption, network, role, RLS, and key-rotation records; * engine/extension upgrade and provider-exit runbooks; * latency, error, WAL, vacuum, storage, and cost data under representative load. --- # Cloud PostgreSQL service map Canonical URL: https://pg.edu.rich/en/docs/cloud/service-map Last reviewed: 2026-08-06 ## Managed community PostgreSQL [#managed-community-postgresql] | Service | Verified capabilities | Validate during selection | | ----------------------------------------------------------------------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------- | | [Amazon RDS for PostgreSQL](https://docs.aws.amazon.com/AmazonRDS/latest/UserGuide/CHAP_PostgreSQL.html) | Automated backup/PITR, Multi-AZ, read replicas, VPC, and TLS | No host access; parameters, privileged capabilities, and extensions come from platform allowlists | | [Cloud SQL for PostgreSQL](https://docs.cloud.google.com/sql/docs/postgres/introduction) | Managed backup, HA/failover, encryption, private/public networking, replicas, and maintenance | Maintenance or configuration may restart instances; check regional, extension, connection, and AI-feature availability | | [Azure Database for PostgreSQL Flexible Server](https://learn.microsoft.com/en-us/azure/postgresql/overview) | Same-zone/zone-redundant HA, PITR, TLS, private networking, managed maintenance, optional built-in PgBouncer | Automated backup retention defaults to 7 days and extends to 35; built-in PgBouncer uses port 6432, so verify pooling mode | | [Alibaba Cloud RDS for PostgreSQL](https://help.aliyun.com/en/rds/apsaradb-rds-for-postgresql/what-is-apsaradb-rds-for-postgresql/) | Basic, High-availability, and Cluster editions; automated/manual backup, read-only instances, and proxy options | HA standbys are not directly readable; sync mode, proxy routing, and backup type vary by architecture | | [TencentDB for PostgreSQL](https://www.tencentcloud.com/document/product/409) | Managed installation, storage, HA, backup, major/minor upgrades, and read-only groups | Official guidance says one read-only instance has no HA/SLA; validate node count, routing, and consistency for production read groups | Cloud providers do not necessarily ship the same engine or extension on the day the community does. Split “supports PostgreSQL 17/18” into three checks: can a new instance use it, can an existing instance upgrade to it, and does the required extension support it? For the two AWS paths — managed community RDS and the Aurora engine — see [AWS RDS for PostgreSQL and Aurora](/en/docs/cloud/aws-rds-aurora). ## PostgreSQL-compatible enhanced engines [#postgresql-compatible-enhanced-engines] | Service | Architecture | Difference to accept | | --------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | [Aurora PostgreSQL-Compatible](https://docs.aws.amazon.com/AmazonRDS/latest/AuroraUserGuide/CHAP_AuroraOverview.html) | Customized PostgreSQL-compatible engine and distributed storage, managed primarily as clusters | Aurora has its own release cadence; extensions come from a [support matrix](https://docs.aws.amazon.com/AmazonRDS/latest/AuroraPostgreSQLReleaseNotes/AuroraPostgreSQL.Extensions.html) and do not automatically upgrade with the community extension | | [AlloyDB for PostgreSQL](https://docs.cloud.google.com/alloydb/docs/overview) | Decoupled compute/storage, cross-zone HA, optional columnar engine, vector and model integrations | Not community-binary-equivalent; validate extensions, parameters, migration tools, analytical paths, and regional capability | “PostgreSQL-compatible” describes a migration starting point, not a test result. Exercise schema migration, critical queries, transactional concurrency, drivers, extensions, and recovery. Per-engine landing pages: [AWS RDS for PostgreSQL and Aurora](/en/docs/cloud/aws-rds-aurora), [AlloyDB for PostgreSQL](/en/docs/cloud/alloydb), and [Alibaba Cloud PolarDB for PostgreSQL](/en/docs/cloud/polardb). ## Developer data platforms [#developer-data-platforms] | Service | Strength | Often-missed production concern | | -------------------------------------------------------------- | ------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------- | | [Neon](https://neon.com/docs/get-started/why-neon) | Separated compute/storage, autoscaling, scale-to-zero, database branching, and pooled connections | Cold starts, changing compute size, branch data governance, and the different purposes of pooled versus direct connections | | [Supabase](https://supabase.com/docs/guides/database/overview) | A full PostgreSQL database per project plus Auth, Storage, Realtime, APIs, and Supavisor | Design RLS correctly before browser-facing data APIs; verify backup/PITR plan, connection budget, and coupling to platform components | For no-cost development environments, use the [free PostgreSQL cloud database guide](/en/docs/cloud/free-postgresql), checked on 2026-08-02. Supabase, Neon, Nhost, and Prisma Postgres have different platform structures and should not be ranked by free storage alone. Per-platform landing pages: [Neon](/en/docs/cloud/neon) and [Supabase](/en/docs/cloud/supabase). Validate recovery granularity, retention, cross-region copies, key dependencies, exportability, and measured restore time. Managed physical backups on platforms such as Azure cannot be exported directly; an exit path usually needs `pg_dump`, logical replication, or a migration service. ## For AI workloads [#for-ai-workloads] When choosing cloud PG for RAG or agent metadata, check: 1. the **exact pgvector version**, HNSW/IVFFlat support, and upgrade cadence; 2. connection limits and pooling for short-lived functions and agents; 3. memory, temporary storage, WAL, and replica lag during vector-index builds; 4. real recall after tenant/ACL filtering, not an unfiltered benchmark; 5. whether the embedding model, vector data, and database share the required residency boundary; 6. whether text, metadata, and embeddings can be exported without pipeline lock-in. Provider model endpoints, automated embedding, or AI assistants can remove glue code, but add privilege, region, model-lifecycle, and cost dimensions. They do not replace database RLS, least-privilege roles, or retrieval evaluation. If a candidate only supports the PostgreSQL protocol or reuses its query layer, continue with [PostgreSQL lineage and compatible databases](/en/docs/reference/postgresql-compatible-databases). --- # Supabase PostgreSQL platform Canonical URL: https://pg.edu.rich/en/docs/cloud/supabase Last reviewed: 2026-08-06 [Supabase](https://supabase.com/docs/guides/database/overview) gives each project a full, dedicated PostgreSQL database and surrounds it with platform services: Auth, Storage, Realtime, auto-generated REST and GraphQL APIs (PostgREST), Edge Functions, and the Supavisor connection pooler. The database is real PostgreSQL — you can connect with any client and use ordinary extensions — but the security model and day-to-day operations are shaped by the platform. For positioning against other options, see the [cloud service map](/en/docs/cloud/service-map). ## What the platform includes [#what-the-platform-includes] * **A complete PostgreSQL instance per project**, with direct connection strings alongside pooled ones. * **Auth** integrated with the database: user identities live in `auth.users`, and JWT claims are available inside SQL. * **Auto-generated data APIs** that expose schema tables to browsers and mobile clients directly. * **Realtime** push built on PostgreSQL logical replication. * **Storage and Edge Functions** for files and server-side logic close to the database. This bundling is the reason to choose Supabase: one platform replaces a backend layer. It is also the main thing to accept — the more of Auth, Storage, Realtime, and the APIs you adopt, the more your application couples to Supabase-specific schemas and services. ## RLS is the security boundary [#rls-is-the-security-boundary] Because browser clients can reach tables through the auto-generated API, [row-level security](https://supabase.com/docs/guides/database/postgres/row-level-security) is not a hardening option — it is the security model: ```sql ALTER TABLE notes ENABLE ROW LEVEL SECURITY; CREATE POLICY "users can read their own notes" ON notes FOR SELECT USING (user_id = auth.uid()); CREATE POLICY "users can insert their own notes" ON notes FOR INSERT WITH CHECK (user_id = auth.uid()); ``` `auth.uid()` reads the user id from the request's JWT, so the frontend can query through the API while the database enforces ownership. Three rules keep this safe: 1. Enable RLS on every table exposed through the API, including new tables added later — the API exposes what the schema exposes. 2. Test policies from the anonymous and authenticated roles, not just with the dashboard's service connection. 3. Never ship the `service_role` key to a client; it bypasses RLS entirely. ## Differences to accept [#differences-to-accept] * **Backup and PITR depend on your plan.** The free plan has no automatic backups or PITR — see the current numbers in the [free PostgreSQL guide](/en/docs/cloud/free-postgresql) — and paid tiers differ in retention and restore granularity. Confirm the [backup documentation](https://supabase.com/docs/guides/platform/backups) matches your RPO before committing data. * **Connection budget is platform-shaped.** Direct connections are limited; serverless and high-concurrency clients go through Supavisor, whose transaction-pooling mode restricts session-level features such as prepared statements. Test your driver and ORM against the pooled string, not only the direct one. * **Realtime consumes database resources.** It streams changes via logical replication, so high-churn tables and large publications create real WAL and replication load on the same instance serving your queries. * **Exit requires unwinding platform pieces.** `pg_dump` gets your data out, but Auth users, Storage objects, Realtime subscriptions, and Edge Functions are platform services that a dump does not capture. ## Validate before production [#validate-before-production] 1. RLS enabled and policy-tested on every API-exposed table, from both anonymous and authenticated contexts. 2. Backup, PITR, and restore granularity confirmed against the plan you actually pay for, with one real restore drill. 3. Driver, ORM, and migration tooling tested against the pooled connection string, including prepared-statement behavior. 4. Realtime load rehearsed on production-shaped write churn. 5. An export drill that covers `pg_dump` plus a plan for Auth, Storage, and functions. Plan limits, backup coverage, and platform features were checked against Supabase documentation on 2026-08-06 and change frequently; re-verify at procurement time. For pooling, monitoring, and recovery drills that apply on any platform, see the [production stack guide](/en/docs/operations/production-stack). --- # PostgreSQL autovacuum and table bloat Canonical URL: https://pg.edu.rich/en/docs/operations/autovacuum-bloat Last reviewed: 2026-08-06 Standard `VACUUM` does more than “free space”: it makes dead row versions reusable, maintains planner statistics and the visibility map, and prevents transaction ID/multixact wraparound. Most systems should leave autovacuum enabled. ## Routine observation [#routine-observation] ```sql SELECT schemaname, relname, n_live_tup, n_dead_tup, last_vacuum, last_autovacuum, vacuum_count, autovacuum_count, last_analyze, last_autoanalyze FROM pg_stat_user_tables ORDER BY n_dead_tup DESC LIMIT 30; ``` Statistics are estimates and can reset, so one `n_dead_tup` threshold cannot prove bloat. Combine table size, update rate, query latency, autovacuum logs, and trends. Inspect active vacuum work: ```sql SELECT pid, datname, relid::regclass AS relation, phase, heap_blks_total, heap_blks_scanned, heap_blks_vacuumed, index_vacuum_count, dead_tuple_bytes, num_dead_item_ids, indexes_total, indexes_processed FROM pg_stat_progress_vacuum; ``` These column names target PostgreSQL 18. Older majors can expose a different progress-view shape, so version-portable monitoring should inspect the target catalog first. ## Why it does not trigger or keep up [#why-it-does-not-trigger-or-keep-up] * a large table makes the default scale factor translate into too many changed rows; * workers, I/O capacity, or maintenance memory are insufficient; * long transactions, prepared transactions, slots, or standby snapshots block reclamation; * conflicting locks repeatedly cancel vacuum; * sustained writes exceed cleanup capacity. Override a measured hot table before making aggressive global changes: ```sql ALTER TABLE app.events SET ( autovacuum_vacuum_scale_factor = 0.02, autovacuum_vacuum_threshold = 1000, autovacuum_analyze_scale_factor = 0.01 ); ``` These are examples, not universal values. Calculate expected trigger frequency from table size and daily changes, then observe I/O, WAL, latency, and completion time. ## Manual maintenance boundary [#manual-maintenance-boundary] ```sql VACUUM (ANALYZE, VERBOSE) app.events; ``` Plain `VACUUM` mainly makes space reusable inside the relation and normally does not shrink the file back to the operating system. `VACUUM FULL` rewrites the table, needs extra disk, and takes `ACCESS EXCLUSIVE`; it is not routine cleanup. ## Measuring and repairing bloat [#measuring-and-repairing-bloat] Dead-tuple ratios from `pg_stat_user_tables` are estimates. [`pgstattuple`](https://www.postgresql.org/docs/18/pgstattuple.html) scans the relation and reports exact dead tuples and free space: ```sql CREATE EXTENSION IF NOT EXISTS pgstattuple; SELECT * FROM pgstattuple('app.events'); -- dead_tuple_count, dead_tuple_percent, free_space, free_percent ``` When space must be returned to the operating system, plain `VACUUM` is not enough and `VACUUM FULL` blocks reads and writes for the whole rewrite. The usual online options: ```bash # pg_repack rebuilds the table in the background with only short locks sudo -u postgres pg_repack -h db.example -U postgres -d commerce --table=events ``` ```sql -- Index-only bloat: rebuild one index without locking the table REINDEX INDEX CONCURRENTLY app.events_pkey; ``` `REINDEX CONCURRENTLY` is built in since PostgreSQL 12. `pg_repack` is a third-party extension with its own release and operational track — install it, pin its version, and rehearse it separately from core features. ### When VACUUM FULL appears stuck [#when-vacuum-full-appears-stuck] It is almost always waiting for `ACCESS EXCLUSIVE` behind existing sessions. Identify the blockers before retrying: ```sql SELECT a.pid, a.usename, a.state, a.wait_event, left(a.query, 160) AS query FROM pg_stat_activity a WHERE a.pid = ANY (pg_blocking_pids(12345)); -- pid of the VACUUM FULL session ``` Cancel or terminate the holders only after confirming they are disposable. Even when it runs, `VACUUM FULL` needs free disk roughly equal to the table size and rewrites every index — another reason to prefer `pg_repack` for routine space reclamation. ## Transaction ID wraparound defense [#transaction-id-wraparound-defense] Transaction IDs are 32-bit: after roughly two billion transactions, old tuples would appear to lie in the future, and PostgreSQL stops accepting writes before that can happen. Autovacuum prevents this by freezing old tuples so their IDs can be reused. Track freeze age per database and per table: ```sql SELECT datname, age(datfrozenxid) AS xid_age FROM pg_database ORDER BY xid_age DESC; SELECT n.nspname, c.relname, age(c.relfrozenxid) AS xid_age, pg_size_pretty(pg_total_relation_size(c.oid)) AS total_size FROM pg_class c JOIN pg_namespace n ON n.oid = c.relnamespace WHERE c.relkind = 'r' AND n.nspname NOT IN ('pg_catalog', 'information_schema') ORDER BY xid_age DESC LIMIT 20; ``` Common alert starting points: a database `xid_age` past one billion transactions deserves attention, and past 1.5 billion is an emergency — calibrate against your transaction rate. If autovacuum cannot freeze fast enough, force a freeze during a low-traffic window and watch it in `pg_stat_progress_vacuum`: ```sql VACUUM (FREEZE, VERBOSE) app.events; ``` Omit the table name to freeze the whole database when database-level age is the problem. Identify the table, phase, wait event, and resource bottleneck first. Disabling autovacuum accumulates dead tuples, stale statistics, and freeze risk; anti-wraparound vacuum can run even when table-level autovacuum is disabled. Read [PostgreSQL 18 Routine Vacuuming](https://www.postgresql.org/docs/18/routine-vacuuming.html) for normative behavior. --- # Backup, recovery, and PITR Canonical URL: https://pg.edu.rich/en/docs/operations/backup-recovery Last reviewed: 2026-08-06 ## Choose the mechanism [#choose-the-mechanism] | Need | Mechanism | Boundary | | ------------------------------------------- | --------------------------------------------------- | ------------------------------------------------- | | One database, portability, object selection | `pg_dump` / `pg_restore` | Excludes cluster roles and tablespace definitions | | Logical cluster objects | `pg_dumpall --globals-only` plus per-database dumps | Slow for large systems; rebuilds indexes | | Fast whole-instance recovery | `pg_basebackup` or a mature backup tool | Stronger version/platform constraints | | Point-in-time recovery | Physical base backup plus continuous WAL archive | WAL continuity must be verified continuously | ### Choosing pgBackRest, WAL-G, or pg\_dump [#choosing-pgbackrest-wal-g-or-pg_dump] | Option | Better fit | Does not prove by itself | | -------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------- | | [`pgBackRest`](https://github.com/pgbackrest/pgbackrest) | Full/differential/incremental backup, parallelism, multiple repositories, WAL, and PITR for self-hosted instances | That target RTO is met or every key and WAL file is usable | | [`WAL-G`](https://github.com/wal-g/wal-g) | Object-storage-oriented physical backup and WAL workflows | Repository retention, deletion protection, or restore correctness | | `pg_dump` / `pg_restore` | Logical migration, object selection, small restores, and cross-version export | Continuous PITR or a low-RTO whole-instance restore | | Cloud-platform backup | Lower infrastructure maintenance | Cross-account, cross-region, off-platform, or complete extension recovery | Do not copy a fixed daily/weekly schedule. Derive backup cadence from RPO, WAL volume, restore bandwidth, retention policy, and measured RTO, and keep at least one copy outside the primary database permission boundary. ## Logical backup [#logical-backup] Custom format supports parallel restore and object selection: ```bash pg_dump \ --format=custom \ --file=commerce-20260802.dump \ --dbname='postgresql://backup@db.example/commerce' pg_restore --list commerce-20260802.dump createdb commerce_restore_test pg_restore \ --dbname=commerce_restore_test \ --jobs=4 \ --exit-on-error \ commerce-20260802.dump ``` `pg_dump` provides a consistent snapshot during export but covers one database. Back up global objects separately: ```bash pg_dumpall --globals-only > globals-20260802.sql ``` Do not put globals files containing password hashes in a general artifact store. ## Physical backup and PITR [#physical-backup-and-pitr] PITR requires a usable base backup, an unbroken WAL stream from that backup, correct recovery configuration, and timeline handling. WAL alone is insufficient; a base backup alone cannot recover to an arbitrary point. An archive command returns zero only after a safe copy and never overwrites an existing file. For object storage, use a mature tool for concurrency, checksums, retention, and encryption rather than an unmonitored shell one-liner. Continuously alert on archive failures, missing WAL, repository capacity, and the latest recoverable time. Let the backup tool calculate dependency-aware retention; do not delete physical backup files solely by date. ### PITR walkthrough [#pitr-walkthrough] A minimal manual restore to a point in time: ```bash sudo systemctl stop postgresql # restore the base backup into an empty data directory tar -xzf /backup/2026-08-01/base.tar.gz -C "$PGDATA" touch "$PGDATA"/recovery.signal ``` ```ini # postgresql.conf (or postgresql.auto.conf) restore_command = 'cp /archive/%f %p' recovery_target_time = '2026-08-01 14:30:00+08' ``` On start, the server replays WAL up to the target and pauses (default `recovery_target_action = 'pause'`). Verify the data, then finish with `SELECT pg_wal_replay_resume();`. Rehearse timeline handling as well: each completed recovery creates a new timeline, and archiving must follow it. ### pgBackRest configuration example [#pgbackrest-configuration-example] A minimal repository setup: ```ini # /etc/pgbackrest/pgbackrest.conf [global] repo1-path=/var/lib/pgbackrest # example; derive retention from RPO and storage repo1-retention-full=4 repo1-cipher-type=aes-256-cbc repo1-cipher-pass= [main] pg1-path=/var/lib/postgresql/18/main ``` ```bash sudo -u postgres pgbackrest --stanza=main stanza-create sudo -u postgres pgbackrest --stanza=main check sudo -u postgres pgbackrest --stanza=main --type=full backup sudo -u postgres pgbackrest --stanza=main --type=incr backup sudo -u postgres pgbackrest --stanza=main info # Point-in-time restore sudo systemctl stop postgresql sudo -u postgres pgbackrest --stanza=main \ --type=time --target='2026-08-01 14:30:00+08' restore sudo systemctl start postgresql ``` `check` verifies WAL archiving end to end; run it after every configuration change. Retention, encryption keys, and repository permissions must all be available at restore time — losing the cipher pass renders the repository useless. ## Restore drill [#restore-drill] Run every drill on an isolated instance — a spare host or a second port on a test machine, never the production data directory: 1. Restore the most recent backup to a scratch path, for example `pgbackrest --stanza=main --pg1-path=/tmp/restore-test restore`. 2. Start a throwaway instance on a separate port: `postgres -D /tmp/restore-test -p 6543`. 3. Run the validation queries and smoke tests below. 4. Confirm the required extensions, roles, and restore keys were available; per-database dumps exclude cluster roles. 5. Stop and delete the scratch instance. Record backup id, start/end, recovery target, server version, required keys, actual RTO, latest recoverable transaction time, validation queries, and anomalies. At minimum inspect: ```sql SELECT count(*) FROM critical_table; SELECT min(created_at), max(created_at) FROM critical_table; SELECT conname, convalidated FROM pg_constraint WHERE NOT convalidated; SELECT indexrelid::regclass, indisvalid FROM pg_index WHERE NOT indisvalid; ``` Then run application-level read-only smoke tests. Matching row counts do not prove relations and privileges are correct. Replication quickly copies accidental deletes, bad updates, and logical corruption. Backups need independent retention, deletion protection, verification, and restore drills. --- # PostgreSQL configuration tuning (postgresql.conf) Canonical URL: https://pg.edu.rich/en/docs/operations/configuration Last reviewed: 2026-08-06 The default `postgresql.conf` ships conservative values so PostgreSQL can start on almost any hardware. That makes the defaults a poor fit for a dedicated server, but it does not make any fixed set of numbers a "best practice." The values below are **starting points for PostgreSQL 18**, derived from common rules of thumb: apply them, then validate against your own workload before and after each change. ## Measure before tuning [#measure-before-tuning] Changing GUCs without a baseline is how a slow system becomes a differently slow system. Before touching anything: * enable `pg_stat_statements` (add it to `shared_preload_libraries`, restart, then `CREATE EXTENSION pg_stat_statements;`) so you can rank queries by total time and by `shared_blks_read` vs `shared_blks_hit`; * record current latency, I/O, and checkpoint behavior, so every change can be attributed; * change one group of settings at a time and re-measure. ```sql SELECT query, calls, total_exec_time, mean_exec_time, shared_blks_hit, shared_blks_read FROM pg_stat_statements ORDER BY total_exec_time DESC LIMIT 20; ``` If `pg_stat_statements` shows the hot queries are missing indexes or doing full scans of small tables, no memory GUC will fix that. Tune what the measurements point at. ## Core memory GUCs [#core-memory-gucs] ### shared\_buffers [#shared_buffers] PostgreSQL's own page cache, allocated from shared memory at startup. Requires a restart. ```sql ALTER SYSTEM SET shared_buffers = '4GB'; ``` A common starting point is **25% of RAM** on a dedicated database server. The frequently quoted "\~25%, never above 40%" ceiling is folklore, not a documented limit — write-heavy workloads can benefit from more, and read-mostly workloads that already fit in shared\_buffers may benefit from less. Validate larger values with `pg_buffercache` hit rates and end-to-end latency instead of assuming more is better. ### effective\_cache\_size [#effective_cache_size] Not an allocation — a planner hint about how much data is likely cached across shared\_buffers plus the OS page cache. It mainly influences index scan vs sequential scan choices. Reloadable. ```sql ALTER SYSTEM SET effective_cache_size = '12GB'; ``` A reasonable starting point is **50–75% of RAM**. Setting it far above actual cacheable memory makes the planner over-optimistic about index scans; measure plan changes before and after. ### work\_mem [#work_mem] Memory limit **per sort, hash, or similar operation** — a single complex query can consume several times this amount, across every concurrent backend. This is the GUC that causes OOM when raised carelessly. ```sql ALTER SYSTEM SET work_mem = '16MB'; ``` * The default 4MB forces many sorts and hashes to spill to disk; `log_temp_files` (see [Logging](#logging-worth-enabling-in-production)) tells you whether this is happening. * Raise it per role or per session for heavy analytical users instead of globally: `ALTER ROLE analytics SET work_mem = '256MB';` * A safe global upper bound is roughly `RAM × 0.25 / max_connections`, and even that assumes every backend runs one big operation at a time. ### maintenance\_work\_mem [#maintenance_work_mem] Used by `VACUUM`, `CREATE INDEX`, `ALTER TABLE ADD FOREIGN KEY`, and similar maintenance operations. ```sql ALTER SYSTEM SET maintenance_work_mem = '1GB'; ``` Larger values speed up index builds and vacuum. Remember that parallel index builds and parallel vacuum workers each allocate their own share, so the total can be several times this value. ## Connections and pooling [#connections-and-pooling] ```sql ALTER SYSTEM SET max_connections = 200; ``` Each PostgreSQL connection is a separate backend process with several MB of memory and its own scheduling cost, so `max_connections` is a budget, not a throughput dial. If the application needs more client concurrency than the database can serve, put a pooler in front rather than raising this number indefinitely — see [PostgreSQL production stack and HA](/en/docs/operations/production-stack) for PgBouncer trade-offs, including what breaks under transaction pooling. ```sql SELECT count(*), state FROM pg_stat_activity GROUP BY state; ``` A large `idle in transaction` count means the application is holding transactions open; fix that in the application before raising any connection limit. `max_connections` requires a restart. ## WAL and checkpoints [#wal-and-checkpoints] ```ini wal_compression = on max_wal_size = 4GB min_wal_size = 1GB checkpoint_timeout = 15min checkpoint_completion_target = 0.9 ``` `max_wal_size` that is too small forces frequent checkpoints and I/O spikes; too large lengthens crash recovery and WAL replay. Values between 4GB and 16GB are common on write-heavy systems — pick from your measured WAL generation rate (`pg_stat_bgwriter`, `pg_stat_wal`), not from a table. `wal_buffers` defaults to `-1` (auto-sized from shared\_buffers) and usually does not need an explicit value. ## Planner and I/O costing [#planner-and-io-costing] ```ini random_page_cost = 1.1 # SSD; the default 4.0 dates from the HDD era effective_io_concurrency = 200 # SSD/NVMe can sustain far more than the default 16 jit = off # JIT helps long analytical queries; often a cost in OLTP ``` The default `random_page_cost = 4.0` assumes random reads are four times more expensive than sequential ones. On SSD/NVMe storage that makes the planner avoid index scans it should take. 1.1 is a widely used SSD value, but confirm with `EXPLAIN (ANALYZE, BUFFERS)` on your own queries rather than trusting any fixed number. `default_statistics_target` defaults to 100 and rarely needs a global change. For specific skewed columns, raise statistics per column instead: `ALTER TABLE t ALTER COLUMN c SET STATISTICS 1000;` followed by `ANALYZE t;`. `jit` defaults to on in PostgreSQL 18. If your workload is dominated by short OLTP queries, disabling it globally is a defensible starting point; enable it per session for reporting queries that actually benefit. ## Three ways to change a setting [#three-ways-to-change-a-setting] ```sql -- 1. ALTER SYSTEM — writes postgresql.auto.conf, survives restarts ALTER SYSTEM SET work_mem = '32MB'; SELECT pg_reload_conf(); -- 2. Edit postgresql.conf directly, then -- sudo systemctl reload postgresql -- 3. Session or transaction scope, for one-off work SET work_mem = '256MB'; SET LOCAL work_mem = '256MB'; -- current transaction only ``` Prefer `ALTER SYSTEM` for persistent changes: it is auditable (`pg_settings` shows `source = 'configuration file'` with the auto.conf path) and keeps hand-edited files out of the change path. Remove a setting with `ALTER SYSTEM RESET name;`. Not every change applies on reload. Check the context before assuming: ```sql SELECT name, setting, context FROM pg_settings WHERE context IN ('postmaster', 'superuser-backend') ORDER BY name; -- context = 'postmaster' requires a full restart ``` Common restart-required GUCs: `shared_buffers`, `max_connections`, `shared_preload_libraries`. ## Starter templates by machine size [#starter-templates-by-machine-size] These are starting points for a dedicated PostgreSQL 18 server on SSD/NVMe storage, to be validated with the measurement workflow above. ### 4 vCPU / 16 GB [#4-vcpu--16-gb] ```ini shared_buffers = 4GB effective_cache_size = 12GB maintenance_work_mem = 1GB work_mem = 16MB max_connections = 100 wal_compression = on max_wal_size = 4GB random_page_cost = 1.1 effective_io_concurrency = 200 max_parallel_workers = 4 max_parallel_workers_per_gather = 2 ``` ### 8 vCPU / 32 GB [#8-vcpu--32-gb] ```ini shared_buffers = 8GB effective_cache_size = 24GB maintenance_work_mem = 2GB work_mem = 32MB max_connections = 200 wal_compression = on max_wal_size = 8GB random_page_cost = 1.1 effective_io_concurrency = 200 max_parallel_workers = 6 max_parallel_workers_per_gather = 4 ``` ### 16 vCPU / 64 GB [#16-vcpu--64-gb] ```ini shared_buffers = 16GB effective_cache_size = 48GB maintenance_work_mem = 4GB work_mem = 64MB max_connections = 500 # with a pooler in front wal_compression = on max_wal_size = 16GB random_page_cost = 1.1 effective_io_concurrency = 200 max_parallel_workers = 12 max_parallel_workers_per_gather = 6 ``` [PGTune](https://pgtune.leopard.in.ua/) generates an equivalent template from machine specs if you want a second opinion to compare against. ## Logging worth enabling in production [#logging-worth-enabling-in-production] ```ini log_min_duration_statement = '500ms' log_checkpoints = on log_connections = on log_disconnections = on log_lock_waits = on log_temp_files = 0 log_autovacuum_min_duration = 0 log_line_prefix = '%t [%p] %u@%d %a ' ``` `log_temp_files = 0` logs every temporary file creation and is the direct signal that `work_mem` is too small for real queries. `log_autovacuum_min_duration = 0` makes autovacuum behavior auditable — combine it with the queries in [autovacuum and table bloat](/en/docs/operations/autovacuum-bloat). `log_connections`/`log_disconnections` are cheap on pooled workloads but can be noisy with very short-lived connections; adjust to taste. See [Monitoring and logging](/en/docs/operations/monitoring-logging) for the full observability setup. ## Parallel query [#parallel-query] ```ini max_worker_processes = 8 # total background workers, including replication max_parallel_workers = 6 # workers available to parallel query max_parallel_workers_per_gather = 4 # workers a single query node can use min_parallel_table_scan_size = '8MB' ``` `max_worker_processes` is the global budget for all background workers (parallel query, logical replication apply workers, and extensions), so it must be at least as large as `max_parallel_workers` plus whatever replication and extensions need; both it and `max_parallel_workers` are typically sized from vCPU count. Parallelism pays off on large scans; small OLTP queries rarely trigger it once `min_parallel_table_scan_size` is at or above the default 8MB. ## Inspect the running configuration [#inspect-the-running-configuration] ```sql -- everything that differs from defaults, and where it came from SELECT name, setting, unit, source FROM pg_settings WHERE source <> 'default' ORDER BY name; -- current value vs the value at last startup (pending restarts) SELECT name, setting, boot_val, pending_restart FROM pg_settings WHERE setting <> boot_val OR pending_restart; ``` `pending_restart = true` means an `ALTER SYSTEM` or config edit is waiting for a restart to take effect — check it before concluding a change "did nothing." ## Prompt: have AI draft a config for your hardware [#prompt-have-ai-draft-a-config-for-your-hardware] Treat the output the same way as the templates on this page: a hypothesis to verify against measurements, not a finished configuration. ## Related pages [#related-pages] * Connection pooling, backup, and HA decisions → [PostgreSQL production stack and HA](/en/docs/operations/production-stack) * Autovacuum tuning and bloat detection → [autovacuum and table bloat](/en/docs/operations/autovacuum-bloat) * Metrics and log pipelines → [Monitoring and logging](/en/docs/operations/monitoring-logging) --- # PostgreSQL memory management and OOM in containers Canonical URL: https://pg.edu.rich/en/docs/operations/container-memory-oom Last reviewed: 2026-08-06 The classic container OOM report looks like this: PostgreSQL restarts every few hours, `dmesg` shows `Out of memory: Killed process ... (postgres)`, yet the team insists the pod has "plenty of memory" because `free -m` inside the container shows gigabytes free. Both observations are true, and they do not contradict each other: `free` reads `/proc/meminfo`, which is **not namespaced** — inside a container it reports the host's memory, not the cgroup limit the kernel actually enforces. The limit that kills you lives in the cgroup, not in `/proc/meminfo`. ## Where the real limit is read from [#where-the-real-limit-is-read-from] On cgroup v2 hosts, the effective limit is in `memory.max` (and the current charge in `memory.current`); on cgroup v1 it is `memory.limit_in_bytes` and `memory.usage_in_bytes`: ```bash # cgroup v2 (Kubernetes 1.31+, most modern distros) cat /sys/fs/cgroup/memory.max cat /sys/fs/cgroup/memory.current # cgroup v1 cat /sys/fs/cgroup/memory/memory.limit_in_bytes ``` See the kernel documentation for [cgroup v2](https://docs.kernel.org/admin-guide/cgroup-v2.html) and the [cgroup v1 memory controller](https://docs.kernel.org/admin-guide/cgroup-v1/memory.html) for the full file list. Any sizing tool or runbook that derives "available memory" from `/proc/meminfo` inside a container is measuring the wrong box. ## The budget that must fit inside the limit [#the-budget-that-must-fit-inside-the-limit] PostgreSQL does not read cgroup limits. It sizes itself from `postgresql.conf`, so the configuration must be derived from the container limit, not from host RAM. The pieces that must fit (see [Resource Consumption](https://www.postgresql.org/docs/18/runtime-config-resource.html) in the PostgreSQL 18 docs): * `shared_buffers` — fixed at startup; * `work_mem` × concurrent sorts/hashes — this is **per operation**, not per session: one query with several hash joins can consume several times `work_mem`, so `max_connections × work_mem` is only a rough ceiling; * `maintenance_work_mem` (or `autovacuum_work_mem`) × concurrent vacuum workers and maintenance commands; * `temp_buffers` per backend, WAL buffers, and per-connection overhead of a few MB per backend process; * plus everything else in the container (sidecars, monitoring agents) and the page cache PostgreSQL depends on. As a starting point, not a rule: on a dedicated PostgreSQL container, `shared_buffers` around 25% of the cgroup limit leaves room for connections and page cache; raise it only after measuring. The full configuration checklist is in [Server configuration](/en/docs/operations/configuration). ## What PostgreSQL counts vs what the cgroup counts [#what-postgresql-counts-vs-what-the-cgroup-counts] Two accounting differences explain most "we were nowhere near the limit" surprises: * **Page cache is charged to the cgroup.** Reads and writes of heap and WAL files are buffered by the kernel and charged to the cgroup that first touches each page. A container whose processes hold little anonymous memory can still sit at its limit because the rest is page cache — which is normal and mostly reclaimable, until a burst of anonymous allocations (a big `work_mem` spike, a new connection wave) arrives while the cache is dirty or actively referenced. * **RSS grows with touched pages.** Backend processes share `shared_buffers`, and tools that sum RSS across processes double-count those shared pages. Do not size the limit by summing `ps` output; size it from `memory.current` under peak load. The kernel only kills when there is no reclaimable memory left at the moment of allocation, which is why OOM kills cluster under load spikes rather than at steady state. ## OOM killer scoring and oom\_score\_adj [#oom-killer-scoring-and-oom_score_adj] When the cgroup (or the host) hits its limit, the kernel picks a victim by `oom_score`, adjustable per process through `/proc//oom_score_adj` (range `-1000` to `1000`; `-1000` exempts the process entirely — see the [proc filesystem documentation](https://docs.kernel.org/filesystems/proc.html)). Two operational facts matter: * Protecting the postmaster with `-1000` is the standard recipe, but **`oom_score_adj` is inherited by child processes**: every backend forked afterwards also becomes unkillable, which is the opposite of what you want. If the postmaster is protected, backends must reset their own score — for example by re-setting `oom_score_adj` from a wrapper, or by accepting that on Kubernetes you control this at pod level instead. * On Kubernetes you cannot set `oom_score_adj` per container; the kubelet assigns it from the pod's QoS class: Guaranteed pods get `-997`, BestEffort pods get `1000`, Burstable pods land in between. This is documented under [node-pressure eviction](https://kubernetes.io/docs/concepts/scheduling-eviction/node-pressure-eviction/). One database per pod is what makes this protection meaningful. Also set `vm.overcommit_memory = 2` considerations aside carefully: the PostgreSQL documentation's [Linux Memory Overcommit](https://www.postgresql.org/docs/18/kernel-resources.html) section explains the trade-off — raising `vm.overcommit_memory` to `2` lowers the chance of the OOM killer being invoked at all, at the cost of failing allocations earlier. ## Kubernetes requests, limits, and QoS [#kubernetes-requests-limits-and-qos] For a database pod, the defensible pattern from the [Kubernetes memory resource documentation](https://kubernetes.io/docs/tasks/configure-pod-container/assign-memory-resource/) and [Pod QoS classes](https://kubernetes.io/docs/concepts/workloads/pods/pod-qos/): * Set `requests.memory` equal to `limits.memory` so the pod is **Guaranteed**: it will not be evicted under node pressure before Burstable/BestEffort pods, and it gets the favorable `oom_score_adj`. * Derive `postgresql.conf` from the limit (previous section), not from the node size — pods get rescheduled to larger nodes and silently change their environment. * Leave headroom in the limit for page cache and spikes: sizing `shared_buffers + connections` to consume nearly 100% of the limit guarantees the next `work_mem` burst is fatal. ## Diagnosing an OOM kill [#diagnosing-an-oom-kill] Work from the kernel evidence inward to the query: ```bash # Kernel record of the kill (host or via kubectl logs of the node) dmesg -T | grep -i -E 'out of memory|oom_kill' journalctl -k | grep -i oom # cgroup v2: oom_kill counter increments per kill in this cgroup cat /sys/fs/cgroup/memory.events # low 0 / high 0 / max / oom / oom_kill # cgroup v1 equivalent cat /sys/fs/cgroup/memory/memory.oom_control ``` `memory.events` also shows `max` events (limit hits that triggered reclaim without a kill) — a rising `max` count with zero `oom_kill` means you are permanently at the ceiling and a kill is a matter of time. On the PostgreSQL side, look for the memory consumers active just before the kill. Temp-file usage is the classic `work_mem` overflow signature; it is tracked per database, not per session: ```sql SELECT datname, temp_files, pg_size_pretty(temp_bytes) AS temp_written FROM pg_stat_database ORDER BY temp_bytes DESC; ``` `temp_bytes` is cumulative, so compare deltas around the incident window. Attribution to specific statements comes from the log: set `log_temp_files = 0` (or a threshold) so every spill records the query that caused it, then correlate with `pg_stat_activity` sessions that were active at the kill time. The monitoring pipeline for this is covered in [Monitoring and logging](/en/docs/operations/monitoring-logging). ## Prevention checklist [#prevention-checklist] * One postmaster per pod; Guaranteed QoS (`requests == limits`). * `shared_buffers`, `work_mem`, `maintenance_work_mem`, and `max_connections` derived from the cgroup limit, not host RAM — re-derive after any limit change. * Connection pooler in front of the database so `max_connections × work_mem` stays bounded. * `log_temp_files = 0` (or a threshold) and alerts on `temp_bytes` growth. * Alert on `memory.events` `oom_kill` increments and on `max` counter growth. * Never trust `free`, `top`, or `htop` inside a container for capacity decisions. Raising the limit without re-deriving `postgresql.conf` just moves the crash. Either the configuration is too large for the limit, or a query-level consumer (`work_mem` spike, connection wave) is unbounded — find which before buying more memory. Treat the output as a hypothesis: apply it in a test pod and watch `memory.current` and `memory.events` under peak load before changing production. ## Related pages [#related-pages] * Deriving `postgresql.conf` values → [Server configuration](/en/docs/operations/configuration) * Temp-file and statement-level evidence → [Monitoring and logging](/en/docs/operations/monitoring-logging) * When metadata itself eats the memory → [Too many tables are a problem](/en/docs/operations/too-many-tables) --- # Enable data checksums online Canonical URL: https://pg.edu.rich/en/docs/operations/data-checksums Last reviewed: 2026-08-06 As of 2026-08, PostgreSQL 19 has not reached GA. Function signatures, states, and view columns below follow the [PostgreSQL 19 documentation](https://www.postgresql.org/docs/19/checksums.html); re-check the final release notes before relying on them in production. Data checksums store a checksum value in every data page, written when the page is flushed and verified when it is read back — the primary mechanism for detecting storage and file-system corruption. Since PostgreSQL 18, `initdb` enables them by default for new clusters. Clusters initialized on older major versions often still run without them, which is exactly the gap this page covers. ## Before PostgreSQL 19: pg\_checksums required downtime [#before-postgresql-19-pg_checksums-required-downtime] Through PostgreSQL 18, changing the checksum state of an existing cluster meant [`pg_checksums`](https://www.postgresql.org/docs/19/app-pgchecksums.html), and `pg_checksums` requires the server to be **shut down cleanly**. Enabling checksums rewrites every relation block in place, so on a multi-terabyte cluster the maintenance window is long, and the cluster must not be started mid-operation. For always-on systems this made retroactively enabling checksums nearly impossible to schedule. ## Enabling checksums online [#enabling-checksums-online] PostgreSQL 19 adds SQL functions that flip checksums on a running cluster while clients keep working, described in [Section 28.2, Data Checksums](https://www.postgresql.org/docs/19/checksums.html): ```sql -- check current state: off / inprogress-on / on / inprogress-off SHOW data_checksums; -- start enabling; cost_delay / cost_limit throttle the rewrite SELECT pg_enable_data_checksums(cost_delay => 10, cost_limit => 1000); ``` What happens after the call: 1. The cluster state becomes `inprogress-on`. From this point checksums are *written* on page flush but not yet *verified* on read. 2. A launcher starts one background worker per database, which walks every relation and marks each page dirty so it is rewritten with a checksum. 3. When all databases finish, the state automatically switches to `on` and read-time verification begins. Prerequisites and stalls to know about: * The process consumes two background worker slots — make sure `max_worker_processes` has headroom. * It waits for all open transactions to finish before starting, and for each database it waits for pre-existing temporary tables to be dropped. Applications with long-lived temp tables can block completion indefinitely; terminating those connections may be necessary. * If the cluster stops while in `inprogress-on`, there is **no resume**: after restart, re-run `pg_enable_data_checksums()` and the rewrite starts over. Plan it for a window without reboots or failovers. ## I/O cost and progress monitoring [#io-cost-and-progress-monitoring] Enabling checksums rewrites every page in the cluster — dirty page flushes plus WAL for the rewrites — so the I/O impact is substantial. The `cost_delay` and `cost_limit` arguments of `pg_enable_data_checksums()` apply the same cost-based throttling model as vacuum; on production systems, start with conservative values and watch latency. Progress is visible in `pg_stat_progress_data_checksums`: one row for the launcher (tracking `databases_total` / `databases_done`) and one row per worker (tracking `relations_total` / `relations_done` and `blocks_total` / `blocks_done` for the current relation): ```sql SELECT pid, datname, phase, databases_total, databases_done, relations_total, relations_done, blocks_total, blocks_done FROM pg_stat_progress_data_checksums; ``` The `phase` column distinguishes `enabling`, `disabling`, `waiting on barrier` (backends acknowledging the state change), and `waiting on temporary tables` — the latter two explain most "stuck" appearances. The processes also show up in `pg_stat_activity` with `backend_type` of `datachecksums launcher` / `datachecksums worker`. Replication note from the documentation: when a standby receives the checksum state change in the WAL stream it forces a restartpoint, which blocks redo until finished and can induce replication lag — on synchronous standbys this also blocks the primary. Reducing `max_wal_size` before starting shortens that restartpoint. ## Disabling checksums online [#disabling-checksums-online] ```sql SELECT pg_disable_data_checksums(); ``` The state moves to `inprogress-off` (checksums still written, no longer verified) and settles to `off` once all backends acknowledge the change. No pages are rewritten, so there is no I/O surge, though checkpoints are still required. Disabling while an enable is in progress aborts the enable. If the cluster stops during `inprogress-off`, it comes back up with checksums `off`. ## Hardware-accelerated checksum computation [#hardware-accelerated-checksum-computation] The checksum stored in data pages is PostgreSQL's own FNV-1a-based algorithm rather than CRC — on x86-64 an AVX2-vectorized implementation is selected at runtime when available. CRC-32C protects WAL records instead, with runtime dispatch to SSE4.2 or AVX-512 instructions on x86 and the CRC32/PMULL extensions on ARMv8 (background on the implementation: [digoal's commit walkthrough](https://github.com/digoal/blog/blob/master/202604/20260406_06.md); WAL's use of CRC-32C is documented in [Section 28.1, Reliability](https://www.postgresql.org/docs/19/wal-reliability.html)). On modern CPUs the steady-state CPU overhead of keeping checksums enabled is small; the dominant cost is the one-time rewrite when enabling, not day-to-day operation. ## Interaction with backup verification [#interaction-with-backup-verification] Checksums compose with the backup toolchain at two different layers: * [`pg_basebackup`](https://www.postgresql.org/docs/19/app-pgbasebackup.html) verifies page checksums while reading the cluster (unless `--no-verify-checksums` is given); a checksum failure produces a non-zero exit status and is counted in `pg_stat_database.checksum_failures`. Enabling checksums therefore makes every base backup a full-cluster corruption scan for free. * [`pg_verifybackup`](https://www.postgresql.org/docs/19/app-pgverifybackup.html) validates a finished backup against the `backup_manifest` file-level hashes. That catches corruption introduced in backup storage or transfer — a different failure domain from page checksums, which guard the live cluster. Both belong in the verification routine described in [backup, recovery, and PITR](/en/docs/operations/backup-recovery); surface `checksum_failures` in the dashboards from [monitoring and logging](/en/docs/operations/monitoring-logging). ## When it is worth enabling [#when-it-is-worth-enabling] * Clusters initialized before PostgreSQL 18's default change that never had checksums on — now there is no downtime excuse. * Compliance or integrity requirements that demand detection of silent storage corruption. * Any system where `pg_basebackup`'s built-in verification doubles as a periodic corruption sweep. Reasons to hold off are narrow: throwaway clusters, or storage stacks that already checksum end-to-end (e.g. ZFS) combined with a high risk tolerance. Note that only PostgreSQL-level checksums are verified by PostgreSQL's own tools. For more on PostgreSQL 19, see the [release overview](/en/docs/postgresql-19). ## AI prompt: plan a checksum rollout [#ai-prompt-plan-a-checksum-rollout] --- # Production operations Canonical URL: https://pg.edu.rich/en/docs/operations Last reviewed: 2026-08-02 ## Define objectives first [#define-objectives-first] | Objective | Question | | ------------ | ------------------------------------------------------------ | | RPO | How much data can be lost? | | RTO | How quickly must service return? | | Capacity | What are peak connections and data/WAL/backup growth? | | Availability | Which failures auto-fail over, and which require judgment? | | Security | Who can connect, read which data, and perform which changes? | “High availability” and “we have backups” are not verifiable without targets. ## Minimal daily view [#minimal-daily-view] ```sql SELECT now(), version(); SELECT state, count(*) FROM pg_stat_activity GROUP BY state ORDER BY state; SELECT datname, age(datfrozenxid) FROM pg_database ORDER BY age(datfrozenxid) DESC; SELECT num_timed, num_requested, num_done, buffers_written, write_time, sync_time FROM pg_stat_checkpointer; SELECT buffers_clean, maxwritten_clean, buffers_alloc FROM pg_stat_bgwriter; ``` Since PostgreSQL 17, checkpoint statistics live in `pg_stat_checkpointer`; background-writer statistics remain in `pg_stat_bgwriter`. See the [PostgreSQL 18 cumulative statistics views](https://www.postgresql.org/docs/18/monitoring-stats.html#MONITORING-PG-STAT-CHECKPOINTER-VIEW) for the field definitions. Also monitor disk, WAL generation/archive, replication lag, backup state, transaction age, lock waits, query latency, autovacuum, and pool saturation. Thresholds come from this system's baseline. ## Change discipline [#change-discipline] 1. Measure lock and duration on representative data. 2. Document rollback and the point of irreversibility. 3. Set `lock_timeout` so a migration does not wait indefinitely then acquire a disruptive lock. 4. Observe locks, WAL, replication lag, and errors during execution. 5. Verify with queries and business signals. Production DDL is not done when the command succeeds. It is an observable, interruptible, verified release. --- # Instant PostgreSQL clones with copy-on-write Canonical URL: https://pg.edu.rich/en/docs/operations/instant-clone Last reviewed: 2026-08-06 Copying a 200 GB database used to mean copying 200 GB. Copy-on-write (CoW) changes the arithmetic: the filesystem creates a second directory tree that shares the same physical extents, and only pages written later by either side consume new space. The clone appears in seconds and starts at near-zero extra disk. PostgreSQL 18 makes this a first-class operation for single databases, and any reflink-capable filesystem makes it possible for whole instances. ## Why instant clones [#why-instant-clones] * **One database per pull request.** CI runs migrations and integration tests against a full-size copy of real data, then throws it away. * **One sandbox per agent task.** An agent can mutate, drop, and retry without touching shared state — a natural fit with agent workflows driven over [MCP](/en/docs/ai/mcp). * **Migration and upgrade drills.** Rehearse the schema migration against the production shape, measure duration, then delete the rehearsal copy. * **Incident debugging.** Give engineers a frozen copy of the broken state instead of poking at production. The common requirement is that the copy be cheap enough to be disposable. `pg_dump` + restore fails that bar at size: it is minutes to hours, doubles storage for the duration, and rebuilds every index. ## The PostgreSQL 18 building blocks [#the-postgresql-18-building-blocks] CoW copies rely on **reflinks**: a file copy where source and destination share extents until either is written. Filesystem support is the hard prerequisite — XFS formatted with `reflink=1` (the default on modern distributions), Btrfs, and APFS all support it; OpenZFS added block cloning in 2.3. A one-line check tells you whether your mount qualifies: ```bash cp --reflink=always /srv/pg/somefile /tmp/reflink-test && rm /tmp/reflink-test # fails with "Operation not supported" on filesystems without reflinks ``` On top of that, PostgreSQL 18 adds the server setting [`file_copy_method`](https://www.postgresql.org/docs/18/runtime-config-resource.html#GUC-FILE-COPY-METHOD) (`copy` by default, `clone` to opt in). With `clone`, the file copies made by [`CREATE DATABASE ... STRATEGY = FILE_COPY`](https://www.postgresql.org/docs/18/sql-createdatabase.html) and by `ALTER DATABASE ... SET TABLESPACE` use `copy_file_range()` on Linux or `copyfile()` on macOS, which become reflinks on capable filesystems and raise an error where cloning is unsupported. Note the boundary: in PostgreSQL 18, `initdb` and `pg_basebackup` still perform plain block-by-block copies — there is no reflink option for them. Instance-level CoW comes from the filesystem (a reflink copy or a snapshot), not from a PostgreSQL flag. ## Clone a whole instance with reflinks [#clone-a-whole-instance-with-reflinks] The consistent path: stop the source cleanly, copy with reflinks, bring the copy up as a separate instance. ```bash # 1. clean stop — fast mode finishes a checkpoint and waits for clients pg_ctl -D /srv/pg/18/main stop -m fast # 2. CoW copy of the entire data directory — seconds, near-zero extra space cp -a --reflink=always /srv/pg/18/main /srv/pg/18/clone-pr482 # 3. start the source again pg_ctl -D /srv/pg/18/main start # 4. adjust what must differ, then start the clone as its own instance echo 'port = 55432' >> /srv/pg/18/clone-pr482/postgresql.conf pg_ctl -D /srv/pg/18/clone-pr482 start ``` What must differ between source and clone: `port` (and any `listen_addresses`/socket directory assumptions), `data_directory` if it is set explicitly, and anything in `postgresql.auto.conf` that pins resources per instance. If the cluster uses additional tablespaces, those directories live outside the data directory and must be reflink-copied too, with the `pg_tblspc` symlinks repointed. If stopping the source is not acceptable, use an **atomic filesystem snapshot** instead of a plain copy — a Btrfs/ZFS/LVM snapshot taken in one instant yields a crash-consistent image, and PostgreSQL replays WAL on first start exactly as after a power failure. A plain `cp` (reflink or not) of a *running* data directory is neither atomic nor consistent: files change while the copy walks the tree, and the result may refuse to start or, worse, start with subtle corruption. Do not treat that as a shortcut. ## Clone one database with CREATE DATABASE [#clone-one-database-with-create-database] Inside one instance, [`CREATE DATABASE ... TEMPLATE`](https://www.postgresql.org/docs/18/sql-createdatabase.html) copies a database at the file level. The `STRATEGY` option (PostgreSQL 15+) plus PostgreSQL 18's `file_copy_method = clone` turns that into a CoW clone: ```sql SET file_copy_method = 'clone'; CREATE DATABASE app_pr482 TEMPLATE app STRATEGY = FILE_COPY; ``` Constraints, per the official documentation: * No other session may be connected to the template database; `CREATE DATABASE` fails if one is, and new connections to the template are locked out until the copy finishes. * `FILE_COPY` forces a checkpoint before and after the copy, which can be noticeable on a busy system. * Database-level configuration (`ALTER DATABASE ... SET`) and database-level `GRANT`s are not copied. * The default strategy `WAL_LOG` copies block by block through WAL — slower and full-size, but unaffected by filesystem reflink support. The result is a normal database on the same instance: same roles, same port, isolated schema and data, sharing unmodified extents with the template until either side writes. ## Choosing a method [#choosing-a-method] | Method | Granularity | Consistency requirement | Time and space | | --------------------------------------------------------------------------- | --------------------- | -------------------------------------------------------------- | ---------------------------------------- | | `pg_dump` / `pg_restore` | one database, logical | online; self-consistent snapshot | full copy; slow at size; indexes rebuilt | | `CREATE DATABASE ... TEMPLATE` (default `WAL_LOG`) | one database | no connections to template | full physical copy inside the instance | | `CREATE DATABASE ... STRATEGY = FILE_COPY` + `file_copy_method = clone` | one database | no connections to template; reflink-capable filesystem (PG 18) | near-instant; extents shared CoW | | reflink copy of the data directory | whole cluster | source stopped, or atomic snapshot | near-instant; extents shared CoW | | [`pg_basebackup`](https://www.postgresql.org/docs/18/app-pgbasebackup.html) | whole cluster | online | full copy; no CoW in PG 18 | Reach for `pg_dump` when the clone must move across versions, platforms, or instances — logical copies are portable in ways file-level copies are not. ## Relation to branching platforms [#relation-to-branching-platforms] Managed branching platforms productize exactly this idea. [Neon](/en/docs/cloud/neon) implements branches in its storage layer: a branch is a copy-on-write fork of the data at a point in time, created via API in seconds and billed for the delta. The self-hosted techniques on this page give you the same primitive without the platform — you trade the API and the billing model for a few shell commands and the operational boundaries listed below. ## Risk checklist [#risk-checklist] * **A clone is not a backup.** Source and clone share physical extents; a storage-level corruption hits every clone at once, and deleting the source frees nothing while clones exist. Real, independent backups remain mandatory — see [Backup, recovery, and PITR](/en/docs/operations/backup-recovery). * **WAL consistency.** Only a clean shutdown copy or an atomic snapshot produces a safe clone. A live plain copy can yield a directory that fails crash recovery or starts with torn state. * **Disk accounting lies.** `df` and PostgreSQL's size functions do not reflect shared extents; capacity alerts calibrated on full copies will misreport. Monitor the filesystem's actual allocation. * **Same filesystem only.** Reflinks cannot cross mounts or hosts — the clone must live on the same filesystem as the source. * **Sharing erodes with writes.** Every first write to a shared extent allocates new blocks on that side. Long-lived, heavily written clones converge back toward full size; design CI lifetimes accordingly. * **Template lock-out.** Per-database clones freeze connections to the template database for the duration; on CoW filesystems that duration is seconds, on plain copies it scales with size. Copy-on-write makes clones cheap enough to be disposable — treat them as ephemeral by design. Anything you cannot afford to lose must exist as an independent copy on different storage, not as one more reflink of the same extents. --- # PostgreSQL monitoring and logs Canonical URL: https://pg.edu.rich/en/docs/operations/monitoring-logging Last reviewed: 2026-08-06 PostgreSQL observability needs at least three evidence layers: **query statistics show where resources go, metrics show when the system leaves its baseline, and logs preserve error and event context**. A single dashboard does not replace these layers. ## 1. Find workload hotspots with pg\_stat\_statements [#1-find-workload-hotspots-with-pg_stat_statements] `pg_stat_statements` is an official PostgreSQL extension. Add it to `shared_preload_libraries`, which normally requires a restart, then create it in each database that needs statistics: ```ini shared_preload_libraries = 'pg_stat_statements' compute_query_id = auto ``` ```sql CREATE EXTENSION IF NOT EXISTS pg_stat_statements; SELECT queryid, calls, total_exec_time, mean_exec_time, rows, shared_blks_hit, shared_blks_read, left(query, 160) AS query FROM pg_stat_statements ORDER BY total_exec_time DESC LIMIT 20; ``` Rank total time, mean time, calls, rows, and I/O separately. The slowest individual call and the largest cumulative consumer are different problems. Statistics can reset, so record sampling windows around deployments, incidents, and configuration changes. Use the [official pg\_stat\_statements documentation](https://www.postgresql.org/docs/current/pgstatstatements.html) for field semantics. Constants are normalized, but logs, DDL, dynamic SQL, and application comments can still expose identifiers or business data. Restrict access to statistics and logs, and define collection, retention, and redaction rules. ## 2. Preserve event context with JSON logs [#2-preserve-event-context-with-json-logs] `jsonlog` makes timestamps, SQLSTATE, backend, database, user, application name, and error context reliably parseable: ```ini logging_collector = on log_destination = 'jsonlog' log_min_duration_statement = '500ms' # Example only; derive from workload baseline log_lock_waits = on deadlock_timeout = '1s' ``` Do not copy one threshold into every environment. Too low creates excessive I/O and sensitive query text; too high misses frequent medium-latency queries. Use `pg_stat_statements` for cumulative hotspots and logs for errors, lock waits, checkpoints, autovacuum, and specific slow requests. See [Error Reporting and Logging](https://www.postgresql.org/docs/current/runtime-config-logging.html). [pgBadger](https://github.com/darold/pgbadger) can process PostgreSQL native logs and `jsonlog`, producing query, connection, error, lock, checkpoint, and autovacuum reports. Stabilize format, rotation, and time zones before adding it to offline analysis. ## 3. Metrics, Prometheus, and Grafana [#3-metrics-prometheus-and-grafana] [postgres\_exporter](https://github.com/prometheus-community/postgres_exporter) fits teams already using Prometheus and Grafana. Prefer `pg_monitor` or the minimum read-only statistics privileges over superuser access: ```sql CREATE ROLE metrics LOGIN; GRANT pg_monitor TO metrics; ``` Verify collectors and privileges against the target version. Upstream still labels multi-target mode Beta, and custom `extend.query-path` queries are deprecated. Prefer built-in collectors or a separate generic SQL exporter for new custom collection rather than growing an unmaintainable query file. ### Minimum signal set [#minimum-signal-set] | Domain | Signal | Context to correlate | | --------------- | ------------------------------------------------- | ----------------------------------------------------- | | Connections | Usage, waits, pool queue | Pool mode, application replicas, reserved connections | | Queries | Latency, calls, rows, I/O | Deployments, plan changes, parameter distribution | | Transactions | Long transactions, idle in transaction, conflicts | Owner, retryability, vacuum impact | | Locks | Wait duration, blocking chain, deadlocks | DDL, batch jobs, business transactions | | WAL/replication | Generation, archive failure, lag, slot retention | RPO, network, free disk | | Maintenance | Dead tuples, freeze age, vacuum/analyze progress | Write rate and autovacuum settings | | Storage | Data/WAL/temp growth and I/O latency | Capacity forecast, checkpoints, query spills | | Recovery | Latest backup, recoverable time, measured RTO | Repository, keys, restore drills | Derive thresholds from normal and peak baselines, and make each alert lead to an actionable diagnostic path. Replication lag in bytes, time, and replay state has different meanings; one global threshold is insufficient. ### Alert threshold starting points [#alert-threshold-starting-points] Calibrate every threshold against your own baseline; these are common starting points, not universal values: | Metric | Warning | Critical | | ------------------------------------------- | --------- | ----------- | | Connections used (% of `max_connections`) | 70% | 90% | | Replication lag (seconds) | 10 | 60 | | Replication lag (bytes) | 1 GB | 10 GB | | Database transaction ID age | 1 billion | 1.5 billion | | Disk used | 75% | 90% | | Deadlocks per second | 0.1 | 1 | | Autovacuum on one table running longer than | 2 h | 6 h | | Dead tuples (% of table) | 20% | 40% | | Buffer cache hit ratio below | 95% | 90% | A warning should not page anyone and a critical should have a runbook; delete or tune any alert that fires without an action. ## 4. Diagnostic SQL templates [#4-diagnostic-sql-templates] Keep these close during an incident; they are read-only and safe to run on production. ```sql -- Connections in use, by state SELECT count(*) AS total, count(*) FILTER (WHERE state = 'active') AS active, count(*) FILTER (WHERE state = 'idle in transaction') AS idle_in_transaction FROM pg_stat_activity; -- Queries running longer than 5 seconds SELECT pid, usename, application_name, now() - query_start AS duration, wait_event, left(query, 200) AS query FROM pg_stat_activity WHERE state = 'active' AND now() - query_start > interval '5 seconds' ORDER BY query_start; -- Blocking chain: who is waiting on whom SELECT blocked.pid AS blocked_pid, blocking.pid AS blocking_pid, left(blocked.query, 120) AS blocked_query, left(blocking.query, 120) AS blocking_query FROM pg_stat_activity blocked JOIN pg_stat_activity blocking ON blocking.pid = ANY (pg_blocking_pids(blocked.pid)); -- Replication lag, seen from the primary SELECT application_name, state, sync_state, pg_wal_lsn_diff(pg_current_wal_lsn(), replay_lsn) AS lag_bytes, replay_lag FROM pg_stat_replication; -- Replication slot retention (unbounded growth fills the primary's disk) SELECT slot_name, active, pg_wal_lsn_diff(pg_current_wal_lsn(), restart_lsn) AS retained_bytes FROM pg_replication_slots; ``` `SELECT pg_cancel_backend(pid);` cancels the running query; `SELECT pg_terminate_backend(pid);` closes the connection. Prefer cancel, and confirm the session is disposable first. ## Adoption order [#adoption-order] 1. Enable and govern `pg_stat_statements` on every production instance. 2. Emit parseable logs and collect SQLSTATE, lock waits, archive failures, and autovacuum events. 3. Add postgres\_exporter and PostgreSQL-specific dashboards when Prometheus already exists. 4. Add pgBadger for log trends; evaluate [pgwatch](https://github.com/cybertec-postgresql/pgwatch) for multiple instances. 5. Evaluate [PoWA](https://github.com/powa-team/powa) for deeper workload analysis and [pg\_activity](https://github.com/dalibo/pg_activity) for interactive incident diagnosis. More tools do not automatically remove blind spots. First standardize the dimensions that connect evidence: instance, database, role, application, query ID, time window, and change event. --- # OS upgrades and silent index corruption Canonical URL: https://pg.edu.rich/en/docs/operations/os-upgrade-index-corruption Last reviewed: 2026-08-06 A B-tree index on a `text` column stores keys in the sort order defined by the collation that was active when each key was inserted. For libc collations, that order comes from the operating system's C library. An OS upgrade that ships a new glibc (or a new ICU library) can change the rules — and PostgreSQL does not revalidate existing indexes against the new rules. Queries keep using the affected indexes and can silently return wrong results: rows missing from range scans and prefix searches, incorrect `ORDER BY ... LIMIT` output, unique checks that fail to spot duplicates. Nothing raises an error. ## How an OS upgrade corrupts indexes [#how-an-os-upgrade-corrupts-indexes] String comparison in a libc collation goes through the operating system (`strcoll_l` and friends). When glibc changes the sort order of a locale, keys inserted after the upgrade are placed according to the new rules while keys inserted before it still sit in old-rule positions. The index remains structurally valid — page links, checksums, and tuple formats are all fine — but it is no longer sorted under one consistent order. Any index scan that relies on ordering can stop early, skip entries, or return them in the wrong sequence. The best-known trigger is glibc 2.28 (2018), which aligned many locales with a new common sorting template. Upgrades that cross it — for example RHEL/CentOS 7 → 8 or Debian 9 → 10 — are the classic breakage scenario, but any distro upgrade can ship changed locale data. ICU collations have the same exposure to ICU library upgrades, independent of glibc. ## Which indexes are at risk [#which-indexes-are-at-risk] * B-tree indexes on `text`, `varchar`, `char`, and domains over them, when the column collation is provided by **libc** and is not `C`/`POSIX` — including indexes that use the database **default** collation when the database itself uses the libc provider. * Unique constraints and primary keys on those columns, since they are backed by such indexes. * ICU-provider collations are affected by ICU library version changes, not by glibc. Not affected: `C` and `POSIX` collations (byte-order comparison, stable everywhere), the builtin provider locales such as `C.UTF-8` (immutable by design, PostgreSQL 17+), hash indexes (no ordering), and indexes on non-collatable types such as `integer`, `bigint`, `uuid`, or `timestamptz`. See [Collation Support](https://www.postgresql.org/docs/18/collation.html) in the PostgreSQL 18 documentation for provider semantics. ## Inventory before the upgrade [#inventory-before-the-upgrade] Before any OS upgrade, list the indexes that depend on libc ordering and save the output — it is your diff baseline afterwards: ```sql SELECT DISTINCT i.indexrelid::regclass AS index_name, i.indrelid::regclass AS table_name, c.collname AS collation, c.collprovider AS provider -- c = libc, d = default, i = icu, b = builtin FROM pg_index i JOIN LATERAL unnest(i.indcollation) AS u(coll_oid) ON true JOIN pg_collation c ON c.oid = u.coll_oid WHERE c.collprovider IN ('c', 'd') AND c.collname NOT IN ('C', 'POSIX') ORDER BY 2, 1; ``` The `default` collation (provider `d`) inherits the database locale, so check what each database actually uses: ```sql SELECT datname, datcollate, datlocprovider, datcollversion FROM pg_database; ``` If `datlocprovider` is `c` and `datcollate` is not `C`/`POSIX`, every default-collation text index in that database depends on libc. The query above covers index key columns; expression indexes can embed additional collations in their expressions and need a manual review of their definitions. ## Detecting damage after the upgrade [#detecting-damage-after-the-upgrade] ### Compare recorded collation versions [#compare-recorded-collation-versions] Since PostgreSQL 10, the catalog records the provider version of each collation in [`pg_collation.collversion`](https://www.postgresql.org/docs/18/catalog-pg-collation.html); PostgreSQL 15 extended this to the database default collation (`pg_database.datcollversion`, together with `ALTER DATABASE ... REFRESH COLLATION VERSION`). When an object is used whose recorded version no longer matches what the OS reports, the session emits a warning with a version-mismatch notice once per collation. You can compare directly without waiting for warnings: ```sql SELECT collname, collprovider, collversion, pg_collation_actual_version(oid) AS os_version FROM pg_collation WHERE collversion IS DISTINCT FROM pg_collation_actual_version(oid); ``` [`pg_collation_actual_version`](https://www.postgresql.org/docs/18/functions-info.html) asks the operating system for the currently installed version. Rows returned by this query are collations whose behavior may have changed under your feet. On PostgreSQL 9.6 and older the version-tracking infrastructure does not exist — there is no warning and nothing to compare, so detection rests entirely on the structural check below or on planned reindexing. ### Verify index structure with amcheck [#verify-index-structure-with-amcheck] The [`amcheck`](https://www.postgresql.org/docs/18/amcheck.html) extension re-evaluates B-tree ordering using the *current* comparison rules. That makes it a direct detector for this failure mode: an index built under old rules can be internally consistent yet fail verification after the upgrade, because verification now expects the new order. ```sql CREATE EXTENSION IF NOT EXISTS amcheck; SELECT bt_index_parent_check('app.orders_customer_name_idx', heapallindexed => true); ``` The function returns no rows when the index is clean and raises an error on the first inconsistency. `bt_index_parent_check` is the stricter variant (it also checks parent/child page relationships); `bt_index_check` is lighter. `bt_index_parent_check` takes a `ShareLock` on the index and its table, blocking concurrent `INSERT`/`UPDATE`/`DELETE`, so run it in a maintenance window; `bt_index_check` takes only an `AccessShareLock` — the same lock a plain `SELECT` takes — and does not block writes. Two limitations to keep in mind: amcheck only reports when stored key order conflicts with the current rules — a collation change that happens not to reorder any existing key passes silently, so a clean run is not proof of safety. Conversely, after an OS upgrade that touched glibc or ICU, an amcheck error on a text index almost certainly means "rebuild", not "hardware fault". ## Rebuilding affected indexes [#rebuilding-affected-indexes] Rebuild without blocking the application: ```sql REINDEX INDEX CONCURRENTLY app.orders_customer_name_idx; -- or every index on a table at once: REINDEX TABLE CONCURRENTLY app.orders; ``` `REINDEX CONCURRENTLY` keeps the table readable and writable, takes longer than a plain `REINDEX`, and cannot run inside a transaction block. A failed or interrupted run can leave an `INVALID` index behind — find it in `pg_index` (`NOT indisvalid`) and drop it before retrying. For an instance-wide incident, rebuild per affected table ordered by index size so the largest windows are scheduled deliberately. After rebuilding, refresh the recorded versions to clear the mismatch warnings: ```sql ALTER COLLATION "de_DE" REFRESH VERSION; ALTER DATABASE app REFRESH COLLATION VERSION; ``` Only refresh versions *after* the dependent indexes have been rebuilt — refreshing first silences the warning while the corruption is still in place. No error is raised, checksums do not fire, and standard monitoring sees nothing. The symptom is users reporting "missing" rows days or weeks after an OS upgrade. Put the inventory query and an amcheck pass into the OS upgrade runbook, not into post-incident analysis. ## Prevention [#prevention] * **Prefer ICU or builtin collations for new databases.** `CREATE DATABASE ... LOCALE_PROVIDER = icu ICU_LOCALE = 'de-DE'` pins ordering to an explicitly versioned ICU rule set instead of whatever glibc the distro ships; where natural-language ordering is not needed, the builtin provider's `C.UTF-8` (PostgreSQL 17+) is immutable across OS upgrades. Existing libc-based databases can move individual columns to ICU collations with `CREATE COLLATION` plus a concurrent index rebuild. * **Separate OS upgrades from PostgreSQL major upgrades.** `pg_upgrade` copies data files without rebuilding indexes, so combining both upgrades in one window doubles the exposure and makes anomalies hard to attribute. Do them in separate windows, with the version comparison and an amcheck pass in between; upgrade paths are covered in [Replication, failover, and upgrades](/en/docs/operations/replication-upgrades). * **Recheck after every glibc/ICU bump.** The version-comparison query is cheap; run it after each OS patch cycle and reindex whatever changed. * **Know what your queries rely on.** The damage surfaces through index scans and ordered output — [Indexes and EXPLAIN](/en/docs/core/indexes-explain) covers how to see which plans depend on the affected indexes. --- # Parallel autovacuum and score-based scheduling Canonical URL: https://pg.edu.rich/en/docs/operations/parallel-autovacuum Last reviewed: 2026-08-06 As of 2026-08, PostgreSQL 19 has not reached GA. Parameter names, view columns, and defaults below follow the [PostgreSQL 19 documentation](https://www.postgresql.org/docs/19/runtime-config-autovacuum.html); re-check the final release notes before relying on them in production. ## Before PostgreSQL 19: serial autovacuum [#before-postgresql-19-serial-autovacuum] Two long-standing limitations shaped autovacuum operations through PostgreSQL 18: * An autovacuum worker processes one table at a time, and the *vacuuming indexes* and *cleaning up indexes* phases walk the table's indexes serially. A wide table with a dozen indexes could hold a worker for hours while other tables waited. * Within a database, the worker processed candidate tables in roughly `pg_class` catalog order. A table approaching transaction ID wraparound had no formal priority over one that barely crossed an analyze threshold. Manual `VACUUM` has supported `PARALLEL` for index work since PostgreSQL 13, but autovacuum could not use it. Closing the gap for a big table meant running manual vacuums by hand — see [autovacuum and table bloat](/en/docs/operations/autovacuum-bloat) for the baseline mechanics. ## Parallel index processing: autovacuum\_max\_parallel\_workers [#parallel-index-processing-autovacuum_max_parallel_workers] [`autovacuum_max_parallel_workers`](https://www.postgresql.org/docs/19/runtime-config-autovacuum.html) sets the maximum number of parallel workers a single autovacuum worker may recruit to process indexes during the index vacuuming and index cleanup phases. It defaults to `0` (disabled), so it is opt-in: ```sql ALTER SYSTEM SET autovacuum_max_parallel_workers = 4; SELECT pg_reload_conf(); ``` It is the autovacuum equivalent of the `PARALLEL` option of manual `VACUUM`. The actual worker count is further limited by `max_parallel_workers`, which is shared with parallel query. A per-table storage parameter caps individual tables — useful for keeping one hot table from consuming the whole parallel budget: ```sql ALTER TABLE app.events SET (autovacuum_parallel_workers = 2); ``` Constraints, per [Section 24.1.7, Parallel Vacuum](https://www.postgresql.org/docs/19/routine-vacuuming.html): * An index participates only if it is larger than `min_parallel_index_scan_size`, and each index gets at most one worker — so a table needs at least two eligible indexes for parallel workers to launch at all. * The calculated number of workers is not guaranteed; a vacuum may run with fewer workers or none. * Parallel workers use the same cost delay parameters as the leader autovacuum worker, so the existing throttling model still applies. ## Score-based scheduling [#score-based-scheduling] Within a database, the autovacuum worker now builds its candidate list and sorts it by score instead of catalog order. The score of a table is the maximum of five component scores, described in [Section 24.1.6.1, Autovacuum Prioritization](https://www.postgresql.org/docs/19/routine-vacuuming.html): * **Transaction ID age** — `age(relfrozenxid)` against `autovacuum_freeze_max_age`; it grows sharply once the age passes `vacuum_failsafe_age`. Weight: `autovacuum_freeze_score_weight`. * **Multixact ID age** — `relminmxid` against `autovacuum_multixact_freeze_max_age`; also grows sharply past `vacuum_multixact_failsafe_age` or when multixact members exceed roughly 2 billion entries. Weight: `autovacuum_multixact_freeze_score_weight`. * **Vacuum** — updated/deleted tuples against the vacuum threshold. Weight: `autovacuum_vacuum_score_weight`. * **Vacuum insert** — inserted tuples against the insert threshold. Weight: `autovacuum_vacuum_insert_score_weight`. * **Analyze** — changed tuples against the analyze threshold. Weight: `autovacuum_analyze_score_weight`. All five weights default to `1.0` (equal treatment) and are reloadable with `pg_reload_conf()`. Two subtleties from the documentation: * Raising a freeze weight above 1.0 does more than multiply the score — the age at which the component starts scaling aggressively is *divided* by the weight, so freeze pressure becomes urgent earlier. * Setting all five weights to `0.0` reverts to the pre-19 strategy of plain catalog order. Database selection is separate: the launcher still prioritizes databases at risk of wraparound, then the least recently processed one. ## Observing scores with pg\_stat\_autovacuum\_scores [#observing-scores-with-pg_stat_autovacuum_scores] The new `pg_stat_autovacuum_scores` view shows the current scores for every table in the current database, turning "why is autovacuum ignoring this table" from guesswork into a query: ```sql SELECT relid::regclass AS relation, round(score::numeric, 1) AS score, round(xid_score::numeric, 1) AS xid_score, round(vacuum_score::numeric, 1) AS vacuum_score, round(vacuum_insert_score::numeric, 1) AS insert_score, round(analyze_score::numeric, 1) AS analyze_score, do_vacuum, do_analyze, for_wraparound FROM pg_stat_autovacuum_scores ORDER BY score DESC LIMIT 20; ``` `score` is the maximum of the five `*_score` components; `do_vacuum` / `do_analyze` show what the table currently qualifies for, and `for_wraparound` marks anti-wraparound pressure. One caveat from the documentation: the view computes scores from the information visible to your session, which can differ from what an autovacuum worker sees when it builds its list — so it is a debugging aid, not a guarantee of processing order. Use it before and after reweighting to confirm the ordering actually changed — for example, making dead-tuple reclamation outrank statistics refresh: ```sql ALTER SYSTEM SET autovacuum_vacuum_score_weight = 2.0; ALTER SYSTEM SET autovacuum_analyze_score_weight = 0.5; SELECT pg_reload_conf(); ``` Add the view to the regular inspection routine described in [monitoring and logging](/en/docs/operations/monitoring-logging), alongside `pg_stat_progress_vacuum`. ## Compared with vacuumdb --jobs on PostgreSQL 18 [#compared-with-vacuumdb---jobs-on-postgresql-18] The classic PostgreSQL 18 workaround for slow maintenance is manual parallelism: ```bash vacuumdb --jobs=4 --analyze dbname ``` `vacuumdb --jobs` runs several connections, each processing a different table — parallelism *across* tables, on a schedule you own, during windows you pick. It does not speed up the index phases of a single large table unless you run `VACUUM (PARALLEL n)` yourself, and it does nothing about processing order inside autovacuum. PostgreSQL 19 covers the complementary axis: parallelism *within* one table's index phases, and urgency-aware ordering, both automatic. Scheduled `vacuumdb` runs remain useful for predictable batch windows and full-cluster `FREEZE` passes; the two are not mutually exclusive. ## When not to enable it [#when-not-to-enable-it] * **I/O-bound systems.** Parallel index vacuuming multiplies concurrent I/O streams. The cost limit is still shared, but latency-sensitive workloads on saturated storage should enable it gradually and watch `pg_stat_io` and query latency. * **Small databases.** If all indexes are under `min_parallel_index_scan_size`, no index ever participates — the setting changes nothing except expectations. * **Tight worker budgets.** Vacuum parallel workers come out of `max_parallel_workers`, the same pool parallel query uses. On small instances, raising both can starve query parallelism. * **Mostly single-index tables.** One worker per index means a table with one eligible index never goes parallel. As with any autovacuum tuning, change one knob at a time and verify with `pg_stat_autovacuum_scores` and the autovacuum log rather than assuming. For more on PostgreSQL 19, see the [release overview](/en/docs/postgresql-19). ## AI prompt: tune autovacuum scoring [#ai-prompt-tune-autovacuum-scoring] --- # PostgreSQL production stack and HA Canonical URL: https://pg.edu.rich/en/docs/operations/production-stack Last reviewed: 2026-08-02 A useful production order is: **prove recovery → protect connection capacity → make failures observable → make changes safe → then automate failover**. High availability does not replace backup, and a replica cannot undo a deletion already replicated to it. As of **2026-08-02**, PostgreSQL 18.4 is the current minor of the newest stable major, while majors 14–18 remain supported. Production systems should run the current minor of their chosen major; being newest does not by itself justify a major upgrade. See the [PostgreSQL versioning policy](https://www.postgresql.org/support/versioning/) and [18.4 release notes](https://www.postgresql.org/docs/current/release-18-4.html). ## Minimum production baseline [#minimum-production-baseline] | Layer | First question | Common choice | | ------------- | --------------------------------------------------------------------------- | ----------------------------------------------------------------- | | Database | How are minor updates, roles, TLS, settings, and extensions controlled? | Official PostgreSQL packages or a verified image | | Connections | Can peak application concurrency exhaust backends? | Application pooling, then PgBouncer where needed | | Recovery | What are the RPO/RTO, and can recovery happen outside the platform? | pgBackRest, WAL-G, or cloud backup plus an independent copy | | Observability | Which query, wait event, metric, and log explains an incident? | `pg_stat_statements`, JSON logs, and metric collection | | Changes | How are DDL locks, backfills, rollback, and compatibility windows tested? | Expand-and-contract, migration linting, and real PostgreSQL tests | | Availability | What are the independent failure domains and post-failover data boundaries? | Managed HA, Patroni, CloudNativePG, or Pigsty | ## PgBouncer: do not assume transaction pooling [#pgbouncer-do-not-assume-transaction-pooling] [PgBouncer](https://www.pgbouncer.org/) multiplexes many client connections onto fewer PostgreSQL server connections. Repeatedly raising `max_connections` is not a capacity plan: every backend consumes memory and scheduling resources, and operations need reserved connections for migrations, monitoring, backup, administration, and replication. | Mode | Server connection returns after | Boundary | | ------------- | ------------------------------- | ----------------------------------------------------------------------------- | | `session` | Client disconnect | Highest compatibility; use when applications depend on session state | | `transaction` | Transaction end | Common for Web/API traffic, but session features must be tested | | `statement` | Every statement | Disallows multi-statement transactions; only for tightly controlled workloads | With transaction pooling, SQL-level `PREPARE`, session advisory locks, `LISTEN`, holdable cursors, and many session-state patterns are unsupported or constrained. Protocol-level prepared statements require an appropriate `max_prepared_statements` configuration and driver testing. Use the [PgBouncer feature map](https://www.pgbouncer.org/features.html) as the compatibility source. Before rollout, test authentication and TLS, prepared statements, temporary tables, migration tooling, connection reset, failover, transaction retries, and any ORM state that may be session-scoped. ## pgBackRest, WAL-G, and cloud backup [#pgbackrest-wal-g-and-cloud-backup] [pgBackRest](https://github.com/pgbackrest/pgbackrest) supports full, differential, and incremental backups, parallel transfer, multiple repositories, WAL archiving, and PITR. It is a strong fit for a complete self-hosted recovery chain. [WAL-G](https://github.com/wal-g/wal-g) is oriented toward object-storage workflows. Neither proves recoverability merely by being installed. Derive schedules from RPO, WAL volume, restore bandwidth, retention requirements, and measured RTO rather than copying a generic weekly/daily calendar. Continuously check archive gaps, deletion protection, encryption keys, cross-host or cross-account copies, and restore into a disposable instance. `pg_dump` remains valuable for logical migration, object selection, and small restores, but cannot provide continuous point-in-time recovery on its own. See [Backup, restore, and PITR](/en/docs/operations/backup-recovery). ## Choosing Patroni, CloudNativePG, or Pigsty [#choosing-patroni-cloudnativepg-or-pigsty] | Environment | Candidate | Adoption requirement | | ---------------------------------- | --------------------------------------------- | ------------------------------------------------------------------------------------------ | | Managed cloud database | Provider multi-zone HA | Verify region, failover, PITR, extensions, and off-platform restore limits | | Independent VMs or bare metal | [Patroni](https://github.com/patroni/patroni) | Multiple failure domains, a reliable DCS, and network/storage operations skills | | Existing Kubernetes platform | [CloudNativePG](https://cloudnative-pg.io/) | A team already capable of operating Kubernetes, storage, networking, and operator upgrades | | Multi-cluster self-hosted platform | [Pigsty](https://github.com/pgsty/pigsty) | Acceptance of its Ansible/VM model and verification of integrated component upgrades | For object-storage backup with CloudNativePG, evaluate the current [Barman Cloud Plugin](https://cloudnative-pg.io/plugin-barman-cloud/docs/intro/) rather than copying deprecated in-tree object-store configuration. Do not introduce Kubernetes solely for one PostgreSQL instance. They share host power, kernel, storage, and networking failures. Automated election improves availability only when node, storage, and quorum boundaries are genuinely independent. ## Adopt in phases [#adopt-in-phases] ### Production baseline [#production-baseline] * Current minor, role separation, TLS, and controlled extensions; * a connection budget and a tested PgBouncer deployment where required; * independently retained backups, continuous WAL/PITR, and restore drills; * query, metric, and log observability; * pre-migration lock tests, timeouts, rollback, and business validation. ### Conditional additions [#conditional-additions] * Add automatic failover only with independent failure domains and an explicit RTO; * prefer CloudNativePG only when Kubernetes operations already exist; * add Pigsty when managing multiple self-hosted clusters justifies a platform; * install pgvector, PostGIS, TimescaleDB, or other extensions only for a measured workload. ### Experimental lane [#experimental-lane] PostgreSQL 19 Beta, newer extensions, and storage engines belong in disposable compatibility environments, not in the production baseline. Keep separate CI lanes for the stable target and the next major; see [Safe migrations and zero-downtime schema changes](/en/docs/operations/safe-migrations). --- # REPACK and REPACK CONCURRENTLY: online debloating Canonical URL: https://pg.edu.rich/en/docs/operations/repack Last reviewed: 2026-08-06 As of 2026-08, PostgreSQL 19 has not reached GA. The syntax, GUC names, and views below follow the [PostgreSQL 19 documentation](https://www.postgresql.org/docs/19/sql-repack.html); re-check the final release notes before relying on them in production. Plain `VACUUM` makes dead space reusable inside a relation but almost never returns it to the operating system. The two classic ways to actually shrink a bloated table — `VACUUM FULL` and `CLUSTER` — both rewrite the table and its indexes under an `ACCESS EXCLUSIVE` lock held for the entire rewrite, which is why teams either schedule maintenance windows or install the third-party `pg_repack` extension. PostgreSQL 19 adds a built-in third option: [`REPACK`](https://www.postgresql.org/docs/19/sql-repack.html), with a `CONCURRENTLY` mode that keeps the table readable and writable while it is rebuilt. ## The lock problem with VACUUM FULL and CLUSTER [#the-lock-problem-with-vacuum-full-and-cluster] `VACUUM FULL` rewrites the live tuples of a table into a new file and rebuilds every index; `CLUSTER` does the same but sorts rows by an index. Both need `ACCESS EXCLUSIVE` from start to finish, so every read and write on the table queues behind them. On a large table the rewrite takes minutes to hours, and the lock wait alone can pile up enough blocked sessions to take an application down. The mechanics of measuring bloat and choosing a target table are covered in [autovacuum and table bloat](/en/docs/operations/autovacuum-bloat); this page is about the rebuild step. ## REPACK: one command, two semantics [#repack-one-command-two-semantics] `REPACK` folds the `VACUUM FULL` and `CLUSTER` behaviors into a single statement: ```sql REPACK [ ( option [, ...] ) ] [ table_and_columns [ USING INDEX [ index_name ] ] ] REPACK [ ( option [, ...] ) ] USING INDEX ``` Options are `VERBOSE`, `ANALYZE`, and `CONCURRENTLY`: * `REPACK t;` — plain rewrite that reclaims disk, the `VACUUM FULL` equivalent. * `REPACK t USING INDEX i;` — additionally reorders rows by the index, the `CLUSTER` equivalent. Without an index name it uses the index previously set with `ALTER TABLE ... CLUSTER ON`. * `REPACK;` — every table and materialized view in the current database that you hold the `MAINTAIN` privilege on. This form cannot run inside a transaction block and cannot be combined with `CONCURRENTLY`. * `REPACK (ANALYZE) t;` — runs `ANALYZE` after the rewrite. Currently only supported for a single, non-partitioned table; the planner statistics reset after a rewrite makes this worth doing either way. You need the `MAINTAIN` privilege on the table. Without `CONCURRENTLY`, `REPACK` still holds `ACCESS EXCLUSIVE` for the whole operation — the lock semantics are unchanged; only the command surface is unified. ## How CONCURRENTLY works [#how-concurrently-works] With `CONCURRENTLY`, `REPACK` copies the live tuples into a new file (plus a new file for each index) while normal reads and writes continue against the old files. Changes made during the copy are captured through logical decoding and applied to the new files; only then does the command take a brief `ACCESS EXCLUSIVE` lock to swap old and new files and drop the old ones. The lock is typically held just for the swap — but if many changes accumulated, they must be replayed *while the lock is held*, so a hot table can still see a noticeable blocking window at the end. Two behaviors from the documentation are worth internalizing: * Rows inserted after the repack started are not ordered, even with `USING INDEX` — clustering is a one-time physical reorder. * `REPACK CONCURRENTLY` can fail if other sessions run DDL on the table while it works. Keep migrations away from an in-flight repack. The documentation warns that `REPACK` with `CONCURRENTLY` is not MVCC-safe (see [MVCC caveats](https://www.postgresql.org/docs/19/mvcc-caveats.html)). Treat it as a maintenance operation you schedule deliberately, not a background no-op. ## Hard limits [#hard-limits] `CONCURRENTLY` is refused in all of these cases: * the table is `UNLOGGED`; * the table is partitioned (plain `REPACK` on a partitioned table works — it repacks each partition — but not concurrently, and not inside a transaction block); * the table has no primary key and no index-based replica identity — logical decoding needs a way to identify rows; * the table is a system catalog or a TOAST table; * `REPACK` runs inside a transaction block; * `max_repack_replication_slots` has no free slot (see below). Disk is the other hard limit, and it applies with or without `CONCURRENTLY`: the rewrite needs a temporary copy of the table plus every index, so free space of at least table size + index sizes. On the sequential-scan-and-sort path a temporary sort file can push peak usage toward double the table size plus indexes (you can force the index-scan path by setting `enable_sort = off` for the session). `CONCURRENTLY` adds more: changes made during the copy are buffered in a temporary file until they can be applied. Give the session a generous `maintenance_work_mem` before starting. ## max\_repack\_replication\_slots and slot isolation [#max_repack_replication_slots-and-slot-isolation] `CONCURRENTLY` needs a replication slot for logical decoding. PostgreSQL 19 gives `REPACK` its own pool for this: [`max_repack_replication_slots`](https://www.postgresql.org/docs/19/runtime-config-replication.html#GUC-MAX-REPACK-REPLICATION-SLOTS) (default **5**, settable only at server start) reserves slots exclusively for `REPACK`, on top of `max_replication_slots`. The isolation cuts both ways: * a repack job can never consume the slots your logical replication publications and subscriptions depend on; * subscribers can never starve repack jobs either; * but only 5 `REPACK CONCURRENTLY` operations can run at once by default — a sixth fails outright. Operationally: serialize repack jobs (one or two at a time is usually right for I/O reasons anyway), and if you legitimately need more parallelism, raise `max_repack_replication_slots` with a planned restart. Do not try to "fix" a failing repack by raising `max_replication_slots` — that pool is not the one being exhausted. ## Core REPACK vs the pg\_repack extension [#core-repack-vs-the-pg_repack-extension] The built-in command and the [pg\_repack extension](https://pgrepack.github.io/pg_repack/) are two separate implementations that solve the same problem. Choosing between them: | | `REPACK` (PostgreSQL 19 core) | `pg_repack` (extension) | | ------------------------ | ----------------------------------------------------- | -------------------------------------------------------------------- | | Availability | PostgreSQL 19+, nothing to install | Extension + CLI; the production answer for PostgreSQL 18 and earlier | | Online change capture | Logical decoding via a dedicated slot | Triggers writing to a log table, replayed before the swap | | Row identity requirement | Primary key or index-based replica identity | Primary key or suitable unique index | | Lock profile | `ACCESS EXCLUSIVE` only for the final swap | Brief exclusive locks at start and end | | Physical reorder | `USING INDEX` | `--order-by` | | Index-only rebuild | No — use `REINDEX CONCURRENTLY` | `--index` / `--only-indexes` | | Operational shape | SQL statement, `pg_stat_progress_repack` for progress | External CLI with its own release cadence and version-matching rules | On PostgreSQL 19 the core command removes the extension dependency, the version-matching burden, and the trigger-based log table. On anything older, `pg_repack` remains the tool — the core command does not exist there. ## Runbook [#runbook] A deliberate repack, end to end: ```sql -- 1. confirm the table is actually bloated (estimates from pg_stat_user_tables are not enough) CREATE EXTENSION IF NOT EXISTS pgstattuple; SELECT * FROM pgstattuple('app.events'); -- look at dead_tuple_percent and free_percent -- 2. confirm CONCURRENTLY is possible: replica identity must be default-with-PK or an index SELECT relreplident FROM pg_class WHERE oid = 'app.events'::regclass; -- 'd' (default, needs a PK) or 'i' (replica identity index) are OK; 'n' and 'f' are not -- 3. confirm disk: free space >= table + index sizes SELECT pg_size_pretty(pg_total_relation_size('app.events')); ``` ```sql -- 4. run it, in its own session SET maintenance_work_mem = '1GB'; REPACK (CONCURRENTLY, ANALYZE, VERBOSE) app.events; ``` ```sql -- 5. watch progress from another session SELECT pid, datname, relid::regclass AS relation, command, phase, heap_blks_scanned, heap_blks_total, heap_tuples_scanned, heap_tuples_inserted, index_rebuild_count FROM pg_stat_progress_repack; ``` Phases run from `initializing` through the heap scan/copy (`seq scanning heap` / `index scanning heap`, `sorting tuples`, `writing new heap`), then `catch-up` (applying the buffered changes), `swapping relation files` (the brief exclusive lock), `rebuilding index`, and `performing final cleanup`. A long `catch-up` phase means the table is hot enough that the final lock window will be noticeable — consider rerunning in a quieter window. The same view also tracks `CLUSTER` and `VACUUM FULL`, distinguished by the `command` column. If the repack fails partway through, the temporary files and the replication slot are cleaned up by the command; a failed attempt leaves the original table untouched. Retry after fixing the cause — most commonly a conflicting DDL, a missing replica identity, or an exhausted slot pool. ## Prompt: plan a debloating campaign [#prompt-plan-a-debloating-campaign] Command reference: [PostgreSQL 19 REPACK](https://www.postgresql.org/docs/19/sql-repack.html) and the [progress reporting view](https://www.postgresql.org/docs/19/progress-reporting.html). See also [PostgreSQL 19 overview](/en/docs/postgresql-19). Independent field reports: [depesz's walkthrough](https://www.depesz.com/2026/03/19/waiting-for-postgresql-19-introduce-the-repack-command/) and [digoal on slot isolation (Chinese)](https://github.com/digoal/blog/blob/master/202604/20260408_02.md). --- # Replication, failover, and upgrades Canonical URL: https://pg.edu.rich/en/docs/operations/replication-upgrades Last reviewed: 2026-08-06 ## Physical versus logical [#physical-versus-logical] | Dimension | Physical streaming | Logical replication | | --------------- | ------------------------------------ | ---------------------------------------------------- | | Unit | WAL / instance | Table changes | | Goal | Same-major standby, HA, read scaling | Selected tables, cross-major migration, distribution | | DDL | Physically present | Usually synchronize schema separately | | Sequences | Physically present | Handle sequence state separately | | Write conflicts | Standby is not writable | Local subscriber writes can conflict | Basic physical observations: ```sql -- primary SELECT application_name, state, sync_state, sent_lsn, write_lsn, flush_lsn, replay_lsn FROM pg_stat_replication; -- standby SELECT pg_is_in_recovery(), pg_last_wal_receive_lsn(), pg_last_wal_replay_lsn(), now() - pg_last_xact_replay_timestamp() AS replay_delay; ``` With no new transactions, time-based replay delay can be null or misleading. Also inspect LSN distance, WAL rate, and service health. ## Replication slots [#replication-slots] Slots prevent required WAL from disappearing too early, but retain disk indefinitely when a consumer stops. Alert on slot lag and `pg_wal` capacity; confirm no consumer depends on a slot before dropping it. ```sql SELECT slot_name, slot_type, active, pg_size_pretty(pg_wal_lsn_diff(pg_current_wal_lsn(), restart_lsn)) AS retained FROM pg_replication_slots; ``` ## Failover is not one command [#failover-is-not-one-command] A runbook confirms the primary is truly unavailable, evaluates unsent WAL, promotes the target, fences the old primary from writes, updates routing, verifies writes and background jobs, and rebuilds redundancy. Without fencing, split brain is possible. ## Upgrade paths [#upgrade-paths] * **Minor**: fixes within a major; generally replace binaries and restart, but read release notes. * **Major**: requires `pg_upgrade`, logical dump/restore, or logical replication; data directories are not forward-compatible. * You can skip intervening majors, but read every intervening release note and verify extension support. ### Major cutover checklist [#major-cutover-checklist] 1. Inventory extensions, collations, types, drivers, and topology. 2. Rehearse on a restored production copy; measure downtime, disk, and `ANALYZE` time. 3. Run application tests, critical-plan comparisons, and data validation. 4. Make source of truth explicit during freeze or dual-write. 5. Before cutover, confirm catch-up, clear long transactions, and retain the rollback window. 6. After cutover, refresh statistics and inspect invalid objects, errors, performance, and backups. For the major currently in testing, use the [PostgreSQL 19 release and 18-to-19 upgrade guide](/en/docs/postgresql-19) to check Beta status, incompatibilities, and the `pg_upgrade --check` workflow. ### Logical replication additions in PostgreSQL 19 [#logical-replication-additions-in-postgresql-19] PostgreSQL 19 (Beta as of this update) plans several changes that directly affect logical-replication upgrades; confirm details against the final release notes. * **Sequence synchronization** closes the classic gap in the table above: `CREATE PUBLICATION ... FOR ALL SEQUENCES` publishes sequences, and `ALTER SUBSCRIPTION ... REFRESH SEQUENCES` reconciles the subscriber's sequence values with the publisher; `pg_get_sequence_data()` reports synchronization state. After cutover this prevents new sequence values on the subscriber from colliding with rows already replicated — the usual cause of primary-key conflicts following a logical-replication major upgrade. See [CREATE PUBLICATION](https://www.postgresql.org/docs/19/sql-createpublication.html) and [ALTER SUBSCRIPTION](https://www.postgresql.org/docs/19/sql-altersubscription.html). * **Publication blacklists**: `FOR ALL TABLES EXCEPT (TABLE ...)` excludes named tables instead of enumerating every included table, which simplifies migrations where only a few tables must stay behind. * **Conflict retention**: the subscription parameters `retain_dead_tuples` and `max_retention_duration` keep dead-tuple information used for conflict detection on the subscriber, with a bounded retention window. * **Read-your-writes on standbys**: the `WAIT FOR LSN` command lets a session wait until a given LSN is replayed on a standby, so a cutover runbook can route reads to a replica without losing just-committed writes. See [WAIT FOR LSN](/en/docs/operations/wait-for-lsn). Each PostgreSQL major is normally supported for five years. New systems should run the newest minor of a supported major. As of August 2026, 18, 17, 16, 15, and 14 are supported; 14 reaches end of support in November 2026. See [Version policy](/en/docs/reference/version-policy). --- # Safe PostgreSQL schema migrations Canonical URL: https://pg.edu.rich/en/docs/operations/safe-migrations Last reviewed: 2026-08-02 Zero downtime is not an intrinsic property of one DDL statement. It depends on **whether old and new application versions can operate together during the migration window**. Safe migration needs compatibility phases, a lock budget, representative-scale tests, observable execution, and explicit rollback boundaries. ## Default to expand-and-contract [#default-to-expand-and-contract] For a renamed or replaced field on a high-traffic table: 1. **Expand**: add the new column or table without removing the old structure; prefer short metadata operations. 2. **Dual compatible**: make the application read both structures and dual-write when required; writes must be idempotent. 3. **Backfill**: use small primary-key or time ranges and bound transaction duration, WAL, and replica lag. 4. **Switch**: change reads first, then stop old writes; verify with business signals and consistency queries. 5. **Contract**: remove the old structure only after a full rollback-compatible release window. Putting “add, backfill, set `NOT NULL`, and drop old” in one long transaction often amplifies lock, WAL, rollback, and replication-lag risk. ## Bound lock waits [#bound-lock-waits] ```sql BEGIN; SET LOCAL lock_timeout = '2s'; SET LOCAL statement_timeout = '15min'; ALTER TABLE app.orders ADD COLUMN IF NOT EXISTS fulfillment_state text; COMMIT; ``` These timeouts are examples, not universal defaults. `lock_timeout` prevents a migration from waiting indefinitely and acquiring a strong lock at an uncontrolled moment; `statement_timeout` bounds execution. On failure, exit and investigate the blocker instead of retrying forever. Constraint scanning can be separated from the short-lock phase: ```sql ALTER TABLE app.orders ADD CONSTRAINT orders_total_nonnegative CHECK (total_cents >= 0) NOT VALID; ALTER TABLE app.orders VALIDATE CONSTRAINT orders_total_nonnegative; ``` Check the exact DDL lock level on the target PostgreSQL version and representative data. `CREATE INDEX CONCURRENTLY` still consumes I/O, WAL, and time, and a failure can leave an invalid index that must be detected. ## Recommended CI pipeline [#recommended-ci-pipeline] ```text Schema / reviewed SQL ↓ Migration generation ↓ Squawk static checks ↓ Disposable PostgreSQL 18.4 ↓ Apply every migration from empty and upgraded states ↓ pgTAP + application integration + RLS negative tests ↓ PostgreSQL 19 Beta compatibility lane ↓ Representative-data rehearsal → staging → production ``` * [Squawk](https://github.com/sbdchd/squawk) detects common hazards such as non-concurrent indexes, constraints added without `NOT VALID`, and selected lock risks; it is not proof of zero downtime. * [Testcontainers for Node.js](https://github.com/testcontainers/testcontainers-node) starts real PostgreSQL in CI for transactions, locks, RLS, JSONB, extensions, and driver behavior. * [pgTAP](https://github.com/theory/pgtap) tests functions, triggers, constraints, and policies inside PostgreSQL. With Drizzle ORM, Drizzle Kit can generate ordinary schema changes before Squawk and human review. Use reviewed native SQL for complex indexes, policies, functions, extensions, and PostgreSQL 19 syntax. An ORM's inability to express a feature does not make the database feature inappropriate. ## Test the upgrade path, not only an empty database [#test-the-upgrade-path-not-only-an-empty-database] CI needs at least two starting states: | Starting point | Problems exposed | | --------------------------------------------------------------- | -------------------------------------------------------- | | Empty database running every migration | Ordering, dependencies, syntax, and bootstrap | | Production schema or masked data running incremental migrations | Locks, backfills, old data, constraints, and performance | Split version tests into two lanes: * **Production gate**: the current production major/minor, such as PostgreSQL 18.4; failures block release. * **Forward compatibility**: PostgreSQL 19 Beta 2; failures may initially be allowed but must be classified, tracked, and cleared before GA adoption. Beta testing does not replace the stable-version gate. Follow version status on the [PostgreSQL 19 topic page](/en/docs/postgresql-19). ## RLS and security objects need negative tests [#rls-and-security-objects-need-negative-tests] A successful migration does not prove correct authorization. For every tenant and role, verify that allowed `SELECT/INSERT/UPDATE/DELETE` operations succeed and forbidden cross-tenant reads and writes fail. The runtime role should not own tables; use `FORCE ROW LEVEL SECURITY` where required. See the [Security baseline](/en/docs/operations/security). Dropped columns, irreversible backfills, narrowed types, and external side effects may not reverse safely. Every release needs a last rollback point, an answer for whether the old app can read the new schema, and criteria for a forward fix. ## When heavier tooling is justified [#when-heavier-tooling-is-justified] | Tool | Useful when | Verify before adoption | | ------------------------------------------------------------------------- | ------------------------------------------------------------------ | ---------------------------------------------------------------- | | [pgroll](https://github.com/xataio/pgroll) | High-frequency changes with explicit compatibility windows | Supported DDL, proxy/connection path, rollback semantics | | [Bytebase](https://github.com/bytebase/bytebase) | Multi-team approval, SQL review, environment, and audit governance | Permission boundaries, deployment model, existing CI integration | | [Database Lab Engine](https://github.com/postgres-ai/database-lab-engine) | Fast clones and migration rehearsal for large databases | Storage, masking, clone lifecycle, and cost | Small teams should first make expand-and-contract, real PostgreSQL tests, lock observation, and recovery drills routine before adding a control plane. --- # Security baseline Canonical URL: https://pg.edu.rich/en/docs/operations/security Last reviewed: 2026-08-06 ## Separate roles [#separate-roles] ```sql CREATE ROLE app_owner NOLOGIN; CREATE ROLE app_runtime LOGIN; CREATE ROLE app_migrator LOGIN NOINHERIT; CREATE SCHEMA app AUTHORIZATION app_owner; GRANT app_owner TO app_migrator; GRANT CONNECT ON DATABASE commerce TO app_runtime, app_migrator; GRANT USAGE ON SCHEMA app TO app_runtime; GRANT SELECT, INSERT, UPDATE, DELETE ON ALL TABLES IN SCHEMA app TO app_runtime; ALTER DEFAULT PRIVILEGES FOR ROLE app_owner IN SCHEMA app GRANT SELECT, INSERT, UPDATE, DELETE ON TABLES TO app_runtime; ALTER DEFAULT PRIVILEGES FOR ROLE app_owner IN SCHEMA app GRANT USAGE, SELECT ON SEQUENCES TO app_runtime; ``` Runtime does not own objects, the migrator uses `SET ROLE app_owner` only during migrations, and the owner cannot log in. `ALTER DEFAULT PRIVILEGES` affects future objects created by the specified creator; it does not repair existing privileges. A read-only tier for BI tools, support access, and AI agents completes the layering: ```sql CREATE ROLE app_readonly LOGIN; ALTER ROLE app_readonly SET default_transaction_read_only = on; GRANT CONNECT ON DATABASE commerce TO app_readonly; GRANT USAGE ON SCHEMA app TO app_readonly; GRANT SELECT ON ALL TABLES IN SCHEMA app TO app_readonly; ALTER DEFAULT PRIVILEGES FOR ROLE app_owner IN SCHEMA app GRANT SELECT ON TABLES TO app_readonly; ``` `default_transaction_read_only` makes accidental writes fail even if a write grant is added later by mistake. ## Authentication and network [#authentication-and-network] * Listen only on required interfaces and restrict sources with firewall/security groups. * Require TLS remotely and verify the server certificate; consider client certificates for sensitive systems. * Use SCRAM for new password authentication and retire MD5 configuration. * Order `pg_hba.conf` from narrow rules to broad ones by network, database, and role; reload and test both allow and deny paths. * Separate administrative and application entry points; do not expose the database port publicly. A `pg_hba.conf` example matching that ordering — the first matching rule wins, so narrow rules come before broad ones: ``` # TYPE DATABASE USER ADDRESS METHOD local all all peer hostssl commerce app_runtime 10.0.1.0/24 scram-sha-256 hostssl commerce app_readonly 10.0.2.0/24 scram-sha-256 host all all 0.0.0.0/0 reject ``` Reload with `SELECT pg_reload_conf();` and test both an allowed and a rejected connection. ## Protect search\_path [#protect-search_path] Do not resolve names through untrusted writable schemas. Revoke default create privilege on `public` and pin paths for sensitive functions: ```sql REVOKE CREATE ON SCHEMA public FROM PUBLIC; CREATE FUNCTION app.current_tenant() RETURNS bigint LANGUAGE sql STABLE SECURITY DEFINER SET search_path = pg_catalog, app AS $$ SELECT current_setting('app.tenant_id')::bigint $$; REVOKE ALL ON FUNCTION app.current_tenant() FROM PUBLIC; GRANT EXECUTE ON FUNCTION app.current_tenant() TO app_runtime; ``` `SECURITY DEFINER` runs with owner privilege. Audit every input, qualified object name, search path, and execute grant. ## Row-level security [#row-level-security] ```sql ALTER TABLE app.orders ENABLE ROW LEVEL SECURITY; ALTER TABLE app.orders FORCE ROW LEVEL SECURITY; CREATE POLICY tenant_orders ON app.orders USING (tenant_id = current_setting('app.tenant_id')::bigint) WITH CHECK (tenant_id = current_setting('app.tenant_id')::bigint); ``` With RLS enabled, no applicable policy means default deny. Superusers, `BYPASSRLS`, and normally table owners can bypass; `FORCE ROW LEVEL SECURITY` subjects owners during ordinary access. Normal object privileges still apply. An administrator's result cannot prove RLS. Test the allowed tenant, another tenant, missing context, inserts, and updates. Confirm the pool sets and clears tenant context on every checkout/return. ## Secrets and logs [#secrets-and-logs] Rotate and shorten credentials and distribute them via a secret manager. Avoid bound values and sensitive DDL in database logs; restrict audit-log access and retention. `pg_stat_activity` can expose query text, so monitoring-view privilege also matters. ## Audit [#audit] The built-in starting point is statement logging: ```sql ALTER SYSTEM SET log_statement = 'ddl'; SELECT pg_reload_conf(); ``` For accountable audit trails use [pgaudit](https://github.com/pgaudit/pgaudit), which must be loaded via `shared_preload_libraries`: ```ini shared_preload_libraries = 'pgaudit' pgaudit.log = 'write, ddl, role' ``` Plan for its boundaries: * pgaudit is an extension, not core; availability and version depend on your distribution or cloud provider. * Broad classes such as `pgaudit.log = 'all'` produce large log volumes on busy systems. Start with `ddl, role` and add classes per compliance requirement; use object auditing (`pgaudit.log_relation`) for a few sensitive tables instead of logging everything. * Session auditing can still capture sensitive values, so the log access, retention, and redaction rules above apply to audit output as well. --- # Why too many tables in one database hurt PostgreSQL Canonical URL: https://pg.edu.rich/en/docs/operations/too-many-tables Last reviewed: 2026-08-06 Databases with tens or hundreds of thousands of tables usually get there by one of three routes: * **Per-tenant schemas**: every customer gets its own set of tables (`tenant_1234.orders`, `tenant_1234.invoices`, ...) because that felt like clean isolation; * **Partition maximalism**: daily partitions kept forever across dozens of tables, each partition carrying indexes, constraints, and statistics; * **ORM and tool churn**: frameworks that create a table per entity revision, per report, or per import job and never drop them. All three produce the same failure shape, and it arrives gradually: nothing breaks at 5,000 tables, something is vaguely slow at 50,000, and at 500,000 you are debugging memory pressure and multi-minute `pg_dump` runs. ## Why too many tables hurt [#why-too-many-tables-hurt] ### Relcache: metadata cached per backend [#relcache-metadata-cached-per-backend] Every backend process keeps its own relation cache — the *relcache* — holding parsed metadata for each relation it has touched: tuple descriptor, indexes, rules, triggers, statistics pointers. This is a per-process cache in backend memory, populated lazily on first access and kept for the life of the backend (see `src/backend/utils/cache/relcache.c` in the PostgreSQL source). Two consequences follow: * the memory cost of a touched table is paid **once per backend**, so a pool of 200 connections that each touch 20,000 tables carries the metadata 200 times; * nothing reclaims it while the connection lives, so long-lived pooler connections in `session` mode grow monotonically — in a container this is exactly the anonymous-memory growth that ends in an [OOM kill](/en/docs/operations/container-memory-oom). PostgreSQL 14+ exposes this through [`pg_backend_memory_contexts`](https://www.postgresql.org/docs/18/view-pg-backend-memory-contexts.html): metadata accumulates under `CacheMemoryContext` and its children, and you can watch it grow as a backend touches more relations. ### System catalogs grow, and everything that walks them slows down [#system-catalogs-grow-and-everything-that-walks-them-slows-down] Each table is not one catalog row. A table with five columns, a primary key, and one index adds rows to `pg_class`, `pg_attribute`, `pg_index`, `pg_constraint`, `pg_depend`, `pg_description`, `pg_statistic`, and more — a dozen or so rows per table at minimum, per the [system catalogs](https://www.postgresql.org/docs/18/catalogs.html) layout. Consequences at scale: * metadata introspection slows down: `\d` in psql, ORM schema reflection at application startup, and GUI tools that enumerate the catalog; * `pg_dump` walks and locks every relation, so backup duration grows with table count even when data volume is flat; * catalogs themselves bloat and need vacuum like any other table — autovacuum on a huge `pg_class` becomes a real workload. ## How many is too many [#how-many-is-too-many] These are field experience values, not documented thresholds — the actual breaking point depends on connection count, columns per table, and how many relations each backend touches: * **a few thousand tables**: fine on any reasonable setup; * **tens of thousands**: noticeable — slower connection warm-up, slower dumps, relcache memory visible in monitoring; * **hundreds of thousands**: incident territory — backends holding gigabytes of relcache, OOM risk in containers, `pg_dump` measured in hours. The multiplier that matters is *tables touched per backend × concurrent backends*, not the raw count. 50,000 tables that each request touches 50 of are far cheaper than 50,000 tables all touched by every request. ## Diagnosing the problem [#diagnosing-the-problem] Count relations by kind, excluding system schemas: ```sql SELECT c.relkind, count(*) FROM pg_class c JOIN pg_namespace n ON n.oid = c.relnamespace WHERE n.nspname NOT IN ('pg_catalog', 'information_schema') GROUP BY c.relkind ORDER BY count(*) DESC; ``` `relkind`: `r` table, `p` partitioned table, `i`/`I` index, `S` sequence, `t` TOAST table, `v`/`m` views. If the count is dominated by indexes, the tables underneath are still the root cause — each one drags its indexes along. Find where the tables concentrate: ```sql SELECT n.nspname, count(*) FILTER (WHERE c.relkind = 'r') AS tables, count(*) FILTER (WHERE c.relkind = 'i') AS indexes FROM pg_class c JOIN pg_namespace n ON n.oid = c.relnamespace WHERE n.nspname NOT IN ('pg_catalog', 'information_schema') GROUP BY n.nspname ORDER BY tables DESC LIMIT 20; ``` Thousands of schemas with identical table names is the per-tenant signature. Measure catalog size and per-backend metadata memory: ```sql -- Largest system catalogs SELECT relname, pg_size_pretty(pg_total_relation_size(oid)) AS total FROM pg_class WHERE relnamespace = 'pg_catalog'::regnamespace AND relkind = 'r' ORDER BY pg_total_relation_size(oid) DESC LIMIT 10; -- Metadata memory of the current backend (PG 14+; -- other backends require superuser or pg_read_all_stats) SELECT name, pg_size_pretty(sum(used_bytes)) AS used FROM pg_backend_memory_contexts GROUP BY name ORDER BY sum(used_bytes) DESC LIMIT 15; ``` A `CacheMemoryContext` that grows with the number of distinct relations a backend has touched — and never shrinks — confirms the relcache cost. ## Remediation paths [#remediation-paths] **Merge per-tenant tables into shared tables with Row-Level Security.** One `orders` table with a `tenant_id` column and an RLS policy preserves the isolation guarantee while collapsing the table count by orders of magnitude; see [Row Security Policies](https://www.postgresql.org/docs/18/ddl-rowsecurity.html) and our [Security](/en/docs/operations/security) page. The migration is mechanical (union the tenant tables in, add the column, backfill) but test RLS plan behavior on your query shapes — policies are applied per query and interact with indexes on `tenant_id`. **Cap partition counts deliberately.** Choose partition granularity from retention and query patterns, not from calendar habit: monthly instead of daily, and drop or detach old partitions instead of keeping history online forever. Details below. **Split by database.** If tenants genuinely need hard separation, a database per tenant bounds the blast radius better than a schema per tenant — catalogs and relcache are per database, and connections are too. The trade-off is connection management and cross-tenant reporting. **Drop what the ORM forgot.** Audit for tables with zero scans and zero tuples over a `pg_stat_user_tables` window and remove them; schema hygiene is cheaper than any of the above. ## The boundary with partitioning [#the-boundary-with-partitioning] Partitioning is not an escape hatch from this problem — **partitions are tables**. Every partition has its own `pg_class` entry, its own relcache footprint per backend that touches it, and usually its own indexes. Declarative partitioning merely automates the routing; the metadata cost is additive. The [partitioning documentation](https://www.postgresql.org/docs/18/ddl-partitioning.html) notes that the planner handles hierarchies of up to a few thousand partitions reasonably well *when pruning eliminates most of them at plan time*; planning time and memory grow when many partitions survive pruning. So the practical limits compose: a table partitioned into 5,000 children, queried without a partition-key filter, from a 200-connection pool, combines worst-case planning cost with worst-case relcache multiplication. Keep partition counts in the hundreds where possible, always query with the partition key, and let retention (`DROP PARTITION` / `DETACH`) — not storage capacity — decide how many partitions stay online. For the data-modeling side of choosing between shared tables, schemas, and partitions, see [Data modeling](/en/docs/core/data-modeling). The relcache is not governed by `work_mem` or `shared_buffers`; it lives in per-backend private memory with no configured ceiling. The only levers are fewer tables, fewer touched tables per query, fewer long-lived backends, or more memory — in that order of preference. Verify the plan on a restored copy first — catalog surgery and RLS rollout are exactly the changes that deserve a rehearsal. ## Related pages [#related-pages] * Row-Level Security setup → [Security](/en/docs/operations/security) * Choosing table layouts → [Data modeling](/en/docs/core/data-modeling) * Connection and memory budgeting → [Server configuration](/en/docs/operations/configuration) * When metadata memory meets the cgroup limit → [PostgreSQL memory management and OOM in containers](/en/docs/operations/container-memory-oom) --- # WAIT FOR LSN: read-your-writes on async replicas Canonical URL: https://pg.edu.rich/en/docs/operations/wait-for-lsn Last reviewed: 2026-08-06 As of 2026-08, PostgreSQL 19 has not reached GA. The syntax and behavior below follow the [PostgreSQL 19 documentation](https://www.postgresql.org/docs/19/sql-wait-for.html); re-check the final release notes before relying on them in production. A common scaling pattern sends writes to the primary and reads to asynchronous replicas. Its weakest point is the read immediately after a write: the client committed on the primary, but the replica has not replayed that WAL yet, so the follow-up read returns the old value. This is the stale read problem, and until PostgreSQL 19 there was no server-side primitive to close the window — only application workarounds. ## The stale read problem [#the-stale-read-problem] With asynchronous streaming replication, `COMMIT` on the primary does not wait for any standby. The WAL record is shipped, written, flushed, and replayed on the replica some milliseconds — or under load, seconds — later. A client that writes and then reads through a replica connection pool can observe its own commit as missing: ```text primary: UPDATE profile SET display_name = 'new' WHERE id = 42; -- COMMIT replica: SELECT display_name FROM profile WHERE id = 42; --> 'old' (replay has not reached the commit record yet) ``` The window is unbounded: replication lag spikes with write bursts, long transactions, and recovery conflicts, so any fixed assumption about "the replica is probably caught up" eventually breaks. ## Workarounds before PostgreSQL 19 [#workarounds-before-postgresql-19] Four patterns are widely deployed, and each pays for correctness somewhere else: * **Fixed sleep after writes.** Delay the next read by N milliseconds. Lag is not constant, so the sleep is either too short (still stale under load) or too long (idle latency on every request when the replica is caught up). * **Sticky reads.** Pin a session's reads to the primary for a while after its writes. Correct, but it pushes read traffic back to the primary for exactly your most active users — the ones you offloaded reads for. * **`synchronous_commit = remote_apply`.** Every `COMMIT` waits until a synchronous standby has replayed the change, making it visible everywhere. This works, but it taxes *all* writes with replica round-trip latency and couples write availability to standby health. It is a cluster-wide durability decision, not a per-request consistency hint. * **Polling `pg_last_wal_replay_lsn()`.** After the write, record `pg_current_wal_insert_lsn()` on the primary, then loop on the replica comparing replay position until it passes your LSN. Semantically correct, but every check is a network round trip, the loop burns connections, and timeout/promotion handling is entirely your code. PostgreSQL 19 replaces the last pattern with a server-side blocking wait. ## WAIT FOR LSN syntax [#wait-for-lsn-syntax] [`WAIT FOR`](https://www.postgresql.org/docs/19/sql-wait-for.html) blocks the session until the server reaches a target LSN, then returns a one-row status: ```sql WAIT FOR LSN '0/306EE20'; WAIT FOR LSN '0/306EE20' WITH (MODE 'standby_flush'); WAIT FOR LSN '0/306EE20' WITH (MODE 'standby_write', TIMEOUT '100ms', NO_THROW); ``` Possible return values are `success`, `timeout`, and `not in recovery`. Without `TIMEOUT` (or with `0`) the command waits indefinitely. On timeout — or on running a standby mode against a server that is not in recovery — it raises an error unless `NO_THROW` is given, in which case the status column tells you what happened. The typical read-your-writes flow: ```sql -- on the primary, immediately after the write SELECT pg_current_wal_insert_lsn(); -- 0/306EE20 -> hand this to the application / pooler -- on the replica, before the dependent read WAIT FOR LSN '0/306EE20' WITH (TIMEOUT '200ms', NO_THROW); -- status = success -> the write is now visible on this replica SELECT display_name FROM profile WHERE id = 42; ``` Use `pg_current_wal_insert_lsn()` (insert position), not the flush position, so the sequence stays correct even when the writing session runs with `synchronous_commit = off`. The documentation notes that the LSN of the last modification should be tracked on the client application or connection pooler side — `WAIT FOR` itself does not remember anything between statements. ## The four MODE values [#the-four-mode-values] `MODE` selects which stage of WAL processing to wait for. The default is `standby_replay`. | MODE | Waits until | Server state | | ---------------- | ------------------------------------------------------------------------------------------------------------ | ------------ | | `standby_replay` | the LSN is replayed (applied) on the standby; afterwards `pg_last_wal_replay_lsn()` is at or past the target | standby only | | `standby_flush` | the WAL is flushed to disk on the standby — durability without waiting for apply | standby only | | `standby_write` | the WAL is written to the OS on the standby — faster than flush, weaker durability | standby only | | `primary_flush` | the WAL is flushed to disk on the primary; afterwards `pg_current_wal_flush_lsn()` is at or past the target | primary only | For read-your-writes, `standby_replay` is the mode that matters: replay is what makes the row visible to queries. `standby_write` and `standby_flush` are also satisfied by WAL the standby already holds from a base backup or archive restore, not only freshly streamed WAL. Using a standby mode on a primary — or `primary_flush` on a standby — is an error. ## Limitations [#limitations] The restrictions in the [command reference](https://www.postgresql.org/docs/19/sql-wait-for.html) shape how you can call it: * **Top-level statement only.** `WAIT FOR` cannot run inside a function, procedure, or `DO` block. Your driver or pool middleware must issue it as its own statement. * **No held snapshot.** The command requires that no active or registered snapshot is held, so it cannot be used in contexts where a snapshot must stay active — including transactions at isolation levels stricter than `READ COMMITTED`. In practice: send it as a standalone autocommit statement, before the read that needs the guarantee. * **Promotion changes the answer.** If the standby is promoted while you wait with a standby mode, the command returns `not in recovery` (or errors without `NO_THROW`). Promotion creates a new timeline, and the LSN you were waiting on may belong to the old one — the application must re-evaluate whether the target is still meaningful. * **Numeric comparison only.** `WAIT FOR` compares LSN values, not timelines. A cascading standby whose upstream was promoted can report `success` once its replay position passes the number, even if that position is on a different timeline. If that distinction matters, validate the timeline yourself. * **Recovery conflicts still apply.** On a standby, the waiting session can be interrupted by recovery conflict resolution — some conflicts (a tablespace drop is the documented example) terminate all backends unconditionally. Callers need retry and fallback paths. ## Degradation after timeout [#degradation-after-timeout] Treat `WAIT FOR` as a best-effort consistency upgrade with a strict budget, not as a correctness primitive. Always pair `TIMEOUT` with `NO_THROW` and branch on the status: * `success` — read from the replica as planned. * `timeout` — fall back: re-issue the read against the primary, or serve the stale read and refresh asynchronously. Record the replay lag at the moment of timeout; a rising timeout rate is an early warning of replica capacity problems. * `not in recovery` — the topology changed. Re-resolve the primary, re-fetch a fresh LSN there if the write path moved, and do not reuse the old LSN across the timeline switch. Size the timeout from your read latency budget (tens to a few hundred milliseconds), not from average lag — the wait exists precisely for the tail. A `WAIT FOR` that frequently hits its timeout is telling you to fix replication lag first; see [Streaming replication upgrades and lag control](/en/docs/operations/replication-upgrades). ## Prompt: design the routing layer [#prompt-design-the-routing-layer] Command reference: [PostgreSQL 19 WAIT FOR](https://www.postgresql.org/docs/19/sql-wait-for.html). For the broader PostgreSQL 19 feature set, see the [PostgreSQL 19 overview](/en/docs/postgresql-19). Field reports with measured behavior: [rednafi's walkthrough](https://rednafi.com/system/wait-for-lsn/) and [digoal's analysis (Chinese)](https://github.com/digoal/blog/blob/master/202601/20260106_01.md). --- # Data modeling and constraints Canonical URL: https://pg.edu.rich/en/docs/core/data-modeling Last reviewed: 2026-08-02 A good PostgreSQL model does not postpone every rule to application code. It teaches the database which states are valid. ## A working order model [#a-working-order-model] ```sql CREATE TABLE customers ( id bigint GENERATED ALWAYS AS IDENTITY PRIMARY KEY, email text NOT NULL, display_name text NOT NULL CHECK (length(trim(display_name)) > 0), created_at timestamptz NOT NULL DEFAULT now(), CONSTRAINT customers_email_unique UNIQUE (email) ); CREATE TABLE orders ( id bigint GENERATED ALWAYS AS IDENTITY PRIMARY KEY, customer_id bigint NOT NULL REFERENCES customers(id), status text NOT NULL DEFAULT 'pending' CHECK (status IN ('pending', 'paid', 'shipped', 'cancelled')), total_cents bigint NOT NULL CHECK (total_cents >= 0), placed_at timestamptz NOT NULL DEFAULT now() ); COMMENT ON COLUMN orders.total_cents IS 'Order total in the smallest currency unit; never a floating-point amount.'; ``` ## Type choices [#type-choices] | Need | Prefer | Avoid | | ------------------- | -------------------------------------------------------------- | --------------------------------------------------- | | Primary key | `bigint GENERATED ... AS IDENTITY` or `uuid` | New designs depending on implicit `serial` behavior | | Money | Smallest unit in `bigint`, or explicit `numeric(p,s)` | `real` / `double precision` | | Instant | `timestamptz` | Storing a real-world instant as text | | Text | `text` plus business constraints | Arbitrary `varchar(255)` without meaning | | State | `CHECK` for small stable sets; reference table for a lifecycle | Unconstrained free text | | Document attributes | `jsonb` | Hiding core relations and foreign keys in JSON | `timestamptz` stores an absolute instant and renders it in the session time zone. It does not retain the input zone name. Store a zone identifier separately when rules such as `Europe/Paris` matter. ## What constraints mean [#what-constraints-mean] * `NOT NULL`: a value must exist. * `CHECK`: each row must satisfy a predicate. * `UNIQUE`: a candidate key is unique; multiple `NULL`s are allowed by default. * `PRIMARY KEY`: unique, non-null row identity. * `FOREIGN KEY`: the target must exist; deletion behavior is a design decision. `ON DELETE CASCADE` says children should disappear with the parent. Use it only for truly dependent lifecycles. Invoices and audit records normally should not cascade. ## Schemas and names [#schemas-and-names] Use an explicit schema for application objects and reduce default privilege: ```sql CREATE SCHEMA app; REVOKE CREATE ON SCHEMA public FROM PUBLIC; ALTER ROLE app_runtime SET search_path = app, pg_catalog; ``` Names that help both humans and models are complete, stable, and low on abbreviations: `customer_id` beats `cid`; `created_at` beats `ctime`. Use `COMMENT ON` for units, state transitions, and sensitivity—not to repeat the column name. ## Verify the model [#verify-the-model] ```sql INSERT INTO customers (email, display_name) VALUES ('ada@example.com', 'Ada') RETURNING id, created_at; -- Expected to fail: totals cannot be negative INSERT INTO orders (customer_id, total_cents) VALUES (1, -100); ``` A model is not verified merely because valid rows work. Representative invalid rows must fail for the intended reason. --- # Graph queries with SQL/PGQ Canonical URL: https://pg.edu.rich/en/docs/core/graph-queries Last reviewed: 2026-08-06 PostgreSQL 19 implements SQL/PGQ (ISO/IEC 9075-16, SQL:2023 Part 16): graph pattern matching over ordinary relational tables. A property graph is a read-only view over tables you already have — no extension, no data copy — and `GRAPH_TABLE` queries are planned through the same relational planner as any join query. SQL/PGQ is new in PostgreSQL 19. PostgreSQL 18 and earlier do not have `CREATE PROPERTY GRAPH` or `GRAPH_TABLE`. See the [PostgreSQL 19 release and upgrade guide](/en/docs/postgresql-19) for version status; syntax below follows the official documentation ([5.15 Property Graphs](https://www.postgresql.org/docs/19/ddl-property-graphs.html), [7.9 Graph Queries](https://www.postgresql.org/docs/19/queries-graph.html)). ## When graph queries pay off [#when-graph-queries-pay-off] * Social relationships: friends of friends, mutual connections. * Data lineage: which source rows and transformations produced a report number. * Fraud and audit: money flows, suspicious paths, compliance tracing. The trade-off is fixed depth. SQL/PGQ in PostgreSQL 19 matches patterns hop by hop; variable-length paths are not supported. Within a few hops the pattern syntax is much clearer than the equivalent join chain; beyond that, use `WITH RECURSIVE` (see [Limits and the WITH RECURSIVE boundary](#limits-and-the-with-recursive-boundary)). ## Define the graph: CREATE PROPERTY GRAPH [#define-the-graph-create-property-graph] The canonical social network: `person` and `knows`. ```sql CREATE TABLE person ( id int PRIMARY KEY, name text NOT NULL, age int, city text ); CREATE TABLE knows ( a int NOT NULL REFERENCES person(id), -- knows whom b int NOT NULL REFERENCES person(id), -- known by whom since int, PRIMARY KEY (a, b) ); ``` Declare it as a graph — vertices are `person`, edges are `knows`, directed from `a` to `b`: ```sql CREATE PROPERTY GRAPH social VERTEX TABLES ( person KEY (id) LABEL person PROPERTIES (id, name, age, city) ) EDGE TABLES ( knows SOURCE KEY (a) REFERENCES person (id) DESTINATION KEY (b) REFERENCES person (id) LABEL knows PROPERTIES (since) ); ``` Read it as a contract: * A vertex table's `KEY` is usually its primary key; vertex tables need one. * `SOURCE KEY ... REFERENCES ...` and `DESTINATION KEY ... REFERENCES ...` state the edge direction: from which column, to which column of which table. * `LABEL` is the name inside the graph (it may differ from the table name); `PROPERTIES` controls which columns are visible in the graph. When table and column names already match what you want as labels and properties, both clauses can be omitted. `CREATE PROPERTY GRAPH` stores metadata only. Data stays in the original tables; creating or dropping the graph never touches business data, and the definition can be recreated at any time. ## Query with GRAPH\_TABLE [#query-with-graph_table] List everyone: ```sql SELECT name FROM GRAPH_TABLE (social MATCH (p IS person) COLUMNS (p.name) ) ORDER BY name; ``` Logically this is `SELECT name FROM person ORDER BY name`. `MATCH (p IS person)` walks every `person` vertex and binds it to `p`; `COLUMNS` projects the output. The result of `GRAPH_TABLE` is an ordinary table: it can be aliased, filtered, and joined like any other `FROM` item. `COLUMNS (p.*)` fails with `"*" is not supported here`. List every output column explicitly, e.g. `COLUMNS (p.id, p.name)`. ### Edge patterns and direction [#edge-patterns-and-direction] ```sql SELECT * FROM GRAPH_TABLE (social MATCH (p IS person)-[IS knows]->(p2 IS person) COLUMNS (p.id, p.name, p2.id, p2.name) ) ORDER BY 1, 2, 3; ``` `(p)-[IS knows]->(p2)` is an edge pattern: * `->` follows the declared direction (SOURCE → DESTINATION); `<-` walks it in reverse. * A bare `-` matches either direction. A bare `-` matches an edge in either direction (an OR of both directions), so every relationship appears twice. Unless the underlying data is genuinely symmetric — one row inserted per direction — write `->` or `<-` explicitly. ### Multiple hops [#multiple-hops] ```sql SELECT * FROM GRAPH_TABLE (social MATCH (a IS person)-[IS knows]-> (b IS person)-[IS knows]->(c IS person) WHERE a.id <> c.id COLUMNS (a.name AS a, b.name AS via, c.name AS c) ) ORDER BY a, c, via; ``` * Each extra hop is another `-[IS knows]->(...)` in the chain. * `WHERE a.id <> c.id` lives inside `MATCH` and filters out cycles such as "Alice → Bob → Alice". Drop it to see all two-hop paths including the circular ones. ## Multiple vertex and edge types [#multiple-vertex-and-edge-types] Real models span several tables. Add companies and employment: ```sql CREATE TABLE company ( id int PRIMARY KEY, name text NOT NULL, industry text NOT NULL ); CREATE TABLE works_at ( pid int NOT NULL REFERENCES person(id), cid int NOT NULL REFERENCES company(id), role text NOT NULL, PRIMARY KEY (pid, cid) ); CREATE PROPERTY GRAPH company_social VERTEX TABLES ( person KEY (id) LABEL person PROPERTIES (id, name, age, city), company KEY (id) LABEL company PROPERTIES (id, name, industry) ) EDGE TABLES ( knows SOURCE KEY (a) REFERENCES person (id) DESTINATION KEY (b) REFERENCES person (id) LABEL knows PROPERTIES (since), works_at SOURCE KEY (pid) REFERENCES person (id) DESTINATION KEY (cid) REFERENCES company (id) LABEL works_at PROPERTIES (role) ); ``` One graph, two vertex types and two edge types, traversed in a single query — "where do Alice's friends work?": ```sql SELECT * FROM GRAPH_TABLE (company_social MATCH (me IS person WHERE me.name = 'Alice') -[IS knows]->(friend IS person) -[IS works_at]->(co IS company) COLUMNS (friend.name AS friend, co.name AS company) ) ORDER BY friend, company; ``` `WHERE` can sit directly inside a vertex pattern. The equivalent plain SQL needs a multi-table join; the graph syntax encodes "which relationship to walk" in the pattern itself. ### No multi-pattern MATCH: join two GRAPH\_TABLEs [#no-multi-pattern-match-join-two-graph_tables] The SQL standard's comma-separated `MATCH (a...), (b...)` is not implemented in PostgreSQL 19. The workaround is the most useful idiom on this page: a `GRAPH_TABLE` result is an ordinary table, so project the IDs and join: ```sql SELECT m.me, m.via, m.coworker, w.company FROM GRAPH_TABLE (company_social MATCH (a IS person)-[IS knows]->(b IS person)-[IS knows]->(c IS person) WHERE a.id <> c.id COLUMNS (a.id AS aid, c.id AS cid, a.name AS me, b.name AS via, c.name AS coworker) ) m JOIN GRAPH_TABLE (company_social MATCH (x IS person)-[IS works_at]->(co IS company) <-[IS works_at]-(y IS person) WHERE x.id <> y.id COLUMNS (x.id AS xid, y.id AS yid, co.name AS company) ) w ON w.xid = m.aid AND w.yid = m.cid ORDER BY me, coworker; ``` "Co-workers who know each other" decomposes exactly like this: one `GRAPH_TABLE` per graph pattern, IDs as the bridge. In a graph with a single edge table, `(a)->(b)` is unambiguous. In a multi-edge graph like `company_social`, an anonymous edge silently unions all edge types — `knows` and `works_at` rows come out together. Name the edge label (`[IS knows]`, `[IS works_at]`) in any graph with more than one edge table. ## Performance: it is planned as joins [#performance-it-is-planned-as-joins] The release notes state that SQL/PGQ queries "are processed like views so are written as standard relational queries". A two-hop pattern becomes a multi-way join and goes through the ordinary planner; there are no graph-specific executor nodes in `EXPLAIN`. That is why performance is predictable: the plan shape follows the pattern shape, and the indexing rules are the ones you already know from joins. The one rule that matters: **the edge table's primary key covers the source side; add an index on the destination column if you traverse in reverse.** * Forward ("whom does 42 know") looks up `knows.a = 42`, which the primary key `(a, b)` covers. * Reverse ("who knows 42") filters on `knows.b`, which that primary key cannot serve — the scan falls back to reading the whole edge table, and the cost gap grows with table size. ```sql CREATE INDEX ON knows (b); ``` Verify with `EXPLAIN (ANALYZE, BUFFERS)` exactly as you would for any join; see [Indexes and EXPLAIN](/en/docs/core/indexes-explain). ## Data lineage pattern [#data-lineage-pattern] The classic compliance question is "where does this number in the report come from?". Model the ETL pipeline as a graph: * One vertex table per layer: `clickstream_source` (raw) → `staging` → `fact` → `report`, each with an `id` primary key and a `name`. * One edge table per transformation type, with `PRIMARY KEY (src, dst)` and foreign keys to the two layers it connects: `loads_into` (straight copy), `aggregates_into` (grouped aggregation), `rollup_into` (further summarization). ```sql CREATE PROPERTY GRAPH lineage VERTEX TABLES ( clickstream_source KEY (id), staging KEY (id), fact KEY (id), report KEY (id) ) EDGE TABLES ( loads_into SOURCE KEY (src) REFERENCES clickstream_source (id) DESTINATION KEY (dst) REFERENCES staging (id), aggregates_into SOURCE KEY (src) REFERENCES staging (id) DESTINATION KEY (dst) REFERENCES fact (id), rollup_into SOURCE KEY (src) REFERENCES fact (id) DESTINATION KEY (dst) REFERENCES report (id) ); ``` Edge labels carry semantics — that is the advantage over "foreign keys plus a recursive CTE": `rollup_into` tells you *what kind of transformation happened*, while a foreign key only says *a relationship exists*. The `CREATE PROPERTY GRAPH` statement itself can be generated from orchestration metadata (dbt, Airflow and similar tools already track upstream/downstream relationships). Four recurring query patterns: ```sql -- 1. Value trace: what fed this report number SELECT * FROM GRAPH_TABLE (lineage MATCH (r IS report)<-[IS rollup_into]-(f IS fact) COLUMNS (r.name AS report, f.name AS fact) ); -- 2. Impact analysis: if this source changes, which reports break? SELECT DISTINCT rpt FROM GRAPH_TABLE (lineage MATCH (s IS clickstream_source WHERE s.name = 'events_raw') -[IS loads_into]->(IS staging) -[IS aggregates_into]->(IS fact) -[IS rollup_into]->(r IS report) COLUMNS (r.name AS rpt) ); -- 3. Full audit trail: same MATCH as 2, but COLUMNS outputs every hop -- (s.name, staging name, fact name, r.name) -- 4. Black holes: source data with no downstream consumer SELECT src.name FROM GRAPH_TABLE (lineage MATCH (s IS clickstream_source) COLUMNS (s.id AS sid, s.name AS name) ) src WHERE NOT EXISTS ( SELECT 1 FROM loads_into l WHERE l.src = src.sid ); ``` Pattern 4 shows the escape hatch: `GRAPH_TABLE` results compose with ordinary SQL, so anti-joins, aggregates, and `EXISTS` all work on top of graph output. ## Limits and the WITH RECURSIVE boundary [#limits-and-the-with-recursive-boundary] PostgreSQL 19's SQL/PGQ does not yet have: * **Variable-length paths**: `(a)-[IS knows]->{1,3}(b)` is rejected. Either unroll hop by hop and `UNION`, or fall back to `WITH RECURSIVE`. * **Multi-pattern MATCH** (comma-separated) — join two `GRAPH_TABLE`s instead, as shown above. * **Path variable binding** (`p = (a)->(b)`) and `ANY SHORTEST` / `ALL SHORTEST` path finding. For chains beyond roughly five hops, a recursive CTE is still the right tool: ```sql WITH RECURSIVE chain AS ( SELECT a, b, 1 AS depth FROM knows WHERE a = 42 UNION ALL SELECT k.a, k.b, c.depth + 1 FROM chain c JOIN knows k ON k.a = c.b WHERE c.depth < 5 ) SELECT * FROM chain; ``` See [Query toolbox](/en/docs/core/queries) for more on CTEs. ## AI prompt: draft a graph query [#ai-prompt-draft-a-graph-query] ## Related [#related] * [PostgreSQL 19 release and upgrade guide](/en/docs/postgresql-19) — version boundary for SQL/PGQ * [Indexes and EXPLAIN](/en/docs/core/indexes-explain) — verifying the reverse-traversal index * [Query toolbox](/en/docs/core/queries) — CTEs and `WITH RECURSIVE` --- # Indexes and EXPLAIN Canonical URL: https://pg.edu.rich/en/docs/core/indexes-explain Last reviewed: 2026-08-06 ## Capture the plan first [#capture-the-plan-first] ```sql EXPLAIN (ANALYZE, BUFFERS, VERBOSE) SELECT id, customer_id, placed_at FROM orders WHERE customer_id = 42 ORDER BY placed_at DESC LIMIT 20; ``` * `EXPLAIN` shows estimates without executing. * `ANALYZE` executes and reports actual rows and timing. Wrap writes in a transaction and roll back. * `BUFFERS` reports shared/local/temp block hits and reads. * Compare estimated `rows` with `actual rows`, loop counts, expensive nodes, and disk sorts. `EXPLAIN ANALYZE DELETE ...` really deletes. Inspect a write with `BEGIN; EXPLAIN (ANALYZE, BUFFERS) ...; ROLLBACK;` only after confirming there are no non-transactional external effects. ## Index the query shape [#index-the-query-shape] The filter and ordering above can use: ```sql CREATE INDEX CONCURRENTLY orders_customer_placed_idx ON orders (customer_id, placed_at DESC) INCLUDE (id); ``` A multicolumn B-tree normally matches from its left side. Actual predicates, range conditions, and ordering determine column order—not a simplistic “most selective first” rule. `INCLUDE` columns do not participate in search ordering but may enable an index-only scan. The visibility map still determines whether heap access is avoidable. ## Common index types [#common-index-types] | Type | Fits | | --------------- | ------------------------------------------------------------------------------------------------------------------------- | | B-tree | Equality, ranges, ordering; the default | | GIN | `jsonb` containment, arrays, full-text search | | GiST | Geometry, ranges, and extension operators | | SP-GiST | Tries, quadtrees, k-d trees, and other partitioned search spaces | | BRIN | Huge tables whose physical order correlates with values, such as append-only time data | | Hash | Equality only; B-tree is usually more versatile | | Bloom extension | Equality across arbitrary combinations of many columns; lossy and rechecked; bundled classes only cover `int4` and `text` | These labels still do not prove an index is usable: the operator class determines exact operators and data types. See [index and storage access methods](/en/docs/reference/index-access-methods) for table access methods, HNSW/IVFFlat, and the full selection map. ## Two high-value patterns [#two-high-value-patterns] A partial index covers only relevant rows: ```sql CREATE INDEX orders_unfinished_idx ON orders (placed_at) WHERE status IN ('pending', 'paid'); ``` An expression index accelerates normalized lookup: ```sql CREATE UNIQUE INDEX customers_email_ci_idx ON customers (lower(email)); ``` The query predicate must match the expression or imply the partial condition for the planner to use it. ## Why an index is not used [#why-an-index-is-not-used] * The table is small and a sequential scan is cheaper. * The query returns a large fraction of rows. * Statistics are stale or miss cross-column correlation. * A function or implicit cast does not match the index expression. * The leftmost prefix of a multicolumn index is not usable. * Cost parameters do not reflect the storage system. Run `ANALYZE orders;` and inspect estimate errors before disabling sequential scans. ## Production creation and cleanup [#production-creation-and-cleanup] `CREATE INDEX CONCURRENTLY` reduces write blocking but takes longer, cannot run in a transaction block, and can leave an invalid index after failure. Inspect with: ```sql SELECT indexrelid::regclass, indisvalid, indisready FROM pg_index WHERE indrelid = 'orders'::regclass; ``` Every index adds write amplification, WAL, cache pressure, and vacuum work. Review unused indexes with `pg_stat_user_indexes` over a representative business cycle. ## Comparing plans across versions [#comparing-plans-across-versions] PostgreSQL 19 (Beta as of this update) adds an `IO` option to `EXPLAIN ANALYZE` that reports asynchronous I/O activity for scan nodes: prefetch-queue distance and capacity, plus issued I/O requests, their average size, I/O waits, and concurrency. `EXPLAIN (ANALYZE, WAL)` additionally reports full-page-write bytes. Both make storage behavior visible without external tooling; see the [PostgreSQL 19 EXPLAIN documentation](https://www.postgresql.org/docs/19/sql-explain.html). JIT is also disabled by default in 19 because its cost model proved unreliable. When comparing plans or runtimes between 18 and 19, set `jit` explicitly on both sides — otherwise a difference may reflect the changed default rather than an optimizer or index change. --- # JSONB, full-text, and semantic retrieval Canonical URL: https://pg.edu.rich/en/docs/core/jsonb-search Last reviewed: 2026-08-02 PostgreSQL can hold relational data and JSON documents, perform lexical full-text search, and add vector retrieval through an extension. Co-location does not mean every concern belongs in one column. ## When JSONB fits [#when-jsonb-fits] Good fits: metadata from heterogeneous sources, optional attributes that change infrequently, and integration payloads that must preserve their original shape. Poor fits: primary and foreign keys, money, authorization boundaries, and core fields used constantly for joins, ordering, or aggregation. ```sql CREATE TABLE products ( id bigint GENERATED ALWAYS AS IDENTITY PRIMARY KEY, sku text NOT NULL UNIQUE, name text NOT NULL, attributes jsonb NOT NULL DEFAULT '{}'::jsonb, CHECK (jsonb_typeof(attributes) = 'object') ); INSERT INTO products (sku, name, attributes) VALUES ('KB-01', 'Keyboard', '{"layout":"75%","wireless":true}'); SELECT id, name FROM products WHERE attributes @> '{"wireless":true}'; ``` ## JSONB indexes [#jsonb-indexes] ```sql CREATE INDEX products_attributes_gin ON products USING gin (attributes); ``` The default GIN operator class supports several key and containment operations. If the workload is almost entirely `@>`, `jsonb_path_ops` is often smaller but supports a narrower operator set. Compare with real queries and distributions. A frequently queried attribute can use an expression index—or graduate into a normal column: ```sql CREATE INDEX products_layout_idx ON products ((attributes ->> 'layout')); ``` ## Built-in full-text search [#built-in-full-text-search] ```sql ALTER TABLE products ADD COLUMN search_document tsvector GENERATED ALWAYS AS ( setweight(to_tsvector('simple', coalesce(name, '')), 'A') || setweight(to_tsvector('simple', coalesce(attributes::text, '')), 'B') ) STORED; CREATE INDEX products_search_gin ON products USING gin (search_document); SELECT id, name, ts_rank(search_document, websearch_to_tsquery('simple', $1)) AS rank FROM products WHERE search_document @@ websearch_to_tsquery('simple', $1) ORDER BY rank DESC LIMIT 20; ``` The built-in `simple` configuration does not fully solve Chinese tokenization. Production Chinese search needs a dedicated segmentation extension, application-side tokenization, or an external search system. ## Semantic retrieval with pgvector [#semantic-retrieval-with-pgvector] Vectors are not a PostgreSQL core type. A common approach is the independent `pgvector` extension: ```sql CREATE EXTENSION IF NOT EXISTS vector; CREATE TABLE document_chunks ( id bigint GENERATED ALWAYS AS IDENTITY PRIMARY KEY, document_id bigint NOT NULL, content text NOT NULL, embedding vector(1536) NOT NULL, embedding_model text NOT NULL ); ``` Dimensions must match the model. Do not mix incomparable embedding models in one index. Store model name, chunking version, and source location so retrieval can be rebuilt and audited. Reduce candidates with keywords, permissions, tenant, and time filters before vector ranking. Enforce access filters in SQL; never rely on the model to remember them. Start with [Install pgvector](/en/docs/ai/pgvector-setup) for a working environment, then read [pgvector production practices](/en/docs/ai/vector-production) before adding approximate indexes. --- # Diagnose PostgreSQL lock waits and deadlocks Canonical URL: https://pg.edu.rich/en/docs/core/locks-deadlocks Last reviewed: 2026-08-02 ## Lock wait versus deadlock [#lock-wait-versus-deadlock] * **Lock wait**: a session waits for another transaction to release a conflicting lock; it may eventually succeed or time out. * **Deadlock**: sessions form a wait cycle and none can progress; PostgreSQL detects it and aborts one transaction. The deadlock SQLSTATE is `40P01`. The transaction must roll back; perform a bounded retry only when the entire business transaction is safe to replay. ## Inspect the blocking graph [#inspect-the-blocking-graph] ```sql SELECT blocked.pid AS blocked_pid, blocker.pid AS blocker_pid, now() - blocked.query_start AS blocked_for, blocked.wait_event_type, blocked.wait_event, left(blocked.query, 120) AS blocked_query, left(blocker.query, 120) AS blocker_query FROM pg_stat_activity AS blocked CROSS JOIN LATERAL unnest(pg_blocking_pids(blocked.pid)) AS b(pid) JOIN pg_stat_activity AS blocker ON blocker.pid = b.pid ORDER BY blocked.query_start; ``` Check whether the blocker is `idle in transaction`, what it changed, and whether the application is alive. Do not terminate a PID on sight. ## Reduce deadlocks [#reduce-deadlocks] 1. Lock resources in the same order on every code path, such as ascending account id. 2. Keep only database work that must be atomic inside the transaction; exclude user input and external APIs. 3. Index predicates used to locate rows for update, reducing scan and lock scope. 4. Chunk bulk work and keep concurrent workers out of overlapping key ranges. 5. Set evidence-based `lock_timeout` and `statement_timeout` values. ```sql BEGIN; SET LOCAL lock_timeout = '1s'; SET LOCAL statement_timeout = '10s'; SELECT id FROM accounts WHERE id = ANY($1::bigint[]) ORDER BY id FOR UPDATE; -- bounded writes COMMIT; ``` ## Before terminating a session [#before-terminating-a-session] `pg_cancel_backend(pid)` requests cancellation of the current statement. `pg_terminate_backend(pid)` ends the session and rolls back its transaction. Confirm that: * the PID still belongs to the target session, not an old screenshot; * rollback may take time and generate additional I/O; * the application will not reconnect and repeat the same blocker immediately; * interrupted work is retryable or has a business compensation path. See [PostgreSQL 18 explicit locking](https://www.postgresql.org/docs/18/explicit-locking.html) for lock modes and the conflict matrix. --- # PostgreSQL MVCC and snapshot visibility Canonical URL: https://pg.edu.rich/en/docs/core/mvcc-snapshots Last reviewed: 2026-08-02 MVCC (multi-version concurrency control) lets ordinary reads avoid blocking writes. An `UPDATE` does not overwrite the value in place for every reader; it creates a new row version, and each query uses its snapshot to decide which version is visible. ## Observe snapshots in two sessions [#observe-snapshots-in-two-sessions] Prepare data: ```sql CREATE TABLE mvcc_demo ( id integer PRIMARY KEY, value text NOT NULL ); INSERT INTO mvcc_demo VALUES (1, 'before'); ``` Session A: ```sql BEGIN ISOLATION LEVEL REPEATABLE READ; SELECT value FROM mvcc_demo WHERE id = 1; -- before ``` Session B: ```sql UPDATE mvcc_demo SET value = 'after' WHERE id = 1; COMMIT; ``` Back in session A: ```sql SELECT value FROM mvcc_demo WHERE id = 1; -- still before COMMIT; SELECT value FROM mvcc_demo WHERE id = 1; -- after ``` At `READ COMMITTED`, each statement gets a new snapshot, so a second query in the same transaction can observe session B's committed value. ## Why long transactions hurt [#why-long-transactions-hurt] While an old snapshot can still see row versions, vacuum cannot treat them as fully reclaimable. Long transactions therefore increase: * dead tuples and table/index bloat; * vacuum work and disk use; * retention pressure from slots, logical decoding, or standbys; * the risk window around transaction ID wraparound. Find sessions holding old transactions or snapshots: ```sql SELECT pid, usename, application_name, state, now() - xact_start AS transaction_age, age(backend_xmin) AS snapshot_xid_age, wait_event_type, wait_event, left(query, 120) AS query FROM pg_stat_activity WHERE xact_start IS NOT NULL OR backend_xmin IS NOT NULL ORDER BY xact_start NULLS LAST; ``` Do not terminate a session merely because it is old. Confirm workload purpose, backup/maintenance activity, retry safety, and termination impact. Row versions solve read visibility. Write conflicts, DDL, foreign-key checks, and explicit locks can still wait. Continue with [lock waits and deadlocks](/en/docs/core/locks-deadlocks). Use [PostgreSQL 18 concurrency control](https://www.postgresql.org/docs/18/mvcc.html) and [transaction isolation](https://www.postgresql.org/docs/18/transaction-iso.html) as authoritative references. --- # Query toolbox Canonical URL: https://pg.edu.rich/en/docs/core/queries Last reviewed: 2026-08-06 ## Shape of a maintainable query [#shape-of-a-maintainable-query] ```sql SELECT o.id, c.email, o.total_cents, o.placed_at FROM orders AS o JOIN customers AS c ON c.id = o.customer_id WHERE o.status = $1 AND o.placed_at >= $2 ORDER BY o.placed_at DESC, o.id DESC LIMIT $3; ``` This query specifies output, join condition, parameters, deterministic ordering, and a bound. The driver binds `$1`, `$2`, and `$3`; never concatenate user input into SQL. ## Choosing a join form [#choosing-a-join-form] | Goal | Use | | ---------------------------------- | ------------------------------------------------- | | Keep matching rows from both sides | `INNER JOIN` / `JOIN` | | Keep every row on the left | `LEFT JOIN` | | Test whether a related row exists | `EXISTS`, often clearer than join-plus-`DISTINCT` | | Find rows without a relation | `NOT EXISTS`, avoiding `NOT IN` null traps | ```sql SELECT c.id, c.email FROM customers AS c WHERE NOT EXISTS ( SELECT 1 FROM orders AS o WHERE o.customer_id = c.id ); ``` ## Aggregates and windows differ [#aggregates-and-windows-differ] `GROUP BY` collapses rows. Window functions retain detail rows while calculating across a window. ```sql SELECT customer_id, id AS order_id, total_cents, row_number() OVER ( PARTITION BY customer_id ORDER BY placed_at DESC, id DESC ) AS recency_rank, sum(total_cents) OVER (PARTITION BY customer_id) AS lifetime_cents FROM orders; ``` ## What CTEs are for [#what-ctes-are-for] A CTE names a stage in a complex query; it is not an automatic optimization switch. ```sql WITH recent_paid AS ( SELECT customer_id, total_cents FROM orders WHERE status = 'paid' AND placed_at >= now() - interval '30 days' ) SELECT customer_id, sum(total_cents) AS paid_cents FROM recent_paid GROUP BY customer_id; ``` ## Pagination [#pagination] Prefer keyset pagination for large result sets: ```sql SELECT id, placed_at, total_cents FROM orders WHERE (placed_at, id) < ($1, $2) ORDER BY placed_at DESC, id DESC LIMIT 50; ``` Unlike a large `OFFSET`, this does not repeatedly skip earlier rows and behaves more predictably under concurrent inserts. The cursor must contain every ordering key. ## PostgreSQL 19 syntax conveniences [#postgresql-19-syntax-conveniences] PostgreSQL 19 (Beta as of this update) plans several small syntax additions. Do not send them to an older server; verify against the final release notes. `GROUP BY ALL` groups by every non-aggregate, non-window target-list item, so the grouping list is not repeated: ```sql SELECT customer_id, status, count(*) FROM orders GROUP BY ALL; ``` Window functions accept `IGNORE NULLS` / `RESPECT NULLS` for `lead()`, `lag()`, `first_value()`, `last_value()`, and `nth_value()`: ```sql SELECT customer_id, placed_at, lag(placed_at) IGNORE NULLS OVER ( PARTITION BY customer_id ORDER BY placed_at ) AS previous_order_at FROM orders; ``` `INSERT ... ON CONFLICT DO SELECT ... RETURNING` makes get-or-create a single atomic statement: the row is either inserted or the conflicting existing row is returned. A `conflict_target` and a `RETURNING` clause are both required for `DO SELECT`, and an optional locking clause (`FOR UPDATE`, `FOR NO KEY UPDATE`, `FOR SHARE`, `FOR KEY SHARE`) locks the conflicting row against concurrent updates. ```sql INSERT INTO customers (email) VALUES ($1) ON CONFLICT (email) DO SELECT RETURNING id; ``` See the [PostgreSQL 19 INSERT documentation](https://www.postgresql.org/docs/19/sql-insert.html). ## Pre-flight checklist [#pre-flight-checklist] * Is the output contract explicit and minimal? * Can any join multiply rows? * Is `NULL` semantics intentional? * Does ordering have a unique final tie-breaker? * Does the driver bind every input value? * Do you need statement timeout and a result limit? --- # Transactions, MVCC, and concurrency Canonical URL: https://pg.edu.rich/en/docs/core/transactions Last reviewed: 2026-08-02 ## A minimal transaction [#a-minimal-transaction] ```sql BEGIN; SELECT balance_cents FROM accounts WHERE id = $1 FOR UPDATE; UPDATE accounts SET balance_cents = balance_cents - $2 WHERE id = $1 AND balance_cents >= $2; COMMIT; ``` A transaction makes several statements one atomic unit. `FOR UPDATE` locks selected rows until commit or rollback; the application must still inspect the update count. ## MVCC intuition [#mvcc-intuition] Multi-version concurrency control means readers normally do not block writers and writers normally do not block ordinary readers. Updates create new row versions. Autovacuum can reclaim old versions after no active snapshot can see them. Long transactions delay that cleanup and increase bloat, WAL retention, and replication-lag risk. Never hold a transaction open across user think time, network retries, or unrelated external API calls. ## Isolation levels [#isolation-levels] | Level | PostgreSQL behavior | Application duty | | ----------------- | -------------------------------------------------------------- | -------------------------------------------------------- | | `READ COMMITTED` | Default; each statement gets a fresh snapshot | Do not assume two reads in one transaction are identical | | `REPEATABLE READ` | Stable transaction snapshot; serialization failure is possible | Retry the whole transaction on `40001` | | `SERIALIZABLE` | Commits only outcomes proven equivalent to serial execution | Implement bounded whole-transaction retry and backoff | PostgreSQL treats `READ UNCOMMITTED` as `READ COMMITTED`. Beyond the SQL standard's minimum, PostgreSQL `REPEATABLE READ` also prevents phantom reads, but serialization failures can still require a whole-transaction retry. Sequence changes such as `nextval()` are not rolled back, so gaps are normal. ## Correct retry boundary [#correct-retry-boundary] After serialization failure or deadlock, the current transaction cannot continue. Roll it back and replay the **whole transaction**, not just the final SQL statement. ```text begin run all reads and writes commit on SQLSTATE 40001 or 40P01 rollback retry whole unit with bounded exponential backoff ``` External effects inside the retry boundary must be idempotent. A transactional outbox is often safer for post-commit work. ## Deadlocks and waits [#deadlocks-and-waits] Reduce deadlocks by locking resources in a consistent order, keeping transactions short, indexing lookup predicates, and setting appropriate `lock_timeout` and `statement_timeout` values. ```sql SET LOCAL lock_timeout = '2s'; SET LOCAL statement_timeout = '10s'; ``` `SET LOCAL` applies only to the current transaction. A client that begins and never ends a transaction retains a snapshot and perhaps locks. Monitor `pg_stat_activity.state = 'idle in transaction'` and consider `idle_in_transaction_session_timeout`. Use the [PostgreSQL 18 transaction-isolation documentation](https://www.postgresql.org/docs/18/transaction-iso.html) as the authoritative behavior reference. Continue with [MVCC and snapshot visibility](/en/docs/core/mvcc-snapshots) for old row versions and long transactions, then [lock waits and deadlocks](/en/docs/core/locks-deadlocks) for blocking diagnostics and safe intervention. --- # Troubleshoot PostgreSQL connection errors Canonical URL: https://pg.edu.rich/en/docs/reference/connection-errors Last reviewed: 2026-08-02 When failure happens before SQL execution, the client may not receive SQLSTATE. Preserve the complete error, timestamp, client version, and target host/port—but never the password. ## Fixed diagnostic order [#fixed-diagnostic-order] ```text DNS resolution → TCP route/firewall/listening port → TLS negotiation and certificate identity → pg_hba.conf match → user authentication → database and CONNECT privilege → instance/pool connection capacity → session initialization settings ``` Resetting passwords or widening privileges before proving the previous layer usually hides the real cause. ## Frequent errors [#frequent-errors] | Error | Meaning | Verify | | -------------------------------- | ----------------------------------------------------------------------- | ------------------------------------------------------------------ | | `could not translate host name` | DNS/hostname cannot resolve | `getent hosts`, `nslookup`, spelling, and private DNS | | `connection refused` | Nothing accepts the target address/port | Service state, `listen_addresses`, port, and container mapping | | `connection timed out` | Network path or firewall drops traffic | Test TCP from the application environment, not a laptop substitute | | `no pg_hba.conf entry` | No rule matches source/database/user/TLS | Inspect server log and rule order; reload after editing | | `password authentication failed` | Credential or authentication method mismatch, commonly SQLSTATE `28P01` | Confirm target instance/user and rotate securely | | `database ... does not exist` | Database absent on this instance, SQLSTATE `3D000` | Connect to `postgres` and inspect `pg_database` | | `too many connections` | Instance/role/database limit exhausted, SQLSTATE `53300` | `pg_stat_activity`, pool size, and reserved administration access | | `certificate verify failed` | CA, hostname, validity, or chain mismatch | `sslmode`, URI host, CA file, and provider rotation notice | ## Client verification [#client-verification] ```bash psql --version psql -X "postgresql://app_reader@db.example.com:5432/commerce?sslmode=verify-full" ``` Immediately after success: ```sql \conninfo SELECT current_database(), current_user, inet_server_addr(), inet_server_port(), current_setting('server_version'); ``` ## Minimal server-side checks [#minimal-server-side-checks] ```sql SELECT datname, datallowconn, datconnlimit FROM pg_database ORDER BY datname; SELECT usename, application_name, client_addr, state, count(*) FROM pg_stat_activity GROUP BY usename, application_name, client_addr, state ORDER BY count(*) DESC; ``` Use OS access only when needed to inspect listening sockets, firewalls, and PostgreSQL logs. For managed databases, use provider connection diagnostics, network flow logs, and audit logs. Changing `pg_hba.conf` to `trust` removes the authentication boundary and does not explain the original failure. Rotate credentials through a controlled channel and identify the matching HBA rule and server log entry. For a secure successful connection, see [`psql` and SSL](/en/docs/setup/psql-connection). For SQL execution failures, use the [SQLSTATE fieldbook](/en/docs/reference/errors). --- # Editorial and verification policy Canonical URL: https://pg.edu.rich/en/docs/reference/editorial-policy Last reviewed: 2026-08-02 ## Accountability [#accountability] PostgreSQL Field Guide is an independent community knowledge base. It is not affiliated with the PostgreSQL Global Development Group or the cloud vendors it covers. The site provides learning paths, engineering explanations, and verifiable examples; the official documentation for the target PostgreSQL version remains normative. ## Source priority [#source-priority] 1. Official manuals, release notes, and versioning pages for supported PostgreSQL versions. 2. Upstream repositories and release notes for extensions such as pgvector. 3. Vendor documentation bound to a specific product, region, and engine version. 4. Reproducible local tests and public technical standards. Community articles can help discover questions, but never stand alone for version, API, security, or recovery claims. ## Publication checks [#publication-checks] * Verify SQL names, catalog views, and settings against the target-version manual instead of recalling interfaces. * Run safely executable examples against an ephemeral PostgreSQL 18 instance where practical. * State risks and prerequisites for writes, disruptive locks, recovery, privilege, and replication. * Publish matching Chinese and English paths; neither language may be an empty placeholder or unreviewed machine translation. * Date cloud-service claims and require readers to re-check region, SKU, and extension versions. * Pass type checking, lint, production build, internal-link validation, and critical HTTP/SEO checks. ## AI assistance disclosure [#ai-assistance-disclosure] AI may assist research organization, translation drafts, example review, and consistency checks. It is not a factual source. Database-behavior claims must resolve to official material or reproducible tests; business semantics, risk acceptance, and production changes remain human decisions. ## Dates and corrections [#dates-and-corrections] The “Last updated” marker represents the latest content-review date. Corrections should update both languages, cross-links, and machine-readable Markdown together, with a traceable note in the project's content audit record. Current site-wide review baseline: **2026-08-02**. --- # SQLSTATE error fieldbook Canonical URL: https://pg.edu.rich/en/docs/reference/errors Last reviewed: 2026-08-02 Applications branch on **SQLSTATE**, not error text that can vary by version and locale. ## Frequent states [#frequent-states] | SQLSTATE | Name | Common meaning | Safe action | | -------- | ----------------------------- | ---------------------------------------- | ------------------------------------------------------- | | `23505` | unique\_violation | Unique key conflict | Return conflict or use explicit `ON CONFLICT` semantics | | `23503` | foreign\_key\_violation | Target absent or still referenced | Fix operation order; do not disable the constraint | | `23502` | not\_null\_violation | Required column absent | Fix input or migration order | | `23514` | check\_violation | `CHECK` rejected the row | Explain boundary and correct the value | | `22P02` | invalid\_text\_representation | Type conversion failed | Validate at the boundary and bind the right type | | `40001` | serialization\_failure | Isolation guarantee cannot be maintained | Roll back and retry the entire transaction | | `40P01` | deadlock\_detected | Wait cycle | Roll back whole transaction; normalize lock order | | `55P03` | lock\_not\_available | `NOWAIT` or lock timeout | Retry later or return conflict | | `57014` | query\_canceled | Statement timeout or cancellation | Distinguish cancel from timeout; optimize or narrow | | `25P02` | in\_failed\_sql\_transaction | Earlier statement failed | `ROLLBACK`; do not continue business SQL | | `42501` | insufficient\_privilege | Action/object not granted | Fix grant/owner; do not jump to superuser | | `42P01` | undefined\_table | Missing table or wrong search path | Check database/schema/migration version | | `42703` | undefined\_column | Column absent | Check schema contract and deploy version | | `53300` | too\_many\_connections | Connection slots exhausted | Inspect pools, leaks, and reserved admin access | | `57P03` | cannot\_connect\_now | Starting, recovering, or shutting down | Bounded backoff and instance inspection | | `08006` | connection\_failure | Connection failed | Determine unknown commit state before safe retry | ## After a transaction error [#after-a-transaction-error] An error inside a transaction normally leaves it aborted: ```text ERROR: current transaction is aborted... SQLSTATE: 25P02 ``` Issue `ROLLBACK`, or roll back to a savepoint created before the failure. More statements do not repair it automatically. ## Retry classes [#retry-classes] * **Retry the whole transaction**: `40001`, `40P01`; bounded attempts, exponential backoff, and jitter. * **Possibly transient**: `55P03`, `57P03`, selected `08***`; verify idempotency and unknown commit state. * **Input/model errors—do not blind retry**: `22***`, `23***`, `42***`, `42501`. * **Resource conditions**: `53300`, disk full, memory pressure; retries amplify the incident. Shed load and repair capacity. ## Diagnostic context [#diagnostic-context] Record SQLSTATE, constraint/table/column fields, database and schema, application and migration versions, transaction/request id, parameter types with sensitive values redacted, and known commit state. Drivers usually expose structured fields; use them directly. ## AI tool response [#ai-tool-response] ```json { "ok": false, "sqlstate": "23505", "category": "constraint", "retryable": false, "constraint": "customers_email_unique", "message_safe": "A customer with this email already exists" } ``` Do not return raw database errors to end users without filtering. They may reveal object names, paths, or data fragments. --- # PostgreSQL extension selection guide Canonical URL: https://pg.edu.rich/en/docs/reference/extensions-ecosystem Last reviewed: 2026-08-06 PostgreSQL extensions can place types, indexes, planner hooks, background workers, and storage behavior inside the database process. They also enter the critical path for backup, replication, recovery, and major upgrades. The selection rule is simple: **do not install an extension without a defined workload and exit path**. ## Six gates before installation [#six-gates-before-installation] 1. The target PostgreSQL major, operating system, and CPU architecture have explicit support and packages. 2. The license fits self-hosting, SaaS, redistribution, and commercial-feature boundaries. 3. Backup, PITR, standbys, logical replication, and recovery environments can load the same version. 4. `pg_upgrade`, extension update, and any required index rebuild have a rehearsed path. 5. The managed-cloud region, SKU, and allowlist provide the required version, not merely an extension with the same name. 6. An export or migration path exists without the extension, limiting accidental platform lock-in. Record `SELECT extname, extversion FROM pg_extension`, and place extension versions beside the database version in deployment manifests and AI context. First check native B-tree/GIN/GiST/SP-GiST/BRIN, full-text search, partitions, FDWs, and materialized views in [PostgreSQL index and storage access methods](/en/docs/reference/index-access-methods). Add an extension only when native behavior fails a measured workload requirement. ## Choose by workload [#choose-by-workload] | Workload | Common candidate | Adoption boundary | | ---------------------------- | ------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------ | | SQL statistics | [`pg_stat_statements`](https://www.postgresql.org/docs/current/pgstatstatements.html) | Official contrib and an observability baseline; govern query-text access | | Vector search / RAG | [pgvector](https://github.com/pgvector/pgvector) | Measure recall, latency, memory, and index build with real filters | | Geospatial | [PostGIS](https://postgis.net/) | Standard GIS choice; verify extension and data-format upgrades | | Time series | [TimescaleDB](https://github.com/timescale/timescaledb) | Evaluate for hypertables, compression, or continuous aggregates; check licensing feature by feature | | Distributed multi-tenancy | [Citus](https://github.com/citusdata/citus) | Add only after measuring a single-node bottleneck and stabilizing a shard key | | BM25 / search | [ParadeDB / pg\_search](https://github.com/paradedb/paradedb) | Verify license, index recovery, replication, and cloud support | | In-PostgreSQL BM25 | [pg\_textsearch](https://github.com/timescale/pg_textsearch) | Upstream currently calls it production ready; independently validate target version and corpus | | Embedded analytics / Parquet | [pg\_duckdb](https://github.com/duckdb/pg_duckdb) | Fit for analytics paths; test transaction boundaries, resource isolation, and object-store credentials | | Iceberg columnstore mirror | [pg\_mooncake](https://github.com/Mooncake-Labs/pg_mooncake) | Maintains a columnstore mirror from logical changes; validate consistency, object storage, pg\_duckdb dependency, and recovery | | Graph queries | [Apache AGE](https://github.com/apache/age) | Adopt only when a graph model and Cypher produce measured value | “Production ready” is an upstream project status, not a certification for your workload, SLA, or cloud platform. ## TimescaleDB quickstart and version boundaries [#timescaledb-quickstart-and-version-boundaries] The adoption boundary above applies first: evaluate TimescaleDB only when hypertables, compression, or continuous aggregates answer a measured requirement. A minimal evaluation sequence on TimescaleDB 2.x: ```sql CREATE EXTENSION IF NOT EXISTS timescaledb; CREATE TABLE metrics ( time timestamptz NOT NULL, device text NOT NULL, value double precision ); SELECT create_hypertable('metrics', by_range('time')); CREATE MATERIALIZED VIEW metrics_hourly WITH (timescaledb.continuous) AS SELECT time_bucket('1 hour', time) AS hour, device, avg(value), max(value) FROM metrics GROUP BY hour, device; ALTER TABLE metrics SET (timescaledb.compress); SELECT add_compression_policy('metrics', INTERVAL '7 days'); SELECT add_retention_policy('metrics', INTERVAL '90 days'); ``` * The `by_range` dimension builder requires TimescaleDB 2.13 or later; earlier releases use `create_hypertable('metrics', 'time')`. * `ALTER TABLE ... SET (timescaledb.compress)` is the legacy compression API. Recent 2.x releases route compression through hypercore with different reloptions and policy names — confirm the syntax against the installed version and the TimescaleDB documentation. * A continuous aggregate does not refresh on a schedule until `add_continuous_aggregate_policy` is attached. * TimescaleDB is distributed under the Timescale License (TSL), not the PostgreSQL License. Check feature-by-feature boundaries before production use. ## Maintenance and data governance [#maintenance-and-data-governance] | Tool | Purpose | Do not mistake it for | | ------------------------------------------------------------------------ | ----------------------------------------------------------------- | ----------------------------------------------------------- | | [HypoPG](https://github.com/HypoPG/hypopg) | Evaluate planner choices with hypothetical indexes | Proof of build cost or production benefit | | [pg\_repack](https://github.com/reorg/pg_repack) | Reorganize tables and indexes with shorter exclusive-lock windows | A replacement for routine autovacuum | | [pg\_partman](https://github.com/pgpartman/pg_partman) | Manage native time/serial partition lifecycles | An automatic fix for a poor partition key | | [pg\_cron](https://github.com/citusdata/pg_cron) | Schedule simple SQL inside PostgreSQL | A general business queue or workflow engine | | [Greenmask](https://github.com/GreenmaskIO/greenmask) | Produce masked and subsetted test data | Permission to copy sensitive production data without review | | [PostgreSQL Anonymizer](https://gitlab.com/dalibo/postgresql_anonymizer) | Declarative static and dynamic masking | Automatic compliance with every requirement | For severe bloat, first find long transactions, autovacuum, write patterns, and fillfactor causes before using pg\_repack. Partitioning helps only where lifecycle and queries can use the partition key. ## PostgreSQL 19 REPACK is not pg\_repack [#postgresql-19-repack-is-not-pg_repack] The core [`REPACK`](https://www.postgresql.org/docs/19/sql-repack.html) in PostgreSQL 19 Beta is a new SQL command. `REPACK (CONCURRENTLY)` uses logical decoding and has constraints around primary keys or replica identity, unlogged/partitioned/system tables, replication slots, and disk space. Third-party **pg\_repack** is an independent extension and command-line tool with its own compatibility matrix, packages, and operational boundaries. Similar names do not make pg\_repack experience, monitoring, or risks directly transferable to core PostgreSQL 19 `REPACK`. As of 2026-08-02, PostgreSQL 19 remains Beta 2. Validate core REPACK semantics and constraints against final GA documentation and a restored copy of your own data. ## PostgreSQL 19 planner-advice modules [#postgresql-19-planner-advice-modules] PostgreSQL 19 Beta adds two contrib modules aimed at plan stabilization. [`pg_plan_advice`](https://www.postgresql.org/docs/19/pgplanadvice.html) lets key planner decisions be described, reproduced, and altered through advice attached to a query; [`pg_stash_advice`](https://www.postgresql.org/docs/19/pgstashadvice.html) stores advice strings in dynamic shared memory keyed by query identifier and applies them automatically. Treat both as an experimental channel, not a production plan-management contract: the advice format and coverage can still change before GA, and pinning plans this way bypasses the same re-validation a rerun of `ANALYZE` or an extension like HypoPG would give you. Evaluate them in the Beta lane alongside your plan-comparison benchmarks, not as a reason to skip them. ## AI, backup, and upgrade checklist [#ai-backup-and-upgrade-checklist] Provide AI agents with `server_version_num`, `extname/extversion`, allowed operators and index methods, cloud limitations, and forbidden syntax. “This is PostgreSQL” is insufficient context. Before every extension upgrade: 1. Read target release notes, SQL update scripts, and known rebuild requirements. 2. Restore a real backup into an isolated environment. 3. Upgrade PostgreSQL and the extension, then run integrity, performance, and RLS tests. 4. Rebuild required indexes and compare plans, recall, or business results. 5. Create a new backup and restore it once to prove the new-version chain. Supabase, Neon, YugabyteDB, CockroachDB, Cloudberry, Gel, and FerretDB do not belong in an extension ranking. They are platforms, forks, independent databases, or protocol translation layers; classify them with [PostgreSQL lineage and compatible databases](/en/docs/reference/postgresql-compatible-databases). --- # PostgreSQL index access methods Canonical URL: https://pg.edu.rich/en/docs/reference/index-access-methods Last reviewed: 2026-08-02 PostgreSQL does not use the everyday MySQL model of selecting InnoDB or MyISAM for ordinary tables. Almost every table uses the core **heap table access method**. Separate index access methods and operator classes determine which queries an index supports. ## A table access method is not a routine tuning switch [#a-table-access-method-is-not-a-routine-tuning-switch] The PostgreSQL [Table Access Method API](https://www.postgresql.org/docs/current/tableam.html) lets extensions or custom builds implement table storage, but ordinary applications still default to `heap`: ```sql CREATE TABLE events ( id bigint GENERATED ALWAYS AS IDENTITY PRIMARY KEY, occurred_at timestamptz NOT NULL, payload jsonb NOT NULL ) USING heap; ``` `USING heap` is normally omitted. A new table access method enters the critical path for WAL, MVCC, VACUUM, backup, replication, extensions, and major upgrades. It is not a query hint that can be switched casually. Inspect access methods exposed by an instance: ```sql SELECT amname, CASE amtype WHEN 't' THEN 'table' WHEN 'i' THEN 'index' ELSE amtype::text END AS access_method_type FROM pg_am ORDER BY amtype, amname; ``` ## Core index access methods [#core-index-access-methods] PostgreSQL 18 core provides B-tree, Hash, GiST, SP-GiST, GIN, and BRIN. `bloom` ships as a module but requires `CREATE EXTENSION bloom`. Use the official [Index Types](https://www.postgresql.org/docs/current/indexes-types.html) as the behavior boundary. | Type | Prefer for | Critical boundary | | --------------- | ------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------- | | B-tree | Equality, ranges, ordering, uniqueness, anchored patterns | Default choice; column order and operator class determine usable queries | | Hash | Single-column equality | Supports only `=`; B-tree is usually more versatile, so require measured benefit | | GIN | JSONB, arrays, full text, and multi-valued content | Higher update/build cost; behavior depends on the operator class | | GiST | Ranges, geometry, PostGIS, and nearest-neighbor search | An extensible framework, not one algorithm; operators/classes must match | | SP-GiST | Tries, quadtrees, k-d trees, and partitioned search spaces | Fits naturally partitionable data; not a general GiST replacement | | BRIN | Very large append-heavy tables correlated with physical order | Stores block-range summaries; weak correlation reads many heap blocks | | Bloom extension | Equality across arbitrary combinations of many columns | Lossy and rechecked; no range, unique, or `NULL` search; bundled operator classes cover only `int4` and `text` | GIN, GiST, SP-GiST, and BRIN are frameworks. The operator class determines supported operators, ordering, and data types. “Uses GIN” is not enough information to reproduce an index design. ## Common workload map [#common-workload-map] ```text Equality / range / order / unique → B-tree JSONB contains / array member → GIN PostgreSQL full-text search → GIN in most cases Range / GIS / nearest neighbor → GiST or a matching SP-GiST class Huge, time-correlated append table → BRIN Vector approximate-nearest-neighbor → pgvector HNSW / IVFFlat ``` HNSW and IVFFlat come from [pgvector](https://github.com/pgvector/pgvector); they are not core PostgreSQL index methods. Test recall, filtering, memory, build time, WAL, replica lag, and extension upgrades independently. Built-in full-text search includes parsers, dictionaries, ranking, highlighting, and GIN/GiST indexing. Built-in configurations do not solve tokenization for every language. Chinese, for example, normally needs an additional tokenizer/extension or application preprocessing; creating a GIN index alone does not prove search quality. ## Match the query before creating an index [#match-the-query-before-creating-an-index] ```sql -- Ordinary filter and order CREATE INDEX CONCURRENTLY orders_customer_time_idx ON orders (customer_id, placed_at DESC); -- JSONB containment: payload @> '{"status":"paid"}' CREATE INDEX CONCURRENTLY events_payload_gin_idx ON events USING gin (payload jsonb_path_ops); -- Large table whose timestamps correlate with physical append order CREATE INDEX CONCURRENTLY events_time_brin_idx ON events USING brin (occurred_at); ``` `jsonb_path_ops` is focused on containment and jsonpath operators such as `@>`, `@?`, and `@@`; it does not support every operator available from the default `jsonb_ops`. Derive index DDL from real query shapes. Record plans and size after creation: ```sql SELECT indexrelname, idx_scan, pg_size_pretty(pg_relation_size(indexrelid)) AS index_size FROM pg_stat_user_indexes WHERE relname = 'events' ORDER BY pg_relation_size(indexrelid) DESC; ``` Then use `EXPLAIN (ANALYZE, BUFFERS)` to compare actual rows, heap blocks, rechecks, sorting, and write cost. See [Indexes and EXPLAIN](/en/docs/core/indexes-explain) for the complete workflow. ## Prefer native capability first [#prefer-native-capability-first] | Need | Validate in PostgreSQL first | Evaluate only when insufficient | | --------------------- | ------------------------------------------------------ | ---------------------------------------- | | Fuzzy search | FTS, `pg_trgm` contrib, expression/GIN/GiST indexes | External search or a BM25 extension | | Work claiming | Transactions, `FOR UPDATE SKIP LOCKED`, advisory locks | Dedicated queue and workflow systems | | Time lifecycle | Native partitions, BRIN, scheduled cleanup | pg\_partman or TimescaleDB | | Cross-database access | `postgres_fdw`, logical replication | CDC platform or separate sync system | | Analytics | Materialized views, partitioning, parallel query | pg\_duckdb, pg\_mooncake, or a warehouse | | Vector search | No core vector type or ANN index | pgvector or a dedicated vector system | Native-first does not reject extensions. It avoids unnecessary binary, license, backup, and upgrade dependencies. See [PostgreSQL extension selection](/en/docs/reference/extensions-ecosystem) for candidates. --- # PostgreSQL field reference Canonical URL: https://pg.edu.rich/en/docs/reference Last reviewed: 2026-08-02 ## psql [#psql] ```bash psql 'postgresql://user@host:5432/database?sslmode=verify-full' psql -X --set ON_ERROR_STOP=on --file migration.sql "$DATABASE_URL" ``` | Command | Purpose | | ---------------- | ----------------------------------------------- | | `\conninfo` | Current connection | | `\l` | Databases | | `\dn` | Schemas | | `\dt app.*` | Tables | | `\d+ app.orders` | Object definition and storage detail | | `\du` | Roles | | `\dx` | Extensions | | `\timing on` | Client-observed duration | | `\x auto` | Expanded output for wide results | | `\gdesc` | Describe result columns without displaying rows | | `\q` | Quit | Scripts use `-X` to ignore a user's `.psqlrc` and `ON_ERROR_STOP` to exit at the first error. ## Current context [#current-context] ```sql SELECT version(), current_database(), current_user, session_user, current_schema(), current_setting('TimeZone') AS timezone, inet_server_addr(), inet_server_port(); ``` ## Object size [#object-size] ```sql SELECT relname, pg_size_pretty(pg_total_relation_size(relid)) AS total, pg_size_pretty(pg_relation_size(relid)) AS heap, pg_size_pretty(pg_indexes_size(relid)) AS indexes FROM pg_catalog.pg_statio_user_tables ORDER BY pg_total_relation_size(relid) DESC LIMIT 20; ``` ## Active sessions and old transactions [#active-sessions-and-old-transactions] ```sql SELECT pid, usename, application_name, state, now() - xact_start AS xact_age, wait_event_type, wait_event, left(query, 160) AS query FROM pg_stat_activity WHERE pid <> pg_backend_pid() ORDER BY xact_start NULLS LAST; ``` ## Blocking graph [#blocking-graph] ```sql SELECT blocked.pid AS blocked_pid, blocker.pid AS blocker_pid, now() - blocked.query_start AS blocked_for, left(blocked.query, 120) AS blocked_query, left(blocker.query, 120) AS blocker_query FROM pg_stat_activity AS blocked CROSS JOIN LATERAL unnest(pg_blocking_pids(blocked.pid)) AS b(pid) JOIN pg_stat_activity AS blocker ON blocker.pid = b.pid; ``` Do not immediately call `pg_terminate_backend` on a blocker. Identify the workload, transaction, retry behavior, and termination impact first. ## Safe session settings [#safe-session-settings] ```sql BEGIN; SET LOCAL statement_timeout = '10s'; SET LOCAL lock_timeout = '2s'; SET LOCAL search_path = app, pg_catalog; -- work COMMIT; ``` ## Diagnostic order [#diagnostic-order] Confirm target and role → record SQLSTATE → inspect transaction state → inspect waits/blockers → capture plan and statistics → reproduce safely → run a verification query after the fix. Before a session exists, use [PostgreSQL connection troubleshooting](/en/docs/reference/connection-errors). For SQL failures after connection, use the [error and SQLSTATE fieldbook](/en/docs/reference/errors). For upgrade boundaries around extensions, maintenance tools, and open-source components, use the [PostgreSQL extensions and ecosystem guide](/en/docs/reference/extensions-ecosystem). See [PostgreSQL index and storage access methods](/en/docs/reference/index-access-methods) for storage concepts, and [PostgreSQL lineage and compatible databases](/en/docs/reference/postgresql-compatible-databases) for forks and protocol compatibility. --- # PostgreSQL 17 features and upgrade notes Canonical URL: https://pg.edu.rich/en/docs/reference/postgresql-17 Last reviewed: 2026-08-06 PostgreSQL 17 reached GA on **2024-09-26** and is supported until **2029-11-08**. As of this review date the current minor is 17.10; see the [version and support policy](/en/docs/reference/version-policy) for the live support snapshot. Every feature claim below is checked against the [PostgreSQL 17 release notes](https://www.postgresql.org/docs/17/release-17.html). ## Streaming I/O framework [#streaming-io-framework] PostgreSQL 17 introduced a streaming I/O interface for sequential reads. Instead of issuing one block request at a time, the executor can keep a stream of read requests in flight, which improves the performance of sequential scans and related bulk reads. PostgreSQL 18 builds directly on this framework for its asynchronous I/O subsystem, so the work done in 17 is the foundation for the larger I/O gains in [PostgreSQL 18](/en/docs/reference/postgresql-18). ## Logical replication: failover slots and pg\_createsubscriber [#logical-replication-failover-slots-and-pg_createsubscriber] * **Replication slots can survive failover.** The replication protocol gained a `failover` property, and `sync_replication_slots` lets a standby synchronize failover-enabled logical slots, so logical subscribers can keep streaming after the publisher fails over to a standby. * **`pg_createsubscriber`** converts a physical standby into a logical replica. This is the standard tool for turning a streaming-replication copy into a logical subscriber, and it is a practical path for low-downtime major-version migration. ## MERGE: WHEN NOT MATCHED BY SOURCE and RETURNING [#merge-when-not-matched-by-source-and-returning] PostgreSQL 17 extended `MERGE` with `WHEN NOT MATCHED BY SOURCE` actions and a `RETURNING` clause (including the `merge_action()` function, which reports which DML action each row took): ```sql MERGE INTO target USING source ON source.id = target.id WHEN MATCHED AND target.deleted = false THEN UPDATE SET ... WHEN NOT MATCHED THEN INSERT ... WHEN NOT MATCHED BY SOURCE THEN DELETE RETURNING merge_action(), *; ``` Together these make `MERGE` usable for full synchronization workloads, not only upserts. ## JSON\_TABLE [#json_table] PostgreSQL 17 added the SQL-standard `JSON_TABLE()` function, which projects JSON data into a relational rowset: ```sql SELECT * FROM JSON_TABLE(jsonb_col, '$.items[*]' COLUMNS ( id bigint PATH '$.id', name text PATH '$.name' )) AS t; ``` This removes a class of hand-written `jsonb_to_recordset` and lateral-extraction queries when JSON documents need to be joined, filtered, or aggregated as rows. ## Incremental backup with pg\_basebackup [#incremental-backup-with-pg_basebackup] PostgreSQL 17 added incremental file-system backup. `pg_basebackup --incremental` produces a backup containing only the blocks changed relative to a previous backup's manifest, and `pg_combinebackup` reconstructs a full backup from a full plus incremental chain: ```bash pg_basebackup --incremental=/path/to/backup_manifest -D ./backup_inc ``` For large databases this shortens backup windows and reduces backup storage, at the cost of a combine step during restore. ## Configurable SLRU buffer pools [#configurable-slru-buffer-pools] The SLRU caches behind subtransactions, commit timestamps, and related subsystems are now individually sizable through settings such as `subtransaction_buffers` and `commit_timestamp_buffers`, and by default they scale with `shared_buffers`. Previously these caches were fixed-size and could become a contention point under heavy concurrency. ## Also in PostgreSQL 17 [#also-in-postgresql-17] * **`COPY ... ON_ERROR ignore`** skips malformed input rows instead of aborting the whole copy. (PostgreSQL 18 later added `REJECT_LIMIT` to bound how many rows may be discarded.) * PostgreSQL 17 removed the `old_snapshot_threshold` setting, the `adminpack` contrib extension, and several other long-deprecated pieces — check the release notes if you migrate from a much older major. ## Upgrade notes for PostgreSQL 17 [#upgrade-notes-for-postgresql-17] * A major upgrade requires `pg_upgrade`, dump/restore, or logical replication; the data directory is not compatible across majors. `pg_createsubscriber` (above) is the logical-replication route introduced in this version. * Read the release notes for every major you cross, not only the target version. * Monitoring queries may need updates: PostgreSQL 17 renamed several `pg_stat_statements` timing columns (for example `blk_read_time` to `shared_blk_read_time`), and `pg_stat_bgwriter` lost its `buffers_backend`/`buffers_backend_fsync` columns (redundant with `pg_stat_io`). * Verify each extension's PostgreSQL 17 support separately; extensions carry their own versions and upgrade scripts. Run the current 17.x minor in production and rehearse any major upgrade on a restored copy of real data before scheduling a cutover. The general checklist lives in the [version and support policy](/en/docs/reference/version-policy). ## Version path [#version-path] * Next major: [PostgreSQL 18 features and upgrade notes](/en/docs/reference/postgresql-18) * In beta as of 2026-08: [PostgreSQL 19 release and 18-to-19 upgrade guide](/en/docs/postgresql-19) * Support lifecycle: [Version and support policy](/en/docs/reference/version-policy) ## Sources [#sources] * [PostgreSQL 17 release notes](https://www.postgresql.org/docs/17/release-17.html) * [PostgreSQL versioning policy](https://www.postgresql.org/support/versioning/) --- # PostgreSQL 18 features and upgrade notes Canonical URL: https://pg.edu.rich/en/docs/reference/postgresql-18 Last reviewed: 2026-08-06 PostgreSQL 18 reached GA on **2025-09-25** and is supported until **2030-11-14**. As of this review date the current minor is 18.4; see the [version and support policy](/en/docs/reference/version-policy) for the live support snapshot. Every feature claim below is checked against the [PostgreSQL 18 release notes](https://www.postgresql.org/docs/18/release-18.html). ## Asynchronous I/O (AIO) [#asynchronous-io-aio] PostgreSQL 18 is the first release with a real asynchronous I/O subsystem, building on the streaming I/O framework introduced in [PostgreSQL 17](/en/docs/reference/postgresql-17). I/O can be issued without blocking the backend on each request, which benefits sequential scans, vacuum, and bitmap heap scans. The behavior is controlled by `io_method`: * `sync` — traditional synchronous I/O; * `worker` — background I/O worker processes; this is the **default**, chosen for broad compatibility; * `io_uring` — Linux io\_uring, available on supported kernels. The [PostgreSQL 18 announcement](https://www.postgresql.org/about/news/postgresql-18-released-3142/) reports benchmark gains of up to 3× in certain scenarios. Treat that as an upper bound, not an expectation: measure your own workload before and after switching `io_method`, and move from `worker` to `io_uring` only after testing. ## Native uuidv7() [#native-uuidv7] ```sql SELECT uuidv7(); ``` UUID v7 embeds a 48-bit millisecond timestamp, so generated values are roughly time-ordered. Compared with random v4 UUIDs, time-ordered keys insert into the right edge of a B-tree, which keeps primary-key indexes compact and cache-friendly under high write rates. New tables can adopt it directly: ```sql CREATE TABLE events ( id uuid PRIMARY KEY DEFAULT uuidv7() ); ``` ## Virtual generated columns [#virtual-generated-columns] ```sql ALTER TABLE orders ADD COLUMN total numeric GENERATED ALWAYS AS (qty * price) VIRTUAL; ``` Before PostgreSQL 18, generated columns were always `STORED` — computed at write time and occupying disk. PostgreSQL 18 adds `VIRTUAL` generated columns, computed when the row is read and occupying no storage, and makes `VIRTUAL` the default; write-time behavior remains available through the explicit `STORED` keyword. ## B-tree skip scan [#b-tree-skip-scan] A multicolumn B-tree index on `(a, b)` used to require a predicate on the leading column `a` to be useful. PostgreSQL 18 can perform skip scans: the executor iterates over the distinct values of `a` and descends into each `b` range, so queries that filter only on a non-leading column can still use the index. Some indexes that existed only to cover the second column may now be redundant — verify with `EXPLAIN (ANALYZE, BUFFERS)` on real data before dropping anything. ## RETURNING OLD / NEW [#returning-old--new] `INSERT`, `UPDATE`, `DELETE`, and `MERGE` can now return both the old and new row versions explicitly: ```sql UPDATE products SET price = price * 1.1 RETURNING OLD.price AS prev_price, NEW.price AS new_price; ``` This replaces the common CTE or self-join workarounds for capturing before/after values in one statement. ## Also in PostgreSQL 18 [#also-in-postgresql-18] * **`EXPLAIN ANALYZE` includes `BUFFERS` output automatically**, so buffer statistics no longer need an explicit option in the common case. * **`COPY FROM` gained `REJECT_LIMIT`** to bound how many invalid rows an `ON_ERROR ignore` copy may discard before failing. (`ON_ERROR ignore` itself was added in [PostgreSQL 17](/en/docs/reference/postgresql-17).) * **`pg_stat_io` reports I/O in bytes and adds WAL I/O rows**; correspondingly, the read/sync columns were removed from `pg_stat_wal` — update dashboards that referenced them. * **`pg_upgrade` now retains optimizer statistics**, shortening the window of degraded plans right after a major upgrade. Transparent Data Encryption was proposed during the PostgreSQL 18 cycle but was not merged, and PostgreSQL 18 ships without built-in page-level encryption. Encryption at rest still requires filesystem or disk-level encryption (for example LUKS), storage-level features, or a vendor distribution that provides TDE. Do not plan around a core TDE feature arriving in this version. ## Upgrade notes for PostgreSQL 18 [#upgrade-notes-for-postgresql-18] * A major upgrade requires `pg_upgrade`, dump/restore, or logical replication; the data directory is not compatible across majors. * **`initdb` now enables data checksums by default.** Because `pg_upgrade` requires matching checksum settings, the new `--no-data-checksums` initdb option exists for upgrading old clusters that were initialized without checksums. * **MD5 password authentication is deprecated** and `ALTER ROLE` emits warnings when setting MD5 passwords; plan the move to SCRAM rather than suppressing the warnings. * As noted above, `pg_stat_wal` lost its read/sync columns to `pg_stat_io`; adjust monitoring collectors. * Verify extension, driver, pool, and backup-tool support for 18 individually; `pg_upgrade` cannot prove third-party modules are compatible. ## Version path [#version-path] * Previous major: [PostgreSQL 17 features and upgrade notes](/en/docs/reference/postgresql-17) * In beta as of 2026-08: [PostgreSQL 19 release and 18-to-19 upgrade guide](/en/docs/postgresql-19) * Support lifecycle: [Version and support policy](/en/docs/reference/version-policy) ## Sources [#sources] * [PostgreSQL 18 release notes](https://www.postgresql.org/docs/18/release-18.html) * [PostgreSQL 18 announcement](https://www.postgresql.org/about/news/postgresql-18-released-3142/) * [PostgreSQL versioning policy](https://www.postgresql.org/support/versioning/) --- # PostgreSQL compatible database guide Canonical URL: https://pg.edu.rich/en/docs/reference/postgresql-compatible-databases Last reviewed: 2026-08-02 “Built on PostgreSQL,” “uses the PostgreSQL protocol,” and “can replace PostgreSQL” are three different claims. Compatibility has at least five layers: ```text A driver connects → pgwire messages work → SQL, types, and functions match → catalogs, extensions, and transactions match → backup, replication, upgrade, and failure semantics match ``` Each deeper layer needs migration and failure testing. A PostgreSQL-compatible label usually describes only part of this stack. ## Developer platforms centered on PostgreSQL [#developer-platforms-centered-on-postgresql] | Platform | Where PostgreSQL sits | Platform additions | Do not assume | | ------------------------------------------------------ | -------------------------------------------- | --------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------- | | [Supabase](https://github.com/supabase/supabase) | Each project runs PostgreSQL | PostgREST, Auth, Realtime, Storage, Functions, Dashboard, and pooling | Cloud and self-hosted operations are identical, or browser APIs are secure automatically | | [Neon](https://github.com/neondatabase/neon) | Compute nodes run the PostgreSQL query layer | Separated compute/storage, page servers, branches, and scale-to-zero | Data directory, WAL, unlogged tables, and recovery behave like ordinary PostgreSQL | | [Nhost](https://github.com/nhost/nhost) | PostgreSQL is the database | Hasura GraphQL, Auth, Storage, and Functions | GraphQL permissions are the complete database security boundary | | [Prisma Postgres](https://www.prisma.io/docs/postgres) | Managed PostgreSQL | PgBouncer, HTTP/edge driver, query cache, temporary databases, and Prisma tooling | Operation billing, pooling, and extensions match every self-hosted installation | These platforms preserve ordinary SQL, migration, and `pg_dump` thinking, but their connection proxies, sleep behavior, backups, extension allowlists, API permissions, and billing belong in the architecture. ## Forks, reused query layers, and new storage [#forks-reused-query-layers-and-new-storage] | Project | Implementation path | Primary target | Migration risk center | | --------------------------------------------------------------------------- | ------------------------------------------------------------- | -------------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------- | | [YugabyteDB](https://github.com/yugabyte/yugabyte-db) | YSQL reuses the PostgreSQL query layer over distributed DocDB | Distributed transactions, horizontal scale, and multi-region | Extensions, locks/isolation, catalogs, DDL, and distributed cost model | | [PolarDB for PostgreSQL](https://github.com/polardb/PolarDB-for-PostgreSQL) | Compute/storage-separated PostgreSQL-lineage fork | Shared storage, one writer/many readers, cloud architecture | Open-source version cadence, specialized storage/HA, and differences from the hosted product | | [Apache Cloudberry](https://github.com/apache/cloudberry) | Greenplum/PostgreSQL-lineage MPP database | Warehousing and massively parallel analytics | OLTP transactions, distribution keys, SQL/extensions, and operational tooling | | [IvorySQL](https://github.com/IvorySQL/IvorySQL) | Oracle-compatible fork tracking PostgreSQL | PL/iSQL, Oracle syntax, packages, and migration | Compatibility mode, Oracle semantics, extension packaging, and upstream synchronization | | [openGauss](https://github.com/opengauss-mirror/openGauss-server) | Independent database kernel with PostgreSQL lineage | Enterprise deployment, parallelism, and its own ecosystem | Long independent evolution means current PostgreSQL compatibility is not a default | | [OrioleDB](https://github.com/orioledb/orioledb) | New PostgreSQL storage engine, generally on a supported build | Undo-based MVCC, copy-on-write/checkpointing, and reduced bloat for selected workloads | Binary/build, WAL/backup, extensions, major upgrade, and failure recovery | Lineage is not a drop-in guarantee. Distributed storage in particular changes transaction retries, hot keys, sequences, foreign keys, locks, and consistency/latency tradeoffs. ## PostgreSQL clients work, but the server is not PostgreSQL [#postgresql-clients-work-but-the-server-is-not-postgresql] | Project | Role of pgwire/PostgreSQL | Actual product | | --------------------------------------------------------------------------------- | -------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------- | | [CockroachDB](https://www.cockroachlabs.com/docs/stable/postgresql-compatibility) | Implements pgwire and much PostgreSQL syntax | Independent distributed SQL database with different extensions, catalogs, and transaction behavior | | [Materialize](https://github.com/MaterializeInc/materialize) | PostgreSQL-compatible drivers query views | Streaming incremental-compute and live-data layer, not general-purpose OLTP PostgreSQL | | [Gel](https://github.com/geldata/gel) | Uses PostgreSQL technology underneath and offers SQL/ecosystem integration | Graph-relational database whose primary model and language are Gel/EdgeQL | Compatible databases may report a PostgreSQL-like `server_version` or expose similar catalogs. A version string helps a driver choose a protocol path; it does not prove the same PostgreSQL kernel is running. ## FerretDB is compatibility in the opposite direction [#ferretdb-is-compatibility-in-the-opposite-direction] [FerretDB 2.x](https://github.com/FerretDB/FerretDB) accepts MongoDB 5.0+ wire-protocol requests, translates them to SQL, and uses PostgreSQL with the DocumentDB extension as its database engine: ```text MongoDB driver → FerretDB proxy → PostgreSQL + DocumentDB extension ``` It is not a PostgreSQL client connecting to a MongoDB-compatible server. It is a MongoDB client using a PostgreSQL-backed document database. Test the MongoDB command/BSON matrix, DocumentDB extension, indexes, transactions, backup, and exact version combination. ## Migration compatibility matrix [#migration-compatibility-matrix] | Layer | Required tests | Insufficient evidence | | ------------ | --------------------------------------------------------------------- | ------------------------------------------------ | | Connection | TLS, SCRAM, startup parameters, prepared statements, pooling | `psql` runs `SELECT 1` | | Schema | Types, identity/sequences, generated columns, constraints, partitions | ORM migration succeeds only on an empty database | | SQL | Functions/operators, JSONB, CTE/windows, collations, full text | Basic CRUD works | | Transactions | Isolation, retries, row locks, deadlocks, advisory locks | A product page says “ACID” | | Extensions | Exact version, operators/index methods, update scripts | The extension name appears on an allowlist | | Operations | Backup/PITR, CDC, replication, catalogs, monitoring | A backup button exists | | Failure | Node/zone failure, connection convergence, RPO/RTO, rollback | Vendor benchmarks or an SLA number | Run application tests and migrations first, restore a masked production copy next, and rehearse cutover and rollback last. For distributed databases such as CockroachDB and YugabyteDB, deliberately induce transaction conflicts, hot partitions, and node failures. ## Choosing a direction [#choosing-a-direction] * For standard PostgreSQL ecology and minimum migration cost, prefer community PostgreSQL or a managed service that explicitly runs it. * For BaaS choose Supabase; for GraphQL-first choose Nhost; for branches and scale-to-zero evaluate Neon. * For multi-region distributed OLTP, evaluate YugabyteDB and CockroachDB as new databases, not as settings. * For an MPP warehouse, evaluate Cloudberry without extrapolating from OLTP benchmarks. * For Oracle migration, evaluate IvorySQL and preserve separate PostgreSQL-mode and Oracle-mode tests. * For continuously updated views, Materialize is a data-layer candidate, not a transparent primary-OLTP replacement. See [free PostgreSQL cloud databases](/en/docs/cloud/free-postgresql) for free tiers and [PostgreSQL extension selection](/en/docs/reference/extensions-ecosystem) for actual extensions. --- # PostgreSQL vs DuckDB Canonical URL: https://pg.edu.rich/en/docs/reference/postgresql-vs-duckdb Last reviewed: 2026-08-06 PostgreSQL is a row-store client/server database built for many concurrent transactions; DuckDB is an in-process columnar analytical engine — a library linked into your application, not a service you connect to. They are compared often because both speak SQL, but they answer different questions: "can a thousand users safely transact at once" versus "how fast can one process aggregate a billion rows". Behavior statements on this page follow the official documentation of both projects as of August 2026: [PostgreSQL](https://www.postgresql.org/docs/current/) and [DuckDB](https://duckdb.org/docs/stable/). Verify version-specific details before relying on them. ## Architecture comparison [#architecture-comparison] | | PostgreSQL | DuckDB | | ------------------- | --------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | Storage layout | Row-oriented (heap), tuned for point lookups and small writes | Columnar, vectorized execution, tuned for scans and aggregation | | Process model | Independent server; clients connect over the network (pgwire) | In-process library or CLI; the database is a file | | Concurrent writes | Many connections under MVCC | Either one process opens the database read-write, or multiple processes open it read-only (`access_mode = 'READ_ONLY'`); writing from multiple processes is not supported ([concurrency docs](https://duckdb.org/docs/stable/connect/concurrency.html)) | | Transactions | Full ACID, configurable isolation levels | ACID transactions within the single attached process | | Deployment | Server you operate, or a managed service | Embedded in Python/R/Java/Wasm/CLI; nothing to run | | Extension ecosystem | Extensions loaded into the server (pgvector, PostGIS, TimescaleDB, …) | Loadable extensions (`postgres`, `parquet`, `iceberg`, …) | | Serving model | Online services, APIs, multi-tenant applications | Local analysis, ETL, data preparation, edge/embedded analytics | The concurrency row is the practical divider: if the workload is "many writers connected from many machines", DuckDB is architecturally out of scope; if it is "one job reading columnar data as fast as the disk allows", a client/server round trip per query is overhead DuckDB simply does not have. ## When to choose which [#when-to-choose-which] ### Choose DuckDB [#choose-duckdb] * Interactive analysis over local or object-storage files (Parquet, CSV, JSON) without loading them anywhere. * Data preparation and transformation steps in a pipeline or notebook. * Edge and embedded analytics where shipping a server is not an option. ### Choose PostgreSQL [#choose-postgresql] * Multi-user transactional workloads with concurrent writes, constraints, and foreign keys. * Anything that serves an API or an application around the clock. * Workloads that need row-level security, logical replication, point-in-time recovery, or the extension ecosystem — see [the extension ecosystem](/en/docs/reference/extensions-ecosystem) and the [cloud service map](/en/docs/cloud/service-map) for what that buys in practice. ### The analytics boundary on the PostgreSQL side [#the-analytics-boundary-on-the-postgresql-side] PostgreSQL executes analytical queries correctly but row-at-a-time relative to a columnar engine; on large scans the gap is structural, not a tuning miss. Three ways to close it without leaving PostgreSQL data behind: * **Columnar extensions**: [pg\_mooncake](https://github.com/Mooncake-Labs/pg_mooncake) maintains a columnstore mirror of PostgreSQL tables in Iceberg and accelerates analytics with DuckDB execution inside PostgreSQL; [Citus](https://github.com/citusdata/citus) offers a columnar storage option alongside distribution. * **Export to DuckDB**: keep PostgreSQL as the system of record and export snapshots to Parquet for analysis — DuckDB reads Parquet natively. * **Push analysis to the warehouse** when the workload outgrows one node entirely. ## Using both together [#using-both-together] ### DuckDB reads PostgreSQL: postgres\_scanner [#duckdb-reads-postgresql-postgres_scanner] DuckDB's official [`postgres` extension](https://duckdb.org/docs/stable/core_extensions/postgres.html) attaches a live PostgreSQL database and runs queries against it, including pushing filters down: ```sql INSTALL postgres; ATTACH 'dbname=app user=analyst host=127.0.0.1' AS pg (TYPE postgres, READ_ONLY); SELECT status, count(*), avg(total_cents) FROM pg.orders WHERE placed_at >= now() - interval '30 days' GROUP BY status; ``` This is the standard pattern for "analyze production data without exporting it": DuckDB pulls the rows it needs, and the heavy aggregation happens in DuckDB's vectorized engine. Attach read-only and use a low-privilege PostgreSQL role so the analysis path cannot write back. ### PostgreSQL reads files: FDW [#postgresql-reads-files-fdw] In the other direction, PostgreSQL's [foreign data wrapper](https://www.postgresql.org/docs/current/ddl-foreign-data.html) mechanism can expose Parquet files as foreign tables (for example with [`parquet_s3_fdw`](https://github.com/pgspider/parquet_s3_fdw)). This suits cases where the files are inputs to relational processing and should be joinable with live tables under PostgreSQL permissions — not a replacement for DuckDB's scan speed. ```sql CREATE EXTENSION parquet_s3_fdw; CREATE SERVER parquet_files FOREIGN DATA WRAPPER parquet_s3_fdw; CREATE FOREIGN TABLE lake_events (...) SERVER parquet_files OPTIONS (dirname 's3://analytics/events/', sorted 'event_time'); ``` Server and table options depend on the FDW version; check its README for the exact syntax. ## In AI data stacks [#in-ai-data-stacks] The two engines typically appear at different stages of the same pipeline: * **DuckDB for preparation**: cleaning, joining, and aggregating raw exports and data-lake files into the documents and tables an AI application will actually serve. No server to run, and Parquet output drops straight into the next stage. * **PostgreSQL for serving**: the online path — multi-tenant transactional state, hybrid retrieval with pgvector and full-text search for [RAG](/en/docs/ai/rag-pipeline), and agent state such as [long-term memory](/en/docs/ai/agent-memory), where concurrency, permissions, and auditability are the point. A rule of thumb: data at rest being *prepared* belongs in DuckDB's reach; data being *served* to users and agents belongs in PostgreSQL. ## AI prompt: pick the engine for a workload [#ai-prompt-pick-the-engine-for-a-workload] ## Related [#related] * [PostgreSQL lineage and compatible databases](/en/docs/reference/postgresql-compatible-databases) — a different comparison: systems that reuse PostgreSQL itself * [Cloud service map](/en/docs/cloud/service-map) — managed PostgreSQL options when the choice lands on a server * [Extension ecosystem](/en/docs/reference/extensions-ecosystem) — pgvector, PostGIS, and the rest of what "choose PostgreSQL" includes * [RAG pipeline](/en/docs/ai/rag-pipeline) — the online retrieval path PostgreSQL serves --- # Version and support policy Canonical URL: https://pg.edu.rich/en/docs/reference/version-policy Last reviewed: 2026-08-02 ## Version meaning [#version-meaning] Since PostgreSQL 10, the first number is the major, such as 18; the dotted number is the minor, such as 18.4. A major arrives roughly yearly with features. A minor contains bug, security, and low-risk fixes. A minor update does not require dump/restore; it usually replaces binaries and restarts, though its release notes still apply. Data directories are incompatible across majors, requiring `pg_upgrade`, logical dump/restore, or logical replication migration. ## Current support snapshot [#current-support-snapshot] As of **2026-08-02**: | Major | Current minor | Status | Final support date | | ----- | ------------: | ------------------- | ------------------ | | 18 | 18.4 | Supported | 2030-11-14 | | 17 | 17.10 | Supported | 2029-11-08 | | 16 | 16.14 | Supported | 2028-11-09 | | 15 | 15.18 | Supported | 2027-11-11 | | 14 | 14.23 | Supported, near EOL | 2026-11-12 | Source: [official PostgreSQL versioning policy](https://www.postgresql.org/support/versioning/). Treat that live page as authoritative. ## Upstream versions are not distribution package versions [#upstream-versions-are-not-distribution-package-versions] ### Sample default postgresql packages across distributions PkgSeek package snapshot; checked: 2026-08-02 15:47:58 UTC. | Distribution | Release | Full version | Repository | Linked advisories | |---|---|---|---|---:| | Alibaba Cloud Linux | 3 | 13.23-3.0.1.al8 | official / updates | 0 | | Alibaba Cloud Linux | 4 | 15.18-1.alnx4 | official / updates | 0 | | AlmaLinux | 10 | 16.14-1.el10_2 | official / AppStream | 0 | | AlmaLinux | 9 | 18.4-2.module_el9.8.0+280+5ad12178 | official / AppStream | 0 | | Arch | rolling | 18.4-3 | official / extra | 0 | | CentOS Stream | 10 | 16.14-1.el10 | official / AppStream | 0 | | CentOS Stream | 9 | 13.23-3.el9 | official / AppStream | 0 | | Debian | trixie | 17+278 | official / main | 3 | | deepin | 25.2 | 16+255 | official / main | 0 | | Fedora | 42 | 16.13-1.fc42 | official / updates | 0 | | Fedora | 43 | 18.3-2.fc43 | official / updates | 0 | | Fedora | 44 | 18.3-2.fc44 | official / updates | 0 | Source: [PkgSeek package lookup](https://pkgseek.com/packages/postgresql). Distribution revisions and backported fixes are part of the complete version identity. Distributions may freeze a major and record packaging revisions or backported security fixes in complete versions such as `16.14-1.el10_2` or `18+290ubuntu1`. Do not decide vulnerability status from the leading `16` or `18` alone. Identify the distribution, release, repository, architecture, and complete package version, then verify the vendor advisory. ## Choosing for a new system [#choosing-for-a-new-system] Default to the latest minor of the latest stable major unless a driver, extension, managed platform, or organizational certification is not ready. A conservative choice is a major with ample support runway that your workload has tested—not a near-EOL version chosen merely because it feels old. PostgreSQL 19 is still in its beta cycle in August 2026 and is not a production default. Test a new major against extensions, collations, backup tools, pools, ORM, plans, and monitoring collectors. See the [PostgreSQL 19 release and 18-to-19 upgrade guide](/en/docs/postgresql-19) for the current Beta status, feature changes, and migration risks. ## Keep the matrix in the repository [#keep-the-matrix-in-the-repository] ```yaml postgresql: supported_majors: [17, 18] tested_minor_floor: 17: 17.10 18: 18.4 extensions: vector: "tested in CI" upgrade_owner: platform-database next_review: 2026-11-01 ``` `latest` is not a deployment strategy. Pin auditable images/packages and let a dependency process advance minors. ## Upgrade principles [#upgrade-principles] * Run the current minor of the chosen major; remaining on an old minor is often riskier than updating. * Read all release notes across the version span. * Test application behavior and plans on a restored real-data copy. * Extensions have separate versions and upgrade scripts; inspect each one. * Refresh statistics and validate backups after a major cutover. --- # Install and connect to PostgreSQL Canonical URL: https://pg.edu.rich/en/docs/setup Last reviewed: 2026-08-02 ## Choose the outcome first [#choose-the-outcome-first] | Outcome | Recommended starting point | | ---------------------------- | --------------------------------------------------------------------------------- | | Learning, tests, CI | Docker: pin the version and remove the whole environment cleanly | | Persistent local development | OS package manager or a trusted installer | | Production | A managed cloud service or team-managed repositories and configuration automation | | Client only | Install `psql`/libpq client packages without running a local server | Do not expose port 5432 publicly or turn `pg_hba.conf` into a global trust rule merely to connect. Verify locally first, then design networking, TLS, authentication, and least privilege. ## Run the same verification everywhere [#run-the-same-verification-everywhere] ```bash psql --version psql -X "postgresql://postgres@localhost:5432/postgres" \ -c "select current_setting('server_version'), current_database(), current_user;" ``` Client and server versions are separate. `psql --version` reports only the client; the SQL query reports the server you actually reached. When versions coexist, record executable path, port, and data directory. Next, read [`psql` connections and SSL](/en/docs/setup/psql-connection), then create an application role that is not a superuser. --- # PostgreSQL Linux packages, versions, and PGDG Canonical URL: https://pg.edu.rich/en/docs/setup/linux-packages Last reviewed: 2026-08-02 “Install PostgreSQL on Linux” is not one stable command. The distribution, release, repository source, CPU architecture, and package name jointly determine the installed major, packaging revision, and security fixes. This page combines changing package evidence from [PkgSeek](https://pkgseek.com/packages/postgresql) with this guide's installation and upgrade rules. ### Default postgresql packages across Linux distributions PkgSeek package snapshot; checked: 2026-08-02 15:47:58 UTC. | Distribution | Release | Full version | Repository | Linked advisories | |---|---|---|---|---:| | Alibaba Cloud Linux | 3 | 13.23-3.0.1.al8 | official / updates | 0 | | Alibaba Cloud Linux | 4 | 15.18-1.alnx4 | official / updates | 0 | | AlmaLinux | 10 | 16.14-1.el10_2 | official / AppStream | 0 | | AlmaLinux | 9 | 18.4-2.module_el9.8.0+280+5ad12178 | official / AppStream | 0 | | Arch | rolling | 18.4-3 | official / extra | 0 | | CentOS Stream | 10 | 16.14-1.el10 | official / AppStream | 0 | | CentOS Stream | 9 | 13.23-3.el9 | official / AppStream | 0 | | Debian | trixie | 17+278 | official / main | 3 | | deepin | 25.2 | 16+255 | official / main | 0 | | Fedora | 42 | 16.13-1.fc42 | official / updates | 0 | | Fedora | 43 | 18.3-2.fc43 | official / updates | 0 | | Fedora | 44 | 18.3-2.fc44 | official / updates | 0 | | Kali Linux | kali-rolling | 18+290 | official / main | 0 | | Kylin OS | V10-SP1 | 12+214kylin0.1 | official / 10.1-main | 0 | | Kylin OS Server | V10-SP3-2403 | 10.5-23.p09.ky10 | official / updates | 0 | | OpenAnolis | 23.4 | 15.18-1.an23 | official / updates | 0 | | OpenAnolis | 8.10 | 12.22-7.0.1.module+an8.10.0+11420+6683745d | official / AppStream | 0 | | openEuler | 24.03-LTS-SP4 | 15.18-1.oe2403sp4 | official / everything | 0 | | openSUSE | 15.6 | 18-150600.17.9.1 | official / update-sle | 0 | | openSUSE | tumbleweed | 18-3.4 | official / oss | 0 | | Oracle Linux | 10 | 16.14-1.0.1.el10_2 | official / appstream | 0 | | Oracle Linux | 8 | 12.22-6.0.1.module+el8.10.0+90932+f6d78e3c | official / appstream | 34 | | Oracle Linux | 9 | 13.23-3.el9_8 | official / appstream | 26 | | Raspberry Pi OS | bookworm | 15+248+deb12u1 | official / bookworm-main | 0 | | Raspberry Pi OS | trixie | 17+278 | official / trixie-main | 0 | | Red Hat Enterprise Linux | 10.2 | 16.14-1.el10_2 | official / AppStream | 0 | | Red Hat Enterprise Linux | 9.8 | 18.4-2.module+el9.8.0+24359+da7fad50 | official / AppStream | 0 | | Rocky Linux | 10 | 16.14-1.el10_2 | official / AppStream | 0 | | Rocky Linux | 9 | 13.23-3.el9_8 | official / AppStream | 13 | | Ubuntu | focal | 12+214 | official / main | 0 | | Ubuntu | jammy | 14+238 | official / main | 0 | | Ubuntu | noble | 16+257build1 | official / main | 0 | | Ubuntu | resolute | 18+290ubuntu1 | official / main | 0 | | Void Linux | rolling | 18_1 | official / current | 0 | Source: [PkgSeek package lookup](https://pkgseek.com/packages/postgresql). Distribution revisions and backported fixes are part of the complete version identity. ## Distinguish four package families [#distinguish-four-package-families] | Family | Common examples | Meaning | | ------------------------ | ------------------------------------------------------- | -------------------------------------------------------------------- | | Distribution metapackage | `postgresql` | Follows the major selected by that distribution release | | Versioned server | `postgresql-18`, `postgresql18-server` | Pins a major; naming differs between DEB and RPM ecosystems | | Client and development | `postgresql-client-18`, `libpq-dev`, `postgresql-devel` | Provides `psql`, libpq headers, or build files; may not run a server | | Extension | `postgresql-18-pgvector`, `pgvector` | Must also match server major, architecture, and extension version | Similar names do not prove identical roles. Verify the client and server separately after installation: ```bash psql --version sudo -u postgres psql -X -d postgres \ -c "select version(), current_setting('server_version_num');" ``` ## Distribution repositories and PGDG [#distribution-repositories-and-pgdg] Distribution repositories generally maintain a selected major. The [PostgreSQL Global Development Group repositories](https://www.postgresql.org/download/linux/) are commonly used for other supported majors. Adding PGDG creates an external software-source dependency, so record repository signing, release support, upgrade policy, and an exit path. ### postgresql-18 in Ubuntu repositories PkgSeek package snapshot; checked: 2026-08-02 15:47:58 UTC. | Distribution | Release | Full version | Repository | Linked advisories | |---|---|---|---|---:| | Ubuntu | resolute | 18.3-1 | official / main | — | Source: [PkgSeek package lookup](https://pkgseek.com/packages/postgresql-18). Distribution revisions and backported fixes are part of the complete version identity. ### postgresql-18 in Ubuntu PGDG PkgSeek package snapshot; checked: 2026-08-02 15:47:58 UTC. > No exact indexed coordinate. This means the current snapshot has no match, not that the package does not exist. Source: [PkgSeek package lookup](https://pkgseek.com/search?q=postgresql-18). Distribution revisions and backported fixes are part of the complete version identity. The tables query exact coordinates in the Ubuntu and PGDG repositories separately. If a card reports “no exact indexed coordinate”, that describes index coverage, not the target repository's contents. Use the target system's `apt-cache policy` and official PostgreSQL repository instructions as the final installation evidence. ```bash apt-cache policy postgresql postgresql-18 postgresql-client-18 apt-cache madison postgresql-18 ``` ## Track PostgreSQL 19 packages [#track-postgresql-19-packages] ### postgresql-19 in distribution repositories PkgSeek package snapshot; checked: 2026-08-02 15:47:58 UTC. > No exact indexed coordinate. This means the current snapshot has no match, not that the package does not exist. Source: [PkgSeek package lookup](https://pkgseek.com/search?q=postgresql-19). Distribution revisions and backported fixes are part of the complete version identity. A missing stable distribution package is expected while PostgreSQL 19 remains in beta. Even after a test build appears, keep the beta compatibility lane separate from production PostgreSQL 18. Use the [PostgreSQL 18-to-19 upgrade guide](/en/docs/postgresql-19) for adoption criteria. ## Use package evidence safely [#use-package-evidence-safely] 1. Record `distro + release + source + repository + architecture + full version`. 2. Determine whether a package contains the server, client, development files, or an extension. 3. For CVEs, prioritize distribution advisories and backport status over an upstream version prefix. 4. After installation, verify the connected server; `psql --version` only identifies the client. 5. A major change requires `pg_upgrade`, dump/restore, or logical replication. Replacing a package does not upgrade a data directory. Package indexes have coverage and refresh boundaries. When this page reports “no exact indexed coordinate”, continue with the upstream repository; do not let a person or model turn an empty array into a definitive answer. ## Minimum contract for an AI agent [#minimum-contract-for-an-ai-agent] ```yaml task: resolve_postgresql_package target: distro: ubuntu release: noble architecture: amd64 source: pgdg requirements: - return exact package coordinates and observed_at - separate indexed fact, inference, and unknown - never interpret missing index data as package absence - ask before repository, package, service, or data changes - verify client and server versions after execution ``` PkgSeek publishes a [read-only MCP tool catalogue](https://pkgseek.com/mcp/tools) and an [OpenAPI 3.1 specification](https://pkgseek.com/openapi.json). See the [AI context contract](/en/docs/ai/context-contract) for the execution boundary. --- # Install PostgreSQL 18 on macOS Canonical URL: https://pg.edu.rich/en/docs/setup/macos Last reviewed: 2026-08-02 The [PostgreSQL macOS download page](https://www.postgresql.org/download/macosx/) lists three common paths: the EDB graphical installer, Postgres.app, and Homebrew. Do not run several installations on port 5432 accidentally; choose one and record its data directory. ## Homebrew [#homebrew] ```bash brew update brew install postgresql@18 brew services start postgresql@18 "$(brew --prefix postgresql@18)/bin/psql" --version ``` Homebrew may not place a versioned client in the default `PATH`. For a persistent shell setup, follow the path printed by `brew info postgresql@18` instead of copying a Cellar path that can change on upgrade. Verify the connection: ```bash "$(brew --prefix postgresql@18)/bin/psql" -X -d postgres \ -c "select version(), current_setting('data_directory');" ``` Use `-U` and `-d` when your local role or database differs. ## Postgres.app or graphical installer [#postgresapp-or-graphical-installer] * Postgres.app fits local developers who want menu-bar start/stop with little system configuration. * The EDB installer includes the server, pgAdmin, and StackBuilder for a guided graphical setup. After installation, do not rely on an icon alone. Run this through the bundled `psql`: ```sql SELECT version(), current_database(), current_user; ``` Confirm that the package or Homebrew prefix matches arm64/amd64. Do not copy a data directory across architectures; prefer logical backup or a tested upgrade path. ## Diagnose multiple versions [#diagnose-multiple-versions] ```bash which -a psql psql --version lsof -nP -iTCP:5432 -sTCP:LISTEN ``` A client/server version difference does not automatically require reinstalling. Confirm the actual target first; a client no older than the server is generally the safer choice. --- # Connect to PostgreSQL with psql and SSL Canonical URL: https://pg.edu.rich/en/docs/setup/psql-connection Last reviewed: 2026-08-06 ## State all five connection dimensions [#state-all-five-connection-dimensions] ```bash psql -X \ --host=db.example.com \ --port=5432 \ --username=app_reader \ --dbname=commerce ``` Host, port, database, user, and TLS parameters identify the target together. A database name alone is insufficient; many instances can contain the same name. Equivalent connection URI: ```bash psql -X "postgresql://app_reader@db.example.com:5432/commerce?sslmode=verify-full" ``` Do not place passwords in command-line URIs, source code, or logs. Use an interactive prompt, secret manager, short-lived identity, `.pgpass`, or a libpq service file. ## pgpass [#pgpass] The Unix default is `~/.pgpass`, restricted to mode `0600`: ```text hostname:5432:database:username:password ``` ```bash chmod 600 ~/.pgpass ``` The Windows default is `%APPDATA%\postgresql\pgpass.conf`. Wildcards broaden where a credential applies; prefer a specific host, database, and user. ## Connection service file [#connection-service-file] A libpq service file keeps host, user, and TLS parameters out of shell history and command lines. The default path is `~/.pg_service.conf`, overridable with `PGSERVICEFILE`: ```text [prod] host=db.example.com port=5432 user=app_reader dbname=commerce sslmode=verify-full ``` ```bash psql -X service=prod ``` Applications can reference the same entry through `PGSERVICE=prod`. Keep the password in `.pgpass` or a secret manager, not in the service file. ## SSL modes [#ssl-modes] | `sslmode` | Behavior | Guidance | | ------------- | ----------------------------------------------- | ------------------------------------------------------------ | | `disable` | No TLS | Controlled local or isolated testing only | | `require` | Encrypts without complete identity verification | Better than plaintext, insufficient against a wrong endpoint | | `verify-ca` | Verifies the certificate chain | Does not verify hostname | | `verify-full` | Verifies chain and hostname | Target for remote production connections | `verify-full` requires the URI host to match the certificate identity and a correct root certificate. Cloud providers can have CA rotation workflows; do not pin an expired certificate forever. ## Confirm immediately after connecting [#confirm-immediately-after-connecting] ```sql \conninfo SELECT current_database(), current_user, session_user, inet_server_addr(), inet_server_port(), current_setting('server_version') AS server_version, current_setting('TimeZone') AS timezone; ``` For scripts: ```bash psql -X --set ON_ERROR_STOP=on --file migration.sql "$DATABASE_URL" ``` `-X` prevents a user's `.psqlrc` from changing automation; `ON_ERROR_STOP` exits on SQL errors. See the [PostgreSQL 18 psql documentation](https://www.postgresql.org/docs/18/app-psql.html) for exit-status semantics. ## Machine-readable output formats [#machine-readable-output-formats] Interactive output is an aligned table; pipelines need unaligned, tuples-only results: ```bash psql -X -A -F, -t -c "SELECT id, email FROM users" > users.csv ``` `-A` disables alignment, `-F,` sets the field separator, and `-t` prints tuples only. Inside a session, `\pset format csv`, `\pset null '[NULL]'`, and `\o /tmp/out.txt` (then `\o` again to return to stdout) provide the same control. ## Useful one-liners [#useful-one-liners] Largest tables by total relation size: ```bash psql -X -c " SELECT schemaname||'.'||relname AS table_name, pg_size_pretty(pg_total_relation_size(relid)) AS total_size FROM pg_stat_user_tables ORDER BY pg_total_relation_size(relid) DESC LIMIT 10;" ``` Currently running queries, oldest first: ```bash psql -X -c " SELECT pid, now() - query_start AS duration, state, query FROM pg_stat_activity WHERE state = 'active' AND query NOT ILIKE '%pg_stat_activity%' ORDER BY query_start LIMIT 20;" ``` Terminate a runaway backend after confirming the PID: ```bash psql -X -c "SELECT pg_terminate_backend(12345);" ``` ## An interactive \~/.psqlrc [#an-interactive-psqlrc] `~/.psqlrc` runs at every interactive startup; `-X` skips it, so automation is unaffected. A minimal template: ```text \timing on \pset null '[NULL]' \x auto \pset linestyle unicode \pset border 2 \set conns 'SELECT pid, usename, application_name, state, query FROM pg_stat_activity WHERE state <> ''idle'';' \set locks 'SELECT pid, mode, locktype, relation::regclass, granted FROM pg_locks WHERE NOT granted;' \set HISTFILE ~/.psql_history- :DBNAME \set HISTCONTROL ignoredups ``` `:conns` and `:locks` then expand to the stored queries, and per-database history files keep statements from one database out of another's prompt. Connection parameters bound session establishment; `statement_timeout` bounds SQL execution; `lock_timeout` only bounds lock waiting. The application also needs a request deadline and must cancel or release the database connection after timeout. On failure, retain the full error and SQLSTATE, then use the [error fieldbook](/en/docs/reference/errors). Do not troubleshoot by disabling TLS or broadening privileges. --- # Install PostgreSQL 18 on Ubuntu Canonical URL: https://pg.edu.rich/en/docs/setup/ubuntu Last reviewed: 2026-08-02 Ubuntu includes PostgreSQL, but the major version is fixed by the Ubuntu release snapshot. If that maintained default is acceptable: ```bash sudo apt update sudo apt install postgresql postgresql-client ``` ### Ubuntu default postgresql metapackage PkgSeek package snapshot; checked: 2026-08-02 15:47:58 UTC. | Distribution | Release | Full version | Repository | Linked advisories | |---|---|---|---|---:| | Ubuntu | focal | 12+214 | official / main | — | | Ubuntu | jammy | 14+238 | official / main | — | | Ubuntu | noble | 16+257build1 | official / main | — | | Ubuntu | resolute | 18+290ubuntu1 | official / main | — | Source: [PkgSeek package lookup](https://pkgseek.com/packages/postgresql). Distribution revisions and backported fixes are part of the complete version identity. The snapshot shows that the default major changes between Ubuntu releases. `postgresql` follows the distribution; it does not always install PostgreSQL 18. See [Linux packages and PGDG](/en/docs/setup/linux-packages) for the complete boundary. ## Install PostgreSQL 18 explicitly [#install-postgresql-18-explicitly] For a specific major, use the Apt repository maintained by the PostgreSQL project. Its automated setup path is: ```bash sudo apt install -y postgresql-common ca-certificates sudo /usr/share/postgresql-common/pgdg/apt.postgresql.org.sh sudo apt update sudo apt install postgresql-18 postgresql-client-18 ``` Read the script output and confirm that your OS release is supported. Treat the [PostgreSQL Ubuntu download page](https://www.postgresql.org/download/linux/ubuntu/) as authoritative for current commands and releases. ### PkgSeek: Ubuntu / PGDG / postgresql-18 PkgSeek package snapshot; checked: 2026-08-02 15:47:58 UTC. > No exact indexed coordinate. This means the current snapshot has no match, not that the package does not exist. Source: [PkgSeek package lookup](https://pkgseek.com/search?q=postgresql-18). Distribution revisions and backported fixes are part of the complete version identity. If the card does not return an exact PGDG coordinate, the empty result cannot prove that the package is unavailable. Confirm against PGDG repository metadata and the official PostgreSQL download page. The documentation keeps the gap visible instead of filling it with an assumption. ## Verify service and cluster [#verify-service-and-cluster] ```bash systemctl status postgresql --no-pager pg_lsclusters sudo -u postgres psql -X -d postgres \ -c "select version(), current_setting('data_directory');" ``` `postgresql.service` is the cluster-management entry point; a concrete instance commonly appears as `postgresql@18-main`. `pg_lsclusters` belongs to Debian/Ubuntu `postgresql-common`, not every Linux distribution. Create a local practice role and database: ```bash sudo -u postgres createuser --pwprompt learner sudo -u postgres createdb --owner=learner learner psql -X -h localhost -U learner -d learner -c "select current_user;" ``` Remote access involves `listen_addresses`, `pg_hba.conf`, firewall rules, and TLS together. Save the original configuration, start with a specific network/database/role rule, and test both allowed and denied paths. ## Upgrade boundary [#upgrade-boundary] `apt upgrade` can install minor fixes inside one major. Moving from 17 to 18 is a major upgrade requiring `pg_upgrade`, logical dump/restore, or logical replication; replacing packages does not make an old data directory compatible. --- # Install PostgreSQL 18 on Windows Canonical URL: https://pg.edu.rich/en/docs/setup/windows Last reviewed: 2026-08-02 The PostgreSQL project links the EDB-certified interactive installer from its [Windows download page](https://www.postgresql.org/download/windows/). The bundle commonly includes the PostgreSQL server, pgAdmin, and StackBuilder. ## Record during installation [#record-during-installation] * major version and install directory; * data directory, outside folders controlled by consumer sync tools; * PostgreSQL service account; * listening port, commonly 5432; * the `postgres` administrator password, stored in a password manager; * locale, with collation behavior tested before production migration. Select only components you need. Extra drivers and extensions in StackBuilder are not PostgreSQL core and should follow project requirements. ## Verify with psql [#verify-with-psql] Open the SQL Shell installed with PostgreSQL, or add its `bin` directory to the current terminal path: ```powershell psql.exe --version psql.exe -X -h localhost -p 5432 -U postgres -d postgres ` -c "select version(), current_database(), current_user;" ``` Inspect the service through Windows Services or PowerShell: ```powershell Get-Service *postgres* ``` If the client is missing, locate `psql.exe` in the installation directory; do not download standalone DLLs from an unknown source. ## Common connection failures [#common-connection-failures] | Symptom | Check first | | ------------------------------ | -------------------------------------------------------------------- | | connection refused | Whether the PostgreSQL service is running and the port is correct | | password authentication failed | Username, target instance, password, and matching `pg_hba.conf` rule | | database does not exist | Whether the database passed with `-d` was created | | unexpected server version | Whether several PostgreSQL services are running on Windows | Local development normally needs no inbound firewall opening. Remote access must restrict source addresses and use appropriate TLS validation. Do not replace diagnosis with a passwordless `trust` rule in `pg_hba.conf`. ## Before uninstalling [#before-uninstalling] Removing software and deleting the data directory are separate operations. Export required content with `pg_dump` plus `pg_dumpall --globals-only`, test a restore, and confirm the exact data directory before deletion.