Skip to content

ADR 001 — Postgres state and serial partition processing

Status: storage, source-worker lifecycle, and relay/view transport implemented and integration-tested · updated 3 October 2026

Context

The reference example must make the relationship between authority, derived state, publication intent, input position, and stale ownership easy to inspect. It is not a throughput benchmark.

Decision

Keep processing state and authoritative source progress in Postgres. Process records serially within each source partition. Lock the partition-state row before mutable reads and hold it through the all-or-nothing state/quarantine/outbox/frontier decision. Install ownership epochs using the same row-lock boundary.

Use two worker instances to exercise actual assignment and recovery. Keep one relay service and a separate durable view materializer. The database fence is effective after the replacement epoch is durably installed; it is not a claim of instantaneous Kafka/database coordination.

The initial client configuration is franz-go v1.22.1 with classic groups, the eager range balancer, disabled auto-commit, explicit external-frontier adjustment, and no automatic offset reset. A bounded spike against Redpanda v26.2.3 established the client observations and remaining cases. The source adapter now pins that dependency; the pure core does not import it.

Offset loading must have an independent cancellation/deadline path. In the tested version, adjustment can delay revocation, and returning context.Canceled can terminate group management. Retire an interrupted client explicitly from outside its callbacks. Keep epoch installation under an assignment lifecycle guard; neither a live callback context nor the broker assignment alone grants a durable fence.

Alternatives

Embedded state plus a changelog could reduce database interaction but would add restoration and progress coordination. Per-key concurrency inside a partition would need safe tracking of partially completed source offsets. Neither is needed to explain the initial correction pattern.

Consequences

The implementation is easier to inspect and can keep state, versions, quarantine, and progress in one transaction. A deep correction may hold a partition lock for a long time, delaying unrelated keys on that partition and ownership installation. Observe refold cost and lock/lag behavior instead of hiding that tradeoff.

The memory store does not establish actual SQL correctness. The initial PostgreSQL adapter now has real lock, rollback, lost-commit-reply, and view recovery tests, described in the storage guide. The source worker guide records the implemented eager-group lifecycle tests. The delivery guide records the implemented relay/view transport, dedicated session guards, and their limits: no broker fencing or automatic takeover after an uncertain old relay. The bounded operational verifier now uses a clean stopped-relay handoff; the process-crash showcase now exercises actual worker death and a higher-epoch recovery.

Eager rebalancing gives the first adapter a simple revoke/reassign boundary at the cost of moving unaffected partitions. Cooperative and server-side assignment remain alternatives requiring their own lifecycle tests. The operation outcomes and retry controls preserve committed, aborted, and unknown observations without committing the application to a general-purpose simulation framework.

Review trigger

Revisit when measured optional workloads show a material limit and a more complex design would demonstrate a useful additional tradeoff. This ADR describes the reference example only; it says nothing about a client's architecture.