feat(mongodb): implement the MongoDB document persistence platform
Implements the mongodb-superpowers-package design: Stable Tasks 1-50 and Advanced Tasks 1-15. The design assumes 19 Stable + 12 Advanced Gradle projects under modules/mongodb*. This repository's fail-closed registry declares exactly 19 leaf identities, so those modules become package boundaries inside the registered leaf :adapter:outbound:persistence-mongo, with the design's module dependency table enforced by ten ArchUnit rules. The mapping and every deviation are recorded in docs/mongodb/repository-adaptation.md. Contract highlights, all enforced by tests rather than convention: - Transaction body retry and commit retry are separate loops. A new session per body attempt; commit-only retry on an unknown commit. The body is never replayed after a commit ambiguity, so a failover cannot become a duplicate. - MongoExecutionOutcome keeps both ambiguous outcomes distinct from success and failure, and MongoFailureContext records only the design-permitted fields. - Failure classification reads server error labels before numeric codes. - BSON representations come from a pinned manifest, never a library default, and a golden type-signature gate fails on any drift. - Index and validator changes go through the manifest and the admin plane; metadata ownership gates every drop. - Every Advanced capability refuses construction unless its flag is enabled. Verified against real servers, not only unit tests. Running the lanes for the first time exposed four defects that a green `check` had hidden: - Four release lanes passed while executing zero tests; the gate now counts executed tests per lane and fails on zero. - The "single replica set" fixture was a standalone, because Testcontainers 2.x needs withReplicaSet(); its test only asserted a connection string. - The three-node fixture was three independent clusters, so no election could occur, and awaitNewPrimary() compared against the post-stop primary. - The migration lease checked modifiedCount, so a same-millisecond refresh read as a lost lease. scripts/verify-mongodb-platform.sh now reports: 9 lanes, 0 skipped, 0 failed, every evidence category produced. scripts/verify-mongodb-advanced.sh reports NOT PROMOTABLE: actual-topology evidence (real sharded cluster, real KMS, real target deployment) is unobtainable here, so it is named rather than assumed. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
co-authored by
Claude Opus 5
parent
3b5aee50e3
commit
d57d2f62a0
@@ -0,0 +1,98 @@
|
||||
---
|
||||
title: Runbook — MongoDB primary failover
|
||||
category: mongodb
|
||||
severity: P2
|
||||
owner: oncall
|
||||
last_updated: 2026-08-13
|
||||
status: active
|
||||
---
|
||||
|
||||
# Runbook: MongoDB primary failover
|
||||
|
||||
Design §29, scenarios `PRIMARY_KILL`, `NETWORK_PARTITION`, `SERVER_SELECTION_TIMEOUT`,
|
||||
`WRITE_RESPONSE_LOSS`.
|
||||
|
||||
## Symptoms
|
||||
|
||||
- `MongoServerSelectionException` / `MongoConnectionException` spike, then recovery within seconds.
|
||||
- `MongoSdamObservationListener` reports a topology change (primary removed, new primary elected).
|
||||
- `MongoPoolObservationListener` shows checkout wait times rising while server-side command duration
|
||||
stays flat — the wait is topology, not query cost.
|
||||
- Latency spike on writes with no corresponding rise in read latency.
|
||||
|
||||
A failover that resolves in under ~15 s and produces no `WRITE_RESULT_UNKNOWN` is normal replica-set
|
||||
behaviour and needs no action beyond confirming it self-healed.
|
||||
|
||||
## Diagnosis
|
||||
|
||||
1. Confirm an election actually happened. SDAM events distinguish an election from "the database got
|
||||
slow"; without them the two are indistinguishable in application metrics.
|
||||
2. Split the failure categories. Metric tag `failureCategory`:
|
||||
- `SERVER_SELECTION` / `CONNECTION` → the driver could not reach a primary. `NOT_SENT`; safe.
|
||||
- `TIMEOUT` with outcome `WRITE_RESULT_UNKNOWN` → a write may have applied. Not safe; see below.
|
||||
- `TRANSACTION_COMMIT_UNKNOWN` → go to [unknown-commit.md](unknown-commit.md) instead.
|
||||
3. Check the election duration against `MongoRetryBudget`. If the election outlasted the budget, the
|
||||
retries were exhausted before a primary existed and callers saw errors that a longer budget would
|
||||
have absorbed.
|
||||
4. Check whether the new primary is in the expected region/AZ. A failover to a distant node changes
|
||||
write latency permanently, not transiently.
|
||||
|
||||
## Action
|
||||
|
||||
**Self-healed (the common case).**
|
||||
Confirm outcome distribution contains no `WRITE_RESULT_UNKNOWN`, record the election in the incident
|
||||
log, and close. Nothing to replay.
|
||||
|
||||
**Writes with `WRITE_RESULT_UNKNOWN`.**
|
||||
These writes may or may not have applied. Do not blind-retry.
|
||||
- Idempotent operation (registered `MongoUpdateOperator` with an `AtomicFilter` precondition): retry.
|
||||
The precondition makes the second application a no-op.
|
||||
- Non-idempotent operation: reconcile by reading the target document and comparing against the
|
||||
intended post-state. Retry only if it does not reflect the write.
|
||||
|
||||
**Server selection never recovers.**
|
||||
The set has lost quorum — two of three nodes are down or partitioned. No client-side action fixes
|
||||
this; escalate to the database owner to restore a majority. The application should be failing closed,
|
||||
not queueing.
|
||||
|
||||
**Elections are frequent (more than one a day, unprompted).**
|
||||
This is an infrastructure symptom, not an application one: check node resource saturation, disk
|
||||
latency on the primary, and network stability between members. Repeated elections cause repeated
|
||||
unknown-outcome windows.
|
||||
|
||||
## Escalation
|
||||
|
||||
- P2 → P1 if server selection has failed for more than 2 minutes, or if any non-idempotent write
|
||||
returned `WRITE_RESULT_UNKNOWN` and cannot be reconciled.
|
||||
- Page the database owner for quorum loss, and the service owner for reconciliation of ambiguous
|
||||
writes.
|
||||
|
||||
## Verification
|
||||
|
||||
The failover lane reproduces this deliberately:
|
||||
|
||||
```bash
|
||||
cd src
|
||||
./gradlew :adapter:outbound:persistence-mongo:mongoFailoverTest --console=plain
|
||||
```
|
||||
|
||||
It starts a real three-node set (`MongoThreeNodeReplicaSet`), stops the primary
|
||||
(`MongoPrimaryController`), and injects network faults through Toxiproxy
|
||||
(`ToxiproxyMongoNetworkFaultController`). A single-node set is not sufficient for the election: it
|
||||
never holds one, so every guarantee that depends on a primary change goes untested.
|
||||
|
||||
The network faults need their own fixture (`MongoProxiedReplicaSetNode`) because a stopped container
|
||||
cannot produce them. Stopping a node tells the client the write did not happen; cutting the *path*
|
||||
while the server keeps running produces a client that cannot tell. `MongoNetworkFaultLaneTest`
|
||||
asserts the difference by reaching the same server twice — once through the proxy, once directly:
|
||||
|
||||
- **Partition**: the proxied client fails, the direct client finds the server healthy and the earlier
|
||||
write intact. The path was cut, not the server.
|
||||
- **Response loss**: the proxied client fails, and the direct client then finds the document
|
||||
*present*. The write applied and only the acknowledgement was lost —
|
||||
`DefaultMongoFailureClassifier` returns `WRITE_RESULT_UNKNOWN`, and a retry would have inserted a
|
||||
second document.
|
||||
|
||||
One detail the lane depends on: the connection is warmed before the toxic is applied. On a cold
|
||||
connection it is the driver's handshake whose response is dropped, so the write is never transmitted
|
||||
— `NOT_SENT`, the opposite of the ambiguity being tested.
|
||||
@@ -0,0 +1,90 @@
|
||||
---
|
||||
title: Runbook — MongoDB change stream history lost
|
||||
category: mongodb
|
||||
severity: P1
|
||||
owner: oncall
|
||||
last_updated: 2026-08-13
|
||||
status: active
|
||||
---
|
||||
|
||||
# Runbook: change stream history lost
|
||||
|
||||
Design §20.3, scenarios `OPLOG_HISTORY_LOSS`, `RESUME_TOKEN_LOSS`.
|
||||
|
||||
The stored resume token predates the oldest entry in the oplog. The events between the checkpoint and
|
||||
now are gone from the server; no retry recovers them. `MongoChangeStreamRecoveryPolicy` returns
|
||||
`halt(HISTORY_LOST, "history-lost")` and the runner stops.
|
||||
|
||||
**The platform will not silently restart from "now".** That looks like a recovery and is actually a
|
||||
permanent, unreported gap in the projection.
|
||||
|
||||
## Symptoms
|
||||
|
||||
- `MongoChangeHistoryLostException`.
|
||||
- `MongoChangeStreamState.HISTORY_LOST`; the consumer is stopped, not looping.
|
||||
- Precedes it: consumer lag approaching the oplog window, or a consumer that was down for a long
|
||||
period (a deploy that failed, a scaled-to-zero worker, a long outage).
|
||||
|
||||
## Diagnosis
|
||||
|
||||
1. **Determine the gap.** The checkpoint's cluster time is the start; the oldest oplog entry is the
|
||||
end of what is unrecoverable. Everything in between was never processed.
|
||||
2. **Determine the oplog window.** `rs.printReplicationInfo()` on the primary gives the first and last
|
||||
oplog timestamps. If the window is materially smaller than it was, the write rate rose or the
|
||||
oplog was resized — the consumer may be fine and the server changed.
|
||||
3. **Determine what the projection is missing.** Which collections and which operations does this
|
||||
projector consume? The gap is bounded by that, not by everything that happened.
|
||||
4. **Check for a second consumer.** If another projector on the same collection is healthy, its
|
||||
checkpoint tells you whether the problem is this consumer or the oplog.
|
||||
|
||||
## Action
|
||||
|
||||
Resuming is not an option. The choices are:
|
||||
|
||||
**Rebuild from source.** If the projection is derivable from the current state of the source
|
||||
collections, rebuild it: stop the consumer, rebuild the projection, then start the stream from the
|
||||
cluster time at which the rebuild snapshot was taken. This is the correct answer whenever the
|
||||
projection is a materialised view rather than an event log, and it is the reason a projection should
|
||||
be derivable.
|
||||
|
||||
**Backfill the gap.** If the source documents carry a timestamp covering the gap, run a bounded
|
||||
backfill for that window through the migration runner (checkpointed, resumable — see
|
||||
[schema-index-migration-guide.md](../schema-index-migration-guide.md) §5), then resume from the
|
||||
current cluster time.
|
||||
|
||||
**Accept the gap explicitly.** Only when the projection is advisory and the business owner says so.
|
||||
Record the window in the incident log and reset the checkpoint. This is a decision someone signs, not
|
||||
a default.
|
||||
|
||||
Never: reset the checkpoint to "now" and restart quietly. That converts a visible P1 into an
|
||||
invisible data-quality defect that surfaces months later as "the report has been wrong since March".
|
||||
|
||||
## Prevention
|
||||
|
||||
- **Alert on lag against the oplog window, not wall-clock.** "Consumer is 30 minutes behind" is fine
|
||||
with a 24-hour oplog and an emergency with a 45-minute one. The threshold that matters is
|
||||
`lag / oplogWindow`.
|
||||
- **Size the oplog for the longest tolerable consumer outage**, including a failed deploy discovered
|
||||
the next morning.
|
||||
- **Checkpoint after processing, never on receipt** — see
|
||||
[change-stream-guide.md](../change-stream-guide.md) §3.
|
||||
- **Back up the checkpoint store.** `RESUME_TOKEN_LOSS` is the same incident reached from the other
|
||||
direction: the oplog is fine, the checkpoint is gone.
|
||||
- **Make the projection rebuildable.** A projection that can only be built by replaying every event
|
||||
has no recovery path once the oplog rolls.
|
||||
|
||||
## Escalation
|
||||
|
||||
- P1 on detection. The consumer is stopped, so lag grows for as long as this is unresolved.
|
||||
- Page the service owner for the rebuild decision, and the database owner if the oplog window shrank
|
||||
unexpectedly.
|
||||
|
||||
## Verification
|
||||
|
||||
```bash
|
||||
cd src
|
||||
./gradlew :adapter:outbound:persistence-mongo:mongoFailoverTest --console=plain
|
||||
```
|
||||
|
||||
`MongoFailoverScenario.OPLOG_HISTORY_LOSS` and `RESUME_TOKEN_LOSS` assert the runner halts and names
|
||||
this runbook rather than restarting from the current position.
|
||||
@@ -0,0 +1,87 @@
|
||||
---
|
||||
title: Runbook — MongoDB unknown transaction commit result
|
||||
category: mongodb
|
||||
severity: P1
|
||||
owner: oncall
|
||||
last_updated: 2026-08-13
|
||||
status: active
|
||||
---
|
||||
|
||||
# Runbook: unknown transaction commit result
|
||||
|
||||
Design §16, decision D-10, scenario `UNKNOWN_TRANSACTION_COMMIT_RESULT`.
|
||||
|
||||
`MongoExecutionOutcome.TRANSACTION_COMMIT_UNKNOWN` means the commit **may have applied**. It is not a
|
||||
failure and must never be reported to a caller as one. The single worst response is to re-run the
|
||||
transaction body: if the commit did apply, the body applies a second time.
|
||||
|
||||
## Symptoms
|
||||
|
||||
- `MongoTransactionCommitUnknownException` in logs.
|
||||
- Metric `failureCategory=TRANSACTION_COMMIT_UNKNOWN`.
|
||||
- Usually accompanies a primary election — see [failover.md](failover.md).
|
||||
- Downstream reports of duplicated effects (double charge, double increment) are the symptom of this
|
||||
being handled wrongly, not of the condition itself.
|
||||
|
||||
## Diagnosis
|
||||
|
||||
1. **Confirm the platform did the right thing automatically.**
|
||||
`MongoTransactionRetryCoordinator` retries the *commit only*, on the same session, within
|
||||
`MongoRetryBudget`. A commit retry against an already-committed transaction is a no-op by design.
|
||||
Most occurrences resolve here and never reach a human.
|
||||
|
||||
2. **If the budget was exhausted, determine the actual state.** The commit either applied or it did
|
||||
not; you must find out which, not guess.
|
||||
- If the transaction body wrote a deterministic marker (an idempotency key, a business id, a
|
||||
revision), read it back. That is exactly what `MongoCommitReconciler` does, and it is the
|
||||
reason the design requires transactions to write one.
|
||||
- If there is no marker: reconstruct from a downstream artefact — an outbox row, an audit record,
|
||||
an external side effect. If nothing exists to compare against, the transaction was not
|
||||
designed to be reconcilable and that is the finding to record.
|
||||
|
||||
3. **Check whether the body was replayed.** Grep for a second execution with the same operation name
|
||||
and correlation id. If the body ran twice, the effects need reversing, and the code path that
|
||||
replayed it is a defect: an ambiguous commit is `COMMIT_ONLY` scope
|
||||
(`MongoRetryScope.COMMIT_ONLY`), never `BODY`.
|
||||
|
||||
## Action
|
||||
|
||||
**Commit applied.** Nothing to do. Record the reconciliation.
|
||||
|
||||
**Commit did not apply.** Re-run the whole operation from the top — a new session, a new body
|
||||
attempt. This is safe precisely because you established the previous attempt left no trace.
|
||||
|
||||
**Cannot determine.** Do not retry. Escalate. A blind retry here is a coin flip between "no effect"
|
||||
and "duplicate effect", and duplicates in a financial or notification path are worse than a delay.
|
||||
Freeze the affected entity if the domain supports it, and hand off with: operation name, correlation
|
||||
id, document id, the time window, and what you checked.
|
||||
|
||||
**Recurring.** More than one an hour means the commit path is racing something structural — a
|
||||
`maxCommitTime` shorter than the observed election duration, an oversized transaction, or an
|
||||
undersized retry budget. Fix the budget or the transaction shape; do not raise the retry count and
|
||||
call it resolved.
|
||||
|
||||
## Prevention
|
||||
|
||||
- Every transaction body writes a deterministic marker that identifies its own commit.
|
||||
- `maxCommitTime` exceeds the observed p99 election duration.
|
||||
- Callers surface the ambiguity to their own callers rather than mapping it to a generic 500 — an
|
||||
ambiguous outcome reported as a failure invites the caller to retry, which is the one thing that
|
||||
must not happen.
|
||||
- Prefer a single-document atomic operation (D-09). A transaction that exists only to wrap one
|
||||
document write has invented this failure mode for nothing.
|
||||
|
||||
## Escalation
|
||||
|
||||
- Always P1 when the state cannot be determined and the operation has an external effect.
|
||||
- Page the service owner immediately; the database owner only if elections are the trigger.
|
||||
|
||||
## Verification
|
||||
|
||||
```bash
|
||||
cd src
|
||||
./gradlew :adapter:outbound:persistence-mongo:mongoFailoverTest --console=plain
|
||||
```
|
||||
|
||||
`MongoFailoverScenario.UNKNOWN_TRANSACTION_COMMIT_RESULT` runs this path against a real three-node
|
||||
set, and the coordinator test asserts the body is never replayed after a commit ambiguity.
|
||||
Reference in New Issue
Block a user