228 KiB
Messaging Production Capability Deep Design
- 작성일: 2026-07-28
- 상태: 상세 설계 승인, 실행 계획 작성·독립 검토 완료, 구현 미착수
- 기준: Java 21, Spring Boot 4.0.0, Gradle 멀티모듈 Clean Architecture
- 현재 outbound leaf:
adapter-outbound-messaging - 미래 inbound leaf:
adapter-inbound-messaging-kafka - 상위 문서: Production Capability Platform Design
0. 문서 상태와 구현 상태
이 문서는 messaging 전체 수명주기를 한 번에 설계하되 구현은 단계적으로 진행하기 위한 정본이다. 여기서 messaging 전체 수명주기는 다음을 뜻한다.
domain event
-> integration event
-> transactional outbox
-> polling 또는 CDC dispatch
-> acknowledgement-aware Kafka producer
-> Kafka consumer
-> inbox + application side effect
-> DLT / replay / reconciliation
이 문서가 검토되었다는 사실은 위 기능이 구현되었거나 production-ready라는 뜻이 아니다. 구현 상태와 향후 추가 가능 범위를 혼동하지 않도록 세 종류의 표현만 사용한다.
| 표현 | 의미 |
|---|---|
현재 구현 |
2026-07-28 repository에서 코드와 테스트로 직접 확인한 범위 |
최초 R2 구축 대상 |
첫 실행 계획에서 실제로 구현하고 real-service 증거를 만들 범위 |
후속 설계 완료 / 미구현 |
경계와 보장은 이 문서에서 결정했지만 코드·설정·증거는 아직 없는 범위 |
0.1 현재 구현
현재 repository에는 다음 기반이 있다.
application-core의 framework-free transactional outbox append/store/publish port;- 비즈니스 쓰기와 같은
TransactionPort.inWrite(...)안에서 outbox event를 append하는 계약; - PostgreSQL
SKIP LOCKED기반 claim,PENDING/IN_FLIGHT/PUBLISHED/FAILED/DEAD상태, retry/backoff, timestamp 기반 aggregate FIFO gate; - broker publish를 DB transaction 밖에서 수행하고 결과 상태만 짧은 transaction으로 갱신하는 relay use case;
app.messaging.broker로 단일MessageBroker를 선택하는 outbound composition;KafkaSender라는 project-supplied seam과 fake 기반 unit test;- 비활성 시 fail-fast하는
DisabledMessagePublisher와DisabledOutboxMessagePublisher; - 일반 publisher의 fail-open과 durable outbox publisher의 fail-closed 구분;
- 확인된 FAILED/DEAD 전이 뒤에만 기록되는 typed
OutboxRelayFailureReport; - outbox backlog/lag/outcome metric과 stub runbook;
sample-portfolio의 WorkLog/Poster integration-event 예시.
이 기반이 증명하지 않는 것은 다음과 같다.
- 실제 Kafka client가 존재한다는 것;
KafkaSender.send()반환이 broker acknowledgement를 뜻한다는 것;- 현재
PUBLISHED상태가 실제 broker ACK 뒤에만 기록된다는 것; - 현재 hand-written JSON envelope가 schema-valid 또는 rolling-compatible하다는 것;
- consumer가 존재하거나 중복을 inbox로 흡수한다는 것;
- TLS/SASL, ACL, topic topology, resource bound, graceful drain이 준비되었다는 것;
- CDC mode나 polling/CDC 전환이 가능하다는 것;
- real Kafka/PostgreSQL/Kafka Connect 장애 시험을 통과했다는 것.
0.2 최초 R2 구축 대상
첫 R2 reference tuple은 다음 하나다.
producer-provider = kafka-spring
producer-semantics = acknowledged-idempotent-v1
outbox-dispatch = postgresql-polling-v2
claim-strategy = postgresql-per-record-jit-claim-v1
wire-format = json-schema-envelope-v1
topic-management = externally-provisioned-and-validated-v1
security = sasl-ssl-scram-sha-512-v1
compression = none-v1
ordering = per-key-normal-path-sequence-detectable-v1
transaction-resource = same-postgresql-transaction-resource-v1
operator-control = authenticated-internal-web-disposition-v1
consumer = disabled
cdc = disabled
첫 구현은 다음 순서로 하나의 실제 경로를 만든다.
- logical destination과 versioned event contract catalog;
- UTF-8 JSON envelope v1과 checked-in JSON Schema;
- immutable
outbox_event, polling-onlyoutbox_delivery, append-only attempt journal; - publication epoch, per-record JIT claim, claim token/valid lease, aggregate sequence, bounded retry/attempt budget;
Spring Kafka의KafkaTemplate/ProducerFactory를 직접 소유하는 outbound provider;- broker ACK를 기다리는 typed outcome;
- TLS/SASL_SSL production profile, finite queue/timeout, readiness와 graceful shutdown;
- application disposition use case + authenticated internal web operator control;
- PostgreSQL + real Kafka 통합·장애·보안 evidence.
첫 R2 구축에는 inbound Kafka consumer, inbox, retry topic, DLT replay, Debezium/Kafka Connect가 들어가지 않는다. 다만 최초 wire/outbox 구조가 그 후속 기능을 갈아 끼우거나 추가할 수 있도록 설계한다.
0.3 후속 설계 완료 / 미구현
| Capability | 문서상 결정 | 현재 구현 |
|---|---|---|
| inbound Kafka | 별도 adapter:inbound:messaging-kafka leaf |
없음 |
| consumer acknowledgement | application commit 뒤 MANUAL_IMMEDIATE |
없음 |
| inbox | (consumerId, eventId) unique + business write와 같은 DB transaction |
없음 |
| consumer retry | 짧고 bounded한 blocking retry가 기본 | 없음 |
| retry topic | ordering을 잃는 opt-in card | 없음 |
| DLT | DLT publish ACK 뒤 원본 offset 진행 | 없음 |
| replay | 별도 group/job, 범위·승인·audit 필수 | 없음 |
| CDC | 외부 Kafka Connect/Debezium, insert-only event source | 없음 |
| polling/CDC 전환 | 같은 production destination에서 상호 배타, 별도 cutover runbook | 없음 |
| Avro/Protobuf | schema-registry와 함께 optional serialization card | 없음 |
| Kafka transaction/EOS | DB-free Kafka consume-process-produce에만 optional | 없음 |
| 대체 broker | 동일 semantic guarantee/evidence를 만족하는 provider card로만 추가 | 없음 |
0.4 상태 ledger
이 표는 구현 진척의 human-readable SSOT다. 이후 구현 작업은 이 표만 갱신하고 완료 표현을 본문 여러 곳에 복제하지 않는다.
| Phase | 산출물 | 2026-07-28 상태 | 허용 표현 |
|---|---|---|---|
| P0 | current truth, design, characterization | IMPLEMENTATION_PLAN_READY |
설계·실행 계획 완료, 구현 미착수 |
| P1 | event contract/catalog/envelope/schema | NOT_STARTED |
미구현 |
| P2 | immutable event + polling delivery v2 | NOT_STARTED |
legacy polling만 존재 |
| P3 | Spring Kafka ACK-aware producer | NOT_STARTED |
Kafka seam R0 |
| P4 | security/observability/fault/real-service R2 evidence | NOT_STARTED |
R2 주장 금지 |
| P5 | inbound Kafka leaf + inbox + DLT/replay | DESIGNED_NOT_IMPLEMENTED |
후속 설계 |
| P6 | PostgreSQL Debezium CDC + cutover | DESIGNED_NOT_IMPLEMENTED |
후속 설계 |
| P7 | Avro/Protobuf/EOS/대체 provider cards | OPTIONAL_BACKLOG |
후보 |
Phase 진행도와 capability readiness는 별도 축이다. 예를 들어 P3 코드가 존재해도 exact selected profile이 real Kafka, 보안, fault, shutdown evidence를 통과하지 않으면 R2가 아니다.
1. 설계 판정
이번 설계는 provider-neutral semantic contract와 Kafka reference implementation을 분리한다.
provider-neutral이라는 말은 모든 broker의 최저 공통분모를 가진 범용
send(topic, key, payload) API를 만들겠다는 뜻이 아니다.
선택한 구조는 다음과 같다.
- Application은 integration event의 의미, logical destination, identity, ordering intent, transactional append와 consumption policy를 소유한다.
adapter:outbound:messaging은 wire contract compilation, JSON envelope, Kafka producer, ACK outcome, producer lifecycle과 provider telemetry를 소유한다.adapter:outbound:persistence-jpa는 same-store outbox/inbox persistence와 claim CAS를 소유한다.- 첫 reference provider는 outbound leaf 내부의 Spring Kafka
DefaultKafkaProducerFactory+KafkaTemplate이다. - 현재
KafkaSender는 legacy/test migration seam으로만 남길 수 있고 R2 provider로 표시하지 않는다. - durable publication의 성공은 broker ACK metadata가 확인된 경우에만 선언한다.
- ACK를 기다리다 deadline/cancellation/connection loss가 발생하면 성공이나 확정 실패가 아니라
INDETERMINATE다. - 첫 serialization은 versioned UTF-8 JSON envelope와 checked-in JSON Schema다.
- event type을 physical Kafka topic으로 직접 사용하지 않는다. logical destination을 deployment binding이 physical topic으로 컴파일한다.
- polling outbox는 immutable event와 mutable delivery control을 분리한다.
- CDC는 첫 R2에 포함하지 않고 같은 immutable event source를 사용하는 후속 dispatch card다.
- consumer를 구현하기 전 inbound Kafka 전용 leaf를 registry에 추가한다.
- 자동 publication은 bounded하므로 무조건적인 eventual delivery를 주장하지 않는다. consumer/inbox까지 구현된 범위만 duplicate-possible delivery와 idempotent effect로 표현한다.
- Kafka producer idempotence나 transaction을 DB와 Kafka 사이의 generic exactly-once로 표현하지 않는다.
- 사용하지 않는 producer/consumer/CDC profile은 connection, thread, scheduler, AdminClient, connector, schema가 0개여야 한다.
이 결정의 핵심은 “Kafka로 먼저 하나를 만든다”와 “나중에 교체 가능하게 한다”를 동시에 만족하는 것이다. 교체 가능성은 빈 SPI 하나가 아니라 stable semantic contract, explicit capability card, exact profile validation, provider별 evidence로 확보한다.
2. 상위 설계와 기존 심화 설계의 관계
상위 Production Capability Platform Design은 다음을 이미 결정했다.
- transactional outbox append는 source-of-truth transaction과 함께한다;
- dispatch는
disabled | polling | cdc로 선택한다; - immutable
outbox_event와 polling-onlyoutbox_delivery를 분리한다; - physical topic은 adapter 설정이고 application event type이 아니다;
- real Kafka producer는 ACK, bounded delivery timeout, idempotence, security를 가져야 한다;
- consumer는 별도 inbound leaf, manual acknowledgement, inbox, DLT/replay를 사용한다;
- DB와 broker를 아우르는 generic exactly-once를 주장하지 않는다.
이번 문서는 그 방향을 구현자가 다시 추론하지 않도록 다음을 추가로 고정한다.
- 첫 R2 exact tuple;
- stable contract와 replaceable capability card의 경계;
- integration event identity와 versioned wire envelope;
- logical destination catalog와 physical topic binding;
- producer ACK/REJECTED/INDETERMINATE state machine;
- Kafka client retry와 relay retry의 combined amplification budget;
- immutable event/polling delivery schema와 claim token;
- polling ACK-to-DB gap, dead resolution, replay와 retention;
- future consumer/inbox/DLT/replay의 정확한 transaction과 acknowledgement 순서;
- future CDC connector, slot/offset/WAL, shadow/cutover/rollback 조건;
- configuration expected-state, readiness card와 evidence fingerprint;
- 보안, privacy, observability, test/CI/no-skip와 runbook 요구.
Redis/FileServer/HTTPClient deep design에서 재사용하는 공통 패턴은 다음이다.
- 현재 truth와 목표를 문서 앞에서 분리한다;
- semantic contract와 provider runtime을 분리한다;
- logical ID와 physical endpoint/topic을 분리한다;
- exact-one provider selection과 disabled resource-0를 사용한다;
- typed outcome으로 definite/indeterminate를 구분한다;
- finite deadline, queue, body/message size와 graceful shutdown을 계약으로 둔다;
- readiness를 capability card와 real-service evidence로 제한한다;
- optional provider/config를 구현되기 전에 존재하는 것처럼 노출하지 않는다.
그 문서에서 그대로 복사하지 않는 부분은 다음이다.
- HTTP mutation의 unknown outcome과 Kafka duplicate 가능성은 비슷하지만 동일하지 않다;
- Redis fail-open cache 정책은 durable messaging에 적용하지 않는다;
- Fileserver의 atomic rename/manifest가 Kafka acknowledgement를 대체하지 않는다;
- Kafka partition ordering을 global FIFO나 distributed lock으로 표현하지 않는다;
- Kafka transaction은 DB outbox transaction을 대체하지 않는다.
세부 내용이 상위 문서의 messaging 요약과 다르면 이 문서가 messaging 범위의 정본이다. 다른 capability 결정은 변경하지 않는다.
2.1 Normative decision ledger
| 결정 | 정본 절 |
|---|---|
| 현재 구현/후속 상태 | §0 |
| 첫 R2 exact tuple | §0.2, §10 |
| guarantee와 readiness 용어 | §7 |
| architecture와 module ownership | §8–§9 |
| capability card와 교체 조건 | §10 |
| identity와 ordering vocabulary | §11 |
| event transformation | §12 |
| contract/destination/topic catalog | §13 |
| JSON envelope와 schema evolution | §14 |
| application outcome contract | §15 |
| ACK-aware Kafka producer | §16–§17 |
| polling outbox state/data | §18–§19 |
| best-effort 경계 | §20 |
| activation/configuration | §21 |
| security/topic governance | §22 |
| observability/readiness | §23 |
| consumer/inbox/DLT/replay | §24–§25 |
| CDC와 mode cutover | §26–§27 |
| test/CI/evidence | §28–§29 |
| migration/status update | §30 |
| 완료/후속 card | §31 |
| runbook | §32 |
예시 YAML, Java pseudocode, migration alias, README, runbook은 이 ledger의 정본 절보다 우선하지 않는다. 두 activation source가 충돌하면 임의 precedence나 fallback을 적용하지 않고 startup을 실패시킨다.
3. 증거 기반 현재 상태
3.1 실제 Kafka production dependency가 없다
adapter:outbound:messaging/build.gradle의 production dependency는 현재 다음뿐이다.
application-core
shared-contract
adapter:outbound:support
spring-boot-autoconfigure
slf4j-api
spring-kafka와 kafka-clients가 없으므로 production KafkaProducer,
ProducerFactory, KafkaTemplate, AdminClient도 없다. app-bootstrap의
testCompileOnly kafka-clients는 architecture test classpath를 위한 것이며 실제 provider가
아니다.
따라서 현재 app.messaging.broker=kafka는 Kafka capability 활성화가 아니라 fork project가
KafkaSender bean을 별도로 제공했을 때 seam을 선택한다는 의미다.
3.2 void KafkaSender.send()는 broker ACK를 표현하지 못한다
현재 contract는 다음과 같다.
void send(OutboundMessage message) throws Exception;
Kafka send()는 일반적으로 local buffer에 record를 넣고 future를 즉시 반환한다. 외부 seam이
future를 기다리는지, 어떤 acks를 쓰는지, delivery timeout이 유한한지 repository는 알 수 없다.
그런데 application의 OutboxEventStatus.PUBLISHED 문서는 broker acknowledgement를 뜻한다고
설명한다.
현재 seam 구현자가 callback/future 완료 전에 정상 반환하면 relay는 ACK가 없는 record를
PUBLISHED로 기록한다. 이는 단순한 구현 누락이 아니라 현재 상태 이름과 실제 evidence의
불일치다.
3.3 event type이 physical topic으로 사용된다
현재 OutboxMessagePublishAdapter는 다음 매핑을 한다.
topic = event.eventType()
key = event.aggregateId()
payload = hand-written envelope
이 구조는 application event naming이 Kafka topic naming, ACL, retention, partition count, replication, environment naming과 결합되게 한다. event type을 동적으로 만들 수 있으면 arbitrary topic publication과 metric cardinality도 열린다.
목표에서는 contractId -> logicalDestination -> physicalTopicBinding의 닫힌 두 단계 mapping을
사용한다.
3.4 envelope가 schema-bound document가 아니다
OutboxEnvelopeJson은 문자열을 직접 이어 붙인다. payload는 “이미 올바른 JSON”이라는 주석
계약만 있고 parser/schema validation 없이 verbatim 삽입된다.
현재 envelope에는 다음도 없다.
- envelope spec version;
- payload schema version;
- contract ID;
- aggregate type와 aggregate sequence;
- causation ID;
- content type;
- logical destination;
- schema hash/compatibility evidence;
- maximum depth/string/array/record bytes;
- rolling producer/consumer compatibility fixture.
따라서 현재 JSON은 example wire shape이지 versioned integration contract가 아니다.
3.5 immutable event와 mutable delivery state가 한 행에 섞여 있다
현재 outbox_event에는 event metadata와 다음 polling state가 함께 있다.
status
attempt_count
next_attempt_at
relay는 같은 row를 PENDING -> IN_FLIGHT -> PUBLISHED/FAILED/DEAD로 UPDATE한다.
Debezium Outbox Event Router는 outbox table을 INSERT-only queue로 기대하고 UPDATE를 비정상
operation으로 분류한다. 현재 table에 connector flag만 켜는 방식으로 CDC를 추가할 수 없다.
3.6 현재 FIFO는 동일 timestamp와 긴 batch에서 불완전하다
현재 claim과 defensive sort는 occurred_at 중심이다. 같은 aggregate에서 같은 timestamp를 가진
두 event의 완전한 tie-breaker나 domain aggregate sequence가 없다. 따라서 strict ordering을
증명할 수 없다.
또한 batch row는 같은 시점에 now + inFlightTimeout으로 claim되고 순차 publish된다. batch의
최악 처리 시간이 in-flight timeout을 넘으면 뒤쪽 row를 첫 worker가 처리 중일 때 다른 worker가
재claim할 수 있다.
목표에서는 aggregate sequence, claim token, per-row remaining-lease validation, attempt budget과 claim lease의 관계를 고정한다.
3.7 polling ACK-to-DB gap은 이미 duplicate를 허용한다
현재 순서는 다음과 같다.
publishPort.publish(event)
-> 별도 DB transaction에서 markPublished(eventId)
broker가 record를 수락한 뒤 process가 종료되거나 markPublished가 실패하면 row는
IN_FLIGHT에 남고 timeout 뒤 재claim된다. 이는 올바른 transactional outbox에서 피할 수 없는
ACK-to-state duplicate gap이며 consumer dedupe가 필요하다.
현재 repository에는 inbound consumer와 inbox가 없으므로 “consumer dedupe가 안전하게 흡수한다”는 runbook 표현은 목표 계약이지 현재 보장이 아니다.
3.8 retry 분류가 모든 exception을 같은 경로로 보낸다
현재 publish exception은 attempt count가 남았으면 FAILED, 소진됐으면 DEAD가 된다. 다음을 구분하지 않는다.
- local schema/size violation처럼 절대 broker에 도달하지 않은 permanent rejection;
- authorization/topic-not-found처럼 configuration/operator 조치가 필요한 failure;
- retriable broker/network failure;
- broker가 수락했을 수도 있는 timeout/cancel/connection loss;
- programming defect;
- stale claim owner가 수행한 결과.
poison event도 max attempt까지 재시도하므로 불필요한 amplification과 aggregate head blocking을 만든다.
3.9 configuration은 topology와 guarantee를 표현하지 못한다
현재 typed settings는 사실상 다음뿐이다.
app.messaging.broker
app.messaging.kafka.brokers
다음이 없다.
- expected state;
- producer semantic profile;
- logical destination binding;
- contract/schema catalog;
- dispatch mode;
- acknowledgement deadline;
- delivery/request/max-block timeout;
- buffer/in-flight/batch/record size;
- security protocol, TLS, SASL, secret reference;
- topic partitions/replication/min ISR/retention expectation;
- readiness requirement;
- shutdown drain;
- selected capability/evidence digest.
host:port regex도 IPv6, duplicate endpoint, port range, blank-normalization과 secret/source policy를 충분히 검증하지 않는다.
3.10 best-effort와 durable success vocabulary가 섞일 수 있다
일반 OutboundMessagePublisher는 exception을 삼키는 fail-open이고 outbox publisher는
exception을 전파하는 fail-closed다. 이 구분 자체는 유용하다.
그러나 둘 다 같은 MessageBroker.send(void)를 호출하므로 다음을 구분하지 못한다.
- local admission;
- producer buffer enqueue;
- broker acknowledgement;
- definite rejection;
- indeterminate result.
일반 publisher의 logSuccess도 broker ACK가 아니라 seam의 정상 반환만 의미할 수 있다.
3.11 consumer/inbox/CDC runtime이 없다
현재 registry에는 19개 leaf만 있고 inbound Kafka leaf가 없다. production source 검색 기준으로 다음도 없다.
@KafkaListener또는 listener container;- consumer group/offset/ack policy;
- deserializer allowlist;
- inbox table/port/executor;
- retry/DLT/replay;
- rebalance/pause/resume/drain;
- Debezium connector/Kafka Connect deployment;
- replication slot/offset/WAL monitoring.
이 기능은 package를 outbound leaf에 추가하지 않고 각각 §24–§27의 단계에서 도입한다.
3.12 현재 test와 runbook이 증명하는 범위
현재 focused messaging test는 fake sender/broker와 settings/composition/report rendering을
검증한다. PostgreSQL outbox integration test는 same-transaction append, row lifecycle,
normal-path two-worker SKIP LOCKED row partitioning과 일부 claim 동작을 검증한다.
현재 test가 증명하지 않는 것은 다음이다.
- real broker ACK metadata;
- leader loss/min ISR/timeout/duplicate;
- real buffer saturation;
- TLS/SASL/ACL;
- topic drift;
- producer close/drain;
- schema compatibility;
- consumer rebalance/inbox;
- CDC restart/slot/offset.
outbox-publish-failed runbook은 이미 제거된 APP_MESSAGING_KAFKA_ENABLED와 존재하지 않는
KafkaOutboxMessagePublishAdapter를 참조한다. outbox-dead-letter runbook은 raw SQL로 상태를
직접 수정하며 operator identity, reason, CAS, audit generation이 없다. 두 문서는 R2 구현과 함께
도구 기반 절차로 교체해야 한다.
4. 범위와 명시적 비범위
4.1 전체 설계 범위
이 문서는 다음을 설계한다.
- domain event와 integration event의 분리;
- framework-free event metadata와 application outbox contract;
- logical destination, contract, payload schema, topic binding catalog;
- JSON envelope v1, schema evolution와 compatibility evidence;
- ACK-aware Kafka producer와 typed certainty;
- producer retry, ordering, batching, compression, resource bound와 lifecycle;
- immutable event + polling delivery state;
- claim token, retry/dead/replay/retention;
- best-effort와 durable publication 구분;
- typed activation, expected state와 capability cards;
- TLS/SASL, ACL, topic governance, secret rotation;
- metrics/tracing/log/readiness;
- future inbound Kafka leaf;
- consumer manual ack, inbox, retry, DLT, replay;
- future PostgreSQL Debezium CDC;
- polling/CDC shadow, cutover, rollback;
- real-service/fault/security/compatibility CI.
4.2 최초 R2 baseline
최초 R2는 다음 common subset만 구현한다.
- Kafka 한 provider;
- polling outbox 한 dispatch mode;
- JSON Schema 한 serialization profile;
- external topic provisioning + startup validation;
- acknowledged idempotent producer;
- production TLS/SASL_SSL와 explicit local plaintext profile;
- single-cluster, single-region producer;
- bounded record, buffer, batch, retry, deadline와 shutdown;
- real PostgreSQL + Kafka qualification.
consumer와 CDC는 설계에 포함되지만 최초 R2 implementation acceptance에는 포함되지 않는다.
4.3 후속 optional capability
다음은 stable boundary 뒤에 추가할 수 있다.
- Kafka manual-ack consumer + PostgreSQL inbox;
- retry-topic/delayed retry;
- Debezium PostgreSQL CDC;
- Avro + schema registry;
- Protobuf + schema registry;
- JSON Schema registry;
- Kafka transactional consume-process-produce;
- Kafka Streams;
- alternative partitioner proven by ordering vectors;
- multi-cluster replication/failover;
- broker/provider 대체;
- contract-specific compaction;
- large-message claim-check pattern with object storage;
- module split 또는 external capability artifact.
optional card는 이름만 등록하지 않는다. code, typed settings, tests, runbook, evidence가 같은 변경에서 존재할 때 registry에 추가한다.
4.4 비범위
다음은 이번 설계의 목표가 아니다.
- arbitrary topic/key/header를 application에 노출하는 Kafka facade;
- 모든 broker를 lowest-common-denominator API로 감싸기;
- DB와 Kafka를 XA/distributed transaction으로 묶기;
- generic exactly-once delivery 주장;
- Kafka partition ownership을 distributed lock으로 사용하기;
- Kafka를 source-of-truth database로 일반화하기;
- unbounded event sourcing platform;
- dynamic user input으로 topic 생성하기;
- payload에 임의 Java class/type header를 넣기;
- active-active multi-region ordering을 보장하기;
- 대용량 binary를 Kafka record에 직접 싣기;
- sample-portfolio business contract를 production leaf에 넣기.
5. HARD invariants
다음 중 하나라도 위반하면 동작하는 코드라도 messaging 설계 구현 완료가 아니다.
domain-core는 Kafka, Spring, JSON library, database, transport 타입을 알지 않는다.application-core는KafkaTemplate,ProducerRecord,ConsumerRecord, offset, partition SDK 타입을 알지 않는다.- business/best-effort/outbox publication producer와 inbound consumer는 같은 leaf에 두지
않는다. inbound processing lifecycle에만 쓰이는 closed retry/DLT publisher는
adapter:inbound:messaging-kafka가 소유할 수 있는 유일한 명시적 예외이며 application publish, outbox relay 또는 arbitrary destination 전송에 재사용하지 않는다. - inbound Kafka leaf는 persistence 또는 outbound messaging leaf에 직접 의존하지 않는다.
- business write와 durable event append는 같은 source-of-truth transaction에 참여한다.
- event type을 physical topic으로 암묵 변환하지 않는다.
- application은 arbitrary topic/header/security property를 전달하지 않는다.
- 모든 publish는 closed contract catalog와 destination binding을 통과한다.
- wire envelope와 payload는 append 전에 versioned schema와 size limit을 통과한다.
- broker ACK metadata 전에는 durable publish 성공을 선언하지 않는다.
- future 반환/local enqueue를 broker ACK라고 부르지 않는다.
- timeout/cancel/late callback race를 definite failure로 축소하지 않는다.
INDETERMINATE는 broker acceptance 가능성을 보존한다.- producer idempotence를 process restart나 relay retry 전체의 dedupe로 표현하지 않는다.
- Kafka transaction을 DB + Kafka atomic commit으로 표현하지 않는다.
- end-to-end는 bounded source/retention 조건 안의 duplicate-possible delivery와 inbox-covered idempotent effect로만 표현한다.
outbox_event는 CDC qualification 전에 insert-only immutable source가 된다.- polling mutable state는
outbox_delivery에만 둔다. - polling과 CDC는 같은 production destination에서 동시에 발행하지 않는다.
- shadow CDC는 격리된 topic과 격리된 consumer group만 사용한다.
- strict aggregate ordering은 timestamp가 아니라 explicit aggregate sequence를 요구한다.
- active worker가 소유한 renew/outcome transition은 current claim token/owner와 unexpired DB-time lease를 CAS 조건으로 검증한다. initial claim, expired reclaim, operator disposition과 generation authority handoff는 §18.6의 각기 다른 fenced predicate를 사용한다.
- Kafka client retry와 relay retry는 하나의 finite amplification budget으로 검증한다.
- timeout, queue, buffer, in-flight, batch, payload, header, retry는 모두 finite다.
- automatic provider fallback과 automatic topic creation은 production에서 금지한다.
- disabled profile은 client, AdminClient, listener, scheduler, connector, refresh thread가 0개다.
- production plaintext와 literal secret은 startup fail-closed다.
- event payload, partition key, tenant, credential, raw headers는 log/metric tag에 넣지 않는다.
- consumer offset은 application transaction commit 뒤에만 진행한다.
APPLIED와 verifiedDUPLICATE만 바로 ACK할 수 있다.- poison record를 DLT로 보낼 때 DLT publish ACK 전에는 원본 offset을 진행하지 않는다.
- retry topic이 ordering을 잃는다는 사실을 숨기지 않는다.
- replay는 별도 audited operation이며 운영 group offset을 임의 rewind하지 않는다.
- R2는 exact selected profile의 real-service/security/fault/shutdown evidence가 있어야 한다.
- required qualification lane은 Docker나 broker 부재를 이유로 silent skip하지 않는다.
- 현재 구현되지 않은 consumer/CDC/schema-registry setting을 live config처럼 추가하지 않는다.
- production leaf는
sample-portfoliocontract나 fixture에 의존하지 않는다. - 모든 dependency edge는
src/config/architecture/modules.json변경과 검증을 통과한다. - nullable tenant scope를 unique/dedupe/order key에 사용하지 않는다.
- legacy/v2/polling/CDC relay authority는 같은 destination에서 항상 정확히 하나다.
- CDC commit order를 aggregate sequence order라고 표현하지 않는다.
- plain unique-violation catch 뒤 rollback-only PostgreSQL/JPA transaction을 계속 사용하지 않는다.
- consumer의 bounded retry가 소진되면 durable HOLD와 partition/container stop 중 하나로 automation을 끝낸다. recoverer 실패로 같은 retry cycle을 무기한 다시 시작하지 않는다.
6. 대안 검토
6.1 현재 KafkaSender seam만 확장
장점:
- broker SDK가 leaf에 없으므로 가볍다;
- fake unit test가 쉽다;
- fork project가 client를 자유롭게 선택할 수 있다.
문제:
- actual producer config와 lifecycle evidence를 template이 소유하지 못한다;
- ACK, timeout, buffer saturation, TLS/SASL, metrics를 검증할 수 없다;
- fork마다 guarantee가 달라져 같은 capability label을 사용할 수 없다.
판정: legacy/test seam으로만 유지한다. void/throws shape는 R2 선택 불가다.
6.2 Apache Kafka client 직접 사용
장점:
KafkaProducerlifecycle, callback, metrics, transaction을 가장 직접 통제한다;- Spring abstraction 없이 Kafka API 의미를 그대로 사용할 수 있다.
문제:
- producer factory, generation rotation, close/drain, observation, transaction cache, Spring lifecycle integration을 모두 직접 소유해야 한다;
- 현재 Spring Boot composition과 중복이 커진다.
판정: future provider candidate다. Spring Kafka가 보장을 막는 구체적 evidence가 있을 때만 추가한다.
6.3 Spring Kafka KafkaTemplate + explicit producer factory
장점:
- 실제 Kafka producer guarantee를 사용하면서 Spring lifecycle/composition과 정렬된다;
- future/ACK metadata, observation, test support를 사용할 수 있다;
- provider config를 outbound leaf가 직접 검증할 수 있다.
주의:
- Boot의 broad auto-configuration/default에 activation을 맡기지 않는다;
KafkaTemplate.send()반환 자체가 ACK가 아니므로 future 완료를 기다려야 한다;- shared producer에서 per-message
flush()를 사용하지 않는다; - provider settings를 Kafka raw map으로 무제한 노출하지 않는다.
판정: 첫 reference provider로 선택한다.
6.4 Spring Cloud Stream/binder를 baseline으로 사용
장점:
- binder 교체와 functional pipeline이 편리하다;
- broker-neutral developer experience를 제공할 수 있다.
문제:
- 현재 요구는 exact producer acknowledgement, outbox state, topic/security drift, client retry와 lifecycle을 직접 증명하는 것이다;
- binder abstraction이 provider-specific guarantee와 evidence ownership을 흐릴 수 있다;
- dependency와 activation 범위가 첫 reference path보다 크다.
판정: 첫 baseline으로 선택하지 않는다. 동일 semantic card와 exact evidence를 만족하는 future provider로 검토할 수 있다.
6.5 consumer를 outbound messaging leaf에 추가
장점:
- Kafka dependency와 설정을 한 곳에서 공유한다.
문제:
- producer는 driven adapter이고 consumer는 driving adapter다;
- outbound leaf가 use case invocation, offset lifecycle, rebalance를 소유하게 된다;
- adapter-to-adapter/persistence 직접 의존 유혹이 생긴다.
판정: 금지한다. consumer 구현 전 별도 inbound leaf를 추가한다.
6.6 CDC를 첫 dispatch로 구현
장점:
- Java polling scheduler와 ACK-to-mark gap을 제거한다;
- database log 기반 확장성이 좋을 수 있다.
문제:
- 현재 mutable table과 호환되지 않는다;
- Kafka Connect, Debezium, replication slot, WAL, offset topic, snapshot과 운영 범위가 함께 필요하다;
- 첫 실제 producer/contract도 없는 상태에서 장애 면적이 너무 크다.
판정: 전체 설계에는 포함하되 첫 구현은 polling이다.
6.7 polling만 설계하고 CDC는 나중에 처음부터 다시 설계
문제:
- mutable row/envelope가 굳으면 CDC 전환 비용이 커진다;
- topic/identity/wire parity가 dispatch 구현마다 갈라진다.
판정: 거부한다. 처음부터 immutable event와 replaceable dispatch card를 둔다.
6.8 JSON Schema, Avro, Protobuf
JSON Schema 장점:
- 현재 JSON 예제와 이행 거리가 짧다;
- schema 파일과 golden vectors를 repository에서 바로 review할 수 있다;
- registry 없이 첫 contract를 세울 수 있다.
JSON Schema 한계:
- compatibility를 schema diff heuristic만으로 완전히 증명할 수 없다;
- binary 효율과 generated type safety는 Avro/Protobuf보다 약할 수 있다.
Avro/Protobuf 장점:
- generated type과 schema registry ecosystem이 강하다;
- compact binary wire를 제공한다.
Avro/Protobuf 비용:
- registry availability/security/compatibility mode와 build generation을 함께 설계해야 한다;
- 현재 skeleton의 첫 R2 경로를 넓힌다.
판정: JSON Schema v1을 먼저 구축하고 Avro/Protobuf는 별도 evidence card로 추가한다.
6.9 Kafka transaction을 모든 durable publish에 사용
Kafka transaction은 Kafka 안의 여러 record와 offset을 원자화할 수 있다. 그러나 PostgreSQL business commit과 Kafka transaction을 하나의 atomic commit으로 만들지 않는다. Spring의 DB/Kafka transaction synchronization도 순차 commit이며 두 번째 commit 실패 가능성이 남는다.
판정: polling outbox producer baseline에서는 사용하지 않는다. DB-free Kafka consume-process-produce에만 optional card로 둔다.
7. Capability readiness와 guarantee vocabulary
7.1 Readiness level
| Level | 의미 |
|---|---|
| R0 | interface/seam/example만 존재; 실제 service guarantee 없음 |
| R1 | deterministic unit/contract/local composition은 검증; production topology/fault 증거 없음 |
| R2 | exact selected profile이 real service, security, fault, lifecycle, compatibility gate 통과 |
| R3 | HA/failover/upgrade/DR/capacity와 운영 rehearsal까지 통과 |
KafkaSender는 R0다. 현재 PostgreSQL polling control plane은 일부 real-DB test가 있으므로 R1
skeleton으로 설명할 수 있지만 end-to-end Kafka publication은 R0다.
7.2 Publication vocabulary
| 용어 | 정확한 의미 |
|---|---|
COMPILED |
contract + destination + provider profile이 startup에 검증됨 |
ADMITTED |
local bounded admission을 통과함 |
ENQUEUED |
producer local buffer가 record를 받음 |
ACKNOWLEDGED |
configured ACK 조건을 만족한 broker metadata를 local process가 관찰함 |
ACKNOWLEDGED_MISMATCH |
ACK metadata가 compiled destination과 달라 misroute incident가 됨 |
REJECTED |
provider가 record 비수락을 확정할 수 있음 |
INDETERMINATE |
broker가 수락했을 수도, 아닐 수도 있음 |
DELIVERY_RECORDED |
ACK 뒤 polling delivery row가 terminal success로 commit됨 |
CONSUMED |
consumer가 record를 읽음; business side effect 완료와 다름 |
APPLIED |
inbox + business effect transaction이 commit됨 |
OFFSET_COMMITTED |
APPLIED/DUPLICATE 뒤 Kafka offset이 진행됨 |
ACKNOWLEDGED와 DELIVERY_RECORDED 사이 crash는 duplicate를 만든다.
APPLIED와 OFFSET_COMMITTED 사이 crash도 duplicate delivery를 만든다. 둘 다 event ID와 inbox가
흡수해야 하는 duplicate-possible 경계다. ACKNOWLEDGED_MISMATCH는 retry 가능한 일반 실패가
아니며 producer admission과 relay scope를 멈추는 fatal misroute incident다.
7.3 허용 guarantee
최초 R2가 주장할 수 있는 표현:
same-store transactional outbox append
+ bounded automatic publish attempts
+ broker ACK 또는 명시적 unresolved/operator disposition
+ 같은 producer generation의 정상 경로에서 stable key별 Kafka order
+ aggregate sequence를 통한 gap/regression 탐지 가능성
+ idempotent Kafka producer within its supported producer session
+ explicit duplicate handling requirement
이는 무조건적인 eventual delivery 또는 failure-path strict FIFO가 아니다. finite budget이 끝난
EXHAUSTED, 승인된 SKIPPED/COMPENSATED, 영구 hold가 존재할 수 있다. 따라서 첫 R2 card의
정확한 표현은 durable acknowledged-or-explicit-disposition publication이다.
future consumer/inbox까지 구현한 뒤 주장할 수 있는 표현:
declared source/retention/disposition horizon 안의 duplicate-possible delivery
+ idempotent application effect for inbox-covered handlers
허용하지 않는 표현:
- exactly-once DB-to-Kafka;
- exactly-once end-to-end;
- global ordering;
- no duplicates;
- no loss across an unqualified CDC slot failover;
- DLT가 곧 성공 처리 또는 데이터 복구라는 표현.
7.4 Evidence identity
R2 evidence는 단순 test 이름이 아니라 다음 fingerprint에 묶인다.
semantic-card-version
provider-card-version
Kafka client version
Spring Kafka version
broker image/version
JDK version
security profile
topic profile
contract catalog hash
schema set hash
settings digest
test scenario version
다른 version/profile에 이전 evidence를 자동 승계하지 않는다.
8. 목표 아키텍처
flowchart LR
DOMAIN[Domain event] --> MAP[Application integration-event mapping]
MAP --> APPEND[OutboxAppendPort]
APPEND --> DBTX[(Business DB transaction)]
DBTX --> EVENT[(immutable outbox_event)]
DBTX --> DELIVERY[(polling outbox_delivery)]
DELIVERY --> RELAY[Application polling relay]
RELAY --> PUBPORT[OutboxMessagePublishPort]
PUBPORT --> CATALOG[Contract + destination compiler]
CATALOG --> KAFKA[Spring Kafka ACK-aware provider]
KAFKA --> TOPIC[(Kafka topic)]
EVENT -. future CDC .-> DBZ[Debezium / Kafka Connect]
DBZ -. same wire contract .-> TOPIC
TOPIC -. future consume .-> INBOUND[adapter:inbound:messaging-kafka]
INBOUND --> CUSE[Application consume use case]
CUSE --> INBOXTX[(Inbox + business + optional outbox transaction)]
INBOXTX --> ACK[Kafka offset ACK]
BOOT[app-bootstrap] -. compile exact profiles .-> CATALOG
BOOT -. compose .-> RELAY
BOOT -. future compose .-> INBOUND
OBS[Metrics / trace / readiness] -. bounded observation .-> KAFKA
OBS -. bounded observation .-> INBOUND
8.1 Stable plane
provider/dispatch가 바뀌어도 다음은 유지한다.
- event ID와 contract ID;
- schema version과 envelope version;
- logical destination;
- aggregate identity/sequence와 partition-key intent;
- publication certainty taxonomy;
- inbox dedupe identity;
- application transaction meaning;
- capability/evidence descriptor shape.
8.2 Replaceable plane
다음은 capability card로 교체할 수 있다.
- Kafka Spring provider / direct Kafka provider / future broker provider;
- polling / CDC;
- JSON Schema / Avro / Protobuf;
- blocking retry / retry topic;
- PostgreSQL inbox / other same-store inbox;
- local plaintext dev / TLS / SASL_SSL security;
- externally provisioned topic validation / future managed provisioning;
- non-transactional idempotent producer / Kafka transactional workflow.
8.3 교체가 아닌 것
다음은 silent fallback이며 금지한다.
- Kafka outage 때 다른 broker로 자동 전송;
- schema validation 실패 때 raw JSON으로 전송;
- CDC 장애 때 polling을 자동으로 동시에 켬;
- TLS secret 실패 때 plaintext로 연결;
- DLT publish 실패 때 원본 offset을 ACK;
- required consumer failure 때 record를 best-effort로 폐기.
9. 모듈과 계층 소유권
9.1 domain-core
소유:
- 순수 domain event와 aggregate invariant;
- aggregate version/sequence가 domain 의미일 때 그 증가 규칙.
금지:
- integration topic/destination;
- JSON/schema;
- outbox/inbox;
- Kafka header/partition;
- retry/DLT.
Domain event는 같은 bounded context 내부의 사실이다. Integration event는 외부 consumer와의 versioned contract이므로 자동으로 동일한 타입을 직렬화하지 않는다.
9.2 application-core
소유:
- integration-event draft의 framework-free metadata;
OutboxAppendPort;- polling relay orchestration;
- provider-neutral publication outcome/certainty;
- provider-neutral late-publication observation drain과 attempt observation port;
- producer generation 교체 시 durable
INDETERMINATE/HOLD, admission 재개와 generation barrier를 조정하는 provider-neutral rotation use case; - outbox requeue/hold/skip/compensate command use case와 authorization/audit policy;
- future
InboxStorePort와MessageConsumptionExecutor; - feature consume use case와 transaction policy;
- retry/dead decision의 application/operational policy.
금지:
- physical topic;
KafkaTemplate, record metadata SDK 타입;- connector/replication slot;
- JPA entity;
- Micrometer/SLF4J.
9.3 adapter:outbound:persistence-jpa
소유:
outbox_event,outbox_delivery, futureinbox_consumptionmigration/entity/repository;- PostgreSQL claim query와 claim-token + unexpired DB-time lease CAS;
- same-store transaction participation;
- attempt observation append와 audited disposition/authority-handoff persistence;
- polling retention query와 operator transition persistence;
- future inbox unique constraint.
금지:
- Kafka producer;
- topic routing;
- event business mapping;
- use-case retry policy.
9.4 adapter:outbound:messaging
소유:
- destination binding compiler;
- envelope/payload schema validation runtime;
- provider-private publish gateway;
- Spring Kafka producer factory/template/AdminClient;
- ACK mapping, finite deadline, buffer/admission;
- late completion을 payload-free bounded queue로 노출하는 application port 구현;
- producer security와 provider-private generation 생성/drain/close/attestation primitive 및 payload-free lifecycle fact;
- adapter-local best-effort publisher;
- structured outbox diagnostics.
금지:
- consumer listener;
- inbox repository;
- JPA entity;
- business route/event policy;
- sample contract.
초기 package shape:
dev.caskeleton.adapter.outbound.messaging
config/
contract/
destination/
envelope/
publication/
kafka/
lifecycle/
observation/
outbox/
package 이름은 예시이고 책임 분리가 정본이다.
9.5 미래 adapter:inbound:messaging-kafka
consumer 구현 전 registry migration으로 새 leaf를 추가한다.
module id: adapter-inbound-messaging-kafka
gradle path: :adapter:inbound:messaging-kafka
source path: src/adapter/inbound/messaging-kafka
allowed production dependencies:
- application-core
- domain-core
- shared-contract
정확한 allowed edge는 그 시점의 modules.json review로 확정한다. persistence/outbound messaging
adapter edge는 추가하지 않는다. app-bootstrap만 inbound listener, application use case,
transaction/inbox provider를 조립한다.
소유:
- Kafka listener container;
- record/header/envelope decode와 application command mapping;
- ack/seek/pause/resume/rebalance;
- consumer-local retry/DLT publisher. §5의 HARD invariant 3에 둔 유일한 예외이며 closed retry/DLT binding 외 destination과 application/outbox publication에는 사용할 수 없음;
- consumer lifecycle/metrics/security.
금지:
- repository 직접 호출;
- JPA entity;
- producer outbox implementation;
- business effect.
9.6 shared-contract
소유 가능:
- skeleton-wide generic envelope JSON Schema;
- framework-free bounded operational descriptor vocabulary;
- error/metric contract에서 truly shared인 값.
금지:
- WorkLog/Poster 등 business event schema;
- Kafka SDK;
- provider setting;
- feature-specific topic.
현재 shared-contract/CLAUDE.md에는 messaging schema resource가 아직 책임으로 등록되어 있지
않다. P1에서 공통 envelope schema를 추가하는 변경은 같은 commit scope에서 해당
CLAUDE.md의 Responsibility를 갱신하고 Java-stdlib-only 규칙과 business-free 검증을
추가해야 한다. 이 정책 갱신 없이 resource만 넣지 않는다.
9.7 app-bootstrap
소유:
- leaf가 bind/validate한 typed settings의 cross-leaf aggregation;
- exact capability tuple selection;
- cross-field/expected-state validation;
- provider, relay, future listener와 health composition;
- required capability readiness aggregation;
- secret reference resolution.
금지:
- event mapping;
- retry/DLT business policy;
- repository/Kafka implementation;
- schema compatibility rule 자체.
broker namespace의 typed settings/value validation은 현재 local SSOT와 같이
adapter:outbound:messaging이 소유한다. app-bootstrap은 raw Kafka map을 다시 bind하지 않고
leaf의 compiled descriptor를 transaction resource, persistence dispatch와 합성한다.
9.8 sample-portfolio
소유:
- sample domain event -> sample integration event mapping;
- sample payload schema/golden vectors;
- sample contract catalog contribution;
- sample consumer fixture가 생길 경우 application consume use case.
Production leaf는 sample module에 의존하지 않는다. 새 프로젝트는 sample contract를 제거하고 자기 feature contract를 같은 확장점에 등록한다.
합법적인 조립 경로는 다음으로 고정한다.
- framework-free
IntegrationEventContractContributionSPI는application-core에 둔다; - feature/sample module은 typed payload record, schema resource와 contribution bean을 제공한다;
- outbound messaging compiler는 application SPI의 bean 목록만 주입받으며 sample class를 import,
scan 또는
Class.forName하지 않는다; - production
app-bootstrap은 sample에 의존하지 않는다. base skeleton은 messagingDISABLED이고 empty catalog가 정상이다; - provider qualification test는 test-source fixture contribution을 사용한다;
- standalone sample이 ACTIVE example을 실행하는 phase에서만
sample-portfolio -> adapter-outbound-messagingruntime edge를modules.json과sample-portfolio/build.gradle에 함께 추가한다. 이 edge는 fixture consumer 방향이며 반대 edge는 금지한다.
Spring bean discovery는 조립 수단일 뿐 contract SSOT가 아니다. 동일 contribution 목록으로 build-time checksum manifest와 runtime compiler를 검증한다.
application SPI의 최소 shape는 다음처럼 closed type token을 포함한다.
interface IntegrationPayload {}
interface IntegrationEventContractContribution<P extends IntegrationPayload> {
ContractId contractId();
int payloadVersion();
Class<P> exactPayloadRecordType();
List<String> canonicalRecordComponentOrder();
SchemaResourceId payloadSchemaResource();
Sha256 payloadSchemaHash();
ContractDescriptor descriptor();
}
이는 API 이름을 고정하는 코드가 아니라 경계를 고정하는 pseudocode다. outbound compiler는
startup에 exactPayloadRecordType()이 final Java record이고 descriptor가 허용한 scalar,
collection, nested-record component만 갖는지 검증한다. runtime payload는 exact class equality로
closed catalog를 찾으며 assignable-type scan, Class.forName, default typing, feature-provided
Jackson serializer를 허용하지 않는다. contribution은 type token/order/schema/hash만 제공하고
JSON mapper, deterministic writer, parser와 schema validator는 계속 outbound messaging leaf가
소유한다. 따라서 feature mapper가 JSON string/tree를 만들거나 messaging leaf가 sample class를
compile-time import할 필요가 없다.
9.9 외부 deployment asset
다음은 application Java leaf가 아니라 deployment/integration-test asset이다.
- Kafka cluster/topic/ACL provisioning;
- Kafka Connect worker;
- Debezium connector JSON;
- PostgreSQL publication/replication-slot procedure;
- connector image/plugin digest;
- dashboard/alert/runbook.
10. Capability card와 exact selection
10.1 두 층의 card
Semantic card는 application이 요구하는 의미를 나타낸다.
| Card | 의미 |
|---|---|
messaging-best-effort-publish.v1 |
persistence/replay 보장 없는 bounded attempt |
messaging-outbox-publish.v1 |
transactional append 뒤 ACK 또는 explicit disposition까지 추적 |
messaging-inbox-consume.v1 |
future inbox-covered idempotent application effect |
messaging-cdc-dispatch.v1 |
future insert-only source log dispatch |
Provider/profile card는 그 의미를 실제로 제공하는 조합이다.
| Card | 초기 상태 |
|---|---|
external-kafka-sender-legacy.v1 |
R0, R2 selection 금지 |
kafka-spring-acknowledged-idempotent.v1 |
최초 R2 목표 |
postgresql-polling-outbox.v2 |
최초 R2 목표 |
json-schema-envelope.v1 |
최초 R2 목표 |
external-topic-validated.v1 |
최초 R2 목표 |
postgresql-per-record-jit-claim.v1 |
최초 R2 목표 |
kafka-sasl-ssl-scram-sha-512.v1 |
최초 production R2 목표 |
kafka-compression-none.v1 |
최초 R2 목표 |
per-key-normal-path-sequence-detectable.v1 |
최초 R2 목표 |
same-postgresql-transaction-resource.v1 |
최초 R2 목표 |
authenticated-internal-web-disposition.v1 |
최초 R2 목표 |
미래 consumer/CDC/EOS/Avro/Protobuf semantic/provider 이름은 §31.5의 설계 extension ledger일 뿐 machine registry row가 아니다. code, exact settings, test와 evidence가 생기는 변경에서만 machine registry에 추가한다.
10.2 Card 필드
각 executable card는 최소 다음을 가진다.
cardId
cardVersion
semanticContractIds
providerId
providerVersion
maturity
guarantees
explicitNonGuarantees
outcomeTaxonomyVersion
orderingProfile
resourceBounds
automaticPublicationAge
sameEventRequeueHorizon
securityProfile
topologyProfile
lifecycleProfile
operatorControlProfile
schemaSetHash
settingsDigest
evidenceFingerprint
evidenceTasks
requiredScenarios
runbookIds
owner
maturity는 정확히 다음 하나다.
not-implemented
implemented-candidate
release-eligible
별도 boolean releaseEligible이나 중복 maturity/readiness 필드를 두지 않는다. R0–R3는
evidence 설명용 level이고 machine selection state를 대신하지 않는다.
10.3 Selection 규칙
required semantic contract/version 일치
AND required guarantees ⊆ provider achieved guarantees
AND outcome/failure policy compatible
AND exact dispatch/serialization/security/topic/operator-control profiles compatible
AND automatic publication/requeue/dedupe horizons compatible
AND profile.maturity = release-eligible
AND current evidence fingerprint = PASS
하나라도 불충족하면 production ACTIVE startup 또는 release gate를 실패시킨다.
10.4 First R2 selected tuple
messaging-outbox-publish.v1
+ kafka-spring-acknowledged-idempotent.v1
+ postgresql-polling-outbox.v2
+ postgresql-per-record-jit-claim.v1
+ json-schema-envelope.v1
+ external-topic-validated.v1
+ kafka-sasl-ssl-scram-sha-512.v1
+ kafka-compression-none.v1
+ per-key-normal-path-sequence-detectable.v1
+ same-postgresql-transaction-resource.v1
+ authenticated-internal-web-disposition.v1
local/development는 별도 local-plaintext.v1 evidence를 가질 수 있지만 production tuple에
승격되지 않는다.
10.5 Configuration과 registry의 역할
src/config/messaging/readiness-cards.yaml: 구현된 card의 maturity와 base scenario;src/config/messaging/profile-compatibility.yaml: wildcard 없는 exact tuple과 interaction scenario;src/config/messaging/release-profile-assertions.yaml: 실제 release configuration digest와 expected selected profile assertion;- deployment binding: 이 deployment가 어떤 exact card와 logical destination을 요구하는지 선언;
- compiled descriptor: 두 입력을 합성한 실제 runtime truth;
- §0 status ledger: 구현 phase 진행의 human truth.
설정에 provider=avro 같은 값을 추가하는 것만으로 optional card가 생기지 않는다. registry에
없는 값은 unknown configuration으로 startup 실패다.
candidate 승격 deadlock을 피하기 위해 isolated test harness에만 QUALIFICATION_ONLY를 둔다.
production과 같은 binder/resolver/resource composition을 사용하되
implemented-candidate를 허용하고, ACTIVE_READY, release assertion 또는 production
descriptor는 절대 만들지 않는다. machine registry 파일은 P1–P4 구현과 함께 생성하며, 현재
설계 extension 이름만 미리 row로 만들지 않는다.
11. Identity, ordering과 vocabulary
11.1 Event ID
eventId는 integration event의 canonical identity다.
- globally unique하고 immutable하다;
- first card의 wire/storage grammar는 1–96자의 canonical US-ASCII
[A-Za-z0-9][A-Za-z0-9._:-]*이며 DB에는VARCHAR(96)+ CHECK로 저장한다; - outbox retry, producer restart, polling/CDC mode가 바뀌어도 동일하다;
- Kafka producer attempt ID나 database row ID와 다르다;
- consumer inbox dedupe의 기본 identity다;
- payload와 함께 생성된 뒤 다시 계산하지 않는다.
현재 idempotencyKey는 권장값이 eventId이고 별도 의미가 불명확하다. 목표 contract에서는
consumer dedupe는 eventId 하나를 사용한다. 원본 command의 idempotency identity가 필요하면
sourceOperationId처럼 의미가 다른 이름으로 보존하며 consumer dedupe key로 자동 대체하지 않는다.
같은 eventId와 다른 exact envelope document hash가 관찰되면 정상 duplicate가 아니라 identity collision 또는
contract violation이다. consumer는 이를 DUPLICATE로 ACK하지 않고 quarantine/operator path로
보낸다.
11.2 Contract ID와 event name
contractId는 version과 분리된 안정적인 semantic name이다.
예:
portfolio.worklog.reserved
portfolio.poster.published
규칙:
- closed code/manifest catalog에 등록한다;
- user/tenant/request 입력으로 동적 생성하지 않는다;
- Java class name과 자동 결합하지 않는다;
- physical topic을 포함하지 않는다;
- metric tag로 사용할 때 catalog cardinality budget을 통과해야 한다.
eventType legacy field는 migration 동안 contract ID alias로 읽을 수 있지만 새 event에는
contractId를 사용한다.
11.3 Envelope version과 payload version
두 version을 분리한다.
envelopeVersion = messaging 공통 metadata shape version
payloadVersion = contractId별 business payload schema version
단일 schemaVersion으로 두 의미를 합치지 않는다.
- envelope version 변경은 모든 producer/consumer/CDC mapping에 영향을 준다;
- payload version 변경은 특정 contract에만 영향을 준다;
- schema file은 version별 immutable하다;
- 같은 version file의 checksum 변경은 CI 실패다.
11.4 Logical destination과 physical topic
logicalDestinationId는 application/contract가 요구하는 delivery class를 나타낸다.
physicalTopic은 deployment binding이다.
contractId
-> logicalDestinationId
-> environment-specific physicalTopic
logical destination은 retention class, ordering class, maximum record size, sensitivity, replay horizon과 같은 semantic/operational intent를 묶는다. topic 이름, cluster bootstrap server, ACL principal은 포함하지 않는다.
11.5 Aggregate identity와 total order
ordering을 요구하는 contract는 다음을 가진다.
aggregateType
aggregateId
aggregateOrder = (aggregateSequence, eventIndex)
aggregateSequence는 domain aggregate version 또는 같은 transaction에서 allocation한 monotonically increasing sequence다;- 하나의 aggregate version에서 여러 integration event가 나오면
eventIndex로 total order를 완성한다; - 더 단순한 구현이 event마다 고유 단조 sequence를 할당하면
eventIndex=0으로 고정할 수 있다; (tenant?, logicalDestinationId, aggregateType, aggregateId, aggregateSequence, eventIndex)는 unique constraint로 보호한다;- sequence를 제공할 수 없는 event는 strict aggregate ordering card를 선택할 수 없다.
timestamp와 random event ID는 strict total order의 대체물이 아니다.
tenant scope는 nullable uniqueness에 맡기지 않는다.
- tenant mode ACTIVE: canonical
tenant_scope는NOT NULL이고 unique key에 포함한다; - tenant mode DISABLED: canonical non-null system scope를 저장하거나 tenant column을 제외한 별도 constraint를 사용한다;
- 일반 PostgreSQL
UNIQUE의 NULL-distinct 동작에 dedupe/order correctness를 의존하지 않는다; NULLS NOT DISTINCT를 선택하면 adopted PostgreSQL version과 migration test에 명시한다.
11.6 Partition key
partition key는 catalog가 정한 deterministic mapping이다.
기본 ordered event:
partitionKeyText =
lowerHex(
SHA-256(
UTF8("ca-skeleton.messaging.partition-key.v1") || 0x00
|| u32be(len(UTF8(tenantScope))) || UTF8(tenantScope)
|| u32be(len(UTF8(logicalDestinationId))) || UTF8(logicalDestinationId)
|| u32be(len(UTF8(aggregateType))) || UTF8(aggregateType)
|| u32be(len(UTF8(aggregateId))) || UTF8(aggregateId)
)
)
partitionKeyBytes = US_ASCII(partitionKeyText)
규칙:
tenantScope는 §11.5의 canonical non-null scope다;u32be는 뒤따르는 UTF-8 byte length의 unsigned 32-bit big-endian 표현이다;- 결과는 정확히 64자의 lowercase hexadecimal text이고 DB에는
VARCHAR(64) NOT NULL+ lowercase-hex CHECK로 저장한다; - 같은 ordering scope는 동일 text/bytes를 만든다;
- raw PII/tenant/user ID를 metric/log에 노출하지 않는다;
- null/blank key는 ordering-required contract에서 startup/runtime rejection이다;
- polling producer는 저장된 text의 US-ASCII bytes를
ByteArraySerializer로 보내고, CDC는 같은 PostgreSQLVARCHAR를 Kafka ConnectStringConverter로 보내 같은 bytes를 만든다; - producer와 CDC가 domain-separated length-prefix golden vector를 공유한다;
- custom partitioner가 key를 무시하면 해당 ordering card는 invalid다.
11.7 Attempt, claim과 generation
다음 identity는 event ID와 다르다.
| Identity | 용도 |
|---|---|
claimToken |
polling row의 현재 owner를 fence하는 opaque token |
deliveryGeneration |
operator replay/requeue가 만든 새 delivery lifecycle |
publicationAttemptId |
한 application-level send attempt 진단 |
producerGeneration |
credential/settings rotation으로 생성된 producer runtime |
consumerId |
inbox effect identity |
replayOperationId |
audited replay 요청 |
attempt/generation을 consumer dedupe event ID로 사용하지 않는다.
11.8 Consumer identity
future inbox의 consumerId는 최소 다음을 compile한다.
logical subscription
+ handler name
+ effect contract version
+ tenant dimension when storage is tenant-isolated
Kafka group ID가 배포 편의 때문에 바뀌어도 의도하지 않은 business effect 재적용이 일어나지
않도록 logical identity를 명시한다. group ID를 consumer identity에 포함해야 하는 deployment는
그 관계를 descriptor에 고정한다. intentional reprocessing은 새 replayGeneration과 승인을
요구한다.
11.9 Clock authority
occurredAt: event가 일어난 application/domain wall-clock fact;createdAt: database insert time;- claim lease, retry due, retention cutoff: database time authority;
- producer deadline: monotonic process clock;
- broker record timestamp: event timestamp policy 또는 broker append time descriptor.
여러 pod의 wall-clock으로 claim lease를 판정하지 않는다. database time을 사용하지 못하면 허용 clock-skew bound와 failure policy를 card에 포함한다.
12. Integration event pipeline
12.1 Domain event와 integration event
Domain event를 그대로 JSON으로 직렬화하지 않는다.
DomainEvent
-> feature application mapper
-> IntegrationEventDraft
-> bounded local contract compiler/encoder
-> ValidatedIntegrationEvent
-> immutable outbox_event
-> WireEnvelope v1
feature application mapper가 소유하는 것은 외부에 공개할 semantic field 선택이다. encoder가 소유하는 것은 UTF-8 JSON encoding, schema validation, byte bound와 checksum이다. mapper에 retry/topic/security 정책을 넣지 않고 encoder에 business rule을 넣지 않는다.
12.2 Draft
개념적인 draft shape는 다음과 같다.
record IntegrationEventDraft<P extends IntegrationPayload>(
EventId eventId,
ContractId contractId,
int payloadVersion,
LogicalDestinationId destinationId,
AggregateIdentity aggregate,
AggregateOrder order,
Instant occurredAt,
CorrelationId correlationId,
Optional<CausationId> causationId,
Optional<TenantId> tenantId,
P featurePayload) {}
이는 구현 이름을 강제하는 Java API가 아니라 ownership을 보여주는 pseudocode다.
featurePayload는 §9.8 contribution의 exact type token으로 등록된 typed immutable Java
record다. JSON tree, Jackson node, raw map/string, Kafka record가 application contract가 되지
않는다.
12.3 Local contract compiler/encoder
Application은 framework-free port를 통해 deterministic local encoder를 사용할 수 있다. 실제 JSON/schema library는 outbound messaging adapter가 소유한다.
encoder는:
- startup에 schema/catalog를 precompile한다;
- runtime remote schema fetch를 하지 않는다;
- bounded CPU/memory 안에서 typed payload를 JSON으로 encode한다;
- envelope/payload schema, duplicate key, depth와 exact UTF-8 bytes를 검증한다;
- immutable serialized document와 schema/catalog digest를 반환한다.
first encoder는 같은 logical event가 같은 exact UTF-8 document를 만들도록 deterministic field order와 scalar rendering을 고정한다. 저장·재발행·CDC의 authority는 이 exact byte document이며 JSONB 재직렬화 결과가 아니다.
business transaction 안에서 호출될 경우 local computation만 수행하고 network, broker, filesystem, secret refresh를 하지 않는다. encoding 비용이 transaction budget을 넘는 event는 transaction 전에 immutable input을 준비하거나 별도 staged workflow를 사용한다.
12.4 Transaction sequence
durable application command의 기본 순서는 다음이다.
1. command/idempotency/authorization validation
2. tx.inWrite begin
3. domain aggregate load + invariant check + mutation
4. domain event -> integration-event draft mapping
5. precompiled local encoder validation
6. business state save
7. outbox_event INSERT
8. polling mode이면 outbox_delivery INSERT
9. commit
4–8 중 하나라도 실패하면 business write도 rollback한다. broker send는 이 transaction 안에서 수행하지 않는다.
same-store는 이름뿐인 가정이 아니다. compiled card는 transactionResourceId를 갖고 business
repository, TransactionPort, OutboxAppendPort, outbox migration이 같은 resolved
DataSource/EntityManagerFactory/PlatformTransactionManager resource에 bind됐는지 startup에
검증한다. multi-datasource deployment는 contract별 resource binding을 명시한다. 다른 resource면
ACTIVE를 거부한다. real rollback test가 이 identity assertion을 보완한다.
dispatch mode 판단은 feature mapper가 하지 않는다. persistence append adapter가 같은
transaction에서 §27의 active publication epoch를 읽고 event에 epoch/authority를 기록한 뒤,
POLLING_V2일 때만 delivery row를 함께 만든다.
12.5 Validated event와 stored event
ValidatedIntegrationEvent는 최소 다음을 가진다.
all stable identities
envelopeVersion
payloadVersion
logicalDestinationId
partitionKeyText and its exact US-ASCII bytes
validated envelope JSON bytes/document
contentType
schemaSetHash
envelopeSha256
envelopeSchemaHash
payloadSchemaHash
contractCatalogRevision
destinationBindingRevision
validated traceparent/tracestate allowlist
publicationEpoch, dispatchAuthority, transactionResourceId, DB-authoritative createdAt은
encoder 결과가 아니다. OutboxAppendPort의 persistence 구현이 caller의 write transaction
안에서 ACTIVE epoch를 읽고 same-store resource identity를 확인한 뒤 이 네 값을 더해
StoredOutboxEvent를 구성한다. 따라서 transaction 전에 만들어 둔 validated bytes가 stale
application setting의 authority를 내장하거나 application clock을 DB creation time으로 가장하지
않는다.
retry 때 payload를 다시 business object에서 직렬화하지 않는다. polling attempt는 저장된 같은 identity와 validated document를 사용한다.
envelopeSha256은 다음 exact input으로 계산한다.
SHA-256(
UTF8("ca-skeleton.messaging.envelope.v1") || 0x00
|| u32be(len(exactEnvelopeBytes))
|| exactEnvelopeBytes
)
여기서 u32be는 §11.6과 같은 unsigned 32-bit big-endian byte length다.
이는 integrity/collision diagnosis용이지 confidentiality control이 아니다. DB/API/log/metric에
노출하지 않고 payload와 같은 access control/retention을 적용한다. 같은 event ID에서 다른
envelope hash는 duplicate가 아니라 collision/quarantine이다. semantic JSON을 JSONB로 round-trip한
뒤 다시 hash하지 않는다.
12.6 Polling/CDC wire parity
polling과 CDC는 같은 logical WireEnvelope v1을 emit한다.
- field와 semantic value가 같아야 한다;
- event ID, contract ID, versions, key가 같아야 한다;
- JSON object member byte ordering 차이를 허용할지 card가 명시한다;
- first baseline은 polling retry에서 exact stored UTF-8 bytes 재사용을 요구한다;
- CDC는 golden semantic equality와 consumer decode equality를 통과한다;
- “같은 contract”를 단순히 비슷한 JSON이라고 표현하지 않는다.
13. Contract catalog, destination binding과 topic
13.1 두 catalog
Contract catalog는 code/repository artifact다.
contractId
payload versions
owner module
logical destination
payload schema resource/hash
serializer id
ordering requirement
partition-key policy
maximum payload/envelope bytes
sensitivity classification
supported producer/consumer version matrix
same-event requeue horizon
Destination binding은 deployment configuration다.
logical destination
Kafka cluster binding
physical topic
expected partitions
minimum replication factor
minimum in-sync replicas
cleanup policy
retention expectation
maximum record bytes
security profile
required readiness
contract가 infrastructure topology를 소유하지 않고 configuration이 business schema를 재정의하지 않는다.
13.2 Compile
startup compiler는 다음을 합성한다.
contract descriptor
+ destination descriptor
+ producer provider descriptor
+ serialization descriptor
+ security descriptor
+ evidence card
= CompiledPublicationBinding
검증:
- contract/destination ID unique;
- 모든 active contract에 정확히 한 destination binding;
- unknown destination/topic 금지;
- ordering-required contract에 nonblank stable key;
- contract maximum bytes <= destination/provider/topic bounds;
- schema/catalog hash가 evidence와 일치;
- production profile과 security profile 호환;
- dispatch mode와 provider 요구 일치;
- required binding은 release-eligible evidence 보유.
13.3 Configuration override 제한
설정은 code contract를 약화하지 못한다.
- code maximum record bytes보다 크게 override할 수 없다;
- ordering-required를
NONE으로 낮출 수 없다; - schema validation을 끌 수 없다;
- production TLS 요구를 plaintext로 바꿀 수 없다;
- required destination을 optional로 바꿀 수 없다;
- unknown compatibility mode를 선택할 수 없다.
더 엄격한 deployment bound는 허용한다.
13.4 Topic naming
physical topic은 operator-owned static value다.
- request/event/tenant 값을 문자열 보간하지 않는다;
- environment prefix/suffix는 binding compiler가 allowlist pattern으로 검증한다;
- producer principal은 production에서 Create/Delete/Alter 권한을 갖지 않는다;
- auto-create를 끈다;
- topic rename은 새 binding/revision과 migration runbook을 요구한다.
13.5 Topic topology attestation
ACTIVE startup 또는 pre-deploy gate는 최소 다음을 확인한다.
- topic 존재;
- expected partition count;
- replication factor가 minimum 이상;
min.insync.replicas가 policy minimum 이상;- cleanup policy;
- retention/replay horizon;
- topic maximum message bytes;
- unclean leader election 관련 cluster/topic policy가 deployment 요구와 호환;
- producer principal의 최소 Describe/Write 동작;
- consumer/DLT profile이 있을 때 대응 Read/Write ACL.
acks=all만 확인하고 replication/min ISR를 보지 않은 상태를 durable topology로 표시하지 않는다.
first tuple은 verification source를 항목별로 고정한다.
| 항목 | Runtime source | Release/provisioning source |
|---|---|---|
| cluster identity, topic existence, partition/leader/ISR/RF | producer principal의 bounded AdminClient Describe |
IaC expected resource identity |
| topic cleanup/retention/max bytes/min ISR/topic override | exact topic 범위 read-only DescribeConfigs |
IaC rendered config/digest |
| broker-level unclean election/default/max bounds/auto-create | runtime에서 과도한 cluster config 권한을 요구하지 않음 | signed/provenance-attested broker policy |
| exact-topic Write와 denied Create/Alter/Delete/other-topic Write | startup에 임의 canary를 만들지 않음 | security release lane의 positive/negative probe |
| ACL/quota owner와 rollback | runtime ACL enumeration 금지 | IaC/security evidence |
release/provisioning evidence는 environment/cluster alias, topic resource identity, rendered config/ACL policy digest, issuer/provenance, generated-at, expires-at와 release assertion digest를 가진다. missing, stale, wrong-cluster, signature/provenance failure 또는 runtime-observed 값과의 mismatch는 production ACTIVE를 fail-closed한다. runtime에서 확인할 수 없는 값을 “검증됨”으로 표시하지 않고 descriptor에 source와 freshness를 함께 노출한다.
13.6 Partition expansion
Kafka default key partitioning에서 partition 수가 바뀌면 같은 key가 다른 partition으로 이동할 수 있다. rolling producer/consumer 기간에는 old/new partition의 event order가 섞일 수 있다.
ordering-required topic은 in-place partition expansion을 일반적인 무중단 변경으로 취급하지 않는다. 기본 절차는 새 topic/binding generation, write cutover watermark, consumer dual-read 또는 drain, order reconciliation과 rollback이다.
event append 시 destinationBindingRevision을 immutable capture한다. retry는 같은 revision을
resolve하며 current config의 새 topic으로 조용히 reroute하지 않는다. backlog를 새 binding으로
옮기려면 audited delivery generation/explicit migration을 사용한다.
13.7 Compaction
first baseline topic은 delete-retention event log다. compaction은 다음이 모두 정의된 contract만 별도 card로 선택한다.
- key가 entity state identity인지;
- tombstone 의미;
- intermediate event 손실 허용 여부;
- consumer bootstrap 의미;
- minimum compaction lag;
- delete retention;
- replay/ordering 영향.
integration event에 compaction을 기본 적용하지 않는다.
14. JSON envelope v1과 schema evolution
14.1 Dialect와 resource ownership
first baseline은 JSON Schema Draft 2020-12를 사용한다.
공통 envelope schema 예시 위치:
src/shared-contract/src/main/resources/contracts/messaging/envelope/v1.schema.json
sample payload schema 예시 위치:
src/sample-portfolio/src/main/resources/contracts/messaging/
portfolio.worklog.reserved/v1.schema.json
실제 production project는 feature owner module에 payload schema를 둔다. 각 schema는:
- explicit
$schema; - immutable absolute
$id; - contract/payload version;
- checked-in checksum manifest;
- local prebundled
$refallowlist; - owner와 compatibility vectors를 가진다.
runtime HTTP/file remote $ref resolution은 금지한다.
14.2 Envelope shape
개념적인 envelope v1:
{
"envelopeVersion": 1,
"eventId": "019...",
"contractId": "portfolio.worklog.reserved",
"payloadVersion": 1,
"logicalDestination": "portfolio-domain-events",
"aggregate": {
"type": "worklog",
"id": "worklog-42",
"sequence": 17,
"eventIndex": 0
},
"occurredAt": "2026-07-28T05:10:30.123Z",
"correlationId": "corr-...",
"causationId": "cause-...",
"contentType": "application/json",
"payload": {
"workLogId": "worklog-42"
}
}
tenant가 실제로 활성인 deployment만 bounded tenantId를 포함한다. causation ID가 없을 때
null로 넣을지 field를 생략할지는 envelope schema가 하나로 고정한다.
envelope에는 다음을 넣지 않는다.
- physical topic/cluster/bootstrap server;
- delivery status/attempt/backoff/claim token;
- Kafka offset/partition;
- credential/security profile;
- Java class name;
- raw exception;
- arbitrary baggage;
- mutable consumer state.
14.3 Envelope/header ownership
authoritative semantic metadata는 envelope다. Kafka header는 transport 기능에 필요한 bounded allowlist만 사용한다.
초기 header allowlist 후보:
id
contract-id
payload-version
traceparent
tracestate
first header의 id는 Debezium Outbox Event Router와 polling producer가 공유하는 event ID다.
envelope와 header에 중복된 identity가 다르면 producer와 consumer 모두 reject한다. header 이름,
개수, key bytes, total bytes를 제한한다. arbitrary inbound header forwarding은 금지한다.
14.4 Strict versioning
first baseline은 STRICT_VERSIONED 정책을 사용한다.
- 같은 schema version file을 수정하지 않는다;
- optional field 추가도 새 payload version을 만든다;
- consumer가 새 version을 지원한 뒤 producer를 배포한다;
- rolling overlap 동안 consumer는 최소 명시된 N/N-1 version allowlist를 가진다;
- producer가 지원 종료된 version을 emit하지 않는다는 release assertion을 둔다;
- unsupported future version은 일반 retry 대상이 아니다.
JSON Schema diff heuristic만으로 backward/full compatibility를 주장하지 않는다. 실제 old/new reader/writer golden vectors가 compatibility evidence다. Kafka/DLT/archive/inbox replay horizon 안에 남아 있는 모든 payload version은 reader support를 유지한다. version retirement는 해당 version의 source/DLT/archive가 더 이상 replay 가능하지 않거나 versioned upcaster가 qualification됐다는 purge proof가 있어야 한다. N/N-1은 replay horizon을 대신하지 않는다.
14.5 Object와 unknown-field policy
- envelope v1은
unevaluatedProperties: false로 닫는다; - payload schema도 first baseline에서 explicit property set을 사용한다;
- additive evolution은 in-place field 추가가 아니라 payload version 증가로 처리한다;
- tolerant-reader card를 나중에 추가할 수 있지만 그때 unknown-field behavior와 rolling vectors를 별도 증명한다.
14.6 Scalar/collection policy
각 schema가 표현할 수 있는 범위는 최소 다음을 명시한다.
required;- null과 missing의 차이;
- string
minLength/maxLength와 Unicode normalization policy; - array maximum items;
- integer/number semantic range;
- enum evolution;
- timestamp string format;
- object property count;
JSON Schema format은 implementation에 따라 annotation일 수 있다. runtime validator에서 format
assertion을 켰다는 evidence를 만들거나 timestamp/UUID 등을 explicit parser로 검증한다.
Draft 2020-12가 직접 표현하지 않는 exact UTF-8 byte 수, nesting depth, numeric precision,
exponential notation/canonical lexical form, parser time/memory는
json-codec-admission-v1의 별도 규칙이다. maxLength를 byte limit으로 오해하거나 custom
keyword 없이 schema가 이 제한을 증명한다고 쓰지 않는다.
14.7 Parser hardening
codec은 다음을 거부한다.
- malformed UTF-8;
- duplicate object member names;
- unpaired surrogate;
- excessive nesting;
- maximum을 넘는 string/array/object;
- resource budget을 넘는 arbitrary-precision number;
- trailing garbage;
- non-finite number;
- remote reference;
- polymorphic Java type metadata.
validation CPU/memory/time budget을 test한다. record size는 Java character 수가 아니라 최종 UTF-8 key + value + header bytes로 계산한다.
schema compiler는 exact offline resource registry만 사용한다.
- duplicate
$id/unknown vocabulary/unknown dialect 거부; - remote URI가 local allowlist resource로 정확히 resolve되지 않으면 거부;
- cyclic/recursive
$ref는 명시적 depth/resource budget 안에서 지원하거나 compile-time 거부; - pathological regular expression/validator recursion adversarial corpus;
- meta-schema와 vocabulary도 pinned local checksum 대상;
- startup precompile 뒤 runtime schema fetch 0.
14.8 Envelope document hash
§12.5의 exact envelope byte hash는:
- same event ID document collision 탐지;
- polling retry exact-document 확인;
- CDC shadow byte parity;
- audit/diagnosis
에 사용한다. algorithm/input은 §12.5 하나만 정본으로 사용한다. hash를 metric tag로 쓰지 않고 restricted storage 밖에 노출하지 않는다.
14.9 Schema registry future card
Avro/Protobuf/JSON Schema registry card가 추가되면 다음을 별도로 설계·검증한다.
- subject naming;
- compatibility mode;
- registry auth/TLS/readiness;
- schema ID cache와 outage behavior;
- generated code ownership;
- rolling compatibility;
- registry unavailable 시 write policy;
- schema deletion/retention;
- cross-cluster replication.
first JSON Schema card에는 registry 설정 key나 placeholder를 추가하지 않는다.
15. Application contract와 publication outcome
15.1 Durable append port
OutboxAppendPort는 provider-neutral same-store append 의미를 유지한다. target input은
legacy raw String payload가 아니라 validated integration event다.
interface OutboxAppendPort {
void append(ValidatedIntegrationEvent event);
}
실제 이름은 implementation plan에서 정하지만 다음은 변하지 않는다.
- application/core 타입;
- same transaction requirement;
- provider/physical topic 없음;
- immutable event identity;
- validation/catalog evidence 포함.
15.2 Publish port
target publish port는 expected technical outcome을 exception 하나로 뭉치지 않는다.
sealed interface PublicationOutcome {
record Acknowledged(PublicationReceipt receipt) implements PublicationOutcome {}
record AcknowledgedMismatch(PublicationReceipt receipt) implements PublicationOutcome {}
record Rejected(PublicationFailure failure) implements PublicationOutcome {}
record Indeterminate(PublicationFailure failure) implements PublicationOutcome {}
}
PublicationReceipt는 Kafka SDK 타입 대신 다음 같은 bounded provider-neutral reference를
가진다.
providerId
logicalDestinationId
providerGeneration
ackObservedAt
opaque bounded providerRecordReference
physical topic/partition/offset가 persistence audit에 필요하면 adapter가 safe bounded string으로
만들며 application이 이를 routing에 사용하지 않는다. ackObservedAt은 broker clock이 아니라
local future-completion 관찰 시각이다.
Late completion observation boundary
동기 publish 결과가 INDETERMINATE로 반환된 뒤 Kafka future가 완료될 수 있으므로 다음
provider-neutral application ports를 둔다.
interface LatePublicationObservationSourcePort {
List<LatePublicationObservation> pollBounded(int maximum);
void acknowledgePersisted(ObservationId id);
void releaseForRetry(ObservationId id);
}
interface OutboxAttemptObservationPort {
void appendLateObservations(List<LatePublicationObservation> observations);
}
- outbound messaging adapter가 bounded payload-free queue로 source port를 구현한다;
- persistence adapter가 append port를 구현한다;
- application의
RecordLatePublicationObservationsUseCase는 다음 순서를 고정한다:poll/lease bounded batch -> tx.inNew(idempotent append) 정상 반환 -> acknowledgePersisted.tx.inNew(...)의 정상 반환은 commit 완료를 뜻하며 source ACK를 transaction callback 안에서 호출하지 않는다; app-bootstrap은 provider ACTIVE일 때만 bounded drain scheduler를 조립한다;- Kafka callback thread는 JPA/repository/transaction을 직접 호출하지 않는다;
- observation identity는
(eventId, deliveryGeneration, publicationAttemptId, LATE_ACK_OBSERVED)이고 DB unique/ON CONFLICT로 duplicate drain을 흡수한다; - callback/timeout은 adapter-local atomic terminal marker로 단 한 synchronous outcome을 결정한다. deadline marker가 먼저 이기고 ACK가 나중에 오면 queue에 late observation 하나만 제안한다;
- DB append 또는 commit 실패는
releaseForRetry로 item을 bounded retry에 되돌리고 delivery state를 바꾸지 않는다; - DB commit 뒤 source ACK 전 process crash는 같은 observation의 duplicate drain을 만들 수
있으며 DB unique/
ON CONFLICT가 이를 흡수한다.
attempt admission과 deadline 시점의 INDETERMINATE outcome journal은 authoritative하고 반드시
Tx B/Tx C에서 영속화한다. 반면 process crash나 queue overflow로 late callback diagnostic 자체를
잃을 수 있으므로 LATE_ACK_OBSERVED capture를 delivery correctness 근거로 사용하지 않는다.
queue overflow/drop은 bounded metric, readiness degradation과 alert 대상이며 capacity
qualification에서는 0이어야 한다. 이 경계 덕분에 messaging→persistence project dependency를
추가하지 않는다.
15.3 Failure stage
최소 stage:
CONTRACT_COMPILE
SERIALIZATION
LOCAL_ADMISSION
METADATA
SEND
BROKER_ACK
DEADLINE
SHUTDOWN
PROVIDER
최소 failure class:
INVALID_CONTRACT
INVALID_PAYLOAD
RECORD_TOO_LARGE
DESTINATION_MISSING
UNAUTHORIZED
AUTHENTICATION_FAILED
TLS_FAILED
TOPIC_POLICY_MISMATCH
BUFFER_EXHAUSTED
BROKER_UNAVAILABLE
THROTTLED
DEADLINE_EXCEEDED
CLIENT_CLOSED
UNKNOWN_PROVIDER_FAILURE
exception class name이나 message를 stable application contract로 사용하지 않는다.
15.4 Certainty와 retry disposition
acceptance certainty와 retryability는 독립 축이다.
acceptanceCertainty = NOT_ACCEPTED | ACCEPTED | INDETERMINATE | ACCEPTED_MISMATCH
retryDisposition = NO_RETRY | RETRY_WITHIN_BUDGET | STOP_PROVIDER | OPERATOR_HOLD
REJECTED는 provider가 broker acceptance가 없음을 확정할 수 있을 때만 사용한다.
예:
- schema/size/local catalog rejection;
- local admission 전 rejection;
- definitive broker authorization rejection;
- startup AdminClient attestation이 send admission 전에 확정한 missing destination.
INDETERMINATE 예:
- send 뒤 deadline;
- ACK response loss;
- connection break after request write;
- callback/cancel race;
- shutdown 중 unresolved in-flight;
NotEnoughReplicasAfterAppendException처럼 append 뒤 실패할 수 있는 broker 응답;- post-admission
UnknownTopicOrPartitionException, retriable/unknown producer exception; - provider가 acceptance를 증명할 수 없는 unknown exception.
Kafka RetriableException이라는 사실은 미수락 증거가 아니다. 분류가 애매하면
INDETERMINATE가 안전한 기본이다. retry 여부는 certainty를 바꾸지 않고 remaining
attempt/elapsed budget과 producer health로 결정한다.
15.5 Relay decision
| Publication outcome | Polling action |
|---|---|
| ACKNOWLEDGED | claim-token/valid-lease CAS로 DELIVERY_RECORDED 기록 |
| ACKNOWLEDGED_MISMATCH | relay scope HOLD, producer readiness DOWN, misroute incident |
| definite transient REJECTED | retry budget이 남으면 RETRY_WAIT |
| definite permanent REJECTED | 즉시 EXHAUSTED/operator disposition |
| INDETERMINATE | duplicate 가능성을 기록하고 bounded retry/reconciliation path |
| programming invariant failure | cycle 실패 + readiness/alert; 일반 transient로 숨기지 않음 |
report는 persisted transition이 성공한 뒤에만 emit한다. report adapter failure는 authoritative state를 바꾸지 않는다.
15.6 Best-effort와 durable port
두 contract는 계속 분리한다.
- best-effort: persistence/replay 없음, bounded attempt 뒤 failure를 삼킬 수 있음;
- durable: transactionally stored event, polling/CDC, explicit terminal disposition.
best-effort가 내부적으로 같은 ACK-aware producer를 사용해도 durable로 승격되지 않는다. durable append를 자동 수행하지도 않는다.
16. Spring Kafka producer protocol
16.1 Runtime ownership
adapter:outbound:messaging이 다음을 직접 만든다.
DefaultKafkaProducerFactory<byte[], byte[]>;KafkaTemplate<byte[], byte[]>;- bounded AdminClient/topology attestor;
- provider generation의 생성/drain/close/attestation primitive owner;
- observation convention;
- resolved credential/certificate material을 한 provider generation에 적용하는 owner.
adapter:outbound:messaging은 durable delivery state나 HOLD 정책을 결정하지 않는다.
application-core의 rotation use case가 provider-neutral lifecycle fact를 받아
INDETERMINATE/HOLD persistence와 generation barrier/admission 재개를 조정하고,
app-bootstrap이 secret refresh와 그 use case invocation을 compose한다.
classpath presence나 generic spring.kafka.bootstrap-servers=localhost:9092 default로 활성화하지
않는다. canonical messaging binding이 ACTIVE일 때만 만든다.
16.2 Adapter-private gateway
provider-private SPI는 ACK를 표현해야 한다.
interface KafkaPublishGateway {
KafkaAttemptOutcome publish(
CompiledKafkaRecord record,
MonotonicDeadline deadline);
}
이 SPI는 outbound adapter 내부 또는 package-private다. KafkaTemplate, SendResult,
RecordMetadata를 application/shared에 노출하지 않는다.
16.3 Send sequence
1. compiled binding lookup
2. immutable envelope/key/header byte verification
3. local admission acquire
4. ProducerRecord<byte[], byte[]> construction
5. KafkaTemplate.send
6. send future를 monotonic deadline까지 await
7. RecordMetadata와 expected destination 검증
8. ACKNOWLEDGED / ACKNOWLEDGED_MISMATCH / REJECTED / INDETERMINATE map
9. admission/resource release
serialization은 prevalidated bytes를 사용하는 Kafka ByteArraySerializer 계열로 단순화한다.
Kafka serializer callback 안에서 business JSON serialization이나 remote schema lookup을 하지
않는다.
16.4 ACK condition
broker ACK 관찰은 다음으로 정의한다.
future completed successfully
AND metadata is present
metadata topic이 compiled topic과 같으면 ACKNOWLEDGED, 다르면
ACKNOWLEDGED_MISMATCH다. deadline 뒤 future가 성공해도 ACK 관찰 사실은 append-only attempt
journal drain이 성공한 경우에만 LATE_ACK_OBSERVED로 남으며, 이미 정한 application
outcome/delivery state를 뒤집지 않는다. drain 전 crash/overflow로 진단 관찰을 잃을 수 있다는
§15.2의 한계가 적용된다. provider generation의 현재 선택 여부도 broker fact 자체를 바꾸지
않는다.
acks=0은 모든 selected profile에서 금지한다. first R2는 acks=all이다. acks=all은 모든
configured replica가 아니라 당시 ISR의 ACK를 뜻하므로 §13.5의 replication/min ISR/unclean
leader policy attestation과 함께 해석한다.
16.5 Deadline/cancellation/late completion
future await deadline이 끝나면:
- application outcome은
INDETERMINATE; cancel()이 broker delivery 취소를 보장한다고 가정하지 않는다;- late callback은 §15.2의 bounded source port에 payload-free observation을 제안하고 application drain이 성공한 경우에만 §18.3 append-only journal에 기록한다. persisted retry/exhausted/HOLD transition을 뒤집지 않는다;
- attempt terminal state는 atomic one-way transition이다;
- late ACK와 다음 retry가 duplicate를 만들 수 있음을 관측한다.
deadline wrapper가 worker thread만 interrupt하고 producer request를 완전히 취소하지 못한다는 한계를 descriptor에 기록한다.
16.6 Effective producer configuration
first R2는 다음을 explicit setting과 startup assertion으로 고정한다.
acks = all
enable.idempotence = true
retries = provider recommended effectively-unbounded/MAX
max.in.flight.requests.per.connection <= 5
delivery.timeout.ms = finite
request.timeout.ms = finite
max.block.ms = finite
buffer.memory = finite
batch.size = finite
linger.ms = finite
max.request.size = finite
그리고 다음 관계를 검증한다.
delivery.timeout.ms >= request.timeout.ms + linger.ms
application attempt budget >=
admission wait budget + max.block.ms + delivery.timeout.ms + callback/transition reserve
claim remaining lease >
application attempt budget + DB transition reserve + clock/scheduling safety margin
Kafka library default가 현재 원하는 값과 같더라도 explicit effective config assertion을 둔다. conflicting property가 idempotence를 끄면 startup을 실패시킨다.
retries를 작은 숫자로 잘라 broker retry를 임의 약화하지 않고 delivery.timeout.ms가 한
physical send의 시간 budget을 지배하게 한다. request.timeout.ms는 selected broker의
replica.lag.time.max.ms와 Kafka 권고 관계를 provisioning evidence로 검증한다.
size는 한 줄 부등식으로 합치지 않는다.
- exact envelope + key + headers + record overhead가 contract record bound 안;
- uncompressed record batch가 producer batch/request 제약 안;
- compressed record batch가 topic
max.message.bytes와 broker bound 안; - 여러 partition batch를 담을 수 있는 request가
max.request.size안.
모든 limit에 protocol/header/batch headroom을 두며 payload와 request/topic candidate를 똑같이 1 MiB로 두지 않는다. adopted serializer/compression의 실제 encoded batch를 real broker에서 검증한다.
exact numeric default와 허용 범위는 implementation plan의 benchmark/fault test로 고정한다. 무한 또는 사실상 운영 shutdown/SLO를 넘는 값은 허용하지 않는다.
16.7 Retry ownership
Kafka client는 delivery.timeout.ms 안에서 같은 producer send를 retry할 수 있다. relay는 하나의
application attempt가 definite/indeterminate failure로 끝난 뒤 새 attempt를 만든다.
physical Kafka retries
inside one publicationAttemptId
relay retries
new publicationAttemptId, same eventId and wire document
Kafka producer idempotence는 supported producer session의 client retries를 보호하지만 다음을 제거하지 않는다.
- producer restart 뒤 relay resend;
- ACK-to-DB gap;
- application deadline 뒤 late ACK + resend;
- polling과 CDC 이중 활성;
- operator replay.
fatal producer exception은 acceptance certainty와 별도로 generation lifecycle을 종료한다. authorization/unsupported-version/out-of-order-sequence 또는 adopted client가 fatal로 정의한 상태는 즉시 new admission 차단, readiness DOWN, bounded close/recreate를 수행한다. 같은 defunct producer를 계속 사용하지 않으며 새 generation이 application resend duplicate를 제거한다고 주장하지 않는다.
16.8 Flush
per-message KafkaTemplate.flush()를 금지한다. shared producer의 다른 batch를 강제로 flush하고
throughput/latency를 결합하기 때문이다. future completion으로 해당 record ACK를 기다린다.
flush는 bounded shutdown/explicit maintenance에서만 사용하고 그 보장과 timeout을 test한다.
16.9 Producer transaction
first polling provider는 Kafka transaction을 사용하지 않는다. Kafka transaction card가 later
추가되면 transactional ID uniqueness, producer fencing, cache size, timeout, abort, rolling deploy,
read_committed consumer까지 별도 evidence를 요구한다.
16.10 Producer generation과 rotation
credential/certificate/settings rotation은 immutable producer generation 교체로 처리한다.
소유권은 둘로 나뉜다. messaging adapter는 old/new provider generation의
pause/drain/create/attest/close primitive와 bounded fact만 제공한다. application rotation use
case는 그 fact를 바탕으로 unresolved attempt의 durable INDETERMINATE/HOLD, DB failure 시 전환
차단, generation barrier와 admission 재개 정책을 소유한다. bootstrap은 secret resolver와
application use case를 연결할 뿐 state policy를 구현하지 않는다.
1. 신규 admission과 claim을 일시 중단
2. old generation의 admitted/in-flight future를 bounded drain
3. drain deadline의 unresolved attempt를 Tx C에서 INDETERMINATE로 기록하고 영향받은 ordering
scope를 HOLD
4. old generation을 bounded close하고 더 이상 callback을 authoritative outcome으로 사용하지 않음
5. 새 secret generation resolve
6. 새 producer compile/start/attest
7. 모든 old attempt가 ACK/REJECTED 또는 durable INDETERMINATE라는 application terminal
observation을 가진 뒤 generation barrier 전환
8. HOLD 없는 scope의 admission 재개; HOLD scope는 audited duplicate-risk disposition 뒤에만 재개
한 producer object의 mutable config를 바꾸지 않는다. old/new generation metric tag는 bounded revision이어야 하며 secret value를 포함하지 않는다. first profile은 old/new generation send를 겹치지 않는 global barrier를 사용한다. 여기서 “resolved”는 broker acceptance가 definitively 밝혀졌다는 뜻이 아니라 state machine이 ACK/REJECTED/INDETERMINATE 중 하나를 durable하게 기록했다는 뜻이다. response loss의 영원한 확정을 기다리지 않는다.
barrier 전환 뒤 reorder-tolerant scope는 card가 허용한 bounded duplicate-aware retry를 자동
재개할 수 있다. ordering-required scope의 indeterminate head는 HOLD를 유지하고
§19.4 operator가 REMEDIATE_AND_REQUEUE, SKIP_WITH_GAP, COMPENSATE 중 하나를 선택한다. DB가
unavailable해 INDETERMINATE/HOLD를 durable하게 기록할 수 없으면 generation 전환과 admission을
계속 막는다. forced crash/indeterminate write 뒤 failure-path strict order는 주장하지 않고
aggregate sequence로 gap/regression을 탐지한다.
persistence가 없는 best-effort/direct caller는 durable HOLD 대상이 아니다. bounded drain 뒤
unresolved outcome을 caller/telemetry에 INDETERMINATE로 확정해 반환하고 새 generation을
전환하되, 자동 replay나 ordering 안전을 주장하지 않는다.
17. Ordering, retry budget, resource와 lifecycle
17.1 Ordering guarantee
Kafka가 제공하는 기본 ordering 범위는 한 partition 안이다. first R2의 정상 경로는 다음을 요구한다.
stable physical topic generation
+ stable non-null partition key
+ idempotent producer-compatible config
+ aggregate total sequence
+ single authoritative aggregate-head claim/admission
+ same ordering scope의 concurrent out-of-order send 금지
+ partitioner.ignore.keys = false
+ unqualified custom partitioner 없음
+ one producer generation barrier
= same generation normal-path key order + failure-path sequence detectability
global order, 여러 topic 사이 order, partition expansion 중 order, operator replay와 live stream 사이 order는 보장하지 않는다. process crash, indeterminate send, forced producer rotation, operator replay 뒤의 strict order도 첫 card 보장이 아니다. sequence metadata만으로 Kafka append order를 강제했다고 주장하지 않는다. strict effect order가 필요하면 future consumer-side sequence gate/reorder card를 추가한다.
17.2 Polling ordering gate
ordering-required contract의 다음 event는 같은 ordering scope의 앞선 delivery가
DELIVERY_RECORDED 또는 audited SKIPPED/COMPENSATED일 때만 claim한다.
EXHAUSTED head는 후행을 block한다. 자동 skip하지 않는다. hot aggregate가 전체 batch를
starve하지 않도록 batch selection은 scope별 head만 후보로 삼고 destination 전체 fairness를
관측한다.
17.3 Combined amplification budget
최악의 wire work는 대략 다음이다.
relayAttempts
× Kafka client physical retries within delivery timeout
× number of destinations
× replay generations
first baseline은 event당 destination 하나다. 설정 compiler는:
- maximum relay attempts;
- delivery-generation DB-created-at 기준 maximum automatic publication age;
- per-attempt deadline;
- backoff/jitter;
- producer internal delivery timeout;
- shutdown budget;
- dead/exhausted transition
을 하나의 descriptor로 계산한다. max attempt만 있고 maximum automatic publication age가 없는 정책은 허용하지 않는다.
17.4 Failure class와 retry
| Failure | 기본 |
|---|---|
| invalid contract/schema/size | retry 없음, writer rejection 또는 operator path |
| auth/ACL/topic policy mismatch | readiness down, 빠른 반복 retry 금지 |
| pre-admission transient metadata/network | definite rejection일 때만 bounded retry |
| post-admission leader/network/retriable | 기본 indeterminate + bounded duplicate-aware retry |
| not-enough-replicas-after-append | indeterminate |
| throttle | broker signal과 remaining budget 안에서 retry |
| local buffer exhausted | bounded admission/backpressure, retry budget 공유 |
| deadline/response loss | indeterminate, duplicate-aware retry |
| application programming defect | fail fast/alert, transient로 숨기지 않음 |
17.5 Record and memory bounds
다음을 별도로 제한한다.
- key bytes;
- value UTF-8 bytes;
- header count/key/value/total bytes;
- uncompressed record bytes;
- compressed batch bytes;
- batch size;
- request size;
- producer buffer memory;
- application admitted in-flight records;
- pending callback/attempt contexts.
Kafka client buffer.memory는 전체 producer memory의 완전한 hard bound가 아니다. compression,
in-flight request, object overhead와 callback context를 포함한 process memory budget을
capacity test로 계산한다.
17.6 Large message
contract maximum을 넘는 payload는 outbox에 append하지 않는다. large payload가 실제 요구되면 object storage에 immutable object를 먼저 publish하고 checksum/size/authorization이 있는 claim-check event를 보내는 별도 design을 사용한다.
object upload와 DB business transaction 사이 atomicity가 없으므로 staged object, outbox, orphan cleanup과 authorization을 함께 설계한다. 단순 URL을 Kafka에 넣는 것은 대안이 아니다.
17.7 Compression
compression은 provider profile이다.
- first profile은
compression.type=none으로 고정한다; - 선택 시 broker/client version 지원과 CPU/memory를 test한다;
- decompression bomb 방어를 위해 consumer는 decoded envelope/payload bound를 별도로 검증한다;
- record limit은 wire/uncompressed 의미를 혼동하지 않는다.
17.8 Admission과 backpressure
producer 내부 buffer만을 application bulkhead로 사용하지 않는다.
application admission semaphore
-> per-record JIT claim
-> Kafka producer buffer
-> broker
- admission wait는 attempt deadline에 포함한다;
- queue는 finite이며 queue timeout을 가진다;
- queue saturation 때 더 많은 outbox row를 claim하지 않는다;
- virtual thread를 사용해도 in-flight/message/memory bound는 유지한다;
- initial polling R2는 per-record JIT claim과 bounded sequential send를 사용한다. 성능 evidence가 필요할 때만 partition-key-aware concurrency card를 추가한다.
17.9 Graceful shutdown
순서:
1. readiness에서 신규 relay admission 제거
2. scheduler/new claim 중단
3. active attempt를 bounded drain
4. 완료 ACK의 delivery transition을 bounded flush
5. unresolved attempt를 indeterminate로 남기거나 lease reclaim 가능하게 종료
6. KafkaTemplate/ProducerFactory/AdminClient close
7. metrics/secret refresh resource close
shutdown timeout이 끝났다고 delivery row를 성공 처리하지 않는다. unresolved row는 lease expiry 뒤 재claim되며 duplicate 가능성이 있다.
17.10 Startup
순서:
1. typed settings bind
2. card/catalog/schema hash compile
3. secret resolve
4. producer runtime create
5. topic/security attestation
6. readiness ACTIVE
7. polling scheduler admission
relay scheduler를 producer/topic readiness보다 먼저 시작하지 않는다.
18. Polling outbox v2
18.1 Immutable event
목표 outbox_event conceptual columns:
event_id VARCHAR(96) PK, canonical US-ASCII
envelope_version
contract_id
payload_version
logical_destination_id
destination_binding_revision
aggregate_type
aggregate_id
aggregate_sequence
aggregate_event_index
partition_key_text VARCHAR(64), canonical lowercase SHA-256 hex
occurred_at
created_at
tenant_scope NOT NULL canonical scope
correlation_id
causation_id nullable
traceparent nullable, validated
tracestate nullable, validated
content_type
envelope_bytes BYTEA, exact UTF-8 wire document
envelope_sha256
envelope_schema_hash
payload_schema_hash
schema_set_hash
contract_catalog_revision
publication_epoch
dispatch_authority VARCHAR + CHECK: LEGACY_POLLING | POLLING_V2 | CDC
transaction_resource_id
규칙:
- INSERT-only after migration cutover;
- identity/order unique constraint;
- event bytes/metadata immutable;
JSONB나 재직렬화 가능한TEXT를 wire authority로 사용하지 않음;BYTEA와 hash가 polling/CDC의 exact byte authority;- polling status/attempt/owner 없음;
- CDC source predicate가 이 table의 INSERT만 받음;
- delete는 retention maintenance뿐이며 connector behavior를 test함.
18.2 Delivery control
목표 outbox_delivery conceptual columns:
event_id FK
delivery_generation
authority_status CURRENT | SUPERSEDED
superseded_by_generation nullable
dispatch_profile_id
destination_binding_revision
state
claim_count
publication_attempt_count
first_attempt_at
next_attempt_at
claim_token
claim_owner
claim_until
last_outcome_certainty
last_failure_class
last_failure_stage
provider_generation
provider_record_reference
delivery_recorded_at
terminal_at
row_version
created_at delivery generation DB creation time
automatic_attempt_deadline DB time, immutable per generation
updated_at
primary identity:
(event_id, delivery_generation)
UNIQUE(event_id) WHERE authority_status='CURRENT'로 event당 authoritative delivery
generation을 정확히 하나만 허용한다. requeue transaction은 current row를 lock하고 audit를
append한 뒤 기존 row를 SUPERSEDED로 바꾸고 delivery_generation + 1, CURRENT, READY row를
삽입한다. 새 row의 created_at과 automatic_attempt_deadline은 같은 DB transaction에서
§18.9대로 계산한다. superseded_by_generation은 새 generation을 가리킨다. update와 insert 중
하나라도 실패하면 transaction 전체가 rollback한다.
first baseline은 event 하나에 logical destination 하나다. multi-destination fan-out card를 나중에 추가하면 destination을 delivery identity와 current-authority constraint에 포함하고 한 destination 성공이 다른 destination을 완료시키지 않도록 별도 설계한다.
18.3 Append-only attempt journal
outbox_delivery_attempt_observation은 delivery control과 별도인 append-only audit다.
event_id
delivery_generation
publication_attempt_id
observation_sequence
observation_type ATTEMPT_ADMITTED | OUTCOME_OBSERVED | LATE_ACK_OBSERVED
claim_token_digest
producer_generation
destination_binding_revision
observed_at
acceptance_certainty
retry_disposition
failure_class
failure_stage
provider_record_reference
흐름:
- local admission permit를 먼저 확보한다;
- 한 짧은 claim transaction에서 한 row만 claim하고 valid remaining lease를 확인한 뒤
claim_count + 1,publication_attempt_id,publication_attempt_count + 1,ATTEMPT_ADMITTED를 함께 기록한다; - 이 durable marker 이후에만 Kafka send를 호출한다;
- marker 뒤 process crash는 실제 send 전이어도 안전하게
INDETERMINATE로 복구한다; - provider outcome observation과 Tx C state transition은 같은 short transaction에서 기록한다;
- deadline 뒤 ACK는 §15.2의 bounded source/drain 경계를 통해 성공적으로 persisted된 경우에만
LATE_ACK_OBSERVED를 append하고 delivery state를 뒤집지 않는다.
raw claim token은 journal/log/metric에 복제하지 않는다. attempt observation retention은 operator
reconciliation과 delivery retention보다 짧을 수 없다.
LATE_ACK_OBSERVED는 성공적으로 capture됐을 때 durable한 진단 사실이지만 late callback capture
자체는 crash-proof하지 않다. admission/outcome journal만 publication state machine의 필수
evidence다.
first profile에는 claim만 commit하고 나중에 queue에서 send하는 중간 상태가 없다. lease reclaim
수와 실제 publication admission 수는 별도 counter로 관측한다.
18.4 State
target polling state:
READY
CLAIMED
RETRY_WAIT
DELIVERY_RECORDED
EXHAUSTED
HOLD
SKIPPED
COMPENSATED
LEGACY_RECORDED_UNVERIFIED
LEGACY_ACCEPTED_UNVERIFIED
EXHAUSTED는 “broker에 절대 전달되지 않았다”는 뜻이 아니다. 정해진 attempt/elapsed budget 안에
delivery recording을 완료하지 못해 자동 처리를 중단했다는 뜻이다. 마지막 certainty가
INDETERMINATE이면 이미 전달되었을 수 있다.
EXHAUSTED와 HOLD는 automation-terminal이지만 ordering/retention 관점에서는 unresolved다.
HOLD는 ACK destination mismatch, invariant violation 또는 audited operator pause 때문에
자동 재시도를 허용하지 않는 상태다.
LEGACY_RECORDED_UNVERIFIED는 current void KafkaSender normal return을 보존하는 migration-only
상태이며 broker ACK, offset, delivery_recorded_at을 채우지 않는다. downstream reconciliation과
operator approval 뒤 LEGACY_ACCEPTED_UNVERIFIED로만 전이할 수 있고 이 상태도 broker ACK를
뜻하지 않는다.
consumer-side Kafka DLT와 producer-side EXHAUSTED를 둘 다 “dead letter”라고 부르지 않는다.
legacy DEAD/OUTBOX_DEAD_LETTER vocabulary는 migration alias로만 유지하고 runbook을 분리한다.
18.5 State machine
stateDiagram-v2
[*] --> READY
READY --> CLAIMED: claim(token, lease)
RETRY_WAIT --> CLAIMED: due + claim(token, lease)
CLAIMED --> DELIVERY_RECORDED: broker ACK + valid-lease token CAS
CLAIMED --> RETRY_WAIT: retryable/indeterminate + token CAS
CLAIMED --> EXHAUSTED: permanent/budget exhausted + token CAS
CLAIMED --> HOLD: ACK mismatch/invariant + token CAS
CLAIMED --> CLAIMED: lease expired + new token reclaim
EXHAUSTED --> NEW_READY: audited supersede + new generation row
HOLD --> NEW_READY: audited supersede + new generation row
EXHAUSTED --> SKIPPED: audited disposition
EXHAUSTED --> COMPENSATED: audited disposition
HOLD --> SKIPPED: audited disposition
HOLD --> COMPENSATED: audited disposition
LEGACY_RECORDED_UNVERIFIED --> LEGACY_ACCEPTED_UNVERIFIED: reconciliation + approval
state "READY (generation + 1)" as NEW_READY
DELIVERY_RECORDED --> [*]
SKIPPED --> [*]
COMPENSATED --> [*]
LEGACY_ACCEPTED_UNVERIFIED --> [*]
실제로 EXHAUSTED/HOLD row를 READY로 UPDATE하지 않는다. operator requeue는 같은 immutable
event를 참조하는 deliveryGeneration + 1 current row와 audit record를 만들고 이전 row를
SUPERSEDED authority로 바꾸는 한 transaction이다. 상태 diagram의 NEW_READY 화살표는 이전
row의 state overwrite가 아니라 이 authority handoff를 뜻한다.
18.6 Claim token와 valid-lease CAS
active worker가 소유한 renew와 outcome transition은 다음 조건을 가진다.
WHERE event_id = ?
AND delivery_generation = ?
AND authority_status = 'CURRENT'
AND state = 'CLAIMED'
AND claim_token = ?
AND claim_owner = ?
AND claim_until > database_now
affected row가 정확히 1이 아니면 stale-owner conflict다. stale worker는 broker ACK를 늦게 받아도
새 owner의 state를 DELIVERY_RECORDED/RETRY_WAIT/EXHAUSTED로 덮지 못한다. 새 worker가 아직
reclaim하지 않았더라도 lease가 만료된 old owner는 terminal state를 기록할 수 없다.
claim token은 추측 불가능한 opaque value이고 metric tag가 아니다. row_version은 JPA optimistic
locking 보조 수단일 뿐 claim token을 대체하지 않는다. worker renew와 worker-owned outcome
transition은 DB time으로 valid lease를 검사한다.
다른 mutation은 active-worker predicate를 흉내 내지 않고 각자 다음 fence를 사용한다.
| Mutation | Required predicate/fence |
|---|---|
| initial claim | CURRENT + READY 또는 due RETRY_WAIT + ordering eligibility + expected row_version; row lock 안에서 새 token/owner/DB-time lease 설정 |
| expired reclaim | CURRENT + CLAIMED + claim_until <= database_now + expected row_version; 이전 token을 새 opaque token으로 교체 |
| worker renew/outcome | 위의 current token/owner + claim_until > database_now predicate |
| operator HOLD/SKIP/COMPENSATE/legacy accept | CURRENT + expected generation/state/row_version + active unexpired claim 없음 + authorization/audit record in same transaction |
| requeue | current row lock + expected generation/state/row_version + active unexpired claim 없음; old authority SUPERSEDED와 new CURRENT/READY insert를 same transaction |
각 update의 affected row는 정확히 1이어야 한다. operator가 live worker의 token을 무시하고 raw status를 덮어쓰지 않는다. 긴급 HOLD가 필요하면 신규 claim을 먼저 fence하고 active worker drain 또는 lease expiry 뒤 operator CAS를 수행한다.
18.7 Claim eligibility
후보:
- READY;
RETRY_WAIT AND next_attempt_at <= database_now;CLAIMED AND claim_until <= database_now.
ordering-required scope에서는 더 작은 aggregate order event의
authority_status=CURRENT generation이
DELIVERY_RECORDED, audited SKIPPED/COMPENSATED 또는 audited
LEGACY_ACCEPTED_UNVERIFIED가 아니면 claim하지 않는다.
EXHAUSTED/HOLD/LEGACY_RECORDED_UNVERIFIED를 단순 terminal로 보고 통과시키지 않는다. query는
stable total order와 FOR UPDATE SKIP LOCKED를 사용한다.
18.8 First claim/lease strategy
first implementation은 postgresql-per-record-jit-claim.v1 하나로 고정한다.
local admission permit reserve
-> one eligible row JIT claim
-> valid remaining lease check
-> attempt admission journal
-> one send/observe/Tx C
-> permit release
- publish 시작 전 remaining lease가 attempt budget보다 작으면 send하지 않는다;
- renew도 claim token CAS다;
- lease expiry 뒤 late sender는 authoritative state를 바꾸지 못한다;
- duplicate publication 가능성은 남으므로 consumer inbox가 필요하다.
claimLease >
admission-after-claim reserve
+ max.block.ms
+ delivery.timeout.ms
+ callback/TxC reserve
+ scheduling safety margin
batch/window와 partition-key-aware concurrent claim은 throughput evidence가 필요할 때 별도 profile로 추가한다.
18.9 Retry time
retry due와 lease는 database time을 사용한다. backoff는 bounded exponential + jitter를 사용할 수 있지만 다음을 descriptor에 고정한다.
- claim count와 publication attempt count;
- delivery-generation DB-created-at 기준 maximum automatic publication age;
- minimum/maximum delay;
- jitter source/range;
- failure-class override;
- operator hold;
- destination backlog capacity.
현재 fixed maxAttempts=3을 영구 정본으로 보지 않는다. real fault/capacity evidence로 first R2
profile 값을 고정한다.
initial generation은 DB-authoritative outbox_event.created_at에서 automatic publication age를
시작한다. audited requeue가 만든 새 generation은 그 delivery row의 DB created_at에서 새롭지만
여전히 finite한 한-generation attempt budget을 시작하되, 원본 event의 same-ID requeue horizon을
넘지 못한다.
initial generation:
automaticAttemptDeadline =
min(
outbox_event.created_at + profile.maximumAutomaticPublicationAge,
outbox_event.created_at + contract.sameEventRequeueHorizon
)
requeue generation:
automaticAttemptDeadline =
min(
outbox_delivery.created_at + profile.maximumAutomaticPublicationAge,
outbox_event.created_at + contract.sameEventRequeueHorizon
)
계산 결과를 outbox_delivery.automatic_attempt_deadline에 immutable하게 저장해 profile reload로
기존 generation의 deadline이 움직이지 않게 한다. claim 시 database time이 deadline 이상이거나
한 full attempt budget이 남지 않으면, 첫 시도 전 backlog row라도 send하지 않고 EXHAUSTED로
fenced transition한다. first_attempt_at은 관측값일 뿐 budget을 새로 시작하지 않는다.
따라서 원본 event의 initial generation이 age로 EXHAUSTED된 뒤라도 requeue horizon 안에서 승인된
operator requeue는 새 generation에 한 번의 bounded automatic window를 부여한다. 그 window도
requeueDeadline에서 잘리며 horizon 뒤 same-ID resend는 여전히 금지한다. 오래된 event를 장애
복구 직후 발행해야 하면 이 audited requeue 또는 새 compensation/corrected event를 선택한다.
18.10 Leader election
PostgreSQL row claim과 token CAS가 correctness를 제공한다. single leader는 scheduler amplification을 줄이는 efficiency option일 수 있지만 correctness의 유일한 근거가 아니다.
현재 OutboxLeaderElectionToken marker와 “leader election” test 이름이 실제 consensus leader를
증명한다고 표현하지 않는다. 여러 instance가 claim에 참여하는 profile이라면
multi-worker-row-partitioning처럼 정확히 이름 붙인다.
19. Polling transaction, crash, disposition과 retention
19.1 Crash matrix
| Crash/failure point | Persisted state | Broker 가능성 | Recovery |
|---|---|---|---|
| business write 전 | 없음 | 없음 | caller retry |
| business write 후 outbox append 전, same tx rollback | 없음 | 없음 | caller retry |
| event/delivery commit 후 claim 전 | READY | 없음 | normal claim |
| claim commit 후 send 전 crash | CLAIMED | 없음 | lease expiry/reclaim |
| send request 뒤 ACK 전 connection loss | CLAIMED | accepted 가능 | indeterminate + reclaim |
| broker ACK 뒤 delivery CAS 전 crash | CLAIMED | accepted | reclaim, duplicate 가능 |
| ACK 뒤 stale token | 새 owner state | accepted | late owner state mutation 거부 |
| definite transient rejection | RETRY_WAIT | 미수락 확정 | due retry |
| permanent definite rejection | EXHAUSTED | 미수락 확정 | operator remediation |
| DELIVERY_RECORDED commit 뒤 process crash | DELIVERY_RECORDED | accepted | no automatic re-send |
| shutdown timeout 중 unresolved send | CLAIMED | accepted 가능 | lease reclaim, duplicate 가능 |
이 표는 “중복 없음”이 아니라 중복 발생 지점과 authoritative recovery를 고정한다.
19.2 Transaction boundaries
Tx A: business write + outbox_event + outbox_delivery
Tx B: one-row JIT claim + valid lease + publication attempt admission observation
No DB Tx: broker publish/ACK wait
Tx C: outcome observation + valid-lease token-CAS DELIVERY_RECORDED/RETRY_WAIT/EXHAUSTED
broker call을 Tx B/C 안에 넣어 DB connection/row lock을 ACK timeout 동안 잡지 않는다.
first relay command invocation은 최대 한 record만 처리해 per-record REQUIRES_NEW loop를 만들지
않는다. scheduler가 bounded rate로 다음 invocation을 요청하고, 각 invocation은 Tx B와 Tx C를
순차로 열되 capacity relation을 test한다.
19.3 Append disabled/misconfigured
R2 deployment에서 active durable contract가 하나라도 있으면 dispatch는 polling 또는 cdc여야
한다.
polling인데 ACK-aware producer/topic binding이 없으면 startup fail;cdc인데 external connector expected-state/evidence가 없으면 deployment gate fail;disabled인데 durable contract binding이 있으면 startup fail;- empty contract catalog + disabled는 resource 0.
현재 .env처럼 relay enabled + broker blank로 모든 row를 DEAD에 보내는 조합은 target에서
허용하지 않는다.
19.4 Exhausted head disposition
strict ordered aggregate head가 EXHAUSTED이면 operator는 다음 중 하나를 선택한다.
- REMEDIATE_AND_REQUEUE: 원인 수정 뒤 새 delivery generation;
- SKIP_WITH_GAP: business owner 승인과 reason/audit 뒤 후행 release;
- HOLD: 후행 계속 차단;
- COMPENSATE: 별도 compensating integration event.
raw SQL로 status=PUBLISHED를 설정해 skip을 숨기지 않는다. 모든 disposition은:
operator identity
authorization
reason code/text bound
incident/change reference
old/new generation
payload/schema hash
affected ordering scope
timestamp
approval when destructive
를 immutable audit로 남긴다.
첫 R2의 operator control surface는 기존 adapter:inbound:web leaf의 인증된 internal HTTP
endpoint 하나로 고정한다.
POST /internal/operations/messaging/outbox/{eventId}/dispositions
permission: outbox:disposition
destructive permission for SKIP/COMPENSATE: outbox:disposition:destructive
required: Idempotency-Key, expected deliveryGeneration, expected rowVersion,
disposition, bounded reason, incident/change reference
- web request/auth principal은 inbound DTO에서 application의
ApplyOutboxDispositionCommand로 mapping하고 application에 web/security 타입을 넘기지 않는다; - application-core는
ApplyOutboxDispositionUseCase와OutboxDispositionPort를 소유한다; - use case는 existing framework-free
@RequiresPermission("outbox:disposition")contract를 사용하고 destructive operation은AuthorizationPort로 추가 permission을 검증한다; - use case가 permission, allowed source state, requeue horizon, ordering impact, destructive approval reference와 compensation event reference를 검증한다;
- persistence adapter는 §18.6 operator CAS, authority handoff와 immutable audit를 한 transaction에 구현한다;
- REQUEUE는 old authority supersede + new generation insert, HOLD/SKIP/COMPENSATE는 expected current row transition이다;
- controller가 repository/entity를 직접 호출하거나 app-bootstrap이 policy를 구현하지 않는다;
- endpoint는 management/public business API와 구분한 internal network policy, strong authentication, rate bound와 audit를 요구하고 OpenAPI/public-path/security snapshot test에 포함한다;
- raw SQL과 writable Actuator endpoint는 대체 control surface가 아니다.
COMPENSATED는 “보상할 예정”이 아니다. feature owner가 만든 immutable compensating event
reference가 같은 transaction에서 검증·audit된 뒤에만 기록한다. SKIP과 COMPENSATE는
destructive permission과 승인 reference 없이는 실패한다.
19.5 Replay/requeue
producer-side requeue는 같은 event ID와 document를 새 delivery generation으로 다시 publish한다.
한 transaction에서 기존 current generation의 authority를 SUPERSEDED로 넘기고 새
CURRENT/READY generation을 만든다. 새 business event를 만들지 않는다. 이미 consumer effect가
적용되었을 수 있으므로 duplicate를 전제로 한다.
same-event requeue는 무기한 허용하지 않는다.
requeueDeadline =
outbox_event.created_at(DB time) + contract.sameEventRequeueHorizon
- horizon은 finite이고 contract/catalog hash에 포함한다;
- required consumer가 존재하면 horizon은 모든 required consumer의 inbox/dedupe archive coverage 중 최솟값 이하여야 한다;
- producer-only first R2는 end-to-end duplicate absorption을 주장하지 않더라도 deadline을 enforce하고 operator에게 downstream dedupe 확인 책임을 노출한다;
- deadline 이후 같은 event ID generation 생성은 fail-closed다;
- 오래된 EXHAUSTED/HOLD row는 audit/retention 때문에 남을 수 있지만 same-ID resend 대상은
아니다. business owner는
SKIP, 검증된 새 compensation/corrected event 또는 별도 durable dedupe-archive card를 선택한다.
payload를 수정해야 하면 기존 event를 바꾸지 않고 새 event ID/contract version을 가진 corrected 또는 compensating event를 만든다.
19.6 Retention
polling retention 조건:
the unique CURRENT authoritative generation resolved as
DELIVERY_RECORDED or audited SKIPPED/COMPENSATED/LEGACY_ACCEPTED_UNVERIFIED
AND no active claim/requeue
AND publication audit retention elapsed
AND operator/legal hold 없음
AND configured replay horizon elapsed
삭제 순서는 delivery/audit FK와 partition strategy가 결정한다. cascade가 audit를 조용히
없애지 않도록 test한다. reaper는 claim/requeue와 CAS로 경합하고 event를 먼저 지우지 않는다.
CURRENT generation이 EXHAUSTED, HOLD, LEGACY_RECORDED_UNVERIFIED이거나 unresolved attempt
observation이 있으면 automation-terminal이어도 삭제하지 않는다. superseded generation과 그
authority-handoff audit도 current generation의 전체 retention 조건이 충족되기 전에 따로
삭제하지 않는다.
Kafka topic retention이 consumer replay source라 해도 outbox event retention과 동일한 기간이라고 가정하지 않는다.
19.7 Partitioning
outbox_event는 occurred/created time 기준 range partition을 사용할 수 있다. 하지만 strict
aggregate ordering query, active delivery FK와 cleanup을 함께 benchmark한다.
closed partition 삭제는:
- polling에서는 terminal/replay 조건;
- CDC에서는 connector checkpoint proof
가 다르다. polling DELIVERY_RECORDED status를 CDC cleanup proof로 재사용하지 않는다.
19.8 Backlog capacity
durable API는 broker outage 중 DB에 event를 안전하게 쌓을 수 있으므로 Kafka 순간 장애만으로 모든 write endpoint readiness를 즉시 내릴 필요는 없다.
대신 다음을 구분한다.
- relay readiness: producer/topic에 의존;
- write admission readiness: DB free space, oldest age, backlog count/growth, retention/SLO;
- direct required producer readiness: Kafka에 직접 의존;
- liveness: 외부 dependency와 무관.
backlog capacity/SLO threshold를 넘으면 새 durable writes를 받을지 degrade할지는 deployment policy로 명시한다.
20. Best-effort publication
20.1 정확한 의미
best-effort는 다음만 보장한다.
closed contract validation
+ bounded local/producer attempt
+ outcome observation
- durable persistence
- automatic replay
- business transaction atomicity
first provider는 같은 ACK-aware Kafka gateway를 사용할 수 있다. failure를 caller에게 전파하지 않더라도 metric/log에는 ACKNOWLEDGED/ACKNOWLEDGED_MISMATCH/REJECTED/INDETERMINATE를 정확히 기록한다.
20.2 Naming
MessagePublisher처럼 durability가 모호한 이름은 migration 동안 유지할 수 있으나 target
application-facing 이름은 BestEffort...를 포함한다. durable event는 Outbox... contract를
사용한다.
20.3 Failure policy
- non-critical telemetry-like side effect만 fail-open을 선택한다;
- failure를 삼킨다고 outbox가 자동으로 대신하지 않는다;
- caller가 같은 semantic event를 best-effort와 outbox로 동시에 보내지 않는다;
- disabled best-effort binding은 호출 시 fail-fast하고 silent no-op이 아니다;
- business correctness가 delivery에 의존하면 best-effort를 선택할 수 없다.
20.4 Async optional card
caller latency를 위해 local enqueue 뒤 즉시 반환하는 truly asynchronous best-effort card를
나중에 추가할 수 있다. 그 card는 결과를 ENQUEUED로만 표현하고 broker ACK/durability를
주장하지 않는다. bounded queue, drop policy, shutdown drain과 loss metric을 별도 evidence로
가져야 한다.
21. Activation, configuration과 expected state
21.1 Canonical target shape
다음은 설계 목표 shape이며 현재 application.yml에 그대로 추가하라는 뜻이 아니다. first R2
구현이 존재할 때 구현된 필드만 live configuration으로 추가한다.
app:
messaging:
expected-state: ACTIVE
publication:
producer-provider: kafka-spring
producer-profile: acknowledged-idempotent-v1
serialization-profile: json-schema-envelope-v1
topic-profile: externally-provisioned-and-validated-v1
ordering-profile: per-key-normal-path-sequence-detectable-v1
compression-profile: none-v1
outbox:
dispatch-mode: polling
polling-profile: postgresql-polling-v2
claim-profile: postgresql-per-record-jit-claim-v1
transaction-resource-id: primary-jpa
operator-control-profile: authenticated-internal-web-disposition-v1
claim-lease: 90s
maximum-relay-attempts: 5
maximum-automatic-publication-age: 15m
same-event-requeue-horizon: 7d
kafka:
cluster-id: primary
bootstrap-servers:
- kafka-1.example.internal:9093
- kafka-2.example.internal:9093
security-profile: kafka-sasl-ssl-scram-sha-512-v1
secret-reference: secret://messaging/kafka/producer
producer:
admission-timeout: 1s
delivery-timeout: 60s
request-timeout: 35s
max-block-timeout: 5s
application-attempt-budget: 68s
buffer-memory-bytes: 33554432
maximum-request-bytes: 4194304
maximum-admitted-records: 1
destinations:
portfolio-domain-events:
binding-revision: portfolio-domain-events-r1
topic: portfolio.domain-events.v1
expected-partitions: 12
minimum-replication-factor: 3
minimum-in-sync-replicas: 2
maximum-record-bytes: 1048576
maximum-envelope-bytes: 786432
required: true
숫자는 설명을 위한 candidate다. implementation plan에서 adopted Kafka/Broker version, Testcontainers/fault/capacity evidence로 기본값과 상한을 고정한다. candidate도 protocol/header headroom과 §16.6 budget 관계를 만족하도록 서로 같은 1 MiB 값을 복제하지 않는다.
21.2 Expected state
DISABLED
ACTIVE
DISABLED:
- active contract/destination 0;
- producer factory/template/AdminClient 0;
- polling scheduler 0;
- listener container 0;
- secret refresh 0;
- network connection 0.
ACTIVE:
- exact selected tuple가 모두 known/release-eligible;
- active contract가 모두 compiled;
- required security material이 resolved;
- runtime state는
STARTING | ACTIVE_NOT_READY | ACTIVE_READY; - required destination/security/topology attestation이 fresh할 때만
ACTIVE_READY.
enabled=true와 provider 이름을 여러 곳에서 조합하지 않는다.
static schema/card/security/deadline conflict는 resource 생성 전 startup failure다. exact tuple은
유효하지만 broker/topic이 일시적으로 unavailable한 경우 first profile은 context를
ACTIVE_NOT_READY로 시작하고 scheduler admission을 막은 채 bounded backoff로 재-attest한다.
credential 누락, plaintext downgrade, unknown topic binding처럼 static/authorization failure를
transient로 숨기지 않는다. QUALIFICATION_ONLY는 test harness mode이지 deployment expected
state가 아니다.
21.3 Cross-field validation
최소 startup failure:
- ACTIVE + provider blank/unknown;
- ACTIVE + empty bootstrap server;
- ACTIVE + empty contract/destination;
- ACTIVE + unqualified card;
- polling + ACK-aware producer 없음;
- polling + delivery schema/claim settings 없음;
- polling + claim profile가 per-record JIT가 아님;
- polling + authenticated operator disposition control 없음;
- business/outbox transaction resource identity 불일치;
- CDC + polling scheduler active;
- DB publication epoch와 expected authority 불일치;
- disabled + active durable contract;
- production + plaintext;
- production + literal credential;
- idempotence와 충돌하는
acks/retries/max.in.flight; - delivery/request/linger deadline 관계 위반;
- attempt budget과 claim lease 관계 위반;
- automatic publication/requeue/inbox dedupe horizon 관계 위반;
- contract bytes > destination/provider/topic bound;
- ordering-required + null key;
- duplicate topic/binding/schema ID;
- schema/catalog/evidence hash mismatch;
- legacy와 target activation이 동시에 설정됨.
21.4 Typed settings
raw Map<String, Object> kafkaProperties를 R2 public config로 노출하지 않는다. first profile이
실제로 support하는 setting만 typed field로 제공한다.
Kafka client upgrade로 새 setting이 필요하면:
- threat/guarantee 영향 검토;
- typed setting/validation;
- effective config assertion;
- fault/security/compatibility test;
- card version 또는 evidence fingerprint update
를 함께 수행한다.
21.5 Legacy migration
현재:
APP_MESSAGING_BROKER
APP_MESSAGING_KAFKA_BROKERS
ca-skeleton.outbox.relay-enabled
목표 migration:
- legacy 값은 R0
external-kafka-sender-legacy.v1descriptor와 endpoint seed로만 해석한다; - target R2 exact tuple은 새 contract/destination/security/dispatch 설정을 모두 명시해야 한다;
- target key와 legacy key가 동시에 존재하면 fail;
broker=kafka가kafka-springR2를 자동 의미하지 않는다;relay-enabled=true/false는dispatch-mode로 대체한다;- warning과 removal release를 명시한다;
- legacy seam 사용은 descriptor에 R0로 노출한다;
- runbook/.env/README/env registry를 같은 변경에서 갱신한다.
조용한 precedence는 없다.
21.6 Environment registry
새 environment key는 docs/registries/env-keys.yaml에:
- owner;
- type/default/allowed values;
- secret classification;
- validation;
- compatibility impact;
- required test
를 등록한다.
logical destination/topic key는 fork가 실제 contract를 추가할 때 등록한다. skeleton은 존재하지 않는 sample production destination을 global env registry에 강제로 추가하지 않는다.
21.7 Secret reference
configuration에는 secret value가 아니라 reference만 둔다.
secret://messaging/kafka/producer
secret resolver가 반환하는 material은:
- char/byte lifecycle을 제한;
- log/toString/config dump에서 redact;
- generation과 expiry만 sanitized descriptor에 노출;
- rotation 실패 시 old generation 사용 가능 기간을 bounded policy로 관리한다.
21.8 Compiled runtime descriptor
startup 뒤 sanitized endpoint는 다음을 보여준다.
expectedState
runtimeState
semantic/provider/dispatch/serialization/topic/security card IDs
provider/client versions
contract catalog hash
schema set hash
destination IDs
destination binding revisions
transaction resource ID
publication epoch/dispatch authority
settings digest
producer generation
readiness level
evidence fingerprint/status
explicit non-guarantees
runbook IDs
bootstrap server, topic이 민감한 deployment에서는 hash/alias만 노출한다. credential, raw headers, payload는 절대 노출하지 않는다.
21.9 Future consumer/CDC settings
consumer와 CDC는 §24–§27의 contract가 구현될 때만 typed setting을 추가한다. 지금 live YAML에:
consumer.enabled
inbox.provider
cdc.enabled
schemaRegistry.url
같은 미구현 switch를 먼저 만들지 않는다.
22. Security와 topic governance
22.1 Threat model
| Threat | Control |
|---|---|
| arbitrary topic publish | closed destination binding |
| broker MITM | TLS hostname verification + trusted CA |
| credential leak | secret reference, redaction, generation rotation |
| over-privileged principal | destination별 least-privilege ACL |
| plaintext downgrade | production startup fail |
| payload/header injection | schema/header allowlist + byte bounds |
| cross-tenant leak | tenant-aware contract/key/auth, no dynamic topic |
| replay abuse | audited replay authorization/rate bound |
| poison/oversized record | writer validation + consumer hardening |
| dependency compromise | lock/SBOM/signature/vulnerability gate |
| topic policy drift | startup/pre-deploy attestation |
| DLT sensitive-data accumulation | restricted ACL, retention, redaction policy |
22.2 Network profile
허용 profile:
local-plaintext-v1 local/dev only
tls-server-auth-v1 controlled non-production or explicit policy
sasl-ssl-scram-v1 production candidate
sasl-ssl-oauth-v1 future/qualified candidate
mtls-v1 deployment requirement가 있을 때
production은 SSL 또는 SASL_SSL만 허용한다. PLAIN/SCRAM credential을 TLS 없이 사용하지
않는다. custom trust-all, hostname verification disable, insecure callback handler를 금지한다.
first production reference는 SASL_SSL + SCRAM-SHA-512로 고정한다. OAuth, mTLS와
server-auth-only TLS는 future profile이며 first tuple의 보장을 자동 상속하지 않는다.
22.3 TLS
- endpoint hostname verification 활성;
- protocol/cipher allowlist는 platform security policy와 정렬;
- truststore/keystore location과 password를 secret material로 취급;
- certificate expiry/chain/hostname negative test;
- rotation generation swap;
- emergency revocation runbook;
- clock skew와 certificate validity 관측;
- local self-signed CA는 explicit dev test profile에만 허용.
22.4 SASL
mechanism은 typed allowlist다. JAAS literal string을 일반 application YAML/log에 노출하지 않는다.
- SCRAM: username/password secret generation과 broker-side iteration/security 정책;
- OAUTHBEARER: issuer/audience/token endpoint TLS, token refresh deadline, secret/key rotation;
- GSSAPI: 실제 platform 요구와 qualification이 있을 때만;
- PLAIN: SASL_SSL에서만 explicit qualification.
auth refresh thread/resource도 disabled profile에서 0이어야 한다.
22.5 ACL
producer principal의 최소 권한:
Describe on required cluster/topic scope
DescribeConfigs on exact production topics
Write on exact production topics
IdempotentWrite/transactional permissions only when adopted version/profile requires
기본 금지:
Create
Delete
Alter
Write to wildcard all topics
consumer Read
Connect internal topic access
AdminClient attestation 때문에 broker-wide config나 ACL enumeration 권한을 요구하지 않는다. §13.5에서 runtime으로 확인할 수 없는 policy는 fresh signed/provenance-attested deployment provisioning evidence로 보완하고 descriptor에 verification source를 기록한다.
consumer, DLT publisher, Connect worker는 서로 다른 principal/ACL을 사용한다.
22.6 Topic provisioning
production topic은 infrastructure-as-code가 만든다.
- topic name/config review;
- partition/RF/min ISR;
- retention/cleanup;
- max message bytes;
- quota;
- ACL;
- ownership/contact;
- change/rollback record.
application startup은 validate하지 create/alter하지 않는다.
22.7 Data classification
contract descriptor는 payload sensitivity를 분류한다.
- credential/token/password를 event payload로 보내지 않는다;
- 필요한 personal data만 최소화;
- tenant/user raw identity를 partition key로 쓸 때 bounded digest/pseudonymization 검토;
- topic/DLT/outbox/inbox retention과 data deletion 법적 요구를 맞춘다;
- encryption-at-rest는 broker/DB/platform control과 evidence로 관리;
- payload encryption field-level card는 key lifecycle과 consumer authorization을 함께 설계할 때만 추가한다.
22.8 Header/tracing security
- W3C
traceparent/tracestategrammar와 size를 검증; - baggage는 default propagation하지 않음;
- inbound credential/auth/cookie/header forwarding 금지;
- exception stack/Java class를 header에 넣지 않음;
- DLT header는 원본 allowlist + safe failure code만;
- repeated retry/DLT로 header가 무한 증식하지 않게 canonical rewrite.
22.9 Tenant isolation
tenant-aware deployment는 다음을 명시한다.
- event에 tenant metadata가 필요한지;
- partition key에 tenant dimension 포함 여부;
- topic을 tenant별로 나눌지 shared로 둘지;
- producer/consumer ACL isolation;
- inbox unique scope;
- metric/log pseudonymization;
- replay authorization.
request tenant input으로 topic을 동적 생성하지 않는다. tenant topic isolation은 finite provisioned catalog로만 허용한다.
23. Observability, health와 readiness
23.1 관측 단위
다음을 분리한다.
business event append
polling claim
logical publication attempt
Kafka physical request/retry
broker acknowledgement
delivery state transition
consumer receive
application effect
offset commit
CDC source/connector checkpoint
한 publish latency에 queue/admission/broker/DB transition을 모두 합쳐 원인을 숨기지 않는다.
23.2 Producer/outbox metrics
최소 후보:
messaging.producer.attempts
messaging.producer.ack.latency
messaging.producer.queue.time
messaging.producer.outcome
messaging.producer.indeterminate
messaging.producer.buffer.available
messaging.producer.inflight
messaging.producer.throttle
messaging.producer.generation
outbox.append.total
outbox.delivery.claim.total
outbox.delivery.claim.conflict
outbox.delivery.lease.expired
outbox.delivery.outcome
outbox.delivery.exhausted
outbox.backlog.count
outbox.backlog.oldest.age
outbox.ordering.blocked.count
outbox.replay.total
exact metric name은 metrics registry naming convention에 맞춰 구현 계획에서 확정한다.
23.3 Metric tag
허용 후보:
provider_id
logical_destination_id
contract_id catalog budget 안에서만
outcome
certainty
failure_stage
failure_class
security_profile
dispatch_profile
금지:
event_id
aggregate_id
partition_key
tenant_id
user_id
correlation_id
physical offset
exception message
payload/schema hash
raw topic when not finite catalog
현재 event_type cardinality_limit=50 문서 값만 있고 runtime enforcement가 없는 상태를
readiness evidence로 보지 않는다. compiled catalog cardinality와 global meter filter를 함께
test한다.
23.4 Tracing
producer span:
logical publish span
-> Kafka client send observation
-> delivery-state DB span
consumer span:
Kafka receive/process span
-> application use-case span
-> inbox/business DB span
-> offset commit observation
Spring Kafka Micrometer Observation을 단일 instrumentation owner로 선택하고 manual trace header writer와 중복하지 않는다. trace propagation은 W3C allowlist를 사용한다. payload, key, tenant, event ID를 span attribute로 기본 기록하지 않는다.
23.5 Logs
정상 record마다 INFO log를 남기지 않는다. structured warning/error의 safe field:
error code/category
provider/logical destination/contract
outcome/certainty/failure stage
attempt count/delivery generation
opaque event ID와 correlation ID는 approved error log에서만
runbook link
payload, raw key/header, credential, full broker config는 금지한다. exception cause는 logging framework throwable로만 연결하고 message-derived arbitrary field를 만들지 않는다.
확인된 persistence transition이 canonical ERROR 한 번을 소유한다. producer callback, relay, report adapter가 같은 failure를 ERROR 세 번 남기지 않는다.
23.6 Audit
다음은 일반 log가 아니라 durable audit가 필요하다.
- destination/topic binding 변경;
- capability/profile/security generation 변경;
- outbox requeue/skip/hold/compensate;
- consumer replay;
- group/consumer identity migration;
- CDC slot/offset reset;
- polling/CDC cutover/rollback;
- ACL/secret emergency action.
23.7 Liveness
Kafka/PostgreSQL/Connect outage가 JVM liveness를 내리지 않는다. liveness는 process/event-loop deadlock 같은 내부 생존성만 본다.
23.8 Startup/readiness
role별 readiness:
| Role | Readiness |
|---|---|
| direct required producer | provider/topic/security가 unavailable이면 DOWN |
| polling relay | producer + DB claim path + contract catalog |
| durable write API | DB/outbox append + backlog capacity; 순간 Kafka outage와 분리 가능 |
| optional best-effort producer | app readiness와 분리, descriptor DEGRADED |
| future required consumer | listener assignment/contract/inbox path |
| future CDC deployment | connector task/slot/offset/WAL/topic |
continuous broker probe 하나로 모든 role을 동시에 DOWN시키지 않는다.
23.9 Readiness hysteresis
단일 transient timeout으로 readiness가 flap하지 않게:
- startup hard failure와 runtime degradation을 구분;
- consecutive failure/success 또는 freshness window;
- last successful metadata/ACK/connector progress timestamp;
- backlog/SLO threshold;
- manual maintenance state;
- recovery proof
를 descriptor에 둔다. 오래된 success를 영구 healthy로 사용하지 않는다.
23.10 Alerts/dashboard
최소 dashboard:
- publish ACK/error/indeterminate rate와 latency;
- producer buffer/admission/throttle;
- outbox backlog/oldest age/state/lease conflict;
- destination/contract별 bounded view;
- consumer phase에는 lag/rebalance/retry/DLT/inbox duplicate;
- CDC phase에는 connector state/LSN lag/WAL retained bytes/offset progress/queue.
alert는 runbook ID와 guarantee impact를 포함한다. stub runbook에 alert 이름만 있는 상태는 operational evidence가 아니다.
24. Future inbound Kafka consumer
24.1 Phase와 module gate
consumer는 first producer/polling R2와 별도 phase/card다. 구현 시작 전에:
modules.json에 inbound Kafka leaf 추가;- settings include/mapping;
- app-bootstrap allowed edge;
- nearest
CLAUDE.md; - architecture tests;
- focused test path
를 먼저 승인한다.
기존 19-leaf topology는 producer/polling phase까지 유지하고 consumer phase에서 정확히 20개로 registry migration한다.
24.2 Baseline listener configuration
첫 consumer card:
record listener
enable.auto.commit = false
AckMode = MANUAL_IMMEDIATE
asyncAcks = false
syncCommits = true
syncCommitTimeout = finite explicit value
max.poll.records = 1
bounded concurrency
bounded fetch/message bytes
finite max.poll.interval.ms
finite session/heartbeat/request timeout
DefaultErrorHandler.ackAfterHandle = false
DefaultErrorHandler.commitRecovered = false
DefaultErrorHandler.resetStateOnRecoveryFailure = false
default logging/no-op recoverer = forbidden
batch listener는 unfinished record를 건너뛰는 offset high-water, partial failure와 memory bound를 별도 증명하기 전에는 baseline이 아니다.
first card는 handler, DB transaction, retry/DLT wait와 Acknowledgment.acknowledge()를 모두
listener/consumer thread에서 동기 실행한다. off-thread worker가 ACK하지 않는다.
MANUAL_IMMEDIATE의 immediate 의미는 listener thread 호출과 explicit synchronous commit
profile에서만 주장한다. async handoff/batch는 별도 card다.
expected decode/application outcomes는 listener가 typed result로 처리하고 §24.5의 explicit
ACK/DLT/HOLD 결정을 수행한다. container DefaultErrorHandler는 listener가 놓친 unexpected
exception의 seek/redelivery safety net일 뿐 disposition owner가 아니다.
setAckAfterHandle(false)와setCommitRecovered(false)를 explicit effective assertion으로 고정한다;- retry exhaustion 뒤 정상 반환하는 Spring default logging recoverer,
CommonLoggingErrorHandler와 no-op recoverer를 금지한다; - unexpected exception의 bounded retry가 소진되면 custom terminal recoverer가 application
HoldUnexpectedConsumerFailureUseCase를 호출해 stable record identity, failed offset, failure-class와 attempt evidence를 durable consumer HOLD로 기록한다. 성공하면 listener thread가 failed offset으로 seek하고 해당 partition을 pause한 뒤 recoverer가 반환한다. assignment callback은 §25.7과 같은 durable HOLD를 재적용하므로 restart/rebalance도 자동 재시작 경로가 아니다; - durable HOLD 기록/seek/pause 중 하나라도 실패하면 recoverer는 예외를 던지고 별도
recoveryFailedlifecycle listener가 container를 bounded stop하며 readiness를 DOWN으로 만든다.resetStateOnRecoveryFailure=false를 effective assertion으로 고정해 stop과 경합해도 전체 backoff cycle을 다시 시작하지 않는다. stop이 lifecycle deadline 안에 완료되지 않으면 process liveness를 fail-closed하고 source offset은 commit하지 않는다; - baseline은 key/value
byte[]deserializer를 사용하므로 content decode failure는 listener 안의 explicit poison path를 탄다. framework-level deserializer를 나중에 쓰면 동일 no-commit invariant를 별도 evidence로 증명한다; - DLT 성공은 error handler의 “recovered” 반환이 아니라 §25.7 ACK-aware DLT gateway 성공 뒤 listener thread의 명시적 source ACK로만 표현한다;
- recoverer/DLT가 실패하거나 indeterminate이면 source commit 0이며, durable HOLD partition 또는 stopped container라는 terminal automation state가 반드시 관찰돼야 한다.
24.3 Receive sequence
1. ConsumerRecord receive
2. key/header/value byte bounds
3. header allowlist + envelope decode
4. envelope/payload schema/version validation
5. destination/subscription/contract allowlist
6. application command mapping
7. consume use case / MessageConsumptionExecutor
8. APPLIED 또는 DUPLICATE commit 확인
9. acknowledgement
10. offset commit result observation
Kafka SDK type은 step 6에서 끝난다. application command는 provider-neutral event metadata와 typed payload만 가진다.
24.4 Deserialization failure
listener method 전에 발생하는 deserializer exception도 다룬다.
- byte[]로 먼저 받고 bounded envelope codec에서 decode하는 방식을 baseline 후보로 한다;
- framework deserializer를 쓰면
ErrorHandlingDeserializer/동등 error path를 명시한다; - trusted Java package/default typing으로 arbitrary class를 만들지 않는다;
- malformed UTF-8/schema/unknown version은 무한 retry하지 않는다;
- DLT publish ACK 전 source offset을 진행하지 않는다.
24.5 Ack result
| Application result | Listener |
|---|---|
| APPLIED | ACK |
| DUPLICATE with same document hash/effect generation | ACK |
| RETRYABLE_FAILURE | no ACK, bounded retry |
| REJECTED/PERMANENT, reorder-tolerant subscription | DLT ACK 뒤 source ACK |
| REJECTED/PERMANENT, strict ordered subscription | idempotent quarantine 뒤 partition HOLD |
| IDENTITY_COLLISION | idempotent quarantine 뒤 subscription policy |
| DB commit outcome unknown | no ACK, retry; inbox로 reconcile |
| listener shutdown/revoke before commit | no ACK |
ack 호출 뒤 offset commit failure도 관측한다. commit failure는 record redelivery를 만들 수 있으며 inbox가 business duplicate를 막아야 한다.
strict ordered subscription은 DLT ACK만으로 gap을 승인하지 않는다. operator가 audited
ADVANCE_WITH_GAP 또는 compensation을 승인한 뒤에만 source offset을 진행한다.
24.6 Bounded processing
- handler + bounded retry/backoff + DB pool/lock/deadlock retry + GC/scheduler reserve +
synchronous offset commit의 worst case가
max.poll.interval.ms안에 들어야 한다; - handler concurrency는 partition ordering, DB pool, executor queue에 맞춘다;
- one in-flight per partition가 first ordered baseline이다;
- first baseline은 async executor를 사용하지 않는다;
- queue saturation 때 container/partition pause로 poll heartbeat를 유지;
- capacity 회복 때 resume;
- pause가 buffer/fetch memory를 무한하게 만들지 않게 monitoring한다.
24.7 Rebalance
first synchronous card에서 rebalance callback은 long-running handler drain 장소가 아니다. handler budget이 poll membership deadline 안에서 끝나야 하며 callback은 finite한 다음 작업만 한다.
- revoked partition의 신규 dispatch 중단;
- 이미 commit된 APPLIED/DUPLICATE offset만 callback budget 안에서 commit 시도;
- commit-failed/rebalance-in-progress는 redelivery로 분류;
- 미완료 work는 ACK하지 않음;
- resource/context 정리와 assignment generation 갱신.
consumer는 thread-safe하다고 가정하지 않는다. cooperative assignor/static membership는 rebalance evidence를 통과한 optional profile이다. off-thread processing을 추가하면 continued polling, per-partition unfinished high-water와 listener-thread ordered ACK handoff를 별도 설계한다.
24.8 Shutdown
readiness DOWN
-> listener pause/new dispatch stop
-> active DB transaction bounded drain
-> eligible ACK/commit
-> unresolved no-ACK
-> container close
-> DLT producer close
shutdown timeout 뒤 unfinished record를 ACK하지 않는다.
24.9 External side effect
consumer handler가 DB inbox transaction 안에서 HTTP/email/object storage side effect를 직접 수행하면 same-store atomicity가 없다.
기본 pattern:
inbox + business state + follow-up outbox intent
same DB transaction
external side effect
별도 durable worker/provider
외부 side effect를 반드시 inline 수행해야 하면 idempotency/reconciliation/compensation을 feature-specific design으로 추가하고 inbox만으로 exactly-once라고 표현하지 않는다.
25. Inbox, retry, DLT, replay와 Kafka EOS
25.1 Inbox contract
application-core는 framework-free InboxStorePort와 MessageConsumptionExecutor를 소유한다.
PostgreSQL provider는 persistence adapter가 구현한다.
conceptual inbox_consumption:
consumer_id
event_id
tenant_scope NOT NULL canonical scope
contract_id
payload_version
document_sha256
effect_contract_version
effect_generation default 0
source_reference safe bounded diagnostic
applied_at
created_at
PK/UNIQUE (consumer_id, effect_generation, tenant_scope, event_id)
tenant-disabled deployment도 non-null canonical system scope를 사용한다. 일반 nullable UNIQUE에 dedupe를 맡기지 않는다.
25.2 Same transaction algorithm
tx.inWrite:
validate handler/contract
INSERT ... ON CONFLICT DO NOTHING RETURNING inbox identity
inserted:
execute business mutation
optional follow-up outbox append
commit -> APPLIED
no returned row:
load existing bounded metadata
same document hash/effect version/generation -> DUPLICATE
mismatch -> IDENTITY_COLLISION
business mutation이 실패하면 inbox insert도 rollback한다. PROCESSING row를 먼저 별도 transaction에
commit해 영구 stuck 상태를 만들지 않는다. plain INSERT unique exception을 catch한 뒤 같은
PostgreSQL/JPA transaction을 계속 사용하지 않는다. native
ON CONFLICT DO NOTHING RETURNING 또는 동일 의미의 검증된 atomic primitive를 사용하고 두
consumer 동시 claim을 real PostgreSQL에서 test한다.
25.3 Crash behavior
| Point | Result |
|---|---|
| inbox insert 전 crash | redelivery, normal apply |
| insert 뒤 business mutation 전 crash/rollback | row 없음, redelivery |
| business + inbox commit 전 crash | rollback, redelivery |
| commit 뒤 ACK 전 crash | redelivery -> DUPLICATE -> ACK |
| ACK 뒤 offset commit response loss | redelivery 가능 -> DUPLICATE |
25.4 Inbox retention
inbox retention은 최소 다음보다 길어야 한다.
Kafka replayable retention
DLT retention
maximum audited replay horizon
maximum producer duplicate/requeue horizon
cross-region/cold-recovery horizon when applicable
inbox를 먼저 지우고 Kafka/DLT record를 다시 replay하면 effect가 재적용된다. cleanup은 consumer contract version, legal retention과 archive policy를 검증한다.
각 consumer card는 finite dedupeHorizon과 source/DLT/archive replay cutoff를 pin한다.
dedupeHorizon은 자신이 소비하는 모든 producer contract의 sameEventRequeueHorizon 이상이어야
하며 release compiler가 compatibility matrix에서 이를 검증한다.
무한 Kafka retention, legal hold 또는 cold archive가 있으면 inbox도 보존하거나 별도 durable
dedupe archive를 제공해야 한다. purge cutoff보다 오래된 replay는 자동 earliest/apply가 아니라
unsupported incident로 fail한다.
25.5 Baseline retry
첫 consumer card는 짧고 bounded한 blocking/seek retry다.
- retryable failure class allowlist;
- small maximum attempts;
- total elapsed bound;
- backoff가 max.poll/rebalance와 호환;
- same partition ordering 유지;
- long dependency outage를 listener thread에서 오래 sleep하지 않음;
- remaining attempts/age가 끝나면 DLT/operator path.
구체 횟수/시간은 handler SLO와 real fault test로 고정한다.
25.6 Retry topic optional card
non-blocking retry topic은 main record를 retry topic으로 publish하고 source offset을 진행한다. Kafka ordering을 잃으므로:
- unordered/explicitly reorder-tolerant contract만;
- original event ID/contract/exact document hash 유지;
- retry generation/attempt metadata bounded;
- retry/DLT topic provisioning/ACL/retention;
- retry publish ACK 뒤 source ACK;
- retry ACK 뒤 source commit crash가 duplicate retry record를 만들므로 stable retry identity/dedupe;
- container transaction과 adopted Spring Kafka version의 제약 검증;
- live/retry stream의 stale effect policy
를 요구한다.
25.7 DLT
consumer DLT는 producer-side outbox EXHAUSTED와 다르다.
future consumer tuple은 별도
kafka-consumer-dlt-acknowledged.v1 provider card를 반드시 선택한다. 이 provider는 inbound
Kafka leaf가 소유하며 outbound messaging leaf에 의존하지 않는다. 이는 §5의 HARD invariant 3에
둔 consumer-processing-local publisher 예외이며 closed DLT/retry binding 외 publish에는 사용할
수 없다. 최소 exact profile:
acks=all
enable.idempotence=true
retries=effectively-unbounded/MAX within finite delivery.timeout.ms
max.in.flight.requests.per.connection<=5
finite admission/max.block/request/delivery/buffer/record/header bounds
ByteArraySerializer key/value with prevalidated DLT bytes
closed pre-provisioned DLT binding, auto-create disabled
SASL_SSL/SCRAM least-privilege Write/Describe ACL
future metadata ACK + expected topic verification
bounded producer generation rotation/shutdown
DLT gateway outcome도 ACKNOWLEDGED, ACKNOWLEDGED_MISMATCH, REJECTED,
INDETERMINATE를 구분한다. mismatch/indeterminate/timeout/close는 source ACK를 허용하지
않는다. application outbox producer의 card/evidence를 이름만 재사용하지 않고 consumer leaf에서
real broker/security/fault evidence를 별도로 만든다. 공통 구현 추출은 §31.5의 module-split
trigger가 실제로 충족될 때만 한다.
DLT record:
- stable
dltIdentity = hash(clusterAlias, topic, partition, offset, consumerId, effectGeneration); - original event ID/contract/version/key/value 또는 approved sanitized representation;
- original topic/partition/offset safe reference;
- bounded failure code/stage;
- first/last failure timestamp;
- consumer/effect contract version;
- replay generation;
- no raw credential;
- no unbounded stacktrace/header chain.
first DLT publish와 source offset commit은 Kafka transaction으로 원자적이지 않다. DLT ACK 뒤
source commit 전 crash/response loss는 duplicate DLT를 만든다. DLT tooling은 dltIdentity로
dedupe하고 이 crash를 test한다.
reorder-tolerant subscription의 source ACK 조건:
DLT producer future ACKNOWLEDGED
AND DLT metadata verified
THEN source acknowledgement
DLT publish가 실패/indeterminate면 source offset을 진행하지 않는다.
quarantine profile은 source와 같거나 더 엄격한 sensitivity ACL, encryption-at-rest, finite byte bound와 retention을 가진다. poison raw key/header/value를 재생 가능하게 보존할지 sanitized non-replayable evidence만 보존할지는 contract별로 하나를 고정한다. sanitized mode는 자동 replay 불가를 descriptor에 노출한다. strict ordered subscription은 successful quarantine 뒤 partition을 listener thread에서 failed offset으로 seek한 뒤 pause/HOLD하고 audited disposition 전 source ACK를 하지 않는다.
strict-order HOLD는 container memory에만 두지 않는다. application-core가
ConsumerPartitionHoldPort와 hold/disposition use case를 소유하고 persistence adapter가 다음
durable control을 구현한다.
logical_subscription_id
consumer_id
effect_generation
cluster_alias
topic_binding_revision
partition
failed_offset
event_id
document_sha256
state HOLD | ADVANCED_WITH_GAP | COMPENSATED | RELEASED_FOR_RETRY
reason/incident/approval
row_version
created_at/updated_at
UNIQUE(logical_subscription_id, effect_generation,
cluster_alias, topic_binding_revision, partition)
- quarantine ACK와 HOLD insert는 동일한 provider transaction이 아니므로 source ACK는 여전히 하지 않으며, 두 결과를 reconciliation 가능한 stable identities로 기록한다;
- assignment callback은 dispatch 전에 application hold query를 호출하고 held partition을 failed offset에 seek/pause한다;
- restart/rebalance/new pod도 durable HOLD를 다시 적용하며 in-memory pause 소실로 poison을 진행하지 않는다;
- operator disposition은 expected row version CAS, permission, approval/audit를 요구한다;
ADVANCED_WITH_GAP또는 verified compensation 뒤에만 listener thread가 failed offset 이후로 명시적 commit/resume한다;- HOLD partition lag는 정상 retry lag와 분리하고 required subscription readiness를 DEGRADED/DOWN 정책에 따라 표시한다.
25.8 Replay
live consumer group offset을 임의 rewind하지 않는다. 별도 replay job/group은:
replayOperationId
source (DLT/topic/archive)
contract/version allowlist
time/partition/offset/event-id scope
target consumer/effect version
reuse or new replay generation
dry-run count/hash
rate/concurrency limit
operator/approver/reason
start/stop/progress/result
를 가진다.
기본 replay는 같은 inbox identity를 사용하므로 이미 APPLIED event는 DUPLICATE가 된다. 의도적으로
effect를 다시 적용하려면 unique key에 참여하는 새 effectGeneration, business owner 승인,
compensation 위험을 명시한다. consumer/effect generation별 durable replay lease는 overlapping
replay job, live replay와 inbox cleanup race를 막는다.
25.9 Offset out of range
topic retention 뒤 offset이 사라졌을 때 자동 earliest/latest reset으로 data gap을 숨기지 않는다.
auto.offset.reset은 profile에 explicit하며 required consumer의 offset out-of-range는 startup 또는
runtime incident다. replay/archive/bootstrap 절차를 선택한다.
25.10 Kafka EOS optional card
DB-free Kafka consume-process-produce는 Kafka transaction으로:
input read_committed
process
output records + source offsets in one Kafka transaction
을 구성할 수 있다.
이는:
- DB write;
- HTTP/email/storage side effect;
- PostgreSQL inbox;
- 다른 non-transactional system
을 포함하지 않는다. 해당 card만 “Kafka transaction 범위의 exactly-once processing”이라고 제한해 표현한다.
Spring의 DB/Kafka transaction synchronization은 commit 순서를 조정할 뿐 distributed atomic commit이 아니다. 두 번째 commit failure compensation을 별도 설계해야 한다.
26. Future PostgreSQL Debezium CDC
26.1 위치
CDC는 application process 안의 scheduler가 아니다.
PostgreSQL logical decoding
-> replication slot/publication
-> Debezium PostgreSQL connector
-> Outbox Event Router
-> Kafka Connect producer
-> Kafka topic
Java repository는 immutable event schema/contract와 deployment expected-state descriptor를 제공한다. connector worker/image/config는 deployment asset이다.
26.2 Prerequisite
first future CDC qualification target는 다음 exact family다.
PostgreSQL 16 + pgoutput
Debezium PostgreSQL/Outbox Event Router 3.6.0.Final
Kafka Connect worker exact patch/image digest pinned by the implementation plan
snapshot.mode = no_data for cutover connectors
publication.autocreate.mode = disabled
production publication = exact outbox table + CDC row filter + INSERT only
partition_key_text VARCHAR + StringConverter key
envelope_bytes BYTEA + Debezium BinaryDataConverter value
header.converter = Kafka SimpleHeaderConverter
binary.handling.mode = bytes
errors.tolerance = none
transforms.outbox.table.op.invalid.behavior = fatal
skipped.operations = t
ordering = commit-order/detectable-sequence, strict aggregate order unsupported
resolved Connect/Kafka/plugin patch와 image digest가 없으면 이 card는
not-implemented다. floating stable/current documentation은 discovery일 뿐 evidence가 아니다.
CDC card를 활성화하기 전:
outbox_event신규 row가 insert-only;- legacy status UPDATE writer 0;
- connector invalid UPDATE behavior가 fatal/alert로 검증;
- event ID/key/envelope/schema fields가 CDC mapping 가능;
- event row에
publication_epoch와dispatch_authority=CDC가 존재; - PostgreSQL production publication이 exact outbox table의
WHERE (dispatch_authority = 'CDC')row filter와publish='insert'만 사용; - polling/CDC wire golden parity;
- PostgreSQL logical replication prerequisites;
- dedicated publication/slot;
- Connect internal topics;
- connector/task security;
- snapshot/cutover/retention runbook;
- real end-to-end evidence
를 모두 만족한다.
현 mutable V3 table에 connector만 붙이는 것은 금지한다. strict aggregate ordering contract도 §26.10의 별도 serialization/reorder card 없이 first CDC profile에 bind하지 못한다.
26.3 Connector mapping
Outbox Event Router mapping은 최소 다음을 고정한다.
event ID column -> canonical `id` header
partition_key_text VARCHAR -> StringConverter -> exact US-ASCII Kafka key bytes
envelope_bytes BYTEA -> Kafka value bytes
occurredAt column -> record timestamp policy
contract/version -> bounded headers when required
logical destination -> closed route mapping
traceparent/tracestate columns -> validated bounded headers
Debezium default aggregateType -> dynamic topic를 그대로 사용하지 않는다. closed logical
destination/topic allowlist와 route regex/replacement를 exact config로 관리한다. first CDC
profile은 finite DB CHECK/catalog value, exact-match route, pre-provisioned topic,
auto-create disabled와 connector ACL을 모두 사용한다. unknown route는 다른 topic으로 fallback하지
않고 connector를 실패시킨다.
EventRouter는 heartbeat/schema/transaction/tombstone 같은 non-outbox record에 적용하지 않는다. exact source-topic/table SMT predicate를 사용한다. production authority filter는 scripting SMT가 아니라 PostgreSQL 16 publication row filter로 고정한다.
CREATE PUBLICATION <pinned_name>
FOR TABLE ONLY <schema>.outbox_event
WHERE (dispatch_authority = 'CDC')
WITH (publish = 'insert', publish_via_partition_root = true);
publication.autocreate.mode=disabled와 pinned publication.name을 사용한다. startup/deployment
attestation은 pg_publication, pg_publication_tables/row-filter catalog를 읽어 exact table,
row filter, pubinsert=true, pubupdate/pubdelete/pubtruncate=false,
publish_via_partition_root=true와 다른 connector publication이 섞이지 않았음을 확인한다.
dispatch_authority는 PostgreSQL user-defined enum이 아니라 bounded VARCHAR + CHECK로 저장해
PostgreSQL row-filter의 built-in type/operator 제약 안에 둔다.
outbox를 실제로 partition하지 않는 implementation에서도 이 값을 pin해 future partition
동작을 조용히 바꾸지 않는다. first profile은 Debezium scripting Filter SMT/plugin을 요구하지
않는다.
UPDATE/TRUNCATE 또는 unexpected operation을 한 설정으로 뭉뚱그리지 않는다.
- UPDATE는
transforms.outbox.table.op.invalid.behavior=fatal로 EventRouter가 connector를 중지하게 한다. INSERT-only publication 때문에 정상적으로 관찰될 수 없고, publication drift에 대한 defense-in-depth다. 기본값warn은 허용하지 않는다; - DELETE/TRUNCATE는 production publication에서 publish하지 않는다. §26.13의 승인된 retention DELETE가 production event/tombstone을 emit하지 않음을 검증한다;
skipped.operations=t도 connector defense-in-depth로 pin한다. runtime role의 DELETE/TRUNCATE privilege 제거와 migration gate/audit가 주 방어선이며, 예상 밖 UPDATE/DELETE/TRUNCATE가 DB audit에서 발견되면 connector를 중지하고DATA_GAP_SUSPECTED로 전이한다;errors.tolerance=none은 converter/SMT가 실제로 throw한 오류를 skip/DLQ로 우회하지 않는 정책이지 UPDATE/TRUNCATE 자체의 분류 설정이 아니다.
polling-era event는 PostgreSQL publication row filter에서 production CDC source에 들어오지 않는다. shadow connector만 별도 insert-only publication/slot과 격리 topic에서 authority 전체를 비교할 수 있다.
26.4 Insert-only behavior
- application은 event row를 UPDATE하지 않는다;
- polling state는 delivery table에만 있다;
- cleanup DELETE는 retention proof 뒤에만 실행하고 INSERT-only publication의 production output 0을 검증;
- update event가 관찰되면
transforms.outbox.table.op.invalid.behavior=fatal로 warning 없이 stop/alert; - CDC connector는 outbox table만 capture하도록 include list/predicate를 제한한다.
- runtime role의 UPDATE/DELETE/TRUNCATE/DDL privilege를 제거하고 migration role만 별도 승인;
- exact publication row filter +
publish=insert,skipped.operations=t, table list, partition-root behavior를 config/catalog attestation으로 pin; - DROP/DETACH/TRUNCATE는 logical decoding alert에만 의존하지 않고 migration gate/audit에서 차단.
26.5 Payload
first CDC profile은 PostgreSQL binary.handling.mode=bytes로 읽은 BYTEA envelope_bytes를
EventRouter 결과 value로 만들고
value.converter=io.debezium.converters.BinaryDataConverter로 exact bytes를 emit한다.
heartbeat 같은 non-outbox record는 BinaryDataConverter가 처리할 수 없으므로 공식
value.converter.delegate.converter.type=org.apache.kafka.connect.json.JsonConverter와
value.converter.delegate.converter.type.schemas.enable=false 설정을 pin한다. delegate가 만든
record는 exact source-table predicate와 closed route에 의해 production event topic으로 들어갈 수
없어야 한다. JsonConverter/String expansion으로 event envelope 자체를 다시 직렬화하는
profile은 first profile이 아니다.
key는 partition_key_text VARCHAR(64)를 EventRouter key field로 선택하고
key.converter=org.apache.kafka.connect.storage.StringConverter로 직렬화한다. polling
producer가 쓰는 US-ASCII bytes와 동일함을 golden vector로 검증한다. event ID의 EventRouter 기본
id header를 canonical header로 채택하며 alias를 하나 더 남기지 않는다.
header.converter=org.apache.kafka.connect.storage.SimpleHeaderConverter를 default에 맡기지 않고
explicit pin한다. event ID/contract/version/trace header source column은 bounded canonical
ASCII/UTF-8 STRING Connect type으로 유지하고 null/optional placement을 exact config로 고정한다.
polling의 exact UTF-8 header bytes와 adopted Kafka Connect version의
SimpleHeaderConverter output을 golden test로 비교한다. key/header binary representation과 full
SMT/converter chain을 golden test로 고정한다.
다음이 polling과 같아야 한다.
- event ID;
- key bytes;
- contract/payload/envelope versions;
- exact value bytes와
envelope_sha256; - required headers;
- record timestamp policy.
malformed JSON은 writer admission에서 거부되어야 한다. CDC converter가 malformed string을 정상 string value로 우회시키는 profile은 qualification 실패다.
26.6 Offset와 internal topics
Kafka Connect distributed mode는 최소:
- config storage topic;
- offset storage topic;
- status storage topic
의 partition/RF/cleanup/ACL을 운영 profile로 고정한다. offset reset/alter는 destructive audited operation이다. connector REST API 노출과 authorization도 제한한다.
connector/task restart 뒤 마지막 committed offset부터 중복 change event가 재방출될 수 있다. consumer inbox가 이를 흡수해야 한다.
source connector producer도 application producer card의 보장을 자동 상속하지 않는다. exact Connect worker/connector profile은 최소 다음을 별도로 pin한다.
acks=all
enable.idempotence=true
max.in.flight.requests.per.connection<=5
finite delivery/request/max.block/buffer/request bounds
SASL_SSL/SCRAM credential and least-privilege exact-topic ACL
topic auto-creation disabled
errors.tolerance=none
converter/SMT failure -> task FAILED, no skip
Connect internal topic producer/consumer 권한과 outbox destination 권한은 분리한다.
26.7 Replication slot와 WAL
monitor:
slot exists/active
confirmed_flush_lsn/restart_lsn
current WAL LSN
retained WAL bytes
connector source lag
last event/heartbeat
database disk free
publication/table inclusion
slot catalog state
low-traffic database는 heartbeat/action query가 WAL progress에 미치는 영향을 exact connector version으로 검증한다.
slot drop/recreate는 이전 LSN history를 복구하지 못할 수 있고 silent gap을 만들 수 있다. 자동 recreate 뒤 healthy 표시를 금지한다.
max_slot_wal_keep_size, database disk free admission과 operator emergency threshold를 finite
deployment 값으로 pin한다. cap 초과로 required WAL segment가 제거되면 slot/card는
DATA_GAP_SUSPECTED이며 resnapshot/reconciliation 전 READY로 복귀하지 않는다. monitoring만
있고 WAL/disk bound가 없는 profile은 R2가 아니다.
26.8 Snapshot
first shadow와 production cutover connector는 snapshot.mode=no_data를 사용한다. 각 connector의
unique logical slot을 writes-frozen boundary에서 생성하고 slot consistent point 이후 INSERT만
stream한다. existing outbox history를 snapshot으로 다시 publish하지 않는다. initial snapshot이
필요한 다른 profile은 historical duplicate scope와 inbox coverage를 별도 card로 증명한다.
snapshot/restart 중 duplicate, schema change, queue saturation, connector crash를 test한다.
26.9 PostgreSQL 16 failover limitation
repository의 first CDC database target은 PostgreSQL 16이다. 이 profile은 PostgreSQL 17+의 failover logical slot continuity를 주장하지 않는다.
- primary failover 시 connector/card는 즉시 NOT_READY;
- 새 primary의 slot, Connect offset과 source LSN을 수동 reconcile;
- missing/ahead/behind 또는 WAL loss가 있으면
DATA_GAP_SUSPECTED; - resnapshot/audited backfill/inbox reconciliation 전 writes 재개 조건을 runbook으로 판단;
- automatic slot recreate와 no-gap 표현 금지.
PostgreSQL 17+ synchronized failover slot은 별도 future card와 real failover evidence가 있을 때만 추가한다.
26.10 Ordering compatibility
logical decoding은 transaction commit order를 emit한다. application이 allocate한
aggregateSequence와 commit order는 두 transaction이 역순 commit하면 다를 수 있다.
따라서 first CDC card는:
- strict aggregate ordering guarantee를 제공하지 않는다;
- key와 aggregate sequence를 보존해 gap/regression을 탐지 가능하게 한다;
- shadow compare도 global/aggregate sequence order가 아니라 source offset, event set, key와 exact bytes를 비교한다.
strict ordered CDC가 필요하면 같은 aggregate transaction serialization으로
sequence order == commit order를 증명하거나 Kafka 전 reorder/consumer sequence gate를 가진
별도 card가 필요하다.
26.11 DDL와 source limitations
qualification은 최소 다음을 다룬다.
- logical decoding이 DDL event를 직접 제공하지 않는 한계;
- outbox schema migration과 rolling application;
- partition root publication behavior;
- primary key/replica identity 변경;
- TOAST/unchanged value behavior가 envelope column에 미치는 영향;
- delete/tombstone;
- TRUNCATE/DROP/DETACH와
skipped.operations; - connector/plugin/JDBC/PostgreSQL version compatibility.
26.12 CDC readiness
CDC readiness는 polling metric을 재사용하지 않는다.
connector/task RUNNING
AND slot/publication valid
AND WAL retained bytes within bound
AND topic/security/schema attested
AND last cutover epoch/watermark reconciled
AND expected connector name/config/image/SMT hash matches
AND expected task count/source producer profile matches
AND last successful source interaction/offset commit is fresh for observed source activity
idle database에서 LSN이 움직이지 않는다는 이유만으로 DOWN시키지 않는다. worker/task liveness,
last source interaction, actual source activity, Kafka offset commit freshness와 measured lag를
분리한다. RUNNING이나 stale prior attestation만으로 no-gap/READY를 주장하지 않는다.
application liveness와 분리한다.
26.13 CDC retention proof
time partition이 오래됐다는 이유만으로 삭제하지 않는다.
closed partition의 모든 source transaction/batch와 destination partition이:
- persisted Connect source offset과 slot checkpoint에서 coverage됐고;
- expected Kafka destination의 event ID/hash manifest에서 관찰됐고;
- replay retention이 지났고;
- connector offset/slot continuity가 검증되었고;
- legal/operator hold가 없다는
구체적 proof를 남긴 뒤 purge한다. 한 sentinel의 한 Kafka partition 관찰이나
confirmed_flush_lsn 하나만으로 다른 destination partition coverage를 추론하지 않는다.
27. Polling/CDC shadow, cutover와 rollback
27.1 Mutual exclusion은 두 층이다
application setting 하나만으로 외부 Connect deployment를 막을 수 없다.
두 층을 모두 사용한다.
- database publication epoch/dispatch authority fence;
- infrastructure authority: deployment replica state, connector state, producer/connector ACL.
같은 production topic에 polling producer principal과 CDC connector principal의 Write authority를 동시에 열지 않는다.
outbox_publication_epoch conceptual control row:
epoch_id monotonic PK
authority LEGACY_POLLING | POLLING_V2 | CDC
state PREPARED | ACTIVE | RETIRED
source_boundary
settings_digest
activated_at
activated_by/approved_by
row_version
정확히 한 ACTIVE row를 partial unique constraint로 보호한다. append adapter는 business
transaction 안에서 ACTIVE row를 FOR SHARE로 읽고 event에 epoch_id/authority를 기록한다.
POLLING_V2면 같은 transaction에서 delivery row를 만들고, CDC면 만들지 않는다. epoch
activation은 같은 row를 FOR UPDATE로 전환하므로 concurrent append와 serializes한다.
old binary가 cached config로 다른 mode를 쓰지 못하도록:
- target append는 DB epoch를 authority로 사용;
- activation 전 old writer/relay binary 0을 deployment fingerprint로 확인;
- DB trigger/constraint가 stale epoch, CDC event의 delivery row, polling event의 delivery 누락을 거부;
- active authority와 맞지 않는 legacy status mutation을 DB guard가 거부한다.
application setting은 expected epoch/authority assertion일 뿐 DB truth를 덮지 않는다.
27.2 Shadow qualification
shadow CDC는:
- production과 다른 topic;
- production consumer가 읽지 않는 group;
- same immutable event source;
- same schema/key mapping;
- event ID/count/exact byte hash/key/source-offset/lag compare;
- bounded retention와 restricted ACL
을 사용한다.
shadow record를 production effect에 적용하지 않는다. shadow 성공은 cutover rehearsal/evidence지 production CDC ACTIVE가 아니다.
topology는 두 connector/두 slot으로 고정한다.
shadow connector + shadow logical slot + shadow topic
production connector + production logical slot + production topic
slot이나 Connect offset을 공유·복사·재사용하지 않는다. shadow connector의 advanced offset을 production topic으로 retarget하지 않는다. production connector/slot은 writes-frozen cutover에서 새로 만든다.
27.3 Cutover invariant
cutover는 다음을 증명해야 한다.
모든 pre-cutover event가 polling 또는 reconciled duplicate로 처리
AND 모든 post-cutover event가 CDC source boundary에 포함
AND 같은 event가 두 authority에서 production topic으로 발행되는 window가 없음
AND rollback boundary가 기록됨
27.4 Polling -> CDC high-level sequence
exact Debezium 3.6.0.Final/PostgreSQL 16/Connect worker runbook이 세부 명령을 소유한다. 고수준 순서:
- 별도 shadow connector/slot/topic을
no_data로 qualification; - bounded write maintenance 시작, 신규 business transaction admission 중단, active write drain;
- polling new claim 중단, active attempt drain, indeterminate/backlog reconcile;
- polling producer production Write ACL revoke, producer close, old relay fence 확인;
- writes가 frozen인 상태에서 exact CDC row filter + INSERT-only인 dedicated production publication과 새 production slot 생성;
- slot creation이 반환한 consistent point/source boundary와 empty Connect offset identity 기록;
snapshot.mode=no_data,publication.autocreate.mode=disabled, pinned publication, exact source-table SMT predicate, StringConverter-key/BinaryDataConverter-value chain인 production connector를 준비하고 exact topic Write ACL만 부여;- DB transaction에서 new CDC epoch ACTIVE 전환과 CDC-authority cutover sentinel INSERT를 원자적으로 수행; sentinel에는 polling delivery row가 생기지 않음;
- production connector를 start/resume하여 unique slot boundary부터 stream;
- production topic에서 sentinel exact event ID/hash를 확인하고 persisted Connect offset, slot LSN과 reconcile;
- target application expected epoch를 확인한 뒤 regular writes 재개;
- event set/key/hash/source lag와 consumer inbox duplicate를 관측;
- rollback window 동안 polling artifacts와 production slot/offset을 보존.
pg_current_wal_lsn()을 application transaction에서 읽은 값이나 shadow connector offset을 commit
boundary로 추정하지 않는다. authoritative boundary는 writes-frozen 상태에서 생성한 production
logical slot consistent point, persisted Connect source offset과 CDC-epoch marker transaction의
correlation evidence다.
27.5 CDC -> polling high-level sequence
- incident/cutback 승인, 신규 business write admission fence와 active write transaction drain;
- ACTIVE epoch를
FOR SHARE로 잡고 있던 writer가 0이고 writes가 frozen임을 DB session/epoch evidence로 확인; - current CDC epoch의 마지막 transaction으로 controlled rollback-boundary sentinel INSERT;
- connector가 sentinel을 Kafka에 emit하고 sentinel transaction end LSN까지의 source offset과 slot high-water coverage가 persisted됐는지 확인;
- connector pause/stop, exact offset/LSN 기록, production Write ACL revoke;
- polling producer/topic/security를 attest하되 scheduler는 아직 정지;
- DB transaction에서 새 POLLING_V2 epoch ACTIVE 전환과 polling sentinel event+delivery INSERT;
- polling producer Write ACL grant와 scheduler start;
- polling sentinel
DELIVERY_RECORDED와 Kafka record 확인; - target application expected epoch 확인 뒤 regular writes 재개;
- suspected CDC gap만 immutable event에서 audited delivery generation으로 backfill;
- duplicate/inbox/reconciliation 확인.
sentinel event ID만 보였다는 이유로 connector를 멈추지 않는다. writes-frozen boundary, sentinel transaction end LSN, persisted Connect source offset와 slot state가 같은 high-water를 가리켜야 한다. maintenance 전에 시작한 transaction이 sentinel 뒤 commit할 수 있는 상태에서는 step 3으로 진행하지 않는다.
CDC-era event 전체를 무조건 polling delivery로 backfill하지 않는다. 이미 emit된 event를 대량 duplicate할 수 있기 때문이다. 범위와 certainty를 계산한 audited backfill만 허용한다.
27.6 Rollback
cutover 실패 시:
- 어느 authority가 마지막으로 production Write를 가졌는지;
- 마지막 confirmed event ID/WAL/offset;
- indeterminate range;
- consumer inbox coverage;
- duplicate-safe replay 범위;
- slot/offset 보존 여부
를 먼저 판정한다. “둘 다 켜서 빨리 복구”는 허용하지 않는다.
27.7 Automatic failover 금지
CDC connector health DOWN을 감지해 application이 polling을 자동 활성화하지 않는다. external
connector와 in-process scheduler 사이 split-brain을 만들기 때문이다. bounded backlog가 source
DB에 남고 required WAL이 finite retention bound 안에 있는 동안 operator-run cutover/rollback을
수행한다. WAL loss가 생기면 automatic failover가 아니라 DATA_GAP_SUSPECTED reconciliation이다.
28. Test strategy
28.1 원칙
messaging readiness는 fake 하나나 happy-path broker 하나로 증명하지 않는다.
pure contract/property
+ application state machine
+ real PostgreSQL
+ real Kafka
+ security/topology
+ fault/crash/lifecycle
+ compatibility
= exact card evidence
unit test는 빠른 feedback이고 real-service test는 실제 guarantee evidence다. 둘 중 하나가 다른 하나를 대체하지 않는다.
28.2 Current characterization
첫 변경 전에 현재 behavior를 고정한다.
- broker blank -> disabled sentinel;
- broker selected + sender 없음 -> startup failure;
- broker ID mismatch -> startup failure;
- current sender normal return -> current PUBLISHED transition;
- current exception -> FAILED/DEAD;
- ACK-to-mark failure -> IN_FLIGHT/reclaim duplicate possibility;
- same-transaction append rollback;
- current timestamp FIFO;
- current runbook/config drift 목록.
characterization은 current behavior를 정당화하는 것이 아니라 migration 중 accidental loss를 방지한다.
28.3 Contract catalog unit/property
- duplicate contract/destination/schema IDs;
- unknown provider/card;
- maximum bound intersection;
- ordering-required + null key;
- deterministic partition key golden vectors;
- tenant/no-tenant scope;
- aggregate sequence/eventIndex uniqueness;
- tenant ACTIVE/DISABLED non-null scope uniqueness;
- stable catalog/schema/settings digest;
- config cannot relax code maximum;
- legacy + target conflict;
- disabled resource-0 descriptor;
- unsupported card rejection.
28.4 JSON Schema v1
최소:
- Draft 2020-12 meta-schema validation;
- immutable
$id와 checksum; - no remote
$ref; - offline meta-schema/vocabulary registry, duplicate
$id, cyclic$ref, pathological regex; - valid/invalid envelope golden corpus;
- contract별 valid/invalid payload;
- required/null/missing;
- unknown property;
- duplicate JSON key;
- invalid UTF-8/unpaired surrogate;
- depth/string/array/object/number bound;
- timestamp/format assertion;
- exact UTF-8 bytes;
- envelope/header mismatch;
- N/N-1 rolling vectors와 replay horizon 전체 version vectors;
- retired/future payload version;
- parser CPU/memory/time bound.
official JSON Schema Test Suite 또는 Bowtie 호환성 증거를 adopted validator에 대해 추가한다. 그 결과가 모든 custom contract compatibility를 대신한다고 주장하지 않는다.
28.5 Application-core
- business + event append same transaction;
- validation failure rolls back business write;
- publication outcome exhaustive mapping;
- ACKNOWLEDGED/ACKNOWLEDGED_MISMATCH/REJECTED/INDETERMINATE;
- acceptance certainty와 retry disposition의 독립 mapping;
- permanent failure no repeated retry;
- combined attempt/elapsed budget;
- report only after persisted transition;
- reporter failure containment;
- exhausted ordered head blocking;
- audited requeue/hold/skip/compensate policy와 authorization;
- same-event requeue horizon cutoff;
- operator endpoint unauthenticated/forbidden/destructive-permission, idempotency와 stale generation/row-version rejection;
- late-completion bounded source/drain race, duplicate drain과 persistence failure;
- no provider/SDK type in contract;
- disabled active-contract failure.
future consumer:
- APPLIED/DUPLICATE/RETRY/REJECTED mapping;
- inbox same transaction;
- payload collision;
- follow-up outbox same transaction;
- external side effect prohibited/defaulted to outbox.
28.6 PostgreSQL polling integration
real PostgreSQL Testcontainers lane:
- forward-only migration from V3;
- business + event + delivery commit/rollback;
- same transaction resource identity mismatch startup rejection;
- immutable event update rejection after cutover;
- exact
BYTEAround trip와 envelope hash; - delivery FK/unique/order constraints;
- tenant scope nullable-dedupe 공격;
- two or more workers disjoint claim;
- same aggregate same timestamp with sequence/eventIndex total order;
- different aggregates parallel progress;
- stale token cannot ACK/FAIL/DEAD;
- expired-but-not-yet-reclaimed token cannot record delivery;
- lease expiry during send;
- same-token renew;
- database time authority;
- initial automatic publication age starts at event DB created_at;
- requeue generation deadline is
min(generation DB created_at + maximum age, event created_at + same-event requeue horizon); - ACK-to-mark crash;
- append-only attempt admission/outcome/late-ACK journal;
- late observation DB commit before source ACK, commit-to-ACK crash duplicate absorption;
- late observation queue overflow/drop does not mutate delivery state;
- claim count와 publication attempt count;
- mark transaction failure;
- exhausted head/fairness/hot aggregate;
- one CURRENT generation partial-unique constraint;
- concurrent requeue generation race와 atomic authority handoff rollback;
- EXHAUSTED/HOLD supersede 뒤 old generation이 다시 claim되지 않음;
- legacy/v2 relay authority epoch and no dual claim/send;
- legacy PUBLISHED -> unverified migration only;
- reaper vs claim/requeue;
- retention partition/FK/audit;
- DB pool/batch capacity.
Docker unavailable이면 required R2 lane는 skip이 아니라 failure다.
28.7 Real Kafka producer integration
real Kafka broker에서:
- successful send returns actual topic/partition/offset metadata;
- metadata destination mismatch fatal incident;
acks=alleffective config;- idempotence conflict startup failure;
- stable key/partition;
- record/header too large;
- missing topic with auto-create disabled;
- unauthorized Write/Describe;
- broker unavailable before send;
- leader move/retriable error;
- response loss/deadline/late ACK;
- local buffer saturation/max-block;
- broker throttle;
- retry ordering;
- old/new producer generation barrier와 forced indeterminate order downgrade;
- fatal/defunct producer generation recreation;
- no per-message flush;
- producer generation rotation;
- graceful drain/forced close;
- unresolved shutdown -> indeterminate;
- duplicate after ACK-to-DB gap;
- low-cardinality metric/trace/log.
single-node Testcontainers Kafka는 HA/min ISR/leader-failover 전체 증거가 아니다. R2 topology profile에 필요한 multi-broker scenario는 별도 lane에서 수행한다.
28.8 Security integration
- trusted TLS success;
- untrusted CA failure;
- hostname mismatch failure;
- expired/not-yet-valid certificate;
- SASL valid/invalid credential;
- least-privilege producer ACL;
- denied Create/Delete/Alter;
- secret absent/expired;
- credential rotation old/new generation;
- production plaintext startup rejection;
- config/log/descriptor secret redaction.
28.9 Topic conformance
- expected partition/RF/min ISR;
- cleanup/retention/max bytes drift;
- topic missing;
- partition expansion mismatch;
- auto-create disabled;
- wrong cluster/topic binding;
- provisioning evidence/card fingerprint;
- readiness recovery after operator correction.
28.10 Fault/crash matrix
process kill/fault injection 지점:
after DB event commit
after claim commit
before Kafka send
after request write
after broker append before ACK receipt
after ACK before DB transition
during DB transition commit
after DB success before scheduler result
during shutdown
각 시나리오는 DB state, broker observed event IDs, duplicates, late callback, claim token과 recovery를 검증한다.
28.11 Capacity and soak
- sustained throughput와 backlog drain;
- hot aggregate;
- many bounded contracts/destinations;
- producer buffer/memory/GC;
- DB claim query/index;
- poll batch vs claim lease;
- broker throttle/outage/recovery storm;
- retry amplification;
- shutdown under load;
- metric cardinality;
- long-running secret generation rotation.
benchmark 숫자는 repository/device 일반 성능 주장으로 사용하지 않고 selected deployment capacity evidence로 기록한다.
28.12 Future consumer integration
inbound card가 구현될 때:
- auto commit disabled/manual immediate ACK;
- max.poll.records=1 synchronous listener-thread ACK, wrong-thread ACK rejection;
ackAfterHandle=false,commitRecovered=false,resetStateOnRecoveryFailure=false, default logging recoverer absent;- unexpected exception retry exhaustion -> durable HOLD + seek/pause + source commit 0;
- recoverer/HOLD persistence failure -> bounded container stop/readiness DOWN, no retry-cycle reset;
- explicit sync commit timeout/failure/redelivery;
- commit 후 ACK;
- commit 뒤 ACK 전 crash duplicate;
- same event duplicate no side effect;
- two consumers concurrent
ON CONFLICT DO NOTHING RETURNING; - tenant scope/effect generation uniqueness;
- same event different hash quarantine;
- malformed bytes/schema/unknown version;
- pause/resume under bounded saturation;
- max.poll interval;
- rebalance revoke/assign while active;
- partition ordering;
- short retry;
- DLT exact ACK-aware provider effective config/metadata mismatch/rotation/shutdown;
- DLT ACK success/failure/indeterminate;
- DLT ACK 뒤 source commit crash의 duplicate DLT identity;
- strict ordered durable quarantine HOLD, restart/reassignment reapply와 approved gap CAS;
- header sanitization/growth;
- inbox retention/replay;
- live/replay overlap, effect-generation lease와 cleanup race;
- offset out of range;
- graceful shutdown;
- TLS/SASL/ACL;
- real consumer lag/metrics/trace.
28.13 Future CDC integration
exact PostgreSQL + Kafka + Connect + Debezium image set:
- insert-only event mapping;
table.op.invalid.behavior=fatalUPDATE stop과skipped.operations=t/DB-audit TRUNCATE policy;- exact CDC-authority row-filtered INSERT-only PostgreSQL publication catalog;
- polling-authority row 0 emission, approved retention DELETE output 0;
- polling/CDC key/envelope golden parity;
partition_key_text/StringConverter key parity,BYTEA/Debezium BinaryDataConverter exact value bytes, canonicalidheader와 alias 0;- pinned SimpleHeaderConverter와 polling/CDC exact header-byte parity;
- non-outbox heartbeat/schema/tombstone predicate;
- unknown route + auto-create disabled + ACL failure;
errors.tolerance=nonepoison/converter failure;- source producer idempotence/security/effective config;
- connector task crash/restart duplicate;
- Connect internal offset commit failure after Kafka append;
- Connect offset commit/replay;
- slot missing/drop/recreate;
- WAL retained/disk threshold;
- heartbeat low traffic;
- snapshot/no-data mode;
- existing history not republished by cutover
no_data; - schema migration during connector lifecycle;
- partitioned table mapping;
- PostgreSQL 16 failover -> NOT_READY/DATA_GAP_SUSPECTED, no no-gap claim;
- reverse commit-order aggregate sequence incompatibility;
- shadow event ID/count/hash;
- cutover sentinel;
- cutover crash/abort after every numbered step;
- polling/CDC Write authority exclusivity;
- publication epoch stale-writer/old-binary fence;
- rollback and audited backfill;
- retention checkpoint proof.
28.14 Compatibility
matrix:
old producer -> new consumer
new producer -> old live consumer
old polling binary -> new DB schema
new polling binary -> compatibility DB schema
old/new Kafka client -> selected broker
old/new connector -> selected PostgreSQL/Kafka
rolling credential/schema/catalog generation
unsupported combination을 명시하고 자동 fallback하지 않는다.
28.15 Observability contract
test는 metric이 “존재한다”뿐 아니라:
- logical vs physical attempt 구분;
- outcome/certainty 정확성;
- ACK latency boundary;
- event/tenant/key/payload tag 부재;
- catalog cardinality enforcement;
- one canonical ERROR;
- trace context validation;
- readiness role 분리;
- stale health/hysteresis;
- sanitized capability descriptor
를 검증한다.
29. Gradle, CI, dependency와 supply chain
29.1 Dependency ownership
| Dependency | Owner |
|---|---|
| Spring Kafka/Kafka producer client | adapter:outbound:messaging |
| JSON/schema validator runtime | outbound messaging; future inbound leaf도 자기 runtime 소유 |
| envelope schema resource/vocabulary | shared-contract when truly generic |
| PostgreSQL/JPA/Flyway | adapter:outbound:persistence-jpa |
| future Kafka listener client | adapter:inbound:messaging-kafka |
| Kafka Connect/Debezium | deployment/integration-test asset |
| Testcontainers Kafka | provider/integration test configuration |
| Testcontainers PostgreSQL | persistence/bootstrap integration test |
| Micrometer/Spring observation composition | adapter/bootstrap ownership에 맞춤 |
application/domain은 Kafka/Jackson/schema validator에 의존하지 않는다.
29.2 Spring Boot BOM
Spring Boot 4.0.0 BOM이 관리하는 Spring Kafka/Kafka client 조합을 first implementation candidate로 사용한다. 실제 resolved version은 dependency lock과 evidence fingerprint로 기록한다.
direct version override는:
- Boot/Spring Kafka compatibility;
- Kafka broker protocol compatibility;
- CVE/license;
- tests/locks/SBOM
을 함께 통과할 때만 허용한다.
29.3 Architecture change
producer/polling phase는 현재 19 leaf를 유지한다.
first R2 operator surface는 existing
adapter-inbound-web -> application-core와
adapter-outbound-persistence-jpa -> application-core edges를 composition root에서 조립한다.
web leaf가 persistence leaf/repository/entity에 직접 의존하지 않으므로 새 project edge가 없다.
standalone sample이 real messaging example을 실행하는 phase에는 leaf 수를 늘리지 않고
sample-portfolio의 allowed dependency/runtime dependency에
adapter-outbound-messaging edge만 추가한다. provider qualification 자체는 adapter
test-source contract로 가능하므로 이 sample edge가 production R2의 필수 전제는 아니다.
consumer phase는:
- registry에 20번째 leaf;
- settings include/mapping;
- app-bootstrap dependency;
- architecture tests/fixtures;
- module lockfile;
- documentation
을 한 migration으로 추가한다. inbound leaf에서 outbound/persistence adapter로 edge를 만들지 않는다.
29.4 Proposed tasks
구현 계획은 repository task naming convention을 확인한 뒤 정확한 이름을 확정한다. 목표 lane:
:adapter:outbound:messaging:test
:application-core:test
:adapter:outbound:persistence-jpa:test
:app-bootstrap:test
verifyMessagingContracts
verifyMessagingJsonSchemaV1
verifyMessagingPollingOutboxR2
verifyMessagingKafkaProducerR2
verifyMessagingSecurityR2
verifyMessagingReleaseProfile
future:
verifyMessagingConsumerR2
verifyMessagingCdcR2
기존 공통 gate:
test
check
verifyCleanArchitectureDependencies
verifyEnvKeys
verifyPublicPathSnapshot
dependency lock/SBOM/vulnerability/license gates
29.5 CI lanes
PR blocking:
- pure/unit/property;
- architecture/dependency/env/schema registry contract;
- JSON schema/golden/N/N-1;
- real PostgreSQL polling baseline;
- pinned real Kafka producer baseline;
- docs/config/runbook drift checks.
production-readiness:
- TLS/SASL/ACL;
- topic conformance;
- selected RF/min ISR의 multi-broker leader loss, below-min-ISR rejection과 recovery;
- multi-worker stale-token/crash matrix;
- response-loss/Toxiproxy;
- bounded shutdown/rotation;
- metrics/readiness artifact.
nightly/R3:
- prolonged leader churn과 repeated multi-broker failure;
- rolling broker/client upgrade;
- partition migration;
- soak/capacity;
- future consumer rebalance storm;
- future CDC restart/slot/failover.
29.6 No silent skip
developer local test는 Docker 없음 등을 명시적으로 보고할 수 있다. 그러나 selected R2 release gate는:
- required service unavailable;
- image pull failure;
- test skipped by assumption;
- credential fixture missing;
- no matching test;
- stale evidence
를 PASS로 변환하지 않는다.
29.7 Evidence artifact
sanitized artifact:
git commit/source digest supplied by human/CI
commands and timestamps
task/test counts
broker/client/Spring/PostgreSQL/Connect/Debezium versions
image digests
effective non-secret settings
topology/security profile
schema/catalog/settings hashes
fault scenarios/results
readiness/card status
skips/failures
unsupported claims
runbook links
CI artifact 없이 문서 표를 수동으로 R2로 바꾸지 않는다.
29.8 Supply chain
- Gradle dependency locks;
- SBOM;
- vulnerability severity/suppression policy;
- license allowlist;
- Kafka/Connect/Debezium container digest pin;
- connector plugin inventory/classloader compatibility;
- JSON schema validator dependency review;
- image signature/provenance where platform supports;
- upgrade cadence/runbook.
30. Implementation and migration sequence
30.1 Phase 0 — Truth and characterization
목표:
- 이 설계 검토/승인;
- current ACK overclaim 표시;
- fake-only/real-service status 정리;
- current tests characterization;
- README/YAML/env/runbook drift 목록;
- implementation plan 작성.
완료 기준:
- §0 status ledger 갱신;
- current code behavior test;
- no implementation/R2 overclaim;
- human-only git policy 유지.
30.2 Phase 1 — Contract, catalog, envelope and JSON Schema
추가:
- application event metadata/value types;
- application-owned contract contribution SPI와 합법적인 sample/test composition;
- logical destination/contract descriptors;
- validated integration-event port;
- envelope v1/payload schema resources;
- schema/catalog checksum manifest;
- JSON codec/validator;
- sample payload contract/golden vectors;
- shared-contract/messaging leaf
CLAUDE.mdownership update; - compatibility/no-remote-ref/resource-bound tests;
- current raw payload/eventType migration adapter.
이 phase만으로 Kafka/polling R2가 아니다.
30.3 Phase 2 — Immutable event and polling delivery v2
기존 V3 migration을 수정하지 않고 새 forward-only migration을 추가한다.
base template의 기본 migration card는
ADDITIVE_IN_PLACE_EMPTY_OR_DRAINED_V3.v1이다. production data가 없거나 verified drain/reset
maintenance로 legacy row가 0인 새 템플릿 출발점을 대상으로 한다. 이 기본 card에서 P2의
rolling-safe 고수준 순서:
- row/data classification preflight가 empty-or-drained 조건을 fail-closed로 확인;
- existing
outbox_event에 immutable v2 metadata를 additive하게 추가하고 legacy columns는 compatibility를 위해 유지; outbox_delivery, append-only attempt journal, disposition audit,outbox_publication_epoch생성;- known legacy event type alias/catalog와 empty/drained evidence를 기록;
- compatibility release가 legacy required columns와 v2 immutable metadata를 채우되
LEGACY_POLLINGcontrol만 사용; retained V3 compatibility projection은event_type=contract_id,payload=exact v1 envelope UTF-8 text,status=PENDING,attempt_count=0,next_attempt_at=occurred_at,idempotency_key=event_id로 고정한다. canonical metadata가 있는 row의 legacy publisher는OutboxEnvelopeJson으로 다시 감싸지 않는다. compiled logical-destination binding, stored partition-key bytes와 immutableenvelope_bytes를 byte-for-byte passthrough한다. canonical metadata가 없는 true legacy row만 기존 v0 wrapper를 사용한다; - compatibility binary의 legacy relay query가 ACTIVE
LEGACY_POLLINGepoch/generation만 처리하도록 fence; - v2 constraint/index/trigger, one-CURRENT authority와 migration rollback을 rehearsal하되 active legacy status mutation을 아직 막지 않음;
- old pre-fence binary를 0으로 만들고 no-dual-writer probe 확인;
LEGACY_POLLINGepoch와 legacy relay를 계속 ACTIVE로 유지;- P3 ACK-aware producer/new relay/operator tool이 scheduler-disabled 상태로 배포·attest되기 전
POLLING_V2authority 전환과 v2 claim을 금지.
P2는 schema/control-plane candidate일 뿐 publish path cutover가 아니다.
live non-empty deployment는 이 base card를 통과시켜 자동 backfill하지 않는다. row volume, active state distribution, lock/replication budget, data classification, maintenance window와 rollback rehearsal을 입력으로 다음 중 하나를 별도 deployment migration design과 승인 gate에서 고른다.
LIVE_ADDITIVE_BACKFILL_IN_PLACE.v1
COPY_AND_CUTOVER_WITH_RECONCILIATION.v1
두 live card 모두 이 문서의 stable event identity, legacy ACK non-overclaim, relay authority fence, same-transaction/no-dual-write invariant를 따라야 하지만 어느 것이 안전한지는 실제 database evidence 없이 base template가 추측하지 않는다. 따라서 이 선택은 구현 계획에 숨겨 둔 선택이 아니라 명시적으로 차단된 deployment-specific gate다.
모든 migration card는:
- dual-write gap 없음;
- same transaction;
- rollback;
- old/new binary compatibility;
- relay authority가 어떤 순간에도 정확히 하나;
- legacy/v2 query generation과 backlog handoff watermark;
- no event identity change;
- backup/rollback
을 증명한다.
30.4 Legacy event migration
현재 pending row는 hand-written envelope/eventType/raw payload다.
base template default:
- verified empty/drained V3이면 legacy payload transform 없이 additive migration;
- drain/reset은 production data 삭제를 뜻하지 않으며 non-production 또는 승인된 maintenance 범위만 대상으로 evidence를 남김.
live-data deployment gate:
- live data면
legacy-envelope-v0read-only publisher; - known contract는 승인된 live migration card에서만 v1로 deterministic transform하되 original hash/audit 보존;
- unknown contract는 자동 topic publish하지 않고 operator quarantine.
이미 발행된 event의 wire contract를 몰래 바꾸지 않는다.
현재 PUBLISHED는 broker ACK가 아니라 void KafkaSender 정상 반환이다. migration은
ackObservedAt, provider metadata 또는 DELIVERY_RECORDED를 조작해 채우지 않는다. 운영자는
event 범위별로:
- downstream reconciliation 뒤 legacy terminal로 수용;
- duplicate-aware requeue;
- quarantine/hold
중 하나를 audit한다. legacy IN_FLIGHT는 old relay fence와 drain/expiry 전 변환하지 않는다.
final fenced reconciliation의 기본 상태 매핑은 다음과 같다.
legacy PENDING -> READY
legacy FAILED -> RETRY_WAIT (preserved finite due/budget, 없으면 reviewed DB-time due)
legacy DEAD -> EXHAUSTED (acceptance certainty를 fabricated definite rejection으로 만들지 않음)
legacy PUBLISHED -> LEGACY_RECORDED_UNVERIFIED
legacy IN_FLIGHT -> HOLD + remaining-indeterminate audit
이미 provisional delivery가 있으면 final legacy state와 expected row version을 대조해 같은 current generation을 migration-only CAS로 맞추고, 없으면 정확히 하나를 insert한다. duplicate CURRENT, unknown status/contract, event/hash mismatch는 자동 추측하지 않고 cutover transaction을 rollback한다. 이 매핑은 migration에서만 허용되며 정상 v2 worker state machine을 우회하는 일반 운영 API가 아니다.
30.5 Phase 3 — Spring Kafka producer and polling reference path
구현:
- Spring Kafka dependency/lock;
- explicit producer factory/template;
- typed provider settings/compiler;
- ACK-aware gateway/outcome;
- topic attestation;
- finite producer/admission/deadline;
- polling relay outcome/state update;
- late-completion source/drain/attempt observation;
- application disposition use case + authenticated inbound-web operator endpoint;
- producer generation/lifecycle;
- best-effort migration;
- config/env/README/runbook update;
- real Kafka/PostgreSQL happy/failure tests.
P3 cutover는 다음 순서를 고정한다.
- ACK-aware producer와 v2 relay를 scheduler-disabled로 배포;
- contract/catalog/schema, producer/topic/security와 transaction-resource preflight;
- operator disposition endpoint의 auth/CAS/audit negative test;
- bounded write maintenance를 시작해 신규 business transaction admission을 막고 active writer와
ACTIVE epoch
FOR SHAREholder가 0이 될 때까지 drain; - old legacy relay 신규 claim 중단, IN_FLIGHT maximum budget drain, remaining indeterminate audit;
- old producer close/Write fence와 DB legacy mutation guard 활성;
- 하나의 cutover DB transaction을 열어 ACTIVE legacy epoch를
FOR UPDATE로 잠그고, writes와 legacy mutation이 fenced된 snapshot에서 fixed legacy handoff watermark를 기록; - 같은 transaction에서 pre-backfill 뒤 watermark까지 생긴 delta를 포함해 모든 legacy event를
final reconcile한다. 각 row는 final legacy state에 대응하는 정확히 한 CURRENT v2 delivery
(
READY/RETRY_WAIT/EXHAUSTED/HOLD/LEGACY_RECORDED_UNVERIFIED)를 가지며 active claim은 0이어야 한다. row count, event ID/hash manifest와 unmapped/duplicate count 0을 assertion; - 같은 transaction에서
LEGACY_POLLING -> POLLING_V2ACTIVE epoch 전환과 v2 cutover sentinel event+delivery INSERT 뒤 commit. reconciliation/manifest/switch 중 하나라도 실패하면 전체 rollback; - v2 relay만 start하고 claim token/sequence/valid-lease CAS 사용;
- canonical append path가 만든 v2 sentinel ACK/
DELIVERY_RECORDED, legacy writer 0, no-dual-send와 fresh readiness를 확인한 뒤 expected FROZEN generation/epoch/evidence CAS로 write admission을OPEN(generation+1)하고 business writes 재개; - rollback window 뒤 obsolete legacy columns 제거는 별도 later forward migration.
write maintenance는 process-local boolean이 아니다. 모든 TransactionPort.inWrite는 같은
transaction에서 PostgreSQL admission singleton을 FOR KEY SHARE로 잡고 OPEN/fence generation을
확인한다. freeze transaction은 그 row를 FOR UPDATE로 잡아 기존 share holder가 commit/rollback할
때까지 기다린 뒤 FROZEN generation을 durable하게 기록한다. 새 writer는 그 뒤 fail-closed
rollback한다. live node lease와 deployment instance inventory가 expected source/fence protocol로
일치하지 않거나 legacy relay/reaper/producer Write의 zero-active/negative-Write probe가 없으면
one-shot precondition evidence를 만들지 않는다.
epoch commit 뒤에도 admission은 자동으로 열리지 않는다. v2 sentinel이
DELIVERY_RECORDED이고, frozen 상태에서 maintenance-only canonical append canary가 exact
envelope/projection/delivery를 만들고, topic/security/readiness가 fresh한 경우에만 application
resume use case가 expected FROZEN generation + POLLING_V2 epoch + evidence digest CAS로
OPEN(generation+1)을 기록한다. ordinary TransactionPort.inWrite는 그 전까지 계속 거부된다.
resume mismatch/replay/failure는 FROZEN을 유지한다. epoch commit 뒤 runner가 실패하면 별도
resume-polling-v2-writes one-shot recovery operation만 같은 증거를 다시 검증할 수 있고 raw SQL
status update는 금지한다.
cutover는 authenticated public/web endpoint가 아니라 non-web one-shot maintenance runner가
opaque approval evidence ID와 expected target/source/epoch를 받아 수행한다. operation ID와
evidence consumption은 DB에서 재실행을 막고, 실패는 non-zero exit와 immutable audit를 남긴다.
epoch commit 전 실패는 DB transaction rollback만으로 끝내지 않는다. 먼저 모든 fence를 유지한
채 fresh evidence/approval deadline 안에서 bounded forward retry할 수 있다. abort-to-legacy를
선택하면 exact epoch가 여전히 LEGACY_POLLING이고 v2 business send/sentinel authority가 없으며
inventory가 보존됐음을 확인한다. 외부 ACL을 바꾸기 전에 DB에서 exact cutover attempt를 잠그고
CUTOVER_PENDING -> RECOVERING_LEGACY로 CAS하면서 recovery operation/lease/evidence digest를
결합한다. 이 전이는 같은 attempt를 모든 forward finalizer에서 원자적으로 무효화하며, lease가
만료돼도 CUTOVER_PENDING으로 돌아가지 않고 recovery-only takeover만 허용한다. 반대로
finalizer는 epoch transaction 안에서 CUTOVER_PENDING -> FINALIZING_V2를 먼저 CAS하고
성공 commit에서만 CONSUMED_V2로 바꾼다. 따라서 두 분기는 동일 attempt에서 함께 진행될 수 없다.
상호 배제 범위는 attempt 하나가 아니라 outbox authority 전체다. 모든 attempt는 상수
OUTBOX_PUBLICATION authority scope, ACTIVE legacy epoch, FROZEN fence generation과 target binding을
저장하고, partial unique constraint는 이 scope에 nonterminal
CUTOVER_PENDING|FINALIZING_V2|RECOVERING_LEGACY row를 정확히 하나만 허용한다. evidence 생성,
finalization, recovery prepare/completion은 모두 write-admission singleton FOR UPDATE 뒤 ACTIVE
epoch FOR UPDATE의 동일 lock order를 사용하고 exact generation/target/sole-attempt를 검증한다.
따라서 recovery 중 별도 attempt를 만들어 v2 epoch를 commit할 수 없다. 같은 attempt뿐 아니라
서로 다른 attempt의 생성/finalization과 recovery 사이 양방향 경쟁도 real PostgreSQL 시험으로
고정한다.
recovery claim이 commit된 뒤에만 외부 provisioning이 legacy principal Write를 재부여하고 fresh
positive probe를 만든다. 그 다음 exact attempt/owner/lease와 ACTIVE
LEGACY_POLLING epoch를 다시 잠가 검증한다. 불일치하면 legacy Write를 즉시 다시 revoke하고 fresh
negative probe를 만든 뒤 business writes를 FROZEN으로 유지한다. 검증이 성공하면 immutable
recovery/ACL audit 아래 in-process legacy Write fence, reaper, relay를 generation-CAS로 다시 열고
마지막에 write admission을 OPEN(generation+1)로 CAS하면서 attempt를
RECOVERED_LEGACY로 끝낸다. 어느 단계든 실패하면 business writes는 FROZEN을 유지하고 이미 연
legacy component를 다시 fence하거나 safe degraded state로 둔 채 recovery-only 재시도한다.
forward-finalization 대 recovery-claim의 양방향 lock-order 경쟁과 각 외부 mutation/crash
경계를 real PostgreSQL 계약 시험으로 고정한다. raw SQL이나 단순 boolean toggle은 금지한다.
epoch commit 뒤에는 business v2 send 여부와 무관하게 reverse epoch/legacy reactivation을 first R2에서 지원하지 않는다. admission을 닫고 backlog/schema/epoch/audit를 보존한 채 forward-fix한다.
qualification fixture의 broker/principal은 qualificationEnvironmentIdentity, 실제 target은
deploymentBindingIdentity로 별도 기록한다. 공통 비교 대상은 capability/profile, supported
broker/client version constraint, settings/catalog/schema와 scenario contract다. 실제 cutover는
target cluster/topic/principal/secret generation에 대한 fresh topology/security/ACL attestation과
legacy principal negative-Write probe를 별도 요구하며 fixture identity를 target evidence로
재사용하지 않는다.
large live-data card는 bulk pre-backfill을 별도 bounded batch로 수행할 수 있지만, step 7
transaction 안의 fixed watermark, final delta, manifest assertion과 authority switch는 한 fenced
atomic unit에 남긴다. 그 transaction의 lock/statement/replication budget이 evidence로 안전하지
않으면 COPY_AND_CUTOVER_WITH_RECONCILIATION.v1 또는 더 긴 maintenance를 다시 승인하며 부분
authority switch를 허용하지 않는다.
이 phase 끝은 R1/R2 candidate다. security/fault/release evidence 없이 R2 완료가 아니다.
30.6 Phase 4 — R2 qualification and rollout
- production TLS/SASL_SSL;
- least-privilege ACL;
- topic topology policy;
- selected RF/min ISR multi-broker failure/recovery;
- secret/certificate rotation;
- fault/crash/late ACK/stale token;
- multi-worker/capacity/shutdown;
- observability/readiness/card registry;
- no-skip CI;
- sanitized evidence artifact;
- canary/rollback rehearsal;
- runbook completion.
exact first tuple만 R2로 승격한다.
30.7 Phase 5 — Inbound Kafka and inbox
- registry 19 -> 20 migration;
- inbound listener leaf;
- consumer contract/catalog decoder;
- manual ACK/bounded processing/rebalance;
- application
MessageConsumptionExecutor; - PostgreSQL inbox;
- DLT/replay tooling;
- security/observability;
- real Kafka/PostgreSQL consumer evidence.
producer/polling R2와 별도 card/status다.
30.8 Phase 6 — CDC
- insert-only source enforcement;
- deployment connector/slot/publication/internal topics;
- exact Debezium mapping;
- shadow topic;
- fault/WAL/offset/failover;
- cutover/rollback rehearsal;
- CDC retention proof;
- separate readiness/evidence.
polling R2를 제거하지 않는다. deployment가 둘 중 하나를 선택하되 같은 production destination에서 동시 활성화하지 않는다.
30.9 Phase 7 — Optional cards
실제 요구와 evidence가 있을 때:
- retry topic;
- Kafka EOS;
- schema registry;
- Avro/Protobuf;
- compaction;
- claim-check/large message;
- multi-cluster;
- alternate provider;
- module/artifact split.
30.10 Rollout
first polling rollout:
- selected migration card preflight와 forward schema migration;
- compatibility append + legacy relay-fence binary;
- contract/catalog/schema artifact;
- provisional delivery backfill/legacy unverified reconciliation;
- ACK-aware producer + v2 relay의 disabled canary startup;
- producer/topic/security/transaction-resource/operator-tool preflight;
- business write admission freeze + active writer drain;
- old relay stop, IN_FLIGHT drain, producer Write/DB legacy mutation fence;
- 한 fenced DB transaction에서 ACTIVE epoch lock + fixed handoff watermark + final delta reconciliation + count/hash manifest assertion;
- 같은 transaction에서 atomic
LEGACY_POLLING -> POLLING_V2authority switch + v2 sentinel insert 뒤 commit; - v2 relay start와 sentinel ACK/
DELIVERY_RECORDED, legacy writer 0 확인 뒤 writes 재개; - bounded subset/destination enable;
- backlog/duplicate/indeterminate/no-dual-authority observation;
- full enable;
- rollback window 뒤 legacy seam/config removal.
rollback은 DB schema를 destructive downgrade하지 않는다. authority epoch, producer Write ACL과 relay fence를 먼저 판정하고 new relay disable/backlog preservation을 사용한다. 이미 v2 record가 Kafka에 갈 수 있는 시점부터 old/new relay를 동시에 켜는 rollback은 금지한다.
30.11 Status update discipline
각 phase가 끝날 때 §0에:
- implemented capability;
- not implemented;
- evidence/card level;
- exact test command/result;
- known limitation;
- next optional candidates
를 갱신한다. 본문 설계 문장을 “구현됨”으로 다시 쓰지 않는다.
31. Completion criteria and extension ledger
31.1 Design complete
설계 완료 조건:
- 사용자 review/approval;
- architecture/module ownership 확정;
- first tuple 확정;
- producer/outbox/consumer/CDC guarantee와 non-guarantee 명시;
- base template migration card 확정과 live-data deployment approval gate 명시;
- migration/test/runbook/evidence 계획;
- unresolved decision이 implementation plan에 숨지 않음;
- LLM Wiki capture.
현재 문서 상태는 사용자 승인에 따라
상세 설계 승인, 실행 계획 작성·독립 검토 완료, 구현 미착수다.
31.2 First R2 implementation complete
정본 완료 판정은 §10.5 machine registry의 exact selected tuple이 모두
release-eligible이고 required scenario/evidence fingerprint가 PASS인 경우뿐이다. 아래는 drift를
막기 위한 human review checklist다.
- raw event type -> topic 제거;
- closed contract/destination catalog;
- envelope/payload schema v1;
- append-before-persist validation;
- immutable event + delivery split;
- aggregate sequence + normal-path key order/non-guarantee;
- claim token + unexpired-lease CAS와 JIT claim;
- ACK-aware Spring Kafka producer;
- ACK/ACK-mismatch/REJECTED/INDETERMINATE와 attempt journal;
- legal late-completion observation drain과 bounded-loss semantics;
- authenticated operator disposition use case/control surface;
- finite effective config;
- combined retry/requeue horizon budget;
- production security/ACL/topic conformance;
- multi-broker RF/min ISR failure evidence;
- disabled resource 0;
- bounded shutdown/rotation;
- real Kafka/PostgreSQL/fault/security tests;
- no-skip release gate;
- descriptor/card/evidence artifact;
- runbook;
- legacy drift/migration;
- Wiki capture.
31.3 Consumer/inbox complete
별도 조건:
- inbound leaf/registry;
- manual ACK after commit;
- inbox + business same transaction;
- duplicate/collision behavior;
- bounded poll/backpressure/rebalance;
- DLT ACK ordering;
- replay audit;
- security/retention;
- real fault evidence.
producer R2만으로 이 조건을 충족했다고 표시하지 않는다.
31.4 CDC complete
별도 조건:
- insert-only source;
- exact connector mapping;
- slot/offset/WAL/security;
- duplicate/restart/failover;
- shadow/cutover/rollback;
- authority exclusivity;
- retention proof;
- real multi-service evidence.
event/delivery table split만으로 CDC ready라고 표시하지 않는다.
31.5 추가 가능한 card
| 추가 요구 | 추가 card/설계 | 안정적으로 유지할 것 |
|---|---|---|
| consumer | inbound Kafka + inbox | event ID, envelope, logical destination |
| CDC | Debezium dispatch | immutable event, key/wire semantics |
| binary schema | Avro/Protobuf registry | contract ID, payload version/outcome |
| delayed retry | retry topic | event ID/inbox, explicit ordering loss |
| Kafka-only workflow | transactional EOS | DB/non-Kafka scope 제외 |
| large payload | object-storage claim check | event identity/schema/security |
| alternate broker | provider card | semantic contract/evidence rule |
| multi-cluster | replication/failover card | no global ordering/exactly-once overclaim |
| topic compaction | contract-specific compacted log | key/tombstone/replay semantics |
outbound messaging leaf 분리는 다음 trigger 전에는 하지 않는다.
- 두 번째 broker provider가 독립 dependency/release lifecycle을 가짐;
- codec/schema runtime을 inbound leaf도 재사용해야 하지만
shared-contractpurity로 수용할 수 없음; - AdminClient/topology attestor가 독립 deployment artifact가 됨;
- producer와 compiler의 dependency/security ownership이 별도 release를 요구함.
trigger가 생기면 registry leaf/edge migration과 동일 semantic card/evidence compatibility를 함께 설계한다. package가 많다는 이유만으로 선제 분리하지 않는다.
31.6 금지하는 완료 표현
다음 표현은 조건 없이 사용하지 않는다.
- “Kafka 지원 완료” — seam인지 real selected card인지 명시;
- “broker ACK” — actual future metadata evidence 필요;
- “exactly once” — boundary를 좁힌 Kafka EOS card 외 금지;
- “중복 안전” — 해당 consumer/inbox evidence 필요;
- “strict FIFO” — sequence/key/topology뿐 아니라 failure/rotation/consumer order gate evidence 필요;
- “CDC ready” — connector/slot/offset/cutover evidence 필요;
- “DLT 처리 완료” — DLT publish와 business remediation 구분;
- “security ready” — TLS/SASL/ACL/rotation negative test 필요;
- “production ready” — exact R2 card/evidence 필요;
- “test passed” — command/result/skip을 기록.
31.7 구현 보고 template
후속 구현 완료 보고는 최소 다음을 포함한다.
이번에 구현된 card/phase
변경 파일
architecture/dependency 변화
실행한 focused/real-service/common gate
test count/result/skip
evidence fingerprint/artifact
아직 미구현인 card
현재 보장과 non-guarantee
runbook
LLM Wiki capture
남은 위험
32. Required runbooks
32.1 First R2 baseline
producer-unavailable-or-unauthorized- DNS/network/TLS/SASL/ACL/topic/min ISR 분기;
- direct/relay/write role별 guarantee 영향;
- safe pause/recovery proof.
outbox-backlog-and-stale-lease- oldest age/count/growth;
- claim owner/token/lease conflict;
- hot/stuck aggregate;
- capacity/scale 한계.
delivery-indeterminate-and-duplicate-burst- ACK loss/late ACK/mark failure;
- suspected event ID range;
- downstream inbox/reconciliation;
- resend duplicate 경고.
schema-poison-or-record-too-large- contract/version/byte diagnosis;
- retry 중단;
- corrected/compensating event;
- payload 원문 log 금지.
terminal-delivery-disposition- hold/requeue/skip/compensate;
- business-owner 승인;
- aggregate 후행 영향;
- audit/rollback.
topic-policy-or-partition-change- topic drift;
- partition expansion order risk;
- new topic migration/cutback.
shutdown-deploy-and-secret-rotation- new claim stop/drain;
- cert/secret expiry;
- old/new generation;
- forced timeout/rollback.
legacy-to-v2-relay-authority-cutover- compatibility binary/fingerprint;
- legacy claim stop/IN_FLIGHT drain;
- DB epoch/legacy mutation fence;
- backlog watermark/no-dual-send probe;
- rollback stop condition.
32.2 Future consumer
- consumer lag/max.poll/rebalance loop;
- retry exhaustion/DLT publish failure;
- DLT duplicate identity/quarantine capacity/strict-order HOLD;
- schema/deserialization poison;
- inbox collision/retention;
- inbox purge vs live/replay effect-generation lease;
- audited replay;
- offset out of range;
- consumer identity/group migration;
- shutdown with active handler.
32.3 Future CDC
- connector/task down;
- WAL growth/disk pressure;
- slot missing/ahead/behind;
- offset reset/alter;
- snapshot restart/schema drift;
- PostgreSQL primary failover;
- polling -> CDC cutover;
- CDC -> polling rollback;
- shadow/production connector와 unique slot ownership;
- cutover 단계별 abort/authority proof;
- PostgreSQL 16 failover DATA_GAP_SUSPECTED recovery;
- WAL emergency without automatic slot drop;
- CDC retention/partition purge.
32.4 Common structure
모든 runbook:
detection/trigger
blast radius
current guarantee degradation
safe first response
diagnosis evidence
non-destructive mitigation
destructive action approval boundary
reconciliation
recovery proof
rollback
audit/post-incident
related metrics/errors/cards
32.5 Existing runbook migration
outbox-publish-failed.md와 outbox-dead-letter.md는 현재 stub다. first implementation에서:
- removed
APP_MESSAGING_KAFKA_ENABLED제거; - 실제 class/settings/table/state 이름;
- producer EXHAUSTED와 consumer DLT 구분;
- “consumer dedupe가 현재 안전 보장” 표현 제거;
- raw SQL status rewrite 제거;
- token/generation/audit operator tool;
- real dashboard/alert link;
- split event/delivery query;
- indeterminate/duplicate 절차
로 갱신한다.
33. Primary references
구현은 실제 BOM/lock에 resolve된 version 문서를 우선한다. 아래는 설계 시 확인한 official primary reference다.
Spring Boot and Spring Kafka
- Spring Boot 4.0 — Apache Kafka Support
- Spring Kafka 4.0 — Sending Messages
- Spring Kafka 4.0 — Message Listener Containers
- Spring Kafka 4.0 — Pausing and Resuming Listener Containers
- Spring Kafka 4.0 — Handling Exceptions
- Spring Kafka 4.0 — DefaultErrorHandler API
- Spring Kafka 4.0 — Retry Topic Pattern
- Spring Kafka 4.0 — Transactions
- Spring Kafka 4.0 — Exactly Once Semantics
- Spring Kafka 4.0 — Monitoring
- Spring Kafka 4.0 — Testing Applications
Apache Kafka
- Kafka 4.1 Producer Configs
- Kafka 4.1 Topic Configs
- Kafka 4.1 Consumer Configs
- Kafka 4.1 Design — Message Delivery Semantics
- Kafka 4.1 Security Overview
- Kafka 4.1 Kafka Connect User Guide
- Kafka 4.1 Kafka Connect Configs
- Kafka 4.1 KafkaProducer API
- Kafka 4.1 UnknownTopicOrPartitionException
- Kafka 4.1 NotEnoughReplicasAfterAppendException
JSON and trace contract
- JSON Schema Draft 2020-12
- JSON Schema Core 2020-12
- JSON Schema Validation 2020-12
- JSON Schema Test Suite
- RFC 8259 — The JavaScript Object Notation Data Interchange Format
- W3C Trace Context
Debezium, Kafka Connect and PostgreSQL CDC
- Debezium 3.6 Release Series
- Debezium 3.6 Outbox Event Router
- Debezium 3.6 PostgreSQL Connector
- PostgreSQL 16 Logical Decoding Concepts
- PostgreSQL 16 Logical Replication Security
- PostgreSQL 16 CREATE PUBLICATION
- PostgreSQL 16 Logical Replication Row Filters
- PostgreSQL 16 INSERT / ON CONFLICT
- PostgreSQL 16 Unique Constraints
Test infrastructure
문서 링크는 executable evidence가 아니다. implementation 시 exact dependency/image version, effective config, tests와 runbook으로 다시 확인한다.
34. Review 이후 다음 단계
base template의 핵심 architecture option과 empty/drained V3 기본 migration card는 확정됐다. 사용자에게 구현 세부 선택을 더 요구하지 않는다. 다만 실제 live non-empty database에 적용하는 시점에는 §30.3의 row/lock/data evidence를 수집한 뒤 live migration card를 별도 승인해야 한다. 그것은 지금 숨겨 둔 선택이 아니라 deployment-specific safety gate다. review에서 방향 수정이 없으면 다음 순서로 진행한다.
- P0–P4 first R2 실행 계획과 독립 review 완료;
superpowers:subagent-driven-development또는superpowers:executing-plans로 계획 실행;- test-first로 contract/schema/polling/producer를 단계 구현;
- exact first tuple R2 evidence 뒤 consumer/inbox 계획;
- consumer evidence 뒤 CDC qualification/cutover 계획.
repository의 human-only commit 정책에 따라 agent는 stage/commit/amend/push하지 않는다.