# ArkLog Architecture

ArkLog is the general distributed append-only log subsystem of Ark Fabric.
BookKeeper is its first durable engine, just as FoundationDB is the first
transaction engine beneath ArkTxn. ArkLog owns the logical log contract;
BookKeeper provides the initial storage engine.

ArkLog provides ordered append, immutable segments, replay, retention,
watermarks, and topology-aware dissemination across independent Ark sites. It
serves ArkTxn mutation replication and any other Ark subsystem that needs a
durable replayable stream.

Replication, placement, repair, and writer failover are specified separately
in [`ark-log-replication.md`](ark-log-replication.md).

## Role And Boundaries

ArkLog owns:

* logical stream and partition identity;
* writer epochs and monotonically ordered positions;
* append durability and immutable segment formation;
* replay cursors, retention watermarks, and gap detection;
* dissemination topology and replica convergence;
* compaction and optional archival of closed segments.

ArkLog does not own ArkTxn conflict detection, object manifests, WAN path
selection, or application-specific consumer groups. ArkTxn chooses which
mutation streams and hot replicas exist. ArkObj stores immutable
archives. ArkNet moves segments across WAN paths.

```mermaid
flowchart TB
    Producers["ArkTxn outbox<br/>and other producers"] --> Gateway["ArkLog append gateway"]
    Gateway --> Catalog["Ark Control Catalog<br/>stream, epoch,<br/>and policy"]
    Gateway --> Active["Site-local BookKeeper<br/>active ordered segments"]
    Active --> Ack["Append acknowledgement"]
    Active --> Scheduler["ArkLog dissemination<br/>scheduler"]
    Scheduler --> Net["ArkNet"]
    Net --> ReplicaA["ArkLog replica A"]
    Net --> ReplicaB["ArkLog replica B"]
    Net --> Consumer["Ordered consumer"]
    Active --> Closed["Closed compacted<br/>segment"]
    Closed -. "optional<br/>long-term archive" .-> Object["ArkObj"]
```

## Logical Model

| Concept | Meaning |
|---|---|
| Stream | Named append namespace with policy and retention. |
| Partition | Independently ordered unit used to scale a stream. |
| Writer epoch | Fencing generation authorizing one writer for a partition. |
| Position | Tuple of stream, partition, writer epoch, and sequence. |
| Record | Immutable producer payload and metadata at one position. |
| Segment | Contiguous immutable range of records stored as a BookKeeper ledger or ledger range. |
| Durable head | Highest contiguous position satisfying minimum durability. |
| Replica watermark | Highest contiguous position received or applied at a site. |
| Retention watermark | Lowest position that all protected consumers and policy pins allow ArkLog to discard. |

A partition has one authoritative writer epoch at a time. Streams scale through
partitions, not concurrent writers racing to assign positions in one partition.
Consumers may receive future segments early, but they apply or expose records
only through a gap-free contiguous watermark.

## Storage Layout

Active ArkLog segments use BookKeeper directly. The foreground append path must
not pass through ArkObj because that would add a manifest layer to
an operation BookKeeper already models directly: fenced, ordered, durable
append.

The storage boundary is:

* the Ark Control Catalog stores stable stream IDs, partition ownership,
  writer epochs, policy roots, and active-segment pointers;
* the BookKeeper metadata service stores physical ledger and ensemble state;
* BookKeeper entries store active records and segments;
* ArkObj may store closed, immutable, compacted segment archives for
  long retention or replica bootstrap.

Physical ledger IDs are not public stream identities. Compaction or archival
may relocate a segment without changing its logical stream positions.

## Append Path

```mermaid
sequenceDiagram
    participant Producer
    participant Edge as ArkLog gateway
    participant Catalog as Ark Control Catalog
    participant BK as BookKeeper sites
    participant Consumer

    Producer->>Edge: append(stream,<br/>partition, producer_id,<br/>producer_seq, payload)
    Edge->>Catalog: validate writer epoch<br/>and policy
    Edge->>BK: append immutable record
    BK-->>Edge: minimum durable<br/>sites reached
    Edge-->>Producer: position and durable token
    Edge-->>Consumer: disseminate after durability
```

Every append carries a producer ID and producer sequence. Repeated submission
of the same producer tuple is idempotent and returns the existing position.
Payload checksums and the previous-segment hash protect integrity across sites.

The default `ARK_REPLICATED` stream acknowledges after two independent sites
are durable, then converges to three sites with a maximum of five. Critical
control streams use `ARK_SYSTEM`, with three sites before acknowledgement and a
desired five, maximum seven. `ARK_LOCAL_ASYNC` is allowed only for explicitly
reconstructible streams that accept a site-loss window.

## ArkTxn Mutation Capture

The first production ArkTxn integration uses a transactional outbox, not
FoundationDB Native CDC.

```mermaid
sequenceDiagram
    participant Client
    participant Txn as ArkTxn authority
    participant FDB as FoundationDB
    participant Publisher as Outbox publisher
    participant Log as ArkLog

    Client->>Txn: transaction request
    Txn->>FDB: application mutations<br/>plus outbox record
    FDB-->>Txn: one atomic durable commit
    Txn-->>Client: commit token
    Publisher->>FDB: read committed outbox<br/>records in order
    Publisher->>Log: append mutation segment<br/>idempotently
    Log-->>Publisher: ArkLog durable token
    Publisher->>FDB: advance exported watermark
```

All writes to an ArkTxn-managed range must pass through the ArkTxn API so the
application mutations and outbox record commit atomically. Publication is
at-least-once; deterministic segment identity makes retries harmless. The
outbox publisher advances its exported watermark only after ArkLog satisfies
the stream durability policy.

FoundationDB 8.0 Native CDC is a future capture adapter. It remains
experimental, is independently acknowledged rather than transactionally
coupled to application writes, and is not required for the first production
correctness path. Adoption requires a stable upstream interface plus proven
gap recovery, redelivery, retention, upgrade, and rollback behavior.

## Mutation Segments

ArkTxn mutation segments preserve transaction boundaries and authoritative
order:

```yaml
arkMutationSegment:
  rangeId: tenant-acme
  ownerEpoch: 91
  firstSequence: 58100
  lastSequence: 58103
  sourceEngine: foundationdb-8
  firstEngineVersion: 1000001
  lastEngineVersion: 1000004
  previousSegmentHash: previous-hash
  segmentHash: segment-hash
  transactions: []
```

A segment belongs to one range and owner epoch. A hot replica rejects an epoch
mismatch, buffers out-of-order segments, and applies only a contiguous
sequence. ArkLog transports authoritative history; it never resolves conflicts
or reruns transactions.

## Dissemination

ArkLog uses deterministic, topology-aware, pipelined dissemination rather than
random gossip. A verified replica may forward a segment to other eligible
sites. ArkLog chooses the dissemination tree; ArkNet chooses the paths for each
tree edge. The placement records, minimum-site acknowledgement rule, replica
watermarks, dissemination algorithm, bootstrap path, and repair invariants are
defined in [`ark-log-replication.md`](ark-log-replication.md).

## Replay And Retention

Consumers checkpoint the last durably processed position. ArkLog retains every
position required by:

* a protected consumer cursor;
* a hot-replica applied watermark;
* a snapshot or bootstrap operation;
* a legal, audit, or operational retention pin;
* the configured time or size retention floor.

A slow consumer cannot force the authority database to retain transaction-log
history indefinitely. It may hold ArkLog retention within policy; when that
limit is exceeded, the consumer must bootstrap from a snapshot and resume from
the snapshot token.

## Compaction And Archive

Closed BookKeeper segments remain immutable. Compaction may copy retained
records into a new segment, publish a position-preserving index, and delete old
ledgers only after all replicas and protected cursors can resolve the new
location. Key-compacted application streams are a higher-level policy; ArkTxn
mutation history is compacted only through a verified range snapshot plus a
resume token.

Long-retention closed segments may be archived to ArkObj. The
archive manifest records the stream range, checksums, predecessor hash, and
ArkLog positions. Active appends and near-term replay continue to use
BookKeeper directly.

## Failure And Recovery

A stale writer is rejected by both BookKeeper fencing and the ArkLog writer
epoch. Recovery verifies the durable head before advancing the epoch and
opening a new segment. Detailed writer failover, gap repair, replica removal,
and rebalancing rules live in
[`ark-log-replication.md`](ark-log-replication.md).

## Non-Goals

ArkLog does not initially provide Kafka-compatible consumer groups, arbitrary
cross-partition transactions, multi-master ordering, or exactly-once external
side effects. Those semantics require higher-level coordination. ArkLog
provides durable ordered records, idempotent append, replay, and explicit
watermarks on which such services can be built.

## Accepted Initial Decisions

* BookKeeper is the active ordered append substrate.
* ArkObj is an optional archive for closed compacted segments, not
  the foreground append path.
* The Ark Control Catalog owns stream epochs and policy roots.
* Transactional outbox capture feeds ArkTxn mutations first.
* Standard streams default to `ARK_REPLICATED`; control streams default to
  `ARK_SYSTEM`.
* Dissemination is deterministic, topology-aware, and pipelined.

Revisit Native CDC when FoundationDB marks it production-ready and Ark testing
proves upgrade, replay, acknowledgement, and retention behavior. Revisit
erasure-coded segment archives only after replicated archive recovery is
operationally proven.
