# ArkTxn Architecture

ArkTxn is Ark Fabric's geographically federated transaction subsystem. Its
contract is a strongly ACID, ordered transactional key-value space backed by
site-local or region-local NewSQL key-value engines. FoundationDB is the first
and reference engine; it is an implementation dependency, not the subsystem's
product name.

The first implementation uses FoundationDB for local ACID execution, MVCC,
conflict detection, transaction logs, storage replication, and recovery.
ArkTxn adds the global layer needed across many independent cells: range
ownership, hierarchical placement, nearby synchronous Ark durability,
asynchronous hot-replica convergence, and global routing. Another engine can
replace FoundationDB only if it satisfies the same transaction, ordering,
durability, version-token, replay, and fencing contracts.

ArkTxn is not FoundationDB running on ArkFS or ArkObj API. FoundationDB hot data,
transaction logs, storage engines, resolvers, proxies, and storage servers
should remain on local infrastructure chosen for FoundationDB latency and I/O
requirements. ArkObj can still hold large values, backups, archives,
snapshots, and bulk payloads referenced by transactions, with ArkObj API and ArkFS
providing access surfaces where appropriate.

## Design Principle

ArkTxn follows one core rule:

```mermaid
flowchart LR
    Txn["Run engine transaction<br/>and conflict logic once"]
    Owner["Authoritative owner<br/>for key range"]
    Stream["Ordered committed<br/>mutation stream"]
    Replicas["Hot replicas<br/>apply stream"]

    Txn --> Owner
    Owner --> Stream
    Stream --> Replicas
```

At the architectural level, ArkTxn follows the same broad durability and
convergence pattern as ArkLog: establish the required durable ordered history,
acknowledge according to policy, and then converge additional replicas
asynchronously. For ArkTxn, native FoundationDB transaction logs and nearby
satellite logs satisfy the foreground durability requirement; ArkLog then
carries the committed mutation stream to independent hot replicas.

```mermaid
flowchart TB
    subgraph LogPattern["ArkLog"]
        LogAppend["Append ordered record"] --> LogMinimum["Minimum durable<br/>ArkLog copies"]
        LogMinimum --> LogAck["Acknowledge append"]
        LogAck --> LogConverge["Converge copies<br/>and consumers"]
    end

    subgraph FDBPattern["ArkTxn"]
        FDBCommit["Commit transaction"] --> FDBMinimum["Authority-cell and<br/>satellite tLog<br/>durability"]
        FDBMinimum --> FDBAck["Acknowledge transaction"]
        FDBAck --> Publish["Publish committed<br/>mutations to ArkLog"]
        Publish --> ReplicaConverge["Converge desired<br/>hot replicas"]
    end
```

ArkObj can replicate immutable bytes without imposing transaction
order across unrelated chunks. ArkTxn must preserve committed transaction
order. Its geographic unit is therefore not a chunk, but a key range plus its
ordered mutation history.

## High-Level Architecture

```mermaid
flowchart TB
    ArkTxn["ArkTxn"] --> Catalog["Ark Control Catalog"]
    Catalog --> Policy["Hierarchical policy map"]
    Policy --> Domains["Transaction domains"]
    Domains --> Ownership["Range / epoch ownership"]

    Ownership --> CellA["Authority Cell A<br/>owns R1, R4"]
    Ownership --> CellB["Authority Cell B<br/>owns R2, R3"]

    CellA --> FDBA["FoundationDB cell<br/>normal ACID txn"]
    CellB --> FDBB["FoundationDB cell<br/>normal ACID txn"]

    FDBA --> LogsA["Local + satellite logs"]
    FDBB --> LogsB["Local + satellite logs"]
    LogsA --> AckA["Client ACK"]
    LogsB --> AckB["Client ACK"]

    FDBA --> OutboxA["Transactional outbox"]
    FDBB --> OutboxB["Transactional outbox"]
    OutboxA --> Segments["ArkLog segments"]
    OutboxB --> Segments
    Segments --> Dissemination["ArkLog dissemination<br/>scheduler"]
    Dissemination --> Transport["ArkNet<br/>QUIC or MPQUIC"]
    Transport --> ReplicaA["Hot FDB replica"]
    Transport --> ReplicaB["Hot FDB replica"]
    Transport --> ReplicaC["Hot FDB replica"]
```

## Components

| Component | Responsibility |
|---|---|
| ArkTxn API | Presents one logical ordered keyspace to applications. |
| Range router | Routes reads and writes to the correct range owner or replica. |
| Ark Control Catalog | Stores the authoritative range / epoch map, domain policy, site identity, promotion leases, and fences. |
| Range / epoch cache | Router-local versioned snapshot of catalog state. |
| Transaction domain | Groups keys that should remain in one authority cell. |
| Authority cell | The FoundationDB cell that executes writes for a range. |
| Satellite durability tier | Nearby transaction-log durability in an independent failure domain. |
| ArkLog | Included Ark Fabric subsystem that durably stores and disseminates ordered mutation segments. |
| Transactional outbox | Captures mutations atomically with the authoritative commit. |
| Replicator | Converts committed outbox records into immutable ArkLog mutation segments. |
| Epidemic scheduler | Chooses which site forwards segments to which replicas. |
| Hot replica | Independent FDB cluster materialized from the ordered stream. |
| Promotion controller | Promotes a caught-up hot replica when ownership moves or fails over. |

## Ark Control Catalog

The global range and epoch map lives in the Ark Control Catalog. The first
implementation is a dedicated, highly available FoundationDB deployment that
is outside the ArkTxn-routed application keyspace. This avoids a bootstrap
cycle in which a router would need a range map from the same routed keyspace it
cannot yet locate.

The catalog stores only compact global control state:

* range boundaries, owner cell, and owner epoch;
* transaction-domain declarations and placement constraints;
* site and authority-cell identities;
* policy roots and storage-class definitions;
* promotion leases, fencing records, and durable-head tokens;
* ArkLog stream IDs, writer epochs, and policy references.

Service metadata, user values, object namespaces, and bulk data do not belong
in the catalog. Routers cache signed or authenticated versioned snapshots and
serve normal requests without a catalog round trip. Ownership mutations,
domain moves, and promotion fencing are transactional catalog operations.

```mermaid
flowchart LR
    Catalog["Ark Control Catalog"] --> Snapshot["Versioned range /<br/>epoch snapshot"]
    Snapshot --> RouterA["Router A cache"]
    Snapshot --> RouterB["Router B cache"]
    RouterA --> Owner["Current authority cell"]
    RouterB --> Owner
    Promotion["Promotion controller"] -->|transactional<br/>epoch advance| Catalog
```

## Authority Model

ArkTxn is active-active globally, but single-primary per key range.

```mermaid
flowchart LR
    R1["Range R1"] --> IST["Owner IST"]
    R2["Range R2"] --> FRA["Owner FRA"]
    R3["Range R3"] --> AMS["Owner AMS"]
    R4["Range R4"] --> IAD["Owner IAD"]
```

All of those owners can commit transactions at the same time. The restriction
is that each range has exactly one write authority for a given epoch.

That rule prevents two sites from independently committing conflicting writes
to the same key while still allowing the global system to scale across many
sites.

## Authority Cell

An authority cell is the transaction execution domain for a key range. It may
be more than one physical site when FoundationDB native high-availability
configuration is used.

Example:

```mermaid
flowchart TB
    Cell["ArkTxn authority<br/>cell EU-1"]
    Cell --> Primary["Primary FDB storage<br/>IST"]
    Cell --> Satellite["Satellite transaction<br/>logs<br/>ANK"]
    Cell --> Secondary["Native FDB secondary<br/>FRA"]

    Cell -. "outside<br/>foreground<br/>commit path" .-> AMS["External hot replica<br/>AMS"]
    Cell -. "outside<br/>foreground<br/>commit path" .-> LON["External hot replica<br/>LON"]
    Cell -. "outside<br/>foreground<br/>commit path" .-> IAD["External hot replica<br/>IAD"]
```

This keeps the foreground commit path small while still allowing a much larger
set of independent hot replicas outside the native FoundationDB cell.

## Write Path

For a write to a key range owned by IST:

```mermaid
sequenceDiagram
    participant Client as Client at any site
    participant Router as ArkTxn router
    participant Owner as Owner authority cell IST
    participant FDB as FoundationDB transaction
    participant Local as Local tLogs
    participant Satellite as Nearby satellite tLogs
    participant Publisher as Outbox publisher
    participant ArkLog
    participant Replica as Hot replicas

    Client->>Router: write request
    Router->>Owner: route by range / epoch
    Owner->>FDB: application mutations<br/>+ outbox record
    FDB->>Local: durable local log
    FDB->>Satellite: durable nearby log
    Local-->>Owner: durable
    Satellite-->>Owner: durable
    Owner-->>Client: ACK after minimum durability
    Publisher->>FDB: read committed outbox<br/>records in order
    Publisher->>ArkLog: append mutation segment<br/>idempotently
    Publisher->>FDB: advance exported watermark<br/>after ArkLog durability
    ArkLog-->>Replica: converge asynchronously
```

The key latency rule is:

```text
Ark commit latency ~= transaction processing
                   + max(local durable log path,
                         near-site durable log path)
```

It should not be a foreground local FDB transaction plus a second foreground
remote FDB transaction.

Replaying into a second full FDB cluster on the foreground path would serialize
two commit paths and harm latency. ArkTxn should acknowledge after the required
nearby durable log tier, then let independent hot replicas materialize
asynchronously.

The outbox record is part of the same FoundationDB transaction as the
application mutations. All writes to an ArkTxn-managed range therefore pass
through the ArkTxn API. Publication is at-least-once and uses deterministic
segment IDs so retries cannot create a second logical mutation segment.

FoundationDB 8.0 Native CDC is not the initial capture path. Upstream marks it
experimental, registration is disabled by default, and acknowledgement is not
atomic with application transactions. ArkTxn may add it as an adapter after
its API and operational lifecycle are production-ready and Ark validation
proves redelivery, gap recovery, retention, upgrade, and rollback behavior.

## Read Path

Reads are routed according to consistency requirement and replica watermark.
`LINEARIZABLE` is the default. The first API exposes all five modes below;
weaker modes require explicit selection.

| Read mode | Behavior |
|---|---|
| `LOCAL` | Read the nearest replica even if it may be stale. |
| `BOUNDED(max_lag)` | Read local if replica lag is within the requested bound. |
| `AT_LEAST(token)` | Read only from a replica that has applied the token. |
| `READ_YOUR_WRITES` | Convenience mode implemented as `AT_LEAST` the client's last commit token. |
| `LINEARIZABLE` | Route to the authoritative owner for the range. |

Every hot replica exposes an applied watermark per range:

```yaml
replicaWatermark:
  range: R1
  epoch: 91
  sequence: 58103
```

If a London replica has not applied the required token, ArkTxn can wait briefly
or route the read to another current replica.

## Transaction Domains

ArkTxn should keep related keys in one authority cell whenever they commonly
participate in the same transaction.

Example key prefixes:

* `/tenants/acme/accounts/*`
* `/tenants/acme/billing/*`
* `/tenants/acme/permissions/*`

These should share a transaction domain:

```yaml
transactionDomain: tenant/acme
```

Then most transactions remain ordinary single-FDB transactions inside one
authority cell. That is much faster than cross-cell distributed transactions.

Design target: more than 99% of write transactions should stay inside one
authority cell.

Transaction domains are declared as catalog policy records over non-overlapping
or properly nested key ranges. The most specific matching policy wins. Catalog
validation enforces these rules:

* every range in one domain has the same owner and owner epoch;
* a policy update cannot split an active domain across owners;
* nested policies cannot contradict the parent's domain assignment;
* a domain migration moves the complete domain under one fenced epoch change;
* routers preflight read and write conflict ranges before sending mutations and
  reject a transaction that spans owners.

Telemetry records cross-domain transaction attempts and commonly co-accessed
ranges. Operators use that evidence to merge, split, or migrate domains rather
than guessing transaction affinity in advance.

## Cross-Owner Transactions

A transaction that writes two independently owned ranges cannot preserve full
FoundationDB serializability without coordination.

ArkTxn should handle this in priority order:

1. Co-locate transactionally related ranges in one transaction domain.
2. Route the whole transaction group to one owner when practical.
3. Reject with `CROSS_AUTHORITY_TRANSACTION` or use an application saga and
   outbox when atomicity is not required.
4. Add a slower cross-owner commit protocol only for exceptional cases that
   require full serializability.

The initial architecture should optimize for the first two strategies and
treat arbitrary cross-cell ACID transactions as a later major subsystem.
The cross-owner protocol enters design only if, after domain redesign and
migration, legitimate cross-owner writes exceed 1% of write transactions or a
required workflow cannot be co-located. The 1% threshold is a review trigger,
not automatic permission to add distributed commit latency to the fast path.

## Relationship To Ark Fabric

ArkTxn belongs to the transaction plane of the wider platform.

```mermaid
flowchart TB
    Fabric["Ark Fabric<br/>containing platform"]
    Fabric -->|includes| Txn["ArkTxn<br/>transaction subsystem"]
    Fabric -->|includes| Log["ArkLog<br/>log subsystem"]
    Fabric -->|includes| Object["ArkObj<br/>object subsystem"]
    Fabric -->|includes| Transport["ArkNet<br/>network subsystem"]
    Fabric -->|presents| Views["Service views<br/>S3, BK-QCOW, ArkFS"]

    Txn -->|publishes<br/>committed mutations| Log
    Txn -. "may reference<br/>large values" .-> Object
    Views -->|use transactions<br/>and metadata| Txn
    Views -->|store immutable<br/>payloads| Object
    Txn -->|uses for<br/>WAN traffic| Transport
    Log -->|uses for<br/>dissemination| Transport
    Object -->|uses for payload<br/>movement| Transport
```

ArkTxn can store:

* namespace metadata;
* indexes;
* locks and leases;
* service-level placement policy references;
* roots and manifest references;
* object version pointers;
* access-control metadata;
* application state.

The Ark Control Catalog, rather than an application range, owns global ArkTxn
range placement and owner epochs.

Large values and bulk payloads should live in ArkObj and may be
accessed through ArkObj API or ArkFS. Examples include:

* large values;
* files;
* checkpoints;
* archival snapshots;
* backup payloads;
* mutation-segment archives when appropriate.

## Non-Goals

ArkTxn does not initially provide:

* native FoundationDB storage-server files on ArkFS;
* every site carrying every range;
* multi-master writes to the same key range;
* arbitrary cross-owner ACID transactions on the fast path;
* multi-source reads of small FoundationDB values as if they were blobs;
* foreground commits that wait for maximum replication factor.

## Key Invariants

* A key range has one authoritative owner per epoch.
* The Ark Control Catalog is authoritative for ownership and epoch changes.
* The owner executes transaction logic and conflict checks exactly once.
* Application mutations and their outbox record commit atomically.
* Replicas apply committed mutations, not transaction logic.
* Mutation segments are immutable and ordered.
* Hot replicas may receive segments out of order but apply them in order.
* A client acknowledgement requires the configured minimum durability tier.
* Desired and maximum hot replicas are convergence targets, not commit-path
  participants.
* Promotion requires a replica to catch up to the durable head and advance the
  range epoch.
* Hierarchical policy must respect validated transaction domains.

## Accepted Initial Decisions

| Topic | Decision |
|---|---|
| Global range and epoch map | Store it in the dedicated Ark Control Catalog, implemented first with an independent highly available FoundationDB deployment. |
| ArkLog substrate | Use BookKeeper for active streams and ArkObj only for optional closed-segment archives. |
| Mutation capture | Use an ArkTxn transactional outbox first; evaluate FoundationDB Native CDC after it is production-ready. |
| System-range policy | Use `ARK_SYSTEM`: three independent durability sites before acknowledgement, desired five hot replicas, maximum seven. |
| Read API | Expose all five documented modes and default to `LINEARIZABLE`. |
| Transaction domains | Declare them as validated catalog range policies and move them atomically under an owner epoch. |
| Cross-owner transactions | Reject or use sagas initially; review a cross-owner commit protocol only after the measured 1% trigger or a non-co-locatable mandatory workflow. |

The focused replication details are in
[`ark-txn-replication.md`](ark-txn-replication.md), and the ordered stream
contract is in [`ark-log-architecture.md`](ark-log-architecture.md).
