# Ark Fabric Architecture

Ark Fabric is the umbrella architecture for the Ark system. It contains
WAN-native transport, geographically distributed object storage, ordered logs,
a transactional key-value system, shared control services, and service views
such as S3, OCI registries, BK-QCOW, ArkFS, and Kubernetes CSI volumes.

`Ark` is the project term for geographic distribution across independent
sites, fault domains, and institutions. A National Ark is an Ark Fabric
deployment operated by a nation state across embassies, consulates, and
national sites. A Corporate Ark is an Ark Fabric deployment operated by a large
geographically distributed corporation across branch offices and corporate
sites.

The design began as a BookKeeper-backed qcow2/NBD concept, but the final
architecture is broader: Ark Fabric is the whole WAN-enabled fabric, not just
the object layer.

## Terminology And System Boundaries

The name **Ark Fabric** always denotes the containing platform. Runtime
behavior belongs to a named subsystem; Ark Fabric is not a synonym for ArkLog,
ArkTxn, ArkObj, or ArkNet.

| Concept | Architectural role |
|---|---|
| Ark Fabric | Umbrella platform that includes all Ark subsystems and service views. |
| ArkTxn | Transaction subsystem for the geographically federated, strongly ACID key-value space. FoundationDB is its first engine. |
| ArkLog | Log subsystem for durable ordered append history, replay, and dissemination. |
| ArkObj | Object subsystem for immutable chunks, manifests, roots, and replicas. |
| ArkNet | Network subsystem for WAN paths, QUIC, MPQUIC, and overlays. |
| Ark Control Catalog | Bootstrap-safe control store for global ownership, epochs, policies, site identity, and promotion fencing. |
| Ark scheduler | Shared control component for placement, source, target, path, and hedge decisions. |
| Service view | User-facing protocol or product interface that consumes one or more Ark subsystems. |

Names follow the kind of boundary they describe. The `Ark` prefix identifies
Ark-owned subsystems and native interfaces. Standards-based surfaces keep
their established names: S3-compatible API, OCI registry, and CSI driver.
ArkFS and BK-QCOW remain project-specific service names. The ArkObj API is a
descriptive interface within ArkObj, not a fifth subsystem.

Subsystem behavior must be compared with the appropriate peer subsystem. For
example, ArkTxn shares ArkLog's durability-and-convergence pattern; it does not
share a pattern "with Ark Fabric," because Ark Fabric contains both systems.

```mermaid
flowchart TB
    subgraph Fabric["Ark Fabric"]
        subgraph Services["Service layer"]
            Views["S3, BK-QCOW<br/>ArkFS, OCI, CSI<br/>backup, CDN"]
        end

        subgraph Core["Core subsystems"]
            Txn["ArkTxn<br/>transaction subsystem"]
            Log["ArkLog<br/>ordered log subsystem"]
            Object["ArkObj<br/>immutable object<br/>subsystem"]
            Transport["ArkNet<br/>WAN network subsystem"]
        end

        subgraph Control["Shared control"]
            Catalog["Ark Control<br/>Catalog<br/>ownership,<br/>epochs, policy<br/>roots, and<br/>fencing"]
            Scheduler["Ark scheduler<br/>placement,<br/>source, path,<br/>and hedge<br/>decisions"]
        end
    end

    subgraph Sites["Independent Ark sites"]
        EdgeA["Ark Edge A"]
        EdgeB["Ark Edge B"]
        EdgeC["Ark Edge C"]
        BKA["BookKeeper cluster A"]
        BKB["BookKeeper cluster B"]
        BKC["BookKeeper cluster C"]
        FDBA["FoundationDB cell A"]
        FDBB["FoundationDB cell B"]
        FDBC["FoundationDB cell C"]
    end

    Views -->|immutable payloads<br/>and manifests| Object
    Views -->|transactions<br/>and metadata| Txn
    Txn -->|committed<br/>mutation stream| Log
    Txn -. "large-value<br/>references" .-> Object
    Scheduler -->|controls<br/>placement| Object
    Scheduler -->|controls replica<br/>dissemination| Log
    Scheduler -->|controls range<br/>placement| Txn
    Scheduler -->|controls<br/>path policy| Transport
    Catalog -->|authorizes<br/>range epochs| Txn
    Catalog -->|registers streams<br/>and policies| Log
    Object -->|payload movement| Transport
    Log -->|segment<br/>dissemination| Transport
    Txn -->|routing and<br/>replication<br/>traffic| Transport
    Transport --> EdgeA
    Transport --> EdgeB
    Transport --> EdgeC
    EdgeA --> BKA
    EdgeB --> BKB
    EdgeC --> BKC
    EdgeA --> FDBA
    EdgeB --> FDBB
    EdgeC --> FDBC
```

BookKeeper is not the product surface. It is a site-local durable substrate
under ArkObj and ArkLog. FoundationDB is not hidden inside ArkFS; it
remains the local ACID substrate for ArkTxn authority cells and hot replicas.

## Goals

Ark Fabric is designed to provide:

* geographically distributed immutable object storage;
* geographically federated transactional key-value metadata through ArkTxn;
* ordered durable mutation and append streams through ArkLog;
* virtual block devices with native qcow2-like copy-on-write semantics;
* S3-compatible object storage with accelerated large-object transfer;
* an OCI registry for container images and other OCI artifacts;
* Kubernetes volumes through a CSI driver backed by supported block or file
  services;
* deterministic multi-source reads without peer discovery;
* multi-WAN path aggregation through ArkNet, QUIC, and MPQUIC;
* MASQUE CONNECT-IP site overlay support where a packet tunnel is required;
* minimum-durability acknowledgement followed by desired-replica convergence;
* cheap snapshots, clones, object copies, and manifest-only updates;
* service views over shared object, transaction, log, and transport substrates.

## Non-Goals

The architecture does not assume:

* a mutable `.qcow2` file as the native runtime format;
* BookKeeper bookies exposed directly to the public Internet;
* FoundationDB storage-server files hosted on ArkFS;
* IP multicast over commodity ISP networks;
* traditional NBD-over-TCP or BookKeeper-over-TCP as the ArkNet data
  plane;
* a single global BookKeeper or FoundationDB cluster as the first-generation
  design;
* multi-master writes to the same ArkTxn key range.

TCP remains appropriate inside low-latency sites. QUIC, MPQUIC, multi-source
scheduling, and application-level fanout are most valuable across
geographically distributed WAN links.

## Architectural Roles

The roles below describe responsibilities. The named implementation makes clear
which included subsystem owns each behavior.

| Role | Primary implementation | Responsibility |
|---|---|---|
| Service layer | S3-compatible API, OCI registry, BK-QCOW, ArkFS, CSI driver, backup, CDN, data lake | User-facing protocols and product interfaces. |
| Transaction plane | ArkTxn | Range ownership, transactions, metadata, indexes, and leases. |
| Log plane | ArkLog | Ordered append, mutation segments, replay, retention, and dissemination. |
| Metadata plane | Initially ArkTxn | Buckets, keys, volume heads, roots, leases, policies, and versions. |
| Object plane | ArkObj | Immutable chunks, object manifests, roots, and replicas. |
| Global control | Ark Control Catalog | Ownership, epochs, policy roots, site identity, and fencing. |
| Shared scheduling | Ark scheduler | Site selection, replica placement, path selection, hedging, and chunk assignment. |
| Data plane | ArkNet | QUIC, MPQUIC, MASQUE CONNECT-IP, multi-source reads, and fanout writes. |
| Site access | Ark Edge | Authentication, authorization, caching, verification, and range serving. |
| Local durability | BookKeeper and FoundationDB | Site-local append storage and ACID transaction storage. |

ArkTxn owns transactional service metadata such as S3 namespace state, ArkFS
hierarchy, leases, roots, indexes, versions, and access-control records.
FoundationDB is the first ArkTxn engine. The separate Ark Control Catalog owns
only bootstrap and global control state that ArkTxn routing itself depends on;
service metadata must not leak into that catalog. BookKeeper is better suited
to durable append payloads, immutable object bytes, and active ArkLog storage.

### Metadata Ownership

| Metadata | Owner | Reason |
|---|---|---|
| S3 buckets, keys, versions, multipart state, ACLs, and object locks | ArkTxn service keyspace | Requires atomic namespace updates and strongly consistent listing. |
| ArkFS directories, inodes, file versions, leases, ACLs, and snapshots | ArkTxn service keyspace | Requires transactional hierarchy and version publication. |
| BK-QCOW volume heads, writer leases, epochs, and retained roots | ArkTxn service keyspace | Requires fencing and atomic head publication. |
| ArkTxn range ownership, owner epochs, transaction-domain policy, promotion fences, and site identity | Ark Control Catalog | Must be available before the routed ArkTxn keyspace can be trusted. |
| ArkLog stream identity, writer epoch, policy root, and active-segment pointer | Ark Control Catalog | Avoids a bootstrap cycle between ArkTxn and ArkLog. |
| BookKeeper ledger ensembles and physical entries | BookKeeper metadata and data services | Remains local substrate state, not an Ark service namespace. |
| Object bytes, manifests, and immutable roots | ArkObj | Keeps bulk immutable state out of the transaction engine. |

The Ark Control Catalog is a small, dedicated, highly available FoundationDB
deployment in the first implementation. It is not reached on every data-plane
operation: routers and controllers cache versioned snapshots, while ownership
changes and fencing remain catalog transactions.

## ArkObj

ArkObj is the immutable object subsystem. It stores chunks, manifests, and
roots; verifies payloads; and enforces object durability and replica placement.
ArkTxn publishes mutable service names and heads that reference those durable
records. ArkNet moves payloads between sites, while the shared scheduler
selects sources and targets within ArkObj policy.

The record model, write and read paths, native ArkObj API, retention rules, and
service integrations are specified in
[`ark-obj-architecture.md`](ark-obj-architecture.md).

### Initial Size Profile

The first implementation uses explicit size classes rather than one universal
unit:

| Unit | Initial default | Purpose |
|---|---:|---|
| ArkNet transfer frame | 64 KiB | Scheduling, retransmission, and path reassignment. |
| General immutable data chunk | 1 MiB | S3, ArkObj API, ArkFS, backup, and artifact payloads. |
| Large-object manifest span | 128 MiB | Bounds manifest fanout and aligns large transfers with multipart work. |
| BK-QCOW logical cluster | 128 KiB | Copy-on-write mapping and allocation. |
| BK-QCOW subcluster | 4 KiB | Guest block changes and zero/unallocated state. |

BK-QCOW may store a 128 KiB cluster directly or aggregate adjacent immutable
clusters into a larger extent without changing the logical mapping. Lengths are
encoded in records, and these defaults are benchmark-tunable deployment
parameters rather than permanent wire-format constants.

## Durability Model

ArkLog, ArkObj, and ArkTxn share a two-stage policy shape: satisfy a
minimum durability contract before acknowledgement, then converge toward the
desired placement. They do not share one generic commit protocol; each
subsystem implements that policy using its own durable unit and ordering rules.

```mermaid
flowchart TB
    ObjectWrite["ArkObj write"] --> ObjectMinimum["Minimum durable<br/>chunk replicas"]
    LogAppend["ArkLog append"] --> LogMinimum["Minimum durable<br/>ordered-log copies"]
    FDBCommit["ArkTxn transaction"] --> FDBMinimum["Authority-cell and<br/>satellite tLog<br/>durability"]

    ObjectMinimum --> ObjectAck["Acknowledge object<br/>operation"]
    LogMinimum --> LogAck["Acknowledge log append"]
    FDBMinimum --> FDBAck["Acknowledge transaction"]

    ObjectAck --> ObjectConverge["Converge desired<br/>object replicas"]
    LogAck --> LogConverge["Converge desired<br/>ArkLog copies<br/>and consumers"]
    FDBAck --> FDBConverge["Publish through ArkLog<br/>and converge hot<br/>FDB replicas"]
```

Example ArkObj or ArkLog policy:

```yaml
minimumDurableSites: 2
desiredSites: 3
maximumSites: 5
```

An object manifest or root must not become visible unless every referenced
payload satisfies the ArkObj minimum durability policy. An ArkLog
append follows ArkLog's ordered-log durability policy. An ArkTxn transaction
follows its authority-cell and satellite-log commit policy before ArkLog
disseminates the resulting committed mutations.

For block storage, this rule is critical:

```text
durability(data referenced by transaction)
    >=
durability(commit for transaction)
```

Losing data after preserving a commit is dangerous. Preserving uncommitted data
only creates garbage that can be collected later.

### Initial Storage Classes

The first release exposes three policy bundles. The same names express common
intent, while each subsystem enforces durability using its own commit unit.

| Class | Acknowledgement rule | Convergence target | Intended use |
|---|---|---|---|
| `ARK_REPLICATED` | Two independent sites | Three sites, maximum five | Default objects and general ArkLog streams. |
| `ARK_LOCAL_ASYNC` | One locally durable site | Three sites, maximum five | Latency-sensitive, reconstructible data that explicitly accepts a site-loss window. |
| `ARK_SYSTEM` | Three independent sites | Five sites, maximum seven | Control records, system ranges, and critical log streams. |

Erasure-coded `ARK_EC` and all-edge `ARK_HOT_GLOBAL` classes are deferred until
the replicated path, repair loop, and failure-domain accounting are validated.
`ARK_LOCAL_ASYNC` must never be selected implicitly for data whose contract
requires survival of local-site loss at acknowledgement.

## Transaction And Log Flow

ArkTxn is the geographically federated transactional key-value layer inside
Ark Fabric. It uses local FoundationDB cells for ACID transaction execution,
range ownership, MVCC, conflict checks, and hot-replica materialization.

ArkLog is the ordered durable log plane used by ArkTxn and by other Ark
subsystems that need replayable append history. It stores immutable mutation
segments and append records in site-local durable stores, with BookKeeper as
the natural first substrate.

ArkTxn and ArkLog are included core subsystems, not service views. ArkTxn owns
transaction execution, key-range authority, and replica-placement intent.
ArkLog owns durable ordered mutation history, replay, and segment
dissemination. ArkNet carries their cross-site traffic.

```mermaid
flowchart TB
    Client["Application or<br/>service view"] --> Router["ArkTxn router"]
    Router --> RangeMap["Ark Control Catalog<br/>range / epoch map"]
    RangeMap --> Owner["Authority<br/>FoundationDB cell"]
    Owner --> Commit["FoundationDB<br/>transaction"]
    Commit --> Durable["Local and satellite<br/>tLogs durable"]
    Durable --> Ack["Client acknowledgement"]
    Durable --> Segment["Committed Ark<br/>mutation segment"]
    Segment --> ArkLog["ArkLog ordered stream"]
    ArkLog --> Scheduler["Ark scheduler"]
    Scheduler --> Transport["ArkNet"]
    Transport --> ReplicaA["Hot FDB replica A"]
    Transport --> ReplicaB["Hot FDB replica B"]
    Transport --> ReplicaC["Hot FDB replica C"]
```

ArkTxn is active-active across ranges, but single-primary per key range and
epoch. The owner executes transaction logic and conflict checks exactly once.
Hot replicas apply committed mutation segments in order; they do not rerun
transaction logic. The first production capture path is a transactional outbox
written atomically with each ArkTxn mutation. FoundationDB Native CDC remains
an experimental future adapter, not the initial correctness dependency.

Detailed ArkTxn design lives in
[`ark-txn-architecture.md`](ark-txn-architecture.md). Replication and placement
details live in [`ark-txn-replication.md`](ark-txn-replication.md).
ArkLog's general stream, durability, replay, and retention design lives in
[`ark-log-architecture.md`](ark-log-architecture.md); its placement, repair,
and dissemination design lives in
[`ark-log-replication.md`](ark-log-replication.md).

## ArkNet

ArkNet is the WAN-native data plane for Ark Fabric. It is object-,
replica-, topology-, and path-aware. It does not reduce storage I/O,
replication, or metadata movement to a single endpoint socket.

ArkNet incorporates the multi-WAN QUIC, MPQUIC, and MASQUE CONNECT-IP
architecture developed by Ark Sites. It is the reusable transport and
site-connectivity layer consumed by ArkObj, ArkTxn, ArkLog, ArkFS, and other
service views.

```mermaid
flowchart LR
    Scheduler["Ark scheduler"] --> SourceChoice["Choose source<br/>and target sites"]
    SourceChoice --> PathChoice["Choose WAN paths"]
    PathChoice --> QUIC["QUIC streams"]
    PathChoice --> MPQUIC["MPQUIC path set"]
    PathChoice --> MASQUE["MASQUE CONNECT-IP<br/>overlay"]
    QUIC --> Edge["Ark Edge"]
    MPQUIC --> Edge
    MASQUE --> Edge
    Edge --> LocalStores["BookKeeper and<br/>FoundationDB"]
```

For reads, the scheduler knows which sites contain each chunk and can fetch
ranges from eligible replicas without DHT, tracker, or peer discovery.

```mermaid
flowchart LR
    Client --> Scheduler["Ark scheduler"]
    Scheduler --> FRA["Frankfurt replica<br/>chunks 0-220"]
    Scheduler --> AMS["Amsterdam replica<br/>chunks 221-390"]
    Scheduler --> IST["Istanbul replica<br/>chunks 391-520"]
    FRA --> Client
    AMS --> Client
    IST --> Client
```

Scheduling should optimize predicted completion time, not only round-trip
time:

```text
completion_time =
  queue_delay
+ RTT_component
+ bytes / estimated_delivery_rate
+ loss_or_retransmission_penalty
```

QUIC and MPQUIC add:

* independent streams for unrelated operations;
* path migration;
* per-path congestion and loss telemetry;
* simultaneous use of multiple paths within a connection;
* application-controlled scheduling across paths.

Multi-source scheduling and multipath scheduling are separate but
complementary.

```mermaid
flowchart TB
    subgraph MultiSource["Multi-source scheduling"]
        SiteA["Site A"] --> ClientA["Client"]
        SiteB["Site B"] --> ClientA
        SiteC["Site C"] --> ClientA
    end

    subgraph Multipath["Multipath scheduling"]
        ClientB["Client"] <-->|WAN 1| SiteD["Site A"]
        ClientB <-->|WAN 2| SiteD
        ClientB <-->|WAN 3| SiteD
    end
```

See [`ark-net-architecture.md`](ark-net-architecture.md) for the
network-specific architecture and the accepted QUIC implementation path.

## Site Edge

BookKeeper bookies and FoundationDB processes should not be directly exposed
to the Internet. Each site should provide a site-local Ark Edge service.

```mermaid
flowchart TB
    WAN["Internet / WAN"] -->|"QUIC, MPQUIC,<br/>MASQUE"| Edge["Ark Edge"]
    Edge --> Auth["Authentication<br/>and authorization"]
    Edge --> Cache["Cache, prefetch,<br/>range serving"]
    Edge --> Verify["Checksum and<br/>policy verification"]
    Edge --> BK["BookKeeper cluster"]
    Edge --> FDB["FoundationDB cell"]
```

The edge is responsible for:

* authentication and authorization;
* object and range serving;
* checksum verification;
* caching and prefetching;
* rate control;
* local BookKeeper access;
* local FoundationDB access through approved ArkTxn roles;
* replication fanout and repair;
* telemetry for scheduler decisions.

## BK-QCOW View

BK-QCOW is a BookKeeper-native qcow2 semantic engine. It is not a conventional
`.qcow2` file stored on BookKeeper.

The runtime representation preserves qcow2 concepts but maps them onto
immutable objects and persistent roots.

| qcow2 concept | BK-QCOW and ArkObj representation |
|---|---|
| Header | Immutable volume descriptor and committed state |
| Host cluster address | Synthetic cluster or object ID |
| Data cluster | Immutable chunk or object |
| L2 table | Immutable or delta mapping object |
| L1 table | Persistent root mapping |
| Refcount table | Asynchronous liveness and reclamation metadata |
| Internal snapshot | Immutable root reference |
| Backing file | Parent volume or snapshot root |
| Zero cluster | Mapping state with no payload |
| Unallocated cluster | Missing mapping or parent fallback |
| Dirty bit | Open generation plus commit-log recovery |
| Persistent bitmap | Versioned bitmap object or deltas |
| Compression | Data-object property |
| COW gate | Unneeded on writes because all writes allocate new objects |

### BK-QCOW Read Path

```mermaid
flowchart LR
    Guest["Guest byte range"] --> Cluster["Cluster /<br/>subcluster numbers"]
    Cluster --> Root["Current root"]
    Root --> Mapping["L1 / L2 mapping"]
    Mapping --> Objects["Object IDs"]
    Objects --> Reads["Replica-aware<br/>object reads"]
    Reads --> Decode["Decode / decompress"]
    Decode --> Response["Block response"]
```

### BK-QCOW Write Path

```mermaid
flowchart LR
    Write["NBD or ublk write"] --> Data["Append immutable<br/>data object"]
    Data --> Mapping["Construct mapping<br/>change"]
    Mapping --> Root["Construct new root"]
    Root --> Commit["Append durable commit"]
    Commit --> Head["Publish new head"]
    Head --> Ack["Acknowledge client"]
```

Writes use a single authoritative volume epoch and ordered commit stream.
BookKeeper fencing or an external lease system prevents stale writers from
publishing conflicting roots.

The initial block path uses durable write-through semantics: `WRITE` and `FUA`
are acknowledged only after the data, mapping update, root, and commit are
durable to the configured acknowledgement quorum. `FLUSH` is a barrier through
the current commit sequence. `TRIM` publishes a hole mapping, and
`WRITE_ZEROES` publishes a zero mapping without storing a zero-filled payload.
A later write-back mode may batch ordinary writes, but it must preserve these
`FLUSH` and `FUA` guarantees.

The native record families are `VOLUME`, `DATA`, `L2_DELTA`, `ROOT`, `COMMIT`,
`SNAPSHOT`, `OBJECT_MOVE`, and `CHECKPOINT`. Runtime state uses these records,
not a mutable qcow2 file. Binary qcow2 import and export are serializer views
over the native state and must not constrain the internal object layout.

Snapshots are root references. Clones refer to parent snapshot roots and write
only their own changed objects. Backing images become parent-root relationships
rather than filename links.

```mermaid
flowchart TD
    Snapshot["Snapshot S1"] --> Root["Root R57"]
    Clone["Clone"] --> Snapshot
    Clone --> ChildRoot["Child root"]
    ChildRoot --> ChangedObjects["Changed objects only"]
```

### Block Frontends

NBD should be a compatibility frontend. A preferred Linux frontend may use
`ublk` or another high-performance userspace block path.

```mermaid
flowchart TB
    Apps["Applications / VMs"] --> NBD["NBD"]
    Apps --> Ublk["ublk"]
    NBD --> Client["BK-QCOW client"]
    Ublk --> Client
    Client --> Scheduler["Ark scheduler"]
    Scheduler --> Transport["ArkNet"]
```

Standard NBD-over-TCP is useful locally or as a compatibility surface, but it
should not define the WAN architecture.

## S3 View

S3 is a natural service view because objects are chunkable, range-readable, and
well aligned with multipart transfer.

```mermaid
flowchart TD
    Key["s3://bucket/key"] --> Version["Object version"]
    Version --> Manifest["Object manifest"]
    Manifest --> C0["Chunk C000"]
    Manifest --> C1["Chunk C001"]
    Manifest --> C2["Chunk C002"]
```

For large `GET` operations, the S3 edge can retrieve chunks from multiple
sites and reassemble the HTTP response. For `PUT` and multipart upload, the
edge can chunk the stream and immediately fan chunks out to durability sites.

S3 metadata should be handled by a strongly ordered metadata plane:

```yaml
buckets:
  example:
    objects:
      video.mkv:
        current: version-8172
objectVersions:
  version-8172:
    size: 4815162342
    etag: example-etag
    contentType: video/x-matroska
    manifestId: manifest-8172
    timestamp: 2026-10-05T00:00:00Z
    tags: {}
    policy: standard
```

Object overwrites become atomic metadata pointer updates after the new manifest
is durable. Deletes and versioned delete markers are metadata operations that
make old chunks eligible for garbage collection.

### S3 Acceleration Modes

```mermaid
flowchart TB
    Standard["Standard S3 client"] --> S3Edge["S3 edge endpoint"]
    S3Edge --> Transport["ArkNet"]
    Transport --> ObjectFabric["ArkObj"]

    App["Application or<br/>local proxy"] --> ArkS3["Ark-S3 client"]
    ArkS3 --> SiteA["Site A"]
    ArkS3 --> SiteB["Site B"]
    ArkS3 --> SiteC["Site C"]
```

A local S3-compatible proxy can let unmodified tools use the enhanced protocol
through a localhost endpoint.

## ArkFS View

ArkFS is the filesystem service view over ArkObj. It presents a
familiar hierarchy of directories, files, permissions, timestamps, and
extended attributes, but the hierarchy itself is metadata. The lower object
layer stores immutable chunks and manifests; it does not contain an innate
filesystem tree.

```mermaid
flowchart TD
    Path["/projects/report.pdf"] --> FileVersion["File version 42"]
    FileVersion --> FileManifest["File manifest"]
    FileManifest --> ChunkA["Chunk A"]
    FileManifest --> ChunkB["Chunk B"]
    FileManifest --> ChunkC["Chunk C"]
    Directory["Directory /projects"] --> Path
    Directory --> Imagery["imagery/ directory<br/>node 18"]
```

ArkFS should support file versioning as a first-class metadata feature. A file
version is a retained pointer to a file manifest plus file metadata. Updating a
file creates a new file version and publishes an updated directory or
filesystem root. Unchanged chunks remain shared.

```mermaid
flowchart LR
    File["/projects/report.pdf"] --> V40["Version 40<br/>chunks A B C"]
    File --> V41["Version 41<br/>chunks A B D"]
    File --> V42["Version 42<br/>chunks A E D"]
```

This model gives ArkFS:

* cheap per-file version history;
* cheap directory and filesystem snapshots;
* copy-on-write updates for changed file ranges;
* metadata-only rename, move, and copy when payload bytes do not change;
* policy-driven retention before old manifests become garbage-collection
  candidates.

Versioning is therefore not automatic in raw ArkObj API operations, but it is native
to the ArkFS metadata model. The required invariant is that every visible file
version references a manifest whose chunks satisfy the configured durability
policy.

## ArkObj API

The ArkObj API is the native internal interface for payload, manifest,
lifecycle, and placement operations. It is an interface of ArkObj, not another
subsystem or service view. ArkTxn and service views supply transactional names,
version publication, and access semantics. The planned operations and
contract boundaries are in
[`ark-obj-architecture.md`](ark-obj-architecture.md#arkobj-api).

## OCI Registry View

An OCI registry is a first-class service view for container images and other
OCI artifacts. ArkObj stores content-addressed layer and configuration blobs
and immutable manifests. ArkTxn owns repository names, tag pointers, and
publication metadata. A pull resolves a tag to an OCI manifest or image index,
then retrieves its referenced blobs. The OCI Distribution API is the external
compatibility contract; registry conformance remains an implementation gate.
See the [ArkObj architecture](ark-obj-architecture.md#oci-registry) for the
storage and publication boundaries.

## Kubernetes CSI Volumes

Kubernetes integration is a CSI driver over supported Ark services. BK-QCOW
provides raw block volumes or block volumes formatted and mounted with a
filesystem. ArkFS may provide filesystem volumes once its mount and
concurrency semantics are implemented. CSI maps provisioning, attachment,
node publication, expansion, and snapshots to those service capabilities.
ArkObj stores immutable backing state, but it does not itself provide a
mountable writable volume. The [ArkObj architecture](ark-obj-architecture.md#kubernetes-volumes)
defines this boundary.

## Additional Service Views

Additional views should reuse the appropriate Ark Fabric subsystems rather than
reimplementing object storage, transactions, logs, or WAN transport.

| Service view | Fit |
|---|---|
| Backup repository | Incremental backups become changed chunks only |
| CDN / HTTP delivery | Edge cache misses can multi-source chunks |
| ML model registry | Large immutable model shards and datasets |
| Data lake | Parquet and ORC range reads over S3-compatible API |
| Language package registry | npm or Maven package protocols can reuse immutable payload storage. |
| WAL/archive store | Closed log segments are immutable objects |

Services with rich transactional semantics, such as SQL databases, LDAP, and
FoundationDB-like key-value stores, should not be forced into the object layer.
They can use ArkObj for large values, checkpoints, archives, and
backups while ArkTxn provides transactional metadata or application keyspace.

## Garbage Collection And Compaction

ArkObj traces reachability from published service roots, retained versions,
protected ArkLog archives, and policy pins. It reclaims unreachable records
after retention and safety windows. Local BookKeeper compaction can move live
records without changing stable object IDs. See
[`ark-obj-architecture.md`](ark-obj-architecture.md#retention-and-compaction)
for the service-specific roots and storage boundary.

## Implementation Priorities

1. Define the content-addressed ArkObj records, native ArkObj API,
   and invariants.
2. Implement the Ark Control Catalog and initial FoundationDB-backed ArkTxn
   service metadata and transactional outbox.
3. Implement site-local Ark Edge services over local BookKeeper clusters.
4. Implement BookKeeper-backed ArkLog append, replay, retention, and segment
   dissemination.
5. Implement ArkTxn range ownership, epochs, routing, domains, and hot-replica
   materialization.
6. Implement the Ark scheduler with standard QUIC and basic site selection.
7. Build transparent S3 mode as the first user-facing production proof.
8. Add garbage collection, compaction, repair, and policy convergence.
9. Build BK-QCOW with NBD compatibility and a path to `ublk`.
10. Add capability-gated MPQUIC, hedged reads, and adaptive multi-source
    scheduling.
11. Add an OCI registry, backup and ArkFS views, then a CSI driver over
    validated block and file volume capabilities.

## Key Invariants

* Visible manifests and roots only reference chunks that satisfy minimum
  durability policy.
* Data is durable before a commit that makes it visible.
* Writes publish new roots; they do not overwrite existing objects.
* A writable BK-QCOW volume has one authoritative writer epoch.
* A key range has one authoritative ArkTxn owner per epoch.
* ArkTxn owners execute transaction logic and conflict checks exactly once.
* Every ArkTxn write produces its outbox record in the same FoundationDB
  transaction as the application mutations.
* ArkLog mutation segments are immutable and ordered.
* Hot replicas apply mutation segments in order.
* Reads identify the required generation or commit token and accept data only
  from replicas valid for that generation or token.
* Object checksums are verified when data crosses trust or site boundaries.
* Physical BookKeeper locations can change without changing logical object IDs.
* Garbage collection starts from retained roots, retained log segments, live
  range watermarks, and policy pins, not synchronous write-path refcounts.
* Visible ArkFS file versions only reference manifests whose chunks satisfy
  durability policy.

## Accepted Initial Decisions

| Topic | Decision |
|---|---|
| Metadata ownership | ArkTxn owns service metadata; the Ark Control Catalog owns bootstrap and global control state. |
| Chunk identity | Content-addressed by default, with opaque IDs allowed by security policy. |
| Durability classes | Start with `ARK_REPLICATED`, `ARK_LOCAL_ASYNC`, and `ARK_SYSTEM`. |
| QUIC implementation | Use upstream Cloudflare quiche for standard QUIC; keep the pinned Ark Sites research fork behind the MPQUIC capability gate. |
| MPQUIC requirement | Never required for correctness or baseline interoperability; use only when both peers, path inventory, policy, and telemetry qualify. |
| Deduplication | Scope by tenant and encryption-key deduplication domain; disable cross-tenant sharing by default. |
| Initial sizes | 64 KiB transfer frame, 1 MiB general chunk, 128 MiB manifest span, 128 KiB BK-QCOW cluster, and 4 KiB subcluster. |
| First proof | Build CAS and ArkObj API internally, then ship transparent S3 as the first user-facing production proof. |
| ArkLog storage | Use BookKeeper directly for active ordered streams; archive closed compacted segments to ArkObj when policy requires. |
| ArkTxn reads | Expose `LINEARIZABLE`, `AT_LEAST(token)`, `READ_YOUR_WRITES`, `BOUNDED(max_lag)`, and explicit stale `LOCAL`; default to `LINEARIZABLE`. |

These decisions are initial implementation contracts, not unresolved
questions. Their measurable revisit triggers live in the focused subsystem
documents.
