# ArkObj Architecture

ArkObj is Ark Fabric's immutable object storage subsystem. It stores chunks,
manifests, and immutable roots across independent sites; verifies their
integrity; and enforces object durability and replica placement policy. It is a
peer of ArkTxn, ArkLog, and ArkNet within Ark Fabric.

Its native interface is the **ArkObj API**.

## Role And Boundaries

| Concern | Owner |
|---|---|
| Immutable payloads, object manifests, immutable roots, replica validity, object durability | ArkObj |
| Transactional bucket keys, file paths, volume heads, versions, leases, and publication pointers | ArkTxn |
| Bootstrap policy roots and global ownership fences | Ark Control Catalog |
| Ordered active streams and replay | ArkLog |
| WAN sessions, paths, and transfer transport | ArkNet |
| Source and target selection within subsystem policy | Ark scheduler |

ArkObj uses site-local BookKeeper storage through Ark Edge. Physical ledger
locations may change without changing stable object IDs. ArkObj may store
closed, immutable ArkLog archives, but active ArkLog appends use BookKeeper
directly. FoundationDB's active transaction files and logs remain on storage
chosen for its latency and I/O needs, not on ArkObj.

```mermaid
flowchart TB
    Views["S3, OCI registry,<br/>ArkFS, BK-QCOW"] --> Obj["ArkObj<br/>immutable objects"]
    Obj --> Edge["Ark Edge<br/>verified site access"]
    Edge --> BK["BookKeeper<br/>local payload storage"]
    Obj --> Net["ArkNet<br/>WAN transfer"]
    Txn["ArkTxn<br/>namespace and heads"] -->|references| Obj
    Log["ArkLog<br/>closed archives"] -.-> Obj
```

## Object Model

ArkObj stores three immutable record kinds:

| Record | Meaning |
|---|---|
| Chunk | Bounded byte range with a stable ID, size, checksum, and generation. |
| Manifest | References chunks, checksums, length, and policy for one logical payload or mapping. |
| Root | Immutable reference to a committed state or set of manifests. |

```mermaid
classDiagram
    class Chunk {
        id
        length
        checksum
        generation
        payload
    }
    class ObjectManifest {
        id
        total_length
        checksum
        policy
        replica_locations
        generation
    }
    class Root {
        id
        manifest_refs
        metadata_refs
        parent_refs
        generation
    }
    ObjectManifest "1" --> "*" Chunk : references
    Root "1" --> "*" ObjectManifest : publishes
    Root "0..*" --> "0..*" Root : derives-from
```

Content-addressed chunk IDs are the default:

```text
chunk_id = hash(canonical_chunk_bytes)
```

The hash algorithm and encoding are versioned. A policy may instead use
opaque random chunk IDs when content equality must not be revealed. Deduplication
is limited to an explicit tenant and encryption-key domain; cross-domain
sharing is disabled by default. Logical IDs never expose mutable physical
BookKeeper locations.

General immutable payloads initially use 1 MiB chunks and a 128 MiB
large-object manifest span. BK-QCOW has its own 128 KiB logical cluster and
4 KiB subcluster sizes. ArkNet's 64 KiB transfer frame is a transport unit,
not an ArkObj record size.

## Write, Publication, And Durability

An ArkObj write becomes eligible for publication only when every referenced
chunk meets the requested minimum site durability. ArkObj then continues
toward the desired replica set in the background. ArkTxn atomically publishes
mutable service names or volume heads that reference the durable manifest or
root. An immutable root alone does not publish an S3 key, file path, or volume
head.

```mermaid
flowchart LR
    Write["Write chunks"] --> Verify["Verify checksum"]
    Verify --> Minimum["Minimum durable sites"]
    Minimum --> Manifest["Commit manifest or root"]
    Manifest --> Publish["ArkTxn publishes name or head"]
    Minimum --> Converge["Converge desired replicas"]
```

The safety rule is that the durability of data referenced by a transaction
must be at least the durability promised for the published transaction. An
unpublished durable chunk can be collected later; a published pointer to lost
data cannot be repaired from that pointer alone. Replica validity is checked
against the requested generation or commit token before serving a read.

## Reads And Placement

ArkObj resolves eligible replicas for each requested chunk or range. The shared
scheduler chooses sources and targets under ArkObj policy. ArkNet carries WAN
traffic; Ark Edge authenticates requests, verifies checksums, and serves
approved local ranges. Large reads may use multiple sites without requiring
peer discovery. Object identity and policy remain stable as local ledgers are
compacted or replicas are repaired.

## ArkObj API

The ArkObj API is the native internal interface to the subsystem. It is a
descriptive interface name, not a separate subsystem or a promise that these
operations have already been implemented.

| Operation group | Planned operations |
|---|---|
| Payload | `put`, `get`, `getRange`, `openStream` |
| Manifest | `commitManifest`, `cloneManifest`, `snapshot` |
| Lifecycle | `pin`, `unpin` |
| Placement | `locations`, `durability`, `prefetch`, `replicate` |

`put` returns a chunk or object ID inside the caller's deduplication domain.
The API does not create S3 keys, file hierarchy, OCI repository tags, or
BK-QCOW volume heads. Those names and visibility rules belong to the service
view and its ArkTxn metadata.

## Service Views And Integrations

The names in this table identify different kinds of surfaces. S3, OCI, and
CSI are standards-based interfaces; ArkFS and BK-QCOW are project-specific
service names. These surfaces can share ArkObj without becoming ArkObj peers.

| Surface | ArkObj use | Other owner or boundary |
|---|---|---|
| S3-compatible API | Object payloads, multipart chunks, and manifests. | ArkTxn owns buckets, keys, versions, and object-lock metadata. |
| OCI registry | Content-addressed layer and artifact blobs, plus stored manifests. | ArkTxn owns repositories, tags, and publication; the registry must follow the OCI Distribution API. |
| BK-QCOW block volumes | Immutable cluster payloads, mapping objects, and roots. | ArkTxn owns volume heads, writer leases, and epochs. |
| ArkFS file service | File payloads, file manifests, and immutable roots. | ArkTxn owns paths, hierarchy, versions, and leases. |
| Kubernetes CSI driver | BK-QCOW block volumes, optionally formatted and mounted as filesystem volumes; ArkFS-backed volumes only after its mount semantics are implemented. | CSI handles volume lifecycle and node publication; Kubernetes owns PersistentVolume and claim orchestration. |
| ArkLog archive | Closed, immutable compacted segments. | ArkLog owns stream positions, replay, and active append storage. |

### OCI Registry

An OCI registry is a first-class service view, not just an incidental artifact
example. The registry maps repository names and tags to OCI manifests or image
indexes; those descriptors reference content-addressed blobs such as image
layers and configuration. ArkObj holds immutable content, while ArkTxn provides
atomic publication of tag pointers and repository metadata. A pull resolves a
tag to a manifest, then reads and verifies the referenced blobs through ArkObj.
The OCI Distribution API is the compatibility contract; support for referrers,
artifact types, and conformance requires explicit implementation and testing.

### Kubernetes Volumes

Kubernetes volume support is a CSI integration, not a new ArkObj storage
subsystem. The CSI driver maps Kubernetes volume lifecycle operations to
BK-QCOW raw block volumes or to filesystems formatted and mounted on those
volumes. ArkFS-backed volumes are a later option once its mount semantics are
implemented. Claims, StorageClasses, attachment, node publication, expansion,
snapshots, and access modes must map to the underlying service's actual
capabilities. ArkObj's immutable chunks alone do not provide a mountable
writable volume. Do not promise multi-writer semantics unless the selected
block or filesystem service implements and validates them. OCI image
distribution and Kubernetes persistent volumes are separate integration
contracts even when deployed for the same workloads.

## Retention And Compaction

Garbage collection starts from published volume heads, S3 versions, OCI
manifests and tags, file roots, snapshots, pinned manifests, retained ArkLog
archives, and other protected references. It traces to manifests and chunks.
Unreachable records are reclaimable after retention and safety windows.
Compaction copies live objects from mostly dead local BookKeeper ledgers and
updates location metadata while retaining stable logical IDs.

## Implementation Order And Checks

1. Define versioned chunk, manifest, root, and ArkObj API contracts.
2. Prove checksums, minimum durability, placement convergence, and repair.
3. Prove ArkTxn publication cannot expose missing referenced chunks.
4. Add S3 first, then the OCI registry, with protocol conformance checks for
   each.
5. Add BK-QCOW and ArkFS service behavior, then expose the supported volume
   capabilities through CSI.
6. Validate retention, garbage collection, and compaction against every
   service view's live roots.

## External Interface References

* [OCI Distribution Specification](https://github.com/opencontainers/distribution-spec/blob/main/spec.md) (reviewed 2026-10-06): registry pull, push, and content distribution contract.
* [OCI Image Manifest Specification](https://github.com/opencontainers/image-spec/blob/main/manifest.md) (reviewed 2026-10-06): image manifests, layers, and descriptors.
* [Kubernetes Persistent Volumes](https://kubernetes.io/docs/concepts/storage/persistent-volumes/) (reviewed 2026-10-06): PV, claim, and raw block concepts.
* [Container Storage Interface specification](https://github.com/container-storage-interface/spec) (reviewed 2026-10-06): driver volume lifecycle and node interface.
