5 ArchiveBox Architecture Diagrams
Nick Sweeting edited this page 2026-09-01 15:09:46 -07:00

ArchiveBox Architecture Diagrams

This page is a map of the current execution and persistence paths. The implementation lives primarily in:

  • archivebox/cli/ for CLI entry points
  • archivebox/services/runner.py for crawl and snapshot execution
  • archivebox/crawls/models.py for the Crawl model and its atomic queue transitions
  • archivebox/core/models.py for Snapshot, ArchiveResult, and Snapshot queue transitions
  • archivebox/services/ for bus event projectors
  • abxpkg and abx-plugins for binary resolution and plugin hooks

High-Level Execution Flow

flowchart TD
    ENTRY["CLI, Web UI, REST API, or scheduler"] --> CRAWL["Create or resume a Crawl row"]
    CRAWL --> RUNNER["run_crawl() / CrawlRunner"]
    RUNNER --> DISCOVER["Create or select Snapshot rows"]
    DISCOVER --> EVENTS["Emit crawl and snapshot lifecycle events"]
    EVENTS --> PLUGINS["Run selected abx-plugin hooks"]
    PLUGINS --> PROCESSES["Persist Process rows and hook output"]
    PROCESSES --> RESULTS["Project ArchiveResult rows"]
    RESULTS --> FILES["Write snapshot output files"]
    RESULTS --> SNAPSTATE["Seal or requeue Snapshot"]
    SNAPSTATE --> CRAWLSTATE["Seal, pause, or continue Crawl"]

    EVENTS --> BINREQ["BinaryRequestEvent"]
    BINREQ --> ABXPKG["abxpkg resolution"]
    ABXPKG --> HOST["Compatible host binary"]
    ABXPKG --> MANAGED["Managed install fallback"]
    HOST --> ENV["Project resolved binary into LIB_DIR/env/bin"]
    MANAGED --> ENV

    CRAWL -.-> DB["SQLite database"]
    PROCESSES -.-> DB
    RESULTS -.-> DB
    FILES -.-> STORAGE["archive/users/... snapshot storage"]

ArchiveBox has one normal crawl execution path. CLI commands and web/API actions create or select database rows, then call the same runner. The runner emits lifecycle events, abx-plugin hooks do the extraction work, and service projectors persist processes and results.

Binary discovery and installation always goes through abxpkg. Compatible host binaries are preferred; managed providers are the fallback. Resolved binaries are projected into LIB_DIR/env/bin before programmatic use. LIB_DIR/bin is only a convenience directory for humans.

Persistent Data

flowchart LR
    DATA["ArchiveBox data directory"] --> DB["index.sqlite3"]
    DATA --> ARCHIVE["archive/users/<user>/snapshots/<date>/<domain>/<uuid>/"]
    DATA --> SOURCES["sources/"]
    DATA --> LOGS["logs/"]
    DATA --> LIB["lib/env/bin/ resolved binaries"]

    ARCHIVE --> PLUGINOUT["Plugin-namespaced outputs"]
    ARCHIVE --> META["Snapshot metadata and indexes"]

The database is the source of truth for model state. Snapshot directories contain captured artifacts and rendered metadata. Older collections may also contain legacy timestamp-named snapshot directories.

Crawl Queue Lifecycle

Implemented directly by Crawl in archivebox/crawls/models.py. The database row is the durable state; the runner claims retry_at with a conditional update before it performs side effects, then calls the model's explicit lifecycle methods. There is deliberately no second in-memory state machine that can drift from the row owned by another process.

stateDiagram-v2
    [*] --> QUEUED
    QUEUED --> STARTED: runner claim and valid URLs
    QUEUED --> QUEUED: claimed but not ready
    QUEUED --> SEALED: all existing snapshots finished
    STARTED --> SEALED: all snapshots finished
    QUEUED --> PAUSED: pause requested
    STARTED --> PAUSED: pause requested
    PAUSED --> QUEUED: resume requested
    PAUSED --> PAUSED: not runnable
    QUEUED --> SEALED: explicit seal
    STARTED --> SEALED: explicit seal
    PAUSED --> SEALED: explicit seal
    SEALED --> [*]

A crawl owns a set of snapshots. The runner creates or discovers those snapshots and projects crawl events while the row is STARTED; sealing waits for their normal lifecycle to finish. Pausing also schedules child snapshots to pause, and resuming returns the crawl to the runnable queue. The archivebox://update sentinel remains control-plane work: the runner invokes database maintenance directly and never turns it into a Snapshot.

Snapshot Queue Lifecycle

Implemented directly by Snapshot in archivebox/core/models.py, using the same conditional retry_at claim protocol as Crawl.

stateDiagram-v2
    [*] --> QUEUED
    QUEUED --> STARTED: runner claim and URL is ready
    QUEUED --> QUEUED: claimed but not ready
    QUEUED --> SEALED: all existing results finished
    STARTED --> SEALED: all hook results finished
    QUEUED --> PAUSED: pause requested
    STARTED --> PAUSED: pause requested
    PAUSED --> QUEUED: resume requested
    PAUSED --> PAUSED: not runnable
    QUEUED --> SEALED: explicit seal
    STARTED --> SEALED: explicit seal
    PAUSED --> SEALED: explicit seal
    SEALED --> [*]

The runner creates one queued ArchiveResult per selected hook, executes those hooks through the shared event bus, and seals the snapshot after every result reaches a final status. The narrow search-index maintenance operation on an already sealed snapshot is the intentional exception; it does not reopen or invent a second general lifecycle path.

ArchiveResult Projection

ArchiveResult is not driven by a separate in-memory state machine. The runner creates queued rows, and ArchiveResultService projects ArchiveResultEvent and ProcessCompletedEvent data into them.

flowchart LR
    QUEUED["queued"] --> STARTED["started"]
    STARTED --> SUCCEEDED["succeeded"]
    STARTED --> FAILED["failed"]
    STARTED --> SKIPPED["skipped"]
    STARTED --> NORESULTS["noresults"]
    STARTED -. recoverable wait .-> BACKOFF["backoff"]
    BACKOFF -. resumed work .-> STARTED

succeeded, failed, skipped, and noresults are final result statuses. Each row identifies the plugin and hook that produced it and stores structured output, file metadata, timing, and error details.