Skip to content

Vault (Log Archive)

Archive suppressed log lines to cloud storage and recover them instantly

What Is Vault?

The Vault is where the suppressed lines go: a complete archive, in your own cloud storage, of exactly the log lines a summary stands for, in open columnar files you can query with any SQL engine. Forwarded lines are not duplicated into it; they already reach your platform. Every summary carries a link that opens its originals from the Vault in one click, and can replay them into your log platform when you need them there.

Key Benefits:

  • Complete history of what was suppressed: every suppressed log line is archived, in full
  • Instant recovery: Retrieve original logs in seconds
  • Cost-effective: Cloud storage archiving at ~$0.02/GB/month
  • Multi-cloud: Supports AWS S3, Google Cloud Storage, and Azure Blob Storage

Need a complete copy of every log line, suppressed or not, in storage you control? See Raw Archive below.

Enabling Vault

Add the vault configuration under the configMap section in your values.yaml:

AWS S3

configMap:
  forwarder:
    vault:
      enabled: true
      provider: "s3"
      bucket: "my-log-archive"
      prefix: "logs"              # Optional: object key prefix
      region: "us-east-1"         # AWS region (default: us-east-1)

Authentication: Uses the node's IAM role or IRSA (recommended). No credentials needed in config. The role must allow s3:PutObject, s3:GetObject, s3:DeleteObject, and s3:ListBucket on the vault bucket.

Google Cloud Storage

configMap:
  forwarder:
    vault:
      enabled: true
      provider: "gcs"
      bucket: "my-log-archive"
      prefix: "logs"

Authentication: Uses Workload Identity (recommended) or node service account. The identity must have the Storage Object Admin role on the vault bucket (covers create, get, delete, and list).

Azure Blob Storage

configMap:
  forwarder:
    vault:
      enabled: true
      provider: "azure"
      bucket: "my-log-container"   # Azure Blob Storage container name
      prefix: "logs"
      azure_storage_account: "mystorageaccount"

Authentication: Uses Managed Identity (recommended) or storage account key. The Managed Identity needs one Azure RBAC role:

  • Storage Blob Data Contributor on the storage account or container: allows Edge to read, write, and delete archived log files

If using a storage account key instead of Managed Identity, only the key is needed, no RBAC roles required. Note: The bucket field maps to an Azure Blob Storage container name (the top-level grouping inside a storage account).

S3-Compatible Storage (MinIO)

Vault also supports any S3-compatible storage backend:

configMap:
  forwarder:
    vault:
      enabled: true
      provider: "s3"
      bucket: "my-log-archive"
      s3_endpoint: "https://minio.internal:9000"
      s3_force_path_style: true    # Required for MinIO
      region: "us-east-1"

Durability and Retention Settings

The defaults archive every suppressed line without tuning. These are the knobs for how the archive behaves while cloud storage is slow or unreachable, and how long it is kept:

configMap:
  forwarder:
    vault:
      flush_interval: "5m"      # longest a batch waits before it is archived
      buffer_size: 10000        # or this many lines, whichever comes first
      spool_max_mb: 1024        # disk per node for lines the Vault has not taken yet
      buffer_wal: true          # lines in a batch survive an Edge pod crash
      cloud_buffer_wal: true    # ...and the loss of the node itself
      cloud_buffer_wal_interval: "30s"
      retention_days: 0         # 0: keep the archive forever
  policy:
    suppression:
      require_archive: true     # suppress only while the Vault is taking writes
  • flush_interval / buffer_size: a backfill link works the moment the summary is written, before its batch is archived, so there is no reason to archive more often than the default; larger batches make fewer, more efficient archive files
  • spool_max_mb: how much the node’s disk holds for the archive while cloud storage is unreachable (the chart’s daemonSet.spoolDir, which survives a pod restart and a node reboot); it is uploaded when storage returns. With require_archive on, suppression has already paused by then and every new line is going to your platform, so the bound only ever concerns lines suppressed before storage went away; if it is reached, the Edge pod says so once, with the amount
  • buffer_wal / cloud_buffer_wal: lines accepted for the archive but not yet uploaded are kept on disk and, every cloud_buffer_wal_interval, copied to the bucket under <prefix>/_wal/, so they outlive the node they were written on; the hourly archive maintenance folds that copy into the archive
  • retention_days: the compactor deletes archived data older than this; unset means the archive grows forever and your bucket’s own lifecycle rules are the only bound

Retrieving Archived Logs

Every suppression summary written by Edge includes a backfill_url field. Click the link to open the Replay Viewer with all original suppressed logs for that pattern and time window.

Every backfill link opens on the lines it names. Auditty suppresses a line only once the Vault can hold it (require_archive, on by default): while cloud storage is unreachable, or the license has expired, Edge pauses suppression and sends every line straight to your log platform, resuming the moment a write to the Vault lands. A summary is only ever written for lines that are archived. Lines accepted for the archive wait on the node’s own disk (up to spool_max_mb, 1 GiB by default) across pod restarts and node reboots, and cloud_buffer_wal keeps a copy in the bucket itself, so the path from suppression to archive to backfill holds end to end.

Example summary (JSON format):

{
  "message": "Health check passed",
  "auditty": {
    "suppressed": 100,
    "window": { "start": "2026-01-15T09:00:00Z", "end": "2026-01-15T09:05:00Z" },
    "duration": "5m0s",
    "fingerprint": "a1b2c3d4",
    "backfill_url": "https://<your-hivemind-url>/backfill/replay?token=<TOKEN>&sig=<SIG>"
  }
}

No extra tooling required. The backfill URL is self-contained and signed. The link format is unchanged regardless of access level; verification is transparent to the person sharing the link.

Access control: Admins can require viewer identity verification (Google/Microsoft sign-in) or full authentication before viewing backfill URLs. See User Management & Access → Backfill Link Access Control for details.

Programmatic / Agentic Backfill API (JSON)

Raw (including suppressed) archived logs are available as JSON for scripts, agents, and compliance tooling, no browser required. For automated/agentic workflows use the versioned public API under /api/v1/backfill/*. These endpoints return JSON only and carry a stability guarantee (see below); build against them rather than the HTML backfill page or the UI’s internal routes.

Authentication is the token + sig pair already embedded in every suppression summary’s backfill_url (signed, tenant-scoped, and a capability in itself). No session or separate API key is required, so an agent can go straight from a Datadog/Splunk record to its originals.

Endpoints

  • GET /api/v1/backfill/status: whether the replay is ready to read. Returns { status } where status is loading | ready | empty | failed | unavailable. Pass bust=1 to force a re-fetch of an empty result.
  • GET /api/v1/backfill/page: paginated raw original records. Params: page (1 to 1000, default 1), search (substring filter on the log line), sort (asc | desc, default desc).
  • GET /api/v1/backfill/histogram returns time-bucketed counts: { histogram: [{ t, count }] }.
  • GET /api/v1/backfill/integrations: enabled restore targets for the tenant: [{ id, platform, name }], with no keys or config. Like replay, it answers 403 unless public replay is enabled.
  • POST /api/v1/backfill/replay: restore the originals into an integration (e.g. Datadog). Body: { token, sig, integrationId, force? }. Idempotent; streams progress when the request sends Accept: text/event-stream.

Read flow: poll status until ready, then page through the records:

# token + sig come from the summary's backfill_url
GET /api/v1/backfill/status?token=<TOKEN>&sig=<SIG>
{ "status": "ready" }            # retry while "loading"

GET /api/v1/backfill/page?token=<TOKEN>&sig=<SIG>&page=1&sort=desc
{
  "records": [
    {
      "timestamp": "2026-01-01T00:00:01.123Z",
      "namespace": "payments",
      "workload": "api",
      "fingerprint": "a1b2c3d4",
      "node_id": "node-1",
      "source_uri": "...",
      "suppressed": true,
      "log_line": "Health check passed",
      "preserved_json": "{...}"
    }
  ],
  "totalCount": 100,
  "hasMore": false,
  "page": 1,
  "pageSize": 100,
  "queryTimeMs": 7
}
# paginate while hasMore === true

Restore flow: list targets, then replay the originals back into your platform:

GET /api/v1/backfill/integrations?token=<TOKEN>&sig=<SIG>
[ { "id": 12, "platform": "datadog", "name": "DD - prod" } ]

POST /api/v1/backfill/replay
{ "token": "<TOKEN>", "sig": "<SIG>", "integrationId": 12 }

{ "totalSent": 100, "totalFailed": 0, "durationMs": 842 }
# already replayed (idempotent): { "alreadyReplayed": true, "totalSent": 100, ... }
# pass "force": true to re-send

Versioning & stability

  • Stable contract: response shapes under /api/v1/backfill/* are frozen. We only make additive, backward-compatible changes (new optional fields).
  • No breaking changes to v1: any incompatible change ships as a new /api/v2/ namespace, and v1 keeps working.
  • Tenant-isolated: the signed token scopes every call to a single tenant; a forged path or fingerprint cannot cross tenant boundaries.
  • POST /api/v1/backfill/replay and GET /api/v1/backfill/integrations require admin-enabled public replay (Settings → Backfill Replay); when disabled they return 403.

Alternative: by fingerprint (no summary needed). To query the archive for any fingerprint and date range, use GET /api/vault/logs?fingerprint=…&startDate=YYYY-MM-DD&page=1. This call is authenticated with your Auditty session (not a backfill token) and returns { status, fingerprint, start, end, records } with the same record shape; { status: "loading", retryAfterMs } (HTTP 202) while a cold fingerprint is being fetched, 503 while the archive is unavailable.

Field Predicates: Filtering by Value

The summary's ids object indexes only searchable correlation keys; bare magnitudes and measurements (latencies.kong, bytes_sent) are archived in full but aren't listed there. A field predicate lets you filter a backfill down to only the rows where a specific field in preserved_json matches a value or comparison, no need to promote the field first, and no need to page through every archived record by hand.

In the Replay Viewer, use the filter box above the table: enter a field name (dotted paths like latencies.kong work directly, no need to escape the dot), pick an operator, and a value.

Supported operators:

  • = / !=: exact match / not-equal (string or numeric)
  • > / >= / < / <=: numeric comparison
  • contains: substring match
Example
latencies.kong > 100    # only rows where the kong measurement exceeded 100

Signed and shareable: the predicate is embedded in the signed backfill token itself (GET /api/backfill/refine mints a fresh token when you set or clear a filter), so a refined link filters identically for anyone who opens it, and can’t be edited client-side to see data outside the token’s original scope. Refining or clearing a filter reuses the already-fetched data; it does not re-run the backfill from scratch.

This works with the paginated table view and the time histogram alike, and with the JSON API (GET /api/v1/backfill/page) once the predicate is baked into the token you pass.

Make a repeated search permanent:

If you find yourself filtering the archive by the same field again and again, promote it: adding the field to promote_numeric_fields indexes it in every summary’s ids object, searchable directly in your logging platform, no archive query needed, and makes it eligible for real-time outlier passthrough.

Global Search: Field Names and Words

The Vault Explorer’s search box accepts free text, not just fingerprint hashes, and one search is answered from two places at once. Click any match to open its archived logs in the Replay Viewer, where a field predicate narrows to the values you are after.

Signalled patterns, by field name

Type a field name (trace_id, latencies.kong, error.code) and Auditty searches the patterns Edge has reported as anomalies (new patterns and rate spikes, kept for 90 days), returning up to 20 that carry that field, within the Explorer’s selected date range, each with its shape, namespace/workload and last-seen date. What a signal carries is the pattern’s structure alone: in JSON and logfmt logs every value, the message included, is a placeholder, and a plain-text line exports no template at all. So a field name matches here; message wording, a host, a request id or a user id does not, because none of it is stored in Auditty.

Every pattern in your cluster, by its words

The same search is also asked of your cluster. Edge keeps, inside the cluster, every pattern its nodes have seen and matches your words against them there; what comes back to Auditty is the matching patterns’ fingerprints, the workloads that emit them and, for structured lines, their field shape, so the words of a pattern stay in the cluster. This finds plain-text patterns by their wording (connection refused, retrying checkout), which is most of what an operator hunts for. The card under the search box appears when a cluster has checked in within the last hour, and says who answered and from how many patterns, or why not:

Card saysMeaning
N patterns in your cluster contain "…"Answered. Click one to see which days of the archive hold it
Your cluster predates pattern searchEdge in this cluster is older than 1.7.5; only the patterns above are searchable until it is upgraded (the Edge API Deployment is what answers)
Your cluster is busy; the search was not askedToo many replays and searches are waiting for this cluster; try again in a minute
Your cluster did not answerNothing in your cluster claimed the search within a minute: check that Edge is running there and can reach Auditty

A repeated search within five minutes is served from the first answer. Right after the Edge API Deployment restarts in a cluster its list of patterns is empty and refills over the following minutes; the card says so.

A pattern that matches neither way is still fully archived and reachable by its fingerprint or via the backfill link on its suppression summaries. To search by a value rather than a pattern, open a pattern and add a field predicate in the Replay Viewer: that reads the archive in your own account directly.

Vault Explorer (Auditty UI)

Beyond following backfill links from summaries, you can browse all archived logs directly in Auditty. Navigate to Vault in the Auditty dashboard:

How It Works:

  1. Open Vault in Auditty
  2. Browse archived files by date and fingerprint; nothing to register
  3. Click any fingerprint to open the Replay Viewer for that pattern and time range
  4. Search, paginate, and inspect original suppressed log lines

Vault data appears automatically after the first Edge upload, no manual credential configuration needed. Just ensure the Edge service account has the required permissions for your cloud provider (see the Enabling Vault section above).

Replay to Platform

The Replay Viewer lets you send archived logs directly back to your logging platform. Supported platforms include Datadog, Splunk, Elasticsearch, Coralogix, Grafana Loki, New Relic, Sumo Logic, and any custom endpoint via Webhook. This is useful when you need the suppressed logs re-ingested for investigation or compliance.

How Replay Works:

  1. Open a backfill URL from a suppression summary or the Vault Explorer
  2. Click the Replay button and select your integration; logs are sent in batches to your platform immediately
  3. A progress bar tracks delivery in real time

No Auditty account needed (once enabled). When public replay is turned on, the signed backfill URL grants replay access; an on-call engineer can click the link from Datadog or Splunk and replay immediately, no login required.

Guardrails:

  • Idempotent: Replaying the same logs to the same integration is automatically detected; you see a confirmation that the records were already sent, with no duplicate ingestion. This works across both backfill URLs and the Vault Explorer.
  • Opt-in, off by default: Public replay (clicking a backfill URL to replay without login, including POST /api/v1/backfill/replay) is disabled by default. An admin enables it per tenant in Settings → Backfill Replay; while off, replay is only available through the Vault Explorer (requires login) and the public endpoint returns 403. Viewing access is controlled separately via Settings → Backfill URL Access (Open, Verified, or Locked).

Query the Archive with SQL (Open Table Format)

The Vault is a standard columnar archive in your bucket under a Hive-style partition layout (tenant/year/month/day/hour), open by construction, no proprietary format. In Vault Explorer, click Query with SQL, pick your engine, and copy the generated setup statement for your registered bucket. The full archive (suppressed originals included) is queryable with SQL, Spark, or any engine that reads open columnar data, with partition pruning and no Auditty components in the read path.

  • Amazon Athena: partition projection: no crawler, no MSCK REPAIR, new vault files are queryable the moment they land
  • Trino / Presto: hive connector external table
  • Apache Iceberg (Spark SQL): creates an Iceberg table and registers the existing vault files in place via add_files (no copy, no rewrite)

Iceberg registrations are snapshots. Schedule a refresh:

Athena and Trino list the bucket at query time, so they are always current with zero maintenance. Iceberg is different: add_files registers the files that exist at registration time. New vault files are not visible until re-registered, and the Auditty compactor periodically merges small files and deletes the originals; a stale Iceberg registration will then error on missing files. Refreshing is cheap (metadata only, no data copied): re-run add_files on a schedule that matches the hourly compaction; files already registered are skipped, so a re-run is safe.

No lock-in, by design:

Your data outlives any vendor decision, including ours. The schema is stable (timestamp, namespace, workload, fingerprint, node_id, source_uri, suppressed, log_line, preserved_json) and identical across Edge-written and compacted files. S3, GCS (gs://), and Azure (az://) layouts are handled per provider. Teams standardized on Apache Iceberg catalogs (Glue, Nessie, REST, Snowflake external tables, BigQuery BigLake) get a real catalog entry over the same files.

Raw Archive (self-managed, full-fidelity)

Some teams need a complete, untouched copy of every log line in storage they control, independent of Auditty, for compliance or long-term retention. The raw archive is an optional, node-local copy: it writes the complete pre-suppression stream (every line, suppressed or not) to plain-text files alongside the reduced stream that flows to your observability backend.

How it differs from Vault:

  • Vault is written and indexed by Auditty into your cloud storage, with backfill and replay built in.
  • The raw archive writes plain-text files to a node-local directory that your agent (e.g. Vector, Fluentd) tails and ships to storage you own: plain text, nothing for Auditty to write, no lock-in.
  • They are independent: use either, both, or neither.
Enable the raw archive
configMap:
  forwarder:
    raw_archive:
      enabled: true
      directory: /var/log/auditty/raw   # node-local base dir (mount as a hostPath/volume)

The directory layout mirrors your source paths, so one recursive glob captures every stream. Point your downstream agent at it:

Vector example
[sources.auditty_raw]
type    = "file"
include = ["/var/log/auditty/raw/**/*.log"]

The raw archive is best-effort and node-local: it runs beside the reduced stream that feeds your observability backend, at no cost to it, and it is rotation-bounded so its size on the node’s disk is fixed. It is Kubernetes-only and holds your pod log stream as it was before suppression, with no summaries in it.

Summary

  • Vault archives every suppressed log line to cloud storage; non-suppressed events already reach your normal destination and are not duplicated into Vault
  • Multi-cloud support: AWS S3, Google Cloud Storage, Azure Blob Storage, and S3-compatible backends
  • Instant recovery via backfill links in suppression summaries or Auditty Vault Explorer
  • Suppression summaries include a signed backfill URL to view original logs instantly
  • One-click replay: Send archived logs directly to your logging platform from the Replay Viewer
  • Suppress with confidence: every suppressed original is archived to your own storage, and while storage is unreachable suppression pauses and your platform receives every line
  • Automatic optimization: Archived files are consolidated in the background for fast retrieval, no configuration needed
  • Raw archive (optional): Write the complete pre-suppression stream to node-local plain-text files for a downstream agent to ship to storage you control, independently of Auditty