Skip to content

Log Intelligence

How Auditty distinguishes signal from noise: searchable ids in summaries, full originals in the Vault

Key Principle

Three things decide what happens to a line, in this order: your rules, then what Auditty has learned about the pattern, then the defaults. A rule you write is final; nothing automatic reverses it. And a suppressed line always has a home: the Vault, or your log platform while the Vault cannot take it (require_archive, on by default).

The Three Layers

1. Your Rules (Always Respected)

Explicit rules you define are always followed:

rules:
  # Suppress known noisy patterns
  - name: suppress_healthchecks
    action: suppress
    match:
      entryIncludes: "GET /health"

  # Preserve audit logs (never summarize)
  - name: preserve_audit
    action: preserve
    match:
      entryIncludes: "audit_event"

If you have a rule, Auditty follows it. A line a suppress rule takes goes to the Vault; while the Vault cannot take it, it goes to your log platform instead.

Bootstrap preserve rules from your existing alerts (all platforms, 1.7.1+):

Go to Integrations in the Auditty dashboard, connect your platform, then click Sync Monitors on its card: Auditty fetches your log-alerting objects (Datadog log monitors, Splunk saved-search alerts, Kibana alert rules, Coralogix log alerts, Grafana Loki ruler rules, New Relic NRQL conditions, Sumo Logic log monitors) and translates each alert’s query into a preserve rule, so every alert keeps firing on the lines that power it. Translation errs broad by design (a preserve rule can only add forwarding); queries it can’t translate faithfully are listed for your review. Review the generated rules, copy the YAML into your Edge config’s rules: list, and re-sync when alerts change. Edge itself stays vendor-agnostic; the integrations live entirely in Auditty, and the output is plain Auditty rules YAML.

2. Smart Defaults (When No Rule Matches)

For logs without explicit rules, Auditty applies intelligent defaults:

Log TypeDefault Behavior
First occurrence (new pattern)Forward
Rate spike (5x+ above baseline)Suppression threshold doubled: more of the pattern is forwarded
Repetitive INFO/DEBUGSummarize
Errors (without rule)Forward

High-Severity Logs Are Preserved by Default

Auditty automatically preserves high-severity logs (WARN, ERROR, CRITICAL, FATAL, PANIC, ALERT, EMERGENCY): they are always forwarded unless you explicitly create a suppress rule for them. It is a promise about what Auditty recognizes: a warning written in a form it does not read is treated as an ordinary line, and a severity rule still works only on levels it reads.

Severity is read from the first line of an entry in every common form: JSON ("level": "error", "severity", and a "status" holding a level word), logfmt (level=error, severity=), bracketed ([ERROR]), dashed (- ERROR -), a bare level word near the start of the line, and the container’s stderr stream when nothing else says. A stack trace joined to its first line takes that line’s level, so a 150-frame ERROR trace is protected whole.

3. The Identity Ledger

When logs are summarized, Auditty indexes the varying correlation keys under a single ids object, one entry per searchable id. An id is a declared identity key, a shape-detected key (uuid, a hex trace id, a prefixed id such as cus_…, internal_ip/public_ip, MAC, JWT), or a value under an id-ish key name. Bare magnitudes and dimensions (bytes_sent, duration_ms, status, method) are context, kept in full in the Vault, not listed in ids. Every id maps directly to an array of its complete distinct values; a window holds up to 50 distinct identity groups, then rolls to a new summary, so ids are enumerated in full rather than sampled. Every id is searchable, and the exact originals, including the excluded fields, stay one click away via backfill_url:

JSON:

{
  "message": "DNS query",
  "auditty": {
    "suppressed": 100,
    "window": { "start": "2026-01-15T09:00:00Z", "end": "2026-01-15T09:05:00Z" },
    "duration": "5m0s",
    "fingerprint": "a1b2c3d4",
    "class": "network",
    "namespace": "prod",
    "workload": "dns-proxy",
    "ids": {
      "dest_ip": ["10.0.10.1"],
      "src_ip":  ["10.0.1.5", "10.0.1.12", "10.0.1.20", "10.0.1.33", "10.0.1.41"]
    },
    "backfill_url": "https://<your-hivemind-url>/backfill/replay?token=<TOKEN>&sig=<SIG>"
  }
}

Reading the summary: 100 DNS queries collapsed into one line. Every suppressed query went to dest_ip: 10.0.10.1 (a single-element array) and originated from five distinct sources; src_ip lists all of them. Each id maps directly to an array of its complete distinct values; nothing is sampled away. Non-id context (query names, byte counts) is not indexed here but is preserved in full; replay the exact originals via the backfill_url.

Plain text:

DNS query for <str> took <decimal> ms [auditty: suppressed 100 over 5m (09:00:00-09:05:00) | namespace:prod | workload:dns-proxy | fingerprint:a1b2c3d4 | dest_ip=10.0.10.1 | src_ip=10.0.1.5,10.0.1.12,10.0.1.20,10.0.1.33,10.0.1.41 | backfill_url:https://…]

What Gets Preserved in Summaries

When multiple logs are compressed into one summary, Auditty keeps:

KeyWhat It Means
messageReadable headline: the masked template with <num>/<decimal>/<str>/<id> placeholders for free-text/logfmt, or an HTTP subject / natural message field for JSON. It reads as the log does, not as a notice
suppressedHow many logs were compressed
window.start / window.endExact time range the compressed logs spanned
durationTime span of the compressed logs
fingerprintPattern ID (use for backfill)
classSemantic classification: http, db, network, auth, cache, security, k8s, files, errors, systemd, ci_cd; absent when nothing matched
idsOne entry per identifier, each mapping directly to an array of its complete distinct values. Searchable via auditty.ids.<name>
backfill_urlOne-click signed link to retrieve original logs from vault

How the ids object reads:

Each id appears once: declared identity keys, shape-detected ids, and values under id-ish key names; magnitudes and dimensions are kept in the Vault, not indexed here. Every id maps directly to an array of its complete distinct values. No occurrence count is attached to a value, so a unique identifier reads as unique:

"ids": {
  "user_id":  ["user_17", "user_42", "user_88"],
  "trace_id": ["abc123", "def456", "ghi789"],
  "src_ip":   ["10.0.0.1", "10.0.0.2", "10.0.0.3"]
}

Every id is searchable: auditty.ids.user_id:user_42, auditty.ids.src_ip:10.0.0.3. Even a unique-per-event id (e.g. a per-request trace id) is enumerated in full across summaries rather than sampled. The complete set for any window is also one click away via the backfill link.

Anomaly Detection

Auditty detects four kinds of anomaly on a node:

AnomalyWhat It Means
first_occurrenceA pattern seen for the first time (or fewer than six times). Not raised in the first 30 minutes after Edge starts on a node (a fresh deploy’s new patterns are not anomalies), nor for a pattern under two minutes old unless it is an error: a new pod starting is not an anomaly, a new error during a rollout is
rate_spikeA pattern running at 2x or more of its learned rate. The score rises with the ratio: 2x, 5x and 10x are the three tiers; from 5x the pattern’s suppression threshold is doubled so more of it is forwarded
dormant_resurrectionA pattern quiet for over an hour (dormant_threshold) that reappears
content_changeA pattern whose emitters changed: the set of workloads printing it tripled, or a different workload became its main source (checked every five minutes)

Important:

Anomaly detection informs the default behavior. It does NOT override your rules.

  • If you have a suppress rule --> Log is suppressed (even if anomaly detected)
  • If you have no rule --> a first occurrence is forwarded, and a spike loosens suppression for that pattern

Pattern Correlation (Precursor → Outcome)

Auditty does not just report anomalies; it remembers which patterns historically precede a rate spike in the same workload. When a new pattern, a pattern back after a long silence or a pattern whose content changed is followed within a minute by a rate spike in that workload, the pair is remembered. The second time it happens the pair is emerging; from the third it is reliable, and the Insights page tells you to treat its next firing as an early warning.

Where you see this

On the Insights page, under the anomaly category, a card reads N× reliable precursor of a rate spike (or emerging) for the workload. The card says what the precursor did, “A pattern seen for the first time preceded a rate spike 7 times since 12 Sep 2026, most recently 3 Oct 2026, average lead time 47s”, and offers to write a preserve rule for the workload, because an early warning is only one while its line reaches your log platform live. At most three such cards show at a time, one per workload, and a pair whose average lead time is under 15 seconds is not shown: that is a symptom, not a warning.

Privacy & retention

What is remembered is which pattern preceded which, by fingerprint alone. Pairs are kept for 180 days from their last observation and are isolated per tenant.

Overriding Defaults Per Workload

While high-severity logs are preserved by default, you can override this for specific workloads using rules. This is useful when certain workloads produce known, noisy errors that are safe to suppress.

ConfigMap: Suppress known noisy errors cluster-wide

rules:
  - name: "suppress-noisy-db-errors"
    action: suppress
    scope:
      workloadIncludes: "health-checker"
    match:
      severity: "error"
    description: "Suppress known noisy errors from health-checker"

Annotation: Override per workload

Application teams can use annotations to preserve their logs even when a platform-wide suppress rule exists, because preserve has higher precedence than suppress.

annotations:
  auditty.ai/rule.keep-errors: |-
    action: preserve
    match:
      severity: "error"

Precedence reminder: preserve > skip > intercept > suppress. A workload annotation with preserve always wins over a ConfigMap suppress rule.

Example Scenarios

ScenarioYour RuleAnomaly?Result
Health check spamsuppressNoSuppressed
Health check spamsuppressYes (spike)Suppressed (your rule wins)
New error patternnoneYes (first)Forwarded (smart default)
Known noisy errorsuppressNoSuppressed
Audit logpreserveN/AAlways forwarded
ERROR (no rule)noneNoForwarded (high-severity default)
ERROR (suppress rule)suppressNoSuppressed (your rule wins)
ERROR (suppress + annotation preserve)bothNoForwarded (preserve > suppress)

Recovering Suppressed Logs

Once the Vault is configured, every summary includes a backfill_url: a signed link that opens the Replay Viewer with all the original suppressed logs for that fingerprint and time window. Just click the URL.

How backfill works:

  1. Edge archives suppressed logs to your cloud storage (S3/GCS/Azure)
  2. Every suppression summary contains a signed backfill_url once the Vault is configured
  3. Open the URL: Auditty validates the link, asks your cluster for the archived lines, and shows the original logs; your cloud-storage credentials stay in your cluster
  4. No extra tooling or credentials needed

Summary

  1. Your rules always win over automation, and a suppressed line always has a home: the Vault, or your log platform while the Vault cannot take it
  2. Smart defaults: first occurrences and every WARN-or-worse line forward when no rule matches, and a spike loosens suppression for its pattern
  3. Measure first: a measure rule, or the shadow-mode switch on the Fleet Rules page, runs all of this over live traffic and reports what it would do without suppressing a line
  4. Identity ledger: Every searchable correlation key indexed under ids in summaries; full originals in the Vault
  5. Every original kept: suppressed logs are archived to your own cloud storage and open from any summary’s backfill link, or by fingerprint