Skip to content

Metrics

Prometheus metrics for monitoring and observability

Accessing Metrics

Every Edge pod exposes Prometheus metrics for what it did to your logs: how many lines it read, forwarded, summarised and archived, per namespace and workload, and how it is doing it. Five of them tell the whole story; the rest are for tuning.

Metrics are exposed on port 2112 at the /metrics endpoint.

From within the pod:

wget -qO- http://localhost:2112/metrics

From Prometheus: Configure Prometheus to scrape the metrics endpoint. Enable prometheusScrape.enabled: true in your Helm values to add scraping annotations automatically.

The 5 Essential Metrics

These five counters tell you everything you need to know about Auditty's core value proposition:

1. Log Volume Ingested (Count)

auditty_events_processed_total

Total number of log events ingested from container runtime since startup.

Labels: namespace, workload

Use: Baseline traffic volume. How many log lines are your workloads producing?

2. Log Volume Ingested (Size)

auditty_event_size_bytes_total

Total bytes ingested from container runtime before processing.

Labels: namespace, workload

Use: Raw data volume. How much bandwidth/storage would you need without Auditty?

3. Log Volume Suppressed (Count)

auditty_events_suppressed_total

Total number of log events suppressed (removed from output).

Labels: namespace, workload, policy, reason

Use: Volume reduction. How many log lines did Auditty eliminate?

Suppression Rate: (events_suppressed / events_processed) * 100%

4. Log Volume Suppressed (Size)

auditty_suppressed_size_bytes_total

Total bytes removed via suppression.

Labels: namespace, workload, policy, reason

Use: Cost savings. How much storage/bandwidth did Auditty save?

Savings Rate: (suppressed_bytes / event_bytes) * 100%

Dollar Savings: Configure your log platform cost per GB under Plan & Usage to see dollar savings on the Dashboard and Metrics pages.

5. Synthetic Events Injected

auditty_synthetic_events_total

Total synthetic events generated (suppression summaries, diagnostics).

Labels: namespace, workload, event_type

Use: Overhead tracking. Auditty replaces N suppressed logs with 1 summary.

Example: Suppress 1,000 logs → emit 1 summary → synthetic_events_total increments by 1

Key Business Metrics

Data Volume Tracking

  • auditty_written_size_bytes_total: bytes that landed at each destination (the container’s log, the Vault, your upstream), by component and destination
  • auditty_synthetic_payload_size_bytes_total: Bytes emitted as synthetic events

Shadow Mode (Measuring)

With shadow_mode: true Edge reads each container’s output without intercepting it and changes nothing: nothing is suppressed, nothing is archived, and kubectl logs shows every line. It runs its rules over every line all the same and counts what they would have suppressed. (Edge 1.7.6; earlier versions intercepted the log and wrote every line back.)

  • auditty_events_shadow_processed_total: lines read in shadow mode, every one of them left in the log
  • auditty_events_shadow_suppressed_total: those of them that would be suppressed live. The two together are the suppression rate shadow mode measured
  • auditty_shadow_processed_size_bytes_total and auditty_shadow_suppressed_size_bytes_total: the same two counts in bytes
  • auditty_events_suppressed_total stays at zero on a node in shadow mode: a measurement is counted as a measurement, not as a suppression

Where you see it in Auditty. The fleet's mode is one line in the header of every page: nothing while every node is live, Shadow mode once every node measures, Shadow mode on N of M when only some do, and Entering shadow mode / Going live with the count of nodes still to switch while a change settles. It is read from what each node last reported, not inferred from the counters.

On the Dashboard, while the nodes measure, a one-line bar under the header says which nodes are measuring, and the headline figures show the projection: the suppression rate shadow mode measured over the lines it read, what it would have suppressed, and what that would save at your configured rate, with nothing is suppressed while measuring where a result's footnote would be. A node on an Edge before 1.7.6 did not count what it read and shows the lines alone, with no rate. A fleet that measured earlier in the window and is live now gets one footnote beside each figure; a live fleet's dashboard says nothing about shadow mode.

On the Fleet Rules page, an admin switches the whole fleet with Measure first / Go live (also from the dashboard bar while it shows). The switch is a fleet rule named measure with the action measure and no scope, so it reaches every node within about a second, is audited and restorable like any fleet rule, and can be narrowed to one cluster. A measure rule with a scope reads only the containers it selects in shadow mode while the rest of the node runs live: the way to try a rule on one workload first. Because shadow mode runs every rule over every line it reads, a suppress rule written while measuring is measured too: under each fleet rule the page says what share of the lines read in shadow mode it would suppress, at what rate, on how many nodes, so a rule is judged on live traffic before it goes live. shadow_mode: true in a node's own configuration is the same thing as a measure rule for everything, and the card says when that, rather than the switch, is what keeps nodes measuring. The Configuration guide has the rule syntax and the troubleshooting steps.

Operational Metrics

These metrics help monitor system health and troubleshoot issues:

Pipeline Health

auditty_events_pending, auditty_intercepted_logs

Error Tracking

auditty_component_errors_total, auditty_interception_failures_total

Performance

auditty_processing_duration_seconds: processing time per line, by component (a sample of one line in 256)

Pattern Management

auditty_patterns_stored, auditty_pattern_evictions_total

Label Architecture

Static Labels (set at startup)

  • environment: Environment name (from config)
  • cluster: Cluster identifier (from config)
  • node_id: Kubernetes node name
  • version: Auditty Edge version
  • execution_id: Unique ID for this process instance

Dynamic Labels (vary by metric)

  • namespace, workload: Kubernetes context (most metrics)
  • component: Component name (forwarder, processor, etc.)
  • policy, reason: Suppression context
  • event_type: Type of synthetic event
  • destination: Output destination (file, s3, etc.)

Data Integrity Verification (Auditty)

Auditty's accounting is arithmetic over the conservation counters above: ingested == forwarded + suppressed, with suppressed events fully represented by both the emitted summaries and the Vault archive. It has no dashboard presence; data integrity is the default, not something to announce. The Auditty team reads it on an integrity view over every tenant, and every release is held to it against a live cluster.

Outlier Passthrough

auditty_outlier_detected_total counts events whose suppression was bypassed because a promoted numeric field’s value fell far outside its fingerprint’s learned normal range; see Suppression → Outlier Passthrough. auditty_suppression_skipped_total{reason="outlier_passthrough"} mirrors the same events from the suppression-decision side, alongside reason="vault_unhealthy" for events that were forwarded because the Vault was not accepting writes; suppression stops rather than summarising into an archive that cannot store the originals.

For operators, the underlying pattern-stability gauges (auditty_fingerprints_active_total / auditty_fingerprints_singleton_total) remain available at the /metrics Prometheus endpoint; a rising singleton share usually means a recent log format change made templating too granular for a workload. See Suppression → Promoting Numeric Fields for the related identity-ledger controls.

Forwarding Auditty Metrics to External Platforms

Auditty Edge emits its own Prometheus metrics (prefixed with auditty_) for its own monitoring and for Auditty. If you are using a Prometheus remote-write pipeline or an agent (e.g. Datadog Agent, Splunk OTEL Collector, Grafana Agent) that scrapes all Prometheus endpoints and forwards them to an external platform, Auditty metrics will be forwarded too.

This can result in unexpected ingestion costs on your external platform, especially with high-cardinality label combinations (namespace × workload × policy).

Recommended actions:

  • Filter by prefix: Exclude metrics matching auditty_* from your remote-write or agent scrape config if you do not need them externally.
  • Datadog: In your Datadog Agent config, add auditty_* to exclude_metrics in the OpenMetrics check, or use autodiscovery annotations to skip the Auditty metrics endpoint.
  • Splunk OTEL: Use a filter processor in your pipeline to drop metrics with the auditty_ prefix.
  • Grafana Agent / Prometheus: Use metric_relabel_configs to drop metrics matching auditty_.*.
  • Keep selectively: If you want some Auditty metrics externally (e.g. for cost-savings dashboards), allowlist only the 5 essential metrics and drop the rest.

Summary

  • 5 essential metrics tell you about ingestion, suppression, and synthetic event overhead
  • Shadow mode measures what your rules would suppress without touching a log
  • Operational metrics help monitor pipeline health and troubleshoot issues
  • Labels provide Kubernetes context for filtering and aggregation
  • Data integrity verification runs continuously behind the scenes in Auditty, confirming that every line ingested was forwarded or summarised, from the conservation counters
  • Metrics are read from any Edge pod, as shown at the top of this page