Skip to content

Installation

Deploy Auditty Edge on Kubernetes, Linux VMs, the CLI, or as a network (OTLP) proxy

Kubernetes (Helm)

Deploy Auditty Edge on any Kubernetes cluster via Helm. The core components are:

  • auditty-edge (DaemonSet): runs on every node to intercept and process container logs. Runs as root with two scoped Linux capabilities (DAC_READ_SEARCH, SYS_PTRACE), not in privileged mode
  • auditty-edge-api (Deployment): non-privileged cluster-level service with RBAC permissions to query the K8s API
  • auditty-edge-compactor (CronJob): created automatically when Vault is enabled (and Edge API is, the default); runs hourly to optimize archived data for fast retrieval

Resource Sizing

Edge runs as a DaemonSet; resources are allocated per node. The right sizing depends on log throughput, not pod count; a single high-volume workload can produce more logs than dozens of lightweight services. The chart’s default (daemonSet.resources) requests 500m CPU and 256Mi and allows up to 1.5 CPU and 768Mi; Edge sizes its own memory to the limit it is given.

ProfileCPU limitMemory limitLog throughput
Standard (chart default)1.5 CPU768MiUp to ~5K lines/sec per node
High throughput2 CPU1Gi5K to 20K lines/sec per node
Extreme3 CPU1.5Gi20K+ lines/sec per node

A container that logs faster than Edge can process on the CPU it has keeps every line: if Edge stays more than five minutes behind, it stops intercepting that container, so its output reaches your platform unreduced, every line in order; says so once in its own log (reason: too_far_behind), and tries again ten minutes later. Seeing that line is the signal to raise the CPU limit or skip the workload.

Prerequisites

  • Helm installed on your system
  • Access to a Kubernetes cluster
  • A namespace to install into: defaults to auditty, but any namespace works via --namespace <name> (add --create-namespace if it does not exist). Edge reads its own namespace at runtime, so self-monitoring is prevented wherever you install it.

Step 1: Prepare values.yaml

Create and configure your values.yaml file. At minimum set hivemind_api_key and cluster_name.

Step 2: Install via Helm

helm upgrade --install auditty-edge oci://ghcr.io/auditty/helm-charts/auditty-edge --version <tag> --namespace auditty --create-namespace --values values.yaml

Use the version tag provided by your CX contact.

Using ArgoCD?

Set repoURL to the full OCI path including the chart name:

repoURL: ghcr.io/auditty/helm-charts/auditty-edge
targetRevision: "<tag>"

What Happens Next

  • auditty-edge pods deploy as DaemonSet (one per node)
  • auditty-edge-api pod deploys as Deployment (one replica)
  • auditty-edge-compactor CronJob is created automatically when Vault is enabled
  • This typically takes a few seconds
  • If pods fail to start, check their logs for more information
  • Log processing begins as soon as the pods are running

Strict / Proxied Clusters

On clusters that force all egress through a forward proxy or enforce a default-deny NetworkPolicy, set both of the following so that Edge can report to Auditty and sign backfill links:

  • proxy.httpsProxy / proxy.httpProxy: routes outbound calls (Vault uploads, Auditty) through your proxy. In-cluster traffic (edge → edge-api, the K8s API) is excluded from the proxy automatically.
  • networkPolicy.enabled: true. Opens the edge → edge-api hop on port 8080, which metrics reporting and backfill links travel over.
  • networkPolicy.restrictEgress stays false unless you have mapped your egress. If you turn it on while using a forward proxy, add the proxy’s CIDR/port to egressHTTPSCIDRs first; proxy traffic uses the proxy port (e.g. 8080), not 443, so the default 443-only rules would block it.

Values.yaml Configuration

The fields you are most likely to set, with the chart’s defaults. Anything you leave out keeps its default; the chart’s own values.yaml (helm show values oci://ghcr.io/auditty/helm-charts/auditty-edge --version <tag>) is the complete reference, including the Vault, suppression and multiline settings documented in their own guides.

Important Configuration

  • hivemind_api_key is mandatory: provide via helm install, values.yaml, or secretRef
  • secretRef: store API key in a Kubernetes Secret instead of ConfigMap
  • cluster_name is required: unique identifier for this cluster in Auditty
  • edgeApi.enabled controls the Edge API component (enabled by default)
  • Configure rules to control which logs are intercepted and processed
# values.yaml for auditty-edge (chart defaults shown)

image:
  repository: public.ecr.aws/o8y0g3i1/space/edge
  tag: ""                      # empty: the chart's appVersion
  pullPolicy: IfNotPresent

commonLabels: {}               # applied to every resource

# Edge configuration, rendered into the auditty-config ConfigMap.
# Most of it hot-reloads on helm upgrade; see the Configuration guide.
configMap:
  create: true
  rollOnConfigChange: false    # true: any ConfigMap change rolls the pods

  cluster_name: ""             # REQUIRED: how this cluster is named in Hivemind
  # env_name: "prod"           # optional: powers the Environment filter in Hivemind
  hivemind_api_key: "placeholder"  # REPLACE, or use secretRef below
  # log_level: info            # debug | info (default) | warn | error
  # shadow_mode: false         # true: measure everything, suppress nothing
  # remote_rules:
  #   enabled: true            # fleet rules from Hivemind, on by default
  enable_rule_annotations: false   # per-workload annotation rules (needs Edge API)

  rules:
    - name: "intercept-prod-databases"
      action: intercept
      scope:
        namespaceIncludes: "prod"
        workloadIncludes: "database"
      description: "Start monitoring all production database workloads"
    - name: "preserve-prod-db-critical-errors"
      action: preserve
      scope:
        namespaceIncludes: "prod"
        workloadIncludes: "database"
      match:
        any:
          - entryIncludes: "CRITICAL"
          - entryIncludes: "FATAL"
          - entryIncludes: "PANIC"
          - entryRegex: "ERROR.*database connection lost"
      description: "Preserve critical errors from production databases"

# Store the API key in a Secret instead of the ConfigMap
secretRef:
  enabled: false
  secretName: "auditty-secret"
  apiKeyName: "hivemind_api_key"

# The per-node DaemonSet
daemonSet:
  serviceAccount:
    create: true
    name: "auditty-edge"
  priorityClassName: ""
  nodeSelector: {}
  annotations: {}
  resources:
    requests:
      memory: "256Mi"
      cpu: "500m"
    limits:
      memory: "768Mi"
      cpu: "1500m"
  # Directory on the node where lines not yet archived wait while cloud storage
  # is unreachable. Survives pod deletion and node reboot; at most spool_max_mb
  # (1 GiB) of disk.
  spoolDir: /var/lib/auditty/edge
  # Needed on AppArmor-enforcing nodes (GKE, Ubuntu); harmless elsewhere.
  apparmorUnconfined: true
  # Needed on SELinux-enforcing nodes (Bottlerocket).
  seLinuxOptions:
    type: spc_t
  affinity: {}
  tolerations: []

# The cluster-level Edge API (keep replicas at 1)
edgeApi:
  enabled: true
  replicas: 1
  resources:
    requests:
      memory: "256Mi"
      cpu: "250m"
    limits:
      memory: "768Mi"
      cpu: "1000m"
  nodeSelector: {}
  tolerations: []
  affinity: {}

# The hourly Vault compactor (created only when the Vault is enabled)
compactor:
  schedule: "0 * * * *"
  resources:
    requests:
      memory: "256Mi"
      cpu: "500m"
    limits:
      memory: "2Gi"
      cpu: "1000m"

# The OTLP / HTTP network proxy, off by default (see "Edge Proxy" below)
edgeProxy:
  enabled: false

prometheusScrape:
  enabled: false               # true: add prometheus.io/scrape annotations

# Egress proxy (optional): for clusters where all outbound traffic must route
# through a forward proxy. Injected into edge, edge-api, and the compactor.
# In-cluster destinations (loopback, metadata IP, RFC1918 ranges, .svc /
# .cluster.local, kubernetes.default) are added to NO_PROXY automatically, so
# internal traffic (edge → edge-api, the K8s API) never goes through the proxy.
proxy:
  httpProxy: ""
  httpsProxy: ""
  noProxy: ""        # extra bypass entries, merged on top of the in-cluster defaults

# NetworkPolicy: REQUIRED on clusters with a default-deny policy. Without it the
# edge → edge-api hop (port 8080) is dropped, silently disabling metrics
# reporting and Vault backfill links.
networkPolicy:
  enabled: false
  # Opt-in egress hardening (default false). When true, egress is scoped to only
  # what Edge needs (DNS, the K8s API server, edge → edge-api on 8080, and HTTPS
  # to Hivemind / cloud storage) instead of allow-all. Left off by default because
  # egress topology is environment-specific; if you run a forward proxy you must
  # add the proxy's CIDR/port below before enabling, or all outbound traffic is
  # blocked (proxy traffic uses the proxy port, not 443).
  restrictEgress: false
  dnsNamespace: kube-system          # namespace where CoreDNS/kube-dns runs
  egressAPIServerCIDRs:              # K8s API server (ports 443/6443)
    - 0.0.0.0/0
  egressHTTPSCIDRs:                  # external HTTPS (Hivemind, cloud storage, proxy)
    - 0.0.0.0/0

Linux VM / Bare Metal (systemd)

Auditty Edge runs as a systemd service (auditty-edge) on any Linux host (Ubuntu, Debian, RHEL, and compatible). It reads your log files in read-only mode and writes the processed output to a separate directory that your existing log collector reads from; your application’s own files are left exactly as they are.

Your App → /var/log/myapp/app.log → Auditty (tail) → /var/log/auditty/var_log_myapp_app.log → Your Log Collector

Step 1: Install

curl -fsSL https://get.auditty.ai/edge | sudo sh -s -- --key <API_KEY>

The installer detects architecture (amd64/arm64), downloads and verifies the binary, installs to /usr/local/bin/auditty-edge, creates /etc/auditty/config.yaml, and starts the systemd service.

Step 2: Configure

# /etc/auditty/config.yaml

hivemind_api_key: "YOUR_API_KEY"
log_level: info

interceptor:
  type: service

cluster_name: my-host          # how Hivemind groups this host (the installer sets the hostname)

service:
  input_dirs:
    - /var/log/myapp
  input_patterns:
    - "*.log"
  input_recursive: true
  output_dir: /var/log/auditty   # where the processed logs are written
  enable_rotation: false          # let logrotate manage output files

Step 3: Start & Verify

sudo systemctl restart auditty-edge   # restart after config changes
systemctl status auditty-edge
journalctl -u auditty-edge -f

Point Your Log Collector at the Output Directory

# Fluentd
<source>
  @type tail
  path /var/log/auditty/*.log
</source>

# Vector
[sources.auditty_logs]
type    = "file"
include = ["/var/log/auditty/*.log"]

# Filebeat
filebeat.inputs:
  - type: log
    paths:
      - /var/log/auditty/*.log

CLI (Local / Batch Processing)

edge-cli runs the full Auditty pipeline against a directory of log files and writes the processed output locally. Useful for one-off processing, backfills, CI pipelines, or testing rules before deploying to production.

Usage

# One-shot: process all logs in /var/logs → /processed/logs
edge-cli -i /var/logs -o /processed/logs

# Watch mode: run continuously like a service
edge-cli -i /var/logs -o /processed/logs -w

# With custom config and debug logging
edge-cli -i /var/logs -o /processed/logs -c edge-config.yaml -l debug

# Validate a config file without processing anything
edge-cli validate-config -c edge-config.yaml
FlagShortDefaultDescription
--input-i(required)Input directory containing log files
--output-o(required)Output directory for processed logs
--config-cConfig file path
--watch-wfalseWatch for new/modified files continuously
--recursive-rtrueProcess subdirectories recursively
--log-level-linfoLog level: debug, info, warn, error
--shadow-mode-sfalseMetrics only, no suppression applied
--metrics-mtrueWrite a CSV metrics report to the output dir
--rulesCustom rules YAML file (overrides config)
--preserve-structuretrueMirror input directory structure in output

Suppression Summaries

When logs are suppressed, edge-cli writes each summary into the same output file as the lines it stands for, exactly as Edge does on Kubernetes. Each summary contains the count, time window, and the searchable ids (IPs, trace_id, user_id, etc.) indexed inline, and every original line stays recoverable from the output or the Vault.

Edge Proxy (OTLP / Network Ingest)

auditty-edge-proxy brings the same dedup intelligence to sources that aren’t Kubernetes pods on a node Auditty can reach: serverless, CI runners, app SDKs, or any OpenTelemetry pipeline. Instead of intercepting files, it receives logs over the network, runs the exact same pipeline (fingerprint → dedup → rules → rate-limit), forwards survivors to a configured upstream, and archives suppressed originals to the Vault.

sender → edge-proxy → [fingerprint → dedup → rules → rate-limit] → survivors → your upstream (OTLP/HTTP or JSON)
                                                              └→ suppressed → Vault (optional)

How non-K8s sources appear in Auditty

Edge Proxy reports through the same metrics/signals/logs pipeline as the DaemonSet. It uses your configured cluster_name as the cluster label and a stable node identity (proxy.node_id, falling back to NODE_NAME, then hostname) as the node; pin proxy.node_id on scaled Deployments so backfill links keep working after a redeploy. OTLP resource attributes (service.name, k8s.namespace.name, …) populate the workload/namespace dimensions.

Step 1: Configure

# /etc/auditty/config.yaml

hivemind_api_key: "YOUR_API_KEY"
cluster_name: "edge-proxy-us-east"   # synthetic cluster label in Hivemind

interceptor:
  type: proxy

proxy:
  listen_addr: ":4318"               # OTLP/HTTP default port
  node_id: "cdn-gateway-1"           # stable node identity (pin on scaled Deployments)
  auth_token: "${EDGE_PROXY_TOKEN}"  # require a bearer token on ingress
  enable_otlp: true                  # POST /v1/logs  (OTLP/HTTP)
  enable_http: true                  # POST /ingest   (generic JSON/NDJSON)
  # timestamp_fields: ["timestamp_ms"]  # extra event-time keys; @timestamp and
  #                                     # EdgeStartTimestamp work out of the box
  upstream:
    endpoint: "https://otlp.example.com/v1/logs"
    protocol: otlp                   # otlp (protobuf) | http (JSON)
    compression: gzip
    headers:
      Authorization: "Bearer ${UPSTREAM_TOKEN}"

forwarder:
  vault:
    enabled: true                    # archive suppressed originals (optional)

Step 2: Point senders at the proxy

  • OpenTelemetry: set your OTLP/HTTP logs exporter endpoint to http://<edge-proxy>:4318 (logs path /v1/logs).
  • Generic JSON/NDJSON: POST to /ingest: a JSON array, NDJSON, or plain text lines. Recognized keys include message, namespace, service, timestamp.
  • Event time: resolved from timestamp/ts/time, with zero-config fallbacks for @timestamp (Elastic) and EdgeStartTimestamp (Cloudflare Logpush); add vendor-specific keys via proxy.timestamp_fields. Batch-delivered sources (CDN → S3 → proxy) are archived under the log’s own time, so backfill windows match the source’s clock.
  • Include the bearer token: Authorization: Bearer $EDGE_PROXY_TOKEN.

Step 3: Forward survivors to Datadog (Vector → Auditty → Datadog)

To place Auditty between Vector and Datadog, point your senders at the proxy and forward the surviving (post-suppression) stream on to Datadog. There is no separate “Datadog sink” to configure; survivors egress over OTLP/HTTP to any upstream with custom auth headers. Two supported paths:

your apps → Vector → edge-proxy → [suppress] → survivors → Datadog
                                              └→ suppressed → Vault (optional)

Option A: keep Vector as your Datadog shipper (recommended). Auditty slots in as a suppression hop and your existing datadog_logs sink (tags, pipelines, API key) is reused unchanged.

# edge-proxy: send survivors back to Vector
proxy:
  upstream:
    endpoint: "http://vector:8080"   # a Vector http_server source
    protocol: http                   # JSON array of records

# vector.yaml: receive survivors, ship to Datadog with your existing sink
sources:
  from_auditty:
    type: http_server
    address: 0.0.0.0:8080
    decoding:
      codec: json
sinks:
  datadog:
    type: datadog_logs
    inputs: [from_auditty]
    default_api_key: "${DD_API_KEY}"

Option B: straight to Datadog (one less hop). Point the upstream at Datadog’s HTTP logs intake and pass your API key as a header. The message field maps to the Datadog log message, a sender’s service (OTLP service.name or a generic service field) maps to Datadog’s reserved service facet, and every other attribute (e.g. application and your custom keys) passes through untouched. Use your site’s intake host (e.g. datadoghq.eu).

proxy:
  upstream:
    endpoint: "https://http-intake.logs.datadoghq.com/api/v2/logs"
    protocol: http
    compression: gzip
    headers:
      DD-API-KEY: "${DD_API_KEY}"

Security & overload

  • Secure by default: set auth_token (bearer) and/or mTLS (tls_client_ca_file, with tls_cert_file/tls_key_file for the proxy’s own certificate). With neither, the proxy refuses to start rather than accept logs from anyone; opt into an unauthenticated listener on a trusted network with proxy.insecure: true.
  • Admission is all-or-nothing per request: under load the proxy returns HTTP 503 with Retry-After having accepted nothing, so a retry duplicates nothing. A 200 means every record landed. Only a proxy stopping in the middle of a request can keep the records it had already accepted and still answer 503; the retry then delivers those again, so the worst case is a duplicate. A client always gets an answer.

Scaling

  • Scale up first. A single instance runs the same processing as the DaemonSet; give one instance more CPU/memory before adding replicas.
  • Scale out for capacity & HA. Run multiple replicas behind a Service/load balancer for raw throughput. Survivors and Vault archival are always correct regardless of replica count.
  • Keep each source on one replica. Suppression is decided per replica, so a source spread round-robin across replicas is suppressed less than it would be on one. The Helm chart pins each sender to a replica with sessionAffinity: ClientIP (below); behind your own load balancer, key it on a stable source attribute (service.name / host.name).

Deploy on Kubernetes (same Helm chart)

The network proxy ships in the same auditty-edge Helm chart as the DaemonSet; enable it alongside (or instead of) the node interceptor with edgeProxy.enabled. It deploys as a horizontally-scalable Deployment + Service, with an optional HPA that scales on CPU/memory.

edgeProxy:
  enabled: true
  listenPort: 4318                # OTLP/HTTP ingress port
  authToken: "${EDGE_PROXY_TOKEN}" # required (or set insecure: true)
  extraEnv:                       # the token comes from a Secret, never the ConfigMap
    - name: EDGE_PROXY_TOKEN
      valueFrom:
        secretKeyRef:
          name: auditty-edge-proxy-token
          key: token
  upstream:
    endpoint: "https://otlp.example.com/v1/logs"
  hpa:
    enabled: true                 # autoscaler owns the replica count
    minReplicas: 2
    maxReplicas: 10
    targetCPUUtilization: 70

Edge API and the proxy each have their own autoscaler, and all three components have their own resource settings, so you can run any combination from one release.

Scaling and suppression: suppression is decided per replica. So that a source is still suppressed as one stream, the proxy Service defaults to sessionAffinity: ClientIP, pinning each sender to one replica. For diverse senders you still get spread across pods; if many sources share one egress IP they pin to a single replica (correct, just less balanced). Set edgeProxy.service.sessionAffinity: None only if you accept per-replica suppression; summary_min_count then applies per replica and the dedup ratio scales sub-linearly with replica count.