Skip to content
English

Retention, purging, and observability

Three operational questions that come up: how long data is kept, how purging happens and what evidence it leaves, how the state of the system is observed, and which figures to use when planning disks.

The policy sits in the “log retention and review” section of the Security Policies page: five retention periods and one ceiling per purge run.

Policy key Shipped value
Operation log retention (days) 0 (kept forever)
Command stream retention (days) 0 (kept forever)
Alert record retention (days) 0 (kept forever)
Session recording retention (days) 90
Checkpoint retention (days) 0 (kept forever)
Deletion ceiling per purge run 100000

0 means kept forever. If an in-force policy group has a requirement for this key, the settings drawer shows it; the page head can apply a group after a preview.

One constraint spans the keys: checkpoint retention cannot fall below the longest of the four data kinds (0 counts as unbounded), and saving the batch is refused when it does. The floor for the per-run ceiling is 5000 rows.

The purge runs once a day on a schedule, deletes in batches, continues the next day when a run does not finish, and blocks no online request. Before you shorten a retention period, the interface asks for confirmation, because a purge is a hard delete.

Audit records use a sealed checkpoint interval as the smallest indivisible unit: while a single row in the interval has yet to expire, the whole stretch waits, so the chain stays continuous and verifiable.

When a recording that has been copied offsite expires, only the local file is cleared and the ledger is marked; the system issues no delete call against the remote object. The bucket’s lifecycle rule carries expiry on the remote side.

GET /metrics serves operational metrics in the Prometheus text format. The path deliberately sits outside /api, and the production external proxy forwards only /api and the connection channels, so a zero-configuration deployment cannot be reached there from outside. To open it to a monitoring system, set METRICS_TOKEN, after which requests carry the matching bearer token; a mismatch returns 401 with no metric name or value in the response.

The minimum metric set covers: active sessions per protocol, active connections, recording storage usage, unreviewed alerts by severity, audit queue depth, cumulative audit drops, the seal state, and cumulative HTTP requests with a latency distribution. With offsite storage on, it also exposes pending uploads, failures, pending remote deletions, integrity mismatches, the age of the oldest pending item, and the last successful upload time.

Expensive metrics are refreshed by a background task; recording storage usage refreshes every 30 seconds, so a value lags by at most one cycle. /metrics stays readable while sealed, and the metrics of a service that has not been built up yet are absent rather than reported as 0.

A user with audit view permission also sees a “recording usage” card on the dashboard. The figure is the sum of the actual file sizes in the recording directory, across every protocol.

These are observations from a single machine. Both are lower bounds taken while idle, for getting the order of magnitude:

Protocol Observation
Text terminal About 30 B/s idle, around 105 KB per session hour; the amplification factor under load is about 1.32x
Graphical connection A static desktop starts at about 10.4 MB per session hour

The arithmetic: disk for text recordings is roughly the terminal output in bytes times 1.32; for graphical recordings it is roughly concurrent sessions times average session length times the growth rate times the retention days.

The resources that saturate first, in order: the graphics proxy container, the write bandwidth and capacity of the recording disk, the database connection pool ceiling, and the connection limits of the target hosts themselves. Each session adds about 16 goroutines in the backend, with an upper bound of about 148 KB of memory per session.

Set an alert on the recording storage usage metric, on a growth rate or a capacity threshold.

.env holds a few more adjustable knobs: the terminal idle timeout and maximum duration, the time and row thresholds for checkpoints, the per-run ceiling for a rekey, and the cluster list timeout. Left empty they take the built-in defaults. The query console allows 4 concurrent sessions per person and 64 across the system.

The backend process runs with the minimum privileges it needs, forbids privilege escalation, writes no core, and locks memory pages. The process is set non-dumpable before start. If that cannot be applied, it stops and writes a log. Account secrets the system holds are cleared after their last use. How long a data key stays in memory, and when a seal clears the cache: Key management.

  • Each purge writes a record with the data kind, the time range purged, the rows actually deleted, and whether it finished only part of the work.
  • A purge of audit records also states the checkpoint sequence numbers or intervals it covered.
  • Purge records are forwarded over syslog as well, so a copy of what was deleted lives on an external system too.
  • These fields let a lawful purge on expiry be told apart from a malicious deletion in the evidence; see The checkpoint chain for how that is decided.