All research systems
VigiloActive

How quickly can host telemetry surface suspicious activity, and does correlation add useful evidence?

Question → hypothesis/design → architecture → evaluation → results → limitations → next questions

Motivation

Signing and other irreversible operations make timely evidence important. Vigilo collects host signals and exposes them to an analyst. Current measurements cover immediate alerts in a bounded container experiment; they do not establish attack prevention, pre-action detection or the effectiveness of LLM correlation.

Hypothesis

Hypothesis to test: correlating host events can identify some multi-step patterns missed by per-event rules at acceptable cost and false-positive rates. The correlation comparison remains unmeasured.

Threat model

  • Key exfiltration: a private key or keystore is read by a process that has no business reading it, then leaves the host
  • Remote code execution: a shell spawned from an application runtime (node, python) as the first foothold
  • Supply-chain compromise: a package install initiated by a running application process rather than a deploy
  • Privilege escalation from an application process toward root
  • Coordinated campaigns: the same sequence appearing across several monitored hosts, invisible to any single host's rules
  • Out of scope: kernel-level rootkits that defeat /proc and inotify, and physical or hypervisor-level compromise

Architecture

  • Go daemon: fsnotify file watcher, /proc process poller, /proc/net connection poller, suppression rules
  • SQLite event buffer with a signal-dedup table, so repeated noise does not become repeated pages
  • Two-tier alerting: immediate push (Slack, Telegram, email, generic webhook) for high/critical events
  • MCP server on :7070 exposing five tools, so the analyst reads events as a first-class interface rather than scraping logs
  • TypeScript analyst agent: multi-daemon aggregation, context compaction, inspector-and-retry loop, running every five minutes

Detection model

  • Tier one, immediate: single events with an unambiguous signature (keystore read by an unfamiliar process, secret written to a world-readable path) page in about a second
  • Tier two, correlated: the analyst reads a five-minute window and looks for shapes, not strings (read-then-connect, spawn-then-install, repeat-across-hosts)
  • Severity is assigned by irreversibility, not by rarity, so anything touching signing material is critical regardless of frequency
  • Suppression is explicit and configured per host, so known-noisy legitimate processes do not train operators to ignore the channel
  • Dedup keys collapse repeated identical signals into one page with a count
  • The analyst never acts. Its output is an explanation and a severity, and every conclusion cites the events it rests on

Implementation

  • Single Go binary per host, no agent framework; state is one SQLite file
  • Collection by polling /proc and /proc/net plus inotify on watched paths, deliberately portable ahead of an eBPF path
  • MCP server is the only read interface; the analyst has no log-scraping fallback, which keeps the event schema honest
  • Analyst runs off-host on a five-minute cadence, aggregating several daemons and compacting context before each pass
  • Alert transports are pluggable: Slack, Telegram, email, generic webhook

Evaluation methodology

The interesting comparison is not detection rate in isolation. It is whether the five-minute correlation window catches chains that per-event rules structurally cannot, without drowning the operator in the process.

Workloads
  • Private key and keystore reads
  • Env dump followed by an outbound connection (exfiltration chain)
  • Shell spawned from node/python (RCE)
  • Package install from an app process (supply chain)
  • The same pattern appearing across multiple monitored servers
Baselines
  • Host metrics and log tailing with alert-on-keyword rules
  • Per-event rules with the LLM analyst disabled
Metrics
  • Time from first hostile syscall to a human-readable alert
  • False-positive rate per monitored host per day
  • False-negative rate against the scripted chain suite
  • Fraction of multi-step chains caught by correlation but missed by single-event rules
  • CPU and memory overhead on the monitored host
  • Event throughput ceiling before the buffer backs up
  • Analyst verdict latency and token cost per pass

Experiments

V1: Detection latency

Scripted attack chains replayed on an instrumented host with timestamps at each stage. Run: real Docker container, real daemon, three immediate-tier chains (keystore write, secret-file write, outbound connection to a known-bad port), 10 repeats each. Analyst-tier (correlated) latency not yet measured, blocked on the same LLM-billing gap as the MemKit resolver work.

Measures Time from first hostile syscall to immediate page, and to correlated analyst verdict.

V2: Correlation gain over single-event rules

The same chains run with the analyst disabled, leaving only per-event rules.

Measures Fraction of chains detected only by correlation, and stage at which each was caught.

V3: False positives over time

Daemon left running on normal production workloads for a multi-week window with no injected attacks. Run: 60-second bounded window of ordinary in-container activity, a real measurement, but at a scale disclosed as a fraction of the multi-week design above, not equivalent to it.

Measures False positives per host per day, before and after suppression tuning; alert volume trend.

V4: Host overhead

Signer-shaped workload benchmarked with the daemon on and off. Run: real docker stats sampling, idle vs a real trigger burst. Event throughput ceiling not yet measured.

Measures CPU and RSS overhead, event throughput ceiling, SQLite write amplification.

V5: Analyst cost and stability

Repeated analyst passes over identical event windows.

Measures Verdict latency, token cost per pass, and verdict consistency across runs on the same input.

Results

Preliminary results
  • Immediate-tier detection latency, real measurements against a real running daemon (n=10 per chain): file-based signals (keystore write, secret-file write) fire in ~55-72ms; a suspicious outbound connection, poll-based rather than event-driven, fires in ~490ms median (854ms p95) at a 1-second poll interval
  • Zero false positives across a 60-second window of ordinary in-container activity, a real but small-scale result, not a claim about a multi-week production window
  • Two real bugs were found in the daemon itself while building this evaluation, both fixed and shipped. The watch_paths entries pointing at a single file (the project's own documented .env example) were silently never watched at all, and an explicit signal_cooldown: 0s ("no cooldown") was silently overridden to 15 minutes by a zero-value bug in the alerter
  • In the measured container: 0.3% CPU / 12.7MB RSS idle, 1.95% CPU / 13.9MB RSS under a real trigger burst. Write-path storage cost is not reported as a clean number, because the daemon runs SQLite in WAL mode, and a short burst without a checkpoint doesn't reflect steady-state per-event cost

This covers the immediate tier only. Analyst-tier (correlated) latency and the V2 correlation-gain comparison are still blocked on the same LLM-billing gap as MemKit's resolver work. The false-positive window is bounded (60s), explicitly not equivalent to V3's own multi-week design. Event throughput ceiling (V4's other sub-metric) isn't measured yet either. Marked preliminary until these land.

Failure cases and risks — not an exhaustive catalogue

  • Polling misses short-lived processes that start and exit between /proc samples
  • A legitimate deploy that installs packages from an app process looks exactly like a supply-chain event
  • Under heavy event volume the analyst's context compaction drops the very events that made a chain legible
  • Suppression rules written during an incident are rarely removed afterwards, quietly widening blind spots

Limitations

  • Linux-only: collection depends on /proc and inotify semantics
  • Single-host SQLite buffer; cross-server correlation happens in the analyst, not the storage layer
  • The LLM analyst is a detection aid, not a control, and it does not block anything
  • False-positive measurement so far is a 60-second bounded window on a dev container, not a long production run or a real signer's noise profile

Current status

Immediate-tier latency and a bounded false-positive window have been measured. The analyst is a separate correlation path; its detection gain and latency are not established by those runs. Signal coverage is not proof that an attack is detected or prevented before harm.

Roadmap

  • eBPF collection to replace polling where the kernel allows it
  • Signed event chains so the buffer itself is tamper-evident
  • Response hooks: quarantine and key-rotation triggers behind human approval

Next questions

  • Where the boundary sits between detection and structural prevention on an irreversible action
  • How much host context an LLM analyst needs before correlation beats a hand-written rule
  • Suppression as a learned rather than configured behaviour