Question → hypothesis/design → architecture → evaluation → results → limitations → next questions
Signing and other irreversible operations make timely evidence important. Vigilo collects host signals and exposes them to an analyst. Current measurements cover immediate alerts in a bounded container experiment; they do not establish attack prevention, pre-action detection or the effectiveness of LLM correlation.
Hypothesis to test: correlating host events can identify some multi-step patterns missed by per-event rules at acceptable cost and false-positive rates. The correlation comparison remains unmeasured.
The interesting comparison is not detection rate in isolation. It is whether the five-minute correlation window catches chains that per-event rules structurally cannot, without drowning the operator in the process.
Scripted attack chains replayed on an instrumented host with timestamps at each stage. Run: real Docker container, real daemon, three immediate-tier chains (keystore write, secret-file write, outbound connection to a known-bad port), 10 repeats each. Analyst-tier (correlated) latency not yet measured, blocked on the same LLM-billing gap as the MemKit resolver work.
Measures Time from first hostile syscall to immediate page, and to correlated analyst verdict.
The same chains run with the analyst disabled, leaving only per-event rules.
Measures Fraction of chains detected only by correlation, and stage at which each was caught.
Daemon left running on normal production workloads for a multi-week window with no injected attacks. Run: 60-second bounded window of ordinary in-container activity, a real measurement, but at a scale disclosed as a fraction of the multi-week design above, not equivalent to it.
Measures False positives per host per day, before and after suppression tuning; alert volume trend.
Signer-shaped workload benchmarked with the daemon on and off. Run: real docker stats sampling, idle vs a real trigger burst. Event throughput ceiling not yet measured.
Measures CPU and RSS overhead, event throughput ceiling, SQLite write amplification.
Repeated analyst passes over identical event windows.
Measures Verdict latency, token cost per pass, and verdict consistency across runs on the same input.
This covers the immediate tier only. Analyst-tier (correlated) latency and the V2 correlation-gain comparison are still blocked on the same LLM-billing gap as MemKit's resolver work. The false-positive window is bounded (60s), explicitly not equivalent to V3's own multi-week design. Event throughput ceiling (V4's other sub-metric) isn't measured yet either. Marked preliminary until these land.
Immediate-tier latency and a bounded false-positive window have been measured. The analyst is a separate correlation path; its detection gain and latency are not established by those runs. Signal coverage is not proof that an attack is detected or prevented before harm.