Articles, short research notes, and talks. The goal is graduate-level reasoning, not marketing.
The long-form argument behind every project here: memory substrates with provenance, capability-scoped authorization, and observability indexed by decision rather than span.
Why retrieval-augmented generation is not enough, and what a real memory substrate for autonomous agents might look like.
Metrics, logs, and traces were built for human operators. Autonomous agents demand a new observability primitive: the reasoning trace.
Static allow-lists will not keep autonomous agents safe. Governance must be dynamic, uncertainty-aware, and grounded in simulated impact.
Coordination, conflict resolution, and shared context are the defining challenges of multi-agent systems.
What a decade of production infrastructure engineering taught me about building agentic systems.
MCP is a promising step toward standardized tool use, but it leaves hard questions about state, identity, and trust unanswered.
Stuffing history into a prompt is a hack, not an architecture.
Benchmarks that measure single-turn correctness miss the point. We need to evaluate agents across episodes, not just turns.
A systems-level view of the infrastructure needed for trustworthy autonomous agents.