Trustworthy infrastructure for autonomous AI systems.
A systems engineer's point of view on what it will take to make agentic AI reliable, observable, and safe enough to run in production.
How I arrived here
I began in systems engineering and moved through networking, cloud engineering, DevOps, security and platform architecture. Across those roles, the aim became consistent: observable, fault-tolerant, self-healing systems that give teams safe self-service capabilities and reduce unnecessary intervention.
At Exodus, from February 2022 to August 2026, I designed and operated Kubernetes and distributed infrastructure across five production EKS clusters. GitOps, stateful workloads, networking and incident response taught me to examine ownership, state, failure domains and recovery together.
My current applied work at BlockOps brings that experience to model serving and deployment across cloud and privately managed GPU infrastructure. Alongside it, Agent Rails investigates the authorization, memory, observability and orchestration that autonomous systems need.
The progression is production systems → distributed and platform architecture → AI inference infrastructure → trustworthy agent infrastructure → applied AI systems for businesses.
Understanding the Stack
I cannot design strong distributed AI systems without understanding the components beneath them. A framework interface can hide the mechanism, but it cannot remove its resource constraints or failure modes.
I am studying inference runtimes and serving systems such as vLLM: model weights and GPU memory, KV-cache growth, prefill and decode, batching, scheduling and admission. These mechanics shape capacity, latency and the boundaries within which recovery is possible.
Above serving sit orchestration, memory, tool execution, authorization and observability. I ask what state each component owns, which operations can be retried, what evidence survives a failure and where one workload can affect another.
This is a learning and design method, not a claim that every layer is already solved. Explain the mechanism, state the assumptions, build a bounded experiment, then use the result to revise the design.
From components to business workflows
BlockOps is the applied infrastructure bridge: I am building a self-service path for organizations to deploy models and inference workloads on cloud or privately managed GPU infrastructure, while going deeper into serving internals.
The SME Agent Template explores the next boundary: can a business use the resulting system without deep infrastructure expertise? Its architecture supports transactional coffee-shop and reservation-based event-space workflows. I am working with businesses to validate workflows and develop it toward an MVP.
Two implemented workflow shapes do not establish adoption or real-world effectiveness. Those are separate questions about setup effort, useful task completion, human approval, recovery and the support an operator still needs.
Why agentic systems?
Agents introduce software that can observe, reason and act over time, often across systems its designers do not fully control. I am interested in where this is useful and what boundaries it requires.
Existing automation already manages infrastructure and data pipelines. Model-driven action introduces additional uncertainty: a system can choose a plausible but unauthorized tool call, misunderstand context or fail to recognize an incomplete operation.
That generality is also the risk. An agent that can do many things can also do many things wrong. It can misinterpret intent, persist bad decisions, or act on outdated knowledge. Making agents useful means making them trustworthy.
Why infrastructure?
Most of the conversation around AI safety focuses on models: alignment, fine-tuning, red-teaming, and evaluation benchmarks. These matter. But they are not sufficient.
A model is one component of a larger system. The system also includes memory, tools, observability, governance, identity, and coordination. Even a perfectly aligned model can fail if the infrastructure around it is brittle or opaque.
Infrastructure is where reliability engineering meets AI research. It is where we ask: How do we know what the agent believes? How do we stop it when we need to? How do we make sure its actions are auditable and reversible? How do multiple agents share context without corrupting each other's state?
These are systems questions. They require distributed systems, security, databases, networking, and observability. They also require a deep understanding of how agents actually behave in production. That intersection is where I want to work.
What is not solved today?
The current state of agent infrastructure is reminiscent of web development in the late 1990s. There are powerful primitives, but the surrounding tooling is immature. We are building skyscrapers on foundations that were designed for smaller structures.
- Memory is mostly retrieval. Most "memory" systems are vector search over history. They lack forgetting, consolidation, causal structure, and provenance.
- Observability is human-centric. Metrics, logs, and traces assume a human operator. They do not capture reasoning, intent, or the agent's model of the world.
- Governance is static. Approval gates and allow-lists do not scale to thousands of autonomous actions per hour, and they ignore uncertainty.
- Coordination is ad-hoc. Multi-agent systems lack shared identity, shared memory, and shared protocols. Each integration is bespoke.
- Evaluation is narrow. Benchmarks measure single-turn correctness, not long-episode reliability, recovery, or side effects.
These gaps are not accidental. They exist because agentic systems are genuinely harder to build than the systems that came before. Solving them is the work of years, not months.
The questions I'm pursuing
My research is organized around a few concrete questions. Each one could occupy a career; together they define a coherent research program.
- How should autonomous agents represent, update, and forget long-term knowledge while remaining efficient and explainable?
- What observability primitives are needed when the consumer of telemetry is another machine rather than a human?
- How can runtime governance be dynamic, uncertainty-aware, and grounded in simulated impact?
- What identity, memory, and trust infrastructure enables reliable multi-agent coordination across organizational boundaries?
- How should autonomous agents be evaluated across long episodes, not just single turns?
These questions are not purely theoretical. They are grounded in the systems I have built and the failures I have observed. The goal is to turn hard-won production experience into generalizable research.
How I work
I believe the strongest research in this area will come from building real systems, failing informatively, and then abstracting the lessons into theory. I do not want to write papers about systems that do not exist.
Agent Rails organizes that work around research problems, with Agent Guard as the flagship. Each project follows a checkable sequence: question → hypothesis/design → architecture → experiment/evaluation → results → limitations → next questions. A working design paper records the synthesis; it is not a publication.
If this vision resonates, I would love to collaborate. You can explore the projects, read the writing, or get in touch.