Platform Engineer · Distributed Systems & AI Infrastructure Research

Franklin Okpako

Building and evaluating infrastructure for reliable autonomous AI systems.

My work connects production systems and distributed architecture to AI inference, trustworthy agent infrastructure and usable business workflows. At BlockOps I build model infrastructure; through Agent Rails I investigate the controls around agents; with businesses I am testing how those components can become useful systems.

Open source·Production systems·Technical writing
The systems stack
01Inference & serving
02Distributed infrastructure
03Authorization · memory · observability
04Usable agent systems
10+ years
systems, networking & production infrastructure
5 production EKS clusters
operated at Exodus, Feb 2022–Aug 2026
35+ services · 10+ teams
org-wide GitOps & self-service platform
Millions of users
wallet and blockchain infrastructure
400-query evaluation
MemKit contradiction benchmark

Built on

AWS
Kubernetes
Linux
Cilium
Terraform
ArgoCD
GitHub
Docker
GitLab
Cloudflare
[ 001 about ]

About the researcher

  1. 01Production systems
  2. 02Distributed and platform architecture
  3. 03AI inference infrastructure
  4. 04Trustworthy agent infrastructure
  5. 05Applied AI systems for businesses

I began in systems engineering and moved through networking, cloud engineering, DevOps, security, platform engineering and infrastructure architecture. My focus increasingly became designing observable, fault-tolerant, self-healing systems that reduce unnecessary operational intervention and give teams safe, self-service platforms to build on.

At Exodus, from February 2022 to August 2026, I designed and operated production Kubernetes and distributed infrastructure across five EKS clusters supporting wallet and blockchain systems at global scale. That work included GitOps architecture, networking, stateful workloads, observability and incident response, and turning recurring failures into automation and more resilient system design.

Today, my applied work at BlockOps focuses on AI infrastructure: model serving, GPU-backed workloads, isolation, observability and simplifying deployment across cloud and privately managed environments. I am going deeper into inference and serving internals so that design decisions follow from how the components actually work.

Alongside that work, I build Agent Rails, an open-source effort exploring authorization, memory, observability, delegated authority, orchestration and evaluation for autonomous AI systems. Agent Guard is the flagship: a concrete place to examine policy, identity, human approval, audit and execution boundaries together.

I am also working with businesses to understand how these components can become usable AI tools. The SME Agent Template supports two different workflow shapes; workflow validation with businesses and development toward an MVP are in progress, not completed real-world validation.

Read the full research vision
[ 002 research systems ]

Agent Rails: questions with running code.

Agent Rails is the umbrella for my work on authorization, memory, observability and orchestration. Agent Guard leads the research, bringing policy, identity, approval, audit and isolation into one inspectable boundary.

[ 005 journal ]

Notes from the workbench.

2026-08-01

August 2026: Building memory isn't enough

The harder problem is deciding what should be remembered, and what should be allowed to fade.

2026-07-05

July 2026: Uncertainty is information, not a bug

The most interesting failures are not when an agent is wrong, but when it is uncertain and acts anyway.

[ 006 open source ]

Agent Rails, organized by research problem.

The repositories are components of one research effort. Agent Guard is the flagship; the other components investigate the memory, evidence and coordination it needs around it.

Memory

What should an agent retain, revise or discard when sources disagree?

Observability / evidence

What happened, what was claimed, and what can the record establish?

Orchestration

How do these boundaries hold across a complete workflow?

[ 007 professional work ]

A decade of production systems, as evidence for the research.

At Exodus, I operated five EKS clusters and built distributed infrastructure and self-service delivery platforms. At BlockOps, I now apply that experience to model serving and deployment infrastructure. The recurring goal is observable, fault-tolerant systems that reduce unnecessary intervention.

AWS · Kubernetes · Linux · Cilium · Terraform · ArgoCD · Cloudflare · reliability · security

[ 008 consulting ]

Hands-on help with agentic AI infrastructure.

Architecture review, security baselines, and implementation support for teams building autonomous systems.

Get in touch

Talk to an engineer, not a sales bot.

Tell us about your stack and what you're trying to automate. We'll reply within one business day with a concrete next step — pilot, architecture review, or a quick async answer.

SLAResponse within 1 business day

By submitting you agree to be contacted about Techbots.