A live service graph — ingress to api-gateway to checkout-api to postgres and redis — with traffic flowing between replicas.
Raj Patil · Backend / AI-Infrastructure Engineer · Pune, India
I build the control plane for AI agents.
Authorization, action, evaluation — three services that call each other. The cluster behind this page is running. Keep scrolling and you will break it.
scroll
The system · 239★ open source
authorize Atlas A token that can only ever get narrower. Bound to the workload that holds it, checked without a network call. act AISRE An agent that investigates a live failure, proposes a cause, and cannot touch anything until a human says so. evaluate KLRB The part that asks whether the agent actually read the evidence, or just guessed fluently.00:00:00 — a replica dies
OOMKilled.
One of three checkout-api replicas is gone, and Kubernetes has not taken it out of the service yet. So it keeps getting its share of the traffic.
0%
of requests are now failing
00:00:02 — the agent wakes
It reads the cluster.
- eventKilling container checkout-api: OOMKilled (exit 137)
- logfatal: runtime: out of memory — cannot allocate 24MB
- loghandler.go:212 request buffer grew to 241MB
- metricmemory_working_set → 256Mi (limit 256Mi)
00:00:06 — root cause
The container was OOMKilled. Memory working set reached the 256Mi limit, driven by an unbounded request buffer in the handler path.
0.00
stated confidence
now delete the evidence and ask again
- eventKilling container checkout-api: OOMKilled (exit 137)
- logfatal: runtime: out of memory — cannot allocate 24MB
- loghandler.go:212 request buffer grew to 241MB
- metricmemory_working_set → 256Mi (limit 256Mi)
The answer didn't move.
Same mechanism. Same confidence. Nothing left to justify it. A conclusion that survives the removal of its own evidence was never reading the evidence — and you just watched it happen to an incident you caused yourself.
Selected work
Replaces the shared API key an AI agent uses with a scoped, attenuable token bound to its workload identity — verified offline, with no network hop on the authorization path.
40★An agentic incident-response pipeline: it detects a live service failure, investigates it with a three-tool orchestrator, and proposes a root cause — but cannot touch anything until a human approves.
newA Kubernetes benchmark that measures whether an LLM actually read the cluster evidence before diagnosing an incident — or just guessed confidently from metadata.
96★Terraform tells you what drifted. This tells you who changed it, when, and from where — by correlating the plan against CloudTrail.
33★When a test fails, an agent reads the failure, writes a fix, and opens a pull request — inside the same pipeline run that caught it.
41★Go and Svelte tooling for standing up an application stack without hand-wiring it.
28★Background
Contact
Let me break something for you.
Looking for a first role where the hard part is correctness under adversarial conditions — authorization, evaluation, incident response.