0 req/s 0.0% err healthy

Raj Patil · Backend / AI-Infrastructure Engineer · Pune, India

I build the control plane for AI agents. 

Authorization, action, evaluation — three services that call each other. The cluster behind this page is running. Keep scrolling and you will break it.

scroll

The system · 239★ open source

authorize Atlas A token that can only ever get narrower. Bound to the workload that holds it, checked without a network call. act AISRE An agent that investigates a live failure, proposes a cause, and cannot touch anything until a human says so. evaluate KLRB The part that asks whether the agent actually read the evidence, or just guessed fluently.

00:00:00 — a replica dies

OOMKilled. 

One of three checkout-api replicas is gone, and Kubernetes has not taken it out of the service yet. So it keeps getting its share of the traffic.

0%

of requests are now failing

00:00:02 — the agent wakes

It reads the cluster.

  • eventKilling container checkout-api: OOMKilled (exit 137)
  • logfatal: runtime: out of memory — cannot allocate 24MB
  • loghandler.go:212 request buffer grew to 241MB
  • metricmemory_working_set → 256Mi (limit 256Mi)

00:00:06 — root cause

The container was OOMKilled. Memory working set reached the 256Mi limit, driven by an unbounded request buffer in the handler path.

0.00

stated confidence

now delete the evidence and ask again

  • eventKilling container checkout-api: OOMKilled (exit 137)
  • logfatal: runtime: out of memory — cannot allocate 24MB
  • loghandler.go:212 request buffer grew to 241MB
  • metricmemory_working_set → 256Mi (limit 256Mi)

The answer didn't move. 

Same mechanism. Same confidence. Nothing left to justify it. A conclusion that survives the removal of its own evidence was never reading the evidence — and you just watched it happen to an incident you caused yourself.

I built a benchmark for this →

Selected work

Atlasauthorize

Replaces the shared API key an AI agent uses with a scoped, attenuable token bound to its workload identity — verified offline, with no network hop on the authorization path.

40★
AISREact

An agentic incident-response pipeline: it detects a live service failure, investigates it with a three-tool orchestrator, and proposes a root cause — but cannot touch anything until a human approves.

new
KLRBevaluate

A Kubernetes benchmark that measures whether an LLM actually read the cluster evidence before diagnosing an incident — or just guessed confidently from metadata.

96★
tf.whyattribute

Terraform tells you what drifted. This tells you who changed it, when, and from where — by correlating the plan against CloudTrail.

33★
Self-Healing CI/CDact

When a test fails, an agent reads the failure, writes a fix, and opens a pull request — inside the same pipeline run that caught it.

41★
AutoStackact

Go and Svelte tooling for standing up an application stack without hand-wiring it.

28★

Background

MIT ADT University, Pune B.Tech, Electronics & Computer Engineering Aug 2022 – May 2026
Certified AWS Certified Solutions Architect – AssociateAWS Certified AI PractitionerAWS Certified Cloud PractitionerGitHub Copilot CertifiedRedis Certified Developer (Python)
Works in Go · Python · Kubernetes · AWS SPIFFE/SPIRE · gRPC · MCP · Terraform Open to backend / AI-infrastructure / platform roles

Contact

Let me break something for you. 

Looking for a first role where the hard part is correctness under adversarial conditions — authorization, evaluation, incident response.