Skip to content
← All work

Evidence-first AI agent for production incidents

AI DevOps Incident Investigator

A multi-tenant SaaS where an AI agent investigates production incidents like an on-call engineer, cites every claim, and waits for human approval before acting.

My role: Design and full-stack development: agent, tools, SaaS layer, UI, eval harness and target service

Visit GitHub ↗
  • LangGraph
  • MCP
  • FastAPI
  • Next.js 15
  • Evals
Screenshot of AI DevOps Incident Investigator, evidence-first ai agent for production incidents, built by Anshuman Verma

Overview

The hard part of an incident agent is not calling a model, it is trust: knowing whether to believe what it tells you. This platform is built around that. The agent forms hypotheses, gathers evidence from logs, metrics, the database, deploy events and Git history, rules wrong causes out visibly, and writes a report in which every claim cites the tool result it came from. It stops before taking any action: opening a GitHub issue requires a signed-in human, and the decision is audited. A six-scenario eval suite publishes how often it is right: mean score 0.83, the "is this even our bug?" call correct 6/6, at about $0.08 per investigation.

The challenge

When production breaks, on-call engineers spend the first hour gathering evidence across logs, metrics, databases, deploys and Git history. AI agents promise to speed this up, but most of them produce confident answers with no way to tell whether they are right, and some will happily take actions on a guess.

This project set out to build an agent an engineer could actually trust: one that shows its reasoning, proves every claim with real tool output, can conclude "this isn't our bug" when a third party is at fault, never writes anything without human approval, and is measured against known answers rather than demos.

How it works

  1. 1Alert / incident
  2. 2Triage
  3. 3Hypotheses
  4. 4Evidence via MCP tools
  5. 5Confirm or refute
  6. 6Cited report → human approval

Architecture

The agent is a LangGraph graph with triage, hypothesis generation, evidence gathering, evaluation and report nodes. One model call judges the last tool result and chooses the next step, keeping a full investigation to 6–9 calls. Tools are exposed through an MCP server and are read-only: log search with error-signature grouping, metrics with change-point detection, SQL guarded by sqlglot (SELECT-only, row-capped, statement timeout, read-only connection), Git history via GitPython and deploy events.

Every claim in the report must cite a source reference a tool actually returned; a validator removes invented citations and downgrades unsupported "confirmed" hypotheses. A cause is confirmed only with evidence from both the symptom side and the change side, and if the model tries to stop early, the loop fetches the missing side itself. Tool output is wrapped as untrusted data with secret and PII redaction. Before its only write action, opening a GitHub issue, the graph interrupts; a signed-in reviewer can edit, approve or reject, and the checkpointed state lets a run resume hours later in another process.

Around the agent is a multi-tenant SaaS on FastAPI and SQLAlchemy 2.0: organisations as the data boundary (cross-tenant requests return 404), per-organisation roles, plan limits enforced with 402, usage metering, hashed API keys for alert ingestion and Stripe Checkout in test mode. A Next.js 15 UI streams each step over SSE and resumes from a sequence number after a refresh. Every model call and tool result is recorded to a cassette, so investigations, the demo and CI evals replay at zero cost. A purpose-built FastAPI checkout service with seeded Git history and a fault injector provides six scenarios with ground truth: bad deploy, config change, rotated secret, N+1 query, feature flag and third-party outage. OpenTelemetry spans carry tokens, cost and latency for every node and call.

What I built

Hypothesis loop, not a pipeline

A LangGraph graph (triage, hypotheses, evidence, evaluation, report) where one model call both judges the last tool result and picks the next move, so a full investigation takes 6–9 calls.

Citations that can't be invented

Every claim must reuse a source reference a tool actually returned; invented citations are dropped and logged, and a "confirmed" hypothesis without real evidence is downgraded.

Two-sided corroboration

A cause is confirmed only with evidence from both the symptom side (logs, metrics, SQL) and the change side (deploys, commits, diffs); missing evidence is fetched deterministically with no extra model call.

Read-only tools over MCP

Log search with error-signature grouping, metrics with change-point detection, guarded SQL (parsed, SELECT-only, row-capped, read-only connection), Git history and deploy events.

Human approval gate

The graph pauses before its only write action; a reviewer can edit the root cause or fix and approve or reject. State is checkpointed, so a run can wait hours and resume in another process.

Live run UI and replay

A Next.js page streams each step over SSE with the tool call behind every entry. Every model call and tool result is recorded, so investigations replay exactly with no API key.

Engineering decisions

  • Accept a root cause only when symptom and change evidence agree, which keeps "this is a third-party outage" a reachable answer.
  • Treat tool output as untrusted data, with secret and PII redaction before it reaches the model, as prompt-injection defence.
  • Build a target service that breaks in known ways, so every investigation has ground truth and the agent's accuracy is measured, not assumed.
  • Cap cost per run in USD with token and time budgets, prompt caching and a small model for triage.

Results

  • Mean eval score 0.83 across six scripted incidents, with failures published alongside passes.
  • The "is this even our bug?" call correct in 6 of 6 scenarios.
  • About $0.08 per investigation, with a hard per-run cost cap.
  • 96 tests across five packages, with a CI gate that fails the build below the score threshold.

Stack

Agent
LangGraph (stateful graph, interrupts, checkpointing), Pydantic structured outputs, Claude / OpenAI / Gemini, prompt caching
Tools
MCP server with read-only tools, sqlglot SQL guard, GitPython
Backend
Python 3.11+, FastAPI, SQLAlchemy 2.0, SSE, background workers, event bus
Frontend
Next.js 15 (App Router), React 19, TypeScript, EventSource streaming
SaaS
Organisations and roles, plan limits and usage metering, Stripe Checkout (test mode), hashed API keys, audit trail
Quality
96 tests, record/replay cassettes, eval harness with CI score gate, OpenTelemetry tracing, Ruff, GitHub Actions

Related services

Need something like this?

I can build a version of this for your product, your data and your stack.

Start a project