Skip to content
← All work

AI quality gate & automated PR security reviewer

Gatekeeper AI

An AI senior security reviewer inside CI/CD: Claude plus static scanners review every pull request, triage false positives and block risky merges.

My role: Design and development of the Python engine, CI/CD integrations and Nuxt dashboard

Discuss a similar project
  • DevSecOps
  • Claude
  • Python 3.12
  • GitHub Actions
  • Nuxt 4
Screenshot of Gatekeeper AI, ai quality gate & automated pr security reviewer, built by Anshuman Verma

Overview

Gatekeeper AI is an enterprise-grade quality gate for GitHub Actions and Bitbucket Pipelines. It reviews the changed code in each pull request with Claude alongside Semgrep, Bandit, Gitleaks and SonarQube, filters out false positives, posts inline comments with ready-to-paste fixes, and enforces PASS/FAIL thresholds so insecure code cannot be merged. A Nuxt 4 dashboard shows pass rates, token spend and risk categories across repositories.

The challenge

Security problems slip through pull request review because reviewers are busy and static analysis tools are noisy. Teams that run Semgrep or SonarQube on every PR quickly learn to ignore the flood of warnings, including the real ones.

Gatekeeper AI set out to act like an automated senior security reviewer in the pipeline: look only at what changed, combine scanners with an LLM that understands context, drop the false positives, and stop a risky merge before it reaches production, without slowing down everything else.

How it works

  1. 1Pull request
  2. 2Changed hunks
  3. 3Scanners + Claude
  4. 4False-positive triage
  5. 5PASS / FAIL gate
  6. 6Inline comments + alerts

Architecture

The core engine and CLI are Python 3.12 with Typer, Pydantic v2, Rich, GitPython and httpx, managed with uv. It runs as a GitHub Action (with SARIF 2.1.0 upload) or a Bitbucket Pipelines pipe (with Code Insights reports). For each pull request it extracts the changed hunks, and every scanner's results are filtered to those lines so untouched legacy code is not re-reported.

Semgrep, Bandit, Gitleaks and the SonarQube API run alongside two Claude 3.5 Sonnet passes, one for security and one for code quality, using tool calling and structured JSON output. Scanner alerts are then sent to Claude with the surrounding code for false-positive triage. Before any code leaves the pipeline, a high-entropy redactor strips secrets, and SHA-256 content-addressed caching skips chunks that were already reviewed.

The gate computes a verdict of PASS, FAIL, ERROR or SKIPPED against policy thresholds, posts a summary and inline comments with suggested fixes, and sends alerts to Slack, Microsoft Teams Adaptive Cards and email. Baselines (.gatekeeper/baseline.json) and inline ignore markers accommodate known debt. Runs are published to a Nuxt 4 dashboard on Neon PostgreSQL with Drizzle that shows pass rates, LLM spend, risk categories and a PR audit drawer. A deterministic mock LLM lets tests and offline CI run without API tokens.

What I built

Automated PR quality gate

Runs inside GitHub Actions and Bitbucket Cloud, analyses only the changed hunks, returns PASS, FAIL, ERROR or SKIPPED, and comments on the exact lines with suggested fixes.

Scanner + AI consensus

Semgrep, Bandit, Gitleaks and SonarQube run alongside two Claude passes, one for security and one for code quality, all filtered to the lines the PR touched.

LLM false-positive triage

Raw scanner alerts go to Claude with surrounding code context, so noise is dismissed before a developer or security lead sees it.

Redaction and caching

A high-entropy secret redactor cleans diffs before they reach the LLM, and SHA-256 content-addressed caching cuts latency and API spend on re-runs.

Alerts everywhere

Failure reports to Slack, Microsoft Teams Adaptive Cards and SMTP email, plus SARIF, HTML and Markdown reports.

Dashboard and CLI

A Nuxt 4 dashboard with gate pass rates, LLM spend, top risk categories and a PR audit drawer; a Typer CLI with audit, init, baseline and validate-config commands.

Engineering decisions

  • Scan only the lines a PR changes, so legacy code does not block new work, with baselines and inline ignore markers for known debt.
  • Redact secrets before any code leaves the pipeline for an external LLM.
  • Ship a deterministic mock LLM so CI and tests run offline without spending API tokens.

Results

  • All six phases of the specification delivered.
  • 85 automated tests with 87% coverage, mypy --strict and Ruff with zero errors.
  • Exact token and USD cost tracked for every run, on the terminal report and the dashboard.
  • Works in both GitHub Actions and Bitbucket Pipelines.

Stack

Core & CLI
Python 3.12, Typer, Pydantic v2, Rich, GitPython, httpx, uv
AI
Claude 3.5 Sonnet (tool calling, structured JSON output), mock LLM client
Static analysis
Semgrep, Bandit, Gitleaks, SonarQube API
CI/CD
GitHub Actions (SARIF 2.1.0), Bitbucket Pipelines (Code Insights)
Dashboard
Nuxt 4, Vue 3, Tailwind CSS, Better Auth, Neon PostgreSQL, Drizzle
Quality
Pytest (85 tests, 87% coverage), mypy --strict, Ruff

Related services

Need something like this?

I can build a version of this for your product, your data and your stack.

Start a project