Manufacturing company · Jul–Sep 2026

A governed AI gateway for the whole company

One sign-in that lets AI assistants answer ERP, business, and IT questions with the same numbers finance would give — and only the data each person's role allows.

Business impactEvery employee can ask the business a question in plain English and get a governed, auditable answer — without a new data-exposure risk.

Role: Set the AI strategy; architect and technical lead from proof of concept to company-wide rollout

golden questions answered correctly (Sonnet)
212 / 212
on a blind holdout scored before any fixes
90.8%
commit to production, gated by evals
~7 min

The problem

People were starting to ask AI assistants about the business — open orders, units built, how the ERP does something — and getting confident, wrong answers. Models will happily invent a query or a revenue number. Leadership needed AI that could reach real company data, give the same number finance would give, respect who is allowed to see what, and be provably correct before every change.

How it works

  1. 01

    One connector, one sign-in

    Users add a single connector and sign in with company SSO. Every domain (ERP, IT, metrics) sits behind one endpoint, so there is one thing to connect and one place to govern.

  2. 02

    Deny-by-default access

    Access follows the person's role. Anything not explicitly granted is refused, and a missing grant fails the build rather than slipping through.

  3. 03

    Served metrics, not generated SQL

    Each business number is a registered metric with a vetted definition and test questions. The gateway returns a finished answer — the model relays the number instead of computing it.

  4. 04

    Knowledge that checks itself

    ERP and IT knowledge lives as small cards, each verified against live data on a schedule. Cards that stop matching reality are flagged stale automatically.

  5. 05

    Release gates

    Every change runs the full evaluation suite before it can deploy. A failed eval blocks the release.

What I built

  • The gateway service, including SSO and sessions that survive redeploys so users aren't logged out.
  • A served-metrics catalog that grew from 10 to roughly 170 business metrics, with financial figures restricted by role.
  • Verified knowledge bases for the ERP (~320 cards) and for IT runbooks and systems, both published through the gateway.
  • The evaluation program: test questions per metric, a retrieval test set, and a 390-question persona eval written by independent authors who never saw the metric catalog.

Results

  • Golden-question accuracy across three model tiers: Haiku 211/212, Sonnet 212/212, Opus 210/212.
  • Persona eval (26 personas, 390 questions) improved from 251/390 to 378/390 through targeted fixes; a blind holdout scored 118/130 (90.8%) before any fixes were applied to it.
  • Retrieval test set at 53/53, holdout 10/10.
  • Rolled out company-wide in September 2026.

What I learned

  • Don't let the model do arithmetic on the business. Serve the number; let the model handle the language.
  • Evals are a management tool, not just a test suite: they let you say yes to change without betting on it.
  • Governance is easiest at one chokepoint. One endpoint, one access policy, many AI clients.

Stack

  • Python
  • Model Context Protocol (MCP)
  • Anthropic Claude
  • SQL Server
  • Docker
  • CI/CD
  • pytest