RWA Chatbot

A FastAPI reasoning agent answering domain queries from validated per-tenant data snapshots, not live queries.

  • 2026–present
  • Professional · RWA

Why there is no measurement

Measured internally at RWA — not mine to publish

No measurement

Why not

This runs in production and its answer quality is tracked internally through the eval harness described below, but those numbers are RWA's rather than mine and are not mine to publish. I would rather show the gap than publish a figure I cannot let anyone check, or leave the project off the site because the result is inconvenient to evidence.

Problem

A chat interface over live data either runs a fresh query per question, which is slow and lets a bad query reach production data directly, or it hallucinates an answer that looks plausible and isn't. Every customer also needed answers scoped to their own tenant and sector, from a shared service, without a prompt alone being trusted to enforce that boundary.

Approach

Snapshots first, reasoning second. Scheduled jobs precompute and validate per-tenant data snapshots ahead of any question being asked, so the agent answers from a checked, versioned dataset rather than a live query it constructed itself. On top of that sits a reasoning loop — intent classification, tool orchestration, session memory — reading a database- backed knowledge store for grounding context, with its own eval harness run against a fixed corpus of domain queries so a change in the prompt or the model has a regression suite to answer to before it reaches production.

Components

  1. Reasoning agent loop

    Evidence: MEASURED INTERNALLY — NOT PUBLISHABLE

    Intent classification and tool orchestration rather than a single prompt-and-respond call, so a query is routed to the tools it actually needs.

  2. Snapshot-backed data layer

    Evidence: MEASURED INTERNALLY — NOT PUBLISHABLE

    Scheduled jobs precompute per-tenant data snapshots ahead of time, so the agent answers from a validated snapshot rather than a live query it constructed itself.

  3. Snapshot validation

    Evidence: MEASURED INTERNALLY — NOT PUBLISHABLE

    A check run on a snapshot before it is allowed to serve, kept separate from the job that produces it so a bad snapshot fails closed rather than reaching a customer.

  4. Multi-tenant sector configuration

    Evidence: MEASURED INTERNALLY — NOT PUBLISHABLE

    JSON-driven catalogs mapping tenants to sectors and modules, read at query time to decide what a given customer's questions can see.

  5. Knowledge store

    Evidence: MEASURED INTERNALLY — NOT PUBLISHABLE

    Internal reference material ingested into the database and retrieved as grounding context for an answer, rather than baked into the prompt by hand.

  6. Feedback-driven correction

    Evidence: MEASURED INTERNALLY — NOT PUBLISHABLE

    A downvoted answer is analysed by its own job, feeding a route back into what the agent is told rather than a rating that goes nowhere.

  7. Eval harness

    Evidence: MEASURED INTERNALLY — NOT PUBLISHABLE

    A corpus-driven regression suite over domain queries, with an LLM-as-judge pass locally and a lighter smoke subset that runs in CI on every change.

  8. Observability

    Evidence: MEASURED INTERNALLY — NOT PUBLISHABLE

    A staff-facing view onto chatbot activity, separate from the customer-facing chat interface itself.

Snapshots before reasoning

The instinct with a data chatbot is to let the model write the query. This service deliberately does not: scheduled jobs produce and validate a per-tenant snapshot of the data ahead of any question being asked, and the reasoning agent answers from that rather than from a query it generated itself against production. It’s a slower path to build and a much smaller surface for a bad question to do damage through.

An eval harness rather than a vibe check

The reasoning layer — intent classification, tool orchestration, a knowledge store for grounding — is the part most likely to regress silently when a prompt or a model version changes. It has a corpus-driven suite behind it precisely so a change has something to answer to before it reaches production, with a lighter subset wired into CI so the full suite isn’t the only gate.

What is deliberately absent here

No client names, no internal metrics, no architecture diagram of an RWA system, and no product name or industry sector, even where the service’s own README states them plainly. The employer boundary on this site is not a formality — this is a live production system belonging to my employer.