Somewhere at your company, someone already wrote down the answer to your question. It's in a Google Doc from 18 months ago, or a Slack thread, or a README three repos deep. You have no idea which one. So you ask in a channel, and the one person who knows gets interrupted for the fifth time that week.
For most of this year I've been building our fix for that at Credit Karma. It's an internal RAG app (retrieval-augmented generation: look up the relevant docs first, then have the model answer from them). It takes in the documentation we already have and answers questions about it in Slack, in the browser, or inside your editor over MCP.
The short version
You point it at some docs. It cuts them into pieces, turns each piece into a vector, and stores them in Postgres. When you ask a question, it finds the most relevant pieces, hands them to Gemini, and gets back an answer based on your docs instead of whatever the model half-remembers. Before you see anything, it grades its own answer for hallucinations and quality.
The stack: Next.js 16 and React 19 on the front, Prisma and Postgres with pgvector for storage, BullMQ on Redis for background jobs, and Vertex AI plus Gemini for the model calls. All model traffic goes through an internal AI gateway that speaks the OpenAI API format.
One image, four programs
My favorite part of the architecture is also the dumbest. The whole thing is one codebase and one Docker image, and a single env var decides which of four programs it boots as. The Dockerfile ends with:
CMD ["bash","-c","npm run ${NPM_CMD:-start}"]start(the default) boots the web UI and HTTP API.start-workerboots the document ingestion worker.start-slack-workerboots the Slack bot.start-mcpboots a standalone MCP server.
That's four Kubernetes deployments from one build pipeline, with no shared library to package and publish. The web app and the MCP server import the same ask() function, so when retrieval gets better, every surface gets it at once. At a bigger company this would be a monorepo with four packages and a publish step. I'm glad every week that it isn't.
Getting documents in
The ingestion worker pulls jobs off a queue. Each job fetches source text from Google Docs, Sheets, GitHub or Slack threads and runs it through the chunker.
The chunker splits every document twice. First into big parent chunks of about 1,500 characters, then each parent into small child chunks of about 300. Only the children get embedded (768-dimension vectors from Vertex's text-embedding-004) and written to pgvector. That second split turned out to be the most important decision in the app. Step 3 below is why.
Chunks are also deduplicated by hash, so boilerplate that shows up in forty docs gets stored and embedded once.
Answering a question
- Embed the question with the same model that embedded the docs.
- Vector search. pgvector's
<=>cosine-distance operator pulls the top 30 child chunks, limited to the tags you asked about. - Return the parents. The same query joins each matching child back to its parent and returns the parent's text. You match on a tight 300-character idea, and the model gets 1,500 characters of context to write from. Match small, read big.
- Hybrid rerank. The 30 candidates also get scored by BM25 keyword search. The two rankings get combined and the top 10 survive. Vector search is bad at exact identifiers like error codes, flag names and service names, and BM25 is good at exactly those. If the keyword service is down, it falls back to vector-only.
- Generate with Gemini 2.5 Pro.
- Judge the answer. Three cheap models grade the expensive one. A hallucination check (an open-source model called LettuceDetect) returns a groundedness score and highlights the exact phrases it couldn't find in the sources. A relevance judge and a quality judge score the rest. There's also a check for whether the answer is just the canned "I don't know."
- Save everything: the answer, every score, the flagged phrases and links to the sources.
Steps 2 and 3 are one raw SQL query. Simplified, it looks like this:
SELECT p.id, p.text, c.embedding <=> $1 AS distance
FROM child_chunks c
JOIN parent_chunks p ON p.id = c.parent_id
WHERE c.tag = ANY($2)
ORDER BY distance
LIMIT 30;Step 7 is the one I'd push hardest on. Every answer the bot has ever given is a database row with a groundedness score and a quality grade attached. That's an evaluation dataset that builds itself. "Did that retrieval change make things better?" turns into a SQL query instead of a gut feeling.
There's also a fast path. Send stream: true and tokens come straight back from Gemini, but it skips the judges and saves nothing. You trade measurement for speed, and you have to ask for it.
The Slack bot
Most people use it here. The Slack worker runs @slack/bolt in socket mode, which means the bot opens an outbound WebSocket and we don't have to expose a public webhook. That keeps the networking simple inside a private cluster.
Ask a question in a channel it watches and it drops a 🤔 on your message right away, so you know it heard you during the ~10 seconds it takes to answer. Then it replies in the thread with the answer, source links and 👍/👎 buttons. The feedback goes into the same table as the web answers.

It reads channel history. Every night, a cron job backfills each watched channel and stores resolved threads as documents, so questions that got answered in conversation become searchable. Teams opt into this, and answers from threads come back as a separate follow-up message. You can always tell "the docs say this" apart from "someone said this in a thread once."

It can use tools. The Slack agent has MCP toolsets, so for some questions it goes and checks a live system instead of quoting a doc. When a tool answered, the thread follow-up stays quiet. Otherwise you'd get a fresh live answer followed by an old thread saying the opposite.
In your editor
The MCP server exposes four tools to any MCP client, like Cursor or Claude Code:
askreturns a full answer.getChunksreturns raw retrieval results, for when the agent wants to do its own reasoning.getTagslists the available tags.inferTagpicks the right tag for a question.
This turned out to be the most useful surface, almost by accident. A coding agent that can look up your internal docs in the middle of a task stops guessing at your conventions. And since it's just another caller of ask(), it cost almost nothing to add.
Tags do most of the work
Every document is tagged, and every question is asked against a tag. A tag sets what gets searched. It also carries its own "I don't know" message, custom prompt rules, the Slack channels it listens to, and whether it reads Slack history.
So one deployment serves many teams without real multi-tenancy. A team creates a tag, points it at their docs and adds the bot to their channel.
There's an admin area too, with document management, tag config, MCP toolset binding, a usage dashboard, a playground for testing docs before you commit them, and the Bull Board queue UI. "Why hasn't my doc shown up yet?" gets answered without anyone opening a terminal.
Deploys and tests
It runs on Kubernetes and deploys through GitOps. There's no deploy button. Merging to main opens a pull request against the cluster config for each environment, and a person merges the one they want. It's slower than I'd like and I've complained about it. But the cluster's state is a file in git I can diff and revert, and I've come around on that.
Testing is the weak spot. There's no unit test framework. We have typecheck, lint, a few assertion scripts for the prompt contracts that break silently, and replayable "validation manifests": JSON files of route-and-expected-content checks that drive the real app in a headless browser. It's cheaper than a real end-to-end suite and a lot better than clicking around and hoping. A proper E2E suite that stands up its own database and exercises the write paths is next on my list.
If you're building one of these
Start with the parent/child split and hybrid search. Those two did more for answer quality than any amount of tuning a single chunk size.
Score and save every answer from the first day. Adding evaluation to a system that's already in production is miserable, and without it "this feels better" can't be checked.
Put every model call behind one OpenAI-compatible endpoint. We've moved the judges between models several times, and each move was a one-line diff.
And ship it where people already ask questions. The web UI is fine. The Slack bot is what gets used.