Skip to main content

RAG development company: answers from your documents, on your server

Boffin Coders is a RAG development company that connects a language model to your own documents, so staff get answers with the source named underneath. Retrieval runs on your server, and the model can too when the documents must not leave it. We start with plain keyword search, which is enough for most first projects, and add a vector database only when real questions show it is needed.

A RAG build starts at $4,000 · $25-40 an hour for shorter work · a developer by the month from $4,000 a month
  • Runs on your server
  • Answers name their source
  • No vector database to start
Your serverIllustration. Figures are invented.

Question

How much leave do part-time staff get?

Searched your documents

  • Leave policy, section 421.4
  • Part-time contracts17.9
  • Holiday calendar 20266.2
Always sent: the rule “never quote salary figures”

Answer

Part-time staff get leave in proportion to their contracted hours.Source: Leave policy, section 4

Proof first

A RAG chatbot we run, measured in public

Before anything about what we would build for you, the one we run on this site, with its numbers. Open the chat in the corner and test it while you read.

The worked example

The assistant on this website, measured

The chat assistant on this site answers from our own content, cut into 216 sections. For each question it scores the sections with BM25, a keyword method that needs no model and no database, and sends the best few along with ten pinned sections that hold our prices and rules. We wrote up why a first RAG chatbot does not need a vector database, with the numbers from this one.

One honest limit: this assistant sends the selected sections to a hosted model through an API, because our own pages are public. For a client whose documents must stay put, the model runs on their server instead, the way IELTS Builder runs its speech model.

Knowledge base
About 71,500 tokens of this site’s content
What it used to send
The first 38,000 characters, 23% of it
What it sends now
About 4,500 tokens a question
Always sent
10 pinned sections, about 2,000 tokens
Search
BM25 on our server, no vector database

Figures from the index this site builds at deploy time.

How we actually build it

Why your RAG chatbot ignores the rules you gave it

Retrieval is where most RAG systems quietly fail: the model answers confidently from whatever part of your content happened to fit, and a rule you wrote can be the part that did not. This is how to measure what yours actually receives, taken from the assistant on this site.

Read the write-up

Check it yourself

context budget ÷ corpus size
Where we are

What is running, and how we build the rest

The stage of each piece of this service, stated plainly, and how the parts built for your documents come together on your server.

Running for ourselves

Retrieval on this website

The site assistant searches 216 sections of our content for every question, with pinned rules that are never dropped. It is the RAG system we can show you working, and the write-up gives the command that prints what it sent.

Running with a client

A model on the client’s server

IELTS Builder runs a Whisper speech model inside its own server, with a switch the client holds for sending audio elsewhere. It is speech, not document search, but it is the private deployment pattern this page sells, running today.

How we approach it

Your documents, searched on your own server

For a client, retrieval runs over your own documents on your own server, tested against fifty of your real questions before anyone relies on it. Fine-tuning comes in only when retrieval cannot do the job. Much of our client work is under NDA, so we walk you through the architecture on a call.

What we build

LLM integration services, built around your documents

Our LLM integration services cover six pieces of work, each one something the systems above already do, built into your own accounts and tested on your own questions.

Answers from your documents

Manuals, policies, contracts and help pages, split into sections and searched for every question. The answer names the section it came from, so a wrong one can be traced to its source instead of argued about.

Rules that are always sent

Prices, prohibitions and the things the model must never say go into every request, outside the search, so retrieval cannot drop them on the question where they matter most.

Model calls inside your software

A summary, a classification or a first draft added to a system you already run. The model returns text to your code and nothing else. What happens next is decided by your code and your people.

Reading scanned documents

Scans and photos turned into searchable text before retrieval sees them. PDF Toolkit, our own app, runs OCR on the phone itself, and the same approach runs on your server.

Search that shows its working

A command that prints, for any question, which sections were sent and which were cut. When a user reports a bad answer, you can see whether the model was wrong or never saw the right page.

A switch for where the model runs

One setting decides whether the model runs on your server or a hosted API under your account. You hold the switch, and the choice is written into the code rather than buried in a vendor contract.

Private LLM deployment

Private LLM deployment, with the trade-off stated

Private LLM deployment in practice: where each part runs, what leaves your server in each option, and what you give up for keeping everything in.

Private LLM deployment

What stays on your server, and what does not

Private means you can name every place your documents go. The documents, the search index, the search itself and the logs stay on your server or in your cloud account. The model is the one choice with a trade-off: an open-weight model on your own server keeps everything in, but it is slower and less capable than the best hosted models unless you pay for a graphics card to run it.

So we test both on your real questions before you choose, and we write the result down. You can keep the model in-house for everything, or send only the question and a few selected sections to a hosted model and keep the rest at home. Either way, the choice is yours and you can change it later.

Always on your server
Documents, index, search and logs
The model, option one
An open-weight model on your server
The model, option two
A hosted API under your own account
Leaves the server in option one
Nothing
Leaves the server in option two
The question and the selected sections

Which option suits you is decided on your own questions during the build, not on a public benchmark.

How it works

From fifty real questions to a system on your server

Four steps, built around a test you can read: the questions your people already ask, with the answers they should get.

Collect the real questions

Fifty questions your staff or customers actually ask, taken from email, tickets or a shared inbox. They become the test the system has to pass, so nobody grades it on a demo.

  • 50 real questions
  • The right answer for each
  • Where each answer lives

Measure what the model would see

Divide the context budget by the size of your documents. If it is under a third, retrieval decides what the model reads, and the rules need pinning. This takes an afternoon and saves weeks.

  • Budget against corpus size
  • The rules to pin
  • A search method chosen

Build search and test it

Sections, keyword scoring and pinned rules first, run against the fifty questions. A vector database is added only where keyword search fails on a real question, not by default.

  • Retrieval on your server
  • Pass rate on the 50
  • A list of failures

Deploy where the documents live

On your server or cloud account, with the model where you decided it should run. The engineers who build it can stay on by the month, or hand it to your team.

  • Code in your repository
  • Model where you chose
  • A running-cost estimate
What it costs

Published, not quoted on the call

In USD. Hourly and monthly rates move with the stack and the seniority, from a junior on routine work to a senior on complex builds. Nothing is quoted outside them, and if the number does not work for you, you have saved yourself a meeting.

A RAG build starts at
$4,000
Scoped once, with the lines shown.
Short pieces of work
$25-40 an hour
By seniority and the work.
A developer by the month
From $4,000 a month
Month to month, a month’s notice either way.
Agency sprint
$2,000 per two-week sprint
Under your brand, invoiced after you see the work.

Every price on one page →

FAQ

Questions to ask a RAG development company

Cost, where the model runs, whether you need a vector database and what it still gets wrong.

A system that answers questions from your own documents. It splits the documents into sections, finds the ones that match each question, and gives only those to a language model with an instruction to answer from them and name the source. The work is mostly in the search and in testing it on your real questions. The model is the easy part.

Not answered here?

Not quite this job?

The other AI jobs, and where they live

If the model has to act in your tools rather than answer questions, that is an AI agent with a person on the send button.

If the answers are for customers on your website rather than for staff, see website chatbots built on your own content.

Not sure which you need? Our AI automation page sorts the four jobs in one screen.

Bring ten real questions to a 20-minute call

Send us the questions your staff ask most and tell us where the documents live. On the call we will say whether retrieval can answer them, where the model should run and what it would cost.