RAG development company: answers from your documents, on your server
Boffin Coders is a RAG development company that connects a language model to your own documents, so staff get answers with the source named underneath. Retrieval runs on your server, and the model can too when the documents must not leave it. We start with plain keyword search, which is enough for most first projects, and add a vector database only when real questions show it is needed.
- Runs on your server
- Answers name their source
- No vector database to start
Question
How much leave do part-time staff get?
Searched your documents
- Leave policy, section 421.4
- Part-time contracts17.9
- Holiday calendar 20266.2
Answer
Part-time staff get leave in proportion to their contracted hours.Source: Leave policy, section 4
A RAG chatbot we run, measured in public
Before anything about what we would build for you, the one we run on this site, with its numbers. Open the chat in the corner and test it while you read.
The assistant on this website, measured
The chat assistant on this site answers from our own content, cut into 216 sections. For each question it scores the sections with BM25, a keyword method that needs no model and no database, and sends the best few along with ten pinned sections that hold our prices and rules. We wrote up why a first RAG chatbot does not need a vector database, with the numbers from this one.
One honest limit: this assistant sends the selected sections to a hosted model through an API, because our own pages are public. For a client whose documents must stay put, the model runs on their server instead, the way IELTS Builder runs its speech model.
- Knowledge base
- About 71,500 tokens of this site’s content
- What it used to send
- The first 38,000 characters, 23% of it
- What it sends now
- About 4,500 tokens a question
- Always sent
- 10 pinned sections, about 2,000 tokens
- Search
- BM25 on our server, no vector database
Figures from the index this site builds at deploy time.
Why your RAG chatbot ignores the rules you gave it
Retrieval is where most RAG systems quietly fail: the model answers confidently from whatever part of your content happened to fit, and a rule you wrote can be the part that did not. This is how to measure what yours actually receives, taken from the assistant on this site.
Read the write-upCheck it yourself
context budget ÷ corpus sizeWhat is running, and how we build the rest
The stage of each piece of this service, stated plainly, and how the parts built for your documents come together on your server.
Retrieval on this website
The site assistant searches 216 sections of our content for every question, with pinned rules that are never dropped. It is the RAG system we can show you working, and the write-up gives the command that prints what it sent.
A model on the client’s server
IELTS Builder runs a Whisper speech model inside its own server, with a switch the client holds for sending audio elsewhere. It is speech, not document search, but it is the private deployment pattern this page sells, running today.
Your documents, searched on your own server
For a client, retrieval runs over your own documents on your own server, tested against fifty of your real questions before anyone relies on it. Fine-tuning comes in only when retrieval cannot do the job. Much of our client work is under NDA, so we walk you through the architecture on a call.
LLM integration services, built around your documents
Our LLM integration services cover six pieces of work, each one something the systems above already do, built into your own accounts and tested on your own questions.
Answers from your documents
Manuals, policies, contracts and help pages, split into sections and searched for every question. The answer names the section it came from, so a wrong one can be traced to its source instead of argued about.
Rules that are always sent
Prices, prohibitions and the things the model must never say go into every request, outside the search, so retrieval cannot drop them on the question where they matter most.
Model calls inside your software
A summary, a classification or a first draft added to a system you already run. The model returns text to your code and nothing else. What happens next is decided by your code and your people.
Reading scanned documents
Scans and photos turned into searchable text before retrieval sees them. PDF Toolkit, our own app, runs OCR on the phone itself, and the same approach runs on your server.
Search that shows its working
A command that prints, for any question, which sections were sent and which were cut. When a user reports a bad answer, you can see whether the model was wrong or never saw the right page.
A switch for where the model runs
One setting decides whether the model runs on your server or a hosted API under your account. You hold the switch, and the choice is written into the code rather than buried in a vendor contract.
Private LLM deployment, with the trade-off stated
Private LLM deployment in practice: where each part runs, what leaves your server in each option, and what you give up for keeping everything in.
What stays on your server, and what does not
Private means you can name every place your documents go. The documents, the search index, the search itself and the logs stay on your server or in your cloud account. The model is the one choice with a trade-off: an open-weight model on your own server keeps everything in, but it is slower and less capable than the best hosted models unless you pay for a graphics card to run it.
So we test both on your real questions before you choose, and we write the result down. You can keep the model in-house for everything, or send only the question and a few selected sections to a hosted model and keep the rest at home. Either way, the choice is yours and you can change it later.
- Always on your server
- Documents, index, search and logs
- The model, option one
- An open-weight model on your server
- The model, option two
- A hosted API under your own account
- Leaves the server in option one
- Nothing
- Leaves the server in option two
- The question and the selected sections
Which option suits you is decided on your own questions during the build, not on a public benchmark.
From fifty real questions to a system on your server
Four steps, built around a test you can read: the questions your people already ask, with the answers they should get.
Collect the real questions
Fifty questions your staff or customers actually ask, taken from email, tickets or a shared inbox. They become the test the system has to pass, so nobody grades it on a demo.
- 50 real questions
- The right answer for each
- Where each answer lives
Measure what the model would see
Divide the context budget by the size of your documents. If it is under a third, retrieval decides what the model reads, and the rules need pinning. This takes an afternoon and saves weeks.
- Budget against corpus size
- The rules to pin
- A search method chosen
Build search and test it
Sections, keyword scoring and pinned rules first, run against the fifty questions. A vector database is added only where keyword search fails on a real question, not by default.
- Retrieval on your server
- Pass rate on the 50
- A list of failures
Deploy where the documents live
On your server or cloud account, with the model where you decided it should run. The engineers who build it can stay on by the month, or hand it to your team.
- Code in your repository
- Model where you chose
- A running-cost estimate
Published, not quoted on the call
In USD. Hourly and monthly rates move with the stack and the seniority, from a junior on routine work to a senior on complex builds. Nothing is quoted outside them, and if the number does not work for you, you have saved yourself a meeting.
- A RAG build starts at
- $4,000
- Scoped once, with the lines shown.
- Short pieces of work
- $25-40 an hour
- By seniority and the work.
- A developer by the month
- From $4,000 a month
- Month to month, a month’s notice either way.
- Agency sprint
- $2,000 per two-week sprint
- Under your brand, invoiced after you see the work.
Questions to ask a RAG development company
Cost, where the model runs, whether you need a vector database and what it still gets wrong.
A system that answers questions from your own documents. It splits the documents into sections, finds the ones that match each question, and gives only those to a language model with an instruction to answer from them and name the source. The work is mostly in the search and in testing it on your real questions. The model is the easy part.
Not answered here?
The other AI jobs, and where they live
If the model has to act in your tools rather than answer questions, that is an AI agent with a person on the send button.
If the answers are for customers on your website rather than for staff, see website chatbots built on your own content.
Not sure which you need? Our AI automation page sorts the four jobs in one screen.
Work we’ve delivered, and what we’ve written
Bring ten real questions to a 20-minute call
Send us the questions your staff ask most and tell us where the documents live. On the call we will say whether retrieval can answer them, where the model should run and what it would cost.