Voice AI agent development that starts with speech to text on your server
Boffin Coders does voice AI agent development starting from the part that decides where a recording goes: speech to text. We run the transcription model on your own server, so audio becomes text without being sent to an AI vendor. IELTS Builder runs this way today, with a teacher approving every grade. A phone agent can be built on top, and it says it is an AI in its first sentence.
- Speech to text on your server
- The agent says it is an AI
- A person takes over on request
Agent, first sentence
“Hello, you are speaking to the automated assistant for the clinic. I can take a message or put you through to a person.”
Self-hosted speech to text, running for a client today
Before a word about phone agents, the voice system we actually run. Self-hosted speech to text is the part that decides where a person's voice ends up, so it is the part we build first.
IELTS Builder transcribes every spoken answer on its own server
Students on IELTS Builder answer speaking questions out loud in the browser. Each answer is converted on the platform’s own server to the format the model needs, transcribed by a Whisper model running in the same process, and only the text goes on to scoring. A teacher approves every grade before a student sees it.
Two things we do not hide. The recording is kept, in the platform’s own private storage, because a failed transcription is retried from it later. And there is a switch that can send audio to a hosted service instead. The client holds that switch, not us and not a vendor.
- Model
- openai/whisper-small
- Runs
- In the API process, on CPU
- Audio the model gets
- 16 kHz, mono, PCM
- Transcriptions at once
- 1, in a queue
- Audio sent to an AI vendor
- None, in local routing
- Submissions
- 4,584 in a thirty-day window
Model settings from the IELTS Builder code.
Your speech-to-text sends the recording. It only needed the words.
A voice system has to turn speech into text before it can do anything, and most of them upload the voice to a vendor to get it. This is how the model runs on your own server instead, taken from IELTS Builder, including what it costs you and why the recording is still kept.
Read the write-upCheck it yourself
ffmpeg -i clip.webm -ar 16000 -ac 1 -c:a pcm_s16le clip.wav && ls -l clip.*What is running, and how we build the rest
Our voice work, stated plainly: one system live with a client, the phone agent we build on top of it, and two principles every voice agent keeps.
Self-hosted speech to text
Inside IELTS Builder, transcribing spoken exam answers on the platform’s own server, with a teacher approving every grade. It is the voice system we can show you, and the write-up below explains every setting.
A phone agent on top
Answering a line, taking a message, booking into a calendar with a confirmation step, handing over to a person. It is built on the transcription above, measured on your own recordings first, and goes live with a person watching every transcript and handover.
An agent that is honest about itself
Every agent we build uses a stock voice and never clones a real person’s voice. It says it is an AI in its first sentence. Both protect the business that runs it from legal and reputational risk, and callers trust an agent that is straight with them.
Voice AI that keeps the recording on your side
Six pieces. The first four run in IELTS Builder today. The summaries and the phone agent are built on top of them when you want them.
Transcription on your server
A Whisper model running in your own server process turns recordings into text. The audio is converted and read on the same machine, and only the text moves on.
Recordings kept where you keep things
The original audio goes to your own private storage, not a vendor’s. You decide how long it is kept and who can play it back, and the policy is written down.
Retries from the original
A transcription that comes back empty is retried later from the stored recording, instead of failing a person’s answer or a caller’s message silently.
A switch you hold
One setting chooses between the local model and a hosted service. It lives in your configuration, set by you, so where audio goes is never decided by a vendor update.
Summaries a person checks
Transcripts turned into a short summary, a list of actions or a draft reply, which a person reads before anything is saved or sent.
A phone agent that says what it is
Built on top when you want one. It says it is an AI in its first sentence, answers from your own information, and puts the caller through to a person when asked.
Voice AI agent development, disclosure first
What the phone agent will do, and what it will not
A phone agent is a speech-to-text model, a language model and a voice, in that order. The first part decides where the caller’s voice goes, which is why we build it first and on your server. The rest is built around a script you approve: what the agent answers, what it takes a message for, and when it hands over.
Disclosure is not an option you switch on. The agent’s first sentence says it is an automated assistant for your business, and a caller who asks for a person gets one or gets a message taken for one. In many places the law expects it, and it is the honest way to run one anyway.
- First sentence
- Says it is an AI assistant
- Answers from
- Your own information, nothing else
- Bookings
- Prepared, then confirmed by your system or a person
- Asks for a person
- Transfers, or takes a message
- Voice
- A stock voice, never a cloned one
The rules every phone agent we build follows, walked through line by line on a call.
From your own recordings to a system on your server
Four steps, starting with audio you already have, so the model is judged on your accents and your line quality rather than on a demo.
Start with recordings you already have
A sample of your real calls, meetings or voice notes, transcribed on a server you control. We measure how well the model handles your accents, your line quality and your vocabulary before anything else is built.
- Your audio, transcribed
- Errors listed by type
- Model size chosen
Decide what is kept
Where the recordings live, how long they are kept, who can play them back and whether the switch to a hosted service is ever used. Written down before the first live recording, not after.
- Storage and retention set
- Access written down
- Routing switch set
Build what reads the text
Summaries, actions or a draft reply from each transcript, with a person reading them. For a phone agent, this is where the script, the disclosure and the handover are written and tested.
- Outputs a person checks
- Script approved by you
- Disclosure in line one
Run it with a person on hand
It goes live with someone watching the transcripts and the handovers, and we fix what they find. The code, the model and the recordings stay on your side when we hand over.
- Handovers reviewed
- Fixes from real use
- Code and model in your name
Published, not quoted on the call
In USD. Hourly and monthly rates move with the stack and the seniority, from a junior on routine work to a senior on complex builds. Nothing is quoted outside them, and if the number does not work for you, you have saved yourself a meeting.
- A voice build starts at
- $4,000
- Scoped once, with the lines shown.
- Short pieces of work
- $25-40 an hour
- By seniority and the work.
- A developer by the month
- From $4,000 a month
- Month to month, a month’s notice either way.
- Agency sprint
- $2,000 per two-week sprint
- Under your brand, invoiced after you see the work.
Voice AI questions before you book
Disclosure, where the audio goes, speed, what is kept, cloning, and how a phone agent is built, step by step.
Yes, always. The agent says it is an automated assistant in its first sentence, and a caller who asks for a person is transferred or has a message taken. We do not build agents that pass for a person, and we do not clone a real person’s voice. Disclosure is the default on every voice agent we build, not a setting for regulated industries.
Not answered here?
Voice is one of four AI jobs
If the questions arrive as typed messages on your website rather than calls, that is a website chatbot on your own content.
If the transcripts then need answering from your documents, see RAG on your own server.
Not sure which? Our AI automation page sorts the four jobs and shows what we run.
Work we’ve delivered, and what we’ve written
Bring a few recordings to a 20-minute call
Tell us what you record and where it goes today. We will say whether self-hosted speech to text fits, what the server would cost and what a phone agent on top would take.