AI app development cost: what you pay to build it, and every month after
- October 5, 2026
- 7 min read
Manoj Sethi
Founder, Boffin Coders

You asked how much it costs to develop an AI app and got a range so wide it told you nothing. That happens because AI app development cost is two numbers, not one. The first is the build: people multiplied by months, paid once. The second is the running cost: every question your users ask is a paid call to a model, and the server, the search index and the person who checks the answers are paid every month as well. The running cost depends on your volume and on three choices you make early. Here is how to put a number on each.
The build: people multiplied by months
Most of an AI app is ordinary software: screens, accounts, a database, the part that finds the right piece of your content for each question, the rules the model must follow, a set of real questions to test it against, and a screen where a person checks what it wrote. The model itself is often one API call in the middle of all that.
So the build is priced like any other build: the people it needs, multiplied by the months they spend. Our published rates for AI and automation work are $25-40 an hour, by seniority, and a dedicated developer starts at $4,000 a month. Mobile work is $20-38 an hour, and a dedicated mobile developer starts at $3,500 a month. Separately, a custom project of any size starts at $4,000 in total, a floor that filters rather than an average.
The months depend on the scope. An assistant added to an app you already have is mostly the search, the rules and the test questions. A new app from nothing adds the months for everything around the AI, scoped as a mobile app build with its own lines. Either way you get the exact figure once the scope is written down, before anyone starts.
Two costs, not one
One is paid once. The other arrives every month.
The build, paid once
- Search over your content
- AI engineer
- Rules and test questions
- AI engineer
- Screens and the review view
- app developer
- Accounts, data, admin
- app developer
- Priced as
- people × months
Running, every month
- Model usage
- per request
- Hosting
- your server
- Search index
- same server, or a service
- Monitoring and logs
- storage and time
- Human review
- staff hours
- For as long as it runs
- set by your volume
The running cost: the bill that never stops
A build quote does not cover this half, because it is not paid to the builder. Five lines keep arriving after launch:
- Model usage. A hosted model charges for every request, by the amount of text going in and coming out.
- Hosting. The server your app runs on. If the model runs on your own server, that server has to be large enough to hold it, and you pay for it whether anyone asks a question or not.
- The search index. A keyword index sits on the same server at no extra cost. A vector database is usually a separate service with its own monthly charge.
- Monitoring. Logs of what was asked, what was sent to the model and what came back, kept so you can find out why an answer was wrong.
- Human review. A person reading the output before it reaches a customer. It is staff time, every month, and it belongs in the budget from the first day.
How to estimate the model bill yourself
You can estimate the line that moves most with volume yourself. It is three numbers multiplied together:
requests a month × tokens a request × the provider's published price per token
A token is a short piece of a word. The tokens in a request are the question, the rules, the pieces of your content sent with it, and the answer. Providers price text sent and text returned separately, so work out each side on its own.
For scale: the assistant on this website sends about 4,500 tokens a question, and about 2,000 of those are ten pinned sections that hold our prices and rules and go with every request. If your app expected 20,000 questions a month at that size, it would send about 90 million tokens a month, plus the answers that come back. Multiply that by the price on the provider's own pricing page on the day you check. Those prices change, which is why no figure for them is printed here.
Estimate it yourself
Three numbers, multiplied
Requests a month × tokens a request × the provider's published price per token.
- 01
Your logs or a pilot
Requests a month
Count questions, not users. Estimate launch and a year out, because the bill follows the busier number.
- 02
Question, rules, content, answer
Tokens a request
The assistant on this site sends about 4,500 a question, about 2,000 of them ten pinned sections of prices and rules.
- 03
The provider's pricing page
Price per token
Text sent and text returned are priced separately. Read the page on the day you estimate.
- 04
Multiply the three
The monthly model bill
The middle number is the one you control. Fewer, better sections cut the bill on every request.
The number you control most is the middle one. Sending three well-chosen sections instead of twenty cuts the bill on every request, for as long as the app runs.
What moves AI app development cost the most
Three decisions change both the build and the running cost more than any hourly rate does.
1. Where the model runs. A hosted model through an API charges per request and sends your text to the provider. A model on your own server has no per-request charge and keeps your data in place, but needs a machine large enough to run it, paid every month. A model on the device needs neither, but only small models fit on a phone. There is a working example of each. The assistant on this site sends its selected sections to a hosted model through an API, because our pages are public anyway. IELTS Builder, a live exam platform we built for a coaching institute, runs a Whisper speech-to-text model on the client's own server, so a student's recording never leaves it. PDF Toolkit, one of our own apps, reads documents with OCR on the phone itself and works fully offline, with no server and no per-request cost.
Choice one, where the model runs
Pay per request, or pay for the machine
| Where the model runs | Cost per request | Your data goes to | What you pay for instead | Working example |
|---|---|---|---|---|
| A hosted model, through an API | Yes, by the text in and out | The model provider | Nothing extra. The provider runs the machine | The assistant on this site, because its pages are public |
| A model on your own server | None | Nowhere. It stays on your server | A server large enough to run the model, every month | IELTS Builder, speech-to-text on the client's server |
| A model on the device | None | Nowhere. It stays on the phone | Nothing monthly, but only small models fit | PDF Toolkit, OCR on the phone, fully offline |
A hosted model, through an API
- Cost per request
- Yes, by the text in and out
- Your data goes to
- The model provider
- What you pay for instead
- Nothing extra. The provider runs the machine
- Working example
- The assistant on this site, because its pages are public
A model on your own server
- Cost per request
- None
- Your data goes to
- Nowhere. It stays on your server
- What you pay for instead
- A server large enough to run the model, every month
- Working example
- IELTS Builder, speech-to-text on the client's server
A model on the device
- Cost per request
- None
- Your data goes to
- Nowhere. It stays on the phone
- What you pay for instead
- Nothing monthly, but only small models fit
- Working example
- PDF Toolkit, OCR on the phone, fully offline
2. Keyword search before a vector database. Most tutorials start with embeddings and a vector database, which adds a paid call for every page and every question and a service to pay for every month. The assistant on this site answers from our own content, cut into 176 sections, with BM25 keyword search and no vector database. Start with keyword search, log a hundred real questions, and add vectors only where keyword search misses. Why a first RAG chatbot does not need a vector database has the numbers.
3. A person approves the output where a mistake is costly. Review time is a running cost, and it is still cheaper than a wrong answer sent in your name. On IELTS Builder the platform took 4,584 module submissions in a thirty-day window, and teachers approve every AI-drafted grade before a student sees it. Reading and Listening are marked against an answer key. Only Writing and Speaking need a teacher, which is about thirty-nine a day, against about 153 a day if every submission were marked by hand. That is the shape to aim for: the model drafts, a person decides, and the budget shows the person's time.
One honest limit, since it bears on your estimate: we have no client RAG deployment to show you yet, we have not fine-tuned a model for a client, and we do not build image or video generation. For an app that answers from your own documents, fine-tuning is rarely the right first step anyway, because it cannot show where an answer came from.
What to have ready before you ask for a number
Five facts turn a range into a number:
- What the app does, in one sentence, and who uses it.
- How many requests a month you expect at launch, and in a year.
- Where your data is allowed to go: a provider, your own server, or nowhere off the device.
- What happens when an answer is wrong, and who would notice.
- Which platforms it needs: web, Android, iOS, or an existing product.
If you already have the product and need the AI part built into it, AI engineers by the month start at $4,000 a month, and you interview them before they start. If your documents must stay on your own server, that is the private LLM and RAG build.
An AI app costs what its people cost to build it, and then what its questions cost to answer, for as long as anyone asks them. The first number is fixed once you know the scope, and the second is the one worth designing for: where the model runs, how much text you send it, and who reads the answer before your customer does. Get those three right at the start and the monthly bill becomes something you chose, rather than something you discover.
Related: Private LLM and RAG development · Hire AI engineers · You do not need a vector database for your first RAG chatbot
Keep reading
Manoj Sethi
Founder, Boffin Coders
Manoj founded Boffin Coders in 2017 and leads its work with agencies: white-label builds, delivery and partnerships. Manoj writes about how agencies can add development capacity without losing control of their clients.
Need this built?
We build websites and apps for businesses, and work behind digital agencies as their development team. Tell us where things stand on a 20-minute call.