Integrating OpenAI, Claude and Gemini into a real app
How to add AI to an app properly: choosing OpenAI, Claude or Gemini per task and keeping cost and errors in check, with lessons from ProAI Training.
By Rokibul Hasan7 min read
- AI
- AI integration
Adding AI to an app has never been easier to start and never been easier to get wrong. Connecting to a model from OpenAI, Anthropic or Google takes an afternoon. Turning that connection into a feature people trust — one that is fast enough, cheap enough to run, and honest about its mistakes — takes real design.
I have built AI into several products. SplitMate uses it to read receipts from a photo. Clearday uses it to write a morning brief and a nightly recap, and to suggest changes the user approves. ProAI Training is built around it: a voice-first app where you rehearse difficult work conversations against an AI character and get scored, specific feedback. This guide shares what those builds taught me about integrating OpenAI, Claude and Gemini into a real app.
Start from the job, not the model
The worst way to add AI is to start with "we should use AI" and look for somewhere to put it. The best way is to start with a job that is slow, tedious or impossible for users today, and ask whether a model can do it well enough.
Each of these is a narrow job:
- Reading a receipt. In SplitMate, you photograph a receipt and the app reads the amount, the merchant and the category, so logging an expense on the go takes seconds.
- Writing a summary from someone's own data. Clearday writes a brief in the morning — what is on, where the clear hours are — and a recap at night of what moved rather than finished.
- Playing a role and judging a performance. ProAI Training holds a character in a live spoken conversation, then marks the transcript and writes back something specific.
A narrow job is easier to test, cheaper to run and far easier to trust than a general assistant that is supposed to do everything.
Not a chat box bolted on
The second principle follows from the first: the model should do one job inside the product, and the app should be built around it. A chat box in the corner of an existing app rarely helps anyone. An AI step that sits exactly where the user needs it — the moment they photograph a receipt, the moment they open the app in the morning — does.
That usually means the AI's output is structured data the app can use, not just text on a screen. A receipt becomes an amount, a merchant and a category in editable fields. A session becomes a score, four sub-scores and two pieces of feedback. Structure is what lets the rest of the app treat AI output like any other data.
Choosing between OpenAI, Claude and Gemini
There is no single best model, and the ranking changes every few months. What matters is matching each job to the model that does it best, and building so you can change your mind.
In ProAI Training there is no single model behind the product. OpenAI, Gemini and Claude each handle different parts of the job, chosen on what that part actually needs rather than on loyalty to one provider. Holding a character in a live conversation, marking a transcript against four criteria and writing back something specific, and generating lesson content are not the same problem, and they do not reward the same model.
Keeping those boundaries explicit has a second benefit: any one model can be swapped when a better option ships, without rewriting the app. In practice that means one place in your code that talks to each provider, with the rest of the app asking for "a transcript score" or "a receipt reading" rather than calling a specific model directly.
When you compare models for a job, test them on your own inputs. Public benchmarks measure someone else's problem.
Write instructions like a brief, not a wish
The instructions you give a model — the prompt — are part of your product, and they deserve the same care as code. The biggest lesson from ProAI Training is about this.
Each character in ProAI Training needs a position, a reason to hold it and the stubbornness to keep holding it after one good answer. Adjectives such as "be skeptical" or "be tough" barely worked. What worked was a brief: a role, an employer and one concrete grievance. A CFO who is skeptical because she lost two contracts last year to late shipments gives the learner something real to find and use. Giving the AI that kind of constraint produced sharper conversations than any amount of instruction about tone.
Practical habits that follow:
- Keep prompts in version control next to the code, and change them deliberately.
- Give concrete context, not adjectives.
- Ask for structured output — a fixed format the app can check — whenever the result feeds into the interface.
- Validate everything that comes back. If the model returns something malformed, retry or fall back; never pass it straight to the user.
Build an evaluation set before you ship
You cannot judge an AI feature by trying it a few times. Before launch, collect a set of real inputs — receipts in different formats, transcripts of good and bad conversations, messy user data — and decide what a good result looks like for each.
Run the set whenever you change a prompt or a model. It turns "it seems better" into "it got these twelve right that it used to get wrong", and it catches regressions before your users do.
Make it fast enough
Users forgive a lot, but not waiting. A few techniques help:
- Stream responses so text appears as it is generated instead of all at once.
- Use a smaller, faster model for steps that do not need the strongest one.
- Do work ahead of time where you can. A morning brief can be prepared before the user opens the app.
Voice raises the bar further. In ProAI Training, voice is the default input because the skill being trained is spoken: typing lets you compose and edit, speaking does not, and that is the point. The character's replies are spoken back through ElevenLabs, because a flat synthetic voice undercuts the whole exercise — you cannot rehearse reading a difficult room if the room has no tone. A keyboard toggle stays available for a quiet office.
Keep the cost under control
AI is billed per use, so cost scales with your success. Plan for it from the start:
- Right-size the model for each job.
- Send only what the job needs. Long context costs money on every request.
- Cache results that do not change.
- Set limits per user and per day, so a bug or an abusive user cannot run up a bill.
- Watch the numbers after launch, per feature, not just in total.
I covered running costs more broadly in what it costs to build an app.
Plan for being wrong
Every model will sometimes be wrong. Good AI features are designed around that fact:
- Let users check and correct the result. A receipt reading should land in fields the user can edit, not straight into the totals.
- Never block the core flow on AI. If the model is slow or down, the user should still be able to do the job by hand.
- Show the evidence. In Clearday, AI suggestions are placed only where there is something to say, and every suggestion shows its evidence: "You skipped reading four nights, always after 10 PM" is a claim the user can check against their own week. That is the difference between advice and a horoscope. Clearday also waits for four weeks of data before the screen offers anything at all.
- Keep the last move with the human for anything that matters. Clearday's assistant proposes and waits for a yes before it touches the calendar. I wrote more about that rule in AI agents for business.
Privacy and safety
Treat everything you send to a model as data leaving your system:
- Send the minimum needed for the job.
- Never put secrets — API keys, passwords, other users' data — into prompts.
- Read the provider's terms for business use, including how data is retained.
- Tell users when AI is involved and what it sees.
- Keep server-side control: the app talks to your server, and only your server talks to the AI provider, so keys never ship inside the app.
ProAI Training runs on iOS, Android and the web from one account, so scoring and session history lived on the server from the start — a session started on a commute can be finished at a desk. A useful side effect: the keys and the scoring logic never ship inside the app.
FAQ
Should I use one AI provider or several?
Start with whichever model does your first job best. Design the integration so adding a second provider later is easy. Once you have several distinct jobs, using the best model for each usually pays off.
Can AI run on the phone instead of the server?
Some smaller models can run on-device, which helps privacy and offline use. For most business features today, server-side calls to a provider remain the practical choice, with on-device models as an option for specific jobs.
How long does it take to add an AI feature to an existing app?
The connection itself is quick. The time goes into the parts that make the feature trustworthy: structured output, an evaluation set, fallbacks, cost limits and a review step where it matters.
If you want AI in your product — reading documents, summarising data, an agent for your customers — send me a short brief describing the job you want done. You can read the ProAI Training case study for a full AI build, or see the AI API integration service.