AIGuide
daLæs på danskHow to Integrate an LLM API Into Your Existing App (OpenAI or Claude)
How to integrate an LLM API into your existing app safely: queues, caching, logging, fallbacks and EU data rules, explained for non-developers.

Freelance full-stack developer
- Published
- Reading time
- 13 min
In this post7
To integrate an LLM API into your existing app without putting the rest of it at risk, keep every call on your backend and route it through one small layer that handles queuing, caching, logging, limits and a fallback for when the model fails. The API call itself is an afternoon's work for a developer. Everything around it decides whether the feature holds up once real users and real data hit it.
The setup is the same for OpenAI's models (the ones behind ChatGPT) and Anthropic's Claude. I've written it for the person signing off on the project, with an EU angle: if you have European users, where the data goes matters as much as the model.
The short answer: six building blocks
A production-ready LLM integration has six parts. A small internal tool can skip some of them at first, but you should know which ones you're skipping and why.
| What it does | Without it | |
|---|---|---|
| Server-side AI layer | Keeps keys, prompts and model choice in one place | Keys leak, and switching providers means edits all over the codebase |
| Queue | Runs slow calls in the background with retries | Users wait, and provider errors break the page |
| Caching | Reuses answers for identical input | You pay for the same answer again and again |
| Logging | Records model, tokens, cost and outcome for every call | No way to control spend or explain a wrong answer |
| Fallback | Decides what happens when the model fails | An outage at your provider becomes an outage in your app |
| Validation and limits | Checks output and restricts what the model can do | Wrong or manipulated answers go straight into your data |
My rule of thumb: the closer the feature sits to money, customers or data you can't afford to lose, the more of these six need to be in place before launch. An internal summary tool can start with the AI layer, logging and a spending cap. A customer-facing feature needs all six.
This guide is about adding AI to an app you already run. If the whole app was built with a tool like Lovable or Cursor and now has to go live, start with my guide to taking a vibe-coded app to production.
Why calling the API directly breaks in production
The quickest integration is a direct call from the code that renders the page: the user clicks, your app asks the model and waits. Fine for a demo. In production it runs into five problems.
- Latency. Responses often take several seconds, and the page hangs in the meantime.
- Rate limits. Every provider caps requests and tokens (the units text is billed in) per minute.
- Outages. Every large service goes down now and then, and your feature goes down with it.
- Runaway cost. A bug that calls the model in a loop can burn through a month's budget overnight.
- Unpredictable output. The same prompt can produce different answers, occasionally in a format your code doesn't expect.
None of these are reasons to hold off. They're reasons to put a thin layer around the call.
How to add an LLM to your existing app, step by step
This is the order I work in. The first two steps are decisions rather than code, but they shape everything after.
1. Pick one narrow, testable task
Start with a single task where you can tell a good answer from a bad one: summarizing a support ticket, drafting a reply for an agent, routing incoming requests or pulling fields out of an invoice. "Make the product smarter" is not a task.
Then collect a few dozen real examples from your system, each paired with an answer you'd accept. That set becomes your test suite (an eval set): whenever you change the prompt or switch models, you rerun it and check that quality holds. Decide up front whether a person approves the output before it's used. If you're short on ideas, I've collected 12 practical AI features for SaaS products and web apps.
2. Decide what data may leave your system
Everything you send to the model is processed by the provider. If it includes personal data, the provider becomes your data processor under GDPR, which means a data processing agreement (DPA) and clarity on where data is processed and for how long. That applies outside the EU too if you serve people in the EU.
The defaults are better than many assume. OpenAI states that API data is not used to train its models unless you opt in, that abuse-monitoring logs are kept for up to 30 days, and that new projects can get European data residency once approved. Anthropic's commercial terms likewise rule out training on what you send through its API. Claude also runs on Amazon Bedrock and Google Cloud, which can be the simpler route if your company already has a DPA with one of them.
My rule is to send the minimum that does the job. Strip national ID numbers, bank details and names if the model doesn't need them. This isn't legal advice, and my guide to GDPR and LLM APIs goes deeper.
3. Put an AI layer on your backend
Route every model call through one place in your code: a small module that holds the API key, the prompts, the model choice and the timeouts. The rest of your app just asks it for "a summary of ticket 1234" without knowing which provider answers.
That buys you three things. The key lives only on the server, never in the browser, where anyone with developer tools can lift it and run up your bill. Switching from OpenAI to Claude, or using both, becomes a change in one place. And prompts get treated like code: version-controlled, numbered and tested against your examples from step 1 before a change ships.
4. Move slow calls onto a queue
A queue lets your app reply "got it" straight away and do the heavy lifting in the background. The user sees "processing", and the result shows up on the page or by email when it's ready. Laravel has queues built in, and most other stacks have an equivalent.
The queue is also where you handle failures. When you hit a rate limit, the job waits and tries again. OpenAI's guidance is to respect the Retry-After header and use exponential backoff: wait as long as the provider asks, and otherwise leave longer, slightly randomized gaps between attempts. Also cap how many AI jobs run at once.
Chat is the exception, because the user is actively waiting. There you stream the response so it appears word by word, but the timeout and error handling still apply.
5. Cache what you can
Many AI features get called with the same input over and over. A product description, a ticket summary or a category label only needs generating once and storing until the source changes, which is faster for users and cheaper for you.
Store each answer with its input, prompt version and model, so a prompt change automatically produces fresh answers. And never share cached answers across users when they're based on personal data, or one customer may end up seeing another's information.
OpenAI and Anthropic also cache prompts on their side, which makes long, repeated instructions cheaper. It works when the fixed part of the prompt comes first and the variable part last.
6. Log every call
Without logs, your AI feature is a black box. For each call, record the timestamp, feature, user, model, prompt version, input and output tokens, calculated cost, latency and outcome (success, error or fallback). If users can rate answers with a thumbs up or down, store that too.
That data tells you what the feature costs per month and why ticket 1234 got a strange answer last Tuesday. Set alerts on error rate and daily spend. If you don't have monitoring yet, start with my rundown of uptime and monitoring tools for web apps.
Be careful with the text itself. Log full prompts and you're logging whatever personal data they contain. Keep the metadata, and keep full text only for a short, defined period.
7. Plan for when the model fails
Fallback is your plan for the days the model is down, slow or wrong. There are three levels:
- Retries in the queue, which cover short blips.
- A backup model: if OpenAI fails, the call goes to Claude, or the other way round. This depends on the AI layer from step 3 and on prompts tested against both models.
- A manual path, where the task lands with a person, or the AI button disappears for a while and the rest of the app carries on.
The ground rule is that your app must keep working without AI. If the provider fails repeatedly, stop calling it for a while instead of hammering away. Developers call this a circuit breaker, and it stops an outage on the provider's side from clogging your queues.
8. Validate output and limit what the model can do
Treat everything the model returns as input from a stranger. Ask for a fixed format, such as JSON with defined fields, and have your code check structure and values before anything is saved. Both providers can force responses into a schema, but keep the check in your own code anyway.
The biggest security risk is prompt injection: text in an email, a document or a user message that steers the model into doing something it shouldn't. OWASP ranks it first in its Top 10 for LLM applications and recommends, among other things, least-privilege access and human approval for privileged actions. In practice, the model may propose a refund, but a person or a hard rule in your code approves it.
Finally, set limits on response length, calls per user and monthly spend. If customers talk to the AI directly, tell them so clearly. The EU AI Act includes transparency rules for this kind of system as well.
What it costs to build and run
Costs fall into two buckets: building the integration and paying the provider for usage.
The build depends more on your existing app than on the AI. With queues, logging and a sensible structure already in place, a first narrow feature is a manageable project. Without them, the foundation comes first, and that's often where most of the hours go. I'd settle it in a small, paid discovery phase before asking anyone for a fixed price.
Usage is billed per token, and a single call usually costs very little. Volume is what makes it expensive: many users, long documents or a large model doing a simple job. You have four levers:
- Use the smallest model that passes your test examples.
- Cache answers and repeated instructions, as in step 5.
- Run anything that isn't urgent as a batch job. OpenAI and Anthropic both discount batch processing by 50%, and Anthropic says most batches finish within an hour.
- Set a spending cap. Anthropic lets you set your own monthly spend limit below your tier's cap, and the budget alert from step 6 covers the rest.
For token prices and a simple way to budget, see my guide to estimating LLM API costs.
When you shouldn't build a custom integration
Here's when I'd point you elsewhere.
If your team mainly wants to write, summarize and translate faster, buy business seats for ChatGPT or Claude. It's usually cheaper than development.
If you run an off-the-shelf CRM, e-commerce platform or accounting tool, check for built-in AI features or marketplace apps before building something the vendor may already be working on.
For a small internal workflow with a handful of calls a day, an automation tool like Zapier or Make with an AI step may be enough, at least to test the idea. And if all you need is a chatbot on your website, there are ready-made products. I compare the options in build vs buy for an AI chatbot.
If nobody dares touch your current codebase because every change breaks something, AI isn't the first problem to solve. Stabilize the app first.
A custom integration is the right call when the AI works with your own data inside your own workflows, and you want control over what gets sent, what it costs and what happens when it fails.
Next steps: from idea to your first AI feature in production
Start small: one task, one provider and a layer that makes switching easy later. Run through the checklist before launch, and again before you open the feature to more users.
Ready to ship your LLM feature
- One task is chosen, with examples of good and bad answers.
- Data flows are settled: what may be sent, a DPA in place and where data is processed.
- The API key lives only on the server, in environment variables.
- Queue and timeouts are set up, with retries for temporary errors.
- Cached answers are reused where input repeats.
- Every call is logged with model, tokens, cost and outcome.
- Fallback is decided: backup model, manual path or hidden feature.
- Output is validated in code, and the model can't act without approval.
- A spending cap and budget alert are in place.
If you have an existing app that needs AI built in, here's how I work on custom web apps and platforms. You deal directly with me as the developer, and you own the code from day one.
Frequently asked questions
Can I use a ChatGPT subscription instead of the API?
No, not for an integration. A ChatGPT subscription only covers the ChatGPT app itself. When your software needs to call the model, you need OpenAI's API, which is billed separately based on usage. Claude works the same way: a Claude subscription and API access are separate products with separate billing.
Should I choose OpenAI or Claude?
Choose based on a test with your own examples, not on hype. For most business tasks both will do the job, and the differences show up on your data. Also consider where your infrastructure already lives: Claude is available through AWS, Google Cloud and Microsoft Foundry, and OpenAI's models through Microsoft Azure. Build the integration so you can switch, and the choice isn't permanent.
What happens when a provider retires a model?
You move to a newer model before a set deadline, which OpenAI and Anthropic both announce in advance. Pin your integration to a specific model version rather than always using the latest, and run your test examples against the new model before switching. Your logs will then show whether latency, cost or quality changed.
Can the model read data from my app?
Yes, through what providers call tool use or function calling. The model asks for data, say a customer's recent orders, and your code decides whether it may have it and then fetches it. Permissions are enforced by your app, not the model. Start with read-only access, and only allow changes to data once logging and approval are in place.