Skip to content

AIPricing guide

daLæs på dansk

LLM API Cost Estimation: What AI Features Really Cost per User (2026)

LLM API cost estimation made practical: worked per-user examples with OpenAI, Claude and Gemini prices in EUR, what inflates the bill and how to budget.

By

Freelance full-stack developer

Published
Reading time
12 min
In this post9

LLM API cost estimation boils down to one formula: tokens per call × calls per user × price per token. Run it for typical AI features and you land somewhere between a few cents and well over €100 per user per month. Which end you hit depends far more on how the feature is designed than on which model you pick.

The numbers below use each provider's list prices, checked in October 2026, converted from USD to EUR at about €0.89 per dollar. Prices move often, so treat the method as the durable part. I'm a freelance developer who builds these features, so keep that bias in mind.

The short answer

Estimated API cost per user per month in EUR (USD in brackets), four worked examples at October 2026 list prices
Small model (GPT-6 Luna)Mid-tier model (Claude Sonnet 5.5)Flagship model (Claude Opus 5.5)
In-app chat, 40 messages€0.02 ($0.02)€0.43 ($0.48)€0.86 ($0.96)
Summarizing 30 documents€0.07 ($0.08)€1.34 ($1.50)€2.67 ($3.00)
AI agent, 100 tasks€1.60 ($1.80)€32 ($36)€64 ($72)
Classifying 10,000 emails or tickets (total)€0.94 ($1.05)€18.70 ($21)€37.40 ($42)

The assumptions behind each row are in the worked examples below. These figures cover the API bill only. Building, testing and maintaining the feature often costs more than running it.

My rule of thumb: if the AI feature is a chat or a tool people trigger a few times a day, the API bill is rarely the problem. If it runs multi-step agents or processes everything that comes in, do the math before you build it.

If you're taking an AI-built prototype to real users, my guide to moving a vibe-coded app to production covers the other things that need to be in place besides a running-cost budget.

How LLM APIs bill you

OpenAI, Anthropic (Claude) and Google (Gemini) all price their APIs per million tokens. A token is a chunk of text, often part of a word. Anthropic's rule of thumb is about 4 characters or 0.75 English words per token, and it varies by language. German, French and the Nordic languages usually produce more tokens per word than English, so add some headroom if your users don't write in English.

Two things drive every bill. Input is everything you send: the user's message, plus your system prompt (standing instructions), earlier turns in the conversation and any documents or data you attach. Output is what the model writes back, and it's the expensive part.

List prices in USD per million tokens, input / output (checked October 2026)
Small modelsMid-tier modelsFlagship models
OpenAIGPT-6 Luna: $0.10 / $0.50GPT-6.1 Sol: $2 / $10GPT-6 Astra: $10 / $50
AnthropicClaude Haiku 4.5: $1 / $5Claude Sonnet 5.5: $2 / $10Claude Opus 5.5: $4 / $20
GoogleGemini 3.5 Flash-Lite: $0.30 / $2.50Gemini 3.8 Flash: $1.50 / $7.50Gemini 3.1 Pro Preview: $2 / $12

Sources: Anthropic's pricing page, OpenAI's API pricing and Google's Gemini API pricing. Gemini 3.8 Flash ran at an introductory price until December 31, 2026 and doubled on January 1, 2027. The table shows the 2027 price. Anthropic's most expensive models, such as Claude Fable 5.1, cost $10 / $50 and are left out here. EUR figures use the ECB euro reference rates from October 2, 2026.

You can't compare per-token prices across providers one to one, because each one splits text into tokens differently. Anthropic notes that its models from Claude 4.7 onward use a tokenizer that produces roughly 30% more tokens for the same text. Two models with identical list prices can produce different bills on the same input. The only reliable comparison is to run your own samples through each and read the usage numbers.

Worked examples: cost per user per month

Four common AI features in a SaaS product or web app. The mid-tier model is Claude Sonnet 5.5 at $2 / $10 per million tokens. GPT-6.1 Sol has the same list price, so the numbers apply to both.

In-app chat or support assistant

Each message sends about 4,000 input tokens: system prompt, relevant snippets from your help docs and the conversation so far. The reply is about 400 tokens. An active user sends 10 messages across 4 conversations a month, so 40 calls.

  • Input: 40 × 4,000 = 160,000 tokens at $2 per million = $0.32.
  • Output: 40 × 400 = 16,000 tokens at $10 per million = $0.16.
  • Total: $0.48, about €0.43 per user per month.

Cache the system prompt and help content (more on that below) and it drops to around €0.26. At 1,000 active users, you're looking at roughly €260-430 a month. Easy to budget.

Document summaries and analysis

Users upload contracts, reports or proposals and get a summary or extracted figures back. A 25-30 page document is roughly 20,000 tokens, with a 1,000-token answer. Thirty documents a month means 600,000 tokens in and 30,000 out.

  • Mid-tier: $1.20 + $0.30 = $1.50, about €1.34 per user per month.
  • Flagship: about €2.67 per user per month.

If results don't need to come back instantly, say an overnight job, use the provider's batch API. Requests are queued and processed when capacity allows, and OpenAI, Anthropic and Google all knock 50% off for waiting.

Multi-step AI agent

An agent decides its own next step: query data, call a tool, read the result, try again. Each step is a new API call, and each call resends everything that happened so far.

Assumptions: 8 calls per task, averaging 15,000 input and 1,500 output tokens per call, including the model's reasoning. A user runs 5 tasks per working day, about 100 a month.

  • Per task: 120,000 tokens in and 12,000 out. On the mid-tier model: $0.24 + $0.12 = $0.36.
  • Per user per month: 100 × $0.36 = $36, about €32.

That's 75 times the chat example on the same model. Swap in GPT-6 Astra at $10 / $50 and you reach about €160 per user per month. At that point the AI bill can exceed what the customer pays you.

Background jobs: classification and extraction

The cheapest workload is often the most useful: a small model reading incoming emails, tickets or form submissions to sort, tag or pull out fields. At about 800 tokens in and 50 out per item, 10,000 items cost around €0.94 on GPT-6 Luna. A mid-tier model charges about €18.70 for the same work, and for simple sorting you rarely get much for the extra money.

Always test the small model first. Move up only when it fails on your own data.

What makes the bill grow

The examples above are averages. In production a handful of patterns push costs up, and most of them are under your control in code.

What inflates LLM API costs in production
Why it costs moreHow to contain it
Long conversationsThe full history is resent on every call, so message 10 costs far more than message 1Cap the history or summarize older turns
Large contextHelp docs, files and data you attach count as input on every callRetrieve first, then send only the relevant snippets
ReasoningModels that think before answering burn extra tokens, billed as outputEnable it only for tasks that need it
Tools and web searchTool definitions add tokens to every call, and search is billed per queryGive the model only the tools the task needs
Abuse and bugsA user pasting a whole book, a bot, or a retry loop calling the model nonstopLimit input length, call count and spend per user
EU data residencyGuaranteed regional processing typically adds about 10%Decide based on your actual data requirements

Conversation history is the one that catches teams out. The model remembers nothing between calls, so your code resends the whole thread each time. A 30-message conversation doesn't cost 3 times a 10-message one. It costs many times more.

Tools and search carry their own costs. With Anthropic, simply enabling tool use adds a few hundred tokens to every request, and web search costs $10 per 1,000 searches. Google charges $14 per 1,000 grounded searches after 5,000 free ones a month.

For European companies, data residency is a real line item. OpenAI charges a 10% uplift for regional processing endpoints on models released from March 2026, and Anthropic applies the same premium on regional endpoints through AWS and Google Cloud. I cover what you can send, and when regional processing is worth paying for, in GDPR and LLM APIs.

How to cut your LLM API costs

  1. Match the model to the task. Use a small model for classification, extraction and short answers, and route only the hard cases to a larger one. Model routing alone can cut a bill to a fraction.
  2. Turn on prompt caching. If every call carries the same instructions and reference text, the provider can cache it. Cached input then costs 90-95% less: 10% of the base price on Claude Sonnet 5.5, 5% on GPT-6.1 Sol. At Anthropic, the first cache write costs 25% more than normal input.
  3. Batch anything that can wait. OpenAI, Anthropic and Google all give 50% off.
  4. Ask for short answers and set a maximum output length. Output is the expensive part.
  5. Reuse repeated answers. If many users ask the same question, store the answer in your own database and skip the call entirely.
  6. Set spending limits, both per user in your own code and as an overall cap or alert with the provider.

For how to structure the integration itself so it stays observable and controllable, see my guide to integrating an LLM API into an existing app.

Budgeting AI into your SaaS pricing

When AI is part of a SaaS product, the number that matters isn't the total bill. It's cost per user against revenue per user. My rule of thumb is to keep AI spend under about 10% of the subscription price. On a €29 monthly plan, that's under €3 per user. The chat example fits comfortably. The agent example doesn't.

What I'd put in the budget:

  • A per-action estimate (tokens in and out), measured on real samples rather than guessed.
  • The usage distribution. Averages mislead, because your heaviest users can account for a large share of the spend.
  • A 30-50% buffer for long conversations, new features, price changes and EUR/USD swings.
  • Quotas in your plans, for example a number of AI tasks per month with the option to buy more.
  • Build and maintenance time. Models get retired, and switching means retesting your prompts. Claude Opus 4 and Sonnet 4, for instance, are already retired on Anthropic's own API.

If a feature is expensive per user, let your pricing reflect it. You can put heavy AI features in a higher tier or charge per use. I go through the trade-offs in my post on SaaS pricing models.

When an AI feature doesn't pay off

I build AI features myself, and they're still not always the right call. Four cases where I'd usually advise against it:

  • A subscription covers it. If your team just needs AI for writing and summarizing, ChatGPT or Claude seats are cheaper and faster than a custom build.
  • The need is generic. A chatbot that answers questions from your website can often be bought off the shelf. See my comparison of building vs buying an AI chatbot.
  • The unit economics don't work. If the feature costs more per user than the user pays, and you can't change the model, the number of steps or the price, it isn't ready.
  • Plain code does the job. Rules, search and templates cost next to nothing to run and give the same answer every time. LLMs earn their keep on the unpredictable.

If you're still deciding where AI actually helps, start with 12 practical AI features for SaaS products.

Next steps: build your own estimate

  1. Pick one concrete AI feature and write down what goes into a typical call.
  2. Measure tokens on 10-20 real samples, using the provider's console or its token-counting endpoint.
  3. Multiply by calls per user and number of users, then add your buffer.
  4. Test the cheapest model that can do the job before reaching for a bigger one.
  5. Set limits and monitoring from day one, and track cost per customer for the first few months.

If you want help building the AI feature into your product with costs you can actually control, see how I work on my page about custom web app and platform development. You work directly with me as the developer, and you own the code from day one.

Frequently asked questions

Is a ChatGPT or Claude subscription the same as API access?

No. A ChatGPT or Claude subscription covers use inside the provider's own apps. When your software talks to a model, it goes through the API, which is billed separately per token on a developer account. You can't run a product feature on a personal subscription, and your API bill scales with how much your users actually use the feature.

Can I test LLM APIs for free?

Yes, within limits. Google offers a free tier for the Gemini API, and Anthropic gives new accounts a small amount of free credit. Note that Google's pricing page says free-tier data may be used to improve its products, which isn't the case on the paid tier. For a European business that matters: keep real customer data out until you're on a paid plan.

Will LLM API prices keep falling?

Often, but not reliably enough to bank on. Anthropic dropped a planned price increase for Claude Sonnet 5 and made its introductory price permanent, while Gemini 3.8 Flash only held its introductory price until the end of 2026 and costs twice as much from 2027. Newer models can also use more tokens for the same text. Keep a buffer and write your code so switching models is cheap.

Is self-hosting an open-source model cheaper?

Rarely for a small business. You avoid per-token fees but pay for GPU servers whether they're busy or idle, plus the work of running, updating and monitoring them. It can pay off at high, steady volume or when data must never leave your own infrastructure. Otherwise the hosted APIs are almost always cheaper and far less work.