Types of developersComparison
daLæs på danskAI Developer vs Data Engineer vs ML Engineer: Who Should Build Your AI Features?
AI developer vs ML engineer vs data engineer: what each one builds, when to hire which, and why most AI features are really API integration work.

Freelance full-stack developer
- Published
- Reading time
- 11 min
In this post8
For most small and mid-sized companies, the AI developer vs ML engineer decision comes down to one question: are you building on an existing model, or training your own? An AI developer connects models from OpenAI, Anthropic or Google to your product through their APIs, an ML engineer trains and runs custom models, and a data engineer builds the pipelines that get your data into shape. A support chatbot, ticket triage or invoice extraction is integration work, and a full-stack developer with LLM experience can build it.
I'm a full-stack developer who builds this kind of integration, so factor that in. I've tried to be just as clear about when you should hire one of the other two instead.
The short answer: who builds what
| AI developer | Data engineer | ML engineer | |
|---|---|---|---|
| What they build | Features on top of existing models via API | Pipelines that collect, clean and move data | Custom models trained on your data |
| Typical work | Support chatbot, ticket triage, document extraction, translation | Data from your shop, CRM and ERP combined in a warehouse for reporting | Demand forecasting, pricing, fraud detection, visual defect detection |
| What you need to bring | Real examples of the task and access to the relevant data | A map of where data lives and what it should answer | Large volumes of historical, labeled data and a clear success metric |
| Time to first result | Fastest: a prototype can be tested early | Once data flows and can be trusted | Slowest: after data work, training and evaluation |
| Rate level | Medium to high | High | High |
| Running costs | Per-call usage fees for the model | Warehouse and tooling subscriptions | Compute for training, serving and retraining |
| Biggest risk | A demo that works but breaks on real data | An expensive warehouse nobody uses | Months spent on something an off-the-shelf model could have done |
| Best for | Most AI features in small companies and SaaS | Companies with data spread across many systems | Products where model accuracy is the product |
Here's the rule I use. If the task sounds like "read this and do something with it" (emails, documents, support tickets, images), start with an AI developer. If the real problem is that your data is scattered and unreliable, you need a data engineer. If a model has to learn patterns in your own numbers that no public model knows, that's ML engineering.
For where these three sit next to front-end, back-end and mobile developers, see my overview of the types of developers and when to hire each.
What an AI developer actually does
An AI developer doesn't train models. They take a model that OpenAI, Anthropic or Google has already trained and connect it to your systems through an API (the interface software uses to talk to other software). It's a lot like adding payments to a web app: you don't build card processing yourself, you plug into Stripe and build everything around it.
Typical AI features for a small company or a SaaS product:
- A chatbot or search that answers from your own docs, terms or help center.
- Triage of incoming support tickets by topic, urgency and language.
- Extracting fields from invoices, purchase orders or contracts so nobody retypes them.
- Summaries of long customer threads or call notes.
- Draft replies, product copy or translations that a person reviews. For a business selling across Europe, translation alone can be a strong first feature.
The docs chatbot usually goes by RAG (retrieval-augmented generation). The idea is simpler than the name: the system finds the relevant passages in your documents and sends them to the model along with the question. No model gets trained.
The job title is new and loosely used, and "AI engineer" and "AI developer" often describe the same thing: a full-stack or back-end developer who has put LLM features into real products. That background matters more than the title, because the hard parts of the job are rarely the AI. I cover when one person can handle front end, back end and integrations in when one full-stack developer is enough.
Where the hours go in a typical AI feature
The call to the model is often a handful of lines. Everything around it is where the time goes. Take a SaaS product with customers across Europe that wants incoming support tickets sorted and routed automatically. The work breaks down roughly like this:
- Data access. Pull tickets from the helpdesk, and write the result back to the right queue.
- Instructions and output format. Write the prompt and define exactly what shape the answer must have so the rest of the system can use it.
- Testing on real examples. Build a fixed set of real tickets with known correct answers and measure how often the model gets it right. These test sets are called evals.
- Failure handling. What happens when the provider is down, the answer is wrong, or a ticket is in a language nobody expected?
- Permissions and GDPR. Who can see what, and which personal data may leave your system? If you serve customers in the EU, GDPR generally applies wherever your company is based.
- Review interface. A screen where an agent can see and correct the suggestion before it's acted on.
- Monitoring. Track usage, cost per call and whether quality drops when the provider updates the model.
Only steps 2 and 3 are AI work in the narrow sense. The rest is ordinary software development: integrations, databases, queues, access control and interfaces. That's why I think a full-stack developer with LLM experience is a better fit than an ML specialist for most small companies. The specialist can train models, but that's rarely what's missing.
OpenAI recommends the same order in its model optimization guide: build evals, iterate on prompts, and only then consider fine-tuning. For the build itself, with queues, logging and a fallback plan, see my guide to integrating an LLM API into an existing app.
When you need a data engineer
A data engineer builds the pipelines that move data from your business systems into one place where it's clean, consistent and trustworthy. The destination is usually a data warehouse (a database designed for analysis and reporting) such as BigQuery or Snowflake.
You need one when scattered data is the actual blocker. Sales live in your shop, customers in the CRM, stock in the ERP, none of the numbers agree, and someone rebuilds the same spreadsheet report every month. A data engineer is also the prerequisite for any ML work, because a model is only as good as the data it learns from.
In a smaller company the job is often smaller than the title suggests. Syncing two systems or feeding a dashboard is something an experienced web developer can build. A dedicated data engineer starts paying off when you have many sources, large volumes, history to preserve and several teams relying on the same numbers.
One thing a docs chatbot doesn't need is a data engineer. It needs documentation that's current and covers what customers actually ask. That's an editorial job, not an engineering one.
When you need an ML engineer
An ML (machine learning) engineer trains, evaluates and runs models built on your own data. That's the right call when the problem is about patterns in numbers and events that no public language model knows: demand forecasting, dynamic pricing, credit or churn risk, fraud detection, or spotting specific defects in images from a production line.
The work doesn't end when the model is trained. It has to be deployed, monitored and retrained as the world changes, for example when customer behavior shifts. That discipline is called MLOps, and it's often bigger than the training itself.
You'll need three things before an ML engineer can help: a lot of historical data, data labeled with the right answer (which orders really were fraud), and a precise definition of "good enough". Without all three, there's nothing for them to work with.
Plenty of founders assume they need an ML engineer to "train the AI on our data". For text tasks that's rarely true, because the relevant data is usually sent along with each request instead. Fine-tuning language models has also become a narrower path. As of October 2026, OpenAI's fine-tuning documentation said its fine-tuning platform was being wound down and no longer accepted new users.
When each option is the wrong hire
When an AI developer isn't enough
- When the task is prediction from your own numbers, like stock levels, pricing or churn. Language models are strong with text, not with patterns across thousands of transactions.
- When the AI is the product and accuracy decides whether customers pay. Someone has to measure and improve the model systematically.
- When data is too sensitive to leave your own infrastructure. Self-hosting an open model takes more specialized operations work than calling an API.
When a data engineer is premature
- When you have two or three systems and a handful of reports. An integration or an off-the-shelf reporting tool is cheaper.
- When nobody has defined the questions the data should answer. A warehouse without questions is an expensive archive.
When an ML engineer is the wrong call
- When you have a few hundred examples, not tens of thousands. There isn't enough to learn from.
- When an existing model with good instructions handles most of the task. Start there and decide later whether the rest justifies a specialist.
- When there's no plan for running the model. One that isn't monitored and retrained gets quietly worse over time.
If you're unsure what kind of developer the project needs in the first place, my decision guide by project type is a better starting point. And if AI is becoming central to your product strategy, a fractional CTO can set the direction before you start hiring specialists.
Next steps: test the idea on existing models first
The cheapest way to find out who you need is a small test on existing models with your own data. If it turns out not to be an integration job, you'll know exactly what a specialist needs to solve.
- Collect real examples. Pull 20-50 emails, documents or tickets and write down the correct result for each.
- Map where the data lives. Which systems, which formats, and who has access?
- Build a small prototype. Run an existing model against your examples and measure how often it gets it right.
- Decide based on the result. If it's accurate enough, build the production integration. If not, you'll see whether the gap is data or model, and whether the next hire is a data engineer or an ML engineer.
I build these integrations with Laravel and React, including inside systems that are already running. Larger projects start with a paid, fixed-price discovery phase so scope and price are agreed before anything gets built, and you own the code from day one. You can see how I work on my page about custom web apps and platforms.
Frequently asked questions
What's the difference between a data scientist and an ML engineer?
A data scientist analyzes data and builds models to answer questions, while an ML engineer makes those models run reliably inside a product. The data scientist asks what the data can tell you, the ML engineer asks how the model holds up in production. In small teams the roles often overlap. If you need analysis and reporting rather than a model in your product, a data scientist or analyst is the better hire.
Can I switch AI providers later?
Yes, if the integration is built for it. Put every model call behind a single layer in your own code, so the rest of the system doesn't care whether the answer comes from OpenAI, Anthropic or Google. Then you can switch provider or model as prices and quality change. Keep your eval set too, so you can measure whether the new model is as good as the old one.
Is it GDPR-compliant to send customer data to OpenAI or Anthropic?
It can be, but you need to decide before you send anything. The major providers typically offer data processing agreements and business terms where your data isn't used for training, and some offer processing in the EU. Send only the data the task needs and strip personal data where you can. This isn't legal advice, so talk to a privacy professional if you handle sensitive data.
Should I hire an AI developer full-time or use a contractor?
For your first AI features, a contractor is usually the better fit. A well-defined feature can be built and tested by an external developer, and after launch the work is mostly tuning and maintenance. A full-time hire makes sense once AI is core to your product and there's daily work on it for years. At that point, in-house knowledge of your data and customers is worth more than flexibility.