Open-source or closed-source AI tools? A decision guide for SMEs
"Which AI tools do you use, open source or closed?" is one of the questions we get most. Our answer is always the same: whichever fits the job and your budget. We publish our own tooling openly, and we recommend closed-source tools where they are clearly better.
That answer is only useful if we explain how we decide. So here is the method. It is the same one you can use yourself.
First, what the words mean
Closed-source here means a model or product you use through someone else's service. You send your data to their API or use their app. You do not see or change how it works. The major model providers, and most AI features built into business software, are closed.
Open-source means the code is published under a licence that lets you run it, change it and keep it. Much of the plumbing around AI is genuinely open source: vector databases, agent frameworks, model gateways, logging and tracing tools.
Open-weight is the middle case that causes most confusion. An open-weight model publishes the trained model so you can run it on your own hardware or cloud. Its licence may attach conditions, and the training data usually is not published. Several well-known model families are open-weight rather than open source in the strict sense. If you plan to build a product on one, read the licence.
The practical point: "open or closed" is not one decision. It is a separate decision for each layer of your stack — the model, the search over your data, the agent code, the gateway, the monitoring.
The five questions
1. How sensitive is the data?
This is the first question because it can end the conversation. If a job involves data you are not allowed to send outside your control — health records, some financial data, anything a contract says stays in a named country — then a closed-source API may be ruled out, or allowed only under specific terms.
Two things to check before assuming it is ruled out. First, what the provider's business terms actually say about retention and training on your data; business plans often differ from consumer ones. Second, whether the provider offers the model inside a cloud region you already use. If neither works, a self-hosted open-weight model is the fallback.
For most everyday business data, a closed-source API under business terms you have read is fine. The mistake is not using one. The mistake is not reading the terms.
2. What does it cost at your volume?
A free licence is not a free system. Someone has to host the model or tool, patch it, secure it and watch it. For a team of five engineers, that time is expensive.
The rough shape of the trade-off:
- Low or unpredictable volume favours closed-source APIs. You pay per use and someone else runs the servers.
- High, steady volume on a simple task can favour a smaller open-weight model you host, because the per-use cost of an API adds up and a small model may be good enough.
- Plumbing that is simple to run — a vector database, a gateway, a tracing tool — often favours open source at any volume, because the running cost is low and you keep control.
Do the sum with your own numbers. Measure what a normal request costs today, multiply by your expected volume, and compare it honestly with the hours needed to run the alternative.
3. Is there a capability gap?
For some jobs, the best closed-source models are clearly ahead. For others — sorting emails, extracting fields from a form, tagging tickets — a smaller open model does the job well.
Do not decide this from benchmarks or opinions, including ours. Decide it with a small test: take thirty to fifty real examples from your own work, with answers you know are right, and run them through each candidate. The one that gets your examples right at a cost you can accept is the answer. That test set is also what lets you change models later with confidence.
4. How hard would it be to leave?
Lock-in is not about whether a tool is open or closed. It is about where your knowledge and logic live, and how much would have to be rewritten to move.
Three habits keep you free to leave, whatever you use:
- Keep your data, prompts and agent logic in your own accounts and repositories, in formats you control.
- Put a thin gateway between your applications and the model providers, so switching a model is a configuration change rather than a rewrite. Open-source gateways such as LiteLLM do this.
- Keep the test set from question 3, so you can check a replacement before you switch.
With those in place, you can use closed-source models heavily and still not be locked in. Without them, an open-source stack can lock you in just as firmly — to the one engineer who understands it.
5. Can your team run it?
The best tool is the one your people can operate at 2am when it breaks. An open-source system with nobody who understands it is a risk, not a saving. A closed-source product your team cannot see inside is also a risk, in a different way.
Be honest about who will own each part after launch. If the answer is "nobody yet", choose the option that needs less running, and make hiring or training part of the plan.
The decision at a glance
| Question | Leans closed-source | Leans open-source |
|---|---|---|
| Data sensitivity | Everyday business data, under business terms you have read | Data that must stay under your control or in a named location |
| Cost at volume | Low or unpredictable volume | High, steady volume on a simple task; plumbing that is cheap to run |
| Capability gap | Hard reasoning, long documents, the best available quality | Narrow, repetitive tasks where a smaller model passes your tests |
| Lock-in | Low risk if you keep a gateway and your own test set | Parts that hold your knowledge and logic, which you want to own |
| Team skills | Little capacity to run infrastructure | Someone who can host, patch and monitor it |
Most rows will not all point the same way. That is normal. The answer is usually a mix, decided layer by layer.
How we usually choose
For a company of 20 to 200 people, this is the pattern we start from. It is a starting point, not a rule.
- Models: closed-source APIs for most work to begin with, behind a gateway, with business terms checked. Open-weight models where data or volume makes the case, proven with a test set first.
- Search over your own data: open-source parts running in your own accounts, because this is where your knowledge lives. Our Hybrid Search RAG example is built this way, on Qdrant and Postgres.
- Agents: open-source frameworks, so the logic is readable and yours. Our Recursive Agentic Improvements tooling works across several of the main ones.
- Guardrails and monitoring: open-source gateways and tracing where the team can run them, a managed version where it cannot. We covered the options in What a sane agent spend cap looks like.
- How we build software with agents: written up in the open in our AI Engineering Cookbook, so you can see the practice before you decide anything.
We publish these because the pieces are commodity and getting cheaper. The part that is not commodity is the judgment about which to use for which job, in which order, and who owns it.
Frequently asked questions
Should a small business use open-source or closed-source AI? Usually both, for different jobs, decided on data sensitivity, cost at volume, capability, lock-in and team skills.
Is open-source AI cheaper? Not automatically. The licence may be free, but hosting and running it is not. Open source tends to win at high, steady volume or for plumbing that is simple to run.
What is the difference between open-source and open-weight models? Open-weight models publish the trained model but may attach licence conditions, and often do not publish training data or code. Read the licence.
How do we avoid lock-in? Keep your data and logic in your own accounts, put a gateway in front of model providers, and keep a test set so you can check a replacement before switching.
Working with us
If you want to see the tools we use and publish, they are all on our Toolkit page. If you would rather have someone make these calls inside your team, that is part of what the AI-Native CTO OS installs.