What Is Self-Hosted AI? A Practical Guide for Teams That Want Control

TLDR: Self-hosted AI means running models and the software around them (inference server, chat interface, retrieval, and workflow tools) on infrastructure you operate: a workstation, your own servers, or GPU instances in your own cloud account. You control where prompts go, which model version runs, and how cost scales, and you own sizing, security, and upkeep. Self-host the layers that touch sensitive data; buy the rest unless volume justifies owning them.
A self-hosted stack is more than a downloaded model. Each section below covers one decision and links to the guide that goes deeper on it.
What counts as self-hosted AI?
Self-hosted AI is defined by who operates the software and controls the data path, not by where the hardware sits. If your team runs the inference server, holds the model files, and decides which requests leave the network, the deployment is self-hosted, whether it runs on a laptop, in your data center, or on GPU instances in a cloud account you administer. Neighboring terms cover parts of the same picture:
| Term | What it emphasizes | Where it is covered |
|---|---|---|
| Self-hosted AI | Operating the whole stack anywhere you control, including your own cloud account | This guide |
| On-premise AI | Hardware in your own facilities, compliance boundaries, and ownership cost | On-premise AI guide |
| Local LLM | One language model on hardware you own | Local LLM overview |
| Self-hosted LLM | A serving stack that holds up under concurrent traffic | Self-hosted LLM serving guide |
| Self-hosted AI search engine | Retrieval plus cited answers over web or private sources | Self-hosted AI search engine guide |
When a case is ambiguous, trace one prompt: if it reaches only machines your team administers, the path is self-hosted. The tool name does not settle it. A signed-in Ollama server can run cloud models such as gpt-oss:120b-cloud through the same local API, per its authentication docs, and the Ollama FAQ documents OLLAMA_NO_CLOUD=1 to turn cloud features off.
What are the layers of a self-hosted AI stack?
Each layer has mature options you can run yourself, and their licenses vary more than teams expect. Open weights are not the same as open source, and several popular tools attach conditions that matter once a pilot becomes a product.
| Layer | What it does | Self-hostable options (license) | Watch for |
|---|---|---|---|
| Model weights | The trained model | Qwen3, gpt-oss, Gemma 4 (Apache 2.0); Llama 3.1 (Llama 3.1 Community License) | Llama 3.1 needs a separate Meta license above 700 million monthly active users, plus a "Built with Llama" notice when distributed |
| Inference server | Loads weights and serves an HTTP API | Ollama (MIT), llama.cpp (MIT), vLLM (Apache 2.0) | Pick by concurrency |
| Interface | Chat UI for people | Open WebUI (Open WebUI License) | You cannot remove its branding above 50 users in a rolling 30 days without written permission or an enterprise license |
| Document retrieval | Embeddings and vector search over your files | Qdrant (Apache 2.0), pgvector (PostgreSQL License) | Access control on the corpus |
| Web retrieval | Current public information | SearXNG (AGPL-3.0) or a search API | Leaves your network by design |
| Orchestration | Workflows, agents, and tool calls | n8n (Sustainable Use License) | Internal business, non-commercial, or personal use only |
| Speech and images | Transcription and image generation | whisper.cpp (MIT), ComfyUI (GPL-3.0) | ComfyUI's GPL terms apply if you modify and distribute it |
Terms come from each project's license file or model card as of September 2026, including the Llama 3.1 license, the Open WebUI license, and the n8n license. If you ruled out Gemma earlier, check again: Gemma 3 shipped under Google's Gemma Terms of Use, while the Gemma 4 model card lists Apache 2.0, the same license as Qwen3 and gpt-oss.
The model and the server are separate choices: the Ollama library carries all four model families above, and the gpt-oss README documents both vLLM and Ollama. Choose the model for the task, using the shortlist in the local AI models guide, and the server for the traffic, using the self-hosted LLM serving guide. Bundles shorten setup: n8n's self-hosted AI starter kit packages n8n, Ollama, Qdrant, and PostgreSQL in one Docker Compose template, though its README says it is not fully optimized for production.
Where can a self-hosted stack run?
All five shapes below are self-hosted. They differ in who owns the hardware and what still leaves your network.
| Shape | You control | What still leaves | Good fit |
|---|---|---|---|
| Single workstation | Hardware, OS, and model files | Model downloads and any web retrieval you enable | One developer, evaluation, offline work |
| Servers you own | Hardware, network, and physical access | Only the integrations you allow | Regulated data, steady volume, existing GPUs |
| GPU instances in your cloud account | Software, network rules, and keys | Traffic you allow out; the provider runs the physical hosts | Teams without a data center, or capacity that must scale |
| Air-gapped | Everything, with no outbound path | Nothing during operation; models and updates arrive by controlled transfer | Strict isolation, with answers limited to data you import |
| Hybrid | The sensitive path | Requests you deliberately route to hosted APIs | Mixed data sensitivity |
The cloud-account shape is where self-hosted and on-premise part ways. Your team runs the software and holds the keys, but the provider owns the hardware, and a compliance review may treat that differently from a rack in your building. The on-premise AI guide covers that boundary and the cost of owning hardware. A hybrid split holds only if the routing rule is enforced in code, not left to whoever writes the next integration.
How much hardware does a self-hosted model need?
Memory decides the hardware, and it has two parts: the weights, which are fixed, and the key-value (KV) cache, which grows with context length and with every concurrent request. A download size tells you only the first.
Take Llama 3.1 8B. Its weights are roughly 16 GB at 16-bit precision (8 billion parameters at 2 bytes each), and Ollama's default build is a 4.9 GB download, per the Ollama library. For the cache, Meta's Llama 3 paper lists 32 layers, a 4,096 model dimension split across 32 attention heads, and 8 key-value heads, which works out to 128 KiB per token at 16-bit precision, Ollama's default cache type according to its FAQ.
| Llama 3.1 8B, 16-bit KV cache | Math | Memory |
|---|---|---|
| Cache per token | 2 (keys and values) × 32 layers × 8 KV heads × 128 dimensions × 2 bytes | 128 KiB |
| One request at Ollama's default 4,096-token context | 4,096 × 128 KiB | 512 MiB |
| One request at the full 128K window | 131,072 × 128 KiB | 16 GiB |
| Eight concurrent requests at 32K tokens each | 8 × 32,768 × 128 KiB | 32 GiB |
A single full-length conversation needs more cache than the 4.9 GB quantized weights take, and concurrency multiplies it. Defaults differ: Ollama uses a 4,096-token context, while llama.cpp's -c and vLLM's --max-model-len take the model's own context length when unset. Cap context at what your prompts need; Ollama can also quantize the cache when Flash Attention is on. The self-hosted LLM serving guide covers how each server schedules this memory under load.
For team-scale reference points, the gpt-oss README says gpt-oss-120b (117 billion parameters, 5.1 billion active) runs on a single 80 GB GPU such as an NVIDIA H100 or AMD MI300X, and gpt-oss-20b (21 billion, 3.6 billion active) runs within 16 GB of memory, because their mixture-of-experts weights were post-trained in MXFP4. The Ollama library lists them at 65 GB and 14 GB, before any cache.
How do the layers talk to each other?
The OpenAI chat-completions format is the common interface. Ollama serves it under /v1 on port 11434 (compatibility docs), llama.cpp's llama-server listens on 127.0.0.1:8080 (server README), and vLLM starts at http://localhost:8000 (quickstart). Open WebUI connects to Ollama or any OpenAI-compatible API. Swapping engines becomes a base-URL change, so a team can start on Ollama and move to vLLM without rewriting the application.
The client below uses only the Python standard library, so it runs on an isolated host with nothing to install. It fails loudly on the conditions that otherwise produce quietly wrong output: an unreachable engine, a rejected key, a non-JSON reply, an empty answer, or a reply cut off at the token limit. It also bypasses proxy settings in the environment, so prompts go straight to the engine.
import json
import os
import urllib.error
import urllib.request
# Documented default addresses. Change them to match your deployment.
ENGINES = {
"ollama": "http://127.0.0.1:11434/v1",
"llamacpp": "http://127.0.0.1:8080/v1",
"vllm": "http://127.0.0.1:8000/v1",
}
# Talk to engines directly, never through a proxy set in the environment.
OPENER = urllib.request.build_opener(urllib.request.ProxyHandler({}))
class EngineError(RuntimeError):
pass
def call(url, payload=None, api_key=None, timeout=120):
headers = {"Content-Type": "application/json"}
if api_key:
headers["Authorization"] = "Bearer " + api_key
data = None if payload is None else json.dumps(payload).encode("utf-8")
request = urllib.request.Request(
url, data=data, headers=headers, method="GET" if data is None else "POST")
try:
with OPENER.open(request, timeout=timeout) as response:
body = json.loads(response.read().decode("utf-8"))
except urllib.error.HTTPError as exc:
raise EngineError("HTTP %d from %s" % (exc.code, url)) from None
except OSError as exc: # refused connection, DNS failure, timeout
raise EngineError("cannot reach %s (%s)" % (url, exc)) from None
except ValueError:
raise EngineError("non-JSON response from %s" % url) from None
if not isinstance(body, dict):
raise EngineError("expected a JSON object from %s" % url)
return body
def list_models(base_url, api_key=None):
body = call(base_url + "/models", api_key=api_key, timeout=10)
return [m["id"] for m in body.get("data") or [] if isinstance(m, dict) and "id" in m]
def chat(base_url, model, prompt, api_key=None, max_tokens=512):
body = call(base_url + "/chat/completions", {
"model": model,
"messages": [{"role": "user", "content": prompt}],
"max_tokens": max_tokens,
"stream": False,
}, api_key=api_key)
choices = body.get("choices") or []
if not choices or not isinstance(choices[0], dict):
raise EngineError("response contained no choices")
if choices[0].get("finish_reason") == "length":
raise EngineError("answer was cut off at max_tokens")
content = (choices[0].get("message") or {}).get("content")
if not isinstance(content, str) or not content.strip():
raise EngineError("engine returned an empty answer")
return content.strip()
if __name__ == "__main__":
base = os.environ.get("ENGINE_URL") or ENGINES[os.environ.get("ENGINE", "ollama")]
key = os.environ.get("ENGINE_API_KEY") # vLLM or llama-server started with --api-key
models = list_models(base, key)
if not models:
raise SystemExit("no models loaded at " + base)
model = os.environ.get("ENGINE_MODEL", models[0])
print(model + ": " + chat(base, model, "In one sentence, what is a KV cache?", key))
Save it as client.py, set ENGINE to ollama, llamacpp, or vllm (or ENGINE_URL for another address), and run python3 client.py. Compatible does not mean identical. Ollama ignores API keys locally, lists tool_choice and logprobs as unsupported, and cannot set context size through the OpenAI interface, so use a Modelfile or OLLAMA_CONTEXT_LENGTH. llama.cpp's /v1/models returns one model whose ID defaults to the file path. vLLM serves one model per server process. Test every engine with the prompts your application actually sends.
What security work moves to you?
Self-hosted engines ship with defaults built for one developer on one machine. Closing the gap to a shared deployment is your job.
- Ollama's local API has no authentication. It stays private only because it binds to 127.0.0.1:11434 by default, per the Ollama FAQ. The official image overrides that with
OLLAMA_HOST=0.0.0.0:11434in its Dockerfile, and the documenteddocker run -p 11434:11434publishes the port beyond the host, which Docker's port publishing docs call insecure by default. Use-p 127.0.0.1:11434:11434instead. - vLLM's API key covers only part of the server. Its security docs say
--api-keyprotects endpoints under the/v1,/v2,/inference, and/cohereprefixes. Other endpoints stay open, and/invocationsoffers the same inference without a key. vLLM recommends a reverse proxy that allowlists endpoints. - llama.cpp is private by default but keyless.
llama-serverbinds to 127.0.0.1 and accepts requests without authentication until you pass--api-key. - Exposure is not hypothetical. A Cisco study published September 1, 2025 used Shodan to find 1,139 publicly reachable Ollama instances, 214 of them answering prompts with live models. It also found that 88.89% of the LLM endpoints it discovered used OpenAI-style routes, which makes exposed servers easy to script against.
Keep inference servers on loopback or a private network, put an authenticating reverse proxy with an endpoint allowlist in front of anything shared, and verify both from outside:
# On the inference host: which addresses are the engines bound to?
# In the Local Address column, expect 127.0.0.1 or a private IP, not 0.0.0.0, [::], or *.
ss -ltn | grep -E ':(11434|8000|8080)\b'
# From a second machine: the gateway must refuse a request without credentials.
# Expect 401 or 403, never 200.
curl -s -o /dev/null -w '%{http_code}\n' https://ai-gateway.internal.example/v1/models
Those checks cover inbound access. For outbound traffic, the on-premise AI guide describes a boundary test that pushes synthetic sensitive data through the stack and inspects every connection it makes.
How does a self-hosted model stay current?
Model weights stop learning at training time, and the serving runtime will not warn you when a question depends on something newer. The fix is retrieval at request time, from two sources with opposite boundary properties.
Internal documents can stay inside: embed them locally (Ollama serves /v1/embeddings), store the vectors in Qdrant or pgvector, and retrieve with no outbound call. Current public information cannot, because the web is outside your network. Self-hosting the search layer changes who runs the intermediary, not who sees the query: the SearXNG API docs state that the query is passed to external search services. The self-hosted search engine guide compares SearXNG with crawler-based options that build their own index.
A search API makes that outbound call explicit and contractual. The You.com Web Search API takes a POST to https://ydc-index.io/v1/search with an X-API-Key header and returns results under results.web and results.news, per the Search API reference. As of September 2026 it costs $5.00 per 1,000 calls, according to the billing docs. Zero Data Retention is available for the Web Search and Answer APIs on enterprise agreements, enabled on request. ZDR limits what is retained; it does not keep the query inside your network, so send only text your policy allows to be public.
If you self-host gpt-oss, part of this wiring exists already. The reference browser tool in OpenAI's gpt-oss repository ships YouComBackend as its default and ExaBackend as the alternative, and OpenAI labels the tool educational, telling production users to build their own equivalent. You.com's gpt-oss integration guide shows the setup. For the full answer pipeline, including separating the private question from the public query and checking citations, see the self-hosted AI search engine guide. For retrieval alone, see how to build RAG with web search.
When is self-hosting worth it?
Decide layer by layer, starting from the data each layer touches:
| Layer | Self-host it when | Buy it when |
|---|---|---|
| Model and inference | Prompts hold data you may not send out, you must pin a model version, or steady volume keeps GPUs busy | Traffic is small or spiky, or the task needs a model with no open-weight equivalent |
| Interface | Users are internal and inference is already self-hosted | Users already work in a hosted assistant your policy allows |
| Document retrieval | The corpus is private or access-controlled | The corpus is public and a hosted index covers it |
| Web retrieval | You can operate metasearch and accept that upstream engines see queries | You want structured results and a retention agreement without operating search infrastructure |
| Orchestration | Workflows call internal systems with internal credentials | Workflows touch only SaaS tools that already hold the data |
Cost follows utilization. Self-hosting turns per-token billing into fixed costs for hardware or GPU rental, power, and engineering time. It pays off when that monthly total, divided by the tokens you actually serve, falls below the hosted price for a model that meets your quality bar. Owned or reserved GPUs cost the same when idle, so low utilization erases the savings. The on-premise AI guide covers ownership costs in detail.
Before committing, measure real traffic and answer five questions in writing:
- Data: What share of requests holds data your policy says cannot leave the network? Near zero means the case for self-hosting inference is weak.
- Volume: How many tokens do you serve per day, and what is peak concurrency? Those numbers size the hardware and pick the serving stack.
- Freshness: What share of questions needs current information? That share is how much of the stack depends on a web retrieval call.
- Quality: How does your strongest candidate model compare with your hosted baseline on 20 to 50 of your own prompts, graded by someone who knows the right answers?
- Ownership: Who owns upgrades, security patches, and on-call? If the answer is nobody, buy.
Starting from zero? Pilot on one machine with Ollama and the client above, then let the five answers choose the shape. The run LLM locally walkthrough covers the single-machine setup.
LI Test
LI Test
Share Article:
Related resources.

How to Use the You.com Web Search API in TypeScript
September 22, 2026
Blog

How to Build a News Search Pipeline With the You.com Web Search API
September 22, 2026
Blog

What Is the You.com Web Search API? Endpoint, Pricing, and Limits
September 22, 2026
Blog

How Authentication Works in the You.com Web Search API
September 21, 2026
Blog

How to Use Date Filters With the You.com Web Search API
September 21, 2026
Blog
