September 18, 2026

What Is Self-Hosted AI? A Practical Guide for Teams That Want Control

What Is Self-Hosted AI? A Practical Guide for Teams That Want Control

TLDR: Self-hosted AI means running models and the software around them (inference server, chat interface, retrieval, and workflow tools) on infrastructure you operate: a workstation, your own servers, or GPU instances in your own cloud account. You control where prompts go, which model version runs, and how cost scales, and you own sizing, security, and upkeep. Self-host the layers that touch sensitive data; buy the rest unless volume justifies owning them.

A self-hosted stack is more than a downloaded model. Each section below covers one decision and links to the guide that goes deeper on it.

What counts as self-hosted AI?

Self-hosted AI is defined by who operates the software and controls the data path, not by where the hardware sits. If your team runs the inference server, holds the model files, and decides which requests leave the network, the deployment is self-hosted, whether it runs on a laptop, in your data center, or on GPU instances in a cloud account you administer. Neighboring terms cover parts of the same picture:

TermWhat it emphasizesWhere it is covered
Self-hosted AIOperating the whole stack anywhere you control, including your own cloud accountThis guide
On-premise AIHardware in your own facilities, compliance boundaries, and ownership costOn-premise AI guide
Local LLMOne language model on hardware you ownLocal LLM overview
Self-hosted LLMA serving stack that holds up under concurrent trafficSelf-hosted LLM serving guide
Self-hosted AI search engineRetrieval plus cited answers over web or private sourcesSelf-hosted AI search engine guide

When a case is ambiguous, trace one prompt: if it reaches only machines your team administers, the path is self-hosted. The tool name does not settle it. A signed-in Ollama server can run cloud models such as gpt-oss:120b-cloud through the same local API, per its authentication docs, and the Ollama FAQ documents OLLAMA_NO_CLOUD=1 to turn cloud features off.

What are the layers of a self-hosted AI stack?

Each layer has mature options you can run yourself, and their licenses vary more than teams expect. Open weights are not the same as open source, and several popular tools attach conditions that matter once a pilot becomes a product.

LayerWhat it doesSelf-hostable options (license)Watch for
Model weightsThe trained modelQwen3, gpt-oss, Gemma 4 (Apache 2.0); Llama 3.1 (Llama 3.1 Community License)Llama 3.1 needs a separate Meta license above 700 million monthly active users, plus a "Built with Llama" notice when distributed
Inference serverLoads weights and serves an HTTP APIOllama (MIT), llama.cpp (MIT), vLLM (Apache 2.0)Pick by concurrency
InterfaceChat UI for peopleOpen WebUI (Open WebUI License)You cannot remove its branding above 50 users in a rolling 30 days without written permission or an enterprise license
Document retrievalEmbeddings and vector search over your filesQdrant (Apache 2.0), pgvector (PostgreSQL License)Access control on the corpus
Web retrievalCurrent public informationSearXNG (AGPL-3.0) or a search APILeaves your network by design
OrchestrationWorkflows, agents, and tool callsn8n (Sustainable Use License)Internal business, non-commercial, or personal use only
Speech and imagesTranscription and image generationwhisper.cpp (MIT), ComfyUI (GPL-3.0)ComfyUI's GPL terms apply if you modify and distribute it

Terms come from each project's license file or model card as of September 2026, including the Llama 3.1 license, the Open WebUI license, and the n8n license. If you ruled out Gemma earlier, check again: Gemma 3 shipped under Google's Gemma Terms of Use, while the Gemma 4 model card lists Apache 2.0, the same license as Qwen3 and gpt-oss.

The model and the server are separate choices: the Ollama library carries all four model families above, and the gpt-oss README documents both vLLM and Ollama. Choose the model for the task, using the shortlist in the local AI models guide, and the server for the traffic, using the self-hosted LLM serving guide. Bundles shorten setup: n8n's self-hosted AI starter kit packages n8n, Ollama, Qdrant, and PostgreSQL in one Docker Compose template, though its README says it is not fully optimized for production.

Where can a self-hosted stack run?

All five shapes below are self-hosted. They differ in who owns the hardware and what still leaves your network.

ShapeYou controlWhat still leavesGood fit
Single workstationHardware, OS, and model filesModel downloads and any web retrieval you enableOne developer, evaluation, offline work
Servers you ownHardware, network, and physical accessOnly the integrations you allowRegulated data, steady volume, existing GPUs
GPU instances in your cloud accountSoftware, network rules, and keysTraffic you allow out; the provider runs the physical hostsTeams without a data center, or capacity that must scale
Air-gappedEverything, with no outbound pathNothing during operation; models and updates arrive by controlled transferStrict isolation, with answers limited to data you import
HybridThe sensitive pathRequests you deliberately route to hosted APIsMixed data sensitivity

The cloud-account shape is where self-hosted and on-premise part ways. Your team runs the software and holds the keys, but the provider owns the hardware, and a compliance review may treat that differently from a rack in your building. The on-premise AI guide covers that boundary and the cost of owning hardware. A hybrid split holds only if the routing rule is enforced in code, not left to whoever writes the next integration.

How much hardware does a self-hosted model need?

Memory decides the hardware, and it has two parts: the weights, which are fixed, and the key-value (KV) cache, which grows with context length and with every concurrent request. A download size tells you only the first.

Take Llama 3.1 8B. Its weights are roughly 16 GB at 16-bit precision (8 billion parameters at 2 bytes each), and Ollama's default build is a 4.9 GB download, per the Ollama library. For the cache, Meta's Llama 3 paper lists 32 layers, a 4,096 model dimension split across 32 attention heads, and 8 key-value heads, which works out to 128 KiB per token at 16-bit precision, Ollama's default cache type according to its FAQ.

Llama 3.1 8B, 16-bit KV cacheMathMemory
Cache per token2 (keys and values) × 32 layers × 8 KV heads × 128 dimensions × 2 bytes128 KiB
One request at Ollama's default 4,096-token context4,096 × 128 KiB512 MiB
One request at the full 128K window131,072 × 128 KiB16 GiB
Eight concurrent requests at 32K tokens each8 × 32,768 × 128 KiB32 GiB

A single full-length conversation needs more cache than the 4.9 GB quantized weights take, and concurrency multiplies it. Defaults differ: Ollama uses a 4,096-token context, while llama.cpp's -c and vLLM's --max-model-len take the model's own context length when unset. Cap context at what your prompts need; Ollama can also quantize the cache when Flash Attention is on. The self-hosted LLM serving guide covers how each server schedules this memory under load.

For team-scale reference points, the gpt-oss README says gpt-oss-120b (117 billion parameters, 5.1 billion active) runs on a single 80 GB GPU such as an NVIDIA H100 or AMD MI300X, and gpt-oss-20b (21 billion, 3.6 billion active) runs within 16 GB of memory, because their mixture-of-experts weights were post-trained in MXFP4. The Ollama library lists them at 65 GB and 14 GB, before any cache.

How do the layers talk to each other?

The OpenAI chat-completions format is the common interface. Ollama serves it under /v1 on port 11434 (compatibility docs), llama.cpp's llama-server listens on 127.0.0.1:8080 (server README), and vLLM starts at http://localhost:8000 (quickstart). Open WebUI connects to Ollama or any OpenAI-compatible API. Swapping engines becomes a base-URL change, so a team can start on Ollama and move to vLLM without rewriting the application.

The client below uses only the Python standard library, so it runs on an isolated host with nothing to install. It fails loudly on the conditions that otherwise produce quietly wrong output: an unreachable engine, a rejected key, a non-JSON reply, an empty answer, or a reply cut off at the token limit. It also bypasses proxy settings in the environment, so prompts go straight to the engine.

import json
import os
import urllib.error
import urllib.request

# Documented default addresses. Change them to match your deployment.
ENGINES = {
    "ollama": "http://127.0.0.1:11434/v1",
    "llamacpp": "http://127.0.0.1:8080/v1",
    "vllm": "http://127.0.0.1:8000/v1",
}

# Talk to engines directly, never through a proxy set in the environment.
OPENER = urllib.request.build_opener(urllib.request.ProxyHandler({}))


class EngineError(RuntimeError):
    pass


def call(url, payload=None, api_key=None, timeout=120):
    headers = {"Content-Type": "application/json"}
    if api_key:
        headers["Authorization"] = "Bearer " + api_key
    data = None if payload is None else json.dumps(payload).encode("utf-8")
    request = urllib.request.Request(
        url, data=data, headers=headers, method="GET" if data is None else "POST")
    try:
        with OPENER.open(request, timeout=timeout) as response:
            body = json.loads(response.read().decode("utf-8"))
    except urllib.error.HTTPError as exc:
        raise EngineError("HTTP %d from %s" % (exc.code, url)) from None
    except OSError as exc:  # refused connection, DNS failure, timeout
        raise EngineError("cannot reach %s (%s)" % (url, exc)) from None
    except ValueError:
        raise EngineError("non-JSON response from %s" % url) from None
    if not isinstance(body, dict):
        raise EngineError("expected a JSON object from %s" % url)
    return body


def list_models(base_url, api_key=None):
    body = call(base_url + "/models", api_key=api_key, timeout=10)
    return [m["id"] for m in body.get("data") or [] if isinstance(m, dict) and "id" in m]


def chat(base_url, model, prompt, api_key=None, max_tokens=512):
    body = call(base_url + "/chat/completions", {
        "model": model,
        "messages": [{"role": "user", "content": prompt}],
        "max_tokens": max_tokens,
        "stream": False,
    }, api_key=api_key)
    choices = body.get("choices") or []
    if not choices or not isinstance(choices[0], dict):
        raise EngineError("response contained no choices")
    if choices[0].get("finish_reason") == "length":
        raise EngineError("answer was cut off at max_tokens")
    content = (choices[0].get("message") or {}).get("content")
    if not isinstance(content, str) or not content.strip():
        raise EngineError("engine returned an empty answer")
    return content.strip()


if __name__ == "__main__":
    base = os.environ.get("ENGINE_URL") or ENGINES[os.environ.get("ENGINE", "ollama")]
    key = os.environ.get("ENGINE_API_KEY")  # vLLM or llama-server started with --api-key
    models = list_models(base, key)
    if not models:
        raise SystemExit("no models loaded at " + base)
    model = os.environ.get("ENGINE_MODEL", models[0])
    print(model + ": " + chat(base, model, "In one sentence, what is a KV cache?", key))

Save it as client.py, set ENGINE to ollama, llamacpp, or vllm (or ENGINE_URL for another address), and run python3 client.py. Compatible does not mean identical. Ollama ignores API keys locally, lists tool_choice and logprobs as unsupported, and cannot set context size through the OpenAI interface, so use a Modelfile or OLLAMA_CONTEXT_LENGTH. llama.cpp's /v1/models returns one model whose ID defaults to the file path. vLLM serves one model per server process. Test every engine with the prompts your application actually sends.

What security work moves to you?

Self-hosted engines ship with defaults built for one developer on one machine. Closing the gap to a shared deployment is your job.

  • Ollama's local API has no authentication. It stays private only because it binds to 127.0.0.1:11434 by default, per the Ollama FAQ. The official image overrides that with OLLAMA_HOST=0.0.0.0:11434 in its Dockerfile, and the documented docker run -p 11434:11434 publishes the port beyond the host, which Docker's port publishing docs call insecure by default. Use -p 127.0.0.1:11434:11434 instead.
  • vLLM's API key covers only part of the server. Its security docs say --api-key protects endpoints under the /v1, /v2, /inference, and /cohere prefixes. Other endpoints stay open, and /invocations offers the same inference without a key. vLLM recommends a reverse proxy that allowlists endpoints.
  • llama.cpp is private by default but keyless. llama-server binds to 127.0.0.1 and accepts requests without authentication until you pass --api-key.
  • Exposure is not hypothetical. A Cisco study published September 1, 2025 used Shodan to find 1,139 publicly reachable Ollama instances, 214 of them answering prompts with live models. It also found that 88.89% of the LLM endpoints it discovered used OpenAI-style routes, which makes exposed servers easy to script against.

Keep inference servers on loopback or a private network, put an authenticating reverse proxy with an endpoint allowlist in front of anything shared, and verify both from outside:

# On the inference host: which addresses are the engines bound to?
# In the Local Address column, expect 127.0.0.1 or a private IP, not 0.0.0.0, [::], or *.
ss -ltn | grep -E ':(11434|8000|8080)\b'

# From a second machine: the gateway must refuse a request without credentials.
# Expect 401 or 403, never 200.
curl -s -o /dev/null -w '%{http_code}\n' https://ai-gateway.internal.example/v1/models

Those checks cover inbound access. For outbound traffic, the on-premise AI guide describes a boundary test that pushes synthetic sensitive data through the stack and inspects every connection it makes.

How does a self-hosted model stay current?

Model weights stop learning at training time, and the serving runtime will not warn you when a question depends on something newer. The fix is retrieval at request time, from two sources with opposite boundary properties.

Internal documents can stay inside: embed them locally (Ollama serves /v1/embeddings), store the vectors in Qdrant or pgvector, and retrieve with no outbound call. Current public information cannot, because the web is outside your network. Self-hosting the search layer changes who runs the intermediary, not who sees the query: the SearXNG API docs state that the query is passed to external search services. The self-hosted search engine guide compares SearXNG with crawler-based options that build their own index.

A search API makes that outbound call explicit and contractual. The You.com Web Search API takes a POST to https://ydc-index.io/v1/search with an X-API-Key header and returns results under results.web and results.news, per the Search API reference. As of September 2026 it costs $5.00 per 1,000 calls, according to the billing docs. Zero Data Retention is available for the Web Search and Answer APIs on enterprise agreements, enabled on request. ZDR limits what is retained; it does not keep the query inside your network, so send only text your policy allows to be public.

If you self-host gpt-oss, part of this wiring exists already. The reference browser tool in OpenAI's gpt-oss repository ships YouComBackend as its default and ExaBackend as the alternative, and OpenAI labels the tool educational, telling production users to build their own equivalent. You.com's gpt-oss integration guide shows the setup. For the full answer pipeline, including separating the private question from the public query and checking citations, see the self-hosted AI search engine guide. For retrieval alone, see how to build RAG with web search.

When is self-hosting worth it?

Decide layer by layer, starting from the data each layer touches:

LayerSelf-host it whenBuy it when
Model and inferencePrompts hold data you may not send out, you must pin a model version, or steady volume keeps GPUs busyTraffic is small or spiky, or the task needs a model with no open-weight equivalent
InterfaceUsers are internal and inference is already self-hostedUsers already work in a hosted assistant your policy allows
Document retrievalThe corpus is private or access-controlledThe corpus is public and a hosted index covers it
Web retrievalYou can operate metasearch and accept that upstream engines see queriesYou want structured results and a retention agreement without operating search infrastructure
OrchestrationWorkflows call internal systems with internal credentialsWorkflows touch only SaaS tools that already hold the data

Cost follows utilization. Self-hosting turns per-token billing into fixed costs for hardware or GPU rental, power, and engineering time. It pays off when that monthly total, divided by the tokens you actually serve, falls below the hosted price for a model that meets your quality bar. Owned or reserved GPUs cost the same when idle, so low utilization erases the savings. The on-premise AI guide covers ownership costs in detail.

Before committing, measure real traffic and answer five questions in writing:

  1. Data: What share of requests holds data your policy says cannot leave the network? Near zero means the case for self-hosting inference is weak.
  2. Volume: How many tokens do you serve per day, and what is peak concurrency? Those numbers size the hardware and pick the serving stack.
  3. Freshness: What share of questions needs current information? That share is how much of the stack depends on a web retrieval call.
  4. Quality: How does your strongest candidate model compare with your hosted baseline on 20 to 50 of your own prompts, graded by someone who knows the right answers?
  5. Ownership: Who owns upgrades, security patches, and on-call? If the answer is nobody, buy.

Starting from zero? Pilot on one machine with Ollama and the client above, then let the five answers choose the shape. The run LLM locally walkthrough covers the single-machine setup.

    Share Article:

  1. LI Test

  2. LI Test

Related resources.

How to Use the You.com Web Search API in TypeScript

How to Use the You.com Web Search API in TypeScript

September 22, 2026

Blog

How to Build a News Search Pipeline With the You.com Web Search API

How to Build a News Search Pipeline With the You.com Web Search API

September 22, 2026

Blog

What Is the You.com Web Search API? A Practical Guide for Developers

What Is the You.com Web Search API? Endpoint, Pricing, and Limits

September 22, 2026

Blog

How Authentication Works in the You.com Web Search API

How Authentication Works in the You.com Web Search API

September 21, 2026

Blog

How to Use Date Filters With the You.com Web Search API

How to Use Date Filters With the You.com Web Search API

September 21, 2026

Blog