Blog
 / 

What Is an OpenAI-Compatible API? What Ports and What Breaks

What Is an OpenAI-Compatible API? What Ports and What Breaks

TLDR: An OpenAI-compatible API accepts requests shaped like OpenAI's Chat Completions API, so you can point the official client at it by changing only base_url and api_key. The endpoint contract ports well; specific parameters do not. vLLM, llama.cpp, and Ollama each support a different subset, and an unsupported parameter gets rejected, silently dropped, or routed around depending on the provider. This guide shows exactly what ports, what breaks, and a tested script to check which behavior you got.

What "OpenAI-Compatible" Actually Means

OpenAI publishes its request and response shapes as part of the Chat Completions API: a POST to /v1/chat/completions with a messages array of role/content pairs, optional tools and tool_choice for function calling, and sampling parameters like temperature and max_tokens. The response is a ChatCompletion object: an id, a choices array where each entry carries a message and a finish_reason (stop, length, tool_calls, or content_filter), and usage counts. When a provider says it is "OpenAI-compatible," this is the contract it is claiming to implement. Because the shape is public, any server can implement it, which is why you can usually swap providers by changing base_url and api_key on the official OpenAI Python or TypeScript client and nothing else in your code.

That claim has a narrower scope than it sounds. "Compatible" almost always means the Chat Completions surface, sometimes the Completions and Embeddings surfaces, and rarely the full OpenAI API. Assistants, fine-tuning, and batch jobs are usually absent even from providers that nail chat completions, and support for the newer Responses API varies.

What Each Serving Stack Actually Implements

The three most common self-hosted servers do not implement the same subset. Their own documentation, not marketing copy, is the source of truth: vLLM's OpenAI-compatible server guide, llama.cpp's server README, and Ollama's OpenAI compatibility page.

EndpointvLLMllama.cpp (llama-server)Ollama
/v1/chat/completionsYes; user is ignored, parallel_tool_calls=false caps replies at one tool callYes, plus an Anthropic-style /v1/messages route (vLLM and Ollama have one too)Yes; no tool_choice, logit_bias, user, or n
/v1/completionsYes; suffix is not supportedYesYes, including suffix; no n, echo, or logit_bias
/v1/embeddingsYes, for embedding models onlyYes, for pooling-enabled modelsYes
/v1/responsesYesYes (translated to chat completions internally)Yes since v0.13.3, non-stateful only: no previous_response_id
LogprobsSupportedSupportedSupported since v0.12.11, though the compatibility page still marks it unsupported
Tool callingSupported, model-dependentSupported "for ~any model" via grammar constraintsSupported, but no forced tool_choice

One field-name split trips up a lot of ported code, documented in the am-i-openai-compatible project's compatibility matrix: OpenAI has no repetition-penalty parameter at all (only frequency_penalty and presence_penalty), vLLM calls its equivalent repetition_penalty, and llama.cpp calls it repeat_penalty. Ollama uses repeat_penalty only in its native API; its /v1 endpoint forwards only frequency_penalty and presence_penalty and ignores both spellings. Send the wrong name and nothing errors. The server drops the unknown key and samples with defaults, and the only symptom is output that looks a little more repetitive than you expected.

The Three Ways an Unsupported Parameter Fails

Every "OpenAI-compatible" layer has to decide what happens when your request includes a field it does not implement. Checking several providers' own compatibility pages side by side shows three distinct, each fully intentional, answers:

  • Reject it. Groq's compatibility docs state that logprobs, top_logprobs, logit_bias, and messages[].name "will result in a 400 error" if supplied, and that n must equal 1. This is the loud failure: your request breaks immediately.
  • Silently ignore it. Google's OpenAI-compatibility docs for the Gemini Enterprise Agent Platform state the opposite rule in one sentence: "If you pass any unsupported parameter, it is ignored." Your request returns 200. The field you set simply had no effect, and there is no error to notice.
  • Route around it. OpenRouter's API reference says that when the selected model does not support a field, "then the parameter is ignored. The rest are forwarded to the underlying model API," so by default OpenRouter behaves like the silent-ignore case. What sets it apart is provider.require_parameters: true, which turns a silent ignore into a routing exclusion: a provider that cannot honor every parameter you sent is skipped rather than called.

The silent-ignore case is the expensive one. A rejected request fails a test on day one. A silently dropped seed means your "reproducible" pipeline is not reproducible, and a silently dropped logit_bias means content you tried to suppress can still appear, and neither shows up until someone notices the output is wrong.

Testing What You Actually Got

Because a 200 response does not mean your parameters were honored, the only reliable check is to send a request that uses the fields your application depends on and inspect the response, not the provider's documentation page. The script below does three checks against a live /v1/chat/completions endpoint: that basic chat works at all, that a tool call comes back with arguments as a JSON string (some non-standard servers return a nested object here, which breaks any client that calls json.loads on it), and how the endpoint behaves when asked for n=2 plus seed, logprobs, and top_logprobs together.

import json
import urllib.error
import urllib.request


def post_json(url, api_key, payload, timeout=15):
    body = json.dumps(payload).encode("utf-8")
    req = urllib.request.Request(
        url, data=body, method="POST",
        headers={"Content-Type": "application/json",
                 "Authorization": "Bearer " + api_key})
    try:
        with urllib.request.urlopen(req, timeout=timeout) as resp:
            raw = resp.read().decode("utf-8")
            return resp.status, json.loads(raw)
    except urllib.error.HTTPError as e:
        return e.code, json.loads(e.read().decode("utf-8") or "{}")


def probe_compat(base_url, api_key, model="test-model"):
    url = base_url.rstrip("/") + "/chat/completions"
    report = {}

    status, body = post_json(url, api_key, {
        "model": model,
        "messages": [{"role": "user", "content": "Say the word OK."}],
    })
    report["basic_chat"] = status == 200 and bool(
        body.get("choices") and body["choices"][0]["message"].get("content"))

    status, body = post_json(url, api_key, {
        "model": model,
        "messages": [{"role": "user", "content": "Weather in Boston?"}],
        "tools": [{"type": "function", "function": {
            "name": "get_weather", "description": "Get current weather.",
            "parameters": {"type": "object",
                            "properties": {"city": {"type": "string"}},
                            "required": ["city"]}}}],
        "tool_choice": "auto",
    })
    calls = (body.get("choices") or [{}])[0].get("message", {}).get("tool_calls")
    report["tool_args_is_string"] = bool(
        calls and isinstance(calls[0]["function"]["arguments"], str))

    status, body = post_json(url, api_key, {
        "model": model,
        "messages": [{"role": "user", "content": "Say the word OK."}],
        "n": 2, "seed": 42, "logprobs": True, "top_logprobs": 3,
    })
    choices = body.get("choices", [])
    if status >= 400:
        report["risky_params"] = "rejected"
    elif len(choices) == 2 and all(c.get("logprobs") for c in choices):
        report["risky_params"] = "honored"
    else:
        report["risky_params"] = "silently_ignored"
    return report

This is not a hypothetical script. It ran offline against local fixture servers shaped like the documented behaviors: the Groq-style fixture returns a 400 when logprobs is present, and the Gemini-style fixture returns two choices without logprobs, because Google lists n as supported but not logprobs. The probe labeled them rejected and silently_ignored; it reports honored only when both choices carry logprobs. One response cannot prove seed works, so repeat a seeded request and compare before you build anything that depends on seed, logprobs, or tool calls.

Connecting From an Agent Framework

LangChain's own documentation gives the pattern directly: pass a custom base_url to ChatOpenAI.

from langchain_openai import ChatOpenAI

model = ChatOpenAI(
    base_url="https://your-provider.com/v1",
    api_key="your-api-key",
    model="provider-model-name",
)

LangChain is explicit about the limits of this pattern: ChatOpenAI "targets official OpenAI API specifications only," so non-standard response fields some providers add, such as a separate reasoning_content field for chain-of-thought, are not extracted or preserved. For a provider-specific parameter that has no OpenAI equivalent, such as vLLM's use_beam_search or LM Studio's model-eviction ttl, LangChain's docs specify the extra_body argument, not model_kwargs, because extra_body nests the field the way the provider expects while model_kwargs merges it into the top level and can trigger a request error. The Vercel AI SDK follows the identical shape on the TypeScript side: construct an OpenAI-compatible provider with a custom baseURL and the standard chat completions call works unchanged.

The Security Check Every Compatibility Guide Skips

Compatibility guides focus on request shape and rarely mention that "OpenAI-compatible" says nothing about how a server is secured. vLLM's own documentation carries a direct warning on this point: its --api-key flag "only authenticates requests to endpoints under the /v1, /v2, and /inference path prefixes," and explicitly calls out /invocations as unauthenticated even when an API key is configured, because it "exposes the same inference capabilities as the /v1 endpoints." If you deploy vLLM behind a load balancer that only checks the documented OpenAI-shaped routes, an unauthenticated path to the same model can still be reachable. vLLM's own recommendation is to not rely on --api-key alone and to put a reverse proxy in front of the whole server. Before you expose any self-hosted OpenAI-compatible server past localhost, enumerate every route it serves, not just the /v1/* ones you intended to expose.

When an OpenAI-Compatible API Is the Wrong Choice

The compatible surface is a lowest common denominator by design, so it excludes anything a provider built that OpenAI's API has no field for, such as custom rerankers or reasoning traces as a first-class response field. A provider can offer both: an OpenAI-compatible endpoint for portability and a richer native API for its own features. The OpenAI Web Search API is the reverse case worth knowing: it is a capability inside OpenAI's own API, not a separate OpenAI-compatible product, so "compatible" providers do not inherit it automatically.

Retrieval is where this tradeoff shows up in practice. Because the You.com Web Search API is not itself trying to be an OpenAI-compatible chat endpoint, it composes cleanly with any of the inference servers above through ordinary tool calling: the model, served through vLLM, llama.cpp, Ollama, or a hosted OpenAI-compatible provider, emits a tool_calls entry, your code resolves it with a search request, and the result goes back as a tool message. This is the same pattern LangChain's web search tool integration and the Vercel AI SDK's web search setup both use, and it works whether the model itself is a hosted frontier model or something you are running yourself; see how to run an LLM locally for the serving side of that setup.

A Pre-Flight Checklist

Before you ship code against a new "OpenAI-compatible" provider, confirm these five things against a real request and response, not the marketing page:

  1. Which endpoints it actually exposes: chat completions almost always; completions, embeddings, and responses vary by provider.
  2. What happens to a parameter it does not support: rejected or silently dropped, using the probe script above.
  3. Whether tool_calls[].function.arguments comes back as a JSON string, not a nested object.
  4. Whether streaming includes a final usage chunk when you request stream_options.include_usage, since some servers omit it.
  5. What is actually reachable on the server besides the documented /v1/* routes.

None of these takes more than a few test requests, and every one of them is the kind of gap that only shows up in production if you skip it.

    Share Article:

  1. LI Test

  2. LI Test

Related resources.

No items found.
No items found.