September 18, 2026

Company Data Enrichment API: Providers, Pricing, and How to Test Them

What Is a Company Data Enrichment API? A Practical Guide for Developers

TLDR: A company data enrichment API takes a domain, company name, or profile URL and returns a company record: industry, headcount, headquarters, founding year, and often funding and parent company. Providers differ most on accepted inputs, no-match handling, billing rules, and record freshness. This guide compares five of them on those points, with prices as of September 2026, plus a script to test them on your own domains.

This guide is for choosing and testing a provider. Related topics have their own pages: resolving a messy name to one company is in the company lookup API guide, the firmographic field set and a public-source build are in the firmographic data API guide, contact records are in the lead enrichment API guide, and chaining providers is in waterfall enrichment.

What does a company data enrichment API return?

You send one identifier and get back one company record. Every provider here accepts a domain, which is usually the most precise input you have. People Data Labs (PDL) also takes a name, social profile URL, ticker, or its own ID, and its endpoint reference says website, ticker, or profile inputs are more likely to match than a name. Apollo's organization enrichment reference takes a domain, LinkedIn URL, or website, and does not accept a name on its own.

Beyond firmographics, providers add funding, detected technologies (see the technographic data API guide), social profiles, and parent or subsidiary links. Two fields matter more than the field count: match confidence and a freshness stamp. PDL returns a likelihood score from 1 to 10, Hunter returns an indexedAt date, and The Companies API returns meta.syncedAt, which its OpenAPI spec describes as the date the data was last synced.

Check confidence before you trust a match. PDL's input parameter docs set the default minimum likelihood at 2 and say a match at that score has roughly a 10 to 30 percent chance of being the company you asked for. Name-only requests score between 2 and 5, and PDL recommends a min_likelihood of 6 or higher when accuracy matters.

How do the main company enrichment APIs compare?

All five providers below document their company endpoints publicly. Prices are as of September 2026 and change often, so confirm them before you budget.

ProviderEndpoint and inputWhat you are billed forEntry pricingWorth knowing
People Data LabsGET /v5/company/enrich with name, website, profile, ticker, or PDL IDEach match; no match returns 404Free tier up to 100 records a month; paid plans from about $100 a month (pricing)Match threshold and required fields are set per request
ApolloGET /api/v1/organizations/enrich with domain, LinkedIn URL, or website1 credit per organization (credit table)Credits come with your Apollo plan; free accounts must register with a work emailParent, ultimate parent, and subsidiary fields; bulk endpoint takes up to 10 companies
HunterGET /v2/companies/find with a domain0.2 credits, only when name, category, location, and size all return (credit rules)Free plan with 50 credits a month; Starter at $49 a month for 2,000 credits (pricing)indexedAt date, parent domains, optional Clearbit-format output
The Companies APIGET /v2/companies/{domain}1 credit per company found; nothing when not found (endpoint docs)500 free credits; Startup at $95 a month for 50,000 credits (pricing)Free simplified profile; refresh=true re-crawls the company for 10 extra credits
CoresignalGET /cdapi/v2/company_multi_source/enrich with a website URL20 credits per successful multi-source record (endpoint table)7-day trial; plans from $49 a month for 2,500 credits (pricing)Monthly series such as employees_count_by_month

At list prices, one enriched company costs about $0.002 on The Companies API's Startup plan ($95 for 50,000 credits), about $0.005 on Hunter's Starter plan (0.2 credits at $49 for 2,000), and about $0.14 for a Coresignal multi-source record on its Growth plan (20 credits at $0.007). That roughly 70-fold spread buys different data, such as Coresignal's monthly headcount history, so price the fields you need rather than the call.

If older code calls Clearbit, plan a migration. According to Clearbit's announcement, HubSpot acquired Clearbit in December 2023, shut down the free Clearbit platform on April 30, 2025, and retired the Logo API on December 1, 2025. Its successor, Breeze Intelligence, enriches records inside HubSpot. Hunter's API reference documents a clearbit_format parameter that returns its company response in Clearbit's schema, which can shorten the port.

What does a no-match look like, and what do you pay for?

Match rates only compare across providers after you normalize what a miss looks like. PDL and Hunter return HTTP 404 when they find nothing. The Companies API's endpoint page says a miss returns an empty object with no charge, while its OpenAPI spec defines a 404 with companyNotFound, so handle both. Apollo returns 422 when a request has no usable identifier, which is a request error, not a miss.

Errors vary as much. Hunter returns 403 at its rate limit and 429 when your plan's usage runs out, so a client that backs off only on 429 gets both cases wrong. The Companies API returns 403 with noCreditsRemaining, and Coresignal returns 402 with "Insufficient credits", per its credits docs. Log any of these as "not found" and your pipeline blanks records the day credits run out, so treat only a documented miss as a miss.

A match is not always a usable record, and billing rules decide who pays for the gap. Hunter charges only when name, category or description, location, and company size all come back. PDL charges per match, but its required parameter makes a response count as a match only when the fields you name are present, so you pay only for those. The Companies API and Coresignal charge for any record found, sparse or complete. Compare cost per usable record, which the bake-off below computes.

Reconcile your estimates against each provider's own counter. The Companies API reports remaining credits in meta.credits on each response, Coresignal returns an x-credits-remaining header, and Hunter's GET /v2/account endpoint reports credits used and remaining.

How do you normalize responses from several providers?

Every provider nests its record differently. PDL puts company fields at the top level rather than inside a data object, Hunter wraps them in data, Apollo in organization, and The Companies API splits them across about, locations, and meta. One adapter per provider that maps into a single internal record, with an explicit matched flag, keeps the rest of your pipeline vendor-neutral.

Headcount needs the most care. PDL's employee_count counts the profiles it links to the company, and its company schema says that may be higher or lower than the real headcount; its size field is a self-reported band. Hunter and The Companies API each return a band and a number, and the bands do not line up: PDL's canonical sizes run 1-10, 11-50, 51-200, while The Companies API's enum runs 1-10, 10-50, 50-200. The adapter keeps each vendor's band, parses it into numeric bounds, and keeps the freshness stamp that came with it.

import json
import urllib.error
import urllib.parse
import urllib.request


def pdl_request(domain, key):
    query = urllib.parse.urlencode({"website": domain, "min_likelihood": 6})
    return urllib.request.Request(
        "https://api.peopledatalabs.com/v5/company/enrich?" + query,
        headers={"X-Api-Key": key})


def hunter_request(domain, key):
    query = urllib.parse.urlencode({"domain": domain})
    return urllib.request.Request(
        "https://api.hunter.io/v2/companies/find?" + query,
        headers={"X-API-KEY": key})


def tca_request(domain, key):
    return urllib.request.Request(
        "https://api.thecompaniesapi.com/v2/companies/" + urllib.parse.quote(domain),
        headers={"Authorization": "Basic " + key})


def fetch(request, opener=urllib.request.urlopen):
    """Return (status, payload). Only a 404 counts as a no-match."""
    try:
        with opener(request, timeout=30) as response:
            body = response.read()
            return response.status, (json.loads(body) if body else {})
    except urllib.error.HTTPError as err:
        if err.code == 404:
            return 404, {}
        raise  # 401, 402, 403, 429: an error to retry or alert on, never "not found"


def parse_band(text):
    """'11-50', '10-50', '500-1k', '10K-50K', '10001+', 'over-10k' -> (low, high)."""
    if not text:
        return None
    t = text.lower().replace("over-", "").replace(",", "")

    def num(s):
        return int(float(s[:-1]) * 1000) if s.endswith("k") else int(s)

    if t.endswith("+"):
        return (num(t[:-1]), None)
    if "-" in t:
        low, high = t.split("-", 1)
        return (num(low), num(high))
    return (num(t), None)


def record(vendor, status, name, domain, band, count, industry, country,
           founded, as_of, confidence=None):
    return {
        "vendor": vendor,
        "matched": status == 200 and bool(name),
        "name": name,
        "domain": domain,
        "employees_band": parse_band(band),
        "employee_count": count,  # PDL counts profiles; others estimate
        "industry": industry,
        "country": country,
        "founded": founded,
        "as_of": as_of,
        "confidence": confidence,
    }


def from_pdl(status, p):
    version = p.get("dataset_version")
    return record("pdl", status, p.get("name"), p.get("website"), p.get("size"),
                  p.get("employee_count"), p.get("industry"),
                  (p.get("location") or {}).get("country"), p.get("founded"),
                  as_of=("release " + version) if version else None,
                  confidence=p.get("likelihood"))


def from_hunter(status, p):
    d = p.get("data") or {}
    metrics = d.get("metrics") or {}
    return record("hunter", status, d.get("name"), d.get("domain"),
                  metrics.get("employees"), metrics.get("employeesCount"),
                  (d.get("category") or {}).get("industry"),
                  (d.get("geo") or {}).get("country"), d.get("foundedYear"),
                  as_of=d.get("indexedAt"))


def from_tca(status, p):
    about = p.get("about") or {}
    hq = (p.get("locations") or {}).get("headquarters") or {}
    return record("tca", status, about.get("name"), (p.get("domain") or {}).get("domain"),
                  about.get("totalEmployees"), about.get("totalEmployeesExact"),
                  about.get("industry"), (hq.get("country") or {}).get("name"),
                  about.get("yearFounded"), as_of=(p.get("meta") or {}).get("syncedAt"))

The fetch function treats only a 404 as a miss and lets every other HTTP error propagate, so a 402 or a rate limit never becomes an empty record; an empty 200 from The Companies API lands as no match because it carries no name. The PDL request sets min_likelihood to 6; add required, for example size AND location, to pay only for records with those fields. PDL's dataset_version is a release number, not a per-record date. Store raw payloads so you can re-score without paying twice.

How do you run a bake-off on your own records?

A vendor's coverage claim describes its database, not your accounts. Pull a few hundred domains from your CRM that match your real mix of regions and company sizes, and hand-label about 50 of them with headquarters country and headcount from each company's own site or a registry filing. Send identical inputs to every candidate and save every raw response.

Score each provider on five numbers: match rate after normalizing misses, correctness on the labeled subset, fill rate per field, cost per usable record, and the age of its freshness stamps. The lead enrichment guide explains why a raw match rate overstates quality, and the waterfall enrichment guide covers ordering several providers by cost per incremental hit once you have these numbers.

from datetime import date

FIELDS = ("employees_band", "industry", "country", "founded")


def billed_units(vendor, rec):
    """Approximates each documented charging rule. Reconcile with the vendor's usage report."""
    if not rec["matched"]:
        return 0  # none of these three bills a clean no-match
    if vendor == "hunter":  # 0.2 credits, only when name, category, location, and size return
        core = ("name", "industry", "country", "employees_band")
        return 0.2 if all(rec[f] is not None for f in core) else 0
    return 1  # PDL: per match; The Companies API: per company found


def in_band(count, band):
    low, high = band
    return count >= low and (high is None or count <= high)


def score(rows, truth, price_per_unit, required=("country", "employees_band"), today=None):
    """rows: [(domain, record)] for one vendor; truth: {domain: {"country", "employees"}}."""
    today = today or date.today()
    matched = [r for _, r in rows if r["matched"]]
    usable = [r for r in matched if all(r[f] is not None for f in required)]
    checks = []
    for domain, r in rows:
        label = truth.get(domain)
        if label and r["matched"]:
            checks.append((r["country"] or "").lower() == label["country"]
                          and r["employees_band"] is not None
                          and in_band(label["employees"], r["employees_band"]))
    stamps = [r["as_of"][:10] for r in matched if r["as_of"] and r["as_of"][:4].isdigit()]
    ages = sorted((today - date.fromisoformat(s)).days for s in stamps)
    spend = sum(billed_units(r["vendor"], r) for _, r in rows) * price_per_unit
    return {
        "match_rate": round(len(matched) / len(rows), 3),
        "fill_rate": {f: round(sum(r[f] is not None for r in matched) / max(len(matched), 1), 3)
                      for f in FIELDS},
        "correct_on_labeled": round(sum(checks) / len(checks), 3) if checks else None,
        "usable_records": len(usable),
        "cost_per_usable": round(spend / len(usable), 4) if usable else None,
        "median_stamp_age_days": ages[len(ages) // 2] if ages else None,
    }

Call score once per provider with its normalized rows and unit price, for example 95 / 50000 for The Companies API's Startup plan. The billed_units function approximates each documented rule, so check its totals against the provider's usage report before you trust the cost column. Sourcing methods, compliance, and contract terms sit outside a bake-off; the B2B data API buyer's guide covers them.

Where does live web data fit?

Every provider in the table serves a stored snapshot. When a freshness stamp is old on a field that changes, such as headcount, funding, or headquarters, or when no provider matches a company at all, a cited lookup against the live web fills the gap. You.com's APIs search and read the web for each request instead of returning a stored company record, so they belong next to an enrichment provider as a verification step, not in place of one.

The You.com Research API returns that lookup as structured JSON when you pass an output_schema: the result arrives in output.content and the pages it used in output.sources. The structured output docs say the API does not add citation fields to your schema, so the schema below gives each field its own source_url and as_of, and makes every field nullable so a missing value returns null rather than an empty string. Structured output needs a research_effort of standard or higher, which costs $50 per 1,000 calls (about $0.05 per profile) with typical latency of 10 to 30 seconds, per the Research API docs as of September 2026. That suits high-value accounts, not bulk enrichment.

import json
import os
import urllib.request

FACT = {
    "type": "object",
    "properties": {
        "value": {"type": ["string", "null"]},
        "source_url": {"type": ["string", "null"]},
        "as_of": {"type": ["string", "null"]},
    },
    "required": ["value", "source_url", "as_of"],
    "additionalProperties": False,
}
FIELDS = ["legal_name", "headquarters_country", "employee_count", "latest_funding_round"]
SCHEMA = {
    "type": "object",
    "properties": {name: {"$ref": "#/$defs/fact"} for name in FIELDS},
    "required": FIELDS,
    "additionalProperties": False,
    "$defs": {"fact": FACT},
}


def company_profile(domain, api_key, opener=urllib.request.urlopen):
    body = {
        "input": (f"Build a company profile for the organization whose website is {domain}. "
                  "For each field, give the value, the URL of the page that states it, and "
                  "the date that page gives. Use null when no source states the value."),
        "research_effort": "standard",  # output_schema returns 422 on lite
        "source_control": {"boost_domains": [domain]},
        "output_schema": SCHEMA,
    }
    request = urllib.request.Request(
        "https://api.you.com/v1/research", method="POST",
        data=json.dumps(body).encode("utf-8"),
        headers={"X-API-Key": api_key, "Content-Type": "application/json"})
    with opener(request, timeout=120) as response:
        return json.loads(response.read())


def needs_review(result):
    """Fields with a value whose source_url is missing from output.sources."""
    output = result["output"]
    if output.get("content_type") != "object":
        raise ValueError("expected structured output")
    cited = {s["url"].rstrip("/") for s in output.get("sources", [])}
    return [field for field, fact in output["content"].items()
            if fact["value"] is not None
            and (fact["source_url"] or "").rstrip("/") not in cited]


if __name__ == "__main__":
    result = company_profile("example.com", os.environ["YDC_API_KEY"])
    print(json.dumps(result["output"]["content"], indent=2))
    print("needs review:", needs_review(result))

The boost_domains setting favors the company's own site without excluding news, and needs_review flags any filled field whose source URL is not among the cited sources, so a person checks it before it reaches your CRM. To re-read known pages, such as an About or Careers page, the Contents API returns Markdown for up to 10 URLs per request at $1 per 1,000 pages as of September 2026, returns null content for pages it cannot fetch, and bypasses its cache when max_age=0. The firmographic data API guide builds a full Search plus Contents pass on that pattern; the Research API guide covers effort levels.

Which company enrichment API should you start with?

Start from the job, then test the two or three providers that fit it.

If you needStart by testingCheck before you commit
Low-cost enrichment of sign-ups or forms from a work-email domainHunter or The Companies APIShare of records meeting Hunter's core-field rule; band alignment; stamp age
Accurate matches from names, tickers, or profile URLsPeople Data LabsLabeled correctness at min_likelihood 6; which required fields to pay for
Parent and subsidiary rollupsApollo, with Hunter's parent domains as a cross-checkKnown subsidiaries resolve to the right parent (edge cases are in the company lookup API guide)
Headcount trends rather than a single numberCoresignal, or PDL's premium employee_count_by_month fieldCredit cost per record at your volume
Recent events, gap-fill, or evidence for a disputed fieldYou.com Research API with a per-field source schemaShare of fields flagged by needs_review; cost per profile

Decide on cost per usable record and labeled correctness, not list price or coverage counts. Run 200 of your domains through the adapters and score them. To add the verification step, get a key on the You.com platform; new accounts start with $100 in free credits, and current rates are on the pricing page.

    Share Article:

  1. LI Test

  2. LI Test

Related resources.

What Is Jev? TypeSafe AI's System One Model, Explained for Developers

What Is Jev? TypeSafe AI's System One Model, Explained for Developers

September 20, 2026

Blog

Best Local LLM for Coding: A Developer's Guide to AI-Powered Programming

Best Local LLM for Coding: A Developer's Guide to AI-Powered Programming

August 20, 2026

Blog

Local LLM: Running Large Language Models on Your Own Infrastructure

Local LLM: Running Large Language Models on Your Own Infrastructure

August 19, 2026

Blog

Lead Enrichment API: Automated Contact and Company Data Enhancement

Lead Enrichment API: Automated Contact and Company Data Enhancement

August 18, 2026

Blog

MAP Violation Monitoring: Automated Brand Protection for Ecommerce

MAP Violation Monitoring: Automated Brand Protection for Ecommerce

August 15, 2026

Blog