Back to blog
Company data APIOctober 11, 2026·13 min read

Company Data API Guide: Build with Official Records

By The SOSfinder Team

Company Data API Guide: Build with Official Records

“Real-time company data” is often sold as if it means real-time registry truth. It doesn't. A fast response can still contain a cached status, an outdated registered agent, or no evidence that a recent filing was retrieved successfully.

That distinction matters when a lookup controls onboarding, KYC, vendor approval, or a legal filing workflow. The useful differentiator isn't the lowest latency alone. It's whether the API shows where each value came from, when it was retrieved, how fresh it is, and what the system couldn't verify.

What a Company Data API Actually Is

A company data API is a read-heavy data system that collects official entity records from state registries, including all 50 U.S. states and the District of Columbia, then maps them into a consistent response model. The engineering challenge isn't sending JSON. It's reconciling different search behaviors, field names, filing formats, status definitions, and availability conditions without hiding the differences that affect a decision.

A name-search endpoint returns a result. A production-grade API should return an evidence-aware result. That means including the jurisdiction, registry identifier, retrieval timestamp, source timestamp where available, cache age, confidence, and meaningful outcomes such as not found, temporarily unavailable, and source blocked.

A diagram illustrating how a Company Data API aggregates information from 50 U.S. state registries into a JSON response.

Formation data isn't operating-company data

Official data also requires careful interpretation. The U.S. Census Bureau's Business Formation Statistics defines a business application as an EIN application submitted through IRS Form SS-4. A business formation is identified through the first instance of payroll-tax liabilities associated with that application.

Those are different events. An EIN application can appear before an entity becomes an employer, so an application isn't automatically an operating company. In December 2024, seasonally adjusted U.S. business applications reached 457,544, while projected employer formations within four quarters totaled 28,834, according to the same Census release. A company-data API should therefore keep application, formation, operating status, and employment milestones separate rather than treating them as interchangeable.

Registry coverage must include nonemployers

Payroll data alone also gives an incomplete view. The SBA Office of Advocacy's 2024 summary reported 34.8 million U.S. small businesses, representing 45.9% of employment. Its detailed estimate counted 34,752,434 small businesses, including 28,477,518 nonemployer firms, or 81.9%, and 6,274,916 firms with paid employees, or 18.1%.

The practical conclusion is simple: legal identity and employment status are different dimensions. Registry records can identify a newly formed or nonemployer business that won't appear in conventional employer datasets. A useful API preserves that distinction instead of presenting a narrow employment-oriented view as the complete company universe.

Core Endpoints and Normalized Schema

The API shape should match the decisions your product needs to make. A unified entity-search endpoint normally accepts a jurisdiction parameter, search terms, and optional identifiers. A separate identifier lookup should use the state registry ID when available, because exact identifiers remove much of the ambiguity associated with common business names.

Name availability deserves its own operation. Formation services need to know whether a proposed name is already in use, while compliance teams may need a broader search that returns similar candidates rather than a single answer. Filing-history retrieval is another distinct endpoint because documents, annual reports, amendments, and status records have different retrieval costs and freshness characteristics.

A practical endpoint set includes:

  • Entity search: Find candidates by name, jurisdiction, address, or related query fields.
  • Identifier lookup: Retrieve a known registry record without repeating fuzzy matching.
  • Name clearance: Check whether a proposed business name appears available in a specific jurisdiction.
  • Filing history: Return filing metadata and available document links for audit or review.
  • Monitoring and webhooks: Notify downstream systems when tracked fields change.

The entity search API guide is useful background for thinking through search behavior, identifiers, and normalized responses.

The normalized record

A normalized response should expose stable field names while preserving source-specific context. At minimum, expect fields such as:

Field Why it matters
Status Supports active, inactive, dissolved, suspended, or other jurisdiction-specific interpretations
Entity type Helps distinguish corporations, LLCs, nonprofits, partnerships, and other structures
Formation or registration date Establishes a timeline for onboarding and historical review
Principal address Supports identity matching and vendor records
Registered agent Provides a service-of-process relationship that may change independently of ownership
Officers or related persons Supplies corporate-role context, subject to state-level availability
Registry identifier Creates a durable key for subsequent retrieval and monitoring
Source and timestamps Show provenance, retrieval time, and freshness

Normalization shouldn't erase uncertainty. A provider may map several state-specific statuses into a common category, but the response should retain the original value and jurisdiction. It should also distinguish a genuinely empty result from a failed upstream request. Returning an empty array for a blocked registry can cause an onboarding system to approve a customer under the false assumption that no record exists.

How to Use Entity and Ownership Fields in Verification

Fast company lookup does not give you decision-grade verification. What decides onboarding quality is whether each entity and ownership field carries usable lineage, retrieval time, and enough structure to show what was confirmed, what was submitted by the customer, and what still needs review.

Under FinCEN's customer due diligence rule, covered U.S. financial institutions must identify and verify beneficial owners of legal-entity customers, including each individual who owns 25% or more of the equity and one individual with significant control or management responsibility. A state registry record helps anchor the legal entity, but it usually does not establish the full beneficial-ownership picture on its own.

Use registry data to confirm the legal name, jurisdiction, registry identifier, formation date, status, registered agent, officers, and filings. Then compare that record with customer-supplied ownership documents, identity checks, and risk-based review. If you want a broader reference for how these records fit together, a deeper guide to entity and ownership data is useful context.

Map fields to decisions

A sound verification flow separates identity, authority, and ownership. Systems that collapse those into one company object create avoidable errors, especially when the registry is current but the ownership package is stale.

  1. Match the legal entity. Compare the submitted legal name and jurisdiction with registry results. If the customer provides a registry identifier, treat that as the strongest match key. For name-based matching, keep alternate candidates and route weak-confidence matches to review instead of forcing a single result.

  2. Check current registry state. Status is evidence of the entity's condition in the registry at retrieval time. It does not prove that the submitted address, signer authority, ownership chart, or customer documents are current.

  3. Reconstruct the timeline. Formation date and filing history help explain whether the entity is newly formed, recently amended, reinstated, or drifting toward a status change. That context improves review quality, but it does not replace ownership verification.

  4. Separate people from entities. Store officers, registered agents, beneficial owners, and control persons as related records with relationship type, source reference, effective date when available, and retrieval timestamp. Freshness-aware schema design matters in this context. Officer data and beneficial-owner attestations often age at different rates.

  5. Create an immutable decision record. Preserve the raw state response and filing documents alongside normalized values. The verification event should record the matched entity identifier, jurisdiction, status, evidence used, reviewer or automated decision, and any discrepancy found.

Practical rule: A current status validates one fact about the registry record. It does not validate every fact about the customer.

This structure also makes corrections possible. If a later review shows an outdated address or an ownership document that conflicts with the registry record, investigators can see exactly what the system relied on at the time.

Performance, Caching, and Error Handling for Multi-State Lookups

A multi-state integration demands evidence-preserving reads. Registry response times, blocking behavior, document retrieval, and field availability differ by jurisdiction, so a single latency average hides the conditions that shape onboarding and compliance decisions. Fast lookup paths help, but speed alone does not tell you whether the record was fetched live, served from cache, or returned from a partially degraded upstream.

Measure performance by region and by operation. Track p50, p95, and p99 latency, throughput, error rate, and cache-hit ratio. Keep synthetic tests separate from real-user monitoring so you can see whether delays come from DNS, connection setup, TLS, time to first byte, total response time, or the registry itself. Identifier lookups usually need a tighter service objective than document retrieval, because the user decision often waits on the first and can tolerate delay on the second.

A flow diagram illustrating performance optimization, caching, and error handling strategies for multi-state lookup API architectures.

Cache for stability, not concealment

Caching cuts repeated upstream calls and smooths latency only if the cache key reflects how the request was made. Include jurisdiction, query type, entity identifier, and the search parameters that affect the result set. A name search and an identifier lookup should not share an entry because the same text appears in both requests.

TTL policy should follow field volatility. Filing-history documents can often be cached on a different schedule from status, registered-agent, or address data. Monitoring jobs should refresh or invalidate records after a relevant change is detected, and every response should expose enough metadata for downstream systems to tell a fresh retrieval from an older cached value. Good lookup caching guidance is useful here, but the key design decision is schema-level: make freshness visible per field or evidence object, not only at the envelope level.

Make failures actionable

Retries need limits. Use bounded exponential backoff, idempotent request handling, and clear retryable versus non-retryable error classes. Uncontrolled retries can turn a registry outage into duplicate work, duplicate billing events, and longer queues for requests that might have succeeded on the first pass.

Dashboards should show:

  • Latency by path: Separate cache hits, cache misses, identifier lookups, searches, and document retrieval.
  • Upstream health: Track timeout and blocking rates by jurisdiction.
  • Freshness: Report stale-response age and the age of the last successful retrieval.
  • Delivery: Monitor webhook success, retry counts, and permanently failed notifications.
  • Data quality: Track normalization errors, missing required fields, and conflicting source values.

The API should distinguish “not found” from “temporarily unavailable” and “source blocked.” Those outcomes drive different compliance actions. An empty result can support a search decision. An unavailable or blocked source should trigger retry, escalation, or a request for supplemental evidence.

Real-World Workflows for Onboarding, Monitoring, and Due Diligence

A product team integrating business verification usually starts with a familiar problem. Its onboarding form collects a legal name and jurisdiction, but the back end must query different registry interfaces, interpret different fields, and handle inconsistent document paths. A unified endpoint lets the team keep one internal workflow while passing the jurisdiction as a parameter.

During onboarding, the application can first perform a name search, present candidate matches when ambiguity exists, and then store the selected registry identifier. The next lookup can use that identifier instead of repeating a broad name search. If the result is unavailable, the workflow can pause rather than automatically treating the absence of data as evidence that the applicant has no registered entity.

Compliance teams can use the same record differently. Before approving a vendor for payment, they can compare the submitted legal name, address, status, registered agent, officers, and filing history. For higher-risk cases, filing documents can be attached to the review record, with the raw response and retrieval metadata retained for later examination. Guidance on combining corporate records with broader due diligence appears in this CDD and KYC resource.

A monitoring workflow watches for corporate drift after approval. Changes to status, registered agent, or address can create a review task, update a vendor master record, or request fresh documentation. The alert should include the changed field, previous value, new value, source, and retrieval time. Without that context, an alert creates noise instead of a usable investigation trail.

Legal filing services and registered-agent providers have another pattern. They retrieve filing histories and available state PDFs for client work, audits, underwriting, and case review. For portfolios, a CSV upload can initiate multi-state verification, while webhooks keep internal systems synchronized when tracked records change.

The common architecture is not “search once and trust forever.” It's lookup, normalize, preserve evidence, monitor material changes, and re-evaluate when freshness or source conditions require it.

Choosing and Validating a Company-Data Provider

Evaluate providers on decision quality, not response speed in isolation. Coverage depth matters, but so do source lineage, freshness controls, document access, monitoring, and error semantics.

Best-match name resolution deserves direct testing. A provider should return match confidence, the registry identifier, and alternative candidates instead of discarding ambiguity without notice. Test common names, punctuation differences, inactive entities, duplicate names across jurisdictions, and searches that produce no result.

Ask these questions before signing:

  • Freshness: Can responses show source timestamps, retrieval timestamps, cache age, and field-level freshness?
  • Caching: Do TTLs vary by jurisdiction, query type, status, address, agent, and document?
  • Errors: Can your system distinguish not found, unavailable, blocked, and malformed upstream responses?
  • Evidence: Are raw responses and filing documents available for audit?
  • Monitoring: Can alerts identify the exact field that changed and deliver reliably through webhooks?
  • Cost: How are searches, documents, monitoring events, retries, and bulk jobs counted?
  • Operations: Are service objectives defined by operation rather than only as a global uptime statement?

Run a proof of concept across target jurisdictions. Test identifier lookups, ambiguous names, document retrieval, webhook delivery, stale-result handling, and cross-jurisdiction field consistency. A provider that performs well in a clean demo but hides unavailable sources or ambiguous matches won't support decision-grade onboarding.

Next Steps for Building with Official Company Records

Start with the workflow, not the endpoint. List the fields your onboarding, KYC, vendor-risk, or legal process uses, then classify each as identity, status, ownership, authority, document evidence, or monitoring data.

Next, define a normalized schema with jurisdiction parameters and durable registry identifiers. Include source, retrieval timestamp, cache age, confidence, completeness, and explicit error states from the beginning. Adding lineage after launch is difficult because earlier decisions may no longer be reproducible.

Then test representative searches and identifier lookups across your target jurisdictions. Include ambiguous names, newly formed entities, inactive records, missing documents, and temporary source failures. Confirm that your application can pause or escalate when evidence is incomplete instead of converting uncertainty into approval.

Finally, enable monitoring and webhook alerts before scaling bulk verification. Store immutable verification events, preserve raw registry responses, and set refresh rules by field volatility and business risk. The strongest company data API integration treats official records as an auditable identity anchor, with normalized structure, freshness-transparent caching, and evidence-preserving decisions.


SOSfinder provides a single REST API for normalized business-entity records across all 50 states and DC, with search, identifier lookups, filing histories, monitoring, webhooks, and bulk verification workflows. If you're building onboarding or vendor-risk infrastructure, visit SOSfinder to evaluate an official-records integration around your required jurisdictions and evidence model.

Company data APIBusiness entity APIRegistered agent lookupCorporate records APIBusiness verification