Skip to content
NPISignal

Methodology

The exact, deterministic rules NPISignal uses to validate NPIs, derive provider status, match exclusion records, score identity, and decide which pages get indexed.

Updated

Every derived fact on NPISignal — meaning anything it computes rather than copies verbatim from a source file — follows a fixed rule applied identically to every record. Nothing here is judged case by case, and nothing is generated by a language model.

NPI format validation

NPISignal checks that an NPI is exactly 10 numeric digits, begins with 1, 2, 3, or 4 (CMS currently issues only 1 and 2; 3 and 4 are reserved for future use but already pass as structurally valid), and carries a valid check digit: the Luhn (mod 10) algorithm applied to the constant 80840 followed by the number's first nine digits. The tenth digit must equal the result. See How Does NPI Number Validation Work? for the calculation worked out digit by digit.

Name normalisation

To make search tolerant of punctuation and diacritics without merging genuinely different names, every name is normalised the same way before comparison:

  • Lower-case the text.
  • Strip diacritics (e.g., José → jose).
  • Delete periods and both straight and curly apostrophes, so “St. John” and “ST JOHN” collapse together and “O'Brien” stays one token.
  • Turn every other non-alphanumeric character into a space.
  • Collapse repeated whitespace and trim.

This exact procedure is implemented once and shared between the database and the application code, so a name that matches in one place matches in the other. For individuals, the surname-first form of a name is also indexed, so “John Smith” and “Smith John” both find the same record.

Active vs. deactivated status

NPPES stores a deactivation date and a reactivation date on each record. NPISignal derives a single status from the two:

ConditionStatus shown
A deactivation date exists, and there is no reactivation date, or the reactivation date is earlier than the deactivation dateDeactivated
Any other case (no deactivation date, or a reactivation date equal to or later than the deactivation date)Active

See What Does NPI Deactivation Mean? for what deactivation does and does not tell you about a provider.

Taxonomy matching

A provider can report up to 15 taxonomy codes, one marked primary. NPISignal resolves each code against the NUCC reference table to show its grouping, classification, specialization, and definition. A taxonomy code that does not resolve against the current reference table — because it has been retired since the provider last updated their record — is shown as reported, without an invented description.

Exclusion matching

Matching a person or organization against an exclusion list is the one place on this site where NPISignal makes a judgement, so the rules are written down in full — and the same rules produce the result on the free search, on a provider profile, and inside a customer’s monitoring run. There is no model. Every point below is a field comparison a human can check against the two records side by side, and every stored match carries the version of these rules that produced it.

What is compared

Both sides are normalised with the procedure above, then compared field by field. Names are compared with Jaro-Winkler for people and trigram overlap for organizations. A documented table of common English diminutives lets Bill and William compare as the same given name — no distance metric bridges those two, and a screening product that misses that pairing is not doing its job.

What each agreement is worth

PointsSignal
+100The NPI on both records is identical
+55Organization names match exactly
+50Organization names match once legal suffixes (LLC, Inc, PLLC) are set aside
+40Dates of birth match exactly
+30First and last name match exactly
+26Last name matches; first names are a documented short form of each other
+22Last names share a component (hyphenated or married name) and first names agree
+22A person’s name matches a business name on the other record (the sole-proprietor case)
+20Last name matches; one record carries only a first initial
+20Organization names are similar but not identical
+15Whole names are a close spelling variant, not an exact match
+12Only the year of birth matches
+10Same state
+10Street number and street word agree
+10Specialty is recognisably the same field of practice
+5Middle name matches in full
+5Same city, within the same state
+3Middle initials match
−10Both records carry a middle initial and they differ
−10Both records carry an NPI and they differ
−35Both records carry a full date of birth and they differ
The name tiers are alternatives, not additions — the strongest one that applies is the one that counts.

What the score means

Below 30, nothing is reported. At 30 an exact full-name agreement clears the floor on its own, and a spelling variant alone does not. At 70 a match is labelled a strong identity match — which is exactly an exact name plus an exact date of birth, or a suffix-insensitive organization name plus a matching state and street address. An NPI agreement is labelled by the identifier rather than by the score, because an NPI is assigned rather than accumulated.

  • NPI match. The exclusion record carries the same NPI. The strongest signal in the public data, and still worth confirming with the source before acting.
  • Strong identity match. Several independent identity fields agree. Warrants review; not a confirmed identification.
  • Possible name match. The names are similar or identical and little else corroborates them. Common names produce results like this routinely.

Evidence against counts too

Two records carrying different full dates of birth are two different people, whatever their names say, and the arithmetic reflects that: a name agreement plus a date-of-birth disagreement falls below the reporting floor and is not raised at all. Where a contradiction does not suppress a match outright it is still shown, alongside the evidence for it, on the review screen.

The limit of what this can establish

NPISignal cannot conclusively verify identity from public data, and does not claim to. The downloadable LEIE contains no Social Security number and no employer identification number; where two people share a name and a date of birth, the public file genuinely cannot separate them. HHS-OIG operates its own verification process for that situation and it is the authoritative route. NPISignal does not automate it, and never presents a name match as a confirmed exclusion.

What a provider profile shows, and why it is narrower

A provider profile reports only identifier-based results: an exclusion record whose NPI is that NPI. It does not publish name-based potential matches, because that page is indexed under a real person’s name and a name match is a lead requiring review, not an identification — publishing one there would attach an allegation to somebody who may simply share a surname with an excluded party. Name-based screening is one click away on the exclusion search, behind a deliberate action, on a page that is never indexed.

Sources that could not be checked

A source NPISignal could not reach reports as unavailable — a distinct state from “no potential match identified”, everywhere it appears: on screen, in a screening run, and in a generated report. A screening in which a source failed is marked incomplete and names the records that were not checked. An unchecked source is not a negative result, and nothing in this product presents one as though it were.

Not every provider page is included in NPISignal's sitemap or told to search engines it may be indexed. Each provider gets a 0–100 score built additively from whether specific fields are present and usable:

PointsCondition
+25Status is active
+15Display name is present and longer than 2 characters
+15Primary taxonomy code is present and resolves in the taxonomy reference table
+15Primary state is present and is a real U.S. jurisdiction
+10Primary city is present
+10Practice address line 1 is present
+5Enumeration date is present
+5An individual has both a first and last name, or an organization has an organization name

A provider page is eligible for indexing only when its score is at least 75 and its status is active. Every other provider still gets a working page — the facts NPISignal has are always shown — but the page is marked noindex,follow rather than submitted for indexing. City directory pages have their own threshold: at least 50 active providers.

What NPISignal will not do

  • Generate descriptive text about a specific provider. Every word about a specific provider is a value copied from a source field, a computed label (like a status), or a link — never composed prose about that individual or organization.
  • Show a provider as “Verified,” “Licensed,” “Credentialed,” or “Certified” on the strength of NPPES data alone. NPPES license fields are labeled provider-reported, because that is what they are.
  • Merge or infer a fact that would require guessing across an ambiguous name match.
  • Use a language model to decide whether two records describe the same person. Exclusion matching is arithmetic over field comparisons, and every point is shown to whoever has to make the decision.
  • Describe a screening result as “compliant”, “cleared” or “passed”. The strongest thing NPISignal can say is that no potential match was identified in the data it holds, as of a stated date.
  • Collect a Social Security number. There is no column for one anywhere in the product, and an uploaded file containing one is refused rather than partially imported.
  • Publish a date of birth. NPISignal imports the one HHS-OIG publishes because matching needs it, and it never appears on a public page, in an export, in a report, or in an email.

Known limitations

  • NPISignal's change history for a provider begins the first time NPISignal itself observed that record. There is no way to reconstruct what a record looked like before its first import, and pages say so rather than implying a longer history than actually exists.
  • A record's “last changed” timestamp updates only when a meaningful field actually changes (name, status, primary taxonomy, or primary location) — a routine re-import that changes nothing does not touch it, so it stays an accurate signal of real change rather than import noise.
  • Every rule above is deterministic and reviewable; none of it is a machine-learning model making a judgment call about a specific provider.
  • Exclusion matching covers two federal sources — HHS-OIG LEIE and SAM.gov. State Medicaid exclusion lists are not covered, and many payer and state contracts require them separately.
  • An exclusion imposed after the copy of a source NPISignal holds was published cannot be found in it. Every result states the publication date of the data it searched, so the gap is visible rather than implied.
  • SAM.gov publishes neither a date of birth nor an NPI, so a SAM name match rarely has anything available to corroborate it and usually stays a possible match on public data alone.