Skip to main content

Match

Identify related records across your data sources using intelligent matching algorithms. The Match module finds duplicates within systems and links records across systems to build a complete view of each entity.

Why Matching Matters

Organizations typically have the same people, accounts, or entities represented in multiple systems. Without matching:
  • Records are fragmented across operational, billing, and administrative systems
  • Duplicate records create confusion and operational gaps
  • Analytics undercount or overcount unique individuals
  • Teams lack the complete information needed for effective coordination
skyMDM’s Match module uses advanced algorithms to accurately identify when two records represent the same real-world entity.

What You Can Do

Configure Match Rules

Define which fields to compare and how much weight each field carries in match scoring.

Choose Match Algorithms

Select from exact matching or fuzzy matching approaches based on your data quality needs.

Review Match Candidates

Examine potential matches with detailed comparison views before accepting.

Set Match Thresholds

Configure score thresholds for auto-match, manual review, and non-match decisions.

Match Types

Each attribute in a rule is compared with a match type you choose, so you can be strict on identifiers and forgiving on names.

Exact

String equality — the right choice for member IDs, SSNs, and email addresses that should be identical

High

Tolerates minor typos and transposed characters, such as “Jon Smith” against “John Smith”

Medium

Tolerates stronger variation, such as “Jon Smyth” against “John Smith”

Low

Broad fuzziness for badly degraded data, such as “Jno Smyth” against “Jonathan Smith”

Nickname

A curated dictionary of common nicknames, so “Bob” matches “Robert” and “Liz” matches “Elizabeth”

Date of Birth

Date matching with tolerance for two-digit year flips, transpositions, and month or day swaps

Two Ways to Configure Matching

Semantic Label Mode

Pick the labels to match on, choose a match type and priority for each, and set one overall confidence threshold. skyMDM applies the rules across every pair of tables.

Cascading Step Mode

Author a multi-step strategy for Member 360, Patient 360, and Resident 360. Step one might be an exact match on member ID; records it resolves are excluded from step two, which can fall back to a fuzzy match on name and date of birth. Each step carries its own rules and thresholds.

Match Attributes

Name Matching

Handle nicknames, misspellings, name changes, and cultural naming conventions

Address Matching

Match despite formatting differences, abbreviations, and address changes

ID Matching

Exact matching on SSN, account IDs, and other unique identifiers

Date Matching

Match dates across different formats with tolerance for data entry errors

Phone/Email

Match contact information with normalization and validation

Custom Fields

Include any field in your matching strategy based on your data

Key Capabilities

Match Scoring

Every potential match receives a confidence score from 0 to 100 percent, calculated as a weighted average of the per-attribute similarity scores. Higher-priority attributes carry more weight, and the score always shows which attributes agreed and how strongly, so no match is a black box.

Null-Safe Scoring

An attribute missing on either side is left out of the average rather than counted as a mismatch. A three-attribute rule where one record has no phone number becomes a two-attribute comparison, so incomplete records are not unfairly penalized.

The Review Band

Set an auto-merge threshold and, optionally, a review band beneath it. Pairs at or above the threshold merge automatically, pairs inside the band go to the Review Queue for a data steward to decide, and anything below is discarded. This is how you keep automation on the confident cases and human judgment on the rest.

Configuration Checks

Saving a match configuration runs a check over it first and reports errors, warnings, and notes — duplicate step names, rules with no effect, unused semantic labels, thresholds that would leave an empty review band. Misconfigurations surface before the job runs, not after.

Transitive Matching

If Record A matches Record B, and Record B matches Record C, skyMDM recognizes that all three may represent the same entity.

Block and Compare

Efficiently process large datasets by first blocking records into candidate groups, then running detailed comparisons only within blocks.

Persistent Cluster Identity

Clusters keep the same identifier across re-runs, so a person’s identity survives new data arriving. Re-running Match carries each cluster’s existing identity and its golden record forward instead of renumbering everything, and a full-rebuild mode is available when you deliberately want to start clean.

Progress, Cancel, and Recovery

Match jobs report progress continuously in the run window, and a run can be cancelled from the same place. Previous runs stay available in the run history with the person who triggered each. A reconciler detects runs that stalled or lost their completion signal and heals the status automatically.

Runs Where Your Data Lives

Matching executes natively on Databricks, Snowflake, or Microsoft Fabric, with the same rules, scores, and clusters on every platform.

Match Audit Trail

Every match decision is logged with the score, contributing fields, and timestamp for compliance and quality review.

Business Impact

Complete View

Link records across systems for a comprehensive view of each individual or entity.

Accurate Counts

Know exactly how many unique individuals or accounts you serve.

Better Coordination

Connect teams with complete information for better decision-making.

Who Benefits

  • Operations Teams: Access complete records for planning and coordination
  • Quality Teams: Accurate attribution for quality measures and reporting
  • Finance Teams: Proper counting for revenue and cost analysis
  • Compliance Teams: Maintain accurate records for regulatory reporting