Match
Identify related records across your data sources using intelligent matching algorithms. The Match module finds duplicates within systems and links records across systems to build a complete view of each entity.Why Matching Matters
Organizations typically have the same people, accounts, or entities represented in multiple systems. Without matching:- Records are fragmented across operational, billing, and administrative systems
- Duplicate records create confusion and operational gaps
- Analytics undercount or overcount unique individuals
- Teams lack the complete information needed for effective coordination
What You Can Do
Configure Match Rules
Define which fields to compare and how much weight each field carries in match scoring.
Choose Match Algorithms
Select from exact matching or fuzzy matching approaches based on your data quality needs.
Review Match Candidates
Examine potential matches with detailed comparison views before accepting.
Set Match Thresholds
Configure score thresholds for auto-match, manual review, and non-match decisions.
Match Types
Each attribute in a rule is compared with a match type you choose, so you can be strict on identifiers and forgiving on names.Exact
String equality — the right choice for member IDs, SSNs, and email addresses that should be identical
High
Tolerates minor typos and transposed characters, such as “Jon Smith” against “John Smith”
Medium
Tolerates stronger variation, such as “Jon Smyth” against “John Smith”
Low
Broad fuzziness for badly degraded data, such as “Jno Smyth” against “Jonathan Smith”
Nickname
A curated dictionary of common nicknames, so “Bob” matches “Robert” and “Liz” matches “Elizabeth”
Date of Birth
Date matching with tolerance for two-digit year flips, transpositions, and month or day swaps
Two Ways to Configure Matching
Semantic Label Mode
Pick the labels to match on, choose a match type and priority for each, and set one overall confidence threshold. skyMDM applies the rules across every pair of tables.
Cascading Step Mode
Author a multi-step strategy for Member 360, Patient 360, and Resident 360. Step one might be an exact match on member ID; records it resolves are excluded from step two, which can fall back to a fuzzy match on name and date of birth. Each step carries its own rules and thresholds.
Match Attributes
Name Matching
Handle nicknames, misspellings, name changes, and cultural naming conventions
Address Matching
Match despite formatting differences, abbreviations, and address changes
ID Matching
Exact matching on SSN, account IDs, and other unique identifiers
Date Matching
Match dates across different formats with tolerance for data entry errors
Phone/Email
Match contact information with normalization and validation
Custom Fields
Include any field in your matching strategy based on your data
Key Capabilities
Match Scoring
Every potential match receives a confidence score from 0 to 100 percent, calculated as a weighted average of the per-attribute similarity scores. Higher-priority attributes carry more weight, and the score always shows which attributes agreed and how strongly, so no match is a black box.Null-Safe Scoring
An attribute missing on either side is left out of the average rather than counted as a mismatch. A three-attribute rule where one record has no phone number becomes a two-attribute comparison, so incomplete records are not unfairly penalized.The Review Band
Set an auto-merge threshold and, optionally, a review band beneath it. Pairs at or above the threshold merge automatically, pairs inside the band go to the Review Queue for a data steward to decide, and anything below is discarded. This is how you keep automation on the confident cases and human judgment on the rest.Configuration Checks
Saving a match configuration runs a check over it first and reports errors, warnings, and notes — duplicate step names, rules with no effect, unused semantic labels, thresholds that would leave an empty review band. Misconfigurations surface before the job runs, not after.Transitive Matching
If Record A matches Record B, and Record B matches Record C, skyMDM recognizes that all three may represent the same entity.Block and Compare
Efficiently process large datasets by first blocking records into candidate groups, then running detailed comparisons only within blocks.Persistent Cluster Identity
Clusters keep the same identifier across re-runs, so a person’s identity survives new data arriving. Re-running Match carries each cluster’s existing identity and its golden record forward instead of renumbering everything, and a full-rebuild mode is available when you deliberately want to start clean.Progress, Cancel, and Recovery
Match jobs report progress continuously in the run window, and a run can be cancelled from the same place. Previous runs stay available in the run history with the person who triggered each. A reconciler detects runs that stalled or lost their completion signal and heals the status automatically.Runs Where Your Data Lives
Matching executes natively on Databricks, Snowflake, or Microsoft Fabric, with the same rules, scores, and clusters on every platform.Match Audit Trail
Every match decision is logged with the score, contributing fields, and timestamp for compliance and quality review.Business Impact
Complete View
Link records across systems for a comprehensive view of each individual or entity.
Accurate Counts
Know exactly how many unique individuals or accounts you serve.
Better Coordination
Connect teams with complete information for better decision-making.
Who Benefits
- Operations Teams: Access complete records for planning and coordination
- Quality Teams: Accurate attribution for quality measures and reporting
- Finance Teams: Proper counting for revenue and cost analysis
- Compliance Teams: Maintain accurate records for regulatory reporting

