Clean
Transform raw data into reliable, standardized information ready for matching and analysis. The Clean module applies intelligent data cleansing rules that fix inconsistencies and standardize formats at scale, so downstream matching compares like with like.Why Data Cleaning Matters
Healthcare data comes from many sources, each with its own formats, conventions, and quality issues. Patient names might be all uppercase in one system and mixed case in another. Phone numbers might include dashes, parentheses, or no formatting at all. These inconsistencies create problems:- Duplicate patient records that fragment care history
- Failed matching that misses related records
- Inaccurate analytics and reporting
- Compliance risks from poor data quality
What You Can Do
Apply Cleaning Rules
Configure rules to trim whitespace, standardize case, format phone numbers, and more.
Turn Rules On and Off
Disable a rule to take it out of the next run without losing its configuration, then switch it back on later.
Chain Multiple Operations
Apply multiple cleaning operations in sequence for comprehensive data standardization.
Track Execution
Monitor cleaning jobs in real-time with progress updates and detailed logs.
Clean Several Sources at Once
Pick multiple source tables for a single run and give each one its own destination table.
Inspect Rule Results
Expand any rule in the job view to see the columns it touched and how many rows it changed.
Cleaning Operations
skyMDM supports a comprehensive set of cleaning operations:Text Normalization
Trim whitespace, convert case (upper, lower, title), remove special characters
Phone Standardization
Strip formatting characters, handle a leading country code, and reformat to a consistent pattern
Date Standardization
Interpret a wide range of input formats and write a single standard date format
Name Standardization
Title-case names and collapse stray whitespace, including hyphenated and all-caps entries
Address Standardization
Title-case street addresses and expand common street-type and directional abbreviations
Value Replacement
Replace specific values, fill nulls with a default, and apply custom regular expressions
Key Capabilities
Rule-Based Configuration
Define cleaning rules visually without writing code. Select columns, choose operations, and set parameters through an intuitive interface.Scalable Processing
Cleaning jobs run natively on the platform that holds your data — Spark on Databricks, Snowpark on Snowflake, or OneLake on Microsoft Fabric — processing millions of records without moving them out of your environment.Rule Conflict Detection
skyMDM checks a rule set before it runs. Two conflicting case transformations on the same column, or a duplicate rule already covering that column, are flagged as you configure them rather than producing a surprising result at run time.Run History and Recovery
A live run window reports progress, elapsed time, and row counts while a job executes, and keeps the history of previous runs for the table. If a run fails or its completion signal is lost, skyMDM reconciles the job status automatically so the stage never sits stuck in “running”.Column Profiling
After cleaning completes, skyMDM automatically profiles your data—showing null counts, unique values, and data types for each column.Business Impact
Higher Match Rates
Standardized data dramatically improves identity matching accuracy.
Reduced Manual Work
Automate repetitive data cleanup that would take analysts weeks.
Trusted Analytics
Clean data produces reliable reports and insights for decision-making.
Who Benefits
- Data Stewards: Maintain data quality standards across the organization
- Clinical Teams: Access accurate patient information for better care decisions
- Revenue Cycle: Reduce claim denials caused by data quality issues
- Compliance Officers: Meet regulatory requirements for data accuracy

