Skip to main content

Clean

Transform raw data into reliable, standardized information ready for matching and analysis. The Clean module applies intelligent data cleansing rules that fix inconsistencies and standardize formats at scale, so downstream matching compares like with like.

Why Data Cleaning Matters

Healthcare data comes from many sources, each with its own formats, conventions, and quality issues. Patient names might be all uppercase in one system and mixed case in another. Phone numbers might include dashes, parentheses, or no formatting at all. These inconsistencies create problems:
  • Duplicate patient records that fragment care history
  • Failed matching that misses related records
  • Inaccurate analytics and reporting
  • Compliance risks from poor data quality
skyMDM’s Clean module addresses these challenges with rule-based, scalable data transformation.

What You Can Do

Apply Cleaning Rules

Configure rules to trim whitespace, standardize case, format phone numbers, and more.

Turn Rules On and Off

Disable a rule to take it out of the next run without losing its configuration, then switch it back on later.

Chain Multiple Operations

Apply multiple cleaning operations in sequence for comprehensive data standardization.

Track Execution

Monitor cleaning jobs in real-time with progress updates and detailed logs.

Clean Several Sources at Once

Pick multiple source tables for a single run and give each one its own destination table.

Inspect Rule Results

Expand any rule in the job view to see the columns it touched and how many rows it changed.

Cleaning Operations

skyMDM supports a comprehensive set of cleaning operations:

Text Normalization

Trim whitespace, convert case (upper, lower, title), remove special characters

Phone Standardization

Strip formatting characters, handle a leading country code, and reformat to a consistent pattern

Date Standardization

Interpret a wide range of input formats and write a single standard date format

Name Standardization

Title-case names and collapse stray whitespace, including hyphenated and all-caps entries

Address Standardization

Title-case street addresses and expand common street-type and directional abbreviations

Value Replacement

Replace specific values, fill nulls with a default, and apply custom regular expressions

Key Capabilities

Rule-Based Configuration

Define cleaning rules visually without writing code. Select columns, choose operations, and set parameters through an intuitive interface.

Scalable Processing

Cleaning jobs run natively on the platform that holds your data — Spark on Databricks, Snowpark on Snowflake, or OneLake on Microsoft Fabric — processing millions of records without moving them out of your environment.

Rule Conflict Detection

skyMDM checks a rule set before it runs. Two conflicting case transformations on the same column, or a duplicate rule already covering that column, are flagged as you configure them rather than producing a surprising result at run time.

Run History and Recovery

A live run window reports progress, elapsed time, and row counts while a job executes, and keeps the history of previous runs for the table. If a run fails or its completion signal is lost, skyMDM reconciles the job status automatically so the stage never sits stuck in “running”.

Column Profiling

After cleaning completes, skyMDM automatically profiles your data—showing null counts, unique values, and data types for each column.

Business Impact

Higher Match Rates

Standardized data dramatically improves identity matching accuracy.

Reduced Manual Work

Automate repetitive data cleanup that would take analysts weeks.

Trusted Analytics

Clean data produces reliable reports and insights for decision-making.

Who Benefits

  • Data Stewards: Maintain data quality standards across the organization
  • Clinical Teams: Access accurate patient information for better care decisions
  • Revenue Cycle: Reduce claim denials caused by data quality issues
  • Compliance Officers: Meet regulatory requirements for data accuracy