Data

gold, ingots, treasure, bullion, gold bars, wealth, gold, gold, gold, gold, gold

Advanced Analytics in the Gold Layer: Leveraging Window Functions for Insights and Synthetic Testing

In this post, we will explore how to use SQL window functions within the Gold layer to solve two common needs: calculating granular aggregates and generating synthetic test data for healthcare scenarios. Modern data engineering has converged on the Medallion Architecture as the gold standard for managing scalable data lakes. By organizing data into Bronze […]

Advanced Analytics in the Gold Layer: Leveraging Window Functions for Insights and Synthetic Testing Read More »

A white sportscar in motion showcasing speed and style on an open road.

AI Acceleration of Master Data Management

Executive Summary Master Data Management (MDM) is the discipline that enables organizations to maintain accurate, consistent, and trusted domain-specific master entities—customers, products, distributors, patients, and more—across all systems. It ensures enterprise operations and analytics are based on a single source of truth. Historically, MDM has been slow and resource-intensive, creating latency between data capture and

AI Acceleration of Master Data Management Read More »

Databricks vs Azure Synapse Analytics: Understanding the Differences for Smarter Data Platform Choices

As a data engineer navigating cloud platforms and analytics ecosystems, I’ve worked extensively with both Databricks and Azure Synapse Analytics. While they often appear side-by-side in Azure environments, they come from different design philosophies and cater to slightly different needs—even when they seem to do many of the same things. Both platforms allow you to

Databricks vs Azure Synapse Analytics: Understanding the Differences for Smarter Data Platform Choices Read More »

recipe, tab, index, cards, dividers, print, food, book, pages, cookies, candy, seafood, beverage, bread, soup, salad, gray food, gray book, gray books, gray bread, gray candy, tab, index, index, index, index, index

The Hidden Cost of Overloaded Data Fields — And How Data Governance Saves the Day

When One Field Tries to Do Too Much Across many organizations, the same pattern repeats. A single data field—perhaps introduced during a system rollout or legacy data migration—is intended to serve a focused purpose. Over time, however, teams start using it to represent different ideas. Marketing redefines the field to support campaigns. Sales uses it

The Hidden Cost of Overloaded Data Fields — And How Data Governance Saves the Day Read More »

Copy Activity vs Data flows in Azure Data Factory: A Practical Look at Two Competing Features

When building cloud data pipelines with Azure Data Factory (ADF), you’re often presented with two paths: use the tried-and-true Copy Activity, or lean into the newer, powerful Mapping Data flows. Both are native ADF features and technically serve the same purpose—moving and transforming data—but the similarities end there. In our case, we started with Data

Copy Activity vs Data flows in Azure Data Factory: A Practical Look at Two Competing Features Read More »

Using SQL UNPIVOT to Unlock Healthcare Provider Insights with Unpivoted Taxonomy Data

I worked many years with the NPI dataset. Transforming the data brought actionable insight. Imagine a conversation between an enterprise IT strategist, a healthcare IT consultant, and a healthcare administrator. The three sit around a conference table reviewing a dataset pulled from the National Plan and Provider Enumeration System (NPPES), full of NPIs, taxonomy codes,

Using SQL UNPIVOT to Unlock Healthcare Provider Insights with Unpivoted Taxonomy Data Read More »

Cardboard boxes labeled keep, donate, and trash for effective home organization.

Dealing with Duplicates: A Practical SQL-Based Deduplication System

Many organizations struggle with duplicate customer sites and accounts scattered across operational systems. These duplicates introduce confusion in invoicing, shipping, reporting, and analytics. The immediate need to resolve this data fragmentation often outpaces the availability of enterprise-wide tooling. This post outlines how we built a pragmatic, SQL-based deduplication framework to identify, cluster, and prepare duplicate

Dealing with Duplicates: A Practical SQL-Based Deduplication System Read More »