Skip to main content

Our Core Technology is now a Granted US Patent (opens in a new tab)

Menu

The Enterprise Semi-Structured Data Assurance Layer.

DataPancake® is a patented structural assurance layer for Snowflake that discovers complex source structure before transformation, captures it as configurable metadata, and generates complete, governed, relational outputs that stay aligned as schemas evolve.

“We looked everywhere... Nothing comes close to what DataPancake delivers.”


C.H. Robinson (opens in a new tab)

$1M+ saved in 6 months

Most Snowflake customers struggle with this issue, and almost all of our SEs have had to deal with it. DataPancake saves data engineers weeks of time, money, and pain.


Cameron Wasilewsky, Technical Lead
Snowflake Startup Accelerator
Powered by Snowflake
Proven at Enterprise Scale

Accelerate Structural Discovery with DataPancake

DataPancake helps teams discover complete source structure, reduce manual pipeline effort, and improve downstream data quality across complex enterprise environments.

10 Minutes

Avg time to scan and discover the complete schema for 1 billion rows.

10x ROI

Demonstrated across lower compute, reduced engineering time, and increased data quality.

10+ Industries

Financial services, healthcare and life sciences, public sector, manufacturing, retail, logistics, energy, and more.

Approved by InfoSec Teams

Built for Sensitive Data Environments Across Industries

DataPancake is designed for Snowflake-centered execution, governed metadata, and security-aware output generation across regulated and operational data teams.

Financial Services
Healthcare
Transportation
Public Sector
Logistics / Supply Chain

Configurable Security Metadata

Configurable and reportable metadata drives row access, column masking, secure views, and policy-aware output generation.

Snowflake-Centered Processing

Data and processing stay aligned with the customer’s Snowflake environment, reducing external processing risk.

Structural assurance is only meaningful if it's verifiable, configurable, and maintainable at enterprise scale.

Every capability below is designed with that constraint in mind, not as a feature added to a tool, but as a component of a complete assurance architecture.

When “Close Enough” Is Not Good Enough

DataPancake is the only patented technology designed to take you from raw, complex, and deeply nested semi-structured data to accurately flattened, normalized, enriched, secure, documented, and GenAI-ready relational Dynamic Tables and Views, with zero technical debt in minutes, not months.

Step 01

Data Loader

Load, stage, inspect, and verify source data.

Discover, Configure, Shape, Generate

Step 02

Schema Discovery

Discover paths, arrays, objects, and type variation.

Step 03

Pipeline Design

Configure relationships and output behavior.

Step 04

Schema Shaping

Refine structure before generation and downstream use.

Step 05

Code Generation

Generate SQL, Dynamic Tables, and Secure Views.

Continuous Control Layer

Schema Drift Monitoring

Monitor new paths, changed types, and structural variation as source data evolves.

Document + Activate

Step 06

Data Dictionary

Document attributes, definitions, and generated structures.

Step 07

Semantic Model

Produce AI-ready semantic context for analytics and AI.

Detailed Workflow

From Raw Source Data to Governed, AI-Ready Snowflake Assets

DataPancake turns semi-structured complexity into a controlled, metadata-driven workflow: load and verify the source, discover the complete structure, configure how it should behave, shape the model, generate governed SQL, monitor drift, document the result, and activate semantic context.

Data Loader

Load and validate complex source files before flattening, normalization, or activation decisions are made. Data Loader helps teams understand what is present, what is malformed, and what needs to be inspected before structural discovery begins.

Pre-flight validation

Stage, chunk, inspect, and validate source data before modeling decisions are made.

Inventory visibility

Capture file-level and source-level statistics that help teams understand completeness and structure.

Exception handling

Surface malformed records, unexpected source behavior, and validation issues early in the workflow.

Discovery-ready inputs

Prepare source data for complete structural discovery, schema shaping, and downstream generation.

Schema Discovery

Recursively scan the full semi-structured source population to discover nested arrays, objects, polymorphic attributes, escaped JSON, and every observed version of each attribute before downstream logic is written.

Complete structural inventory

Discover nested paths, arrays, objects, sparse fields, and low-frequency attributes across the full source.

Polymorphic type detection

Detect primitive, array, object, and mixed type variations that conventional inference can miss.

Stringified JSON discovery

Identify escaped or embedded JSON within string fields so hidden structure can be modeled correctly.

Snowflake type inference

Infer destination data types, including datetime formats, to support accurate generated outputs.

Pipeline Design

Configure how each generated SQL DDL pipeline should behave before code is created. Pipeline Design turns discovered metadata into reviewable decisions about relationships, transformations, governance, and output structure.

Relationship configuration

Configure foreign key relationships for nested arrays and post-normalized table joins.

Column-level logic

Apply column-level transformation logic during the materialization process.

Virtual attributes

Create derived fields, semantic metrics, facts, filters, and other configured virtual attributes.

Governance configuration

Configure row access, column masking, secure views, and policy-aware generated outputs.

Schema Shaping

Refine complex source structure before generation. Schema Shaping lets teams apply targeted transformation decisions to specific discovered paths before those decisions are reflected in generated Snowflake outputs.

Path-targeted transformations

Apply transformations to specific source paths so structural refinement can happen before SQL generation and downstream use.

SQL Code Generation

Generate SQL DDL to materialize operationalized sources into relational Dynamic Tables and policy-aware views in Snowflake, based on the configured metadata model.

Dynamic Table SQL

Generate Snowflake Dynamic Table SQL DDL using DataPancake ITDCs™ for durable pipelines.

ITDC: Immutably Typed Derived Columns.

Secure views

Generate views that incorporate row access, column masking, and additional transformations.

Configured relationships

Reflect configured transformations, foreign keys, and virtual attributes in generated outputs.

Operational metadata

Generate supporting metadata structures to track Dynamic Table processing and updates.

Schema Drift Monitoring

Monitor structural change as source systems evolve. DataPancake helps teams detect new attributes, changed data types, structural variation, and schema drift before downstream assumptions break.

Continuous monitoring

Monitor and alert when semi-structured data source schemas change over time.

Type and structure changes

Flag changes in data types, structures, nested paths, and newly discovered attributes.

Review before regeneration

Give teams a controlled way to configure new attributes before regenerating pipeline code.

ITDC evolution

Generate new immutably typed derived columns as source variation changes without accumulating unnecessary debt.

Data Dictionary Builder

Generate a governed data dictionary from discovered and configured metadata, including definitions, synonyms, sample values, datasource descriptions, nested array descriptions, and attribute-level documentation.

Attribute documentation

Create definitions, synonyms, and sample values for discovered and generated attributes.

LLM-assisted dictionary building

Use preferred LLM inside Snowflake to assist with dictionary building.

Custom GenAI Prompt Context

Extend our system prompt with your own custom context for greater clarity and improved responses.

Semantic model input

Feed semantic generation with governed documentation aligned to the generated Snowflake model.

Semantic Model Generator

Generate Cortex Analyst semantic for semi-structured data sources flattened and normalized by DataPancake.

Cortex Analyst-ready YAML

Generate semantic model files that define the complete governed model for natural language analysis.

Relationship metadata

Automatically include relationship metadata based on selected and configured columns.

Virtual metrics and filters

Use virtual attributes to configure custom metrics, facts, filters, and semantic behavior.

Verified queries and instructions

Add verified queries and custom instructions to improve semantic quality and business usability.

Streaming Data

Pancake Your Streaming Data With Ease

Apache KafkaTopics streamed into Snowflake

DataPancake includes native support for Kafka topics streamed into Snowflake, giving teams control over complex, high-volume data in motion without adding unnecessary pipeline overhead.

Discovery of Stringified JSON

Accurate schema discovery of complex escaped message strings, so embedded structure can be discovered before downstream generation.

Deduplication Logic

Includes configurable logic to materialize the most recent message per message key.

Flattened Message Metadata

Normalized Dynamic Tables and views can include flattened Kafka-specific metadata, such as topic, partition, offset, and timestamp.

Incremental Scanning

Recurring scans use the last processed message timestamp for incremental data scanning, reducing compute cost.

Blog

The art and science of nested, polymorphic JSON transformation in Snowflake

Read how we leveraged the Snowflake Native App framework to help you dynamically transform JSON data into actionable insights in a fraction of the usual time.

Read Article (opens in a new tab)

Semi-Structured Data Pipelines

How to Mitigate Risk and Accelerate Innovation with Snowflake

Explore the unknown challenges that contribute to downstream impact and risk, and how to identify and solve them with DataPancake from TDAA inside of Snowflake.

Read Article (opens in a new tab)

“DataPancake has proven to be a superior tool. Its adoption has lead to substantial savings in engineering resources, cloud compute costs, and many other areas and has accelerated our adoption of GenAI.”


C.H. Robinson (opens in a new tab)

Scott Hamilton
Data Engineering Manager

Ready to make your most valuable semi-structure data analytics and AI ready in days, not weeks or months?