The Enterprise Semi-Structured Data Assurance Layer.
DataPancake® is a patented structural assurance layer for Snowflake that discovers complex source structure before transformation, captures it as configurable metadata, and generates complete, governed, relational outputs that stay aligned as schemas evolve.
“We looked everywhere... Nothing comes close to what DataPancake delivers.”
$1M+ saved in 6 months
Most Snowflake customers struggle with this issue, and almost all of our SEs have had to deal with it. DataPancake saves data engineers weeks of time, money, and pain.
Snowflake Startup Accelerator

Accelerate Structural Discovery with DataPancake
DataPancake helps teams discover complete source structure, reduce manual pipeline effort, and improve downstream data quality across complex enterprise environments.
10 Minutes
Avg time to scan and discover the complete schema for 1 billion rows.
10x ROI
Demonstrated across lower compute, reduced engineering time, and increased data quality.
10+ Industries
Financial services, healthcare and life sciences, public sector, manufacturing, retail, logistics, energy, and more.
Built for Sensitive Data Environments Across Industries
DataPancake is designed for Snowflake-centered execution, governed metadata, and security-aware output generation across regulated and operational data teams.
Configurable Security Metadata
Configurable and reportable metadata drives row access, column masking, secure views, and policy-aware output generation.
Snowflake-Centered Processing
Data and processing stay aligned with the customer’s Snowflake environment, reducing external processing risk.
Structural assurance is only meaningful if it's verifiable, configurable, and maintainable at enterprise scale.
Every capability below is designed with that constraint in mind, not as a feature added to a tool, but as a component of a complete assurance architecture.
When “Close Enough” Is Not Good Enough
DataPancake is the only patented technology designed to take you from raw, complex, and deeply nested semi-structured data to accurately flattened, normalized, enriched, secure, documented, and GenAI-ready relational Dynamic Tables and Views, with zero technical debt in minutes, not months.
Step 01
Data Loader
Load, stage, inspect, and verify source data.
Step 02
Schema Discovery
Discover paths, arrays, objects, and type variation.
Step 03
Pipeline Design
Configure relationships and output behavior.
Step 04
Schema Shaping
Refine structure before generation and downstream use.
Step 05
Code Generation
Generate SQL, Dynamic Tables, and Secure Views.
Continuous Control Layer
Schema Drift Monitoring
Monitor new paths, changed types, and structural variation as source data evolves.
Step 06
Data Dictionary
Document attributes, definitions, and generated structures.
Step 07
Semantic Model
Produce AI-ready semantic context for analytics and AI.
From Raw Source Data to Governed, AI-Ready Snowflake Assets
DataPancake turns semi-structured complexity into a controlled, metadata-driven workflow: load and verify the source, discover the complete structure, configure how it should behave, shape the model, generate governed SQL, monitor drift, document the result, and activate semantic context.
Data Loader
Load and validate complex source files before flattening, normalization, or activation decisions are made. Data Loader helps teams understand what is present, what is malformed, and what needs to be inspected before structural discovery begins.
Pre-flight validation
Stage, chunk, inspect, and validate source data before modeling decisions are made.
Inventory visibility
Capture file-level and source-level statistics that help teams understand completeness and structure.
Exception handling
Surface malformed records, unexpected source behavior, and validation issues early in the workflow.
Discovery-ready inputs
Prepare source data for complete structural discovery, schema shaping, and downstream generation.
Schema Discovery
Recursively scan the full semi-structured source population to discover nested arrays, objects, polymorphic attributes, escaped JSON, and every observed version of each attribute before downstream logic is written.
Complete structural inventory
Discover nested paths, arrays, objects, sparse fields, and low-frequency attributes across the full source.
Polymorphic type detection
Detect primitive, array, object, and mixed type variations that conventional inference can miss.
Stringified JSON discovery
Identify escaped or embedded JSON within string fields so hidden structure can be modeled correctly.
Snowflake type inference
Infer destination data types, including datetime formats, to support accurate generated outputs.
Pipeline Design
Configure how each generated SQL DDL pipeline should behave before code is created. Pipeline Design turns discovered metadata into reviewable decisions about relationships, transformations, governance, and output structure.
Relationship configuration
Configure foreign key relationships for nested arrays and post-normalized table joins.
Column-level logic
Apply column-level transformation logic during the materialization process.
Virtual attributes
Create derived fields, semantic metrics, facts, filters, and other configured virtual attributes.
Governance configuration
Configure row access, column masking, secure views, and policy-aware generated outputs.
Schema Shaping
Refine complex source structure before generation. Schema Shaping lets teams apply targeted transformation decisions to specific discovered paths before those decisions are reflected in generated Snowflake outputs.
Path-targeted transformations
Apply transformations to specific source paths so structural refinement can happen before SQL generation and downstream use.
SQL Code Generation
Generate SQL DDL to materialize operationalized sources into relational Dynamic Tables and policy-aware views in Snowflake, based on the configured metadata model.
Dynamic Table SQL
Generate Snowflake Dynamic Table SQL DDL using DataPancake ITDCs™ for durable pipelines.
ITDC: Immutably Typed Derived Columns.
Secure views
Generate views that incorporate row access, column masking, and additional transformations.
Configured relationships
Reflect configured transformations, foreign keys, and virtual attributes in generated outputs.
Operational metadata
Generate supporting metadata structures to track Dynamic Table processing and updates.
Schema Drift Monitoring
Monitor structural change as source systems evolve. DataPancake helps teams detect new attributes, changed data types, structural variation, and schema drift before downstream assumptions break.
Continuous monitoring
Monitor and alert when semi-structured data source schemas change over time.
Type and structure changes
Flag changes in data types, structures, nested paths, and newly discovered attributes.
Review before regeneration
Give teams a controlled way to configure new attributes before regenerating pipeline code.
ITDC evolution
Generate new immutably typed derived columns as source variation changes without accumulating unnecessary debt.
Data Dictionary Builder
Generate a governed data dictionary from discovered and configured metadata, including definitions, synonyms, sample values, datasource descriptions, nested array descriptions, and attribute-level documentation.
Attribute documentation
Create definitions, synonyms, and sample values for discovered and generated attributes.
LLM-assisted dictionary building
Use preferred LLM inside Snowflake to assist with dictionary building.
Custom GenAI Prompt Context
Extend our system prompt with your own custom context for greater clarity and improved responses.
Semantic model input
Feed semantic generation with governed documentation aligned to the generated Snowflake model.
Semantic Model Generator
Generate Cortex Analyst semantic for semi-structured data sources flattened and normalized by DataPancake.
Cortex Analyst-ready YAML
Generate semantic model files that define the complete governed model for natural language analysis.
Relationship metadata
Automatically include relationship metadata based on selected and configured columns.
Virtual metrics and filters
Use virtual attributes to configure custom metrics, facts, filters, and semantic behavior.
Verified queries and instructions
Add verified queries and custom instructions to improve semantic quality and business usability.
Pancake Your Streaming Data With Ease
DataPancake includes native support for Kafka topics streamed into Snowflake, giving teams control over complex, high-volume data in motion without adding unnecessary pipeline overhead.
Discovery of Stringified JSON
Accurate schema discovery of complex escaped message strings, so embedded structure can be discovered before downstream generation.
Deduplication Logic
Includes configurable logic to materialize the most recent message per message key.
Flattened Message Metadata
Normalized Dynamic Tables and views can include flattened Kafka-specific metadata, such as topic, partition, offset, and timestamp.
Incremental Scanning
Recurring scans use the last processed message timestamp for incremental data scanning, reducing compute cost.
Blog
The art and science of nested, polymorphic JSON transformation in Snowflake
Read how we leveraged the Snowflake Native App framework to help you dynamically transform JSON data into actionable insights in a fraction of the usual time.
Read Article (opens in a new tab)Semi-Structured Data Pipelines
How to Mitigate Risk and Accelerate Innovation with Snowflake
Explore the unknown challenges that contribute to downstream impact and risk, and how to identify and solve them with DataPancake from TDAA inside of Snowflake.
Read Article (opens in a new tab)“DataPancake has proven to be a superior tool. Its adoption has lead to substantial savings in engineering resources, cloud compute costs, and many other areas and has accelerated our adoption of GenAI.”
Scott Hamilton
Data Engineering Manager