Blog Posts

Snowflake Salesforce Integration: How to Connect Salesforce and Snowflake

The vast majority of Salesforce data is heavily underutilized, being stuck inside the customer relationship management (CRM) platform with limited querying options, problematic joining with other sources, and restricted via API limits that make most large-scale analyses impractical. Moving that Salesforce data into Snowflake allows it to be used for advanced analytics among other use cases.

However, the integration is rarely as straightforward as it might seem at first. Synchronization, schema drift, deleted records, and data connector pricing models are all issues that only appear after the integration decision has been made. 

In this guide, we aim to cover how the Snowflake Salesforce integration works, what challenges to expect, and how to evaluate the connectors and architectural options available on the market in order to help engineering and data teams make the right call before committing to a specific approach.

Table of Contents

How does data flow from Salesforce to Snowflake?

Data transfer from Salesforce to Snowflake follows a structured Extract, Transform, Load process (ETL). The Snowflake Salesforce integration utilizes Salesforce APIs to pull records that are then staged and loaded into the data warehouse.

What challenges arise when moving data from Salesforce into Snowflake?

The Snowflake Salesforce integration introduces several technical challenges that teams need to account for before building or selecting a pipeline:

  • API governor limits — Salesforce enforces daily API call limits which vary by edition; high-volume syncs can exhaust these limits and stall pipelines mid-run
  • Deleted and merged records — Salesforce does not surface deleted records in standard queries; capturing hard deletes requires querying the recycle bin or using the Bulk API with specific parameters
  • Field type mismatches — Salesforce data types (such as picklists, formula fields, and multi-select fields) do not map cleanly to Snowflake column types and require transformation logic
  • Schema changes — Salesforce admins can add, rename, or remove custom fields at any time, which causes schema drift that breaks downstream queries if not handled automatically
  • Compound and encrypted fields — certain Salesforce fields, such as address compounds and Shield-encrypted fields, require special handling before they can be loaded into a data warehouse

Should data flow from Salesforce to Snowflake, from Snowflake to Salesforce, or in both directions?

In most implementations, the data flows in one direction: from Salesforce to Snowflake. Salesforce remains the source of truth for CRM activity, and Snowflake is just an analytical layer for storing and accessing that information. 

Because this is a one-way data flow, it’s extremely easy to engineer as well as maintain visibility into it. It even eliminates the chance for one of the two systems having a different opinion on which version of a record is the correct one.

Snowflake-to-Salesforce flow is not as common but it’s by no means absent, though it typically registers as reverse ETL instead of a full-blown pipeline in the other direction. What often goes back isn’t raw Salesforce data but its analytically useful interpretation: a lead-score calculated in Snowflake, a churn-risk flag, a product-usage summary, or a rollup metric that a sales rep needs. The actual tools making these movements possible (Census, Hightouch, Hevo) simply synchronize specific fields on a schedule without mirroring entire tables every time.

Actual bidirectional data synchronization where both systems read/write constantly is extremely rare and is rarely used unless it’s absolutely necessary for a specific business goal. Sending data back into Salesforce can lead to the overwriting of user-entered data and a chain of unrelated workflow automations. If sending information back is unavoidable, narrow scope can keep the consequences manageable.

How does a data warehouse architecture work with Salesforce and Snowflake?

A data warehouse architecture incorporating Salesforce and Snowflake separates the concerns of data capture and data analysis across two purpose-built environments. That way, operational CRM activity is handled by Salesforce, while Snowflake database works as the analytical layer where that data is stored, modeled, and queried.

How does Salesforce data cloud differ from a traditional data warehouse?

Salesforce Data Cloud is Salesforce’s own customer data platform — not a general-purpose data warehouse. The distinction between the two is important to know about before committing to an integration architecture.

DimensionSalesforce Data CloudSnowflake (Data Warehouse)
Primary purposeUnify customer profiles within SalesforceStore and analyze data from any source
Query languageSalesforce-native (limited SQL)Full ANSI SQL
Data scopeCustomer and engagement dataAny structured or semi-structured data
BI tool supportLimited to Salesforce ecosystemBroad — Tableau, Looker, Power BI, etc.
Cost modelSalesforce licensingCompute and storage consumption

The primary purpose of Data Cloud is to enrich Salesforce workflows, not replace a warehouse. Organizations in need of cross-functional analysis (a combination of CRM data and finance, product, or operational data) are still going to have to use Snowflake as their primary analytical platform.

Salesforce Data Cloud vs Snowflake: When Should You Use Data Movement, Federation, or Zero-Copy Access?

Those three approaches are all part of the spectrum of coupling between Salesforce and Snowflake. The most prominent factors between them are answers to two questions:

  • How much data actually needs to move?
  • How tightly the two systems stay in sync?

Data movement creates an independent copy of business data in the warehouse by replicating Salesforce records into Snowflake using an ETL or ELT pipeline. It involves keeping a discrete replica of your data within your warehouse permanently. If you have a Snowflake Salesforce integration set up and wish to perform full SQL queries and retain data historically that needs to be joined with data pulled from other systems — it makes sense to use this default strategy for Snowflake Salesforce integration.

Federation keeps data at the source and queries it on the fly without copying: such as exposing Snowflake tables as external objects within Salesforce, or querying Salesforce data directly from Snowflake without the use of a staging layer. Federation can be a good approach for when freshness is more important than speed and duplication of storage is not a major concern.

Zero-copy access uses Snowflake’s native data sharing capability to provide any other Snowflake account read access to dataset as if it were local so no copying is needed. This works best for data modeled in Snowflake and then shared to partners or other departments that exist in separate Snowflake accounts. That said, this method doesn’t include the Salesforce-to-Snowflake leg itself because Salesforce isn’t a Snowflake account.

The default Salesforce to Snowflake integration in most cases is data movement. Both federation and zero-copy access are targeting specific use cases like data freshness or cross-account distribution, but they cannot replace the standard pipeline.

Where Does GRAX Fit in a Snowflake-Salesforce Data Integration Strategy?

The majority of connector types explored in this article involve moving current-state Salesforce data into Snowflake with some degree of regularity. GRAX approaches this issue in a completely different way. It’s a Salesforce-native platform developed from the ground up and focused primarily on archiving, backup, and historical preservation of data, with the Snowflake integration being one of many means to make archived information usable for analytics.

And the difference is crucial because Salesforce by design doesn’t keep a comprehensive history of every change made. Previous updates can be replaced by newer entries during field updates while removed records simply cease to be discoverable by normal queries. GRAX automatically records that entire history, with every version of every record, to store it in an independent storage. The last point is particularly important because it avoids Salesforce’s storage limits and API access constraints simultaneously.

For a Snowflake Salesforce integration, this creates a few specific use cases:

  • Querying what a record looked like on a specific past data for point-in-time analysis.
  • Keeping a complete change history for compliance and audit retention purposes.
  • Assessing and recovering records that were permanently removed from Salesforce.
  • Reducing Salesforce API load by maintaining and processing a separate data copy.

GRAX usually positions as complementary to a common sync connector, not as a replacement. When comparing it against the attributes we described earlier, the relevant questions shift slightly towards retention depth, recovery guarantees, and how cleanly the archived history meshes with your existing data in Snowflake.

Sync moves data. It doesn’t keep history.

See how GRAX preserves what standard syncs overwrite.

Learn More

How does data sync between Salesforce and Snowflake work?

The synchronization process allows Snowflake to stay up-to-date with what is going on in Salesforce. The Snowflake Salesforce integration offers different synchronization types to choose from, each with different trade-offs in terms of latency, cost, and complexity.

What is the difference between batch and real-time synchronization?

Batch and real-time data sync are two fundamentally different approaches to transferring data to Snowflake from Salesforce. The best option for a specific company depends on how fresh the data has to be and what the downstream use cases actually need.

DimensionBatch SyncReal-Time Sync
How it worksExtracts records on a fixed schedule (hourly, daily)Streams changes as they occur using Change Data Capture or webhooks
LatencyMinutes to hoursSeconds to minutes
ComplexityLower — simpler pipelines, easier to debugHigher — requires streaming infrastructure
CostGenerally lowerGenerally higher
Best forReporting, historical analysis, overnight dashboardsOperational use cases, live dashboards, alerts
Salesforce API impactConcentrated API usage during sync windowsDistributed but continuous API consumption

How Do Salesforce API Limits Affect Continuous Data Ingestion into Snowflake?

Salesforce API access limits are refreshed once every 24 hours, with the limits themselves scaling depending on Salesforce edition and the number of licenses in an org. Meaning, a sync occurring smoothly in one org may deplete the daily allocation limits in another org that only has half as many licenses. 

Continuous and near-real-time ingestion are the ones usually creating a high amount of API calls compared to batch data syncing that is condensed into one-two runs in a day. The solution isn’t so much about lowering the sync frequency but more about picking an appropriate API for the tasks at hand. 

REST API would consume one API call per request, irrespective of how many records that request returns; not a good solution for high-volume continuous extraction. 

The Bulk API handles large record sets with fewer, larger API calls while Change Data Capture publishes record-level changes as events instead of asking the pipeline to constantly poll for updates. Both these options consume daily allocation limits far more effectively than REST API used for continuous polling.

Which Tools Can Connect Snowflake to Salesforce Reliably at Scale?

The Snowflake Salesforce integration is properly represented by a range of purpose-built solutions that can be separated into two broad categories:

  • Managed connectors that handle extraction, loading, and schema management out of the box
  • Transformation-focused tools that assume there already is a connector, focusing on modeling data once it has already been transferred into Snowflake

The same trade-offs discussed earlier, like cost or scalability, are still applicable when choosing a specific tool for connecting Salesforce to Snowflake together. 

How does data sharing work between Salesforce and Snowflake?

Information sharing in the context of Snowflake Salesforce integration cannot be classified as bidirectional data sharing — it typically refers to the controlled data flow from Salesforce into Snowflake, with the latter making the records from the former becoming available for cross-functional analysis alongside other data sources.

The mechanism behind that sharing (be it a managed connector, a custom pipeline, or the native sharing feature of Snowflake) determines how up-to-date, reliable, and accessible that that is going to be for downstream customers.

How Does Salesforce Connect Expose Snowflake Data as External Objects Without Copying It?

When using Salesforce Connect, you can see data stored outside the Salesforce database in your Salesforce UI without importing or copying said data beforehand, as a standard Salesforce object. The concept of an external object allows this data to appear almost like a regular Salesforce object in the UI: it’ll show up in list views, related lists, and reports but pulls its data in real-time from an external source at query time.

Linking these objects to Snowflake requires an intermediary, as Salesforce Connect doesn’t directly communicate to a data warehouse. A common configuration will be an OData adapter sitting in front of Snowflake and translating requests from Salesforce into queries that Snowflake can execute while returning the results in a format Salesforce Connect understands. There are a couple of points that can be derived from that design:

  • Data displayed through external objects is always up-to-date.
  • Write support is limited compared to native objects and relies on what the OData layer uncovers.
  • Query performance directly depends on Snowflake’s response time.
  • Field-level security and sharing rules still apply through the adapter layer.

This pattern works best when you want to expose Snowflake-modeled data to Salesforce users without creating an entire reverse-sync pipeline. At its core, this is a viewing mechanism that can’t replace the write-back patterns covered earlier in the article.

How does Snowflake data sharing differ from traditional pipelines?

Traditional pipelines extract data from a source, transform it, and then load its copy into a destination. Snowflake’s native data sharing capability operates differently — granting another Snowflake account read access to data that lives in the original one without the need to create a physical copy of said information. The shared data remains in one place and is always up-to-date, eliminating the possibility of sync lag that traditional pipelines are known to have.

DimensionTraditional PipelineSnowflake Native Data Sharing
Data movementData is copied to destinationNo copy — access is granted to original data
LatencyDepends on sync frequencyAlways current
Infrastructure requiredETL tools, schedulers, monitoringNone — managed within Snowflake
CostStorage duplicated across systemsSingle storage location, consumer pays compute
Use caseInternal analytics, transformationCross-account or cross-organization data access

Traditional pipeline tools remain the primary method for most modern Salesforce to Snowflake workflows. Native Snowflake data sharing only becomes relevant when processed or modeled Salesforce data has to be distributed to partners, subsidiaries, or other internal Snowflake accounts without the prerequisite of rebuilding pipelines from scratch for each new consumer.

Can Salesforce Users View, Edit, Report On, and Automate Snowflake Data Without Leaving the CRM?

Viewing Snowflake data from within Salesforce is the strongest of the four capabilities. Once Salesforce Connect is configured with external objects, users can view Snowflake records in list views and related lists just as if they were records in Salesforce itself. No separate login. No export process.

Editing has its limitations. The possibility of the user writing back to Snowflake through an external object depends exclusively on what the OData adapter exposes. Many implementations configure external objects as read-only by default in an attempt to avoid the write-back risks that we went over before.

Reporting also works with some caveats. External objects can be included in standard Salesforce reports, however the performance of the report depends entirely on Snowflake’s live query response time rather than Salesforce’s indexing speed. It typically feels slower to run reports from an external object compared to one built entirely on standard objects.

Automation is the least robust of the four, because Salesforce automation technologies like Flow might have the ability to pull in external object data but the deeper connections needed for most admins (declarative automation features, certain trigger contexts) don’t fully extend to external objects. Any automation built against Snowflake-sourced fields needs to be tested carefully rather than assumed to behave identically.

How does a Snowflake connector work with Salesforce data?

A Snowflake connector is a purpose-built tool that helps manage the extraction of Salesforce data and its feeding into Snowflake. The connector takes on the technical burden of API communication, data typing, and loading so that engineering teams would not have to build and maintain the infrastructure in question themselves.

Snowflake Connector vs Direct Connector vs Output Connector: Which Salesforce-Native Option Should You Use?

There are three natively-branded options for connecting Salesforce and Snowflake and the naming conventions used on these options are dangerously close to each other, causing confusion during tool selection. Yet, they all address a distinctly different issue.

  • The Output Connector (also known as Sync Out to Snowflake) pushes Salesforce data to Snowflake regularly, sending your Salesforce data directly to the warehouse without needing an intervening third-party pipeline.
  • The Direct Connector does the reverse and works very differently — giving your live, zero-copy, on-demand access to Snowflake data directly from within Salesforce’s Analytics Studio, no imports or duplications necessary.
  • The Snowflake Connector utilizes the Salesforce Data Pipelines platform to copy data into Salesforce’s own pipeline platform and supports both private key as well as OAuth 2.0 authentication.

The choice between the three generally comes down to freshness and direction:

  • Output Connector for moving Salesforce data into Snowflake per schedule
  • Direct Connector for live visibility into Snowflake without import
  • Snowflake Connector for feeding data back into Salesforce-native pipeline

How does the Snowflake connector handle schema changes?

Schema changes in Salesforce are one of the most common causes of pipeline failures in a Snowflake to Salesforce integration — be it because of new custom fields, renamed fields, or removed fields. The way a connector handles schema drift varies substantially between tools, and is usually a very important factor for evaluation.

Most managed connectors approach schema changes in one of the following ways:

  • Auto-detection and column addition — the connector detects new fields in Salesforce and automatically adds the corresponding column to the Snowflake table, which is the most seamless approach
  • Schema versioning — the connector creates a new table version when breaking changes occur, preserving historical data while accommodating the new structure
  • Alerts without auto-resolution — the connector flags the schema change and pauses the pipeline until a human reviews and approves the change
  • Silent failure — lower-quality connectors may skip changed fields without alerting, which causes data loss that is difficult to detect

What limitations exist when using a Snowflake connector?

Even purpose-built connectors for the Snowflake Salesforce integration carry inherent limitations that teams should understand before committing to a tool, such as:

  • Salesforce API dependency — all connectors are subject to Salesforce’s API governor limits, which means high-volume syncs can consume a significant portion of the organization’s daily API allocation
  • Limited transformation support — most connectors are designed for extraction and loading, not transformation; complex data modeling still requires a separate tool such as dbt
  • Connector-specific object support — not all connectors support every Salesforce object, particularly newer or less common ones such as Salesforce Inbox or Experience Cloud data
  • Latency ceilings — even connectors that advertise near-real-time sync typically introduce some lag; true sub-second delivery is generally not achievable through connector-based architectures
  • Vendor lock-in risk — switching connectors later requires remapping pipelines, which validating data continuity across the transition adds significant migration overhead

Troubleshooting and common pitfalls in Snowflake Salesforce Integration

Even the most well-designed Snowflake Salesforce integrations might encounter failures, this is practically inevitable. As such, the main focus should be not to try and prevent all failures imaginable, but to develop a robust detection and troubleshooting Salesforce environment that can find failures quickly and resolve them cleanly.

What are typical connectivity and authentication errors and how do you fix them?

The first category of errors encountered by most teams are connectivity and authentication failures — and they tend to reappear whenever credentials are rotated, IP allowlists are updated, or Salesforce security policies change. Most of those issues can be quickly resolved once the root cause has been identified correctly; the problem here is that Salesforce API’s error messaging is not always specific enough to directly point at the source of the issue.

The table below maps the most common error types to their likely causes and recommended fixes:

Error TypeLikely CauseFix
INVALID_LOGINExpired password or locked integration user accountReset credentials; enable “Password Never Expires” on the integration user
REQUEST_LIMIT_EXCEEDEDDaily API call limit reachedReduce sync frequency; request a limit increase from Salesforce; batch requests more efficiently
INVALID_SESSION_IDOAuth token expired mid-syncImplement token refresh logic; check connected app session timeout settings
IP_RESTRICTEDConnector IP not in Salesforce trusted IP rangeAdd connector egress IPs to Salesforce network access settings
UNABLE_TO_LOCK_ROWRecord-level locking conflict during high-volume syncReduce concurrency settings in the connector; schedule syncs outside peak Salesforce usage hours
SSL/TLS handshake failureCertificate mismatch or outdated TLS versionVerify connector TLS version compatibility; update root certificates

The best prevention strategy is a dedicated monitoring check that can validate the connectivity and authentication status of the Salesforce API before each sync run begins — to help avoid discovering failures partway through a load job.

Why might data be missing or duplicated after sync?

Both missing and duplicated data are symptoms of the same underlying issue — that the sync mechanism doesn’t have a complete or accurate picture of what changed in Salesforce since the last run. The causes differ, but neither issue triggers an obvious failure, making both of them particularly problematic in production pipelines. 

Missing data can be traced back to timestamp-based incremental sync logic in most cases. Connectors that rely on LastModifiedDate or SystemModstamp as a detection mark are going to miss records modified via bulk operations, some automation flows, or back-end data corrections that do not affect the timestamp field. Deleted records are another common issue — as standard Salesforce queries don’t return soft-deleted records, and connectors that don’t explicitly query the recycle bin will just drop those changes from the warehouse altogether.

Duplicated data mostly originates from the retry logic. Once a sync job fails mid-run and restarts, connectors without idempotent loading are going to re-insert records that were already successfully loaded before the failure happened. This can be fixed by ensuring that the loading process uses upsert logic keyed on Salesforce record IDs instead of relying on regular inserts — such an option is available in most managed connectors, but it’s not always enabled by default.

Missing records, no alarm bells.

GRAX captures every change, deletes included.

Try GRAX for free

FAQs

What limitations should you expect from Snowflake Salesforce integration?

Salesforce API governor limits place a hard ceiling on how much data can be extracted per day, which means high-volume integrations require careful planning around sync frequency and object prioritization. Schema drift, timestamp-based sync gaps, and the handling of deleted records are persistent operational challenges that no connector eliminates entirely — they can only be managed with the right tooling and monitoring in place. 

How viable are open-source frameworks for building and maintaining production-grade Snowflake—Salesforce data pipelines at scale?

Open-source tools like Singer, Airbyte, and Apache Airflow can support production Salesforce to Snowflake pipelines, but their viability depends heavily on the engineering capacity available to build, maintain, and extend them over time. The total cost of ownership for a custom open-source pipeline — including maintenance, incident response, and keeping pace with Salesforce API changes — frequently exceeds the cost of a managed connector at any meaningful scale. 

What are the long-term trade-offs of developing and maintaining a custom Snowflake–Salesforce connector?

A custom connector gives engineering teams full control over extraction logic, schema handling, and sync behavior. However, that control comes with permanent ownership of a system that Salesforce’s evolving API and data model will continuously pressure to change. Most organizations that build custom connectors underestimate the ongoing maintenance burden and find themselves allocating disproportionate engineering time to pipeline upkeep rather than higher-value data work.

See all

Join the best
with GRAX Enterprise.

Be among the smartest companies in the world.