Skip to main content

Reverse ETL vs. Salesforce Data Cloud Zero Copy

Emily Winks, Data Governance Expert, Atlan
Data Governance Expert
Updated:
|
Published:
18 min read

Key takeaways

  • Reverse ETL tools like Hightouch and Census copy warehouse data into Salesforce; Zero Copy queries it in place instead.
  • Zero copy federation runs three ways, and one of them, Cached Acceleration, does store a copy inside Data 360.
  • Census was acquired by Fivetran in 2025 and rebranded as Fivetran Activations.
  • Neither sync mechanism verifies whether the data Agentforce reads is current, certified, or accurate.

What is reverse ETL for Salesforce, and how does it compare to Zero Copy?

Reverse ETL tools like Hightouch and Census copy warehouse-computed fields into Salesforce on a schedule, while Salesforce Data 360's zero copy federation queries that same warehouse data in place, across Snowflake, BigQuery, Redshift, Databricks, and object stores. Both get data in front of Agentforce, but neither verifies it is current or correct once it arrives. Census was acquired by Fivetran in 2025, and one of the three zero copy modes, Cached Acceleration, does keep a cached copy inside Data 360 on a refresh schedule you configure.

What the comparison comes down to:

  • Mechanism reverse ETL copies data on a schedule; Zero Copy runs Live Query and File Federation without duplicating it, and Cached Acceleration with a temporary copy
  • Freshness Live Query is resolved at read time; a Cached Acceleration refresh runs on a configured window, minutes to days
  • Trust gap neither approach certifies, traces, or verifies the data before Agentforce acts on it

See how ready your data is for Agentforce

Take the Maturity Assessment

Hightouch, Census, Salesforce’s own Data Cloud Zero Copy, and Atlan each solve a different piece of the same problem: getting warehouse-computed data into Salesforce so Agentforce and other AI agents for sales can act on it. Reverse ETL physically copies a transformed field, a churn score, a propensity model, an enriched segment, from your warehouse into a Salesforce object on a schedule or trigger. Zero copy federation queries that same warehouse in place and registers it in Data 360, with one of its three modes keeping a temporary cached copy. Both are mature, well-documented approaches; the decision between them usually comes down to cost, latency, and who owns the transformation logic, not which one is “right.”

Once an AI agent is the one reading the result, the stakes of that decision change. A stale churn score doesn’t just look bad on a dashboard; it can feed a decision an agent makes on its own, like offering a retention discount to a customer who already churned. This page compares reverse ETL and Zero Copy on their own mechanical terms, then looks at the separate question of what makes the data either one delivers trustworthy enough for Agentforce to act on.

Dimension Reverse ETL Salesforce Data Cloud Zero Copy
What it is Tools (Hightouch, Census, Fivetran) that copy warehouse data into Salesforce Salesforce’s federated query layer into Snowflake, BigQuery, Redshift, Databricks, and object stores
What it does Syncs transformed fields into Salesforce objects on a schedule or trigger Lets Data 360 read warehouse data in place through Live Query, File Federation, or Cached Acceleration
Who owns it Data/analytics engineering, usually SQL or dbt-based Salesforce admins and the warehouse team jointly
Key strength Full control over transformation logic and sync cadence Query pushdown to the source, and no permanent second copy under Live Query or File Federation
Best for Complex dbt models, high-frequency syncs, keeping logic in SQL Teams already inside the Salesforce/Data 360 ecosystem wanting fewer pipelines
Questions it answers How do I get this exact warehouse field into a Salesforce object? Can Salesforce query my warehouse without copying it first?
Cost model Per-tool subscription, scales with rows/destinations synced Bundled into Data 360 consumption, plus the source platform’s own costs
Complexity level Moderate, one more pipeline to build and monitor Lower for federated sources, higher for anything that needs a connector built

Reverse ETL vs. Salesforce Data Cloud: what’s the difference?

Reverse ETL moves data; Zero Copy queries it where it already lives. Reverse ETL tools like Hightouch, Census, and Fivetran Activations take a field you’ve already computed in the warehouse and copy it into a Salesforce object, on a schedule ranging from near-real-time streaming to nightly batch. Zero copy federation, Salesforce’s current term for the approach, lets Data 360 query a source in place and register it for use inside Data 360. Salesforce names the mechanism behind it: advanced query pushdown “delegates data retrieval to the originating data warehouse or lake using an optimized pushdown query that retrieves just the data needed.”

According to Salesforce Ben, reverse ETL emerged because warehouses became the place where real analytical work happens, dbt transformations, ML scoring, enrichment, while Salesforce still needed that output to act on it. Salesforce’s own Data 360 Interoperability decision guide frames federation as the way to avoid maintaining another copy and another pipeline for data that already sits somewhere governed.

Both sides market their approach as the modern one, but they optimize for different failure modes: reverse ETL trades pipeline maintenance for control over exactly what gets written where, while Zero Copy trades some flexibility for eliminating duplicate copies drifting out of sync. That distinction matters more once an enterprise copilot is reading whichever version lands in Salesforce, rather than a person checking a report once a week, though the two mechanisms remain equally capable of delivering that data correctly.


What is reverse ETL for Salesforce?

Reverse ETL for Salesforce syncs data already modeled or scored in a warehouse, most often Snowflake, BigQuery, or Databricks, into Salesforce objects so sales, service, and AI agents can act on it inside the CRM. It exists because Salesforce was never meant to be the system of record for complex transformations; that work lives in the warehouse, and reverse ETL delivers the result.

The category is dominated by a handful of purpose-built tools. Hightouch positions itself as a “composable CDP”, emphasizing broad destination coverage and a visual audience builder for marketing teams; its founders have described the goal plainly as giving every team access to warehouse data in the tool they already use, without relying on engineers to build a custom pipeline for it. Census, acquired by Fivetran in 2025 and rebranded into Fivetran Activations, leans into dbt-native, SQL-first syncs for data engineering teams that want granular control over which model feeds which field. Fivetran increasingly markets Activations to reuse the same connectors for both ingestion and reverse sync, cutting the number of tools a team has to run.

Core components of a reverse ETL pipeline


  • Warehouse-side transformation: dbt models or SQL that compute the field, a churn score, an LTV estimate, a segment assignment, before it reaches Salesforce.
  • Sync engine: the tool, Hightouch, Census, Fivetran Activations, that reads the warehouse table and writes to Salesforce’s API.
  • Destination mapping: the configuration connecting a warehouse column to a specific Salesforce object and field.
  • Sync frequency or trigger: batch, streaming, or event-triggered, depending on how fresh the field needs to be.
  • Monitoring and alerting: catching failed syncs or schema drift before someone acts on stale or low-quality data.

Get the AI Context Stack

A practical brief on what enterprise AI agents actually need underneath the pipeline, whichever sync mechanism you choose.

Get the AI Context Stack

What is Salesforce Data Cloud Zero Copy?

Zero Copy is Salesforce’s approach to eliminating the sync step: instead of copying warehouse data into Salesforce, Data 360 queries it where it lives. Salesforce’s Data 360 Interoperability decision guide names three modes, and they behave differently. Live Query runs queries directly against the external system “with no data duplication.” File Federation gives “direct, read-only access to large-scale datasets in object stores,” reading Apache Iceberg tables on S3, GCS, or ADLS. Cached Acceleration “temporarily stores cached copies of federated data in Salesforce Data 360,” for a configurable duration Salesforce describes as minutes to days, and Salesforce calls that caching “an optional optimization, not a requirement.” So the shorthand that the data never moves holds for two modes out of three. Named federated sources include Snowflake, Google BigQuery, Redshift, and Databricks.

Rahul Auradkar, EVP and GM for Data Cloud at Salesforce, has framed this as the point of the architecture: “In this new era of AI and agents, customer data and metadata are the new gold for the enterprise… because Data Cloud is the foundation of Salesforce, companies can act on data to create the most personalized and meaningful customer experiences,” he said in Salesforce’s own announcement. The pitch: no duplicate copies drifting out of sync, no pipeline to monitor, and a federated table registered in Data 360 where Data 360’s own policies already apply.

Two caveats worth knowing before choosing Zero Copy. It depends on the source being one Salesforce federates, so anything else still needs a connector or a pipeline to reach Data 360. And freshness depends on which mode a given table uses: Salesforce describes Live Query as suited to infrequent or ad-hoc queries “where freshness is critical,” while a Cached Acceleration table is only as current as its configured refresh, which matters for genuinely time-sensitive data on either side of this comparison.

Core components of zero copy federation


  • Live Query: queries run straight against the external system, with no duplication.
  • File Federation: read-only access to Iceberg tables sitting in S3, GCS, or ADLS object stores.
  • Cached Acceleration: a temporary cached copy held inside Data 360, refreshed incrementally on a schedule from minutes to days.
  • Advanced query pushdown: the mechanism that delegates retrieval to the source warehouse or lake and pulls back only the data needed.
  • Registration in Data 360: federated tables become usable inside Data 360 without persisting a copy, and Data 360 policies apply to them the same way they apply everywhere else in Data 360.

Reverse ETL vs. Zero Copy: head-to-head comparison

The sharpest differences between reverse ETL and Zero Copy show up in who controls the transformation logic, how fresh the data is, and what breaks when something goes wrong, not in which one is objectively better.

Dimension Reverse ETL Zero Copy
Primary focus Delivering a specific transformed field into a Salesforce object Federating warehouse queries without duplicating data
Representative tools Hightouch, Census, Fivetran Activations Salesforce Data 360 zero copy federation: Live Query, File Federation, Cached Acceleration
Transformation logic location Warehouse (dbt/SQL), fully under data team control Warehouse or Data 360, depending on the use case
Sync or refresh latency Sub-second streaming to hourly/nightly batch, tool-dependent Resolved at read time under Live Query; set by the configured refresh under Cached Acceleration
Cost model Per-tool subscription, scales with destinations and sync volume Bundled into Data 360 consumption plus the source platform’s own costs
CRM data footprint Data duplicated into Salesforce storage None under Live Query or File Federation; a temporary cached copy under Cached Acceleration
Lock-in and portability Sync logic is tool-specific, destination data is portable Tied to the sources Salesforce federates and to Data 360 architecture
Failure mode Silent sync failures, schema drift, stale fields nobody notices Federation gaps for sources Salesforce does not reach, dependence on Salesforce uptime
Best for Complex, warehouse-owned transformations landing as real Salesforce fields Teams standardized on a federated source wanting fewer pipelines to run

Example: a churn score reaching Agentforce two different ways. A RevOps team computes a churn-risk score in Snowflake using a dbt model combining usage data, support tickets, and billing history. With reverse ETL, Census syncs that score into a custom Salesforce field hourly, and Agentforce reads it when deciding whether to flag an account for outreach. With Zero Copy in Live Query mode, the score stays in Snowflake; Data 360 queries it there, and Agentforce reads it through the federated layer. Both paths get the score in front of the agent on the schedule each mechanism promises. What happens if that dbt model failed to run this morning is a separate question from which mechanism moved the number, and it’s the same question either way.


How do reverse ETL and Zero Copy work together for Agentforce?

Reverse ETL and Zero Copy aren’t mutually exclusive, and most enterprises past a certain size run both for different fields, depending on which warehouse holds the source data and how time-sensitive the use case is.

Signal Points to reverse ETL Points to Zero Copy Points to both
Warehouse platform Any warehouse, including ones Salesforce does not federate Snowflake, BigQuery, Redshift, Databricks, or an Iceberg object store Multi-warehouse environments
Transformation complexity Heavy dbt/SQL logic the team wants to keep warehouse-side Simpler federated lookups, less custom transformation Some fields simple, others complex
Freshness requirement Batch is acceptable, or the tool’s streaming tier covers it Genuinely time-sensitive reads where duplication is the bigger risk Mixed requirements across use cases
Pipeline appetite Comfortable owning and monitoring another sync tool Wants to reduce the number of pipelines running Depends on which fields fall on which side

When to start with reverse ETL: Salesforce doesn’t federate the source warehouse, the transformation logic needs full SQL control, or the use case tolerates batch latency.

When to start with Zero Copy: the warehouse is one Salesforce federates, the team wants fewer pipelines, and the field is a relatively direct federated lookup.

When to invest in both: most mature Agentforce deployments end up here, a pattern also visible in how teams weigh Agentforce against building agents in-house more broadly, using Zero Copy for simple lookups and reverse ETL for heavier warehouse-owned models. Whichever mix a team lands on, that’s a decision about mechanics, not about whether the data pipeline feeding Agentforce is trustworthy once it’s built, which is a separate question this page turns to next.

Take the AI Agent Context Readiness Checklist

Score whether the data feeding your Salesforce agents, however it gets there, is actually ready for an agent to act on.

Take the Readiness Checklist

How Atlan approaches Agentforce’s governed-context problem

Neither reverse ETL nor Zero Copy solves what actually determines whether Agentforce works reliably: whether the field an agent reads is current, certified, and traceable back to its source. A synced churn score and a federated one fail identically when the underlying model breaks silently, or when nobody documents that “churn risk” changed definitions last quarter, the same gap covered in Agentforce data governance: a data contract and metadata management problem sitting underneath both sync mechanisms, not a limitation of either one specifically.

Atlan doesn’t compete with Hightouch, Census, Fivetran, or Zero Copy; it sits above whichever mechanism moves the data. Atlan’s context graph spans the warehouse, the transformation pipeline, and Salesforce itself, closer to a context layer than a data catalog or semantic layer alone, reconstructing lineage across that boundary regardless of whether a field arrived via a scheduled sync or a federated query. Certified business definitions travel with the field: once “churn risk” is defined and owned by a specific team, that definition, its freshness, and its certification status stay visible wherever the field shows up next, the same knowledge architecture principle that underpins a semantic layer for AI agents.

Delivered through Atlan’s MCP server for Salesforce, built the way any enterprise MCP server is, Agentforce gains a way to call for context that lives outside Salesforce and outside Data Cloud’s federated reach, ownership, lineage, and definitions, at inference time, the same case made for why MCP matters for AI agents generally and for giving agents access to enterprise data specifically. That’s the piece MCP delivers that neither reverse ETL nor Zero Copy was built to provide: not another way to move or federate data, but a way to know whether the data an agent is about to act on deserves to be trusted, which is also why enterprise-ready agents need this layer regardless of which CRM or warehouse sits underneath them. Teams building that layer from scratch typically start with how to implement an enterprise context layer for AI or how to build an AI agent harness, depending on whether the gap is in the agent’s scaffolding or the data underneath it.


See Governed Context Reach Agentforce

Watch how certified definitions and lineage travel with a field into Salesforce, no matter which sync mechanism got it there.

Watch the Live Demos

What actually decides whether Agentforce can trust its Salesforce data

The reverse-ETL-versus-Zero-Copy decision is real, worth taking seriously on the cost, latency, and pipeline-ownership terms this page compared it on. Picking correctly there won’t stop an agent from acting on a number that’s simply wrong, because getting a field into Salesforce faster or with less duplication says nothing about whether it’s still accurate once it arrives. As more enterprises settle into hybrid patterns, reverse ETL for the complex warehouse-owned models, Zero Copy for the simpler federated lookups, the ones getting reliable agent accuracy out of Salesforce are treating governed context as its own line item, budgeted and built alongside the sync decision rather than assumed to come free with it.


FAQs about reverse ETL for Salesforce AI agents

1. What is reverse ETL for Salesforce?


Reverse ETL for Salesforce is the practice of syncing data already modeled or scored in a warehouse, typically Snowflake, BigQuery, or Databricks, into Salesforce objects and fields so sales, service, and AI agents can act on it directly. Tools like Hightouch, Census, and Fivetran Activations handle the sync on a schedule ranging from streaming to nightly batch. It exists because complex transformations happen in the warehouse, not in Salesforce itself.

2. What is the difference between reverse ETL and Salesforce Data Cloud Zero Copy?


Reverse ETL physically copies a transformed field from your warehouse into a Salesforce object. Zero copy federation queries that same warehouse in place, and Salesforce runs it three ways: Live Query against the external system with no duplication, File Federation reading Iceberg tables in object stores, and Cached Acceleration, which temporarily stores a cached copy inside Data 360 on a refresh window you configure. Reverse ETL gives you full control over transformation logic and sync cadence; zero copy trades some of that control for fewer duplicate, drifting copies of the data.

3. Does Agentforce require Salesforce Data Cloud?


Agentforce is built to read from and write to Salesforce Data Cloud, which functions as its primary grounding source for customer and operational data. It can also reference data delivered through reverse ETL into standard Salesforce objects, or context served through external sources like an MCP server. Data Cloud is the native path, but it isn’t the only way to get relevant data in front of an Agentforce agent.

4. Can reverse ETL and Zero Copy be used together?


Yes, and most mature Salesforce deployments end up doing exactly this. Zero Copy tends to cover simpler, high-frequency federated lookups against a supported warehouse, while reverse ETL handles complex, warehouse-owned scoring models that need to land as durable Salesforce fields. Which one a specific field uses usually comes down to the warehouse platform and how heavily transformed the data is before it reaches Salesforce.

5. What causes Agentforce to act on stale or wrong data?


Agentforce reads whatever value currently sits in the Salesforce field or federated source it’s pointed at, and neither reverse ETL nor Zero Copy verifies that value’s freshness or correctness before the agent uses it. A reverse ETL sync can fail silently or run on a schedule too slow for the decision being made; a field served from Cached Acceleration can sit behind whatever refresh window someone configured for it. In both cases, the agent has no built-in way to know the data is stale unless something outside the sync mechanism is tracking it.

6. Is reverse ETL still needed once Data Cloud has native connectors?


For the sources Salesforce federates natively, Snowflake, BigQuery, Redshift, Databricks, and object stores on S3, GCS, or ADLS, zero copy reduces the need for reverse ETL on simpler federated use cases. But reverse ETL remains the more practical choice for complex dbt-modeled fields, sources Salesforce does not federate, or teams that want transformation logic to stay entirely in SQL they control. The two approaches address different points on the complexity and control spectrum rather than one strictly replacing the other.

7. How do Hightouch, Census, and Fivetran differ for Salesforce syncs?


Hightouch positions itself as a composable CDP with broad destination coverage and a visual audience builder aimed at marketing teams. Census, now part of Fivetran, leans into dbt-native, SQL-first syncs favored by data engineering teams that want granular control over sync logic. Fivetran markets its own Activations as a way to reuse the same connectors for both ingestion and reverse sync, reducing the total number of tools a data team has to run.

8. Does a context layer replace the need to choose between reverse ETL and Zero Copy?


No. A context layer doesn’t remove the decision between reverse ETL and Zero Copy; it determines whether whichever mechanism you pick delivers data Agentforce can actually trust. Certified definitions, lineage, and freshness signals sit above the sync layer, applying the same governed context to a field whether it arrived through a scheduled sync or a federated query.


Sources

  1. Salesforce Architects, “Data 360 Interoperability” (decision guide). https://architect.salesforce.com/docs/architect/decision-guides/guide/data-360-interoperability.html
  2. Salesforce, “Zero Copy Connectivity.” https://www.salesforce.com/data/connectivity/zero-copy/
  3. Salesforce Newsroom, “Data Cloud Momentum, Agentforce” (Rahul Auradkar, EVP and GM for Data Cloud). https://www.salesforce.com/news/press-releases/2024/09/17/data-cloud-momentum-agentforce/
  4. Salesforce Ben, “What Is Reverse ETL?” https://www.salesforceben.com/what-is-reverse-etl/
  5. Hightouch, “What Is a Composable CDP?” https://hightouch.com/blog/composable-cdp
  6. MarTech, “More Consolidation Among Data Tools, as Fivetran Acquires Census.” https://martech.org/more-consolidation-among-data-tools-as-fivetran-acquires-census/

Share this article

signoff-panel-logo

Atlan is the Context Layer for AI. It translates business knowledge, including data definitions, working procedures, and governance policies, into context AI can actually use. This knowledge lives in a single Enterprise Data Graph that every team and AI agent can reach.

In Atlan's AI Labs benchmark, adding this context improved AI's text-to-SQL accuracy by 38%.

Atlan is recognized as a Leader across multiple Gartner reports and Forrester Waves, and is trusted by over 400 enterprises representing $10T+ in market cap, including Mastercard, Workday, General Motors, CME Group, HubSpot, FOX, Virgin Media O2, and Elastic.

Bridge the context gap.
Ship AI that works.