Data integration blog with guides on sync, ETL and APIs
Practical writing on data integration from the team building production-grade adapters:
ETL vs ELT, webhooks vs polling, data mapping rules, and what point-to-point scripts
really cost. No filler, plenty of field names.
Informatica publishes no price per IPU on its own site, but it lists IDMC on AWS Marketplace at 120 IPUs a month for $131,760 on a 12-month contract, about $91.50 per IPU-month at list. The conversion from PowerCenter is metered against the same pool. Here is every line of a PowerCenter to IDMC budget, which ones Informatica prices, and which ones it does not.
Most MySQL types have an obvious Snowflake target and a connector will pick it without complaint. The dangerous ones convert, load, reconcile perfectly on row counts, and change what your data means. MySQL BOOL is TINYINT and arrives as a number, so a column holding 2 is excluded by a = TRUE filter in MySQL and included by the same filter in Snowflake. Here is every mapping, plus the one line query that proves each risky one is safe, run against MySQL before the load.
Snowflake publishes two Oracle to Snowflake type mappings and they disagree about the most common numeric column in an Oracle schema. A NUMBER declared with no precision becomes NUMBER(38,19) under one and NUMBER(38,18) under the other, and that one digit decides whether a twenty digit value loads or overflows. Here is every mapping, every disagreement adjudicated against Oracle's own reference, and the query that proves each risky conversion is safe, run against Oracle before the load.
Snowflake publishes two separate PostgreSQL type mappings and they disagree on timestamps, arrays and currency. One of them sends a money column to FLOAT, which is the exact thing the PostgreSQL manual tells you not to do. Here is every mapping, plus the query that proves each risky one is safe, run against Postgres before the load.
Most SQL Server types have an obvious Snowflake target, and a converter will pick it without complaint. Roughly a third of those conversions then produce valid DDL, load without an error, reconcile on row counts, and change what your data means. Here is every mapping, plus the one line query that proves each risky one is safe, run against SQL Server before the load.
Most T-SQL converts to PostgreSQL mechanically. TOP becomes LIMIT, IIF becomes a CASE expression, and a converter handles the bulk of it. Then comes the part that costs teams a fortnight of debugging after go-live: six translations that every conversion guide publishes, which compile, run, return a value, and return a different value than the statement they replaced.
Converters handle most of an Informix code base: DEFINE, LET, FOREACH and ON EXCEPTION all have PL/pgSQL equivalents. Six SPL behaviors do not survive the trip, from ON EXCEPTION WITH RESUME to DATETIME variables declared as DATE, and each one compiles in PostgreSQL and returns a different answer.
Converters handle most of an SAP ASE code base: @@rowcount, raiserror, *= joins and getdate() all have PL/pgSQL equivalents. Six ASE behaviors do not survive the trip, from blank strings stored as a space to set rowcount on DELETE, and each one compiles in PostgreSQL and returns a different answer.
Four ways to keep Snowflake current from IBM Db2, what each one needs switched on in Db2, and the cost nobody quotes: how often you merge decides how long the Snowflake warehouse runs, and that bill can be larger than the replication tool.
Most Db2 types cross into PostgreSQL unchanged. DECFLOAT, long timestamps, FOR BIT DATA, oversized LOBs and identity columns do not. A full type mapping, the defaults to reject, and the SYSCAT query to run on Db2 that proves each risky mapping safe before a single row loads.
Converting Oracle DDL to MySQL looks mechanical until you meet NUMBER, DATE and the empty string. A full type mapping for every common Oracle column, the four decisions no converter can make for you, and the query to run on Oracle that proves each risky mapping safe before a single row loads.
Most SQL Server to MySQL type conversions are obvious. About ten are not, and they deploy cleanly while changing what the data means: a rowversion mapped to a date, MONEY to a whole number, datetime2 rounded across midnight. The full mapping table, plus the query to run on SQL Server that proves each risky mapping safe before the load.
Most MySQL to PostgreSQL type conversions are obvious and a converter gets them right. About a dozen are not, and those produce a schema that deploys without a single error and quietly means something different. The full mapping table, plus the part nobody publishes: the one line query that proves each risky mapping is safe, run against MySQL before the load rather than PostgreSQL after it.
AWS publishes an automation rating for every Oracle feature area in its own migration playbook, and two of them are rated no automation at all: MERGE statements and database links. Ora2Pg calls its own PL/SQL conversion basic. Here is what actually converts, what you rewrite by hand, and the shorter list that worries experienced teams most: the constructs that convert at the highest rating and quietly return different results.
The tool is almost never the expensive part of a data migration. Cloud services bill per replication hour, ELT vendors bill per row, consultancies bill per day, and all of it is dwarfed by the time people spend mapping fields and proving the result is correct. Here is how each pricing model behaves when a project runs long, and how to build an estimate that survives contact with the data.
A NetSuite account on the Standard service tier gets five concurrent web services requests in total, shared by every integration connected to it. That single number breaks more NetSuite projects than any rate limit, and the error it produces is HTTP 400 rather than the 429 every retry library watches for. The real limits by tier, the six failures teams experience as one error, and the fixes.
Stripe allows 100 requests a second in live mode and 25 in a sandbox. Those are the numbers everyone quotes. The reason Stripe integrations break is that six genuinely different problems all return HTTP 429, and their fixes contradict each other. The current limits, how to tell the six apart in one header, and seven changes that fix them.
QuickBooks throttles at 500 requests a minute per company file and allows ten concurrent connections to it. Almost every 429 we get asked about is the second limit, not the first, which is why adding workers makes a slow QuickBooks sync slower. The current numbers, the five failures that all look like one problem, and the six changes that fix them.
Shopify does not count your requests, it prices them. The GraphQL Admin API meters calculated query cost in points per second and your plan sets the rate, which is why the same connector behaves like a different product on Standard and on Plus. The current numbers, and the six changes that fix a throttled sync.
Almost every stalled HubSpot integration is stalled for one of two reasons, and neither is the daily call limit everybody worries about. The current numbers by tier, what a 429 actually costs you, and the six changes that fix a rate-limited sync.
Enterprise, Unlimited and Performance orgs get five of these at no extra cost, and most of them sit unclaimed while an integration runs as somebody's admin account. What the license permits, what it costs, where it gets in your way, and the setup order that trips up almost every first attempt.
Zero-ETL is described in two wrong ways: as marketing for a faster connector, or as a universal replacement for ETL. It is neither. Here is exactly what AWS replicates for you, what you give up in exchange, and the failures that show up quietly in week one.
The commands are three lines long and teams still end up throttled by a quota nobody read, money columns that lost their cents, and a pipeline that stopped a month ago. Every realistic path into BigQuery, with the numbers from Google's own documentation.
The commands are short and almost every team still ends up with slow loads, money columns that lost their cents, and a pipeline that stopped three weeks ago. Every realistic path into Snowflake, with the numbers from Snowflake's own documentation.
Two completely different things get called connecting Snowflake to Salesforce, and picking the wrong one costs a month. Here is the difference between a write that lands on records and zero copy federation that does not, plus the API limits and type mismatches that break the first sync.
PostgreSQL ships real logical replication, and it is the right answer whenever both ends are Postgres. It also silently refuses to replicate your schema, your sequences and your large objects. Here is the setup, the documented restrictions, and the slot that fills your disk.
PostgreSQL gives you real log-based CDC through logical decoding, and it gives you a replication slot that will retain write-ahead log files until your disk is full. Here is how the pieces fit, what to set on RDS, and when a watermark query is the better answer.
Stripe charges into QuickBooks with the processor fee on its own account. A Shopify order that decrements NetSuite inventory before the warehouse oversells it. Twelve integrations teams actually run, and the three disagreements that make each one hard.
PostgreSQL folds unquoted identifiers to lowercase while SQL Server keeps the case you typed, and the MONEY type on both sides means different things. Here is the phased migration that lets you prove the totals match before you move the application.
MONEY columns that land as a float lose cents nobody notices until the quarter is closed, and SQL Server change data capture holds the transaction log open when nothing is consuming it. Here is the incremental migration that survives both.
BigQuery bills by bytes scanned, so an unpartitioned copy of a large MySQL table turns a cheap dashboard into a recurring line item. Here is the incremental load that keeps analysts off production and keeps the query bill flat.
A big-bang UNLOAD and COPY means a frozen reporting window nobody signs off on, and Redshift folds identifiers to lowercase while Snowflake folds them to uppercase. Here is the parallel-run migration that lets you validate totals before you cut over.
The strongest HubSpot to Snowflake route is HubSpot's own Data Share, and it is available on one subscription tier only. Below that tier, Snowflake's connector lands a JSON payload with a view of fields it considers common. Here is every option scored on the three questions that decide the purchase: custom objects, deleted records and property history.
Every MongoDB to Snowflake tool reads the same change stream, so the pipeline is rarely what you are buying. What separates them is what happens to a nested document on arrival. Snowflake's own connector lands it whole in one column and leaves the modeling to you. Here is which tools do the flattening, and what each does when a field changes shape.
Estuary, Debezium, DMS, Airbyte and Fivetran all sync MongoDB into MySQL, and two of them need a server setting MySQL ships disabled. The bigger risk is MySQL's own upsert, which matches on any unique index, not only the key.
AWS DMS cannot read DynamoDB, so most AWS teams end up on Glue. The export connector spares production read capacity, a run costs a few dollars, and Glue's own transform leaves every number as a string. Here is the job, the bill, and the alternatives.
Fivetran reads DynamoDB with a full scan, then DynamoDB Streams, and by default lands every item as one JSON column. Here is what that means in Snowflake, what the sync bills under 2026 pricing, what happens after a 24-hour outage, and which alternatives beat it for which job.
MongoDB lands in BigQuery as JSON, with dates, longs and decimals wrapped in $date, $numberLong and $numberDecimal. Google changed the Datastream default on 16 September 2026, so a date now sits at a different path in new streams. Here is the SQL that reads both shapes and the query that shows which one each table holds.
The popular MongoDB to PostgreSQL tools guess your schema from a sample and skip what the sample missed. The fix is nine short mongosh queries run against MongoDB before the load. Here is the mapping, the checks, and what each one catches.
Google's HubSpot connector for BigQuery is free while it is in Preview, and it reloads all of your HubSpot data on every run. That makes it free today and the one whose bill grows with portal size once Google starts charging slot-hours. Here is what every route costs on the same portal, including the HubSpot API allowance nobody invoices.
Google's own Salesforce connector for BigQuery is paid and publicly priced: $0.06 per slot-hour, with a planning ceiling of about $1.20 per hour of run time. That makes the schedule, not the row count, the thing that sets your bill. Here is what each route costs on the same Salesforce org.
There is no single price for moving Salesforce into Snowflake, because there is no single product. Snowflake says its own connector is built for teams who do not own Salesforce Data Cloud, so a licensing fact picks your architecture before any feature comparison. Here is what each route actually bills you for.
Analysts querying the production MySQL replica slow down the app at month end, and hand-rolled load scripts double revenue after a retry. Here is how to replicate MySQL into Snowflake incrementally with type mapping that survives contact with real money columns.
A mysqldump is stale the second it finishes, and MySQL types do not land cleanly in Postgres: TINYINT(1) is a boolean in practice, unsigned integers overflow, and 0000-00-00 has no equivalent. Here is the migration that runs both databases in parallel until the numbers agree.
The Square API pages by cursor and returns money as integer cents in a nested object, and its reports cannot join to your inventory or accounting tables. Here is how to replicate payments into Postgres incrementally so reporting queries transaction data with plain SQL.
The PayPal Transaction Search API only returns a bounded date window per call and splits gross, fee, and net across objects, so a naive pull drops rows and breaks reconciliation. Here is how to replicate transactions into Postgres incrementally and tie out to the payout.
Querying a Snowflake table on every page load is slow and expensive, so teams cache it by hand and the copy drifts. Here is how to push modeled data into Postgres incrementally so your app reads it with a fast local SQL query.
Qlik retired Talend Open Studio and moved Talend onto capacity-based cloud tiers with no public list price. Here is how the capacity model works, what contracts actually run, and what pushes the number up.
A big-bang export and reload freezes reporting and hides type bugs until someone spots a wrong number. Here is how to migrate BigQuery to Snowflake incrementally, handle nested fields, validate both warehouses in parallel, and cut over safely.
Calling the QuickBooks Online API live is slow and throttled, and its reports cannot join to your product or CRM data. Here is how to replicate the ledger into Postgres incrementally so finance dashboards query it with plain SQL.
The Xero API caps you at 60 calls a minute, and its reports cannot join to your product or billing tables. Here is how to replicate the ledger into Postgres incrementally so reporting and app features query it with plain SQL.
A big-bang unload and reload freezes reporting and hides type bugs until someone spots a wrong number. Here is how to migrate Snowflake to BigQuery incrementally, validate both warehouses in parallel, and cut over safely.
Calling the Shopify Admin API live is slow and rate limited. Here is how to replicate orders, line items, customers, and refunds into Postgres incrementally so your dashboards and app features query store data with plain SQL.
Matillion sells the Data Productivity Cloud on a credit-based consumption model with no public list price. Here is how credits work, what the Developer, Teams, and Scale editions include, what a deployment runs, and what pushes the bill up.
Xero reports are per organization and cannot join to product or ad data. Here is how to load Xero into BigQuery incrementally so revenue, cash, and margin models read one clean ledger across every entity.
PayPal Activity exports cannot join to your orders or ad spend. Here is how to load PayPal into BigQuery incrementally so revenue, fee, and net-settlement models read one clean table you can trust.
SnapLogic sells three packages with unlimited pipelines and no posted prices. Here is what the pricing page states, what its $125,000 AWS Marketplace listing works out to, and the questions that make a quote comparable.
Square reports are per location and cannot join to other channels. Here is how to load Square into Snowflake incrementally so sales, fee, and margin models read one clean table across every location.
Querying NetSuite live through SuiteTalk hits governance limits. Here is how to replicate NetSuite into Postgres incrementally so your apps, tools, and reports read live ERP data without throttling the API.
Informatica sells IDMC through sales on an IPU consumption model with no public list price. Here is how the meter works, the one price Informatica publishes, what PowerCenter adds, and what inflates the bill.
Querying Salesforce live burns API calls against a daily cap. Here is how to replicate Salesforce into Postgres incrementally so your apps, tools, and reports read live CRM data with plain SQL.
HubSpot reports cannot join to product usage or ad spend. Here is how to load HubSpot into BigQuery incrementally so pipeline, attribution, and funnel models read one clean, partitioned history.
Jitterbit publishes no prices, but it does publish how its plans are sized. Here are the connection, agent, and environment limits for each plan, where App Builder fits, and what to ask for in a quote.
Stripe dashboards show today, not history you can model. Here is how to load Stripe into BigQuery incrementally so revenue, fee, and MRR models read one clean table.
Shopify reports cannot join to ad spend or fees, and CSV exports go stale. Here is how to load Shopify into BigQuery incrementally so cohort, LTV, and margin models run on a full order history.
Tray.ai publishes plans but no prices, and bills usage in Tasks. Here is what each plan includes, what its $150,000 AWS Marketplace listing works out to, why the Task counting rule matters, and what to ask sales.
Monthly CSV exports give you a snapshot, not a live model. Here is how to load QuickBooks Online into Snowflake incrementally so finance models read today's numbers.
HubSpot reports cannot join to product or billing data, and list exports go stale. Here is how to load HubSpot into Snowflake incrementally so attribution and pipeline models run on a full history.
Fivetran charges each distinct primary key once a month. Airbyte charges the volume moved, and a full refresh recharges every row. At high volume that one difference decides the winner.
Airbyte Cloud Standard is $20 a month with 5 credits, and every extra credit is $5. Here is what each plan includes, what a million rows costs, and what happens when the credits run out.
Monthly CSV exports give you a snapshot, not a live model. Here is how to load QuickBooks Online into BigQuery incrementally so finance dashboards read today's numbers.
The Stripe API paginates 100 objects at a time and rate-limits, so backfilling by hand is brittle. Here is how to load charges, fees, and payouts into Snowflake so revenue models tie out.
n8n cloud bills by workflow execution and the community edition is free to download. Here is what each tier actually covers, what pushes the meter up, and what self-hosting really costs to run.
Nightly full dumps rewrite whole tables and inflate BigQuery cost. Here is how to load Postgres incrementally, map the types correctly, and handle updates, deletes, and schema drift.
Saved-search exports are manual and stale, and SuiteAnalytics Connect carries its own license. Here is how to load NetSuite transactions into Snowflake incrementally so the numbers tie out.
Celigo prices by endpoints and flows across Standard, Professional, and Enterprise, and publishes no list price. Here is what each edition includes and what to ask before you sign.
Full table pulls burn Salesforce API calls and land stale data in the warehouse. Here is how to load accounts, opportunities, and activity into Snowflake incrementally and keep it fresh.
Square pays you net of fees, so the deposit never equals gross sales. Here is how to post Square sales, refunds, and fees into NetSuite so each payout reconciles on the first try.
A CSV dump lands accounts in the wrong type and duplicates your contacts. Here is how to move customers, the chart of accounts, and open balances from QuickBooks to Xero cleanly.
Sales closes a deal in HubSpot and finance rekeys it into QuickBooks. Here is how to sync contacts, turn closed-won deals into invoices, and push payment status back to the deal.
Boomi publishes one price, Pay-As-You-Go at $99 a month plus $0.05 per message, and quotes every edition above it. Here is how both meters work, what drives the bill, and where a flat price wins.
PayPal takes its fee before it pays you, so the deposit never equals the sale. Here is how to post sales, fees, and refunds so QuickBooks reconciles on the first try.
Fivetran charges by Monthly Active Rows and prints no rate per million on its pricing page. Here is what it does publish: the Free allowance, the $5 base charge, deletes counting since January 2026, annual discounts, and how to get a comparable quote.
Shopify pays out net of fees, so the deposit never equals the order total. Here is how to post orders, fees, refunds, and sales tax so QuickBooks reconciles on the first try.
Self-serve platforms run tens to hundreds of dollars a month, enterprise suites run six figures a year, and in-house looks free until you price the maintenance. The arithmetic, in full.
The build is the cheap part. Auth, pagination, retries, backfills, and every upstream API version bump are the bill you pay forever. A framework for deciding which side you are on.
iPaaS is a hosted platform that moves data between your apps so you do not write and babysit point-to-point scripts. Here is how it works and when it pays off.
An API adapter sits between two systems and translates one data shape into another. It is the piece of glue code every team rewrites, and the piece we productized.
Most sync failures are mapping failures. Nine rules, from picking a stable match key to versioning every change, that keep records landing where they should.
Every pair of connected systems adds another script someone has to own. At ten systems you are maintaining dozens of fragile links. The math gets ugly fast.
7 min read
Done reading about glue code? Retire it
Connect your first pair and watch records land where they should. Flat price from $49 a month.