Skip to content
adapters.io

AWS Glue DynamoDB to Postgres migration, AWS DMS DynamoDB to Postgres, the cost of each run, and the casts Glue leaves to you

8 min read Buying guides The Adapters team

Field mapping auto-plugged · tap a port to rewire

5 sample records ready

AWS Glue can migrate DynamoDB to PostgreSQL. AWS DMS cannot, because DynamoDB is not one of its supported sources. A Glue Spark job reads the table through the export connector, which uses no read capacity but needs point-in-time recovery, and writes to Postgres over JDBC. It bills $0.44 per DPU-hour in US East. The catch is types: Glue's own documented transform hands DynamoDB numbers over as strings, so the casts are code you write.

Key takeaways

  • DMS is out. DynamoDB is a DMS target only, so the AWS-native path is an S3 export read by Glue.
  • Use the export connector. It skips the table's read capacity and AWS says it is faster past 80 GB. It needs PITR.
  • Cast every number. simplifyDDBJson removes the type descriptors and leaves numbers and number sets as strings.
  • The compute is cheap. Engineering time for the schema, the casts and the change replay is the real cost.

Can AWS DMS migrate DynamoDB to Postgres?

No. This is the first thing most AWS teams try, and it stops at the endpoint screen. The AWS DMS list of supported source endpoints covers Oracle, SQL Server, MySQL, MariaDB, PostgreSQL, MongoDB, SAP ASE, IBM Db2, Amazon DocumentDB and Amazon S3. DynamoDB appears only in the target list, so DMS can load a DynamoDB table but cannot read one. Every DynamoDB to PostgreSQL route therefore starts from one of three AWS feeds: an export to S3, DynamoDB Streams, or a direct table scan. Glue can use the first and the last. The full list of tools built on those feeds is on our DynamoDB to Postgres migration tools page.

How does AWS Glue read DynamoDB?

Through one of two connectors, and the choice matters more than anything else in the job. The ETL connector scans the live table. The export connector asks DynamoDB for a point-in-time export to S3 and reads the files. AWS's own guidance favors the export connector for large tables, saying it "performs better than the ETL connector when the DynamoDB table size is larger than 80 GB", and it removes two settings you would otherwise have to tune.

Glue option How it reads Table capacity Needs Use it for
ETL connector Scans the live table Uses read capacity, 0.5 of the table by default dynamodb.splits tuned to your workers Small tables, or no PITR
Export connector Calls ExportTableToPointInTime, reads the S3 files No table read capacity Point-in-time recovery on the table Almost every migration
Export connector, s3 mode Reuses a past export already in S3 None, and no new export charge A completed export at a known prefix Re-runs while you fix casts

Two defaults on the ETL connector catch people. dynamodb.throughput.read.percent is 0.5, so the job tries to consume half the table's read capacity while your application is using the same table. And dynamodb.splits defaults to 1, which AWS describes as "no parallelism". On an on-demand table, Glue treats capacity as 40,000 read units, which is a lot of reads to aim at production. For a migration, the export connector avoids both problems. The s3 mode is worth knowing too: it reuses an export you already have, so you can rerun the job ten times while you fix casts without paying for ten exports.

Does AWS Glue keep DynamoDB number types?

Not by default, and this is the detail that decides how much work the migration is. The export writes each value inside a type descriptor, such as {"N":"600"}, because DynamoDB sends every number across the wire as a string. Glue offers dynamodb.simplifyDDBJson to strip those descriptors, and AWS's documented example of its output shows a number attribute as updatedAt: string and a number set as an array whose elements are strings. The descriptors go, the strings stay.

Written to PostgreSQL as-is, that frame creates text columns. Sorting then runs on characters, so "9" lands after "10", and MAX() on a price column returns the wrong row without an error. The fix is an explicit cast per attribute before the JDBC write: numeric with a fixed scale for money, bigint for counts up to 19 digits and numeric past that (DynamoDB numbers carry up to 38 digits), and timestamptz for epoch fields once you know whether each one is seconds or milliseconds. Avoid double precision for anything that has to tie out, since PostgreSQL only promises 15 digits of precision for it.

What else breaks between Glue and PostgreSQL?

Three PostgreSQL rules that DynamoDB never had to follow. First, the character with code zero cannot be stored in any PostgreSQL text column, and jsonb rejects \u0000 for the same reason, while a DynamoDB string is any UTF-8. One such value can fail a write deep into a long job, so strip it in the transform. Second, identifiers stop at 63 bytes by default, and flattened paths from nested maps can be longer, so two columns can truncate to the same name. Third, unquoted identifiers fold to lower case, and DynamoDB attribute names are case-sensitive, so customerId and customerid become one column unless you rename them.

Lists and sets need a decision rather than a default. A list of line items belongs in a child table with its position kept. A string set belongs in a text[] column, sorted on load, because DynamoDB does not preserve set order. Keep the whole item in a jsonb column as well, so an attribute nobody mapped can still be recovered after DynamoDB is switched off.

How much does AWS Glue cost for a DynamoDB to Postgres migration?

Less than the people running it. A standard Glue Spark job in US East costs $0.44 per DPU-hour, billed per second with a 1-minute minimum, and Flex execution costs $0.29. As an illustration of the arithmetic, a job on 10 DPUs that runs for 30 minutes uses 5 DPU-hours, or $2.20 at the standard rate. Add the DynamoDB export, which AWS bills by table size for a full export and by data processed for an incremental one (with a 10 MB minimum per incremental export), plus S3 storage. Even run fifty times during testing, the compute rarely matters. What costs money is writing the typed schema, the cast for every attribute, and the change replay that keeps PostgreSQL current until cutover.

How do I keep Postgres in sync until cutover?

With incremental exports or DynamoDB Streams, applied as upserts. An incremental export covers a window of 15 minutes to 24 hours and marks each item as an insert, update or delete by its shape: a delete arrives as keys only, or keys plus the old image. Glue can read those files like any other export, but its JDBC writer is not an upsert, so the usual pattern is to load a staging table and run INSERT ... ON CONFLICT from it, deleting rows whose record carries keys only. DynamoDB Streams is fresher but keeps changes for 24 hours, so a consumer that stops over a long weekend needs a fresh export to recover. Before the final switch, it pays to trace which dashboards and jobs read each table, so nothing downstream is still pointed at DynamoDB the morning after. The wider cutover checklist is in our data migration tools guide.

What are the alternatives to AWS Glue for DynamoDB to Postgres?

Four are worth pricing against Glue, and each wins a different kind of project.

Option Pick it when What it bills
AWS Glue Your team writes Spark and wants the job inside AWS with IAM and no new vendor. $0.44 per DPU-hour standard, $0.29 Flex, per second with a 1-minute minimum
S3 export plus a script One or two small tables, a single cutover, and an engineer for a day. The export charge and the engineer's time
Airbyte You want a UI and only need append loads, because deletes are not replicated. Free self-hosted, credits on cloud
Estuary You need Postgres kept current for weeks with little setup. $0.50 per GB plus connector instances
Adapters You want typed columns and child tables from nested items, kept in sync until cutover. Flat monthly plan from $49, not metered by rows

When is AWS Glue the right choice?

When your team already writes PySpark, wants everything inside one AWS account, and is happy to own a job. Glue handles the export, the parallel read and the JDBC write, and the IAM story is simple because no outside vendor touches the table. It is the wrong choice when nobody on the team wants to maintain Spark code after the migration, or when the source is a single-table design that has to be split into several PostgreSQL tables by key prefix, which turns the job into a small application.

Can I get typed Postgres tables from DynamoDB without writing a Glue job?

Yes, by declaring the mapping instead of coding it. Adapters has you name the attribute paths that matter once, casts each number with the scale you state, converts epoch fields with the unit you state, sends lists to child tables and keeps the raw item in jsonb beside them, then keeps PostgreSQL current, deletes included, until you switch writes. The demo at the top of this page maps a real DynamoDB order onto PostgreSQL columns. Plans are flat, from $49 a month. For other sources feeding the same database, see Postgres ETL tools, and if MongoDB is also on the way out, the equivalent guide is MongoDB to PostgreSQL migration tools.

DynamoDB into PostgreSQL columns you named

Declare the attributes that matter, cast them on the way in, split single-table designs by entity and keep Postgres current until cutover. Flat $49 a month, not metered by rows.

The live demo needs no card, and Starter is $49 a month.

Get started