Meltano’s pipelinewise tap-mysql does not correctly handle multiple GTIDs from a MySQL DB
Our setup:
Source: MySQL (Master → Replica)
Replication: Replica receives transactions from Master and also allows its own transactions (not read-only)
Extractor: Meltano tap-mysql (TransferWise variant) pointing to the replica
Target: PostgreSQL
In this setup, the replica contains GTIDs from:
- Master transactions
- Its own local transactions
And running SELECT @@GLOBAL.gtid_executed; on replica returns multiple comma-separated GTID sets.
Like: uuid1:1-100,uuid2:1-50
Expected
State should store full GTID set (all comma-separated values).
Actual
Only one GTID is stored.
Impact
Our goal of using GTID-based replication is to allow the tap to switch to a new primary and resume seamlessly without relying on binlog file/position (which become invalid after failover).
However, since only a single GTID is stored:
Transactions from other GTIDs are missing from the state
On failover, the tap cannot correctly resume and may require a historical sync
Even during regular incremental runs, those missing GTIDs are not tracked, causing those transactions to be reprocessed in every sync
This leads to inefficient incremental syncs and can become costly or infeasible for large datasets.
Additional Note
We would be happy to raise a PR with a proposed fix if aligned with the maintainers.
Meltano’s pipelinewise tap-mysql does not correctly handle multiple GTIDs from a MySQL DB
Our setup:
Source: MySQL (Master → Replica)
Replication: Replica receives transactions from Master and also allows its own transactions (not read-only)
Extractor: Meltano tap-mysql (TransferWise variant) pointing to the replica
Target: PostgreSQL
In this setup, the replica contains GTIDs from:
And running SELECT @@GLOBAL.gtid_executed; on replica returns multiple comma-separated GTID sets.
Like: uuid1:1-100,uuid2:1-50
Expected
State should store full GTID set (all comma-separated values).
Actual
Only one GTID is stored.
Impact
Our goal of using GTID-based replication is to allow the tap to switch to a new primary and resume seamlessly without relying on binlog file/position (which become invalid after failover).
However, since only a single GTID is stored:
Transactions from other GTIDs are missing from the state
On failover, the tap cannot correctly resume and may require a historical sync
Even during regular incremental runs, those missing GTIDs are not tracked, causing those transactions to be reprocessed in every sync
This leads to inefficient incremental syncs and can become costly or infeasible for large datasets.
Additional Note
We would be happy to raise a PR with a proposed fix if aligned with the maintainers.