Integration / Apache Pulsar Interview questions
How do you troubleshoot message duplication in a Pulsar producer?
First confirm whether producer-side deduplication is actually enabled on the namespace or topic - it's off by default in many configurations, so what looks like a "bug" causing duplicates may simply be expected behavior from retried publishes with dedup disabled.
If dedup is enabled but duplicates still appear, check whether the application is creating a new producer instance, with an effectively different producer name each time, on every retry rather than reusing one stable producer - deduplication tracks state per producer name, so a fresh producer name defeats it even with the feature turned on.
Check the deduplication window/snapshot interval configuration - the broker only remembers recent sequence IDs for a bounded time or entry count, and a retry arriving long after that window has rolled off won't be recognized as a duplicate even with dedup correctly enabled and the producer name unchanged.
Distinguish producer-level duplicates from application-level duplicates: if the same logical event is published twice under two legitimately new sequence IDs, because application-layer retry logic sits above the Pulsar client and calls publish twice, no amount of Pulsar-side deduplication will catch that - the fix belongs in the application's own idempotency handling, not Pulsar configuration.
Finally, check consumer-side handling isn't the actual source: a message redelivered after a nack or ack timeout is a different failure than true producer-side duplication, and easy to mistake for one if consumer logs aren't inspected for negative-acknowledgment or timeout events alongside the "duplicate" message IDs.
More Related questions...