Pub/Sub and Event-Driven Systems: Where Loose Coupling Pays Off
Event-driven architecture is attractive because producers can publish facts without knowing which consumers will use them. That decoupling can make systems easier to evolve, scale, and integrate, but only if the design accepts the realities of distributed messaging: retries, duplicates, delayed processing, ordering limits, schema evolution, and partial failure.
Google Cloud Pub/Sub is a natural service to understand for the Professional Cloud Architect exam and the Google Professional Cloud Architect role because it sits between application architecture and operations. The important design skill is not memorizing how to create a topic. It is knowing when asynchronous messaging reduces coupling and when it simply moves complexity into harder-to-see places.
Loose coupling pays off when time, ownership, scaling, or availability should be separated between components. It is not automatically the right answer for every request.
Use events to separate responsibility, not to avoid APIs
A synchronous API is appropriate when the caller needs an immediate answer from a specific capability. An event is stronger when the producer only needs to state that something happened and independent consumers can react on their own schedule.
For example, an order service may need a synchronous response that payment authorization succeeded. After the order is accepted, events can trigger analytics, notifications, fulfillment, fraud enrichment, and downstream data updates without making checkout wait for all of them.
This pattern is common in serverless architecture because short-lived consumers can scale in response to work. The value, however, comes from the boundary between responsibilities, not from the serverless label.
Design consumers for at-least-once delivery
Pub/Sub uses at-least-once delivery by default. A message can therefore be delivered again, including after the subscriber has already done some work. Consumers should be idempotent or have a deduplication strategy for operations that cannot safely repeat.
Idempotency can be implemented with a business key, processed-message record, conditional write, or state transition that rejects duplicates. The right technique depends on whether processing is a pure calculation, an external side effect, or a transactional update.
Treat duplicate handling as a normal code path. If the design assumes duplicates are exceptional, the system will eventually create duplicate emails, double-counted metrics, repeated payments, or conflicting state.
Ordering should be scoped to the business requirement
By default, Pub/Sub does not promise ordering across messages. Ordering can be enabled for messages that share an ordering key, provided the publisher follows the relevant regional behavior. This is useful when related events truly require sequence, such as changes to the same customer or record.
Do not impose global ordering when only per-entity ordering is needed. Global order reduces parallelism and can make one slow or failed message block unrelated work. Most event systems scale better when ordering is scoped to the smallest unit that requires it.
Also ask whether the consumer can be designed around state rather than exact message order. Sometimes a consumer can fetch the current record after receiving an event and avoid reconstructing truth from a perfect event sequence.
Exactly-once is a tool, not permission to ignore side effects
Pub/Sub supports exactly-once delivery for supported subscription patterns, but delivery semantics alone do not make an entire business workflow exactly once. A consumer can acknowledge a message and still fail while calling another system, or perform an external side effect that cannot participate in the same transaction.
Architects should define the atomic boundary. If processing writes to a database, can the state update and deduplication marker be committed together? If it calls a payment provider or email service, what reconciliation process handles an ambiguous outcome?
Strong messaging features reduce one class of duplicates, but correctness still requires application-level thinking.
Retries need an exit path
Transient failures should retry, but permanent failures should not cycle forever. Use retry policies and dead-letter handling so poison messages become visible and can be diagnosed without blocking healthy throughput.
This is part of mature DevOps operations: the delivery pipeline includes production failure handling, not just deployment. Teams need dashboards for backlog, oldest unacked message, retry rates, dead-letter volume, consumer errors, and end-to-end processing latency.
A dead-letter topic is not a trash bin. Assign ownership, include enough message context for diagnosis, and define whether repaired messages can be replayed safely.
Backlog is both a buffer and a warning
One advantage of asynchronous systems is that a producer can continue while consumers are briefly slow. The backlog absorbs the mismatch. But a growing backlog is also a capacity signal: the system is accumulating work faster than it completes it.
Scale consumers using measures that reflect work, not only CPU. Message age and backlog size can be more useful than instance utilization. A consumer that is blocked on an external API can have low CPU while the queue is becoming operationally dangerous.
Define a maximum acceptable event age for each subscription. That turns backlog into a service objective rather than an aesthetic dashboard number.
Schemas are contracts even when publishers are decoupled
Loose coupling does not mean no contract. Consumers still depend on field meaning, data types, identifiers, and event semantics. A producer that changes these without compatibility can break many services at once.
Cloud-native teams often discover that cloud-native engineering depends as much on interface discipline as on containers or managed services. Events need versioning rules, optional-field conventions, ownership, and a deprecation process.
Prefer additive schema change where possible. New consumers can use new fields while old consumers continue operating. Breaking changes deserve explicit versions or parallel topics rather than silent reinterpretation of an existing event.
Know when messaging is the wrong abstraction
Do not put a queue between two components simply because asynchronous design sounds modern. If the caller needs an immediate authoritative response and there is one clear owner, a synchronous API can be simpler to understand, secure, test, and troubleshoot.
Messaging also adds operational state: subscriptions, retention, retries, permissions, schemas, dead-letter handling, and consumer lag. That complexity is justified when it buys meaningful independence in scaling, availability, release cadence, or integration.
A strong cloud engineering practice chooses the smallest architecture that meets the requirement. Event-driven patterns are most valuable at real boundaries between teams, domains, or timing needs.
Replay behavior deserves the same attention as normal delivery. If a consumer is repaired after a bug or outage, operators need to know whether older messages can be replayed safely, how far back the retained history reaches, and how duplicate side effects will be prevented. Idempotency keys, business-level deduplication, and transactional boundaries should be designed around the effect that matters, such as charging a card or creating an order, not only around message identifiers. Schema evolution also needs a rollout sequence: producers and consumers may run different versions at the same time, so compatible fields and tolerant readers are often more valuable than tightly synchronized releases. During incident recovery, monitor backlog age as well as backlog size; a modest queue containing very old messages can be more damaging than a large queue that is draining quickly. These details determine whether loose coupling remains an advantage under failure.
Design the event flow as an end-to-end system
A useful architecture review follows one event from the business action that creates it to every important consumer. Identify the publisher’s transaction boundary, event identifier, schema, topic, subscription type, retry policy, dead-letter route, consumer side effects, acknowledgment point, monitoring, and replay procedure.
Then test uncomfortable scenarios. What if the same event arrives twice? What if event B arrives before event A? What if the consumer is down for six hours? What if a schema field is missing? What if a downstream payment or email call times out after the remote system actually completed it? Those questions expose the real architecture.
Security belongs in the flow as well. Limit who can publish and subscribe, separate environments, avoid placing secrets or unnecessary personal data in events, and understand that a topic can become a high-value data distribution point.
Pub/Sub is powerful because it lets time and ownership separate. The producer can finish its job without coordinating every consumer, and consumers can scale or recover independently. The trade is that the system must become comfortable with eventual work, repeated delivery, and observable queues. When those properties match the business process, loose coupling is a major architectural advantage.
Replay deserves its own design. Retained messages or downstream event stores can help rebuild a consumer after a defect, but replaying historical traffic may trigger side effects that were safe only once. A repaired consumer should be able to distinguish state-reconstruction work from actions such as sending notifications or charging customers.
Event identity is therefore useful beyond deduplication. Include a stable event ID, occurrence time, producer or domain, and version so operators can trace one business event across topics and consumers. Correlation becomes especially important when a single action fans out to many systems and one branch fails hours later.
Capacity planning should consider subscriber failure as well as normal growth. If a consumer is offline for several hours, it may need to process backlog faster than the ordinary publish rate when it returns. The architecture should know whether downstream databases and APIs can tolerate that recovery surge or whether consumers need controlled catch-up rates.
Subscriber concurrency should be bounded by the slowest dependency in the path. If a database can process five hundred writes per second, allowing thousands of consumers to retry simultaneously can turn a recoverable backlog into a downstream outage. Backpressure, quotas, and controlled concurrency are architecture tools, not just tuning details.
Ownership should follow subscriptions rather than topics alone. A topic may serve many teams, but each subscription has its own backlog, retry behavior, and service objective. Naming an owner for every subscription makes stale consumers and forgotten integrations easier to find before they become reliability or cost problems.