Building Resilient Microservices with Dapr: Pub/Sub, State and Service Invocation Patterns

Every microservices team eventually writes the same infrastructure code: a retry wrapper around HTTP calls, a Kafka producer with just the right configuration, a Redis client with optimistic locking, a service-discovery lookup, some mTLS certificate handling. Then the team next door writes it again in a different language, slightly differently.

Dapr (the Distributed Application Runtime) is a CNCF graduated project built to stop that repetition. It runs as a sidecar next to each service and exposes distributed-systems capabilities (messaging, state, service invocation, workflows, secrets and more) through simple HTTP and gRPC APIs. Your code calls localhost. Dapr handles the broker, the database, retries, encryption and tracing.

This article covers five patterns that make Dapr-based microservices resilient in practice, with code and configuration you can adapt.

How the sidecar model helps

Before getting into the patterns, it helps to see why the sidecar model matters for resilience. When your service publishes a message through Dapr, it isn't linked against a Kafka or RabbitMQ client library. It sends a request to its sidecar, and the sidecar talks to whichever broker is configured in a component file. That gives you three benefits:

  1. Consistent behaviour across languages. A Go service and a Python service get identical retry, timeout and tracing semantics.
  2. Swappable infrastructure. Moving from Redis Streams in development to Kafka in production is a YAML change, not a code change.
  3. Central policy. Resilience and security settings live in configuration that platform teams can review and version.

Pattern 1: Asynchronous communication with pub/sub

Synchronous call chains are the most common source of cascading failure. If Service A calls B, which calls C, and C is slow, everyone is slow. Wherever the business process allows, switch to events.

Publishing an event from Python:

import json

from dapr.clients import DaprClient

order = {"orderId": "4711", "total": 129.90}

with DaprClient() as client:

    client.publish_event(

        pubsub_name="orderpubsub",

        topic_name="orders",

        data=json.dumps(order),

        data_content_type="application/json",

    )

Subscribing declaratively, so the subscriber's code has no broker dependency at all:

apiVersion: dapr.io/v2alpha1

kind: Subscription

metadata:

  name: order-subscription

spec:

  pubsubname: orderpubsub

  topic: orders

  routes:

    default: /orders

scopes:

– shipping-service

Dapr wraps messages in the CloudEvents format, delivers at least once, and supports dead-letter topics for messages that keep failing. Because delivery is at least once, make your handlers idempotent. Store the event ID of processed messages, or design updates so that applying them twice has no extra effect.

Pattern 2: Safe concurrent updates with ETags

When several instances of a service update the same record, the classic bug is the lost update: two instances read version 1, both write, and one change silently disappears. Dapr's state API supports optimistic concurrency with ETags on stores that implement it.

from dapr.clients import DaprClient

from dapr.clients.grpc._state import StateOptions, Concurrency

with DaprClient() as client:

    item = client.get_state(store_name="statestore", key="cart-42")

    cart = json.loads(item.data or "{}")

    cart["items"] = cart.get("items", []) + ["sku-123"]

    client.save_state(

        store_name="statestore",

        key="cart-42",

        value=json.dumps(cart),

        etag=item.etag,

        options=StateOptions(concurrency=Concurrency.first_write),

    )

If another instance updated the cart in between, the save fails with an ETag mismatch instead of overwriting their change. Your code can re-read and retry. The same API works whether the underlying store is PostgreSQL, Redis, Cosmos DB or DynamoDB.

Pattern 3: The transactional outbox

This is the pattern that separates hobby systems from production ones. A service needs to update its database and publish an event, for example "order saved" plus "OrderCreated" for downstream services. If it does these as two separate operations, a crash between them leaves the system inconsistent: either the order exists but nobody hears about it, or an event announces an order that was never saved.

The usual fix is the transactional outbox: write the event into an outbox table in the same database transaction as the business data, and have a relay publish it afterwards. Building that relay yourself is tedious. Dapr has outbox support built into the state API. You configure it on the state store component:

apiVersion: dapr.io/v1alpha1

kind: Component

metadata:

  name: orderstore

spec:

  type: state.postgresql

  version: v1

  metadata:

  – name: connectionString

    secretKeyRef:

      name: pg-conn

      key: connectionString

  – name: outboxPublishPubsub

    value: orderpubsub

  – name: outboxPublishTopic

    value: orders

Any state transaction against orderstore now publishes the corresponding event reliably, with no extra code in your service. The team behind Dapr at Diagrid has written in depth about the dual-write problem and how the outbox solves it.

Pattern 4: Secure, observable service invocation

Some interactions really are request/response: fetching a price, validating a token, checking inventory before confirming an order. For those, Dapr's service invocation API gives you service discovery, mTLS encryption and distributed tracing without touching your HTTP code.

with DaprClient() as client:

    resp = client.invoke_method(

        app_id="inventory-service",

        method_name="reserve",

        data=json.dumps({"sku": "sku-123", "qty": 1}),

        http_verb="POST",

    )

The caller addresses the target by its app ID, not by hostname and port. Dapr resolves the location, encrypts the call with mutual TLS using workload identities issued by its Sentry service, and propagates W3C trace context so the call appears in your tracing backend. Access-control policies can restrict which app IDs may call which operations, which is a practical first step toward zero trust inside the cluster. The same identity and mTLS model is central to how durable execution platforms built on Dapr secure service and agent traffic.

Pattern 5: Declarative resiliency policies

Here's where Dapr removes the most boilerplate. Instead of writing retry loops and circuit breakers in every service, you declare them in a Resiliency resource and apply them to apps, components and actors:

apiVersion: dapr.io/v1alpha1

kind: Resiliency

metadata:

  name: checkout-resiliency

spec:

  policies:

    timeouts:

      fast: 3s

    retries:

      standard:

        policy: exponential

        maxInterval: 15s

        maxRetries: 5

    circuitBreakers:

      protectInventory:

        maxRequests: 1

        timeout: 30s

        trip: consecutiveFailures >= 5

  targets:

    apps:

      inventory-service:

        timeout: fast

        retry: standard

        circuitBreaker: protectInventory

    components:

      orderpubsub:

        outbound:

          retry: standard

With this in place, every call to inventory-service times out after three seconds, retries with exponential backoff, and trips a circuit breaker after five consecutive failures, which gives the struggling service room to recover instead of piling more load onto it. The Dapr resiliency documentation covers the full policy language, including per-operation overrides.

A word of caution: retries multiply load. If three services in a chain each retry five times, one user request can turn into 125 calls to the bottom service. Retry at one layer (usually the caller closest to the failure), keep timeouts tight, and rely on circuit breakers to fail fast.

Don't forget secrets and configuration

Resilience isn't only about retries. A surprising number of outages start with a rotated database password that one service never picked up, or a connection string copied into three places. Dapr's secrets API and secret-store components (Kubernetes secrets, HashiCorp Vault, AWS Secrets Manager, Azure Key Vault and others) let components reference credentials by name, as the secretKeyRef in the outbox example does, instead of embedding them. Rotate the secret in one place and every component that references it picks up the change on its next reload. Combine that with component scoping so each service can read only the secrets and components it actually needs.

When the process spans many steps: add workflows

Pub/sub, state and service invocation cover most interactions, but some business processes span many services and take minutes or days: a checkout saga with payment, inventory, shipping and compensation logic, for example. Coordinating that with events alone ("choreography") gets hard to reason about quickly. Dapr Workflow lets you write the orchestration as durable code, with each step automatically checkpointed and retried, alongside the building blocks described above.

Putting it together

A resilient Dapr-based order system might look like this:

  • The checkout service saves the order through a state store with outbox enabled, so the order and the OrderCreated event are atomic.
  • The shipping and notification services subscribe to orders with idempotent handlers and a dead-letter topic.
  • Checkout calls the inventory service synchronously through service invocation, protected by timeouts, retries and a circuit breaker.
  • All traffic between services is encrypted with mTLS and traced end to end.
  • The multi-step payment and fulfilment process runs as a durable workflow.

None of these services contains broker SDKs, retry libraries or certificate-handling code. That's the real payoff: less infrastructure code in each service, and resilience behaviour that's consistent, reviewable and centrally managed.

Getting started

Install the Dapr CLI, run dapr init to start a local environment with Redis and Zipkin, and try one pattern at a time. The outbox and resiliency policies are good places to start, because they fix real production failure modes with almost no application code. Once you're running on Kubernetes, invest early in operational tooling such as control-plane upgrades, certificate rotation and metrics. Those day-2 concerns are what keep a resilient design resilient over time.